Coding method, decoding method, bitstream, coder, decoder and storage medium

By utilizing the statistical results of the prediction mode set of decoded or encoded blocks in the video coding standard, multiple intra-frame prediction modes are selected for weighted prediction, which solves the problem of high computational complexity of intra-frame prediction modes and achieves more efficient encoding/decoding and video compression performance.

WO2025260253A1PCT designated stage Publication Date: 2025-12-26GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/099956
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

In existing video coding standards, intra-frame prediction mode has high computational complexity, resulting in low encoding and decoding efficiency and difficulty in meeting the bandwidth requirements of high-resolution video.

Method used

By determining the statistical results in the set of prediction modes of decoded or encoded blocks, multiple intra-frame prediction modes are selected for weighted prediction, reducing computational complexity and improving encoding and decoding efficiency.

Benefits of technology

It reduces computational complexity, improves encoding and decoding efficiency, enhances video compression performance, reduces bitstream overhead, and improves prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024099956_26122025_PF_FP_ABST
    Figure CN2024099956_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a coding method, a decoding method, a bitstream, a coder, a decoder and a storage medium. The decoding method comprises: determining one or more preset blocks that have been decoded; on the basis of prediction modes respectively used by the one or more preset blocks, determining statistical results of at least two candidate intra-frame prediction modes in a first mode set; on the basis of the statistical results, determining from the first mode set at least two intra-frame prediction modes for a current block; and performing prediction on the current block on the basis of the at least two intra-frame prediction modes, so as to determine a predicted value of the current block. Thus, the compression performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Encoding / decoding methods, bitstreams, encoders, decoders, and storage media Technical Field

[0001] This application relates to the field of video encoding and decoding technology, and in particular to an encoding and decoding method, a bitstream, an encoder, a decoder, and a storage medium. Background Technology

[0002] As people's demands for video display quality have increased, high-resolution video, such as HD and UHD, has emerged. However, high-resolution video typically contains more information, thus requiring more bandwidth. To reduce bandwidth requirements, video coding standards involving video compression have been introduced.

[0003] In video coding standards, for more complex intra-prediction modes, they all utilize the reconstruction regions adjacent to the current block for processing and analysis to deduce the intra-prediction mode used in the current block, and they all support combinations of multiple intra-prediction modes. However, these more complex intra-prediction modes require a certain amount of computation to deduce the intra-prediction mode used in the current block, thus increasing computational complexity and affecting encoding and decoding efficiency.

[0004] Summary of the Invention

[0005] This application provides an encoding / decoding method, a bitstream, an encoder, a decoder, and a storage medium, which can reduce computational complexity, improve encoding / decoding efficiency, and thus enhance compression performance.

[0006] The technical solution of this application embodiment can be implemented as follows:

[0007] In a first aspect, embodiments of this application provide a decoding method applied to a decoder, the method comprising:

[0008] Identify one or more pre-defined blocks that have been decoded;

[0009] Based on the prediction modes used by one or more preset blocks, determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set;

[0010] Based on the statistical results, at least two intra-prediction modes for the current block are determined from the first mode set;

[0011] The current block is predicted based on at least two intra-frame prediction modes to determine the predicted value of the current block.

[0012] Secondly, embodiments of this application provide an encoding method applied to an encoder, the method comprising:

[0013] Identify one or more pre-defined blocks that have already been encoded;

[0014] Based on the prediction modes used by one or more preset blocks, determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set;

[0015] Based on the statistical results, at least two intra-prediction modes for the current block are determined from the first mode set;

[0016] The current block is predicted based on at least two intra-frame prediction modes to determine the predicted value of the current block.

[0017] Thirdly, embodiments of this application provide a bitstream generated by bit encoding according to the encoding method described in the second aspect; wherein the information to be encoded in the encoding method includes at least one of the following: the transform coefficients of the current block, the transform kernel group index of the current block, the transform kernel index of the current block, and the value of a first syntax element, wherein the first syntax element is used to indicate whether the current block uses a first prediction mode.

[0018] Fourthly, embodiments of this application provide an encoder, which includes a first determining unit and a first predicting unit, wherein:

[0019] The first determining unit is configured to determine one or more pre-coded blocks; and to determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set according to the prediction modes used by each of the one or more pre-coded blocks.

[0020] The first determining unit is further configured to determine at least two intra-prediction modes of the current block in the first mode set based on statistical results.

[0021] The first prediction unit is configured to predict the current block based on at least two intra-frame prediction modes to determine the predicted value of the current block.

[0022] Fifthly, embodiments of this application provide an encoder, which includes a first memory and a first processor, wherein:

[0023] A first memory for storing computer programs that can run on a first processor;

[0024] A first processor is configured to perform the steps of the encoding method as described in the second aspect when running the computer program.

[0025] Sixthly, embodiments of this application provide a decoder, which includes a second determining unit and a second predicting unit, wherein:

[0026] The second determining unit is configured to determine one or more pre-defined blocks that have been decoded; and to determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by each of the one or more pre-defined blocks.

[0027] The second determining unit is further configured to determine at least two intra-prediction modes of the current block from the first mode set based on statistical results.

[0028] The second prediction unit is configured to predict the current block based on at least two intra-frame prediction modes and determine the predicted value of the current block.

[0029] In a seventh aspect, embodiments of this application provide a decoder, which includes a second memory and a second processor, wherein:

[0030] The second memory is used to store computer programs that can run on the second processor;

[0031] The second processor is configured to perform the steps of the decoding method as described in the first aspect when running the computer program.

[0032] Eighthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the decoding method as described in the first aspect, or the steps of the encoding method as described in the second aspect.

[0033] In a ninth aspect, embodiments of this application provide a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the decoding method as described in the first aspect, or the steps of the encoding method as described in the second aspect.

[0034] In a tenth aspect, embodiments of this application provide a computer-readable storage medium having a bitstream stored thereon, the bitstream being generated by performing the steps of the encoding method as described in the second aspect.

[0035] This application provides an encoding / decoding method, a bitstream, an encoder, a decoder, and a storage medium. At the encoding end, one or more pre-defined blocks that have been encoded are determined; based on the prediction modes used by each of the one or more pre-defined blocks, statistical results of at least two candidate intra-frame prediction modes in a first mode set are determined; based on the statistical results, at least two intra-frame prediction modes for the current block are determined in the first mode set; and the current block is predicted based on the at least two intra-frame prediction modes to determine the predicted value of the current block. At the decoding end, one or more pre-defined blocks that have been decoded are determined; based on the prediction modes used by each of the one or more pre-defined blocks, statistical results of at least two candidate intra-frame prediction modes in a first mode set are determined; based on the statistical results, at least two intra-frame prediction modes for the current block are determined in the first mode set; and the current block is predicted based on the at least two intra-frame prediction modes to determine the predicted value of the current block. In this way, by using the relevant blocks that have been encoded and decoded to perform statistics, the statistical results of these relevant blocks are obtained, and at least two intra-prediction modes used in the current block are derived based on the statistical results. Since it is not necessary to explicitly indicate which intra-prediction mode to use, the overhead in the bitstream can be saved. Moreover, using at least two intra-prediction modes for weighted prediction can also improve the accuracy of prediction and reduce prediction errors to a certain extent. In addition, deriving at least two intra-prediction modes used in the current block based on statistical results can also avoid the need for intra-prediction mode derivation through gradient calculation, such as in DIMD mode, thus reducing computational complexity, thereby improving encoding and decoding efficiency and improving compression performance. Attached Figure Description

[0036] Figure 1 is a flowchart of a hybrid coding framework;

[0037] Figure 2 is a schematic diagram of template matching for the current block;

[0038] Figure 3 is a schematic diagram of a reference sample of the current block;

[0039] Figure 4 is a schematic diagram of multiple reference rows in the current block;

[0040] Figure 5 is a schematic diagram of multiple prediction modes corresponding to one type of intra-frame prediction.

[0041] Figure 6 is a schematic diagram of multiple prediction modes corresponding to one type of intra-frame prediction.

[0042] Figure 7 is a schematic diagram of multiple prediction modes corresponding to one type of intra-frame prediction.

[0043] Figure 8 is a schematic diagram of multiple prediction modes corresponding to one type of intra-frame prediction;

[0044] Figure 9 is a schematic diagram of screen content encoding;

[0045] Figure 10 is a schematic diagram of the prediction process of a MIP mode;

[0046] Figure 11 is a schematic diagram of a template for the current block and a template reference area;

[0047] Figure 12 is a schematic diagram of a gradient and intra-frame prediction mode;

[0048] Figure 13 is a schematic diagram of weighted fusion of three intra-frame prediction modes;

[0049] Figure 14 is a schematic diagram of the search range of the ITMP mode;

[0050] Figure 15 is a schematic diagram of the weights of multiple modes under a GPM mode;

[0051] Figure 16A shows a typical filter schematic diagram.

[0052] Figure 16B is a typical schematic diagram of a filter.

[0053] Figure 16C is a typical schematic diagram of a filter.

[0054] Figure 17A is a schematic diagram of a reconstructed sample region used for training a filter;

[0055] Figure 17B is a schematic diagram of a reconstructed sample region used for training filters.

[0056] Figure 17C is a schematic diagram of a reconstructed sample region used for training filters.

[0057] Figure 18 is a schematic diagram of a DCT transformation;

[0058] Figure 19 is a schematic diagram of the base image for a DCT transform;

[0059] Figure 20 is a flowchart of a method without LFNST transformation;

[0060] Figure 21 is a schematic diagram of a process with LFNST transformation;

[0061] Figure 22 is a detailed flowchart of an LFNST transformation;

[0062] Figure 23 is a schematic diagram of a base image of multiple transform kernel groups;

[0063] Figure 24 is a schematic diagram of the base image for an NSPT transform;

[0064] Figure 25 is a schematic diagram of a video encoding and decoding network architecture provided in an embodiment of this application;

[0065] Figure 26 is a schematic block diagram of an encoder provided in an embodiment of this application;

[0066] Figure 27 is a schematic block diagram of a decoder system provided in an embodiment of this application;

[0067] Figure 28 is a schematic flowchart of a decoding method provided in an embodiment of this application;

[0068] Figure 29 is a schematic diagram of the spatial relationship between adjacent and non-adjacent blocks of the current block according to an embodiment of this application;

[0069] Figure 30 is a schematic diagram of the spatial relationship between a current block and a preset range provided in an embodiment of this application;

[0070] Figure 31 is a schematic flowchart of a decoding method provided in an embodiment of this application;

[0071] Figure 32 is a schematic diagram of the structure of a candidate sample provided in an embodiment of this application;

[0072] Figure 33 is a schematic flowchart of a decoding method provided in an embodiment of this application;

[0073] Figure 34 is a schematic flowchart of a decoding method provided in an embodiment of this application;

[0074] Figure 35 is a schematic flowchart of an encoding method provided in an embodiment of this application;

[0075] Figure 36 is a schematic flowchart of an encoding method provided in an embodiment of this application;

[0076] Figure 37 is a schematic diagram of the composition structure of an encoder provided in an embodiment of this application;

[0077] Figure 38 is a schematic diagram of the specific hardware structure of an encoder provided in an embodiment of this application;

[0078] Figure 39 is a schematic diagram of the composition structure of a decoder provided in an embodiment of this application;

[0079] Figure 40 is a schematic diagram of the specific hardware structure of a decoder provided in an embodiment of this application;

[0080] Figure 41 is a schematic diagram of the composition structure of an encoding / decoding system provided in an embodiment of this application. Detailed Implementation

[0081] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.

[0082] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0083] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0084] It should also be noted that the terms "first, second, and third" used in the embodiments of this application are only used to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0085] In video images, a first color component, a second color component, and a third color component are generally used to represent a coding block (CB). These three color components are a luma component, a blue chroma component, and a red chroma component, respectively. Specifically, the luma component is usually represented by the symbol Y, the blue chroma component is usually represented by the symbol Cb or U, and the red chroma component is usually represented by the symbol Cr or V. Thus, video images can be represented in YCbCr format or YUV format.

[0086] Before providing a further detailed description of the embodiments of this application, the nouns and terms used in the embodiments of this application will be explained. The nouns and terms used in the embodiments of this application shall be interpreted as follows:

[0087] H.265 / High Efficiency Video Coding (HEVC);

[0088] H.266 / Versatile Video Coding (VVC);

[0089] VVC's reference software testing platform (VVC Test Model, VTM);

[0090] The platform that improves compression performance after VVC (Enhanced Compression Model, ECM);

[0091] Joint Video Experts Team (JVET);

[0092] Coding Unit (CU);

[0093] Coding Tree Unit (CTU);

[0094] Largest Coding Unit (LCU);

[0095] Motion Vector (MV);

[0096] Prediction Unit (PU);

[0097] Transform Unit (TU);

[0098] Merge technology;

[0099] Skip technique;

[0100] Quantization parameter (QP);

[0101] The technique of fusion with motion vector difference (MMVD) is used.

[0102] Motion Vector Prediction (MVP);

[0103] Temporal Motion Vector Prediction (TMVP);

[0104] Discrete Cosine Transform (DCT);

[0105] Discrete Sine Transform (DST);

[0106] Multiple Transform Selection (MTS);

[0107] Low Frequency Non-Separable Transform (LFNST);

[0108] Non-Separable Primary Transform (NSPT);

[0109] Context-based Adaptive Binary Arithmetic Coding (CABAC).

[0110] Currently, common video coding and decoding standards all employ a block-based hybrid coding framework. Each image, sub-image, or frame in the video is divided into square maximum coding units (MCUs) or coding tree units of the same size (e.g., 256×256, 128×128, 64×64, etc.). Each MCU or coding tree unit can be further divided into rectangular coding units according to rules. Coding units may also be further divided into prediction units, transform units, etc. Specifically, as shown in Figure 1, the hybrid coding framework 10 may include a pre-encode filtering module 11, an intra-prediction module 12, an inter-prediction module 13, a motion estimation module 14, a transform module 15, a quantization module 16, an entropy coding module 17, an inverse quantization module 18, an inverse transform module 19, a loop filter module 20, and a decoded picture buffer module 21. The predictions here can include intra-prediction and inter-prediction, and inter-prediction can include motion estimation and motion compensation. Since there is a strong correlation between adjacent samples in an image of a video, intra-prediction is used in video coding and decoding technology to eliminate spatial redundancy between adjacent samples. Furthermore, since there is a strong similarity between adjacent images in a video, inter-image prediction is used in video coding and decoding technology to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency. It should be noted that a sample can also be called a pixel, and a sample includes both location information and value.

[0111] The basic workflow of a video codec is as follows: At the encoding end, an image is divided into blocks. Intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block. The prediction block is subtracted from the initial block of the current block to obtain a residual block. The residual block is transformed and quantized to obtain a quantization coefficient matrix. The quantization coefficient matrix is ​​entropy-encoded and output to the bitstream. At the decoding end, intra-frame prediction or inter-frame prediction is used on the current block to generate a prediction block. On the other hand, the bitstream is parsed to obtain a quantization coefficient matrix. The quantization coefficient matrix is ​​inverse-quantized and inverse-transformed to obtain a residual block. The prediction block and the residual block are added to obtain a reconstructed block. The reconstructed blocks form a reconstructed image. Loop filtering is performed on the reconstructed image based on the image or based on the blocks to obtain a decoded image. The encoding end also needs similar operations to the decoding end to obtain the decoded image. The decoded image can be used as a reference image for subsequent images in inter-frame prediction. The block division information, prediction, transform, quantization, entropy coding, loop filtering, and other mode information or parameter information determined at the encoding end need to be output to the bitstream if necessary. The decoding end determines the same block partitioning information, prediction, transform, quantization, entropy coding, loop filtering, and other mode or parameter information as the encoding end by parsing the bitstream and analyzing existing information. This ensures that the decoded image obtained by the encoding end is the same as that obtained by the decoding end. The decoded image obtained by the encoding end is usually called the reconstructed image. During prediction, the current block can be divided into prediction units, and during transform, the current block can be divided into transform units. The division of prediction units and transform units can be different. The above is the basic flow of a video codec under a block-based hybrid coding framework. With the development of technology, some modules or steps of this framework or flow may be optimized. The embodiments of this application are applicable to the basic flow of a video codec under this block-based hybrid coding framework, but are not limited to this framework and flow.

[0112] Furthermore, in this embodiment, the current block (CB) can be the current coding unit, the current prediction unit, or the current transform unit, etc. Due to the need for parallel processing, an image can be divided into slices, etc. Slices within the same image can be processed in parallel, meaning they have no data dependency. A "frame" is a commonly used term, generally understood as an image. In this embodiment, the term "frame" can also be replaced with an image or a slice, etc.

[0113] The following sections will provide a detailed introduction to the relevant solutions for prediction techniques.

[0114] (a) Template matching.

[0115] Template matching (TM) was first used in inter-frame prediction. It leverages the correlation between adjacent samples to use regions surrounding the current block as templates. During the encoding and decoding of the current block, its left and top sides are already encoded and decoded according to the encoding order. However, in existing hardware decoder implementations, it's not guaranteed that the left and top sides will be decoded before the current block begins decoding. This applies to inter-frame blocks; for example, in HEVC, inter-coded blocks do not require surrounding reconstructed samples when generating prediction blocks, allowing the prediction process for inter-frame blocks to run in parallel. Intra-coded blocks, however, always require reconstructed samples from the left and top sides as reference samples. Theoretically, the left and top sides are available, meaning that hardware design adjustments can be made to achieve this. Relatively speaking, the right and bottom sides are not available under current standards like VVC's encoding order.

[0116] As shown in Figure 2, the rectangular regions on the left and top sides of the current block are set as templates. The height of the template on the left side is generally the same as the height of the current block, and the width of the template on the top side is generally the same as the width of the current block, although they can be different. The optimal matching position of the template is found in the reference image to determine the motion information or motion vector of the current block. This process can be roughly described as follows: in a reference image (Ref0), starting from a starting position, a search is performed within a certain range around the current block. Search rules, such as the search range and search step size, can be pre-defined. Each time the block moves to a new position, the matching degree between the template at that position and the templates surrounding the current block is calculated. The matching degree can be measured by some distortion costs, such as the Sum of Absolute Difference (SAD), the Sum of Absolute Transformed Difference (SATD), and the Mean-Square Error (MSE). Generally, SATD uses the Hadamard transform. Furthermore, the smaller the values ​​of SAD, SATD, and MSE, the higher the matching degree. The cost is calculated using the predicted block of the template corresponding to the current location and the reconstructed blocks of the templates surrounding the current block. Besides searching the entire sample location, a search can also be performed on sub-sample locations. The motion information of the current block is determined based on the location with the highest matching degree. Utilizing the correlation between adjacent samples, motion information suitable for the template may also be suitable for the current block. Of course, template matching may not be applicable to all blocks, so methods can be used to determine whether the template matching method is used for the current block, such as using a control switch to indicate whether template matching is used. This template matching method is called Decoder-Side Motion Vector Derivation (DMVD). Both the encoder and decoder can use templates to search and derive motion information or find better motion information based on existing motion information. It does not require transmitting specific motion vectors or motion vector differences; instead, both the encoder and decoder perform the same search rules to ensure consistency between encoding and decoding. Template matching can improve compression performance, but it requires a "search" at the decoding end, thus introducing some decoding complexity.

[0117] (ii) Intra-frame prediction.

[0118] It is understandable that there is a strong spatial correlation between adjacent parts or samples within an image. Intra-frame prediction is a prediction method that utilizes the spatial correlation between the already encoded / decoded samples around the current block and the samples within the current block. For example, as shown in Figure 3, the 4×4 white-filled sample represents the current block, and the grid-filled samples in the left column and top row of the current block are the reference samples for the current block. Intra-frame prediction uses these reference samples to predict the current block. These reference samples may all be available, i.e., all have been encoded / decoded. Some may also be unavailable; for example, if the current block is the leftmost part of the entire frame, then the reference samples on the left side of the current block are unavailable. Or, if the lower left part of the current block has not yet been encoded / decoded, then the reference samples in the lower left are also unavailable. In cases where reference samples are unavailable, available reference samples, certain values, or certain methods can be used to fill the gaps, or no filling can be performed.

[0119] Multiple reference line (MRL) intra-frame prediction methods can use more reference samples, thereby improving coding efficiency. Figure 4 shows a schematic diagram of using four reference rows / columns.

[0120] Intra-frame prediction has several prediction modes, as shown in Figure 5. These are the nine modes used in H.264 for intra-frame prediction of a 4×4 block. Mode 0 (Vertical Mode) copies the values ​​of the samples above the current block vertically as the predicted values. Mode 1 (Horizontal Mode) copies the values ​​of the reference samples on the left horizontally as the predicted values. Mode 2 (DC Mode) uses the average of points A-D and I-L as the predicted values ​​for all points. Modes 3-8 copy the reference sample values ​​to the corresponding positions in the current block at a specific angle. Because some positions in the current block may not perfectly correspond to the reference samples, a weighted average of the reference samples, or interpolated subsamples of the reference samples, may be needed.

[0121] In addition to these, there are modes such as PLANE and PLANA. With technological advancements and block size increases, the number of angle prediction modes is also increasing. For example, HEVC uses 35 intra-frame prediction modes, including PLANA, DC, and 33 angle modes (see Figure 6). VVC uses 67 intra-frame prediction modes, including PLANA, DC, and 65 angle modes (see Figure 7). Of course, besides the 67 modes mentioned above, VVC also provides wide-angle modes for some rectangular blocks with significant differences in length and width, as shown by the dashed lines in Figure 8, namely the -14 to -1 and 67 to 80 ranges. These modes replace some conventional modes (see Figure 8).

[0122] (iii) Intra Block Copy (IBC).

[0123] IBC significantly improves the compression efficiency of Screen Content Coding (SCC), and therefore, IBC is used for screen content coding from HEVC to VVC. Screen content differs from camera-captured content; it is computer-generated, noise-free, and contains text, computer graphics, etc., with clear boundaries. Screen content contains a large amount of repetitive content, as shown in Figure 9.

[0124] In the embodiments of this application, IBC can be considered as applying the inter-frame prediction method to intra-frame prediction. Inter-frame prediction uses reference blocks from a reference image to generate the prediction block for the current block; the reference image is not the current image. IBC, on the other hand, finds reference blocks from the already encoded or reconstructed portion of the current image to generate the prediction block for the current block. IBC can also be called intra-picture block compensation or Current Picture Referencing (CPR).

[0125] IBC uses a block vector (BV) to represent the positional difference between the current block and the reference block, similar to the MV in inter-frame prediction. The encoder determines the best matching block for the current block within the search range using block matching and encodes the BV. There are various methods for encoding the BV, such as using merge mode, which is similar to inter-frame prediction and will not be elaborated upon here.

[0126] IBC can be considered a type of intra-frame prediction method, or it can be considered a separate prediction method independent of intra-frame and inter-frame prediction. IBC is highly efficient at encoding screen content and can also improve compression efficiency in natural sequences captured by the camera.

[0127] (iv) Matrix-based Intra Prediction (MIP).

[0128] For MIP, this is a special intra-prediction mode, which can also be called Matrix weighted Intra Prediction in some places.

[0129] As shown in Figure 10, to predict a block with width W and height H, MIP requires H reconstructed samples from the left column and W reconstructed samples from the top row of the current block as input. MIP generates the prediction block in three steps: (a) reference sample averaging, (b) matrix vector multiplication, and (c) interpolation. Matrix multiplication is considered the core of MIP. It can be viewed as the process of generating the prediction block from the input samples (reference samples) using a matrix multiplication method. MIP provides various matrices; the different prediction methods are reflected in the different matrices used. The same input samples will yield different results using different matrices. Reference sample averaging and interpolation are a trade-off between performance and complexity. For larger blocks, reference sample averaging can achieve an approximate downsampling effect, allowing the input to fit a smaller matrix, while interpolation achieves an upsampling effect. This eliminates the need to provide MIP matrices for every block size, instead providing only one or a few specific matrices. With increasing demands for compression performance and improvements in hardware capabilities, the next generation of standards may feature more complex MIPs.

[0130] MIP is somewhat similar to PLANA, but MIP is clearly more complex and flexible than PLANA.

[0131] (v) Template-based Intra Mode Derivation (TIMD).

[0132] As shown in Figure 11, for the current block, a region to its left and top is used as the template. Except for boundary cases, theoretically, reconstructed values ​​can be obtained from the left and top of the current block during encoding and decoding. This is the basis for many template adaptation methods. TIMD uses the diagonally filled region shown in Figure 11 as the template, and the template reference region in Figure 11 is the template's reference sample (grid-filled region). The decoder can use a certain intra-prediction mode to predict on the template and compare the predicted value with the reconstructed value to obtain the cost of that intra-prediction mode on the template. Examples include SAD, SATD, and SSE. Since the template and the current block are adjacent and correlated, the performance of a prediction mode on the template can be used to estimate its performance on the current block. TIMD predicts several candidate intra-prediction modes on the template, obtains their costs on the template, and selects the one or two intra-prediction modes with the lowest costs as the intra-prediction modes for the current block.

[0133] Research has found that if the cost difference between two intra-frame prediction modes on the template is not significant, weighted averaging of the prediction values ​​of the two intra-frame prediction modes can improve compression performance. The weights of the prediction values ​​of the two prediction modes are related to the aforementioned cost; in the current version, this weight is inversely proportional to the cost.

[0134] In summary, TIMD uses the prediction performance of intra-prediction modes on a template to select the intra-prediction mode, and it can weight two intra-prediction modes based on the cost on the template. The advantage of TIMD is that if the current block selects the TIMD mode, it doesn't need to specify which intra-prediction mode is used; instead, the decoder derives it through the above process, saving overhead to some extent. Using two intra-prediction modes for weighting also improves compression performance, and theoretically, it can use more than two intra-prediction modes for weighting.

[0135] (vi) Decoder-side Intra Mode Derivation (DIMD).

[0136] DIMD derives a prediction pattern using reconstructed samples to the left and top of the current block, but instead of making predictions on the template, it analyzes the gradients of the reconstructed samples.

[0137] As shown in Figure 12, DIMD analyzes the gradient of black dots, such as horizontal and vertical gradients, and adapts an intra-frame prediction mode based on its gradient.

[0138] In this embodiment of the application, the gradient can be calculated using the Sobel operator. For example, the Sobel operator is specifically as follows:

[0139] Operators for horizontal gradients:

[0140] Vertical gradient operator:

[0141] Thus, assuming the sample value at position (x, y) in the prediction block is P x,y Then the horizontal gradient grad x and vertical gradient grad y The calculation is as follows: grad x =P x+1,y-1 +2*P x+1,y +P x+1,y+1 -P x-1,y-1 -2*P x-1,y -P x-1,y+1 (1) grad y =P x-1,y+1+2*P x,y+1 +P x+1,y+1 -P x-1,y-1 -2*P x,y-1 -P x+1,y -1 (2)

[0142] Here, the gradient intensity is denoted as amp, and for example, amp = abs(grad x )+abs(grad y ).

[0143] According to grad x and grad y Exporting a virtual intra-prediction mode can be achieved through a table lookup. For example, if abs(grad...) x ) equals 0 and abs(grad) y If abs(grad) is not equal to 0, it indicates the presence of horizontal texture, corresponding to intra-prediction mode 18 in VVC. y ) equals 0 and abs(grad) x If a value is not equal to 0, it indicates the presence of vertical texture, corresponding to intra-prediction mode 50 in VVC. In abs(grad x ) and abs(grad y If abs(grad) are not equal to 0, then... x ) equals abs(grad y ), and grad x and grad y If the symbols are the same, then it corresponds to intra-prediction mode 34 in VVC; if abs(grad x ) equals twice abs(grad) y ), and grad x and grad y If the symbols are the same, then it corresponds to intra-prediction mode 40 in VVC. Other cases can be determined by looking up a table using the same principle.

[0144] Analyzing all the points that need to be checked yields a result similar to the bar chart below. This represents the statistics of the number of points matched for each intra-prediction mode. Of course, the bar chart is only for understanding; in practice, it can be implemented in various simple forms, such as using an array. DIMD now selects the two highest-ranking intra-prediction modes from the bar chart, and adds the PLANAR mode, resulting in a weighted average of the predicted values ​​for three intra-prediction modes. The weights are related to the analysis results. For example, as shown in Figure 13, the three intra-prediction modes include M1 mode, M2 mode, and PLANAR mode. The predicted values ​​obtained for these three intra-prediction modes are set as Pred1, Pred2, and Pred3, respectively, and their weights are set as w1, w2, and w3, respectively. The specific calculation formula is as follows:

[0145] The final prediction block can be as follows:

[0146] In summary, DIMD uses gradient analysis of reconstructed samples to select intra-prediction modes, and can weight two intra-prediction modes plus a planar vector based on the analysis results. The advantage of DIMD is that if the current block selects the DIMD mode, it doesn't need to specify which intra-prediction mode is used; instead, the decoder derives this information automatically through the above process, saving overhead to some extent. Specifically, using multiple intra-prediction modes for weighting also improves compression performance, and theoretically, it can use even more intra-prediction modes for weighting. An example is using up to five intra-angle prediction modes weighted with a planar vector, or using up to five intra-angle prediction modes weighted with IBC or ITMP.

[0147] (vii) Intra Template Matching Prediction (ITMP).

[0148] ITMP can be considered a technique combining IBC and TM. As mentioned above, applying TM to inter-frames can reduce the overhead of encoding MV; similarly, using TM in IBC can reduce the overhead of encoding BV. One example is to eliminate the need to encode BV and directly use the matching block found by TM as the prediction block of the current block's ITMP mode.

[0149] Figure 14 shows an example of ITMP. The inverted L-shaped region at the top left corner of the current block serves as a template. The search is performed within the point-filled search area, which is the reconstructed region. The point-filled area shown in Figure 14 includes the current CTU in R1, the CTU on the top left of R2, the CTU on the top of R3, and the left-side CTU in R4. This is just an example; the search area will differ in actual applications. This example finds the best-matching block in R2.

[0150] (viii) Spatial Geometric Partitioning Mode (SGPM).

[0151] In the VVC video codec standard, there is an inter-frame prediction mode called Geometric Partitioning Mode (GPM). In the AVS3 video codec standard, there is an inter-frame prediction mode called Angular Weighted Prediction (AWP). Although these two modes have different names and implementations, they share common principles.

[0152] Traditional unidirectional prediction uses only one reference block of the same size as the current block. Traditional bidirectional prediction uses two reference blocks of the same size, and the sample value of each point in the prediction block is the average of the corresponding positions in the two reference blocks, meaning that all points in each reference block account for 50%. Bidirectional weighted prediction allows the proportions of the two reference blocks to be different; for example, all points in the first reference block may account for 75%, and all points in the second reference block may account for 25%. However, all points within the same reference block have the same proportion. Other optimization methods, such as decoder-side motion vector refinement (DMVR) and bidirectional optical flow (BIO), may cause some changes to the reference or prediction samples, but these are unrelated to the principles described above. BIO can also be abbreviated as BDOF. GPM or AWP also uses two reference blocks of the same size as the current block. However, some sample locations use 100% of the sample values ​​from the first reference block, while others use 100% from the second reference block. In the boundary or transition region, the sample values ​​from both reference blocks are used in proportion. The weights in the boundary region also gradually transition. The specific allocation of these weights is determined by the GPM or AWP pattern. The weight of each sample location is determined according to the GPM or AWP pattern. Of course, in some cases, such as when the block size is very small, some GPM or AWP patterns may not guarantee that some sample locations will use 100% of the sample values ​​from the first reference block, while others will use 100% from the second reference block. Alternatively, GPM or AWP can be considered to use two reference blocks of different sizes than the current block, i.e., taking a portion of each as a reference block. That is, the portion with a non-zero weight is used as the reference block, while the portion with a zero weight is discarded.

[0153] Figure 15 shows the weighting diagram of GPM in VVC for 64 modes on a square block. Black represents a weight of 0% for the position corresponding to the first reference block, white represents a weight of 100% for the position corresponding to the first reference block, and gray areas, depending on their shade, represent a weight value greater than 0% and less than 100% for the position corresponding to the first reference block. The weight value for the position corresponding to the second reference block is 100% minus the weight value for the position corresponding to the first reference block.

[0154] GPM and AWP use different methods to derive weights. GPM determines the angle and offset for each mode and then calculates the weight matrix for each mode. AWP first draws a one-dimensional weight line and then uses a method similar to intra-frame angle prediction to fill the entire matrix with the one-dimensional weight line.

[0155] It's important to note that earlier codec standards only used rectangular partitioning, whether for CU, PU, ​​or TU. GPM and AWP, however, achieved non-rectangular partitioning for prediction without such partitioning. GPM and AWP use a weighted mask of two reference blocks, i.e., the weight map or weight matrix mentioned above. This mask determines the weights of the two reference blocks when generating the prediction block. Simply put, part of the prediction block's location comes from the first reference block, and part comes from the second. The blending area is obtained by weighting the corresponding positions of the two reference blocks, resulting in a smoother transition. Since GPM and AWP do not divide the current block into two CUs or PUs along a dividing line, the residual transformations, quantizations, inverse transforms, and inverse quantizations after prediction are all treated as a single unit.

[0156] It should also be noted that GPM is an inter-frame technique in VVC, but it can also use intra-frame prediction. Both prediction modes of GPM can be inter-frame prediction modes, one can be inter-frame prediction mode and one can be intra-frame prediction mode, or both can be intra-frame prediction modes.

[0157] The SGPM mode in ECM uses this weighted mask, or weight matrix, in intra-frame prediction. It combines the prediction values ​​of two intra-frame prediction modes into the prediction value of SGPM using the weight matrix, which can produce more complex textures compared to a single prediction mode.

[0158] Because one "partition" mode and two intra-prediction modes are needed, the general logic would be to write the syntax elements indicating these three modes into the bitstream, as shown in Table 1: partition_mode_idx, intra_pred_mode0_idx, and intra_pred_mode1_idx. However, to reduce the overhead of this information, SGPM uses the templates surrounding the current block to sort the combinations of the three modes, obtaining a candidate list of combined modes. In the bitstream, only the candidate index of SGPM needs to be written. At the decoding end, SGPM can construct the candidate list of combined modes and derive one "partition" mode and two intra-prediction modes, partition_mode_idx, intra_pred_mode0_idx, and intra_pred_mode1_idx, based on the candidate index of SGPM, such as sgpm_cand_idx in Table 2.

[0159] Table 1

[0160] Table 2

[0161] Here, for sgpm_cand_idx,

[0162] (ix) Extrapolation filter based Intra Prediction (EIP).

[0163] In ECM, there is an EIP mode. Figures 16A, 16B, and 16C show several typical filters for the EIP mode. The sample positions in the grid-filled areas represent the input samples (input of EIP), and the sample positions in the white-filled areas represent the output samples (output of EIP). Taking the first square filter as an example, the predicted value of the sample in the lower right corner is usually required. The values ​​of the 15 samples to its left, upper left, and top are known. Each grid-filled position in the filter has a preset coefficient. By multiplying the sample value at each input position by the corresponding coefficient, summing all the multiplication results, and then normalizing, the predicted value at the desired position can be obtained.

[0164] The EIP mode may provide pre-trained filters or filters trained based on reconstructed sample regions surrounding the current block. Figures 17A, 17B, and 17C illustrate that the EIP mode can select from three types of reconstructed sample regions for filter training. Assuming the current block and the reconstructed sample regions used to train the filter coefficients follow similar patterns, filters trained using these reconstructed sample regions can be applied to the current block.

[0165] Furthermore, the following section provides a detailed introduction to transformation techniques.

[0166] During encoding, commonly used hybrid coding frameworks first perform prediction, using spatial or temporal correlation properties to obtain an image that is the same as or similar to the current block. While it's possible for a single block to be completely identical to the current block, it's difficult to guarantee this for all blocks in a video, especially natural videos or videos captured by a camera. Irregular motion, distortion, occlusion, brightness variations, etc., in videos are difficult to predict completely. Therefore, hybrid coding frameworks subtract the predicted image from the original image of the current block to obtain a residual image, or vice versa. Residual blocks are usually much simpler than the original image, thus prediction can significantly improve compression efficiency. The residual block is not directly encoded but is usually transformed first. The transformation converts the residual image from the spatial domain to the frequency domain, removing correlations. After transforming the residual image to the frequency domain, since most energy is concentrated in the low-frequency region, the non-zero coefficients are mostly concentrated in the upper left corner. Quantization is then used for further compression. Furthermore, because the human eye is not sensitive to high frequencies, a larger quantization step size can be used in the high-frequency region.

[0167] Figure 18 is a schematic diagram of a DCT transform. As shown in Figure 18, after the DCT transform, only the upper left corner of the original image has non-zero coefficients. Of course, this example performs a DCT transform on the entire image. In video encoding and decoding, images are processed by dividing them into blocks, so the transform is also performed on a block-by-block basis.

[0168] Transformations are very useful in typical video compression, but not all blocks need to be transformed. In some cases, transformation is not as effective as no transformation. Therefore, in some standards such as VVC, the encoder can choose whether to use a transformation for the current block. Among them, DCT2 (DCT-II) is the most commonly used transformation in video compression standards, and its base image is shown in Figure 19.

[0169] In addition, DCT8 type (DCT-VIII) and DST7 type (DST-VII) can also be used in VVC. The basic formulas for these transformations are shown in Table 3, which shows the basic transformation formulas for DCT2, DCT8 and DST7 with N-point input.

[0170] Table 3

[0171] Since images are two-dimensional, and the computational and memory overhead of directly performing two-dimensional transformations was unacceptable given the hardware limitations at the time, the DCT2, DCT8, and DST7 transformations used in the standard were all split into two steps: a horizontal transformation and a vertical transformation. For example, the horizontal transformation was performed first, followed by the vertical transformation, or vice versa.

[0172] (a) Multiple Transformation Selection MTS.

[0173] VVC supports transform cores such as DCT2, DCT8, and DST7. The DCT2, DCT8, and DST7 transform cores used in VVC are horizontally and numerically separable transform cores, which can be applied independently to horizontal or vertical transforms. For a block, the encoder can select a suitable transform core and transmit the index to the bitstream; the decoder determines the inverse transform core based on the index. Different transform cores can be selected for the horizontal and vertical directions, such as using DCT8 for the horizontal direction and DST7 for the vertical direction. This technique is generally called MTS.

[0174] VVC uses a syntax element `mts_idx` to determine the transform kernel of the basic transform. As shown in Table 4, `trTypeHor` represents the transform kernel for horizontal transforms, and `trTypeVer` represents the transform kernel for vertical transforms. A value of 0 for `trTypeHor` and `trTypeVer` indicates a DCT2 type transform, 1 indicates a DST7 type transform, and 2 indicates a DCT8 type transform. If `mts_idx` does not exist, its value is assumed to be 0.

[0175] Table 4

[0176] (ii) Low-frequency inseparable transformation LFNST.

[0177] The transformation methods described above are effective for horizontal and vertical textures, but less so for diagonal textures. Indeed, horizontal and vertical textures are the most common, making these transformation methods very useful for improving compression efficiency. As the demand for compression efficiency continues to increase, further improvements could be made if diagonal textures could be handled more effectively.

[0178] To more effectively handle residuals in diagonal textures, VVC uses the LFNST transform. Transformations such as DCT2, DCT8, and DST7 are called primary transforms. At the VVC encoder, LFNST is used after the DCT2 transform and before quantization. At the VVC decoder, LFNST is used after inverse quantization and before the inverse DCT2 transform. Because it's a transformation based on DCT2 (the primary transform), LFNST is a quadratic transform. Figure 20 shows the encoding / decoding process without LFNST (quadratic transform). As shown in Figure 20, the encoder's process is as follows: Transform 211 ---- Quantization 212 ---- Entropy encoding 213, writing the obtained encoded bits into the bitstream; the decoder's process is as follows: Entropy decoding 214 ---- Inverse quantization 215 ---- Inverse transform 216. Figure 21 is a schematic diagram of the encoding and decoding process with LFNST (secondary transform). As shown in Figure 21, the process at the encoding end is as follows: DCT2 transform (basic transform) 221 ---- LFNST transform (secondary transform) 222 ---- quantization 223 ---- entropy encoding 224, writing the obtained encoded bits into the bitstream; the process at the decoding end is as follows: entropy decoding 225 ---- dequantization 226 ---- inverse LFNST transform (secondary transform) 227 ---- inverse DCT2 transform (basic transform) 228. It should be noted that the encoding end can also directly dequantize the stored quantization coefficients without entropy decoding, because entropy encoding is lossless.

[0179] Figure 22 is a detailed encoding / decoding flowchart involving LFNST (Secondary Transform). At the encoding end, LFNST performs a secondary transformation on the low-frequency coefficients in the upper left corner after the basic transformation. The basic transformation concentrates energy in the upper left corner by decorrelating the image. The secondary transformation further decorrelates the low-frequency coefficients of the basic transformation, as shown in Figure 22. At the encoding end, 16 coefficients are input to a 4×4 LFNST, outputting 8 coefficients; 48 coefficients are input to an 8×8 LFNST, outputting 8 coefficients for the 8×8 block and 16 coefficients for other blocks. At the decoding end, 8 coefficients are input to a 4×4 inverse LFNST, outputting 16 coefficients; 8 coefficients are input to the 8×8 block, 16 coefficients to other blocks, and the coefficients are input to the 8×8 inverse LFNST, outputting 48 coefficients.

[0180] Figure 23 shows some basis images of LFNST in VVC. Only the two lowest-frequency basis images of each transform kernel in each transform kernel group are shown in Figure 23. Some obvious diagonal textures can be observed. Besides transform kernels optimized for certain diagonal textures, LFNST also has transform kernels optimized for flat gradient textures, such as transform kernel group 0 of LFNST in VVC.

[0181] Angle prediction uses a specified angle to tile the values ​​of the reference sample onto the current block as the predicted value. This means that the predicted block will have obvious directional texture, and the residual of the current block after angle prediction will also statistically show obvious angular characteristics. Therefore, the transform kernel selected by LFNST can be bound to the intra-prediction mode. That is, after determining the intra-prediction mode, LFNST can only use a set of transform kernels corresponding to the intra-prediction mode.

[0182] Specifically, the LFNST in VVC has a total of 4 transform kernels, with 2 transform kernels selectable in each group. Table 5 shows the correspondence between intra-frame prediction modes and transform kernel groups. Note that the cross-component prediction modes used for chroma intra-frame prediction are 81 to 83, while luma intra-frame prediction does not have these modes. The LFNST transform kernels can be transposed to handle more angles with a single transform kernel group. For example, modes 13 to 23 and 45 to 55 both correspond to transform kernel group 2, but modes 13 to 23 are clearly closer to horizontal modes while modes 45 to 55 are clearly closer to vertical modes.

[0183] Table 5

[0184] VVC's LFNST has four transform kernels, and which one is used is specified based on the intra-prediction mode. This utilizes the correlation between the intra-prediction mode and the LFNST transform kernels, thereby reducing the transmission of the selected LFNST transform kernel in the bitstream. Whether the current block will use LFNST, and if so, whether to use the first or second kernel in a group, needs to be determined by the bitstream and certain conditions.

[0185] In subsequent ECM technology evolution, LFNST was further expanded. LFNST has more transform kernel groups, 35 in ECM. The correspondence between the transform kernel group index (LFNST set index) and the intra prediction mode (Intra pred.mode) is shown in Table 6. Each transform kernel group is more efficient for textures at the corresponding angle. Here, each transform kernel group can select 3 transform kernels.

[0186] Table 6

[0187] (iii) Inseparable basic transformation NSPT.

[0188] LFNST is a horizontally and vertically inseparable transformation. Because it involves a quadratic transformation, DCT2 can be called the fundamental transformation. Performing DCT2 first and then LFNST can be considered a compromise between performance and complexity. Directly performing the inseparable fundamental transformation is more efficient, but it has higher complexity, such as higher computational cost and higher storage space for the transformation kernel.

[0189] In ECM10, some small patches can use NSPT, while larger patches still use DCT2+LFNST. Patch sizes include 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 8×32, and 32×8. In ECM10, NSPT also matches transform kernel groups according to the intra-frame prediction mode, using a matching method similar to LFNST. Each transform kernel group has three selectable transform kernels. For example, an 8×8 base image of an NSPT in ECM10 is shown in Figure 24. Figure 24 corresponds to inter-frame angle prediction mode 7, demonstrating its superior handling of textures at the corresponding angles.

[0190] In summary, for more complex intra-prediction modes (such as TIMD and DIMD), TIMD and DIMD share some similarities. Both utilize the reconstructed regions adjacent to the current block for analysis and derivation of the intra-prediction mode to be used. These adjacent regions have a strong correlation with the current block, thus enabling the derivation of a reasonable intra-prediction mode in many cases. Furthermore, both support combinations of multiple intra-prediction modes. Combining multiple intra-prediction modes can reduce prediction errors because they do not require syntax elements in the bitstream to indicate which intra-prediction mode is used. In other words, even using an additional intra-prediction mode for weighting does not incur additional indication costs.

[0191] However, in related technologies, both TIMD and DIMD require a certain amount of computation to derive the intra-prediction mode used in the current block. For example, TIMD needs to generate prediction values ​​on the template and calculate the matching cost, while DIMD needs to calculate the gradient, which increases the computational complexity and affects the encoding and decoding efficiency.

[0192] Based on this, embodiments of this application provide an encoding method that determines one or more pre-encoded blocks; determines statistical results of at least two candidate intra-frame prediction modes in a first mode set based on the prediction modes used by each of the one or more pre-encoded blocks; determines at least two intra-frame prediction modes for the current block in the first mode set based on the statistical results; and predicts the current block based on the at least two intra-frame prediction modes to determine the predicted value of the current block. Embodiments of this application also provide a decoding method that determines one or more pre-decoded blocks; determines statistical results of at least two candidate intra-frame prediction modes in a first mode set based on the prediction modes used by each of the one or more pre-encoded blocks; determines at least two intra-frame prediction modes for the current block in the first mode set based on the statistical results; and predicts the current block based on the at least two intra-frame prediction modes to determine the predicted value of the current block.

[0193] In this way, by using the relevant blocks that have been encoded and decoded to perform statistics, the statistical results of these relevant blocks are obtained, and at least two intra-prediction modes to be used in the current block are deduced based on the statistical results. Since it is not necessary to explicitly indicate which intra-prediction mode to use, the overhead in the bitstream can be saved. Moreover, using at least two intra-prediction modes for weighted prediction can also improve the accuracy of prediction and reduce prediction error to a certain extent. In addition, deducing at least two intra-prediction modes to be used in the current block based on statistical results can also avoid the need for gradient calculation to derive intra-prediction modes, such as in DIMD mode, thus reducing computational complexity, thereby improving encoding and decoding efficiency and improving compression performance.

[0194] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0195] Figure 25 is a schematic diagram of a video encoding and decoding network architecture provided in an embodiment of this application. As shown in Figure 25, the network architecture includes one or more electronic devices 31 to 3N and a communication network 01, wherein the electronic devices 31 to 3N can perform video interaction through the communication network 01. The electronic devices can be various types of devices with video encoding and decoding capabilities, such as mobile phones, tablets, personal computers, personal digital assistants, navigators, digital phones, video phones, televisions, sensing devices, servers, etc., and this embodiment of the application does not limit the scope of the application.

[0196] This application provides a network architecture for a video encoding / decoding system that includes decoding and encoding methods. The decoder or encoder in this application can be the aforementioned electronic device. That is, the electronic device in this application has video encoding / decoding capabilities and generally includes a video encoder and a video decoder.

[0197] Figure 26 is a schematic block diagram of an encoder system according to an embodiment of this application. As shown in Figure 26, the encoder 100 may include: a segmentation unit 101, a prediction unit 102, a first adder 107, a transform unit 108, a quantization unit 109, an inverse quantization unit 110, an inverse transform unit 111, a second adder 112, a filtering unit 113, a Decoded Picture Buffer (DPB) unit 114, and an entropy coding unit 115. Here, the input of the encoder 100 can be a video composed of a series of images or a single static image, and the output of the encoder 100 can be a bitstream (also called a "bitstream") representing a compressed version of the input video.

[0198] The segmentation unit 101 segments the images in the input video into one or more Coding Tree Units (CTUs). The segmentation unit 101 divides the image into multiple tiles, and can further divide a tile into one or more bricks. Here, a tile or a brick can include one or more complete and / or partial CTUs. Additionally, the segmentation unit 101 can form one or more slices, where a slice can include one or more tiles arranged in raster order in the image, or one or more tiles covering a rectangular area of ​​the image. The segmentation unit 101 can also form one or more sub-images, where a sub-image can include one or more slices, tiles, or bricks.

[0199] During the encoding process of encoder 100, segmentation unit 101 transmits the CTU to prediction unit 102. Typically, prediction unit 102 may consist of block segmentation unit 103, motion estimation (ME) unit 104, motion compensation (MC) unit 105, and intra-prediction unit 106. Specifically, block segmentation unit 103 iteratively uses quadtree segmentation, binary tree segmentation, and ternary tree segmentation to further divide the input CTU into smaller coding units (CUs). Prediction unit 102 can use ME unit 104 and MC unit 105 to obtain inter-frame prediction blocks of the CUs. Intra-prediction unit 106 can use various intra-prediction modes, including MIP modes, to obtain intra-frame prediction blocks of the CUs. In the example, rate-distortion optimized motion estimation can be invoked by ME unit 104 and MC unit 105 to obtain inter-frame prediction blocks, and rate-distortion optimized mode determination can be invoked by intra-prediction unit 106 to obtain intra-frame prediction blocks. Prediction unit 102 outputs the predicted block of the CU. First adder 107 calculates the difference between the CU in the output of segmentation unit 101 and the predicted block of the CU, i.e., the residual CU. Transform unit 108 reads the residual CU and performs one or more transform operations on the residual CU to obtain coefficients. Quantization unit 109 quantizes the coefficients and outputs quantization coefficients (i.e., levels). Inverse quantization unit 110 performs scaling operations on the quantization coefficients to output reconstructed coefficients. Inverse transform unit 111 performs one or more inverse transforms corresponding to the transforms in transform unit 108 and outputs the reconstructed residual. Second adder 112 calculates the reconstructed CU by adding the reconstructed residual to the predicted block of the CU from prediction unit 102. Second adder 112 also sends its output to prediction unit 102 as an intra-frame prediction reference. After all CUs in the image or sub-image are reconstructed, filtering unit 113 performs loop filtering on the reconstructed image or sub-image. Here, the filtering unit 113 includes one or more filters, such as a deblocking filter, a Sample Adaptive Offset (SAO) filter, an Adaptive Loop Filter (ALF), a Luma Mapping with Chroma Scaling (LMCS) filter, and a neural network-based filter. Alternatively, when the filtering unit 113 determines that the CU is not used as a reference for encoding other CUs, the filtering unit 113 performs loop filtering on one or more target samples in the CU. The output of the filtering unit 113 is a decoded image or sub-image, which is buffered in the DPB unit 114. The DPB unit 114 outputs the decoded image or sub-image according to timing and control information.Here, the image stored in DPB unit 114 can also be used as a reference for prediction unit 102 to perform inter-frame prediction or intra-frame prediction. Finally, entropy coding unit 115 converts the parameters (such as control parameters and supplementary information) necessary for decoding the image from encoder 100 into binary form, and writes such binary form into the bitstream according to the syntax structure of each data unit, which is the final output bitstream of encoder 100.

[0200] Furthermore, encoder 100 may be a first memory having a first processor and a computer program for recording. When the first processor reads and runs the computer program, encoder 100 reads the input video and generates a corresponding bitstream. Alternatively, encoder 100 may also be a computing device having one or more chips. These units, implemented as integrated circuits on the chips, have connection and data exchange functions similar to the corresponding units in Figure 26.

[0201] Figure 27 is a schematic block diagram of a decoder system according to an embodiment of this application. As shown in Figure 27, the decoder 200 may include: a parsing unit 201, a prediction unit 202, an inverse quantization unit 205, an inverse transform unit 206, an adder 207, a filtering unit 208, and a decoded image buffer unit 209. Here, the input of the decoder 200 is a bitstream representing a compressed version of a video or a still image, and the output of the decoder 200 may be a decoded video composed of a series of images or a decoded still image.

[0202] The input bitstream to the decoder 200 can be the bitstream generated by the encoder 100. The parsing unit 201 parses the input bitstream and obtains the values ​​of the syntax elements. The parsing unit 201 converts the binary representation of the syntax elements into numerical values ​​and sends these values ​​to units in the decoder 200 to obtain one or more decoded images. The parsing unit 201 can also parse one or more syntax elements from the input bitstream to display the decoded images.

[0203] During the decoding process of decoder 200, parsing unit 201 sends the values ​​of syntax elements and one or more variables set or determined based on the values ​​of syntax elements for obtaining one or more decoded images to units in decoder 200. Prediction unit 202 determines the prediction block for the current decoded block (e.g., CU). Here, prediction unit 202 may include motion compensation unit 203 and intra-frame prediction unit 204. Specifically, when an inter-frame decoding mode is indicated for decoding the current decoded block, prediction unit 202 passes relevant parameters from parsing unit 201 to motion compensation unit 203 to obtain inter-frame prediction blocks; when an intra-frame prediction mode (including MIP mode indicated based on MIP mode index value) is indicated for decoding the current decoded block, prediction unit 202 passes relevant parameters from parsing unit 201 to intra-frame prediction unit 204 to obtain intra-frame prediction blocks. Dequantization unit 205 has the same function as dequantization unit 110 in encoder 100. Dequantization unit 205 performs scaling operations on quantization coefficients (i.e., levels) from parsing unit 201 to obtain reconstruction coefficients. The inverse transform unit 206 has the same function as the inverse transform unit 111 in the encoder 100. The inverse transform unit 206 performs one or more transform operations (i.e., the inverse operations of one or more transform operations performed by the inverse transform unit 111 in the encoder 100) to obtain the reconstruction residual. The adder 207 performs an addition operation on its inputs (the predicted block from the prediction unit 202 and the reconstruction residual from the inverse transform unit 206) to obtain the reconstruction block of the current decoded block. The reconstruction block is also sent to the prediction unit 202 as a reference for other blocks encoded in intra-frame prediction mode.

[0204] After all CUs in an image or sub-image are reconstructed, filtering unit 208 performs loop filtering on the reconstructed image or sub-image. Filtering unit 208 includes one or more filters, such as deblocking filters, sampling adaptive compensation filters, adaptive loop filters, luminance mapping and chroma scaling filters, and neural network-based filters. Alternatively, when filtering unit 208 determines that a reconstructed block is not used as a reference for decoding other blocks, filtering unit 208 performs loop filtering on one or more target samples in the reconstructed block. Here, the output of filtering unit 208 is a decoded image or sub-image, which is buffered in DPB unit 209. DPB unit 209 outputs the decoded image or sub-image based on timing and control information. The image stored in DPB unit 209 can also be used as a reference for performing inter-frame prediction or intra-frame prediction by prediction unit 202.

[0205] Furthermore, the decoder 200 can be a second memory having a second processor and a computer program for recording. When the first processor reads and runs the computer program, the decoder 200 reads the input bitstream and generates the corresponding decoded video. Alternatively, the decoder 200 can also be a computing device having one or more chips. These units, implemented as integrated circuits on the chips, have connection and data exchange functions similar to the corresponding units in Figure 27.

[0206] It should also be noted that when the embodiments of this application are applied to the encoder 100, the "current block" specifically refers to the block to be encoded in the video image (which can also be simply referred to as the "encoded block"); when the embodiments of this application are applied to the decoder 200, the "current block" specifically refers to the block to be decoded in the video image (which can also be simply referred to as the "decoded block").

[0207] In one embodiment of this application, Figure 28 is a schematic flowchart of a decoding method provided in this application. As shown in Figure 28, the method may include:

[0208] S2801, determine one or more preset blocks that have been decoded.

[0209] In this embodiment, the method is applied to the decoder. Specifically, based on the structure of the decoder 200 shown in Figure 27, the decoding method is mainly applied to the intra-prediction unit 204 and the inverse transform unit 206 in Figure 27. When the current block uses Occurrence-Based Intra Prediction (OBIP) mode, the intra-prediction mode used by the current block can be derived by statistically analyzing the prediction modes used by related blocks in the decoded sequence. Thus, compared to TIMD and DIMD modes, which require a certain amount of computation to derive the intra-prediction mode, the OBIP mode saves computation and improves compression performance.

[0210] It should also be noted that when using OBIP mode in the current block, it is first necessary to determine one or more preset blocks that have been decoded, so as to facilitate the subsequent determination of the statistical results of at least two candidate intra-prediction modes in the first mode set.

[0211] In one possible implementation, determining one or more pre-defined blocks that have been decoded may include: determining at least one candidate block that has been decoded, and determining one or more pre-defined blocks based on at least one candidate block.

[0212] In the embodiments of this application, a candidate block may include: adjacent blocks of the current block, and / or non-adjacent blocks of the current block. That is, one or more preset blocks may include only adjacent blocks of the current block; or, may include only non-adjacent blocks of the current block; or, may be composed of both adjacent and non-adjacent blocks of the current block.

[0213] Understandably, when the current block uses OBIP mode, it is possible to statistically analyze the adjacent and non-adjacent blocks of the current block. Figure 29 is a schematic diagram of the spatial relationship between adjacent and non-adjacent blocks of the current block according to an embodiment of this application. As shown in Figure 29, the black-filled blocks represent the current block, blocks numbered 1 to 7 are adjacent blocks of the current block, and blocks numbered 8 to 25 are non-adjacent blocks of the current block. Adjacent blocks refer to blocks spatially adjacent to the current block, and non-adjacent blocks refer to blocks spatially not adjacent to the current block.

[0214] For example, suppose the top-left corner of the current block is (x, y), the width of the current block is width, and the height is height. Then, for the adjacent blocks of the current block, adjacent block 1 contains the block with coordinates (x-1, y-1), adjacent block 2 contains the block with coordinates (x, y-1), adjacent block 3 contains the block with coordinates (x-1, y), adjacent block 4 contains the block with coordinates (x+width-1, y-1), adjacent block 5 contains the block with coordinates (x+width, y-1), adjacent block 6 contains the block with coordinates (x-1, y+height-1), and adjacent block 7 contains the block with coordinates (x-1, y+height).

[0215] Non-adjacent blocks can also be positioned according to certain rules. For example, the position of a non-adjacent block can be calculated based on (x, y), width, and height. In Figure 29, the distance between each cell in the vertical direction is height, and the distance between each cell in the horizontal direction is width. Therefore, for the non-adjacent blocks of the current block, non-adjacent block 8 contains the coordinates (x-width-1, y-height-1), non-adjacent block 9 contains the coordinates (x+2*width-1, y-height-1), non-adjacent block 10 contains the coordinates (x-width-1, y+2*height-1), non-adjacent block 11 contains the coordinates (x-2*width-1, y-2*height-1), non-adjacent block 12 contains the coordinates (x+width / 2, y-2*height-1), and so on. It should be noted that the block containing a certain coordinate refers to the CU or PU containing that coordinate.

[0216] Thus, in this embodiment of the application, for a decoded candidate block, it may be possible to use only adjacent blocks, or only non-adjacent blocks, or both adjacent and non-adjacent blocks, or specific blocks from these adjacent and non-adjacent blocks, without any limitations.

[0217] In another possible implementation, determining one or more decoded preset blocks may include: determining a preset range based on the current block; and determining one or more preset blocks based on all decoded blocks within the preset range.

[0218] In this embodiment, the preset range can be arbitrarily set within the decoded region. For example, assuming the top-left corner of the current block is (x, y), the width of the current block is width, and the height is height, the preset range can be set based on the current block's position and its width and height.

[0219] It can also be understood that when the current block uses OBIP mode, all blocks within a preset range can be statistically analyzed. Figure 30 is a schematic diagram of the spatial relationship between the current block and the preset range provided by an embodiment of this application. As shown in Figure 30, the black-filled block represents the current block, and the dotted-filled part represents the preset range. For example, a possible preset range is: all decoded ranges within a rectangular area with the upper left corner at (x-2*width, y-2*height), the upper right corner at (x+3*width-1, y-2*height), and the lower left corner at (x-2*width, y+3*height-1). If the minimum size of the CU or PU is 4×4, then all 4×4 blocks within the search range can be statistically analyzed. For example, the first 4×4 block in the upper left corner is a block containing the coordinates (x-2*width, y-2*height). Increasing by 4 in the horizontal direction allows scanning all 4×4 blocks in the same row, and increasing by 4 in the vertical direction allows scanning all columns. It should be noted that the block containing a certain coordinate here refers to the CU or PU containing that coordinate.

[0220] Thus, in this embodiment of the application, after determining the preset range that has been decoded, all blocks found within the preset range can be identified as one or more preset blocks.

[0221] S2802, based on the prediction modes used by one or more preset blocks, determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set.

[0222] In this embodiment, the first mode set may include at least two candidate intra-prediction modes. The intra-prediction modes may include many modes, such as the PLANA mode, DC mode, 65 angle prediction modes, and MIP mode already existing in VVC, and the DIMD mode, TIMD mode, SGPM mode, ITMP mode, and OBIP mode newly added in ECM. A first mode set can be defined here, and in some embodiments, the first mode set includes at least one of the following: angle prediction mode, DC mode, and PLANA mode.

[0223] For example, the first mode set may contain only angle prediction modes, or it may contain both angle prediction modes and PLANAR modes, or it may contain angle prediction modes, DC modes, and PLANAR modes, etc., without any limitation. That is to say, one possible implementation is that the first mode set contains only angle prediction modes, and the number of these modes may be 65, or it may be more due to the existence of more granular angle prediction modes. Another possible implementation is that the first mode set contains only angle prediction modes, DC modes, and PLANAR modes. The following is a detailed description of the example where the first mode set contains angle prediction modes, DC modes, and PLANAR modes. In this case, the first mode set corresponds exactly to the 67 intra-frame prediction modes 0 to 66 in VVC.

[0224] In this embodiment, for the OBIP mode, a data structure can be constructed to record the occurrence frequency of each candidate intra-frame prediction mode in the first mode set. For example, this data structure can be an array, denoted as occurrence

[0067] . It should be noted that this example uses a first mode set with 67 modes; if the number of modes is different, the array length will be adjusted accordingly. Furthermore, all values ​​in occurrence

[0067] are initialized to 0.

[0225] In this embodiment, after determining one or more decoded preset blocks, the occurrence of modes in the first mode set within the one or more preset blocks can be statistically analyzed, i.e., the statistical results of at least two candidate intra-frame prediction modes in the first mode set can be determined. In a specific embodiment, this may include: determining the prediction mode used by each of the one or more preset blocks; performing candidate intra-frame prediction mode statistics based on the prediction modes used by each of the one or more preset blocks, and determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set.

[0226] It should be noted that, in the embodiments of this application, after determining a preset block (e.g., CU or PU) containing certain coordinates within the decoded region, the prediction mode used by the preset block can be determined. For example, if the preset block is an intra-frame prediction block, then the intra-frame prediction mode used by the preset block can be determined; if the preset block is an inter-frame prediction block, then the inter-frame prediction mode used by the preset block can be determined. Thus, by performing candidate intra-frame prediction mode statistics based on the prediction modes used by each of these one or more preset blocks, the statistical results of at least two candidate intra-frame prediction modes in the first mode set can be determined.

[0227] It should also be noted that, in the embodiments of this application, taking the first preset block as an example, for determining the statistical results of at least two candidate intra-prediction modes in the first mode set, the method may include: determining the intra-prediction mode used by the first preset block; when the intra-prediction mode is a candidate intra-prediction mode in the first mode set, performing cumulative calculation on the intra-prediction mode to determine the statistical results of the intra-prediction mode in the first mode set.

[0228] Understandably, in the embodiments of this application, the first preset block can be any one of one or more preset blocks. Taking the first preset block as an example, the first preset block can be an inter-frame prediction block or an intra-frame prediction block. The processing situation where the first preset block is an inter-frame prediction block will be described in detail below.

[0229] In one possible implementation, for determining the statistical results of at least two candidate intra-prediction modes in the first mode set, the method may include: when the first preset block is an inter-prediction block, not counting the first preset block.

[0230] In this embodiment, the first preset block can be any one of one or more preset blocks. If the first preset block is an inter-frame prediction block, it means that the first preset block does not use the modes in the first mode set for prediction. In this case, the first preset block can be ignored (or the first preset block can be skipped). That is, no cumulative calculation is performed on any mode in the first mode set. Instead, candidate intra-frame prediction mode statistics can be performed based on the prediction modes used by other preset blocks outside the first preset block, thereby determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set.

[0231] In another possible implementation, for the statistical results of determining at least two candidate intra-prediction modes in the first mode set, as shown in Figure 31, the method may include:

[0232] S3101, when the first preset block is an inter-frame prediction block, determine the first intra-frame prediction mode derived from the candidate samples of the first preset block.

[0233] S3102, perform cumulative calculation on the first intra-frame prediction mode to determine the statistical results of the first intra-frame prediction mode in the first mode set.

[0234] In this embodiment, the first preset block can be any one of one or more preset blocks. If the first preset block is an inter-frame prediction block, it indicates that the first preset block does not use the modes in the first mode set for prediction. In this case, the first intra-frame prediction mode of the first preset block can be derived based on the candidate samples of the first preset block. Then, by accumulating the first intra-frame prediction mode, the statistical result of the first intra-frame prediction mode in the first mode set can be determined. Specifically, the current cumulative value of the first intra-frame prediction mode can be determined by accumulating the cumulative value corresponding to the first intra-frame prediction mode, and the cumulative value corresponding to the first intra-frame prediction mode can be updated based on the current cumulative value. The judgment of the next preset block continues until all one or more preset blocks are traversed, and the final cumulative value corresponding to the first intra-frame prediction mode is determined as the statistical result of the first intra-frame prediction mode in the first mode set.

[0235] In some embodiments, the method may further include: when the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, performing an accumulation operation on the first intra-frame prediction mode to determine the statistical result of the first intra-frame prediction mode in the first mode set. That is, after deriving the first intra-frame prediction mode based on candidate samples from the first preset block, if the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, then an accumulation operation can be performed on the first intra-frame prediction mode to determine the statistical result of the first intra-frame prediction mode in the first mode set.

[0236] In some embodiments, the method may further include: determining candidate samples of the first preset block based on at least a portion of the samples within the reconstruction block of the first preset block; or, determining candidate samples of the first preset block based on at least a portion of the samples within the prediction block of the first preset block.

[0237] In this embodiment, when the first preset block is an inter-frame prediction block, gradient statistics can be performed based on the candidate samples of the first preset block to determine the first intra-frame prediction mode of the first preset block. Here, the candidate samples can be at least a portion of the samples in the reconstructed block or prediction block of the block. That is, gradient statistics can be performed based on at least a portion of the samples in the reconstructed block or prediction block of the block to derive the first intra-frame prediction mode. Figure 32 is a schematic diagram of the structure of a candidate sample provided in this embodiment. As shown in Figure 32, an 8×8 block is provided, and the first intra-frame prediction mode can be derived using the 6×6 candidate samples within it. That is, all points in the reconstructed block or prediction block except for the outermost row and column can be used as candidate samples (filled with a grid).

[0238] Thus, if the first preset block is an inter-frame prediction block, since it does not use any modes from the first mode set for prediction, it can be directly ignored from the statistics, meaning no cumulative calculation is performed on any modes from the first mode set. Alternatively, the first intra-frame prediction mode can be derived from the reconstructed or predicted blocks of this block using gradient statistics, and the derived first intra-frame prediction mode belongs to one of the candidate intra-frame prediction modes within the first mode set. Assuming the mode index corresponding to the first intra-frame prediction mode is X, then the occurrence[X] can be cumulatively calculated.

[0239] It is also understandable that, in the embodiments of this application, the first preset block is still taken as an example, and the processing of the first preset block as an intra-frame prediction block will be described in detail below.

[0240] In one possible implementation, for determining the statistical results of at least two candidate intra-prediction modes in the first mode set, the method further includes: when the first preset block is an intra-prediction block, determining the second intra-prediction mode used by the first preset block; when the second intra-prediction mode is a candidate intra-prediction mode in the first mode set, performing an accumulation operation on the second intra-prediction mode to determine the statistical results of the second intra-prediction mode in the first mode set.

[0241] In this embodiment, if the first preset block is an intra-prediction block, and the second intra-prediction mode used by the first preset block is a candidate intra-prediction mode in the first mode set, assuming the mode index corresponding to the second intra-prediction mode is X, then the accumulation operation is also performed on occurrence[X]. Here, the second intra-prediction mode is one of the candidate intra-prediction modes in the first mode set.

[0242] In another possible implementation, if the second intra-frame prediction mode is not a candidate intra-frame prediction mode in the first mode set, and the statistical results for determining at least two candidate intra-frame prediction modes in the first mode set are as follows: when the second intra-frame prediction mode is a candidate intra-frame prediction mode in the second mode set, determine at least two third intra-frame prediction modes corresponding to the second intra-frame prediction mode; when at least two third intra-frame prediction modes are at least two candidate intra-frame prediction modes in the first mode set, perform cumulative calculation on the at least two third intra-frame prediction modes to determine the statistical results of at least two third intra-frame prediction modes in the first mode set.

[0243] In this application embodiment, the second mode set is different from the first mode set, and the second mode set includes modes that use at least two intra-frame prediction modes for combined prediction.

[0244] In one specific embodiment, the mode that uses at least two intra-frame prediction modes for combined prediction can be DIMD mode, TIMD mode, SGPM mode and OBIP mode; wherein, the intra-frame prediction mode includes, but is not limited to, DC mode, PLANA mode and angle prediction mode.

[0245] In other words, in a more specific embodiment, the second set of modes may include at least one of the following: DIMD mode, TIMD mode, SGPM mode, and OBIP mode.

[0246] In this embodiment, if the first preset block uses DIMD mode, TIMD mode, SGPM mode, OBIP mode, etc., it may use multiple candidate intra-prediction modes from the first mode set. In this case, each candidate intra-prediction mode is accumulated. For example, assuming that the mode uses candidate intra-prediction modes X1 and X2 from the first mode set, then occurrence[X1] and occurrence[X2] are accumulated respectively.

[0247] In another possible implementation, the second intra-prediction mode is not a candidate intra-prediction mode in the first mode set. For the statistical results of determining at least two candidate intra-prediction modes in the first mode set, the method further includes: when the second intra-prediction mode is a candidate intra-prediction mode in the third mode set, the first preset block is not counted.

[0248] In another possible implementation, if the second intra-prediction mode is not a candidate intra-prediction mode in the first mode set, and the statistical results for determining at least two candidate intra-prediction modes in the first mode set are as follows: when the second intra-prediction mode is a candidate intra-prediction mode in the third mode set, determine a fourth intra-prediction mode derived from candidate samples of the first preset block; perform an accumulation operation on the fourth intra-prediction mode to determine the statistical results of the fourth intra-prediction mode in the first mode set. Alternatively, the method further includes: when the second intra-prediction mode is a candidate intra-prediction mode in the third mode set, determine a fourth intra-prediction mode derived from candidate samples of the first preset block; when the fourth intra-prediction mode is a candidate intra-prediction mode in the first mode set, perform an accumulation operation on the fourth intra-prediction mode to determine the statistical results of the fourth intra-prediction mode in the first mode set.

[0249] In the embodiments of this application, the third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode for prediction by copying intra-frame blocks, a mode for prediction by using extrapolation filters, and a mode for prediction by using matrix operations.

[0250] In one specific embodiment, the mode for predicting duplicate intra-blocks can be either ITMP mode or IBC mode.

[0251] In one specific embodiment, the mode for prediction using an extrapolation filter may specifically be the EIP mode.

[0252] In one specific embodiment, the prediction mode using matrix operations can be the MIP mode.

[0253] In other words, in a more specific embodiment, the third set of modes may include at least one of the following: ITMP mode, IBC mode, EIP mode, and MIP mode.

[0254] In this embodiment, if the first preset block is an intra-prediction block, and the first preset block uses MIP mode, ITMP mode, EIP mode, IBC mode, etc., since it does not use candidate intra-prediction modes in the first mode set for prediction, the first preset block can be skipped directly, i.e., the first preset block is not counted, and no cumulative calculation is performed on any mode in the first mode set; alternatively, the fourth intra-prediction mode can be derived from the reconstructed or predicted block of this block using gradient statistics, and the derived fourth intra-prediction mode belongs to one of the candidate intra-prediction modes in the first mode set. Assuming the mode index corresponding to the fourth intra-prediction mode is X, then the occurrence[X] can be cumulatively calculated.

[0255] For example, if the first preset block is an intra-prediction block, then taking the derivation of the fourth intra-prediction mode from the first preset block as an example, the fourth intra-prediction mode is accumulated. Specifically, the accumulated value corresponding to the fourth intra-prediction mode is accumulated to determine the current accumulated value of the fourth intra-prediction mode, and the accumulated value corresponding to the fourth intra-prediction mode is updated based on the current accumulated value. The judgment of the next preset block continues until all one or more preset blocks are traversed, and the accumulated value corresponding to the fourth intra-prediction mode obtained is determined as the statistical result of the fourth intra-prediction mode in the first mode set.

[0256] It can also be understood that in the embodiments of this application, each intra-frame prediction mode actually represents a texture feature. Therefore, the intra-frame prediction mode index is also a texture feature index. For example, DC mode and PLANA mode correspond to gradient texture features, and a certain angle prediction mode corresponds to the texture feature of that angle. Texture feature index can avoid the occurrence of intra-frame prediction modes "between frames" on the one hand, and it is also more conducive to possible expansion on the other hand. For example, an intra-frame prediction mode can correspond to multiple texture features, such as DC mode, which can correspond to horizontal gradient texture, vertical gradient texture, diagonal gradient texture, etc.

[0257] In the embodiments of this application, when determining the candidate samples based on the first preset block to derive the intra-prediction mode, the number of candidate samples used to derive the intra-prediction mode (i.e., "candidate texture feature index") can be at least one, such as 1, 2, 3 or more.

[0258] In this embodiment, the number of candidate samples can be determined based on the size parameter of the first preset block. That is, when deriving one or more candidate texture feature indices based on the candidate samples, the number of candidate samples used can be determined by the size parameter of the first preset block. For example, if the size of the first preset block is small, all available samples can be counted; if the size of the first preset block is large, downsampling of the first preset block can be used for counting, such as counting one sample from every 2, 4, or 8 samples in the horizontal and / or vertical directions. Alternatively, if the size of the first preset block in a horizontal or vertical direction is less than or equal to 8, all available samples in that direction are counted; otherwise, if the size of the first preset block in a horizontal or vertical direction is less than or equal to 16, one sample from every 2 samples in that direction is counted; otherwise, one sample from every 4 samples in that direction is counted. No specific limitation is made here.

[0259] In this embodiment of the application, taking the derivation of a first intra-frame prediction mode based on candidate samples of a first preset block as an example, correspondingly, in some embodiments, determining the first intra-frame prediction mode derived from candidate samples of a first preset block can include: determining the horizontal gradient value and vertical gradient value of the candidate sample; determining the texture feature index and gradient intensity value corresponding to the candidate sample based on the horizontal gradient value and vertical gradient value of the candidate sample; determining a texture feature statistics table based on the texture feature index and gradient intensity value corresponding to the candidate sample; and determining the first intra-frame prediction mode derived from the first preset block based on the texture feature statistics table.

[0260] It should be noted that, in the embodiments of this application, when determining the texture feature index and gradient intensity value corresponding to the candidate sample based on the horizontal gradient value and vertical gradient value of the candidate sample, it may include: performing angle mapping based on the horizontal gradient value and vertical gradient value of the candidate sample to determine the texture feature index corresponding to the candidate sample; and performing gradient intensity calculation based on the horizontal gradient value and vertical gradient value of the candidate sample to determine the gradient intensity value corresponding to the candidate sample.

[0261] In one specific embodiment, determining the texture feature index corresponding to a candidate sample by performing angle mapping based on the horizontal and vertical gradient values ​​of the candidate sample may include: determining the texture feature index corresponding to the candidate sample using a preset lookup table based on the horizontal and vertical gradient values ​​of the candidate sample.

[0262] In this embodiment, the horizontal gradient value of the candidate sample can be represented by grad. x This means that the vertical gradient value of a candidate sample can be expressed as grad. y This indicates that, according to grad... x and grad y The texture feature index (or "virtual intra-prediction mode") can be derived by looking up a table.

[0263] For example, if abs(grad x ) equals 0 and abs(grad) y If abs(grad) is not equal to 0, then there is horizontal texture, corresponding to intra-prediction mode 18 in VVC. y ) equals 0 and abs(grad) x If abs(grad) is not equal to 0, then there is vertical texture, corresponding to intra-prediction mode 50 in VVC. x ) and abs(grad y If abs(grad) are not equal to 0, then... x ) equals abs(grad y ), and grad x and grad y The symbols are the same, corresponding to intra-prediction mode 34 in VVC. If abs(grad x ) equals twice abs(grad) y ), and grad x and grad y The symbols are the same, corresponding to intra-prediction mode 40 in VVC. Other cases can be determined by looking up a table using the same principle.

[0264] In one specific embodiment, calculating the gradient intensity based on the horizontal and vertical gradient values ​​of the candidate sample to determine the gradient intensity value corresponding to the candidate sample may include: performing an addition operation on the absolute values ​​of the horizontal and vertical gradient values ​​to determine the gradient intensity value corresponding to the candidate sample.

[0265] Here, the gradient intensity value corresponding to the candidate sample can be denoted as amp. For example, amp = abs(grad... x )+abs(grad y ).

[0266] It should be noted that, in this embodiment, the horizontal and vertical gradient values ​​of the candidate samples can be calculated using the Sobel operator. For example, the Sobel operator is as follows:

[0267] Operators for horizontal gradient values:

[0268] Operator for vertical gradient values:

[0269] Thus, assuming the sample value at the sample location (x, y) of the reconstructed or predicted block is P x,y Then the horizontal gradient value grad x and vertical gradient value grad y The calculation is as follows: grad x =P x+1,y-1 +2*P x+1,y +P x+1,y+1 -P x-1,y-1 -2*P x-1,y -P x-1,y+1 (7) grad y =P x-1,y+1 +2*P x,y+1 +P x+1,y+1 -P x-1,y-1 -2*P x,y-1 -P x+1,y-1 (8)

[0270] It should also be noted that, in the embodiments of this application, considering that the Sobel operator needs to use the samples in the top, bottom, left, right and one row and one column of the current sample, the embodiments of this application can be set to calculate the gradient of all samples in the reconstruction block or prediction block except for the outermost row and one column.

[0271] In some embodiments, determining a texture feature statistics table based on the texture feature index and gradient intensity value corresponding to the candidate sample may include: when the number of candidate samples is at least one, determining at least one texture feature index and at least one corresponding gradient intensity value; determining at least one reference texture feature index with distinct characteristics based on the at least one texture feature index; and calculating the cumulative gradient intensity values ​​belonging to the same reference texture feature index based on the at least one gradient intensity value to determine the cumulative gradient intensity value corresponding to the at least one reference texture feature index; and determining the texture feature statistics table based on the at least one reference texture feature index and the cumulative gradient intensity value corresponding to the at least one reference texture feature index.

[0272] In other words, in this embodiment, taking at least some samples in the reconstructed block or prediction block as candidate samples, the gradients of all or some samples in the reconstructed block or prediction block are calculated. Generally, horizontal and vertical gradient values ​​can be calculated, and the Sobel operator can be used to calculate the gradient values. For a given sample, the texture direction can be inferred based on its horizontal and vertical gradient values. For example, if the horizontal gradient value is non-zero and the vertical gradient value is zero, then the texture of the sample is vertical. Conversely, if the horizontal gradient value is zero and the vertical gradient value is non-zero, then the texture of the sample is horizontal. For example, if the horizontal and vertical gradient values ​​are equal and not zero, then the texture of the sample is 45 degrees. Of course, in this embodiment, there are many other cases where both the horizontal and vertical gradient values ​​are not zero, and the direction of the texture of the sample can be determined based on their ratio. In this way, the gradient intensity value of each sample can be mapped to the corresponding texture feature index. A texture feature statistics table is constructed, and the gradient intensity value of each calculated sample is accumulated into the corresponding texture feature index item in the statistics table to obtain the final texture feature statistics table. Then, the first intra-frame prediction mode derived from the first preset block can be determined based on the texture feature statistics table.

[0273] In some embodiments, a first intra-frame prediction mode derived from a first preset block is determined based on a texture feature statistics table. The method may include: determining the reference texture feature index corresponding to the highest gradient intensity accumulation value in the texture feature statistics table as the first intra-frame prediction mode derived from the first preset block.

[0274] In other words, in this embodiment, when performing gradient statistics based on candidate samples, assuming there are 67 intra-frame prediction modes, the texture feature statistics table can also be set as an array, such as the array AmpAccumulate

[0067] , and each item of AmpAccumulate

[0067] is initialized to 0. Taking the first preset block as an example, if the index (or "reference texture feature index") of the intra-frame prediction mode derived from a certain sample position is X, then the amp calculated for that sample position is accumulated into AmpAccumulate[X], and the mode with the highest value in the array AmpAccumulate

[0067] is selected as the first intra-frame prediction mode derived by gradient statistics for the reconstruction block or prediction block of this block. It should be noted that the first intra-frame prediction mode derived here belongs to one of the candidate intra-frame prediction modes in the first mode set.

[0275] It is also understood that, in the embodiments of this application, the process of using gradient statistics to derive one of the candidate intra-prediction modes in the first mode set for the reconstruction block or prediction block of this block can be performed when the block is decoded, and the derived candidate intra-prediction mode is saved. In this way, it is not necessary to wait until the current block is to be used to derive gradient statistics, thereby saving computation.

[0276] It can also be understood that, in the embodiments of this application, when it is determined that the intra-prediction mode used or derived by the first preset block is a candidate intra-prediction mode in the first mode set, the intra-prediction mode can be accumulated.

[0277] In one possible implementation, performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set may include: performing cumulative calculations on the first intra-frame prediction modes based on the size parameter value of the first preset block to determine the statistical results of the first intra-frame prediction modes in the first mode set.

[0278] In this embodiment, taking a first preset block as an example, assuming the width of the first preset block is width and the height of the first preset block is height, then the size parameter value of the first preset block is width × height. If the cumulative calculation at this time can take into account the size of the preset block, then occurrence[X] = occurrence[X] + width * height. Here, the first mode set can be represented by the array occurrence[L], where L is the number of modes in the first mode set, i.e., the length of the array; and X is the mode index of the first intra-frame prediction mode in the first mode set.

[0279] In another possible implementation, performing a cumulative operation on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set may include: performing a cumulative operation of incrementing the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set.

[0280] In this embodiment, the cumulative operation may not consider the size of the preset block. In this case, it can also be a cumulative operation of incrementing by one, i.e., occurrence[X] = occurrence[X] + 1, where X represents the mode index of the first intra-prediction mode in the first mode set. That is, after determining the first intra-prediction mode used or derived by the first preset block, if the first intra-prediction mode is a candidate intra-prediction mode in the first mode set, then the cumulative value of the first intra-prediction mode can be incremented by one to obtain the statistical result of the first intra-prediction mode in the first mode set.

[0281] Thus, after the above operation process, after traversing one or more preset blocks, occurrence

[0067] has been statistically completed, that is, the statistical results of at least two candidate intra-frame prediction modes in the first mode set can be obtained.

[0282] S2803, Based on the statistical results, determine at least two intra-prediction modes for the current block in the first mode set.

[0283] In this embodiment of the application, based on the obtained statistical results, at least two intra-frame prediction modes of the current block can be determined from the first mode set.

[0284] In some embodiments, determining at least two intra-prediction modes of the current block in the first mode set based on statistical results may include: sorting the candidate intra-prediction modes in the first mode set in descending order according to the statistical results, and determining the at least two candidate intra-prediction modes with the highest ranking as at least two intra-prediction modes of the current block.

[0285] In this embodiment of the application, based on the statistical results of occurrence

[0067] , at least two intra-prediction modes of the current block can be determined from the first mode set. For example, the top two candidate intra-prediction modes with the largest values ​​in occurrence

[0067] can be selected as the at least two intra-prediction modes of the current block.

[0286] For example, if two intra-prediction modes for the current block are determined, then the candidate intra-prediction mode M1 with the largest value in occurrence

[0067] and the candidate intra-prediction mode M2 ​​with the second largest value can be used as the two intra-prediction modes for the current block.

[0287] S2804, predict the current block based on at least two intra-frame prediction modes to determine the predicted value of the current block.

[0288] In some embodiments, predicting the current block based on at least two intra-frame prediction modes to determine the predicted value of the current block may include: determining the weight values ​​of each of the at least two intra-frame prediction modes based on the statistical results corresponding to the at least two intra-frame prediction modes; and performing weighted prediction on the current block based on the at least two intra-frame prediction modes and their respective weight values ​​to determine the predicted value of the current block.

[0289] In this embodiment of the application, based on the statistical results, not only can at least two intra-prediction modes of the current block be determined from the first mode set, but also the weight values ​​of each of these at least two intra-prediction modes can be determined. Specifically, the weight values ​​of each of the at least two intra-prediction modes can be determined based on the statistical results corresponding to these at least two intra-prediction modes.

[0290] For example, suppose that two intra-prediction modes for the current block can be derived from the statistical results, specifically: intra-prediction mode M1 and intra-prediction mode M2. When performing weighted prediction on the current block, weighted prediction can be performed based on intra-prediction mode M1, intra-prediction mode M2, and the PLANAR mode, with weights W1, W2, and W3 respectively. The specific calculation formula is as follows:

[0291] Where occurrence[M1] represents the cumulative value corresponding to intra-prediction mode M1, and occurrence[M2] represents the cumulative value corresponding to intra-prediction mode M2.

[0292] Furthermore, when using OBIP mode for the current block, assuming the predicted value obtained by using intra-prediction mode M1 to predict the current block is denoted as Pred1, the predicted value obtained by using intra-prediction mode M2 ​​to predict the current block is denoted as Pred2, and the predicted value obtained by using PLANA mode to predict the current block is denoted as Pred3; then the predicted value obtained by using OBIP mode to predict is denoted as Pred1. OBIP The specific calculation formula is as follows:

[0293] In this embodiment, intra-prediction modes from two first mode sets and the PLANAR mode are used for weighted prediction. It is understood that more intra-prediction modes from the first mode set can also be used, such as 3, 4, or 5. Furthermore, the PLANAR mode can be replaced with other modes, such as the ITMP mode; no limitation is made here.

[0294] It is also understood that, in the embodiments of this application, whether the current block uses the first prediction mode can be indicated by the first syntax element. Referring to Figure 33, for step S2801, the method may further include:

[0295] S3301, decode the bitstream and determine the value of the first syntax element.

[0296] S3302, when the first syntax element indicates that the current block uses the first prediction mode, determine one or more preset blocks that have been decoded.

[0297] In this embodiment of the application, after step S3302, steps S2802 to S2804 are executed. That is, when the first syntax element indicates that the current block uses the first prediction mode, one or more preset blocks that have been decoded are determined; then, the statistical results of determining at least two candidate intra-prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks are executed; based on the statistical results, at least two intra-prediction modes of the current block are determined in the first mode set; and the predicted value of the current block is determined by predicting the current block based on the at least two intra-prediction modes.

[0298] In this embodiment of the application, the first syntax element is used to indicate whether the current block uses a first prediction mode. The method may further include: if the value of the first syntax element is a first value, then the first syntax element indicates that the current block uses the first prediction mode; if the value of the first syntax element is a second value, then the first syntax element indicates that the current block does not use the first prediction mode.

[0299] In the embodiments of this application, the first syntax element can be a block-level syntax element, and the first value is different from the second value. Specifically, the first value can be 1 and the second value can be 0; or, the first value can be true and the second value can be false; or, the first value can be 0 and the second value can be 1; or, the first value can be false and the second value can be true, without any limitations.

[0300] Furthermore, in this embodiment, the first prediction mode can refer to the OBIP mode, and the first syntax element can be represented by cu_obip_flag. That is, in this embodiment, a block-level (e.g., CU or PU level) syntax element can be set to indicate whether the current block uses the OBIP mode. Specifically, if the value of cu_obip_flag is 1, it indicates that the current block uses the OBIP mode for prediction; if the value of cu_obip_flag is 0, it indicates that the current block does not use the OBIP mode for prediction. If cu_obip_flag does not appear, it can be inferred that its value is 0.

[0301] In other words, in this embodiment of the application, if the value of cu_obip_flag obtained by decoding is 1, it means that the first syntax element indicates that the current block uses the first prediction mode. At this time, one or more preset blocks that have been decoded can be determined, which makes it easier to determine the statistical results of at least two candidate intra-prediction modes in the first mode set, and then determine the at least two intra-prediction modes used by the current block.

[0302] It is also understood that, in the embodiments of this application, the decoding method can also be applied to the Multiple Transform Set Selection (MTSS) method. NSPT and LFNST both handle texture transformations at various angles, and they may have multiple transformation kernels, one of which may be specifically optimized for a particular angle texture. Of course, in addition to angle textures, NSPT and LFNST also include transformation kernels for handling gradient textures. In fact, these transformation kernels can also be considered as trained Karhunen-Loeve Transforms (KLT). That is, NSPT and LFNST each have multiple transformation kernels, each designed for a specific texture, including angle textures, gradient textures, etc. Furthermore, gradient textures can be further extended to include horizontal gradient textures, vertical gradient textures, diagonal gradient textures, etc. Moreover, the MTSS method is not limited to non-separable transformations like NSPT and LFNST; the MTSS method can also be applied to separable transformations optimized for specific textures.

[0303] It should also be noted that in the embodiments of this application, the current block can have multiple transform kernel groups, and each intra-prediction mode can correspond to one transform kernel group. In other words, each intra-prediction mode actually represents a texture feature. Therefore, the intra-prediction mode index is also a texture feature index. For example, DC and PLANAR correspond to gradient texture features, and a certain angle prediction mode corresponds to the texture feature of that angle. Texture feature indexes can avoid intra-prediction modes from appearing "between frames," and they are also more conducive to possible expansion. For example, one intra-prediction mode can correspond to multiple texture features, such as the DC mode, which can correspond to horizontal gradient textures, vertical gradient textures, diagonal gradient textures, etc.

[0304] Specifically, in the embodiments of this application, considering that some intra-prediction modes are not simple texture features, but may contain two or more texture features, the MTSS technique can be used here. In the MTSS technique, if the intra-prediction mode of the current block is a certain special intra-prediction mode, then there is more than one selectable transform kernel group, such as the first transform kernel group corresponding to intra-prediction mode M1 and the second transform kernel group corresponding to intra-prediction mode M2.

[0305] In some embodiments, for the transformation process of the current block, referring to Figure 34, the method may include:

[0306] S3401, determine the transform kernel of the current block.

[0307] In one possible implementation, determining the transform kernel of the current block may include: determining the transform kernel group index of the current block; determining the transform kernel group of the current block based on the transform kernel group index; and determining the transform kernel of the current block based on the transform kernel group.

[0308] It should be noted that, in this embodiment, the transform kernel set index is used to indicate the number of the transform kernel set of the current block among at least two candidate transform kernel sets. This transform kernel set index can be represented by lfnst_nspt_set_index. The transform kernel set index of the current block can be an integer greater than or equal to zero, such as 0, 1, 2, etc. For example, if the value of lfnst_nspt_set_index is equal to 0, it indicates that the first candidate transform kernel set is selected as the transform kernel set of the current block; if the value of lfnst_nspt_set_index is equal to 1, it indicates that the second candidate transform kernel set is selected as the transform kernel set of the current block.

[0309] In some embodiments, the method for determining the transform kernel group index of the current block may include: decoding the bitstream to determine the transform kernel group index of the current block; or, the method may also include: decoding the bitstream to determine the value of the second syntax element; and determining the transform kernel group index of the current block based on the value of the second syntax element.

[0310] In other words, in the embodiments of this application, the transform kernel group index of the current block can be determined by directly decoding the bitstream, or it can be determined by decoding the value of the second syntax element.

[0311] It should also be noted that, in the embodiments of this application, the second syntax element can be represented by lfnst_nspt_set_index. The value of the second syntax element can be used to indicate the transform kernel group index of the current block, specifically the number of the transform kernel group of the current block among at least two candidate transform kernel groups. The value of the second syntax element can be an integer greater than or equal to zero, such as 0, 1, 2, etc.

[0312] Here, for OBIP mode, because it uses multiple intra-prediction modes for weighting—for example, it uses two intra-prediction modes (M1, M2) from the first mode set and the PLANAR mode for weighting—its predicted values ​​contain characteristics of both M1 and M2. Its residuals may also contain characteristics of M1 and M2, and of course, other characteristics, but the residuals are correlated with both M1 and M2. Therefore, for OBIP mode, when determining the transform kernel (LFNST or NSPT) based on the prediction mode, an additional option can be added. For example, a syntax element can be used to indicate whether the transform kernel (group) is determined based on M1 or M2. For instance, a second syntax element, `lfnst_nspt_set_index`, can be set. If `lfnst_nspt_set_index` is 0, the first candidate transform kernel group is used; if `lfnst_nspt_set_index` is 1, the second candidate transform kernel group is used. The first candidate transform kernel group is determined by the first intra-prediction mode (M1), and the second candidate transform kernel group is determined by the second intra-prediction mode (M2). Optionally, in special cases, such as when the transform kernel groups determined by the first intra-prediction mode (M1) and the second intra-prediction mode (M2) are exactly the same, other intra-prediction modes can be used, such as the third intra-prediction mode or the PLANAR mode involved in the prediction.

[0313] It is also understood that, in the embodiments of this application, the decoding end can also construct a first candidate list. In some embodiments, the method may further include: determining a first candidate list for the current block, the first candidate list indicating at least two candidate transform kernel groups; and determining the transform kernel group for the current block based on the first candidate list and the transform kernel group index.

[0314] It should be noted that, in the embodiments of this application, when the current block uses OBIP mode, the first candidate list may include at least two candidate texture feature indices, or the first candidate list may include at least two candidate transform kernel groups. Here, each candidate texture feature index corresponds to a candidate transform kernel group. Therefore, it can be said that the first candidate list indicates at least two candidate transform kernel groups.

[0315] For example, the transform kernel group determined by the first intra-prediction mode (M1) can be used as the first candidate transform kernel group in the first candidate list. If the transform kernel group determined by the second intra-prediction mode (M2) is different from the first candidate transform kernel group, then the transform kernel group determined by the second intra-prediction mode (M2) is determined as the second candidate transform kernel group in the first candidate list; otherwise, if the transform kernel group determined by the second intra-prediction mode (M2) is the same as the first candidate transform kernel group, then the transform kernel groups determined by the third intra-prediction mode, the fourth intra-prediction mode, etc., used in the current block are checked sequentially to see if they are the same as the first candidate transform kernel group in the first candidate list. If they are different, then the transform kernel group is determined as the second candidate transform kernel group in the first candidate list. If a second candidate transform kernel group is still not determined after checking all the intra-prediction modes used in the current block, then a default transform kernel group can be used as the second candidate transform kernel group, for example, the transform kernel group determined by the PLANA mode can be determined as the second candidate transform kernel group.

[0316] In another possible implementation, determining the transform kernel of the current block after determining the transform kernel of the current block may include: determining the transform kernel index of the current block; and determining the transform kernel of the current block based on the transform kernel group and the transform kernel index.

[0317] In another possible implementation, when the first candidate list indicates the transform kernels included in at least two candidate transform kernel groups, determining the transform kernel of the current block may include: determining the transform kernel index of the current block; and determining the transform kernel of the current block based on the first candidate list and the transform kernel index.

[0318] In this embodiment, for the first candidate list, assuming the first candidate list indicates the transform cores included in two candidate transform core groups, if the first candidate transform core group includes 3 candidate transform cores and the second candidate transform core group includes 2 candidate transform cores, then the first candidate list can indicate 5 candidate transform cores; if the first candidate transform core group includes 3 candidate transform cores and the second candidate transform core group includes 3 candidate transform cores, then the first candidate list can indicate 6 candidate transform cores. In this case, after determining the transform core index of the current block, the transform core of the current block can be determined from the first candidate list according to the transform core index.

[0319] It should also be noted that, in the embodiments of this application, the transform kernel index of the current block can be a positive integer, such as 1, 2, 3, 4, 5, 6, etc. The transform kernel index of the current block can be determined by directly decoding the bitstream, or it can be determined by decoding the value of the third syntax element.

[0320] For example, one possible implementation is to decode the bitstream and determine the transform kernel index of the current block. Alternatively, another possible implementation is to decode the bitstream, determine the value of the third syntax element, and when the third syntax element indicates that the current block uses the first transform mode, determine the transform kernel index of the current block based on the value of the third syntax element.

[0321] It should also be noted that, in the embodiments of this application, the first transformation mode can be LFNST / NSPT, and the third syntax element can be represented by lfnst_nspt_index. The third syntax element can be used to indicate whether the current block uses the first transformation mode, and the corresponding transformation kernel index when the current block uses the first transformation mode.

[0322] It should also be noted that, in the embodiments of this application, if the value of the third syntax element is a third value, it is determined that the current block does not use the first transformation mode; if the value of the third syntax element is a fourth value, it is determined that the current block uses the first transformation mode and the corresponding transformation kernel index. The third value can be set to 0, and the fourth value can be set to a non-zero value, such as 1, 2, 3, 4, 5, 6, etc.

[0323] In other words, in this embodiment of the application, the transform kernel index of the current block for LFNST / NSPT can also be represented by lfnst_nspt_index. Here, lfnst_nspt_index of 0 indicates that the current block does not use LFNST / NSPT. Each transform kernel group of LFNST / NSPT in the ECM has 3 transform kernels, so a value of lfnst_nspt_index of 1, 2, or 3 indicates that the current block uses the first, second, or third transform kernel of the selected transform kernel group of LFNST / NSPT.

[0324] In this embodiment, the current block uses OBIP mode, in which case it has more than one selectable transform core group. For example, it has two selectable transform core groups, then the possible values ​​of lfnst_nspt_index are 0, 1, 2, 3, 4, 5, 6. Among them, 1, 2, and 3 correspond to the 3 transform cores of the first candidate transform core group, and 4, 5, and 6 correspond to the 3 transform cores of the second candidate transform core group.

[0325] For example, instead of adding new syntax elements, you can add possible values ​​for existing syntax elements. For instance, `lfnst_nspt_index` can indicate the transform kernel of LFNST or NSPT. Given that each transform kernel group of LFNST / NSPT has three transform kernels, for a typical prediction mode, the possible values ​​for `lfnst_nspt_index` are 0, 1, 2, and 3. Specifically, if `lfnst_nspt_index` is 0, it means LFNST or NSPT is not used; if it is 1, it means the first transform kernel in the selected transform kernel group is used; if it is 2, it means the second transform kernel in the selected transform kernel group is used; and if it is 3, it means the third transform kernel in the selected transform kernel group is used. For OBIP mode, the possible values ​​for `lfnst_nspt_index` are 0, 1, 2, 3, 4, 5, and 6. Specifically, if `lfnst_nspt_index` is 0, it means LFNST or NSPT is not used; if `lfnst_nspt_index` is 1, it means the first transform kernel in the first candidate transform kernel group is used; if `lfnst_nspt_index` is 2, it means the second transform kernel in the first candidate transform kernel group is used; if `lfnst_nspt_index` is 3, it means the third transform kernel in the first candidate transform kernel group is used; if `lfnst_nspt_index` is 4, it means the first transform kernel in the second candidate transform kernel group is used; if `lfnst_nspt_index` is 5, it means the second transform kernel in the second candidate transform kernel group is used; and if `lfnst_nspt_index` is 6, it means the third transform kernel in the second candidate transform kernel group is used. This determines the transform kernel for the current block.

[0326] S3402, transform the transformation coefficients of the current block according to the transformation kernel to determine the residual block of the current block.

[0327] It should be noted that, in the embodiments of this application, the method may further include: decoding the bitstream to determine the quantization coefficients of the current block; and dequantizing the quantization coefficients of the current block to determine the transform coefficients of the current block.

[0328] It should also be noted that, in the embodiments of this application, when transforming the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block, it may include: performing an inseparable fundamental transformation on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block; or, performing a low-frequency inseparable transformation on the transform coefficients of the current block according to the transform kernel to determine the transform block of the current block, and performing a discrete cosine transform on the transform block of the current block to determine the residual block of the current block.

[0329] In one specific embodiment, if the size parameters of the current block satisfy the first condition, then an inseparable fundamental transformation is performed on the transformation coefficients of the current block according to the transformation kernel to determine the residual block of the current block; if the size parameters of the current block satisfy the second condition, then a low-frequency inseparable transformation is performed on the transformation coefficients of the current block according to the transformation kernel to determine the transformation block of the current block, and a discrete cosine transformation is performed on the transformation block of the current block to determine the residual block of the current block.

[0330] Here, the size parameters of the current block satisfy the first condition, including: the size parameters of the current block are relatively small, for example, the size parameters of the current block are less than a certain threshold. That is, for smaller blocks, the NSPT transform kernel is used here, that is, the inverse NSPT transform is performed on the transform coefficients of the current block according to the transform kernel to determine the residual block of the current block.

[0331] Here, the size parameters of the current block satisfy the second condition, including: the size parameters of the current block are relatively large, for example, the size parameters of the current block are greater than a certain threshold. That is, for larger blocks, the LFNST transform kernel is used here, that is, the inverse LFNST transform is performed on the transform coefficients of the current block according to the transform kernel to determine the transform block of the current block; and the inverse DCT2 transform is performed on the transform block of the current block to determine the residual block of the current block.

[0332] It should also be noted that, in the embodiments of this application, the "inverse transformation" of the transform coefficients at the decoding end can also be referred to as "transformation" in the standard text. In this document, "transformation" and "inverse transformation" correspond to two opposite processes. For example, "transformation" converts spatial domain values ​​to frequency domain coefficients, while "inverse transformation" converts frequency domain coefficients back to spatial domain values. "Inverse" is relative to "forward," and both are essentially transformations. It should be noted that if the standard only specifies decoding, then the "transformation" in the standard text refers to the decoding part, specifically the "inverse transformation" in this document. The "inverse transformation" of the transform coefficients at the decoding end can also be referred to as "transformation" in the standard text.

[0333] S3403, Determine the reconstruction value of the current block based on the residual value and the predicted value of the current block.

[0334] In this embodiment of the application, after determining the residual value of the current block, the predicted value of the current block and the residual value of the current block can be added together to determine the reconstructed value of the current block.

[0335] For example, after the decoder obtains the quantization coefficients from the bitstream through entropy decoding, if it is an NSPT transform, the quantization coefficients are dequantized to obtain the decoded transform coefficients. The decoded transform coefficients are then subjected to an inverse NSPT transform to obtain the decoded residual block. Finally, the reconstructed block is obtained based on the decoded residual block and the prediction block. Conversely, if it is an LFNST transform, the quantization coefficients are dequantized to obtain the decoded transform coefficients. The decoded transform coefficients are then subjected to an inverse LFNST transform, followed by an inverse DCT2 transform to obtain the decoded residual block. Finally, the reconstructed block is obtained based on the decoded residual block and the prediction block.

[0336] This application provides a decoding method, specifically an intra-frame prediction mode and its transformation scheme based on occurrence statistics. First, one or more preset blocks that have already been decoded are determined. Then, based on the prediction modes used by each of the preset blocks, the statistical results of at least two candidate intra-frame prediction modes in a first mode set are determined. Next, based on the statistical results, at least two intra-frame prediction modes for the current block are determined from the first mode set. Finally, the current block is predicted based on these at least two intra-frame prediction modes to determine the predicted value of the current block. Thus, by using statistics from related decoded blocks to obtain the statistical results of these related blocks, and deriving the at least two intra-frame prediction modes used by the current block based on the statistical results, the overhead in the bitstream can be saved. Furthermore, using at least two intra-frame prediction modes for weighted prediction can improve prediction accuracy and reduce prediction errors to some extent. In addition, deriving the at least two intra-frame prediction modes used by the current block based on statistical results can avoid the need for gradient calculation to derive intra-frame prediction modes, as is required in DIMD modes, reducing computational complexity and thus improving encoding / decoding efficiency, thereby enhancing compression performance.

[0337] In another embodiment of this application, Figure 35 is a schematic flowchart of an encoding method provided by an embodiment of this application. As shown in Figure 35, the method may include:

[0338] S3501, determine one or more pre-coded blocks.

[0339] In this embodiment, the method is applied to an encoder. Specifically, based on the structure of the encoder 100 shown in Figure 26, this encoding method is mainly applied to the intra-prediction unit 106, the transform unit 108, and the inverse transform unit 111 in Figure 26. When the current block uses OBIP mode, the intra-prediction mode used by the current block can be derived by statistically analyzing the prediction modes used by related blocks in the encoded region. Thus, compared to TIMD and DIMD modes, which require a certain amount of computation to derive the intra-prediction mode, OBIP mode saves computation and improves compression performance.

[0340] In this embodiment, it is first necessary to determine whether the current block uses the first prediction mode. In some embodiments, the method may include: determining multiple candidate prediction modes, wherein the multiple candidate prediction modes include at least the first prediction mode; calculating the encoding cost of the current block according to the multiple candidate prediction modes, and determining the cost result corresponding to each of the multiple candidate prediction modes; determining the minimum cost result among the cost results corresponding to each of the multiple candidate prediction modes; and determining whether the current block uses the first prediction mode according to the candidate prediction mode corresponding to the minimum cost result.

[0341] In the embodiments of this application, the cost calculation here can be determined based on the cost result of Rate Distortion Optimization (RDO), or based on the cost result of Sum of Absolute Difference (SAD), or even based on the cost result of Sum of Absolute Transformed Difference (SATD), but no limitation is made here.

[0342] In this embodiment of the application, determining whether the current block uses the first prediction mode may include: when the candidate prediction mode corresponding to the minimum cost result is the first prediction mode, determining that the current block uses the first prediction mode; when the candidate prediction mode corresponding to the minimum cost result is not the first prediction mode, determining that the current block does not use the first prediction mode.

[0343] Furthermore, in some embodiments, the method may further include: determining the value of a first syntax element; encoding the value of the first syntax element; and writing the obtained encoded bits into a bitstream.

[0344] It should be noted that, in the embodiments of this application, the first syntax element is used to indicate whether the current block uses the first prediction mode. Determining the value of the first syntax element may include: determining the value of the first syntax element to be a first value when the candidate prediction mode corresponding to the minimum cost result is the first prediction mode; and determining the value of the first syntax element to be a second value when the candidate prediction mode corresponding to the minimum cost result is not the first prediction mode. That is, if the current block uses the first prediction mode, then the value of the first syntax element is determined to be the first value; if the current block does not use the first prediction mode, then the value of the first syntax element is determined to be the second value.

[0345] It should also be noted that, in the embodiments of this application, the first syntax element can be a block-level syntax element, and the first value is different from the second value. Specifically, the first value can be 1 and the second value can be 0; or, the first value can be true and the second value can be false; or, the first value can be 0 and the second value can be 1; or, the first value can be false and the second value can be true, without any limitations.

[0346] Furthermore, in this embodiment, the first prediction mode can refer to the OBIP mode, and the first syntax element can be represented by cu_obip_flag. That is, in this embodiment, a block-level (e.g., CU or PU level) syntax element can be set to indicate whether the current block uses the OBIP mode. Specifically, if the current block uses the OBIP mode for prediction, then the value of cu_obip_flag can be determined to be 1; if the current block does not use the OBIP mode for prediction, then the value of cu_obip_flag can be determined to be 0. If cu_obip_flag does not appear in the bitstream, then its value can also be inferred to be 0.

[0347] Furthermore, in some embodiments, the method may further include: when the current block uses a first prediction mode, performing the step of determining the first or more preset blocks that have been encoded.

[0348] In other words, in the embodiments of this application, if the current block uses the first prediction mode, one or more preset blocks that have been encoded can be determined, which makes it easier to count the occurrence of at least two candidate intra-frame prediction modes in the first mode set in one or more preset blocks.

[0349] In one possible implementation, determining one or more pre-defined blocks that have been encoded may include: determining at least one candidate block that has been encoded, and determining one or more pre-defined blocks based on at least one candidate block.

[0350] In the embodiments of this application, a candidate block may include: adjacent blocks of the current block, and / or non-adjacent blocks of the current block. That is, one or more preset blocks may include only adjacent blocks of the current block; or, may include only non-adjacent blocks of the current block; or, may be composed of both adjacent and non-adjacent blocks of the current block.

[0351] Understandably, when using OBIP mode for the current block, it's possible to statistically analyze the adjacent and non-adjacent blocks. As shown in Figure 29, the black-filled block is the current block, blocks numbered 1-7 are the adjacent blocks, and blocks numbered 8-25 are the non-adjacent blocks. Adjacent blocks are those spatially adjacent to the current block, while non-adjacent blocks are those spatially not adjacent to the current block.

[0352] For example, suppose the top-left corner of the current block is (x, y), the width of the current block is width, and the height is height. Then, for the adjacent blocks of the current block, adjacent block 1 is the block containing coordinates (x-1, y-1), adjacent block 2 is the block containing coordinates (x, y-1), adjacent block 3 is the block containing coordinates (x-1, y), adjacent block 4 is the block containing coordinates (x+width-1, y-1), adjacent block 5 is the block containing coordinates (x+width, y-1), adjacent block 6 is the block containing coordinates (x-1, y+height-1), and adjacent block 7 is the block containing coordinates (x-1, y+height).

[0353] Non-adjacent blocks can also be positioned according to certain rules. For example, the position of a non-adjacent block can be calculated based on (x, y), width, and height. In Figure 29, the distance between each cell in the vertical direction is height, and the distance between each cell in the horizontal direction is width. Therefore, for the non-adjacent blocks of the current block, non-adjacent block 8 contains the coordinates (x-width-1, y-height-1), non-adjacent block 9 contains the coordinates (x+2*width-1, y-height-1), non-adjacent block 10 contains the coordinates (x-width-1, y+2*height-1), non-adjacent block 11 contains the coordinates (x-2*width-1, y-2*height-1), non-adjacent block 12 contains the coordinates (x+width / 2, y-2*height-1), and so on. It should be noted that the block containing a certain coordinate refers to the CU or PU containing that coordinate.

[0354] Thus, in this embodiment of the application, for the encoded candidate block, it may be possible to use only adjacent blocks, or only non-adjacent blocks, or both adjacent and non-adjacent blocks, or designated blocks from these adjacent and non-adjacent blocks, without any limitations.

[0355] In another possible implementation, determining one or more pre-coded blocks may include: determining a pre-defined range based on the current block; and determining one or more pre-defined blocks based on all coded blocks within the pre-defined range.

[0356] In this embodiment, the preset range can be any range that has been coded and set. For example, suppose the top-left corner of the current block is (x, y), the width of the current block is width, and the height is height. Then the preset range can be set according to the position of the current block and its width and height.

[0357] It can also be understood that when the current block uses OBIP mode, all blocks within a preset range can be statistically analyzed. As shown in Figure 30, the black-filled blocks represent the current block, and the dotted-filled parts represent the preset range. For example, a possible preset range is: all encoded ranges within a rectangular area with the top left corner at (x-2*width, y-2*height), the top right corner at (x+3*width-1, y-2*height), and the bottom left corner at (x-2*width, y+3*height-1). If the minimum size of the CU or PU is 4×4, then all 4×4 blocks within the search range can be statistically analyzed. For example, the first 4×4 block in the top left corner contains the coordinates (x-2*width, y-2*height). Increasing by 4 horizontally allows scanning all 4×4 blocks in the same row, and increasing by 4 vertically allows scanning all columns. It should be noted that the block containing a certain coordinate refers to the CU or PU containing that coordinate.

[0358] Thus, in this embodiment of the application, after determining the encoded preset range, all blocks found within the preset range can be identified as one or more preset blocks.

[0359] S3502, based on the prediction modes used by one or more preset blocks, determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set.

[0360] In this embodiment, the first mode set may include at least two candidate intra-prediction modes. The intra-prediction modes may include many modes, such as the PLANA mode, DC mode, 65 angle prediction modes, and MIP mode already existing in VVC, and the DIMD mode, TIMD mode, SGPM mode, ITMP mode, and OBIP mode newly added in ECM. A first mode set can be defined here, and in some embodiments, the first mode set includes at least one of the following: angle prediction mode, DC mode, and PLANA mode.

[0361] In this embodiment, the first mode set includes at least two candidate intra-prediction modes. Exemplarily, the first mode set may contain only angle prediction modes, or it may contain angle prediction modes and PLANAR modes, or it may contain angle prediction modes, DC modes, and PLANAR modes, etc., without any limitation. That is, one possible implementation is that the first mode set contains only angle prediction modes, and the number of these modes may be 65, or it may be more due to the existence of finer-grained angle prediction modes. Another possible implementation is that the first mode set contains only angle prediction modes, DC modes, and PLANAR modes. The following detailed description uses the example of a first mode set containing angle prediction modes, DC modes, and PLANAR modes, in which case the first mode set corresponds exactly to the 67 intra-prediction modes 0-66 in VVC.

[0362] In this embodiment, for the OBIP mode, a data structure can be constructed to record the occurrence frequency of each candidate intra-frame prediction mode in the first mode set. For example, this data structure can be an array, denoted as occurrence

[0067] . It should be noted that this example uses a first mode set with 67 modes; if the number of modes is different, the array length will be adjusted accordingly. Furthermore, all values ​​in occurrence

[0067] are initialized to 0.

[0363] In this embodiment, after determining one or more pre-coded blocks, the occurrence of modes in the first mode set within the one or more pre-coded blocks can be statistically analyzed, i.e., the statistical results of at least two candidate intra-prediction modes in the first mode set can be determined. In a specific embodiment, this may include: determining the prediction mode used by each of the one or more pre-coded blocks; performing candidate intra-prediction mode statistics based on the prediction modes used by each of the one or more pre-coded blocks, and determining the statistical results of at least two candidate intra-prediction modes in the first mode set.

[0364] It should be noted that, in the embodiments of this application, after determining a preset block (e.g., CU or PU) containing certain coordinates within the encoded region, the prediction mode used by the preset block can be determined. For example, if the preset block is an intra-frame prediction block, then the intra-frame prediction mode used by the preset block can be determined; if the preset block is an inter-frame prediction block, then the inter-frame prediction mode used by the preset block can be determined. Thus, by performing candidate intra-frame prediction mode statistics based on the prediction modes used by each of these one or more preset blocks, the statistical results of at least two candidate intra-frame prediction modes in the first mode set can be determined.

[0365] It should also be noted that, in the embodiments of this application, taking the first preset block as an example, for determining the statistical results of at least two candidate intra-prediction modes in the first mode set, the method may include: determining the intra-prediction mode used by the first preset block; when the intra-prediction mode is a candidate intra-prediction mode in the first mode set, performing cumulative calculation on the intra-prediction mode to determine the statistical results of the intra-prediction mode in the first mode set.

[0366] Understandably, in the embodiments of this application, the first preset block can be any one of one or more preset blocks. That is, taking the first preset block as an example, the first preset can be an inter-frame prediction block or an intra-frame prediction block. The processing situation where the first preset block is an inter-frame prediction block will be described in detail below.

[0367] In one possible implementation, for determining the statistical results of at least two candidate intra-prediction modes in the first mode set, the method may include: when the first preset block is an inter-prediction block, not counting the first preset block.

[0368] In this embodiment, the first preset block can be any one of one or more preset blocks. If the first preset block is an inter-frame prediction block, it means that the first preset block does not use the modes in the first mode set for prediction. In this case, the first preset block can be ignored (or the first preset block can be skipped). That is, no cumulative calculation is performed on any mode in the first mode set. Instead, candidate intra-frame prediction mode statistics can be performed based on the prediction modes used by other preset blocks outside the first preset block, thereby determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set.

[0369] In another possible implementation, for determining the statistical results of at least two candidate intra-prediction modes in the first mode set, the method may include: when the first preset block is an inter-prediction block, determining the first intra-prediction mode derived from the candidate samples of the first preset block; performing cumulative calculation on the first intra-prediction mode to determine the statistical results of the first intra-prediction mode in the first mode set.

[0370] In this embodiment, the first preset block can be any one of one or more preset blocks. If the first preset block is an inter-frame prediction block, it indicates that the first preset block does not use the modes in the first mode set for prediction. In this case, the first intra-frame prediction mode of the first preset block can be derived based on the candidate samples of the first preset block. Then, by accumulating the first intra-frame prediction mode, the statistical result of the first intra-frame prediction mode in the first mode set can be determined. Specifically, the current cumulative value of the first intra-frame prediction mode can be determined by accumulating the cumulative value corresponding to the first intra-frame prediction mode, and the cumulative value corresponding to the first intra-frame prediction mode can be updated based on the current cumulative value. The judgment of the next preset block continues until all one or more preset blocks are traversed, and the final cumulative value corresponding to the first intra-frame prediction mode is determined as the statistical result of the first intra-frame prediction mode in the first mode set.

[0371] In some embodiments, the method may further include: when the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, performing an accumulation operation on the first intra-frame prediction mode to determine the statistical result of the first intra-frame prediction mode in the first mode set. That is, after deriving the first intra-frame prediction mode based on candidate samples from the first preset block, if the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, then an accumulation operation can be performed on the first intra-frame prediction mode to determine the statistical result of the first intra-frame prediction mode in the first mode set.

[0372] In some embodiments, the method may further include: determining candidate samples of the first preset block based on at least a portion of the samples within the reconstruction block of the first preset block; or, determining candidate samples of the first preset block based on at least a portion of the samples within the prediction block of the first preset block.

[0373] In this embodiment, when the first preset block is an inter-frame prediction block, gradient statistics can be performed based on the candidate samples of the first preset block to determine the first intra-frame prediction mode of the first preset block. Here, the candidate samples can be at least a portion of the samples in the reconstructed block or prediction block of the block. That is, gradient statistics can be performed based on at least a portion of the samples in the reconstructed block or prediction block of the block to derive the first intra-frame prediction mode. As shown in Figure 32, an 8×8 block is provided, and the first intra-frame prediction mode can be derived using the 6×6 candidate samples within it; that is, all points within the reconstructed block or prediction block except for the outermost row and column can be used as candidate samples.

[0374] Thus, if the first preset block is an inter-frame prediction block, since it does not use any modes from the first mode set for prediction, it can be directly ignored from the statistics, meaning no cumulative calculation is performed on any modes from the first mode set. Alternatively, the first intra-frame prediction mode can be derived from the reconstructed or predicted blocks of this block using gradient statistics, and the derived first intra-frame prediction mode belongs to one of the candidate intra-frame prediction modes within the first mode set. Assuming the mode index corresponding to the first intra-frame prediction mode is X, then the occurrence[X] can be cumulatively calculated.

[0375] It is also understandable that, in the embodiments of this application, the first preset block is still taken as an example, and the processing of the first preset block as an intra-frame prediction block will be described in detail below.

[0376] In one possible implementation, for determining the statistical results of at least two candidate intra-prediction modes in the first mode set, the method further includes: when the first preset block is an intra-prediction block, determining the second intra-prediction mode used by the first preset block; when the second intra-prediction mode is a candidate intra-prediction mode in the first mode set, performing an accumulation operation on the second intra-prediction mode to determine the statistical results of the second intra-prediction mode in the first mode set.

[0377] In this embodiment, if the first preset block is an intra-prediction block, and the second intra-prediction mode used by the first preset block is a candidate intra-prediction mode in the first mode set, assuming the mode index corresponding to the second intra-prediction mode is X, then the accumulation operation is also performed on occurrence[X]. Here, the second intra-prediction mode is one of the candidate intra-prediction modes in the first mode set.

[0378] In another possible implementation, if the second intra-frame prediction mode is not a candidate intra-frame prediction mode in the first mode set, and the statistical results for determining at least two candidate intra-frame prediction modes in the first mode set are as follows: when the second intra-frame prediction mode is a candidate intra-frame prediction mode in the second mode set, determine at least two third intra-frame prediction modes corresponding to the second intra-frame prediction mode; when at least two third intra-frame prediction modes are at least two candidate intra-frame prediction modes in the first mode set, perform cumulative calculation on the at least two third intra-frame prediction modes to determine the statistical results of at least two third intra-frame prediction modes in the first mode set.

[0379] In this application embodiment, the second mode set is different from the first mode set, and the second mode set includes modes that use at least two intra-frame prediction modes for combined prediction.

[0380] In one specific embodiment, the mode that uses at least two intra-frame prediction modes for combined prediction can be DIMD mode, TIMD mode, SGPM mode and OBIP mode; wherein, the intra-frame prediction mode includes, but is not limited to, DC mode, PLANA mode and angle prediction mode.

[0381] In other words, in a more specific embodiment, the second set of modes may include at least one of the following: DIMD mode, TIMD mode, SGPM mode, and OBIP mode.

[0382] In this embodiment, if the first preset block uses DIMD mode, TIMD mode, SGPM mode, OBIP mode, etc., it may use multiple candidate intra-prediction modes from the first mode set. In this case, each candidate intra-prediction mode is accumulated. For example, assuming that the mode uses candidate intra-prediction modes X1 and X2 from the first mode set, then occurrence[X1] and occurrence[X2] are accumulated respectively.

[0383] In another possible implementation, the second intra-prediction mode is not a candidate intra-prediction mode in the first mode set. For the statistical results of determining at least two candidate intra-prediction modes in the first mode set, the method further includes: when the second intra-prediction mode is a candidate intra-prediction mode in the third mode set, the first preset block is not counted.

[0384] In another possible implementation, if the second intra-prediction mode is not a candidate intra-prediction mode in the first mode set, and the statistical results for determining at least two candidate intra-prediction modes in the first mode set are as follows: when the second intra-prediction mode is a candidate intra-prediction mode in the third mode set, determine a fourth intra-prediction mode derived from candidate samples of the first preset block; perform an accumulation operation on the fourth intra-prediction mode to determine the statistical results of the fourth intra-prediction mode in the first mode set. Alternatively, the method further includes: when the second intra-prediction mode is a candidate intra-prediction mode in the third mode set, determine a fourth intra-prediction mode derived from candidate samples of the first preset block; when the fourth intra-prediction mode is a candidate intra-prediction mode in the first mode set, perform an accumulation operation on the fourth intra-prediction mode to determine the statistical results of the fourth intra-prediction mode in the first mode set.

[0385] In the embodiments of this application, the third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode for prediction by copying intra-frame blocks, a mode for prediction by using extrapolation filters, and a mode for prediction by using matrix operations.

[0386] In one specific embodiment, the mode for predicting duplicate intra-blocks can be either ITMP mode or IBC mode.

[0387] In one specific embodiment, the mode for prediction using an extrapolation filter may specifically be the EIP mode.

[0388] In one specific embodiment, the prediction mode using matrix operations can be the MIP mode.

[0389] In other words, in a more specific embodiment, the third set of modes may include at least one of the following: ITMP mode, IBC mode, EIP mode, and MIP mode.

[0390] In this embodiment, if the first preset block is an intra-prediction block, and the first preset block uses MIP mode, ITMP mode, EIP mode, IBC mode, etc., since it does not use candidate intra-prediction modes in the first mode set for prediction, the first preset block can be skipped directly, i.e., the first preset block is not counted, and no cumulative calculation is performed on any mode in the first mode set; alternatively, the fourth intra-prediction mode can be derived from the reconstructed or predicted block of this block using gradient statistics, and the derived fourth intra-prediction mode belongs to one of the candidate intra-prediction modes in the first mode set. Assuming the mode index corresponding to the fourth intra-prediction mode is X, then the occurrence[X] can be cumulatively calculated.

[0391] For example, if the first preset block is an intra-prediction block, then taking the derivation of the fourth intra-prediction mode from the first preset block as an example, the fourth intra-prediction mode is accumulated. Specifically, the accumulated value corresponding to the fourth intra-prediction mode is accumulated to determine the current accumulated value of the fourth intra-prediction mode, and the accumulated value corresponding to the fourth intra-prediction mode is updated based on the current accumulated value. The judgment of the next preset block continues until all one or more preset blocks are traversed, and the accumulated value corresponding to the fourth intra-prediction mode obtained is determined as the statistical result of the fourth intra-prediction mode in the first mode set.

[0392] It can also be understood that in the embodiments of this application, each intra-frame prediction mode actually represents a texture feature. Therefore, the intra-frame prediction mode index is also a texture feature index. For example, DC mode and PLANA mode correspond to gradient texture features, and a certain angle prediction mode corresponds to the texture feature of that angle. Texture feature index can avoid the occurrence of intra-frame prediction modes "between frames" on the one hand, and it is also more conducive to possible expansion on the other hand. For example, an intra-frame prediction mode can correspond to multiple texture features, such as DC mode, which can correspond to horizontal gradient texture, vertical gradient texture, diagonal gradient texture, etc.

[0393] In the embodiments of this application, when determining the candidate samples based on the first preset block to derive the intra-prediction mode, the number of candidate samples used to derive the intra-prediction mode (i.e., "candidate texture feature index") can be at least one, such as 1, 2, 3 or more.

[0394] In this embodiment, the number of candidate samples can be determined based on the size parameter of the first preset block. That is, when deriving one or more candidate texture feature indices based on the candidate samples, the number of candidate samples used can be determined by the size parameter of the first preset block. For example, if the size of the first preset block is small, all available samples can be counted; if the size of the first preset block is large, downsampling of the first preset block can be used for counting, such as counting one sample from every 2, 4, or 8 samples in the horizontal and / or vertical directions. Alternatively, if the size of the first preset block in a horizontal or vertical direction is less than or equal to 8, all available samples in that direction are counted; otherwise, if the size of the first preset block in a horizontal or vertical direction is less than or equal to 16, one sample from every 2 samples in that direction is counted; otherwise, one sample from every 4 samples in that direction is counted. No specific limitation is made here.

[0395] In this embodiment of the application, taking the derivation of a first intra-frame prediction mode based on candidate samples of a first preset block as an example, correspondingly, in some embodiments, determining the first intra-frame prediction mode derived from candidate samples of a first preset block can include: determining the horizontal gradient value and vertical gradient value of the candidate sample; determining the texture feature index and gradient intensity value corresponding to the candidate sample based on the horizontal gradient value and vertical gradient value of the candidate sample; determining a texture feature statistics table based on the texture feature index and gradient intensity value corresponding to the candidate sample; and determining the first intra-frame prediction mode derived from the first preset block based on the texture feature statistics table.

[0396] It should be noted that, in the embodiments of this application, when determining the texture feature index and gradient intensity value corresponding to the candidate sample based on the horizontal gradient value and vertical gradient value of the candidate sample, it may include: performing angle mapping based on the horizontal gradient value and vertical gradient value of the candidate sample to determine the texture feature index corresponding to the candidate sample; and performing gradient intensity calculation based on the horizontal gradient value and vertical gradient value of the candidate sample to determine the gradient intensity value corresponding to the candidate sample.

[0397] In one specific embodiment, determining the texture feature index corresponding to a candidate sample by performing angle mapping based on the horizontal and vertical gradient values ​​of the candidate sample may include: determining the texture feature index corresponding to the candidate sample using a preset lookup table based on the horizontal and vertical gradient values ​​of the candidate sample.

[0398] In this embodiment, the horizontal gradient value of the candidate sample can be represented by grad. x This means that the vertical gradient value of a candidate sample can be expressed as grad. y This indicates that, according to grad... x and grad y The texture feature index (or "virtual intra-prediction mode") can be derived by looking up a table.

[0399] For example, if abs(grad x ) equals 0 and abs(grad) y If abs(grad) is not equal to 0, then there is horizontal texture, corresponding to intra-prediction mode 18 in VVC. y ) equals 0 and abs(grad) x If abs(grad) is not equal to 0, then there is vertical texture, corresponding to intra-prediction mode 50 in VVC. x ) and abs(grad y If abs(grad) are not equal to 0, then... x ) equals abs(grad y ), and grad x and grad y The symbols are the same, corresponding to intra-prediction mode 34 in VVC. If abs(grad x ) equals twice abs(grad) y ), and grad x and grad y The symbols are the same, corresponding to intra-prediction mode 40 in VVC. Other cases can be determined by looking up a table using the same principle.

[0400] In one specific embodiment, calculating the gradient intensity based on the horizontal and vertical gradient values ​​of the candidate sample to determine the gradient intensity value corresponding to the candidate sample may include: performing an addition operation on the absolute values ​​of the horizontal and vertical gradient values ​​to determine the gradient intensity value corresponding to the candidate sample.

[0401] Here, the gradient intensity value corresponding to the candidate sample can be denoted as amp. For example, amp = abs(grad... x )+abs(grad y ).

[0402] It should be noted that, in the embodiments of this application, the horizontal and vertical gradient values ​​of the candidate samples can be calculated using the Sobel operator. For example, taking the Sobel operator as an example, suppose the sample value at position (x, y) of the reconstructed or predicted block is P. x,y Then the horizontal gradient value grad x and vertical gradient value grad y The calculation is shown in equations (7) and (8) above.

[0403] It should also be noted that, in the embodiments of this application, considering that the Sobel operator needs to use the samples in the top, bottom, left, right and one row and one column of the current sample, the embodiments of this application can be set to calculate the gradient of all samples in the reconstruction block or prediction block except for the outermost row and one column.

[0404] In some embodiments, determining a texture feature statistics table based on the texture feature index and gradient intensity value corresponding to the candidate sample may include: when the number of candidate samples is at least one, determining at least one texture feature index and at least one corresponding gradient intensity value; determining at least one reference texture feature index with distinct characteristics based on the at least one texture feature index; and calculating the cumulative gradient intensity values ​​belonging to the same reference texture feature index based on the at least one gradient intensity value to determine the cumulative gradient intensity value corresponding to the at least one reference texture feature index; and determining the texture feature statistics table based on the at least one reference texture feature index and the cumulative gradient intensity value corresponding to the at least one reference texture feature index.

[0405] In other words, in this embodiment, taking at least some samples in the reconstructed block or prediction block as candidate samples, the gradients of all or some samples in the reconstructed block or prediction block are calculated. Generally, horizontal and vertical gradient values ​​can be calculated, and the Sobel operator can be used to calculate the gradient values. For a given sample, the texture direction can be inferred based on its horizontal and vertical gradient values. For example, if the horizontal gradient value is non-zero and the vertical gradient value is zero, then the texture of the sample is vertical. Conversely, if the horizontal gradient value is zero and the vertical gradient value is non-zero, then the texture of the sample is horizontal. For example, if the horizontal and vertical gradient values ​​are equal and not zero, then the texture of the sample is 45 degrees. Of course, in this embodiment, there are many other cases where both the horizontal and vertical gradient values ​​are not zero, and the direction of the texture of the sample can be determined based on their ratio. In this way, the gradient intensity value of each sample can be mapped to the corresponding texture feature index. A texture feature statistics table is constructed, and the gradient intensity value of each calculated sample is accumulated into the corresponding texture feature index item in the statistics table to obtain the final texture feature statistics table. Then, the first intra-frame prediction mode derived from the first preset block can be determined based on the texture feature statistics table.

[0406] In some embodiments, a first intra-frame prediction mode derived from a first preset block is determined based on a texture feature statistics table. The method may include: determining the reference texture feature index corresponding to the highest gradient intensity accumulation value in the texture feature statistics table as the first intra-frame prediction mode derived from the first preset block.

[0407] In other words, in this embodiment, when performing gradient statistics based on candidate samples, assuming there are 67 intra-frame prediction modes, the texture feature statistics table can also be set as an array, such as the array AmpAccumulate

[0067] , and each item of AmpAccumulate

[0067] is initialized to 0. Taking the first preset block as an example, if the index (or "reference texture feature index") of the intra-frame prediction mode derived from a certain sample position is X, then the amp calculated for that sample position is accumulated into AmpAccumulate[X], and the mode with the highest value in the array AmpAccumulate

[0067] is selected as the first intra-frame prediction mode derived by gradient statistics for the reconstruction block or prediction block of this block. It should be noted that the first intra-frame prediction mode derived here belongs to one of the candidate intra-frame prediction modes in the first mode set.

[0408] It is also understood that, in the embodiments of this application, the process of using gradient statistics to derive one of the candidate intra-prediction modes in the first mode set for the reconstruction block or prediction block of this block can be performed when this block is encoded, and the derived candidate intra-prediction mode is saved. In this way, it is not necessary to wait until the current block is to be used to derive gradient statistics, thereby saving computation.

[0409] It can also be understood that, in the embodiments of this application, when it is determined that the intra-prediction mode used or derived by the first preset block is a candidate intra-prediction mode in the first mode set, the intra-prediction mode can be accumulated.

[0410] In one possible implementation, performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set may include: performing cumulative calculations on the first intra-frame prediction modes based on the size parameter value of the first preset block to determine the statistical results of the first intra-frame prediction modes in the first mode set.

[0411] In this embodiment, taking a first preset block as an example, assuming the width of the first preset block is width and the height of the first preset block is height, then the size parameter value of the first preset block is width × height. If the cumulative calculation at this time can take into account the size of the preset block, then occurrence[X] = occurrence[X] + width * height. Here, the first mode set can be represented by the array occurrence[L], where L is the number of modes in the first mode set, i.e., the length of the array; and X is the mode index of the first intra-frame prediction mode in the first mode set.

[0412] In another possible implementation, performing a cumulative operation on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set may include: performing a cumulative operation of incrementing the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set.

[0413] In this embodiment, the cumulative operation may not consider the size of the preset block. In this case, it can also be a cumulative operation of incrementing by one, i.e., occurrence[X] = occurrence[X] + 1, where X represents the mode index of the first intra-prediction mode in the first mode set. That is, after determining the first intra-prediction mode used or derived by the first preset block, if the first intra-prediction mode is a candidate intra-prediction mode in the first mode set, then the cumulative value of the first intra-prediction mode can be incremented by one to obtain the statistical result of the first intra-prediction mode in the first mode set.

[0414] Thus, after the above operation process, after traversing one or more preset blocks, occurrence

[0067] has been statistically completed, that is, the statistical results of at least two candidate intra-frame prediction modes in the first mode set can be obtained.

[0415] S3503, based on the statistical results, determine at least two intra-prediction modes for the current block in the first mode set.

[0416] In this embodiment of the application, based on the obtained statistical results, at least two intra-frame prediction modes of the current block can be determined from the first mode set.

[0417] In some embodiments, determining at least two intra-prediction modes of the current block in the first mode set based on statistical results may include: sorting the candidate intra-prediction modes in the first mode set in descending order according to the statistical results, and determining the at least two candidate intra-prediction modes with the highest ranking as at least two intra-prediction modes of the current block.

[0418] In this embodiment of the application, based on the statistical results of occurrence

[0067] , at least two intra-prediction modes of the current block can be determined from the first mode set. For example, the top two candidate intra-prediction modes with the largest values ​​in occurrence

[0067] can be selected as the at least two intra-prediction modes of the current block.

[0419] For example, if two intra-prediction modes for the current block are determined, then the candidate intra-prediction mode M1 with the largest value in occurrence

[0067] and the candidate intra-prediction mode M2 ​​with the second largest value can be used as the two intra-prediction modes for the current block.

[0420] S3504, predict the current block based on at least two intra-frame prediction modes to determine the predicted value of the current block.

[0421] In some embodiments, predicting the current block based on at least two intra-frame prediction modes to determine the predicted value of the current block may include: determining the weight values ​​of each of the at least two intra-frame prediction modes based on the cumulative values ​​corresponding to the at least two intra-frame prediction modes; and performing weighted prediction on the current block based on the at least two intra-frame prediction modes and their respective weight values ​​to determine the predicted value of the current block.

[0422] In this embodiment of the application, based on the statistical results, not only can at least two intra-prediction modes of the current block be determined from the first mode set, but also the weight values ​​of each of these at least two intra-prediction modes can be determined. Specifically, the weight values ​​of each of the at least two intra-prediction modes can be determined based on the cumulative values ​​corresponding to these at least two intra-prediction modes.

[0423] For example, suppose that two intra-prediction modes for the current block can be derived from the statistical results, specifically: intra-prediction mode M1 and intra-prediction mode M2. When performing weighted prediction on the current block, weighted prediction can be performed based on intra-prediction mode M1, intra-prediction mode M2, and PLANAR mode, with weights of W1, W2, and W3 for each mode, respectively. The specific calculation formulas are shown in equations (9), (10), and (11) above. Wherein, occurrence[M1] represents the cumulative value corresponding to intra-prediction mode M1, and occurrence[M2] represents the cumulative value corresponding to intra-prediction mode M2.

[0424] Furthermore, when using OBIP mode for the current block, assuming the predicted value obtained by using intra-prediction mode M1 to predict the current block is denoted as Pred1, the predicted value obtained by using intra-prediction mode M2 ​​to predict the current block is denoted as Pred2, and the predicted value obtained by using PLANA mode to predict the current block is denoted as Pred3; then the predicted value obtained by using OBIP mode to predict is denoted as Pred1. OBIP The specific calculation formula is shown in the aforementioned formula (12).

[0425] In this embodiment, intra-prediction modes from two first mode sets and the PLANAR mode are used for weighted prediction. It is understood that more intra-prediction modes from the first mode set can also be used, such as 3, 4, or 5. Furthermore, the PLANAR mode can be replaced with other modes, such as the ITMP mode; no limitation is made here.

[0426] It is also understood that, in the embodiments of this application, this encoding method can also be applied to the Multiple Transform Set Selection (MTSS) method. NSPT and LFNST both handle texture transformations at various angles, and they may have multiple transformation kernels, one of which may be specifically optimized for a particular angle texture. Of course, in addition to angle textures, NSPT and LFNST also include transformation kernels for handling gradient textures. In fact, these transformation kernels can also be considered as trained Karhunen-Loeve Transforms (KLT). That is, NSPT and LFNST each have multiple transformation kernels, each designed for a specific texture, including angle textures, gradient textures, etc. Furthermore, gradient textures can be further extended to include horizontal gradient textures, vertical gradient textures, diagonal gradient textures, etc. Moreover, the MTSS method is not limited to non-separable transformations like NSPT and LFNST; the MTSS method can also be applied to separable transformations optimized for specific textures.

[0427] It should also be noted that in the embodiments of this application, the current block can have multiple transform kernel groups, and each intra-prediction mode can correspond to one transform kernel group. In other words, each intra-prediction mode actually represents a texture feature. Therefore, the intra-prediction mode index is also a texture feature index. For example, DC and PLANAR correspond to gradient texture features, and a certain angle prediction mode corresponds to the texture feature of that angle. Texture feature indexes can avoid intra-prediction modes from appearing "between frames," and they are also more conducive to possible expansion. For example, one intra-prediction mode can correspond to multiple texture features, such as the DC mode, which can correspond to horizontal gradient textures, vertical gradient textures, diagonal gradient textures, etc.

[0428] Specifically, in the embodiments of this application, considering that some intra-prediction modes are not simple texture features, but may contain two or more texture features, the MTSS technique can be used here. In the MTSS technique, if the intra-prediction mode of the current block is a certain special intra-prediction mode, then there is more than one selectable transform kernel group, such as the first transform kernel group corresponding to intra-prediction mode M1 and the second transform kernel group corresponding to intra-prediction mode M2.

[0429] In some embodiments, for the transformation process of the current block, referring to Figure 36, the method may include:

[0430] S3601, determine the transform kernel of the current block.

[0431] In one possible implementation, determining the transform kernel of the current block may include: determining a transform kernel group for the current block; and determining the transform kernel of the current block based on the transform kernel group.

[0432] In some embodiments, determining the transform kernel group of the current block may include: calculating the encoding cost of the current block based on at least two candidate transform kernel groups, determining the cost result corresponding to each of the at least two candidate transform kernel groups; determining the minimum cost result among the cost results corresponding to each of the at least two candidate transform kernel groups, and determining the candidate transform kernel group corresponding to the minimum cost result as the transform kernel group of the current block.

[0433] Furthermore, in some embodiments, the method further includes: determining the transform kernel group index of the current block; encoding the transform kernel group index of the current block and writing the obtained encoded bits into the bitstream.

[0434] It should be noted that, in this embodiment, the transform kernel set index can be used to indicate the number of the transform kernel set of the current block among at least two candidate transform kernel sets. This transform kernel set index can be represented by `lfnst_nspt_set_index`. The transform kernel set index of the current block can be an integer greater than or equal to zero, such as 0, 1, 2, etc. For example, if the value of `lfnst_nspt_set_index` is equal to 0, it indicates that the first candidate transform kernel set is selected as the transform kernel set of the current block; if the value of `lfnst_nspt_set_index` is equal to 1, it indicates that the second candidate transform kernel set is selected as the transform kernel set of the current block.

[0435] It should also be noted that, in the embodiments of this application, the transform kernel group index of the current block can be directly written into the bitstream, or it can be represented by a second syntax element, and then the value of the second syntax element is written into the bitstream. In a specific embodiment, the method further includes: determining the transform kernel group index of the current block; determining the value of the second syntax element according to the transform kernel group index of the current block; encoding the value of the second syntax element, and writing the obtained encoded bits into the bitstream.

[0436] In this embodiment, the second syntax element can be represented by `lfnst_nspt_set_index`. The second syntax element can be used to indicate the transform kernel group index of the current block, specifically the number of the transform kernel group of the current block among at least two candidate transform kernel groups. The value of the second syntax element can be an integer greater than or equal to zero, such as 0, 1, 2, etc. For example, if the value of the second syntax element is equal to 0, it indicates that the first candidate transform kernel group is selected as the transform kernel group of the current block; if the value of the second syntax element is equal to 1, it indicates that the second candidate transform kernel group is selected as the transform kernel group of the current block; if the value of the second syntax element is equal to 2, it indicates that the third candidate transform kernel group is selected as the transform kernel group of the current block.

[0437] Here, for OBIP mode, because it uses multiple intra-prediction modes for weighting—for example, it uses two intra-prediction modes (M1, M2) from the first mode set and the PLANAR mode for weighting—its predicted values ​​contain characteristics of both M1 and M2. Its residuals may also contain characteristics of M1 and M2, and of course, other characteristics, but the residuals are correlated with both M1 and M2. Therefore, for OBIP mode, when determining the transform kernel (LFNST or NSPT) based on the prediction mode, an additional option can be added. For example, a syntax element can be used to indicate whether the transform kernel (group) is determined based on M1 or M2. For instance, a second syntax element, `lfnst_nspt_set_index`, can be set. If `lfnst_nspt_set_index` is 0, the first candidate transform kernel group is used; if `lfnst_nspt_set_index` is 1, the second candidate transform kernel group is used. The first candidate transform kernel group is determined by the first intra-prediction mode (M1), and the second candidate transform kernel group is determined by the second intra-prediction mode (M2). Optionally, in special cases, such as when the transform kernel groups determined by the first intra-prediction mode (M1) and the second intra-prediction mode (M2) are exactly the same, other intra-prediction modes can be used, such as the third intra-prediction mode or the PLANAR mode involved in the prediction.

[0438] It is also understood that, in the embodiments of this application, the encoding end can also construct a first candidate list. In some embodiments, the method may further include: determining a first candidate list for the current block, wherein the first candidate list indicates at least two candidate transform kernel groups. Here, the transform kernel group index of the current block is specifically the number of the transform kernel group of the current block in the first candidate list.

[0439] It should be noted that, in the embodiments of this application, when the current block uses OBIP mode, the first candidate list may include at least two candidate texture feature indices, or the first candidate list may include at least two candidate transform kernel groups. Here, each candidate texture feature index corresponds to a candidate transform kernel group. Therefore, it can be said that the first candidate list indicates at least two candidate transform kernel groups.

[0440] For example, the transform kernel group determined by the first intra-prediction mode (M1) can be used as the first candidate transform kernel group in the first candidate list. If the transform kernel group determined by the second intra-prediction mode (M2) is different from the first candidate transform kernel group, then the transform kernel group determined by the second intra-prediction mode (M2) is determined as the second candidate transform kernel group in the first candidate list; otherwise, if the transform kernel group determined by the second intra-prediction mode (M2) is the same as the first candidate transform kernel group, then the transform kernel groups determined by the third intra-prediction mode, the fourth intra-prediction mode, etc., used in the current block are checked sequentially to see if they are the same as the first candidate transform kernel group in the first candidate list. If they are different, then the transform kernel group is determined as the second candidate transform kernel group in the first candidate list. If a second candidate transform kernel group is still not determined after checking all the intra-prediction modes used in the current block, then a default transform kernel group can be used as the second candidate transform kernel group, for example, the transform kernel group determined by the PLANA mode can be determined as the second candidate transform kernel group.

[0441] In one possible implementation, after determining the transform kernel group for the current block, the transform kernel for the current block can be further determined. This method may include: determining at least two candidate transform kernels included in the transform kernel group; calculating the encoding cost of the current block based on the at least two candidate transform kernels, and determining the cost result corresponding to each of the at least two candidate transform kernels; determining the minimum cost result among the cost results corresponding to each of the at least two candidate transform kernels, and determining the candidate transform kernel corresponding to the minimum cost result as the transform kernel for the current block.

[0442] In another possible implementation, when the first candidate list indicates the transform kernels included in at least two transform kernel groups, the method may further include: determining at least two candidate transform kernels indicated by the first candidate list; calculating the encoding cost of the current block based on the at least two candidate transform kernels, and determining the cost result corresponding to each of the at least two candidate transform kernels; determining the minimum cost result among the cost results corresponding to each of the at least two candidate transform kernels, and determining the candidate transform kernel corresponding to the minimum cost result as the transform kernel of the current block.

[0443] In the embodiments of this application, the cost calculation here can be determined based on the cost result of Rate Distortion Optimization (RDO), or based on the cost result of Sum of Absolute Difference (SAD), or even based on the cost result of Sum of Absolute Transformed Difference (SATD), but no limitation is made here.

[0444] In one specific embodiment, calculating the encoding cost of the current block based on at least two candidate transform kernels and determining the cost result corresponding to each of the at least two candidate transform kernels may include: transforming and quantizing the residual block of the current block based on a first candidate transform kernel to determine the first candidate quantization coefficient of the current block, and performing entropy encoding processing on the first candidate quantization coefficient to determine the first generation value of the first candidate transform kernel; performing inverse quantization and inverse transform on the first candidate quantization coefficient to determine the first candidate residual block of the current block, and determining the first candidate prediction block of the current block based on the first candidate residual block; calculating the cost based on the first candidate prediction block and the original image of the current block to determine the second generation value of the first candidate transform kernel; and determining the cost result corresponding to the first candidate transform kernel based on the first generation value and the second generation value of the first candidate transform kernel; wherein the first candidate transform kernel is any one of the at least two candidate transform kernels.

[0445] In this embodiment, for the first candidate list, assuming the first candidate list indicates the transform cores included in two candidate transform core groups, if the first candidate transform core group includes 3 candidate transform cores and the second candidate transform core group includes 2 candidate transform cores, then the first candidate list can indicate 5 candidate transform cores; if the first candidate transform core group includes 3 candidate transform cores and the second candidate transform core group includes 3 candidate transform cores, then the first candidate list can indicate 6 candidate transform cores. In this case, after determining the transform core index of the current block, the transform core of the current block can be determined from the first candidate list according to the transform core index.

[0446] In other words, in this embodiment, the transform kernel index of the current block can be used to indicate the number of the transform kernel of the current block in the transform kernel group or the first candidate list of the current block. The transform kernel index of the current block can be a positive integer, such as 1, 2, 3, 4, 5, 6, etc. This transform kernel index number can be directly written into the code stream, or it can be written into the code stream through the value of a third syntax element.

[0447] In one possible implementation, the transform kernel index of the current block is determined; the transform kernel index of the current block is encoded, and the resulting encoded bits are written into the bitstream.

[0448] In another possible implementation, the value of the third syntax element is determined; wherein the third syntax element is used to indicate whether the current block uses the first transform mode and the corresponding transform kernel index; the value of the third syntax element is encoded, and the resulting encoded bits are written into the bitstream.

[0449] It should also be noted that, in the embodiments of this application, the first transformation mode can be LFNST / NSPT, and the third syntax element can be represented by lfnst_nspt_index. The third syntax element can be used to indicate whether the current block uses the first transformation mode, and the corresponding transformation kernel index when the current block uses the first transformation mode. In this case, the value of the third syntax element can be 0, 1, 2, 3, 4, 5, 6, etc.

[0450] In one specific embodiment, if the value of the third syntax element is a third value, it is determined that the current block does not use the first transformation mode; if the value of the third syntax element is a fourth value, it is determined that the current block uses the first transformation mode and the corresponding transformation kernel index. The third value can be set to 0, and the fourth value can be set to a non-zero value, such as 1, 2, 3, 4, 5, 6, etc.

[0451] In other words, in this embodiment of the application, the transform kernel index of the current block for LFNST / NSPT can also be represented by lfnst_nspt_index. Here, lfnst_nspt_index of 0 indicates that the current block does not use LFNST / NSPT. Each transform kernel group of LFNST / NSPT in the ECM has 3 transform kernels, so a value of lfnst_nspt_index of 1, 2, or 3 indicates that the current block uses the first, second, or third transform kernel of the selected transform kernel group of LFNST / NSPT.

[0452] In this embodiment, the current block uses OBIP mode, in which case it has more than one selectable transform core group. For example, it has two selectable transform core groups, then the possible values ​​of lfnst_nspt_index are 0, 1, 2, 3, 4, 5, 6. Among them, 1, 2, and 3 correspond to the 3 transform cores of the first candidate transform core group, and 4, 5, and 6 correspond to the 3 transform cores of the second candidate transform core group.

[0453] For example, in this case, the encoder can also add possible values ​​for existing syntax elements instead of adding new syntax elements. For instance, `lfnst_nspt_index` can be used to indicate the transform kernel of LFNST or NSPT. It is known that each transform kernel group of LFNST / NSPT has 3 transform kernels. For a typical prediction mode, the possible values ​​of `lfnst_nspt_index` are 0, 1, 2, and 3. Specifically, if `lfnst_nspt_index` is 0, it means LFNST or NSPT is not used; if `lfnst_nspt_index` is 1, it means the first transform kernel in the selected transform kernel group is used; if `lfnst_nspt_index` is 2, it means the second transform kernel in the selected transform kernel group is used; and if `lfnst_nspt_index` is 3, it means the third transform kernel in the selected transform kernel group is used. For OBIP mode, the possible values ​​of `lfnst_nspt_index` are 0, 1, 2, 3, 4, 5, and 6. Specifically, if `lfnst_nspt_index` is 0, it means LFNST or NSPT is not used; if `lfnst_nspt_index` is 1, it means the first transform kernel in the first candidate transform kernel group is used; if `lfnst_nspt_index` is 2, it means the second transform kernel in the first candidate transform kernel group is used; if `lfnst_nspt_index` is 3, it means the third transform kernel in the first candidate transform kernel group is used; if `lfnst_nspt_index` is 4, it means the first transform kernel in the second candidate transform kernel group is used; if `lfnst_nspt_index` is 5, it means the second transform kernel in the second candidate transform kernel group is used; and if `lfnst_nspt_index` is 6, it means the third transform kernel in the second candidate transform kernel group is used. This determines the transform kernel for the current block.

[0454] S3602, determine the residual value of the current block based on the predicted value of the current block, and transform the residual value of the current block according to the transformation kernel to determine the transformation coefficient of the current block.

[0455] It should be noted that, in the embodiments of this application, after determining the predicted value of the current block, the method may further include: performing a subtraction operation on the initial value of the current block and the predicted value of the current block to determine the residual value of the current block.

[0456] It should also be noted that, in the embodiments of this application, when transforming the residual value of the current block according to the transform kernel to determine the transform coefficients of the current block, it may include: performing an inseparable fundamental transform on the residual value of the current block according to the transform kernel to determine the transform coefficients of the current block; or, performing a discrete cosine transform on the residual value of the current block to determine the transform block of the current block; and performing a low-frequency inseparable transform on the transform block of the current block according to the transform kernel to determine the transform coefficients of the current block.

[0457] In one specific embodiment, if the size parameters of the current block satisfy a first condition, then an inseparable fundamental transformation is performed on the residual value of the current block according to the transformation kernel to determine the transformation coefficients of the current block; if the size parameters of the current block satisfy a second condition, then a discrete cosine transform is performed on the residual value of the current block to determine the transformed block of the current block; and a low-frequency inseparable transform is performed on the transformed block of the current block according to the transformation kernel to determine the transformation coefficients of the current block.

[0458] Here, the size parameters of the current block satisfy the first condition, including: the size parameters of the current block are relatively small, for example, the size parameters of the current block are less than a certain threshold. That is, for smaller blocks, the NSPT transform kernel is used here, that is, the NSPT transform is performed on the residual block of the current block according to the transform kernel to determine the transform coefficients of the current block.

[0459] Here, the size parameters of the current block satisfy the second condition, including: the size parameters of the current block are relatively large, for example, the size parameters of the current block are greater than a certain threshold. That is, for larger blocks, the LFNST transform kernel is used here, that is, first the DCT2 basic transform is performed on the residual value of the current block, and then the LFNST transform is performed on the transformed block of the current block according to the transform kernel to determine the transform coefficients of the current block.

[0460] S3603 encodes the transform coefficients of the current block and writes the resulting encoded bits into the bitstream.

[0461] It should be noted that, in the embodiments of this application, when encoding the transform coefficients of the current block, the method may include: quantizing the transform coefficients of the current block to determine the quantization coefficients of the current block; encoding the quantization coefficients of the current block and writing the obtained encoded bits into the bitstream.

[0462] In simple terms, after determining the prediction block, the encoder derives candidate texture feature indices from the prediction block, and then determines the transform kernel group of NSPT / LFNST based on the candidate texture feature indices. If a transform kernel group has multiple transform kernels to choose from, the encoder tries each transform kernel in that transform kernel group.

[0463] For the NSPT transform, the residual block can be forward-transformed using NSPT to obtain transform coefficients. These transform coefficients are then quantized to obtain quantized coefficients, which are then entropy-encoded. Entropy encoding yields the overhead cost in the bitstream under this transform kernel. The quantized coefficients are inverse-quantized to obtain the decoded transform coefficients, and the decoded transform coefficients are then subjected to inverse NSPT to obtain the decoded residual block. The decoded transform coefficients may differ from the original transform coefficients because quantization is lossy; similarly, the decoded residual block may also differ from the original residual block. A reconstructed block is obtained from the decoded residual block and the predicted block. The distortion cost is obtained from the reconstructed block and the original image of the current block. The cost of encoding using the current NSPT transform kernel is the overhead cost plus the distortion cost. The costs of several transform kernels are compared, and the minimum cost is selected as the optimal NSPT for the current block.

[0464] If using the LFNST transform, a forward transform of the residual block can be performed using DCT2, followed by a forward transform using LFNST to obtain transform coefficients. These transform coefficients are then quantized to obtain quantized coefficients, which are then entropy-encoded. Entropy encoding provides the overhead cost in the bitstream under this transform kernel. Inverse quantization of the quantized coefficients yields the decoded transform coefficients. An inverse LFNST transform is then performed on these decoded transform coefficients, followed by an inverse DCT2 transform to obtain the decoded residual block. The decoded transform coefficients may differ from the original transform coefficients because quantization is lossy; similarly, the decoded residual block may also differ from the original residual block. A reconstructed block is obtained from the decoded residual block and the predicted block. The distortion cost is obtained from the reconstructed block and the original image of the current block. The cost of encoding using the current NSPT transform kernel is the overhead cost plus the distortion cost. The costs of several transform kernels are compared, and the minimum cost is selected as the optimal NSPT for the current block.

[0465] It should be noted that, in the embodiments of this application, the residual block of the current block may include the residual value of at least one pixel in the current block, and can therefore be simply referred to as the "residual value of the current block"; the prediction block of the current block may include the predicted value of at least one pixel in the current block, and can therefore be simply referred to as the "predicted value of the current block"; similarly, the reconstruction block of the current block may include the reconstruction value of at least one pixel in the current block, and can therefore be simply referred to as the "reconstructed value of the current block".

[0466] It should also be noted that, in the embodiments of this application, the "transformation" of the residual block at the encoding end can also be called a "positive transform," specifically referring to the transformation from the spatial domain to the frequency domain to remove the correlation of the residual. It should be noted that if the standard only specifies decoding, then the "transformation" in the standard text refers to the decoding part, specifically the "inverse transform" in this document.

[0467] In another embodiment of this application, a bitstream is provided, wherein the bitstream is generated by bit encoding according to the encoding method described in the foregoing embodiments. The information to be encoded in this encoding method may include at least one of the following: transform coefficients of the current block, transform kernel group index of the current block, transform kernel index of the current block, value of a first syntax element, value of a second syntax element, and value of a third syntax element. Here, this information to be encoded is encoded to be written into the bitstream.

[0468] In this embodiment, the first syntax element indicates whether the current block uses the first prediction mode, the value of the second syntax element indicates the transform kernel group index of the current block, and the value of the third syntax element indicates whether the current block uses the first transform mode and the corresponding transform kernel index when the current block uses the first transform mode. For example, taking the third syntax element as an example, if the value of the third syntax element is 0, it is determined that the current block does not use the first transform mode; if the value of the second syntax element is non-zero, it is determined that the current block uses the first transform mode and the corresponding transform kernel index. For example, if the value of the second syntax element is 1, it is determined that the current block uses the first transform kernel; if the value of the second syntax element is 2, it is determined that the current block uses the second transform kernel, etc., without any limitation.

[0469] This application provides an encoding method, specifically an intra-frame prediction mode and its transformation scheme based on occurrence statistics. The method involves: determining one or more pre-coded blocks; determining the statistical results of at least two candidate intra-frame prediction modes in a first mode set based on the prediction modes used by each of the pre-coded blocks; determining at least two intra-frame prediction modes for the current block in the first mode set based on the statistical results; and predicting the current block based on the at least two intra-frame prediction modes to determine the predicted value of the current block. In this way, statistical analysis is performed on the relevant encoded blocks to obtain their statistical results, and the at least two intra-frame prediction modes used by the current block are derived from these results. Since it is not necessary to explicitly indicate which intra-frame prediction mode to use, overhead in the bitstream can be saved. Furthermore, using at least two intra-frame prediction modes for weighted prediction can improve prediction accuracy and reduce prediction errors to some extent. In addition, deriving the at least two intra-frame prediction modes used by the current block based on statistical results avoids the need for gradient calculation in intra-frame prediction mode derivation, as is required in DIMD mode, reducing computational complexity and thus improving encoding / decoding efficiency and compression performance.

[0470] In another embodiment of this application, based on the encoding / decoding method of the foregoing embodiments, the OBIP mode statistically analyzes the intra-prediction modes used by some related blocks and derives the intra-prediction mode used by the current block based on the intra-prediction modes of the related blocks. Because each intra-prediction block stores the intra-prediction mode information it uses, compared to the DIMD mode which requires gradient calculation for derivation, the OBIP mode can save this step.

[0471] In simple terms, intra-prediction modes include many types, such as the PLANA mode, DC mode, 65 angle prediction modes, and MIP mode already present in VVC. ECM has added new modes like DIMD, TIMD, SGPM, ITMP, and OBIP. Here, we define a first mode set. One possible implementation is that the first mode set only contains angle prediction modes, which could be 65, or even more due to the availability of finer-grained angle prediction modes. Another possible implementation is that the first mode set only contains angle prediction modes, DC mode, and PLANA mode. We will use the second method as an example, where the first mode set corresponds to the 67 intra-prediction modes (0-66) in VVC.

[0472] The OBIP mode can derive several candidate intra-prediction modes from a first-mode set, which are then used to generate a weighted prediction value. During the weighted prediction value generation, PLANA or ITMP modes can also be added. The following section describes how to derive several candidate intra-prediction modes from the first-mode set for the current block.

[0473] Here, a data structure can be constructed to record the occurrence count of various candidate intra-prediction modes in the first mode set. This data structure can be an array, denoted as occurrence

[0067] . It should be noted that this is because the first mode set has 67 intra-prediction modes as an example; if the number is different, the length of the array will be set accordingly. Each number in occurrence

[0067] is initialized to 0.

[0474] (1) Data source.

[0475] Method 1: Statistical analysis using adjacent and non-adjacent blocks.

[0476] The OBIP pattern can count adjacent and non-adjacent blocks. As shown in Figure 29, adjacent blocks are labeled 1-7, and non-adjacent blocks are labeled 8-25. One possible implementation is to use only adjacent blocks. Another possible implementation is to use only non-adjacent blocks. Yet another possible implementation is to use both adjacent and non-adjacent blocks.

[0477] Assume the top-left corner of the current block is at (x, y), the width of the current block is width, and the height is height. Then, for the adjacent blocks of the current block, adjacent block 1 contains the block with coordinates (x-1, y-1), adjacent block 2 contains the block with coordinates (x, y-1), adjacent block 3 contains the block with coordinates (x-1, y), adjacent block 4 contains the block with coordinates (x+width-1, y-1), adjacent block 5 contains the block with coordinates (x+width, y-1), adjacent block 6 contains the block with coordinates (x-1, y+height-1), and adjacent block 7 contains the block with coordinates (x-1, y+height).

[0478] Non-adjacent blocks can also be positioned according to certain rules. For example, the position of a non-adjacent block can be calculated based on (x, y), width, and height. In Figure 29, the distance between each cell in the vertical direction is height, and the distance between each cell in the horizontal direction is width. Therefore, for the non-adjacent blocks of the current block, non-adjacent block 8 contains coordinates (x-width-1, y-height-1), non-adjacent block 9 contains coordinates (x+2*width-1, y-height-1), non-adjacent block 10 contains coordinates (x-width-1, y+2*height-1), non-adjacent block 11 contains coordinates (x-2*width-1, y-2*height-1), non-adjacent block 12 contains coordinates (x+width / 2, y-2*height-1), and so on.

[0479] It is important to note that the block containing a certain coordinate refers to the CU or PU containing that coordinate. After identifying the CU or PU containing a certain coordinate, the intra-prediction mode of that CU or PU can be obtained.

[0480] Method 2: All blocks within the preset range.

[0481] Here you can set a range. Let's assume the top-left corner of the current block is at (x, y), the current block's width is `width`, and its height is `height`. You can set a preset range based on the current block's position and its width and height.

[0482] As shown in Figure 30, a possible preset range is the range of all encoded and decoded blocks within a rectangular area with the top left corner at (x-2*width, y-2*height), the top right corner at (x+3*width-1, y-2*height), and the bottom left corner at (x-2*width, y+3*height-1). If the minimum size of the CU or PU is 4×4, then all 4×4 blocks within the search range can be counted. For example, the first 4×4 block in the top left corner contains the coordinates (x-2*width, y-2*height). Increasing the horizontal direction by 4 allows scanning all 4×4 blocks in the same row, and increasing the vertical direction by 4 allows scanning all columns. Here, "block containing a certain coordinate" refers to the CU or PU containing that coordinate. After determining the CU or PU containing a certain coordinate, the intra-frame prediction mode of that CU or PU can be obtained.

[0483] (2) Statistical data.

[0484] For all blocks found in the data source.

[0485] (a) If this block is an intra-coded block:

[0486] If this block uses a candidate intra-prediction mode from the first mode set, and its mode number is X, then accumulate occurrence[X].

[0487] If this block uses DIMD mode, TIMD mode, SGPM mode, OBIP mode, etc., it may use multiple candidate intra-prediction modes in the first mode set. Then each candidate intra-prediction mode is accumulated separately. For example, if it uses mode X1 and X2, then occurrence[X1] and occurrence[X2] are accumulated separately.

[0488] If this block uses MIP, ITMP, EIP, IBC, or other modes, and it does not use any candidate intra-prediction modes from the first mode set for prediction, one possible implementation is to not accumulate any modes from the first mode set. Another possible implementation is to use gradient statistics to derive a candidate intra-prediction mode from the first mode set, with mode number X, for the reconstructed or predicted block of this block, and then accumulate occurrence[X].

[0489] (b) If this block is an inter-frame coded block:

[0490] This block does not use any modes from the first mode set for prediction. One possible implementation is to not accumulate any modes from the first mode set. Another possible implementation is to use gradient statistics to derive a candidate intra-prediction mode from the first mode set for the reconstructed or predicted block of this block, with mode number X, and then accumulate occurrence[X]. The process of deriving a candidate intra-prediction mode from the first mode set for the reconstructed or predicted block of this block using gradient statistics can be performed during the encoding and decoding of this block, and the derived candidate intra-prediction mode from the first mode set can be saved. This way, the derivation does not have to be performed when the current block needs to use it, which can save computational workload.

[0491] (3) Method for accumulating occurrence statistics.

[0492] For pattern X, one possible accumulation method is occurrence[X] = occurrence[X] + 1. This method can be used for both data source method one (using adjacent and non-adjacent blocks for statistics) and data source method two (all blocks within a preset range). Another possible accumulation method is to consider the size of the block, i.e., occurrence[X] = occurrence[X] + width * height.

[0493] (4) Use gradient statistics to derive a candidate intra-frame prediction mode within a first mode set for a reconstructed or predicted block of a block.

[0494] For example, the gradient can be computed using the Sobel operator. An example of the Sobel operator is as follows:

[0495] Operators for horizontal gradient values:

[0496] Operator for vertical gradient values:

[0497] Thus, assuming the sample value at the sample location (x, y) of the reconstructed or predicted block is P x,y Then the horizontal gradient value grad x and vertical gradient value grad y The calculation is shown in equations (7) and (8) above.

[0498] An example of the gradient intensity denoted as amp is amp = abs(grad). x )+abs(grad y ).

[0499] According to grad x and grad y Exporting a virtual intra-prediction mode can be achieved through a table lookup. For example, if abs(grad...) x ) equals 0 and abs(grad) y If abs(grad) is not equal to 0, then there is horizontal texture, corresponding to intra-prediction mode 18 in VVC. y ) equals 0 and abs(grad) x If abs(grad) is not equal to 0, then there is vertical texture, corresponding to intra-prediction mode 50 in VVC. x ) and abs(grad y If abs(grad) are not equal to 0, then... x ) equals abs(grad y ), and grad x and grady The symbols are the same, corresponding to intra-prediction mode 34 in VVC. If abs(grad x ) equals twice abs(grad) y ), and grad x and grad y The symbols are the same, corresponding to intra-prediction mode 40 in VVC. Other cases can be determined by looking up a table using the same principle.

[0500] For example, the above derivation is performed on all points within a reconstructed or predicted block except for the outermost row and column. Figure 32 shows an 8×8 block; the above derivation can be performed on the 6×6 samples within it. Here, an array AmpAccumulate

[0067] is used, with each item initialized to 0. If the intra-prediction mode derived from a sample location is X, the amp calculated for that sample location is accumulated into AmpAccumulate[X]. The mode with the highest value in AmpAccumulate

[0067] is taken as the candidate intra-prediction mode derived from gradient statistics for this block's reconstructed or predicted block.

[0501] (5) Determine the intra-prediction mode and weights used by OBIP.

[0502] After the above operations, occurrence

[0067] has been statistically analyzed. Based on occurrence

[0067] , the intra-prediction mode and weights used by OBIP can be determined.

[0503] One possible implementation is similar to the weighted prediction method of DIMD patterns in related technologies, selecting the two patterns with the largest values ​​in occurrence

[0067] as M1 and M2, and weighting M1, M2 and PLANAR. The weights of M1, M2 and PLANAR are W1, W2 and W3, respectively, as shown in the aforementioned equations (9), (10) and (11).

[0504] Furthermore, let the predicted values ​​of M1, M2, and PLANA be Pred1, Pred2, and Pred3, respectively. Then the predicted value of the OBIP mode is Pred. OBIP Specifically, as shown in equation (12) above.

[0505] The example above uses two intra-prediction modes from the first mode set and the PLANAR mode for weighted calculation. Understandably, more intra-prediction modes from the first mode set could also be used, such as 3, 4, or 5; additionally, the PLANAR mode could be replaced with other modes, such as the ITMP mode.

[0506] (6) OBIP control.

[0507] Here, a block-level syntax element, such as CU or PU, can be set to indicate whether the current block uses OBIP mode, for example, cu_obip_flag. If cu_obip_flag has a value of 1, it means that the current CU uses OBIP mode for prediction; if cu_obip_flag has a value of 0, it means that the current CU does not use OBIP mode for prediction. If cu_obip_flag does not appear, then its value is inferred to be 0.

[0508] In one specific embodiment, the workflow of OBIP mode prediction includes: statistically analyzing the occurrence of candidate intra-prediction modes in a first mode set in a specified block or region; determining the intra-prediction modes and their corresponding weights based on the statistical results; and generating prediction values ​​based on the determined intra-prediction modes and their weights.

[0509] It should be noted that the data source for the OBIP mode is the same for both the encoder and decoder. Therefore, the encoder and decoder will obtain the same intra-frame prediction mode and weights, and generate the same prediction values ​​using the OBIP mode.

[0510] At the encoding end, the encoder attempts to encode using the OBIP mode and determines the encoding cost of using the OBIP mode, such as rate-distortion cost. The encoder also attempts to encode using other modes and determines their encoding costs. The encoder selects the mode with the lowest encoding cost as the prediction mode for the current block. For example, the encoder can write the value of cu_obip_flag into the bitstream.

[0511] On the decoding side, the decoder will parse the value of cu_obip_flag. If the value of cu_obip_flag is 1, then OBIP mode can be used for prediction.

[0512] (6) Adaptation of transformation mode.

[0513] In VVC and ECM, the transform kernel group of LFNST or NSPT is derived by default from the intra-prediction mode of the current block. This setting is reasonable for using only a single intra-prediction mode, such as PLANA, DC, or angle prediction modes. However, for OBIP mode, because it uses multiple prediction modes weighted together (e.g., weighted by two intra-prediction modes from the first mode set and the PLANA mode), its predicted values ​​contain characteristics of both M1 and M2. Its residuals may also contain characteristics of M1 and M2, and possibly other characteristics as well, but the residuals are correlated with both M1 and M2. Therefore, for OBIP mode, when determining the transform based on the prediction mode (LFNST or NSPT), an additional option can be added. A syntax element can be used to indicate whether the transform kernel (group) is determined based on M1 or M2. For example, a syntax element `lfnst_nspt_set_index` can be set. If lfnst_nspt_set_index is 0, the first candidate transform kernel group is used; if lfnst_nspt_set_index is 1, the second candidate transform kernel group is used. The first candidate transform kernel group is determined by the first intra-prediction mode (M1), and the second candidate transform kernel group is determined by the second intra-prediction mode (M2). Optionally, for special cases, such as when the transform kernel groups determined by the first intra-prediction mode (M1) and the second intra-prediction mode (M2) are exactly the same, other intra-prediction modes can be used, such as the third intra-prediction mode involved in the prediction, or the PLANAR mode, etc.

[0514] Alternatively, instead of adding new syntax elements, you can add possible values ​​for existing syntax elements. For example, `lfnst_nspt_index` can be used to indicate the transform kernel of LFNST or NSPT. Given that each transform kernel group of LFNST / NSPT has three transform kernels, for the general mode, the possible values ​​for `lfnst_nspt_index` are 0, 1, 2, and 3. A value of 0 indicates that LFNST or NSPT is not used; a value of 1 indicates that the first transform kernel in the selected transform kernel group is used; a value of 2 indicates that the second transform kernel in the selected transform kernel group is used; and a value of 3 indicates that the third transform kernel in the selected transform kernel group is used. For OBIP mode, the possible values ​​for `lfnst_nspt_index` are 0, 1, 2, 3, 4, 5, and 6. The value of lfnst_nspt_index is 0, indicating that LFNST or NSPT is not used; the value of lfnst_nspt_index is 1, indicating that the first transform kernel in the first candidate transform kernel group is used; the value of lfnst_nspt_index is 2, indicating that the second transform kernel in the first candidate transform kernel group is used; the value of lfnst_nspt_index is 3, indicating that the third transform kernel in the first candidate transform kernel group is used; the value of lfnst_nspt_index is 4, indicating that the first transform kernel in the second candidate transform kernel group is used; the value of lfnst_nspt_index is 5, indicating that the second transform kernel in the second candidate transform kernel group is used; and the value of lfnst_nspt_index is 6, indicating that the third transform kernel in the second candidate transform kernel group is used.

[0515] In a specific embodiment, taking the decoding end as an example, the workflow of OBIP mode prediction plus transform includes: parsing the bitstream to determine whether the current block uses OBIP mode; if the current block uses OBIP mode, statistically analyzing the occurrence of candidate intra-prediction modes in the first mode set in a specified block or region; determining the intra-prediction mode and its corresponding weight based on the statistical results; generating a prediction value based on the determined intra-prediction mode and its weight; if the current block uses LFNST / NSPT for transform, parsing the bitstream to determine the transform kernel group index; determining the transform kernel group based on the transform kernel group index, performing transform using the transform kernel in the transform kernel group to determine the residual value; and generating a reconstructed value based on the prediction value and the residual value.

[0516] At the encoding end, the encoder attempts to encode in OBIP mode. After generating the predicted value, if there is a prediction residual, the encoder tries various transformation methods. For example, as mentioned above, LFNST or NSPT can have two transform kernel groups to choose from, and each transform kernel group has three transform kernels to choose from. The encoder then tries these six possible LFNST or NSPT transforms and selects the transform kernel with the lowest rate-distortion cost. If the encoder determines that the current block uses OBIP mode and that the current block uses LFNST / NSPT for transformation, then the transform kernel group index needs to be written to the bitstream.

[0517] In this application embodiment, the specific implementation of the aforementioned embodiments is described in detail through the above embodiments. It can be seen that, according to the technical solution of the aforementioned embodiments, the intra-prediction mode of the relevant blocks in the encoded and decoded area is statistically analyzed to obtain relevant information and deduce the intra-prediction mode used by the current block. Since it is not necessary to explicitly indicate which intra-prediction mode to use, the overhead in the bitstream can be saved. Moreover, using multiple intra-prediction modes for weighted prediction can also improve the accuracy of prediction and reduce prediction error to a certain extent. Overall, it can improve compression performance.

[0518] In another embodiment of this application, based on the same inventive concept as the foregoing embodiments, FIG37 is a schematic diagram of the composition structure of an encoder provided in an embodiment of this application. As shown in FIG37, the encoder 370 may include a first determining unit 3701 and a first predicting unit 3702, wherein:

[0519] The first determining unit 3701 is configured to determine one or more pre-coded blocks; and to determine statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by each of the one or more pre-coded blocks.

[0520] The first determining unit 3701 is further configured to determine at least two intra-prediction modes of the current block in the first mode set based on statistical results.

[0521] The first prediction unit 3702 is configured to predict the current block based on at least two intra-frame prediction modes and determine the predicted value of the current block.

[0522] In some embodiments, the first determining unit 3701 is further configured to perform the following steps when the current block uses a first prediction mode: determining one or more pre-coded blocks; determining statistical results of at least two candidate intra-prediction modes in a first mode set based on the prediction modes used by the one or more pre-coded blocks respectively; determining at least two intra-prediction modes of the current block in the first mode set based on the statistical results; and predicting the current block based on the at least two intra-prediction modes to determine the predicted value of the current block.

[0523] In some embodiments, the first determining unit 3701 is further configured to determine multiple candidate prediction modes, wherein the multiple candidate prediction modes include at least a first prediction mode; calculate the encoding cost of the current block according to the multiple candidate prediction modes, and determine the cost result corresponding to each of the multiple candidate prediction modes; determine the minimum cost result among the cost results corresponding to each of the multiple candidate prediction modes; and determine whether the current block uses the first prediction mode according to the candidate prediction mode corresponding to the minimum cost result.

[0524] In some embodiments, the first determining unit 3701 is further configured to determine that the current block uses the first prediction mode when the candidate prediction mode corresponding to the minimum cost result is the first prediction mode; and to determine that the current block does not use the first prediction mode when the candidate prediction mode corresponding to the minimum cost result is not the first prediction mode.

[0525] In some embodiments, referring to FIG37, the encoder 370 may further include an encoding unit 3703, wherein: the first determining unit 3701 is further configured to determine the value of a first syntax element, the first syntax element being used to indicate whether the current block uses a first prediction mode; the encoding unit 3703 is configured to encode the value of the first syntax element and write the obtained encoded bits into the bitstream.

[0526] In some embodiments, the first determining unit 3701 is further configured to determine at least one encoded candidate block, wherein the candidate block includes: neighboring blocks of the current block, and / or, non-neighboring blocks of the current block; and to determine one or more preset blocks based on at least one candidate block.

[0527] In some embodiments, the first determining unit 3701 is further configured to determine a preset range based on the current block; and to determine one or more preset blocks based on all encoded blocks within the preset range.

[0528] In some embodiments, the first mode set includes at least one of the following: angle prediction mode, DC mode, and PLANA mode.

[0529] In some embodiments, referring to FIG37, the encoder 370 may further include a first statistics unit 3704, configured to not count the first preset block when the first preset block is an inter-frame prediction block; wherein the first preset block is any one of one or more preset blocks.

[0530] In some embodiments, the first statistical unit 3704 is further configured to, when the first preset block is an inter-frame prediction block, determine a first intra-frame prediction mode derived from candidate samples of the first preset block; perform cumulative calculation on the first intra-frame prediction mode to determine the statistical result of the first intra-frame prediction mode in the first mode set; wherein the first preset block is any one of one or more preset blocks.

[0531] In some embodiments, the first statistical unit 3704 is further configured to perform cumulative calculations on the first intra-frame prediction mode when the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, and determine the statistical result of the first intra-frame prediction mode in the first mode set.

[0532] In some embodiments, the first determining unit 3701 is further configured to determine candidate samples of the first preset block based on at least a portion of the samples in the reconstruction block of the first preset block; or, to determine candidate samples of the first preset block based on at least a portion of the samples in the prediction block of the first preset block.

[0533] In some embodiments, the first determining unit 3701 is further configured to perform cumulative calculations on the first intra-frame prediction mode based on the size parameter value of the first preset block, and determine the statistical results of the first intra-frame prediction mode in the first mode set.

[0534] In some embodiments, the first determining unit 3701 is further configured to perform a cumulative operation of incrementing the first intra-frame prediction mode by one to determine the statistical result of the first intra-frame prediction mode in the first mode set.

[0535] In some embodiments, the first determining unit 3701 is further configured to sort the candidate intra-prediction modes in the first mode set in descending order according to the statistical results, and determine the at least two candidate intra-prediction modes that are ranked first as at least two intra-prediction modes of the current block.

[0536] In some embodiments, the first prediction unit 3702 is further configured to determine the weight values ​​of at least two intra-frame prediction modes based on the statistical results corresponding to at least two intra-frame prediction modes; and to perform weighted prediction on the current block based on the at least two intra-frame prediction modes and their respective weight values ​​to determine the predicted value of the current block.

[0537] In some embodiments, referring to FIG37, the encoder 370 may further include a first transformation unit 3705, wherein: the first determining unit 3701 is further configured to determine the transformation kernel of the current block and determine the residual value of the current block based on the prediction value of the current block; the first transformation unit 3705 is configured to transform the residual value of the current block according to the transformation kernel to determine the transformation coefficients of the current block; the encoding unit 3703 is further configured to encode the transformation coefficients of the current block and write the obtained encoded bits into the bit stream.

[0538] In some embodiments, the first determining unit 3701 is further configured to determine the transform kernel group of the current block; and to determine the transform kernel of the current block based on the transform kernel group.

[0539] In some embodiments, the first determining unit 3701 is further configured to perform encoding cost calculation on the current block based on at least two candidate transform core groups, determine the cost result corresponding to each of the at least two candidate transform core groups; and determine the minimum cost result among the cost results corresponding to each of the at least two candidate transform core groups, and determine the candidate transform core group corresponding to the minimum cost result as the transform core group of the current block.

[0540] In some embodiments, the first determining unit 3701 is further configured to determine the transform kernel group index of the current block; wherein the transform kernel group index is used to indicate the number of the transform kernel group of the current block in at least two candidate transform kernel groups; the encoding unit 3703 is further configured to encode the transform kernel group index of the current block and write the obtained encoded bits into the code stream.

[0541] In some embodiments, the first determining unit 3701 is further configured to determine at least two candidate transform cores included in the transform core group; calculate the encoding cost of the current block based on the at least two candidate transform cores, determine the cost result corresponding to each of the at least two candidate transform cores; and determine the minimum cost result among the cost results corresponding to each of the at least two candidate transform cores, and determine the candidate transform core corresponding to the minimum cost result as the transform core of the current block.

[0542] In some embodiments, the first determining unit 3701 is further configured to determine the transform kernel index of the current block, wherein the transform kernel index is used to indicate the number of the transform kernel of the current block in the transform kernel group, and the first candidate list indicates the transform kernels included in at least two candidate transform kernel groups; the encoding unit 3703 is further configured to encode the transform kernel index of the current block and write the obtained encoded bits into the bit stream.

[0543] Understandably, in the embodiments of this application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular one. Furthermore, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional module.

[0544] In another embodiment of this application, FIG38 is a schematic diagram of the specific hardware structure of an encoder provided in an embodiment of this application. As shown in FIG38, the encoder 370 may include: a first communication interface 3801, a first memory 3802, and a first processor 3803; the various components are coupled together through a first bus system 3804. It is understood that the first bus system 3804 is used to realize the connection and communication between these components. In addition to a data bus, the first bus system 3804 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all buses are labeled as the first bus system 3804 in FIG38.

[0545] The first communication interface 3801 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;

[0546] The first memory 3802 is used to store computer programs that can run on the first processor 3803;

[0547] The first processor 3803 is configured to, when running the computer program, execute:

[0548] Identify one or more pre-defined blocks that have been encoded; determine the statistical results of at least two candidate intra-prediction modes in a first mode set based on the prediction modes used by each of the one or more pre-defined blocks; determine at least two intra-prediction modes for the current block in the first mode set based on the statistical results; predict the current block based on the at least two intra-prediction modes to determine the predicted value of the current block.

[0549] It is understood that the first memory 3802 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The first memory 3802 of the system and method described in this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0550] The first processor 3803 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the first processor 3803 or by instructions in software form. The first processor 3803 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the first memory 3802. The first processor 3803 reads the information in the first memory 3802 and completes the steps of the above method in conjunction with its hardware.

[0551] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof. For software implementation, the technology described in this application can be implemented through modules (e.g., procedures, functions, etc.) that perform the functions described in this application. Software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0552] Alternatively, as another embodiment, the first processor 3803 is further configured to perform the method described in any of the foregoing embodiments when running the computer program.

[0553] This embodiment provides an encoder in which statistical analysis is performed on the encoded related blocks to obtain the statistical results of these related blocks, and at least two intra-prediction modes used in the current block are derived based on the statistical results. Since it is not necessary to explicitly indicate which intra-prediction mode to use, the overhead in the bitstream can be saved. Moreover, using at least two intra-prediction modes for weighted prediction can improve the accuracy of prediction and reduce prediction errors to a certain extent. In addition, deriving the at least two intra-prediction modes used in the current block based on the statistical results can avoid the need for intra-prediction mode derivation through gradient calculation, such as in DIMD mode, thus reducing computational complexity, thereby improving encoding and decoding efficiency and improving compression performance.

[0554] In another embodiment of this application, based on the same inventive concept as the foregoing embodiments, FIG39 is a schematic diagram of the composition structure of a decoder provided in an embodiment of this application. As shown in FIG39, the decoder 390 may include a second determining unit 3901 and a second predicting unit 3902, wherein:

[0555] The second determining unit 3901 is configured to determine one or more preset blocks that have been decoded; and to determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set according to the prediction modes used by each of the one or more preset blocks.

[0556] The second determining unit 3901 is further configured to determine at least two intra-frame prediction modes of the current block in the first mode set based on statistical results.

[0557] The second prediction unit 3902 is configured to predict the current block based on at least two intra-frame prediction modes and determine the predicted value of the current block.

[0558] In some embodiments, the second determining unit 3901 is further configured to determine at least one candidate block that has been decoded, wherein the candidate block includes: neighboring blocks of the current block, and / or, non-neighboring blocks of the current block; and to determine one or more preset blocks based on at least one candidate block.

[0559] In some embodiments, the second determining unit 3901 is further configured to determine a preset range based on the current block; and to determine one or more preset blocks based on all decoded blocks within the preset range.

[0560] In some embodiments, the first mode set includes at least one of the following: angle prediction mode, DC mode, and PLANA mode.

[0561] In some embodiments, referring to FIG39, the decoder 390 may further include a second statistics unit 3903, configured to not count the first preset block when the first preset block is an inter-frame prediction block; wherein the first preset block is any one of one or more preset blocks.

[0562] In some embodiments, the second statistical unit 3903 is further configured to, when the first preset block is an inter-frame prediction block, determine a first intra-frame prediction mode derived from candidate samples of the first preset block; perform cumulative calculation on the first intra-frame prediction mode to determine the statistical result of the first intra-frame prediction mode in the first mode set; wherein the first preset block is any one of one or more preset blocks.

[0563] In some embodiments, the second statistical unit 3903 is further configured to perform cumulative calculations on the first intra-frame prediction mode when the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, and determine the statistical result of the first intra-frame prediction mode in the first mode set.

[0564] In some embodiments, the second determining unit 3901 is further configured to determine candidate samples of the first preset block based on at least a portion of the samples in the reconstruction block of the first preset block; or, to determine candidate samples of the first preset block based on at least a portion of the samples in the prediction block of the first preset block.

[0565] In some embodiments, the second determining unit 3901 is further configured to perform cumulative calculations on the first intra-frame prediction mode based on the size parameter value of the first preset block, and determine the statistical results of the first intra-frame prediction mode in the first mode set.

[0566] In some embodiments, the second determining unit 3901 is further configured to perform a cumulative operation of incrementing the first intra-frame prediction mode by one to determine the statistical result of the first intra-frame prediction mode in the first mode set.

[0567] In some embodiments, the second statistical unit 3903 is further configured to: when the first preset block is an intra-prediction block, determine the second intra-prediction mode used by the first preset block; when the second intra-prediction mode is a first candidate intra-prediction mode in the first mode set, perform cumulative calculation on the second intra-prediction mode to determine the statistical result of the second intra-prediction mode in the first mode set; wherein the first preset block is any one of one or more preset blocks.

[0568] In some embodiments, the second statistical unit 3903 is further configured to: when the second intra-frame prediction mode is a candidate intra-frame prediction mode in the second mode set, determine at least two third intra-frame prediction modes corresponding to the second intra-frame prediction mode; when at least two third intra-frame prediction modes are at least two candidate intra-frame prediction modes in the first mode set, perform cumulative calculations on the at least two third intra-frame prediction modes respectively, and determine the statistical results of at least two third intra-frame prediction modes in the first mode set; wherein the second mode set is different from the first mode set, and the second mode set includes modes that use at least two intra-frame prediction modes for combined prediction.

[0569] In some embodiments, the second statistics unit 3903 is further configured to not count the first preset block when the second intra-frame prediction mode is a candidate intra-frame prediction mode in the third mode set; wherein the third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode that performs prediction by copying intra-frame blocks, a mode that performs prediction by using an extrapolation filter, and a mode that performs prediction by using matrix operations.

[0570] In some embodiments, the second statistical unit 3903 is further configured to, when the second intra-frame prediction mode is a candidate intra-frame prediction mode in the third mode set, determine a fourth intra-frame prediction mode derived from candidate samples of the first preset block; perform cumulative calculation on the fourth intra-frame prediction mode to determine the statistical result of the fourth intra-frame prediction mode in the first mode set; wherein the third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode that performs prediction by copying intra-frame blocks, a mode that performs prediction by using an extrapolation filter, and a mode that performs prediction by using matrix operations.

[0571] In some embodiments, the second determining unit 3901 is further configured to sort the candidate intra-prediction modes in the first mode set in descending order according to the statistical results, and determine the at least two candidate intra-prediction modes that are ranked first as at least two intra-prediction modes of the current block.

[0572] In some embodiments, the second prediction unit 3902 is further configured to determine the weight values ​​of at least two intra-frame prediction modes based on the statistical results corresponding to at least two intra-frame prediction modes; and to perform weighted prediction on the current block based on the at least two intra-frame prediction modes and their respective weight values ​​to determine the predicted value of the current block.

[0573] In some embodiments, referring to FIG39, the decoder 390 may further include a decoding unit 3904 configured to decode a bitstream, determine the value of a first syntax element; when the first syntax element indicates that the current block uses a first prediction mode, perform the following steps: determine one or more preset blocks that have been decoded; determine the statistical results of at least two candidate intra-prediction modes in a first mode set based on the prediction modes used by each of the one or more preset blocks; determine at least two intra-prediction modes of the current block in the first mode set based on the statistical results; and predict the current block based on the at least two intra-prediction modes to determine the predicted value of the current block.

[0574] In some embodiments, referring to FIG39, the decoder 390 may further include a second transformation unit 3905, wherein: the second determination unit 3901 is further configured to determine the transformation kernel of the current block; the second transformation unit 3905 is configured to transform the transformation coefficients of the current block according to the transformation kernel to determine the residual block of the current block; the second determination unit 3901 is further configured to determine the reconstructed value of the current block according to the residual value of the current block and the predicted value of the current block.

[0575] In some embodiments, the decoding unit 3904 is further configured to decode the bitstream and determine the transform kernel group index of the current block; the second determining unit 3901 is further configured to determine the transform kernel group of the current block according to the transform kernel group index; and to determine the transform kernel of the current block according to the transform kernel group.

[0576] In some embodiments, the decoding unit 3904 is further configured to decode the bitstream and determine the transform kernel index of the current block; the second determining unit 3901 is further configured to determine the transform kernel of the current block based on the transform kernel group and the transform kernel index.

[0577] In some embodiments, the second determining unit 3901 is further configured to determine a first candidate list for the current block, the first candidate list indicating at least two candidate transform kernel groups; and to determine the transform kernel group for the current block based on the first candidate list and the transform kernel group index.

[0578] In some embodiments, when the first candidate list indicates the transform kernels included in the at least two candidate transform kernel groups, the decoding unit 3904 is further configured to decode the bitstream and determine the transform kernel index of the current block; the second determining unit 3901 is further configured to determine the transform kernel of the current block based on the first candidate list and the transform kernel index.

[0579] Understandably, in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular component. Furthermore, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.

[0580] In another embodiment of this application, FIG40 is a schematic diagram of the specific hardware structure of a decoder provided in an embodiment of this application. As shown in FIG40, the decoder 390 may include: a second communication interface 4001, a second memory 4002, and a second processor 4003; the various components are coupled together through a second bus system 4004. It is understood that the second bus system 4004 is used to realize the connection and communication between these components. In addition to a data bus, the second bus system 4004 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all buses are labeled as the second bus system 4004 in FIG40.

[0581] The second communication interface 4001 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;

[0582] The second memory 4002 is used to store computer programs that can run on the second processor 4003;

[0583] The second processor 4003 is configured to perform the following when running the computer program:

[0584] Identify one or more pre-defined blocks that have been decoded; determine the statistical results of at least two candidate intra-prediction modes in the first mode set based on the prediction modes used by each of the one or more pre-defined blocks; determine at least two intra-prediction modes for the current block in the first mode set based on the statistical results; predict the current block based on the at least two intra-prediction modes to determine the predicted value of the current block.

[0585] Alternatively, as another embodiment, the second processor 4003 is also configured to execute the method described in any of the foregoing embodiments when running the computer program.

[0586] It is understood that the second memory 4002 has similar hardware functions to the first memory 3802, and the second processor 4003 has similar hardware functions to the first processor 3803; these will not be elaborated here.

[0587] This embodiment provides a decoder that uses previously decoded related blocks to perform statistics, obtains the statistical results of these related blocks, and derives at least two intra-prediction modes used in the current block based on the statistical results. Since it is not necessary to explicitly indicate which intra-prediction mode to use, it can save overhead in the bitstream. Moreover, using at least two intra-prediction modes for weighted prediction can improve the accuracy of prediction and reduce prediction errors to a certain extent. In addition, deriving at least two intra-prediction modes used in the current block based on statistical results can avoid the need for intra-prediction mode derivation through gradient calculation, as is required in DIMD mode, thus reducing computational complexity, thereby improving encoding and decoding efficiency and improving compression performance.

[0588] In another embodiment of this application, FIG41 is a schematic diagram of the composition structure of an encoding and decoding system provided in an embodiment of this application. As shown in FIG41, the encoding and decoding system 410 may include an encoder 4101 and a decoder 4102.

[0589] In this embodiment, encoder 4101 can be any of the encoders described in the foregoing embodiments, and decoder 4102 can be any of the decoders described in the foregoing embodiments.

[0590] In some embodiments, this application provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor (e.g., a first processor or a second processor), the computer program implements the steps of the encoding / decoding method as described in the foregoing embodiments.

[0591] In some embodiments, this application provides a computer program product, including a computer program or instructions. When executed by a processor (e.g., a first processor or a second processor), the computer program or instructions implement the steps of the encoding / decoding method as described in the foregoing embodiments.

[0592] In some embodiments, this application provides a computer program that, when executed by a processor (e.g., a first processor or a second processor), implements the steps of the encoding / decoding method as described in the foregoing embodiments.

[0593] In some embodiments, this application also provides a computer-readable storage medium storing a bitstream thereon. This bitstream is generated by performing the steps of the encoding method as described in the foregoing embodiments.

[0594] In this embodiment of the application, the information to be encoded in the encoding method may include at least one of the following: the transform coefficients of the current block, the transform kernel group index of the current block, the transform kernel index of the current block, the value of the first syntax element, the value of the second syntax element, and the value of the third syntax element. Here, this information to be encoded is processed to write it into the bitstream.

[0595] In this embodiment, the first syntax element indicates whether the current block uses the first prediction mode, the value of the second syntax element indicates the transform kernel group index of the current block, and the value of the third syntax element indicates whether the current block uses the first transform mode and the corresponding transform kernel index when the current block uses the first transform mode. For example, taking the third syntax element as an example, if the value of the third syntax element is 0, it is determined that the current block does not use the first transform mode; if the value of the second syntax element is non-zero, it is determined that the current block uses the first transform mode and the corresponding transform kernel index. For example, if the value of the second syntax element is 1, it is determined that the current block uses the first transform kernel; if the value of the second syntax element is 2, it is determined that the current block uses the second transform kernel, etc., without any limitation.

[0596] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0597] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0598] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0599] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0600] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0601] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0602] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0603] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0604] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0605] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0606] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0607] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims. Industrial applicability

[0608] In this embodiment, whether at the encoding or decoding end, after determining one or more pre-defined blocks that have been encoded / decoded, the statistical results of at least two candidate intra-prediction modes in the first mode set are determined based on the prediction modes used by each of the pre-defined blocks. Then, based on the statistical results, at least two intra-prediction modes for the current block are determined in the first mode set. Finally, the current block is predicted based on the at least two intra-prediction modes to determine the predicted value of the current block. In this way, statistical analysis is performed on the relevant encoded / decoded blocks to obtain the statistical results of these relevant blocks, and the at least two intra-prediction modes used by the current block are derived based on the statistical results. Since it is not necessary to explicitly indicate which intra-prediction mode to use, the overhead in the bitstream can be saved. Moreover, using at least two intra-prediction modes for weighted prediction can improve the accuracy of prediction and reduce prediction errors to a certain extent. In addition, deriving the at least two intra-prediction modes used by the current block based on the statistical results can avoid the need for intra-prediction mode derivation through gradient calculation, as is required in DIMD mode, thus reducing computational complexity and improving encoding and decoding efficiency, thereby enhancing compression performance.

Claims

1. A decoding method applied to a decoder, the method comprising: Identify one or more pre-defined blocks that have been decoded; Based on the prediction modes used by each of the one or more preset blocks, determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set; Based on the statistical results, at least two intra-frame prediction modes for the current block are determined from the first mode set; The current block is predicted based on the at least two intra-frame prediction modes to determine the predicted value of the current block.

2. The method according to claim 1, wherein, The determination of one or more preset blocks that have been decoded includes: Determine at least one candidate block that has been decoded, wherein the candidate block includes: neighboring blocks of the current block, and / or, non-neighboring blocks of the current block; Based on the at least one candidate block, determine the one or more preset blocks.

3. The method according to claim 1, wherein, The determination of one or more preset blocks that have been decoded includes: A preset range is determined based on the current block; The one or more preset blocks are determined based on all decoded blocks within the preset range.

4. The method according to claim 1, wherein, The first set of modes includes at least one of the following: angle prediction mode, DC mode, and PLANA mode.

5. The method according to claim 1, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an inter-frame prediction block, the first preset block is not counted. Wherein, the first preset block is any one of the one or more preset blocks.

6. The method according to claim 1, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an inter-frame prediction block, a first intra-frame prediction mode derived from candidate samples based on the first preset block is determined. The statistical results of the first intra-frame prediction modes in the first mode set are determined by accumulating the first intra-frame prediction modes. Wherein, the first preset block is any one of the one or more preset blocks.

7. The method according to claim 6, wherein, The method further includes: When the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, the first intra-frame prediction mode is accumulated to determine the statistical result of the first intra-frame prediction mode in the first mode set.

8. The method according to claim 6, wherein, The method further includes: Based on at least a portion of the samples within the reconstructed block of the first preset block, candidate samples for the first preset block are determined; or, Candidate samples for the first preset block are determined based on at least a portion of the samples within the prediction block of the first preset block.

9. The method according to claim 6 or 7, wherein, The step of performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set includes: The first intra-frame prediction mode is cumulatively calculated based on the size parameter value of the first preset block to determine the statistical result of the first intra-frame prediction mode in the first mode set.

10. The method according to claim 6 or 7, wherein, The step of performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set includes: The first intra-frame prediction mode is incremented by one to accumulate the results and determine the statistical results of the first intra-frame prediction mode in the first mode set.

11. The method according to claim 1, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an intra-prediction block, determine the second intra-prediction mode used by the first preset block; When the second intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, the second intra-frame prediction mode is accumulated to determine the statistical result of the second intra-frame prediction mode in the first mode set. Wherein, the first preset block is any one of the one or more preset blocks.

12. The method according to claim 11, wherein, The method further includes: When the second intra-frame prediction mode is a candidate intra-frame prediction mode in the second mode set, at least two third intra-frame prediction modes corresponding to the second intra-frame prediction mode are determined. When at least two third intra-frame prediction modes are at least two candidate intra-frame prediction modes in the first mode set, cumulative calculations are performed on the at least two third intra-frame prediction modes respectively to determine the at least two third intra-frame prediction modes in the first mode set. Statistical results of the pattern; The second mode set is different from the first mode set, and the second mode set includes modes that use at least two intra-frame prediction modes for combined prediction.

13. The method according to claim 11, wherein, The method further includes: When the second intra-frame prediction mode is a candidate intra-frame prediction mode in the third mode set, the first preset block is not counted. The third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode that performs prediction by copying intra-frame blocks, a mode that performs prediction by using an extrapolation filter, and a mode that performs prediction by using matrix operations.

14. The method according to claim 11, wherein, The method further includes: When the second intra-frame prediction mode is a candidate intra-frame prediction mode in the third mode set, a fourth intra-frame prediction mode derived from the candidate samples of the first preset block is determined. The statistical results of the fourth intra-frame prediction mode in the first mode set are determined by performing cumulative calculation on the fourth intra-frame prediction mode. The third mode set is different from the first mode set, and the third mode set includes at least one of the following: a mode that performs prediction by copying intra-frame blocks, a mode that performs prediction by using an extrapolation filter, and a mode that performs prediction by using matrix operations.

15. The method according to claim 1, wherein, The step of determining at least two intra-prediction modes for the current block in the first mode set based on the statistical results includes: Based on the statistical results, the candidate intra-prediction modes in the first mode set are sorted in descending order, and the top two candidate intra-prediction modes are determined as the at least two intra-prediction modes of the current block.

16. The method according to claim 15, wherein, The step of predicting the current block based on the at least two intra-frame prediction modes to determine the predicted value of the current block includes: Based on the statistical results corresponding to the at least two intra-frame prediction modes, determine the weight values ​​of each of the at least two intra-frame prediction modes; The predicted value of the current block is determined by performing a weighted prediction on the current block based on the at least two intra-frame prediction modes and their respective weight values.

17. The method according to any one of claims 1 to 16, wherein, The method further includes: Decode the bitstream to determine the value of the first syntax element; When the first syntax element indicates that the current block uses a first prediction mode, the following steps are performed: determining one or more preset blocks that have been decoded; determining statistical results of at least two candidate intra-prediction modes in a first mode set based on the prediction modes used by each of the one or more preset blocks; determining at least two intra-prediction modes of the current block in the first mode set based on the statistical results; and predicting the current block based on the at least two intra-prediction modes to determine the predicted value of the current block.

18. The method according to any one of claims 1 to 16, wherein, The method further includes: Determine the transform kernel of the current block; The transformation coefficients of the current block are transformed according to the transformation kernel to determine the residual block of the current block; The reconstruction value of the current block is determined based on the residual value of the current block and the predicted value of the current block.

19. The method according to claim 18, wherein, Determining the transform kernel of the current block includes: Decode the bitstream and determine the transform kernel group index of the current block; The transform kernel group of the current block is determined based on the transform kernel group index; The transformation kernel of the current block is determined based on the transformation kernel set.

20. The method according to claim 19, wherein, Determining the transform kernel of the current block based on the transform kernel set includes: Decode the bitstream and determine the transform kernel index of the current block; The transform kernel of the current block is determined based on the transform kernel group and the transform kernel index.

21. The method according to claim 19, wherein, Determining the transform kernel group of the current block based on the transform kernel group index includes: Determine a first candidate list for the current block, the first candidate list indicating at least two candidate transform kernel groups; The transform kernel group of the current block is determined based on the first candidate list and the transform kernel group index.

22. The method according to claim 21, wherein, When the first candidate list indicates the transform kernels included in the at least two candidate transform kernel groups, determining the transform kernel of the current block includes: Decode the bitstream and determine the transform kernel index of the current block; The transform kernel of the current block is determined based on the first candidate list and the transform kernel index.

23. An encoding method applied to an encoder, the method comprising: Identify one or more pre-defined blocks that have already been encoded; Based on the prediction modes used by each of the one or more preset blocks, at least two candidate intra-frame prediction modes in the first mode set are determined. Statistical results of the formula; Based on the statistical results, at least two intra-frame prediction modes for the current block are determined from the first mode set; The current block is predicted based on the at least two intra-frame prediction modes to determine the predicted value of the current block.

24. The method according to claim 23, wherein, The method further includes: When the current block uses a first prediction mode, the steps of determining one or more pre-coded blocks; determining statistical results of at least two candidate intra-prediction modes in a first mode set based on the prediction modes used by each of the one or more pre-coded blocks; determining at least two intra-prediction modes of the current block in the first mode set based on the statistical results; and predicting the current block based on the at least two intra-prediction modes to determine the predicted value of the current block.

25. The method according to claim 24, wherein, The method further includes: Multiple candidate prediction modes are determined, wherein the multiple candidate prediction modes include at least the first prediction mode; The encoding cost of the current block is calculated based on the multiple candidate prediction modes, and the cost result corresponding to each of the multiple candidate prediction modes is determined. Determine the minimum cost result among the cost results corresponding to each of the multiple candidate prediction modes; Based on the candidate prediction pattern corresponding to the minimum cost result, determine whether the current block uses the first prediction pattern.

26. The method of claim 25, wherein, Determining whether the current block uses the first prediction mode based on the candidate prediction mode corresponding to the minimum cost result includes: When the candidate prediction mode corresponding to the minimum cost result is the first prediction mode, it is determined that the current block uses the first prediction mode; When the candidate prediction mode corresponding to the minimum cost result is not the first prediction mode, it is determined that the current block does not use the first prediction mode.

27. The method according to claim 25, wherein, The method further includes: Determine the value of the first syntax element, which is used to indicate whether the current block uses the first prediction mode; The value of the first syntax element is encoded, and the resulting encoded bits are written into the bitstream.

28. The method according to claim 23, wherein, The determination of one or more pre-defined encoded blocks includes: Determine at least one candidate block that has been encoded, wherein the candidate block includes: neighboring blocks of the current block, and / or, non-neighboring blocks of the current block; Based on the at least one candidate block, determine the one or more preset blocks.

29. The method according to claim 23, wherein, The determination of one or more pre-defined encoded blocks includes: A preset range is determined based on the current block; The one or more preset blocks are determined based on all encoded blocks within the preset range.

30. The method according to claim 23, wherein, The first set of modes includes at least one of the following: angle prediction mode, DC mode, and PLANA mode.

31. The method according to claim 23, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an inter-frame prediction block, the first preset block is not counted. Wherein, the first preset block is any one of the one or more preset blocks.

32. The method according to claim 23, wherein, The step of determining the statistical results of at least two candidate intra-frame prediction modes in the first mode set based on the prediction modes used by the one or more preset blocks includes: When the first preset block is an inter-frame prediction block, a first intra-frame prediction mode derived from candidate samples based on the first preset block is determined. The first intra-frame prediction mode is accumulated to determine the statistical results of the first intra-frame prediction mode in the first mode set; Wherein, the first preset block is any one of the one or more preset blocks.

33. The method according to claim 32, wherein, The method further includes: When the first intra-frame prediction mode is a candidate intra-frame prediction mode in the first mode set, the first intra-frame prediction mode is accumulated to determine the statistical result of the first intra-frame prediction mode in the first mode set.

34. The method according to claim 32, wherein, The method further includes: Based on at least a portion of the samples within the reconstructed block of the first preset block, candidate samples for the first preset block are determined; or, Candidate samples for the first preset block are determined based on at least a portion of the samples within the prediction block of the first preset block.

35. The method according to claim 32 or 33, wherein, The step of performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set includes: The first intra-frame prediction mode is cumulatively calculated based on the size parameter value of the first preset block to determine the statistical result of the first intra-frame prediction mode in the first mode set.

36. The method according to claim 32 or 33, wherein, The step of performing cumulative calculations on the first intra-frame prediction modes to determine the statistical results of the first intra-frame prediction modes in the first mode set includes: The first intra-frame prediction mode is incremented by one to accumulate the results and determine the statistical results of the first intra-frame prediction mode in the first mode set.

37. The method according to any one of claims 23 to 36, wherein, The step of determining at least two intra-prediction modes for the current block in the first mode set based on the statistical results includes: Based on the statistical results, the candidate intra-prediction modes in the first mode set are sorted in descending order, and the top two candidate intra-prediction modes are determined as the at least two intra-prediction modes of the current block.

38. The method according to claim 37, wherein, The step of predicting the current block based on the at least two intra-frame prediction modes to determine the predicted value of the current block includes: Based on the statistical results corresponding to the at least two intra-frame prediction modes, determine the weight values ​​of each of the at least two intra-frame prediction modes; The predicted value of the current block is determined by performing a weighted prediction on the current block based on the at least two intra-frame prediction modes and their respective weight values.

39. The method according to any one of claims 23 to 38, wherein, The method further includes: Determine the transform kernel of the current block; The residual value of the current block is determined based on the predicted value of the current block, and the residual value of the current block is transformed according to the transformation kernel to determine the transformation coefficient of the current block; The transform coefficients of the current block are encoded, and the resulting encoded bits are written into the bitstream.

40. The method according to claim 39, wherein, Determining the transform kernel of the current block includes: Determine the transform kernel set of the current block; The transformation kernel of the current block is determined based on the transformation kernel set.

41. The method according to claim 40, wherein, Determining the transform kernel set of the current block includes: The encoding cost of the current block is calculated based on at least two candidate transform kernel groups, and the cost result corresponding to each of the at least two candidate transform kernel groups is determined. The minimum cost result is determined from the cost results corresponding to each of the at least two candidate transformation kernel groups, and the candidate transformation kernel group corresponding to the minimum cost result is determined as the transformation kernel group of the current block.

42. The method according to claim 41, wherein, The method further includes: Determine the transform kernel group index of the current block; wherein the transform kernel group index is used to indicate the number of the transform kernel group of the current block in the at least two candidate transform kernel groups; The transform kernel group index of the current block is encoded, and the resulting encoded bits are written into the bit stream.

43. The method according to claim 40, wherein, Determining the transform kernel of the current block based on the transform kernel set includes: Determine at least two candidate transform kernels included in the transform kernel group; The encoding cost of the current block is calculated based on the at least two candidate transform kernels, and the cost result corresponding to each of the at least two candidate transform kernels is determined. The minimum cost result is determined from the cost results corresponding to the at least two candidate transformation kernels, and the candidate transformation kernel corresponding to the minimum cost result is determined as the transformation kernel of the current block.

44. The method according to claim 43, wherein, The method further includes: Determine the transform kernel index of the current block, wherein the transform kernel index is used to indicate the number of the transform kernel of the current block in the transform kernel group; The transform kernel index of the current block is encoded, and the resulting encoded bits are written into the bitstream.

45. A bitstream, wherein, The bitstream is generated by bit encoding according to the encoding method described in any one of claims 23 to 44; wherein the information to be encoded in the encoding method includes at least one of the following: The transform coefficients of the current block, the transform kernel group index of the current block, the transform kernel index of the current block, and the value of the first syntax element, wherein the first syntax element is used to indicate whether the current block uses the first prediction mode.

46. ​​An encoder, the encoder comprising a first determining unit and a first predicting unit, wherein: The first determining unit is configured to determine one or more pre-coded blocks; and to determine statistical results of at least two candidate intra-frame prediction modes in a first mode set based on the prediction modes used by each of the one or more pre-coded blocks. The first determining unit is further configured to determine at least two intra-frame prediction modes for the current block in the first mode set based on the statistical results. The first prediction unit is configured to predict the current block based on the at least two intra-frame prediction modes and determine the predicted value of the current block.

47. An encoder, the encoder comprising a first memory and a first processor, wherein: The first memory is used to store computer programs that can run on the first processor; The first processor is configured to, when running the computer program, execute the steps of the encoding method as described in any one of claims 23 to 44.

48. A decoder, the decoder comprising a second determining unit and a second predicting unit, wherein: The second determining unit is configured to determine one or more preset blocks that have been decoded; And based on the prediction modes used by each of the one or more preset blocks, determine the statistical results of at least two candidate intra-frame prediction modes in the first mode set; The second determining unit is further configured to determine at least two intra-frame prediction modes for the current block in the first mode set based on the statistical results. The second prediction unit is configured to predict the current block based on the at least two intra-frame prediction modes and determine the predicted value of the current block.

49. A decoder, the decoder comprising a second memory and a second processor, wherein: The second memory is used to store computer programs that can run on the second processor; The second processor is configured to execute the steps of the decoding method as described in any one of claims 1 to 22 when running the computer program.

50. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the decoding method as described in any one of claims 1 to 22, or the steps of the encoding method as described in any one of claims 23 to 44.

51. A computer-readable storage medium having a bitstream stored thereon, wherein, The bitstream is generated by performing the steps of the encoding method as described in any one of claims 23 to 44.

Citation Information

Patent Citations

  • Method and device for video coding and decoding, storage medium and computer equipment

    CN115103183A

  • Intra prediction method, encoder, decoder, and storage medium

    CN117676133A

  • Video coding and decoding method, device, system and storage medium

    CN117981307A

  • Improvement method of precision inspection inside water supply and sewage pipes

    KR1020230010920A

  • Multiple transform selection-based image coding method and device therefor

    WO2020050651A1