Encoding method, decoding method, encoder, decoder, and storage medium

By decomposing the current block during the video encoding and decoding process, determining its size is smaller than the first residual block of the current block, and transforming it, the problem of increasing the computational complexity caused by the large size of the current block is solved, and a more efficient encoding and decoding process is achieved.

WO2025112031A1PCT designated stage expired Publication Date: 2025-06-05GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/135808
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

During the video encoding and decoding process, the larger the size of the current block, the larger the size of the corresponding residual block, resulting in an increase in the computational complexity of the transformation and inverse transformation processes.

Method used

By determining the second residual block corresponding to the current block, its size is equal to the size of the current block, and then the first residual block is determined according to the second residual block, its size is smaller than the size of the current block, and the first residual block is transformed.

Benefits of technology

The calculation complexity of the transformation and inverse transformation processes is reduced and the encoding and decoding efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023135808_05062025_PF_FP_ABST
    Figure CN2023135808_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an encoding method, a decoding method, an encoder, a decoder, and a storage medium. The decoding method comprises: performing inverse transform on a transform coefficient, and determining a first residual block corresponding to the current block, wherein the size of the first residual block is less than the size of the current block; determining a second residual block on the basis of the first residual block, wherein the size of the second residual block is equal to the size of the current block; and on the basis of the second residual block, determining a reconstruction block corresponding to the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Coding and decoding method, codec and storage medium Technical Field

[0001] The present application relates to the field of video coding and decoding technology, and in particular to a coding and decoding method, a codec, and a storage medium. Background Art

[0002] The larger the current block size, the larger the corresponding residual block size. During the encoding and decoding process, a larger residual block size increases the computational complexity of the transform and inverse transform. Therefore, reducing the computational complexity of the transform and inverse transform processes is a problem that needs to be solved.

[0003] Summary of the Invention

[0004] The present application provides a coding and decoding method, a codec, and a storage medium. The following introduces various aspects of the present application.

[0005] In a first aspect, a decoding method is provided, which is applied to a decoder, and the decoding method includes: performing an inverse transformation on the transformation coefficients to determine a first residual block corresponding to the current block, where the size of the first residual block is smaller than the size of the current block; determining a second residual block based on the first residual block, where the size of the second residual block is equal to the size of the current block; and determining a reconstructed block corresponding to the current block based on the second residual block.

[0006] In a second aspect, a coding method is provided, which is applied to an encoder, and the coding method includes: determining a second residual block corresponding to a current block, where the size of the second residual block is equal to the size of the current block; determining a first residual block based on the second residual block, where the size of the first residual block is smaller than the size of the current block; and transforming the first residual block.

[0007] In a third aspect, a decoder is provided, including: a first determination unit, configured to perform an inverse transformation on the transformation coefficients, determine a first residual block corresponding to the current block, and the size of the first residual block is smaller than the size of the current block; a second determination unit, configured to determine a second residual block based on the first residual block, and the size of the second residual block is equal to the size of the current block; and a third determination unit, configured to determine a reconstructed block corresponding to the current block based on the second residual block.

[0008] In a fourth aspect, a decoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method of the first aspect when running the computer program.

[0009] In a fifth aspect, an encoder is provided, including: a first determination unit, configured to determine a second residual block corresponding to a current block, the size of the second residual block being equal to the size of the current block; a second determination unit, configured to determine a first residual block based on the second residual block, the size of the first residual block being smaller than the size of the current block; and a transformation unit, configured to transform the first residual block.

[0010] In a sixth aspect, an encoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method of the second aspect when running the computer program.

[0011] In a seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the method of the first aspect or the second aspect is implemented.

[0012] In an eighth aspect, a computer program product is provided, comprising a computer program, which implements the method of the first aspect or the second aspect when the computer program is executed.

[0013] In a ninth aspect, a non-volatile computer-readable storage medium for storing a bit stream is provided, wherein the bit stream is generated by an encoding method using an encoder, or the bit stream is decoded by a decoding method using a decoder, wherein the decoding method is the method of the first aspect and the encoding method is the method of the second aspect.

[0014] In a tenth aspect, a code stream is provided, comprising a code stream generated according to the method of the first aspect or a code stream generated according to the method of the second aspect.

[0015] In the embodiment of the present application, during the encoding and decoding process, blocks of smaller sizes are transformed or inversely transformed, which helps to reduce the computational complexity of the transformation or inverse transformation process. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] FIG1 is a structural diagram illustrating an example of a video encoder to which an embodiment of the present application may be applied.

[0017] FIG2 is a diagram showing an example structure of a video decoder to which an embodiment of the present application can be applied.

[0018] FIG3 is an example diagram of an intra-frame prediction mode in the related art.

[0019] FIG4A is an image that has not been transformed.

[0020] FIG4B is an image obtained after the image shown in FIG4A is subjected to discrete cosine transform (DCT).

[0021] Figure 5 shows the base image of DCT2.

[0022] FIG6A is a schematic diagram of a coding and decoding process in which only one transformation is performed.

[0023] FIG6B is a schematic diagram of a coding and decoding process based on basic transformation and secondary transformation.

[0024] FIG7 is a diagram showing an example of input and output data of a low frequency non-separable transform (LFNST).

[0025] FIG8 is an example diagram of the transformation kernel group of LFNST.

[0026] FIG9 is an example diagram of a base image of a non-separable primary transform (NSPT).

[0027] FIG10 is a flow chart of a decoding method according to an embodiment of the present application.

[0028] FIG11 is a flow chart of the encoding method provided in an embodiment of the present application.

[0029] FIG12 is a schematic diagram of the up / down sampling process of the residual block provided by the implementation of this application.

[0030] FIG13 is a schematic diagram of the structure of a decoder provided in one embodiment of the present application.

[0031] FIG14 is a schematic diagram of the structure of a decoder provided in another embodiment of the present application.

[0032] FIG15 is a schematic diagram of the structure of an encoder provided in one embodiment of the present application.

[0033] FIG16 is a schematic diagram of the structure of an encoder provided in another embodiment of the present application. DETAILED DESCRIPTION

[0034] The technical solution in this application will be described below with reference to the accompanying drawings.

[0035] FIG1 is a schematic block diagram of a video encoder according to an embodiment of the present application.

[0036] It should be understood that the video encoder 100 can be used to perform lossy compression or lossless compression on an image. The lossless compression can be visually lossless compression or mathematically lossless compression.

[0037] The video encoder 100 can be applied to image data in a luminance and chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents brightness (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) for describing color and saturation. For example, in terms of color format, 4:2:0 means that every 4 pixels have 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that every 4 pixels have 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 represents full pixel display (YYYYCbCrCbCrCbCrCbCr).

[0038] For example, the video encoder 100 reads video data and, for each image in the video data, divides the image into a number of coding tree units (CTUs). In some examples, a CTU may be referred to as a "tree block," "largest coding unit" (LCU) or "coding tree block" (CTB). Each CTU may be associated with a pixel block of equal size within the image. Each pixel may correspond to one luminance (luminance or luma) sample and two chrominance (chroma) samples. Therefore, each CTU may be associated with one luminance sample block and two chrominance sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU may be further divided into a number of coding units (CUs) for encoding. A CU may be a rectangular block or a square block. A CU may correspond to a prediction unit (PU) and a transform unit (TU).

[0039] In some embodiments, as shown in FIG1 , the video encoder 100 may include a prediction module 110, a residual module 120, a transform / quantization module 130, an inverse transform / quantization module 140, a reconstruction module 150, a loop filter module 160, a decoded image buffer 170, and an entropy coding module 180. It should be noted that the video encoder 100 may include more, fewer, or different functional components.

[0040] Optionally, in this application, the current block may be referred to as the current coding unit (CU). The prediction block may also be referred to as a predicted image block or an image prediction block, and the reconstructed image block may also be referred to as a reconstructed block or an image reconstruction block. Due to the need for parallel processing, an image may be divided into slices. Slices in the same image may be processed in parallel, meaning that there is no data dependency between them. The term "frame" is commonly used, and it can generally be understood that a frame is an image. The term "frame" herein may also be replaced by "image" or "slice," etc.

[0041] In some embodiments, the prediction module 110 includes an inter-frame prediction module 111 and an intra-frame prediction module 112. Because there is a strong correlation between adjacent pixels in a video image, intra-frame prediction is used in video coding and decoding technologies to eliminate spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent images in a video, inter-frame prediction is used in video coding and decoding technologies to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency.

[0042] The inter-frame prediction module 111 can be used for inter-frame prediction. Inter-frame prediction can include motion estimation and motion compensation. It can refer to image information from different images. Inter-frame prediction uses motion information to find a reference block from a reference image and generate a prediction block based on the reference block to eliminate temporal redundancy. Inter-frame prediction uses motion information to find a reference block from a reference image and generate a prediction block based on the reference block. Motion information includes the reference image list in which the reference image is located, the reference image index, and the motion vector. The motion vector can be integer pixel or fractional pixel. If the motion vector is fractional pixel, interpolation filtering is required to generate the required fractional pixel block in the reference image. Here, the integer pixel or fractional pixel block in the reference image found based on the motion vector is called a reference block. Some technologies directly use the reference block as the prediction block, while others further process the reference block to generate a prediction block. Reprocessing the reference block to generate a prediction block can also be understood as using the reference block as the prediction block and then processing the prediction block to generate a new prediction block.

[0043] The intra-frame prediction module 112 only refers to information of the same image to predict pixel information within the current code image block to eliminate spatial redundancy.

[0044] Intra-frame prediction has multiple prediction modes. For example, the H-series international digital video coding standard H.264 / AVC has eight angular prediction modes and one non-angular prediction mode. H.265 / HEVC expands this to 33 angular prediction modes and two non-angular prediction modes. High-efficiency video coding (HEVC) uses planar, direct current (DC), and 33 angular modes for a total of 35 intra-frame prediction modes. Versatile video coding (VVC) uses planar, DC, and 65 angular modes for a total of 67 intra-frame prediction modes.

[0045] It should be noted that with the increase of angle modes, intra-frame prediction will be more accurate and more in line with the needs of high-definition and ultra-high-definition digital video development.

[0046] Residual module 120 may generate a residual block for a CU based on the pixel block of the CU and the prediction block of the CU. For example, residual module 120 may generate a residual block for the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the CU.

[0047] The transform / quantization module 130 may quantize the transform coefficients. The transform / quantization module 130 may quantize the transform coefficients associated with the CU based on a quantization parameter (QP) value associated with the CU. The video encoder 100 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.

[0048] The inverse transform / quantization module 140 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficients to reconstruct a residual block from the quantized transform coefficients.

[0049] Reconstruction module 150 can add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by prediction module 110 to generate a reconstructed image block associated with the CU. By reconstructing each sample block of the CU in this manner, video encoder 100 can reconstruct the pixel blocks of the CU.

[0050] The loop filter module 160 is used to process the inverse transformed and inverse quantized pixels to compensate for distortion information and provide a better reference for subsequent pixel encoding. For example, it can perform a deblocking filtering operation to reduce the blocking effect of pixel blocks associated with the CU.

[0051] In some embodiments, the loop filtering module 160 includes a deblocking filtering module and a sample adaptive offset / adaptive loop filtering (SAO / ALF) module, wherein the deblocking filtering module is used to remove blocking effects, and the SAO / ALF module is used to remove ringing effects.

[0052] The decoded image buffer 170 may store the reconstructed pixel blocks. The inter prediction module 111 may use a reference image containing the reconstructed pixel blocks to perform inter prediction on PUs of other images. In addition, the intra prediction module 112 may use the reconstructed pixel blocks in the decoded image buffer 170 to perform intra prediction on other PUs in the same image as the CU.

[0053] The entropy encoding module 180 may receive the quantized transform coefficients from the transform / quantization module 130. The entropy encoding module 180 may perform one or more entropy encoding operations on the quantized transform coefficients to generate entropy-encoded data.

[0054] FIG2 is a schematic block diagram of a video decoder according to an embodiment of the present application.

[0055] 2 , video decoder 200 includes an entropy decoding module 210, a prediction module 220, an inverse quantization / transformation module 230, a reconstruction module 240, a loop filter module 250, and a decoded image buffer 260. It should be noted that video decoder 200 may include more, fewer, or different functional components.

[0056] The video decoder 200 may receive a bitstream. The entropy decoding module 210 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding module 210 may parse the entropy-encoded syntax elements in the bitstream. The prediction module 220, the inverse quantization / transformation module 230, the reconstruction module 240, and the loop filter module 250 may decode the video data based on the syntax elements extracted from the bitstream, thereby generating decoded video data.

[0057] In some embodiments, the prediction module 220 includes an intra-frame prediction module 222 and an inter-frame prediction module 221 .

[0058] The intra prediction module 222 may perform intra prediction to generate a prediction block for a PU. The intra prediction module 222 may use an intra prediction mode to generate a prediction block for the PU based on pixel blocks of spatially neighboring PUs. The intra prediction module 222 may also determine the intra prediction mode for the PU based on one or more syntax elements parsed from the codestream.

[0059] The inter-frame prediction module 221 may construct a first reference picture list (List 0) and a second reference picture list (List 1) based on syntax elements parsed from the codestream. Furthermore, if a PU is encoded using inter-frame prediction, the entropy decoding module 210 may parse the motion information of the PU. The inter-frame prediction module 221 may determine one or more reference blocks for the PU based on the motion information of the PU. The inter-frame prediction module 221 may generate a prediction block for the PU based on the one or more reference blocks of the PU.

[0060] The inverse quantization / transform module 230 may inversely quantize (ie, dequantize) the transform coefficients associated with the TU. The inverse quantization / transform module 230 may use the QP value associated with the CU of the TU to determine the degree of quantization.

[0061] After inverse quantizing the transform coefficients, inverse quantization / transform module 230 may apply one or more inverse transforms to the inverse quantized transform coefficients in order to generate a residual block associated with the TU.

[0062] Reconstruction module 240 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, reconstruction module 240 can add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU to obtain a reconstructed image block.

[0063] The loop filtering module 250 may perform a deblocking filtering operation to reduce blocking artifacts of pixel blocks associated with a CU.

[0064] The video decoder 200 may store the reconstructed image of the CU in the decoded image buffer 260. The video decoder 200 may use the reconstructed image in the decoded image buffer 260 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.

[0065] The basic process of video encoding and decoding is as follows: At the encoder end, an image is divided into blocks. For the current block, the prediction module 110 uses intra-frame prediction or inter-frame prediction to generate a prediction block for the current block. The residual module 120 calculates a residual block based on the predicted block and the original block of the current block. This residual block is the difference between the predicted block and the original block of the current block. This residual block can also be referred to as residual information. This residual block undergoes transformation and quantization by the transform / quantization module 130, removing information that is insensitive to the human eye and eliminating visual redundancy. Optionally, the residual block before transformation and quantization by the transform / quantization module 130 can be referred to as a time-domain residual block, and the time-domain residual block after transformation and quantization by the transform / quantization module 130 can be referred to as a frequency residual block or a frequency-domain residual block. The entropy coding module 180 receives the quantized change coefficients output by the transform and quantization module 130 and performs entropy coding on these quantized change coefficients to output a bitstream. For example, the entropy coding module 180 can eliminate character redundancy based on the target context model and the probability information of the binary bitstream.

[0066] At the decoding end, the entropy decoding module 210 can parse the code stream to obtain the prediction information, quantization coefficient matrix, etc. of the current block. The prediction module 220 uses intra-frame prediction or inter-frame prediction on the current block based on the prediction information to generate a prediction block for the current block. The inverse quantization / transformation module 230 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction module 240 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed blocks constitute a reconstructed image, and the loop filtering module 250 performs loop filtering on the reconstructed image based on the image or block to obtain a decoded image. The encoding end also requires similar operations as the decoding end to obtain a decoded image. The decoded image can also be called a reconstructed image, and the reconstructed image can be used as a reference image for inter-frame prediction of subsequent images.

[0067] It should be noted that the block division information determined by the encoder, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering, etc., are carried in the bitstream when necessary. The decoder parses the bitstream and analyzes the existing information to determine the same block division information, prediction, transform, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoder, thereby ensuring that the decoded image obtained by the encoder and the decoder are identical.

[0068] The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. This application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.

[0069] The above describes in detail the encoding and decoding framework provided by the embodiments of the present application. As mentioned above, during encoding, prediction is usually performed first. The so-called prediction refers to obtaining a predicted block that is identical or similar to the current block based on the spatial or temporal correlation performance of the image. There are many prediction methods, such as intra-frame prediction and inter-frame prediction. Intra-frame prediction can also include multiple prediction modes, such as angle prediction mode. The following introduces an intra-frame prediction mode in related technology, namely matrix-based intra prediction (MIP). MIP can also be called matrix-weighted intra prediction (MIP). As shown in Figure 3, in order to predict a block with a width of W and a height of H, MIP requires H reconstructed pixels in a column to the left of the current block and W reconstructed pixels in a row above the current block as input. MIP can generate a prediction block in the following three steps: reference pixel averaging, matrix multiplication, and interpolation. MIP can be understood as a process of generating a prediction block based on input pixels (reference pixels) through matrix multiplication. MIP provides a variety of matrices, and the differences in MIP prediction methods are reflected in the different matrices used. In other words, in MIP, the same input pixel will produce different results using different matrices. The process of reference pixel averaging and interpolation is a design that compromises performance and complexity. For larger coding blocks, reference pixel averaging can achieve an effect similar to downsampling, allowing the input to fit into a smaller matrix. Interpolation achieves an upsampling effect. This eliminates the need to provide a MIP matrix for every coding block size, but only provides matrices of one or several specific sizes. With the increasing demand for compression performance and the improvement of hardware capabilities, more complex MIPs may appear in the next generation of standards. MIP and planar mode types, but MIP is obviously more complex and more flexible than planar.

[0070] For a coding block, it's possible for the predicted block to be identical to the original block. However, it's difficult to guarantee that the predicted blocks for all coding blocks in a video are identical to the original blocks. In particular, for natural video, or video captured by a camera, there are often differences between the predicted and original blocks. Irregular motion, distortion, occlusion, brightness changes, and other variations in video are difficult to fully predict. Therefore, the hybrid coding framework subtracts the predicted block from the original block to produce a residual block. The residual block is typically much simpler than the original block, so using prediction can significantly improve compression efficiency. After obtaining the residual block, the encoder typically doesn't encode it directly, but instead performs a transform first. This transform converts the residual block from the spatial domain to the frequency domain. After the residual block is transformed to the frequency domain, since most of the energy is concentrated in the low-frequency region, the non-zero coefficients are concentrated in the upper left corner. Quantization can then be used to further compress these non-zero coefficients. Furthermore, since the human eye is less sensitive to high-frequency regions, a larger quantization step size can be used for the quantized coefficients in high-frequency regions.

[0071] The following uses Figures 4A and 4B to illustrate the transformation process in more detail. Figure 4A shows the original image. Figure 4B shows the image after DCT transformation. As can be seen from Figure 4B, after the DCT transformation of the original image, only the upper left corner region has non-zero coefficients. Of course, Figures 4A and 4B use the DCT transformation of the entire image as an example. In video encoding and decoding, images are typically processed in blocks, so the transformation is also performed on a block-by-block basis.

[0072] Generally speaking, transforms are very useful in video compression. However, not all coded blocks need to be transformed. In some cases, transforming a coded block can actually lead to poor compression. Therefore, in some standards (such as VVC), the encoder can choose whether to perform a transform operation on the current block.

[0073] Next, we will introduce the transformation methods commonly used in the encoding and decoding process.

[0074] DCT2 is the most commonly used transform in video compression standards. Its transformed base image is shown in Figure 5. VVC also uses DCT8 and discrete sine transform (DST)7. The basic formulas for these transforms are shown in Table 1.

[0075] Table 1

[0076] Images are generally two-dimensional. Directly transforming an image in two dimensions would incur significant computational and memory overhead, and current hardware is currently unable to support such transformations. Therefore, current standards utilize one-dimensional transformations. This involves splitting the two-dimensional transformation into one-dimensional transformations in the horizontal and vertical directions, each performed in two steps. For example, the horizontal transformation can be performed first, followed by the vertical transformation. Alternatively, the vertical transformation can be performed first, followed by the horizontal transformation.

[0077] In VVC, multiple transformation modes such as DCT2, DCT8 and DST7 are supported. In each transformation mode, multiple transformation cores can be included. The transformation core can be applied separately to the transformation in the horizontal or vertical direction. For a coding block, the encoder can select a suitable transformation core and transmit the index of the transformation core to the bit stream. The decoder can determine the transformation core of the inverse transformation based on the index. Different transformation cores can be selected for horizontal and vertical transformations. For example, DCT8 can be used for transformation in the horizontal direction, and DST7 can be used for transformation in the vertical direction. The technology of selecting a suitable transformation mode from the multiple transformation modes mentioned above is usually called multiple transform selection (MTS).

[0078] In VVC, syntax elements such as mts_idx are used to determine the transform kernel of the base transform. As shown in Table 2, trTypeHor represents the transform kernel of the horizontal transform; trTypeVer represents the transform kernel of the vertical transform. In Table 2, the values ​​of trTypeHor and trTypeVer are as follows: 0 represents a DCT2 transform; 1 represents a DCT7 transform; and 2 represents a DCT8 transform.

[0079] Table 2

[0080] DCT-based transformations work well for horizontal and vertical image textures. However, DCT-based transformations for diagonal image textures are less effective. Since horizontal and vertical image textures are the most common, DCT is very useful for improving compression efficiency. As video codecs continue to demand higher compression efficiency, more efficient processing of diagonal image textures could further improve compression efficiency.

[0081] Therefore, LFNST is introduced in VVC to more effectively process oblique image textures. LFNST is described in detail below with reference to Figures 6A and 6B. Figure 6A shows the normal encoding and decoding process without LFNST. As shown in Figure 6A, during the encoding process, only one primary transform is performed. Correspondingly, during the decoding process, only one inverse transform of the primary transform is performed. Figure 6B shows the encoding and decoding process after the introduction of LFNST. As shown in Figure 6B, during the encoding process, before the quantization operation is performed, a DCT2-based transform is performed, followed by an LFNST-based transform. During the decoding process, after the inverse quantization operation is performed, an LFNST-based inverse transform is performed, followed by an DCT2-based inverse transform. Based on the above description, it can be seen that LFNST is a transform performed again based on the primary transform (e.g., DCT2), so LFNST is often called a secondary transform. The primary transforms mentioned above can include transforms such as DCT2, DCT8, and DST7.

[0082] As described above, on the encoder side, LFNST performs a secondary transform on the low-frequency coefficients after the base transform. The base transform decorrelates the image, concentrating the energy in the upper left corner. LFNST then transforms the low-frequency coefficients after the base transform again to remove correlation. The following, combined with Figure 7, provides a detailed description of the input and output of LFNST during the encoding and decoding process.

[0083] As shown in Figure 7, at the encoding end, if 16 coefficients are input to a 4×4 LFNST, 8 coefficients can be output; if 64 coefficients are input to an 8×8 LFNST, 16 coefficients can be output. Correspondingly, at the decoding end, if 8 coefficients are input to a 4×4 inverse LFNST, 16 coefficients can be output; if 16 coefficients are input to an 8×8 inverse LFNST, 64 coefficients can be output.

[0084] Figure 8 shows some base images for LFNST in VVC. As can be seen from Figure 8, the base images corresponding to transformation kernel groups 1 through 3 have some noticeable diagonal textures. This means that diagonal image textures are better transformed using transformation kernel groups 1 through 3. On the other hand, flat gradient image textures are better transformed using transformation kernel group 0.

[0085] The transform kernel selected by LFNST can be determined based on the intra-frame prediction mode. Specifically, angular prediction tiles the reference pixels to the current block at a specified angle as the prediction value, which means that the predicted block will have obvious directional texture. The residual block after the current block is angularly predicted will also statistically exhibit obvious angular characteristics. Therefore, the transform kernel selected by LFNST can be tied to the intra-frame prediction mode. In other words, after determining the intra-frame prediction mode, LFNST can use a set of transform kernels corresponding to the intra-frame prediction mode.

[0086] The LFNST in VVC uses four groups of transform kernels, each containing two transform kernels. Table 3 shows the correspondence between intra prediction modes and transform kernel groups. It is important to understand that in this table, prediction modes 81 through 83 are cross-component prediction modes used for intra prediction of chroma blocks, but are not used for intra prediction of luma blocks.

[0087] The LFNST transform kernel can be transposed, allowing a single transform kernel group to handle more angles. For example, prediction modes 13 to 23 and prediction modes 45 to 55 can both correspond to transform kernel group 2, but prediction modes 13 to 23 are clearly closer to horizontal prediction modes, while prediction modes 45 to 55 are clearly closer to vertical prediction modes.

[0088] Table 3

[0089] As described above, VVC uses the correlation between intra-frame prediction modes and LFNST kernels to determine which kernel group to use for LFNST, thereby reducing the number of kernels transmitted in the bitstream. Furthermore, the bitstream and certain conditions can be used to determine whether a coding block will use LFNST. If LFNST is used, it can also be used to determine whether the coding block uses the first or second kernel in a kernel group.

[0090] In related art, the transform kernel groups of LFNST have been further expanded to support more transform kernel groups. Specifically, LFNST in related art has 35 transform kernel groups (each of which can include 3 transform kernels). The correspondence between transform kernel groups and intra prediction modes is shown in Table 4. With the introduction of more transform kernel groups, each transform kernel group can process the corresponding texture more efficiently.

[0091] Table 4

[0092] The above describes the basic transform and the secondary transform. However, while the combination of a basic transform (e.g., DCT2) and a secondary transform (e.g., LFNST) offers excellent transform performance, multiple transforms introduce more complex computations. Therefore, the basic transform plus secondary transform can be considered a compromise between performance and complexity. Compared to the basic transform plus secondary transform, NSPT not only handles oblique textures but also eliminates the need for multiple transforms. The following describes NSTP, as provided in the related art.

[0093] In related art, NSPT and LFNST use the same method for matching transform kernel groups: matching transform kernel groups based on intra-frame prediction modes. The specific matching method can be referenced by LFNST, where each transform kernel group includes three transform kernels. An 8×8 base image for NSPT in related art is shown in Figure 9. Figure 9 corresponds to inter-frame angular prediction mode 7. As can be seen from Figure 9, many base images for NSPT have significant diagonal texture, making NSPT very effective in processing diagonal texture in images.

[0094] On the encoder side, NSPT transforms the spatial domain residuals into frequency domain coefficients. The input to the NSPT-based transform is all possible residual values. For example, if a 4×8 block has 32 pixels, then the encoder's NSPT has 32 inputs. For another example, if an 8×16 block has 128 pixels, then the encoder's NSPT has 128 inputs. The encoder's output can be smaller than the input, which reduces the upper limit on the number of coefficients written into the bitstream, thereby reducing the computational complexity of the transform process and the storage space of the transform kernel. Of course, this will also cause distortion, but because of quantization, this can achieve a better cost-effectiveness.

[0095] Accordingly, at the decoding end, NSPT transforms frequency-domain coefficients into spatial-domain residuals. The output of the inverse transform based on NSPT is the residual value. Table 5 shows the transform kernel information of NSPT at the decoding end in related art. Referring to Table 5, a 4×8 block has 32 pixels, so the NSPT at the decoding end has 32 outputs. An 8×16 block has 128 pixels, so the NSPT at the decoding end has 128 outputs. The input to the decoding end can be less than the output.

[0096] Table 5

[0097] As can be seen from Table 5, the number of transform core groups and the number of transform cores in each group corresponding to NSPT are the same as those of LFNST. Of course, this may be changed in future designs, and the embodiments of this application do not specifically limit this. The number of inputs at the decoding end is determined by the specific transform core design, which may also be changed in future designs, and the embodiments of this application do not specifically limit this. For example, for a 4×8 block, as shown in Table 5, the current design is 20 inputs, and it can also be designed to be 16 inputs in the future. The number of outputs is determined by the size of the block. Specifically, the NSPT inverse transform process at the decoding end is a matrix multiplication process.

[0098] The transform kernel is a matrix of (number of inputs × number of outputs). Therefore, the larger the (number of inputs × number of outputs), the larger the storage space required for the transform kernel and the greater the amount of computation. Taking computational complexity into consideration, in related technologies, NSPT is only applied to some small-sized current blocks, while large-sized current blocks still use DCT2+LFNST. The small sizes mentioned here include, for example, 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, 32×4, 8×32, and 32×8. The following describes the application of NSPT in the encoding and decoding process.

[0099] During the encoding process, the encoder first generates a residual block, which can be obtained by subtracting the predicted value from the original value of the current block.

[0100] The residual block is then transformed. If the size of the current block meets the application conditions of NSPT, NSPT can be used for the current block; otherwise, a combination of a basic transform and a secondary transform (such as LFNST) can be used for the current block. The encoder has multiple options for transforming the residual block. For example, only a basic transform (such as DCT2, DST7, and DCT8, etc.) can be used. Alternatively, a combination of NSPT or a basic transform and a secondary transform (such as LFNST) can be used depending on the size of the current block. For example, in related art, block sizes of 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, 32×4, 8×32, and 32×8 can use NSPT, while other block sizes can use DCT2+LFNST. Of course, the current block can also skip the transformation process, that is, no transformation is performed. The encoder can select an appropriate transformation mode for transformation.

[0101] Finally, the transformed coefficients are quantized and entropy coded.

[0102] During the decoding process, the decoder first performs entropy decoding and dequantization to obtain coefficients.

[0103] The coefficients are then inversely transformed to obtain the residual block. The decoder can determine the transform mode based on the indication of the code stream and known information. For example, the selectable transform modes include DCT2, DST7, DCT8, DCT2+LFNST and NSPT. The decoder can distinguish between NSPT or a combination of basic transform plus secondary transform (such as LFNST) according to the block size. For example, in the related art, NSPT can be used for block sizes of 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, 32×4, 8×32 and 32×8, and DCT2+LFNST can be used for other block sizes.

[0104] Finally, the residual block and the predicted block are added together to obtain the reconstructed block.

[0105] According to the above introduction, the larger the size of the current block, the larger the size of the corresponding residual block. During the encoding and decoding process, the larger the size of the residual block, the greater the computational complexity of the transformation and inverse transformation. It is for this reason that many transformation methods with better transformation performance are only applied to current blocks of smaller size. Taking the NSPT mentioned above as an example, in theory, NSPT can provide higher compression efficiency than the basic transform (DCT2) + secondary transform (LFNST). However, precisely because of the high complexity of applying NSPT on large blocks, the relevant technology only uses NSPT on small blocks, and still uses DCT2+LFNST on large blocks.

[0106] In response to the above problem, an embodiment of the present application provides an encoding method, including: determining a second residual block corresponding to a current block, the size of the second residual block being equal to the size of the current block; determining a first residual block based on the second residual block, the size of the first residual block being smaller than the size of the current block; and transforming the first residual block.

[0107] In addition, an embodiment of the present application also provides a decoding method, including: performing an inverse transform on the transform coefficients to determine a first residual block corresponding to the current block, the size of the first residual block being smaller than the size of the current block; determining a second residual block based on the first residual block, the size of the second residual block being equal to the size of the current block; and determining a reconstructed block corresponding to the current block based on the second residual block.

[0108] In the embodiment of the present application, during the encoding and decoding process, blocks of smaller sizes are transformed or inversely transformed, which helps to reduce the computational complexity of the transformation or inverse transformation process.

[0109] The decoding method of the embodiment of the present application is described in detail below with reference to FIG10 .

[0110] Figure 10 is a flowchart of a decoding method provided by an embodiment of the present application. The method of Figure 10 can be applied to a decoder.

[0111] 10 , in step S1010 , the transform coefficients are inversely transformed to determine a first residual block corresponding to the current block. For example, the code stream may be parsed to determine the transform coefficients; then, the transform coefficients may be inversely transformed to determine the first residual block.

[0112] The current block may refer to the current block to be decoded. In some implementations, the current block is a luminance block. In other implementations, the current block may also be a chrominance block.

[0113] The size of the first residual block is smaller than the size of the current block. In some implementations, the size of the first residual block may be associated with the size of the current block. In other words, the size of the first residual block may be determined based on the size of the current block. For example, the size of the first residual block may be proportional to the size of the current block. Assuming that the ratio of the width to the height of the first residual block is a first ratio and the ratio of the width to the height of the current block is a second ratio, the first ratio may be equal to the second ratio.

[0114] As an example, if the size of the current block is 16×16, the size of the first residual block is 8×8; if the size of the current block is 32×32, the size of the first residual block is 8×8; if the size of the current block is 64×64, the size of the first residual block is 8×8; if the size of the current block is 8×32, the size of the first residual block is 4×16; if the size of the current block is 32×8, the size of the first residual block is 16×4; if the size of the current block is 8×64, the size of the first residual block is 4×32; if the size of the current block is 64×8, the size of the first residual block is 32×4; if the size of the current block is 32×64, the size of the first residual block is 8×16; if the size of the current block is 64×32, the size of the first residual block is 16×8.

[0115] Continuing to refer to Figure 10, in step S1020, a second residual block is determined based on the first residual block. The size of the second residual block is equal to the size of the current block. In other words, the size of the second residual block is larger than the size of the first residual block. For example, the second residual block can be determined by upsampling the first residual block. There are many ways to upsample the first residual block. For example, for the first residual block, upsampling is first performed in the horizontal direction, and then upsampling is performed in the vertical direction. For another example, for the first residual block, upsampling is first performed in the vertical direction, and then upsampling is performed in the horizontal direction. Upsampling along a one-dimensional direction can reduce computational complexity. Of course, in some implementations, the first residual block can also be upsampled in two dimensions, which helps to improve sampling accuracy.

[0116] There are many ways to upsample. For example, the first residual block can be upsampled based on linear interpolation. In another example, the first residual block can be upsampled based on interpolation filtering.

[0117] In step S1030, a reconstructed block corresponding to the current block is determined based on the second residual block. For example, the current block may be predicted to determine the prediction block corresponding to the current block. Then, the reconstructed block may be determined based on the second residual block and the prediction block. For example, determining the reconstructed block based on the second residual block and the prediction block may include determining the sum of pixels of the second residual block and the prediction block as the reconstructed block.

[0118] In the decoding method provided in the embodiment of the present application, a smaller residual block (smaller in size relative to the current block) is obtained after the inverse transform operation. That is, the inverse transform operation is an inverse transform operation performed on a smaller block, which helps to reduce the computational complexity of the inverse transform process.

[0119] The embodiment of the present application does not specifically limit the type of inverse transform operation performed in step S1010. In some implementations, the embodiment of the present application can be applied to scenarios where the complexity of the inverse transform operation is significantly positively correlated with the size of the block. For example, as mentioned above, the NSPT scheme is currently only applied to smaller blocks. This is because the larger the block size, the greater the computational complexity of the NSPT. Therefore, the embodiment of the present application can be applied to NSPT-based encoding and decoding scenarios to reduce the computational complexity of the transform / inverse transform operation.

[0120] The embodiments of the present application are applicable to current blocks of any size. In some implementations, if the size of the current block belongs to a specific size or a specific size set, the decoding method shown in Figure 10 is used; if the size of the current block does not belong to a specific size or a specific size set, a residual block with the same size as the current block can be directly determined by inverse transformation. For example, if the current block is a small block, a residual block with the same size as the small block can be obtained by inverse transforming the transform coefficients, and the reconstructed block corresponding to the current block is determined based on the residual block; if the current block is a large block, a first residual block with a smaller size than the large block can be determined by inverse transforming the transform coefficients. Then, a second residual block with the same size as the large block is determined based on the first residual block, and the reconstructed block corresponding to the current block is determined based on the second residual block. Two possible implementations are given below.

[0121] Implementation method 1:

[0122] If the size of the current block belongs to the first size set, the transform coefficients can be inversely transformed based on NSPT to determine the first residual block (smaller in size than the current block). Then, the second residual block can be determined based on the first residual block, and the reconstructed block corresponding to the current block can be determined based on the second residual block. The method for determining the second residual block and the reconstructed block is described in Figure 10 and will not be repeated here. The first size set mentioned here can, for example, include at least one of the following sizes: 32×32, 64×64, 8×32, 32×8, 8×64, 64×8, 32×64, 64×32.

[0123] If the size of the current block does not belong to the first size set, the third residual block (the size is equal to the current block) can be determined directly by inverse transformation based on NSPT. Then, the reconstructed block corresponding to the current block can be determined based on the third residual block. For example, the reconstructed block can be determined based on the third residual block and the prediction block. For example, if the size of the current block is one of the following sizes, the third residual block can be determined directly by inverse transformation based on NSPT: 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, 32×4.

[0124] In the above, whether the current block belongs to the first size set can be determined based on a preset threshold. For example, if the size of the current block is greater than or equal to the first threshold, the size of the current block belongs to the first size set.

[0125] According to the above description, in the first implementation, regardless of whether the current block size belongs to the first size set, an inverse transform is performed based on NSPT. This can fully utilize the high compression efficiency of NSPT and improve the decoding performance as a whole.

[0126] Implementation method 2:

[0127] If the size of the current block belongs to the first size set, the transform coefficients can be inversely transformed based on NSPT to determine the first residual block (smaller in size than the current block). Then, the second residual block can be determined based on the first residual block, and the reconstructed block corresponding to the current block can be determined based on the second residual block. The method for determining the second residual block and the reconstructed block is described in Figure 10 and will not be repeated here. The first size set mentioned here may include blocks of larger sizes. For example, the first size set may include at least one of the following sizes: 32×32, 64×64, 8×32, 32×8, 8×64, 64×8, 32×64, 64×32.

[0128] If the size of the current block belongs to the second size set, the third residual block (with a size equal to the current block) can be determined directly by inverse transformation based on NSPT. Then, the reconstructed block corresponding to the current block can be determined based on the third residual block. For example, the reconstructed block can be determined based on the third residual block and the prediction block. The second size set mentioned here may include blocks of relatively small sizes. For example, the second size set may include at least one of the following sizes: 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, 32×4.

[0129] If the size of the current block belongs to the third size set, an inverse transform can be performed based on the base transform plus the secondary transform or the base transform to determine the fourth residual block (the size is equal to the current block). Then, the reconstructed block corresponding to the current block can be determined based on the fourth residual block. For example, the reconstructed block can be determined based on the fourth residual block and the prediction block. The third size set mentioned here may include, for example, other sizes in addition to the sizes in the first size set and the second size set. For example, the third size set may include 128×64.

[0130] In the above, whether the current block belongs to the first size set can be determined based on a preset threshold. For example, if the size of the current block is greater than or equal to the first threshold, the size of the current block belongs to the first size set.

[0131] Of course, the current block may also be judged based on a preset threshold whether it belongs to the second size set or the third size set. For example, if the size of the current block is less than or equal to the second threshold, the size of the current block belongs to the second size set.

[0132] As can be seen from the above description, the second implementation method provides more inverse transformation methods, which can be flexibly selected according to actual conditions.

[0133] As mentioned above, the second residual block can be determined by upsampling the first residual block. The upsampling method is described below with a more specific example.

[0134] Taking upsampling based on linear interpolation as an example, if the first residual block is upsampled in one dimension, assuming a 1:2 upsampling ratio, the value of the newly interpolated point can be the average of the values ​​of its two adjacent points. Assuming a 1:4 upsampling ratio, the value of the newly interpolated point can be the weighted average of the values ​​of its two adjacent original points, with the weight being inversely proportional to the distance, i.e., 3:1, 2:2, and 1:3, respectively.

[0135] Taking horizontal upsampling of the first residual block as an example, let the width before upsampling be widthDwn and the width after upsampling be widthUp, with parameter xUp = widthUp / widthDwn. The residual array before upsampling is resDwn[widthDwn], and the residual array after upsampling is resUp[widthUp]. The relationship between the upsampling parameters can be expressed by the following formulas (1), (2), and (3).

[0136] Where, for x from 0 to widthDwn-1: resUp[(x+1)×xUp-1]=resDwn[x] (1)

[0137] Assume resUp[-1]=resDwn[0], where resUp[-1] is not counted in the residual block and is only used to assist calculation, similar to padding.

[0138] For every m from 0 to widthDwn-1, let xHor = m*xUp-1; for every dX from 1 to xUp-1. sum = (xUp-dX) × resUp[xHor] + dX × resUp[xHor+xUp] (2) resUp[xHor+dX] = (sum+xUp / 2) / xUp (3)

[0139] The above sampling is based on interpolation filtering as an example. For example, the 8-tap interpolation filter used in VVC for sub-pixel motion vector motion compensation can be reused. Of course, to reduce complexity, a smaller number of taps can be used, such as a 4-tap interpolation filter.

[0140] In some implementations, the method of FIG10 may further include: parsing the bitstream to determine quantization coefficients; performing inverse quantization on the quantization coefficients to determine transform coefficients. After obtaining the transform coefficients, the transform coefficients may be inversely transformed to determine a first residual block of the current block.

[0141] The decoding method provided by the embodiment of the present application is described in detail above in conjunction with Figure 10. The encoding method provided by the embodiment of the present application is described in detail below in conjunction with Figure 11.

[0142] Figure 11 is a flow chart of an encoding method provided by an embodiment of the present application. The method of Figure 11 can be applied to an encoder.

[0143] 11 , in step S1110 , a second residual block corresponding to the current block is determined. The size of the second residual block is equal to the size of the current block.

[0144] The current block may refer to the current block to be encoded. In some implementations, the current block is a luminance block. In other implementations, the current block may also be a chrominance block.

[0145] In some implementations, before executing step S1110, the current block may be predicted to determine a prediction block corresponding to the current block; then, a second residual block may be determined based on the prediction block. Determining the second residual block based on the prediction block may, for example, include determining the difference between the current block and the prediction block as the second residual block.

[0146] In step S1120, the first residual block is determined based on the second residual block. The size of the first residual block is smaller than the size of the current block. In other words, the size of the second residual block is larger than the size of the first residual block. For example, the first residual block can be determined by downsampling the second residual block. There are many ways to downsample the second residual block. For example, for the second residual block, downsampling is first performed in the horizontal direction, and then downsampled in the vertical direction. For another example, for the second residual block, downsampling is first performed in the vertical direction, and then downsampled in the horizontal direction. Downsampling along a one-dimensional direction can reduce computational complexity. Of course, in some implementations, the second residual block can also be downsampled in two dimensions, which helps to improve sampling accuracy.

[0147] There are many ways to downsample. For example, the second residual block can be downsampled based on an average value. In another example, the second residual block can be downsampled based on a random value.

[0148] The size of the first residual block is smaller than the size of the current block. In some implementations, the size of the first residual block may be associated with the size of the current block. Alternatively, the size of the first residual block may be determined based on the size of the current block. Assuming that the ratio of the width to the height of the first residual block is a first ratio and the ratio of the width to the height of the current block is a second ratio, the first ratio may be equal to the second ratio.

[0149] As an example, if the size of the current block is 16×16, the size of the first residual block is 8×8; if the size of the current block is 32×32, the size of the first residual block is 8×8; if the size of the current block is 64×64, the size of the first residual block is 8×8; if the size of the current block is 8×32, the size of the first residual block is 4×16; if the size of the current block is 32×8, the size of the first residual block is 16×4; if the size of the current block is 8×64, the size of the first residual block is 4×32; if the size of the current block is 64×8, the size of the first residual block is 32×4; if the size of the current block is 32×64, the size of the first residual block is 8×16; if the size of the current block is 64×32, the size of the first residual block is 16×8.

[0150] Continuing with FIG11 , in step S1130 , the first residual block is transformed. For example, the first residual block may be transformed to determine transform coefficients of the current block. The transform coefficients may then be quantized to determine quantization coefficients, which are then encoded.

[0151] In the encoding method provided in the embodiment of the present application, the residual block obtained after step S1120 is smaller in size (smaller in size than the current block). Therefore, the transformation operation in step S1130 is a transformation operation performed on the smaller block, which helps to reduce the computational complexity of the transformation process.

[0152] The embodiment of the present application does not specifically limit the type of transformation operation performed in step S1130. In some implementations, the embodiment of the present application can be applied to scenarios where the complexity of the transformation operation is significantly positively correlated with the size of the block. For example, as mentioned above, the NSPT scheme is currently only applied to smaller blocks. This is because the larger the block size, the greater the computational complexity of the NSPT. Therefore, the embodiment of the present application can be applied to NSPT-based encoding and decoding scenarios to reduce the computational complexity of the transformation / inverse transformation operations.

[0153] The embodiments of the present application are applicable to current blocks of any size. In some implementations, if the size of the current block belongs to a specific size or a specific size set, the encoding method shown in Figure 11 is used; if the size of the current block does not belong to a specific size or a specific size set, the second residual block can be directly transformed to determine the transformation coefficient of the current block. For example, if the current block is a small block, a residual block with the same size as the small block can be determined, and the transformation is performed based on the residual block; if the current block is a large block, a second residual block with the same size as the large block can be determined. Then, based on the second residual block, a first residual block smaller than the size of the large block is determined, and the transformation is performed according to the first residual block. Two possible implementations are given below.

[0154] Implementation method 1:

[0155] If the size of the current block belongs to the first size set, the first residual block can be transformed based on the NSPT to determine the transform coefficients of the current block. The transform coefficients can then be quantized to determine the quantization coefficients, and the quantization coefficients can be encoded. The first size set mentioned here can include, for example, at least one of the following sizes: 32×32, 64×64, 8×32, 32×8, 8×64, 64×8, 32×64, and 64×32.

[0156] If the size of the current block does not belong to the first size set, the second residual block can be directly transformed based on the NSPT to determine the transform coefficients of the current block. The transform coefficients can then be quantized to determine the quantization coefficients, and the quantization coefficients can be encoded. For example, if the size of the current block is one of the following sizes, the second residual block can be directly transformed based on the NSPT to determine the transform coefficients of the current block: 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, 32×4.

[0157] In the above, whether the current block belongs to the first size set can be determined based on a preset threshold. For example, if the size of the current block is greater than or equal to the first threshold, the size of the current block belongs to the first size set.

[0158] According to the above description, in the first implementation, regardless of whether the current block size belongs to the first size set, the transformation is performed based on NSPT. This can fully utilize the advantage of NSPT's high compression efficiency and improve the overall coding performance.

[0159] Implementation method 2:

[0160] If the size of the current block belongs to the first size set, the first residual block can be transformed based on the NSPT to determine the transform coefficients of the current block. The transform coefficients can then be quantized to determine the quantization coefficients, and the quantization coefficients can be encoded. The first size set mentioned here can include, for example, at least one of the following sizes: 32×32, 64×64, 8×32, 32×8, 8×64, 64×8, 32×64, 64×32.

[0161] If the size of the current block belongs to the second size set, the second residual block can be directly transformed based on the NSPT to determine the transform coefficients of the current block. The transform coefficients can then be quantized to determine the quantization coefficients; and the quantization coefficients can be encoded. The second size set mentioned here may include blocks of relatively small sizes. For example, the second size set may include at least one of the following sizes: 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, and 32×4.

[0162] If the size of the current block belongs to the third size set, the second residual block can be directly transformed based on a base transform plus a secondary transform or a base transform to determine the transform coefficients of the current block. The transform coefficients can then be quantized to determine quantization coefficients, and the quantization coefficients can be encoded. The third size set mentioned here can, for example, include sizes other than those in the first size set and the second size set. For example, the third size set can include 128×64.

[0163] In the above, whether the current block belongs to the first size set can be determined based on a preset threshold. For example, if the size of the current block is greater than or equal to the first threshold, the size of the current block belongs to the first size set.

[0164] Of course, the current block may also be judged based on a preset threshold whether it belongs to the second size set or the third size set. For example, if the size of the current block is less than or equal to the second threshold, the size of the current block belongs to the second size set.

[0165] As can be seen from the above description, the second implementation provides more transformation methods, so you can flexibly choose according to the actual situation.

[0166] As mentioned above, the first residual block can be determined by downsampling the second residual block. The downsampling method is described below with a more specific example.

[0167] Taking downsampling the second residual block based on the average as an example, if the second residual block is downsampled in one dimension, assuming a 2:1 downsampling ratio, the average of two adjacent points can be used as the value of the downsampled point. Assuming a 4:1 downsampling ratio, the average of four adjacent points can be used as the value of the downsampled point.

[0168] Taking horizontal downsampling as an example, let the width before downsampling be widthOrg and the width after downsampling be widthDwn. The parameter xDwn = widthOrg / widthDown. The residual value array before downsampling is resOrg[widthOrg], and the residual value array after downsampling is resDwn[widthDwn]. The relationship between the downsampling parameters can be expressed by the following formula (4).

[0169] Among them, for x from 0 to widthDwn-1:

[0170] In some implementations, the method of FIG. 11 may further include: quantizing the transform coefficient determined based on the first residual block to determine a quantization coefficient; and encoding the quantization coefficient.

[0171] The above describes the decoding method and encoding method provided by the embodiments of the present application from the perspectives of the decoding end and the encoding end, respectively. The following describes a more specific embodiment of the present application in conjunction with the encoding and decoding process.

[0172] Coding process

[0173] The encoder generates a residual block. This residual block can be obtained by subtracting the predicted value from the original value of the current block. The size of the residual block is equal to the size of the current block.

[0174] The current block size may be the size of the CU. If the chroma component is currently being encoded, the current block size may be the size of the chroma component block. If a CU is divided into multiple TUs, the current block size may also be the size of the TU.

[0175] Transform the residual block. The encoder has multiple options for transforming the residual block. For example, it can use only the basic transforms such as DCT2, DCT7, or DCT8. Another example is using NSPT. Of course, it is also possible to skip the transform, meaning no transformation is performed. The encoder can select the appropriate transform method for the transformation.

[0176] Taking the encoder's selection of NSPT for transform as an example, NSPT can be directly used to transform the residual block for block sizes of 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, and 32×4. For larger current blocks, such as 32×32, 64×64, 8×32, 32×8, 8×64, 64×8, 32×64, and 64×32, the residual block can be downsampled to the appropriate block size before NSPT is performed on the reduced residual block.

[0177] As shown in Figure 12, both the encoding and decoding ends involve a block size conversion. We refer to the block size to which the transform is applied, or the size of the transform block in the embodiment of the present application, as the adapted block size. Before the transform, the encoder downsamples the residual block to the adapted block size, performs the transform using the adapted block size, i.e., the transform block (tb transform block) size is equal to the adapted block size, and performs subsequent quantization and entropy coding.

[0178] In the embodiment of the present application, a corresponding relationship between the current block size and the adapted block size may be preset. The specific corresponding relationship may be shown in Table 6, for example.

[0179] Table 6

[0180] Decoding process

[0181] The decoder obtains the transform coefficients through entropy decoding and inverse quantization.

[0182] The transform coefficients are inversely transformed to obtain the residual block. The decoder can determine the transform method based on the instructions of the bitstream and known information.

[0183] For example, if the encoder selects NSPT for inverse transform, NSPT can be used to directly inverse transform the residual block for block sizes of 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, and 32×4. For larger current blocks, such as 32×32, 64×64, 8×32, 32×8, 8×64, 64×8, 32×64, and 64×32, which meet the NSPT adaptation criteria, the adapted block size can be determined based on Table 6. An NSPT-based inverse transform is performed on the transform coefficients to obtain a residual block of the adapted block size. The adapted residual block is then upsampled to obtain the residual block of the current block.

[0184] The residual block and the predicted block are added together to obtain the reconstructed block.

[0185] In the above introduction, NSPT is also used for the current block of larger size, but instead of directly using a larger transform kernel, a small transform kernel is applied to the current block of larger size. Specifically, at the encoding end, for a residual block of a certain size, it is downsampled to a residual block of a smaller size, and then transformed using NSPT. Correspondingly, at the decoding end, the coefficients are inversely transformed using NSPT to obtain a residual block, and the residual block is upsampled to obtain a residual block of a larger size. The key point is to adapt NSPT using upsampling and downsampling in the spatial domain. In this way, the transform kernel originally used for the current block of small size can also be used on the current block of large size. Of course, upsampling and downsampling are lossy and will lose a certain amount of high-frequency information, but when performing lossy encoding, NSPT and quantization themselves will also lose a certain amount of high-frequency information, so the distortion generated at this step is acceptable.

[0186] It is understandable that the "inverse transformation" of the transform coefficients at the decoding end may also be referred to as "transformation" in the standard text. The "transformation" and "inverse transformation" in the embodiments of the present application correspond to two opposite processes. For example, if the "transformation" converts the numerical values ​​in the spatial domain to the coefficients in the frequency domain, then the "inverse transformation" converts the coefficients in the frequency domain to the numerical values ​​in the spatial domain. If the standard only stipulates decoding, then the "transformation" in the standard text is the decoding part, which refers to the "inverse transformation" in this article. The "inverse transformation" of the transform coefficients at the decoding end may also be referred to as "transformation" in the standard text.

[0187] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 12. The device embodiment of the present application is described in detail below in conjunction with Figures 13 to 16. It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.

[0188] FIG13 is a schematic diagram of the structure of a decoder provided by an embodiment of the present application. As shown in FIG13 , the decoder 1300 includes: a first determining unit 1310 , a second determining unit 1320 , and a third determining unit 1330 .

[0189] The first determining unit 1310 is configured to perform an inverse transform on the transform coefficients to determine a first residual block corresponding to the current block, where the size of the first residual block is smaller than the size of the current block.

[0190] The second determining unit 1320 is configured to determine a second residual block according to the first residual block, where the size of the second residual block is equal to the size of the current block.

[0191] The third determining unit 1330 determines a reconstructed block corresponding to the current block according to the second residual block.

[0192] In some implementations, the first determining unit 1310 is configured to: perform an inverse transform on the transform coefficients based on a first transform method, where the first transform method is a non-separable basic transform.

[0193] In some implementations, the decoder 1300 further includes: a fourth determination unit configured to determine whether the size of the current block belongs to a first size set; and the first determination unit 1310 is configured to: determine the first residual block if the size of the current block belongs to the first size set.

[0194] In some implementations, the decoder 1300 further includes: a fifth determination unit configured to, if the size of the current block does not belong to the first size set, perform an inverse transform on the transform coefficients to determine a third residual block, where the size of the third residual block is equal to the size of the current block.

[0195] In some implementations, the fifth determining unit is configured to: perform an inverse transformation on the transformation coefficients based on a first transformation method, where the first transformation method is a non-separable basic transformation.

[0196] In some implementations, the decoder 1300 further includes: an inverse transform unit configured to: if the size of the current block belongs to a second size set, perform an inverse transform on the transform coefficients based on a first transform mode; and / or if the size of the current block belongs to a third size set, perform an inverse transform on the transform coefficients based on a second transform mode; wherein the first transform mode is different from the second transform mode.

[0197] In some implementations, the first transform mode is an inseparable basic transform, and the second transform mode is a combination of a basic transform and a secondary transform.

[0198] In some implementations, the first size set includes at least one of the following sizes: 32×32, 64×64, 8×32, 32×8, 8×64, 64×8, 32×64, 64×32.

[0199] In some implementations, the second size set includes at least one of the following sizes: 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, 32×4.

[0200] In some implementations, a size of the first residual block is determined based on a size of the current block.

[0201] In some implementations, a ratio of the width to the height of the first residual block is a first ratio, a ratio of the width to the height of the current block is a second ratio, and the first ratio is equal to the second ratio.

[0202] In some implementations, if the size of the current block is 16×16, the size of the first residual block is 8×8; or, if the size of the current block is 32×32, the size of the first residual block is 8×8; or, if the size of the current block is 64×64, the size of the first residual block is 8×8; or, if the size of the current block is 8×32, the size of the first residual block is 4×16; or, if the size of the current block is 32×8, the size of the first residual block is 16×4; or, if the size of the current block is 8×64, the size of the first residual block is 4×32; or, if the size of the current block is 64×8, the size of the first residual block is 32×4; or, if the size of the current block is 32×64, the size of the first residual block is 8×16; or, if the size of the current block is 64×32, the size of the first residual block is 16×8.

[0203] In some implementations, the second determining unit 1320 is configured to: upsample the first residual block to determine the second residual block.

[0204] In some implementations, the second determination unit 1320 is configured to: for the first residual block, first upsample along the horizontal direction, and then upsample along the vertical direction; or, for the first residual block, first upsample along the vertical direction, and then upsample along the horizontal direction; or, for the first residual block, perform two-dimensional upsampling.

[0205] In some implementations, the decoder 1300 further includes: a sixth determination unit configured to: parse the code stream to determine the quantization coefficient; and a seventh determination unit configured to: perform inverse quantization on the quantization coefficient to determine the transform coefficient.

[0206] In some implementations, the third determining unit 1330 is configured to: predict the current block to determine a prediction block corresponding to the current block; and determine the reconstructed block according to the second residual block and the prediction block.

[0207] In some implementations, the first determining unit 1310 is configured to: determine the first residual block based on an inseparable basic transform, and a transform kernel group of the inseparable basic transform is determined based on an intra prediction mode of the prediction block.

[0208] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.

[0209] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0210] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the decoder 1200. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the decoding method in the first embodiment.

[0211] Based on the composition of the above-mentioned decoder 1300 and the computer-readable storage medium, refer to Figure 14, which shows a specific hardware structure diagram of the decoder 1400 provided in an embodiment of the present application. As shown in Figure 14, the decoder 1400 may include: a communication interface 1410, a memory 1420 and a processor 1430; each component is coupled together through a bus system 1440. It can be understood that the bus system 1440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 1440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 1440 in Figure 14. Among them,

[0212] The communication interface 1410 is used to receive and send signals when sending and receiving information with other external network elements.

[0213] The memory 1420 is used to store computer programs.

[0214] Processor 1430 is configured to, when running the computer program, perform the following steps: performing an inverse transform on the transform coefficients to determine a first residual block corresponding to a current block, where a size of the first residual block is smaller than a size of the current block; determining a second residual block based on the first residual block, where a size of the second residual block is equal to a size of the current block; and determining a reconstructed block corresponding to the current block based on the second residual block.

[0215] It is understood that the memory 1420 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1420 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0216] The processor 1430 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1430. The above-mentioned processor 1430 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1420 , and the processor 1430 reads the information in the memory 1420 and completes the steps of the above method in combination with its hardware.

[0217] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0218] Optionally, as another embodiment, the processor 1430 is further configured to execute the decoding method described in the above embodiment when running the computer program.

[0219] FIG14 is a schematic diagram of the structure of an encoder provided by an embodiment of the present application. As shown in FIG15 , the encoder 1500 includes: a first determination unit 1510 , a second determination unit 1520 , and a transformation unit 1530 .

[0220] The first determining unit 1510 is configured to determine a second residual block corresponding to the current block, where the size of the second residual block is equal to the size of the current block.

[0221] The second determining unit 1520 is configured to determine a first residual block according to the second residual block, where the size of the first residual block is smaller than the size of the current block.

[0222] The transform unit 1530 is configured to transform the first residual block.

[0223] In some implementations, the transform unit 1530 is configured to transform the first residual block based on a first transform mode, where the first transform mode is a non-separable basic transform.

[0224] In some implementations, the encoder 1500 further includes: a third determination unit configured to determine whether the size of the current block belongs to a first size set; the second determination unit 1520 is configured to determine the first residual block based on the second residual block if the size of the current block belongs to the first size set.

[0225] In some implementations, the encoder 1500 further includes: a first transform unit configured to: if the size of the current block does not belong to the first size set, transform the second residual block.

[0226] In some implementations, the first transform unit is configured to: transform the second residual block based on a first transform mode, where the first transform mode is a non-separable basic transform.

[0227] In some implementations, the encoder 1500 further includes: a second transformation unit configured to: transform the second residual block based on a first transformation mode if the size of the current block belongs to a second size set; and / or transform the second residual block based on a second transformation mode if the size of the current block belongs to a third size set; wherein the first transformation mode is different from the second transformation mode.

[0228] In some implementations, the first transform mode is an inseparable basic transform, and the second transform mode is a combination of a basic transform and a secondary transform.

[0229] In some implementations, the first size set includes at least one of the following sizes: 32×32, 64×64, 8×32, 32×8, 8×64, 64×8, 32×64, 64×32.

[0230] In some implementations, the second size set includes at least one of the following sizes: 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, 16×8, 4×32, 32×4.

[0231] In some implementations, a size of the first residual block is determined based on a size of the current block.

[0232] In some implementations, a ratio of the width to the height of the first residual block is a first ratio, a ratio of the width to the height of the current block is a second ratio, and the first ratio is equal to the second ratio.

[0233] In some implementations, if the size of the current block is 16×16, the size of the first residual block is 8×8; or, if the size of the current block is 32×32, the size of the first residual block is 8×8; or, if the size of the current block is 64×64, the size of the first residual block is 8×8; or, if the size of the current block is 8×32, the size of the first residual block is 4×16; or, if the size of the current block is 32×8, the size of the first residual block is 16×4; or, if the size of the current block is 8×64, the size of the first residual block is 4×32; or, if the size of the current block is 64×8, the size of the first residual block is 32×4; or, if the size of the current block is 32×64, the size of the first residual block is 8×16; or, if the size of the current block is 64×32, the size of the first residual block is 16×8.

[0234] In some implementations, the second determining unit 1520 is configured to: downsample the second residual block to determine the first residual block.

[0235] In some implementations, the second determination unit 1520 is configured to: for the first residual block, first downsample along the horizontal direction, and then downsample along the vertical direction; or, for the first residual block, first downsample along the vertical direction, and then downsample along the horizontal direction; or, for the first residual block, perform two-dimensional downsampling.

[0236] In some implementations, the encoder 1500 further includes: a fourth determination unit configured to quantize the transform coefficient determined based on the first residual block to determine a quantization coefficient; and an encoding unit configured to encode the quantization coefficient.

[0237] In some implementations, the first determining unit 1510 is configured to: predict the current block to determine a prediction block corresponding to the current block; and determine the second residual block based on the prediction block.

[0238] In some implementations, the transform unit 1530 is configured to transform the first residual block based on an inseparable basic transform, where a transform kernel group of the inseparable basic transform is determined based on an intra prediction mode of the prediction block.

[0239] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.

[0240] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0241] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 1500. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the decoding method in the first embodiment.

[0242] Based on the composition of the above-mentioned encoder 1500 and the computer-readable storage medium, refer to Figure 16, which shows a specific hardware structure diagram of the encoder 1600 provided in an embodiment of the present application. As shown in Figure 16, the encoder 1600 may include: a communication interface 1610, a memory 1620 and a processor 1630; each component is coupled together through a bus system 1640. It can be understood that the bus system 1640 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1640 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 1640 in Figure 16. Among them,

[0243] The communication interface 1610 is used to receive and send signals when sending and receiving information with other external network elements.

[0244] The memory 1620 is used to store computer programs.

[0245] Processor 1630 is configured to, when running the computer program, perform the following steps: determining a second residual block corresponding to a current block, where a size of the second residual block is equal to a size of the current block; determining a first residual block based on the second residual block, where a size of the first residual block is smaller than a size of the current block; and transforming the first residual block.

[0246] It is understood that the memory 1620 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1620 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0247] The processor 1630 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1630. The above-mentioned processor 1630 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1620 , and the processor 1630 reads the information in the memory 1620 and completes the steps of the above method in combination with its hardware.

[0248] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0249] Optionally, as another embodiment, the processor 1630 is further configured to execute the encoding method described in the above embodiment when running the computer program.

[0250] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0251] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0252] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0253] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0254] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0255] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A decoding method, applied to a decoder, comprising: Performing an inverse transform on the transform coefficients to determine a first residual block corresponding to the current block, wherein the size of the first residual block is smaller than the size of the current block; Determining a second residual block according to the first residual block, wherein the size of the second residual block is equal to the size of the current block; Determining a reconstructed block corresponding to the current block according to the second residual block.

2. The method according to claim 1, wherein, the performing an inverse transform on the transform coefficients includes: Performing an inverse transform on the transform coefficients based on a first transform method, and the first transform method is a non-separable basic transform.

3. The method according to claim 1 or 2, wherein, the method further includes: Determining whether the size of the current block belongs to a first size set; the determining the first residual block corresponding to the current block includes: If the size of the current block belongs to the first size set, then determining the first residual block.

4. The method according to claim 3, wherein, the method further includes: If the size of the current block does not belong to the first size set, then performing an inverse transform on the transform coefficients to determine a third residual block, and the size of the third residual block is equal to the size of the current block.

5. The method according to claim 4, wherein, the performing an inverse transform on the transform coefficients to determine the third residual block includes: Performing an inverse transform on the transform coefficients based on the first transform method to determine the third residual block, and the first transform method is a non-separable basic transform.

6. The method according to claim 3, wherein, the method further includes: If the size of the current block belongs to a second size set, then performing an inverse transform on the transform coefficients based on the first transform method; and / or If the size of the current block belongs to a third size set, then performing an inverse transform on the transform coefficients based on a second transform method; wherein the first transform method is different from the second transform method.

7. The method according to claim 6, wherein, the first transform method is a non-separable basic transform, and the second transform method is a combination of a basic transform and a secondary transform or a basic transform.

8. The method according to any one of claims 3 to 7, wherein, the determining whether the size of the current block belongs to the first size set includes: If the size of the current block is greater than or equal to a first threshold, then the size of the current block belongs to the first size set.

9. The method according to claim 8, wherein, the first size set includes at least one of the following sizes: 32×32,64×64,8×32,32×8,8×64,64×8,32×64,64×32。 10. The method according to claim 6 or 7, wherein, the if the size of the current block belongs to the second size set includes: If the size of the current block is less than or equal to a second threshold, then the size of the current block belongs to the second size set.

11. The method according to claim 10, wherein, the second size set includes at least one of the following sizes: 4×4,4×8,8×4,8×8,4×16,16×4,8×16,16×8,4×32,32×4。 12. The method according to any one of claims 1 to 11, wherein, the size of the first residual block is determined based on the size of the current block.

13. The method according to claim 12, wherein, The ratio of the width to the height of the first residual block is a first ratio, the ratio of the width to the height of the current block is a second ratio, and the first ratio is equal to the second ratio.

14. The method according to any one of claims 1 to 13, wherein: if the size of the current block is 16×16, the size of the first residual block is 8×8; or, if the size of the current block is 32×32, the size of the first residual block is 8×8; or, if the size of the current block is 64×64, the size of the first residual block is 8×8; or, if the size of the current block is 8×32, the size of the first residual block is 4×16; or, if the size of the current block is 32×8, the size of the first residual block is 16×4; or, if the size of the current block is 8×64, the size of the first residual block is 4×32; or, if the size of the current block is 64×8, the size of the first residual block is 32×4; or, if the size of the current block is 32×64, the size of the first residual block is 8×16; or, if the size of the current block is 64×32, the size of the first residual block is 16×8.

15. The method according to any one of claims 1 to 14, wherein, determining a second residual block according to the first residual block includes: performing upsampling on the first residual block to determine the second residual block.

16. The method according to claim 15, wherein, performing upsampling on the first residual block includes: for the first residual block, first performing upsampling in the horizontal direction and then performing upsampling in the vertical direction; or, for the first residual block, first performing upsampling in the vertical direction and then performing upsampling in the horizontal direction; or, for the first residual block, performing two-dimensional upsampling.

17. The method according to any one of claims 1 to 16, wherein, the method further includes: parsing the bitstream to determine quantization coefficients; performing inverse quantization on the quantization coefficients to determine transform coefficients.

18. The method according to any one of claims 1 to 17, wherein, determining a reconstructed block corresponding to the current block according to the second residual block includes: performing prediction on the current block to determine a prediction block corresponding to the current block; determining the reconstructed block according to the second residual block and the prediction block.

19. The method according to any one of claims 1 to 18, wherein, the first residual block is determined based on an inseparable basic transform, and the transform kernel group of the inseparable basic transform is determined based on the intra prediction mode of the prediction block.

20. An encoding method, applied to an encoder, including: determining a second residual block corresponding to a current block, the size of the second residual block being equal to the size of the current block; determining a first residual block according to the second residual block, the size of the first residual block being smaller than the size of the current block; performing a transform on the first residual block.

21. The method according to claim 20, wherein, performing a transform on the first residual block includes: Transform the first residual block based on a first transformation method, where the first transformation method is a non-separable basic transformation.

22. The method according to claim 20 or 21, wherein, the method further includes: determine whether the size of the current block belongs to a first size set; the determining the first residual block according to the second residual block includes: if the size of the current block belongs to the first size set, determine the first residual block according to the second residual block.

23. The method according to claim 22, wherein, the method further includes: if the size of the current block does not belong to the first size set, transform the second residual block.

24. The method according to claim 23, wherein, the transforming the second residual block includes: transform the second residual block based on a first transformation method, where the first transformation method is a non-separable basic transformation.

25. The method according to claim 22, wherein, the method further includes: if the size of the current block belongs to a second size set, transform the second residual block based on a first transformation method; and / or if the size of the current block belongs to a third size set, transform the second residual block based on a second transformation method; wherein the first transformation method is different from the second transformation method.

26. The method according to claim 25, wherein, the first transformation method is a non-separable basic transformation, and the second transformation method is a combination of a basic transformation and a quadratic transformation or a basic transformation.

27. The method according to any one of claims 22 to 26, wherein, the determining whether the size of the current block belongs to the first size set includes: if the size of the current block is greater than or equal to a first threshold, the size of the current block belongs to the first size set.

28. The method according to claim 27, wherein, the first size set includes at least one of the following sizes: 32×32,64×64,8×32,32×8,8×64,64×8,32×64,64×32。 29. The method according to claim 25 or 26, wherein, the if the size of the current block belongs to the second size set includes: if the size of the current block is less than or equal to a second threshold, the size of the current block belongs to the second size set.

30. The method according to claim 29, wherein, the second size set includes at least one of the following sizes: 4×4,4×8,8×4,8×8,4×16,16×4,8×16,16×8,4×32,32×4。 31. The method according to any one of claims 20 to 30, wherein, the size of the first residual block is determined based on the size of the current block.

32. The method according to claim 31, wherein, the ratio of the width to the height of the first residual block is a first ratio, and the ratio of the width to the height of the current block is a second ratio, and the first ratio is equal to the second ratio.

33. The method according to any one of claims 20 to 32, wherein: if the size of the current block is 16×16, the size of the first residual block is 8×8; or, if the size of the current block is 32×32, the size of the first residual block is 8×8; or, If the size of the current block is 64×64, the size of the first residual block is 8×8; or, If the size of the current block is 8×32, the size of the first residual block is 4×16; or, If the size of the current block is 32×8, the size of the first residual block is 16×4; or, If the size of the current block is 8×64, the size of the first residual block is 4×32; or, If the size of the current block is 64×8, the size of the first residual block is 32×4; or, If the size of the current block is 32×64, the size of the first residual block is 8×16; or, If the size of the current block is 64×32, the size of the first residual block is 16×8.

34. The method according to any one of claims 20 to 33, wherein, determining the first residual block according to the second residual block includes: downsampling the second residual block to determine the first residual block.

35. The method according to claim 34, wherein, downsampling the second residual block includes: for the first residual block, first downsampling along the horizontal direction and then along the vertical direction; or, for the first residual block, first downsampling along the vertical direction and then along the horizontal direction; or, for the first residual block, performing two-dimensional downsampling.

36. The method according to any one of claims 20 to 35, wherein, the method further includes: quantizing the transform coefficients determined based on the first residual block to determine quantization coefficients; encoding the quantization coefficients.

37. The method according to any one of claims 20 to 36, wherein, determining the second residual block corresponding to the current block includes: predicting the current block to determine the predicted block corresponding to the current block; determining the second residual block according to the predicted block.

38. The method according to any one of claims 20 to 37, wherein, transforming the first residual block includes: transforming the first residual block based on a non-separable base transform, and the transform kernel group of the non-separable base transform is determined based on the intra prediction mode of the predicted block.

39. A decoder, comprising: a first determination unit configured to perform an inverse transform on transform coefficients to determine a first residual block corresponding to a current block, the size of the first residual block being smaller than the size of the current block; a second determination unit configured to determine a second residual block according to the first residual block, the size of the second residual block being equal to the size of the current block; a third determination unit configured to determine a reconstructed block corresponding to the current block according to the second residual block.

40. A decoder, comprising: a memory for storing a computer program; a processor for executing the method according to any one of claims 1 to 19 when running the computer program.

41. An encoder, comprising: a first determination unit configured to determine a second residual block corresponding to a current block, the size of the second residual block being equal to the size of the current block; A second determination unit, configured to determine a first residual block according to the second residual block, wherein a size of the first residual block is smaller than a size of the current block; A transformation unit, configured to perform a transformation on the first residual block.

42. An encoder, comprising: a memory, configured to store a computer program; a processor, configured to execute the method according to any one of claims 20 to 38 when running the computer program.

43. A non-volatile computer-readable storage medium storing a bitstream, the bitstream being generated by using an encoding method of an encoder or being decoded by using a decoding method of a decoder, wherein the decoding method is the method according to any one of claims 1 to 19, and the encoding method is the method according to any one of claims 20 to 38.

44. A bitstream, the bitstream comprising a bitstream generated by the method according to any one of claims 1 to 19 or a bitstream generated by the method according to any one of claims 20 to 38.

Citation Information

Patent Citations

  • Video image processing method, encoder and computer readable storage medium

    CN112055210A

  • Sample value clipping on mip reduced prediction

    CN113966617A

  • Transformation method, encoder, decoder, and storage medium

    CN114830664A

  • Video encoding / decoding apparatus and mehod using direction of prediction

    KR1020100009718A

  • Apparatus and method for video coding / decoding using adaptive intra prediction

    KR1020140079882A