Decoding method and apparatus, encoding method and apparatus, and corresponding encoder and decoder

By using the cross-component linear model (CCLM) mode to derive luminance values ​​from reference block reconstructed data and obtain lookup table indexes, the complexity of intra-frame prediction is reduced, and the encoding and decoding efficiency is improved, thus solving the problem of high intra-frame prediction complexity in existing technologies.

WO2026157419A1PCT designated stage Publication Date: 2026-07-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-11-03
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing encoding and decoding methods have high intra-frame prediction complexity, which affects encoding and decoding efficiency.

Method used

Intra-frame prediction is performed using the Cross-Component Linear Model (CCLM) mode. The brightness value is derived by obtaining the reconstructed data of the reference block, and the CCLM parameters are obtained by using a lookup table index, which reduces the complexity of intra-frame prediction and improves the prediction accuracy.

Benefits of technology

While reducing the complexity of intra-frame prediction, it improves the efficiency of intra-frame prediction and encoding/decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025132104_30072026_PF_FP_ABST
    Figure CN2025132104_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the technical field of encoding and decoding. Provided are a decoding method and apparatus, an encoding method and apparatus, and a corresponding encoder and decoder. The decoding method comprises: acquiring reconstructed data of a reference block corresponding to a block to be decoded in an image to be decoded; then, on the basis of the reconstructed data of the reference block, obtaining parameters of a cross-component linear model (CCLM); next, on the basis of the parameters of the CCLM, performing intra prediction on the block to be decoded, so as to obtain predicted data of the block to be decoded, wherein the parameters of the CCLM are determined by using M luma values, which are derived on the basis of the reconstructed data of the reference block, and M is an integer greater than 2; and finally, on the basis of the predicted data, decoding a bitstream of the image to be decoded, so as to obtain reconstructed data of the block to be decoded. The complexity of the method during intra prediction is relatively low, and thus the intra prediction efficiency can be improved, thereby improving the decoding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Decoding methods, encoding methods and devices, and corresponding encoders and decoders

[0001] This application claims priority to Chinese patent application filed on January 24, 2025, with application number 202510121319.8 and entitled "Decoding method, encoding method and apparatus, corresponding encoder and decoder", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of encoding and decoding technology, and in particular to a decoding method, encoding method and apparatus, and corresponding encoder and decoder. Background Technology

[0003] Digital video capabilities can be incorporated into a wide variety of devices, including digital television, digital live broadcasting systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones (so-called "smartphones"), video conferencing devices, video streaming devices, and the like. Digital video devices implement video compression techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 High-Level Video Coding (AVC), the H.265 / HEVC video coding standard, and extensions to such standards. By implementing such video compression techniques, video devices can transmit, receive, encode, decode, and / or store digital video information more efficiently.

[0004] Video compression techniques perform spatial (intra-image) prediction to reduce or remove redundancy inherent in video sequences. For block-based video coding, video stripes (i.e., video frames or portions of video frames) can be divided into several image blocks, which may also be referred to as tree blocks, coding units (CUs), and / or coding nodes.

[0005] Current encoding and decoding methods have high complexity in intra-frame prediction and need improvement. Summary of the Invention

[0006] To address the aforementioned technical problems, this application provides a decoding method, an encoding method and apparatus, and corresponding encoders and decoders. This method has lower complexity when performing intra-frame prediction, which can improve intra-frame prediction efficiency and thus improve encoding and decoding efficiency.

[0007] In a first aspect, embodiments of this application provide a decoding method, which includes: acquiring reconstructed data of a reference block corresponding to the block to be decoded in an image to be decoded; then, obtaining parameters of a cross-component linear model (CCLM) based on the reconstructed data of the reference block; next, performing intra-frame prediction on the block to be decoded based on the parameters of the CCLM to obtain prediction data of the block to be decoded, wherein the parameters of the CCLM are determined by M luminance values ​​derived from the reconstructed data of the reference block, where M is an integer greater than 2; and finally, decoding the bitstream of the image to be decoded based on the prediction data to obtain reconstructed data of the block to be decoded.

[0008] In this embodiment of the application, when performing intra-frame prediction on a block, if the intra-frame prediction mode is CCLM mode, the method of this embodiment can derive the values ​​of three or more luminance components (hereinafter referred to as luminance values) based on the reconstruction data of the reference block corresponding to the current block in the CCLM mode in order to obtain the parameters of the CCLM model. Then, the CCLM parameters are obtained based on the obtained luminance component values. This method has low complexity when performing intra-frame prediction, which can improve the efficiency of intra-frame prediction and thus improve the decoding efficiency.

[0009] Based on the first aspect, in some possible implementations, the parameters of CCLM are obtained based on the index of a lookup table, which is based on M luminance values.

[0010] In the embodiments of this application, the index of the lookup table is obtained based on three or more luminance values ​​derived from the derivation, and then the parameters of CCLM are obtained. This provides a method for obtaining CCLM parameters that is different from the traditional method, thereby providing more alternative solutions for obtaining the index for CCLM mode prediction.

[0011] Based on the first aspect, in some possible implementations, the above index is obtained based on the relative positional relationship between the above M brightness values.

[0012] This relative positional relationship can indicate the magnitude relationship between the M brightness values, as well as the relative distance between the M brightness values. The above index can be obtained, and then the parameters of CCLM can be obtained. The index obtained by this method fully refers to the relative distance and magnitude relationship between the M brightness values ​​derived, thus providing a more accurate way to obtain the parameters of CCLM.

[0013] Based on the first aspect, in some possible implementations, the index can be obtained based on the weight of each of the M brightness values ​​and the M brightness values, where the weight of each of the M brightness values ​​is obtained based on the relative positional relationship between the M brightness values.

[0014] In this embodiment, the weight of each brightness value can be obtained based on the relative positional relationship between the M brightness values, and then the index can be obtained by combining the weights of the brightness values. This ensures that the index is obtained by fully considering the relative distance between the M brightness values ​​and the spatial positional relationship, which can make the obtained index more accurate and thus improve the accuracy of obtaining CCLM parameters.

[0015] Based on the first aspect, in some possible implementations, the index is obtained by performing a weighted calculation based on the weight of each of the M brightness values ​​and the M brightness values.

[0016] In this embodiment of the application, an index can be obtained by performing a weighted operation on M brightness values ​​and their weights, thereby providing an alternative way to obtain the index.

[0017] Based on the first aspect, in some possible implementations, the value of M for the aforementioned M brightness values ​​is 3.

[0018] In the embodiments of this application, three luminance values ​​can be derived to obtain an index, and then the parameters of CCLM can be obtained. This can reduce the complexity of intra-frame prediction while improving the calculation accuracy of CCLM parameters.

[0019] Based on the first aspect, in some possible implementations, when M of the M brightness values ​​is 3, the aforementioned M brightness values ​​are represented by Y1, Y2, and Y3, where Y1≤Y2≤Y3, the distance between Y2 and Y1 is greater than or equal to the distance between Y2 and Y3, and the index is obtained based on Y1, Y2, Y3 and the weights w1 of Y1, w2 of Y2, and w3 of Y3, where w1<w2<w3, where w1 is a negative integer, and w2 and w3 are both positive integers.

[0020] Based on the first aspect, in some possible implementations, w1 = -3, w2 = 1, w3 = 2.

[0021] Thus, the index = -3*Y1 + Y2 + 2*Y3, where "*" represents multiplication.

[0022] In this embodiment, three luminance values ​​can be derived to obtain an index, which in turn yields the parameters of CCLM. This reduces the complexity of intra-frame prediction while improving the accuracy of CCLM parameter calculation. Different relative distances between the three luminance values ​​result in different weights, ensuring the accuracy of the obtained CCLM parameters.

[0023] Based on the first aspect, in some possible implementations, when M of the M brightness values ​​is 3, the M brightness values ​​are represented by Y1, Y2, and Y3, where Y1≤Y2≤Y3, the distance between Y2 and Y1 is less than or equal to the distance between Y2 and Y3, and the index is obtained based on Y1, Y2, Y3 and the weights w1 of Y1, w2 of Y2, and w3 of Y3, where w1<w2<w3, and w1 and w2 are both negative integers, and w3 is a positive integer.

[0024] Based on the first aspect, in some possible implementations, w1 = -2, w2 = -1, w3 = 3.

[0025] Thus, the index = -2*Y1-Y2+3*Y3, where "*" represents multiplication.

[0026] In this embodiment, three luminance values ​​can be derived to obtain an index, which in turn yields the parameters of CCLM. This reduces the complexity of intra-frame prediction while improving the accuracy of CCLM parameter calculation. Different relative distances between the three luminance values ​​result in different weights, ensuring the accuracy of the obtained CCLM parameters.

[0027] Based on the first aspect, in some possible implementations, the above M brightness values ​​are derived from the N brightness values ​​in the reconstructed data of the above reference block, where N≥M and N is an integer.

[0028] In some embodiments, N brightness values ​​can be selected based on the reconstructed data of the reference block described above, and then M brightness values ​​can be derived based on the N brightness values, where M is an integer greater than or equal to 3.

[0029] In the embodiments of this application, N brightness values ​​can be selected from the reconstructed data of the reference block, and M brightness values ​​can be derived based on the N brightness values. The M brightness values ​​can be the N brightness values, or they can be different from the N brightness values. This application does not limit this.

[0030] Based on the first aspect, in some possible implementations, the sample points corresponding to the above N brightness values ​​in the reference block are adjacent to the block to be decoded, where adjacent means the left side and / or the top side.

[0031] In this embodiment, when selecting N luminance values ​​from the reconstructed data of the reference block, N sample points adjacent to the block to be decoded can be selected in the reference block, and the luminance values ​​of these N sample points can be used as the N luminance values. Considering that the luminance values ​​of sample points adjacent to the block to be decoded are closer to the luminance values ​​of the block to be decoded, the accuracy of the CCLM prediction data can be improved.

[0032] Based on the first aspect, in some possible implementations, the lookup table has less than 2 P The size of the array is P, where P is a positive integer.

[0033] In some embodiments, the larger the value of P, the higher the decoding complexity and the higher the decoding accuracy; the smaller the value of P, the lower the decoding complexity and the lower the accuracy. The value of P can be flexibly set according to the complexity and accuracy requirements of the application scenario.

[0034] Based on the first aspect, in some possible implementations, P = 4.

[0035] This allows for image or video encoding and decoding with low complexity while ensuring high decoding accuracy.

[0036] Based on the first aspect, in some possible implementations, the above lookup table can be represented as a table TABLE[index].

[0037] Where TABLE[index] = {63,63,61,57,54,51,49,47,45,43,41,39,38,37,35,34}, and 0 ≤ index < 16.

[0038] The data structure of TABLE[index] can be a table or an array.

[0039] Based on the first aspect, in some possible implementations, the prediction data of the block to be decoded is obtained based on the reconstructed value of the luminance value of the block to be decoded and the parameters of CCLM.

[0040] Based on the first aspect, in some possible implementations, the reference block is the adjacent block of the block to be decoded in the image to be decoded.

[0041] In this embodiment of the application, one or more reference blocks adjacent to the current block to be decoded in the image to be decoded are selected as reference blocks for the block to be decoded. Since the brightness between adjacent blocks is relatively close, this is used for CCLM parameter acquisition, which can improve the acquisition speed of CCLM parameters.

[0042] Secondly, embodiments of this application provide an encoding method, which includes: acquiring reconstructed data of a reference block corresponding to the block to be encoded in an image to be encoded; obtaining parameters of a CCLM based on the reconstructed data of the reference block; performing intra-frame prediction on the block to be encoded based on the parameters of the CCLM to obtain prediction data of the block to be encoded, wherein the parameters of the CCLM are determined by using M luminance values ​​derived from the reconstructed data of the reference block, where M is an integer greater than 2; and encoding the block to be encoded based on the prediction data to obtain a bitstream of the image to be encoded.

[0043] In this embodiment of the application, when performing intra-frame prediction on a block, if the intra-frame prediction mode is CCLM mode, the method of this embodiment can derive the values ​​of three or more luminance components (hereinafter referred to as luminance values) based on the reconstruction data of the reference block corresponding to the current block in the CCLM mode in order to obtain the parameters of the CCLM mode. This method has low complexity when performing intra-frame prediction, which can improve the efficiency of intra-frame prediction and thus improve the coding efficiency.

[0044] Based on the second aspect, in some possible implementations, the parameters of CCLM are obtained based on the index of a lookup table, which is based on M luminance values.

[0045] Based on the second aspect, in some possible implementations, the above index is obtained based on the relative positional relationship between the M brightness values.

[0046] Based on the second aspect, in some possible implementations, the index is obtained based on the weight of each of the M brightness values ​​and the M brightness values, and the weight of each of the M brightness values ​​is obtained based on the relative positional relationship between the M brightness values.

[0047] Based on the second aspect, in some possible implementations, the above index is obtained by weighting the M brightness values ​​and the M brightness values ​​together.

[0048] Based on the second aspect, in some possible implementations, M is 3.

[0049] Based on the second aspect, in some possible implementations, the aforementioned M brightness values ​​are represented by Y1, Y2, and Y3, where Y1 ≤ Y2 ≤ Y3, the distance between Y2 and Y1 is greater than or equal to the distance between Y2 and Y3, and the index is obtained based on Y1, Y2, Y3, and the weights w1 of Y1, w2 of Y2, and w3 of Y3, where w1 < w2 < w3, where w1 is a negative integer, and w2 and w3 are both positive integers. In some possible implementations, w1 = -3, w2 = 1, and w3 = 2.

[0050] Based on the second aspect, in some possible implementations, the M brightness values ​​are represented by Y1, Y2, and Y3, where Y1 ≤ Y2 ≤ Y3, the distance between Y2 and Y1 is less than or equal to the distance between Y2 and Y3, and the index is obtained based on Y1, Y2, Y3, and the weights w1, w2, and w3 of Y1 and Y3, where w1 < w2 < w3, and w1 and w2 are negative integers, while w3 is a positive integer. In some possible implementations, w1 = -2, w2 = -1, and w3 = 3.

[0051] Based on the second aspect, in some possible implementations, the M luminance values ​​are derived from the N luminance values ​​in the reconstructed data of the reference block, where N ≥ M and N is an integer.

[0052] Based on the second aspect, in some possible implementations, the sample points corresponding to the N luminance values ​​in the reference block are adjacent to the block to be encoded.

[0053] Based on the second aspect, in some possible implementations, M is 3.

[0054] Based on the second aspect, in some possible implementations, the lookup table has less than 2 P The size of the array is P, where P is a positive integer.

[0055] Based on the second aspect, in some possible implementations, P = 4.

[0056] Based on the second aspect, in some possible implementations, the prediction data of the block to be encoded is obtained based on the reconstructed value of the luminance value of the block to be encoded and the parameters of CCLM.

[0057] Based on the second aspect, in some possible implementations, the reference block is the adjacent block of the block to be encoded in the image to be encoded.

[0058] The effect of the encoding method is similar to that of the first aspect and any of the implementation methods described above, and will not be repeated here.

[0059] Thirdly, embodiments of this application provide a decoding apparatus, which may include: an acquisition module for acquiring reconstructed data of a reference block corresponding to the block to be decoded in an image to be decoded; a prediction module for obtaining parameters of a CCLM based on the reconstructed data of the reference block; the prediction module is further configured to perform intra-frame prediction on the block to be decoded based on the parameters of the CCLM to obtain prediction data of the block to be decoded, wherein the parameters of the CCLM are determined by using M luminance values ​​derived based on the reconstructed data of the reference block, where M is an integer greater than 2; and a decoding module for decoding the bitstream of the image to be decoded based on the prediction data to obtain reconstructed data of the block to be decoded.

[0060] Fourthly, embodiments of this application provide an encoding apparatus, comprising: an acquisition module for acquiring reconstructed data of a reference block corresponding to the block to be encoded in an image to be encoded; a prediction module for obtaining parameters of a CCLM based on the reconstructed data of the reference block; the prediction module is further configured to perform intra-frame prediction on the block to be encoded based on the parameters of the CCLM to obtain prediction data of the block to be encoded, wherein the parameters of the CCLM are determined by using M luminance values ​​derived from the reconstructed data of the reference block, where M is an integer greater than 2; and an encoding module for encoding the block to be encoded based on the prediction data to obtain a bitstream of the image to be encoded.

[0061] Fifthly, a system for distributing bitstreams is provided, the system comprising: at least one storage medium for storing a bitstream generated according to the method of the first or second aspect and any embodiment thereof; and a streaming media device for acquiring the bitstream from the at least one storage medium and transmitting the bitstream, wherein the streaming media device includes a content server or a content distribution server.

[0062] In a sixth aspect, a transcoding system is provided, comprising: at least one storage medium for storing a bitstream generated according to the method of the second aspect or any embodiment thereof; and a transcoding device for acquiring the bitstream from the at least one storage medium and transcoding the bitstream.

[0063] In one embodiment, the transcoding device can convert the bitstream into MPEG-4 Part 4 (MP4), Matroska Video (MKV), Audio Video Interleave (AVI), Digital Audio Video (DAV), etc., without limitation.

[0064] In a seventh aspect, this application provides an apparatus comprising: one or more processors; a memory for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect or any embodiment of the first aspect.

[0065] Eighthly, this application provides an apparatus comprising: one or more processors; a memory for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the second aspect or any embodiment of the second aspect above.

[0066] Ninthly, this application provides a computer-readable storage medium including a computer program that, when executed on a device, causes the device to perform the method described in the first aspect or any embodiment of the first aspect.

[0067] In a tenth aspect, this application provides a computer-readable storage medium including a computer program that, when executed on a device, causes the device to perform the method described in the second aspect or any of the embodiments of the second aspect.

[0068] In one aspect, this application provides a computer program that, when executed by a device, is used to perform the method described in the first aspect or any embodiment of the first aspect, or to perform the method described in the second aspect or any embodiment of the second aspect.

[0069] In a twelfth aspect, this application provides a computer program product comprising computer program code that, when executed on a device, causes the device to perform the method described in the first aspect or any embodiment of the first aspect, or to perform the method described in the second aspect or any embodiment of the second aspect.

[0070] In a thirteenth aspect, this application provides a bitstream generated according to the method in the second aspect or any embodiment of the second aspect described above.

[0071] In a fourteenth aspect, this application provides a computer-readable storage medium storing a bitstream generated according to the method in the second aspect or any embodiment of the second aspect described above.

[0072] In a fifteenth aspect, an apparatus for storing a bitstream is provided, the apparatus comprising: a transceiver unit and a storage unit, the transceiver unit being configured to receive a bitstream generated according to the method described in accordance with the second aspect or any embodiment thereof, and the storage unit being configured to store the bitstream.

[0073] In a sixteenth aspect, an apparatus for transmitting a bitstream is provided, the apparatus comprising: a storage unit and a transceiver unit, the storage unit being configured to store a bitstream generated according to the method described in accordance with the second aspect or any embodiment thereof, and the transceiver unit being configured to transmit the bitstream.

[0074] In a seventeenth aspect, an encoder is provided. Encoding includes: processing circuitry that implements the steps described in the second aspect or any embodiment of the second aspect.

[0075] In an eighteenth aspect, a decoder is provided. Encoding includes: processing circuitry that implements the steps described in the first aspect or any embodiment of the first aspect. Attached Figure Description

[0076] Figure 1A is a schematic block diagram of a video encoding and decoding system provided in an embodiment of this application;

[0077] Figure 1B is a schematic block diagram of a video decoding system provided in an embodiment of this application;

[0078] Figure 2 is a schematic block diagram of an encoder provided in an embodiment of this application;

[0079] Figure 3 is a schematic block diagram of a decoder provided in an embodiment of this application;

[0080] Figure 4 is a schematic diagram of the structure of a video decoding device provided in an embodiment of this application;

[0081] Figure 5 is a schematic diagram of the structure of a device provided in an embodiment of this application;

[0082] Figure 6 is a schematic block diagram of an encoder based on wavelet transform provided in an embodiment of this application;

[0083] Figure 7 is a schematic block diagram of a wavelet transform-based decoder provided in an embodiment of this application;

[0084] Figure 8 is a schematic block diagram of an encoder provided in an embodiment of this application;

[0085] Figure 9A is a schematic diagram of sub-graph division provided in an embodiment of this application;

[0086] Figure 9B is a schematic diagram of sub-graph partitioning provided in an embodiment of this application;

[0087] Figure 10 is a schematic diagram of wavelet transform provided in an embodiment of this application;

[0088] Figure 11 is a schematic block diagram of a decoder provided in an embodiment of this application;

[0089] Figure 12 is a schematic block diagram of an encoder provided in an embodiment of this application;

[0090] Figure 13A is a schematic diagram of the structure of an image bitstream provided in an embodiment of this application;

[0091] Figure 13B is a schematic diagram of the structure of an image bitstream provided in an embodiment of this application;

[0092] Figure 14A is a schematic diagram of the structure of an image bitstream provided in an embodiment of this application;

[0093] Figure 14B is a schematic diagram of the structure of an image bitstream provided in an embodiment of this application;

[0094] Figure 15 is a schematic diagram of the structure of an image bitstream provided in an embodiment of this application;

[0095] Figures 16A to 16C are schematic diagrams of the structure of an image bitstream provided in an embodiment of this application;

[0096] Figures 17A and 17B are schematic diagrams of the structure of an image bitstream provided in an embodiment of this application;

[0097] Figure 18 is a schematic block diagram of a decoder provided in an embodiment of this application;

[0098] Figure 19 is a schematic block diagram of a decoder provided in an embodiment of this application;

[0099] Figure 20a is a schematic diagram of an edge-cloud system provided in an embodiment of this application;

[0100] Figure 20b is a schematic diagram of block data obtained by dividing an image according to an embodiment of this application;

[0101] Figure 21a is a schematic diagram of an encoding method provided in an embodiment of this application;

[0102] Figure 21b is a schematic diagram of an encoding method provided in an embodiment of this application;

[0103] Figure 22 is a schematic diagram of an encoding method provided in an embodiment of this application;

[0104] Figure 23 is a schematic diagram of an encoding method provided in an embodiment of this application;

[0105] Figure 24a is a schematic diagram of an application scenario provided by an embodiment of this application;

[0106] Figure 24b is a schematic diagram of an application scenario provided by an embodiment of this application;

[0107] Figure 24c is a schematic diagram of an application scenario provided by an embodiment of this application;

[0108] Figure 25a is a schematic diagram of an application scenario provided by an embodiment of this application;

[0109] Figure 25b is a schematic diagram of an application scenario provided by an embodiment of this application;

[0110] Figure 25c is a schematic diagram of an application scenario provided by an embodiment of this application;

[0111] Figure 26 is a schematic diagram of an application scenario provided by an embodiment of this application;

[0112] Figure 27a is a schematic diagram of a decoding method provided in an embodiment of this application;

[0113] Figure 27b is a schematic diagram of a decoding method provided in an embodiment of this application;

[0114] Figure 28a is a schematic diagram of a decoding device provided in an embodiment of this application;

[0115] Figure 28b is a schematic diagram of a decoding device provided in an embodiment of this application. Detailed Implementation

[0116] The embodiments of this application are described below with reference to the accompanying drawings. In the following description, reference is made to the accompanying drawings, which form part of this application and illustrate specific aspects of the embodiments of this application or to which specific aspects of the embodiments of this application may be used. It should be understood that the embodiments of this application may be used in other aspects and may include structural or logical variations not depicted in the drawings. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of this application is defined by the appended claims. For example, it should be understood that the disclosure of the described methods is equally applicable to corresponding devices or systems for performing the methods, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units, such as functional units, to perform the described one or more method steps (e.g., one unit performs one or more steps, or multiple units, each performing one or more of multiple steps), even if such one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a specific apparatus is described based on one or more units such as functional units, the corresponding method may include a step to perform the functionality of one or more units (e.g., a step to perform the functionality of one or more units, or multiple steps, each of which performs the functionality of one or more units among a plurality of units), even if such one or more steps are not explicitly described or illustrated in the accompanying drawings. Furthermore, it should be understood that, unless otherwise expressly stated, features of the various exemplary embodiments and / or aspects described herein can be combined with each other.

[0117] The terminology used in the implementation section of this application is for the purpose of explaining specific embodiments of this application only, and is not intended to limit this application.

[0118] The following is a brief introduction to some concepts that may be involved in the embodiments of this application. These concepts are only used to explain the specific embodiments of this application and are not intended to limit this application.

[0119] YUV is a color space model used to represent colors, widely used in image encoding and decoding, video encoding and decoding, digital image processing, television broadcasting, and other fields. YUV separates the luminance information from the chrominance information of an image.

[0120] The three components of YUV:

[0121] Y (luminance) component: Y represents the luminance information of an image, that is, the brightness or darkness of the image. It is obtained by weighting the red, green, and blue color channels according to certain weights. The Y component plays an important role in the sharpness and detail of an image.

[0122] U (chromaticity) component: U represents the chromaticity information of the image, indicating the offset of the blue channel relative to the luminance Y. It measures the change in the blue component.

[0123] V (chromaticity) component: V represents the chromaticity information of the image, indicating the offset of the red channel relative to the luminance Y.

[0124] The residual is the difference between the reconstructed value (or actual value) of a sample or data element and its predicted value.

[0125] A residual block is an M×N residual matrix composed of the residuals corresponding to the coded blocks.

[0126] Dequantization is the process of scaling the quantized residual to obtain the reconstructed residual value.

[0127] A partition divides a set into subsets. Each element in the set belongs to one and only one subset.

[0128] Partition type: The way the subsets obtained from the partition are organized.

[0129] A decoded picture is an image reconstructed by the decoder based on the bitstream.

[0130] Prediction is the specific implementation of the prediction process.

[0131] The prediction process uses previously decoded samples to obtain the predicted value for the current sample.

[0132] Syntax element: The result of parsing data units in a bitstream.

[0133] A bitstream is a binary data stream that encodes all or part of an image sample.

[0134] Video coding generally refers to the processing of a sequence of images that form a video or video sequence. In the field of video coding, the terms "picture," "frame," or "image" can be used synonymously. Video coding is performed on the source side and typically involves processing (e.g., by compression) the raw video images to reduce the amount of data required to represent them, thus enabling more efficient storage and / or transmission. Video decoding is performed on the destination side and typically involves inverse processing relative to the encoder to reconstruct the video images. The combination of encoding and decoding is also known as encoding and decoding.

[0135] A video sequence consists of a series of pictures, which are further divided into slices, and slices into blocks. Video coding is performed on a block-by-block basis. In some newer video coding standards, the concept of a block has been further expanded. For example, the H.264 standard uses macroblocks (MBs), which can be further divided into multiple prediction blocks (partitions) for predictive coding. The High Efficiency Video Coding (HEVC) standard uses basic concepts such as coding units (CUs), prediction units (PUs), and transform units (TUs) to functionally divide various block units, and employs a novel tree-based structure for description. For instance, a CU can be divided into smaller CUs using a quadtree, and these smaller CUs can be further divided, forming a quadtree structure. The CU is the basic unit for partitioning and encoding the image. Similar tree structures exist for PUs and TUs. A PU corresponds to a prediction block and is the basic unit for predictive coding. CUs are further divided into multiple PUs according to partitioning patterns. TU can correspond to a transform block, which is the basic unit for transforming the prediction residual. However, regardless of CU, PU, ​​or TU, they all essentially belong to the concept of block data (or image blocks).

[0136] For example, in HEVC, the CTU is split into multiple CUs using a quadtree structure represented as a coding tree. At the CU level, a decision is made on whether to use inter-picture (temporal) or intra-picture (spatial) prediction to encode picture regions. Each CU can be further split into one, two, or four PUs based on the PU splitting type. The same prediction process is applied within a PU, and relevant information is transmitted to the decoder based on the PU. After obtaining residual blocks by applying the prediction process based on the PU splitting type, the CU can be segmented into transform units (TUs) according to other quadtree structures similar to the coding tree used for CUs. In the latest developments in video compression technology, quadtree and binary tree (QTBT) frame segmentation is used to divide coding blocks. In the QTBT block structure, CUs can be square or rectangular in shape.

[0137] In this paper, for ease of description and understanding, the block of data to be processed in the current image is referred to as the current block. For example, in encoding, the current block refers to the block currently being encoded; in decoding, the current block refers to the block currently being decoded. The block in the reference image used to predict the current block is called the reference block; that is, the reference block is the block that provides the reference signal for the current block, where the reference signal represents the pixel value within the image block. In the inter-frame prediction process, the block in the reference image that provides the prediction signal for the current block is called the prediction block, where the prediction signal represents the pixel value, sampled value, or sampled signal within the prediction block. For example, after traversing multiple reference blocks, an optimal reference block is found, and this optimal reference block will provide the prediction for the current block; this block is called the prediction block.

[0138] In lossless video coding, the original video image can be reconstructed, meaning the reconstructed video image has the same quality as the original (assuming no transmission loss or other data loss during storage or transmission). In lossy video coding, further compression is performed, for example, through quantization, to reduce the amount of data required to represent the video image. However, the decoder cannot fully reconstruct the video image, meaning the quality of the reconstructed video image is lower or worse than the original video image.

[0139] The encoding / decoding method of this application embodiment encodes and decodes images or videos on a block-by-block basis.

[0140] In some embodiments, the block to be encoded or the block to be decoded may be an image block or a video block obtained from an image or video.

[0141] In some embodiments, the block to be encoded or the block to be decoded may be an image block or video block within a subgraph obtained by dividing the image or video into subgraphs.

[0142] In some embodiments, the block to be encoded or the block to be decoded can be a block obtained from the transformed data of the image or video to be encoded after transformation processing (e.g., wavelet transform (also known as wavelet forward transform, etc., without limitation).

[0143] This application does not restrict the specific division method of the blocks to be encoded or decoded in an image or video, nor does it restrict the method of obtaining such blocks.

[0144] On the encoding side, the image to be encoded can be a frame from an image or video, a sub-image obtained by dividing a frame from an image or video, an image obtained after transformation (such as wavelet transform), or a sub-image from the transformed image; there are no restrictions here. On the encoding side, "current frame" can represent "image to be encoded," and "current block" can represent "block to be encoded." A frame is a frame of image to be displayed.

[0145] Similarly, the decoding side corresponds to the encoding side. On the decoding side, the image to be decoded can be a frame in an image or video, a sub-image obtained by dividing a frame in an image or video, an image obtained after transformation (such as wavelet transform), or a sub-image in the transformed image; there are no restrictions here. On the decoding side, "current frame" can represent "image to be decoded," and "current block" can represent "block to be decoded." A frame is a frame of image to be displayed.

[0146] Whether on the encoding or decoding side, the reconstructed data (also called the reconstructed value) of a block can be described as a reconstructed block, and the already encoded or decoded block referenced when encoding or decoding the current block can be described as a "reference block".

[0147] Whether on the encoding or decoding side, the prediction data for the current block is also referred to as the prediction block or the prediction value for the current block.

[0148] Whether on the encoding or decoding side, reconstruction is also referred to as refactoring.

[0149] Whether on the encoding or decoding side, the bitstream is also described as a bitstream, etc.

[0150] Whether on the encoding or decoding side, the residual is also referred to as a residual block.

[0151] Whether on the encoding or decoding side, the actual value of the current block is also expressed as the original value of the current block, or the value to be encoded in the current block.

[0152] In the accompanying drawings of the various embodiments described below, dashed boxes or dashed arrows indicate that the step or the data transmitted indicated by the arrow is optional.

[0153] The system architecture used in the embodiments of this application is described below. Referring to FIG1A, FIG1A provides an exemplary schematic block diagram of a video encoding and decoding system 10 used in the embodiments of this application. As shown in FIG1A, the video encoding and decoding system 10 may include a source device 12 and a destination device 14. The source device 12 generates encoded video data, and therefore, the source device 12 may be referred to as a video encoding device. The destination device 14 can decode the encoded video data generated by the source device 12, and therefore, the destination device 14 may be referred to as a video decoding device. Various embodiments of the source device 12, the destination device 14, or both may include one or more processors and memory coupled to the one or more processors. The memory may include, but is not limited to, RAM, ROM, EEPROM, flash memory, or any other media that can be used to store desired program code in the form of computer-accessible instructions or data structures. The source device 12 and the destination device 14 may include a variety of devices, including desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, handsets, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, wireless communication devices, or the like.

[0154] Source device 12 and destination device 14 can communicate via link 13, through which destination device 14 can receive encoded video data from source device 12. Link 13 may include one or more media or devices capable of transmitting encoded video data from source device 12 to destination device 14. In one example, link 13 may include one or more communication media enabling source device 12 to transmit encoded video data to destination device 14 in real time. In this example, source device 12 may modulate the encoded video data according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated video data to destination device 14. The one or more communication media may include wireless and / or wired communication media, such as radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network, such as a local area network, wide area network, or global network (e.g., the Internet). The one or more communication media may include routers, switches, base stations, or other devices facilitating communication from source device 12 to destination device 14.

[0155] The source device 12 includes an encoder 20. Optionally, the source device 12 may also include an image source 16, an image preprocessor 18, and a communication interface 22. In specific implementations, the encoder 20, image source 16, image preprocessor 18, and communication interface 22 may be hardware components or software programs within the source device 12. These are described below:

[0156] Image source 16 may include or be any type of image capture device for, for example, capturing real-world images, and / or any type of image or commentary (for screen content encoding, some text on the screen is also considered as an image to be encoded or part of an image) generation device, such as a computer graphics processor for generating computer-animated images, or any type of device for acquiring and / or providing real-world images, computer-animated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). Image source 16 may be a camera for capturing images or a memory for storing images. Image source 16 may also include any type of (internal or external) interface for storing previously captured or generated images and / or acquiring or receiving images. When image source 16 is a camera, image source 16 may be, for example, a local or integrated camera integrated into a source device; when image source 16 is a memory, image source 16 may be a local or integrated memory integrated into a source device. When the image source 16 includes an interface, the interface may be, for example, an external interface for receiving images from an external video source, such as an external image capture device, like a camera, external storage, or an external image generation device, such as an external computer graphics processor, computer, or server. The interface can be any type of interface according to any proprietary or standardized interface protocol, such as a wired or wireless interface, or an optical interface.

[0157] An image can be viewed as a two-dimensional array or matrix of pixels. Pixels in the array are also called sampling points. The number of sampling points in the array or image along the horizontal and vertical directions (or axes) defines the image's size and / or resolution. To represent color, three color components are typically used; that is, an image can be represented as or contain three sampling arrays. For example, in RBG format or color space, an image includes corresponding red, green, and blue sampling arrays. However, in video coding, each pixel is typically represented in a luma / chroma format or color space. For example, for a YUV format image, this includes a luma component indicated by Y (sometimes also indicated by L) and two chroma components indicated by U and V. The luma component Y represents the brightness or grayscale level intensity (e.g., both are the same in a grayscale image), while the two chroma components U and V represent chroma or color information components. Accordingly, a YUV format image includes a luma sampling array of luma sample values ​​(Y) and two chroma sampling arrays of chroma values ​​(U and V). An RGB format image can be converted or transformed to YUV format, and vice versa; this process is also called color transformation or conversion. If the image is black and white, it may only include a luminance sampling array. In this embodiment, the image transmitted from image source 16 to image processor can also be referred to as raw image data 17.

[0158] Image preprocessor 18 is configured to receive raw image data 17 and perform preprocessing on the raw image data 17 to obtain a preprocessed image 19 or preprocessed image data 19. For example, the preprocessing performed by image preprocessor 18 may include retouching, color format conversion (e.g., from RGB format to YUV format), color correction, or noise reduction.

[0159] Encoder 20 (or video encoder 20) is used to receive preprocessed image data 19 and process the preprocessed image data 19 using a relevant prediction mode (such as the prediction mode in the various embodiments herein) to provide encoded image data 21. In some embodiments, encoder 20 may be used to perform the various embodiments described below to implement the encoding method described herein on the encoding side.

[0160] Communication interface 22 can be used to receive encoded image data 21 and transmit the encoded image data 21 via link 13 to destination device 14 or any other device (such as a memory) for storage or direct reconstruction. The other device can be any device used for decoding or storage. Communication interface 22 can, for example, be used to encapsulate the encoded image data 21 into a suitable format, such as data packets, for transmission over link 13.

[0161] Destination device 14 includes decoder 30. Optionally, destination device 14 may also include communication interface 28, image post-processor 32, and display device 34. These are described below:

[0162] Communication interface 28 can be used to receive encoded image data 21 from source device 12 or any other source, such as a storage device, for example, an encoded image data storage device. Communication interface 28 can be used to transmit or receive encoded image data 21 via link 13 between source device 12 and destination device 14 or via any type of network, such as a wired or wireless connection, any type of network, such as a wired or wireless network or any combination thereof, or any type of private and public network, or any combination thereof. Communication interface 28 can be used, for example, to decapsulate data packets transmitted by communication interface 22 to obtain encoded image data 21.

[0163] Both communication interface 28 and communication interface 22 can be configured as unidirectional or bidirectional communication interfaces, and can be used, for example, to send and receive messages to establish connections, acknowledge and exchange any other information related to the communication link and / or data transmission, such as encoded image data transmission.

[0164] Decoder 30 (or video decoder 30) is used to receive encoded image data 21 and provide decoded image data 31 or decoded image 31 (the structural details of decoder 30 will be further described below based on FIG3, FIG4 or FIG5). In some embodiments, decoder 30 can be used to perform the various embodiments described below to implement the decoding method described in this application on the decoding side.

[0165] Image post-processor 32 is used to perform post-processing on decoded image data 31 (also referred to as reconstructed image data) to obtain post-processed image data 33. The post-processing performed by image post-processor 32 may include: color format conversion (e.g., from YUV format to RGB format), color correction, retouching or resampling, or any other processing, and may also be used to transmit the post-processed image data 33 to display device 34.

[0166] Display device 34 is used to receive post-processed image data 33 to display an image to, for example, a user or viewer. Display device 34 can be or may include any class of displays for presenting reconstructed images, such as integrated or external displays or monitors. For example, displays may include liquid crystal displays (LCDs), organic light emitting diode (OLED) displays, plasma displays, projectors, micro-LED displays, liquid crystal on silicon (LCoS), digital light processors (DLP), or any other class of displays.

[0167] Although Figure 1A illustrates source device 12 and destination device 14 as separate devices, device embodiments may also include the functionality of both source device 12 and destination device 14, or both; that is, the functionality of source device 12 or its corresponding features and the functionality of destination device 14 or its corresponding features. In such embodiments, the same hardware and / or software, or separate hardware and / or software, or any combination thereof, may be used to implement the functionality of source device 12 or its corresponding features and the functionality of destination device 14 or its corresponding features.

[0168] Both encoder 20 and decoder 30 can be implemented as any of a variety of suitable circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology is implemented in part in software, the device can store the software instructions in a suitable non-transitory computer-readable storage medium, and one or more processors can be used to execute the instructions in hardware to perform the technology of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) can be considered as one or more processors.

[0169] In some cases, the video encoding and decoding system 10 shown in Figure 1A is merely an example, and the technology of this application can be applied to video encoding setups (e.g., video encoding or video decoding) that do not necessarily involve any data communication between the encoding and decoding devices. In other instances, data may be retrieved from local storage, streamed over a network, etc. The video encoding device may encode the data and store it in storage, and / or the video decoding device may retrieve the data from storage and decode it. In some instances, encoding and decoding are performed by devices that do not communicate with each other but only encode data to storage and / or retrieve data from storage and decode the data.

[0170] Referring to FIG1B, FIG1B is an illustrative diagram of an example of a video decoding system 40 including the encoder 20 of FIG2 and / or the decoder 30 of FIG3 according to an exemplary embodiment. The video decoding system 40 can implement various combinations of technologies of the embodiments of this application. In the illustrated embodiment, the video decoding system 40 may include an imaging device 41, an encoder 20, a decoder 30 (and / or a video encoder / decoder implemented by logic circuitry of a processing unit 46), an antenna 42, one or more processors 43, one or more memories 44, and a display device 45.

[0171] As shown in Figure 1B, the imaging device 41, antenna 42, processing unit 46, logic circuit, encoder 20, decoder 30, processor 43, memory 44, and display device 45 are capable of communicating with each other. As discussed, although encoder 20 and decoder 30 are used as examples to describe the video decoding system 40, in different instances, the video decoding system 40 may contain only encoder 20 or only decoder 30.

[0172] In some instances, antenna 42 can be used to transmit or receive encoded video data streams. Additionally, in some instances, display device 45 can be used to present video data. In some instances, logic circuitry can be implemented using processing unit 46. Processing unit 46 can include an ASIC, graphics processor, general-purpose processor, etc. Video decoding system 40 can also include an optional processor 43, which can similarly include an ASIC, graphics processor, general-purpose processor, etc. In some instances, logic circuitry can be implemented in hardware, such as dedicated video encoding hardware, while processor 43 can be implemented in general-purpose software, operating system, etc. Furthermore, memory 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting instance, memory 44 can be implemented using cache memory. In some instances, logic circuitry can access memory 44 (e.g., for implementing an image buffer). In other instances, the logic circuitry and / or processing unit 46 may include memory (e.g., cache, etc.) for implementing image buffers, etc.

[0173] In some instances, the encoder 20 implemented via logic circuitry may include (e.g., implemented via processing unit 46 or memory 44) an image buffer and (e.g., implemented via processing unit 46) a graphics processing unit. The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the encoder 20 implemented via logic circuitry to implement various modules discussed with reference to Figure 2 and / or any other encoder system or subsystem described herein. The logic circuitry may be used to perform various operations discussed herein.

[0174] In some instances, decoder 30 may be implemented via logic circuitry in a similar manner to implement the various modules discussed in reference to decoder 30 of Figure 3 and / or any other decoder system or subsystem described herein. In some instances, the logic circuitry-implemented decoder 30 may include an image buffer (implemented via processing unit 46 or memory 44) and a graphics processing unit (e.g., implemented via processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include decoder 30 implemented via logic circuitry to implement the various modules discussed in reference to Figure 3 and / or any other decoder system or subsystem described herein.

[0175] In some instances, antenna 42 can be used to receive an encoded stream of video data. As discussed herein, the encoded stream may contain data related to encoded video frames, indicators, index values, mode selection data, etc., such as data related to code segmentation (e.g., transform coefficients or quantized transform coefficients, optional indicators, and / or data defining code segmentation). Video decoding system 40 may also include a decoder 30 coupled to antenna 42 for decoding the encoded stream. Display device 45 is used to display the video frames.

[0176] It should be understood that, referring to the examples described for encoder 20 in the embodiments of this application, decoder 30 can be used to perform the reverse process. Regarding signaling syntax elements, decoder 30 can be used to receive and parse such syntax elements, and accordingly decode the associated video data. In some examples, encoder 20 can entropy-encode syntax elements into an encoded video stream. In such instances, decoder 30 can parse such syntax elements and accordingly decode the associated video data.

[0177] It should be noted that the encoding and decoding method described in the embodiments of this application is mainly used for the encoding and decoding process of video or images. This process exists in both the encoder 20 and the decoder 30. The encoder 20 and decoder 30 in the embodiments of this application can be, for example, the encoder / decoder corresponding to video standard protocols such as H.263, H.264, HEVV, MPEG-2, MPEG-4, VP8, VP9, ​​or next-generation video standard protocols (such as H.266).

[0178] Referring to Figure 2, which is a schematic / conceptual block diagram of an exemplary example of encoder 20, encoder 20 includes a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a buffer 216, a loop filter unit 220, a decoded picture buffer (DPB) 230, a prediction processing unit 260, and an entropy coding unit 270. Prediction processing unit 260 may include inter-frame prediction unit 244, intra-frame prediction unit 254, and mode selection unit 262. Inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Encoder 20 shown in Figure 2 may also be referred to as a hybrid video encoder or a video encoder based on a hybrid video codec.

[0179] Referring to Figure 3, which is a schematic / conceptual block diagram of an example of a decoder 30, the decoder 30 is used to receive, for example, encoded image data (e.g., encoded bitstream) 21 encoded by encoder 20 to obtain a decoded image 331. During the decoding process, the decoder 30 receives video data from encoder 20, such as encoded video bitstreams representing image blocks of encoded video stripes and associated syntax elements.

[0180] In the example of Figure 3, decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., a summer 314), a buffer 316, a loop filter 320, a decoded image buffer 330, and a prediction processing unit 360. Prediction processing unit 360 may include an inter-frame prediction unit 344, an intra-frame prediction unit 354, and a mode selection unit 362. In some instances, decoder 30 may perform a decoding process that is generally the inverse of the encoding process described in video encoder 20 of Figure 2.

[0181] Referring to Figure 4, Figure 4 is a structural schematic diagram of a video decoding device 400 (e.g., a video encoding device 400 or a video decoding device 400) provided in an embodiment of this application. The video decoding device 400 is adapted to implement the embodiments described herein. In one embodiment, the video decoding device 400 may be a video decoder (e.g., decoder 30 of Figure 1A) or a video encoder (e.g., encoder 20 of Figure 1A). In another embodiment, the video decoding device 400 may be one or more components of the decoder 30 of Figure 1A or the encoder 20 of Figure 1A.

[0182] The video decoding device 400 includes: an input port 410 and a receiving unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an output port 450 for transmitting data; and a memory 460 for storing data. The video decoding device 400 may also include photoelectric conversion components and electro-optical (EO) components coupled to the input port 410, receiver unit 420, transmitter unit 440, and output port 450 for the input or output of optical or electrical signals.

[0183] Processor 430 is implemented in both hardware and software. Processor 430 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 430 communicates with ingress port 410, receiver unit 420, transmitter unit 440, egress port 450, and memory 460. Processor 430 includes either an encoding module 470 or a decoding module 470. The encoding / decoding module 470 implements the embodiments disclosed herein to implement the encoding or decoding methods provided in the embodiments of this application. For example, the encoding / decoding module 470 implements, processes, or provides various encoding operations. Therefore, the encoding / decoding module 470 provides a substantial improvement to the functionality of the video decoding device 400 and affects the transitions of the video decoding device 400 to different states. Alternatively, the encoding / decoding module 470 can be implemented with instructions stored in memory 460 and executed by processor 430.

[0184] Memory 460 includes one or more disks, tape drives, and solid-state drives, which can be used as overflow data storage devices to store programs while they are selectively executed, and to store instructions and data read during program execution. Memory 460 can be volatile and / or non-volatile, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).

[0185] Referring to FIG5, FIG5 is a simplified block diagram of an apparatus 500 that can be used as either or both of the source device 12 and destination device 14 in FIG1A according to an exemplary embodiment. The apparatus 500 can implement the technology of this application. In other words, FIG5 is a schematic block diagram of an implementation of an encoding or decoding apparatus (hereinafter referred to as decoding apparatus 500) according to an embodiment of this application. The decoding apparatus 500 may include a processor 510, a memory 530, and a bus system 550. The processor and the memory are connected via the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory. The memory of the decoding apparatus stores program code, and the processor can call the program code stored in the memory to execute various video encoding or decoding methods described in this application. To avoid repetition, detailed descriptions are not provided here.

[0186] In this embodiment, the processor 510 may be a central processing unit (CPU), or it may be other general-purpose processors, DSPs, ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0187] The memory 530 may include a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may also be used as memory 530. Memory 530 may include code and data 531 accessed by processor 510 using bus 550. Memory 530 may further include an operating system 533 and an application program 535, which includes at least one program that allows processor 510 to execute the encoding or decoding methods described in this application. For example, application program 535 may include applications 1 to N, which further include video encoding or decoding applications that execute the encoding or decoding methods described in this application.

[0188] In addition to the data bus, the bus system 550 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 550 in the diagram.

[0189] Optionally, the decoding device 500 may also include one or more output devices, such as a display 570. In one example, the display 570 may be a haptic display that combines a display with a haptic unit capable of operatively sensing touch input. The display 570 may be connected to the processor 510 via a bus 550.

[0190] For example, commonly used transform methods in image coding include discrete cosine transform and wavelet transform. Wavelet transform is a local transform method that can perform localized, multi-scale analysis of images, focusing on the details of signal changes, making it very suitable for image coding tasks.

[0191] Referring to Figure 6, which is an exemplary schematic / conceptual block diagram of a wavelet transform-based encoder, the encoder 60 includes, but is not limited to, a wavelet forward transform unit 610, a quantization unit 620, and an entropy coding unit 630. Specifically, the encoder acquires the original image 601 through an interface unit (also called an input interface, not shown in the figure). The original image 601 is input to the wavelet forward transform unit 610, and after wavelet transform (also called wavelet forward transform), wavelet transform coefficients 602 are obtained, which can also be simply referred to as wavelet coefficients in this embodiment. The quantization unit 620 quantizes the wavelet transform coefficients 602 to obtain quantization coefficients 603. The entropy coding unit 630 entropy codes the quantization coefficients 603 to obtain the image bitstream 604 of the original image, which can also be called a compressed bitstream or bitstream.

[0192] It should be noted that the wavelet transform-based encoder architecture in Figure 6 is for illustrative purposes only, and other instances may include more units or modules. For example, it may include, but is not limited to, prediction units, transform units, etc., and this application does not impose any limitations.

[0193] Referring to Figure 7, which is an exemplary schematic / conceptual block diagram of a wavelet transform-based decoder, the decoder 70 in the example of Figure 7 includes, but is not limited to, an entropy decoding unit 710, an inverse quantization unit 720, and an inverse wavelet transform unit 730. Specifically, at the decoding end, the decoder interfaces with the image bitstream 701 through an input interface (also called an interface unit, not shown in the figure). The entropy decoding unit 710 performs entropy decoding on the image bitstream 701 to obtain quantization coefficients 702. The image bitstream 701 can be a bitstream 604 generated based on the encoder 60 in Figure 6. The inverse quantization unit 720 performs inverse quantization on the quantization coefficients 702 to obtain inverse quantization coefficients 703 (dequantized coefficient(s)), which can also be called reconstructed coefficients, reconstructed wavelet coefficients, etc. The inverse wavelet transform unit 730 performs inverse wavelet transform on the inverse quantization coefficients 703 to obtain the reconstructed image 704.

[0194] It should be noted that the wavelet transform-based decoder architecture in Figure 7 is for illustrative purposes only, and other instances may include more units or modules. For example, it may include, but is not limited to, prediction units, inverse transform units, etc., and this application does not impose any limitations.

[0195] For example, the wavelet coefficients obtained after wavelet transform include wavelet coefficients of the high-frequency subband and wavelet coefficients of the low-frequency subband. In the architecture shown in Figures 6 and 7, the codec performs the same processing on the wavelet coefficients of the high-frequency subband and the low-frequency subband, which has high processing complexity. In the embodiments of this application, the wavelet coefficients of the high-frequency subband can be simply referred to as the high-frequency subband, and the wavelet coefficients of the low-frequency subband can be simply referred to as the low-frequency subband. It can be understood that the high-frequency subband obtained after wavelet transform can optionally be the set of wavelet coefficients of the high-frequency subband, and the low-frequency subband obtained after wavelet transform can optionally be the set of wavelet coefficients of the low-frequency subband.

[0196] For example, the frequency of the high-frequency sub-band is greater than the frequency of the low-frequency sub-band. This can be understood as the frequency of various coefficients corresponding to the high-frequency sub-band in this application (such as wavelet coefficients, reconstruction coefficients, inverse quantization coefficients, etc. involved in this application) being greater than the frequency of the coefficients of the low-frequency sub-band.

[0197] This application provides a wavelet transform-based codec that can independently encode and decode low-frequency and high-frequency sub-bands, effectively reducing encoding and decoding complexity and improving efficiency. For example, an image undergoes wavelet transform to obtain low-frequency and high-frequency sub-bands, which are then encoded to generate low-frequency and high-frequency sub-band bitstreams, respectively. The low-frequency sub-band can be understood as a sub-image representing the low-frequency signal (or low-frequency information) of the original image, and the high-frequency sub-band can be understood as a sub-image representing the high-frequency signal (or high-frequency information) of the original image.

[0198] Referring to Figure 8, which is a schematic / conceptual block diagram of an encoder as an example, the encoder 80 includes, but is not limited to, a sub-graph partitioning unit 810, a wavelet forward transform unit 820, a low-frequency sub-band processing path 830, and a high-frequency sub-band processing path 840. Optionally, the wavelet forward transform unit and the wavelet inverse transform unit in this embodiment can also be collectively referred to as wavelet transform units, which can be understood as performing both wavelet forward transform processing and wavelet inverse transform processing. In this application, the low-frequency sub-band processing path can also be referred to as a low-frequency sub-band processing unit, and the high-frequency sub-band processing path can also be referred to as a high-frequency sub-band processing unit.

[0199] The low-frequency subband processing path 830 includes, but is not limited to: a transform / quantization unit 831 (also known as a low-frequency subband transform / quantization unit) and a low-frequency subband entropy coding unit 832.

[0200] The high-frequency subband processing path 840 includes, but is not limited to: a transform / quantization unit 841 (also known as a high-frequency subband transform / quantization unit) and a high-frequency subband entropy coding unit 842.

[0201] Specifically, encoder 80 receives image 801. Image 801 can be an image in an image sequence that forms a video or video sequence. Image 801 can be referred to as the current image or the image to be encoded, and image 801 can be the original image or an image obtained by processing the original image.

[0202] The sub-image partitioning unit 810 is used to acquire the current image and partition it to obtain at least one sub-image. Specifically, the sub-image partitioning unit 810 partitions the current image into N sub-images according to a sub-image partitioning method, where N is an integer greater than 0 (or an integer greater than 1). The sub-image partitioning method can include, but is not limited to, at least one of the following:

[0203] The width and / or height of the subgraph are multiples of 128;

[0204] The maximum width of the subimage is 1024 pixels;

[0205] The minimum height and / or width of the subimage is 256 pixels;

[0206] The original image resolution is less than or equal to 1080p, and N is an integer greater than 1 and less than or equal to 8; or,

[0207] The original image has a length greater than or equal to 4320 pixels, a width greater than or equal to 2160 pixels, and N is an integer greater than 1 and less than or equal to 16; or,

[0208] The original image has a length greater than or equal to 7680 pixels, a width greater than or equal to 4320 pixels, and N is an integer greater than 1 and less than or equal to 32.

[0209] For example, common video resolutions include: 1280x720, 1280x1080, 1440x1080, 1920x1080 (1080p), 2048x1080, 2048x1556, 3840x2160 (4K), 4096x2160, 5120x2700, 6144x3240, 7680x4320 (8K), and 8192x4320. The above values ​​are only illustrative examples and can be set according to actual needs.

[0210] For example, the subgraph partitioning unit 810 can set specific subgraph specifications (including width and height) based on the above subgraph partitioning method. The specific values ​​can be set according to actual needs within the range specified by the subgraph partitioning method.

[0211] Referring to Figure 9A, which is an exemplary schematic diagram of sub-image partitioning, in the example of Figure 9A, the sub-image partitioning unit 810 can partition the image 801 into m*n sub-images according to the sub-image partitioning method. Optionally, in this example, the width and height of each sub-image satisfy a multiple of 128.

[0212] Referring to Figure 9B, which is an exemplary schematic diagram of sub-image partitioning, in the example of Figure 9B, the sub-image partitioning unit 810 can partition the image into m*n sub-images according to the sub-image partitioning method. When the sub-image partitioning unit 810 partitions the boundary portion of the current image, the width and / or height of the boundary sub-images may optionally not satisfy a multiple of 128. Figure 9B only illustrates this by showing that the width of some boundaries does not satisfy a multiple of 128. Referring to Figure 9B, in this embodiment, for sub-images whose width and height do not satisfy a multiple of 128, the sub-image partitioning unit 810 can pad these sub-images, such as the gray portion in Figure 9B. Taking sub-image 1_n as an example, the height of sub-image 1_n satisfies a multiple of 128, but its width does not. The sub-image partitioning unit 810 pads sub-image 1_n.

[0213] In one possible implementation, the width and / or height of the filled subgraph can optionally be less than or equal to the subgraph size set by the subgraph partitioning method. For example, the size (i.e., width and height) of the filled subgraph 1_n can be the same as that of subgraph 1_1. As another example, the size of the filled subgraph 1_n can be less than the size of subgraph 1_1, where the width of subgraph 1_1 is a, the height is b, the width of subgraph 1_n before filling is c, the height is b, c is less than b, and c is not a multiple of 128. The subgraph partitioning unit 810 fills the subgraph 1_n, and the width of the filled subgraph 1_n is d, and the height is b. Here, d can optionally be a multiple of 16, and d is greater than c and less than or equal to a.

[0214] In this way, by padding the sub-images, each macroblock can meet the 8*8 (pixel) partitioning requirement during the block division process of the encoding. Furthermore, by finding the least common multiple of the width and / or height of the sub-image before padding and 16, sub-images that do not meet the multiple of 16 are padded. This ensures that the size of the sub-image is a multiple of 16 while minimizing the size of the padded sub-image, thereby reducing the encoding complexity of the padded sub-image.

[0215] Optionally, the sub-image partitioning method can be the same for each image in the same video sequence.

[0216] Optionally, the sub-image division unit 810 can divide the image into sub-images along the horizontal and vertical directions, starting from the upper left corner, according to a preset sub-image size. The division order can be set according to actual needs. Typically, the sub-images that need to be filled are boundary sub-images, as shown in Figure 9B.

[0217] In this embodiment, each sub-image is encoded and decoded independently. Specifically, the image is divided into N sub-images, and the encoder 80 encodes each of the N sub-images separately. During the encoding process, a sub-image can also be referred to as the current sub-image or the sub-image to be encoded.

[0218] Referring again to Figure 8, the wavelet forward transform unit 820 is used to perform wavelet transform (also called wavelet forward transform) on the subgraph to obtain low-frequency subband and high-frequency subband. The low-frequency subband includes low-frequency signals in the subgraph that satisfy the low-frequency filter coefficients, and the high-frequency subband includes high-frequency signals in the subgraph that have been decomposed by the high-frequency filter in the wavelet transform.

[0219] Referring to Figure 10, which is an exemplary schematic diagram of wavelet transform, in the example of Figure 10, the wavelet forward transform unit 820 acquires the current sub-image, for example, sub-image 1_1. The wavelet forward transform unit 820 performs a wavelet transform on the current sub-image, wherein the wavelet transform includes one horizontal wavelet transform and one vertical wavelet transform to obtain the wavelet coefficients of the low-low (LL) sub-band (abbreviated as LL sub-band), the wavelet coefficients of the low-high (LH) sub-band (abbreviated as LH sub-band), the wavelet coefficients of the high-high (HH) sub-band (abbreviated as HH sub-band), and the wavelet coefficients of the high-low (HL) sub-band (abbreviated as HL sub-band).

[0220] In this embodiment, the low-frequency subband includes an LL subband, and the high-frequency subband includes an LH subband, an HH subband, and an HL subband. Optionally, the LL subband, LH subband, HH subband, and HL subband have the same dimensions (including width and height).

[0221] Referring again to Figure 8, the low-frequency subband processing path 830 is used to encode the low-frequency subband to obtain low-frequency subband encoded data 806, which can also be referred to as coded low-frequency subband. The high-frequency subband processing path 840 is used to encode the high-frequency subband to obtain high-frequency subband encoded data 809, which can also be referred to as coded high-frequency subband. In this application, the high-frequency subband encoded data can also be referred to as second encoded data.

[0222] Specifically, the wavelet forward transform unit 820 outputs wavelet coefficients 803 of the low-frequency subband to the transform / quantization unit 831, and outputs wavelet coefficients 804 of the high-frequency subband to the transform / quantization unit 841.

[0223] For example, the transform / quantization unit 831 is used to obtain the wavelet coefficients 803 of the low-frequency sub-band, perform quantization processing on the wavelet coefficients 803 of the low-frequency sub-band, or perform transform and quantization processing, and output the quantized coefficients 805 of the low-frequency sub-band, or the quantized wavelet coefficients of the low-frequency sub-band.

[0224] The transform / quantization unit 831 outputs the quantization coefficients 805 of the low-frequency subband to the low-frequency subband entropy encoding unit 832. Optionally, the transform / quantization unit 831 may include, but is not limited to, a transform processing unit and a quantization processing unit (not shown in the figure).

[0225] The transform processing unit is used to transform the low-frequency subband and output transform coefficients, also known as low-frequency subband transform coefficients. The quantization processing unit is used to quantize the subband or the transformed transform coefficients and output quantized coefficients (also known as quantization results).

[0226] The low-frequency subband entropy coding unit 832 is used to acquire the data to be encoded and perform entropy coding on the data to be encoded to obtain low-frequency subband encoded data. The data to be encoded may include, but is not limited to, the quantization coefficients 805 of the low-frequency subband. In some instances, the data to be encoded may also include parameters or information (or syntax elements) input from other modules not shown in the figure. The low-frequency subband encoded data may also be referred to as the first encoded data in this application.

[0227] Referring again to Figure 8, the transform / quantization unit 841 is used to quantize the wavelet coefficients 804 of the high-frequency subband, or to perform both transform and quantization processing, to obtain the quantized coefficients 807 of the high-frequency subband, which can also be referred to as the quantized wavelet coefficients of the high-frequency subband. The transform / quantization unit 841 outputs the quantized coefficients 807 of the high-frequency subband to the high-frequency subband entropy coding unit 842. For parts not described, please refer to the low-frequency subband transform / quantization unit 831; further details are omitted here.

[0228] The high-frequency subband entropy coding unit 842 is used to acquire the data to be encoded and perform entropy coding on the data to be encoded to obtain high-frequency subband encoded data 808, which can also be called encoded high-frequency subband. Optionally, the data to be encoded may include, but is not limited to, the quantization coefficients 807 of the high-frequency subband. In some instances, the data to be encoded may also include parameters or information (or syntax elements) input from other modules not shown in the figure.

[0229] Optionally, the entropy coding unit (including the low-frequency subband entropy coding unit 832 and the high-frequency subband entropy coding unit 842) is used to encode the data to be encoded using an entropy coding algorithm or scheme. The aforementioned entropy coding scheme may be, for example, at least one of the following: variable-length coding (VLC), context-adaptive VLC (CAVLC), arithmetic coding, and context-adaptive binary arithmetic coding (CABAC).

[0230] Optionally, the encoder 80 may also include, but is not limited to, a combining unit (not shown in the figure), which may also be called a multiplexer (MUX). The combining unit is used to generate an image bitstream based on the low-frequency subband coded data 806 and the high-frequency subband coded data 808.

[0231] Specifically, the combining unit writes low-frequency subband coded data 806 into the image bitstream and writes high-frequency subband coded data 808 into the image bitstream. In this embodiment, by encoding the low-frequency subband and high-frequency subband separately, the low-frequency subband coded data and high-frequency subband coded data can be decoded independently. That is, at the decoding end, it can independently decode the low-frequency subband coded data and high-frequency subband coded data in the image bitstream, thereby improving decoding efficiency.

[0232] As mentioned above, a single image can be divided into multiple sub-images, for example, N sub-images. Each sub-image can be decomposed into wavelet coefficients of a high-frequency sub-band and wavelet coefficients of a low-frequency sub-band after wavelet forward transform. Encoder 80 independently encodes the low-frequency and high-frequency sub-bands of each sub-image to obtain the low-frequency sub-band encoded data and high-frequency sub-band encoded data of a single sub-image. In this way, the combining unit can obtain the high-frequency sub-band encoded data and low-frequency sub-band encoded data of each of the N sub-images, that is, obtain N high-frequency sub-band encoded data (optionally, the high-frequency sub-band encoded data of each sub-image includes the HL sub-band encoded data corresponding to the HL sub-band, the HH sub-band encoded data corresponding to the HH sub-band, and the LH sub-band encoded data corresponding to the LH sub-band) and N low-frequency sub-band encoded data. The combining unit can optionally write the low-frequency sub-band encoded data and high-frequency sub-band encoded data of the N sub-images into the image bitstream. The format (or structure) of the image bitstream will be described in detail below.

[0233] Referring to Figure 11, which is a schematic / conceptual block diagram of an exemplary decoder, decoder 110 is used to receive, for example, a bitstream encoded by an encoder to obtain a decoded image, also referred to as a reconstructed image 1108. During the decoding process, the decoder receives a bitstream (also referred to as the bitstream of the original image, or the image data of the original image) from the encoder.

[0234] In the example shown in Figure 11, the decoder 110 includes, but is not limited to: a high-frequency subband processing path 1120, a low-frequency subband processing path 1110, a wavelet inverse transform unit 1130, and a subgraph stitching unit 1140.

[0235] The low-frequency subband processing path 1110 includes, but is not limited to: a low-frequency subband entropy decoding unit 1111 and an inverse quantization / inverse transform unit 1112 (also referred to as a low-frequency subband inverse quantization / inverse transform unit). The high-frequency subband decoding path 1120 includes, but is not limited to: a high-frequency subband entropy decoding unit 1121 and an inverse quantization / inverse transform unit 1122 (also referred to as a high-frequency subband inverse quantization / inverse transform unit).

[0236] For example, decoder 110 receives an image bitstream and obtains a reconstructed image 1108 (also known as a decoded image) based on high-frequency subband coded data 1102 and low-frequency subband coded data 1101 in the image bitstream.

[0237] For example, the low-frequency subband entropy decoding unit 1111 is used to acquire low-frequency subband encoded data 1101 in the image bitstream to obtain entropy decoding data, which includes, but is not limited to, the quantization coefficients 1103 of the low-frequency subband. Specifically, the low-frequency subband entropy decoding unit 1111 performs entropy decoding processing on the low-frequency subband encoded data 1101 to obtain entropy decoding data. The low-frequency subband entropy decoding unit 1111 outputs the quantization coefficients 1103 of the low-frequency subband to the inverse quantization / inverse transform unit 1112.

[0238] The inverse quantization / inverse transform unit 1112 is used to obtain the quantization coefficients 1103 of the low-frequency subband to obtain the reconstruction coefficients 1105 of the low-frequency subband, which can also be simply referred to as the reconstructed low-frequency subband. In some instances, it can also be called the reconstructed value of the low-frequency subband, the reconstruction coefficient of the low-frequency subband, or the reconstructed wavelet coefficients of the low-frequency subband. Specifically, the low-frequency inverse quantization / inverse transform unit performs inverse quantization and / or inverse transform processing on the quantization coefficients to obtain the reconstructed low-frequency subband, such as the reconstructed LL subband. Optionally, the inverse quantization / inverse transform unit 1112 may include, but is not limited to, an inverse quantization processing unit and an inverse transform processing unit (not shown in the figure). The inverse quantization processing unit can be used to perform inverse quantization processing on the input parameters to obtain inverse quantization coefficients. The inverse transform unit can be used to perform inverse transform processing on the input parameters to obtain inverse transform coefficients, which can also be called inverse transform coefficients. For example, the inverse transform processing unit can perform inverse transform processing on the inverse quantization coefficients output by the inverse quantization processing unit to obtain inverse transform inverse quantization coefficients.

[0239] The high-frequency subband entropy decoding unit 1121 is used to acquire the high-frequency subband encoded data 1102 in the image bitstream to obtain entropy-decoded data. The entropy-decoded data includes, but is not limited to, the quantization coefficients 1104 of the high-frequency subband. Specifically, the high-frequency subband entropy decoding unit 1121 performs entropy decoding processing on the high-frequency subband encoded data 1102 to obtain the entropy-decoded data. The high-frequency subband entropy decoding unit 1121 outputs the quantization coefficients 1104 of the high-frequency subband to the inverse quantization / inverse transform unit 1122.

[0240] The inverse quantization / inverse transform unit 1122 is used to obtain the quantization coefficients 1104 of the high-frequency subband, resulting in the reconstructed coefficients 1106 of the high-frequency subband, which can be simply referred to as the reconstructed high-frequency subband. In some instances, it can also be called the reconstructed value of the high-frequency subband, the reconstructed wavelet coefficients of the high-frequency subband, etc. The inverse quantization / inverse transform unit 1122 outputs the reconstructed high-frequency subband to the wavelet inverse transform unit 1130. Specifically, the inverse quantization / inverse transform unit 1122 performs inverse quantization processing on the quantization coefficients 1104 of the high-frequency subband, or performs inverse quantization and inverse transform processing, to obtain the reconstructed high-frequency subband. For example, this includes reconstructing the HH subband (e.g., the reconstructed coefficients of the HH subband), reconstructing the HL subband (e.g., the reconstructed coefficients of the HL subband), and reconstructing the LH subband (e.g., the reconstructed coefficients of the LH subband).

[0241] For example, the wavelet inverse transform unit 1130 is used to obtain the reconstruction coefficients 1105 of the low-frequency sub-band and the reconstruction coefficients 1106 of the high-frequency sub-band to obtain the reconstructed sub-map 1107. Specifically, the wavelet inverse transform unit 1130 can obtain the reconstructed low-frequency sub-band and the reconstructed high-frequency sub-band of each sub-map. The wavelet inverse transform unit performs wavelet inverse transform processing on the obtained reconstructed low-frequency sub-band and reconstructed high-frequency sub-band corresponding to the same sub-map to obtain the reconstructed sub-map. Its inverse transform process is the inverse of the wavelet forward transform process, which will not be described in detail here.

[0242] The sub-image stitching unit 1140, also known as an image reconstruction unit or sub-image combination unit, is used to obtain reconstructed sub-images of the original image to obtain a reconstructed image of the original image. Specifically, the sub-image stitching unit 1140 obtains N reconstructed sub-images corresponding to the current decoded image. Based on the N reconstructed sub-images, the sub-image stitching unit 1140 can obtain the reconstructed image (also known as the decoded image or the decoded image) of the current image.

[0243] Based on the codecs shown in Figures 8 and 11, there can be more variations of the codec, such as including more processing units or modules.

[0244] Referring to Figure 12, which is a schematic / conceptual block diagram of an encoder as an example, the encoder in the example of Figure 12 includes, but is not limited to, a sub-graph partitioning unit 1210, a wavelet forward transform unit 1220, a low-frequency sub-band processing path 1230, and a high-frequency sub-band processing path 1240.

[0245] The descriptions of the subgraph partitioning unit 1210 and the wavelet forward transform unit 1220 can be found in the relevant content in Figure 8, and will not be repeated here.

[0246] The low-frequency subband processing path 1230 is used to obtain the wavelet coefficients 1203 of the low-frequency subband to obtain the low-frequency subband encoded data 1213. The low-frequency subband processing path 1230 includes, but is not limited to: a block partitioning unit 1231 (also called the low-frequency subband block partitioning unit 1231 or the first block partitioning unit 1231), a residual calculation unit 1232, a prediction unit 1237, a control unit 1238, a transform / quantization unit (also called the low-frequency subband transform / quantization unit 1233 or the first transform / quantization unit), an inverse quantization / inverse transform unit 1234 (also called the low-frequency subband inverse quantization / inverse transform unit 1234 or the first inverse transform / inverse quantization unit), a low-frequency subband wavelet coefficient 1203 splicing unit 1246, a low-frequency subband wavelet coefficient 1203 splicing unit 1246, and a low-frequency subband entropy coding unit 1239, etc.

[0247] The high-frequency subband processing path 1240 is used to acquire high-frequency subbands to obtain high-frequency subband encoded data. The high-frequency subband processing path 1240 includes, but is not limited to: block partitioning unit 1231 (also referred to as high-frequency subband block partitioning unit 1231 or second block partitioning unit 1231), transform / quantization unit (also referred to as transform / quantization unit 1242 or second transform / quantization unit), high-frequency subband entropy coding unit 1243, etc.

[0248] Alternatively, in some instances, the encoder may include more or fewer units or modules than in the structure shown in Figure 12.

[0249] The image 1201 encoding method provided in this application will be described in detail below with reference to the encoder shown in Figure 12:

[0250] The codec receives image 1201. A description of image 1201 can be found above and will not be repeated here.

[0251] Sub-image partitioning unit 1210 partitions image 1201 into sub-images and outputs N sub-images. N is an integer greater than 0. In this embodiment, each sub-image is encoded and decoded independently. During the encoding process, sub-image 1202 can be referred to as the current sub-image or the sub-image to be encoded.

[0252] Wavelet forward transform unit 1220 performs wavelet forward transform on the current sub-image to obtain wavelet coefficients 1203 (hereinafter referred to as low-frequency sub-band) and wavelet coefficients 1214 (hereinafter referred to as high-frequency sub-band) of the current sub-image. The wavelet coefficients 1203 of the low-frequency sub-band include the wavelet coefficients of the LL sub-band, and the high-frequency sub-band includes the wavelet coefficients of the LH, HL, and HH sub-bands. In this embodiment, each sub-image of the image can be independently encoded and decoded, and the high-frequency sub-band and low-frequency sub-band of each sub-image are independently encoded and decoded. The LH, HL, and HH sub-bands in the high-frequency sub-band can also be independently encoded and decoded.

[0253] The block partitioning unit 1231 (which may be called the low-frequency sub-band block partitioning unit) is used to obtain the wavelet coefficients 1203 of the low-frequency sub-band of the current sub-graph 1202, so as to obtain at least one macroblock 1204 of the low-frequency sub-band of the sub-graph 1202, wherein the macroblock can also be understood as a set of partial coefficients in the wavelet coefficients of the low-frequency sub-band.

[0254] Specifically, the block partitioning unit 1231 partitions the wavelet coefficients 1203 of the low-frequency sub-band of the current subgraph 1202 into blocks based on the block partitioning method, obtaining at least one macroblock 1204 of the low-frequency sub-band of the current subgraph, for example, M macroblocks, where M is an integer greater than 0 (or greater than 1). The low-frequency block partitioning unit 1231 outputs the macroblocks 1204 of the wavelet coefficients 1203 of the current low-frequency sub-band one by one to the residual calculation unit 1232 and the control unit 1238.

[0255] In this embodiment, macroblock 1204 is a basic encoding / decoding unit. During the encoding process, macroblock 1204 can also be referred to as the current block, current image 1201 block, macroblock 1204 to be encoded, block to be encoded, image 1204 block to be encoded, etc.

[0256] Alternatively, the block partitioning method includes, but is not limited to:

[0257] The wavelet coefficients 1203 of both the high-frequency subband and the low-frequency subband are divided into basic coding units of 8*8 macroblocks 1204 (unit is pixels).

[0258] For example, as described above, each subband uses macroblock 1204 as the basic coding unit. The macroblock 1204 currently to be encoded is referred to as the current macroblock 1204. Specifically, the low-frequency subband processing path 1230 encodes each macroblock 1204 of the wavelet coefficients 1203 of the low-frequency subband block by block. For example, encoding and prediction are performed on each macroblock 1204. The following description only uses the encoding process of the current macroblock 1204; the processing processes for other macroblocks are the same, and will not be described in detail here. For example, in encoding, it refers to the macroblock currently being encoded; in decoding, it refers to the macroblock currently being decoded. The decoded macroblock in the reference image used for predicting the current macroblock 1204 is called the reference block (i.e., the low-frequency subband reconstruction block 1209 in the figure). The reference block is the block that provides the reference signal for the current block, where the reference signal represents the pixel value within the macroblock 1204. The block in the reference image that provides the prediction signal for the current block can be called prediction block 1205, where the prediction signal represents the pixel value, sample value, or sample signal within prediction block 1205. For example, after traversing multiple reference blocks, an optimal reference block is found, which will provide the prediction for the current block; this block is called prediction block 1205.

[0259] Specifically, referring to Figure 12, the residual calculation unit 1232 is used to obtain the current macroblock 1204 and the prediction block 1205 (further details of the prediction block 1205 are provided below) to obtain the residual block 1206. Specifically, the residual calculation unit performs residual calculations on the current macroblock 1204 and the prediction block 1205 to obtain the residual block 1206. The residual calculation unit 1232 outputs the residual block 1206 to the transform / quantization unit 1233.

[0260] The transform / quantization unit 1233 is used to acquire the residual block 1206 to obtain the residual coefficients 1207. Specifically, the transform / quantization unit 1233 performs transform and / or quantization processing on the residual block 1206 to obtain the residual coefficients 1207, which can also be called the quantization coefficients of the residual block, or the quantized residual block. The transform / quantization unit 1233 outputs the residual coefficients 1207 to the inverse quantization single / inverse transform unit 1234 and the low-frequency subband entropy coding unit 1239.

[0261] The inverse quantization / inverse transform unit 1234, also known as the inverse quantization / inverse transform unit, is used to obtain the residual coefficients 1207 to obtain the residual reconstruction block 1208. Specifically, the inverse quantization / inverse transform unit 1234 performs inverse quantization and / or inverse transform processing on the residual coefficients 1207 to obtain the residual reconstruction block 1208, which can also be referred to as the inverse quantization coefficients of the residual block, the inverse quantized residual block, etc. The inverse quantization / inverse transform unit 1234 outputs the residual reconstruction block 1208 to the low-frequency subband splicing unit 1236.

[0262] The dequantization / inverse transform unit 1234 may include a dequantization unit and an inverse transform unit (not shown in the figure). The dequantization unit is used to dequantize the input coefficients, and the inverse transform unit is used to inverse transform the input coefficients.

[0263] The low-frequency subband reconstruction unit 1235 is used to obtain a low-frequency subband reconstruction block 1209 based on the prediction block 1205 and the residual reconstruction block 1208. Specifically, the low-frequency subband reconstruction unit 1235 adds the residual reconstruction block 1208 to the prediction block 1205 to obtain the low-frequency subband reconstruction block 1209, which can also be referred to as the reconstructed low-frequency subband macroblock. Optionally, the low-frequency subband reconstruction unit 1235 outputs the low-frequency subband reconstruction block 1209 to the prediction unit 1237 and the low-frequency subband splicing unit 1236. Optionally, the low-frequency subband reconstruction unit 1235 outputs the low-frequency subband reconstruction block 1209 to the control unit 1238.

[0264] The low-frequency subband stitching unit 1236 is used to obtain the reconstructed low-frequency subband 1211, which can also be referred to as the reconstructed value or reconstructed data of the low-frequency subband, based on the low-frequency subband reconstruction block 1209. Optionally, the low-frequency subband stitching unit 1236 outputs the reconstructed low-frequency subband 1211 to the prediction unit 1237. Optionally, the low-frequency subband stitching unit 1236 outputs the reconstructed low-frequency subband 1211 to the control unit 1238.

[0265] Specifically, as described above, the low-frequency subband uses macroblocks as the basic coding unit, and the low-frequency subband splicing unit 1236 can obtain M low-frequency subband reconstruction blocks of a low-frequency subband. The low-frequency subband splicing unit 1236 can reconstruct the corresponding low-frequency subband based on the M low-frequency subband reconstruction blocks, that is, obtain the reconstructed low-frequency subband 1211.

[0266] The control unit 1238, also known as the mode selection unit, is used to determine the syntax element 1213 based on the macroblock 1204 (i.e. the current block); or to determine the syntax element 1213 based on the current macroblock 1204, the low-frequency subband reconstruction block 1209, and the reconstructed low-frequency subband 1211.

[0267] Syntax element 1213 includes at least one syntax element, such as pattern information. Pattern information, also known as prediction mode information, is used to indicate the prediction mode (or prediction method) of prediction unit 1237, such as inter-frame or intra-frame prediction mode. Control unit 1238 can output syntax element 1213 to prediction unit 1237 and low-frequency subband entropy coding unit 1239.

[0268] Prediction unit 1237, also known as prediction processing unit, is used to acquire syntax element 1213 and perform prediction processing based on syntax element 1213. Specifically, prediction unit 1237 may select a prediction mode based on syntax element 1213 (e.g., mode information in the syntax element). In one example, prediction unit 1237 may acquire low-frequency subband reconstruction block 1209 based on syntax element 1213 to acquire prediction block 1205. Specifically, prediction unit 1237 may perform intra-frame prediction based on low-frequency subband reconstruction block 1209 to acquire prediction block 1205. In another example, prediction unit 1237 may acquire reconstructed low-frequency subband 1211 based on syntax element 1213 to acquire prediction block 1205.

[0269] The prediction unit 1237 outputs prediction block 1205 to the residual calculation unit 1232 and the low-frequency sub-band splicing unit 1236.

[0270] The low-frequency subband entropy coding unit 1239 is used to obtain low-frequency subband encoded data 1213, also known as encoded low-frequency subband, based on residual coefficients 1207 and syntax elements 1213. Specifically, the low-frequency subband entropy coding unit 1239 uses an entropy coding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC (CAVLC) scheme, arithmetic coding scheme, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to entropy encode the residual coefficients 1207 and syntax elements 1213 to obtain low-frequency subband encoded data 1213 output in the form of, for example, an encoded bitstream.

[0271] Referring again to Figure 12, the block partitioning unit 1241, also known as the high-frequency subband block partitioning unit, is used to obtain the high-frequency subband of the current subgraph 1202 to obtain at least one macroblock 1215 of the high-frequency subband of the subgraph 1202. For a detailed description, please refer to the low-frequency subband section; it will not be repeated here. Specifically, the block partitioning unit 1241 partitions the high-frequency subband 1214 of the current subgraph 1202 (hereinafter referred to as the current high-frequency subband) into blocks based on the block partitioning method, obtaining at least one macroblock 1215 of the current subgraph 1202, for example, M macroblocks, where M is an integer greater than 0 (or an integer greater than 1). Other undescribed parts can be referred to the relevant description of the block partitioning unit 1231; it will not be repeated here.

[0272] Block partitioning unit 1241 outputs the macroblocks of the current high-frequency subband one by one to quantization / conversion unit 1242.

[0273] The transform / quantization unit 1242 is used to transform and / or quantize the macroblock 1215 to obtain the quantization coefficients 1216 of the high-frequency subband block (i.e., the quantization coefficients of the current macroblock). The transform / quantization unit 1242 outputs the quantization coefficients 1216 of the high-frequency subband block to the high-frequency subband entropy coding unit 1243.

[0274] The high-frequency subband entropy coding unit 1243 is used to perform entropy coding on the data to be encoded to obtain high-frequency subband encoded data 1217. The data to be encoded may include, but is not limited to, the quantization coefficients and syntax elements of each high-frequency subband block. The high-frequency subband encoded data 1217 includes, but is not limited to, HH subband encoded data, HL subband encoded data, and LH subband encoded data.

[0275] The encoding and decoding method provided in this application supports two scenarios: full I-frame configuration and I / P frame alternating encoding configuration. The encoder architecture shown in Figure 12 adds relevant modules required for the prediction process based on the wavelet transform architecture shown in Figure 8, which can improve the compression efficiency of I / P frame alternating encoding for scenarios such as fixed camera positions and slow camera movement.

[0276] Optionally, the encoder may also include a combination unit (not shown in the figure) for generating an image bitstream based on low-frequency subband coded data and high-frequency subband coded data. The specific process can be referred to the relevant description in Figure 8 above.

[0277] The bitstream output by the encoder in the embodiments of this application will be described in detail below. The bitstream structure described below can be applied to the encoders shown in Figures 8 and 12, and of course, it can also be applied to other encoder variations based on Figures 8 or 12.

[0278] For example, as described above, the low-frequency subband entropy coding unit 1239 and the high-frequency subband entropy coding unit 1243 output low-frequency subband coded data 1213 and high-frequency subband coded data 1217, respectively. This can be understood as the encoder 120 independently encoding the low-frequency and high-frequency subbands of each subgraph, outputting low-frequency and high-frequency subband coded data corresponding to the current subgraph.

[0279] Referring to Figure 13A, which is an exemplary schematic diagram of an image bitstream structure, the image bitstream in the example of Figure 13A includes, but is not limited to, image header information and image data (also referred to as an image data region).

[0280] For example, the image data includes at least one image data region (also referred to as an image data sub-region), such as, but not limited to, a first image data region and a second image data region. During the encoding process, the encoder (e.g., through a combination unit) writes high-frequency subband encoded data and low-frequency subband encoded data into the image bitstream. Specifically, the encoder writes high-frequency subband encoded data into the first image data region and low-frequency subband encoded data into the second image data region.

[0281] For example, image header information includes, but is not limited to, offset information and image size information.

[0282] For example, image size information is used to indicate the size of the original image. As mentioned above, during the encoding process, some sub-images may be padded during sub-image partitioning to ensure that the length and width of each sub-image are multiples of 16. Thus, during decoding, the size of the reconstructed image obtained by the decoder may be larger than the original image size. The decoder can process the reconstructed image based on the image size information to remove the padded portions.

[0283] For example, offset information is used to indicate the position of a data region in the image bitstream, and can also be understood as indicating the position of independently decodeable coded data in the image bitstream. When decoding coded data (i.e., the image bitstream) according to this application, the offset information in the image header information can be used to obtain independently decodeable coded data, and decoding operations can be performed on the coded data. The independently decodeable coded data (e.g., low-frequency subband coded data and high-frequency subband coded data) can be decoded synchronously during decoding to improve decoding efficiency.

[0284] In one example, the offset information can be the length of the image data region containing adjacent, independently decodeable encoded data in the image bitstream.

[0285] In another example, the offset information can be the offset (i.e., the difference) between the starting position of the image data region where the independently decoded encoded data is located and the ending position of the image header information.

[0286] Referring to Figure 13B, which is an exemplary schematic diagram of an image bitstream structure, the bitstream in the example of Figure 13B includes, but is not limited to, image header information and image data. The image header information includes, but is not limited to, offset information, etc., detailed concepts of which can be found above and will not be repeated here. The image data includes, but is not limited to, low-frequency subband coded data and high-frequency subband coded data. The low-frequency subband coded data is written to the first image data area, and the high-frequency subband coded data is written to the second image data area. Other descriptions can be found in Figure 13A and will not be repeated here.

[0287] It should be noted that the embodiments in this application only use the image bitstream of a single image as an example for illustration, that is, the bitstream includes only one image data. In the process of encoding video images, the encoder can generate an encoded image bitstream for each image, that is, the image bitstream includes multiple image data, and each image data carries the encoded data of the corresponding image.

[0288] Referring to Figure 14A, which is an exemplary schematic diagram of an image bitstream structure, in the example of Figure 14A, as described above, the high-frequency subband coding data of each sub-image further includes: HH subband coding data, HL subband coding data, and LH subband coding data. Accordingly, in this example, the image data includes, but is not limited to: HH subband coding data, HL subband coding data, and LH subband coding data of each sub-image of the image.

[0289] For example, in the example shown in Figure 14A, LL subband coded data is written to the second image data region, HL subband coded data is written to the first image data subregion, HH subband coded data is written to the second image data subregion, and LH subband coded data is written to the dotted image data subregion. The writing order of each high-frequency subband is only an illustrative example and will not be repeated below.

[0290] Optionally, in the example shown in Figure 14A, the offset information of each segment of independently decodeable encoded data in the image header information may include, but is not limited to: the offset of the starting position of the first image data sub-region relative to the ending position of the image header information (this data is usually measured in n bytes), the offset of the starting position of the second image data sub-region relative to the ending position of the image header information, and the offset of the starting position of the third image data sub-region relative to the ending position of the image header information. After obtaining these offsets by decoding the image header information, the starting position of each segment of independently decodeable encoded data can be obtained, so that the LL subband encoded data, HH subband encoded data, HL subband encoded data, and LH subband encoded data can be decoded independently during decoding.

[0291] Optionally, in the example shown in Figure 14A, the offset information may include, but is not limited to: the length information L0 of the first image data region (this length is typically measured in n bytes), the length information L1 of the first image data sub-region, and the length information L2 of the second image data sub-region. L0 is the offset of the starting position of the first image data sub-region relative to the ending position of the image header information. L0 + L1 yields the offset of the starting position of the second image data sub-region relative to the ending position of the image header information. L0 + L1 + L2 yields the offset of the starting position of the third image data sub-region relative to the ending position of the image header information. After obtaining the offset information by decoding the image header information, the starting position of each independently decodeable segment of encoded data can be obtained by addition calculation, so that the LL subband encoded data, HH subband encoded data, HL subband encoded data, and LH subband encoded data can be decoded independently during decoding.

[0292] Referring to Figure 14B, which is an exemplary schematic diagram of an image bitstream structure, in the example of Figure 14B, low-frequency subband coded data is written to the first image data region, and high-frequency subband coded data is written to the second image data region. Further details can be found in Figure 14A, and will not be repeated here.

[0293] Referring to Figure 15, which is an exemplary schematic diagram of an image bitstream structure, in the example of Figure 15, the high-frequency subband encoded data is written to the corresponding image data region according to the type of subband. Specifically, the description of the first image data region and its sub-regions, and the second image data region, can be referred to above and will not be repeated here. Taking HL subband encoded data as an example, the HL subbands of each sub-image (e.g., sub-image 1 to sub-image N) of the image are written to the first image data sub-region; the order is only illustrative.

[0294] In this context, for example, sub-image 1-HL in the figure represents the HL subband encoded data of sub-image 1 of the image. The HL subband encoded data of each sub-image further includes the encoded macroblocks of the HL subband of that sub-image. For example, sub-image 1-HL includes, but is not limited to: sub-image 1-HL-MB0 to sub-image 1-HL-MBm. Sub-image 1-HL-MBx represents the encoded macroblock MBx in the HL subband of sub-image 1. The HH subband encoded data and LH subband encoded data are similar to the HL subband encoded data and will not be described further here.

[0295] Referring again to Figure 15, the LL subband coding data includes, but is not limited to, the LL subband coding data of each sub-image of the image. For example, sub-image 1-LL to sub-image N-LL. Sub-image 1-LL represents the LL subband coding data of sub-image 1 of the image. Each sub-image-LL further includes, but is not limited to, the coding data of each macroblock of that sub-image, i.e., the coded macroblock. For example, sub-image 1-LL includes, but is not limited to, sub-image 1-LL-MB0 to sub-image 1-LL-MBn. Sub-image 1-LL-MBx represents the coding data of macroblock MBx of the LL subband of sub-image 1 of the image.

[0296] For example, during the encoding process, each sub-band of each subgraph is encoded independently; correspondingly, during the decoding process, each encoded sub-band of each subgraph can be decoded independently. In the example shown in Figure 15, it can be understood that the four encoded sub-bands of each subgraph can be decoded independently. For instance, during decoding, the decoder can obtain the HL sub-band encoded data, LH sub-band encoded data, HH sub-band encoded data, and LL sub-band encoded data of subgraph 1 based on the offset information, and perform independent decoding to obtain the decoded subgraph 1.

[0297] In this example, the coded subbands and their coded macroblocks of each sub-image in the figure can also be written into the corresponding image data sub-regions, which are not shown in the figure and will not be repeated below. Accordingly, offset information can be used to indicate the position of the image region to which the independently coded subbands of each sub-image belong. Thus, during decoding, the decoder can obtain the positions of the four coded subbands (LL subband coded data, HH subband coded data, HL subband coded data, and LH subband coded data) of a single sub-image in the image data based on the offset information, and perform decoding on them to obtain the decoded subbands of the corresponding sub-image (including LL subband decoded data, HH subband decoded data, HL subband decoded data, and LH subband decoded data). For example, during decoding, the decoder obtains sub-image 1-HL (including each coded macroblock contained in its sub-region, the same below, and will not be repeated), sub-image 1-HH, sub-image 1-LH, and sub-image 1-LL. In this way, the decoder can obtain the decoded sub-graph 1 based on the above encoded sub-bands without having to decode other sub-graphs one by one according to the order of the bitstream. This can improve decoding efficiency while reducing the occupation of the decoding buffer (e.g., DPB) and reducing the hardware processing pressure and storage burden on the decoding end.

[0298] In one possible implementation, on the encoding side, the low-frequency subband entropy encoding unit can output low-frequency subband encoded data (e.g., LL subband bitstream), and the high-frequency subband entropy encoding unit can output high-frequency subband encoded data (e.g., including HH subband decoded data, HL subband decoded data, and LH subband decoded data). The combining unit writes this data into the image data to generate an image bitstream.

[0299] In another possible implementation, the low-frequency subband entropy coding unit and the high-frequency subband entropy coding unit can also output the coded macroblock of the coded subband to the combining unit after each macroblock is encoded. The combining unit can generate the bitstream structure as shown in the figures (e.g., Figures 15, 16A-16C, 17A-17B, etc.) according to the specified order of the coded macroblocks of the image data.

[0300] Optionally, the subgraph and bitstream structure in Figure 15 can also be applied to the bitstream structures shown in Figures 13B and 14B, and will not be illustrated in detail here.

[0301] In the embodiments of this application, the subbands in the high-frequency subband coded data of the image data can be interleaved and sorted. The interleaving and sorting can be at the sub-image granularity, the subband type granularity, or the macroblock (MB) granularity. Several interleaving and sorting methods are provided below. It should be noted that the interleaving methods shown in the embodiments of this application are only illustrative examples. In other embodiments, other interleaving methods can also be set according to the encoding and decoding requirements.

[0302] Referring to Figure 16A, which is an exemplary schematic diagram of an image bitstream structure, the interleaving method in the example of Figure 16A is interleaving at the subband type of the subimage. Specifically, as shown in Figure 16A, the HL subband encoded data, encoded HH, and encoded LH of each subimage of the image are continuously written into the image bitstream. The subimage order and macroblock order shown in the figure are merely illustrative examples. The format of the low-frequency subband is not shown in the figure; please refer to Figure 15 and its related description, which will not be repeated here.

[0303] For example, as shown in Figure 16A, sub-images 1-HL, 1-HH, and 1-LH are continuously written into the image bitstream. Sub-image 1-HL represents the HL subband encoded data of sub-image 1, which includes, but is not limited to, all encoded macroblocks (MB) of the HL subband of sub-image 1, such as MB0-HL to MBm-HL. Only the bitstream format of sub-image 1 is shown in the figure; other sub-images are similar and will not be described individually here. The LL subband encoded data can be referred to above and will not be repeated here.

[0304] In this example, similar to the description in Figure 15, each coded subband of each sub-image can be decoded independently. Accordingly, offset information can be used to indicate the location of the image region to which the independently coded subband of each sub-image belongs. This allows, during decoding, the positions of the four coded subbands (LL subband coded data, HH subband coded data, HL subband coded data, and LH subband coded data) of a single sub-image can be obtained in the image data based on the offset information, and decoding can be performed on them to obtain the decoded subbands of the corresponding sub-image (including LL subband decoded data, HH subband decoded data, HL subband decoded data, and LH subband decoded data). Details not described herein can be found in Figure 15 and will not be repeated here.

[0305] Referring to Figure 16B, which is an exemplary schematic diagram of the image bitstream structure, the interleaving method in the example of Figure 16B is interleaving at the macroblock level for each sub-image.

[0306] Specifically, in the example shown in Figure 16B, the high-frequency subband encoded data (including LH subband encoded data, HH subband encoded data, and HL subband encoded data) of each sub-image of the image are continuously written into the first image data region. A description of the second image data region can be found in Figure 15, and will not be repeated here.

[0307] For example, as shown in Figure 16B, sub-images 1-HL-MB0, 1-HH-MB0, and 1-LH-MB0 are consecutively written into the first image data region. Here, 1-HL-MB0 represents the encoded macroblock MB0 of the HL subband of 1, 1-HH-MB0 represents the encoded macroblock MB0 of the HH subband of 1, and 1-LH-MB0 represents the encoded macroblock MB0 of the LH subband of 1. The figure only shows the encoded data structure of 1 in the bitstream; the other sub-images are similar and will not be illustrated individually here.

[0308] In this example, during decoding, the decoding end can decode the high-frequency subband encoded data according to the sub-image order, that is, each sub-image in the first image data region is decoded independently. The low-frequency subband encoded data is also decoded according to the sub-image order, that is, each sub-image in the second image data region is decoded independently. Unlike Figures 15, 16A, and 16C, when the bitstream structure in the above figures is applied to decoding, the decoding end decodes the different types of encoded subbands of each sub-image separately. This can also be understood as needing to simultaneously decode the four sub-bitstreams (each sub-bitstream corresponds to one type of encoded subband) of each sub-image to obtain four decoded subbands of an image at the output end, thereby obtaining the decoded sub-image and storing additional data. For the bitstream structure shown in Figure 16B, the decoding end only needs to decode the high-frequency subband and low-frequency subband of each sub-image, that is, two bitstreams. For example, as shown in Figure 16B, when the decoding end decodes the first image data region, it can decode each encoded macroblock one by one according to the encoded macroblock order of each sub-image in the region. That is, the three high-frequency subband encoded data of sub-Figure 1 are written continuously into the first image data area. Therefore, during decoding, the three high-frequency subband encoded data of sub-Figure 1 can be decoded one by one to obtain the decoded high-frequency subband. The processing of LL subband encoded data can be referred to Figure 15, and will not be described in detail here.

[0309] In the example shown in Figure 16B, the offset information is used to indicate the position of the first image data region and the position of the second image data region. That is, the example shown in Figure 16B includes two sub-bitstreams that can be decoded independently. Compared with the framework and decoding methods in Figures 15, 16A, and 16C, it has lower hardware performance requirements, only requiring simultaneous decoding of the two sub-bitstreams of the sub-image (corresponding to low-frequency sub-band encoded data and high-frequency sub-band encoded data).

[0310] Referring to Figure 16C, which is an exemplary schematic diagram of an image bitstream structure, the interleaving method in Figure 16C is based on the sub-bands of each sub-image. Specifically, as shown in Figure 16C, each sub-image of the image is continuously written into the first image data region according to its high-frequency sub-band type. For example, taking sub-image 1 as an example, the coded macroblocks of the HL sub-band encoded data of sub-image 1 are continuously written into the first image data region, that is, sub-image 1-HL-MB0 to sub-image 1-HL-MBm of sub-image 1 are continuously written into the image bitstream. The coded macroblocks of the HH sub-band encoded data of sub-image 1 are continuously written into the first image data region, that is, sub-image 1-HH-MB0 to sub-image 1-HH-MBm of sub-image 1 are continuously written into the image bitstream. The coded macroblocks of the LH sub-band encoded data of sub-image 1 are continuously written into the first image data region, that is, sub-image 1-LH-MB0 to sub-image 1-LH-MBm of sub-image 1 are continuously written into the image bitstream. Other subgraphs are similar and will not be described in detail here. The order of subband types is for illustrative purposes only.

[0311] In this example, similar to the description in Figure 15, each coded subband of each sub-image can be decoded independently. Accordingly, offset information can be used to indicate the location of the image region to which the independently coded subband of each sub-image belongs. This allows, during decoding, the positions of the four coded subbands (LL subband encoded data, HH subband encoded data, HL subband encoded data, and LH subband encoded data) of a single sub-image can be obtained in the image data based on the offset information, and decoding can be performed on them to obtain the decoded subbands of the corresponding sub-image (including LL decoded data, HH decoded data, HL decoded data, and LH decoded data). Parts not described herein can be referred to Figure 15 and will not be repeated here.

[0312] For example, in the example of Figure 16C, during decoding, the decoder obtains the HH subband encoded data, LH subband encoded data, HL subband encoded data, and LL subband encoded data of sub-Figure 1 from the image bitstream based on offset information (not shown in Figure 16C, see Figure 15). The decoder decodes the encoded data of the four subbands of sub-Figure 1 to obtain the reconstructed sub-Figure 1, which can also be called decoded sub-Figure 1 or decoded sub-Figure 1.

[0313] In the embodiments of this application, multiple independently decoded encoded data can be decoded simultaneously, or one or more high-frequency subbands can be decoded simultaneously, and the number of simultaneous decodes depends on the decoder hardware performance.

[0314] Referring to Figure 17A, which is an exemplary schematic diagram of an image bitstream structure, the interleaving method in the example of Figure 17A is granular at the low-frequency and high-frequency subbands of the subgraph. Only the interleaving method of subgraph 1 is shown; other subgraphs are similar and will not be described in detail here. For example, the structure of the low-frequency subband encoded data of the subgraph can be seen in Figure 15, and will not be repeated here. In one example, the structure of the high-frequency subband encoded data of the subgraph can use any of the interleaving methods for high-frequency subband encoded data in Figures 16A to 16C.

[0315] Referring to Figure 17B, which is an exemplary schematic diagram of the bitstream structure, in the example of Figure 17B, the splicing order of the sub-graphs can be low-frequency subband coded data followed by high-frequency subband coded data. Other descriptions can be found in Figure 17A, and will not be repeated here.

[0316] Optionally, for the examples shown in Figures 17A and 17B, the offset information in the image header information is also used to indicate each encoded data that can be decoded independently. The indication method can be seen in Figures 15 and 16A to 16C, which will not be repeated here.

[0317] Referring to Figure 18, which is a schematic / conceptual block diagram of an exemplary decoder, in the example of Figure 18, the decoder receives, for example, an image bitstream encoded by an encoder to obtain a decoded image of the original image, also referred to as a decoded image, reconstructed image, etc. During the decoding process, the decoder receives the image bitstream from the encoder, including, but not limited to, image header information and image data. The image bitstream can be any of the bitstream formats shown in Figures 13A to 17B.

[0318] In the example shown in Figure 18, the decoder includes, but is not limited to: low-frequency subband processing path 1810, high-frequency subband processing path 1820, wavelet inverse transform unit 1830, image combination unit 1840 (also known as image stitching unit), etc.

[0319] For example, the low-frequency subband processing path 1810 is used to acquire low-frequency subband encoded data to obtain reconstructed low-frequency subband 1806 (also known as decoded low-frequency subband). The low-frequency subband processing path includes, but is not limited to: low-frequency subband entropy decoding unit 1811, inverse quantization / inverse transform unit 1812 (also known as low-frequency subband inverse quantization / inverse transform unit), low-frequency subband reconstruction unit 1813, low-frequency subband splicing unit 1815, prediction unit 1814, etc.

[0320] The high-frequency subband processing path 1820 is used to acquire high-frequency subband encoded data to obtain the reconstructed high-frequency subband 1831, which can also be called the reconstructed value of the high-frequency subband or the reconstructed data of the high-frequency subband, including but not limited to: high-frequency subband entropy decoding unit 1821, inverse quantization / inverse transform unit 1821 (also called high-frequency subband inverse quantization / inverse transform unit), high-frequency subband reconstruction unit 1831, etc.

[0321] In some instances, the decoder shown in Figure 18 can perform a decoding process that is largely the reverse of the encoding process described with reference to the encoder in Figure 12.

[0322] The decoding method in the embodiments of this application will be described in detail below with reference to the decoder 180 shown in Figure 18.

[0323] For example, the decoder 180 can obtain high-frequency subband encoded data and low-frequency subband encoded data in the image bitstream based on the image header information in the image bitstream. Furthermore, as described above, during the encoding process, the encoder uses macroblocks as the basic encoding unit, and correspondingly, during the decoding process, the decoder also uses macroblocks (e.g., encoded macroblocks) as the basic decoding unit for decoding.

[0324] For example, the low-frequency subband entropy decoding unit 1811 performs entropy decoding on the low-frequency subband encoded data 1801 in the image bitstream, using macroblocks as the basic decoding unit, to obtain the quantization coefficients 1802 (i.e., the quantization coefficients of the current macroblock) and syntax elements 1807 of the low-frequency subband block. The description of the quantization coefficients 1802 of the low-frequency subband can be found on the encoder side and will not be repeated here. Specifically, the low-frequency subband entropy decoding unit 1811 obtains the encoded macroblocks (i.e., the encoded data of the macroblocks) of the low-frequency subbands (e.g., LL subbands) of each sub-image in the image bitstream, and performs entropy decoding on each encoded macroblock to obtain the quantization coefficients 1802 (which can be simply referred to as the quantization coefficients of the macroblock of the low-frequency subband) and syntax elements 1807 of the corresponding low-frequency subband for each encoded macroblock. During the decoding process, the currently decoded encoded macroblock can be called the current block.

[0325] The low-frequency subband decoding unit is used to output the quantization coefficients 1802 of the low-frequency subband block to the inverse quantization / inverse transform unit 1812, and to output the syntax elements 1807 to the prediction unit 1814.

[0326] The inverse quantization / inverse transform unit 1812 is used to obtain the quantization coefficients 1802 of the low-frequency subband block to obtain the inverse quantization coefficients 1803 of the low-frequency subband block. Alternatively, it can be the inverse transform coefficients of the current block of the low-frequency subband (depending on whether inverse transform processing was performed). Specifically, the inverse quantization / inverse transform unit 1812 performs inverse quantization on the quantization coefficients of the current block of the low-frequency subband, or performs both inverse quantization and inverse transform, to obtain the inverse quantization coefficients of the current block of the low-frequency subband. The inverse quantization / inverse transform unit 1812 outputs the inverse quantization coefficients 1803 of the low-frequency subband block to the low-frequency subband reconstruction unit 1813, for example, the inverse quantization coefficients of the current block of the low-frequency subband.

[0327] The low-frequency subband reconstruction unit 1813 is used to obtain the low-frequency subband reconstruction block 1804, which can also be called the reconstruction coefficient of the low-frequency subband block, based on the quantization coefficients 1803 and prediction block 1805 of the low-frequency subband. Specifically, the low-frequency subband reconstruction unit 1813 adds a prediction block to the inverse quantization coefficients of the current block of the low-frequency subband to obtain the low-frequency subband reconstruction block 1804 corresponding to the current macroblock.

[0328] The prediction unit 1814 is used to obtain syntax element 1807 and perform corresponding prediction processing according to syntax element 1807. For example, intra-frame prediction can be performed based on low-frequency subband reconstruction block 1804, or inter-frame prediction can be performed based on reconstructed low-frequency subband 1806. Its execution method can be referred to the coding side, and will not be elaborated here. The prediction unit 1814 outputs prediction block 1805 to the low-frequency subband reconstruction block 1804 unit.

[0329] For example, the high-frequency subband entropy decoding unit 1821 acquires the high-frequency subband encoded data 1802 in the image bitstream, and, using macroblocks as the basic decoding unit, acquires the quantization coefficients 1808 (which are the quantization coefficients of the current macroblock) of each high-frequency subband block. Specifically, the high-frequency subband entropy decoding unit 1821 performs entropy decoding on the current block of the high-frequency subband encoded data 1802 to obtain the quantization coefficients of the current block of the high-frequency subband. Optionally, based on entropy decoding, the syntax elements corresponding to the current block can also be obtained. The high-frequency subband entropy decoding unit 1821 outputs the quantization coefficients 1808 of the high-frequency subband block to the inverse quantization / inverse transform unit 1822.

[0330] The inverse quantization / inverse transform unit 1822, also known as the high-frequency subband inverse quantization / inverse transform unit, is used to obtain the quantization coefficients 1808 of the high-frequency subband block to obtain the reconstruction coefficients 1809 of the high-frequency subband block. The reconstruction coefficients can be either inverse quantization coefficients after inverse quantization processing, or inverse transform coefficients after inverse quantization and inverse transform processing.

[0331] The high-frequency subband reconstruction unit 1823 (also known as the high-frequency subband splicing unit) is used to obtain the reconstruction coefficients 1809 of the high-frequency subband block to obtain the reconstructed high-frequency subband 1831, which can also be referred to as the reconstructed value or reconstructed data of the high-frequency subband. Specifically, the high-frequency subband reconstruction unit 1823 can obtain the reconstruction coefficients corresponding to each macroblock of the high-frequency subband, that is, reconstruct the high-frequency subband block. The high-frequency subband reconstruction unit 1823 can splice the obtained multiple macroblocks to obtain the corresponding high-frequency subband. Among them, the reconstructed high-frequency subband may optionally include reconstructing the HL subband (e.g., the reconstruction coefficients of the HL subband), reconstructing the HH subband (e.g., the reconstruction coefficients of the HH subband), and reconstructing the LH subband (e.g., the reconstruction coefficients of the LH subband).

[0332] The inverse wavelet transform unit 1830 is used to acquire the reconstructed high-frequency subband 1831 and the reconstructed low-frequency subband 1806 to obtain the reconstructed sub-image 1832. Specifically, the inverse wavelet transform unit 1830 acquires the reconstructed low-frequency subband 1806 output by the low-frequency subband stitching unit 1815 and the reconstructed high-frequency subband 1831 output by the high-frequency subband reconstruction unit 1823, and performs an inverse wavelet transform on the reconstructed low-frequency subband 1806 and the reconstructed high-frequency subband 1831 to obtain the reconstructed sub-image 1832. The inverse wavelet transform unit 1830 outputs the reconstructed sub-image 1832 to the image combining unit (which may also be called the image stitching unit, etc.).

[0333] Image combining unit 1840 is used to acquire reconstructed sub-images 1832 to obtain a reconstructed image 1833 of the original image, which can also be called a decoded image or a decoded image, etc. Specifically, image combining unit 1840 can acquire N reconstructed sub-images (N is an integer greater than 0) of the image (referring to the original image), and stitch (or combine) the N reconstructed sub-images according to the division method (including size and position) of each reconstructed sub-image during encoding to obtain the reconstructed image 1833.

[0334] Optionally, after obtaining the reconstructed image, the image combining unit 1840 can determine whether the reconstructed image contains a padding portion based on the image size information in the image header information and the size information of the current reconstructed image. In one example, if the size of the current reconstructed image is the same as the size indicated by the image size information (i.e., the same as the original image size), the image combining unit 1840 can send the reconstructed image to the display device. In this case, the sizes of the displayed image, the original image, and the reconstructed image are all the same. In another example, if the size of the current reconstructed image is different from the size indicated by the image size information (e.g., larger than the original image size), the image combining unit 1840 can remove the padding portion of the current reconstructed image based on the size indicated by the image size information to obtain the displayed image. The size of the displayed image is the same as the size of the original image. Specifically, in this embodiment, when the encoding side performs sub-image division, the division order is preset, usually from left to right in the horizontal direction and from top to bottom in the vertical direction. Correspondingly, at least one padding sub-image is usually located at the edge of the image, as shown in Figure 9B. For example, the image combining unit 1840 may crop the vertical and / or horizontal edges of the image according to the size information to remove the fill portion of the sub-image at the edge.

[0335] Optionally, in some instances, the image reconstruction unit may also perform the above-mentioned operation of removing the padding portion during the process of acquiring the reconstructed image, so that the size of the reconstructed image is the same as the size of the original image.

[0336] Optionally, the decoder is used, for example, to output a reconstructed image via the decoder's output port (or output interface) for presentation to or viewing by the user.

[0337] Other variations of the decoder can be used to decode compressed image bitstreams.

[0338] For example, as described above, each sub-image in the image bitstream may include two or four independently decodeable encoded data. For instance, in the examples shown in Figures 15, 16A, and 16C, the encoded HH sub-band, encoded LH sub-band, encoded HL sub-band, and encoded LL sub-band of each sub-image can be independently decoded. In the above scenario, the high-frequency sub-band entropy decoding unit can obtain each independently decodeable image data region based on the offset information in the image header information, and obtain the encoded data within the image data region. The high-frequency sub-band entropy decoding unit can simultaneously decode the encoded data of one or more independently decoded image data regions.

[0339] Optionally, in the above scenario, the high-frequency subband processing path may include at least one high-frequency processing subpath (not shown in the figure). For example, the high-frequency subband processing path may include three high-frequency processing subpaths to process three coded subbands of a subgraph simultaneously. Of course, in some instances, there may be more than three or fewer high-frequency subband processing subpaths. The more high-frequency subband processing paths there are, the higher the decoding efficiency. The fewer the paths, the lower the hardware design complexity requirements.

[0340] The decoding method is illustrated below using the bitstream structure shown in Figure 16A. In the example shown in Figure 16A, the low-frequency entropy decoding unit 1801 acquires the LL subband encoded data and performs decoding based on macroblocks. Taking sub-Figure 1 as an example, the low-frequency entropy decoding unit 1801 acquires the LL subband encoded data of sub-Figure 1 and decodes each encoded macroblock. The low-frequency subband decoding path 1810 processes the current block (inverse quantization / inverse transform, prediction, reconstruction, etc.) to output the reconstructed low-frequency subband of sub-Figure 1 to the wavelet inverse transform unit.

[0341] The high-frequency subband entropy decoding unit 1821 acquires LH subband encoded data, HL subband encoded data, and HH subband encoded data, and performs decoding based on the encoded macroblocks. Taking sub-image 1 as an example, specifically, the high-frequency entropy decoding unit 1821 acquires the LH subband encoded data, HL subband encoded data, and HH subband encoded data of sub-image 1 in the image bitstream based on offset information. The acquisition order can be in the order of the bitstream or according to actual needs. The high-frequency subband processing path 1820 processes the decoded macroblocks of sub-image 1 output by the high-frequency subband entropy decoding unit. In one example, the high-frequency subband entropy decoding unit can optionally decode the LH subband encoded data, HL subband encoded data, and HH subband encoded data of sub-Figure 1 simultaneously. The high-frequency subband processing path obtains multiple decoded high-frequency subbands output by the high-frequency subband entropy decoding unit and processes each macroblock in the multiple decoded high-frequency subbands simultaneously to output the reconstructed HH subband, reconstructed HL subband, and reconstructed LH subband of sub-Figure 1 to the wavelet inverse transform unit.

[0342] In another example, the high-frequency subband entropy decoding unit can optionally decode the LH, HL, and HH subband encoded data of sub-Figure 1 simultaneously. The high-frequency subband processing path acquires multiple decoded high-frequency subbands output by the high-frequency subband entropy encoding unit. It can process each macroblock in at least one decoded high-frequency subband, and the decoded data of other received but unprocessed high-frequency subbands can be cached in storage. The processing order can follow the order in the bitstream or be set according to actual needs. Similarly, the high-frequency subband processing path outputs to the wavelet inverse transform unit after acquiring a reconstructed high-frequency subband of sub-Figure 1.

[0343] After the wavelet inverse transform unit obtains the four reconstructed high-frequency subbands of sub-figure 1, including the reconstructed LL subband, reconstructed HL subband, reconstructed HH subband, and reconstructed LH subband, it performs wavelet inverse transform to obtain the reconstructed sub-figure 1.

[0344] Optionally, before acquiring the four reconstructed high-frequency subbands of sub-graph 1, the wavelet inverse transform unit can cache each reconstructed subband acquired. After acquiring all the reconstructed high-frequency subbands of sub-graph 1, it can acquire the cached reconstructed high-frequency subbands of sub-graph 1 and perform the wavelet inverse transform.

[0345] It can be understood that, in the embodiments of this application, when the high-frequency subband processing path and the low-frequency subband processing path process the encoded data of the sub-graph, the order of the sub-graphs and their macroblocks in each independent decoded data is the same. For example, as shown in Figure 15, the high-frequency subband entropy decoding unit can obtain the HH subband encoded data, HL subband encoded data, and LH subband encoded data of a single sub-graph (e.g., sub-graph 1) from the offset. In this way, when decoding, the high-frequency subband entropy decoding unit can obtain each encoded subband and its encoded macroblock of sub-graph 1 to decode sub-graph 1. Correspondingly, the high-frequency subband processing path can process the decoded macroblocks corresponding to the high-frequency subbands of sub-graph 1 to obtain each reconstructed high-frequency subband of sub-graph 1 and output it to the wavelet inverse transform unit so that the wavelet inverse transform unit can output the reconstructed sub-graph 1. In this way, the wavelet inverse transform unit only needs to buffer the decoded data of sub-graph 1 during the processing. If the high-frequency processing unit processes the encoded data in the code stream as shown in Figure 15, the wavelet inverse transform unit will cache the reconstructed HL subbands of other sub-graphs before obtaining the reconstructed HH subband of sub-graph 1, which increases the storage burden and requires a large hardware cache space, affecting the complexity of hardware design.

[0346] For example, in the example shown in Figure 16B, the high-frequency subband encoded data in the image bitstream is interleaved at the MB granularity. That is, during decoding, the high-frequency subband encoded data and low-frequency subband encoded data of each sub-image can be decoded independently. The high-frequency subband processing path can process each macroblock according to the order of the encoded macroblocks of the sub-image in the image bitstream, that is, the reconstructed subbands of each sub-image are obtained according to the order of the sub-images, which can reduce the hardware design complexity of the decoding end.

[0347] Referring to Figure 19, which is an exemplary schematic / conceptual block diagram of a decoder, the decoder 190 includes, but is not limited to: a low-frequency subband processing path 1910, a high-frequency subband processing path 1920, a wavelet inverse transform unit 1930, a sub-image combination unit 1940, and an image combination unit 1950. The low-frequency subband processing path 1910 includes, but is not limited to: a low-frequency subband entropy coding unit 1911, an inverse quantization / inverse transform unit 1912, a low-frequency subband reconstruction unit 1913, a low-frequency subband stitching unit 1915, and a prediction unit 1914. Detailed descriptions can be found in Figure 18, and will not be repeated here. The descriptions of the input and output coefficients or data of each unit (such as low-frequency subband coding data 1901, low-frequency subband quantization coefficients 1902, low-frequency subband inverse quantization coefficients 1903, low-frequency subband reconstruction block 1904, reconstructed low-frequency subband 1906, prediction block 1905, and syntax element 1907) can be found in Figure 18, and will not be repeated here.

[0348] The high-frequency subband processing path 1920 includes, but is not limited to: high-frequency subband entropy coding unit 1921, inverse quantization / inverse transform unit 1922, and high-frequency subband reconstruction unit 1923.

[0349] The high-frequency subband entropy coding unit 1921 obtains the quantization coefficients 1908 of the high-frequency subband block based on the high-frequency subband coded data 1902. The inverse quantization / inverse transform unit 1922 obtains the reconstruction coefficients 1909 of the high-frequency subband block based on the quantization coefficients 1908 of the high-frequency subband block, which can also be called the high-frequency subband reconstruction block.

[0350] The wavelet inverse transform unit 1903 is used to obtain the reconstruction coefficients 1909 of the high-frequency subband block, i.e., the high-frequency subband reconstruction block, and the low-frequency subband reconstruction block 1904 output by the low-frequency subband reconstruction block unit 1913. The wavelet inverse transform is performed on the high-frequency subband reconstruction block (e.g., including HH subband reconstruction block, HL subband reconstruction block, LH subband reconstruction block) and the low-frequency subband reconstruction block 1904 to obtain the reconstruction block 1931, which is the reconstruction block of the current subgraph. It can also be called the reconstruction data of the current block of the current subgraph or the reconstruction value of the current block of the current subgraph.

[0351] The wavelet inverse transform unit 1930 outputs a reconstructed block 1931 to the subgraph combination unit 1940. The subgraph combination unit 1940 can obtain the reconstructed subgraph of the current subgraph based on at least one reconstructed block corresponding to the current subgraph, which can also be referred to as the reconstructed value or reconstructed data of the current subgraph.

[0352] Subgraph combining unit 1940 outputs reconstructed subgraph 1932 to image combining unit 1950.

[0353] The image combination unit 1950 is used to acquire the reconstructed sub-image 1932 to obtain the reconstructed image 1933 of the original image, or the reconstructed value or reconstructed data of the original image, etc. The undescribed parts of Figure 19 can be referred to Figure 18, and will not be repeated here.

[0354] Figure 20a is a schematic diagram of the framework of an edge-cloud collaborative system provided in an embodiment of this application. The edge-cloud collaborative system may include: a central server, edge servers, and clients; wherein, a central server may connect to one or more edge servers, and an edge server may connect to one or more clients. This application does not limit the number of edge servers and clients; the specific number can be flexibly set according to the application scenario.

[0355] For example, a client can access the network through wireless access points such as base stations or Wi-Fi access points and communicate with the edge server through the network, or the client and the edge server can also communicate through a wired connection. Similarly, the edge server can also access the network through wireless access points such as base stations or Wi-Fi access points and communicate with the central server through the network, or the edge server and the central server can also communicate through a wired connection.

[0356] For example, the central server can be a single server, a server cluster consisting of multiple servers, or other distributed systems; this application does not impose any restrictions on this.

[0357] The client can be software, applications, browsers, in-vehicle systems, terminal devices, etc. When the client is implemented as a terminal device, it can include, but is not limited to, the following as shown in Figure 20a: mobile phone, personal computer (PC), virtual reality (VR) device, augmented reality (AR) device, tablet computer, laptop computer, etc. The client can also be a smart TV, mobile internet device (MID), wearable device (such as smartwatch, smart glasses, or smart helmet), smart car, wireless terminal device in industrial control, wireless terminal device in self-driving, wireless terminal device in remote medical surgery, wireless terminal device in smart grid, wireless terminal device in transportation safety, wireless terminal device in smart city, wireless terminal device in smart home, etc. The following embodiments do not impose special limitations on the specific form of the client.

[0358] In one possible scenario, some clients may connect directly to the central server instead of the edge server; in another possible scenario, all clients may connect directly to the central server instead of the edge server.

[0359] Furthermore, the edge-cloud collaborative system framework shown in Figure 20a is only an example of the edge-cloud collaborative system framework of this application. In the edge-cloud collaborative system of this application, the central server and the edge server can also be the same server; or the edge-cloud collaborative system of this application does not include edge servers, but the central server connects with each client. This application does not impose any restrictions on this.

[0360] As shown in Figure 20a, the end-to-cloud collaborative system can be applied to various image encoding and decoding scenarios of end-to-cloud collaboration, such as cloud gaming, cloud exhibitions, 3D cloud conferencing, 3D scenes, interior decoration, clothing design, architectural design and other multi-end collaborative image encoding and decoding scenarios. This application does not limit this.

[0361] In some embodiments, this application provides an encoding method and a decoding method. The functionality of the encoding method can be implemented by at least one of a server and a client, and the functionality of the decoding method can be implemented by at least one of a server and a client.

[0362] This server can be implemented through software or hardware.

[0363] In the first example, when the functionality of the server is implemented through software, the server may be, for example, an application running on a computing instance (such as an encoder or decoder, or an application that implements the encoding and decoding methods of this application), which may be, for example, a virtual machine, a container, or a host.

[0364] In the second example, when the server functionality is implemented through hardware, the server can be implemented through at least one physical device including a processor. This physical device can be a server (e.g., a central server or an edge server), a base station, a relay device, satellite equipment, etc., and there are no restrictions on this.

[0365] The processor can be a central processing unit (CPU) or a graphics processing unit (GPU), or it can be any type of processor or any combination thereof, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a system-on-chip (SoC), a software-defined infrastructure (SDI) chip, an AI chip, or a data processing unit (DPU).

[0366] Furthermore, the number of processors included in the server can be arbitrary, and the types of processors included can be one or more. The specific number and types of processors can be set according to the actual business needs of the application, and this application does not impose any restrictions on this.

[0367] In the third example, when the server-side functionality is implemented through hardware, the server can also be a computing cluster comprising multiple computing nodes. Furthermore, these multiple computing nodes can communicate through the at least one switching node. Exemplarily, a computing node can be a computing server including an accelerator card. This accelerator card can be, for example, a deep-learning processing unit (DPU), a GPU, a neural-network processing unit (NPU), or a tensor processing unit (TPU), or other types of accelerator cards. Alternatively, a computing node can be a computing server including a general-purpose processor (such as a CPU).

[0368] The encoding method provided in this application can be applied to any encoder, and the decoding method provided in this application can be applied to any decoder. The encoder and decoder can be standard video encoders and standard video decoders. These standard video encoders and decoders can be implemented according to industry video compression standards, such as ITU-T H.264, H.265 / HEVC, H.266 / VVC, AVS2, AVS3 standards, or extensions of such standards. The solutions in this application are not limited to any specific encoding / decoding standard.

[0369] The encoding and decoding method provided in this application embodiment can be combined with the encoding and decoding method of any of the above embodiments. In addition, the encoding and decoding method of this application embodiment can be applied to the encoding and decoding of images and videos in any intra-frame prediction scenario.

[0370] The following example illustrates the encoding and decoding methods of this application by using an encoder to execute the encoding method and a decoder to execute the decoding method. However, this application does not limit the subject that executes the encoding and decoding methods of this application.

[0371] The method in this application embodiment can encode and decode the current block in a video or image in intra-frame prediction mode.

[0372] In some embodiments, the intra-frame prediction mode may include at least one of the following: direct current (DC) mode, vertical mode, horizontal mode, planar mode, and cross-component linear model (CCLM) mode.

[0373] In some embodiments, each sample point in the block may include three components: Y, UV, and chromaticity. The Y component is the luminance component, and the UV components are the chromaticity components.

[0374] In some embodiments, the intra-frame prediction mode for the luminance component may include at least one of the following: DC mode, vertical mode, horizontal mode, and Planar mode.

[0375] In some embodiments, the intra-frame prediction mode of the chroma component may include at least one of the following: DC mode, vertical mode, horizontal mode, and CCLM mode.

[0376] In some embodiments, when the current block is encoded and decoded using intra-frame prediction mode, this application provides an encoding and decoding method that differs from traditional encoding and decoding methods.

[0377] The encoding method of this application embodiment will be illustrated below with different examples.

[0378] Example A2

[0379] In some embodiments, when the first intra-frame prediction mode used by the chroma component of the current block is CCLM mode, the following embodiments will describe in detail how to use the CCLM mode of the present application to encode the current block to obtain the above-mentioned bitstream.

[0380] In some embodiments, when the prediction mode of the first frame of the current block is the CCLM mode of the present application, the encoding can be implemented by any of the following methods.

[0381] In some embodiments, the method may include, but is not limited to, the following steps:

[0382] Step 1: Obtain the reconstruction data of the reference block corresponding to the block to be encoded in the image to be encoded.

[0383] Step 2: Obtain the parameters of CCLM based on the reconstruction data of the reference block.

[0384] Because this method involves intra-frame prediction, the block to be encoded and the reference block belong to the image to be encoded.

[0385] CCLM can be represented by linear model 1: U = α1Y + β1, and linear model 2: V = α2Y + β2.

[0386] In the two linear models above, Y is the reconstructed value (also called reconstructed data) of the Y component of the current block, U in linear model 1 is the predicted value of the U component of the current block, and V in linear model 2 is the predicted value of the V component of the current block.

[0387] In other words, the intra-frame prediction principle of CCLM mode can be to use the reconstructed value of the Y component of the current block to obtain the predicted value of the UV component of the current block, thereby realizing the prediction of the UV component of the current block.

[0388] The two linear models of CCLM require two corresponding parameters (α1, β1) and (α2, β2). This step can obtain the two parameters of each of the two linear models based on the reconstruction value of the reference block corresponding to the current block in the current frame, so as to obtain the parameters of CCLM.

[0389] Step 3: Based on the parameters of the cross-component linear model, perform intra-frame prediction on the block to be encoded to obtain the prediction data of the block to be encoded.

[0390] The parameters of CCLM are determined by M luminance values ​​derived from the reconstructed data based on the reference block, where M is an integer greater than 2.

[0391] Based on the parameters α1, β1, α2, and β2 of the two linear models mentioned above, intra-frame prediction of the chromaticity components (UV) of the current block can be performed using the CCLM mode to obtain the predicted values ​​of the UV components of the current block. Furthermore, the predicted value of the luminance component of the current block can also be obtained using the intra-frame prediction mode corresponding to the luminance component of the current block, thus obtaining the prediction data for the current block.

[0392] To obtain the parameters of CCLM, three or more luminance values ​​can be derived using the reconstructed data of the reference block.

[0393] Then, the parameters of CCLM can be obtained by using three or more luminance values ​​derived from the derivation.

[0394] In some embodiments, some or all of the three or more brightness values ​​may not be brightness values ​​directly selected from the reconstructed data of the reference block (i.e., reconstructed brightness values), but rather brightness values ​​obtained by performing a series of derivation calculations based on the reconstructed brightness values ​​of some sample points selected from the reference block.

[0395] Step 4: Based on the predicted data, encode the block to be encoded to obtain the bitstream of the image to be encoded.

[0396] For example, the actual value of the current block and the predicted value of the current block can be used to calculate the residual, and then the residual can be encoded to obtain the bitstream of the image to be encoded.

[0397] For example, the image to be encoded may include multiple blocks of data, which can be encoded in the order of their encoding to obtain the bitstream described above.

[0398] In an optional embodiment, Figure 20b shows multiple blocks of data in the image to be encoded, arranged in the encoding order as block 1->block 2->block 3->block 4->block 5->block 6->block 7>block 8->block 9. Of course, this application does not limit the block division method of the image to be encoded, nor the number of blocks obtained through block division.

[0399] In some embodiments, the size of each block is 8*8. In some embodiments, the size of the block can also be 8*4 or 4*4, and there is no limitation here.

[0400] The strategy for selecting a reference block for the current block within the current frame may also differ depending on the intra-prediction mode.

[0401] In CCLM mode, the reference block corresponding to the current block in CCLM mode can be the left-hand adjacent block of the current block, and / or the top-hand adjacent block of the current block.

[0402] Taking Figure 20b as an example, when the current block is block 5, for example, if the chroma component of block 5 is intra-frame predicted in CCLM mode, then in the above embodiment, the reference block of block 5 can be at least one of block 2 located to the left of block 5 and block 4 located above block 5.

[0403] Similarly, the encoding order and decoding order can be the same, so the decoder can decode each block in the following order: block 1->block 2->block 3->block 4->block 5->block 6->block 7>block 8->block 9.

[0404] In this embodiment of the application, when performing intra-frame prediction on a block, if the intra-frame prediction mode is CCLM mode, the method of this embodiment can derive the values ​​of three or more luminance components (hereinafter referred to as luminance values) based on the reconstruction data of the reference block corresponding to the current block in the CCLM mode in order to obtain the parameters of the CCLM mode. Based on the obtained luminance values, the parameters of the CCLM mode are obtained. This method has low complexity when performing intra-frame prediction, which can improve the efficiency of intra-frame prediction and thus improve the coding efficiency.

[0405] In some embodiments, this encoding method can be applied to a wavelet transform encoder. Thus, the current block on the encoding side can be block data from the low-frequency subband obtained by wavelet transform. Further implementation details regarding wavelet transform encoders can be found in the above introduction to wavelet transform-based encoders, and will not be repeated here.

[0406] Considering that the size of the low-frequency subband output by wavelet transform is only 1 / 4 of the image before wavelet transform (e.g., the image to be encoded), intra-frame prediction in CCLM mode on the block data in the low-frequency subband can reduce the amount of data encoded compared to encoding the image to be encoded, thereby improving the intra-frame prediction efficiency and coding efficiency. Considering that the data distribution of the low-frequency subband is more uniform than that of the high-frequency subband output by wavelet transform, intra-frame prediction and encoding of the low-frequency subband using the method of this application embodiment can enable the decoded low-frequency subband to be used for display. For example, the image to be displayed can be a thumbnail of the image to be encoded (e.g., the original image to be encoded).

[0407] In some embodiments, the current block on the coding side can also be block data in the high-frequency subband obtained by wavelet transform. Intra-frame prediction and corresponding coding are performed on the block data in the high-frequency subband using CCLM mode. The process is similar to the process of intra-frame prediction and coding in CCLM mode for the block data in the low-frequency subband, and will not be described in detail here.

[0408] In some embodiments, the intra-prediction mode used for intra-prediction of the current block can be encoded into the bitstream as a syntax element.

[0409] In some embodiments, as shown in FIG12, the intra-frame prediction mode information may be a specific example of the syntax element output by the mode decision unit (e.g., the mode decision unit shown in FIG12) to the prediction unit. For example, the syntax element may be a field indicating the CCLM mode.

[0410] The prediction unit, also known as the prediction processing unit, is used to acquire syntax elements and perform prediction processing based on these elements. Specifically, the prediction unit can select a prediction mode based on the syntax elements. In one example, the prediction unit can acquire a low-frequency subband reconstruction block (which is the reconstructed data of the reference block corresponding to the block to be coded) based on the syntax elements to obtain the prediction block. Specifically, the prediction unit can perform intra-frame prediction based on the low-frequency subband reconstruction block to obtain the prediction block.

[0411] Based on the encoding method of any of the above embodiments, the encoding method of this application in the scenario where the intra-frame prediction mode is CCLM mode will be described below with reference to different embodiments.

[0412] Figure 21a illustrates a flowchart of an encoding method according to this application. As shown in Figure 21a, the method flow may include, but is not limited to, the following steps:

[0413] S101. Determine the prediction model.

[0414] In an optional embodiment, the encoder determines that the prediction mode for encoding the current block is the intra-frame prediction mode.

[0415] In an optional embodiment, the encoder determines that the prediction mode for the UV components of the current block is CCLM mode.

[0416] S102. Obtain the reconstruction value of the reference block and the reconstruction value of the Y component of the current block.

[0417] In an optional embodiment, the reference block is a sub-block that has already been encoded. After the encoder encodes the reference block, it can reconstruct the data of the encoded reference block to obtain the reconstructed value of the reference block. The encoder can read the reconstructed value of the reference block.

[0418] In some embodiments, the Y component is also referred to as the luminance component, and the U and V components are also referred to as the chromaticity components.

[0419] In an optional embodiment, before encoding the UV components of the current block, the encoder may first encode the Y component of the current block to obtain encoded Y component data. Then, the encoder may decode the encoded Y component data of the current block to obtain the reconstructed Y component value of the current block and store the reconstructed Y component value. In this step, the encoder can read the stored reconstructed Y component value of the current block.

[0420] Following S102 and S101, the method may further include:

[0421] S103. Based on the reconstructed value of the reference block and the reconstructed value of the Y component of the current block, perform intra-frame prediction on the UV component of the current block through CCLM mode to obtain the predicted value of the U component and the predicted value of the V component of the current block.

[0422] S104. Obtain the actual value of the U component and the actual value of the V component of the current block.

[0423] In an optional embodiment, the encoder stores the actual values ​​of the U component and the actual values ​​of the V component of the current block, and the encoder obtains the actual values ​​of the U component and the actual values ​​of the V component of the current block from the encoder.

[0424] It should be noted that the embodiments of this application do not restrict the execution order between S103 and S104. That is, the embodiments of this application do not limit the execution order between the process of obtaining the actual values ​​of the U component and V component of the current block and the process of obtaining the predicted values ​​of the U component and V component of the current block.

[0425] S105. Determine the residuals of the U component and the V component of the current block.

[0426] In an optional embodiment, the encoder determines the residual of the U component of the current block based on the actual value of the U component and the predicted value of the U component of the current block. For example, the encoder determines the difference between the actual value and the predicted value as the predicted value of the U component of the current block.

[0427] In an optional embodiment, the encoder determines the residual of the V component of the current block based on the actual value of the V component and the predicted value of the V component of the current block. For example, the encoder determines the difference between the actual value and the predicted value as the predicted value of the V component of the current block.

[0428] Optionally, as shown in Figure 21a, the method may further include: obtaining the residual of the Y component of the current block.

[0429] In an optional embodiment, the encoder can determine the residual of the Y component of the current block based on the actual value of the Y component of the current block and the predicted value of the Y component of the current block.

[0430] In some embodiments, as shown in FIG21b, FIG21b is a schematic diagram of another encoding method provided by the embodiments of this application. The process describes the process of the encoder encoding the Y component of the current block.

[0431] The encoder determines the intra-prediction mode for the Y component of the current block as a first mode (e.g., one of DC mode, horizontal mode, vertical mode, or Planar mode). After obtaining the reconstructed value of the Y component of the reference block, the encoder performs intra-prediction on the Y component of the current block based on the prediction mode of the current block's luma component to obtain the predicted value of the Y component of the current block. The encoder obtains the actual value of the Y component of the current block and determines the residual of the Y component of the current block based on the actual value and the predicted value. For example, the encoder determines the difference between the actual value and the predicted value as the residual of the Y component of the current block. The residual of the Y component of the current block and the intra-prediction mode used by the Y component of the current block are then encoded into the bitstream.

[0432] Returning to Figure 21a, S106. Encode the residuals of the Y component, U component, and V component of the current block to obtain the bitstream.

[0433] In an optional embodiment, the encoder encodes the syntax element indicating the intra-prediction mode, the residual of the Y component of the current block, the residual of the U component of the current block, and the residual of the V component of the current block to obtain a bitstream.

[0434] In some embodiments, the bitstream may include a syntax element indicating that the intra-prediction mode of the chroma component of the current block is CCLM mode, and residuals including the chroma (UV) component of the current block and the residuals of the luminance (Y) component of the current block.

[0435] In an optional embodiment, the specific implementation of S103 shown in FIG21a can be achieved through the process shown in FIG22. As shown in FIG22, the process may include, but is not limited to, the following steps:

[0436] S1031. Obtain the parameters of the linear model based on the reconstructed values ​​of the Y component, U component, and V component of the reference block.

[0437] In some embodiments, the linear model may be a cross component linear model (CCLM), and the parameters of the linear model may also be referred to as the parameters of the CCLM mode.

[0438] The reconstructed values ​​of the three YUV components of the reference block are obtained by S102 as shown in Figure 21a.

[0439] In an optional embodiment, the linear model parameters may include parameters (α1 and β1) of linear model 1 for predicting the U component, and may also include parameters (α2 and β2) of linear model 1 for predicting the V component.

[0440] S1032. Based on the parameters of the linear model and the reconstructed value of the Y component of the current block, obtain the predicted value of the U component and the predicted value of the V component of the current block.

[0441] The reconstructed value of the Y component of the current block can be obtained from S102 as shown in Figure 21a.

[0442] In an optional embodiment, the encoder determines the linear model based on the parameters of the linear model. For example, the encoder determines linear model 1 as U = α1Y + β1 based on α1 and β1, where Y is the reconstructed value of the Y component of the current block, α1 and β1 are both parameters of linear model 1, and U is the predicted value of the U component of the current block; for example, the encoder determines linear model 2 as V = α2Y + β2 based on α2 and β2, where α2 and β2 are both parameters of linear model 2, and V is the predicted value of the V component of the current block.

[0443] In an optional embodiment, the encoder substitutes the reconstructed Y component value of the current block into the linear model 1 described above for calculation to determine the predicted U component value of the current block. The encoder substitutes the reconstructed Y component value of the current block into the linear model 2 described above for calculation to determine the predicted V component value of the current block.

[0444] In some embodiments, for example, if the size of the current block is 8*8, that is, it has 64 sample points (also called sampling points), then the YUV components of the current block can be sampled to obtain the YUV data of the current block, so that the ratio of the three components of Y, U, and V of the current block is 4:4:4.

[0445] For example, if the size of the current block is 8*8, then as shown in Figure 24a, the size of the Y component of the current block is 8*8, the size of the U component of the current block is 8*8, and the size of the V component of the current block is 8*8. This application embodiment does not limit this.

[0446] For each sample point in the current block, the encoder substitutes the reconstructed value of the Y component of the sample point into the linear model 1 above to obtain the predicted value of the U component of the sample point; the encoder substitutes the reconstructed value of the Y component of the sample point into the linear model 2 above to obtain the predicted value of the V component of the sample point.

[0447] In a specific embodiment, as shown in Figure 24a(2), the reconstructed value of the Y component of sample point P is Ya1. As shown in Figure 24a(1), the encoder substitutes the reconstructed value Ya1 into the above linear model 1 to obtain the predicted value PUa1 = α1 × Ya1 + β1 of the U component of sample point P. As shown in Figure 24a(3), the encoder substitutes the reconstructed value Ya1 into the above linear model 2 to obtain the predicted value PVA1 = α2 × Ya1 + β2 of the V component of sample point P.

[0448] In other embodiments, as shown in Figure 24b, for example, if the size of the current block is 8*8, that is, it has 64 sample points (also called sampling points), then the YUV data of the current block can be obtained by downsampling the UV components of the current block respectively, so that the ratio of the Y, U and V components of the current block is 4:2:2.

[0449] For example, as shown in Figure 24b, the size of the Y component of the current block is 8*8, the size of the U component of the current block is 8*4, and the size of the V component of the current block is 8*4. This application embodiment does not limit this.

[0450] The encoder determines a first reconstructed value based on the reconstructed value of the Y component corresponding to the sample point. The encoder substitutes the first reconstructed value into the linear model 1 described above to obtain the predicted value of the U component of the sample point; the encoder substitutes the first reconstructed value into the linear model 2 described above to obtain the predicted value of the V component of the sample point. For example, the encoder determines the average value of the reconstructed values ​​of the Y component corresponding to the sample point as the first reconstructed value.

[0451] In a specific embodiment, as shown in Figure 24b(2), the reconstructed values ​​of two sample points P4 and P3 of the Y component of the current block are Ya1 and Ya2, respectively. Based on the reconstructed values ​​Ya1 and Ya2, the encoder determines the first reconstructed value Yax = (Ya1 + Ya2) / 2. As shown in Figure 24b(1), the encoder substitutes the first reconstructed value Yax into the above linear model 1 to obtain the predicted value PUa1 = α1 × [(Ya1 + Ya2) / 2] + β1 of the U component of sample point P1. As shown in Figure 24b(3), the encoder substitutes the first reconstructed value Yax into the above linear model 2 to obtain the predicted value PVA1 = α2 × [(Ya1 + Ya2) / 2] + β2 of the V component of sample point P2.

[0452] In some embodiments, as shown in Figure 24c, for example, if the size of the current block is 8*8, that is, it has 64 sample points (also called sampling points), then the UV components of the current block can be downsampled to obtain the UV data of the current block, so that the ratio of the three components of Y, U, and V of the current block is 4:2:0.

[0453] For example, as shown in Figure 24c, the size of the Y component of the current block is 8*8, the size of the U component of the current block is 4*4, and the size of the V component of the current block is 4*4. This application embodiment does not limit this.

[0454] The encoder determines a first reconstructed value based on the reconstructed value of the Y component corresponding to the sample point. The encoder substitutes the first reconstructed value into the linear model 1 described above to obtain the predicted value of the U component of the sample point; the encoder substitutes the first reconstructed value into the linear model 2 described above to obtain the predicted value of the V component of the sample point. For example, the encoder determines the average value of the reconstructed values ​​of the Y component corresponding to the sample point as the first reconstructed value.

[0455] In a specific embodiment, as shown in Figure 24c(2), the reconstructed values ​​of the four sample points P7, P8, P9, and P10 of the Y component of the current block are Ya1, Ya2, Ya3, and Ya4, respectively. Based on the reconstructed values ​​Ya1, Ya2, Ya3, and Ya4, the encoder determines the first reconstructed value Yax = (Ya1 + Ya2 + Ya3 + Ya4) / 4. As shown in Figure 24c(1), the encoder substitutes the first reconstructed value Yax into the above linear model 1 to obtain the predicted value PUa1 = α1 × [(Ya1 + Ya2 + Ya3 + Ya4) / 4] + β1 of the U component of sample point P5. As shown in Figure 24c(3), the encoder substitutes the first reconstructed value Yax into the above linear model 2 to obtain the predicted value PVA1 = α2 × [(Ya1 + Ya2 + Ya3 + Ya4) / 4] + β2 of the V component of sample point P6.

[0456] In some embodiments, when obtaining the parameters of a linear model, as shown in Figure 23, the process may include, but is not limited to, the following steps:

[0457] S301. Obtain the reconstructed values ​​of the Y component of N1 samples, the U component of N2 samples, and the V component of N3 samples of the reference block.

[0458] N1, N2, and N3 are all positive integers greater than or equal to 3. M is less than or equal to N1, M is less than or equal to N2, and M is less than or equal to N3.

[0459] In some embodiments, the N1 samples are adjacent to the Y component of the current block. The N2 samples are adjacent to the U component of the current block, and the N3 samples are adjacent to the V component of the current block.

[0460] Each of the N1, N2, and N3 sample points obtained above is also called a reference sample point.

[0461] In an optional embodiment, as shown in Figures 25a, 25b, and 25c, taking Figure 24b as an example, the ratio of the three components Y, U, and V of the current block is 4:2:2. Thus, the size of the Y component 11 of the current block shown in Figure 25a(1) is 8*8, the size of the U component 12 of the current block shown in Figure 25a(2) is 8*4, and the size of the V component 13 of the current block shown in Figure 25a(3) is 8*4. This embodiment of the application does not impose any limitations on this.

[0462] As shown in Figure 25b(1), the size of the Y component 11 of the current block is 8*8; as shown in Figure 25b(2), the size of the U component 12 of the current block is 8*4; as shown in Figure 25b(3), the size of the V component 13 of the current block is 8*4. This application embodiment does not limit these.

[0463] As shown in Figure 25c(1), the size of the Y component 11 of the current block is 8*8, the size of the U component 12 of the current block is 8*4, and the size of the V component 13 of the current block is 8*4 as shown in Figure 25c(2). This application embodiment does not limit these.

[0464] As shown in Figure 25a, the sample points on the top and left sides of the current block are sample points of the reference block.

[0465] For example, sample points h1 and h2 correspond to sample point a. U 1 and sample point a V 1. Sample points h3 and h4 correspond to sample point a. U 2 and sample point a V 2. Sample points h5 and h6 correspond to sample point a. U 3 and sample point aV 3. Sample points h7 and h8 correspond to sample point a. U 4 and sample point a V 4.

[0466] In an optional embodiment, the sample point in the Y component of the reference block is called the first sample point, the sample point in the U component of the reference block is called the second sample point, and the sample point in the V component of the reference block is called the third sample point.

[0467] In specific embodiments, the current block has a reference block corresponding to the CCLM mode. For example, the reference block is the adjacent and encoded block to the left of the current block, and the reference block is the adjacent and encoded block to the top of the current block.

[0468] As shown in Figure 20b, the current block can be block 5, block 6, block 8, or block 9, etc.

[0469] To illustrate with a specific example 1, please refer to Figure 25a(1). Taking the current block as block 5 as shown in Figure 20b, the encoder determines that the reference blocks corresponding to the current block in the CCLM mode are blocks 2 and 4. The first sample points of block 2 adjacent to the current block are sample points h1 to h8, and the first sample points of block 4 adjacent to the current block are sample points v1 to v8. The encoder can select reference sample points h1, h7, h8, v1, and v8 from sample points h1 to h8 and v1 to v8. The encoder can obtain the reconstructed Y component values ​​of sample points h1, h7, h8, v1, and v8. Here, N1 shown in Figure 23 is 5. However, it is not limited to selecting 5 luminance values; it is sufficient to select more than or equal to 3 luminance values ​​from the reference blocks corresponding to the CCLM mode of the current block. Optionally, the reconstructed Y component values ​​of N1 samples can be selected from the samples adjacent to the current block within the reference block.

[0470] Please refer to Figure 25a(2). The encoder determines the second sample point of block 2 adjacent to the current block as sample point a. U 1 to sample point a U 4. The second sample point of block 4, which is adjacent to the current block, is b. U 1 to sample point b U 8. The encoder can obtain data from sample point a. U 1 to sample point a U 4. Sample point b U 1 to sample point b U In 8, a reference sample point is selected as sample point a. U 1. Sample point a U 4. Sample point b U 1 and sample point b U 8. The encoder can acquire sample point a U 1. Sample point a U 4. Sample point bU 1 and sample point b U The reconstructed values ​​of the U components are 8. Here, N2 in Figure 23 is 4. However, it is not limited to selecting 4 U component chromaticity values; any selection of 3 or more U component chromaticity values ​​from the reference block corresponding to the CCLM mode in the current block is acceptable. Optionally, the reconstructed values ​​of the U components of N2 samples can be selected from the samples adjacent to the current block within the reference block.

[0471] Please refer to Figure 25a(3). The encoder determines the third sample point of block 2 adjacent to the current block as sample point a. V 1 to sample point a V 4. The third sample point of block 4, which is adjacent to the current block, is b. V 1 to sample point b V 8. The encoder can obtain data from sample point a. V 1 to sample point a V 4. Sample point b V 1 to sample point b V In 8, a reference sample point is selected as sample point a. V 1. Sample point a V 4. Sample point b V 1 and sample point b V 8. The encoder can acquire sample point a V 1. Sample point a V 4. Sample point b V 1 and sample point b V The reconstructed V component values ​​of 8. Here, N3 shown in Figure 23 is 4. However, it is not limited to selecting 4 chromaticity values; any chromaticity value greater than or equal to 3 V component values ​​can be selected from the reference block corresponding to the CCLM mode in the current block. Optionally, the reconstructed V component values ​​of N3 samples can be selected from the samples adjacent to the current block in the reference block.

[0472] As shown in Figure 25b, the sample point on the top side of the current 8*8 block is the sample point of the reference block.

[0473] In a specific embodiment, the current block has a reference block corresponding to the CCLM mode. For example, the reference block is the adjacent and already encoded block above the current block.

[0474] As shown in Figure 20b, the current block can be block 4 or block 7, etc.

[0475] To illustrate with a specific example 2, please refer to Figure 25b(1). Taking the current block as block 4 as shown in Figure 20b as an example, the encoder determines that the reference block corresponding to the CCLM mode for the current block is block 1, and the first sample point of block 1 adjacent to the current block is sample point h1 to sample point h8. The encoder can select reference sample points h1, h2, h5 to h8 from sample points h1 to sample points h8, and the encoder can obtain the reconstructed values ​​of the Y components of sample points h1, h2, h5 to h8. Here, N1 shown in Figure 23 is 6. However, it is not limited to selecting 6 luminance values, as long as more than or equal to 3 luminance values ​​are selected from the reference block corresponding to the CCLM mode of the current block. Optionally, the reconstructed values ​​of the Y components of N1 sample points can be selected from the sample points adjacent to the current block in the reference block.

[0476] Please refer to Figure 25b(2). The encoder determines the second sample point of block 1 adjacent to the current block as sample point a. U 1 to sample point a U 4. The encoder can obtain data from sample point a. U 1 to sample point a U In section 4, a reference sample point is selected as sample point a. U 1. Sample point a U 3 and sample point a U 4. The encoder can acquire sample point a U 1. Sample point a U 3 and sample point a U The reconstructed value of the U component is 4. Here, N2 in Figure 23 is 3. However, it is not limited to selecting 3 chromaticity values; any U component chromaticity value greater than or equal to 3 can be selected from the reference block corresponding to the CCLM mode in the current block. Optionally, the reconstructed values ​​of the U components of N2 samples can be selected from the samples adjacent to the current block in the reference block.

[0477] Please refer to Figure 25b(3). The encoder determines the third sample point of block 1 adjacent to the current block as sample point a. V 1 to sample point a V 4. The encoder can obtain data from sample point a. V 1 to sample point a V In section 4, a reference sample point is selected as sample point a. V 1. Sample point a V 3 and sample point a V 4. The encoder can acquire sample point a V 1. Sample point a V 3 and sample point a VThe reconstructed V component values ​​are 4. Here, N3 in Figure 23 is 3. However, it is not limited to selecting 3 chromaticity values; any chromaticity value greater than or equal to 3 V component values ​​can be selected from the reference block corresponding to the CCLM mode in the current block. Optionally, the reconstructed V component values ​​of N3 samples can be selected from the samples adjacent to the current block in the reference block.

[0478] As shown in Figure 25c, the sample points on the left side of the current block are the sample points of the reference block.

[0479] In a specific embodiment, there is a reference block in the current frame that corresponds to the CCLM mode of the current block. For example, the reference block is an adjacent and encoded block located to the left of the current block in the current frame.

[0480] As shown in Figure 20b, the current block can be block 2 or block 3, etc.

[0481] To illustrate with a specific example 3, please refer to Figure 25c(1). Taking the current block as block 3 as shown in Figure 20b, the encoder determines that the reference block corresponding to the CCLM mode for the current block is block 2, and the first sample points of block 2 adjacent to the current block are sample points v1 to v8. The encoder can select reference sample points v1, v5, and v8 from sample points v1 to v8, and the encoder can obtain the reconstructed values ​​of the Y components of sample points v1, v5, and v8. Here, N1 shown in Figure 23 is 3. However, it is not limited to selecting 3 luminance values, as long as 3 or more luminance values ​​are selected from the reference block corresponding to the CCLM mode of the current block. Optionally, the reconstructed values ​​of the Y components of N1 sample points can be selected from the sample points adjacent to the current block in the reference block.

[0482] Please refer to Figure 25c(2). The encoder determines the second sample point of block 2, which is adjacent to the current block, as sample point b. U 1 to sample point b U 8. The encoder can obtain data from sample point b. U 1 to sample point b U In 8, the reference sample point is selected as sample point b. U 1. Sample point b U 5 and sample point b U 8. The encoder can acquire sample point b. U 1. Sample point b U 5 and sample point b U The reconstructed values ​​of the U component of 8. Here, N2 in Figure 23 is 3. However, it is not limited to selecting 3 chromaticity values; any U component chromaticity value greater than or equal to 3 can be selected from the reference block corresponding to the CCLM mode in the current block. Optionally, the reconstructed values ​​of the U components of N2 samples can be selected from the samples adjacent to the current block in the reference block.

[0483] Please refer to Figure 25c(3). The encoder determines the third sample point of block 2 adjacent to the current block as sample point b. V 1 to sample point b V 8. The encoder can obtain data from sample point b. V 1 to sample point b V In 8, the reference sample point is selected as sample point b. V 1. Sample point b V 5 and sample point b V 8. The encoder can acquire sample point b. V 1. Sample point b V 5 and sample point b V The reconstructed V component values ​​of 8. Here, N3 in Figure 23 is 3. However, it is not limited to selecting 3 chromaticity values; any chromaticity value greater than or equal to 3 V component values ​​can be selected from the reference block corresponding to the CCLM mode in the current block. Optionally, the reconstructed V component values ​​of N3 samples can be selected from the samples adjacent to the current block in the reference block.

[0484] It should be noted that the embodiments of this application do not impose any restrictions on how to select sample points from the reference block. For example, the embodiments of this application do not restrict the position of the selected sample points in the reference block, nor do they restrict the specific number of selected sample points, as long as the number of selected sample points is greater than or equal to 3.

[0485] In other embodiments, the selected N1, N2, and N3 samples may also be partially or entirely non-adjacent to the current block. The encoder selects N1, N2, and N3 samples from the reference block of the current block. From these samples, the encoder selects the reconstructed Y component values ​​of the N1 samples, the reconstructed U component values ​​of the N2 samples, and the reconstructed V component values ​​of the N3 samples.

[0486] Returning to Figure 23, after S301, the method may further include:

[0487] S302. Based on the reconstructed values ​​of the Y components of N1 samples, the U components of N2 samples, and the V components of N3 samples of the reference block, the values ​​of M (M is greater than or equal to 3, for example, M=3) Y components, M U components, and M V components are derived.

[0488] In an optional embodiment, the values ​​of the three Y components derived are represented by Y1, Y2, and Y3, the values ​​of the three U components derived are represented by U1, U2, and U3, and the values ​​of the three V components derived are represented by V1, V2, and V3.

[0489] In some embodiments, the relative positional relationship between the three Y components can indicate the magnitude relationship between Y1, Y2, and Y3, or the relative distance between them, etc., without limitation.

[0490] In an optional embodiment, the encoder derives the values ​​of the three Y components based on the reconstructed values ​​of the Y components of the N1 samples of the reference block.

[0491] In some embodiments, N1>3, the encoder calculates the reconstructed values ​​of N1 samples selected from the Y component of the reference block to obtain three luminance values ​​Y1, Y2, and Y3.

[0492] In conjunction with the above specific example 1, please refer to Figure 25a. As a specific example, as shown in Figure 25a(1), the encoder can perform calculations on 5 sample points of the Y component selected from the reference block to obtain the brightness values ​​Y1, Y2, and Y3 of the 3 Y components.

[0493] In some embodiments, the reconstructed values ​​of adjacent sample points h1 and v1 as shown in FIG25a(1) can be averaged to obtain the brightness value Y1; the reconstructed values ​​of adjacent sample points h7 and h8 as shown in FIG25a(1) can be averaged to obtain the brightness value Y2; the reconstructed value of adjacent sample point v8 as shown in FIG25a(1) can be determined as the brightness value Y3.

[0494] For example, the encoder determines Y1 using Formula 1 based on the reconstructed values ​​of sample h1 and sample v1, where Formula 1 is: Y1 = (H1y + V1y) / 2, H1y is the reconstructed value of sample h1, V1y is the reconstructed value of sample v1, and " / " represents division. The encoder determines Y2 using Formula 2 based on the reconstructed values ​​of sample h7 and sample h8, where Formula 2 is: Y2 = (H7y + H8y) / 2, H7y is the reconstructed value of sample h7, and H8y is the reconstructed value of sample h8.

[0495] Thus, in the example of Figure 25a(1), the encoder can average the brightness values ​​of adjacent samples among the five selected gray reference samples to obtain three brightness values ​​Y1, Y2, and Y3.

[0496] In conjunction with the above specific example 2, please refer to Figure 25b. As a specific example, as shown in Figure 25b(1), the encoder can calculate the brightness values ​​Y1, Y2, and Y3 of the three Y components by selecting 6 sample points of the Y component from the reference block.

[0497] In some embodiments, the reconstructed values ​​of adjacent sample points h1 and h2 as shown in FIG25b(1) can be averaged to obtain the brightness value Y1; the reconstructed values ​​of adjacent sample points h5 and h6 as shown in FIG25b(1) can be averaged to obtain the brightness value Y2; the reconstructed values ​​of adjacent sample points h7 and h8 as shown in FIG25b(1) can be averaged to obtain the brightness value Y3.

[0498] For example, the encoder determines Y1 using Formula 3 based on the reconstructed values ​​of sample h1 and sample h2, where Formula 3 is: Y1 = (H1y + H2y) / 2, and H1y is the reconstructed value of sample h1, and H2y is the reconstructed value of sample h2. The encoder determines Y2 using Formula 4 based on the reconstructed values ​​of sample h5 and sample h6, where Formula 4 is: Y2 = (H5y + H6y) / 2, and H5y is the reconstructed value of sample h5, and H6y is the reconstructed value of sample h6. The encoder determines Y3 using Formula 5 based on the reconstructed values ​​of sample h7 and sample h8, where Formula 3 is: Y2 = (H7y + H8y) / 2, and H7y is the reconstructed value of sample h7, and H8y is the reconstructed value of sample h8.

[0499] Thus, in the example of Figure 25a(1), the encoder can average the brightness values ​​of adjacent samples among the five selected gray reference samples to obtain three brightness values ​​Y1, Y2, and Y3.

[0500] In other embodiments, N1 samples selected from the luminance components of the reference block are 3 samples, i.e., N1=3. The encoder directly uses the reconstructed Y component values ​​of the 3 selected reference samples as the M luminance values ​​derived above, M=3, so that the values ​​of Y1, Y2, and Y3 are the same as the reconstructed Y component values ​​of the 3 selected reference samples.

[0501] Referring to the specific example 3 above, and referring to Figure 25c, as a specific example, as shown in Figure 25c(1), the encoder can use the reconstructed values ​​of the Y components of the three reference samples (sample v1, sample v5, and sample v8) selected from the reference block as the three luminance values ​​Y1, Y2, and Y3 derived above, respectively. For example, the encoder determines the reconstructed value of the Y component of sample v1 as Y1; the encoder determines the reconstructed value of the Y component of sample v5 as Y2; and the encoder determines the reconstructed value of the Y component of sample v8 as Y8.

[0502] In an optional embodiment, the encoder derives the values ​​of the three U components based on the reconstructed values ​​of the U components of the N1 samples of the reference block.

[0503] In some embodiments, N2>3, the encoder calculates the reconstructed values ​​of N2 samples selected from the U component of the reference block to obtain three chromaticity values ​​U1, U2, and U3.

[0504] In conjunction with the above specific example 1, please refer to Figure 25a. As a specific example, as shown in Figure 25a(2), the encoder can perform calculations on 4 sample points of the U component selected from the reference block to obtain the chromaticity values ​​U1, U2, and U3 of the 3 U components.

[0505] In some embodiments, adjacent sample points a as shown in FIG25a(2) can be used. U Reconstructed values ​​of 1 and sample b U The average value of the reconstructed values ​​of sample a is calculated to obtain the chromaticity value U1; the sample points a can be used to calculate the chromaticity value U1. U The reconstructed value of 4 is used as the derived chromaticity value U2; the sample point b can be used as the chromaticity value. V The reconstructed value of 8 is used as the derived chromaticity value U3.

[0506] For example, the encoder is based on sample point a U Reconstructed values ​​of 1 and sample b U The reconstructed value of 1 is determined using Formula 6, where Formula 6 is: U1 = (A U 1U+B U 1U) / 2, A U 1U represents sample point a U The reconstruction value of 1, B U 1U represents sample point b U The reconstruction value is 1.

[0507] In other embodiments, the N2 samples reselected from the chromaticity components of the reference block are 3 samples, i.e., N2=3. The encoder directly uses the reconstructed U component values ​​of the 3 selected samples as the M chromaticity values ​​derived above, M=3, so that the values ​​of U1, U2, and U3 are the same as the reconstructed U component values ​​of the 3 selected reference samples.

[0508] Referring to the specific example 2 above, and as shown in Figure 25b, as a specific example, the encoder can select 3 reference samples (sample a) from the reference block. U 1. Sample point a U 3 and sample point a U 4) The reconstructed U component values ​​are used as the three chromaticity values ​​U1, U2, and U3 derived above. For example, the encoder will use sample point a U The reconstructed value of the U component of 1 is determined as U1; the encoder will convert sample point a U The reconstructed value of the U component of 3 is determined as U2; the encoder will convert sample point a U The reconstructed value of the U component of 4 is determined to be U3.

[0509] Referring to the specific example 3 above, and as shown in Figure 25c, as a specific example, the encoder can select 3 reference samples (sample b) from the reference block. U 1. Sample point b U 3 and sample point b U 4) The reconstructed U component values ​​are used as the three chromaticity values ​​U1, U2, and U3 derived above. For example, the encoder will use sample point b U The reconstructed value of the U component of 1 is determined as U1; the encoder will convert sample point b U The reconstructed value of the U component of 5 is determined as U2; the encoder will convert sample point b U The reconstructed value of the U component of 8 is determined to be U3.

[0510] In an optional embodiment, the encoder derives the values ​​of the three V components based on the reconstructed values ​​of the V components of the N3 samples of the reference block.

[0511] In some embodiments, N3>3, the encoder calculates the reconstructed values ​​of N3 samples selected from the V component of the reference block to obtain three chromaticity values ​​V1, V2, and V3.

[0512] In conjunction with the above specific example 1, please refer to Figure 25a. As a specific example, as shown in Figure 25a(3), the encoder can perform calculations on 4 sample points of the V component selected from the reference block to obtain the chromaticity values ​​V1, V2, and V3 of the 3 V components.

[0513] In some embodiments, adjacent sample points a as shown in Figure 25a(3) can be... V Reconstructed values ​​of 1 and sample b V The average value of the reconstructed value of sample a is calculated to obtain the chromaticity value V1; the sample points a can be used to calculate the chromaticity value V1. V The reconstructed value of 4 is used as the derived chromaticity value V2; the encoder will use sample point b V The reconstructed value of 8 is used as the derived chromaticity value V3.

[0514] For example, the encoder is based on sample point a V Reconstructed values ​​of 1 and sample b V The reconstructed value of 1 is determined using Formula 6, where Formula 6 is: V1 = (A V 1V+B V 1V) / 2, A V 1V represents sample point a V The reconstruction value of 1, B V 1V represents sample point b V The reconstruction value is 1.

[0515] In other embodiments, the N3 samples reselected from the chromaticity components of the reference block are 3 samples, i.e., N3 = 3. The encoder directly uses the reconstructed V component values ​​of the selected 3 samples as the M chromaticity values ​​derived above, M = 3, so that the values ​​of V1, V2, and V3 are the same as the reconstructed V component values ​​of the selected 3 reference samples.

[0516] Referring to the specific example 2 above, and as shown in Figure 25b, as a specific example, as shown in Figure 25b(3), the encoder can select 3 reference samples (sample a) from the reference block. V 1. Sample point a V 3 and sample point a V 4) The reconstructed V component values ​​are used as the three chromaticity values ​​V1, V2, and V3 derived above. For example, the encoder will use sample point a V The reconstructed value of 1 is determined as V1; the encoder will convert sample point a V The reconstructed value of 3 is determined as V2; the encoder will convert sample point a V The reconstruction value of 4 is determined to be V3.

[0517] Referring to the specific example 3 above, and considering Figure 25c, as a specific example, as shown in Figure 25c(3), the encoder can select 3 reference samples (sample b) from the reference block. V 1. Sample point b V 3 and sample point b V 4) The reconstructed V component values ​​are used as the three chromaticity values ​​V1, V2, and V3 derived above. For example, the encoder will use sample point b V The reconstructed value of 1 is determined as V1; the encoder will convert sample point b V The reconstructed value of 5 is determined as V2; the encoder will convert sample point b V The reconstruction value of 8 is determined to be V3.

[0518] Returning to Figure 23, the method also includes:

[0519] S303. Based on the derived values ​​of the three Y components, three U components, and three V components, obtain the parameters of the linear model.

[0520] In some embodiments, the parameters of the linear model may include parameters (α1 and β1) of linear model 1 for predicting the U component and parameters (α2 and β2) of linear model 2 for predicting the V component.

[0521] In a specific embodiment, an index and a first parameter Du are determined based on the three Y components. A second parameter α is determined based on this index. The encoder determines parameter α1 based on the first parameter Du and the second parameter α. For example, the encoder determines parameter α1 as the product of the first parameter Du and the second parameter α, that is, the parameter α1 determined by the encoder is α × Du.

[0522] In some embodiments, the index is determined by a weighted calculation based on the three Y components and their respective weights. For example, the weights corresponding to the three Y components are w1, w2, and w3. For example, the index S satisfies the formula: S = w1 × Y1 + w2 × Y2 + w3 × Y3.

[0523] In some embodiments, w1 is a negative integer, and w2 and w3 are both positive integers. For example, w1 = -3, w2 = 1, and w3 = 2.

[0524] In some embodiments, w1 and w2 are both negative integers, and w3 is a positive integer. For example, w1 = -2, w2 = -1, and w3 = 3.

[0525] In some embodiments, the second parameter α is determined based on the index. For example, the second parameter α corresponding to the index can be obtained from a lookup table based on the index.

[0526] In some embodiments, the first parameter Du is determined based on the three U components and their respective weights. It should be understood that the weights corresponding to the three U components are the same as the weights corresponding to the three Y components. For example, Du = w1×U1 + w2×U2 + w3×U3.

[0527] In some embodiments, the first value, the second value, parameter α1, and parameter β1 satisfy a first prediction formula. The first prediction formula is: P1 = α1 × A + β1, where P1 is the first value and A is the second value. Details regarding the first and second values ​​are provided below. Thus, parameter β1 can be obtained.

[0528] For example, the average of U1 and U2 can be determined as the first value, and the average of Y1 and Y2 can be determined as the second value. That is, the first value is (U1+U2) / 2, and the second value is (Y1+Y2) / 2.

[0529] For example, the average of U2 and U3 can be determined as the first value, and the average of Y2 and Y3 can be determined as the second value. That is, the first value is (U2+U3) / 2, and the second value is (Y2+Y3) / 2.

[0530] Similarly, the parameters (α² and β²) of linear model 2 can be obtained based on the three Y components and three V components. This will not be elaborated further here. It should be understood that the weights used to obtain the parameters of linear model 2 are the same as those used to obtain the parameters of linear model 1.

[0531] In an optional embodiment, the encoder determines the relative relationship of the three derived Y components (Y1, Y2, Y3). Based on this relative relationship (also called relative positional relationship), the parameters of the CCLM are obtained.

[0532] In a specific embodiment, the encoder determines the relative relationship between Y1, Y2, and Y3 based on their magnitude relationships.

[0533] In one embodiment, the encoder determines that Y1≤Y2≤Y3 and (2×Y2)≥(Y1+Y3), that is, Y2 is closer to Y3 of Y1 and Y3 (or the distance between Y2 and Y1 is the same as the distance between Y2 and Y3). The encoder determines the relative relationship between Y1, Y2, and Y3 as a first relationship, which is used to indicate that Y1≤Y2≤Y3 and the distance between Y2 and Y1 is greater than or equal to the distance between Y2 and Y3.

[0534] In a specific embodiment, the encoder determines the weight w1 of Y1, the weight w2 of Y2, and the weight w3 of Y3 based on the relative positional relationship between the three Y components. The encoder then determines the index based on Y1, Y2, Y3, w1, w2, and w3.

[0535] In an optional embodiment, the encoder performs a weighted operation based on Y1, Y2, Y3, w1, w2, and w3 to obtain the index. In a specific embodiment, the encoder determines the index based on Y1, Y2, Y3, w1, w2, and w3 using Formula 7. The formula for calculating the index is: S = w1 × Y1 + w2 × Y2 + w3 × Y3, Formula 7. Here, S is the index, and w1 < w2 < w3.

[0536] The encoder uses U1, U2, U3, w1, w2, and w3 to determine the first parameter using a first parameter formula. The first parameter formula is: Du = w1 × U1 + w2 × U2 + w3 × U3 (Formula 8). Here, Du is the first parameter.

[0537] In some embodiments, when the relative relationship of the three Y components is a first relationship, w1 is a negative integer, and w2 and w3 are both positive integers. For example, w1 = -3, w2 = 1, w3 = 2, the encoder determines the index S = -3×Y1 + Y2 + 2×Y3, and the encoder determines the first parameter Du = -3×U1 + U2 + 2×U3.

[0538] In an optional embodiment, the encoder determines a first value and a second value based on the relative relationship of the three Y components. Based on the first value, the second value, and parameter α1, the encoder uses a first prediction formula to determine parameter β1. The first prediction formula is: P1 = α1 × A + β1 (Formula 9). Here, P1 is the first value, and A is the second value.

[0539] In some embodiments, when the relative relationship of the three Y components is a first relationship (e.g., the case where Y2 is closer to Y1 as described above), the encoder determines the average value of U1 and U2 as a first value, and the encoder determines the average value of Y1 and Y2 as a second value. That is, the first value determined by the encoder is (U1+U2) / 2, and the second value determined by the encoder is (Y1+Y2) / 2.

[0540] In an optional embodiment, the encoder determines the second and third values ​​based on the relative relationship of the three Y components. Based on the second and third values ​​and parameter α2, the encoder uses a second prediction formula to determine parameter β2. The second prediction formula is: P2 = α2 × A + β2 (Formula 9). Here, A is the second value, and P2 is the third value.

[0541] In some embodiments, when the relative relationship of the three Y components is a first relationship (e.g., the case where Y2 is closer to Y1 as described above), the encoder determines the average value of V1 and V2 as a third value, and the encoder determines the average value of Y1 and Y2 as a second value. That is, the third value determined by the encoder is (V1+V2) / 2, and the second value determined by the encoder is (Y1+Y2) / 2.

[0542] In another embodiment, the encoder determines that Y1≤Y2≤Y3, and the encoder determines that (2×Y2)≤(Y1+Y3), that is, Y2 is closer to Y1 of Y1 and Y3 (or the distance between Y2 and Y1 is the same as the distance between Y2 and Y3). The encoder determines the relative relationship between Y1, Y2, and Y3 as a second relationship, which is used to indicate that Y1≤Y2≤Y3, and the distance between Y2 and Y1 is less than or equal to the distance between Y2 and Y3.

[0543] In other embodiments, when the relative relationship of the three Y components is a second relationship, w1 and w2 are both negative integers, and w3 is a positive integer. For example, w1 = -2, w2 = -1, w3 = 3, the encoder determines the index S = -2×Y1 - Y2 + 3×Y3, and the encoder determines the first parameter Du = -2×U1 - U2 + 3×U3.

[0544] In other embodiments, when the relative relationship of the three Y components is a second relationship (e.g., the case where Y2 is closer to Y3 as described above), the encoder determines the average value of U2 and U3 as a first value, and the encoder determines the average value of Y2 and Y3 as a second value. That is, the first value determined by the encoder is (U2+U3) / 2, and the second value determined by the encoder is (Y2+Y3) / 2.

[0545] In other embodiments, when the relative relationship of the three Y components is a second relationship (e.g., the case where Y2 is closer to Y3 as described above), the encoder determines the average value of V2 and V3 as the third value, and the encoder determines the average value of Y2 and Y3 as the second value. That is, the third value determined by the encoder is (V2+V3) / 2, and the second value determined by the encoder is (Y2+Y3) / 2.

[0546] It should be noted that there are many other relative relationships between Y1, Y2, and Y3. This application does not limit these relationships. The following embodiments are all illustrated using Y1≤Y2≤Y3 as an example. The specific implementation process for other cases is the same as that for Y1≤Y2≤Y3. This application will not elaborate on these details.

[0547] In an optional embodiment, the encoder obtains α1 and β1 based on the three Y components, the three U components, and the relative relationship of the three Y components.

[0548] In a specific embodiment, the encoder determines an index and a first parameter Du based on the three Y components. The encoder then determines a second parameter α based on this index. Finally, the encoder determines parameter α1 based on the first parameter Du and the second parameter α. For example, the encoder determines parameter α1 as the product of the first parameter Du and the second parameter α; that is, the parameter α1 determined by the encoder is α × Du.

[0549] In an optional embodiment, the encoder obtains α2 and β2 based on the three Y components, the three V components, and the relative relationship of the three Y components.

[0550] In a specific embodiment, the encoder determines the third parameter Dv based on the three Y components, and the encoder determines the parameter α2 based on the third parameter Dv and the second parameter α. For example, the encoder determines the parameter α2 as the product of the third parameter Dv and the second parameter α, that is, the parameter α2 determined by the encoder is α × Dv.

[0551] In some embodiments, the encoder determines the second parameter α based on an index and a lookup table. Optionally, the lookup table has a value less than 2. P The size of P is a positive integer, and this application does not limit this. For example, P = 4. In other embodiments, P can be a positive integer greater than or less than 4.

[0552] In an optional embodiment, the lookup table includes an index and entries corresponding to the indexes, where one index corresponds to one entry. The encoder searches the lookup table based on the index to determine the entry corresponding to that index, and then identifies that entry as the second parameter α corresponding to that index.

[0553] In one example, the lookup table is shown in Table 1 below:

[0554] Table 1

[0555] For example, if the index is "9", the encoder will look up the lookup table as shown in Table 1 based on "9" and determine that the entry corresponding to the index is "43". The encoder will then determine "43" as the second parameter α corresponding to the index "9".

[0556] In some embodiments, the lookup table can be represented as TABLE[index].

[0557] Where TABLE[index] = {63,63,61,57,54,51,49,47,45,43,41,39,38,37,35,34}, and 0 ≤ index < 16.

[0558] It should be noted that the above embodiments use the values ​​of the three derived Y components, three U components, and three V components, as well as the relative relationship between the three obtained Y components (the three derived Y components, M=3), as examples to illustrate how to obtain the parameters of the linear model. The embodiments of this application do not limit the specific number M of the derived Y components, U components, and V components.

[0559] Based on the N1 Y components selected from the reference block, the M luminance values ​​derived are 3 luminance values ​​(i.e., Y1, Y2, Y3 mentioned above). For example, Y1 < Y2 < Y3, or Y2 is closer to Y3 than Y1. Then, the index S used to look up the above lookup table (shown in Table 1) can be obtained by the above formula S = -3 × Y1 + Y2 + 2 × Y3.

[0560] For example, if Y1 < Y2 < Y3, and Y2 is closer to Y1 than Y3, then the index S used to look up the above lookup table (as shown in Table 1) can be obtained by the above formula S = -2 × Y1 - Y2 + 3 × Y3.

[0561] For example, if Y1 < Y2 < Y3, or if the distance between Y2 and Y3 is the same as the distance between Y2 and Y1, then the index S used to look up the above lookup table (shown in Table 1) can be obtained by using the above formula S = -2 × Y1 - Y2 + 3 × Y3, or the above formula S = -3 × Y1 + Y2 + 2 × Y3.

[0562] For example, if at least two of the brightness values ​​Y1, Y2, and Y3 are the same, such as Y2 being closer to Y3 than Y1, then the index S used to look up the above lookup table (as shown in Table 1) can be obtained by the above formula S = -3 × Y1 + Y2 + 2 × Y3.

[0563] For example, if at least two of the brightness values ​​Y1, Y2, and Y3 are the same, such as Y2 being closer to Y1 than Y2, then the index S used to look up the above lookup table (as shown in Table 1) can be obtained by the above formula S = -2 × Y1 - Y2 + 3 × Y3.

[0564] In the above method, M Y components are derived from N1 Y components, M U components are derived from N2 U components, and M V components are derived from N3 V components. In some embodiments, the M Y components, M U components, and M V components can be the reconstructed values ​​of M pixels in the reference block. In the above embodiment, M = 3.

[0565] In other embodiments, in S302 above, the value of M in the M Y components derived from the reconstructed values ​​of N1 Y components can be an integer greater than 3. For example, if the number of the M Y components, M U components, and M V components derived above is all 4, i.e., M = 4, then the weight w of each of the four derived Y components (represented as Y1, Y2, Y3, and Y4) can be different from the weight w value in the scenario where M = 3. However, the weight values ​​of the four Y components are still based on the relative positional relationship between the four Y components (e.g., the size relationship between the four Y components and the relative distance between the four Y components).

[0566] In an optional embodiment, Figure 26(1) is a schematic diagram of the linear model provided in the embodiment of this application. The linear model is the linear model 1 obtained based on the derived values ​​of the three luminance values ​​(Y1, Y2, Y3) and the values ​​of the three U components (U1, U2, U3). The linear model 1 can be the model corresponding to the parameters (α1 and β1) of CCLM. As shown in Figure 26(1), the horizontal axis represents the luminance value (Y component), the vertical axis represents the chrominance value (U component), and the linear model 1 can be expressed as U = α1Y + β1. The slope of the linear model 1 is α1, and the intercept of the linear model 1 is β1.

[0567] In some embodiments, the encoder determines α1 based on U1, U2, U3, Y1, Y2, and Y3. Therefore, the line U = α1Y + β1 is close to points P1, P2, and P3. The encoder determines β1 based on the average of U1 and U2 and the average of Y1 and Y2. Therefore, point (0, β1) is the midpoint of P1(Y1, U1) and P2(Y2, U2).

[0568] In an optional embodiment, Figure 26(2) is a schematic diagram of another linear model provided in this application embodiment. This linear model is obtained based on the derived values ​​of the above three luminance values ​​(Y1, Y2, Y3) and the values ​​of the three V components (V1, V2, V3). The linear model 2 can be the model corresponding to the parameters (α2 and β2) of CCLM. As shown in Figure 26(2), the horizontal axis represents the luminance value (Y component), the vertical axis represents the chrominance value (V component), and the linear model 2 can be expressed as V = α2Y + β2. The slope of the linear model 2 is α2, and the intercept of the linear model 1 is β2.

[0569] In some embodiments, the encoder determines α2 based on V1, V2, V3, Y1, Y2, and Y3. Therefore, the line V = α2Y + β2 is close to points P1, P2, and P3. The encoder determines β2 based on the average of V1 and V2 and the average of Y1 and Y2. Therefore, point (0, β2) is the midpoint of P1(Y1, V1) and P2(Y2, V2).

[0570] Example B2

[0571] Example B2 can be various embodiments of the decoding method corresponding to Example A2.

[0572] In some embodiments, when the first intra-frame prediction mode used by the chroma component of the current block is CCLM mode, the following embodiments will describe in detail how to use the CCLM mode of the present application to decode the bitstream in order to obtain the reconstructed data of the current block.

[0573] In some embodiments, when the prediction mode of the first frame of the current block is the CCLM mode of the present application embodiment, the bitstream decoding can be implemented by any of the following embodiments to obtain the reconstructed data of the current block.

[0574] In some embodiments, when the intra-prediction mode of the current block is CCLM mode, the decoder can decode the syntax elements from the bitstream to determine that the intra-prediction mode of the current block is CCLM mode.

[0575] In some embodiments, as shown in FIG18 or FIG19, the intra-frame prediction mode information may be a specific example of the low-frequency subband decoding unit outputting syntax elements to the prediction unit.

[0576] The prediction unit shown in Figure 18 is used to acquire syntax elements and perform corresponding prediction processing according to the syntax elements. For example, the prediction unit shown in Figure 18 can perform intra-frame prediction on the current block based on the low-frequency sub-band reconstruction block (i.e., the reconstruction data of the reference block corresponding to the intra-frame prediction mode of the block to be decoded).

[0577] Therefore, this application also provides a decoding method, which may include the following steps:

[0578] Step 1: Obtain the reconstruction data of the reference block corresponding to the block to be decoded in the image to be decoded.

[0579] In some embodiments, the reconstructed data of the reference block is obtained based on the bitstream of the image to be decoded.

[0580] The reference block is the CCLM mode-corresponding reference block for the current block in the image to be decoded. This reference block has been decoded before the current block, and its reconstructed data can be obtained from the cache.

[0581] Step 2: Obtain the parameters of CCLM based on the reconstruction data of the reference block.

[0582] Because this method involves intra-frame prediction, the block to be encoded and the reference block belong to the image to be encoded.

[0583] CCLM can be represented by linear model 1: U=α1Y+β1 and linear model 2: V=α2Y+β2.

[0584] In the two linear models above, Y is the reconstructed value (also called reconstructed data) of the Y component of the current block, U in linear model 1 is the predicted value of the U component of the current block, and V in linear model 2 is the predicted value of the V component of the current block.

[0585] In other words, the intra-frame prediction principle of CCLM mode can be to use the reconstructed value of the Y component of the current block to obtain the predicted value of the UV component of the current block, thereby realizing the prediction of the UV component of the current block.

[0586] Both linear models of CCLM require two corresponding parameters. This step can obtain the two parameters of each of the two linear models based on the reconstruction value of the reference block corresponding to the current block in the current frame, so as to obtain the parameters of CCLM.

[0587] Step 3: Based on the parameters of the cross-component linear model, perform intra-frame prediction on the block to be decoded to obtain the prediction data of the block to be decoded.

[0588] The parameters of CCLM are determined by M luminance values ​​derived from the reconstructed data based on the reference block, where M is an integer greater than 2.

[0589] Based on the parameters α1, β1, α2, and β2 of the two linear models mentioned above, intra-frame prediction of the chromaticity components (UV) of the current block can be performed using the CCLM mode to obtain the predicted values ​​of the UV components of the current block. Furthermore, the predicted value of the luminance component of the current block can also be obtained using the intra-frame prediction mode corresponding to the luminance component of the current block, thus obtaining the prediction data for the current block.

[0590] To obtain the parameters of CCLM, the reconstructed data of the reference block can be used to derive M (i.e., 3 or more) luminance values.

[0591] Then, the parameters of the CCLM mode can be obtained by using the three or more brightness values ​​obtained from the derivation.

[0592] In some embodiments, some or all of the three or more brightness values ​​may not be brightness values ​​directly selected from the reconstructed data of the reference block (i.e., reconstructed brightness values), but rather brightness values ​​obtained after a series of calculations based on the reconstructed brightness values ​​of some sample points selected from the reference block.

[0593] Step 4: Based on the prediction data, decode the bitstream of the image to be decoded to obtain the reconstructed data of the block to be decoded.

[0594] This prediction data is also called the prediction value of the current block. For example, the residual of the current block obtained from decoding the bitstream can be added to the prediction value of the current block to obtain the reconstructed data of the block to be decoded.

[0595] For example, the image to be decoded may include multiple blocks of data. The bitstream can be decoded according to the decoding order of the block data to obtain the reconstructed data of the current block.

[0596] In an optional embodiment, Figure 20b shows multiple blocks of data in the image to be decoded, arranged in the encoding order as block 1->block 2->block 3->block 4->block 5->block 6->block 7>block 8->block 9. Therefore, the decoding order could also be block 1->block 2->block 3->block 4->block 5->block 6->block 7>block 8->block 9.

[0597] In some embodiments, the size of each block is 8*8. In some embodiments, the size of the block can also be 8*4 or 4*4, and there is no limitation here.

[0598] The strategy for selecting a reference block for the current block within the current frame also differs under different intra-prediction modes.

[0599] In CCLM mode, the reference block corresponding to the current block in CCLM mode can be the left-hand adjacent block of the current block, and / or the top-hand adjacent block of the current block.

[0600] Taking Figure 20b as an example, when the current block is block 5, for example, if the chroma component of block 5 is intra-frame predicted in CCLM mode, then in the above embodiment, the reference block of block 5 can be at least one of block 2 located to the left of block 5 and block 4 located above block 5.

[0601] Similarly, the encoding order and decoding order can be the same, so the decoder can decode each block in the following order: block 1->block 2->block 3->block 4->block 5->block 6->block 7>block 8->block 9.

[0602] In this embodiment of the application, when performing intra-frame prediction on a block, if the intra-frame prediction mode is CCLM mode, the method of this embodiment can derive the values ​​of three or more luminance components (hereinafter referred to as luminance values) based on the reconstruction data of the reference block corresponding to the current block in the CCLM mode in order to obtain the parameters of the CCLM mode. This method has low complexity when performing intra-frame prediction, which can improve the efficiency of intra-frame prediction and thus improve the decoding efficiency.

[0603] In some embodiments, the decoder implemented by this decoding method can be adapted to a wavelet transform encoder. Thus, the current block on the decoding side can be block data from a low-frequency subband obtained during encoding via wavelet transform. Further implementation details regarding wavelet transform encoders can be found in the above introduction to wavelet transform-based encoders, and will not be repeated here.

[0604] Considering that the size of the low-frequency subband output by wavelet transform is only 1 / 4 of the image before wavelet transform (e.g., the image to be encoded), performing intra-frame prediction in CCLM mode on the block data in the low-frequency subband reduces the amount of data decoded compared to decoding the entire image to be decoded, thereby improving intra-frame prediction efficiency and decoding efficiency. Considering that the data distribution of the low-frequency subband is more uniform than that of the high-frequency subband output by wavelet transform, performing intra-frame prediction and decoding on the low-frequency subband using the method of this application embodiment allows the decoded low-frequency subband to be used for display. For example, the image to be displayed can be a thumbnail of the image to be decoded (e.g., the original image to be encoded).

[0605] In some embodiments, the current block on the decoding side can also be block data in the high-frequency subband obtained by wavelet transform. Intra-frame prediction and corresponding decoding are performed on the block data in the high-frequency subband in CCLM mode. This process is similar to the process of intra-frame prediction and decoding in CCLM mode on the block data in the low-frequency subband, and will not be described in detail here.

[0606] The decoding method of this application is described below with reference to different embodiments, based on any of the above embodiments.

[0607] In some embodiments, corresponding to FIG21a above, FIG27a exemplarily illustrates a flowchart of a decoding method of this application. As shown in FIG27a, the method flow may include, but is not limited to, the following steps:

[0608] S201. Based on the bitstream, determine the intra-frame prediction mode of the current block, the residual of the U component of the current block, and the residual of the V component of the current block.

[0609] In an optional embodiment, the decoder decodes the bitstream to obtain syntax elements indicating the intra-frame prediction mode, the residual of the U component of the current block, and the residual of the V component of the current block.

[0610] In some embodiments, the syntax element may be a field whose value indicates that the intra-prediction mode of the chroma component of the current block is CCLM mode.

[0611] In this embodiment, the intra-frame prediction mode of the current block is CCLM mode.

[0612] S202. Obtain the reconstruction value of the reference block and the reconstruction value of the Y component of the current block.

[0613] In an optional embodiment, the reference block is a sub-block that has already been decoded. After the decoder decodes the reference block, it can reconstruct the data of the decoded reference block to obtain the reconstructed value of the reference block. The decoder can read the reconstructed value of the reference block.

[0614] In an optional embodiment, the decoder may decode the Y component of the current block before decoding the UV components, thus obtaining the decoded Y component data of the current block. Then, the decoder may reconstruct the decoded Y component data of the current block to obtain the reconstructed Y component value of the current block, and store the reconstructed Y component value of the current block. In this step, the decoder can read the stored reconstructed Y component value of the current block.

[0615] In an optional embodiment, the decoder can determine the reconstructed value of the Y component of the current block based on the reconstructed value of the Y component of the reference block and the residual of the Y component of the current block.

[0616] In some embodiments, as shown in FIG27b, which is a schematic diagram of another decoding method provided by an embodiment of this application, the decoder decodes the bitstream and determines the intra-frame prediction mode for the Y component of the current block as a first mode and the residual of the Y component of the current block. After obtaining the reconstructed value of the Y component of the reference block, the decoder performs intra-frame prediction on the Y component of the current block based on the first mode to obtain the predicted value of the Y component of the current block. The decoder obtains the reconstructed value of the Y component of the current block based on the residual of the Y component of the current block and the predicted value of the Y component of the current block. For example, the decoder determines the reconstructed value of the Y component of the current block as the sum of the residual and the predicted value.

[0617] Returning to Figure 27a, after S201 and S202, the method may further include:

[0618] S203. Based on the reconstructed value of the reference block and the reconstructed value of the Y component of the current block, perform intra-frame prediction on the UV component of the current block using CCLM mode to obtain the predicted value of the U component and the predicted value of the V component of the current block.

[0619] The specific implementation process of S203 is the same as that of S103 in Example A2.

[0620] S204. Based on the residuals of the U component of the current block, the residuals of the V component of the current block, the predicted values ​​of the Y component of the current block, and the predicted values ​​of the U component of the current block, determine the reconstructed values ​​of the U component and the V component of the current block.

[0621] In an optional embodiment, the decoder determines the reconstructed value of the U component of the current block based on the residual of the U component of the current block and the predicted value of the U component of the current block. For example, the decoder determines the reconstructed value of the U component of the current block as the sum of the residual of the U component of the current block and the predicted value of the U component of the current block.

[0622] In an optional embodiment, the decoder determines the reconstructed value of the V component of the current block based on the residual of the V component of the current block and the predicted value of the V component of the current block. For example, the decoder determines the reconstructed value of the V component of the current block as the sum of the residual of the V component of the current block and the predicted value of the V component of the current block.

[0623] It should be noted that the embodiments of this application do not restrict the execution order between S202 and S204. That is, the embodiments of this application do not limit the execution order between the process of obtaining the reconstruction value of the Y component of the reference block and the process of obtaining the reconstruction values ​​of the U component and V component of the current block.

[0624] S205. Determine the reconstruction value of the current block based on the reconstruction values ​​of the Y component, the U component, and the V component of the current block.

[0625] In an optional embodiment, the decoder reconstructs the reconstructed values ​​of the Y component, the U component, and the V component of the current block to obtain the reconstructed value of the current block.

[0626] The decoder can also implement any of the following embodiments.

[0627] In some embodiments, the parameters of CCLM are obtained based on the index of a lookup table, which is based on the M brightness values ​​mentioned above.

[0628] In some embodiments, the index of the lookup table can be obtained based on the above M brightness values, and then the lookup table can be used based on the index (e.g., to find the entry corresponding to the index from the lookup table, specifically α as described above) to obtain the parameters of CCLM.

[0629] In some embodiments, the index is obtained based on the relative positional relationship between the M brightness values.

[0630] In some embodiments, the index can be obtained based on the relative positional relationship between the M luminance values ​​derived from the derivation.

[0631] In some embodiments, the index may be obtained based on the weight of each of the M brightness values ​​and the M brightness values, wherein the weight of each of the M brightness values ​​is obtained based on the relative positional relationship between the M brightness values.

[0632] In some embodiments, the weight of each of the M brightness values ​​can be obtained based on the relative positional relationship between the M brightness values; then, an index can be obtained based on the weight of each of the M brightness values ​​and the M brightness values ​​(for example, by weighting the brightness values ​​and the weights to obtain the index).

[0633] In some embodiments, the index is obtained by weighting the M brightness values ​​and the M brightness values ​​together.

[0634] In some embodiments, an index can be obtained by performing a weighted calculation based on the weight of each of the M brightness values ​​and the M brightness values ​​mentioned above.

[0635] In some embodiments, the value of M for the M brightness values ​​is 3.

[0636] In some embodiments, when M of the M brightness values ​​is 3, the M brightness values ​​are represented by Y1, Y2, and Y3, where Y1 ≤ Y2 ≤ Y3, the distance between Y2 and Y1 is greater than or equal to the distance between Y2 and Y3, and the index is obtained based on Y1, Y2, Y3 and the weights w1 of Y1, w2 of Y2, and w3 of Y3, where w1 < w2 < w3, w1 is a negative integer, and w2 and w3 are both positive integers.

[0637] In some embodiments, w1 = -3, w2 = 1, w3 = 2.

[0638] Thus, the index = -3*Y1 + Y2 + 2*Y3, where "*" represents multiplication.

[0639] In some embodiments, when M of the M brightness values ​​is 3, the M brightness values ​​are represented by Y1, Y2, and Y3, where Y1≤Y2≤Y3, the distance between Y2 and Y1 is less than or equal to the distance between Y2 and Y3, and the index is obtained based on Y1, Y2, Y3 and the weights w1 of Y1, w2 of Y2, and w3 of Y3, where w1<w2<w3, w1 and w2 are both negative integers, and w3 is a positive integer.

[0640] In some embodiments, w1 = -2, w2 = -1, w3 = 3.

[0641] Thus, the index = -2*Y1-Y2+3*Y3, where "*" represents multiplication.

[0642] In some embodiments, the M luminance values ​​are derived from the N1 luminance values ​​in the reconstructed data of the reference block, where N1 ≥ M and N1 is an integer.

[0643] In some embodiments, N1 luminance values ​​can be selected based on the reconstructed data of the reference block described above. Then, M luminance values ​​can be derived based on the N1 luminance values, where M is an integer greater than or equal to 3.

[0644] In some embodiments, the sample points corresponding to the N1 luminance values ​​in the reference block are adjacent to the block to be decoded.

[0645] In some embodiments, the lookup table has less than 2 P The size of the array is P, where P is a positive integer.

[0646] The larger the value of P, the higher the decoding complexity and the higher the decoding accuracy. The smaller the value of P, the lower the decoding complexity and the lower the accuracy. The value of P can be flexibly set according to the needs of the application scenario.

[0647] In some embodiments, P = 4. This allows for image or video encoding and decoding with low complexity while ensuring high encoding and decoding accuracy.

[0648] In some embodiments, the lookup table includes the data from Table 1.

[0649] In some embodiments, the lookup table can be represented as TABLE[index].

[0650] Where TABLE[index] = {63,63,61,57,54,51,49,47,45,43,41,39,38,37,35,34}, and 0 ≤ index < 16.

[0651] The data structure of TABLE[index] can be a table or an array.

[0652] In some embodiments, the prediction data for the block to be decoded is obtained based on the reconstructed value of the luminance value of the block to be decoded and the parameters of CCLM.

[0653] In some embodiments, intra-frame prediction can be performed based on the reconstructed data of the luminance values ​​of the block to be decoded and the parameters of the CCLM obtained above, so as to obtain the prediction data of the current block.

[0654] In some embodiments, the reference block is the adjacent block of the block to be decoded in the image to be decoded.

[0655] Regarding the specific implementation process of S203 in Example B2, you can refer to any embodiment of Example A2 in which the encoding side uses the reconstructed value of the reference block to perform inter-frame prediction on the UV components of the current block to obtain the predicted value of the UV components of the current block. For example, you can refer to any embodiment of S103 in Figure 21a, Figure 22, Figure 23, Figures 24a to 24c, Figures 25a to 25c, Figure 26, etc. The only difference is that the current block on the encoding side is the block to be encoded, while the current block on the decoding side is the block to be decoded. It will not be described again here.

[0656] In the above embodiments, the encoding and decoding methods of the embodiments of this application are mostly illustrated using the current block as an example. In other embodiments, the encoding and decoding of the various embodiments of this application can also be implemented by taking sub-graphs or images as units. The methods are similar and will not be described again here.

[0657] Based on the same concept as the above method, as shown in FIG28a, this application embodiment also provides a decoding device 1000, which includes: an acquisition module 1001, used to acquire the reconstruction data of the reference block corresponding to the block to be decoded in the image to be decoded; a prediction module 1002, used to obtain the parameters of CCLM based on the reconstruction data of the reference block; the prediction module 1002 is also used to perform intra-frame prediction on the block to be decoded based on the parameters of CCLM to obtain the prediction data of the block to be decoded, wherein the parameters of CCLM are determined by using M luminance values ​​derived based on the reconstruction data of the reference block, where M is an integer greater than 2; and a decoding module 1003, used to decode the bitstream of the image to be decoded based on the prediction data to obtain the reconstruction data of the block to be decoded.

[0658] Based on the same concept as the above method, as shown in FIG28b, this application embodiment also provides an encoding device 2000, which includes an acquisition module 2001 for acquiring the reconstruction data of the reference block corresponding to the block to be encoded in the image to be encoded; a prediction module 2002 for obtaining the parameters of CCLM based on the reconstruction data of the reference block; the prediction module 2002 is further used to perform intra-frame prediction on the block to be encoded based on the parameters of CCLM to obtain the prediction data of the block to be encoded, wherein the parameters of CCLM are determined by using M luminance values ​​derived from the reconstruction data of the reference block, where M is an integer greater than 2; and an encoding module 2003 for encoding the block to be encoded based on the prediction data to obtain the bitstream of the image to be encoded.

[0659] The prediction module 1002 of the aforementioned decoding device can be applied to the intra-frame prediction process at the decoding end. Specifically, at the decoding end, the prediction module 1002 can be applied to the intra-frame prediction unit or prediction unit of the aforementioned decoder.

[0660] The prediction module 2002 of the aforementioned encoding apparatus can be applied to the intra-frame prediction process at the encoding end. Specifically, at the encoding end, the prediction module 2002 can be applied to the intra-frame prediction unit or prediction unit of the aforementioned encoder.

[0661] The specific implementation process of the encoding device and the decoding device can be referred to the relevant descriptions of the encoding method and decoding method embodiments, and will not be repeated here for the sake of brevity.

[0662] The steps of the methods or algorithms described in conjunction with the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0663] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0664] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.

Claims

1. A decoding method, characterized in that, The method includes: Obtain the reconstructed data of the reference block corresponding to the block to be decoded in the image to be decoded; Based on the reconstructed data of the reference block, the parameters of the cross-component linear model are obtained; Based on the parameters of the cross-component linear model, intra-frame prediction is performed on the block to be decoded to obtain the prediction data of the block to be decoded. The parameters of the cross-component linear model are determined by using M luminance values ​​derived from the reconstructed data of the reference block, where M is an integer greater than 2. Based on the predicted data, the bitstream of the image to be decoded is decoded to obtain the reconstructed data of the block to be decoded.

2. The method according to claim 1, characterized in that, The parameters of the cross-component linear model are obtained based on the index of a lookup table, which is based on the M luminance values.

3. The method according to claim 2, characterized in that, The index is obtained based on the relative positional relationship between the M brightness values.

4. The method according to claim 3, characterized in that, The index is obtained based on the weight of each of the M brightness values ​​and the M brightness values ​​themselves. The weight of each of the M brightness values ​​is obtained based on the relative positional relationship between the M brightness values.

5. The method according to claim 4, characterized in that, The index is obtained by weighting the M brightness values ​​together with their respective weights.

6. The method according to any one of claims 1 to 5, characterized in that, M is 3.

7. The method according to claim 4 or 5, characterized in that, The M brightness values ​​are represented by Y1, Y2, and Y3, where Y1 ≤ Y2 ≤ Y3, and the distance between Y2 and Y1 is greater than or equal to the distance between Y2 and Y3. The index is obtained based on Y1, Y2, Y3, and the weights w1 of Y1, w2 of Y2, and w3 of Y3, where w1 < w2 < w3, and w1 is a negative integer, while w2 and w3 are both positive integers.

8. The method according to claim 7, characterized in that, w1 = -3, w2 = 1, w3 = 2.

9. The method according to claim 4 or 5, characterized in that, The M brightness values ​​are represented by Y1, Y2, and Y3, where Y1 ≤ Y2 ≤ Y3, and the distance between Y2 and Y1 is less than or equal to the distance between Y2 and Y3. The index is obtained based on Y1, Y2, Y3, and the weights w1 of Y1, w2 of Y2, and w3 of Y3, where w1 < w2 < w3, and w1 and w2 are negative integers, while w3 is a positive integer.

10. The method according to claim 9, characterized in that, w1 = -2, w2 = -1, w3 = 3.

11. The method according to any one of claims 2 to 10, characterized in that, The M luminance values ​​are derived from the N luminance values ​​in the reconstructed data of the reference block, where N ≥ M and N is an integer.

12. The method according to claim 11, characterized in that, The sample points corresponding to the N brightness values ​​in the reference block are adjacent to the block to be decoded.

13. The method according to any one of claims 2 to 12, characterized in that, The lookup table has less than 2 P The size of the array is P, where P is a positive integer.

14. The method according to claim 13, characterized in that, P=4。 15. The method according to any one of claims 2 to 14, characterized in that, The lookup table includes the following data:

16. The method according to any one of claims 1 to 15, characterized in that, The prediction data for the block to be decoded is obtained based on the reconstructed luminance values ​​of the block to be decoded and the parameters of the cross-component linear model.

17. The method according to any one of claims 1 to 16, characterized in that, The reference block is the adjacent block of the block to be decoded in the image to be encoded.

18. An encoding method, characterized in that, The method includes: Obtain the reconstructed data of the reference block corresponding to the block to be encoded in the image to be encoded; Based on the reconstructed data of the reference block, the parameters of the cross-component linear model are obtained; Based on the parameters of the cross-component linear model, intra-frame prediction is performed on the block to be coded to obtain the prediction data of the block to be coded. The parameters of the cross-component linear model are determined by using M luminance values ​​derived from the reconstructed data of the reference block, where M is an integer greater than 2. Based on the predicted data, the block to be encoded is encoded to obtain the bitstream of the image to be encoded.

19. The method according to claim 18, characterized in that, The parameters of the cross-component linear model are obtained based on the index of a lookup table, which is based on the M luminance values.

20. The method according to claim 19, characterized in that, The index is obtained based on the relative positional relationship between the M brightness values.

21. The method according to claim 20, characterized in that, The index is obtained based on the weight of each of the M brightness values ​​and the M brightness values ​​themselves. The weight of each of the M brightness values ​​is obtained based on the relative positional relationship between the M brightness values.

22. The method according to claim 21, characterized in that, The index is obtained by weighting the M brightness values ​​together with their respective weights.

23. An encoder, characterized in that, include: A memory and a processor, wherein the memory is coupled to the processor; The memory stores program instructions that, when executed by the processor, cause the processor to perform the steps of the method as described in any one of claims 18 to 22.

24. A decoder, characterized in that, include: A memory and a processor, wherein the memory is coupled to the processor; The memory stores program instructions that, when executed by the processor, cause the processor to perform the steps of the method as described in any one of claims 1 to 17.

25. An encoder, characterized in that, include: A processing circuit that implements the steps of the method as claimed in any one of claims 18 to 22.

26. A decoder, characterized in that, include: A processing circuit that implements the steps of the method as claimed in any one of claims 1 to 17.

27. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed on a computer or processor, causes the computer or processor to perform the method as described in any one of claims 1 to 17, or causes the computer or processor to perform the method as described in any one of claims 18 to 22.

28. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a computer or processor, cause the steps of the method as described in any one of claims 1 to 17 to be performed, or cause the steps of the method as described in any one of claims 18 to 22 to be performed.

29. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a bitstream generated according to the method described in any one of claims 18 to 22.