Image encoding method, image decoding method, and related apparatus and system

By determining a matching MIP size identifier to eliminate upsampling, the method addresses the computational resource issues in MIP mode, enhancing efficiency in image and video encoding and decoding.

JP7857857B2Active Publication Date: 2026-05-13GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
Filing Date
2019-12-10
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

The current Matrix-based Intra Prediction (MIP) mode in video compression requires significant computational resources and storage space due to the need for upsampling processes to match prediction block sizes with current block sizes.

Method used

The method determines an appropriate MIP size identifier to ensure the MIP prediction block matches the current block size, eliminating the need for upsampling and reducing computational complexity by deriving intra-prediction blocks directly.

Benefits of technology

This approach significantly reduces computational complexity and improves efficiency in image and video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007857857000011
    Figure 0007857857000011
  • Figure 0007857857000012
    Figure 0007857857000012
  • Figure 0007857857000013
    Figure 0007857857000013
Patent Text Reader

Abstract

The present invention relates to a method for encoding an image, and in some embodiments, the method includes: (i) determining a width and a height of a coding block in the image; (ii) determining a matrix-based intra-prediction (MIP) size identifier if the width and height are equal to N, where the MIP size identifier indicates that the MIP prediction size is equal to N, where N is a positive integer power of 2; (iii) deriving a group of reference samples for the coding block; and (iv) deriving a MIP prediction value for the coding block based on the group of reference samples and an MIP matrix corresponding to the MIP size identifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of telecommunication technologies, and particularly to an image (such as picture or video) encoding method and an image decoding method.

Background Art

[0002] Versatile Video Coding (VVC) is the next-generation video compression standard that replaces current standards such as the High Efficiency Video Coding (H.265 / HEVC) standard. Compared with current standards, the VVC coding standard provides higher coding quality. To achieve this goal, various intra prediction and inter prediction modes have been investigated. By using these prediction modes, videos can be compressed, thereby reducing the data transmitted in the bitstream (binary format). Matrix-based Intra Prediction (MIP) is one of the above modes. MIP is an intra prediction mode. When implemented under the MIP mode, an encoder (or coder) or decoder can derive an intra prediction block based on the current block (e.g., a group of bits or numbers that are transmitted as a unit and encoded and / or decoded together). However, deriving such a prediction block requires a large amount of computing resources and additional storage space. Therefore, an improved method to solve this problem is desirable and necessary.

Summary of the Invention

[0003] Under the current MIP mode, the size of the prediction block is smaller than the size of the current block in order to generate a prediction block for the current block. For example, an 8x8 current block can have a 4x4 prediction block. Under the current MIP mode, a MIP prediction block smaller than the current block is derived by performing a matrix calculation. This matrix calculation consumes fewer computational resources compared to performing the matrix calculation using a larger block. After the matrix calculation, an upsampling process is performed on the MIP prediction block to derive an intra-prediction block that is the same size as the current block. For example, an 8x8 intra-prediction block can be derived from a 4x4 MIP prediction block by calling an interpolation and / or extrapolation upsampling process. The present invention provides a method for implementing the MIP mode without performing an upsampling process, thereby significantly reducing computational complexity and improving overall efficiency. More specifically, when implementing MIP mode, this method determines an appropriate size identifier (or MIP size identifier) ​​so that the size of the MIP prediction block (e.g., "8x8") is the same as the size of the current block ("8x8"). This eliminates the need for an upsampling process.

[0004] In embodiments of the present invention, an image encoding method is provided. The method can further be applied to encoding a video comprising a sequence of images. The method includes, for example, (i) determining the width and height of an encoding block in an image (e.g., an encoding block), (ii) determining an intra-predicted (MIP) size identifier based on a matrix if the width and height are equal to N (where N is a positive integer power of 2), the MIP size identifier indicating that the MIP predicted size is equal to N, (iii) deriving a group of reference samples for the encoding block (e.g., by utilizing adjacent samples of the encoding block), (iv) deriving a MIP predicted value for the encoding block using the group of reference samples and a MIP weighting matrix according to the MIP size identifier, and (v) setting the predicted value for the encoding block equal to the MIP predicted value for the encoding block. In some embodiments, the method further includes generating a bitstream based on the predicted value for the encoding block.

[0005] In another aspect of the present invention, an image decoding method is provided. The method may include the following: for example, (a) determining the width, height and prediction mode (e.g., whether the bitstream indicates that the MIP mode was used) of a decoding block (e.g., a decoding block) by analyzing a bitstream; (b) determining a MIP size identifier if the width and height are equal to N and the MIP mode was used, the identifier indicating that the predicted MIP size is equal to N (where N is a positive integer power of 2); (c) deriving a group of reference samples of the decoding block (e.g., by using adjacent samples of the decoding block); (d) deriving a predicted MIP value of the decoding block based on the group of reference samples and the MIP matrix according to the MIP size identifier; and (e) setting the predicted value of the decoding block to be equal to the predicted MIP value of the decoding block.

[0006] In some embodiments, the MIP prediction values ​​may include values ​​from an "N×N" prediction sample (e.g., "8×8"). In some embodiments, the MIP matrix may be selected from a predefined group of MIP matrices.

[0007] In another aspect of the present invention, a system for encoding / decoding images and video is included. The system includes an encoding subsystem (or encoder) and a decoding subsystem (or decoder). The encoding subsystem includes a partitioning unit, first Intra The subsystem includes a prediction unit and an entropy coding unit. The splitting unit is configured to receive the input video and split the input video into one or more coding units (CUs). The first intra-prediction unit is configured to generate prediction blocks corresponding to each CU and MIP size identifiers derived from the coding to the input video. The entropy coding unit is configured to transform parameters for writing the predicted values ​​to a bitstream. The decoding subsystem includes an analysis unit and a second intra-prediction unit. The analysis unit is configured to obtain numerical values ​​(e.g., values ​​associated with one or more CUs) by analyzing the bitstream. The second intra-prediction unit is configured to transform the above numerical values ​​into an output video, at least in part, based on the MIP size identifiers.

[0008] The width and height of the CU may be equal to N, where N is a positive integer raised to the power of 2. The MIP size identifier indicates that the MIP prediction size used by the first intra-prediction unit to generate the MIP prediction block is N. A MIP size identifier equal to "2" indicates that the MIP prediction size is "8 × 8". [Brief explanation of the drawing]

[0009] To further clarify the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments are briefly introduced below. Clearly, the drawings described below are used solely for the purpose of describing the present invention and are not intended to limit the present invention. Those skilled in the art can obtain other drawings from these drawings without any creative effort. [Figure 1] Figure 1 is a schematic diagram showing a system according to an embodiment of the present invention. [Figure 2] Figure 2 is a schematic diagram showing an encoding system according to an embodiment of the present invention. [Figure 3] Figure 3 is a schematic diagram showing the derivation of an intra-prediction block utilizing the MIP mode according to an embodiment of the present invention. [Figure 4] Figure 4 is a schematic diagram showing a decoding system according to an embodiment of the present invention. [Figure 5] Figure 5 is a schematic diagram showing an apparatus (e.g., an encoder) according to an embodiment of the present invention. [Figure 6] Figure 6 is a schematic diagram showing an apparatus (for example, a decoder) according to an embodiment of the present invention. [Figure 7] Figure 7 is a schematic diagram showing a communication system according to an embodiment of the present invention. [Modes for carrying out the invention]

[0010] To facilitate understanding of the present invention, the present invention will be described in more detail below with reference to the accompanying drawings.

[0011] Figure 1 is a schematic diagram showing a system 100 according to an embodiment of the present invention. The system 100 can encode, transmit, and decode images. The system 100 can also be applied to encode, transmit, and decode videos containing sequences of images. More specifically, the system 100 can receive an input image, process the input image, and generate an output image. The system 100 includes an encoding device 100a and a decoding device 100b. The encoding device 100a includes a splitting unit 101, a first intra-prediction unit 103, and an entropy encoding unit 105. The decoding device 100b includes an analysis unit 107 and a second intra-prediction unit 109.

[0012] The splitting unit 101 is configured to receive the input video 10 and then split the input video 10 into one or more coding tree units (CTUs) or coding units (CUs) 12. The CUs 12 are transmitted to the first intra-prediction unit 103. The first intra-prediction unit 103 is configured to derive a prediction block for each CU 12 by performing an MIP process. Based on the size of the CU 12, the MIP process has different approaches to handle CUs 12 of different sizes. Each type of CU 12 has a designated MIP size identifier (e.g., 0, 1, 2, etc.). The MIP size identifier is used to derive the size of the MIP prediction block (i.e., the variable "predSize"), the number of reference samples from the upper or left boundary of the CU (i.e., the variable "boundarySize"), and to select an MIP matrix from a plurality of predefined MIP matrices. For example, if the MIP size identifier is "0", the size of the MIP prediction block is "4x4" (for example, "predSize" is set to equal 4) and "boundarySize" is set to equal 2. If the MIP size identifier is "1", "predSize" is set to equal 4 and "boundarySize" is set to equal 2. If the MIP size identifier is "2", "predSize" is set to equal 8 and "boundarySize" is set to equal 4.

[0013] The first intra-prediction unit 103 first determines the width and height of CU12. For example, the first intra-prediction unit 103 can determine that the height of CU12 is "8" and the width is "8". In this example, the width and height are "8". Accordingly, the first intra-prediction unit 103 determines that the MIP size identifier of CU12 is "2", which indicates that the size of the MIP prediction is "8 × 8". The first intra-prediction unit 103 further derives a group of reference samples for CU12 (for example, by utilizing the adjacent samples of CU12, such as the above-neighboring sample or the left-neighboring sample, as detailed in Figure 3). Next, the first intra-prediction unit 103 derives the MIP prediction value for CU12 based on the group of reference samples and the corresponding MIP matrix. The first intra-prediction unit 103 can use the MIP prediction value as the intra-prediction value 14 for CU12. The intra-prediction value 14 and the parameters for deriving the intra-prediction value 14 are sent to the entropy coding unit 105 for further processing.

[0014] The entropy coding unit 105 is configured to convert the parameters for deriving the intra-predicted value 14 into binary form. Accordingly, the entropy coding unit 105 generates a bitstream 16 based on the intra-predicted value 14. In some embodiments, the bitstream 16 can be transmitted over a communication network or stored on disk or a server.

[0015] The decoding device 100b receives bitstream 16 as input bitstream 17. The analysis unit 107 analyzes input bitstream 17 (in binary format) and converts it into numerical values ​​18. These numerical values ​​18 indicate the characteristics (color, brightness, depth, etc.) of the input video 10. The numerical values ​​18 are transmitted to the second intra-prediction unit 109. The second intra-prediction unit 109 can then convert these numerical values ​​18 into output video 19 (based on processing similar to that performed by the first intra-prediction unit 103, for example; relevant embodiments are described in detail with reference to Figure 4). The output video 19 can then be stored, transmitted, and / or rendered by an external device (e.g., storage, transmitter, etc.). The stored video can then be displayed on a display.

[0016] Figure 2 is a schematic diagram showing an encoding system 200 according to an embodiment of the present invention. The encoding system 200 is configured to encode, compress, and / or process an input image 20 to produce an output bitstream 21 in binary format. The encoding system 200 includes a splitting unit 201 configured to split the input image 20 into one or more encoding tree units (CTUs) 22. In some embodiments, the splitting unit 201 can split the image into slices, tiles, and / or bricks. Each brick can contain one or more whole and / or partial CTUs 22. In some embodiments, the splitting unit 201 can also form one or more subimages, each subimage can contain one or more slices, tiles, or bricks. The splitting unit 201 sends the CTUs 22 to a prediction unit 202 for further processing.

[0017] The prediction unit 202 is configured to generate prediction blocks 23 for each CTU 22. The prediction blocks 23 can be generated based on one or more inter or intra prediction methods by utilizing various interpolation and / or extrapolation schemes. As shown in Figure 2, the prediction unit 202 may further include a block partitioning unit 203, a motion estimation (ME) unit 204, a motion compensation (MC) unit 205, and an intra prediction unit 206. The block partitioning unit 203 is configured to partition the CTU 22 into smaller coding units (CUs) or coding blocks (CBs). In some embodiments, CUs can be generated from the CTU 22 via various methods such as quadtree partitioning, binary partitioning, and ternary partitioning. The ME unit 204 is configured to estimate changes caused by the motion of an object shown in the input image 20 or by the motion of an image capture device generating the input image 20. The MC unit 205 is configured to adjust and compensate for the changes caused by the above motion. Both the ME unit 204 and the MC unit 205 are configured to derive interpretation blocks of the CU (for example, at different time points). In some embodiments, the ME unit 204 and the MC unit 205 can utilize a rate-distortion optimized motion estimation method to derive interpretation blocks.

[0018] The intra-prediction unit 206 is configured to derive intra-prediction blocks (e.g., same-time) of the CU (or part of the CU) using various intra-prediction modes, including the MIP mode. Details of the deriving of intra-prediction blocks using the MIP mode (hereinafter referred to as the "MIP process") will be explained with reference to Figure 3. During the MIP process, the intra-prediction unit 206 first derives one or more reference samples from the CU's neighboring samples by, for example, directly using neighboring samples as reference samples, downsampling neighboring samples, or directly extracting from neighboring samples (e.g., step 301 in Figure 3).

[0019] Next, the intra-prediction unit 206 uses the reference sample, the MIP matrix, and the shifting parameter to derive predicted samples at multiple sample locations within the CU. The sample locations can be pre-defined sample locations within the CU. For example, the sample locations can be locations within the CU where the horizontal and vertical coordinate values ​​are odd (e.g., x=1, 3, 5, etc., y=1, 3, 5, etc.). The shifting parameter includes a shifting offset parameter and a shifting number parameter, which can be used for the shift operation when generating the predicted samples. With this arrangement, the intra-prediction unit 206 can generate predicted samples in the CU (i.e., "MIP prediction" or "MIP prediction block" refers to a set of such predicted samples) (e.g., step 302 in Figure 3). In some embodiments, the sample locations can be locations within the CU where the horizontal and vertical coordinate values ​​are even.

[0020] Thirdly, the intra prediction unit 206 can derive prediction samples at the remaining positions of the CUs (for example, they are not sample positions) (for example, step 303 in FIG. 3). In some embodiments, the intra prediction unit 206 can derive prediction samples at the remaining positions by using an interpolation filter. Through the above process, the intra prediction unit 206 can generate a prediction block 23 of the CU in the CTU 22.

[0021] Referring to FIG. 2, the prediction unit 202 outputs a prediction block 23 to the adder 207. The adder 207 calculates the difference (for example, the residual R) between the output of the splitting unit 201 (for example, the CU in the CTU 22) and the output of the prediction unit 202 (that is, the prediction block 23 of the CU). The conversion unit 208 reads out the residual R and performs one or more conversion operations on the prediction block 23 to obtain the coefficient 24 for further use. The quantization unit 209 can quantize the coefficient 24 and output a quantized coefficient 25 (for example, a level) to the inverse quantization unit 210. The inverse quantization unit 210 performs an inverse quantization operation on the quantized coefficient 25 to output a reconstructed coefficient 26 to the inverse conversion unit 211. The inverse conversion unit 211 performs one or more inverse conversions corresponding to the conversion in the conversion unit 208 and outputs a reconstructed residual 27.

[0022] Next, adder 212 calculates the reconstructed CU by adding reconstruction residual 27 and the predicted block 23 of the CU from prediction unit 202. Adder 212 also transfers output 28 to prediction unit 202 so that output 28 is used as an intra-prediction reference. After all CUs in CTU 22 are reconstructed, filtering unit 213 can perform in-loop filtering on the reconstructed image 29. Filtering unit 213 includes one or more filters for suppressing coding distortion or enhancing the coding quality of the image. Examples of filters include a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a luma mapping with chroma scaling (LMCS) filter, a neural-network-based filter, and other suitable filters.

[0023] Next, filtering unit 213 can send the decoded image 30 (or sub-image) to the decoded picture buffer (DPB) 214. DPB 214 outputs the decoded image 31 based on control information. Also, the image 31 stored in DPB 214 can be used as a reference image for prediction unit 202 to perform inter-prediction or intra-prediction.

[0024] Entropy encoding unit 215 is configured to convert the image 31, parameters from units in encoding system 200, and supplementary information (e.g., information for controlling or communicating with system 200) into binary format. Entropy encoding unit 215 can accordingly generate output bitstream 21.

[0025] In some embodiments, the encoding system 200 may be a computing device comprising a processor and a storage medium having one or more encoding programs. When the processor reads and executes an encoding program, the encoding system 200 can receive an input image 20 and generate an output bitstream 21 accordingly. In some embodiments, the encoding system 200 may be a computing device comprising one or more chips. Units or elements of the encoding system 200 may be implemented as integrated circuits on a chip.

[0026] Figure 3 is a schematic diagram showing an MIP process according to an embodiment of the present invention. The MIP process can be implemented by an intra-prediction unit (e.g., intra-prediction unit 206). As shown in Figure 3, the intra-prediction unit may include a prediction module 301 and a filtering module 302. As shown in Figure 3, the MIP process includes three steps: 301, 302, and 303. The MIP process can generate a prediction block based on the current block or coded block 300 (such as a CU or a portion of a CU).

[0027] Step 301

[0028] In step 301, the intra-prediction unit can generate reference samples 32 and 34 using the adjacent samples 31 and 33 of the coding block 300. In the shown embodiment, adjacent sample 31 is the upper adjacent sample, and adjacent sample 33 is the left adjacent sample. The intra-prediction unit 206 can calculate the average value of the values ​​for every two adjacent samples 31 and the average value of the values ​​for every two adjacent samples 33, and set these average values ​​as the values ​​for reference samples 32 and 34, respectively. In some embodiments, the intra-prediction unit 206 can select one value from two adjacent samples 31 or one value from two adjacent samples 33 as the value for reference sample 32 or 34. In the shown embodiment, the intra-prediction unit 206 derives four reference samples 32 from the eight upper adjacent samples 31 of the coding block 300, and derives another four reference samples 34 from the eight left adjacent samples 33 of the coding block 300.

[0029] In step 301, the intra-prediction unit determines the width and height of the encoded block 300, denoting them as the variables "cbWidth" and "cbHeight," respectively. In some embodiments, the intra-prediction unit 206 may utilize a rate-distortion optimization mode determination process to determine the intra-prediction mode (e.g., whether the MIP mode is used). In such embodiments, the encoded block 300 may be divided into one or more transform blocks, the width and height of these transform blocks being denoted as the variables "nTbW" and "nTbH," respectively. If the MIP mode is used as the intra-prediction mode, the intra-prediction unit determines the MIP size identifier (denoted as the variable "mipSizeId") based on the following conditions A to C.

[0030] [Condition A] If both "nTbW" and "nTbH" are 4, "mipSizeId" is set to 0.

[0031] [Condition B] Otherwise, and if either "cbWidth" or "cbHeight" is 4, "mipSizeId" is set to 1.

[0032] [Condition C] Otherwise, "mipSizeId" is set to 2.

[0033] For example, if the size of the encoding block 300 is "8x8" (i.e., both "cbWidth" and "cbHeight" are 8), then "mipSizeId" is set to 2. As another example, if the size of the transform block of encoding block 300 is "4x4" (i.e., both "nTbW" and "nTbH" are 4), then "mipSizeId" is set to 0. As yet another example, if the size of encoding block 300 is "4x8", then "mipSizeId" is set to 1.

[0034] In the embodiment shown, there are three types of "mipSizeId": "0", "1", and "2". Each type of MIP size identifier (i.e., the variable "mipSizeId") corresponds to a specific method of performing the MIP process (e.g., using different MIP matrices). In other embodiments, there may be more than three types of MIP size identifiers.

[0035] Based on the MIP size identifier, the intra-prediction unit can determine the variables "boundarySize" and "predSize" based on the following conditions D to F.

[0036] [Condition D] If "mipSizeId" is 0, "boundarySize" is set to 2 and "predSize" is set to 4.

[0037] [Condition E] If "mipSizeId" is 1, "boundarySize" is set to 4 and "predSize" is set to 4.

[0038] [Condition F] If "mipSizeId" is 2, "boundarySize" is set to 4 and "predSize" is set to 8.

[0039] In the embodiment shown, "boundarySize" represents the number of reference samples 32, 34 derived from the upper adjacent sample 31 and left adjacent sample 33 of the coded block 300, respectively. The variable "predSize" is used in a later calculation (i.e., the following formula) (C) It is used for the purpose of ( ).

[0040] In some embodiments, the intra-prediction unit may derive a variable "isTransposed" to indicate the order of reference samples 32, 34 stored in the temporal array. For example, "isTransposed" being 0 indicates that the intra-prediction unit should present reference sample 32, derived from the upper neighbor sample 31 of the coded block 300, before reference sample 34, derived from the left neighbor sample 33. Selectively, "isTransposed" being 1 indicates that the intra-prediction unit should present reference sample 34, derived from the left neighbor sample 33 of the coded block 300, before reference sample 32, derived from the upper neighbor sample 31. In the implementation of the coding system 200, the value of "isTransposed" is sent to the entropy coding unit (e.g., entropy coding unit 215) as one of the parameters of the MIP process that is coded and written to a bitstream (e.g., output bitstream 21). In response to this, in the implementation of the decoding system 400 shown in Figure 4 of the present invention, the value of "isTransposed" can be received from an analysis unit (e.g., analysis unit 401) by analyzing the input bitstream (which can be the output bitstream 21).

[0041] The intra-prediction unit can further determine the variable "inSize" to indicate the number of reference samples 32, 34 used to derive the MIP prediction value. The value of "inSize" is determined by the following formula (A). In this invention, the meaning and operation of all operators in the formula are the same as the corresponding operators defined in the ITU-T H.265 standard.

[0042] inSize = ( 2 * boundarySize) - (mipSizeId = = 2) ? 1:0; (A)

[0043] For example, "= =" is the relational operator "equal". For example, if "mipSizeId" is 2, then "inSize" is 7 (calculated as (2*4)-1). If "mipSizeId" is 1, then "inSize" is 8 (calculated as (2*4)-0).

[0044] The intra-prediction unit can invoke the following processes to derive groups of reference samples 32 and 34 stored in the array p[x] (where "x" ranges from "0" to "inSize-1"). The intra-prediction unit can derive "nTbW" samples from the upper neighbor samples 31 of the coded block 300 (and store them in the array "refT"), and derive "nTbH" samples from the left neighbor samples 33 of the coded block 300 (and store them in the array "refL").

[0045] The intra-prediction unit can start downsampling on "refT" to obtain a "boundarySize" sample, and can also store the "boundarySize" sample in "refT". Similarly, the intra-prediction unit 206 can start downsampling on "refL" to obtain a "boundarySize" sample, and can also store the "boundarySize" sample in "refL".

[0046] In some embodiments, the intra-prediction unit can align the arrays "refT" and "refL" into a single array "pTemp" based on the order indicated by the variable "isTransposed". The intra-prediction unit can derive "isTransposed" to indicate the order of the reference samples stored in the temporal array "pTemp". For example, if "isTransposed" is 0 (or FALSE), the intra-prediction unit instructs to present the reference sample 32 derived from the upper neighbor sample 31 of the coded block 300 before the reference sample 34 derived from the left neighbor sample 33. In other cases, if "isTransposed" is 1 (or TRUE), the intra-prediction unit instructs to present the reference sample 34 derived from the left neighbor sample 33 of the coded block 300 before the reference sample 32 derived from the upper neighbor sample 31. In some embodiments, the implementation of the coding system 200 can determine the value of "isTransposed" using a rate-distortion optimization method. In some embodiments, in the implementation of the coding system 200, the intra-prediction unit can determine the value of "isTransposed" based on the comparison and / or interrelationship between adjacent samples 32, 34 and the coding block 300. In the implementation of the coding system 200, the value of "isTransposed" can be transferred to the entropy coding unit (e.g., entropy coding unit 215) as one of the parameters of the MIP process written to the bitstream (e.g., output bitstream 21). Correspondingly, in the implementation of the decoding system 400 shown in Figure 4 of the present invention, the value of "isTransposed" can be received from the analysis unit (e.g., analysis unit 401) by analyzing the input bitstream (which can be the output bitstream 21).

[0047] The intra-prediction unit can determine the array p[x] (where x ranges from "0" to "inSize-1") based on the following conditions G and H.

[0048] [Condition G] If "mipSizeId" is 2, then p[x] = pTemp[x+1] - pTemp[0].

[0049] [Condition H] Otherwise (for example, if "mipSizeId" is less than 2), p[0] = pTemp[0] - (1 << (BitDepth - 1)) and p[x] = pTemp[x] - pTemp[0] (where x is from 1 to "inSize - 1").

[0050] In the above condition H, "BitDepth" is the bit depth of the color component (e.g., the Y component) of the sample in coding block 300. The symbol "<<" is the bit shifting symbol used in the ITU-T H.265 standard.

[0051] Selectively, the intra-prediction unit can derive an array p[x] (where x ranges from 0 to "inSize-1") based on the following conditions I and J.

[0052] [Condition I] If "mipSizeId" is 2, then p[x] = pTemp[x+1] - pTemp[0].

[0053] [Condition J] Otherwise (for example, if "mipSizeId" is less than 2), p[0]=(1<<(BitDepth-1))-pTemp[0] and p[x]=pTemp[x]-pTemp[0] (where x is from 1 to "inSize-1").

[0054] In some embodiments, the intra-prediction unit can determine the value of array p[x] using a unified calculation method without determining the value of "mipSizeId". For example, the intra-prediction unit can add "(1<<(BitDepth-1))" as an additional element to "pTemp" and calculate p[x] as "pTemp[x]- pTemp[0]".

[0055] Step 302

[0056] In step 302, the intra-prediction unit (or prediction module 301) derives the MIP prediction value for the coded block 300 by utilizing the group of reference samples 32 and 34 and the MIP matrix. The MIP matrix is ​​selected from a predefined group of MIP matrices based on the corresponding MIP mode identifier (i.e., the variable "mipModeId") and MIP size identifier (i.e., the variable "mipSizeId").

[0057] The MIP prediction values ​​derived by the intra-prediction unit include partial prediction samples 35 for all or some of the sample locations in the coded block 300. The MIP prediction is denoted as "predMip[x][y]".

[0058] In the embodiment shown in Figure 3, the partial prediction sample 35 is the sample marked as a gray rectangle in the current block 300. The reference samples 32 and 34 of the array p[x] derived in step 301 are used as input to the prediction module 301. The prediction module 301 calculates the partial prediction sample 35 using the MIP matrix and shift parameters. The shift parameters include a shift offset parameter and a shift number parameter. In some embodiments, the prediction module 301 derives the partial prediction sample 35 with coordinates (x, y) based on the following equations (B) and (C).

[0059] TIFF0007857857000001.tif10147

[0060] TIFF0007857857000002.tif16147

[0061] In equation (B) above, the parameter "fO" is the shift offset parameter used to determine the parameter "oW". The parameter "sW" is the shift number parameter. "p[i]" is the reference sample. The symbol ">>" is the binary right shifting operator as defined in the H.265 standard.

[0062] The above formula (C) In this, "mWeight[i][j]" is an MIP weighting matrix whose matrix elements are fixed constants for both encoding and decoding. Selectively, in some embodiments, the implementation of the encoding system 200 utilizes an adaptive MIP matrix. For example, the MIP weighting matrix can be updated through various training methods, such as using one or more encoded images as input or using images provided to the encoding system 200 by external means. Once the MIP mode is determined, the intra-prediction unit can transfer "mWeight[i][j]" to the entropy encoding unit (e.g., entropy encoding unit 215). The entropy encoding unit can then write "mWeight[i][j]" to a bitstream, for example, to one or more special data units in the bitstream containing the MIP data. Accordingly, in some embodiments, the implementation of the decoding system 400 utilizing an adaptive MIP matrix can update the MIP matrix using a training method, for example, the output of which is one or more encoded images or blocks, or images from other bitstreams provided to the decoder 200 by external means, or images obtained from the analysis unit 401 by analyzing special data units in the input bitstream containing MIP matrix data.

[0063] The prediction unit 301 can determine the values ​​of "sW" and "fO" based on the current size of the block 300 and the MIP mode used for the current block 300. In some embodiments, the prediction unit 301 can obtain the values ​​of "sW" and "fO" by using a look-up table. For example, Table 1 below can be used to determine "sW".

[0064] [Table 1]

[0065] Selectively, Table 2 below can also be used to determine "sW".

[0066] [Table 2]

[0067] In some embodiments, the prediction module can set "sW" to a constant. For example, for blocks of various sizes utilizing different MIP modes, the prediction module can set "sW" to "5". As another example, for blocks of various sizes utilizing different MIP modes, the prediction module 301 can set "sW" to "6". As yet another example, for blocks of various sizes utilizing different MIP modes, the prediction module can set "sW" to "7".

[0068] In some embodiments, the prediction unit 301 can determine "fO" in Table 3 or Table 4 below.

[0069] [Table 3]

[0070] [Table 4]

[0071] In some embodiments, the prediction module 301 can directly set "fO" to a constant (for example, a value between 0 and 100). For example, for blocks of various sizes utilizing different MIP modes, the prediction module 301 can set "fO" to "46". As another example, the prediction module 301 can set "fO" to "56". As yet another example, the prediction module 301 can set "fO" to "66".

[0072] In some embodiments, the intra-prediction unit can perform a "clipping" operation on the values ​​of the MIP prediction samples stored in the array "predMip". If "isTransposed" is 1 (or TRUE), the array "predMip[x][y]" of "predSize x preSize" (where x is from 0 to "predSize-1" and y is from 0 to "predSize-1") is transposed to "predTemp[y][x] = predMip[x][y]", and then "predMip = predTemp".

[0073] More specifically, if the size of the coded block 303 is "8x8" (i.e., both "cbWidth" and "cbHeight" are 8), the intra-prediction unit can derive an "8x8" "predMip" array.

[0074] Step 303

[0075] In step 303 of Figure 3, the intra prediction unit is a portion of the coding block 300. prediction The predicted sample 37 is derived from the remaining samples, excluding sample 35. As shown in Figure 3, the intra prediction unit uses the filtering module 302 to determine the portion of the coded block 300. predictionThe predicted sample 37 can be derived from the remaining samples other than sample 35. The input to the filtering module 302 is part of step 302. prediction This could be sample 35. The filtering module 302 uses one or more interpolation filters to process a portion of the encoding block 300. prediction Predicted samples 37 can be derived for the remaining samples other than sample 35. The intra prediction unit (or filtering module 302) can generate prediction blocks (containing multiple predicted samples 37) of the coded block 300 based on the following conditions K and L, and store the predicted samples 37 in the array "predSamples[x][y]" (where x is from 0 to "nTbW-1" and y is from 0 to "nTbH-1").

[0076] [Condition K] If the intra-prediction unit determines that "nTbW" is greater than "predSize" or "nTbH" is greater than "predSize", the intra-prediction unit starts an upsampling process to derive "predSamples" based on "predMip".

[0077] [Condition L] Otherwise, the intra prediction unit sets the predicted value of coding block 300 to the MIP predicted value of coding block.

[0078] In other words, an intra-prediction unit can set "predSamples[x][y]" (where x ranges from 0 to "nTbW-1" and y ranges from 0 to "nTbH-1") to be equal to "predMip[x][y]". For example, an intra-prediction unit can set the "predSamples" of an encoded block whose size is equal to "8x8" (i.e., where both "cbWidth" and "cbHeight" are 8) to its "predMip[x][y]".

[0079] Through steps 301-303, the intra-prediction unit can generate predicted values ​​for the current block 300. The generated predicted values ​​can be used for further processing (for example, in the prediction block 23 shown in Figure 2).

[0080] Figure 4 is a schematic diagram showing a decoding system 400 according to an embodiment of the present invention. The decoding system 400 is configured to receive and process an input bitstream 40 and convert it into an output video 41. The input bitstream 40 can be a bitstream representing a compressed / encoded image / video. In some embodiments, the input bitstream 40 can come from an output bitstream (e.g., output bitstream 21) generated by an encoding system (e.g., encoding system 200).

[0081] The decoding system 400 includes an analysis unit 401 configured to analyze the input bitstream 40 in order to obtain the values ​​of syntax elements from the input bitstream 40. The analysis unit 401 further converts the binary representation of the syntax elements into numerical values ​​(i.e., decoding blocks 42) and transfers the numerical values ​​to the prediction unit 402 (for example, for decoding). In some embodiments, the analysis unit 401 may also transfer one or more variables and / or parameters to the prediction unit 402 for decoding the numerical values.

[0082] The prediction unit 402 is configured to determine the prediction block 43 of the decoding block 42 (e.g., a CU or part of a CU, such as a transform block). If the inter-decoding mode is instructed to be used to decode the decoding block 42, the motion compensation (MC) unit 403 of the prediction unit 402 can receive relevant parameters from the analysis unit 401 and decode accordingly under the inter-decoding mode. If the intra-prediction mode (e.g., MIP mode) is instructed to be used to decode the decoding block 42, the intra-prediction unit 404 of the prediction unit 402 can receive relevant parameters from the analysis unit 401 and decode accordingly under the instructed intra-decoding mode. In some embodiments, the intra-prediction mode (e.g., MIP mode) can be identified by a specific flag (e.g., MIP flag) embedded in the input bitstream 40.

[0083] For example, if the MIP mode is identified, the intra-prediction unit 404 will use the following method (as described in Figure 3) Steps 301-303 Based on the similarity to the above, the prediction block 43 (containing multiple prediction samples) can be determined.

[0084] First, the intra-prediction unit 404 derives one or more reference samples from the adjacent samples of the decoding block 42 (in Figure 3). Step 301 (Similar to the above). For example, the intra-prediction unit 404 can generate a reference sample by downsampling an adjacent sample or by directly extracting a portion from an adjacent sample.

[0085] Next, the intra-prediction unit 404 can use the reference sample, MIP matrix, and shift parameters to derive a partial prediction sample in the decoding block 42 (as shown in Figure 3). Step 302(Similar to the above). In some embodiments, the position of a partial prediction sample can be preset in the decoding clock 42. For example, the position of a partial prediction sample can be a position in the decoding block where the horizontal and vertical coordinate values ​​are odd. Shift parameters may include a shift offset parameter and a shift number parameter, which can be used for a shifting operation when generating partial prediction samples.

[0086] Finally, if a partial prediction sample for decoding block 42 is derived, the intra-prediction unit 404 derives the prediction samples for the remaining samples in decoding block 42 other than the partial prediction sample (Figure 3). Step 303 (Similar to the above). For example, the intra-prediction unit 404 can derive predicted samples using an interpolation filter. Partially predicted samples and adjacent samples are used as inputs to the interpolation filter.

[0087] The decoding system 400 includes a scaling unit 405 having a function similar to that of the inverse quantization unit 210 of the encoding system 200. The scaling unit 405 performs a scaling operation on the quantization coefficients 44 (e.g., levels) from the analysis unit 401 to generate reconstruction coefficients 45.

[0088] The transformation unit 406 has a function similar to that of the inverse transformation unit 211 in the encoding system 200. The transformation unit 406 performs one or more transformation operations (for example, the inverse operations of one or more transformation operations performed by the inverse transformation unit 211) to obtain the reconstruction residual 46.

[0089] The adder 407 adds the prediction block 43 from the prediction unit 402 and the reconstruction residual 46 from the transformation unit 406 to obtain the reconstruction block 47 of the decoding block 42. The reconstruction block 47 is also sent to the prediction unit 402 to be used as a reference (for example, for other blocks encoded under intra-prediction mode).

[0090] After all the decoding blocks 42 in the image or sub-image have been reconstructed (i.e., the reconstructed blocks 48 have been formed), the filtering unit 408 then processes the reconstructed blocks 48 In-loop filtering can be performed on the target pixels. The filtering unit 408 includes one or more filters, examples of which include deblocking filters, sample-adaptive offset (SAO) filters, adaptive loop filters (ALF), chroma-scaling lumens mapping (LMCS) filters, and neural network-based filters. In some embodiments, the filtering unit 408 can perform in-loop filtering on only one or more target pixels in the reconstruction block 48.

[0091] Next, the filtering unit 408 transmits the decoded image 49 (or image) or sub-image to the decoded image buffer (DPB) 409. Based on timing and control information, the DPB 409 outputs the decoded image as output video 41. The decoded image 49 stored in the DPB 409 can also be used as a reference image when the prediction unit 402 performs interpretation or intraprediction.

[0092] In some embodiments, the decoding system 400 may be a computing device comprising a processor and a storage medium for recording one or more decoding programs. Once the processor reads and executes a decoding program, the decoding system 400 can receive an input video bitstream and generate a corresponding decoded video.

[0093] In some embodiments, the decoding system 400 may be a computing device comprising one or more chips. Units or elements of the decoding system 400 may be implemented as integrated circuits on a chip.

[0094] Figure 5 is a schematic diagram showing an apparatus 500 according to an embodiment of the present invention. The apparatus 500 may be a “transmitting” apparatus. More specifically, the apparatus 500 is configured to acquire, encode, and store / transmit one or more images. The apparatus 500 includes an acquisition unit 501, an encoder 502, and a storage / transmitting unit 503.

[0095] The acquisition unit 501 is configured to acquire or receive an image and transfer that image to the encoder 502. The acquisition unit 501 may also be configured to acquire or receive a video containing a sequence of images and transfer that video to the encoder 502. In some embodiments, the acquisition unit 501 may be a device including one or more cameras (e.g., an image camera, a depth camera, etc.). In some embodiments, the acquisition unit 501 may be a device capable of partially or completely decoding a video bitstream to generate an image or video. The acquisition unit 501 may also include one or more elements for capturing an audio signal.

[0096] The encoder 502 is configured to encode an image from the acquisition unit 501 and generate a video bitstream. The encoder 502 is also configured to encode video from the acquisition unit 501 and generate a bitstream. In some embodiments, the encoder 502 can be implemented as the encoding system 200 described in Figure 2. In some embodiments, the encoder 502 may include one or more audio encoders to encode an audio signal and generate an audio bitstream.

[0097] The storage / transmission unit 503 is configured to receive video and / or audio bitstreams from the encoder 502. The storage / transmission unit 503 can encapsulate the video bitstream together with the audio bitstream to form a media file (e.g., an ISO-based media file) or a transport stream. In some embodiments, the storage / transmission unit 503 can write or store the media file or transport stream in a storage unit, such as a hard drive, disk, DVD, cloud storage, or portable memory device. In some embodiments, the storage / transmission unit 503 can transmit the video / audio bitstream to an external device over a transmission network (e.g., the Internet, a wired network, a cellular network, a wireless local area network, etc.).

[0098] Figure 6 is a schematic diagram showing an apparatus 600 according to an embodiment of the present invention. The apparatus 600 may be a “target” apparatus. More specifically, the apparatus 600 is configured to receive, decode, and render an image or video. The apparatus 600 includes a receiving unit 601, a decoder 602, and a rendering unit 603.

[0099] The receiving unit 601 is configured to receive, for example, a media file or transmission stream from a network or storage device. The media file or transmission stream includes a video bitstream and / or an audio bitstream. The receiving unit 601 can separate the video bitstream and the audio bitstream. In some embodiments, the receiving unit 601 can generate a new video / audio bitstream by extracting the video / audio bitstream.

[0100] Decoder 602 includes one or more video decoders, such as those in the decoding system 400 described above. Decoder 602 may also include one or more audio decoders. Decoder 602 decodes the audio bitstream and / or video bitstream from the receiving unit 601 to obtain one or more decoded audio files and / or decoded video files (corresponding to one or more channels).

[0101] The rendering unit 603 receives the decoded video / audio file and processes it to obtain a suitable video / audio signal for display / playback. These adjustment / reconstruction operations may include one or more of the following: denoising, compositing, color space conversion, upsampling, downsampling, etc. The rendering unit 603 can improve the quality of the decoded video / audio file.

[0102] Figure 7 is a schematic diagram showing a communication system 700 according to an embodiment of the present invention. The communication system 700 includes a source device 701, a storage medium or transmission network 702, and a target device 703. In some embodiments, the source device 701 may be the device 500 shown in Figure 5. The source device 701 transmits a media file to the storage medium or transmission network 702 for storage or transmission. The target device 703 may be the device 600 shown in Figure 6. The communication system 700 is configured to encode a media file, transmit or store the encoded media file, and then decode the encoded media file. In some embodiments, the source device 701 may be a first smartphone, the storage medium 702 may be cloud storage, and the target device may be a second smartphone.

[0103] The above embodiments are merely illustrative of some embodiments of the present invention, and their descriptions are specific and detailed. The above embodiments should not be construed as limiting the present invention. It should be noted that many modifications and alterations are possible by those skilled in the art, as long as they do not depart from the spirit and scope of the present invention. Accordingly, the scope of the present invention should be limited to the appended claims.

Claims

1. An image decoding method, The width, height, and prediction mode of the transformation block are determined by analyzing the bitstream, The aforementioned prediction mode indicates that a matrix-based intra-prediction (MIP) mode is used to decode the transformation block, and determines the MIP size identifier. To derive a group of reference samples for the aforementioned transformation block, This includes deriving the MIP prediction value of the transformation block based on the group of reference samples and the MIP matrix corresponding to the MIP size identifier, The MIP size identifier indicates that the predicted MIP size is equal to N, where N is a positive integer raised to the power of 2, and determining the MIP size identifier includes determining the MIP size identifier to 0 if the width and height are equal to 4. Deriving the predicted MIP value of the transformation block based on the group of reference samples and the MIP matrix corresponding to the MIP size identifier is: This includes deriving the MIP prediction value of the conversion block based on the following formula, x ranges from 0 to "predSize-1", and y ranges from 0 to "predSize-1", "sW" represents the shift number parameter, where sW is equal to 6; "fO" represents the shift offset parameter; "inSize" represents a variable indicating the number of reference samples used to derive the MIP prediction value; "p[i]" represents the reference sample; "predMip[x][y]" represents the MIP prediction value; "mWeight[i][j]" represents the MIP weighting matrix; "predSize" represents the size of the MIP prediction block; "pTemp[0]" represents the 0th value in the reference sample buffer; the symbol "<<" represents the binary left shift operator; the symbol ">>" represents the binary right shift operator; based on sW being equal to 6, 1<<(sW-1) is equal to 32; different MIP size identifiers correspond to the same shift offset parameter, and the shift offset parameter is a constant. An image decoding method characterized by the following:

2. The aforementioned image decoding method further, Obtaining the reference sample buffer by downsampling the group of reference samples in the conversion block, This includes determining the input sample value based on the reference sample in the reference sample buffer, the MIP size identifier, and the bit depth of the luminance component. The reference sample buffer includes a group of the downsampled reference samples of the transformation block. The image decoding method according to feature 1.

3. Determining the input sample value based on the reference sample in the reference sample buffer, the MIP size identifier, and the bit depth of the luminance component is: This includes deriving the input sample values ​​based on the following formula, If the aforementioned MIP size identifier is equal to 2, then p[x] = pTemp[x+1] - pTemp[0], If the MIP size identifier is less than 2, "p[x]" represents the reference sample, "pTemp[x]" represents the x-th value in the reference sample buffer, and "BitDepth" represents the bit depth of the luminance component. The image decoding method according to feature 2.

4. The image decoding method further includes deriving the group of reference samples of the transformation block based on adjacent samples, The adjacent sample includes the upper adjacent sample and / or the left adjacent sample. The image decoding method according to feature 1.

5. The further includes setting the predicted value of the conversion block to be equal to the MIP predicted value of the conversion block. The image decoding method according to feature 1.

6. An image encoding method, To determine the width and height of the transformation block in the aforementioned image, Determining the intra-prediction (MIP) size identifier based on the matrix, To derive a group of reference samples for the aforementioned transformation block, This includes deriving the MIP prediction value of the transformation block based on the group of reference samples and the MIP matrix according to the MIP size identifier, The MIP size identifier indicates that the predicted MIP size is equal to N, where N is a positive integer raised to the power of 2, and determining the MIP size identifier includes determining the MIP size identifier to 0 if the width and height are equal to 4. Deriving the predicted MIP value of the transformation block based on the group of reference samples and the MIP matrix according to the MIP size identifier is: This includes deriving the MIP prediction value of the conversion block based on the following formula, x ranges from 0 to "predSize-1", and y ranges from 0 to "predSize-1", "sW" represents the shift number parameter, where sW is equal to 6; "fO" represents the shift offset parameter; "inSize" represents a variable indicating the number of reference samples used to derive the MIP prediction value; "p[i]" represents the reference sample; "predMip[x][y]" represents the MIP prediction value; "mWeight[i][j]" represents the MIP weighting matrix; "predSize" represents the size of the MIP prediction block; "pTemp[0]" represents the 0th value in the reference sample buffer; the symbol "<<" represents the binary left shift operator; the symbol ">>" represents the binary right shift operator; based on sW being equal to 6, 1<<(sW-1) is equal to 32; different MIP size identifiers correspond to the same shift offset parameter, and the shift offset parameter is a constant. An image encoding method characterized by the following.

7. The aforementioned image encoding method further, Obtaining the reference sample buffer by downsampling the group of reference samples in the conversion block, This includes determining the input sample value based on the reference sample in the reference sample buffer, the MIP size identifier, and the bit depth of the luminance component. The reference sample buffer includes a group of the downsampled reference samples of the transformation block. The image coding method according to feature 6.

8. Determining the input sample value based on the reference sample in the aforementioned reference sample buffer, the MIP size identifier, and the bit depth of the luminance component is: This includes deriving the input sample values ​​based on the following formula, If the aforementioned MIP size identifier is equal to 2, then p[x] = pTemp[x+1] - pTemp[0], If the MIP size identifier is less than 2, "p[x]" represents the reference sample, "pTemp[x]" represents the x-th value in the reference sample buffer, and "BitDepth" represents the bit depth of the luminance component. The image coding method according to feature 7.

9. A computer-readable storage medium in which a computer program and a bitstream are stored, When the computer program is executed by the processor, the image encoding method according to any one of claims 6 to 8 is executed to generate the bitstream. A computer-readable storage medium characterized by the following features.