Matrix weighted intra prediction of video signals

By simplifying the matrix weighted intra prediction method, the problem of high-definition video storage and transmission needs is solved, and more efficient video encoding and decoding is achieved.

CN120455680APending Publication Date: 2025-08-08HFI INNOVATION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510812753.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-08-30
Filing Date
2020-07-28
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

When existing video encoding and decoding technologies process high storage and transmission requirements, resulting in excessive bandwidth and storage space requirements, making it difficult to efficiently compress and decode.

Method used

The simplified matrix-weighted intra prediction method is adopted to determine the classification of target blocks, generate matrix-weighted intra prediction signals, simplify the calculation process, and unify the prediction process of different sizes and blocks.

Benefits of technology

Improves video encoding efficiency, reduces the need for storage space and transmission bandwidth, and is suitable for compression and decoding of high-definition videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455680A_ABST
    Figure CN120455680A_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for performing simplified matrix weighted intra prediction. The method may include: determining a classification of a target block; and generating a matrix weighted intra prediction (MIP) signal based on the classification, in which determining the classification of the target block comprises: in response to the target block being 4 * 4 in size, determining that the target block belongs to a first class; or in response to the fact that the size of the target block is 8 * 8, 4 * N or N * 4, and N is an integer between 8 and 64, determining that the target block belongs to the second type.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims priority to U.S. Provisional Application No. 62 / 894,489, filed on August 30, 2019, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present disclosure relates generally to video processing and, more particularly, to methods and systems for performing simplified matrix-weighted intra prediction of video signals. Background Art

[0003] A video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, the video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are many video codec formats that use standardized video codec technologies, the most common of which are based on prediction, transform, quantization, entropy coding, and loop filtering. Video codec standards, such as the High Efficiency Video Codec (HEVC / H.265) standard and the Versatile Video Codec (VVC / H.266) standard (AVS), specify specific video codec formats and are developed by standardization organizations. As more and more video standards adopt advanced video codec technologies, the coding and decoding efficiency of new video codec standards is also getting higher and higher. Summary of the Invention

[0004] An embodiment of the present invention provides a method for simplifying matrix-weighted intra prediction. The method may include: determining a classification of a target block; and generating a matrix-weighted intra prediction (MIP) signal based on the classification, wherein determining the classification of the target block includes: in response to the size of the target block being 4×4, determining that the target block belongs to a first class; or in response to the size of the target block being 8×8, 4×N, or N×4, where N is an integer between 8 and 64, determining that the target block belongs to a second class.

[0005] An embodiment of the present disclosure also provides a system for performing simplified matrix-weighted intra-frame prediction. The system may include: a memory for storing an instruction set; and at least one processor, the at least one processor being configured to execute the instruction set to cause the system to perform: determining a classification of a target block; and generating a matrix-weighted intra-frame prediction (MIP) signal based on the classification, wherein determining the classification of the target block includes: in response to the size of the target block being 4×4, determining that the target block belongs to the first category; or in response to the size of the target block being 8×8, 4×N, or N×4, where N is an integer between 8 and 64, determining that the target block belongs to the second category.

[0006] An embodiment of the present disclosure also provides a non-transitory computer-readable medium storing an instruction set executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content. The method may include: determining a classification of a target block; and generating a matrix-weighted intra-frame prediction (MIP) signal based on the classification, wherein determining the classification of the target block includes: in response to the size of the target block being 4×4, determining that the target block belongs to the first category; or in response to the size of the target block being 8×8, 4×N, or N×4, where N is an integer between 8 and 64, determining that the target block belongs to the second category. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings.The various features shown in the drawings are not drawn to scale.

[0008] Figure 1 The structure of an exemplary video sequence consistent with embodiments of the present disclosure is shown.

[0009] Figure 2A A schematic diagram illustrating an exemplary encoding process performed by a hybrid video codec system consistent with an embodiment of the present disclosure is shown.

[0010] Figure 2B A schematic diagram illustrating another exemplary encoding process performed by a hybrid video coding system consistent with an embodiment of the present disclosure is shown.

[0011] Figure 3A A schematic diagram illustrating an exemplary decoding process performed by a hybrid video codec system consistent with an embodiment of the present disclosure is shown.

[0012] Figure 3B A schematic diagram illustrating another exemplary decoding process performed by a hybrid video coding system consistent with an embodiment of the present disclosure is shown.

[0013] Figure 4 is a block diagram of an exemplary apparatus for encoding or decoding video consistent with embodiments of the present disclosure.

[0014] Figure 5 An exemplary schematic diagram of matrix-weighted intra prediction consistent with an embodiment of the present disclosure is shown.

[0015] Figure 6 Illustrated is a table including three exemplary categories used in matrix-weighted intra prediction consistent with an embodiment of the present disclosure.

[0016] Figure 7 An exemplary lookup table for determining the offset “sO” consistent with embodiments of the present disclosure is illustrated.

[0017] Figure 8An exemplary lookup table for determining the offset “sW” consistent with embodiments of the present disclosure is illustrated.

[0018] Figure 9 An exemplary matrix illustrating an exemplary exclusion operation is shown consistent with embodiments of the present disclosure.

[0019] Figure 10 Another exemplary matrix illustrating another exemplary exclusion operation consistent with embodiments of the present disclosure is illustrated.

[0020] Figure 11 is a flow chart of an exemplary method for processing video content consistent with embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, and unless otherwise specified, the same numbers in different figures represent the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with the aspects related to the present invention as described in the appended claims. Unless otherwise specifically stated, the term "or" encompasses all possible combinations unless not feasible. For example, if a component is stated to include A or B, then, unless otherwise clearly stated or not feasible, the component may include A, or B, or A and B. As a second example, if a component is stated to include A, B, or C, then, unless otherwise clearly stated or not feasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0022] Video coding systems are commonly used to compress digital video signals, for example, to reduce storage space consumption or transmission bandwidth consumption associated with such signals. As high-definition (HD) video (e.g., with a resolution of 1920×1080 pixels) becomes increasingly popular in various video compression applications, such as online video streaming, video conferencing, or video surveillance, the demand for developing coding tools that efficiently compress video data continues to increase.

[0023] For example, video surveillance applications are becoming increasingly widespread in many scenarios (such as security, traffic, and environmental monitoring), and the number and resolution of surveillance devices are rapidly increasing. Many video surveillance applications prefer to provide users with high-definition video to capture more information. HD video has more pixels per frame to capture this information. However, HD video streams can have high bit rates, requiring high bandwidth transmission and large storage space. For example, a surveillance video stream with an average resolution of 1920×1080 can require up to 4 Mbps of bandwidth for real-time transmission. Furthermore, video surveillance typically operates 24 / 7, posing a significant challenge to storage systems when it comes to video data storage. Therefore, the high bandwidth and large storage requirements of HD video have become major constraints to its large-scale deployment in video surveillance.

[0024] A video is a set of static images (or "frames") arranged in a temporal sequence to store visual information. Video capture devices (such as cameras) can be used to capture and store these images in a temporal sequence, and video playback devices (such as televisions, computers, smartphones, tablets, video players, or any end-user terminal with a display function) can be used to display these images in a temporal sequence. In addition, in some applications, video capture devices can transmit the captured video in real time to video playback devices (such as computers with monitors) for monitoring, conferencing, live broadcasting, and other purposes.

[0025] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented through software and executed by a processor (e.g., a processor of a general-purpose computer) or dedicated hardware. The compression module is generally referred to as an "encoder" and the decompression module is generally referred to as a "decoder." Encoders and decoders can be collectively referred to as "codecs." Encoders and decoders can be implemented in any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, a codec can decompress video from a first coding standard and recompress the decompressed video using a second coding standard. In this case, the codec can be referred to as a "transcoder."

[0026] Video encoding processes can identify and retain useful information that can be used to reconstruct an image, while ignoring less important information for reconstruction. If the ignored, less important information cannot be fully reconstructed, the encoding process is called "lossy." Otherwise, it is called "lossless." Most encoding processes are lossy as a trade-off to reduce required storage space and transmission bandwidth.

[0027] Useful information about the image being encoded (called the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include changes in pixel position, brightness, or color, with position changes being of greatest interest. A change in the position of a group of pixels representing an object can reflect the object's motion between the reference image and the current image.

[0028] Depending on whether the reference image is the current image itself or another image, the encoding of the current image can be divided into "inter-frame prediction" and "intra-frame prediction". Intra-frame prediction can exploit spatial redundancy (e.g., correlation between pixels within a frame) by interpolating predicted values from already encoded pixels. Inter-frame prediction can exploit temporal differences (e.g., motion vectors) between adjacent frames (e.g., a reference frame and a target frame), thereby enabling the codec of the target frame. The present disclosure relates to techniques for intra-frame prediction.

[0029] The present disclosure provides a method, apparatus, and system for performing simplified matrix-weighted intra-frame prediction of video signals. By removing the extra matrix omission operations in the MIP prediction process, the prediction process for blocks of different sizes can be unified and the calculation process can be simplified.

[0030] Figure 1 The structure of an example video sequence 100 consistent with embodiments of the present disclosure is illustrated. Video sequence 100 can be live video or video that has been captured and archived. Video 100 can be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). Video sequence 100 can be received from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored on a storage device), or a video provider interface (e.g., a video broadcast transceiver) that receives video from a video content provider.

[0031] like Figure 1 As shown, video sequence 100 may include a series of images arranged in time along a time axis, including images 102, 104, 106, and 108. Images 102-106 are consecutive, and there may be more images between images 106 and 108. Figure 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108, as indicated by the arrow. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately before or after the picture. For example, the reference picture of picture 104 may be the picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are only examples, and the embodiments of the present disclosure that do not talk about reference pictures are not limited to Figure 1 Example shown.

[0032] Typically, video codecs do not encode or decode an entire image at once due to the computational complexity of such tasks. Instead, they may divide the image into elementary segments and encode or decode the image segment by segment. Such elementary segments are referred to in this disclosure as Basic Processing Units ("BPUs"). For example, Figure 1 Structure 110 in. Figure 1 An example structure of an image (e.g., any of images 102-108) of video sequence 100 is shown. In structure 110, the image is divided into 4×4 basic processing units, whose boundaries are represented by dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit in an image can have variable sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of arbitrary shape and size. The size and shape of the basic processing unit for the image can be selected based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit.

[0033] A basic processing unit may be a logical unit that may include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color image may include a luma component (Y) representing non-color luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit may be repeated for each of its luma and chroma components.

[0034] Video encoding has multiple stages of operation, examples of which are given in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large to be processed, so it can be further divided into segments, referred to as "basic processing sub-units" in this disclosure. In some embodiments, the basic processing sub-unit may be referred to as a "block" in some video coding standards (e.g., MPEG series, H.261, H.263 or H.264 / AVC), or as a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-unit may have the same or smaller size as the basic processing unit. Similar to the basic processing unit, the basic processing sub-unit is also a logical unit that may include a set of different types of video data (e.g., Y, Cb, Cr associated with syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing sub-unit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to a deeper level according to processing needs. It should also be noted that different stages can use different schemes to divide the basic processing units.

[0035] For example, in the mode decision phase (an example of which will be Figure 2B ), the encoder can decide what prediction mode to use (e.g., intra-image prediction or inter-image prediction) for a basic processing unit, which may be too large to make this decision. The encoder can split the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and decide the prediction type for each individual basic processing sub-unit.

[0036] As another example, in the prediction phase (an example of which will be Figure 2A ), the encoder can perform prediction operations at the level of basic processing sub-units (e.g., CUs). However, in some cases, the basic processing sub-units may still be too large to process. The encoder can further split the basic processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and prediction operations can be performed at these levels.

[0037] As another example, in the transformation phase (an example of which will be Figure 2A(Detailed in H.265 / HEVC or H.266 / VVC), the encoder may perform transform operations for the residual basic processing sub-unit (e.g., CU). However, in some cases, the basic processing sub-unit may still be too large to be processed. The encoder may further split the basic processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), on which transform operations may be performed. It should be noted that the division schemes for the same basic processing sub-unit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0038] In the structure 110 of the figure, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, whose boundaries are represented by dotted lines. Different basic processing units of the same image can be divided into different basic processing sub-units in different schemes.

[0039] In some embodiments, to provide parallel processing and error resilience for video encoding and decoding, an image can be divided into multiple regions for processing, so that the encoding or decoding process for one region of the image does not need to depend on information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving coding efficiency. Furthermore, if data in one region is corrupted during processing or network transmission is lost, the codec can correctly encode or decode other regions of the same image without relying on the corrupted or lost data, thereby providing error resilience. In some video coding standards, an image can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different images in video sequence 100 may have different partitioning schemes for dividing the image into multiple regions.

[0040] For example, in Figure 1 In FIG, the structure 110 is divided into three regions 114, 116 and 118, whose boundaries are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that Figure 1 The basic processing units, basic processing sub-units, and regions of the structure 110 are merely examples, and the present disclosure is not limited to the embodiments thereof.

[0041] Figure 2A FIG2 illustrates a schematic diagram of an example encoding process 200A consistent with an embodiment of the present disclosure. An encoder may encode a video sequence 202 into a video bitstream 228 according to the process 200A. Figure 1 The video sequence 100 in FIG. 2 may include a set of images (referred to as “original images”) arranged in time sequence. Figure 1 In the structure 110 in FIG. 1 , each original image of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can perform process 200A at the basic processing unit level for each original image of the video sequence 202. For example, the encoder can perform process 200A in an iterative manner, where in one iteration of process 200A, the encoder can encode a basic processing unit. In some embodiments, the encoder can perform process 200A in parallel for each region (e.g., regions 114-118) of the original image of the video sequence 202.

[0042] Reference in Figure Figure 2A In the example of FIG200A, the encoder may provide the basic processing units (referred to as "original BPUs") of the original images of the video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a "forward path." In process 200A, after the quantization stage 214, the encoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that the encoder and decoder use the same reference data for prediction.

[0043] The encoder can iteratively perform process 200A to encode each original BPU of the original image (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original image (in the reconstruction path). After encoding all the original BPUs of the original image, the encoder can continue to encode the next image in the video sequence 202.

[0044] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or in any way inputting data.

[0045] In the prediction phase 204, in the current iteration, the encoder may receive the original BPU and the prediction reference 224 and perform a prediction operation to generate the prediction data 206 and the predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of the previous iteration of the process 200A. The purpose of the prediction phase 204 is to reduce information redundancy by extracting the prediction data 206, which can be used to reconstruct the original BPU into the predicted BPU 208 from the prediction data 206 and the prediction reference 224.

[0046] Ideally, the predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 typically differs slightly from the original BPU. To account for this difference, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate a residual BPU 210. For example, the encoder can subtract the pixel value (e.g., grayscale value or RGB value) of the predicted BPU 208 from the corresponding pixel value of the corresponding original BPU. Each pixel of the residual 210 can have a residual value as a result of this subtraction between the corresponding pixel of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significantly reducing quality. Therefore, the original BPU is compressed.

[0047] To further compress the residual BPU 210, during the transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," each of which is associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency-varying component of the residual BPU 210 (e.g., the frequency of luma variations). A basis pattern cannot be replicated from any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is analogous to a discrete Fourier transform of a function, where the basis patterns are analogous to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are analogous to the coefficients associated with the basis functions.

[0048] Different transform algorithms can use different base patterns. Various transform algorithms can be used in the transform stage 212, such as discrete cosine transform, discrete sine transform, etc. The transform at the transform stage 212 is reversible. That is, the encoder can restore the residual BPU 210 by performing the inverse operation of the transform (referred to as an "inverse transform"). For example, to restore the pixels of the residual BPU 210, the inverse transform may involve multiplying the values of the corresponding pixels of the base pattern by their associated coefficients and adding the products to produce a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same base pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct the residual BPU 210 based on these coefficients without receiving the base pattern from the encoder. Compared to the residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significantly reducing quality. As a result, the residual BPU 210 is further compressed.

[0049] The encoder can further compress the transform coefficients during the quantization stage 214. During the transform process, different basis patterns can represent different frequencies of change (e.g., the frequency of changes in brightness). Because the human eye is generally better at detecting low-frequency changes, the encoder can ignore information about high-frequency changes without significantly degrading decoding quality. For example, during the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. This operation can convert some transform coefficients of high-frequency basis patterns to zero, while transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore zero-valued quantized transform coefficients 216, which further compress the transform coefficients. The quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed as transform coefficients in the inverse operation of quantization (called "inverse quantization").

[0050] Because the encoder ignores the remainder of this division during rounding operations, quantization level 214 is lossy. Typically, quantization level 214 contributes the most information loss in process 200A. The greater the information loss, the fewer bits are required to quantize the transform coefficients 216. To achieve different levels of information loss, the encoder can use different values for the quantization parameter or any other parameter of the quantization process.

[0051] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type in the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packaged for network transmission.

[0052] Referring to the reconstruction path of process 200A, at the inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. At the inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0053] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may be omitted. Figure 2A one or more stages in a process.

[0054] Figure 2B A schematic diagram of another example encoding process 200B consistent with embodiments of the present disclosure is shown. Process 200B can be modified from process 200A. For example, process 200B can be used by encoders compliant with hybrid video coding standards (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and separates the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0055] In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-frame prediction") can use pixels from one or more coded adjacent BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") can use regions from one or more coded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded images. Temporal prediction can reduce the temporal redundancy inherent in the image.

[0056] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra-frame prediction. For an original BPU of the picture being encoded, the prediction reference 224 may include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder may generate a predicted BPU 208 by interpolating the neighboring BPUs. Interpolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, and the like. In some embodiments, the encoder may perform interpolation at the pixel level, for example, by interpolating the values of the corresponding pixels of each pixel of the predicted BPU 208. The neighboring BPUs used for interpolation may be positioned relative to the original BPU in various directions, such as vertically (e.g., on top of the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below the left, below the right, or above the original BPU), or in any direction defined in the video coding standard being used. For intra prediction, the prediction data 206 may include, for example, the position (eg, coordinates) of the used neighboring BPUs, the size of the used neighboring BPUs, interpolation parameters, the direction of the used neighboring BPUs relative to the original BPU, and the like.

[0057] As another example, during the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference image may be a BPU encoding and a reconstructed BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When generating all reconstructed BPUs for the same image, the encoder may generate the reconstructed image as a reference image. The encoder may perform a "motion estimation" operation to search for a matching region within a range of the reference image (referred to as a "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a location in the reference image with the same coordinates as the original BPU in the current image and may extend outward a predetermined distance. When the encoder identifies an area in the search window that is similar to the original BPU (e.g., using a pixel recursive algorithm, a block matching algorithm, etc.), the encoder may determine such an area as a matching region. The matching region may have a different size than the original BPU (e.g., smaller, equal, larger, or in a different shape). Because the reference image and the current image are temporally separated on the timeline (e.g., as Figure 1 As shown in the figure, the matching area can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of this movement as a "motion vector". When using multiple reference images (e.g., Figure 1 The encoder can search for matching areas and determine the motion vector associated with each reference image. In some embodiments, the encoder can assign weights to the pixel values of the matching areas of each matching reference image.

[0058] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.

[0059] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may move a matching area of the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When multiple reference images are used (e.g., Figure 1The encoder may move the matching region of the reference image based on the motion vectors and the average pixel value of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of the respective matching reference images, the encoder may add the weighted sum of the pixel values of the moved matching regions.

[0060] In some embodiments, inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference pictures in the same temporal direction as the current picture. For example, Figure 1 The picture 104 in is a unidirectional inter-frame predicted picture, where the reference picture (i.e., picture 102) precedes the picture 104. Bidirectional inter-frame prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, Figure 1 Picture 106 is a bi-directional inter-predicted picture, where the reference pictures (ie, pictures 104 and 108 ) are relative to picture 104 in both temporal directions.

[0061] Still referring to the forward path of process 200B, after the spatial prediction 2042 and temporal prediction stages 2044, in the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction and inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode to minimize the value of a cost function based on the bit rate of the candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode. Based on the selected prediction mode, the encoder can generate a corresponding predicted BPU 208 and predicted data 206.

[0062] In the reconstruction path of process 200B, if intra-prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current picture), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolation of the next BPU of the current picture). If inter-prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current picture in which all BPUs have been encoded and reconstructed), the encoder can provide the prediction reference 224 to the loop filtering stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortion introduced by inter-prediction (e.g., blocking artifacts). The encoder can apply various loop filtering techniques at the loop filtering stage 232, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The loop-filtered reference picture can be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., as an inter-prediction reference picture for a future picture of the video sequence 202). The encoder may store one or more reference pictures in a buffer 234 for use in a temporal prediction stage 2044. In some embodiments, the encoder may encode loop filtering parameters (e.g., loop filter strength) in a binary encoding stage 226, along with the quantized transform coefficients 216, prediction data 206, and other information.

[0063] Figure 3A A schematic diagram of an example decoding process 300A consistent with an embodiment of the present disclosure is illustrated. Process 300A may correspond to Figure 2A In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder may decode the video bitstream 228 into a video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to the compression and decompression processes (e.g., Figures 2A-2B The information in the quantization stage 214 in the video stream 304 is lost, and typically the video stream 304 is different from the video sequence 202. Figures 2A-2B 200A and 200B, the decoder may perform process 300A at a basic processing unit (BPU) level for each picture encoded in the video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for each region (e.g., regions 114-118) of each picture encoded in the video bitstream 228.

[0064] refer to Figure 3A, the decoder may provide a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as a "coded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to the prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may provide the predicted reference 224 to the prediction stage 204 to perform a prediction operation in the next iteration of process 300A.

[0065] The decoder can iteratively perform process 300A to decode each coded BPU of the coded picture and generate prediction reference 224 for encoding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder can output the picture to video stream 304 for display and continue decoding the next coded picture in the video bitstream 228.

[0066] In the binary decoding stage 302, the decoder can implement the binary encoding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, the transform type, quantization process parameters (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted over the network in packet form, the decoder can depacketize the video bitstream 228 before providing it to the binary decoding stage 302.

[0067] Figure 3B A schematic diagram of another example decoding process 300B consistent with embodiments of the present disclosure is illustrated. Process 300B can be modified from process 300A. For example, process 300B can be used by decoders compliant with hybrid video coding standards (e.g., the H.26x series). Compared to process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0068] In process 300B, for a coding basic processing unit (referred to as a "current BPU") of a decoded coded image (referred to as a "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data, depending on the prediction mode used by the encoder to encode the current BPU. For example, if the encoder encodes the current BPU using intra-frame prediction, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-frame prediction, parameters for the intra-frame prediction operation, and the like. The parameters for the intra-frame prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as a reference, the size of the neighboring BPUs, interpolation parameters, the orientation of the neighboring BPUs relative to the original BPU, and the like. For another example, if the encoder encodes the current BPU using inter-frame prediction, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-frame prediction, parameters for the inter-frame prediction operation, and the like. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with the reference images respectively, the positions (e.g., coordinates) of one or more matching regions in the reference images respectively, one or more motion vectors associated with the matching regions respectively, etc.

[0069] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. The details of performing such spatial prediction or temporal prediction are described in Figure 2B After performing such spatial prediction or temporal prediction, the decoder can generate a predicted BPU 208. The decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a predicted reference 224, as shown in FIG. Figure 3A As described in.

[0070] In process 300B, the decoder may provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-frame prediction in the spatial prediction stage 2042, then after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolation of the next BPU of the current picture). If the current BPU is decoded using inter-frame prediction in the temporal prediction stage 2044, then after generating the prediction reference 224 (e.g., a reference picture in which all BPUs have been decoded), the encoder may feed the prediction reference 224 to the loop filtering stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may Figure 2BThe loop filter is applied to the prediction reference 224 in the manner described in

[15] . The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., as an inter-frame prediction reference picture for a future coded picture in the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that the current BPU is encoded using inter-frame prediction, the prediction data may further include parameters for the loop filtering (e.g., loop filtering strength).

[0071] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding video consistent with an embodiment of the present disclosure. Figure 4 As shown, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 can become a special-purpose machine for video encoding or decoding. The processor 402 can be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chip (SoCs), application-specific integrated circuits (ASICs), and the like. In some embodiments, the processor 402 may also be a group of processors grouped into a single logical component. For example, as shown in the figure. As Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0072] The apparatus 400 may also include a memory 404 configured to store data (eg, a set of instructions, computer code, intermediate data, etc.). Figure 4As shown, the stored data may include program instructions (e.g., program instructions for implementing stages in process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 may access program instructions and data for processing (e.g., via bus 410) and execute program instructions to perform operations or management on the data for processing. Memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, memory 404 may include any number of random access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a group of memories grouped into a single logical component ( Figure 4 not shown).

[0073] The bus 410 may be a communication device for transmitting data between components within the apparatus 400 , such as an internal bus (eg, a CPU-memory bus), an external bus (eg, a Universal Serial Bus port, a Peripheral Component Interconnect Express port), and the like.

[0074] For ease of explanation and to avoid ambiguity, in this disclosure, processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry." The data processing circuitry may be implemented entirely in hardware or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single standalone module or may be fully or partially integrated into any other component of apparatus 400.

[0075] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, any combination of wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.

[0076] In some embodiments, the apparatus 400 may optionally further include a peripheral interface 408 to provide a connection to one or more peripheral devices. Figure 4 As shown, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), etc.

[0077] It should be noted that a video codec (e.g., a codec that performs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of software or hardware modules in apparatus 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of apparatus 400, such as program instructions that can be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of apparatus 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, NPUs, etc.).

[0078] The present disclosure provides a method that can be performed by an encoder and / or decoder to simplify matrix-weighted intra prediction (MIP). The MIP method is a new intra prediction technique in VVC. The MIP mode applies to blocks with an aspect ratio of max(width, height) / min(width, height) less than or equal to 4 and only applies to the luma component. The MIP flag is signaled in parallel with the intra sub-partition mode, multi-reference line intra prediction mode, or most probable mode.

[0079] When encoding a block using MIP mode, similar to traditional intra prediction mode, a row of reconstructed neighboring boundary samples on the left side of the block and a row of reconstructed neighboring boundary samples on the top of the block are used as input to predict the block. When reconstructed neighboring boundary samples are not available, they can be generated as in traditional intra prediction. To predict the luma sample, the reconstructed neighboring boundary samples are first averaged to generate a reduced boundary vector neighbor red , neighbor red [0] represents the first element of the reduced boundary vector. Then the reduced boundary vector neighbor is used red Generate input vector input red , used for the matrix-vector multiplication process, and can obtain the reduced prediction signal pred red Finally, the reduced prediction signal pred red Perform bilinear interpolation to generate the output of the MIP prediction signal pred. The MIP prediction process is shown in Figure 5 Shown in.

[0080] MIP blocks are divided into three categories based on the width (W) and height (H) of the block: Class 0: for W = H = 4 (i.e., 4×4 blocks); Class 1: for max{W,H}=8 (i.e., 4×8, 8×4, 8×8 blocks); and Class2: For max{W,H}>8.

[0081] like Figure 6 As shown in Table 6, the differences between the three categories are the modulus, the number of matrices, the matrix size, and the size of the input vector. red (Input to the matrix-vector multiplication process), reduced prediction signal input red The size of (the output of the matrix-vector multiplication process).

[0082] In the following description, Class0, Class1, and Class2 are represented as S0, S1, and S2, respectively. i Represents the matrix set S i The number of matrices in (i=0,1,2). For MIP mode k, the matrix For the matrix multiplication process, where Represents the matrix set S i The j-th matrix in ,j is derived using the following equation (1).

[0083] It should be noted that when MIP mode k is greater than or equal to N i When generating the reduced boundary vector neighbor red and the reduced prediction signal pred red The swap operation and the transposition operation are performed in the steps of . The details of these two operations are described above. In addition, the following equations (2) and (3) are used to determine whether a swap operation or a transposition operation is required.

[0084] As mentioned above, the generation of the output prediction signal pred of the current block is based on the following three steps, namely averaging, matrix-vector multiplication, and linear interpolation. The details of these steps are described below.

[0085] Among the boundary samples, 4 samples of Class 0 and 8 samples of Class 1 and Class 2 are extracted by averaging. For example, the reduced boundary vector neighbor can be generated by averaging the reconstructed neighboring boundary samples according to the following rule red .

[0086] Class0: Take the average of every two samples. Reduce the size of the boundary vector neighbor red It is 4×1.

[0087] Class1 and Class2: For the neighboring boundary samples reconstructed above the current block, average every W / 4 samples. For the neighboring boundary samples reconstructed to the left of the current block, average every H / 4 samples. Reduced boundary vector neighbor red The size is 8×1.

[0088] Reduced boundary vector neighbor red is the vector obtained by averaging the adjacent boundary samples reconstructed above the current block The vector obtained by averaging the adjacent boundary samples reconstructed on the left side of the current block As mentioned above, for MIP mode k is greater than or equal to N i , perform a swap operation. For example, and The concatenation order of the two vectors can be interchanged, as shown in the following formula (4):

[0089] Then, the input vector of the matrix-vector multiplication is red Generate as follows.

[0090] For Class0 and Class1: input red [0] = neighbor red [0]-(1<<(bitDepth-1)), input red [j] = neighbor red [j]-neighbor red [0],j=1,…,size(neighbor red )-1, Equation (5)

[0091] For Class2 input red [j] = neighbor red [j+1]-neighborr red [0],j=0,…,size(neighborr red )-2, equation (6)

[0092] In the above equations (5) and (6), neighbor red [0] represents the vector neighborr red The first element of . According to equations (5) and (6), the input of Class0, Class1, and Class2 redThe inSize sizes are 4, 8, and 7 respectively.

[0093] Vector input red Performs matrix-vector multiplication as input. The result is a reduced prediction signal pred on the subsampled set of samples in the current block red For example, from the reduced input vector input red In the above example, a reduced prediction signal pred can be generated. red , which is the width W red and height H red Here, W red and H red It is defined by the following equations (7) and (8):

[0094] As mentioned above, when the variable “isTransposed” is equal to 1, the reduced prediction signal pred red Transpose. Assume that the final reduced prediction signal pred red The size of W red ×H red , then the untransposed W′ red ×H′ red The size of is derived from the following equations (9) and (10):

[0095] The reduced prediction signal pred′ is calculated by computing the matrix-vector product according to the following equation (11): red Vector: pred′ red =M·input red +neighbor red [0] Equation (11)

[0096] Then, the vector pred′ red Arranged in raster scan order in matrices pred of size 4×4, 4×8, 8×4, and 8×8 red In the example, the vector pred′ of size 16×1 is red Arranged in a 4×4 matrix pred red The example in is as follows: pred′ red =[A,B,C,D,F,F,G,H,I,J,K,L,M,N,O,P] T

[0097] In other words, for x=0...W′ red -1 and y=0...H' red -1, the matrix pred can be calculated using the following equation (12) red

[0098] In the above equation (12), the variable "inSize" is the input vector input red The size of, as described above, is used to generate the reduced prediction signal pred red The matrix M is obtained from one of three matrix sets S0, S1, S2 according to the block size classification and MIP mode k.

[0099] The variables oW and sW are two predefined values, depending on the matrix used for the matrix-vector multiplication. The variable oW is used to limit the precision of each element in the matrix to 7 bits, so that all elements are greater than or equal to 0. For example, the factor oW can be defined as the following equation (13):

[0100] In equation (13) above, sO is an offset and is derived from a lookup table. For example, Figure 7 Table 7 shows an exemplary lookup table for sO according to some disclosed embodiments.

[0101] Additionally, the variable "sW" is derived using another lookup table. Figure 8 Table 8 shows an exemplary lookup table for sW according to some disclosed embodiments.

[0102] In equation (12), two variables "inch" and "incW" for leaving out half of the matrix row of 4×16 and 16×4 blocks are defined as the following equations (14) and (15):

[0103] In the above equation (15), the variable predC is used to convert the reduced prediction signal pred red Arranged in W′ red ×H′ red Matrix, and is defined as the following equation (16):

[0104] Input vector input red , where inSize is equal to 7, a matrix of 64 rows and 7 columns is used for the blocks belonging to Class2, generating a vector of 64 elements. However, according to the equation , 4×16 and 16×4 blocks only require 32 elements. This is because the reduced prediction signal pred of Class2 red It can be arranged in 8×8, which exceeds the short side of the 4×16 or 16×4 block. Therefore, the omission operation is performed.

[0105] Figure 9 illustrative exclusion operations according to some disclosed embodiments are shown. Figure 9 As shown, for 4×16 blocks with isTransposed=0 and 16×4 blocks with isTransposed=1, the second row of every two rows is excluded from the matrix. red Contains 32 elements arranged in a 4×8 size.

[0106] Figure 10 illustratively illustrates exemplary omission operations according to some disclosed embodiments. Figure 10 As shown, for 16×4 blocks with isTransposed=0 and 4×16 blocks with isTransposed=1, the last 8 rows of every 16 rows are excluded from the matrix. Therefore, the reduced prediction signal pred red Contains 32 elements arranged in a 8×4 size.

[0107] The output prediction signals for the remaining positions are obtained by linear interpolation from the subsampled set pred red The reduced prediction signal on is generated by a single-step linear interpolation in each direction.

[0108] As mentioned above, the MIP mode prediction is generated using a three-step process involving averaging adjacent reconstructed samples, matrix-vector multiplication, and bilinear interpolation. The MIP mode prediction process differs from that of traditional intra-frame prediction. While the MIP mode improves coding efficiency, its design can be complex in two ways.

[0109] Regarding the first aspect, the exclusion operation for 4×16 and 16×4 blocks during matrix-vector multiplication may be problematic for the following three reasons. First, the exclusion operation not only adds extra operations but also makes the prediction process inconsistent because the exclusion operation only applies to 16×4 and 4×16 blocks. Second, for 16×4 and 4×16 blocks, the reduced prediction signal pred red The size of may be different before and after transposition. Therefore, it is necessary to red and H′ red ) size is additionally derived. Third, for 16×4 and 4×16 blocks, the reduced prediction signal pred redThe size of may be different, as shown in equation (17) below.

[0110] For the second aspect, in order to limit the precision of each element in the matrix to 7 bits and ensure that all elements are non-negative, an offset sO is added to the reduced prediction signal during the matrix multiplication process. However, it may be unnecessary and complicated for the following three reasons. First, additional memory is required to store the table of offsets sO. The table contains a total of 34 elements, each element is 7 bits. Therefore, a total of 238 bits of memory are required. Second, a table lookup operation is required to determine the value of sO using the class index and matrix number. Third, additional multiplication and addition operations are required to generate the reduced prediction signal pred red In addition to calculating the matrix M and the input vector input red In addition to the matrix-vector multiplication between sO and the input vector input red For a 4×4 block, the total number of multiplications per sample required to generate the prediction signal increases to 5.

[0111] The present disclosure provides methods to solve these problems without affecting the bit rate. In some exemplary methods, 8 bits are used instead of 7 bits to store the matrix. Then, equation (12) can be reformulated as the following equation (18).

[0112] In equation (18) above, the offset sO is subtracted from all elements in the matrix. The above storage and calculation method of the matrix-vector product can produce the same result. However, the number of bits of the storage matrix increases by 4882 (i.e., 5120×1-34×7), and the bit width of the multiplication operation is expanded to 8 bits.

[0113] The following describes a method that removes the elimination operation and the sO lookup table.

[0114] In order to remove the extra exclusion operation of the matrix in the MIP prediction process, two methods are provided. According to the first method of removing the exclusion operation, the unreasonable classification method in the traditional MIP method can be modified, and the 4×16 and 16×4 blocks are moved from Class2 to Class1, so that the generated reduced prediction signal pred red Do not exceed the short side limit.

[0115] In one exemplary embodiment, the MIP classification rules are modified as follows: Class0: 4×4; Class1: 4×N, 8×8, N×4, where N is an integer between 8 and 64; and Class2: Others.

[0116] With this modification, blocks of size 8×8, 4×8, 4×16, 4×32, 4×64, 8×4, 16×4, 32×4, or 64×4 can be moved from Class2 to Class1. Therefore, blocks of size 8×8, 4×8, 4×16, 4×32, 4×64, 8×4, 16×4, 32×4, or 64×4 can use matrices from set S1, each with 16 rows and 8 columns, during matrix multiplication. In this way, only 16 elements need to be generated to form a 4×4 reduced prediction signal pred red , thus removing the exclude operation. This modification can be indicated by double strikethrough or italic highlighting of the change as follows: If cbWidth and cbHeight are both equal to 4, MipSizeId[x][y] is set equal to 0. Otherwise, if If cbWidth*cbHeight is less than or equal to 64, MipSizeId[x][y] is set to 1. Otherwise, MipSizeId[x][y] is set equal to 2.

[0117] This solution has at least three benefits.

[0118] First, the matrix multiplication process has been simplified and unified. For blocks of size 4×16 or 16×4, the exclusion operation has been removed. Consequently, the checks for whether the exclusion operation is performed and the two variables "inch" and "incW" can be eliminated. Furthermore, during matrix-vector multiplication, no additional matrix operations are required for all blocks. Thus, the matrix multiplication process is unified.

[0119] Secondly, the number of multiplications and additions for 4×16 and 16×4 blocks during matrix multiplication is reduced. In some embodiments, for blocks of size 4×16 or 16×4, a 32×7 matrix can be used for matrix-vector multiplication, while in the provided embodiments, a 16×8 matrix can be used. As a result, the number of multiplications and additions for 4×16 or 16×4 blocks can be reduced.

[0120] Third, pred red The derivation of is simplified and unified. For all blocks, the reduced prediction signal pred red The size of is the same before and after transposition. Therefore, W′ is deleted. red and H′ red For example, pred red The derivation of the size of can be simplified to the following equations (19) and (20).

[0121] The reduced prediction signal pred red The size of is unified by the following formula (21):

[0122] According to the second method of removing the exclusion operation, the MIP classification rules are modified as follows: Class0: 4×4; Class 1: 4×8, 8×4, 4×16, and 16×4; and Class2: Others.

[0123] 4×16 and 16×4 blocks (italicized in the modified MIP classification above) are moved to Class 1, and 8×8 blocks are moved to Class 2. With this modification, the exclusion operation is removed and the MIP classification rules are further simplified. In some embodiments, this modification can be expressed below with changes highlighted by double strikethrough or italics. If cbWidth and cbHeight are both equal to 4, MipSizeId[x][y] is set equal to 0. otherwise, If Min(cbWidth,cbHeight) is equal to 4, then MipSizeId[x][y] is set to 1. Otherwise, MipSizeId[x][y] is set equal to 2.

[0124] In order to remove the table of offset sO, an embodiment of the present disclosure provides a method of modifying the value of the matrix M and the offset sO without looking up the table.

[0125] In the first exemplary embodiment, the offset s0 is replaced by the matrix for Class0 The first element in (j=0…17), the matrix for Class1 The first element in and the matrix of Class2 The seventh element in the matrix. The i-th element in the matrix represents the i-th number counted in raster scan order starting from the top left corner of the matrix. By doing this, you can remove Figure 7 The lookup table of the offset s0 depending on the class index and the matrix number is shown in Table 7. This saves 238 bits of storage space.

[0126] In the second exemplary embodiment, for all classes, the offset s0 is replaced by the first element in each matrix. In addition, since the matrix for Class2 The first element of Figure 7 The corresponding sO in Table 7 iThere are relatively large differences, so Figure 7 , for x=0…6, y=0…63, the elements except the first element are modified using the following equation (22).

[0127] Storing the modified matrix instead of the original one does not add any extra operations during encoding and decoding. This has at least two benefits. First, the offset sO table, which depends on the class index and matrix number, can be removed, saving 238 bits of storage space. Second, the process of extracting the offset from the matrix is unified for all classes.

[0128] In the third exemplary embodiment, the first element of each matrix is replaced by Figure 7 The corresponding offset sO in Table 7 is used. This allows the sO table to be deleted. When performing matrix-vector multiplication, the offset is taken from the first element of each matrix. By storing the modified matrix instead of the original matrix, no additional operations are added during encoding and decoding.

[0129] In the fourth exemplary embodiment, the offset s0 is replaced with a fixed value. Thus, the lookup table can be eliminated without any offset derivation process. In one example, the fixed value is 66, which is the minimum value among all matrices. All matrices are modified as follows. M′=M-sO+66 Equation (23)

[0130] In the above equation (23), sO is obtained from Table 7 ( Figure 7 ). Then, the matrix-vector multiplication process in equation (12) can be modified as follows.

[0131] By storing the modified matrix M' instead of the original matrix, no extra operations are added during encoding and decoding.

[0132] For another example, the fixed value is 64. All matrices are modified as follows. M′=M-sO+64 Eq.(25)

[0133] In equation (25) above, sO is derived from Table 7. Then, the matrix-vector multiplication process can be modified as follows.

[0134] In addition, the negative numbers in the modified matrix need to be changed to 0. When implementing the embodiment of the present invention, only one value is changed from -2 to 0. The modified matrix M' is saved instead of the original matrix, and no additional operations are added during the encoding and decoding process. The multiplication by 64 operation can be replaced by a shift operation. Therefore, sO is equal to the input vector input red The multiplications between can be replaced by left shift operations. For a 4×4 block, the total number of multiplications per sample required to generate the prediction signal is reduced from 5 to 4.

[0135] In the third example, the fixed value is 128. All matrices are modified as follows. M′=M-sO+128 (27)

[0136] In the above equation (27), sO is derived from Table 7. Then the matrix-vector multiplication process can be modified as follows.

[0137] By storing the modified matrix M' instead of the original matrix, no additional operations are added during encoding and decoding. The lookup table is removed, and the multiplication of the offset by the input vector is replaced by a left shift operation. For a 4×4 block, the total number of multiplications required per sample to generate the prediction signal is reduced to 4. Furthermore, coding performance is unchanged.

[0138] Figure 11 is a flow chart of an exemplary method 1100 for processing video content consistent with an embodiment of the present disclosure. The method 1100 may be performed by a codec (e.g., using Figures 2A-2B The encoding process 200A and 200B of the encoder or use Figures 3A-3B The decoding processes 300A and 300B of the apparatus may be performed by a decoder). For example, a codec may be implemented as one or more software or hardware components of an apparatus (e.g., apparatus 400) for encoding or transcoding a video sequence. In some embodiments, the video sequence may be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence being decoded (e.g., video stream 304). In some embodiments, the video sequence may be a monitoring video sequence, which may be processed by a monitoring device (e.g., processor 402) associated with a processor of the apparatus. Figure 4 The video sequence may include multiple images. The apparatus may perform method 1100 at the image level. For example, in method 1100, the apparatus may process one image at a time. For another example, in method 1100, the apparatus may process multiple images at a time. Method 1100 may include the following steps.

[0139] At step 1102, a classification of the target block may be determined. In some embodiments, the classification may include a first class (e.g., Class 0), a second class (e.g., Class 1), and a third class (e.g., Class 2). For a given block, the classification of the given block may be determined based on the size of the given block. For example, the first class may be associated with blocks of size 4×4, and the second class may be associated with blocks of size 8×8, 4×N, or N×4, where N may be an integer between 8 and 64. For example, N is equal to 8, 16, 32, or 64. That is, the second class may include blocks of size 8×8, 4×8, 4×16, 4×32, 4×64, 8×4, 16×4, 32×4, or 64×4. The third class may be associated with the remaining blocks.

[0140] In some embodiments, in response to the size of the target block not being 4×4, 8×8, 4×N, or N×4, it may be determined that the target block belongs to the third category.

[0141] At step 1104, a matrix-weighted intra prediction (MIP) signal may be generated based on the classification. In some embodiments, a first intra prediction signal for the target block may be generated based on the input vector, the matrix, and the classification of the target block, and bilinear interpolation may be performed on the target block using the first intra prediction signal to generate the MIP signal.

[0142] For example, in order to generate an input vector, adjacent reconstructed samples of the target block may be averaged according to the classification of the target block. As described above, for a first-class block, every two adjacent reconstructed samples of the block may be averaged to generate a reduced boundary vector as the input vector. For example, the size of the first-class input vector may be 4×1, the second-class may be 8×1, and the third-class may be 7×1. And for a block of the second or third class (e.g., having a size of M×N), every M / 4 adjacent reconstructed samples above the block and every N / 4 adjacent reconstructed samples to the left of the block may be averaged.

[0143] Different from the input vector, the matrix may be selected from a set of matrices (eg, matrix set S0, S1, or S2) according to the classification and MIP mode index of the target block.

[0144] Then, a first intra-frame prediction signal can be generated by performing a matrix-vector multiplication on the matrix and the input vector. In some embodiments, the first intra-frame prediction signal is further associated with a first offset and a second offset. For example, as discussed in equation (12), the reduced prediction signal can be further offset by a first offset (e.g., oW) and a second offset (e.g., oS). In some embodiments, the first offset and the second offset can be determined based on a matrix index in the matrix set. For example, the first offset and the second offset can be determined with reference to Table 7 and Table 8, respectively.

[0145] The classification of the target block is also related to the size of the first intra-frame prediction signal. For example, in response to the target block belonging to the first or second category, the size of the first intra-frame prediction signal is determined to be 4×4; in response to the target block belonging to the third category, the size of the first intra-frame prediction signal is determined to be 8×8.

[0146] In some embodiments, a non-transitory computer-readable storage medium comprising instructions is also provided, and the instructions can be executed by a device for performing the above method (e.g., the disclosed encoder and decoder). Common forms of non-transitory media include, for example, floppy disks, disks, hard disks, solid-state drives, tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with hole patterns, RAM, PROM and EPROM, FLASH-EPROM or any other flash memory, NVRAM, caches, registers, any other memory chips or cassettes, and network versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.

[0147] The embodiments may be further described using the following terms: 1. A computer-implemented method for processing video content, comprising: determining a classification of the target block; and generating a matrix-weighted intra prediction (MIP) signal based on the classification, Determining the classification of the target block includes: In response to the size of the target block being 4×4, determining that the target block belongs to the first category; or In response to the size of the target block being 8×8, 4×N, or N×4, where N is an integer between 8 and 64, it is determined that the target block belongs to the second category. 2. The method of clause 1, wherein generating the MIP signal comprises: generating a first intra prediction signal for the target block, wherein the generation of the first intra prediction signal is based on an input vector, a matrix, and a classification of the target block; and The target block is bilinearly interpolated using the first intra-frame prediction signal to generate the MIP signal. 3. The method according to clause 1 or 2, further comprising: According to the classification of the target block, adjacent reconstructed samples of the target block are averaged to generate the input vector. 4. The method according to clause 2 or 3, wherein the size of the input vector is 4×1 when the target block belongs to the first class, or the size of the input vector is 8×1 when the target block belongs to the second class. 5. A method according to any of clauses 2-4, wherein the matrix is selected from a set of matrices according to the classification of the target block and a MIP mode index. 6. The method of clause 5, wherein the first intra prediction signal is generated by performing a matrix-vector multiplication on the matrix and the input vector. 7. The method according to clause 6, wherein the first intra prediction signal is generated based on one or more offsets associated with a matrix. 8. The method of clause 7, wherein the one or more offsets are determined based on an index into a matrix in a lookup table. 9. The method according to any one of clauses 1 to 8, wherein determining the classification of the target block further comprises: In response to the size of the target block being other than 4×4, 8×N, 4×N, and N×4, it is determined that the target block belongs to the third category. 10. The method of clause 9, wherein generating the first intra prediction signal for the target block comprises: In response to the target block belonging to the first class or the second class, determining the size of the first intra prediction signal to be 4×4; and In response to the target block belonging to the third category, the size of the first intra prediction signal is determined to be 8×8. 11. A method according to any one of clauses 1 to 10, wherein N is equal to 8, 16, 32 or 64. 12. A video content processing system, comprising: a memory for storing an instruction set; and at least one processor configured to execute a set of instructions to cause the system to: determining a classification of the target block; and generating a matrix-weighted intra prediction (MIP) signal based on the classification, When determining the classification of the target block, the at least one processor is further configured to execute the instruction set so that the system further performs: In response to the size of the target block being 4×4, determining that the target block belongs to the first category; or In response to the size of the target block being 8×8, 4×N, or N×4, where N is greater than 4, it is determined that the target block belongs to the second category. 13. The system of clause 12, wherein when generating the MIP signal, the at least one processor is further configured to execute the set of instructions to cause the system to further perform: generating a first intra prediction signal for the target block, wherein the generation of the first intra prediction signal is based on an input vector, a matrix, and a classification of the target block; and The target block is bilinearly interpolated using the first intra-frame prediction signal to generate the MIP signal. 14. A system according to clause 12 or 13, wherein the at least one processor is further configured to execute the set of instructions so that the system further performs: According to the classification of the target block, adjacent reconstructed samples of the target block are averaged to generate the input vector. 15. A system according to clause 13 or 14, wherein the size of the input vector is 4×1 when the target block belongs to the first class, or the size of the input vector is 8×1 when the target block belongs to the second class. 16. A system according to any of clauses 13-15, wherein the matrix is selected from a set of matrices according to the classification of the target block and a MIP mode index. 17. The system of clause 16, wherein the first intra prediction signal is generated by performing a matrix-vector multiplication on the matrix and the input vector. 18. The system of clause 17, wherein the first intra prediction signal is generated based on one or more offsets associated with the matrix. 19. The system of clause 18, wherein the one or more offsets are determined based on an index into a matrix in a lookup table. 20. A non-transitory computer-readable medium having stored thereon a set of instructions executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising: determining a classification of the target block; and generating a matrix-weighted intra prediction (MIP) signal based on the classification, Determining the classification of the target block includes: In response to the size of the target block being 4×4, determining that the target block belongs to the first category; or In response to the size of the target block being 8×8, 4×N, or N×4, where N is an integer between 8 and 64, it is determined that the target block belongs to the second category.

[0148] It should be noted that the relational terms such as "first" and "second" in this document are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "include", "have", "include" and "including" and other similar forms are equivalent in meaning and are open-ended, in that the one or more items following any of these words are not intended to be an exhaustive list of such items or items, or to be limited to the listed items.

[0149] As used herein, unless specifically stated otherwise, the term "or" encompasses all possible combinations unless not feasible. For example, if a database is stated to contain either A or B, then, unless explicitly stated otherwise or not feasible, the database may contain either A, or B, or A and B. As a second example, if a database is stated to contain either A, B, or C, then, unless explicitly stated otherwise or not feasible, the database may contain either A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.

[0150] It will be appreciated that the above embodiments may be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it may be stored in the above-mentioned computer-readable medium. The software may execute the disclosed method when executed by a processor. The computing units and other functional units described in the present disclosure may be implemented by hardware, or software, or a combination of hardware and software. It will also be appreciated by those skilled in the art that multiple of the above-mentioned modules / units may be combined into one module / unit, and each of the above-mentioned modules / units may be further divided into multiple submodules / subunits.

[0151] In the foregoing description, embodiments have been described with reference to numerous specific details that may vary depending on the implementation. Certain modifications and variations may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. The description and embodiments are to be considered exemplary only, with the true scope and spirit of the invention being indicated by the following claims. The order of steps shown in the figures is also intended to be illustrative only and is not intended to be limiting to any particular order of steps. Therefore, those skilled in the art will appreciate that the steps may be performed in a different order while implementing the same method.

[0152] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications may be made to these embodiments. Therefore, although specific terms are used, they are used in a general and descriptive sense only and not for the purpose of limitation.

Claims

1. A video encoding device, characterized in that: include: a memory for storing instructions; and one or more processors configured to execute the instructions so that the apparatus performs the following operations: Determine the classification of the target block; and Generating a matrix-weighted intra prediction (MIP) signal based on the classification, wherein generating the MIP signal comprises: generating a first intra prediction signal for the target block, and performing bilinear interpolation on the target block using the first intra prediction signal to generate the MIP signal; Determining the classification of the target block includes: In response to the size of the target block being 4×4, determining that the target block belongs to the first category; or In response to the size of the target block being 8×8, 4×8, 8×4, 4×N, or N×4, where N is 16, 32, or 64, it is determined that the target block belongs to the second category.

2. The apparatus of claim 1 , wherein generating the MIP signal comprises: The generation of the first intra prediction signal is based on an input vector, a matrix and a classification of the target block.

3. The device according to claim 2, wherein When the target block belongs to the first category, the size of the input vector is 4×1, and when the target block belongs to the second category, the size of the input vector is 8×1.

4. The device according to claim 2, wherein The matrix is selected from a set of matrices according to the classification and MIP mode index of the target block.

5. The apparatus according to claim 1 , wherein determining the classification of the target block further comprises: In response to the size of the target block being not 4×4, 8×8, 4×8, 8×4, 4×N, and N×4, where N is 16, 32, or 64, it is determined that the target block belongs to the third category.

6. The apparatus of claim 1 , wherein generating the first intra prediction signal for the target block comprises: In response to the target block belonging to the first category or the second category, determining the size of the first intra prediction signal to be 4×4; and In response to the target block belonging to the third category, the size of the first intra prediction signal is determined to be 8×8.

7. A video decoding device, characterized in that: include: a memory for storing instructions; and one or more processors configured to execute the instructions so that the apparatus performs the following operations: Determine the classification of the target block; and generating a matrix-weighted intra prediction (MIP) signal based on the classification, wherein generating the MIP signal comprises: generating a first intra prediction signal for the target block, and performing bilinear interpolation on the target block using the first intra prediction signal to generate the MIP signal; When determining the classification of the target block, the at least one processor is further configured to execute the instruction set so that the system further performs: In response to the size of the target block being 4×4, determining that the target block belongs to the first category; or In response to the size of the target block being 8×8, 4×8, 8×4, 4×N, or N×4, where N is 16, 32, or 64, it is determined that the target block belongs to the second category. 8 . The apparatus of claim 7 , wherein, in generating the MIP signal, generation of the first intra prediction signal is based on classification of an input vector, a matrix, and a target block.

9. The device according to claim 8, wherein When the target block belongs to the first category, the size of the input vector is 4×1, or when the target block belongs to the second category, the size of the input vector is 8×1.

10. The device according to claim 7, wherein The matrix is selected from a set of matrices according to the classification and MIP mode index of the target block. 11 . The apparatus of claim 10 , wherein the first intra prediction signal is generated by performing matrix-vector multiplication on the matrix and an input vector.

12. The device according to claim 11, wherein The first intra prediction signal is generated based on one or more offsets associated with the matrix.

13. The apparatus of claim 12, wherein the one or more offsets are determined based on an index into the matrix in a lookup table.

14. A method for storing a bit stream, characterized in that Execute the following video encoding method to generate a bit stream, and store the bit stream: Determine the classification of the target block; and Generating a matrix-weighted intra prediction (MIP) signal based on the classification, wherein generating the MIP signal comprises: generating a first intra prediction signal for the target block, and performing bilinear interpolation on the target block using the first intra prediction signal to generate the MIP signal; Determining the classification of the target block includes: In response to the size of the target block being 4×4, determining that the target block belongs to the first category; or In response to the size of the target block being 8×8, 4×8, 8×4, 4×N, or N×4, where N is 16, 32, or 64, it is determined that the target block belongs to the second category.

15. A method for transmitting a bit stream, characterized in that: The following video encoding method is executed to generate a bit stream, and the bit stream is transmitted: Determine the classification of the target block; and Generating a matrix-weighted intra prediction (MIP) signal based on the classification, wherein generating the MIP signal comprises: generating a first intra prediction signal for the target block, and performing bilinear interpolation on the target block using the first intra prediction signal to generate the MIP signal; Determining the classification of the target block includes: In response to the size of the target block being 4×4, determining that the target block belongs to the first category; or In response to the size of the target block being 8×8, 4×8, 8×4, 4×N, or N×4, where N is 16, 32, or 64, it is determined that the target block belongs to the second category.

16. A non-transitory computer readable medium having computer instructions and a bit stream stored thereon, characterized in that: When the computer instructions are executed by a processor, the following video encoding method is implemented to generate the bitstream, the method comprising: determining a classification of the target block; and Generating a matrix-weighted intra prediction (MIP) signal based on the classification, wherein generating the MIP signal comprises: generating a first intra prediction signal for the target block, and performing bilinear interpolation on the target block using the first intra prediction signal to generate the MIP signal; Determining the classification of the target block includes: In response to the size of the target block being 4×4, determining that the target block belongs to the first category; or In response to the size of the target block being 8×8, 4×8, 8×4, 4×N, or N×4, where N is 16, 32, or 64, it is determined that the target block belongs to the second category.