Matrix weighted intra prediction of video signals
By simplifying the matrix-weighted intra prediction method, video blocks are classified and matrix-weighted intra prediction signals are generated, which solves the problems of high-definition videos with high storage and high bandwidth, improves encoding efficiency and simplifies the calculation process.
Patent Information
- Application Number
- CN202080053454.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-30
- Filing Date
- 2020-07-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-07-28
AI Technical Summary
The existing video encoding and decoding technology has high storage and transmission requirements when processing high-definition video, resulting in high bandwidth and large storage requirements, making it difficult to efficiently compress and decode.
The simplified matrix-weighted intra prediction method is adopted to classify the target blocks and generate the matrix-weighted intra prediction signals based on the classification, simplify the calculation process and unify the prediction process of different sizes and blocks.
Improves video encoding efficiency, reduces the need for storage space and transmission bandwidth, simplifies the computing process, and reduces the encoding complexity.
Smart Images

Figure CN114145016B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This disclosure claims the benefit of priority to U.S. Provisional Application No. 62 / 894,489, filed on August 30, 2019, the entire content of which is incorporated herein by reference. Technical field
[0003] This disclosure generally relates to video processing, and more particularly, to methods and systems for performing simplified matrix - weighted intra - prediction of video signals. Background art
[0004] Video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding and decoding techniques, and the most common ones are based on prediction, transformation, quantization, entropy coding, and loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, specify specific video coding formats and are developed by standardization organizations. As more and more video standards adopt advanced video coding and decoding techniques, the coding and decoding efficiency of new video coding standards is also getting higher and higher. Summary of the invention
[0005] Embodiments of the present invention provide a method for simplified matrix - weighted intra - prediction. The method may include: determining the classification of a target block; and generating a matrix - weighted intra - prediction (MIP) signal based on the classification, wherein determining the classification of the target block includes: in response to the size of the target block being 4×4, determining that the target block belongs to a first category; or in response to the size of the target block being 8×8, 4×N, or N×4, where N is an integer between 8 and 64, determining that the target block belongs to a second category.
[0006] Embodiments of the present disclosure also provide a system for performing simplified matrix - weighted intra - prediction. The system may include: a memory for storing an instruction set; and at least one processor configured to execute the instruction set to cause the system to perform: determining the classification of a target block; and generating a matrix - weighted intra - prediction (MIP) signal based on the classification, wherein determining the classification of the target block includes: in response to the size of the target block being 4×4, determining that the target block belongs to a first category; or in response to the size of the target block being 8×8, 4×N, or N×4, where N is an integer between 8 and 64, determining that the target block belongs to a second category.
[0007] Embodiments of the present disclosure also provide a non - transitory computer - readable medium that stores a set of instructions executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content. The method may include: determining a classification of a target block; and generating a matrix - weighted intra - prediction (MIP) signal based on the classification, wherein determining the classification of the target block includes: in response to the size of the target block being 4×4, determining that the target block belongs to a first class; or in response to the size of the target block being 8×8, 4×N, or N×4, where N is an integer between 8 and 64, determining that the target block belongs to a second class. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings. The various features shown in the drawings are not drawn to scale.
[0009] Figure 1 The structure of an exemplary video sequence consistent with embodiments of the present disclosure is shown.
[0010] Figure 2A A schematic diagram of an exemplary encoding process performed by a hybrid video codec system consistent with embodiments of the present disclosure is shown.
[0011] Figure 2B A schematic diagram of another exemplary encoding process performed by a hybrid video codec system consistent with embodiments of the present disclosure is shown.
[0012] Figure 3A A schematic diagram of an exemplary decoding process performed by a hybrid video codec system consistent with embodiments of the present disclosure is shown.
[0013] Figure 3B A schematic diagram of another exemplary decoding process performed by a hybrid video codec system consistent with embodiments of the present disclosure is shown.
[0014] Figure 4 is a block diagram of an exemplary apparatus for encoding or decoding video consistent with embodiments of the present disclosure.
[0015] FIG. 5 shows an exemplary schematic diagram of matrix - weighted intra - prediction consistent with embodiments of the present disclosure.
[0016] Figure 6 A table illustrating three exemplary classes used in matrix - weighted intra - prediction consistent with embodiments of the present disclosure is shown.
[0017] Figure 7 A lookup table for determining an offset “sO” consistent with embodiments of the present disclosure is shown.
[0018] Figure 8Illustrated is an exemplary lookup table for determining an offset “sW” consistent with an embodiment of the present disclosure.
[0019] FIG. 9 shows an exemplary matrix consistent with an embodiment of the present disclosure, which shows an exemplary exclusion operation.
[0020] FIG. 10 illustrates another exemplary matrix consistent with an embodiment of the present disclosure that shows another exemplary exclusion operation.
[0021] Figure 11 is a flowchart of an exemplary method for processing video content consistent with an embodiment of the present disclosure. Detailed Description
[0022] Reference will now be made in detail to exemplary embodiments, which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in all cases, unless otherwise noted, and the same numbers in different drawings represent the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with aspects related to the present invention recited in the appended claims. Unless otherwise specifically noted, the term “or” encompasses all possible combinations, unless infeasible. For example, if it is stated that a component may include A or B, then, unless otherwise explicitly stated or infeasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then, unless otherwise explicitly stated or infeasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0023] Video coding systems are typically used to compress digital video signals, for example to reduce the storage space consumption or transmission bandwidth consumption associated with such signals. As high-definition (HD) video (e.g., with a resolution of 1920×1080 pixels) becomes increasingly popular in various applications of video compression, such as online video streaming, video conferencing, or video surveillance, the need to develop coding tools for improving the compression efficiency of video data is constantly increasing.
[0024] For example, video surveillance applications are becoming increasingly widespread in many scenarios (such as security, traffic, environmental monitoring, etc.), and the number and resolution of surveillance devices are growing rapidly. Many video surveillance application scenarios tend to provide users with high-definition videos to capture more information, and high-definition videos have more pixels per frame to capture this information. However, high-definition video bitstreams may have high bitrates, requiring high-bandwidth transmission and large storage space. For example, a surveillance video stream with an average resolution of 1920×1080 may require a bandwidth of up to 4 Mbps for real-time transmission. Additionally, video surveillance is generally continuous 7×24 monitoring, and if video data is to be stored, this poses a great challenge to the storage system. Therefore, the high bandwidth and large storage requirements of high-definition videos have become the main constraints for their large-scale deployment in video surveillance.
[0025] Video is a set of static images (or "frames") arranged in a time sequence to store visual information. Video capture devices (such as cameras) can be used to capture and store these images in a time sequence, and video playback devices (such as televisions, computers, smartphones, tablets, video players, or any end-user terminal with a display function) can be used to display such images in chronological order. In addition, in some applications, video capture devices can transmit the captured video in real time to a video playback device (such as a computer with a monitor) for monitoring, conferencing, live streaming, etc.
[0026] To reduce the storage space and transmission bandwidth required for such applications, videos can be compressed before storage and transmission and decompressed before display. Compression and decompression can be achieved through software and executed by a processor (such as the processor of a general-purpose computer) or dedicated hardware. The compression module is generally called an "encoder", and the decompression module is generally called a "decoder". Encoders and decoders can be collectively referred to as "codecs". Encoders and decoders can be implemented as any one of a variety of suitable hardware, software, or combinations thereof. For example, the hardware implementation of encoders and decoders can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be achieved through various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, a codec can decompress a video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be called a "transcoder".
[0027] The video encoding process can identify and preserve useful information that can be used to reconstruct an image, while ignoring unimportant information for reconstruction. If the ignored, unimportant information cannot be fully reconstructed, such an encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy, which is a trade-off to reduce the required storage space and transmission bandwidth.
[0028] The useful information of the image being encoded (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include changes in the position of pixels, changes in brightness, or changes in color, with the change in position being the most concerned. The change in the position of a group of pixels representing an object can reflect the movement of the object between the reference image and the current image.
[0029] Depending on whether the reference image is the current image itself or another image, the encoding of the current image can be divided into "inter-frame prediction" and "intra-frame prediction". Intra-frame prediction can utilize spatial redundancy (e.g., the correlation between pixels within a frame) by calculating a predicted value through interpolation from already encoded pixels. Inter-frame prediction can utilize the temporal difference (e.g., motion vectors) between adjacent frames (e.g., the reference frame and the target frame), thereby enabling the codec for the target frame. The present disclosure relates to techniques for intra-frame prediction.
[0030] The present disclosure provides methods, apparatuses, and systems for performing simplified matrix weighted intra-frame prediction of video signals. By removing the additional omission operations of the matrix in the MIP prediction process, the prediction processes for different-sized blocks can be unified, and the calculation process can also be simplified.
[0031] Figure 1 The structure of an example video sequence 100 consistent with an embodiment of the present disclosure is illustrated. The video sequence 100 can be a live video or a video that has been captured and archived. The video 100 can be a real video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., a real video with augmented reality effects). The video sequence 100 can be received from a video capture device (e.g., a camera), a video archive containing previously captured videos (e.g., a video file stored in a storage device), or a video providing interface (e.g., a video broadcast transceiver) that receives videos from a video content provider.
[0032] As Figure 1 shown, the video sequence 100 can include a series of images arranged in time along a time axis, including images 102, 104, 106, and 108. Images 102 - 106 are consecutive, and there are more images between images 106 and 108. At Figure 1In [the figure], image 102 is an I image, and its reference image is image 102 itself. Image 104 is a P image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B image, and its reference images are images 104 and 108, as indicated by the arrow. In some embodiments, the reference image of an image (e.g., image 104) may not be immediately before or after that image. For example, the reference image of image 104 may be an image before image 102. It should be noted that the reference images of images 102-106 are only examples, and the embodiments of the present disclosure regarding reference images are not limited to the Figure 1 example shown.
[0033] Generally, due to the computational complexity of such tasks, video codecs do not encode or decode an entire image at once. Instead, they can divide the image into basic segments and encode or decode the image segment by segment. Such a basic segment is referred to as a basic processing unit ("BPU") in the present disclosure. For example, Figure 1 the structure 110 in [the figure]. Figure 1 shows an example structure of an image (e.g., any one of images 102-108) of the video sequence 100. In structure 110, the image is divided into 4×4 basic processing units, and their boundaries are indicated by dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in an image can have variable sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing unit for an image can be selected based on the balance between coding efficiency and the level of detail to be maintained in the basic processing unit.
[0034] The basic processing unit can be a logical unit that can include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, the basic processing unit of a color image can include a luminance component (Y) representing non-color luminance information, one or more chrominance components representing color information (e.g., Cb and Cr), and associated syntax elements, where the luminance and chrominance components can have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components can be referred to as "coding tree blocks" ("CTB"). Any operation performed on the basic processing unit can be repeated for each of its luminance and chrominance components.
[0035] Video encoding has multiple operational stages, examples of which will be detailed in Figures 2A - 2B and FIGS. 3A-3B. For each stage, the size of the basic processing unit may still be too large to process, so it can be further divided into segments, which are referred to as "basic processing subunits" in the present disclosure. In some embodiments, the basic processing subunit may be referred to as a "block" in some video encoding standards (e.g., the MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding unit" ("CU") in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same or a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr associated with syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be carried out to a deeper level according to processing needs. It should also be noted that different stages may use different schemes to divide the basic processing unit.
[0036] For example, in the mode decision stage (examples of which will be detailed in Figure 2B ), the encoder may decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, and the basic processing unit may be too large to make this decision. The encoder may split the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC) and decide the prediction type for each individual basic processing subunit.
[0037] As another example, in the prediction stage (examples of which will be detailed in Figure 2A ), the encoder may perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder may further split the basic processing subunit into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and prediction operations can be performed at these levels.
[0038] As another example, in the transform stage (examples of which will be detailed in Figure 2AAs detailed in [description], the encoder can perform transformation operations for the residual basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further split the basic processing subunit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and transformation operations can be performed on these segments. It should be noted that the partitioning scheme of the same basic processing subunit in the prediction stage and the transformation stage can be different. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU can have different sizes and quantities.
[0039] In the structure 110 of the figure, the basic processing unit 112 is further divided into 3×3 basic processing subunits, and their boundaries are shown as dotted lines. Different basic processing units of the same image can be divided into different basic processing subunits according to different schemes.
[0040] In some embodiments, to provide parallel processing and fault tolerance capabilities for video encoding and decoding, an image can be divided into multiple regions for processing, such that for one region of the image, the encoding or decoding process does not have to depend on information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving the encoding efficiency. And when the data of one region is damaged or lost during network transmission during processing, the codec can correctly encode or decode other regions of the same image without relying on the damaged or lost data, thereby providing fault tolerance capabilities. In some video coding standards, an image can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slice" and "tile". It should also be noted that different images of the video sequence 100 can have different partitioning schemes for dividing the image into multiple regions.
[0041] For example, in Figure 1 , the structure 110 is divided into three regions 114, 116, and 118, and their boundaries are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that Figure 1 the basic processing units, basic processing subunits, and regions of the structure 110 in [[ ]] are only examples, and the present disclosure is not limited to the embodiments therein.
[0042] Figure 2A Illustrates a schematic diagram of an example encoding process 200A consistent with an embodiment of the present disclosure. The encoder can encode the video sequence 202 into a video bitstream 228 according to process 200A. Similar toFigure 1 In the video sequence 100, the video sequence 202 may include a set of images arranged in chronological order (referred to as "original images"). Similar to Figure 1 the structure 110 in, each original image of the video sequence 202 may be divided by the encoder into basic processing units, basic processing subunits, or regions for processing. In some embodiments, the encoder may perform process 200A at the basic processing unit level for each original image of the video sequence 202. For example, the encoder may perform process 200A in an iterative manner, where in one iteration of process 200A, the encoder may encode a basic processing unit. In some embodiments, the encoder may perform process 200A in parallel for regions (e.g., regions 114 - 118) of each original image of the video sequence 202.
[0043] In the figure reference Figure 2A , the encoder may provide the basic processing units (referred to as "original BPUs") of the original images of the video sequence 202 to the prediction stage 204 to generate prediction data 206 and prediction BPUs 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to the transform stage 212 and the quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to the binary coding stage 226 to generate a video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". In process 200A, after the quantization stage 214, the encoder may provide the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of process 200A. The components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path". The reconstruction path may be used to ensure that the encoder and the decoder use the same reference data for prediction.
[0044] The encoder may iteratively perform process 200A to encode each original BPU (in the forward path) of the original image and generate a prediction reference 224 for encoding the next original BPU (in the reconstruction path) of the original image. After encoding all the original BPUs of the original image, the encoder may continue to encode the next image in the video sequence 202.
[0045] Referring to process 200A, the encoder can receive video sequence 202 generated by a video capture device (e.g., a camera). The term "receive" as used herein can refer to any action of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or otherwise inputting data.
[0046] In prediction stage 204, in the current iteration, the encoder can receive the original BPU and prediction reference 224, and perform a prediction operation to generate prediction data 206 and prediction BPU 208. The prediction reference 224 can be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206, which can be used to reconstruct the original BPU into prediction BPU 208 from prediction data 206 and prediction reference 224.
[0047] Ideally, prediction BPU 208 can be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, prediction BPU 208 is typically slightly different from the original BPU. To record this difference, after generating prediction BPU 208, the encoder can subtract it from the original BPU to generate residual BPU 210. For example, the encoder can subtract the pixel value of prediction BPU 208 (e.g., grayscale value or RGB value) from the corresponding pixel value of the original BPU. Each pixel of residual 210 can have a residual value as the result of this subtraction between the corresponding pixels of the original BPU and prediction BPU 208. Compared to the original BPU, prediction data 206 and residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significantly degrading the quality. Thus, the original BPU is compressed.
[0048] To further compress residual BPU 210, in transform stage 212, the encoder can reduce its spatial redundancy by decomposing residual BPU 210 into a set of two-dimensional "basis patterns", each basis pattern associated with a "transformation coefficient". The "basis patterns" can have the same size (e.g., the size of residual BPU 210). Each basis pattern can represent a frequency component of the variation of residual BPU 210 (e.g., the frequency of brightness variation). The basis patterns cannot be replicated from any combination (e.g., linear combination) of any other basis templates. In other words, the decomposition can decompose the variation of residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the basis patterns are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basis functions.
[0049] Different transformation algorithms can use different basic patterns. Various transformation algorithms can be used in the transformation stage 212, such as discrete cosine transform, discrete sine transform, etc. The transformation at the transformation stage 212 is reversible. That is, the encoder can recover the residual BPU 210 through the inverse operation of the transformation (referred to as "inverse transformation"). For example, to recover the pixels of the residual BPU 210, the inverse transformation can be to multiply the values of the corresponding pixels of the basic pattern by their respective associated coefficients and sum the products to produce a weighted sum. For video coding standards, both the encoder and the decoder can use the same transformation algorithm (and thus have the same basic pattern). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 based on these coefficients without receiving the basic pattern from the encoder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significantly reducing the quality. Therefore, the residual BPU 210 is further compressed.
[0050] The encoder can further compress the transformation coefficients in the quantization stage 214. During the transformation process, different basic patterns can represent different change frequencies (e.g., brightness change frequency). Since the human eye is usually better at recognizing low-frequency changes, the encoder can ignore the information of high-frequency changes without significantly degrading the decoding quality. For example, in the quantization stage 214, the encoder can generate the quantized transformation coefficients 216 by dividing each transformation coefficient by an integer value (referred to as "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transformation coefficients of the high-frequency basic pattern can be converted to zero, while the transformation coefficients of the low-frequency basic pattern can be converted to smaller integers. The encoder can ignore the zero-valued quantized transformation coefficients 216 to further compress the transformation coefficients through the zero-valued quantized transformation coefficients. The quantization process is also reversible, where the quantized transformation coefficients 216 can be reconstructed as transformation coefficients in the inverse operation of quantization (referred to as "inverse quantization").
[0051] Because the encoder ignores the remainder of this division in the rounding operation, the quantization stage 214 is lossy. Generally, the quantization stage 214 contributes the most information loss in the process 200A. The greater the information loss, the fewer bits are required for the quantized transformation coefficients 216. To obtain different levels of information loss, the encoder can use different values of the quantization parameter or any other parameter of the quantization process.
[0052] In the binary encoding stage 226, the encoder may use binary encoding techniques to encode the prediction data 206 and the quantized transform coefficients 216. These binary encoding techniques include, for example, entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context - adaptive binary arithmetic encoding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type in the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bit - rate control parameters), etc. The encoder may use the output data of the binary encoding stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0053] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate the reconstructed transform coefficients. In the inverse transform stage 220, the encoder may generate the reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction reference 224 that will be used in the next iteration of process 200A.
[0054] It should be noted that other variables of process 200A may be used to encode the video sequence 202. In some embodiments, the stages of process 200A may be executed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, one or more stages of process 200A may be omitted Figure 2A from the process.
[0055] Figure 2B FIG. shows a schematic diagram of another exemplary encoding process 200B consistent with embodiments of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0056] Generally, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-frame prediction") can use pixels from one or more encoded neighboring BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") can use regions from one or more encoded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce the temporal redundancy inherent in the image.
[0057] Referring to process 200B, in the forward path, the encoder performs prediction operations at the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, at the spatial prediction stage 2042, the encoder can perform intra-frame prediction. For the original BPU of the image being encoded, the prediction reference 224 can include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder can generate a predicted BPU 208 by interpolating the neighboring BPUs. The interpolation technique can include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder can perform interpolation at the pixel level, for example, by performing an interpolation operation on the corresponding pixel values of each pixel of the predicted BPU 208. The neighboring BPUs used for interpolation can be located relative to the original BPU in various directions, such as in the vertical direction (e.g., on top of the original BPU), horizontal direction (e.g., to the left of the original BPU), diagonal direction (e.g., bottom left, bottom right, top left of the original BPU), or any direction defined in the video coding standard being used. For intra-frame prediction, the prediction data 206 can include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, the parameters of the interpolation, the direction of the neighboring BPUs relative to the original BPU, etc.
[0058] As another example, in the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed by the BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the reconstructed BPU. When all the reconstructed BPUs of the same image are generated, the encoder may generate the reconstructed image as a reference image. The encoder may perform an operation of "motion estimation" to search for a matching region within the range of the reference image (referred to as the "search window"). The position of the search window in the reference image may be determined according to the position of the original BPU in the current image. For example, the search window may be centered at the position in the reference image that has the same coordinates as the original BPU in the current image and may extend outward a predetermined distance. When the encoder identifies (e.g., by using a pixel recursion algorithm, a block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder may determine such a region as the matching region. The matching region may have a different size (e.g., smaller than, equal to, larger than, or of a different shape) from the original BPU. Since the reference image and the current image are temporally separated on the timeline (e.g., as Figure 1 shown), it can be considered that the matching region "moves" to the position of the original BPU over time. The encoder may record the direction and distance of this motion as a "motion vector". When multiple reference images are used (e.g., as in the image 106 in Figure 1 ), the encoder may search for the matching region and determine its associated motion vector for each reference image. In some embodiments, the encoder may assign weights to the pixel values of the matching regions of the respective matching reference images.
[0059] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.
[0060] To generate the prediction BPU 208, the encoder may perform an operation of "motion compensation". Motion compensation can be used to reconstruct the prediction BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder may move the matching region of the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When multiple reference images are used (e.g., as in Figure 1In the image 106), the encoder can move the matching region of the reference image according to each motion vector and the average pixel value of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of each matching reference image, the encoder can add the weighted sums of the pixel values of the moved matching regions.
[0061] In some embodiments, the inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same time direction as the current image. For example, Figure 1 the image 104 in is an unidirectional inter-frame prediction image, where the reference image (i.e., image 102) is before image 104. Bidirectional inter-frame prediction can use one or more reference images in two time directions relative to the current image. For example, Figure 1 the image 106 in is a bidirectional inter-frame prediction image, where the reference images (i.e., images 104 and 108) are in two time directions relative to image 104.
[0062] Still referring to the forward path of process 200B, after the spatial prediction 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform rate-distortion optimization techniques, where the encoder can select a prediction mode to minimize the value of a cost function based on the bit rate of the candidate prediction modes and the distortion of the reconstructed reference image under the candidate prediction modes. According to the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.
[0063] In the reconstruction path of process 200B, if an intra prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolation of the next BPU of the current image). If an inter prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current image in which all BPUs have been encoded and reconstructed), the encoder can provide the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion (e.g., blocking artifacts) introduced by inter prediction. The encoder can apply various loop filtering techniques at the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The loop filtered reference image can be stored in the buffer 234 (or "decoded picture buffer") for later use (e.g., as an inter prediction reference image for future images of the video sequence 202). The encoder can store one or more reference images in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder can encode the parameters of the loop filter (e.g., loop filter strength) in the binary coding stage 226, along with the quantized transform coefficients 216, prediction data 206, and other information.
[0064] Figure 3A FIG. illustrates a schematic diagram of an example decoding process 300A consistent with embodiments of the present disclosure. The process 300A can be a decompression process corresponding to Figure 2A the compression process 200A in. In some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., Figures 2A - 2B the quantization stage 214 in), generally, the video stream 304 is different from the video sequence 202. Similar to Figures 2A - 2B the processes 200A and 200B in, the decoder can perform the process 300A for each image encoded in the video bitstream 228 at the basic processing unit (BPU) level. For example, the decoder can perform the process 300A in an iterative manner, where the decoder can decode the basic processing unit in one iteration of the process 300A. In some embodiments, the decoder can perform the process 300A in parallel for each region (e.g., regions 114-118) of each image encoded in the video bitstream 228.
[0065] Refer to Figure 3A, the decoder can provide a portion of the video bitstream 228 associated with the basic processing unit of the encoded image (referred to as "encoded BPU") to the binary decoding stage 302. At the binary decoding stage 302, the decoder can decode this portion into prediction data 206 and quantized transform coefficients 216. The decoder can provide the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder can provide the prediction data 206 to the prediction stage 204 to generate a predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded image buffer in a computer memory). The decoder can provide the prediction reference 224 to the prediction stage 204 to perform prediction operations in the next iteration of process 300A.
[0066] The decoder can iteratively execute process 300A to decode each encoded BPU of the encoded image and generate a prediction reference 224 for encoding the next encoded BPU of the encoded image. After decoding all the encoded BPUs of the encoded image, the decoder can output the image to the video stream 304 for display and continue to decode the next encoded image in the video bitstream 228.
[0067] At the binary decoding stage 302, the decoder can perform the binary encoding techniques used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information at the binary decoding stage 302, such as the prediction mode, the parameters of the prediction operation, the transform type, the quantization processing parameters (e.g., quantization parameter), the encoder control parameters (e.g., bitrate control parameter), etc. In some embodiments, if the video bitstream 228 is transmitted through a network in a packetized form, the decoder can unpack the video bitstream 228 before providing it to the binary decoding stage 302.
[0068] Figure 3B A schematic diagram illustrating another example decoding process 300B consistent with an embodiment of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0069] In process 300B, for the encoded basic processing unit (referred to as "current BPU") of the decoded encoded image (referred to as "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data, depending on what prediction mode the encoder uses to encode the current BPU. For example, if the encoder uses intra prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating the intra prediction, parameters of the intra prediction operation, and the like. The parameters of the intra prediction operation can include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as references, the sizes of the adjacent BPUs, interpolation parameters, the directions of the adjacent BPUs relative to the original BPU, and the like. Again, for example, if the encoder uses inter prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating the inter prediction, parameters of the inter prediction operation, and the like. The parameters of the inter prediction operation can include, for example, the number of reference images associated with the current BPU, the weights respectively associated with the reference images, the positions (e.g., coordinates) of one or more matching regions in the respective reference images, one or more motion vectors respectively associated with the matching regions, and the like.
[0070] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. The details of performing such spatial prediction or temporal prediction are described in Figure 2B and will not be elaborated further below. After performing such spatial prediction or temporal prediction, the decoder can generate a predicted BPU 208. The decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in Figure 3A below.
[0071] In process 300B, the decoder can provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolation of the next BPU of the current image). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference image in which all BPUs have been decoded), the encoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can Figure 2BApply the loop filter to the prediction reference 224 in the manner described. The reference image for loop filtering can be stored in buffer 234 (e.g., the decoded picture buffer in a computer memory) for later use (e.g., as the inter-prediction reference picture for future coded pictures of the video bitstream 228). The decoder can store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction is used to encode the current BPU, the prediction data can further include the parameters of the loop filter (e.g., loop filter strength).
[0072] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding video that is consistent with embodiments of the present disclosure. As Figure 4 shown, the apparatus 400 can include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 can become a special-purpose machine for video encoding or decoding. The processor 402 can be any type of circuit capable of manipulating or processing information. For example, the processor 402 can include any combination of any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), system-on-chips (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, the processor 402 can also be a group of processors grouped as a single logical component. For example, as shown. As Figure 4 shown, the processor 402 can include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0073] The apparatus 400 can also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as Figure 4As shown, the stored data may include program instructions (e.g., program instructions for implementing the stages in processes 200A, 200B, 300A, or 300B) and data to be processed (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data to be processed (e.g., via bus 410) and execute the program instructions to perform operations on or manage the data to be processed. The memory 404 may include high-speed random access storage devices or non-volatile storage devices. In some embodiments, the memory 404 may include any number of random access memories (RAMs), read-only memories (ROMs), optical discs, magnetic disks, hard disk drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a set of memories grouped as a single logical component ( Figure 4 not shown in
[0074] The bus 410 may be a communication device for transferring data between components inside the device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., universal serial bus port, peripheral component interconnect express port), etc.
[0075] For ease of illustration and without causing ambiguity, in the present disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuits may be implemented entirely in hardware or as a combination of software, hardware, or firmware. In addition, the data processing circuits may be a single independent module or may be fully or partially combined into any other component of the device 400.
[0076] The device 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, any combination of wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0077] In some embodiments, optionally, the device 400 may further include a peripheral interface 408 to provide connections to one or more peripheral devices. As shown. As Figure 4 shown, the peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touch screen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to video archives), etc.
[0078] It should be noted that a video codec (e.g., the codec that executes processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0079] The present disclosure provides methods that can be executed by an encoder and / or a decoder to simplify matrix weighted intra prediction (MIP). The MIP method is a newly added intra prediction technique in VVC. The MIP mode is applicable to blocks with an aspect ratio max(width, height) / min(width, height) less than or equal to 4, and is only applicable to the luminance component. The MIP flag is signaled in parallel with the intra sub-partition mode, the multi-reference line intra prediction mode, or the most probable mode.
[0080] When encoding a block using the MIP mode, similar to traditional intra prediction modes, a row of reconstructed neighboring boundary samples on the left side of the block and a row of reconstructed neighboring boundary samples on the top of the block are used as inputs to predict the block. When the reconstructed neighboring boundary samples are not available, they can be generated as in traditional intra prediction. To predict a luminance sample, first the reconstructed neighboring boundary samples are averaged to generate a reduced boundary vector neighbor red , neighbor red [0] represents the first element of the reduced boundary vector. Then the reduced boundary vector neighbor red is used to generate an input vector input red , which is used in a matrix-vector multiplication process, and a reduced prediction signal pred red can be obtained. Finally, bilinear interpolation is performed on the reduced prediction signal pred red to generate the output of the MIP prediction signal pred. An illustration of the MIP prediction process is shown in FIG. 5.
[0081] MIP blocks are divided into three categories according to the width (W) and height (H) of the block:
[0082] Class0: For W = H = 4 (i.e., 4×4 blocks);
[0083] Class1: For max{W,H} = 8 (i.e., 4×8, 8×4, 8×8 blocks); and
[0084] Class 2: For max{W, H}>8.
[0085] As Figure 6 shown in Table 6 below, the differences among the three classes are the modulus, the number of matrices, the matrix size, the size of the input vector input red (input for the matrix-vector multiplication process), and the size of the reduced prediction signal input red (output of the matrix-vector multiplication process).
[0086] In the following description, Class 0, Class 1, and Class 2 are denoted as S0, S1, and S2 respectively. N i represents the number of matrices in the matrix set S i (i = 0, 1, 2). For the MIP mode k, the matrix is used for the matrix multiplication process, where represents the j-th matrix in the matrix set S i , and j is derived using the following equation (1).
[0087]
[0088] It should be noted that when the MIP mode k is greater than or equal to N i , swap operations and transpose operations are performed respectively in the steps of generating the reduced boundary vector neighbor red and the reduced prediction signal pred red . The details of these two operations are described above. In addition, the following equations (2) and (3) are used to determine whether a swap operation or a transpose operation is required.
[0089]
[0090]
[0091] As mentioned above, the generation of the output prediction signal pred of the current block is based on the following three steps, namely averaging, matrix-vector multiplication, and linear interpolation. The details of these steps are described below.
[0092] Among the boundary samples, 4 samples of Class 0 and 8 samples of Class 1 and Class 2 are extracted by averaging. For example, the reduced boundary vector neighbor red can be generated by averaging the reconstructed adjacent boundary samples according to the following rules.
[0093] Class 0: Take the average of every two samples. The size of the reduced boundary vector neighbor red is 4×1.
[0094] Class1 and Class2: For the adjacent boundary samples reconstructed above the current block, average every W / 4 samples. For the adjacent boundary samples reconstructed to the left of the current block, average every H / 4 samples. The reduced boundary vector neighbor red is of size 8×1.
[0095] The reduced boundary vector neighbor red is the vector obtained by averaging the adjacent boundary samples reconstructed above the current block concatenated with the vector obtained by averaging the adjacent boundary samples reconstructed to the left of the current block As described above, for MIP mode k greater than or equal to N i , perform the swap operation. For example, and the concatenation order of the two vectors can be swapped, as shown in Equation (4) below
[0096]
[0097] Then, the input vector input red for the matrix-vector multiplication is generated as follows.
[0098] For Class0 and Class1:
[0099] input red [0] = neighbor red [0] - (1 << (bitDepth - 1)),
[0100] input red [j] = neighbor red [j] - neighbor red [0], j = 1, …, size(neighbor red ) - 1,
[0101] Equation (5)
[0102] For Class2
[0103] input red [j] = neighbor red [j + 1] - neighbor red [0], j = 0, …, size(neighbor red ) - 2, Equation (6)
[0104] In the above Equations (5) and (6), neighbor red [0] represents the vector neighbor redThe first element. According to equations (5) and (6), the inSize of Class0, Class1, and Class2 for input red is 4, 8, and 7 respectively.
[0105] Using the vector input red as input for matrix-vector multiplication. The result is the reduced prediction signal pred red on the subsampled sample set in the current block. For example, from the reduced input vector input red , the reduced prediction signal pred red, can be generated, which is a signal on the downsampled block of width W red and height H red . Here, W red and H red are defined by the following equations (7) and (8) as:
[0106]
[0107]
[0108] As mentioned before, when the variable "isTransposed" is equal to 1, the reduced prediction signal pred red is transposed. Assuming the size of the final reduced prediction signal pred red is W red ×H red , then the size of the untransposed W′ red ×H′ red is derived from the following equations (9) and (10):
[0109]
[0110]
[0111] The vector of the reduced prediction signal pred′ red is calculated by computing the matrix-vector product according to the following equation (11):
[0112] pred′ red = M · input red + neighbor red [0] Equation (11)
[0113] Then, the vector pred′ red is arranged in raster scan order in the matrices pred red of sizes 4×4, 4×8, 8×4, and 8×8. The vector pred′ red of size 16×1Arranged in a 4×4 matrix pred red The examples in it are shown as follows:
[0114] pred′ red =[A,B,C,D,E,F,G,H,I,J,K,L,M,N,O,P] T
[0115]
[0116] In other words, for x = 0…W′ red -1 and y = 0…H′ red -1, the matrix pred can be calculated using the following equation (12) red
[0117]
[0118] In the above equation (12), the variable “inSize” is the size of the input vector input red and, as described above, is used to generate the reduced prediction signal pred red The matrix M is obtained from one of the three matrix sets S0, S1, S2 according to the block size classification and the MIP mode k.
[0119] The variables oW and sW are two predefined values, depending on the matrix used for matrix-vector multiplication. The variable oW is used to limit the precision of each element in the matrix to 7 bits, so that all elements are greater than or equal to 0. For example, the factor oW can be defined as the following equation (13):
[0120]
[0121] In the above equation (13), sO is an offset and is derived from a lookup table. For example, Figure 7 Table 7 of shows an exemplary lookup table for sO according to some disclosed embodiments.
[0122] In addition, the variable “sW” is derived using another lookup table. Figure 8 Table 8 of shows an exemplary lookup table for sW according to some disclosed embodiments.
[0123] In equation (12), the two variables “inch” and “incW” for leaving out half of the matrix rows for 4×16 and 16×4 blocks are defined as the following equations (14) and (15):
[0124]
[0125]
[0126] In the above equation (15), the variable predC is used to arrange the reduced prediction signal pred red into a W′ red ×H′ red matrix and is defined as the following equation (16):
[0127]
[0128] The input vector input red , where the size of inSize is equal to 7, and a 64-row by 7-column matrix is used for blocks belonging to Class2 to generate a 64-element vector. However, according to the equation
[0129]
[0130]
[0131] , 4×16 and 16×4 blocks only require 32 elements. This is because the reduced prediction signal pred red of Class2 can be arranged in an 8×8 manner, which exceeds the short side of a 4×16 or 16×4 block. Therefore, the omission operation is performed.
[0132] Figure 9 shows an exemplary exclusion operation according to some disclosed embodiments. As shown in Figure 9, for a 4×16 block with isTrnsposed = 0 and a 16×4 block with isTrnsposed = 1, the second row of every two rows is excluded from the matrix. Therefore, the reduced prediction signal pred red contains 32 elements, which are arranged in a 4×8 size.
[0133] Figure 10 shows an exemplary omission operation according to some disclosed embodiments. As shown in Figure 10, for a 16×4 block with isTrnsposed = 0 and a 4×16 block with isTranspose = 1, the last 8 rows of every 16 rows are excluded from the matrix. Therefore, the reduced prediction signal pred red contains 32 elements, which are arranged in an 8×4 size.
[0134] The output prediction signal at the remaining positions is generated from the reduced prediction signal on the subsampled set pred red by linear interpolation, which is a single-step linear interpolation in each direction.
[0135] As described above, the prediction in the MIP mode is generated using three steps, including averaging adjacent reconstructed samples, matrix-vector multiplication, and bilinear interpolation. The prediction process in the MIP mode is different from that of the traditional intra prediction mode. Although the MIP mode improves the coding efficiency, the design may be complex in the following two aspects.
[0136] Regarding the first aspect, there may be problems with the exclusion operations on 4×16 and 16×4 blocks during the matrix-vector multiplication process for the following three reasons. First, the exclusion operations not only add extra operations but also make the prediction process inconsistent because the exclusion operations are only applicable to 16×4 and 4×16 blocks. Second, for 16×4 and 4×16 blocks, the size of the reduced prediction signal pred red may be different before and after transposition. Therefore, additional derivation of the size of the non-transposed reduced prediction signals (such as W′ red and H′ red ) is required. Third, for 16×4 and 4×16 blocks, the size of the reduced prediction signal pred red may be different, as shown in the following equation (17).
[0137]
[0138] Regarding the second aspect, in order to limit the precision of each element in the matrix to 7 bits and ensure that all elements are non-negative, an offset sO is added to the reduced prediction signal during the matrix multiplication process. However, it may be unnecessary and complex for the following three reasons. First, additional memory is required to store the table of the offset sO. This table contains a total of 34 elements, each element being 7 bits. Therefore, a total of 238 bits of memory are required. Second, a table lookup operation is required to determine the value of sO using the class index and matrix number. Third, additional multiplication and addition operations are required to generate the reduced prediction signal pred red . In addition to calculating the matrix-vector multiplication between the matrix M and the input vector input red , a multiplication between sO and the input vector input red is further performed. For 4×4 blocks, the total number of multiplications per sample required to generate the prediction signal increases to 5.
[0139] The present disclosure provides methods to solve these problems without affecting the bit rate. In some exemplary methods, 8 bits instead of 7 bits are used to store the matrix. Then, equation (12) can be re-expressed as the following equation (18).
[0140]
[0141] In the above equation (18), all elements in the matrix are subtracted by the offset sO. The above storage and calculation methods for the matrix-vector product can produce the same-bit result. However, the number of bits for storing the matrix increases by 4882 (i.e., 5120×1 - 34×7) bits, and the bit width of the multiplication operation expands to 8 bits.
[0142] The method for removing the exclusion operation and the sO lookup table is described below.
[0143] To remove the extra exclusion operations of the matrix in the MIP prediction process, two methods are provided. According to the first method for removing the exclusion operation, the unreasonable classification method in the traditional MIP method can be modified, and the 4×16 and 16×4 blocks are moved from Class2 to Class1, so that the generated reduced prediction signal pred red does not exceed the short side limit.
[0144] In an exemplary embodiment, the MIP classification rule is modified as follows:
[0145] Class0: 4×4;
[0146] Class1: 4×N, 8×8, N×4, where N is an integer between 8 and 64; and
[0147] Class2: Others.
[0148] Through this modification, blocks of size 8×8, 4×8, 4×16, 4×32, 4×64, 8×4, 16×4, 32×4, or 64×4 can be moved from Class2 to Class1. Therefore, blocks of size 8×8, 4×8, 4×16, 4×32, 4×64, 8×4, 16×4, 32×4, or 64×4 can use the matrices in set S1, and in the matrix multiplication process, each has 16 rows and 8 columns. In this way, only 16 elements need to be generated to form a 4×4 reduced prediction signal pred red , thus the exclusion operation can be removed. This modification can be represented by the changes highlighted in strikethrough or italics as follows:
[0149] If both cbWidth and cbHeight are equal to 4, then MipSizeId[x][y] is set to be equal to 0.
[0150] Otherwise, if , cbWidth * cbHeight is less than or equal to 64, then MipSizeId[x][y] is set to be equal to 1.
[0151] Otherwise, MipSizeId[x][y] is set to be equal to 2.
[0152] This solution has at least three benefits.
[0153] First, the matrix multiplication process is simplified and unified. For blocks of size 4×16 or 16×4, the exclusion operation is removed. Therefore, the check of whether to perform the exclusion operation and the two variables "inch" and "incW" can be deleted. In addition, during the matrix-vector multiplication process, all blocks do not require additional operations on the matrix. Therefore, the matrix multiplication process is unified.
[0154] Second, the number of multiplications and additions for 4×16 and 16×4 blocks in the matrix multiplication process is reduced. In some embodiments, for blocks of size 4×16 or 16×4, a 32×7 matrix can be used for matrix-vector multiplication, while a 16×8 matrix can be used in the provided embodiments. Therefore, the number of multiplications and additions for 4×16 or 16×4 blocks can be reduced.
[0155] Third, the derivation of pred red is simplified and unified. For all blocks, the size of the reduced prediction signal pred red is consistent before and after transposition. Therefore, the additional derivations of W′ red and H′ red are deleted. For example, the derivation of the size of pred red can be simplified to the following equations (19) and (20).
[0156]
[0157]
[0158] And the size of the reduced prediction signal pred red is unified by the following equation (21):
[0159]
[0160] According to the second method of removing the exclusion operation, the MIP classification rules are modified as follows:
[0161] Class0: 4×4;
[0162] Class1: 4×8, 8×4, 4×16, and 16×4; and
[0163] Class2: Others.
[0164] The 4×16 and 16×4 blocks (italicized in the above modified MIP classification) are moved to Class1, and the 8×8 block is moved to Class2. Through this modification, the exclusion operation is deleted and the MIP classification rules are further simplified. In some embodiments, this modification can be expressed by the changes highlighted with a strikethrough or italicized below.
[0165] If both cbWidth and cbHeight are equal to 4, then MipSizeId[x][y] is set to be equal to 0.
[0166] Otherwise, if Min(cbWidth, cbHeight) is equal to 4, then MipSizeId[x][y] is set to be equal to 1.
[0167] Otherwise, MipSizeId[x][y] is set to be equal to 2.
[0168] To remove the table of the offset sO, embodiments of the present disclosure provide a method for modifying the value of the matrix M and the offset sO without a lookup table.
[0169] In the first exemplary embodiment, the offset sO is replaced with the first element in the matrix for Class0 the first element in the matrix for Class1 and the seventh element in the matrix for Class2 The i-th element in the matrix represents the i-th number counted in raster scan order starting from the upper left corner of the matrix. By doing so, the lookup table of the offset s0 depending on the class index and the matrix number shown in Figure 7 Table 7 can be deleted. In this way, 238 bits of storage space can be saved.
[0170] In the second exemplary embodiment, for all classes, the offset s0 is replaced with the first element in each matrix. Additionally, since the first element in the matrix for Class2 has a relatively large difference from Figure 7 the corresponding sO in i Table 7, therefore, in Figure 7 for x = 0…6, y = 0…63, the elements other than the first element are modified using the following equation (22).
[0171]
[0172] Storing the modified matrix instead of the original matrix adds no extra operations during the encoding and decoding processes. There are at least two benefits as follows. First, the lookup table of the offset sO depending on the class index and the matrix number can be deleted, thus saving 238 bits of storage space. Second, the process of extracting the offset from the matrix is unified for all classes.
[0173] In the third exemplary embodiment, the first element of each matrix is replaced with the one according to Figure 7The corresponding offset sO of Table 7. In this way, the sO table can be deleted. When performing matrix-vector multiplication, the offset comes from the first element of each matrix. Instead of storing the original matrix, the modified matrix is stored, and no additional operations are added during the encoding and decoding processes.
[0174] In the fourth exemplary embodiment, the offset s0 is replaced with a fixed value. Therefore, the lookup table can be deleted without any offset derivation process. In one example, the fixed value is 66, which is the minimum value in all matrices. All matrices are modified as follows.
[0175] M′ = M - sO + 66 Equation (23)
[0176] In the above equation (23), sO is derived from Table 7 ( Figure 7 ). Then, the matrix-vector multiplication process in Equation (12) can be modified as follows.
[0177]
[0178] Instead of storing the original matrix, the modified matrix M' is stored, and no additional operations are added during the encoding and decoding processes.
[0179] For another example, the fixed value is 64. All matrices are modified as follows.
[0180] M′ = M - sO + 64 Eq.(25)
[0181] In the above equation (25), sO is derived from Table 7. Then, the matrix-vector multiplication process can be modified as follows.
[0182]
[0183] In addition, the negative numbers in the modified matrix need to be modified to 0. When implementing the embodiments of the present invention, only one value is changed from -2 to 0. Instead of storing the original matrix, the modified matrix M' is saved, and no additional operations are added during the encoding and decoding processes. The operation of multiplying by 64 can be replaced by a shift operation. Therefore, the multiplication between sO and the input vector input red can be replaced by a left shift operation. For a 4×4 block, the total number of multiplications required for each sample to generate the prediction signal is reduced from 5 to 4.
[0184] In the third example, the fixed value is 128. All matrices are modified as follows.
[0185] M′ = M - sO + 128 (27)
[0186] In the above equation (27), sO is derived from Table 7. Then the matrix-vector multiplication process can be modified as follows.
[0187]
[0188] Instead of storing the original matrix, the modified matrix M' is stored, and no additional operations are added during the encoding and decoding processes. The lookup table is removed, and the multiplication of the offset and the input vector is replaced by a left shift operation. For a 4×4 block, the total number of multiplications required for each sample to generate the prediction signal is reduced to 4 times. Additionally, the encoding performance remains unchanged.
[0189] Figure 11 is a flowchart of an exemplary method 1100 for processing video content consistent with an embodiment of the present disclosure. Method 1100 can be performed by a codec (e.g., an encoder using Figures 2A - 2B for the encoding processes 200A and 200B or a decoder using Figures 3A - 3B for the decoding processes 300A and 300B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding or transcoding a video sequence. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a decoded compressed video sequence (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence, which can be captured by a surveillance device (e.g., Figure 4 the video input device in
[0190] associated with the processor (e.g., processor 402) of the device). The video sequence can include multiple images. The device can perform method 1100 at the image level. For example, in method 1100, the device can process one image at a time. Or for another example, in method 1100, the device can process multiple images at a time. Method 1100 can include the following steps.
[0191] In step 1102, the classification of the target block can be determined. In some embodiments, the classification can include a first class (e.g., Class0), a second class (e.g., Class1), and a third class (e.g., Class2). For a given block, the classification of the given block can be determined based on the size of the given block. For example, the first class can be associated with a block of size 4×4, the second class can be associated with blocks of size 8×8, 4×N, or N×4, where N can be an integer between 8 and 64. For example, N is equal to 8, 16, 32, or 64. That is, the second class can include blocks of size 8×8, 4×8, 4×16, 4×32, 4×64, 8×4, 16×4, 32×4, or 64×4. The third class can be associated with the remaining blocks.
[0191] In some embodiments, in response to the size of the target block not being 4×4, 8×8, 4×N, or N×4, it can be determined that the target block belongs to the third class.
[0192] In step 1104, a matrix weighted intra prediction (MIP) signal can be generated based on classification. In some embodiments, a first intra prediction signal for a target block can be generated based on an input vector, a matrix, and the classification of the target block, and a bilinear interpolation can be performed on the target block using the first intra prediction signal to generate the MIP signal.
[0193] For example, to generate the input vector, the neighboring reconstructed samples of the target block can be averaged according to the classification of the target block. As described above, for the first type of block, every two neighboring reconstructed samples of the block can be averaged to generate a reduced boundary vector as the input vector. For example, the size of the first type of input vector can be 4×1, the second type is 8×1, and the third type is 7×1. And for the second or third type of block (e.g., having a size of M×N), every M / 4 neighboring reconstructed samples above the block and every N / 4 neighboring reconstructed samples to the left of the block can be averaged.
[0194] Different from the input vector, a matrix can be selected from a set of matrices (e.g., matrix sets S0, S1, or S2) according to the classification of the target block and the MIP mode index.
[0195] Then, the first intra prediction signal can be generated by performing a matrix-vector multiplication on the matrix and the input vector. In some embodiments, the first intra prediction signal is also associated with a first offset and a second offset. For example, as discussed in equation (12), the reduced prediction signal can be further offset by a first offset (e.g., oW) and a second offset (e.g., oS). In some embodiments, the first offset and the second offset can be determined based on the matrix index in the matrix set. For example, the first offset and the second offset can be determined with reference to Table 7 and Table 8 respectively.
[0196] The classification of the target block is also related to the size of the first intra prediction signal. For example, in response to the target block belonging to the first or second type, the size of the first intra prediction signal is determined to be 4×4; in response to the target block belonging to the third type, the size of the first intra prediction signal is determined to be 8×8.
[0197] In some embodiments, a non - transitory computer - readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) for performing the above - mentioned method. Common forms of non - transitory media include, for example, floppy disks, hard disks, solid - state drives, magnetic tapes, or any other magnetic data - storage media, CD - ROMs, any other optical data - storage media, any physical media with a pattern of holes, RAM, PROM, and EPROM, FLASH - EPROM, or any other flash memory, NVRAM, caches, registers, any other memory chips or cartridges, and their network versions. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0198] The embodiments can be further described using the following terms:
[0199] 1. A computer - implemented method for processing video content, comprising:
[0200] Determining a classification of a target block; and
[0201] Generating a matrix - weighted intra - prediction (MIP) signal based on the classification.
[0202] Wherein determining the classification of the target block includes:
[0203] In response to the size of the target block being 4×4, determining that the target block belongs to a first class; or
[0204] In response to the size of the target block being 8×8, 4×N, or N×4, where N is an integer between 8 and 64, determining that the target block belongs to a second class.
[0205] 2. The method according to clause 1, wherein generating the MIP signal includes:
[0206] Generating a first intra - prediction signal for the target block, the generation of the first intra - prediction signal being based on an input vector, a matrix, and the classification of the target block; and
[0207] Using the first intra - prediction signal to perform bilinear interpolation on the target block to generate the MIP signal.
[0208] 3. The method according to clause 1 or 2, further comprising:
[0209] Averaging adjacent reconstructed samples of the target block according to the classification of the target block to generate the input vector.
[0210] 4. The method according to clause 2 or 3, wherein when the target block belongs to the first category, the size of the input vector is 4×1, or when the target block belongs to the second category, the size of the input vector is 8×1.
[0211] 5. The method according to any one of clauses 2-4, wherein the matrix is selected from a set of matrices according to the classification of the target block and the MIP mode index.
[0212] 6. The method according to clause 5, wherein the first intra prediction signal is generated by performing a matrix-vector multiplication on the matrix and the input vector.
[0213] 7. The method according to clause 6, wherein the first intra prediction signal is generated based on one or more offsets associated with the matrix.
[0214] 8. The method according to clause 7, wherein the one or more offsets are determined based on the index of the matrix in a lookup table.
[0215] 9. The method according to any one of clauses 1-8, wherein determining the classification of the target block further includes:
[0216] In response to the size of the target block not being 4×4, 8×N, 4×N, or N×4, determining that the target block belongs to the third category.
[0217] 10. The method according to clause 9, wherein generating the first intra prediction signal for the target block includes:
[0218] In response to the target block belonging to the first or second category, determining that the size of the first intra prediction signal is 4×4; and
[0219] In response to the target block belonging to the third category, determining that the size of the first intra prediction signal is 8×8.
[0220] 11. The method according to any one of clauses 1-10, wherein N is equal to 8, 16, 32, or 64.
[0221] 12. A video content processing system, comprising:
[0222] A memory for storing a set of instructions; and
[0223] At least one processor configured to execute a set of instructions to cause the system to perform:
[0224] Determine the classification of a target block; and
[0225] Generate a matrix weighted intra prediction (MIP) signal based on the classification.
[0226] Wherein, when determining the classification of the target block, the at least one processor is further configured to execute the instruction set to cause the system to further perform:
[0227] In response to the size of the target block being 4×4, determining that the target block belongs to the first category; or
[0228] In response to the size of the target block being 8×8, 4×N or N×4, where N is greater than 4, determining that the target block belongs to the second category.
[0229] 13. The system according to clause 12, wherein when generating the MIP signal, the at least one processor is further configured to execute the instruction set to cause the system to further perform:
[0230] Generating a first intra prediction signal for the target block, the generation of the first intra prediction signal being based on an input vector, a matrix, and the classification of the target block; and
[0231] Using the first intra prediction signal to perform bilinear interpolation on the target block to generate the MIP signal.
[0232] 14. The system according to clause 12 or 13, wherein the at least one processor is further configured to execute the instruction set to cause the system to further perform:
[0233] Averaging the adjacent reconstructed samples of the target block according to the classification of the target block to generate the input vector.
[0234] 15. The system according to clause 13 or 14, wherein when the target block belongs to the first category, the size of the input vector is 4×1, or when the target block belongs to the second category, the size of the input vector is 8×1.
[0235] 16. The system according to any one of clauses 13-15, wherein the matrix is selected from a set of matrices according to the classification of the target block and the MIP mode index.
[0236] 17. The system according to clause 16, wherein the first intra prediction signal is generated by performing matrix-vector multiplication on the matrix and the input vector.
[0237] 18. The system according to clause 17, wherein the first intra prediction signal is generated based on one or more offsets associated with the matrix.
[0238] 19. The system according to clause 18, wherein the one or more offsets are determined based on the index of the matrix in a lookup table.
[0239] 20. A non-transitory computer-readable medium having an instruction set stored thereon, the instruction set being executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising:
[0240] Determining a classification of a target block; and
[0241] Generating a matrix weighted intra prediction (MIP) signal based on the classification,
[0242] wherein determining the classification of the target block comprises:
[0243] In response to the size of the target block being 4×4, determining that the target block belongs to a first class; or
[0244] In response to the size of the target block being 8×8, 4×N or N×4, where N is an integer between 8 and 64, determining that the target block belongs to a second class.
[0245] It should be noted that the relational terms "first", "second", etc. in this article are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "comprising", "having", "including", and "include" and other similar forms are equivalent in meaning and are open-ended, because one or more items after any of these words are not intended to exhaustively list such items or items, or be limited to the listed items.
[0246] As used herein, unless otherwise specifically stated, the term "or" covers all possible combinations, unless infeasible. For example, if it is stated that a database may contain A or B, then, unless otherwise explicitly stated or infeasible, the database may contain A, or B, or A and B. As a second example, if it is stated that a database may contain A, B, or C, then, unless otherwise explicitly stated or infeasible, the database may contain A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0247] It can be understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. The software can execute the disclosed method when executed by a processor. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art can also understand that multiple of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0248] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary according to implementation. Certain modifications and alterations to the described embodiments can be made. Other embodiments will be apparent to those skilled in the art in view of the specification and practice of the invention disclosed herein. The specification and embodiments are considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims. The order of steps shown in the figures is also intended for illustrative purposes only and is not intended to be limited to any particular order of steps. Thus, those skilled in the art will understand that these steps can be performed in a different order while achieving the same method.
[0249] Exemplary embodiments have been disclosed in the accompanying drawings and the specification. However, many variations and modifications can be made to these embodiments. Accordingly, although specific terms are used, they are used in a general and descriptive sense only and not for purposes of limitation.
Claims
1. A computer-implemented method for processing video content, comprising: Determining a classification of a target block; And Generating a matrix weighted intra prediction (MIP) signal based on the classification, wherein generating the MIP signal includes: generating a first intra prediction signal for the target block, performing bilinear interpolation on the target block using the first intra prediction signal, and generating the MIP signal; Wherein determining the classification of the target block includes: In response to the size of the target block being 4×4, determining that the target block belongs to a first class; or In response to the size of the target block being 8×8, 4×8, 8×4, 4×N or N×4, where N is 16, 32 or 64, determining that the target block belongs to a second class.
2. The method according to claim 1, wherein generating the MIP signal includes: Generating the first intra prediction signal based on an input vector, a matrix, and the classification of the target block.
3. The method as claimed in claim 1, further comprising: Averaging adjacent reconstructed samples of the target block according to the classification of the target block to generate an input vector.
4. The method according to claim 2, wherein When the target block belongs to the first class, the size of the input vector is 4×1, and when the target block belongs to the second class, the size of the input vector is 8×1.
5. The method according to claim 2, wherein, Selecting the matrix from a set of matrices according to the classification of the target block and a MIP mode index.
6. The method according to claim 5, characterized in that, The first intra prediction signal is generated by performing matrix-vector multiplication on the matrix and the input vector.
7. The method according to claim 6, wherein generating the first intra prediction signal is based on one or more offsets associated with the matrix.
8. The method according to claim 7, wherein Determining the one or more offsets based on an index of the matrix in a lookup table.
9. The method according to claim 1, wherein determining the classification of the target block further includes: In response to the size of the target block not being 4×4, 8×8, 4×8, 8×4, 4×N and N×4, where N is 16, 32 or 64, determining that the target block belongs to a third class.
10. The method according to claim 1, wherein generating the first intra prediction signal for the target block includes: In response to the target block belonging to the first class or the second class, determining that the size of the first intra prediction signal is 4×4; And In response to the target block belonging to the third class, determining that the size of the first intra prediction signal is 8×8.
11. A video content processing system, comprising: A memory for storing an instruction set; And At least one processor configured to execute the instruction set to cause the system to perform: Determining a classification of a target block; And Generating a matrix weighted intra prediction (MIP) signal based on the classification, wherein generating the MIP signal includes: generating a first intra prediction signal for the target block, performing bilinear interpolation on the target block using the first intra prediction signal, and generating the MIP signal; Wherein, when determining the classification of the target block, the at least one processor is further configured to execute the instruction set to cause the system to further perform: In response to the size of the target block being 4×4, determining that the target block belongs to a first class; or In response to the size of the target block being 8×8, 4×8, 8×4, 4×N or N×4, where N is 16, 32 or 64, determine that the target block belongs to the second category.
12. The system according to claim 11, wherein when generating the MIP signal, the generation of the first intra-prediction signal in the first frame is based on an input vector, a matrix, and the classification of the target block.
13. The system according to claim 11, wherein, The at least one processor is further configured to execute the instruction set to cause the system to further perform: Average adjacent reconstructed samples of the target block according to the classification of the target block to generate an input vector.
14. The system according to claim 12, wherein When the target block belongs to the first category, the size of the input vector is 4×1, or when the target block belongs to the second category, the size of the input vector is 8×1.
15. The system according to claim 11, wherein, Select the matrix from a set of matrices according to the classification of the target block and the MIP mode index.
16. The system according to claim 15, wherein the first intra-prediction signal is generated by performing matrix-vector multiplication on the matrix and the input vector.
17. The system according to claim 16, wherein Generate the first intra-prediction signal based on one or more offsets associated with the matrix.
18. The system according to claim 17, wherein the one or more offsets are determined based on the index of the matrix in a look-up table.
19. A non-transitory computer-readable medium storing an instruction set executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising: Determine the classification of a target block; and Generate a matrix-weighted intra-prediction (MIP) signal based on the classification, wherein generating the MIP signal includes: generating a first intra-prediction signal for the target block, performing bilinear interpolation on the target block using the first intra-prediction signal, and generating the MIP signal; wherein determining the classification of the target block includes: In response to the size of the target block being 4×4, determine that the target block belongs to the first category; or In response to the size of the target block being 8×8, 4×8, 8×4, 4×N or N×4, N being 16, 32 or 64, determine that the target block belongs to the second category.