Method and system for performing combined inter- and intra-prediction - Patents.com

JP2024534322A5Pending Publication Date: 2025-09-25ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024513949
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-13
Filing Date
2022-09-22
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in achieving high compression efficiency, particularly with the development of advanced standards like VVC, where accurate prediction methods are needed to improve encoding and decoding processes.

Method used

The implementation of combined inter-prediction and intra-prediction (CIIP) along with luma mapping and chroma scaling (LMCS) techniques, including methods like OBMC and CIIP_PDPC, to enhance prediction accuracy and efficiency in video encoding and decoding.

Benefits of technology

Improves the accuracy and efficiency of video encoding and decoding processes, particularly in handling blocks with different motion vectors, leading to enhanced compression performance and reduced bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for video processing in which combined inter-prediction and intra-prediction (CIIP) and luma mapping and chroma scaling (LMCS) are applied, the method includes: obtaining an inter prediction signal, an intra prediction signal and an overlapped block motion compensation (OBMC) prediction signal, obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal, and obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal, where both the intermediate weighted prediction signal and the second prediction signal are in a mapping domain or an original domain.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 247,078, filed September 22, 2021, and U.S. Provisional Patent Application No. 17 / 931,676, filed September 13, 2022, which are incorporated by reference in their entireties into this specification.

[0002] Technical Field

[0002] The present disclosure relates generally to video processing, and more specifically to methods and systems for performing combined inter-prediction and intra-prediction. [Background technology]

[0003] background

[0003] A video is a set of still pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transformation, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, are developed by standardization organizations that specify specific video coding formats. With the adoption of increasingly advanced video coding techniques in video standards, the coding efficiency of new video coding standards is becoming higher and higher. Summary of the Invention

[0004] Disclosure Summary

[0004] An embodiment of the present disclosure provides a method for video processing, in which combined inter prediction and intra prediction (CIIP) and luma mapping and chroma scaling (LMCS) are applied. The method includes: obtaining an inter prediction signal, an intra prediction signal and an overlapped block motion compensation (OBMC) prediction signal; obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; and obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal, where both the intermediate weighted prediction signal and the second prediction signal are in a mapping domain or an original domain.

[0005]

[0005] An embodiment of the present disclosure provides an apparatus for performing video data processing. Combined inter-prediction and intra-prediction (CIIP) and luma mapping and chroma scaling (LMCS) are applied, and the apparatus includes a memory illustrated for storing instructions, and one or more processors, the one or more processors are configured to execute instructions to make the apparatus execute the following: obtain an inter-prediction signal, an intra-prediction signal and an overlapped block motion compensation (OBMC) prediction signal; obtain an intermediate weighted prediction signal by weighting the inter-prediction signal and a first prediction signal among the intra-prediction signal and the OBMC prediction signal; and obtain a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra-prediction signal and the OBMC prediction signal, where both the intermediate weighted prediction signal and the second prediction signal are in a mapping domain or an original domain.

[0006]

[0006] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing a bitstream of an image for processing according to a method, wherein combined inter and intra prediction (CIIP) and luma mapping and chroma scaling (LMCS) are applied to the image, and the method includes obtaining an inter prediction signal in an original domain, an intra prediction signal in a mapping domain and an overlapped block motion compensation (OBMC) prediction signal in the original domain, obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and the OBMC prediction signal in the original domain, transforming the intermediate weighted prediction signal from the original domain to the mapping domain, and obtaining a final prediction signal by weighting the transformed intermediate weighted prediction signal and the intra prediction signal in the mapping domain.

[0007] BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which various features are not necessarily drawn to scale. [Brief description of the drawings]

[0008] [Figure 1]

[0008] FIG. 1 is a schematic diagram illustrating a structure of an exemplary video sequence according to some embodiments of the present disclosure. [Figure 2A]

[0009] 1 is a schematic diagram illustrating an example encoding process for a hybrid video encoding system consistent with embodiments of the present disclosure. [Figure 2B]

[0010] FIG. 2 is a schematic diagram illustrating another example encoding process for a hybrid video encoding system, consistent with embodiments of the present disclosure. [Figure 3A]

[0011] 1 is a schematic diagram illustrating an example decoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3B]

[0012] FIG. 2 is a schematic diagram illustrating another example decoding process for a hybrid video coding system, consistent with embodiments of the present disclosure. [Figure 4]

[0013] 1 is a block diagram of an example apparatus for encoding or decoding video in accordance with some embodiments of the present disclosure. [Diagram 5]

[0014] 1 illustrates angular intra prediction modes in VVC, according to some embodiments of this disclosure. [Figure 6]

[0015] 1 illustrates example neighboring blocks used in deriving a general Most Probable Mode (MPM) list, according to some embodiments of the present disclosure. [Figure 7]

[0016] 1 illustrates example pixels used to calculate gradients in decoder-side intra-mode derivation (DIMD), in accordance with some embodiments of the present disclosure. [Figure 8]

[0017] 1 illustrates a predictive blending process for DIMD, according to some embodiments of the present disclosure. [Figure 9]

[0018] 1 illustrates an example template and its reference sample used during template-based intra-mode derivation (TIMD) in accordance with some embodiments of the present disclosure. [Figure 10]

[0019] 1 illustrates top and left neighboring blocks used in combined inter and intra prediction (CIIP) weight derivation according to some embodiments of this disclosure. [Figure 11]

[0020] 1 illustrates an example flowchart of an enhanced CIIP mode using position-dependent intra-prediction combining (PDPC), in accordance with some embodiments of the present disclosure. [Figure 12A]

[0021] 1 illustrates an example sub-block with overlapped block motion compensation (OBMC) applied, according to some embodiments of the present disclosure. [Figure 12B]1 illustrates an example sub-block with overlapped block motion compensation (OBMC) applied, according to some embodiments of the present disclosure. [Figure 13]

[0022] 1 illustrates an exemplary luma mapping and chroma scaling (LMCS) architecture including luma and chroma components from a decoder perspective, in accordance with some embodiments of the present disclosure. [Figure 14]

[0023] 1 illustrates an exemplary flowchart of a method for generating intra predictors in a CIIP according to some embodiments of the present disclosure. [Figure 15A]

[0024] 1 illustrates a flowchart of an example method for obtaining a final predicted signal in a mapping domain, according to some embodiments of the present disclosure. [Figure 15B]

[0025] 15B shows a table illustrating a method for obtaining a final predicted signal in the mapping domain shown in FIG. 15A according to some embodiments of the present disclosure. [Figure 16A]

[0026] 13 shows a flowchart of another example method for obtaining a final predicted signal in a mapping domain, according to some embodiments of the present disclosure. [Figure 16B]

[0027] 15C shows a table illustrating a modification of a method for obtaining a final predicted signal in the mapping domain shown in FIG. 15B according to some embodiments of the present disclosure. [Figure 17A]

[0028] 13 shows a flowchart of another example method for obtaining a final predicted signal in a mapping domain, according to some embodiments of the present disclosure. [Figure 17B]

[0029] 15C shows a table illustrating a modification of a method for obtaining a final predicted signal in the mapping domain shown in FIG. 15B according to some embodiments of the present disclosure. [Figure 18A]

[0030] 13 shows a flowchart of another example method for obtaining a final predicted signal in a mapping domain, according to some embodiments of the present disclosure. [Figure 18B]

[0031] 15C shows a table illustrating a modification of a method for obtaining a final predicted signal in the mapping domain shown in FIG. 15B according to some embodiments of the present disclosure. [Figure 19A]

[0032] 13 shows a flowchart of another example method for obtaining a final predicted signal in a mapping domain, according to some embodiments of the present disclosure. [Figure 19B]

[0033] 15C shows a table illustrating a modification of a method for obtaining a final predicted signal in the mapping domain shown in FIG. 15B according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Detailed Description

[0034] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations described in the following description of exemplary embodiments do not represent all implementations consistent with the present invention. Instead, they are merely examples of apparatuses and methods consistent with aspects related to the present invention as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall control.

[0010]

[0035] The ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while halving the bandwidth.

[0011]

[0036] To achieve the same subjective quality as HEVC / H.265 with half the bandwidth, JVET is developing a technique that goes beyond HEVC using the Joint Search Model (JEM) reference software. With the coding technique incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.

[0012]

[0037] The VVC standard is a recent development and continues to incorporate more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0013]

[0038] A video is a set of still pictures (or "frames") arranged in a time sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store those pictures in a time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with a display capability) can be used to display such pictures in a time sequence. In some applications, the video capture device can also transmit the captured video to a video playback device (e.g., a computer with a monitor) in real time, such as for surveillance, conferencing, or live broadcasting.

[0014]

[0039] To reduce the storage space and transmission bandwidth required by such applications, video may be compressed before storage and transmission, and decompressed before display. Compression and decompression may be performed by software executed by a processor (e.g., a processor of a general-purpose computer) or specialized hardware. The module for compression is commonly called an "encoder" and the module for decompression is commonly called a "decoder." The encoders and decoders may collectively be referred to as a "codec." The encoders and decoders may be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoders and decoders may include program code, computer executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed on a computer-readable medium. Video compression and decompression may be performed by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, or the like. In some applications, a codec can recover video from a first encoding standard and recompress the recovered video using a second encoding standard, in which case the codec may be called a "transcoder."

[0015]

[0040] A video coding process can identify and keep useful information that can be used to reconstruct a picture and ignore information that is not important for the reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such a coding process may be called "lossy". Otherwise, it may be called "lossless". Most coding processes are lossy, which is a tradeoff to reduce the required storage space and transmission bandwidth.

[0016]

[0041] Useful information of the picture being coded (called the "current picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes may include changes in pixel position, changes in luminance, or changes in color, among which position changes are of most importance. Changes in position of a group of pixels representing an object can reflect the motion of the object between the reference picture and the current picture.

[0017]

[0042] A picture that is coded without reference to another picture (i.e., the picture is its own reference picture) is called an "I-picture". If some or all of the blocks of a picture (e.g., blocks that are generally considered to be part of a video picture) are predicted using intra- or inter-prediction with one reference picture (e.g., uni-prediction), the picture is called a "P-picture". If at least one block of a picture is predicted with two reference pictures (e.g., bi-prediction), the picture is called a "B-picture".

[0018]

[0043] 1 illustrates the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. Video sequence 100 may be live video or captured and archived video. Video 100 may be live action video, computer generated video (e.g., computer game video), or a combination thereof (e.g., live action video with augmented reality effects). Video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored in a storage device), or a video supply interface for receiving video from a video content provider (e.g., a video broadcast transceiver).

[0019]

[0044] As shown in FIG. 1, a video sequence 100 may include a series of pictures arranged in time along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with additional pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture whose reference picture is picture 102 itself. Picture 104 is a P-picture whose reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture whose reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately preceding or following that picture. For example, the reference picture of picture 104 may be a picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the implementation of the reference pictures to the exact example shown in FIG.

[0020]

[0045] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such a task. Rather, video codecs may divide a picture into multiple elementary segments and encode or decode the picture segment by segment. Such elementary segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 110 of FIG. 1 illustrates an example structure of a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing units are also referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) and as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing unit can have variable size pictures, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, etc., or can have pixels of any shape and size. The size and shape of a basic processing unit can be selected based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit for a picture.

[0021]

[0046] A basic processing unit may be a logical unit and may include a group of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb, Cr) representing color information, and related syntax elements, and the luma and chroma components may have the same size basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" ("CTBs") in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be repeatedly performed on each of its luma and chroma components.

[0022]

[0047] Video coding has multiple operation stages, examples of which are shown in Figures 2A, 2B and 3A, 3B. At each stage, the size of the basic processing unit may still be too large to process, and therefore may be further divided into multiple segments, referred to as "basic processing sub-units" in this disclosure. In some embodiments, the basic processing sub-unit is also called a "block" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), and a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-unit may have the same or smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing sub-unit may also be a logical unit and may include a group of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit can be performed repeatedly on each of its luma and chroma components. Note that such division can be performed to further levels depending on the processing needs. Note also that at different stages, different schemes can be used to divide the basic processing unit.

[0023]

[0048] For example, in the mode decision stage (an example of which is shown in FIG. 2B), the encoder can decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can split the basic processing unit into multiple basic processing sub-units (e.g., CUs, as in H.265 / HEVC or H.266 / VVC) and decide a prediction type for each individual basic processing sub-unit.

[0024]

[0049] As another example, in the prediction stage (examples of which are shown in FIG. 2A and FIG. 2B), the encoder can perform prediction operations at the level of a basic processing sub-unit (e.g., CU). However, in some cases, the basic processing sub-unit may still be too large to process. The encoder can further divide the basic processing sub-unit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.

[0025]

[0050] As another example, in the transform stage (examples of which are shown in FIG. 2A and FIG. 2B), the encoder can perform transform operations on residual elementary processing sub-units (e.g., CUs). However, in some cases, the elementary processing sub-units may still be too large to process. The encoder can further divide the elementary processing sub-units into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform transform operations at that level. Note that the division scheme of the same elementary processing sub-unit may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0026]

[0051] 1, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, the boundaries of which are shown by dotted lines. Different basic processing units of the same picture may be divided into multiple basic processing sub-units in different schemes.

[0027]

[0052] In some implementations, to provide parallel processing and error resilience capabilities for video encoding and decoding, a picture can be divided into multiple regions for processing, so that the encoding or decoding process of one region of a picture does not have to depend on information from any other region of the picture. In other words, each region of a picture can be processed independently. By doing so, the codec can process different regions of a picture in parallel, thus improving the coding efficiency. Also, when data of a region is corrupted during processing or lost during network transmission, the codec can also correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience capabilities. In some video coding standards, a picture can be divided into multiple regions of different types. For example, H.265 / HEVC and H.266 / VVC provide two region types: "slice" and "tile". It should also be noted that different pictures of the video sequence 100 may have different partitioning schemes for dividing the picture into multiple regions.

[0028]

[0053] For example, in Figure 1, structure 110 is divided into three regions 114, 116, 118, whose boundaries are shown by solid lines inside structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.

[0029]

[0054] FIG. 2A illustrates a schematic diagram of an exemplary encoding process 200A consistent with embodiments of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As illustrated in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the video sequence 100 of FIG. 1, the video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in a temporal order. Similar to the structure 110 of FIG. 1, each original picture of the video sequence 202 may be divided by the encoder into multiple basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform the process 200A at the level of the basic processing unit for each original picture of the video sequence 202. For example, the encoder may perform the process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in one iteration of the process 200A. In some embodiments, the encoder may perform process 200A in parallel for a region of each original picture of video sequence 202 (eg, regions 114-118).

[0030]

[0055] In Figure 2A, the encoder may provide a basic processing unit of an original picture (referred to as an "original BPU") of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to generate a video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, 228 may be referred to as a "forward path." During process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to a prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of process 200A. The components 218, 220, 222, 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.

[0031]

[0056] The encoder may iteratively perform the process 200A to encode each original BPU of the original picture (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all the original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.

[0032]

[0057] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any act in any manner for inputting data.

[0033]

[0058] In the prediction stage 204, in the current iteration, the encoder may receive the original BPU and a prediction reference 224, perform a prediction operation, and generate predicted data 206 and a predicted BPU 208. The prediction reference 224 may be generated from a reconstruction path in a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting the predicted data 206, which can be used to reconstruct the original BPU from the predicted data 206 and the prediction reference 224 as a predicted BPU 208.

[0034]

[0059] Ideally, the predicted BPU 208 may be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating the predicted BPU 208, the encoder may subtract the predicted BPU 208 from the original BPU to generate the residual BPU 210. For example, the encoder may subtract values ​​(e.g., grayscale or RGB values) of pixels of the predicted BPU 208 from values ​​of corresponding pixels of the original BPU. Each pixel of the residual BPU 210 may have a residual value as a result of such subtraction between the pixels of the predicted BPU 208 and the corresponding pixels of the original BPU. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, which can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.

[0035]

[0060] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns", where each basis pattern is associated with a "transform coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variation frequency (e.g., frequency of luminance variation) component of the residual BPU 210. No basis pattern can be reconstructed from a combination (e.g., a linear combination) of other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, where the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0036]

[0061] Different transform algorithms may use different basis patterns. Various transform algorithms may be used in transform stage 212, such as, for example, discrete cosine transform, discrete sine transform, or the like. The transform in transform stage 212 is invertible. That is, the encoder may recover the residual BPU 210 by inverting the transform (called an "inverse transform"). For example, to recover a pixel of the residual BPU 210, the inverse transform may multiply the value of the corresponding pixel of the basis pattern by each associated coefficient and add the products to generate a weighted sum. For a video coding standard, both the encoder and the decoder may use the same transform algorithm (and therefore the same basis pattern). Thus, the encoder may record only the transform coefficients from which the decoder may reconstruct the residual BPU 210 without receiving the basis pattern from the encoder. Compared to the residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0037]

[0062] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns may represent different variation frequencies (e.g., luminance variation frequencies). Because the human eye is generally good at recognizing low-frequency variations, the encoder can ignore information of high-frequency variations without significant quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization scale factor") and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of the high-frequency basis pattern may be converted to zero, and the transform coefficients of the low-frequency basis pattern may be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, which further compresses the transform coefficients. The quantization process can also be inverted, in which case the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (called "dequantization").

[0038]

[0063] The quantization stage 214 may be lossy because the encoder ignores such division remainders in rounding operations. Typically, the quantization stage 214 may result in the greatest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain different levels of information loss, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0039]

[0064] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, besides the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, a type of transformation in the transformation stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0040]

[0065] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0041]

[0066] It should be noted that other variations of the process 200A may also be used to encode the video sequence 202. In some embodiments, the stages of the process 200A may be performed by the encoder in a different order. In some embodiments, one or more stages of the process 200A may be combined into a single stage. In some embodiments, a single stage of the process 200A may be split into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, the process 200A may include additional stages. In some embodiments, the process 200A may omit one or more stages in FIG. 2A.

[0042]

[0067] 2B shows a schematic diagram of another exemplary encoding process 200B consistent with an embodiment of the present disclosure. The process 200B may be a modification of the process 200A. For example, the process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared to the process 200A, the forward path of the process 200B additionally includes a mode decision stage 230 and splits the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of the process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0043]

[0068] In general, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use pixels from one or more already coded neighboring BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use regions from one or more already coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include already coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0044]

[0069] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. For an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs that are encoded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the encoder may perform extrapolation at a pixel level, such as by extrapolating, for each pixel of the predicted BPU 208, the value of the corresponding pixel. The neighboring BPUs used for extrapolation can be positioned from various directions relative to the original BPU, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction defined in the video coding standard used. In the case of intra prediction, the prediction data 206 may include, for example, the locations (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientations of the neighboring BPUs used relative to the original BPU, or the like.

[0045]

[0070] As another example, in the temporal prediction stage 2044, the encoder may perform inter prediction. For the original BPU of the current picture, the prediction reference 224 may include one or more pictures (called "reference pictures") that are coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. Once all the reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a certain range (called a "search window") of the reference picture. The location of the search window in the reference picture may be determined based on the location of the original BPU in the current picture. For example, the search window may be centered on a location in the reference picture that has the same coordinates as the coordinates of the original BPU in the current picture and extend to a predefined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., by using a pixel recursion algorithm, a block matching algorithm, or the like), the encoder can determine such a region as a matching region. The matching region may have different dimensions (e.g., smaller than, equal to, larger than, or a different shape) than the original BPU. Because the reference picture and the current picture are temporally separated in the timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" toward the location of the original BPU as time progresses. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., as in picture 106 in FIG. 1), the encoder can search for the matching region for each reference picture and determine its associated motion vector.In some embodiments, the encoder may assign weights to pixel values ​​of the matching regions of each matching reference picture.

[0046]

[0071] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, zooming, or the like. In the case of inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, a number of reference pictures, weights associated with the reference pictures, or the like.

[0047]

[0072] To generate the predicted BPU 208, the encoder can perform a "motion compensation" operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder can move the matching regions of the reference picture according to the motion vectors, so that the encoder can predict the original BPU of the current picture. If multiple reference pictures are used (e.g., like picture 106 in FIG. 1), the encoder can move the matching regions of the reference pictures according to their respective motion vectors and average the pixel values ​​of the matching regions. In some embodiments, if the encoder assigns weights to the pixel values ​​of the matching regions of each matching reference picture, the encoder can add a weighted sum of the pixel values ​​of the moved matching regions.

[0048]

[0073] In some embodiments, inter prediction can be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures that are in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter predicted picture in which a reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures that are in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter predicted picture in which reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.

[0049]

[0074] Still referring to the forward path of the process 200B, after the spatial prediction 2042 and temporal prediction stages 2044, in a mode decision stage 230, the encoder can select a prediction mode (e.g., intra-prediction or inter-prediction) for the current iteration of the process 200B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode to minimize the value of a cost function depending on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate a corresponding prediction BPU 208 and prediction data 206.

[0050]

[0075] If an intra prediction mode is selected in the forward path, in the reconstruction path of the process 200B, the encoder may generate a prediction reference 224 (e.g., a current BPU encoded and reconstructed in a current picture) and then directly provide the prediction reference 224 to a spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). The encoder may provide the prediction reference 224 to a loop filter stage 232, where the encoder may apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced during the encoding of the prediction reference 224. The encoder may apply various loop filter techniques in the loop filter stage 232, such as, for example, deblocking, sample adaptive offset, adaptive loop filter, or the like. The loop filtered reference picture may be stored in a buffer 234 (or a “decoded picture buffer”) for later use (e.g., for use as an inter prediction reference picture for a future picture of the video sequence 202). The encoder may store one or more reference pictures in a buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) in the binary encoding stage 226 along with the quantized transform coefficients 216, the prediction data 206, and other information.

[0051]

[0076] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A consistent with an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may decode video bitstream 228 into video stream 304 according to process 300A. Video stream 304 may be similar to video sequence 202. However, due to information loss in the compression and decompression process (e.g., quantization stage 214 in FIGS. 2A and 2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A and 2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each of the coded pictures in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions (e.g., regions 114-118) of each encoded picture in video bitstream 228.

[0052]

[0077] In FIG. 3A, the decoder may provide a portion of the video bitstream 228 associated with a basic processing unit of a coded picture (referred to as a "coded BPU") to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in a computer memory). The decoder may provide the prediction reference 224 to a prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.

[0053]

[0078] The decoder may iteratively perform the process 300A to decode each of the coded BPUs of the coded picture and generate a prediction reference 224 for coding the next coded BPU of the coded picture. After decoding all the coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.

[0054]

[0079] In the binary decoding stage 302, the decoder may perform the inverse operation of the binary encoding technique used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context-adaptive binary arithmetic encoding, or any other lossless compression algorithm). In some embodiments, besides the prediction data 206 and the quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, a type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, if the video bitstream 228 is transmitted in the form of packets over the network, the decoder may depacketize the video bitstream 228 before providing it to the binary decoding stage 302.

[0055]

[0080] 3B shows a schematic diagram of another exemplary decoding process 300B consistent with an embodiment of the present disclosure. The process 300B may be a modification of the process 300A. For example, the process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared to the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0056]

[0081] In process 300B, for a coded elementary processing unit (referred to as a "current BPU") of a coded picture being decoded (referred to as a "current picture"), prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, or the like. The parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, orientation of the neighboring BPUs relative to the original BPU, or the like. As another example, if inter prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, or the like. Parameters of the inter prediction operation may include, for example, the number of reference pictures associated with the current BPU, weights respectively associated with the reference pictures, locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors respectively associated with the matching regions, or the like.

[0057]

[0082] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details regarding performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a prediction BPU 208. The decoder may add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in FIG. 3A.

[0058]

[0083] In the process 300B, the decoder can provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform a prediction operation in the next iteration of the process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, the decoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture) after generating the prediction reference 224 (e.g., the decoded current BPU). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, the decoder can provide the prediction reference 224 to the loop filter stage 232 after generating the prediction reference 224 (e.g., the reference picture that all the BPUs are decoded) to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., for use as an inter-prediction reference picture for a future encoded picture of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data may further include loop filter parameters (e.g., loop filter strength). In some embodiments, the prediction data includes loop filter parameters if the prediction mode indicator of the prediction data 206 indicates that inter prediction was used to encode the current BPU.

[0059]

[0084] FIG. 4 is a block diagram of an example device 400 for encoding or decoding video consistent with an embodiment of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes instructions described herein, the device 400 may be a specialized machine for video encoding or decoding. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, the processor 402 may include any number and combination of a central processing unit (or "CPU"), a graphics processing unit (or "GPU"), a neural processing unit ("NPU"), a microcontroller unit ("MCU"), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property core (IP core), a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a system on a chip (SoC), an application specific integrated circuit (ASIC), or the like. In some embodiments, processor 402 may be a set of processors grouped together as a single logical component. For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0060]

[0085] The device 400 may also include a memory 404 configured to store data (e.g., an instruction set, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for performing steps in a process 200A, 200B, 300A, or 300B) and data for processing (e.g., the video sequence 202, the video bitstream 228, or the video stream 304). The processor 402 may access (e.g., via a bus 410) the program instructions and data for processing, execute the program instructions, and perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any number and combination of random access memory (RAM), read only memory (ROM), optical disks, magnetic disks, hard drives, solid state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, or the like. Memory 404 may also be a group of memories (not shown in FIG. 4) grouped together as a single logical component.

[0061]

[0086] Bus 410 may be a communication device that transfers data between components internal to apparatus 400, such as an internal bus (e.g., a CPU / memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.

[0062]

[0087] For ease of explanation and without creating ambiguity, the processor 402 and other data processing circuitry are collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, independent module or may be fully or partially combined with any other component of the device 400.

[0063]

[0088] The device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communications network, or the like). In some embodiments, the network interface 406 may include any number or combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a near field communication ("NFC") adapter, a cellular network chip, or the like.

[0064]

[0089] In some embodiments, the apparatus 400 may optionally further include a peripheral interface 408 for providing a connection to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera, or an input interface coupled to a video archive), or the like.

[0065]

[0090] It should be noted that a video codec (e.g., a codec that executes processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules within device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions that may be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as specialized data processing circuitry (e.g., FPGA, ASIC, NPU, or the like).

[0066]

[0091] In VVC, multiple intra prediction modes are provided. Figure 5 shows angular intra prediction modes in VVC according to some embodiments of the present disclosure. As shown in Figure 5, in order to capture any edge direction presented in natural video, the number of angular intra prediction modes in VVC is extended from 33 used in HEVC to 65, and the directional modes not found in HEVC are depicted with dotted arrows.

[0067]

[0092] The VVC standard implements two non-angular intra prediction modes: DC and Planar modes (similar to HEVC). In DC intra prediction mode, the average sample value of the reference samples for a block is used to generate the prediction. In VVC, only the reference samples along the long side of a rectangular block are used to calculate the average value, while for a square block, the reference samples from the left and top sides are used. In planar modes, the predicted sample value is obtained as a weighted average of four reference sample values: the reference sample in the same row or column as the current sample, and the reference samples in the bottom-left and top-right positions relative to the current block. The 65 angular modes and the two non-angular modes may be referred to as regular intra prediction modes.

[0068]

[0093] In some embodiments, a most probable mode (MPM) list is proposed. As discussed above, in VVC, there are 67 kinds of angle modes. If the prediction mode of each block is coded separately, 7 bits are required to code the 67 kinds of modes. Therefore, in VVC, a method is adopted to build the MPM list. In image and video coding, adjacent blocks are usually highly correlated, and therefore the probability that the intra-prediction modes of adjacent blocks are the same or similar is high. Therefore, the MPM list is built based on the intra-prediction modes of the left adjacent block and the upper adjacent block. In VVC, the length of the MPM list is 6. To keep the complexity of MPM list generation low, an intra-prediction mode coding method with 6 MPMs is used, which is derived from two available nearby intra-prediction modes.

[0069]

[0094] The unified 6MPM list, also called Primary MPM (PMPM) list, is used for intra blocks, regardless of whether MRL (Multiple Reference Lines) and ISP (Intra Subpartition) coding tools are applied. The MPM list is constructed based on the intra modes of the left and top neighboring blocks. Assuming that the intra mode of the left block is denoted as "Left" and the intra mode of the top block is denoted as "Above", the unified 6MPM list is constructed as follows: If no neighboring blocks are available, the intra prediction mode is set to "Planar" by default. If both the Left and Above modes are non-angular modes, the MPM list is set to {Planar, DC, V, H, V-4, V+4}, where V refers to the vertical mode and H refers to the horizontal mode. If one of the Left and Above modes is an angle mode and the other is a non-angle mode, then the Max mode is set as the larger of Left and Above, and the MPM list is set to {Planar, Max, Max-1, Max+1, Max-2, Max+2}. If both the Left and Above modes are angle modes and they are different, then the "Max" mode is set as the larger of Left and Above, and the "Min" mode is set as the smaller of Left and Above. If Max-Min is equal to 1, then the MPM list is set to {Planar, Left, Above, Min-1, Max+1, Min-2}, and if Max-Min is equal to or greater than 62, then the MPM list is set to {Planar, Left, Above, Min+1, Max-1, Min+2}. If Max-Min is equal to 2 then the MPM list is set to {Planar,Left,Above,Min+1,Min-1,Max+1} else the MPM list is set to {Planar,Left,Above,Min-1,Min+1,Max-1} If both Left and Above modes are angle modes and they are the same then the MPM list is set to {Planar,Left,Left-1,Left+1,Left-2,Left+2}.Moreover, the first bin of the MPM index codeword is CABAC (context-based adaptive binary arithmetic coding) context coded. In total, three contexts are used, which correspond to whether the current intra block is MRL-enabled, ISP-enabled, or a normal intra block. For entropy coding of the 61 non-MPM modes, TBC (truncated binary code) is used.

[0070]

[0095] In some embodiments, a secondary MPM method can be used. The primary MPM (PMPM) list consists of six entries, and the secondary MPM (SMPM) list contains sixteen entries. A general MPM list with 22 entries is first constructed, of which the first six entries are included in the PMPM list, and the remaining entries form the SMPM list. The first entry of the general MPM list is a planar mode. Then, the intra-prediction modes of the neighboring blocks are added to the list. FIG. 6 illustrates exemplary neighboring blocks used in deriving a general MPM list according to some embodiments of the present disclosure. As illustrated in FIG. 6, the intra-prediction modes of the left (L), top (A), bottom-left (BL), top-right (AR) and top-left (AL) neighboring blocks are used. If the CU block is vertically oriented, the order of the neighboring blocks is A, L, BL, AR, AL. Otherwise, i.e., if the CU block is horizontally oriented, the order of the neighboring blocks is L, A, BL, AR, AL. Then, the two decoder-side intra-prediction modes are added to the list. Then, an angle mode derived by adding an offset from the first two available angle modes in the list is added to the list. Finally, if the list is not complete, a default mode is added until the list is complete (i.e., has 22 entries). According to some embodiments of the present disclosure, the default mode list is defined as {DC, V, H, V-4, V+4, 14, 22, 42, 58, 10, 26, 38, 62, 6, 30, 34, 66, 2, 48, 52, 16}.

[0071]

[0096] For a decoder, the PMPM flag is parsed first. If the PMPM flag is equal to 1, the PMPM index is parsed to determine which entry in the PMPM list to select. Otherwise, the SMPM flag is parsed to determine if the SMPM index should be parsed for the remaining modes.

[0072]

[0097] In some embodiments, position-dependent intra-prediction combining (PDPC) is provided. In VVC, the result of intra-prediction is further modified by the PDPC method. PDPC is applied to the following intra-prediction modes without signaling: planar, DC, intra-angle modes below horizontal mode, and intra-angle modes above vertical mode. If the current block is in BDPCM (block-based delta pulse code modulation) mode or the MRL index is greater than 0, PDPC is not applied.

[0073]

[0098] The prediction sample pred(x',y') is predicted using an intra prediction mode (e.g., DC, planar or angular mode) and a linear combination of reference samples based on the following equation: pred(x',y')=Clip(0,(1<<BitDepth)-1,(wL×R-1,y’+wT×Rx’,-1+(64-wL-wT)×pred(x’,y’)+32)> >6) where Rx',-1 and R-1,y' respectively represent the reference samples located at the top and left boundaries of the current sample (x',y'). The PDPC weights and scale factors depend on the prediction mode and block size.

[0074]

[0099] Further, a decoder-side intra-prediction mode derivation (DIMD) method is provided. In the DIMD method, no luma intra-prediction mode is transmitted via the bitstream. Instead, a texture gradient process is performed to derive the two best modes. The same scheme is used at the encoder and decoder sides. The predictors of the derived two modes and the planar mode are computed as usual, and the weighted average of the three predictors is used as the final predictor for the current block.

[0075]

[0100] DIMD mode is used as an alternative intra-prediction mode, and for each block, a flag is signaled to indicate whether DIMD mode should be used or not. If the flag is true (e.g., the flag is equal to 1), the DIMD mode is used for the current block, and the BDPCM flag, MIP (Matrix Weighted Intra Prediction) flag, ISP flag, and MRL index are inferred to be 0. In this case, the entire intra-prediction mode analysis is also skipped. If the flag is false (e.g., the flag is equal to 0), the DIMD mode is not used for the current block, and the analysis of other intra-prediction modes continues as usual.

[0076]

[0101] To derive the two intra prediction modes and determine the weights for each mode, a histogram is created by performing texture gradient processing.

[0077]

[0102] 7 shows example samples used to calculate gradients in DIMD according to some embodiments of the present disclosure. As shown in FIG. 7, to create a DIMD histogram for a block, gradient analysis is performed on samples 710 of an L-shaped template of a second neighborhood line surrounding the block. For each available reconstructed sample of the template, horizontal gradients Gx and vertical gradients Gy are performed by applying horizontal and vertical Sobel filters as follows:

number

[0078]

[0103] For each sample of the template for which a horizontal gradient Gx and a vertical gradient Gy have been calculated, the gradient magnitude (G) and direction (O) are further calculated using Gx and Gy as follows:

number

[0079]

[0104] The orientation of the gradient O is converted to the closest intra-angle prediction mode and used to index a histogram that is initially initialized to zero. The histogram value for that intra-angle prediction mode is increased by G. After all samples of the template have been processed, the histogram may contain the accumulated values ​​of gradient magnitude for each intra-angle prediction mode. For the next prediction fusion process, two modes with the largest and second largest amplitude values ​​are selected and marked as M1 and M2, respectively. If the largest amplitude value of the histogram is 0, the planar mode is selected as the intra prediction mode for the current block.

[0080]

[0105] In DIMD, the two intra prediction angle modes corresponding to the two largest histogram amplitude values ​​M1 and M2 are combined with the planar mode to generate the final prediction value of the current block.

[0081]

[0106] Predictive blending is applied as a weighted average of the above three predictors. The weight of the planar mode is fixed at 21 / 64 (approximately equal to 1 / 3). The remaining weight of 43 / 64 (approximately equal to 2 / 3) is shared between M1 and M2 in proportion to their amplitude values. Figure 8 illustrates the predictive blending process of DIMD according to some embodiments of the present disclosure. As shown in Figure 8, ampl(M1) and ampl(M2) represent the amplitude values ​​of M1 and M2, respectively.

[0082]

[0107] DIMD mode is used only for luma blocks. If the current luma block selects DIMD mode, the intra prediction mode of the current block is stored as M1 for the selection of the Low Frequency Non-Separable Transform (LFNST) set of the current block, for the derivation of the Most Probable Mode (MPM) list of neighboring luma blocks, and for the derivation of the Direct Mode (DM) of co-located chroma blocks.

[0083]

[0108] Further, in some embodiments, another decoder-side intra-prediction mode derivation method (such as template-based intra-mode derivation (TIMD) using MPM) can be used. Instead of being signaled, the intra-prediction mode of a CU is derived using a template-based method at both the encoder and decoder sides. Candidates are built from the MPM list, and the candidate modes are 67 intra-prediction modes similar to VVC or extended to 131 intra-prediction modes. FIG. 9 shows an example template and reference samples used in TIMD according to some embodiments of the present disclosure. As shown in FIG. 9, the prediction sample of the template 910 is generated using the template reference sample 920 for each candidate mode. The value is calculated as the sum of absolute transform difference (SATD) between the template prediction sample and the reconstructed sample. The intra-prediction mode with the minimum value of SATD is selected as the TIMD mode and used for intra-prediction of the current CU.

[0084]

[0109] The TIMD mode is used as an additional intra prediction method for a CU. A flag to enable / disable TIMD is signaled in the sequence parameter set (SPS). If the flag is true (e.g., the flag is equal to 1), a CU-level flag is signaled to indicate whether TIMD is used. The TIMD flag is signaled after the MIP flag. If the TIMD flag is true (e.g., the TIMD flag is equal to 1), all remaining syntax elements related to the luma intra prediction mode are skipped, including the MRL, ISP, and normal parsing phase for the luma intra prediction mode.

[0085]

[0110] In TIMD, the number of intra prediction modes is extended to 131, so when storing the intra prediction mode for the current block, a table is used to map the 131 modes in VVC to the original 67 intra prediction modes.

[0086]

[0111] Combined inter and intra prediction (CIIP) is provided. In VVC, when a CU is coded in merge mode, if the CU contains at least 64 luma samples (i.e., the product of the CU width and CU height is 64 or more), and if the CU width and CU height are both less than 128 luma samples, an additional flag is signaled to indicate whether the CIIP mode applies to the current CU. CIIP prediction combines an inter predictor with an intra predictor. The inter predictor P of the CIIP mode inter is derived using the same inter prediction process as applied in the normal merge mode, and the intra predictor P intra is derived according to the normal intra prediction process in planar mode. Figure 10 shows the upper neighboring block and the left neighboring block used in CIIP weight derivation according to some embodiments of the present disclosure. The intra predictor and the inter predictor are combined using a weighted average, where the weight value is calculated according to the coding mode of the upper and left neighboring blocks (as shown in Figure 10).

[0087]

[0112] The weights (wIntra, wInter) for the intra and inter predictors are adaptively set as follows: If the upper and left neighboring blocks are both intra-coded, then (wIntra, wInter) are set equal to (3, 1); if one of these blocks is intra-coded, then the weights are the same, i.e., set equal to (2, 2); if neither the upper nor left neighboring blocks are intra-coded, then the weights are set equal to (1, 3). The CIIP predictor is formed based on the following: P CIIP =(wInter * P inter +wIntra * P intra +2)>>2

[0088]

[0113] For the chroma components, the DM mode is applied without any additional signaling.

[0089]

[0114] In some embodiments, multi-hypothesis prediction for intra and inter modes can be used. In merge CU, one flag is signaled for merge mode to select intra prediction mode from intra candidate list when the flag is true. For luma component, intra candidate list is derived from four intra prediction modes including DC, planar, horizontal and vertical modes. One intra prediction mode selected by intra prediction mode index and one inter prediction mode selected by merge index are combined using weighted average. Weights for combining predictions are described as follows: If DC or planar mode is selected or CU width or height is less than 4, equal weights are applied. For CU with CU width and height 4 or more, if horizontal / vertical mode is selected, one CU is first divided vertically or horizontally into four equal size regions. Each set of weights is denoted as (wIntrai, wInteri), where i is 1 to 4, and (wIntra1, wInter1) = (6, 2), (wIntra2, wInter2) = (5, 3), (wIntra3, wInter3) = (3, 5), and (wIntra4, wInter4) = (2, 6), and is applied to the corresponding region. (wIntra1, wInter1) is for the region closest to the reference sample, and (wIntra4, wInter4) is for the region farthest from the reference sample. The combined predictor can be calculated by summing the two weighted predictors and right shifting by a number of bits, which is given by the logarithm of the sum of the two weights. In this example, the combined predictor is obtained by summing the two weighted predictors and right shifting by 3 bits. In some embodiments, if the sum of the two weights is equal to 1, then the combined predictor may be obtained by directly summing the two weighted predictors since the logarithm of 1 is 0. No right shift is required.For example, each set of weights may be (wIntra1,wInter1)=(6 / 8,2 / 8), (wIntra2,wInter2)=(5 / 8,3 / 8), (wIntra3,wInter3)=(3 / 8,5 / 8) and (wIntra4,wInter4)=(2 / 8,6 / 8) and applied to a corresponding region, where (wIntra1,wInter1) is for the region closest to the reference sample and (wIntra4,wInter4) is for the region furthest from the reference sample.

[0090]

[0115] In some embodiments, a CIIP_PDPC mode can be used. In CIIP_PDPC, the prediction of the normal merge mode is improved using the reconstructed samples on the upper side (Rx, -1) and the left side (-1, Ry). This improvement inherits the position-dependent prediction combination (PDPC) scheme. Figure 11 shows an example flowchart of CIIP_PDPC according to some embodiments of the present disclosure. With reference to Figure 11, WT and WL are weighting values ​​that depend on the sample position in the block as defined in PDPC.

[0091]

[0116] The CIIP_PDPC mode is signaled along with the CIIP mode. If the CIIP flag is true, another flag (i.e., the CIIP_PDPC flag) is further signaled to indicate whether CIIP_PDPC should be used or not.

[0092]

[0117] In some embodiments, overlapped block motion compensation (OBMC) is proposed in H.263. When OBMC is used in the enhanced compression model (ECM), OBMC is performed for all motion compensation (MC) block boundaries except the right and bottom boundaries of the CU. Moreover, OBMC is applied to both luma and chroma components. In ECM, an MC block corresponds to a coding block. When a CU is coded in a sub-CU mode (such as SbTMVP (sub-block based temporal motion vector prediction) or affine mode), each sub-block of the CU is an MC block. To uniformly handle CU boundaries, OBMC is performed at the sub-block level for all MC block boundaries, and the sub-block size is set equal to 4×4. When OBMC is applied to a current sub-block, besides the current motion vector, the motion vectors of four connected neighboring sub-blocks (if available and not identical to the current motion vector) are also used to derive a prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate a final prediction signal for the current sub-block.

[0093]

[0118] The predicted block based on the motion vector of the neighboring subblock is P N where N denotes the index for the neighboring sub-blocks above, below, left and right. The predicted block based on the motion vector of the current sub-block is P C It is shown as P N If P is based on the motion information of a neighboring sub-block that contains the same motion information as the current sub-block, N OBMC will not be executed from P N All samples of P C are added to the same sample of P N P C Before being added to NNote that weighting factors can be applied to the upper boundary P. Figures 12A and 12B show example sub-blocks with OBMC applied, according to some embodiments of the present disclosure. For example, Figure 12A shows motion vectors used in OBMC for sub-blocks at the CU / PU boundary for the current CU. As shown in Figure 12A, N1 For a sub-block located at N1 Used in the OBMC of the left boundary P N2 For a sub-block in, the motion vector of the left neighboring sub-block 1202 is P N2 Used in the OBMC, upper left corner P N3 For a sub-block in, the motion vectors of the left neighboring sub-block 1203 and the upper neighboring sub-block 1204 are P N3 FIG. 12B shows the motion vectors used in OBMC for subblocks in ATMVP (alternative temporal motion vector prediction) for the current CU. In ATMVP, each subblock is an MC block. As shown in FIG. 12B, the motion vectors used in OBMC for subblocks in ATMVP (alternative temporal motion vector prediction) for the current CU. N In the case where the motion vectors of the four neighboring sub-blocks 1205, 1206, 1207, and 1208 are P N Used by the OBMC.

[0094]

[0119] In VVC, a coding tool called luma mapping and chroma scaling (LMCS) is added as a new processing block before the loop filter. LMCS has two main components: 1) for the luma component, in-loop mapping of the luma component based on an adaptive piecewise linear model, and 2) for the chroma component, luma-dependent chroma residual scaling is applied. Figure 13 shows an example LMCS architecture including luma and chroma components from the decoder's perspective according to some embodiments of the present disclosure. With reference to Figure 13, processing blocks 1302, 1304, 1306 show where processing is applied in the mapping domain, including inverse quantization and inverse transform 1302, luma intra prediction 1304, and reconstruction of the luma component by adding luma prediction and luma residual 1306. Processing blocks 1310-1318 indicate where processing is applied in the original (i.e., non-mapped) domain, including loop filter 1310 (e.g., deblocking, ALF, and SAO), motion compensated prediction 1312, chroma intra prediction 1314, reconstruction of chroma components by adding chroma prediction and chroma residual 1316, and storing the decoded picture as a reference picture in the decoded picture buffer (DPB) 1318. Processing blocks 1322, 1324, 1326 are new LMCS function blocks, including forward mapping of luma signal 1322, backward mapping of luma signal 1324, and luma-dependent chroma scaling process 1326. Like most other tools in VVC, LMCS can be enabled / disabled at the sequence level using the SPS flag.

[0095]

[0120] In the current design, the CIIP mode is only applicable to the merge mode. That is, the inter prediction of the CIIP mode can only be derived from the merge mode. However, the prediction of the merge mode may not be accurate, especially for blocks whose motion is different from that of their neighboring blocks. In this case, the performance of the CIIP mode may be degraded.

[0096]

[0121] To improve the prediction accuracy of the CIIP mode, this disclosure proposes a method for applying the CIIP mode to the normal inter mode.

[0097]

[0122] FIG. 14 illustrates an exemplary flowchart of a method 1400 for generating intra predictors in a CIIP according to some embodiments of the present disclosure. The method 1400 may be performed by a system such as an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B), a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B), or may be performed by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, one or more processors (e.g., processor 402 of FIG. 4) may perform the method 1400. In some embodiments, the method 1400 may be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions (e.g., program code) executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 14, the method 1400 may include the following steps 1402-1408.

[0098]

[0123] In step 1402, the system may determine whether to enable a CIIP mode for the target block. In some embodiments, a flag is signaled in the bitstream to indicate the mode used for the target block. For example, when it is determined that the block is coded in a normal inter mode and not in a merge mode, a first flag is signaled to indicate whether the block is coded in a CIIP mode. A block coded in a normal inter mode and a CIIP mode, but not in a merge mode, is called a CIIP_INTER mode. When the CIIP_INTER mode is applied to a block, some other inter prediction modes, such as a local illumination compensation mode, a bi-prediction mode with CU-level weights, an affine mode, a symmetric motion vector difference mode, an adaptive motion vector resolution mode, or a multi-hypothesis inter prediction mode, may be disabled.

[0099]

[0124] In some embodiments, the first flag is decoded at the end of the entire normal inter mode syntax structure. In this example, the first flag is signaled only if local illumination compensation mode, bi-prediction mode with CU level weights, affine mode, or multi-hypothesis inter prediction mode is disabled.

[0100]

[0125] In some embodiments, the first flag is decoded at the beginning of the entire normal inter mode syntax structure. If the first flag indicates that the CIIP_INTER mode is applied (e.g., the first flag is true), the local illumination compensation mode, the bi-prediction mode with CU-level weights, the affine mode, or the multi-hypothesis inter prediction mode are disabled, and the associated syntax of these modes is not signaled.

[0101]

[0126] In step 1404, a motion vector obtained using the motion vector predictor and the signaled motion vector differential is used to generate an inter predictor (also called an inter prediction signal) for the CIIP mode. Thus, the CIIP inter prediction for a block that has a different motion from its neighboring blocks can be more accurate.

[0102]

[0127] In some embodiments, the CIIP mode is combined with a merge mode using motion vector differentials (called MMVD). The inter prediction of the CIIP mode is predicted in MMVD mode (called CIIP_MMVD). In some embodiments, for a block coded in MMVD mode, in step 1402, a second flag is decoded indicating whether the block is coded in CIIP_MMVD mode. In some embodiments, for a block coded in CIIP mode, in step 1404, a third flag is decoded indicating whether the inter prediction is predicted using CIIP_MMVD mode.

[0103]

[0128] In step 1406, an intra predictor (also called an intra prediction signal) for CIIP is generated. In some embodiments, the intra prediction for CIIP mode is derived using the TIMD or DIMD method instead of the planar mode. In some embodiments, when CIIP_INTER or CIIP_MMVD mode is enabled, the intra prediction is derived using the TIMD or DIMD mode. That is, the TIMD or DIMD method is used to derive the intra prediction mode. Moreover, the intra prediction mode is propagated to neighboring blocks that are coded in CIIP mode.

[0104]

[0129] In step 1408, a final predictor (also called a final predicted signal) for the target block is obtained by weighting the inter predictor and the intra predictor.

[0105]

[0130] Therefore, the CIIP mode can be applied with various modes, and the inter prediction of CIIP for a block whose motion is different from its neighboring blocks can be more accurate, improving the performance of CIIP.

[0106]

[0131] In some embodiments, if a block is coded in CIIP and LMCS modes, the OBMC mode is always applied. FIG. 15A illustrates a flowchart of an example method 1500 for obtaining a final prediction signal in a mapping domain according to some embodiments of the present disclosure. The method 1500 may be performed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B), a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 1500. In some embodiments, the method 1500 may be implemented by a computer program product embodied in a computer-readable medium, the computer program product including computer-executable instructions (e.g., program code) executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 15A, the method 1500 may include the following steps 1502-1510.

[0107]

[0132] In step 1502, an inter prediction signal and an intra prediction signal are obtained. The inter prediction signal is obtained using the same inter prediction process as that applied to the normal merge mode. Referring to FIG. 14, the inter prediction signal and the intra prediction signal can be obtained by the same steps as the steps of obtaining inter predictors and intra predictions in steps 1402-1406. Similar to FIG. 13, the inter prediction signal is obtained in the original domain (e.g., block 1312) and the intra prediction signal is obtained in the mapping domain (e.g., block 1304).

[0108]

[0133] In step 1504, the inter-prediction signal is transformed from the original domain to the mapping domain, so that both the inter-prediction signal and the intra-prediction signal are in the same domain (i.e., the mapping domain).

[0109]

[0134] In step 1506, a weighted prediction signal is obtained by weighting the intra prediction signal and the inter prediction signal in the mapping domain.

[0110]

[0135] In step 1508, an OBMC prediction signal is obtained. The OBMC prediction signal is obtained using motion from neighboring blocks, and the OBMC prediction signal is in the original domain.

[0111]

[0136] In step 1510, a final prediction signal is obtained by weighting the weighted prediction signal and the OBMC prediction signal.

[0112]

[0137] Method 1500 may also be shown in Table 1 of Figure 15B. As shown in Figures 15A and 15B, in method 1500, in step 1506, a weighted prediction signal is obtained in the mapping domain, while OBMC is obtained in the original domain. The OBMC prediction signal is not transformed to the mapping domain before weighting. Thus, the final prediction signal is obtained by weighting the two signals in different domains.

[0113]

[0138] In order to improve the weighting process for blocks coded in CIIP, OBMC and LMCS modes, the present disclosure proposes a method for improving the weighting process, for example, the prediction signals used in the weighting process are transformed to the same domain, for example, both prediction signals are in the mapping domain, or both prediction signals are in the original domain.

[0114]

[0139] In some embodiments, two prediction signals among the inter prediction signal, the intra prediction signal and the OBMC prediction signal can be weighted to obtain an intermediate prediction signal. Then, the intermediate prediction signal can be further weighted with a third prediction signal to obtain a final prediction signal. The final prediction signal is in the mapping domain.

[0115]

[0140] For example, in the first weighting process, the inter prediction signal and the intra prediction signal are weighted first to obtain an intermediate prediction signal (e.g., a weighted prediction signal). Then, in the second weighting process, the intermediate prediction signal (e.g., a weighted prediction signal) and the OBMC prediction signal are weighted to obtain a final prediction signal. In another example, in the first weighting process, the inter prediction signal and the OBMC prediction signal are weighted first to obtain an intermediate prediction signal (e.g., an improved prediction signal). Then, in the second weighting process, the intermediate prediction signal (e.g., an improved prediction signal) and the intra prediction signal are weighted to obtain a final prediction signal. In another example, in the first weighting process, the intra prediction signal and the OBMC prediction signal are weighted first to obtain an intermediate prediction signal (e.g., an improved prediction signal). Then, in the second weighting process, the intermediate prediction signal (e.g., an improved prediction signal) and the inter prediction signal are weighted to obtain a final prediction signal. The predicted signal can be transformed between the original domain and the mapped domain such that the two weighted predicted signals in the first and second weighting processes are in the same domain.

[0116]

[0141] Further details are further explained below.

[0117]

[0142] In some embodiments, the OBMC prediction signal is transformed to the mapping domain before weighting. FIG. 16A shows a flowchart of another exemplary method 1600 for obtaining a final prediction signal in the mapping domain according to some embodiments of the present disclosure. The method 1600 can be performed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B), a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can perform the method 1600. In some embodiments, the method 1600 can be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions (such as program code) executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 16A, the method 1600 may include the following steps 1602-1610.

[0118]

[0143] In step 1602, an inter prediction signal and an intra prediction signal are obtained. The inter prediction signal is obtained using the same inter prediction process as that applied to the normal merge mode. For example, referring to FIG. 14, the inter prediction signal and the intra prediction signal can be obtained by the same steps as obtaining the inter predictor and intra prediction in steps 1402-1406. The inter prediction signal is obtained in the original domain, and the intra prediction signal is obtained in the mapping domain.

[0119]

[0144] In step 1604, the inter-prediction signal is transformed to be in the mapping domain with respect to the original domain, so that both the inter-prediction signal and the intra-prediction signal are in the same domain (i.e., the mapping domain).

[0120]

[0145] In step 1606, a weighted prediction signal is obtained by weighting the intra prediction signal and the inter prediction signal in the mapping domain.

[0121]

[0146] In step 1608, an OBMC prediction signal is obtained and transformed into the mapping domain. The OBMC prediction signal is obtained using motion from neighboring blocks in the original domain. After transformation, the OBMC prediction signal is also in the mapping domain.

[0122]

[0147] In step 1610, a final prediction signal is obtained by weighting the weighted prediction signal and the OBMC prediction signal in the mapping domain. By using the method 1600, the weighted prediction signal and the OBMC prediction signal used in the weighting process are both in the mapping domain. Because the weighted prediction signal and the OBMC prediction signal are in the same domain, the weighting process can be more efficient and accurate.

[0123]

[0148] Method 1600 may also be illustrated in Table 2 of FIG. 16B, where differences from Table 1 shown in FIG. 15B are indicated in italics, bold, and / or strikethrough.

[0124]

[0149] In some embodiments, instead of performing the weighting in the mapping domain, it is proposed to perform the weighting in the original domain. FIG. 17A shows a flowchart of another exemplary method 1700 for obtaining a final prediction signal in the mapping domain according to some embodiments of the present disclosure. The method 1700 can be performed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B), a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can perform the method 1700. In some embodiments, the method 1700 can be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions (such as program code) executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 17A, the method 1700 may include the following steps 1702-1710.

[0125]

[0150] In step 1702, an inter prediction signal and an intra prediction signal are obtained. The inter prediction signal is obtained using the same inter prediction process as that applied to the normal merge mode. For example, referring to FIG. 14, the inter prediction signal and the intra prediction signal can be obtained by the same steps as obtaining the inter predictor and intra prediction in steps 1402-1406. The inter prediction signal is obtained in the original domain, and the intra prediction signal is obtained in the mapping domain.

[0126]

[0151] In step 1704, the intra prediction signal is transformed from the mapping domain to the original domain. Thus, the intra prediction signal is in the original domain. Both the inter prediction signal and the intra prediction signal are in the same domain (i.e., in this example, the original domain).

[0127]

[0152] In step 1706, a weighted prediction signal is obtained by weighting the intra prediction signal and the inter prediction signal in the original domain. Since the inter prediction signal and the intra prediction signal are both in the original domain, the weighted prediction signal is also obtained in the original domain.

[0128]

[0153] In step 1708, an OBMC prediction signal is obtained. The OBMC prediction signal is obtained using motion from neighboring blocks, and the OBMC prediction signal is in the original domain.

[0129]

[0154] In step 1710, a final prediction signal is obtained by weighting the weighted prediction signal and the OBMC prediction signal in the original domain, and transformed to the mapping domain.

[0130]

[0155] In the method 1700, instead of converting the inter prediction signal from the original domain to the mapping domain, the intra prediction signal is converted from the mapping domain to the original domain to ensure that the inter prediction signal and the intra prediction signal are both in the same domain. Thus, the weighted prediction signal is obtained in the original domain, just like the OBMC prediction signal. Because the weighted prediction signal and the OBMC prediction signal are in the same domain, the weighting process can be more efficient and accurate. The final prediction signal is obtained in the original domain and then converted to the mapping domain.

[0131]

[0156] Method 1700 may also be shown in Table 3 of FIG. 17B, where differences from Table 1 shown in FIG. 15B are indicated in italics, bold, and / or strikethrough).

[0132]

[0157] In some embodiments, a method is proposed for transforming the weighted prediction signal from the mapping domain to the original domain before weighting. FIG. 18A shows a flowchart of another exemplary method 1800 for obtaining a final prediction signal in the mapping domain according to some embodiments of the present disclosure. The method 1800 can be performed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B), a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can perform the method 1800. In some embodiments, the method 1800 can be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions (such as program code) executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 18A, the method 1800 may include the following steps 1802-1810.

[0133]

[0158] In step 1802, an inter prediction signal and an intra prediction signal are obtained. The inter prediction signal is obtained using the same inter prediction process as that applied to the normal merge mode. For example, referring to FIG. 14, the inter prediction signal and the intra prediction signal can be obtained by the same steps as obtaining the inter predictor and intra prediction in steps 1402-1406. The inter prediction signal is obtained in the original domain, and the intra prediction signal is obtained in the mapping domain.

[0134]

[0159] In step 1804, the inter-prediction signal is transformed from the original domain to the mapping domain, so that both the inter-prediction signal and the intra-prediction signal are in the same domain (i.e., the mapping domain).

[0135]

[0160] In step 1806, a weighted prediction signal is obtained by weighting the intra prediction signal and the inter prediction signal in the mapping domain, and is transformed to the original domain.

[0136]

[0161] In step 1808, an OBMC prediction signal is obtained. The OBMC prediction signal is obtained using motion from neighboring blocks, and the OBMC prediction signal is in the original domain. Thus, both the weighted prediction signal and the OBMC prediction signal are in the same domain (i.e., the original domain).

[0137]

[0162] In step 1810, a final prediction signal is obtained by weighting the weighted prediction signal and the OBMC prediction signal in the original domain and transformed into the mapping domain. By using the method 1800, the weighted prediction signal and the OBMC prediction signal used in the weighting process are both in the original domain. After the final prediction signal is obtained in the original domain, the final prediction signal is transformed into the mapping domain. Because the weighted prediction signal and the OBMC prediction signal are in the same domain, the weighting process can be more efficient and accurate.

[0138]

[0163] Method 1800 may also be shown in Table 4 of FIG. 18B, with differences from Table 1 shown in FIG. 15B shown in italics, bold, and / or strikethrough.

[0139]

[0164] In some embodiments, a method is proposed for weighting an inter prediction signal with an OBMC prediction signal before an intra prediction signal. FIG. 19A shows a flowchart of another exemplary method 1900 for obtaining a final prediction signal in a mapping domain according to some embodiments of the present disclosure. The method 1900 can be performed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B), a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can perform the method 1900. In some embodiments, the method 1900 can be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions (such as program code) executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 19A, the method 1900 may include the following steps 1902-1908.

[0140]

[0165] In step 1902, an inter prediction signal and an intra prediction signal are obtained. The inter prediction signal is obtained using the same inter prediction process as that applied to the normal merge mode. For example, referring to FIG. 14, the inter prediction signal and the intra prediction signal can be obtained by the same steps as obtaining the inter predictor and intra prediction in steps 1402-1406. The inter prediction signal is obtained in the original domain, and the intra prediction signal is obtained in the mapping domain.

[0141]

[0166] In step 1904, an OBMC prediction signal is obtained. The OBMC prediction signal is obtained using motion from neighboring blocks, and the OBMC prediction signal is in the original domain.

[0142]

[0167] In step 1906, an improved inter prediction signal is obtained by weighting the inter prediction signal and the OBMC prediction signal in the original domain and transformed into the mapping domain. Thus, the improved inter prediction is first obtained in the original domain and then transformed into the mapping domain.

[0143]

[0168] In step 1908, a final prediction signal is obtained by weighting the improved inter prediction signal and the intra prediction signal in the mapping domain. Since the improved inter prediction signal has been transformed into the mapping domain in step 1906, both the improved inter prediction signal and the intra prediction signal are in the same domain (i.e., the mapping domain). Therefore, the weighting process can be more efficient and accurate.

[0144]

[0169] The method 1900 may also be illustrated in Table 5 of FIG. 19B, where differences from Table 1 shown in FIG. 15B are indicated in italics, bold, and / or strikethrough.

[0145]

[0170] In some embodiments, for a block coded in CIIP mode, the OBMC mode is disabled. For example, when a block is coded in CIIP, the flag indicating whether the OBMC mode applies is not signaled and is inferred to be false (i.e., the OBMC mode is disabled). In some embodiments, the flag indicating whether the OBMC mode applies is signaled with the value of the flag being false. LMCS can also be enabled for the block.

[0146]

[0171] In some embodiments, when LMCS is enabled for a block, frame, picture, slice, or tile, CIIP is disabled. Thus, LMCS and CIIP cannot be enabled at the same time. For example, a flag indicating whether CIIP mode applies is not signaled and is inferred to be false (i.e., CIIP mode is disabled). In some embodiments, a flag indicating whether CIIP mode applies is signaled with a value of the flag being false. OBMC can be enabled for the current block, frame, picture, slice, or tile.

[0147]

[0172] The above described embodiments may be combined in any combination.

[0148]

[0173] The embodiments may be further described using the following clauses. 1. A method for video processing in which combined inter- and intra-prediction (CIIP) and luma mapping and chroma scaling (LMCS) are applied, comprising: Obtaining an inter prediction signal, an intra prediction signal and an overlapped block motion compensation (OBMC) prediction signal; obtaining an intermediate weighted prediction signal by weighting an inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal; Including, A method, wherein the intermediate weighted prediction signal and the second prediction signal are both in the mapping domain or the original domain. 2. Obtaining an inter prediction signal, an intra prediction signal and an OBMC prediction signal, Obtaining an inter prediction signal in an original domain; Obtaining an intra prediction signal in a mapping domain; Obtaining an OBMC prediction signal in the original domain 2. The method of claim 1, further comprising: 3. Obtaining an intermediate weighted prediction signal by weighting an inter prediction signal and a first prediction signal among an intra prediction signal and an OBMC prediction signal; Transforming the inter prediction signal from an original domain to a mapping domain; weighting the transformed inter prediction signal and the intra prediction signal in the mapping domain to obtain an intermediate weighted prediction signal; Including, obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal; Transforming the OBMC prediction signal from an original domain to a mapping domain; Obtaining a final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the mapping domain. 3. The method of claim 2, further comprising: 4. Obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; Transforming the inter prediction signal from an original domain to a mapping domain; weighting the transformed inter prediction signal and the intra prediction signal in the mapping domain to obtain an intermediate weighted prediction signal; Including, obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal; Transforming the intermediate weighted prediction signal from the mapping domain to an original domain; Obtaining an intermediate final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the original domain; and obtaining a final prediction signal by transforming the intermediate final prediction signal from an original domain to a mapping domain. 3. The method of claim 2, further comprising: 5. Obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; Transforming the intra prediction signal from the mapping domain to the original domain; weighting the inter prediction signal and the transformed intra prediction signal in the original domain to obtain an intermediate weighted prediction signal; Including, obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal; Obtaining an intermediate final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the original domain; and obtaining a final prediction signal by transforming the intermediate final prediction signal from an original domain to a mapping domain. 3. The method of claim 2, further comprising: 6. Obtaining an intermediate weighted prediction signal by weighting an inter prediction signal and a first prediction signal among an intra prediction signal and an OBMC prediction signal; Obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and the OBMC prediction signal in the original domain. Including, obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal; Transforming the intermediate weighted prediction signal from an original domain to a mapping domain; obtaining a final prediction signal by weighting the transformed intermediate weighted prediction signal and the intra prediction signal in the mapping domain; 3. The method of claim 2, further comprising: 7. The method according to any one of clauses 1 to 6, wherein the inter prediction signal is obtained using an inter prediction process applied in normal merge mode. 8. The method according to any one of clauses 1 to 7, wherein the OBMC prediction signal is obtained using motion from neighbouring blocks. 9. A method for image processing, comprising: Receiving a bitstream associated with a target block; determining whether the target block is coded in a combined inter-prediction and intra-prediction (CIIP) mode; determining that if the current block is coded in a CIIP mode, an overlapped block motion compensation (OBMC) mode is not applied; In CIIP mode, processing the target block using luma mapping and chroma scaling (LMCS) A method comprising: 10. A method for image processing, comprising: Receiving a bitstream associated with a target block; determining whether luma mapping and chroma scaling (LMCS) is enabled for the target block; determining, in response to the LMCS being enabled, that a combined inter-prediction and intra-prediction (CIIP) mode is not applied; Processing the target block in overlapped block motion compensation (OBMC) mode using LMCS; A method comprising: 11. A method for image processing, comprising: determining that a combined inter-prediction and intra-prediction (CIIP) mode is enabled for a target block; determining an inter predictor using a motion vector, the motion vector being obtained using the motion vector predictor and the signaled motion vector differential; determining an intra-predictor; Obtaining a final predictor for the target block by weighting the inter predictor and the intra predictor. A method comprising: 12. The method of claim 11, wherein the target block is coded in normal inter mode and not in merge mode. 13. The method of claim 12, wherein inter modes other than the normal inter mode are disabled. 14. The method of claim 13, wherein the inter mode includes one or more of a local illumination compensation mode, a bi-prediction mode using coding unit level weights, an affine mode, a symmetric motion vector differential mode, an adaptive motion vector resolution mode, or a multi-hypothesis inter prediction mode. 15. Determining whether CIIP mode is enabled for the target block includes: Determine whether CIIP mode is enabled for the target block based on the flag. 15. The method according to any one of clauses 12 to 14, further comprising: 16. The method of claim 15, wherein the flag is decoded at the beginning of a normal inter mode syntax structure. 17. The method of claim 15, wherein the flag is decoded at the end of a normal inter mode syntax structure. 18. The method of any one of clauses 11 to 17, wherein the inter predictor is generated using a merge mode with motion vector difference (MMVD) method. 19. The method of any one of clauses 11 to 18, wherein the intra predictor is generated using a template-based intra mode derivation (TIMD) method or a decoder-side intra mode derivation (DIMD) method. 20. A non-transitory computer-readable medium for storing a bitstream, the bitstream comprising: 1. A non-transitory computer-readable medium comprising: a flag associated with encoded video data indicating that combined inter and intra prediction (CIIP) is used for the encoded video data, a target block being encoded in a normal inter mode and not in a merge mode, the flag configured to cause a decoder to decode the target block using a CIIP mode and to disable inter modes other than the normal inter mode. 22. The non-transitory computer-readable medium of clause 20, wherein the flag is at the beginning of a normal intermode syntax structure. 23. The non-transitory computer-readable medium of clause 20, wherein the flag is located at the end of a normal intermode syntax structure. 24. A non-transitory computer-readable medium for storing a bitstream, the bitstream comprising: 1. A non-transitory computer-readable medium comprising: a flag associated with encoded video data, the flag indicating that a target block is encoded in a combination of combined inter and intra prediction (CIIP) and merge mode with motion vector differential (MMVD) modes, the flag configured to cause a decoder to decode the target block using the CIIP mode and obtain an inter predictor using the MMVD mode. 25. A non-transitory computer-readable medium for storing a bitstream, the bitstream comprising: 1. A non-transitory computer-readable medium comprising: a flag associated with encoded video data, the flag indicating that a combination of combined inter and intra prediction (CIIP) and merge mode with motion vector differential (MMVD) modes is used for inter prediction, the flag configured to cause a decoder to decode a target block using the CIIP mode and obtain an inter predictor using the MMVD mode. 26. A non-transitory computer-readable medium for storing a bitstream, the bitstream comprising: a first flag associated with the encoded video data, the first flag indicating whether a combined inter-prediction and intra-prediction (CIIP) mode is enabled; a second flag associated with the coded video data that indicates whether overlapped block motion compensation (OBMC) mode is applied; Including, A non-transitory computer readable medium, wherein if the first flag is true, then the second flag is inferred to be false. 27. A non-transitory computer-readable medium for storing a bitstream, the bitstream comprising: a first flag associated with the encoded video data, the first flag indicating whether luma mapping and chroma scaling (LMCS) is enabled; a second flag associated with the encoded video data, the second flag indicating whether combined inter-prediction and intra-prediction (CIIP) is enabled; Including, A non-transitory computer readable medium, wherein if the first flag is true, then the second flag is inferred to be false. 28. A non-transitory computer-readable medium for storing a bitstream, the bitstream comprising: a first flag associated with the encoded video data, the first flag indicating whether luma mapping and chroma scaling (LMCS) is enabled; a second flag associated with the encoded video data, the second flag indicating whether combined inter-prediction and intra-prediction (CIIP) is enabled; Including, If the first flag and the second flag are both true, the decoder: Obtaining an inter prediction signal in an original domain, an intra prediction signal in a mapping domain, and an overlapped block motion compensation (OBMC) prediction signal in the original domain; Obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and the OBMC prediction signal in the original domain; Transforming the intermediate weighted prediction signal from an original domain to a mapping domain; obtaining a final prediction signal by weighting the transformed intermediate weighted prediction signal and the intra prediction signal in the mapping domain; 23. A non-transitory computer-readable medium configured to execute the steps of: 29. An apparatus for performing video data processing in which combined inter-prediction and intra-prediction (CIIP) and luma mapping and chroma scaling (LMCS) are applied, comprising: A memory configured to store instructions; One or more processors wherein one or more processors: Obtaining an inter prediction signal, an intra prediction signal and an overlapped block motion compensation (OBMC) prediction signal; obtaining an intermediate weighted prediction signal by weighting an inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal; configured to execute instructions to cause the device to execute The apparatus, wherein the intermediate weighted prediction signal and the second prediction signal are both in the mapping domain or the original domain. 30. One or more processors: Obtaining an inter prediction signal in an original domain; Obtaining an intra prediction signal in a mapping domain; Obtaining an OBMC prediction signal in the original domain 30. The apparatus of clause 29, further configured to execute instructions to cause the apparatus to perform. 31. One or more processors: Transforming the inter prediction signal from an original domain to a mapping domain; weighting the transformed inter prediction signal and the intra prediction signal in the mapping domain to obtain an intermediate weighted prediction signal; Transforming the OBMC prediction signal from an original domain to a mapping domain; Obtaining a final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the mapping domain. 31. The apparatus of claim 30, further configured to execute instructions to cause the apparatus to perform. 32. One or more processors: Transforming the inter prediction signal from an original domain to a mapping domain; weighting the transformed inter prediction signal and the intra prediction signal in the mapping domain to obtain an intermediate weighted prediction signal; Transforming the intermediate weighted prediction signal from the mapping domain to an original domain; Obtaining an intermediate final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the original domain; and obtaining a final prediction signal by transforming the intermediate final prediction signal from an original domain to a mapping domain. 31. The apparatus of claim 30, further configured to execute instructions to cause the apparatus to perform. 33. One or more processors: Transforming the intra prediction signal from the mapping domain to the original domain; obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and the transformed intra prediction signal in the original domain; Obtaining an intermediate final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the original domain; and obtaining a final prediction signal by transforming the intermediate final prediction signal from an original domain to a mapping domain. 31. The apparatus of claim 30, further configured to execute instructions to cause the apparatus to perform. 34. One or more processors: Obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and the OBMC prediction signal in the original domain; Transforming the intermediate weighted prediction signal from an original domain to a mapping domain; obtaining a final prediction signal by weighting the transformed intermediate weighted prediction signal and the intra prediction signal in the mapping domain; 31. The apparatus of claim 30, further configured to execute instructions to cause the apparatus to perform. 35. The apparatus of any one of clauses 29 to 34, wherein the inter prediction signal is obtained using an inter prediction process applied in normal merge mode. 36. An apparatus according to any one of clauses 29 to 35, wherein the OBMC prediction signal is obtained using motion from neighbouring blocks. 37. An apparatus for performing video data processing, comprising: A memory illustrated storing instructions; One or more processors wherein one or more processors: Receiving a bitstream associated with a target block; determining whether the target block is coded in a combined inter-prediction and intra-prediction (CIIP) mode; determining that if the current block is coded in a CIIP mode, an overlapped block motion compensation (OBMC) mode is not applied; In CIIP mode, processing the target block using luma mapping and chroma scaling (LMCS) The apparatus is further configured to execute instructions to cause the apparatus to perform the 38. An apparatus for performing video data processing, comprising: A memory illustrated storing instructions; One or more processors wherein one or more processors: Receiving a bitstream associated with a target block; determining whether luma mapping and chroma scaling (LMCS) is enabled for the target block; determining, in response to the LMCS being enabled, that a combined inter-prediction and intra-prediction (CIIP) mode is not applied; Processing the target block in overlapped block motion compensation (OBMC) mode using LMCS; The apparatus is further configured to execute instructions to cause the apparatus to perform the 39. An apparatus for performing video data processing, comprising: A memory illustrated storing instructions; One or more processors wherein one or more processors: determining that a combined inter-prediction and intra-prediction (CIIP) mode is enabled for a target block; determining an inter predictor using a motion vector, the motion vector being obtained using the motion vector predictor and the signaled motion vector differential; determining an intra-predictor; Obtaining a final predictor for the target block by weighting the inter predictor and the intra predictor. The apparatus is further configured to execute instructions to cause the apparatus to perform the 40. The apparatus of clause 39, wherein the target block is coded in normal inter mode and not in merge mode. 41. The apparatus of clause 40, wherein inter modes other than the normal inter mode are disabled. 42. The apparatus of clause 41, wherein the inter mode includes one or more of a local illumination compensation mode, a bi-prediction mode using coding unit level weights, an affine mode, a symmetric motion vector differential mode, an adaptive motion vector resolution mode, or a multi-hypothesis inter prediction mode. 43. One or more processors: Determine whether CIIP mode is enabled for the target block based on the flag. 43. An apparatus as described in any one of clauses 40 to 42, further configured to execute instructions to cause the apparatus to perform the 44. The apparatus of clause 43, wherein the flag is decoded at the beginning of a normal inter mode syntax structure. 45. The apparatus of clause 44, wherein the flag is decoded at the end of a normal inter mode syntax structure. 46. ​​The apparatus of any one of clauses 39 to 45, wherein the inter predictor is generated using a merge mode with motion vector difference (MMVD) method. 47. The apparatus of any one of clauses 39 to 46, wherein the intra predictor is generated using a template-based intra mode derivation (TIMD) method or a decoder-side intra mode derivation (DIMD) method.

[0149]

[0174] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided. In some embodiments, the medium may store all or a portion of a video bitstream having one or more flags indicating an applied prediction mode, such as the modes applied with respect to Figures 14-19. In some embodiments, the medium may store instructions that can be executed by a device (such as the disclosed encoders and decoders) to perform the methods described above. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid state drive, a magnetic tape or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with a pattern of holes, a RAM, a PROM, an EPROM, a FLASH-EPROM or any other flash memory, a NVRAM, a cache, a register, any other memory chip or cartridge, and network-connected versions thereof. The device may include one or more processors (CPUs), input / output interfaces, a network interface, and / or memory.

[0150]

[0175] It should be noted that relational terms used herein, such as "first" and "second," are used merely to distinguish one entity or operation from another, and do not require or imply an actual relationship or order between those entities or operations. Moreover, the words "comprising," "including," "having," "containing," and other similar forms are intended to be equivalent and open-ended, in that the item or items following any of these words are not intended to be an exhaustive listing of such item or items, nor are they intended to be limited to only the listed item or items.

[0151]

[0176] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations unless infeasible. For example, if it is stated that a database may include A or B, then the database may include A, B, A and B, unless specifically stated otherwise or infeasible. As a second example, if it is stated that a database may include A, B or C, then the database may include A, B, C, A and B, A and C, B and C, A and B and C, unless specifically stated otherwise or infeasible.

[0152]

[0177] It will be understood that the above described embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in a computer-readable medium as described above. The software, when executed by a processor, can perform the disclosed method. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that more than one of the above described modules / units can be combined into one module / unit, and each of the above described modules / units can be further divided into multiple sub-modules / sub-units.

[0153]

[0178] In the foregoing specification, the embodiments have been described with reference to many specific details that may vary from implementation to implementation. Certain adaptations and modifications of the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims. It is also intended that the order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to the particular order of steps. Thus, one skilled in the art can appreciate that steps can be performed in different orders while performing the same method.

[0154]

[0179] In the drawings and specification, illustrative embodiments are disclosed. However, many variations and modifications to these embodiments may be made. Accordingly, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. 1. A method for encoding a video sequence in which combined inter- and intra-prediction (CIIP) and luma mapping and chroma scaling (LMCS) are applied, the method comprising: receiving a video sequence; and The video sequence Obtaining an inter prediction signal, an intra prediction signal and an overlapped block motion compensation (OBMC) prediction signal; obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal; By encoding Including, The method, wherein the intermediate weighted prediction signal and the second prediction signal are both in the mapping domain or the original domain.

2. Obtaining the inter prediction signal, the intra prediction signal, and the OBMC prediction signal, obtaining the inter prediction signal in the original domain; obtaining the intra prediction signal in a mapping domain; obtaining the OBMC prediction signal in the original domain; The method of claim 1 further comprising:

3. obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and the first prediction signal among the intra prediction signal and the OBMC prediction signal, transforming the inter-predicted signal from the original domain to the mapping domain; weighting the transformed inter prediction signal and the intra prediction signal in the mapping domain to obtain the intermediate weighted prediction signal; Including, obtaining the final prediction signal by weighting the intermediate weighted prediction signal and the second prediction signal among an intra prediction signal and the OBMC prediction signal; transforming the OBMC prediction signal from the original domain to the mapping domain; obtaining the final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the mapping domain; The method of claim 2 further comprising:

4. obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and the first prediction signal among the intra prediction signal and the OBMC prediction signal, transforming the inter-predicted signal from the original domain to the mapping domain; weighting the transformed inter prediction signal and the intra prediction signal in the mapping domain to obtain the intermediate weighted prediction signal; Including, obtaining the final prediction signal by weighting the intermediate weighted prediction signal and the second prediction signal among an intra prediction signal and the OBMC prediction signal; transforming the intermediate weighted prediction signal from the mapping domain to the original domain; obtaining an intermediate final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the original domain; obtaining the final predicted signal by transforming the intermediate final predicted signal from the original domain to the mapping domain; The method of claim 2 further comprising:

5. obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; transforming the intra-prediction signal from the mapping domain to the original domain; obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and the transformed intra prediction signal in the original domain; Including, obtaining the final prediction signal by weighting the intermediate weighted prediction signal and the second prediction signal among an intra prediction signal and the OBMC prediction signal; obtaining an intermediate final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the original domain; obtaining the final predicted signal by transforming the intermediate final predicted signal from the original domain to the mapping domain; The method of claim 2 further comprising:

6. obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and the OBMC prediction signal in the original domain; Including, obtaining the final prediction signal by weighting the intermediate weighted prediction signal and the second prediction signal among an intra prediction signal and the OBMC prediction signal; transforming the intermediate weighted prediction signal from the original domain to the mapping domain; obtaining the final prediction signal by weighting the transformed intermediate weighted prediction signal and the intra prediction signal in the mapping domain; The method of claim 2 further comprising:

7. The method of claim 1 , wherein the inter-prediction signal is obtained using an inter-prediction process applied in a normal merge mode.

8. The method of claim 1 , wherein the OBMC prediction signal is obtained using motion from neighboring blocks.

9. 1. A method for decoding a bitstream in which combined inter-prediction and intra-prediction (CIIP) and luma mapping and chroma scaling (LMCS) are applied, comprising: receiving a bitstream; decoding the bitstream to output a video sequence; wherein said decoding comprises: Obtaining an inter prediction signal, an intra prediction signal and an overlapped block motion compensation (OBMC) prediction signal; obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal; Including, The method, wherein the intermediate weighted prediction signal and the second prediction signal are both in the mapping domain or the original domain.

10. Obtaining the inter prediction signal in an original domain; obtaining the intra prediction signal in a mapping domain; obtaining the OBMC prediction signal in the original domain; 10. The method of claim 9, further comprising:

11. Transforming the inter-predicted signal from the original domain to the mapping domain; weighting the transformed inter prediction signal and the intra prediction signal in the mapping domain to obtain the intermediate weighted prediction signal; transforming the OBMC prediction signal from the original domain to the mapping domain; obtaining the final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the mapping domain; The method of claim 10 further comprising:

12. The method of claim 11, further comprising: transforming the inter-predicted signal from the original domain to the mapping domain; weighting the transformed inter prediction signal and the intra prediction signal in the mapping domain to obtain the intermediate weighted prediction signal; transforming the intermediate weighted prediction signal from the mapping domain to the original domain; obtaining an intermediate final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the original domain; obtaining the final predicted signal by transforming the intermediate final predicted signal from the original domain to the mapping domain; The method of claim 10 further comprising:

13. The method of claim 12, further comprising: transforming the intra-predicted signal from the mapping domain to the original domain; obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and the transformed intra prediction signal in the original domain; obtaining an intermediate final prediction signal by weighting the intermediate weighted prediction signal and the OBMC prediction signal in the original domain; obtaining the final predicted signal by transforming the intermediate final predicted signal from the original domain to the mapping domain; The method of claim 10 further comprising:

14. Obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and the OBMC prediction signal in the original domain; transforming the intermediate weighted prediction signal from the original domain to the mapping domain; obtaining the final prediction signal by weighting the transformed intermediate weighted prediction signal and the intra prediction signal in the mapping domain; The method of claim 10 further comprising:

15. The method of claim 9 , wherein the inter-prediction signal is obtained using an inter-prediction process applied in a normal merge mode.

16. The method of claim 9 , wherein the OBMC prediction signal is obtained using motion from neighboring blocks.

17. A method for signaling a bitstream, comprising: receiving a video sequence, wherein combined inter and intra prediction (CIIP) and luma mapping and chroma scaling (LMCS) are applied to the video sequence; The video sequence Obtaining an inter prediction signal, an intra prediction signal and an overlapped block motion compensation (OBMC) prediction signal; obtaining an intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; obtaining a final prediction signal by weighting the intermediate weighted prediction signal and a second prediction signal among the intra prediction signal and the OBMC prediction signal; and encoding the signaling a bitstream generated based on said encoding. Including, The method, wherein the intermediate weighted prediction signal and the second prediction signal are both in the mapping domain or the original domain.

18. Obtaining the inter prediction signal, the intra prediction signal, and the OBMC prediction signal, obtaining the inter prediction signal in the original domain; obtaining the intra prediction signal in a mapping domain; obtaining the OBMC prediction signal in the original domain; 20. The method of claim 17, further comprising:

19. obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and a first prediction signal among the intra prediction signal and the OBMC prediction signal; obtaining the intermediate weighted prediction signal by weighting the inter prediction signal and the OBMC prediction signal in the original domain; Including, obtaining the final prediction signal by weighting the intermediate weighted prediction signal and the second prediction signal among the intra prediction signal and the OBMC prediction signal; transforming the intermediate weighted prediction signal from the original domain to the mapping domain; obtaining the final prediction signal by weighting the transformed intermediate weighted prediction signal and the intra prediction signal in the mapping domain; 20. The method of claim 18, further comprising: