Intra- and inter-coding of sub-blocks
By employing sub-block-based inter prediction modes and SbTMVP, the patent addresses inefficiencies in existing video coding technologies, enhancing compression and decoding performance through optimized sub-block reconstruction.
Patent Information
- Application Number
- JP2024538733
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-06-08
- Filing Date
- 2023-06-13
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-06-13
AI Technical Summary
Existing video coding technologies face inefficiencies in compressing video data due to limitations in intra- and inter-prediction methods, particularly in handling sub-blocks within coding units, leading to suboptimal compression and decoding performance.
The implementation of sub-block-based inter prediction modes, including affine modes and regression-based inter prediction, along with sub-block-based temporal motion vector prediction (SbTMVP), allows for more efficient reconstruction of video blocks by distinguishing and processing sub-blocks within a coding unit (CU) based on intra or inter prediction, using flags and specific reconstruction orders.
This approach enhances video compression efficiency by optimizing the reconstruction of sub-blocks within CUs, improving data reduction and quality in video encoding and decoding processes.
Smart Images

Figure 0007775483000004 
Figure 0007775483000005 
Figure 0007775483000006
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to U.S. Patent Application No. 18 / 207,578, entitled "SUBBLOCK INTRA AND INTER CODING," filed June 8, 2023, which claims the benefit of priority to U.S. Provisional Application No. 63 / 359,155, entitled "Subblock Intra and Inter Coding," filed July 7, 2022. The entire disclosure of the prior application is incorporated by reference.
[0002] [Technical field] This disclosure describes embodiments that relate generally to video coding. [Background technology]
[0003] The background art discussion provided herein is intended to generally present the context for the present disclosure, and the inventors' work, to the extent described in this background art section, as well as aspects of the description that are not admitted as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.
[0004] Image / video compression can help transmit image / video files across different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, video codecs can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, video codecs can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation is commonly represented by a motion vector (MV). Summary of the Invention
[0005] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit receives a coding bitstream carrying at least a picture including a block including a plurality of sub-blocks, and determines, based on a value of a first syntax element in the coding bitstream, that a current coding unit (CU) in the picture is coded in a sub-block-based inter prediction mode, and determines that one or more first sub-blocks in the current CU coded in the sub-block-based inter prediction mode are coded by intra prediction. The processing circuit reconstructs one or more second sub-blocks of the current CU by inter prediction based on the sub-block-based inter prediction mode, where the one or more second sub-blocks do not overlap with the one or more first sub-blocks in the current CU. The processing circuit reconstructs the one or more first sub-blocks of the current CU by intra prediction while the current CU is coded in the sub-block-based inter prediction mode.
[0006] In some examples, the subblock-based inter prediction mode includes at least one of an affine mode, a regression-based inter prediction mode, or a subblock-based temporal motion vector prediction (SbTMVP) mode.
[0007] In some examples, a first flag indicating whether at least one sub-block in a current CU is coded by intra prediction is determined, and then second flags associated with the sub-blocks in the current CU are decoded from the bitstream, and the second flags associated with the sub-blocks in the current CU indicate whether the sub-blocks are coded by intra prediction. In one example, the first flag is decoded from the bitstream.
[0008] In some examples, the first flag is derived without signaling. In one example, the first flag is derived as a true value in response to determining that a co-located block for a sub-block in the current CU does not have a valid motion vector when the current CU is coded in sub-block-based temporal motion vector prediction (SbTMVP) mode. In another example, the first flag is derived as a false value in response to determining that each sub-block in the current CU has a valid motion vector when the current CU is coded in SbTMVP mode.
[0009] In some examples, one or more second sub-blocks of the current CU are reconstructed before one or more first sub-blocks of the current CU, and the one or more first sub-blocks of the current CU are reconstructed according to a specific order determined based on the positions of the one or more second sub-blocks.
[0010] In one example, one or more second sub-blocks of the current CU are reconstructed according to a raster scan order, and one or more first sub-blocks of the current CU are reconstructed according to a raster scan order after the one or more second sub-blocks are reconstructed.
[0011] In some examples, for a first sub-block in one or more first sub-blocks, the first sub-block is predicted according to adjacent inter-coding sub-blocks and adjacent intra-coding sub-blocks in the reconstructed current CU.
[0012] In one example, the transform unit size is determined to be the size of the sub-block in response to a first flag being true, the first flag indicating whether at least one sub-block in the current CU is coded by intra prediction. In another example, the transform unit size is determined to be the size of the current CU in response to the first flag being false.
[0013] In some examples, an inverse transform is performed to obtain a transform unit of the size of the current CU, and residual values corresponding to one or more second sub-blocks are determined according to the transform unit. Then, the one or more second sub-blocks predicted based on inter prediction and the residual values are combined to reconstruct the one or more second sub-blocks.
[0014] In some examples, an inverse transform is performed to obtain one or more transform units of a size corresponding to one or more first sub-blocks in the current CU, where the one or more transform units correspond to one or more first sub-blocks, respectively. Residual values corresponding to the one or more first sub-blocks are determined according to the one or more transform units. Then, the one or more first sub-blocks predicted based on intra prediction and the residual values are combined to reconstruct the one or more first sub-blocks.
[0015] In some examples, a first sub-block within one or more first sub-blocks is reconstructed based on an intra-prediction mode selected from at least one of a set of most probable modes as a selected set of smooth modes or a set of non-directional modes.
[0016] In some examples, the first sub-blocks within the one or more first sub-blocks are rearranged according to at least one of one or more right columns, one or more bottom rows, one or more corner samples, one or more top right rows, or one or more bottom left columns.
[0017] In some examples, for each sub-block in the one or more first sub-blocks, both the luma and chroma components of the sub-block are reconstructed by intra prediction.
[0018] In some examples, inter prediction is determined according to which one or more first sub-blocks in the current CU are coded by intra prediction, along with constraints for reconstructing one or more second sub-blocks. In one example, uni-prediction is used for inter prediction.
[0019] In some examples, a predictor for an intra prediction mode is determined from neighboring pixels of the current CU, and the intra prediction mode includes at least one of a most probable mode (MPM), a decoder-side intra mode derivation (DIMD) mode, a template-based intra mode derivation (TIMD) mode, or a directional intra prediction mode.
[0020] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding. [Brief explanation of the drawings]
[0021] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100). [Figure 2]FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4] 1 illustrates the locations of spatial merge candidates according to one embodiment of the present disclosure. [Figure 5] 1 illustrates candidate pairs considered for redundancy checking of spatial merge candidates, according to one embodiment of the present disclosure. [Figure 6] 10 illustrates exemplary motion vector scaling for temporal merge candidates. [Figure 7] 10 shows exemplary candidate positions (eg, C0 and C1) for temporal merge candidates for the current CU. [Figure 8] 1 illustrates an example of a search process in some examples. [Figure 9] Examples of search points in some examples are shown. [Figure 10] 1 shows a diagram illustrating the direction of improvement in some examples. [Figure 11] 1 illustrates an exemplary subblock-based temporal motion vector prediction (SbTMVP) process in some examples. [Figure 12] 1 illustrates an exemplary subblock-based temporal motion vector prediction (SbTMVP) process in some examples. [Figure 13] 1 illustrates a diagram of a coding block according to some embodiments of the present disclosure. [Figure 14] 1 illustrates a diagram of a coding block according to some embodiments of the present disclosure. [Figure 15] 10 shows a diagram of using intra prediction for sub-blocks, according to some embodiments of this disclosure. [Figure 16] 10 shows a diagram of using intra prediction for sub-blocks, according to some embodiments of this disclosure. [Figure 17]1 shows a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 18] 1 shows a flowchart outlining another process according to some embodiments of the present disclosure. [Figure 19] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0022] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is an example of an application of the disclosed subject matter, a video encoder and video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.), etc.
[0023] The video processing system (100) includes a video source (101) and a capture subsystem (113) that can include, for example, a digital camera, creating a stream of uncompressed video pictures (102). In one example, the video picture stream (102) includes samples captured by the digital camera. The video picture stream (102), shown with a thick line to emphasize its high data volume compared to the coded video data (104) (or coded video bitstream), can be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The coded video data (104) (or coded video bitstream), shown with a thin line to emphasize its low data volume compared to the video picture stream (102), can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as client subsystems 106 and 108 in FIG. 1, can access the streaming server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 may include, for example, a video decoder 110 within an electronic device 130. The video decoder 110 decodes an input copy 107 of the encoded video data and creates an output stream 111 of video pictures that can be rendered on a display 112 (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data 104, 107, and 109 (e.g., a video bitstream) can be encoded according to several video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, a developing video coding standard is informally known as Versatile Video Coding (VVC).The disclosed subject matter may be used in the context of VVC.
[0024] It is noted that electronic devices 120 and 130 may include other components (not shown). For example, electronic device 120 may include a video decoder (not shown), and electronic device 130 may include a video encoder (not shown).
[0025] 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used in place of the video decoder (110) in the example of FIG. 1.
[0026] The receiver (231) may receive one or more coded video sequences to be decoded by the video decoder (210). In one embodiment, one coded video sequence may be received at a time, with the decoding of each coded video sequence being independent of the decoding of the other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (231) may receive the coded video data along with other data (e.g., coded audio data and / or auxiliary data streams), which may be forwarded to respective using entities (not shown). The receiver (231) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as the "parser" (220)). In certain applications, the buffer memory (215) is part of the video decoder (210). In other cases, it may be external to the video decoder (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder 210, for example, to prevent network jitter, and there may be yet another buffer memory 215 within the video decoder 210, for example, to handle playback timing. If the receiver 231 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory 215 may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer memory 215 may be needed and may be relatively large, advantageously adaptively sized, and may be implemented at least in part in an operating system and similar elements (not shown) external to the video decoder 210.
[0027] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and potentially include information for controlling a rendering device, such as a rendering device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but can be coupled to the electronic device (230), as shown in FIG. 2. The rendering device control information may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a set of subgroup parameters for at least one subgroup of pixels in a video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.
[0028] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to generate symbols (221).
[0029] The reconstruction of the symbols (221) can involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (220). The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.
[0030] In addition to the functional blocks described above, the video decoder (210) can be conceptually subdivided into multiple functional units, as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:
[0031] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as symbols (221) from the parser (220), along with control information (including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc.) The scalar / inverse transform unit (251) can output blocks containing sample values that can be input to an aggregator (255).
[0032] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks may relate to blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates blocks of the same size and shape as the block being reconstructed using surrounding, already reconstructed information retrieved from the current picture buffer (258). The current picture buffer (258), for example, buffers the partially reconstructed and / or fully reconstructed current picture. In some cases, the aggregator (255) adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).
[0033] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (253) may access the reference picture memory (257) to retrieve samples used for prediction. After motion-compensating the retrieved samples according to the symbols (221) associated with the block, these samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) retrieves prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (253), for example, in the form of symbols (221) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values retrieved from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.
[0034] The output samples of the aggregator 255 may be subjected to various loop filtering techniques in a loop filter unit 256. The video compression techniques may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called a coded video bitstream) and made available to the loop filter unit 256 as symbols 221 from the parser 220. The video compression may also be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, as well as to previously reconstructed loop-filtered sample values.
[0035] The output of the loop filter unit (256) can be a sample stream that can be output to a rendering device (212) and stored in a reference picture memory (257) for use in future inter-picture prediction.
[0036] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.
[0037] The video decoder (210) may perform decoding operations according to a given video compression technology or standard, such as ITU-T Rec. H.265. A coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence conforms to both the syntax and profile of the video compression technology or standard as documented in the video compression technology or standard. Specifically, a profile can select certain tools from all tools available in the video compression technology or standard as the only tools for use under that profile. Compliance may also require the complexity of the coded video sequence to fall within a range defined by a level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further constrained through a hypothetical reference decoder (HRD) specification and metadata about HRD buffer management conveyed in the coded video sequence.
[0038] In one embodiment, the receiver (231) may receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0039] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.
[0040] The video encoder (303) may receive video samples from a video source (301) (not part of the electronic device (320) in the example of FIG. 3) that may capture video images to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).
[0041] The video source (301) may provide a source video sequence to be encoded by the video encoder (303) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media presentation system, the video source (301) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. Video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following discussion will focus on samples.
[0042] According to one embodiment, the video encoder (303) may encode and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other required time constraints. Achieving an appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) can control and is operatively coupled to other functional units, as described below. Coupling is not shown for clarity. Parameters set by the controller (350) can include rate control-related parameters (e.g., picture skip, quantization, lambda value for rate-distortion optimization techniques), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other appropriate functionality associated with the video encoder (303) optimized for a particular system design.
[0043] In some embodiments, the video encoder (303) is configured to operate in an encoding loop. As a very simplified description, in one example, the encoding loop can include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data similar to that generated by the (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-for-bit accurate results independent of the location of the decoder (local or remote), the contents in the reference picture memory (334) are also bit-for-bit accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (including the resulting drift when synchronization cannot be maintained, for example, due to channel errors) is used in several related techniques as well.
[0044] The operation of the "local" decoder (333) can be the same as the "remote" decoder (210), such as the video decoder (210), already described in detail above in connection with Figure 2. However, briefly referring to Figure 2, because symbols are available and the encoding / decoding of symbols into an encoded video sequence by the entropy coder (345) and parser (220) can be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (433).
[0045] In one embodiment, decoder technology, excluding analysis / entropy decoding, present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the subject matter of the disclosure focuses on decoder operation. A description of the encoder technology can be omitted, as it is the reverse of the decoder technology, which is described generically. In certain areas, more detailed descriptions are provided below.
[0046] During operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.
[0047] The local video decoder (333) may decode the coded video data of pictures that may be designated as reference pictures based on symbols generated by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence may typically be a replica of the source video sequence, with some errors. The local video decoder (333) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (334). In this way, the video encoder (303) may locally store copies of reconstructed reference pictures that have common content as reconstructed reference pictures (without transmission errors) obtained by the far-end video decoder.
[0048] The predictor (335) may perform a predictive search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or specific metadata (reference picture motion vectors, block shapes, etc.), which may serve as suitable prediction references for the new picture. The predictor (335) may operate sample block-by-pixel block to find suitable prediction references. In some cases, the input picture determined by the search results obtained by the predictor (335) may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).
[0049] The controller (350) may manage the encoding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0050] The output of all the above functional units may undergo entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0051] The transmitter (340) may buffer the coded video sequence produced by the entropy coder (345) and prepare it for transmission over a communication channel (360), which may be a hardware or software link to a storage device that stores the coded video data. The transmitter (340) may merge the coded video data from the video encoder (330) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (not shown).
[0052] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign each coded picture a particular coding picture type. The coding picture type may affect the coding technique that may be applied to each picture. For example, a picture may be assigned as one of the following picture types:
[0053] An intra picture (I-picture) may be one that can be coded and decoded without using other pictures in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will recognize these variations of I-pictures and their respective uses and characteristics.
[0054] A predicted picture (P picture) may be one that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values of each block.
[0055] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0056] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0057] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.
[0058] In one embodiment, the transmitter (340) may transmit additional data along with the coded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other types of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0059] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0060] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. Bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, both of which precede the current picture in decoding order (but may also be past and future, respectively, in display order) in a video. A block in the current picture can be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block can be predicted by a combination of the first and second reference blocks.
[0061] Furthermore, to improve coding efficiency, merge mode techniques can be used in inter-picture prediction.
[0062] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU can be partitioned into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the CU's prediction type, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation of coding (encoding / decoding) is performed on a prediction block basis. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0063] It is noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.
[0064] Aspects of this disclosure provide techniques for sub-block intra- and inter-coding.
[0065] VVC allows various inter-prediction modes to be used. For an inter-predicted CU, motion parameters may include an MV, one or more reference picture indexes, a reference picture list usage index, and further information about specific coding features to be used for inter-predicted sample generation. Motion parameters may be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU may be associated with a PU and may not have significant residual coefficients, coded motion vector or MV differences (e.g., MVDs), or reference picture indices. A merge mode may be specified, in which motion parameters for the current CU are obtained from neighboring CUs, including spatial and / or temporal candidates, and optionally further information as introduced in VVC. The merge mode may be applied not only to skip mode but also to inter-predicted CUs. In one example, an alternative to the merge mode is explicit transmission of motion parameters, in which the MV, the corresponding reference picture index for each reference picture list, and a reference picture list usage flag, as well as other information, are explicitly signaled for each CU.
[0066] In one VVC-like embodiment, the VVC Test model (VTM) reference software supports enhanced merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8x8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder-side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), and geometric partitioning mode (GPM). Inter-prediction and related methods are described in more detail below.
[0067] In some examples, enhanced merge prediction can be used. In one example, such as VTM4, a merge candidate list is constructed by including, in order, five types of candidates: spatial motion vector predictor (MVP) from spatially neighboring CUs, temporal MVP from co-located CUs, history-based MVP (HMVP) from a first-in-first-out (FIFO) table, pairwise average MVP, and zero MV.
[0068] The size of the merge candidate list can be signaled in the slice header. In one example, the maximum allowed size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate (e.g., merge index) can be coded using truncated unary binarization (TU). The first bin of the merge index can be coded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.
[0069] Some examples of the generation process for each category of merge candidates are provided below. In one embodiment, spatial candidates are derived as follows: The derivation of spatial merge candidates in VVC can be the same as that in HEVC. In one example, up to four merge candidates are selected from the candidates in the positions shown in Figure 4.
[0070] 4 shows the positions of spatial merge candidates according to an embodiment of the present invention. Referring to FIG. 4, the derivation order is B1, A1, B0, A0, and B2. Position B2 is only considered when none of the CUs in positions A0, B0, B1, and A1 are available (e.g., because the CUs belong to another slice or another tile) or are intra-coded. After the candidate in position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures that candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.
[0071] To reduce computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only the pairs connected by arrows in Figure 5 are considered, and a candidate is added to the candidate list only if the corresponding candidates used in the redundancy check do not have the same motion information.
[0072] 5 illustrates candidate pairs considered for spatial merge candidate redundancy checking according to an embodiment of the present disclosure. Referring to FIG. 5, the pairs connected by arrows are A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Thus, candidates at positions B1, A0, and / or B2 can be compared with the candidate at position A1, and candidates at positions B0 and / or B2 can be compared with the candidate at position B1.
[0073] In one embodiment, temporal candidates are derived as follows: In one example, only one temporal merge candidate is added to the candidate list. Figure 6 shows exemplary motion vector scaling for temporal merge candidates. To derive a temporal merge candidate for a current CU (611) in a current picture (601), a scaled MV (621) (e.g., indicated by a dotted line in Figure 6) can be derived based on a co-located CU (612) belonging to a co-located reference picture (604). The reference picture list used to derive the co-located CU (612) can be explicitly signaled in the slice header. The scaled MV (621) for the temporal merge candidate can be obtained as indicated by a dotted line in Figure 6. The scaled MV (621) can be scaled from the MV of the co-located CU (612) using picture order count (POC) distances tb and td. The POC distance tb may be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td may be defined as the POC difference between the co-located reference picture (604) of the co-located picture (603) and the co-located picture (603). The reference picture index of the temporal merge candidate may be set to zero.
[0074] 7 shows exemplary candidate positions (e.g., C0 and C1) for temporal merge candidates for the current CU. The position of the temporal merge candidate can be selected from candidate positions C0 and C1. Candidate position C0 is located at the bottom right corner of the co-located CU (710) of the current CU. Candidate position C1 is located at the center of the co-located CU (710) of the current CU. If the CU at candidate position C0 is unavailable, intra-coded, or outside the current row of the CTU, candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, inter-coded, and located in the current row of the CTU, candidate position C0 is used to derive the temporal merge candidate.
[0075] In one embodiment, a merge with motion vector difference (MMVD) mode is used in VVC, etc., where implicitly derived motion information can be used to predict samples of a CU (e.g., the current CU). The MMVD mode is used in either skip mode or merge mode with a motion vector representation method. An MMVD merge flag can be signaled to specify whether the MMVD mode is used for a CU, for example, after signaling a skip flag or a merge flag.
[0076] In some examples, MMVD reuses merge candidates. A candidate is selected from among the merge candidates and can be further extended by a motion vector representation. MMVD provides a motion vector representation with simplified signaling. In some examples, the motion vector representation includes a starting point, a motion magnitude, and a motion direction.
[0077] In some instances (e.g., VVC), the MMVD technique can use a merge candidate list to select candidate starting points, but in one instance, only candidates that are the default merge type (MRG_TYPE_DEFAULT_N) are considered for MMVD expansion.
[0078] In some examples, a base candidate index is used to define a starting point. The base candidate index indicates the best candidate among the candidates in a list such as that shown in Table 1. For example, the list is a merge candidate list with a motion vector predictor (MVP). The base candidate index can indicate the best candidate in the merge candidate list. [Table 1]
[0079] Note that in one example, if the number of base candidates is equal to 1, the base candidate IDX is not signaled.
[0080] In MMVD mode, after a merge candidate (also called an MV base or MV starting point) is selected, the merge candidate can be refined by further information, such as signaled MVD information. The further information can include an index used to specify the magnitude of the motion (a distance index, e.g., mmvd_distance_idx[x0][y0], etc.) and an index used to indicate the direction of the motion (a direction index, e.g., mmvd_direction_idx[x0][y0], etc.). In MMVD mode, one of the first two candidates in the merge list can be selected as the MV base. For example, a merge candidate flag (e.g., mmvd_cand_flag[x0][y0]) indicates one of the first two candidates in the merge list. The merge candidate flag can be signaled to indicate (e.g., specify) which of the first two candidates is selected. The further information can indicate the MVD (or motion offset) relative to the MV base. For example, the motion magnitude indicates the magnitude of the MVD, and the motion direction indicates the direction of the MVD.
[0081] In one example, a merge candidate selected from the merge candidate list is used to provide a starting point or MV starting point in a reference picture. The motion vector of the current block can be expressed by a starting point and a motion offset (or MVD) including the magnitude and direction of motion relative to the starting point. At the encoder side, the selection of the merge candidate and the determination of the motion offset can be based on a search process (evaluation process) as shown in Figure 8. At the decoder side, the selected merge candidate and the motion offset can be determined based on signaling from the encoder side.
[0082] Figure 8 shows an example of a search process (800) in MMVD mode. Figure 9 shows example search points in MMVD mode. In some examples, a subset or the entire set of search points in Figure 9 are used in the search process (800) of Figure 9. For example, by performing the search process (800) on the encoder side, additional information can be determined for a current block (801) in a current picture (or current frame), including a merge candidate flag (e.g., mmvd_cand_flag[x0][y0]), a distance index (e.g., mmvd_distance_idx[x0][y0]), and a direction index (e.g., mmvd_direction_idx[x0][y0]).
[0083] A first motion vector (811) and a second motion vector (821) belonging to a first merging candidate are shown. The first motion vector (811) and the second motion vector (821) are MV starting points used in the search process (800). The first merging candidate may be a merging candidate on the merging candidate list constructed for the current block (801). The first motion vector (811) and the second motion vector (821) may be associated with two reference pictures (802) and (803) in the reference picture lists L0 and L1, respectively. Referring to Figures 8 to 9, the first motion vector (811) and the second motion vector (821) may point to two starting points (911) and (921) in the reference pictures (802) and (803), respectively, as shown in Figure 9.
[0084] Referring to Figure 9, two starting points (911) and (921) in Figure 9 can be determined in the reference pictures (802) and (803). In one example, based on the starting points (911) and (921), multiple predetermined points can be evaluated within the reference pictures (802) and (803) extending vertically (represented by +Y or -Y) or horizontally (represented by +X and -X) from the starting points (911) and (921). In one example, pairs of points that mirror each other with respect to their respective starting points 911 or 921, such as the pair of points 914 and 924 (e.g., as indicated by a shift of 1S in FIG. 8) or the pair of points 915 and 925 (e.g., as indicated by a shift of 2S in FIG. 8), can be used to determine pairs of motion vectors (e.g., MVs 813 and 823 in FIG. 8) that can form motion vector predictor candidates for the current block 801. The motion vector predictor candidates (e.g., MVs 813 and 823 in FIG. 8) determined based on predetermined points around the starting point 911 or 921 can be evaluated.
[0085] The distance index (e.g., mmvd_distance_idx[x0][y0]) specifies motion magnitude information and may indicate a predetermined offset (e.g., 1S or 2S in FIG. 8) from the starting point indicated by the merge candidate flag. Note that in one example, the predetermined offset is also referred to as the MMVD step.
[0086] Referring to Figure 8, an offset (e.g., MVD(812) or MVD(822)) can be applied (e.g., added) to the horizontal or vertical component of the starting MV (e.g., MV(811) or (821)). An example relationship between the distance index (IDX) and the predetermined offset is specified in Table 2. When full-pel MMVD is off, e.g., when a full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 0, the predetermined offset for MMVD can range from 1 / 4 luma sample to 32 luma samples. When full-pel MMVD is off, the predetermined offset can have a non-integer value, such as a fraction of a luma sample (e.g., 1 / 4 pixel or 1 / 2 pixel). When full-pel MMVD is on, for example, when a full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 1, the predetermined offset of the MMVD can range from 1 luma sample to 128 luma samples. In one example, when full-pel MMVD is on, the predetermined offset has only integer values, such as one or more luma samples. [Table 2]
[0087] The direction index can represent the direction of the MVD (or the direction of motion) relative to the starting point. In one example, the direction index represents one of the four directions shown in Table 3. The meaning of the MVD code in Table 3 may change according to the information of the starting MV. In one example, when the starting MV is a uni-predictive MV or a bi-predictive MV and both reference lists point to the same side of the current picture (for example, when the POCs of the two reference pictures are both greater than the POC of the current picture, or when the POCs of the two reference pictures are both less than the POC of the current picture), the MVD code in Table 3 specifies the code of the MV offset (or MVD) to be added to the starting MV.
[0088] When the starting MV is a bi-predictive MV and the two MVs point to different sides of the current picture (e.g., the POC of one reference picture is greater than that of the current picture and the POC of the other reference picture is less than that of the current picture), the MVD code in Table 3 specifies the sign of the MV offset (or MVD) added to the MV component of list 0 of the starting MV, and the MVD code for the MV of list 1 has the opposite value. Referring to Figure 8, the starting MVs (811) and (821) are bi-predictive MVs and the two MVs (811) and (821) point to different sides of the current picture. The POC of the L1 reference picture (803) is greater than that of the current picture, and the POC of the L0 reference picture (802) is less than that of the current picture. The MVD sign (e.g., the x-axis sign "+") indicated by the direction index (e.g., 00) in Table 2 specifies the sign (e.g., the x-axis sign "+") of the MVD (e.g., MVD(812)) that is added to the MV components of list 0 of the starting MV (e.g., (811)), and the MVD sign of MVD(822) for the MV components of list 1 of the starting MV (e.g., (821)) has the opposite value, such as the sign "-" opposite to the sign "+" of MVD(812).
[0089] Referring to Table 3, direction index 00 indicates the positive direction of the x-axis, direction index 01 indicates the negative direction of the x-axis, direction index 10 indicates the positive direction of the y-axis, and direction index 11 indicates the negative direction of the y-axis. [Table 3]
[0090] The syntax element mmvd_merge_flag[x0][y0] can be used to represent the MMVD merge flag of the current CU. In one example, an MMVD merge flag equal to 1 (e.g., mmvd_merge_flag[x0][y0]) specifies that the MMVD mode is used to generate inter prediction parameters for the current CU. An MMVD merge flag equal to 0 (e.g., mmvd_merge_flag[x0][y0]) specifies that the MMVD mode is not used to generate inter prediction parameters. The array indexes x0 and y0 can specify the position (x0, y0) of the top-left luma sample of the considered coding block (e.g., the current CB) relative to the top-left luma sample of the picture (e.g., the current picture).
[0091] When the MMVD merge flag (eg, mmvd_merge_flag[x0][y0]) is not present for the current CU, the MMVD merge flag (eg, mmvd_merge_flag[x0][y0]) can be inferred to be equal to 0 for the current CU.
[0092] In some examples, the MMVD flag is signaled immediately after sending the skip and merge flags. If the skip and merge flags are true, the MMVD flag is parsed. If the MMVD flag is equal to 1, the MMVD syntax is parsed. However, if the MMVD flag is not 1, in some examples, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, the AFFINE mode is used for decoding. However, if the AFFINE flag is not 1, in some examples, the skip / merge index is parsed for the VTM's skip / merge mode.
[0093] According to one aspect of the present disclosure, template matching based candidate reordering can be performed for MMVD and affine MMVD.
[0094] In some examples, the MMVD offset is extended to more positions for MMVD and affine MMVD modes.
[0095] Figure 10 shows a diagram illustrating the directions in which refinement directions can be added for MMVD, where further refinement positions are added along the k x π / 8 diagonal, where k is an integer. Position 1001 corresponds to the base candidate and can be the starting point, and positions 1011 through 1014 are in the directions of 0, π / 2, π, and 3π / 2, respectively. Additional directions are added. For example, positions 1021 through 1024 are in the directions of π / 4, 3π / 4, 5π / 4, and 7π / 4, respectively, and positions 1031 through 1038 are in the directions of π / 8, 3π / 8, 5π / 8, 7π / 8, 9π / 8, 11π / 8, 13π / 8, and 15π / 8, respectively. Thus, the number of directions increases from 4 to 16. Furthermore, in one example, each direction can have 6 MMVD refinement positions. The total number of possible MMVD refinement positions is 16 x 6.
[0096] According to one aspect of the present disclosure, the SAD cost between the current template (e.g., one row above and one column to the left of the current block) and the reference template can be calculated for each refinement position. Based on the refinement position's SAD cost, all possible MMVD refinement positions (16x6) for each base candidate are sorted. Then, the upper part of the refinement positions, such as the top 1 / 8 refinement positions (e.g., 12), such as those with the minimum template SAD cost, are retained as available positions for the resulting MMVD index coding. The MMVD index is binarized by a Rice code with a parameter equal to 2.
[0097] In some examples, the refinement positions for affine MMVD can be increased, and template matching-based candidate reordering can be applied to the affine MMVD reordering. For example, the affine MMVD refinement positions are along the k × π / 4 diagonal, such as eight directions: 0, π / 4, π / 2, 3π / 4, π, 5π / 4, 3π / 2, and 7π / 4. Each direction can have six affine MMVD refinement positions. The total number of possible affine MMVD refinement positions is 8 × 6. In one example, the SAD cost between the current template (e.g., one row above and one column to the left of the current block) and the reference template can be calculated for each refinement position. Based on the refinement position's SAD cost, all possible affine MMVD refinement positions (8 × 6) for each base candidate are reordered. Then, the upper portion of the refinement positions, such as the top half of the refinement positions (e.g., 24), such as those with the smallest template SAD cost, are retained as available positions for the resulting affine MMVD index coding.
[0098] To improve coding efficiency and reduce MV transmission overhead, subblock-level MV refinement can be applied to extend CU-level temporal motion vector prediction (TMVP). In one example, subblock-based TMVP (SbTMVP) mode enables subblock-level motion information inheritance from co-located reference pictures. Each subblock of a current CU (e.g., a current CU with a large size) in a current picture can have its own motion information without explicitly transmitting a block partition structure or its own motion information. In SbTMVP mode, the motion information of each subblock can be obtained, for example, in three steps as follows: In the first step, the displacement vector (DV) of the current CU can be derived. In the second step, the availability of SbTMVP candidates can be checked, and the central motion (e.g., the central motion of the current CU) can be derived. In the third step, subblock motion information can be derived from the corresponding subblock in the co-located block using the DV. The three steps may be combined into one or two steps and / or the order of the three steps may be adjusted.
[0099] Unlike TMVP candidate derivation, which derives temporal MVs from co-located blocks in a reference frame or picture, in SbTMVP mode, DVs (e.g., DVs derived from the MVs of the current CU's left-neighboring CUs) can be applied to locate corresponding sub-blocks in the co-located picture for each sub-block in the current CU in the current picture. If the corresponding sub-block is not inter-coded, the motion information of the current sub-block can be set to the central motion of the co-located block.
[0100] The SbTMVP mode can be supported by various video coding standards, including, for example, VVC. Similar to the TMVP mode, for example, in HEVC, in the SbTMVP mode, a motion field (also called a motion information field or MV field) in a co-located picture can be used to improve MV prediction and merging for a CU in a current picture. In one example, the same co-located picture used by the TMVP mode is used in the SbTVMP mode. In one example, the SbTMVP mode differs from the TMVP mode in the following aspects: (i) the TMVP mode predicts motion information at the CU level, while the SbTMVP mode predicts motion information at the sub-CU level; and (ii) the TMVP mode retrieves temporal MV from a co-located block in the co-located picture (e.g., the co-located block is the bottom-right or center block relative to the current CU), and the SbTMVP mode can apply a motion shift before retrieving temporal motion information from the co-located picture. In one example, the motion shift used in the SbTMVP mode is obtained from the MV of one of the spatially neighboring blocks of the current CU.
[0101] 11-12 show an exemplary SbTMVP process used in SbTMVP mode. The SbTMVP process can predict the motion vector (MV) of a sub-CU (e.g., a sub-block) in a current CU (e.g., a current block) (1101) in a current picture (1211), for example, in two steps. In the first step, the spatial neighbors (e.g., A1) of the current block (1101) in FIGS. 11-12 are examined. If the spatial neighbor (e.g., A1) has a motion vector (MV) (1221) that uses the co-located picture (1212) as its reference picture, the motion vector (MV) (1221) can be selected to be the motion shift (or DV) to be applied to the current block (1101). If no such motion vector (e.g., a MV that uses the co-located picture (1212) as its reference picture) is identified, the motion shift or DV can be set to a zero motion vector (e.g., (0, 0)). In some examples, if no such MV is identified for spatial neighbor A1, MVs in further spatial neighbors such as A0, B0, B1, etc. are examined.
[0102] In a second step, the motion shift or DV (1221) identified in the first step is applied to the current block (1101) (e.g., the DV (1221) is added to the coordinates of the current block) to obtain sub-CU level motion information (e.g., including MV and reference index) from the co-located picture (1212). In the example shown in FIG. 12, the motion shift or DV (1221) is set to be the MV of the spatial neighbor A1 (e.g., block A1) of the current block (1101). For each sub-CU or sub-block (1231) in the current block (1101), the motion information of the corresponding co-located block (1201) in the co-located picture (1212) (e.g., the motion information of the minimum motion grid covering the center sample of the co-located block (1201)) can be used to derive the motion information of the sub-CU or sub-block (1231). After the motion information of the co-located sub-CU (1232) in the co-located block (1201) is identified, the motion information of the co-located sub-CU (1232) can be converted into motion information (e.g., MV and one or more reference indices) of the current sub-CU (1231) using a scaling method, such as a method similar to the TMVP process used in HEVC, and temporal motion scaling is applied to align the reference picture of the temporal MV with the reference picture of the current CU.
[0103] The motion field of the current block (1101) derived based on the DV (1221) can include motion information of each sub-block (1231) within the current block (1101), such as the motion vector (MV) and one or more associated reference indexes. The motion field of the current block (1101) is also called an SbTMVP candidate and corresponds to the DV (1221).
[0104] 12 shows an example of a motion field or SbTMVP candidate for the current block 1101. The motion information of the bi-predicted sub-block (1231(1)) includes a first motion vector (MV), a first index indicating a first reference picture in reference picture list 0 (L0), a second motion vector (MV), and a second index indicating a second reference picture in reference picture list 1 (L1). In one example, the motion information of the uni-predicted sub-block (1231(2)) includes a motion vector (MV) and an index indicating a reference picture in L0 or L1.
[0105] In one example, the DV (1221) is applied to the center position of the current block (1101) to locate the displaced center position in the co-located picture (1212). If the block containing the displaced center position is not inter-coded, an SbTMVP candidate is considered unavailable. Alternatively, if the block containing the displaced center position (e.g., the co-located block (1201)) is inter-coded, motion information of the center position of the current block (1101), referred to as the central motion of the current block (1101), can be derived from motion information of the block containing the displaced center position in the co-located picture (1212). In one example, a scaling process can be used to derive the central motion of the current block (1101) from the motion information of the block containing the displaced center position in the co-located picture (1212). When SbTMVP candidates are available, the DV (1221) can be applied to find the corresponding sub-block (1231) in the co-located picture (1212) for each sub-block (1232) of the current block (1101). The motion information of the corresponding sub-block (1232) can be used to derive motion information for the sub-block (1231) in the current block (1101), such as in the same manner as used to derive the central motion of the current block (1101). In one example, if the corresponding sub-block (1232) is not inter-coded, the motion information of the current sub-block (1231) is set to be the central motion of the current block (1101).
[0106] In some examples (e.g., VVC and ECM), sub-block-based motion modes (also called sub-block-based inter prediction modes), such as SbTMVP mode, affine mode, etc., code (encode / decode) all sub-blocks in a coding block using the sub-block-based motion mode with the inter prediction mode. According to one aspect of the present disclosure, due to the dynamic characteristics of video content, not all sub-blocks in a coding block necessarily follow a consistent motion model, and some sub-blocks may be better coded using intra prediction modes, especially when different sub-blocks belong to different objects.
[0107] Some aspects of this disclosure provide techniques that allow one or more sub-blocks within a coding block coded in a sub-block-based motion mode to be coded in an intra-prediction mode.
[0108] According to one aspect of the present disclosure, when a coding block is coded using a sub-block-based inter-prediction mode (e.g., SbTMVP mode, affine mode, etc.), one or more of the sub-blocks within the coding block can be coded using an intra-prediction mode.
[0109] Figure 13 shows a diagram of a coding block (1300) according to some embodiments of the present disclosure. The coding block (1300) is coded in a sub-block-based inter prediction mode. The coding block (1300) includes a plurality of sub-blocks, such as 16 sub-blocks as shown in Figure 13. Of the 16 sub-blocks, 13 sub-blocks are coded in the inter prediction mode, for example, based on the motion vectors indicated by the arrows in Figure 13, and three sub-blocks (1301) to (1303) are coded in the intra prediction mode.
[0110] It is noted that the sub-block-based inter prediction mode can be any suitable sub-block-based inter prediction mode, such as affine mode, regression-based inter prediction mode, SbTMVP mode, etc. In some examples, the regression-based inter prediction mode can derive a motion vector for a sub-block within a coding block using motion vectors of neighboring sub-blocks based on linear regression.
[0111] In some embodiments, to determine whether sub-blocks in a coding block are coded in an intra-prediction mode, a first flag may be signaled or implicitly derived to indicate whether at least one sub-block in the coding block is coded in an intra-prediction mode. For example, when the first flag is 0, all of the sub-blocks in the coding block are coded in an inter-prediction mode, and when the first flag is 1, at least one sub-block in the coding block is coded in an intra-prediction mode.
[0112] In one embodiment, a second flag is signaled for each sub-block to indicate whether the associated sub-block is coded by an intra-prediction mode.
[0113] Note that the first flag can be derived implicitly without signaling in some examples. In one example, when a coding block is coded in SbTMVP mode and there is a sub-block whose corresponding co-located block does not have a valid motion vector, the first flag is implicitly set to 1. In another example, when a coding block is coded in SbTMVP mode and the corresponding co-located blocks of all sub-blocks in the coding block have valid motion vectors, the first flag is implicitly set as 0.
[0114] In one embodiment, the second flag is conditionally signaled for each sub-block, for example, when the sub-block does not have an associated motion vector in the motion vector field of the reference picture, the second flag is signaled to indicate whether the sub-block is coded by an intra-prediction mode.
[0115] According to one aspect of the present invention, when a coding block is coded in a sub-block-based inter-prediction mode and at least one sub-block in the coding block is coded in an intra-prediction mode, the sub-block coded in the inter-prediction mode is reconstructed first, and then the remaining sub-blocks coded in the intra-prediction mode are reconstructed according to a specific coding order.
[0116] In some examples, the particular coding order is determined by the position of the blocks coded by the inter prediction mode.
[0117] In some examples, for a current sub-block, neighboring intra-coded sub-blocks reconstructed before the current sub-block, along with all neighboring inter-coded sub-blocks, can be used as reference samples for the current sub-block that is intra-coded.
[0118] 14 shows a diagram of a coding block (1400) according to some embodiments of this disclosure. The coding block (1400) is coded in a sub-block-based inter prediction mode and includes an intra-coded sub-block (1401). The sub-block (1401) is surrounded by other inter-coded sub-blocks (as indicated by the gray sub-blocks). Reconstruction of the gray sub-blocks can be used to predict the sub-block (1401), which is an intra-coded sub-block.
[0119] In some examples, an at_least_one_intra_flag (also referred to as a first flag in some examples) associated with a coding block is signaled to indicate whether at least one sub-block in the coding block is coded using an intra-prediction mode. When at_least_one_intra_flag is false, no sub-block in the coding block is coded using intra-prediction, and the transform unit (TU) size is equal to the coding unit (CU) size (coding block size). Otherwise (when at_least_one_intra_flag is true), the TU size is the same as the sub-block size. Then, the first one or more sub-blocks (in the coding block) coded using an inter-prediction mode are first reconstructed in raster scan order, and then the second one or more sub-blocks (in the coding block) coded using an intra-prediction mode are reconstructed in raster scan order.
[0120] According to one aspect of the present disclosure, when at least one sub-block is coded by an intra-prediction mode, the residual samples of all sub-blocks coded by an inter-prediction mode are put together and a transform is applied thereon, and some residuals are filled in to apply a forward transform and an inverse transform for the intra-coded sub-block positions. In one example, any suitable value can be used to fill in as the residual at the position of the intra-coded sub-block to form a transform unit.
[0121] According to one aspect of the present disclosure, when at least one sub-block is coded using an intra-prediction mode, each sub-block that is intra-coded, predicted, transformed, and residual coded is performed at the sub-block level.
[0122] According to one aspect of the present disclosure, when a sub-block is coded using intra-prediction modes, the available intra-prediction modes may be selected from a limited set, such as a most probable mode (MPM), or a selected smooth mode or non-directional mode. Smooth modes, such as smooth, smooth vertical, and smooth horizontal modes, can generate a very smooth surface constructed by filtered samples using linear interpolation.
[0123] According to one aspect of the present disclosure, when a sub-block is coded using an intra-prediction mode, the intra-prediction mode used for sub-block prediction can be expanded to use more available reference samples in addition to the row above and the column to the left. In some examples, one or more right columns of available reference samples can be used in intra-prediction for the sub-block. In some examples, one or more bottom rows of available reference samples can be used in intra-prediction for the sub-block. In some examples, corner samples adjacent to the sub-block can be used in intra-prediction of the sub-block. In some examples, multiple rows and / or multiple columns of reference samples may be used.
[0124] 15 shows a diagram of using intra prediction for a sub-block (1501) according to some embodiments of this disclosure. The sub-block (1501) is one of the sub-blocks in a coding block (not shown), where the coding block is coded in a sub-block-based inter prediction mode and the sub-block (1501) is coded in an intra prediction mode.
[0125] In some examples, the intra-prediction mode used for sub-block prediction of the sub-block (1501) can be expanded to use more available reference samples in addition to the row above (e.g., the row immediately above among the rows above (1510)) and the column to the left (e.g., the column immediately to the left among the columns to the left (1530)). In one example, multiple upper rows of available reference samples (1510) can be used in intra-prediction of the sub-block (1501). In another example, multiple upper rows of available reference samples (1510) can be used in intra-prediction of the sub-block (1501). In another example, one or more right columns of right columns of available reference samples (1540) can be used in intra-prediction for the sub-block (1501). In another example, one or more lower rows of available lower rows of reference samples (1520) can be used in intra-prediction of the sub-block (1501). In another example, corner samples adjacent to the subblock (1501) may be used in intra-prediction of the subblock (1501), such as reference samples of one or more available corners among corners (1551) to (1554).
[0126] 16 shows a diagram of using intra prediction for a sub-block (1601) according to some embodiments of this disclosure. The sub-block (1601) is one of the sub-blocks in a coding block (not shown), where the coding block is coded in a sub-block-based inter prediction mode and the sub-block (1601) is coded in an intra prediction mode.
[0127] Similar to the example shown in FIG. 16, the intra-prediction mode used for sub-block prediction of sub-block (1601) can be expanded to use more available reference samples in addition to the row above (e.g., the row immediately above among the rows above (1610)) and the column to the left (e.g., the column immediately to the left among the columns to the left (1630)). In some examples, one or more columns in the bottom right column (1680) of available reference samples can be used in intra-prediction of sub-block (1601). In some examples, one or more rows in the top right row (1660) of available reference samples can be used in intra-prediction of sub-block (1601).
[0128] In some examples, when a sub-block within a coding block in a sub-block-based inter prediction mode is coded using an intra prediction mode, both the luma component and the chroma component of the sub-block are coded using the intra prediction mode.
[0129] In some examples, when at least one sub-block in a coding block of a sub-block-based inter prediction mode is coded by an intra prediction mode, other sub-blocks coded by an inter prediction mode may have restrictions on inter prediction to reduce computational complexity. In one example, to reduce computational complexity, inter prediction is limited to uni-prediction.
[0130] According to one aspect of the present disclosure, one or more predictors for a sub-block coded in an intra-prediction mode (e.g., for predicting the intra-prediction mode) are determined from neighboring pixels of a current coding unit (CU), such as in the intra-prediction process in VVC and ECM reference software. Supported intra-prediction modes include, but are not limited to, most probable mode (MPM), decoder-side intra-mode derivation (DIMD), template-based intra-mode derivation (TIMD), directional intra-prediction mode, etc. In some examples, a predictor for a sub-block coded in an intra-prediction mode (e.g., a first sub-block) can be decoded in parallel with a second sub-block coded in an inter-prediction mode without sub-block-level data dependency. The decoding of the second sub-block is independent of the first sub-block, and the decoding of the predictor for the first sub-block is independent of the second sub-block.
[0131] 17 shows a flowchart outlining a process (1700) according to one embodiment of the present disclosure. The process (1700) can be used in a video encoder. In various embodiments, the process (1700) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (1700) is implemented with software instructions, and thus, the processing circuit performs the process (1700) when it executes the software instructions. The process begins at (S1701) and proceeds to (S1710).
[0132] At (S1710), a current coding unit (CU) in a picture is determined for coding in a sub-block-based inter prediction mode.
[0133] At (S1720), one or more first sub-blocks in the current CU are determined for coding by intra prediction.
[0134] At (S1730), the current CU is coded into a bitstream, one or more second sub-blocks of the current CU are coded by inter prediction, and one or more first sub-blocks of the current CU are coded by intra prediction.
[0135] In some examples, the subblock-based inter prediction mode may be any suitable inter prediction mode for encoding CUs at the subblock level. In one example, the subblock-based inter prediction mode is an affine mode. In another example, the subblock-based inter prediction mode is a regression-based inter prediction mode. In another example, the subblock-based inter prediction mode is a subblock-based temporal motion vector prediction (SbTMVP) mode.
[0136] In some examples, a first flag is coded into the bitstream, and the first flag indicates that at least one sub-block in the current CU is coded by intra-prediction mode. Further, a second flag is coded into the bitstream, and the second flag is associated with each sub-block in the current CU. The second flag associated with the sub-block in the current CU indicates whether the sub-block is coded by intra-prediction.
[0137] In some examples, the first flag is not coded into the bitstream. In one example, the first flag is derived as a true value in response to determining that a co-located block for a sub-block in the current CU does not have a valid motion vector when the current CU is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode. In another example, the first flag is derived as a false value in response to determining that each sub-block in the current CU has a valid motion vector when the current CU is coded in an SbTMVP mode.
[0138] In some examples, one or more second sub-blocks of the current CU are reconstructed before one or more first sub-blocks of the current CU, and the one or more first sub-blocks of the current CU are reconstructed according to a specific order determined based on the positions of the one or more second sub-blocks. In one example, the one or more second sub-blocks of the current CU are reconstructed according to a raster scan order, and the one or more first sub-blocks of the current CU are reconstructed according to a raster scan order after the one or more second sub-blocks are reconstructed. In some examples, for a first sub-block in one or more first sub-blocks, the first sub-block is reconstructed according to adjacent inter-coding sub-blocks and adjacent intra-coding sub-blocks in the current CU.
[0139] In some examples, the transform unit size of the current CU is determined to be the size of the sub-block in response to a first flag being true, the first flag indicating whether at least one sub-block in the current CU is coded using an intra-prediction mode. In another example, the transform unit size is determined to be the size of the current CU in response to the first flag being false.
[0140] In some examples, when a sub-block in a current CU is coded by intra prediction, the residual samples of all sub-blocks coded by inter prediction mode (e.g., the second one or more sub-blocks) and the padded residual samples of the first one or more sub-blocks are combined and a transform is applied thereon. For the first one or more sub-blocks that are intra-coded, any suitable residual value can be used for padding to apply the transform.
[0141] In some examples, when at least one sub-block is coded using an intra-prediction mode, each sub-block that is intra-coded, predicted, transformed, and residual coded is performed at the sub-block level.
[0142] In some examples, a first sub-block within one or more first sub-blocks is reconstructed based on an intra-prediction mode selected from at least one of a set of most probable modes as a selected set of smooth modes or a set of non-directional modes.
[0143] In some examples, the first sub-blocks within the one or more first sub-blocks are rearranged according to at least one of one or more right columns, one or more bottom rows, one or more corner samples, one or more top right rows, or one or more bottom left columns.
[0144] In one example, for an intra-coded sub-block within the one or more first sub-blocks by intra prediction, both a luma component and at least one chroma component of the intra-coded sub-block are coded by intra prediction.
[0145] In some examples, a constraint is applied to inter prediction for reconstructing one or more second sub-blocks in response to one or more first sub-blocks in the current CU being coded by intra prediction. In one example, uni-prediction is used for inter prediction, and bi-prediction is not available.
[0146] In some examples, a predictor of an intra-prediction mode for intra-prediction is determined from neighboring pixels of the current CU. The intra-prediction mode can be one of a most probable mode (MPM), a decoder-side intra mode derivation (DIMD), a template-based intra mode derivation (TIMD), or a directional intra-prediction mode.
[0147] The process then proceeds to (S1799) and ends.
[0148] The process 1700 may be adapted as appropriate. Steps in the process 1700 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0149] 18 shows a flowchart outlining a process (1800) according to one embodiment of the present disclosure. The process (1800) can be used in a video decoder. In various embodiments, the process (1800) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), etc. In some embodiments, the process (1800) is implemented with software instructions, and thus, the processing circuit performs the process (1800) when it executes the software instructions. The process begins at (S1801) and proceeds to (S1810).
[0150] At (S1810), a bitstream carrying at least a picture is received.
[0151] At (S1820), it is determined that a current coding unit (CU) in a picture is coded in a sub-block-based inter prediction mode.
[0152] At (S1830), it is determined that one or more first sub-blocks in the current CU are coded by intra prediction.
[0153] At (S1840), one or more second sub-blocks of the current CU are reconstructed based on inter prediction.
[0154] At (S1850), one or more first sub-blocks of the current CU are reconstructed based on intra prediction.
[0155] In some examples, the subblock-based inter prediction mode includes at least one of an affine mode, a regression-based inter prediction mode, or a subblock-based temporal motion vector prediction (SbTMVP) mode.
[0156] In some examples, a first flag indicating whether at least one sub-block in a current CU is coded by intra prediction is determined, and then second flags associated with the sub-blocks in the current CU are decoded from the bitstream, and the second flags associated with the sub-blocks in the current CU indicate whether the sub-blocks are coded by intra prediction. In one example, the first flag is decoded from the bitstream.
[0157] In some examples, the first flag is derived without signaling. In one example, the first flag is derived as a true value in response to determining that a co-located block for a sub-block in the current CU does not have a valid motion vector when the current CU is coded in sub-block-based temporal motion vector prediction (SbTMVP) mode. In another example, the first flag is derived as a false value in response to determining that each sub-block in the current CU has a valid motion vector when the current CU is coded in SbTMVP mode.
[0158] In some examples, one or more second sub-blocks of the current CU are reconstructed before one or more first sub-blocks of the current CU, and the one or more first sub-blocks of the current CU are reconstructed according to a specific order determined based on the positions of the one or more second sub-blocks.
[0159] In one example, one or more second sub-blocks of the current CU are reconstructed according to a raster scan order, and one or more first sub-blocks of the current CU are reconstructed according to a raster scan order after the one or more second sub-blocks are reconstructed.
[0160] In some examples, for a first sub-block in the one or more first sub-blocks, the first sub-block is predicted according to neighboring inter-coding sub-blocks and neighboring intra-coding sub-blocks in the current CU.
[0161] In one example, the transform unit size is determined to be the size of the sub-block in response to a first flag being true, the first flag indicating whether at least one sub-block in the current CU is coded by intra prediction. In another example, the transform unit size is determined to be the size of the current CU in response to the first flag being false.
[0162] In some examples, an inverse transform is performed to obtain a transform unit of the size of the current CU, and residual values corresponding to one or more second sub-blocks are determined according to the transform unit. Then, the one or more second sub-blocks predicted based on inter prediction and the residual values are combined to reconstruct the one or more second sub-blocks.
[0163] In some examples, an inverse transform is performed to obtain one or more transform units of a size corresponding to one or more first sub-blocks in the current CU, where the one or more transform units correspond to one or more first sub-blocks, respectively. Residual values corresponding to the one or more first sub-blocks are determined according to the one or more transform units. Then, the one or more first sub-blocks predicted based on intra prediction and the residual values are combined to reconstruct the one or more first sub-blocks.
[0164] In some examples, a first sub-block within one or more first sub-blocks is reconstructed based on an intra-prediction mode selected from at least one of a set of most probable modes as a selected set of smooth modes or a set of non-directional modes.
[0165] In some examples, the first sub-blocks within the one or more first sub-blocks are rearranged according to at least one of one or more right columns, one or more bottom rows, one or more corner samples, one or more top right rows, or one or more bottom left columns.
[0166] In some examples, for each sub-block in the one or more first sub-blocks, both the luma and chroma components of the sub-block are reconstructed by intra prediction.
[0167] In some examples, inter prediction is determined according to which one or more first sub-blocks in the current CU are coded by intra prediction, along with constraints for reconstructing one or more second sub-blocks. In one example, uni-prediction is used for inter prediction.
[0168] In some examples, a predictor for an intra prediction mode is determined from neighboring pixels of the current CU, and the intra prediction mode includes at least one of a most probable mode (MPM), a decoder-side intra mode derivation (DIMD) mode, a template-based intra mode derivation (TIMD) mode, or a directional intra prediction mode.
[0169] The process then proceeds to (S1899) and ends.
[0170] The process 1800 may be adapted as appropriate. Steps in the process 1800 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0171] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 19 illustrates a computer system (1900) suitable for implementing certain embodiments of the disclosed subject matter.
[0172] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly or through interpretation, microcode execution, etc.
[0173] The instructions may be executed in various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0174] 19 for computer system (1900) are exemplary in nature and are not intended to suggest any limitation regarding the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components shown in the exemplary embodiment of computer system (1900).
[0175] The computer system 1900 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface input devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic vision).
[0176] The human interface input devices may include one or more of a keyboard (1901), a mouse (1902), a trackpad (1903), a touch screen (1910), a data glove (not shown), a joystick (1905), a microphone (1906), a scanner (1907), and a camera (1908) (only one of each is shown).
[0177] The computer system (1900) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1910), data gloves (not shown), or joystick (1905), although there may also be haptic feedback devices that do not function as input devices), audio output devices (e.g., speakers (1909), headphones (not shown), etc.), visual output devices (e.g., screens (1910), including CRT, LCD, plasma, and OLED screens, each with or without touchscreen input capability and each with or without haptic feedback capability, some of which may provide two-dimensional visual output or greater than three-dimensional output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0178] The computer system (1900) may also include human-accessible storage devices and their associated media, such as optical media or similar media (1921), including CD / DVD ROM / RW (1920) with CDs / DVDs, thumb drives (1922), removable hard drives or solid-state drives (193), legacy magnetic media such as tape and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0179] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0180] The computer system (1900) may also include interfaces (1954) to one or more communications networks (1955). Networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet and WLAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial TV; and vehicular and industrial networks including CANBus and the like. Certain networks typically require an external network interface adapter (e.g., a USB port on the computer system (1900)) attached to a particular general-purpose data port or peripheral bus (1949); others are typically integrated into the core of the computer system (1900) by attachment to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system), as described below. Using any of these networks, the computer system 1900 can communicate with other entities. Such communications can be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., from a particular CANbus to a particular CANbus device), or bidirectional to other computer systems, using, for example, local or wide-area digital networks. As noted above, specific protocols and protocol stacks can be used in each of these networks and network interfaces.
[0181] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces can be attached to the core (1940) of the computer system (1900).
[0182] The core (1940) may include one or more central processing units (CPUs) (1941), graphics processing units (GPUs) (1942), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1943), task-specific hardware accelerators (1944), graphics adapters (1950), etc. These devices may be connected through a system bus (1948), along with read-only memory (ROM) (1945), random access memory (RAM) (1946), and internal mass storage (1947), such as an internal non-user-accessible hard drive or SSD. In some computer systems, the system bus (1948) is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1948) or through a peripheral bus (1949). In one example, a screen (1910) may be connected to the graphics adapter (1950). Peripheral bus architectures include PCI, USB, and the like.
[0183] The CPU (1941), GPU (1942), FPGA (1943), and accelerator (1944) can execute specific instructions, which, in combination, can constitute the computer code described above. The computer code can be stored in ROM (1945) or RAM (1946). Temporary data can be stored in RAM (1946), while permanent data can be stored, for example, in internal mass storage (1947). High-speed storage and retrieval from any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more of the CPU (1941), GPU (1942), mass storage (1947), ROM (1945), RAM (1946), etc.
[0184] The computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0185] By way of example and not limitation, a computer system having the architecture (1900) and particularly the core (1940) can provide functionality as a result of the processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage, as introduced above, as well as media associated with the core's (1940) specific storage of a non-transitory nature, such as the core's internal mass storage (1947) or ROM (1945). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1940). The computer-readable media can include one or more memory devices or chips according to particular needs. The software can cause the core (1940) and particularly the processor therein (including a CPU, GPU, FPGA, etc.) to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM (1946) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerators (1944)) that may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that stores software for execution, circuitry embodying logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0186] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise various systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.
Claims
1. 1. A method for video decoding, comprising: receiving a coding bitstream carrying at least a picture including a block including a plurality of sub-blocks; determining, based on a value of a first syntax element in the coding bitstream, that a current coding unit (CU) in the picture is coded in a sub-block-based inter prediction mode; determining that one or more first sub-blocks in the current CU coded in the sub-block-based inter prediction mode are coded by intra prediction; reconstructing one or more second sub-blocks of the current CU by inter prediction based on the sub-block-based inter prediction mode, wherein the one or more second sub-blocks do not overlap with the one or more first sub-blocks in the current CU; reconstructing the one or more first sub-blocks of the current CU by intra prediction while the current CU is coded in the sub-block-based inter prediction mode; Including, The step of determining that the one or more first sub-blocks in the current CU are coded by intra prediction includes: determining a first flag indicating whether at least one sub-block in the current CU is coded by intra prediction; decoding, from the coding bitstream, second flags associated with the sub-blocks in the current CU, each second flag indicating whether the sub-block is coded by intra prediction; A method comprising:
2. The method of claim 1 , wherein the sub-block-based inter prediction mode includes at least one of an affine mode, a regression-based inter prediction mode, or a sub-block-based temporal motion vector prediction (SbTMVP) mode.
3. The step of determining the first flag includes: The method of claim 1 , further comprising the step of decoding the first flag from the coding bitstream.
4. The step of determining the first flag includes: deriving the first flag to be a true value in response to determining that a co-located block for a sub-block in the current CU does not have a valid motion vector when the current CU is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode; or deriving the first flag to be a false value in response to determining that each sub-block in the current CU has a valid motion vector when the current CU is coded in the SbTMVP mode; The method of claim 1 further comprising:
5. 2. The method of claim 1, wherein the one or more second sub-blocks of the current CU are reconstructed before the one or more first sub-blocks of the current CU, and the one or more first sub-blocks of the current CU are reconstructed according to a specific order determined based on the positions of the one or more second sub-blocks.
6. reconstructing the one or more second sub-blocks of the current CU according to a raster scan order; After the one or more second sub-blocks are reconstructed, reconstructing the one or more first sub-blocks of the current CU according to the raster scan order; The method of claim 1 further comprising:
7. 6. The method of claim 5, further comprising: for a first sub-block in the one or more first sub-blocks, reconstructing the first sub-block according to adjacent inter-coding sub-blocks and adjacent intra-coding sub-blocks in the current CU.
8. The method of claim 7, further comprising: determining, in response to the first flag being true, that a transform unit size is a size of a sub-block; determining, in response to the first flag being false, that the transform unit size is the size of the current CU; The method of claim 1 further comprising:
9. performing an inverse transform to obtain a transform unit of the size of the current CU; determining residual values corresponding to the one or more second sub-blocks according to the transform unit; The method of claim 1 further comprising:
10. performing an inverse transform to obtain one or more transform units of a size of a sub-block in the current CU, the one or more transform units corresponding to the one or more first sub-blocks, respectively; determining residual values corresponding to the one or more first sub-blocks according to the one or more transform units; The method of claim 1 further comprising:
11. The step of reconstructing the one or more first sub-blocks includes:
2. The method of claim 1, further comprising: reconstructing a first sub-block in the one or more first sub-blocks based on an intra-prediction mode selected from at least one of a set of most probable modes as a selected set of smooth modes or a set of non-directional modes.
12. The step of reconstructing the one or more first sub-blocks includes:
2. The method of claim 1, further comprising: rearranging first sub-blocks within the one or more first sub-blocks according to at least one of one or more right columns, one or more bottom rows, one or more corner samples, one or more top right rows, or one or more bottom left columns.
13. The step of reconstructing the one or more first sub-blocks includes: The method of claim 1 , further comprising reconstructing a luma component and at least one chroma component of a sub-block within the one or more first sub-blocks by the intra prediction.
14. The step of reconstructing the one or more first sub-blocks includes:
2. The method of claim 1, further comprising: determining the inter prediction together with a constraint for reconstructing the one or more second sub-blocks in accordance with the one or more first sub-blocks in the current CU being coded by the intra prediction.
15. The step of determining inter prediction includes: The method of claim 14 , further comprising determining to use uni-prediction for the inter-prediction.
16. The step of reconstructing the one or more first sub-blocks includes:
2. The method of claim 1, further comprising: determining a predictor of an intra-prediction mode from neighboring pixels of the current CU, wherein the intra-prediction mode includes at least one of a most probable mode (MPM), a decoder-side intra-mode derivation (DIMD) mode, a template-based intra-mode derivation (TIMD) mode, or a directional intra-prediction mode.
17. 1. A video processing apparatus including a processing circuit, 17. Apparatus, wherein the processing circuitry is configured to perform a method according to any one of claims 1 to 16.
18. A computer program causing a computer to carry out the method of any one of claims 1 to 16.
Citation Information
Patent Citations
A split-block encoding method for video encoding, a split-block decoding method for video decoding, and a recording medium for implementing the same.
JP2012518940A
Constrained block-level optimization and signaling for video coding tools
JP2019512965A
Method, device and computer program for decoding a video sequence
JP2021519050A
Video signal processing method and apparatus using current picture reference
JP2022513857A
Video decoding method, apparatus and computer program thereof
JP2022522398A