Method and apparatus for video coding

JP2025106484A5Pending Publication Date: 2026-03-17TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently managing motion vector prediction, particularly in sub-block-based scenarios, leading to sub-optimal compression ratios and increased data requirements.

Method used

The method involves determining a parameter based on decoded prediction information to set the maximum number of candidates in a sub-block-based merge candidate list, constrained by a flag indicating the enabled/disabled state of sub-block-based temporal motion vector prediction, allowing for more precise motion vector estimation.

Benefits of technology

This approach enhances video coding efficiency by optimizing motion vector prediction, reducing data requirements, and improving compression ratios in video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method for setting the range of values of the number of subblock-based merge candidates.SOLUTION: In an apparatus for video decoding including a receiving circuit and a processing circuit, the processing circuit determines a parameter on the basis of prediction information decoded from a coded video bitstream. The parameter is in a range that depends on a flag indicative of an enabled / disabled status of subblock-based temporal motion vector prediction. Then, the processing circuit calculates the maximum number of candidates in the subblock-based merge candidate lists on the basis of the parameter, and in response to a current block in a subblock-based prediction mode, reconstructs samples of the current block on the basis of a candidate selection from a constructed subblock-based merge candidate list of the current block. The constructed subblock-based merge candidate list of the current block is constrained by the maximum number.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Incorporation by Reference] This application claims the benefit of priority of U.S. Patent Application No. 17 / 217,595, filed Mar. 30, 2021, "METHOD AND APPARATUS FOR VIDEO CODING", which claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 005,511, filed Apr. 6, 2020, "METHOD OF SETTING NUMBER OF SUBBLOCK MERGING CANDIDATES". The entire disclosure of the prior application is hereby incorporated by reference in its entirety.

[0002] This disclosure describes embodiments generally related to video coding.

Background Art

[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The work of the present inventors is not to be recognized as prior art to the present disclosure, either explicitly or implicitly, to the extent that the work is described in this background art section and to the aspects of the description that may not otherwise be eligible as prior art at the time of filing.

[0004] Video coding and decoding can be performed using intra-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having, for example, spatial dimensions of 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have, for example, a fixed or variable picture rate (informally also called the frame rate) of 60 pictures per second or 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920x1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage area of more than 600 GB.

[0005] One purpose of video coding and decoding can be the reduction of redundancy in the input video signal by compression. Compression can, in some cases, help reduce the bandwidth or storage area requirements by more than two orders of magnitude. Both lossless compression and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to techniques that can reconstruct an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for its intended application. In the case of video, lossy compression is widely used. The amount of acceptable distortion depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect the following: higher acceptable / tolerable distortion can result in a higher compression ratio.

[0006] Motion compensation can be a Rossy compression technique and can be associated with a technique in which blocks of sample data from a previously reconstructed image or a part thereof (reference image) are spatially shifted in the direction indicated by a motion vector (hereinafter MV) and then used for prediction of a newly reconstructed image or a part thereof. In some cases, the reference image can be the same as the image being currently reconstructed. The MV can have two dimensions of X and Y, or three dimensions, where the third dimension is the display of the reference image in use (the latter can indirectly be the time dimension).

[0007] In some video compression techniques, the MV applicable to a certain region of sample data can be predicted from other MVs, for example, from an MV associated with another region of sample data that is spatially adjacent to the region being reconstructed and that precedes that MV in decode order. This can significantly reduce the amount of data required for coding the MV, thereby removing redundancy and increasing compression. MV prediction can effectively function because, for example, when coding an input video signal derived from a camera (known as natural video), there is a statistical likelihood that regions larger than the region to which a single MV is applied move in a similar direction, and thus, in some cases, a similar motion vector derived from the MVs of adjacent regions can be used for prediction. As a result, for a given region, an MV predicted from surrounding MVs is found to be similar or identical, and it can be represented with a smaller number of bits than would be used if the MV were directly coded after entropy coding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., the sample stream). In other cases, the MV prediction itself can be lossy, for example, due to rounding errors when calculating predictors from several surrounding MVs.

[0008] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, “High Efficiency Video Coding”, December 2016). Among the many MV prediction mechanisms provided by H.265, the one described in this specification is a technique hereinafter referred to as “spatial merge”.

[0009] Referring to FIG. 1, the current block (101) can be predicted from a previous block of the same size that has been spatially shifted, including samples found by the encoder during motion search processing. Instead of directly coding the MV, the MV can be derived from the latest (in decoding order) reference picture using an MV associated with any of five surrounding samples, for example, A0, A1, and B0, B1, B2 (102 to 106 respectively), which are associated with one or more reference pictures. In H.265, MV prediction can use a predictor from the same reference picture that adjacent blocks are using. SUMMARY OF THE INVENTION

[0010] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video decoding includes a receiving circuit and a processing circuit. For example, the processing circuit determines a parameter based on decoded prediction information from a coded video bitstream. The parameter is within a range that depends on a flag indicating an active / inactive state of sub-block-based temporal motion vector prediction. Next, the processing circuit calculates a maximum number of candidates in a sub-block-based merge candidate list based on the parameter, and reconstructs samples of the current block based on a candidate selection from the configured sub-block-based merge candidate list of the current block in response to the current block in a sub-block-based prediction mode. The configured sub-block-based merge candidate list of the current block is constrained by the maximum number of candidates in the sub-block-based merge candidate list.

[0011] In some examples, the processing circuit determines the maximum number of candidates in a sub-block-based merge candidate list by subtracting a parameter from a default number. In one example, the default number is 5.

[0012] In some embodiments, the upper limit of the range depends on a flag indicating the enabled / disabled state of sub-block-based temporal motion vector prediction.

[0013] In one example, the processing circuit receives a parameter signaled within a coded video bitstream. In another example, in response to the parameter not being signaled in the coded video bitstream, the processing circuit estimates the parameter based on a default number and a flag indicating the enabled / disabled state of sub-block-based temporal motion vector prediction.

[0014] In some examples, the flag indicates the enabled / disabled state of sub-block-based temporal motion vector prediction at the sequence parameter set (SPS) level.

[0015] In some embodiments, the parameter is within a range that depends on a first flag indicating the enabled / disabled state of sub-block-based temporal motion vector prediction at the sequence parameter set (SPS) level and a second flag indicating the enabled / disabled state of temporal motion vector prediction at the picture header (PH) level. In some examples, in response to the parameter not being signaled in the coded video bitstream, the processing circuit can estimate the parameter based on a default number, a first flag indicating the enabled / disabled state of sub-block-based temporal motion vector prediction at the SPS level, and a second flag indicating the enabled / disabled state of temporal motion vector prediction at the PH level.

[0016] Also, aspects of the present disclosure provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a method for video decoding.

Brief Description of the Drawings

[0017] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

[0018]

Figure 1

[0019]

Figure 2

[0020]

Figure 3

[0021]

Figure 4

[0022]

Figure 5

[0023]

Figure 6

[0024]

Figure 7

[0025]

Figure 8A

[0026]

Figure 8B

[0027]

Figure 9

[0028]

Figure 10

[0029]

Figure 11

[0030]

Figure 12

[0031]

Figure 13

[0032]

Figure 14

[0033]

Figure 15

[0034]

Figure 16

[0035]

Figure 17

[0036]

Figure 18

[0037]

Figure 19

Mode for Carrying Out the Invention

[0038] FIG. 2 shows a simplified block diagram of a communication system (200) according to an embodiment of the present disclosure. The communication system (200) includes a plurality of terminal devices that can communicate with each other via, for example, a network (250). For example, the communication system (200) includes a pair of first terminal devices (210) and (220) interconnected via a network (250). In the example of FIG. 2, the pair of first terminal devices (210) and (220) performs one-way data transmission. For example, the terminal device (210) can code video data (e.g., a stream of video images captured by the terminal device (210)) for transmission to another terminal device (220) via the network (250). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (220) can receive the coded video data from the network (250), decode the coded video data to restore a video image, and display the video image according to the restored video data. One-way data transmission is common in media providing applications and the like.

[0039] In another example, the communication system (200) includes a pair of second terminal devices (230) and (240) that perform bidirectional transmission of coded video data that may occur, for example, during a video conference. For bidirectional data transmission, in one example, each of the terminal devices (230) and (240) may code video data (e.g., a stream of video images captured by the terminal device) for transmission to the other terminal device of the terminal devices (230) and (240) via the network (250). Each of the terminal devices (230) and (240) may also receive the coded video data transmitted by the other terminal device of the terminal devices (230) and (240), decode the coded video data to restore a video image, and display the video image on an accessible display device according to the restored video data.

[0040] In the example of FIG. 2, the terminal devices (210), (220), (230) and (240) may be shown as a server, a personal computer and a smartphone, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure find application in laptop computers, tablet computers, media players and / or dedicated video conferencing devices. The network (250) represents any number of networks that transmit coded video data between the terminal devices (210), (220), (230) and (240), including, for example, wired and / or wireless communication networks. The communication network (250) may exchange data within a circuit-switched and / or packet-switched channel. Representative networks include communication networks, local area networks, wide area networks and / or the Internet. For the purposes of this description, the architecture and topology of the network (250) are not important for the operation of the present disclosure, unless otherwise described below.

[0041] FIG. 3 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter can be similarly applicable to other video-enabled applications including, for example, storage of compressed video on digital media including video conferencing, digital TV, CD, DVD, memory stick, etc.

[0042] A streaming system may include, for example, a video source (301) that generates a stream of uncompressed video images (302), and a capture subsystem (313) that may include, for example, a digital camera. In one example, the stream of video images (302) includes samples taken by a digital camera. The stream of video images (302), depicted as a thick line highlighting a high data volume when compared to encoded video data (304) (or a coded video bitstream), can be processed by an electronic device (320) that includes a video encoder (303) coupled to the video source (301). The video encoder (303) can include hardware / software, or a combination thereof, to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video data (304) (or encoded video bitstream (304)) is shown as a thin line to highlight a lower data volume when compared to the stream of video images (302) and can be stored at a streaming server (305) for future use. One or more streaming client subsystems, such as client subsystems (306) and (308) of FIG. 3, can access the streaming server (305) to retrieve copies (307) and (309) of the encoded video data (304). The client subsystem (306) can include, for example, a video decoder (310) within an electronic device (330). The video decoder (310) decodes an input copy (307) of the encoded video data and generates an output stream of video images (311) that can be rendered on a display (312) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (304), (307), and (309) (e.g., video bitstreams) can be encoded according to specific video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265.For example, the video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0043] Note that the electronic devices (320) and (330) can include other components (not shown). For example, the electronic device (320) can include a video decoder (not shown), and the electronic device (330) can also include a video encoder (not shown).

[0044] FIG. 4 shows a block diagram of a video decoder (410) according to an embodiment of the present disclosure. The video decoder (410) can be included in an electronic device (430). The electronic device (430) can include a receiver (431) (e.g., a receiving circuit). The video decoder (410) can be used instead of the video decoder (310) in the example of FIG. 3.

[0045] The receiver (431) may receive one or more coded video sequences to be decoded by the video decoder (410); in the same or another embodiment, it is one coded video sequence at a time, and the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequence may be received from a channel (401), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (431) may receive the encoded video data together with other data that may be transferred to respective using entities (not shown), such as coded audio data and / or auxiliary data streams. The receiver (431) may separate the coded video sequence from other data. To counter network jitter, a buffer memory (415) may be coupled between the receiver (431) and the entropy decoder / parser (420) (hereinafter, “parser (420)”). In certain applications, the buffer memory (415) is part of the video decoder (410). In others, it can be outside the video decoder (410) (not shown). In yet others, for example, to counter network jitter, there can be a buffer memory outside the video decoder (not shown), and, for example, to handle playback timing, another buffer memory (415) inside the video decoder (410). If the receiver (431) is receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (415) may be unnecessary or can be made small. For use in a best-effort packet network such as the Internet, the buffer memory (415) may be required, can be relatively large, can be advantageously sized adaptively, and can be implemented at least partially in an operating system or similar element outside the video decoder (410) (not shown).

[0046] The video decoder (410) may include a parser (420) for reconstructing symbols (421) from the coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (410) and potential information for controlling a rendering device (such as a display screen) (412) that is not an essential part of the electronic device (430) but can be coupled to the electronic device (430), as shown in FIG. 4. The control information for the rendering device(s) can be in the form of supplementary enhancement information (SEI message) or a video usability information (VUI) parameter set fragment (not shown). The parser (420) may perform syntax analysis / entropy decoding on the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (420) may extract a set of subgroup parameters for at least one of the subgroups of pixels within the video decoder based on at least one parameter corresponding to the group from the coded video sequence. The subgroups can include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (420) may also extract quantization parameter values, motion vectors, etc. from the coded video sequence information such as transform coefficients.

[0047] The parser (420) may perform entropy decoding / syntax analysis operations on the video sequence received from the buffer memory (415) to generate symbols (421).

[0048] The reconstruction of the symbol (421) can include a plurality of different units depending on the type of the coded video image or a portion thereof (e.g., inter and intra pictures, inter and intra blocks), and other factors. Which units are involved and how can be controlled by subgroup control information parsed from the coded video sequence by the parser (420). Such a flow of subgroup control information between the parser (420) and the following plurality of units is not shown for clarity.

[0049] In addition to the functional blocks already described, the video decoder (410) can conceptually be divided into several functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can at least partially be integrated with each other. However, for the purpose of explaining the disclosed subject matter, it is appropriate to conceptually subdivide it into the following functional units.

[0050] The first unit is the scaler / inverse transform unit (451). The scaler / inverse transform unit (451) receives the quantized transform coefficients as symbols (plural) (421) from the parser (420) together with control information including the transform to be used, block size, quantization coefficients, quantization scaling matrix, etc. The scaler / inverse transform unit (451) can output a block including sample values that can be input to the aggregator (455).

[0051] In some cases, the output samples of the scaler / inverse transform (451) can be associated with blocks that are intra-coded, i.e., blocks that do not use prediction information from a previously reconstructed image but can use prediction information from a previously reconstructed part of the current image. Such prediction information can be provided by the intra-image prediction unit (452). In some cases, the intra-image prediction unit 452 can use the already reconstructed surrounding information fetched from the current image buffer 458 to generate blocks of the same size and shape as the block being reconstructed. The current image buffer (458) buffers, for example, the partially reconstructed current image and / or the fully reconstructed current image. The aggregator (455) can, in some cases, add, sample by sample, the prediction information generated by the intra prediction unit (452) to the output sample information as provided by the scaler / inverse transform unit (451).

[0052] In other cases, the output samples of the scaler / inverse transform unit (451) can be related to inter-coded, potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (453) can access the reference image memory (457) to fetch the samples to be used for prediction. After motion-compensating the fetched samples according to the symbol (421) associated with the block, these samples can be added by the aggregator (455) to the output of the scaler / inverse transform unit (451) (in this case, called the residual samples or residual signal) to generate the output sample information. The address in the reference image memory (457) from which the motion compensation prediction unit (453) fetches the prediction samples can be controlled by the motion vectors available to the motion compensation prediction unit (453) in the form of, for example, a symbol (421) having X, Y, and a reference image component. Also, motion compensation can include interpolation of the sample values fetched from the reference image memory (457) when exact sub-sample motion vectors are used, a motion vector prediction mechanism, etc.

[0053] The output samples of the aggregator (455) can be subject to various loop filtering techniques within the loop filter unit (456). The video compression technique is controlled by parameters included in a coded video sequence (also referred to as a coded video bitstream), and is made available to the loop filter unit (456) as symbols (421) from the parser (420). However, it can respond to meta information obtained during the decoding of a previous (in decoding order) portion of the coded image or coded video sequence, and can also respond to previously reconstructed and loop filtered sample values, and can include in-loop filtering techniques.

[0054] The output of the loop filter unit (456) can be output to the rendering device (412), and can be a sample stream that can be stored in the reference image memory (457) for use in future inter-image prediction.

[0055] Once a particular coded image is fully reconstructed, it can be used as a reference image for future prediction. For example, when the coded image corresponding to the current image is fully reconstructed and the coded image is identified as a reference image (e.g., by the parser (420)), the current image buffer (458) can become part of the reference image memory (457), and a new current image buffer can be reallocated before starting the reconstruction of the next coded image.

[0056] The video decoder (410) may perform a decoding operation according to a predetermined video compression technique of a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile can select specific tools as the only tools available for use under that profile from all the tools available in the video compression technique or standard. Also, what is required for compliance may be that the complexity of the coded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits, for example, the maximum image size, the maximum frame rate, the maximum reconstruction sample rate (measured in megasamples per second), the maximum reference image size, etc. The limits set by the level can, in some cases, be further restricted through the virtual reference decoder (HRD) specification and the metadata for HRD buffer management signaled in the coded video sequence.

[0057] In one embodiment, the receiver (431) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (410) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.

[0058] FIG. 5 shows a block diagram of a video encoder (503) according to an embodiment of the present disclosure. The video encoder (503) is included in an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmission circuit). The video encoder (503) can be used instead of the video encoder (303) in the example of FIG. 3.

[0059] The video encoder (503) may receive video samples from a video source (501) (not part of the electronic device (520) in the example of FIG. 5) that can capture video image(s) to be coded by the video encoder (503). In another example, the video source (501) is part of the electronic device (520).

[0060] The video source (501) can provide a source video sequence to be coded by the video encoder (503) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media supply system, the video source (501) can be a storage device that stores pre-prepared video. In a video conferencing system, the video source (501) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual images that convey motion when viewed as a sequence. The images themselves can be configured as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0061] According to one embodiment, the video encoder (503) can code and compress the images of the source video sequence into a coded video sequence (543) in real time or under any other time constraints required by the application. Implementing an appropriate coding speed is one function of the controller (550). In some embodiments, the controller (550) controls other functional units and is functionally coupled to other functional units as described below. The couplings are not shown for clarity. The parameters set by the controller (550) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (550) can be configured to have other appropriate functions related to the video encoder (503) optimized for a particular system design.

[0062] In some embodiments, the video encoder (503) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop can include a source coder (530) (which is responsible for generating symbols such as a symbol stream based on an input image and a reference image to be coded), and a (local) decoder (533) embedded in the video encoder (503). The decoder (533) reconstructs the symbols to generate sample data (since any compression between the symbols and the coded video bitstream is lossless in the video compression techniques contemplated by the disclosed subject matter, as would be the case for a (remote) decoder). The reconstructed sample stream (sample data) is input into the reference image memory (534). Since the decoding of the symbol stream yields a bit-exact result that is independent of the decoder location (local or remote), the content in the reference image memory (534) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder "sees" the same sample values as reference image samples that the decoder "sees" when using prediction during decoding. This basic principle of reference image synchronization (and the drift that results, for example, if synchronization cannot be maintained due to channel errors) is also used in some related arts.

[0063] The operation of the "local" decoder (533) can be the same as that of a "remote" decoder such as the video decoder (410), which has already been described above in connection with FIG. 4. However, referring briefly to FIG. 4 as well, since the symbols are available and the encoding / decoding of the symbols to the coded video sequence by the entropy encoder (545) and the parser (420) can be lossless, the entropy decoding portion of the video decoder (410) including the buffer memory (415) and the parser (420) may not be fully implemented in the local decoder (533).

[0064] At this point, the observation that can be made is that any decoder technology other than the syntax analysis / entropy decoding existing within the decoder must also exist in substantially the same functional form within the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. Since the description of encoder technology is the reverse of the comprehensively described decoder technology, it can be omitted. Only in specific fields is a more detailed description necessary, which is provided below.

[0065] During operation, in some examples, the source coder (530) may perform motion-compensated predictive coding that predictively codes the input image with respect to one or more previously coded images from the video sequence designated as the "reference image". In this way, the coding engine (532) codes the difference between the pixel block of the input image and the pixel block of the reference image(s) that can be selected as the prediction reference(s) for the input image.

[0066] The local video decoder (533) may decode the coded video data of the image that can be designated as the reference image based on the symbols generated by the source coder (530). The operation of the coding engine (532) may advantageously be a lossy process. If the coded video data can be decoded by a video decoder (not shown in FIG. 5), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (533) replicates the decoding process that can be performed by the video decoder on the reference image and stores the reconstructed reference image in the reference image cache (534). In this way, the video encoder (503) may locally store a copy of the reconstructed reference image having common content as the reconstructed reference image obtained by the remote video decoder (in the absence of transmission errors).

[0067] Predictor (535) may perform predictive search for the coding engine (532). That is, for a new image to be coded, the predictor (535) may search the reference image memory (534) for specific metadata such as reference image motion vectors, block shapes, or sample data (as candidate reference pixel blocks) that may serve as appropriate predictive references for the new image. The predictor (535) may operate on a sample block-by-pixel block basis to find an appropriate predictive reference. In some cases, the input image may have a predictive reference drawn from a plurality of reference images stored in the reference image memory (534) as determined by the search results obtained by the predictor (535).

[0068] The controller (550) may manage the coding operations of the source coder (530), including, for example, setting parameters and subgroup parameters used to encode video data.

[0069] The outputs of all the aforementioned functional units may be subject to entropy coding in the entropy coder (545). The entropy coder (545) converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0070] The transmitter (540) can buffer video sequence(s) coded as generated by the entropy coder (545) and prepare for transmission via a communication channel (560), which can be a hardware / software link to a storage device storing the coded video data. The transmitter (540) can merge the coded video data from the video coder (503) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (not shown).

[0071] The controller (550) can manage the operation of the video encoder (503). During coding, the controller (550) can assign a specific coded image type to each coded image, which can affect the coding technique applicable to each image. For example, an image is often assigned as one of the following image types:

[0072] An intra picture (I picture) can be coded and decoded without using other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh (「IDR」) pictures. Those skilled in the art are aware of these variations of I pictures, as well as their respective uses and characteristics.

[0073] A predicted picture (P picture) can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0074] The bidirectional prediction image (B image) can be coded and decoded using intra prediction or inter prediction that uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple prediction images can use more than two reference images and associated metadata for the reconstruction of one block.

[0075] The source image is typically spatially divided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and coded block by block. The blocks can be predictively coded by referring to other (already coded) blocks as determined by the coding assignment applied to each image of the block. For example, blocks of an I image can be coded non-predictively, or they can be predictively coded by referring to already coded blocks of the same image (spatial prediction or intra prediction). Pixel blocks of a P image can be predictively coded via spatial prediction or via temporal prediction by referring to one previously coded reference image. Blocks of a B image can be predictively coded via spatial prediction or via temporal prediction by referring to one or two previously coded reference images.

[0076] The video encoder (503) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec. H.265. In its operation, the video encoder (503) can perform various compression operations including predictive coding operations that utilize the temporal and spatial redundancy in the input video sequence. Accordingly, the coded video data can conform to the syntax specified by the video coding technology or standard being used.

[0077] In one embodiment, the transmitter (540) may transmit additional data along with the coded video. The source coder (530) may include such data as part of a coded video sequence. The additional data may include, for example, temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0078] Video may be captured as a plurality of source images (video pictures) in a temporal sequence. Intra picture prediction (often abbreviated as intra prediction) exploits spatial correlations within a given picture, while inter picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture under encoding / decoding, called the current picture, is partitioned into blocks. If a block of the current picture is similar to a reference block of a reference picture of the video that has been previously coded and is still buffered, the block of the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block of the reference picture and can have three dimensions to identify the reference picture if multiple reference pictures are used.

[0079] In some embodiments, bi-prediction techniques can be used in inter picture prediction. According to the bi-prediction technique, two reference pictures such as a first reference picture and a second reference picture that both precede the current picture in the video in decoding order (however, in display order, they can be past and future respectively) are used. A block of the current picture can be coded by a first motion vector pointing to a first reference block of the first reference picture and a second motion vector pointing to a second reference block of the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0080] Furthermore, to improve coding efficiency, merge mode techniques can be used for inter picture prediction.

[0081] According to some embodiments of the present disclosure, predictions such as inter-image prediction and intra-image prediction are performed in units of blocks (block units). For example, according to the HEVC standard, images in a sequence of video images are divided into coding tree units (CTUs) for compression, and the CTUs in an image have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU can be recursively divided into one or more coding units (CUs) by a quadtree. For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (prediction units) (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, a prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels, such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0082] FIG. 6 shows a diagram of a video encoder (603) according to another embodiment of the present disclosure. The video encoder (603) is configured to receive a processing block (e.g., a prediction block) of sample values in a current video image within a sequence of video images and encode the processing block into a coded image that is part of a coded video sequence. In one example, the video encoder (603) is used instead of the video encoder (303) in the example of FIG. 3.

[0083] In the example of HEVC, the video encoder (603) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples. The video encoder (603) determines whether the processing block is best coded using, for example, rate distortion optimization, in an intra mode, an inter mode, or a bi-prediction mode. If the processing block is to be coded in the intra mode, the video encoder (603) may use an intra prediction technique to encode the processing block into the coded picture; if the processing block is to be coded in the inter mode or the bi-prediction mode, the video encoder (603) may use an inter prediction technique or a bi-prediction technique, respectively, to encode the processing block into the coded picture. In certain video coding techniques, the merge mode can be an inter-picture prediction sub-mode in which the motion vector is derived from one or more motion vector predictors without the benefit of the coded motion vector components outside the predictor. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (603) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0084] In the example of FIG. 6, the video encoder (603) includes an inter encoder (630), an intra encoder (622), a residue calculator (623), a switch (626), a residue encoder (624), a general controller (621), and an entropy encoder (625) that are coupled together as shown in FIG. 6.

[0085] The inter-encoder (630) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference image (e.g., blocks in a previous image and a subsequent image), generate inter-prediction information (e.g., a description of redundant information by inter-encoding techniques, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference image is a decoded reference image decoded based on the encoded video information.

[0086] The intra-encoder (622) is configured to receive samples of a current block (e.g., a processing block) and, optionally, compare the block with blocks already coded in the same image, generate quantized coefficients after transformation, and optionally also generate intra-prediction information (e.g., intra-prediction direction information according to one or more intra-encoding techniques). In one example, the intra-encoder (622) also calculates an intra-prediction result (e.g., a predicted block) based on reference blocks and intra-prediction information in the same image.

[0087] A general controller (621) is configured to determine general control data and control other components of the video encoder (603) based on the general control data. In one example, the general controller (621) determines the mode of a block and provides a control signal to a switch (626) based on that mode. For example, when the mode is the intra mode, the general controller (621) controls the switch (626) to select the intra mode result used by the residual calculator (623), controls the entropy encoder (625) to select the intra prediction information, and includes the intra prediction information in the bitstream; when the mode is the inter mode, the general controller (621) controls the switch (626) to select the inter prediction result used by the residual calculator (623), controls the entropy encoder (625) to select the inter prediction information, and includes the inter prediction information in the bitstream.

[0088] The residual calculator (623) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra-encoder (622) or the inter-encoder (630). The residual encoder (624) operates based on the residual data and is configured to encode the residual data to generate conversion coefficients. In one example, the residual encoder (624) is configured to convert the residual data from the spatial domain to the frequency domain and generate conversion coefficients. The conversion coefficients are then subjected to quantization processing to obtain quantized conversion coefficients. In various embodiments, the video encoder (603) also includes a residual decoder (628). The residual decoder (628) is configured to perform inverse conversion and generate decoded residual data. The decoded residual data can be appropriately used by the intra-encoder (622) and the inter-encoder (630). For example, the inter-encoder (630) can generate a decoded block based on the decoded residual data and the inter-prediction information, and the intra-encoder (622) can generate a decoded block based on the decoded residual data and the intra-prediction information. The decoded block is appropriately processed to generate a decoded image, and the decoded image is buffered in a memory circuit (not shown) and can be used as a reference image in some examples.

[0089] The entropy encoder (625) is configured to format the bitstream to include the encoded block. The entropy encoder (625) is configured to include various information according to an appropriate standard such as the HEVC standard. In one example, the entropy encoder (625) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Note that there is no residual information when coding a block in either the merge sub-mode of the inter mode or the bi-prediction mode according to the disclosed subject matter.

[0090] Figure 7 shows a diagram of a video decoder (710) according to another embodiment of the present disclosure. The video decoder (710) is configured to receive a coded image that is part of a coded video sequence and decode the coded image to generate a reconstructed image. In one example, the video decoder (710) is used in place of the video decoder (310) of the example of FIG. 3.

[0091] In the example of FIG. 7, the video decoder (710) includes an entropy decoder (771), an inter decoder (780), a residual decoder (773), a reconstruction module (774), and an intra decoder (772) that are coupled together as shown in FIG. 7.

[0092] The entropy decoder (771) can be configured to reconstruct from the coded image specific symbols that represent the syntax elements from which the coded image is composed. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi-prediction mode, merge sub-mode or the latter two in another sub-mode), prediction information (e.g., intra prediction information or inter prediction information, etc.) that can identify specific samples or metadata used for prediction by each of the intra decoder (772) or the inter decoder (780), and residual information in the form of, for example, quantized transform coefficients. In one example, when the prediction mode is inter or bi-prediction mode, the inter prediction information is provided to the inter decoder (780); when the prediction type is intra prediction type, the intra prediction information is provided to the intra decoder (772). The residual information can undergo inverse quantization and is provided to the residual decoder (773).

[0093] The inter decoder (780) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0094] The intra decoder (772) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0095] The residual decoder (773) is configured to perform inverse quantization to extract de-quantized transform coefficients and process the de-quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require certain control information (including quantization parameter (QP)), and that information may be provided by the entropy decoder (771) (the data path is not shown because it is only low-volume control information).

[0096] The reconstruction module (774) is configured to combine, in the spatial domain, the residual as the output by the residual decoder (773) and the prediction result (optionally as the output by the inter or intra prediction module) to form a reconstructed block, and this reconstructed block may be part of a reconstructed image, and this part of the reconstructed image may be part of a reconstructed video. Note that other appropriate operations such as a deblocking operation can be performed to improve the visual quality.

[0097] Note that the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using any appropriate technology. In one embodiment, the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using one or more processors that execute software instructions.

[0098] Aspects of the present disclosure provide techniques in the field of inter prediction in advanced video codecs. This technique can be used to set the number of candidates in a candidate list, which can be referred to as a sub-block merge candidate list.

[0099] In various embodiments, for an inter prediction CU, motion parameters including a motion vector, a reference picture index, a reference picture list use index, and / or other additional information can be used for generating an inter prediction sample. The inter prediction can include uni-prediction, bi-prediction, and / or the like. In uni-prediction, a reference picture list (e.g., the first reference picture list or list 0 (L0) or the second reference picture list or list 1 (L1)) can be used. In bi-prediction, both L0 and L1 can be used. The reference picture list use index can indicate that the reference picture list(s) include L0, L1, or both L0 and L1.

[0100] The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded in skip mode, the CU can be associated with one PU, cannot include significant residual coefficients (e.g., the residual coefficients are zero), cannot include a coded motion vector difference (MVD), or cannot include a reference picture index.

[0101] The merge mode can be used such that the motion parameters for the current CU can be obtained from adjacent CUs including spatial and temporal merge candidates and optionally other merge candidates. The merge mode can be applied to inter prediction CUs and can be used for skip mode. Alternatively, the motion parameters can be explicitly transmitted or signaled. For example, a motion vector, a corresponding reference picture index for each reference picture list, a reference picture list use flag, and other information can be explicitly signaled for each CU.

[0102] In some examples (e.g., VVC), one or more of the following inter-prediction coding tools are used: (1) extended merge prediction, (2) merge mode with motion vector difference (MMVD), (3) symmetric MVD (SMVD) signaling, (4) affine motion compensation prediction, (5) sub-block-based temporal motion vector prediction (SbTMVP), (6) adaptive motion vector resolution (AMVR), (7) motion field storage: 1 / 16 luma sample MV storage and 8×8 motion field compression, (8) dual prediction with CU-level weighting, (9) bidirectional optical flow (BDOF), (10) decoder side motion vector refinement (DMVR), (11) geometric partitioning mode (GPM), and (12) combined inter and intra prediction (CIIP).

[0103] According to one aspect of the present disclosure, some inter-prediction coding tools may operate based on a sub-block-based merge candidate list. In one example, affine motion compensation prediction can be performed in an affine merge mode (also called sub-block-based merge mode in some examples). In the affine merge mode, prediction can be made based on an affine merge candidate list, which is called a sub-block-based merge candidate list in some examples. In another example, sub-block-based temporal motion vector prediction (SbTMVP) can also operate based on a sub-block-based merge candidate list.

[0104] In some examples (e.g., HEVC), for affine motion compensation prediction, only the translational motion model is applied to motion compensation prediction (MCP). The real world has many types of motions such as zoom in / out, rotation, perspective motion, and other irregular motions. In some examples (e.g., VVC), block-based affine transform motion compensation prediction is applied.

[0105] Figures 8A - 8B show the affine motion model. Figure 8A shows the affine motion field of a block described by the motion information of two control points CP0 and CP1 (4 - parameter affine model), and Figure 8B shows the affine motion field of a block described by three control points CP0, CP1, and CP2 (6 - parameter affine model).

[0106] In some embodiments, for a 4 - parameter affine motion model, the motion vector (mv x , mv y ) at the sample position (x, y) of the block can be derived as (Equation 1), and for a 6 - parameter affine motion model, the motion vector at the sample position (x, y) within the block can be derived as (Equation 2):

Equation

[0107] To simplify motion - compensation prediction, block - based affine - transformation prediction is applied.

[0108] Figure 9 shows an example of an affine MV field per sub-block. In one example, the current CU910 (e.g., 16×16 luma samples) is divided into 4×4 luma sub-blocks (each sub-block can be 4×4 luma samples). To derive the motion vector for each 4×4 luma sub-block, as shown in Figure 9, the motion vector of the central sample of each sub-block is calculated according to the above equations (Equation 1) and (Equation 2). The motion vector can be rounded, for example, to 1 / 16 fraction accuracy. Next, a motion compensation interpolation filter is applied to generate a prediction for each sub-block with the derived motion vector. In some examples, since the sub-block size of the chroma component can also be set to 4×4, a 4×4 chroma sub-block includes four corresponding 4×4 luma sub-blocks. The MV of the 4×4 chroma sub-block is calculated, in one example, as the average of the MVs of the four corresponding 4×4 luma sub-blocks.

[0109] Note that the sub-block can be defined to have other suitable numbers of luma samples. Note also that in some examples, the sub-block is called a sub-CU.

[0110] For translational motion inter prediction, two affine motion inter prediction modes called the affine merge (AF_MERGE) mode and the affine advanced MVP (affine AMVP) mode can be used.

[0111] Regarding affine merge prediction, in one example, the AF_MERGE mode can be applied to CUs where both the width and height are 8 or more. In the AF_MERGE mode, the control point motion vector (CPMV) of the current CU is generated based on the motion information of spatially adjacent CUs. In one example, the affine merge candidate list (also called the sub-block based merge candidate list) can include up to five control point motion vector predictor (CPMVP) candidates, and an index is signaled to indicate which one will be used for the current CU. In one example, three types of CPMV candidates are used to form the affine merge candidate list. The first type of CPMV candidate is an inherited affine merge candidate extrapolated from the CPMV of adjacent CUs. The second type of CPMV candidate is a constructed affine merge candidate CPMV derived using the translational MV of adjacent CUs. The third type of CPMV candidate uses a zero MV.

[0112] In some examples such as VVC, up to two inherited affine candidates can be used. In one example, two inherited affine candidates are derived from the affine motion models of adjacent blocks, one from the left adjacent CU (referred to as the left predictor) and one from the upper adjacent CU (referred to as the upper predictor). Using the adjacent blocks shown in FIG. 1 as an example, for the left predictor, the scan order is A0->A1, and for the upper predictor, the scan order is B0->B1->B2. In one example, only the first inherited candidate available from both sides is selected. In some examples, a pruning check is not performed between the two inherited candidates. When an adjacent affine CU is identified, the control point motion vector of the adjacent affine CU is used to derive the CPMV candidates in the affine merge candidate list of the current CU.

[0113] Figure 10 shows an example for determining the inherited control point motion vectors in the affine merge mode. As shown in Figure 10, when the adjacent lower left sub-block A is coded in the affine mode, the motion vectors mv2, mv3, and mv4 of the upper left corner, upper right corner, and lower left corner of the CU including sub-block A can be obtained. When sub-block A is coded with a 4-parameter affine model, two CPMVs of the current CU are calculated according to mv2 and mv3. When sub-block A is coded with a 6-parameter affine model, three CPMVs of the current CU are calculated according to mv2, mv3, and mv4.

[0114] In some examples, the constructed affine candidates are constructed by combining the adjacent translational motion information of each control point. The motion information of the control point can be derived from the specified spatially adjacent ones (spatial neighbors) and temporally adjacent ones (temporal neighbor).

[0115] Figure 11 shows examples of spatially adjacent ones (e.g., sub-blocks A0 to A2 and B0 to B3) and temporally adjacent ones (e.g., T) according to some embodiments of the present disclosure. In one example, CPMV k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV1, the B2->B3->A2 block is checked (-> is used for the check order), and the MV of the first available block is used as CPMV1. For CPMV2, the B1->B0 block is checked, and the MV of the first available block is used as CPMV2. For CPMV3, the A1->A0 block is checked, and the MV of the first available block is used as CPMV3. In the case of TMVP, T is checked, and if the MV of block T is available, it is used as CPMV4.

[0116] After obtaining the MVs of the four control points CPMV1 - CPMV4, affine merge candidates are constructed based on their motion information. The following combinations of control point MVs are used to construct in the order of {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.

[0117] Combinations of three CPMVs can form 6 - parameter affine merge candidates, and combinations of two CPMVs can form 4 - parameter affine merge candidates. In one example, in order to avoid the motion scaling process, when the reference indices of the control points are different, the related combinations of control point MVs can be discarded.

[0118] In one example, after checking the inherited affine merge candidates and the constructed affine merge candidates, if the candidate list is not yet full, zero MVs are inserted at the end of the list.

[0119] For affine AMVP prediction, the affine AMVP mode can be applied to CUs where both the width and height are 16 or more. In some examples, a CU - level affine flag is signaled in a bitstream (e.g., a coded video bitstream) to indicate whether the affine AMVP mode is used in the CU, and then another flag is signaled to indicate whether 4 - parameter affine or 6 - parameter affine is used. In the affine AMVP mode, the difference between the CPMVs of the current CU and their predictor CPMVPs can be signaled in the bitstream. The size of the affine AMVP candidate list is 2, and the affine AMVP candidate list is generated using the following four types of CPMV candidates in order: (1) inherited affine AMVP candidates extrapolated from the CPMVs of adjacent CUs, (2) constructed affine AMVP candidates derived using the translational MVs of adjacent CUs, (3) translational MVs from adjacent CUs, (4) zero MVs.

[0120] In some examples, the inspection order of the inherited affine AMVP candidates is the same as that of the inherited affine merge candidates. In one example, the only difference between the affine merge prediction and the affine AMVP prediction is that for AMVP candidates, only the affine CUs that have the same reference picture as the current block are considered. In one example, pruning is not applied when inserting the inherited affine motion predictor into the candidate list.

[0121] In some examples, the constructed AMVP candidates can be derived from the specific spatially adjacent ones shown in FIG. 11. In one example, the same checking order as that performed in candidate construction for affine merge prediction is used. In addition, the reference picture indices of the adjacent blocks are also checked. The first block in the checking order that is inter-coded and has the same reference picture as the current CU is used. If the current CU is coded in 4-parameter affine mode and the motion vectors of both of the two control points (e.g., {CPMV1, CPMV2}) are available, the motion vectors of the two control points are added as one candidate to the affine AMVP list. If the current CU is coded in 6-parameter affine mode and all three motion vectors of the control point CPMV (e.g., {CPMV1, CPMV2, CPMV3}) are available, they are added as one candidate to the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.

[0122] After the inherited AMVP candidates and the constructed AMVP candidates are checked, if the number of candidates in the affine AMVP list is still less than 2, CPMV1, CPMV2, and CPMV3 are added in order as translational MVs for predicting all control point MVs of the current CU if available. Finally, if the affine AMVP list is not yet full, zero MVs are used to fill the affine AMVP list.

[0123] According to some aspects of the present disclosure, the motion information can be stored in a suitable buffer such as a local buffer, an image line buffer, etc. The local buffer is used to store motion information at the CTU level, such as the motion vectors of 4×4 blocks within a CTU. For example, when a CU in a CTU is decoded based on inter prediction, the motion vectors of each 4×4 block of the CU can be stored in the local buffer and used for decoding subsequent CUs. The image line buffer is used to store the motion information of CTUs above the current CTU, such as the motion vectors of 4×4 blocks at the lower part of the upper CTU. The CTU above the current CTU can be referred to as the upper CTU line.

[0124] In some examples (e.g., VVC), the CPMV of an affine CU is stored separately from the motion vectors of 4×4 blocks. In one example, the local buffer includes a first part for storing the motion vectors of 4×4 blocks in a CTU and a second part for storing the CPMV of the affine CU in the CTU. The CPMV stored in the second part of the local buffer can be used to generate the inherited CPMVP in the affine merge mode and the affine AMVP mode for the most recently coded CU. The sub-block MVs derived from the CPMV are used for motion compensation, MV derivation for the merge / AMVP list of translational MVs, and block release.

[0125] In some embodiments, the picture line buffer does not store the additional CPMV of the affine CUs in the upper CTU line. In some examples, the inheritance of affine motion data from CUs from the upper CTU is treated differently from the inheritance from normal adjacent CUs in the same CTU line. If the candidate CU for affine motion data inheritance is within the upper CTU line, instead of the CPMV, the lower left and lower right sub-block MVs in the picture line buffer are used for affine MVP derivation. Thus, in some examples, the CPMV is stored only in the local buffer, not in the picture line buffer. In an example where the candidate CU is 6-parameter affine coded, the affine model can be degraded to a 4-parameter model.

[0126] FIG. 12 shows a diagram illustrating the use of motion vectors for affine motion data inheritance in some examples. In FIG. 12, each small square represents a 4×4 sub-block, and the motion vector of the sub-block can be the motion vector at the center of the sub-block. Further, the current CU is located at the vertex position of the current CTU. As shown in FIG. 12, among the adjacent CUs of the current CU, CU-E and CU-D are affine coded. CU-D is in the same CTU line as the current CU, and CU-E is in the upper CTU line of the current CU. The CPMV of CU-D can be stored in the local buffer. For example, for a 4-parameter affine model, mv D0 and mv D1 are stored in the local buffer, and the CPMV of the current CU (e.g., mv0 and mv1) can be calculated according to the corresponding positions of the control points for mv D0 and mv D1 , as well as mv D0 and mv D1 .

[0127] In one example, the picture line buffer stores the motion vectors of the sub-blocks at the lower part of the upper CTU line. The CPMV of CU-E, as indicated by mv E0 or mv E1 , is not stored in the picture line buffer. In one example, mvLE0 and mv LE1 The motion vectors of the lower left sub-block and the lower right sub-block of CU-E as shown by LE1 are used for the affine inheritance of the current CU. For example, the CPMV of the current CU S (e.g., mv0 and mv1) is mv LE0 and mv LE1 and can be calculated according to the corresponding center positions of the two sub-blocks.

[0128] In some embodiments, prediction refinement by optical flow (PROF) (also referred to as the PROF method) can be implemented to improve sub-block-based affine motion compensation and achieve finer-grained motion compensation without increasing the memory access bandwidth for motion compensation. In one embodiment (e.g., VVC), after sub-block-based affine motion compensation is performed, the difference (or refinement values, refinement, prediction refinement) derived based on the optical flow equation can be added to the prediction sample (e.g., the luma-predicted sample, or the luma prediction sample) to obtain the refined prediction sample.

[0129] FIG. 13 shows a schematic diagram of an example of the PROF method according to an embodiment of the present disclosure. The current block (1310) can be divided into four sub-blocks (1312, 1314, 1316, and 1318). Each of the sub-blocks (1312, 1314, 1316, and 1318) can have a size of 4×4 pixels or samples. The sub-block MV (1320) for the sub-block (1312) can be derived according to the CPMV of the current block 1310 using, for example, the center position of the sub-block (1312) and the affine motion model (e.g., the 4-parameter affine motion model, the 6-parameter affine motion model). The sub-block MV (1320) can indicate the reference sub-block (1332) in the reference image. The initial sub-block prediction sample can be determined according to the reference sub-block (1332).

[0130] In some examples, as described by sub - block MV (1320), the translational motion from a reference sub - block (1332) to a sub - block (1312) may not accurately predict the sub - block (1312). In addition to the translational motion described by sub - block MV (1320), the sub - block (1312) may also experience non - translational motion (e.g., rotation as seen in FIG. 13). Referring to FIG. 13, a sub - block (1350) within a reference image having shaded samples (e.g., sample (1332a)) can correspond to and be used to reconstruct samples within the sub - block (1312). The shaded sample (1332a) is shifted by a pixel MV (1340) to accurately reconstruct a sample (1312a) within the sub - block (1312). Thus, in some examples, when non - translational motion occurs, an appropriate prediction refinement method can be applied to the affine motion model as described below to improve the accuracy of prediction.

[0131] In one example, the PROF method is implemented using the following four steps. In step (1), sub - block - based affine motion compensation can be performed to generate a prediction for the current sub - block (e.g., sub - block (1312)), such as an initial sub - block prediction I(i,j), where i and j are coordinates corresponding to samples at a position (i,j) (also referred to as a sample position or sample location) within the current sub - block (1312).

[0132] In step (2), gradient calculation can be performed, where the spatial gradients g x (i,j) and g y (i,j) of the initial sub - block prediction I(i,j) at each sample position (i,j) can be calculated, for example, using a 3 - tap filter [-1,0,1] according to equations 3 and 4 below: g x(i,j) = I(i + 1,j) - I(i - 1,j) (Equation 3) g y (i,j) = I(i,j + 1) - I(i,j - 1) (Equation 4) For sub - block prediction, for gradient calculation, it can be extended by one pixel on both sides. In some embodiments, to reduce memory bandwidth and complexity, the pixels on the extended boundary can be copied from the nearest integer pixel position within the reference image (e.g., the reference image including sub - block (1332)). Thus, additional interpolation for the padding region can be avoided.

[0133] In step (3), the prediction refinement ΔI(i,j) can be calculated as follows by Equation 5 (e.g., the optical flow equation). ΔI(i,j) = g x (i,j)×Δmv x (i,j)+g y (i,j)×Δmv y (i,j) (Equation 5) Here, Δmv(i,j) (e.g., Δmv(1342)) is the difference MV between the pixel MV or sample MV mv(i,j) (e.g., pixel MV(1340)) for the sample position (i,j) and the sub - block MV Mv SB (e.g., sub - block MV(1320)) of the sub - block (e.g., sub - block (1312)) where the sample position (i,j) is located. Δmv(i,j) can also be called the MV refinement (MVR) for the sample position (i,j) or the sample at (i,j). Δmv(i,j) can be determined using Equation 6 as follows. Δmv(i,j) = mv(i,j)-mv SB (Equation 6) Δmv x (i,j) and Δmv y (i,j) are the x - component (horizontal component) and y - component (vertical component) of the difference MV Δmv(i,j), respectively.

[0134] Since the pixel position and the affine model parameters with respect to the sub-block center position do not change from one sub-block to another, Δmv(i,j) can be calculated for the first sub-block (e.g., sub-block (1312)) and reused for other sub-blocks (e.g., sub-blocks (1314), (1316), and (1318)) within the same current block (1310). In some examples, x and y represent the horizontal and vertical shifts of the sample position (i,j) with respect to the center position of sub-block (1312), and Δmv(i,j) (e.g., including Δmv x (i,j) and Δmv y (i,j)) can be derived by Equation 7 as follows:

Equation

[0135] For example, for a 4-parameter affine motion model, the parameters a~d are described by (Equation 1). For a 6-parameter affine motion model, the parameters a~d are described by (Equation 2) as described above.

[0136] In step (4), a prediction refinement ΔI(i,j) (e.g., luma prediction refinement) can be added to the initial sub-block prediction I(i,j) to generate another prediction such as the refined prediction I’(i,j). The refined prediction I’(i,j) can be generated for the sample (i,j) using Equation 8 as follows: I’(i,j)=I(i,j)+ΔI(i,j) (Equation 8).

[0137] In some cases, PROF is not applied to the affine-coded CUs. In one example, all control point MVs are the same, which indicates that only the CU has translational motion and PROF is not applied. In another example, the affine motion parameters are larger than the specified limits, and then PROF is applied. In the second case, the sub-block-based affine motion compensation degrades to CU-based motion compensation to avoid large memory access bandwidth requirements.

[0138] In some embodiments, a fast encoding method can be applied to reduce the coding complexity of affine motion estimation by PROF. In the fast encoding method, PROF is not applied in the affine motion estimation stage in the following two situations. In the first situation, when the current CU is not the root block and its parent block does not select the affine mode as the best mode, the probability of selecting the affine mode as the best mode for the current CU is low, so PROF is not applied. In the second situation, when the magnitudes of all four affine parameters (a to d) are smaller than the predefined thresholds and the current image is not a low-delay image, the improvement brought by PROF in this situation is small, so PROF is not applied. In this way, the affine motion estimation by PROF can be accelerated.

[0139] In some examples (e.g., VVC), sub-block based temporal motion vector prediction (SbTMVP) can be used. Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field of a collocated picture to improve motion vector prediction and merge mode for a current picture's CU. In some examples, the same collocated pictures used in TMVP are used for SbTMVP. SbTMVP differs from TMVP in two aspects. In the first aspect, TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level. In the second aspect, TMVP fetches a temporal motion vector from a collocated block within a collocated picture in the collocated picture (the collocated block is the bottom-right or central block relative to the current CU), and SbTMVP applies a motion shift before fetching temporal motion information from the collocated picture. The motion shift is obtained from a motion vector from one of the spatially adjacent blocks of the current CU.

[0140] Figures 14-15 show examples of the SbTMVP process according to some embodiments of the present disclosure. SbTMVP predicts the motion vectors of sub-CUs within a current CU in two steps. In the first step, the spatially adjacent A1 shown in FIG. 14 is considered. If the spatially adjacent A1 has a motion vector using a collocated picture as its reference picture, the motion vector is selected to be the motion shift that will be applied. If no such motion is identified, the motion shift is set to (0,0).

[0141] In the second step, the motion shift identified in the first step is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vectors and reference indices) from the collocated image as shown in FIG. 15. In the example of FIG. 15, the motion vector of A1 is set as the motion shift (1510). Next, for each sub-CU, the motion information of the corresponding block (the smallest motion grid covering the central sample) in the collocated image is used to derive the motion information for the sub-CU. After the motion information of the collocated sub-CU is identified, it is converted to the motion vector and reference index of the current sub-CU in a similar way to the TMVP process of HEVC. For example, temporal motion scaling is applied to align the reference image of the temporal motion vector with the reference image of the current CU.

[0142] In some examples, such as in VVC, a sub-block-based merge candidate list is used for signaling the sub-block-based merge mode. The sub-block-based merge candidate list can include both SbTMVP candidates and affine merge candidates, and in some examples is called a combined sub-block-based merge candidate list. The SbTMVP mode is enabled / disabled by a flag such as a sequence parameter set (SPS) flag. When enabling the SbTMVP mode, in one example, the SbTMVP predictor is added as the first entry of the combined sub-block-based merge candidate list, followed by the affine merge candidates. In some examples (e.g., VVC), the maximum allowable size of the combined sub-block-based merge candidate list is 5. Note that the maximum allowable size of the combined sub-block-based merge candidate list can be other suitable numbers.

[0143] In one example, the sub-CU size used in SbTMVP is fixed at 8×8, and the SbTMVP mode is applicable only to CUs with both width and height of 8 or more, as performed in the affine merge mode.

[0144] In some embodiments, the encoding logic for additional SbTMVP merge candidates is the same as for other merge candidates. For example, for each CU in a P or B slice, an additional rate-distortion check is performed to determine whether to use the SbTMVP candidate.

[0145] According to some aspects of the present disclosure, the maximum number of candidates in a combined sub-block-based merge candidate list can be signaled.

[0146] FIG. 16 shows an example syntax table (1600) of a sequence parameter set (SPS) in some examples. The SPS contains information applicable to a series of consecutive coded video pictures (also called a coded video sequence).

[0147] In the example syntax table (1600), a flag sps_temporal_mvp_enabled_flag is signaled as shown in (1610). A flag sps_temporal_mvp_enabled_flag equal to 1 specifies that temporal motion vector predictors can be used in the coded video; a flag sps_temporal_mvp_enabled_flag equal to 0 specifies that temporal motion vector predictors are not used in the coded video. In some examples, the coded video can be referred to as a coded layer video sequence (CLVS), which is a group of pictures belonging to the same layer that starts from a random access point, where pictures that may depend on each other and random access point pictures follow.

[0148] In an example of a related syntax table (1600), when the flag sps_temporal_mvp_enabled_flag is equal to 1, two flags, sps_sbtmvp_enabled_flag and sps_affine_enabled_flag, are signaled as shown in (1620) and (1630). The flag sps_sbtmvp_enabled_flag equal to 1 specifies that the sub-block based temporal motion vector predictor can be used for decoding an image in a slice that has a slice type not equal to I (intra-coded) in the coded video. The flag sps_sbtmvp_enabled_flag equal to 0 specifies that the sub-block based temporal motion vector predictor is not used in the coded video. In the example, when the flag sps_sbtmvp_enabled_flag is not signaled, it can be assumed that the flag sps_sbtmvp_enabled_flag is equal to 0.

[0149] The flag sps_affine_enabled_flag specifies whether affine model based motion compensation can be used for inter prediction. When the flag sps_affine_enabled_flag is equal to 0, in some examples, the syntax is constrained such that affine model based motion compensation is not used in the coded video. Otherwise (sps_affine_enabled_flag is equal to 1), affine based motion compensation can be used in the coded video.

[0150] In the example of the syntax table (1600), when the flag sps_affine_enabled_flag is equal to 1, parameters such as five_minus_max_num_subblock_merge_cand can be signaled. The parameter five_minus_max_num_subblock_merge_cand specifies the result of subtracting the maximum number of sub-block-based merge candidates supported by the SPS from 5. The value of five_minus_max_num_subblock_merge_cand is in the range from 0 to 5, including several examples. For example, when the value of five_minus_max_num_subblock_merge_cand is 2, the maximum number of candidates in the combined sub-block-based merge candidate list is 3 (subtracting 2 from 5).

[0151] In some examples, the temporal motion vector predictor can be enabled / disabled at the picture header level. FIG. 17 shows an example of the syntax table of the picture header structure in some examples (1700).

[0152] In the example of the syntax table (1700), when the SPS level flag sps_temporal_mvp_enabled_flag is equal to 1, the flag ph_temporal_mvp_enabled_flag is signaled as shown by (1710). The flag ph_temporal_mvp_enabled_flag specifies whether a temporal motion vector predictor can be used for inter prediction of slices associated with an image header. When ph_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the slice associated with the image header are constrained such that the temporal motion vector predictor is not used for decoding of the slice. Otherwise (ph_temporal_mvp_enabled_flag is equal to 1), the temporal motion vector predictor may be used for decoding of the slice associated with the image header. If it does not exist, in one example, the value of ph_temporal_mvp_enabled_flag is assumed to be equal to 0. When the reference image in the decoded picture buffer does not have the same spatial resolution as the current image, the value of ph_temporal_mvp_enabled_flag becomes equal to 0.

[0153] The maximum number of subblock-based merge candidates can be derived based on the signaled or inferred flags and parameters. In one example, the variable MaxNumSubblockMergeCand is used to indicate the maximum number of subblock-based merge candidates. In one example, when sps_affine_enabled_flag is equal to 1, MaxNumSubblockMergeCand is derived according to (Equation 9), and when sps_affine_enabled_flag is equal to 0, MaxNumSubblockMergeCand is derived according to (Equation 10):

Number

[0154] In some examples, the value of MaxNumSubblockMergeCand is in the range of 0 to 5 (including 0 and 5).

[0155] According to one aspect of the present disclosure, when sps_affine_enabled_flag is signaled as 1, MaxNumSubblockMergeCand is derived from five_minus_max_num_subblock_merge_cand as described in (Equation 9). In some examples, a scenario is permitted where sps_affine_enabled_flag is signaled as 1 and five_minus_max_num_subblock_merge_cand is signaled as equal to 5. In this scenario, the maximum number of subblock-based merge candidates MaxNumSubblockMergeCand is derived as 0, which can turn off the affine merge mode in the same way as SbTMVP regardless of the SbTMVP enabling flag, and the SbTMVP enabling flag may cause a conflict when indicating that SbTMVP is enabled.

[0156] Aspects of the present disclosure provide a technique for setting a range of values (also referred to as the maximum number of subblock-based merge candidates) of the number of subblock-based merge candidates according to the default number of subblock-based merge candidates (e.g., denoted as N) and relevant high-level usage flags for affine and / or SbTMVP coding tools. For example, when the SbTMVP enabling flag indicates that SbTMVP is enabled, the maximum number of subblock-based merge candidates is not 0.

[0157] In some embodiments, the parameter five_minus_max_num_subblock_merge_cand has a negative correlation with the maximum number of subblock-based merge candidates, and the upper limit of the parameter five_minus_max_num_subblock_merge_cand is determined based on the SbTMVP enabling flag.

[0158] In one embodiment, the parameter five_minus_max_num_subblock_merge_cand specifies the maximum number of sub-block based merge motion vector prediction candidates supported by the SPS subtracted from N. Further, the value of the parameter five_minus_max_num_subblock_merge_cand is restricted to the range from 0 to N - sps_sbtmvp_enabled_flag (including 0 and N - sps_sbtmvp_enabled_flag). The upper limit of the parameter five_minus_max_num_subblock_merge_cand depends on the flag sps_sbtmvp_enabled_flag.

[0159] In some examples, the default number N is 5. When the flag sps_sbtmvp_enabled_flag is 0 (SbTMVP is disabled), the value of the parameter five_minus_max_num_subblock_merge_cand can be in the range from 0 to 5 (including 0 and 5). However, when the flag sps_sbtmvp_enabled_flag is 1 (SbTMVP is enabled), the value of the parameter five_minus_max_num_subblock_merge_cand can be in the range from 0 to 4 (including 0 and 4). As an example, on the encoder side, when the flag sps_sbtmvp_enabled_flag is 1 and the calculated value of the parameter five_minus_max_num_subblock_merge_cand exceeds the upper limit of 5, the value signaled for the parameter five_minus_max_num_subblock_merge_cand in the coded video bitstream is restricted to 4 in the range from 0 to 4 (including 0 and 4).

[0160] In some examples, when the value of the parameter five_minus_max_num_subblock_merge_cand is equal to the upper limit of the range, the parameter five_minus_max_num_subblock_merge_cand may not be signaled from the encoder side in the coded video bitstream. On the decoder side, if the decoder detects that the parameter five_minus_max_num_subblock_merge_cand does not exist in the coded video bitstream, the decoder can assume that the value of the parameter five_minus_max_num_subblock_merge_cand is at the upper limit of the range. The upper limit of the range can be determined based on the SbTMVP enable flag. For example, the value of five_minus_max_num_subblock_merge_cand is assumed to be equal to N - sps_sbtmvp_enabled_flag. In one example, the default number N is 5, and when sps_sbtmvp_enabled_flag is 0 (SbTMVP is disabled), the value of the parameter five_minus_max_num_subblock_merge_cand can be assumed to be 5. However, when the flag sps_sbtmvp_enabled_flag is 1 (SbTMVP is enabled), the value of the parameter five_minus_max_num_subblock_merge_cand is assumed to be 4.

[0161] In another embodiment, the upper limit of the parameter five_minus_max_num_subblock_merge_cand is determined based on a combination of a plurality of SbTMVP enable flags, such as the first flag sps_sbtmvp_enabled_flag at the SPS level and the second flag ph_temporal_mvp_enabled_flag at the picture header level. In one example, the value of the parameter five_minus_max_num_subblock_merge_cand is restricted to the range from 0 to N-(sps_sbtmvp_enabled_flag && ph_temporal_mvp_enabled_flag) (including 0 and N-(sps_sbtmvp_enabled_flag && ph_temporal_mvp_enabled_flag)). If the parameter five_minus_max_num_subblock_merge_cand does not exist in the coded video bitstream, the value of five_minus_max_num_subblock_merge_cand is presumed to be equal to N-(sps_sbtmvp_enabled_flag && ph_temporal_mvp_enabled_flag).

[0162] In some examples, the default number N is 5. When at least one of the first flag sps_sbtmvp_enabled_flag and the second flag ph_temporal_mvp_enabled_flag is 0 (SbTMVP is disabled), the value of the parameter five_minus_max_num_subblock_merge_cand can be in the range from 0 to 5 (including 0 and 5). However, when both the first flag sps_sbtmvp_enabled_flag and the second flag ph_temporal_mvp_enabled_flag are 1 (SbTMVP is enabled), the value of the parameter five_minus_max_num_subblock_merge_cand can be in the range from 0 to 4 (including 0 and 4). In one example, on the encoder side, when both the first flag sps_sbtmvp_enabled_flag and the second flag ph_temporal_mvp_enabled_flag are 1 and the calculated value of the parameter five_minus_max_num_subblock_merge_cand exceeds the upper limit of 5, the value of the parameter five_minus_max_num_subblock_merge_cand signaled in the coded video bitstream is restricted to 4 in the range from 0 to 4 (including 0 and 4).

[0163] In some examples, when the value of the parameter five_minus_max_num_subblock_merge_cand is equal to the upper limit of the range, the parameter five_minus_max_num_subblock_merge_cand may not be signaled from the encoder side in the coded video bitstream. On the decoder side, if the decoder detects that the parameter five_minus_max_num_subblock_merge_cand does not exist in the coded video bitstream, the decoder can assume that the value of the parameter five_minus_max_num_subblock_merge_cand is the upper limit of the range. The upper limit of the range can be determined based on, for example, an appropriate combination of the first flag sps_sbtmvp_enabled_flag and the second flag ph_temporal_mvp_enabled_flag. For example, the value of five_minus_max_num_subblock_merge_cand is assumed to be equal to N - (sps_sbtmvp_enabled_flag && ph_temporal_mvp_enabled_flag). In one example, the default number N is 5, and when at least one of the first flag sps_sbtmvp_enabled_flag and the second flag ph_temporal_mvp_enabled_flag is 0 (SbTMVP is disabled), the value of the parameter five_minus_max_num_subblock_merge_cand can be assumed to be 5. However, when both the first flag sps_sbtmvp_enabled_flag and the second flag ph_temporal_mvp_enabled_flag are 1 (SbTMVP is enabled), the value of the parameter five_minus_max_num_subblock_merge_cand can be assumed to be 4.

[0164] FIG. 18 shows a flowchart outlining a process (1800) according to an embodiment of the present disclosure. The process (1800) can be used for block reconstruction and thus generates prediction blocks for blocks being reconstructed. In various embodiments, the process (1800) is executed by processing circuits such as those of terminal devices (210), (220), (230) and (240), a processing circuit that executes the functions of video encoder (303), a processing circuit that executes the functions of video decoder (310), a processing circuit that executes the functions of video decoder (410), a processing circuit that executes the functions of video encoder (503). In some embodiments, the process (1800) is implemented by software instructions, and thus when the processing circuit executes the software instructions, the processing circuit executes the process (1800). The process starts at (S1801) and proceeds to (S1810).

[0165] (In S1810), a parameter (e.g., five_minus_max_num_subblock_merge_cand indicating the maximum number of candidates in a sub-block based merge candidate list) is determined based on prediction information decoded from the coded video bitstream. The parameter is within a range that depends on a flag indicating the enabled / disabled state of sub-block based temporal motion vector prediction. In some examples, the upper limit of the range depends on a flag indicating the enabled / disabled state of sub-block based temporal motion vector prediction. In one example, the flag indicates the enabled / disabled state of sub-block based temporal motion vector prediction at the sequence parameter set (SPS) level.

[0166] In one embodiment, the value of the parameter is signaled in the coded video bitstream. In another example, if the value of the parameter is not signaled in the coded video bitstream, the value of the parameter can be assumed to be the upper limit of the range. For example, the parameter can be estimated based on a default number and a flag indicating the enabled / disabled state of sub-block based temporal motion vector prediction in response to a parameter not signaled in the coded video bitstream.

[0167] In some embodiments, the parameter is within a range that depends on a first flag indicating the enabled / disabled state of sub-block based temporal motion vector prediction at the sequence parameter set (SPS) level and a second flag indicating the enabled / disabled state of temporal motion vector prediction at the picture header (PH) level. In some examples, in response to the parameter not being signaled in the coded video bitstream, the parameter can be estimated based on a default number, a first flag indicating the enabled / disabled state of sub-block based temporal motion vector prediction at the SPS level, and a second flag indicating the enabled / disabled state of temporal motion vector prediction at the PH level.

[0168] (S1820), the maximum number of candidates in the sub-block based merge candidate list is calculated based on the parameter. In some examples, the maximum number of candidates in the sub-block based merge candidate list is calculated by subtracting the parameter from a default number, such as using (Equation 9). In one example, the default number is 5.

[0169] (S1830), in response to the current block in the sub-block based prediction mode, the samples of the current block are reconstructed based on the candidate selection from the constructed sub-block based merge candidate list of the current block. The constructed sub-block based merge candidate list of the current block is restricted by the maximum number of candidates in the sub-block based merge candidate list.

[0170] Thereafter, this process proceeds to (S1899) and ends.

[0171] The above-described technology can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 19 shows a computer system (1900) suitable for implementing a particular embodiment of the disclosed subject matter.

[0172] The computer software can be coded using any suitable machine code or computer language that can be the subject of assembly, compilation, linking, or similar mechanisms to create code including instructions that can be executed directly, or through interpretation, such as through microcode execution, by one or more computer central processing units (CPUs), graphics processing units (GPUs), and the like.

[0173] The instructions can be executed on various types of computers or their components including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0174] The components shown in FIG. 19 for the computer system (1900) are exemplary in nature and are not intended to suggest any limitations regarding the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiment of the computer system (1900).

[0175] The computer system (1900) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, movements of a data glove), voice input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not shown). Also, the human interface device can be used to capture specific media that is not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still camera), and video (e.g., 2D video, 3D video including stereoscopic images).

[0176] The input human interface device may include one or more of a keyboard (1901), a mouse (1902), a trackpad (1903), a touch screen (1910), a data glove (not shown), a joystick (1905), a microphone (1906), a scanner (1907), and a camera (1908).

[0177] The computer system (1900) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., a touch screen (1910), a data glove (not shown), a joystick (1905) with tactile feedback, although there may also be a tactile feedback device that does not function as an input device), audio output devices (e.g., speakers (1909), headphones (not shown)), visual output devices (screens (1910) including CRT screens, LCD screens, plasma screens, OLED screens, each of which may or may not have touch screen input capabilities and may or may not have tactile feedback capabilities - some of these can output more than three-dimensional output through means such as two-dimensional visual output or stereoscopic image output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and may include printers (not shown).

[0178] The computer system (1900) may also include a memory device accessible by humans, and optical media including a CD / DVD ROM / RW (1920) having a CD / DVD or similar medium (1921), a thumb drive (1922), a removable hard drive or solid state drive (1923), legacy magnetic media such as tapes and floppy (registered trademark) disks (not shown), and related media such as specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0179] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0180] The computer system (1900) can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet (registered trademark), wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., cellular networks including cable TV, satellite TV, and terrestrial broadcast TV, TV wired or wireless wide area digital networks, vehicle and industrial including CAN bus, etc. A particular network generally requires an external network interface adapter that attaches to a particular general-purpose data port or peripheral bus (1949) (e.g., the USB port of the computer system (1900)). Others are generally incorporated into the core of the computer system (1900) by attaching to the system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1900) can communicate with other entities. Such communication can be unidirectional, receive only (e.g., broadcast TV), dedicated unidirectional transmission (e.g., CAN bus to a particular CAN bus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Specific protocols and protocol stacks can be used with each of those networks and network interfaces as described above.

[0181] The aforementioned human interface device, human-accessible memory device, and network interface can be attached to the core (1940) of the computer system (1900).

[0182] The core (1940) can include one or more central processing units (CPUs) (1941), a graphics processing unit (GPU) (1942), a special programmable processing unit in the form of a field programmable gate area (FPGA) (1943), a hardware accelerator (1944) for specific tasks, etc. These devices can be connected via a system bus (1948) together with a read-only memory (ROM) (1945), a random access memory (1946), an internal mass storage device such as an internal non-user-accessible hard drive, SSD, etc. (1947). In some computer systems, the system bus (1948) can be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core's system bus (1948) or via a peripheral bus (1949). The architecture of the peripheral bus includes PCI, USB, etc.

[0183] The CPU (1941), GPU (1942), FPGA (1943), and accelerator (1944) can be combined to execute specific instructions that can constitute the above-mentioned computer code. The computer code can be stored in the ROM (1945) or RAM (1946). Transient data can also be stored in the RAM (1946), while permanent data can be stored, for example, in the internal mass storage device (1947). Fast storage and retrieval to any memory device can be enabled through the use of cache memory, which can be closely associated with one or more CPUs (1941), GPUs (1942), mass storage devices (1947), ROM (1945), RAM (1946), etc.

[0184] A computer-readable medium can have thereon computer code for performing operations implemented on various computers. The medium and the computer code can be specially designed and configured for the purposes of this disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0185] As an example, and not by way of limitation, a computer system having an architecture (1900), specifically a core (1940), can provide functionality as a result of software embodied on one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be media associated with a user-accessible mass storage device as described above, as well as a specific storage device (1940) of the core (1940) that is of a non-transitory nature, such as a core-internal mass storage device (1947) or ROM (1945). The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1940). The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software includes defining data structures stored in RAM (1946) and modifying such data structures according to processes defined by the software, and causing a specific process or a specific part of a process described herein to be executed by the core (1940), specifically a processor (including a CPU, GPU, FPGA, etc.) therein. Additionally, or alternatively, the computer system can provide functionality as a result of logic wired within a circuit (e.g., an accelerator (1944)) or otherwise embodied, which can operate in place of or in conjunction with software for executing a specific process or a specific part of a specific process described herein. References to software include logic and, if necessary, vice versa. References to computer-readable media can include a circuit (such as an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video User Interface Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CAN bus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: Solid State Drive IC: Integrated Circuit CU: Coding Unit

[0186] Although several exemplary embodiments have been described, there are changes, substitutions, and various alternative equivalents within the scope of the present disclosure. Accordingly, it will be understood by those skilled in the art that, although not explicitly shown or described herein, many systems and methods that embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure can be devised.

Claims

1. A video encoding method performed by a video encoder: The steps include: calculating the maximum number of candidates in a subblock-based merge candidate list for coding video data; A step of determining a parameter corresponding to the coded video data based on the calculated maximum number of candidates, wherein the parameter is between 0 and an upper limit determined based on sps_sbtmvp_enabled_flag, specifying that if sps_sbtmvp_enabled_flag is equal to 1, the subblock-based temporal motion vector predictor is enabled and can be used in the coded video data, and specifying that if sps_sbtmvp_enabled_flag is equal to 0, the subblock-based temporal motion vector predictor is not used in the coded video data; The steps include: encoding a sample of the current block based on candidate selection from a configured subblock-based merge candidate list for the current block, in response to the current block being in a subblock-based prediction mode, wherein the configured subblock-based merge candidate list for the current block is constrained by the maximum number of candidates in the subblock-based merge candidate list; method.

2. The aforementioned decision-making steps further include: The process further includes the step of determining the parameter by subtracting the maximum number of calculated candidates in the subblock-based merge candidate list from a default number. The method according to claim 1.

3. The default number is 5. The method according to claim 2.

4. The further step includes transmitting the determined parameters in the coded video data, The method according to any one of claims 1 to 3.

5. The step further includes implying the parameter based on a default number, the sps_sbtmvp_enabled_flag, and not transmitting the determined parameter in the coded video data, The method according to claim 1.

6. The sps_sbtmvp_enabled_flag flag is included in the sequence parameter set (SPS) level of the coded video data. The method according to any one of claims 1 to 5.

7. The sps_sbtmvp_enabled_flag is included in the sequence parameter set (SPS) level of the coded video data, and the upper limit of the range is further based on a flag indicating the enabled / disabled state of temporal motion vector prediction at the image header (PH) level of the coded video data. The method according to any one of claims 1 to 5.

8. The steps further include implying the parameter based on a default number, the sps_sbtmvp_enabled_flag, the flag indicating the enabled / disabled state of the temporal motion vector prediction, and not transmitting the parameter in the coded video data. The method according to claim 7.

9. The step further includes, when the determined parameter is equal to the upper limit of the range, the determined parameter and the determined parameter not to be transmitted in the coded video data, The method according to claim 1.

10. A device for video encoding: A processing circuit configured to perform the method described in any one of claims 1 to 9, Device.

11. A non-temporary computer-readable storage medium for storing instructions, wherein, when executed by a computer for video encoding, the instructions cause the computer to perform the method according to any one of claims 1 to 9.

12. A method for storing a coded bitstream, performed by a video encoder: The steps include: coding a video bitstream to generate the coded bitstream; The steps include: storing the coded bitstream in a non-temporary computer-readable storage medium; The step of coding the bitstream of the video to generate the coded bitstream is: The steps include: calculating the maximum number of candidates in a subblock-based merge candidate list for coding the bitstream; A step of determining a parameter corresponding to the coded bitstream based on the calculated maximum number of candidates, wherein the parameter is between 0 and an upper limit determined based on sps_sbtmvp_enabled_flag, specifying that if sps_sbtmvp_enabled_flag is equal to 1, the subblock-based temporal motion vector predictor is enabled and can be used in the coded bitstream, and specifying that if sps_sbtmvp_enabled_flag is equal to 0, the subblock-based temporal motion vector predictor is not used in the coded bitstream; The steps include: encoding a sample of the current block based on candidate selection from a configured subblock-based merge candidate list for the current block, in response to the current block being in a subblock-based prediction mode, wherein the configured subblock-based merge candidate list for the current block is constrained by the maximum number of candidates in the subblock-based merge candidate list; method.

13. A method of transmitting a coded bitstream, performed by a video encoder: The steps include: coding a video bitstream to generate the coded bitstream; The steps include: transmitting the coded bitstream; The step of coding the bitstream of the video to generate the coded bitstream is: The steps include: calculating the maximum number of candidates in a subblock-based merge candidate list for coding the bitstream; A step of determining a parameter corresponding to the coded bitstream based on the calculated maximum number of candidates, wherein the parameter is between 0 and an upper limit determined based on sps_sbtmvp_enabled_flag, specifying that if sps_sbtmvp_enabled_flag is equal to 1, the subblock-based temporal motion vector predictor is enabled and can be used in the coded bitstream, and specifying that if sps_sbtmvp_enabled_flag is equal to 0, the subblock-based temporal motion vector predictor is not used in the coded bitstream; The steps include: encoding a sample of the current block based on candidate selection from a configured subblock-based merge candidate list for the current block, in response to the current block being in a subblock-based prediction mode, wherein the configured subblock-based merge candidate list for the current block is constrained by the maximum number of candidates in the subblock-based merge candidate list; method.