Video decoding method, video decoding device, computer program, and video encoding method

By simplifying the construction of motion vector prediction lists for small blocks in video coding, the proposed method reduces computational complexity and improves decoding efficiency, addressing the challenges faced by existing techniques.

JP2025071096APending Publication Date: 2025-05-02TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025009224
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-08-29
Filing Date
2025-01-22
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Existing video coding techniques face challenges in efficiently constructing motion vector prediction lists, particularly for small blocks, which increases the number of operations and redundancy checks, leading to higher computational complexity.

Method used

The proposed solution involves simplifying the process of building a motion vector prediction list by reducing the number of motion vector predictions and redundancy checks for small blocks. This is achieved by considering the area of the target block and dividing it into smaller blocks if necessary, and constructing the motion vector prediction list based on whether the block area is below a threshold.

Benefits of technology

This approach reduces the computational complexity and improves the efficiency of video decoding by minimizing the number of operations required to construct the motion vector prediction list, especially for small blocks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025071096000001_ABST
    Figure 2025071096000001_ABST
Patent Text Reader

Abstract

To provide a method and device for video encoding / decoding.SOLUTION: Aspects of the disclosure provide methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video decoding includes processing circuit. The processing circuit decodes prediction information for a target block within a target picture. The processing circuit determines whether an area of the target block is smaller than or equal to a threshold. The processing circuit constructs a motion vector prediction list that includes a number of motion vector predictors. The number of motion vector predictors is based on whether the area of the target block is determined to be smaller than or equal to the threshold. The processing circuit reconstructs the target block on the basis of the motion vector predictor list.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Incorporation by Reference This application claims the benefit of priority to U.S. Provisional Application No. 62 / 777,735, entitled "Constructing a Simplified Merge List for Small Coding Blocks," filed on December 10, 2018, which claims the benefit of priority to U.S. Provisional Application No. 16 / 555,549, entitled "Constructing a Simplified Merge List for Small Coding Blocks," filed on August 29, 2019. The entire disclosure of the prior application is incorporated herein by reference in its entirety.

[0002] This disclosure describes embodiments generally related to video coding. [Background technology]

[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The work of the presently named inventors, to the extent that that work is described in this Background section, as well as aspects of the description that would not otherwise qualify as prior art at the time of filing, are not admitted, expressly or impliedly, as prior art to the present disclosure.

[0004] Video encoding and decoding can be done using inter-picture prediction with motion compensation. Uncompressed digital video can include a sequence of pictures, each with spatial dimensions of, for example, 1920x1080 luma samples and associated chroma samples. The sequence of pictures can have a fixed or variable picture rate (also informally called frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has significant bitrate requirements. For example, at 8 bits per sample, 1080p60 4:2:0 video (1920x1080 luma sample resolution at 60 Hz frame rate) requires a bandwidth approaching 1.5 Gbits / s. One hour of such video requires more than 600 Gbytes of storage space.

[0005] One goal of video encoding and decoding may be to reduce redundancy in the input video signal through compression. Compression may help reduce the aforementioned bandwidth or storage space requirements, in some cases by more than two orders of magnitude. Both lossless and lossy compression, as well as combinations of them, may be used. Lossless compression refers to techniques that allow an exact replica of the original signal to be reconstructed from the compressed original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for the intended application. For video, lossy compression is widely adopted. The amount of acceptable distortion varies by application, for example, users of some consumer streaming applications can tolerate higher distortion than users of television distribution applications. The achievable compression ratio may reflect that higher compression ratios can be obtained with higher acceptable / tolerable distortion.

[0006] Video encoders and decoders can employ techniques in several broad categories, such as, for example, motion compensation, transformation, quantization, and entropy coding.

[0007] Video codec techniques may include a technique called intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, the picture may be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, may be used to reset the state of the decoder and therefore may be used as the first picture of a coded video bitstream and video session or as a still image. Samples of an intra block may undergo a transform, and the transform coefficients may be quantized before entropy coding. Intra prediction may be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after the transform and the smaller the AC coefficients, the fewer bits are needed to represent the block after entropy coding for a given quantization step size.

[0008] Conventional intra-coding, for example as known from MPEG-2 generation encoding techniques, does not use intra-prediction. However, some new video compression techniques include approaches that attempt to do so from, for example, surrounding sample data and / or metadata obtained during encoding / decoding that is spatially nearby and preceding in the decoding order of the data block. Such approaches are hereinafter referred to as "intra-prediction" approaches. It should be noted that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed and not from reference pictures.

[0009] Intra prediction can take a variety of forms. If more than one such technique is available for a given video coding technique, the technique used can be coded as an intra prediction mode. In some cases, the mode can include sub-modes and / or parameters, which can be coded separately or included in the mode codeword. The codeword used for a given mode / sub-mode / parameter combination can affect the coding efficiency achieved by intra prediction, and can also affect the entropy coding technique used to convert the codeword into a bitstream.

[0010] Certain modes of intra prediction were introduced in H.264, improved in H.265, and further refined in new coding techniques such as the Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). A prediction block may be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of the neighboring samples are replicated in the prediction block according to a direction. The reference to the direction in use may be coded in the bitstream, or it may be predicted itself.

[0011] Referring to FIG. 1A, shown at the bottom right is a subset of nine prediction directions known from the 33 possible prediction directions (corresponding to the 33 angle modes of the 35 intra modes) of H.265. The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the top right, at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the bottom left of sample (101), at an angle of 22.5 degrees from the horizontal.

[0012] Still referring to FIG. 1A, a square block of 4×4 samples (104) is shown at the top left (indicated by a dashed bold line). The square block (104) contains 16 samples. Each sample is labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample of the block (104) in both the Y and X dimensions. Since the size of the block is 4×4 samples, S44 is at the bottom right. Additionally, reference samples are shown that follow a similar numbering scheme. The reference samples are labeled with R and their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, since the predicted samples are neighbors of the block being reconstructed, there is no need to use negative values.

[0013] Intra-picture prediction may work by copying reference sample values ​​from neighboring samples according to the signaled prediction direction. For example, assuming that the coded video bitstream contains a signal for this block indicating a prediction direction that matches the arrow (102), the sample is predicted from one or more prediction samples to the upper right and at an angle of 45 degrees from the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In some cases, particularly when the orientation is not evenly divisible by 45 degrees, the values ​​of multiple reference samples may be combined, for example by interpolation, to calculate the reference sample.

[0015] As video coding techniques develop, the number of possible directions is increasing. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of disclosure. Experiments have been performed to identify the most likely directions, and some techniques of entropy coding are used to represent these possible directions with a small number of bits, accepting some penalty for less likely directions. Furthermore, there may be cases where the direction itself is predicted from nearby directions used in nearby, already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (105) showing 65 intra prediction directions with JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits in the encoded video bitstream representing the directions may vary from one video coding technique to another, ranging, for example, from a simple direct mapping of prediction directions to complex adaptation schemes involving intra-prediction modes, codewords, most likely modes, and similar techniques. In all cases, however, there may be some directions that are statistically less likely to occur in the video content than some other directions. Since the goal of video compression is to reduce redundancy, in a well-performing video coding technique, these less likely directions are represented by more bits than the more likely directions.

[0018] Motion compensation may be a lossy compression technique, in which blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are spatially shifted in a direction indicated by a motion vector (hereafter MV) and then used to predict a newly reconstructed picture or part of a picture. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third indicating the reference picture in use (the latter may indirectly be the temporal dimension).

[0019] In some video compression techniques, the MV applicable to a region of sample data can be predicted from other MVs, e.g., from MVs associated with other regions of sample data that are spatially adjacent to the region being reconstructed and that precede that MV in the decoding order. Doing so can significantly reduce the amount of data required to encode the MV, thus eliminating redundancy and improving compression ratios. For example, when encoding an input video signal derived from a camera (called natural video), MV prediction can work effectively because regions larger than the region to which a single MV is applicable have a statistical likelihood to move in similar directions and therefore can, in some cases, be predicted using similar motion vectors derived from MVs of nearby regions. As a result, the MV found for a given region is similar or identical to the MV predicted from the surrounding MVs, and then, after entropy encoding, can be represented with fewer bits than would be used to directly encode the MV. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself can be lossy, e.g., due to rounding errors when computing a predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). Of the many MV prediction mechanisms that H.265 offers, the Advanced Motion Vector Prediction (AMVP) mode and the Merge mode are described here.

[0021] In AMVP mode, the motion information of spatially and temporally neighboring blocks of the target block may be used to predict the motion information of the target block, and the prediction residual is further coded. Examples of spatially and temporally neighboring candidates are shown in FIG. 1C and FIG. 1D, respectively. Two candidate motion vector prediction lists are formed. The first candidate prediction is from the first available motion vector of two blocks A0 (112), A1 (113) at the lower left corner of the target block (111), as shown in FIG. 1C. The second candidate prediction is from the first available motion vector of three blocks B0 (114), B1 (115), and B2 (116) above the target block (111). If no valid motion vector is found from the checked locations, the candidate is not put into the list. If two available candidates have the same motion information, only one candidate is kept in the list. If the list is not full, i.e., there are not two distinct candidates in the list, the temporally co-located motion vector (after scaling) from C0 (122) at the bottom right corner of the co-located block (121) in the reference picture is used as another candidate, as shown in FIG. 1D. If motion information at C0 (122) position is not available, the central position C1 (123) of the co-located block in the reference picture is used instead. In the above derivation, if there are still not enough motion vector prediction candidates, a zero motion vector is used to fill the list. Two flags mvp_l0_flag and mvp_l1_flag are signaled in the bitstream, indicating the AMVP index (0 or 1) for the MV candidate lists L0 and L1, respectively.

[0022] In the merge mode of inter-picture prediction, when the merge flag (including the skip flag) is signaled as TRUE, a merge index is signaled to indicate which candidate in the merge candidate list is used to indicate the motion vector of the current block. In the decoder, the merge candidate list is constructed based on the spatial and temporal neighborhood of the current block. As shown in FIG. 1C, up to four MVs derived from five spatial neighboring blocks (A0-B2) are added to the merge candidate list. In addition, as shown in FIG. 1D, up to one MV from two temporally co-located blocks (C0 and C1) in the reference picture is added to the list. The additional merge candidates include combined bi-prediction candidates and zero motion vector candidates. Before obtaining the motion information of a block as a merge candidate, a redundancy check is performed to check whether it is identical to the elements in the current merge candidate list. If it is different from each element in the current merge candidate list, it is added to the merge candidate list as a merge candidate. MaxMergeCandsNum is defined as the size of the merge candidate list in terms of the candidate number. In HEVC, MaxMergeCandsNum is signaled in the bitstream. Skip mode can be thought of as a special merge mode with zero residual.

[0023] In VVC, the sub-block-based temporal motion vector prediction (SbTMVP) method can use the motion fields of the co-located pictures to improve the motion vector prediction and merge mode of CUs in the target picture, similar to the temporal motion vector prediction (TMVP) in HEVC. The same co-located pictures used in TMVP are used in SbTVMP. SbTMVP differs from TMVP in two main ways: (1) TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level; and (2) TMVP fetches temporal motion vectors from co-located blocks in the co-located picture (the co-located blocks are the bottom-right or center blocks with respect to the target CU), while SbTMVP applies a motion shift before fetching the temporal motion information from the co-located picture, and the motion shift is obtained from the motion vector from one of the spatial neighboring blocks to the target CU.

[0024] The SbTVMP process is illustrated in Figure 1E and Figure 1F. SbTMVP predicts motion vectors of sub-CUs in a target CU in two steps. In the first step, the spatial neighborhood of the target block (131) is examined in the following order: A1 (132), B1 (133), B0 (134), and A0 (135), as shown in Figure 1E. If the first available spatial neighboring block with a motion vector that uses the co-located picture as a reference picture is identified, this motion vector is selected as the motion shift to be applied. If no such motion vector is identified from the spatial neighborhood, the motion shift is set to (0, 0).

[0025] In a second step, as shown in FIG. 1F, the motion shift identified in the first step is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (e.g., motion vector and reference index) from the co-located picture. The example of FIG. 1F assumes that the motion shift (149) is set to the motion vector of the spatial neighbor block A1 (143). Then, for a current sub-CU (e.g., sub-CU (144)) in the current block (142) of the current picture (141), the motion information of the corresponding co-located sub-CU (e.g., co-located sub-CU (154)) in the co-located block (152) of the co-located picture (151) is used to derive the motion information of the current sub-CU. The motion information of a corresponding co-located sub-CU (e.g., co-located sub-CU (154)) is converted to a motion vector and reference index for the target sub-CU (e.g., sub-CU (144)) in a manner similar to HEVC's TMVP process, which applies temporal motion scaling to align the reference picture of the temporal motion vector with the reference picture of the target CU.

[0026] In VVC, a combined subblock-based merge list containing both SbTVMP and affine merge candidates can be used in the subblock-based merge mode. The SbTVMP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction is added as the first entry in the subblock-based merge list, followed by the affine merge candidates. In some applications, the maximum allowed size of the subblock-based merge list is 5. The subCU size used in SbTMVP is fixed, for example, 8x8. As in the affine merge mode, the SbTMVP mode can be applied to a CU only if both its width and height are 8 or more.

[0027] The coding logic of the additional SbTMVP merging candidate is the same as that of the other merging candidates: for each CU in a P or B slice, an additional rate-distortion (RD) check is performed to determine whether to use the SbTMVP candidate.

[0028] In VVC, the history-based MVP (HMVP) method includes HMVP candidates defined as the motion information of previously coded blocks. A table containing multiple HMVP candidates is maintained during the encoding / decoding process. Whenever a new slice is detected, the table is emptied. Whenever there is an inter-coded non-affine block, the associated motion information is added to the last entry of the table as a new HMVP candidate. The coding flow of the HMVP method is shown in Figure 1G.

[0029] The table size,S,is set to 6, which indicates that up to six HMVP candidates can be added to the table.,When inserting a new motion candidate into the table, a constrained FIFO rule is used, such that a,redundancy check is first applied to determine whether an identical,HMVP is in the table. If found, the identical HMVP is removed from the table,,and then all HMVP candidates are moved forward, i.e., their index is,decreased by one. Figure 1H shows an example of inserting a new,motion candidate into the HMVP table.

[0030] The HMVP candidates may be used in the process of building a merge candidate list. The most recent few HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidate. Pruning is applied to the HMVP candidates into spatial or temporal merge candidates except for sub-block motion candidates (i.e., ATMVP).

[0031] To reduce the number of pruning operations, the number of HMVP candidates to be checked (denoted as L) is set as L=(N<=4)?M:(8-N), where N denotes the number of available non-subblock merge candidates in the table and M denotes the number of available HMVP candidates. Furthermore, the process of constructing a merge candidate list from the HMVP list ends when the total number of available merge candidates reaches the notified maximum allowed merge candidate minus 1. Furthermore, the number of pairs of combined bi-predictive merge candidate derivations is reduced from 12 to 6.

[0032] HMVP candidates can also be used in the AMVP candidate list construction process. The motion vectors of the last K HMVP candidates in the table are inserted after the TMVP candidates. Only HMVP candidates with the same reference picture as the AMVP target reference picture are used to construct the AMVP candidate list. Pruning is applied to the HMVP candidates. In some applications, K is set to 4, but the size of the AMVP list is not changed, which is 2.

[0033] The pairwise average candidate is generated by averaging pairs of predefined candidates in the current merge candidate list. In VVC, the number of pairwise average candidates is 6, and the predefined pairs are defined as {(0,1), (0,2), (1,2), (0,3), (1,3), (2,3)}, where the numbers indicate the merge index into the merge candidate list. The averaged motion vector is calculated separately for each reference list. If both motion vectors are available in one list, these two motion vectors are averaged even if they point to different reference pictures. If only one motion vector is available, one motion vector is used directly. If no motion vector is available, this list is considered invalid. The pairwise average candidate can be replaced by the combined candidate in the HEVC standard.

[0034] Multi-hypothesis prediction may be used to improve the single prediction of the AMVP mode. One flag is signaled to enable or disable multi-hypothesis prediction. In addition, one additional merge index is signaled if the flag is true. In this way, multi-hypothesis prediction turns the single prediction into a bi-prediction, where one prediction is obtained using the original syntax elements of the AMVP mode, and the other prediction is obtained using the merge mode. The final prediction combines these two predictions using a 1:1 weight, as in the case of bi-prediction. The merge candidate list is first derived from the merge mode, where the sub-CU candidates (such as affine, alternative temporal motion vector prediction (ATMVP)) are excluded. Then, the merge candidate list is split into two separate lists, one for list 0 (L0) containing all L0 motions from the candidates, and another for list 1 (L1) containing all L1 motions. After removing redundancies and filling the spaces, two merge lists are generated for L0 and L1, respectively. There are two constraints when applying multi-hypothesis prediction to improve the AMVP mode. First, it is valid for CUs with Luminance Coding Block (CB) regions equal to or greater than 64. Second, it only applies to L1 for low latency B pictures. Summary of the Invention [Problem to be solved by the invention]

[0035] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a receiving circuit and a processing circuit.

[0036] The processing circuit decodes prediction information for a target block in a target picture that is part of an encoded video sequence. The processing circuit determines whether an area of ​​the target block is less than or equal to a threshold. The processing circuit constructs a motion vector prediction list including a number of motion vector predictions. The number of motion vector predictions is based on whether the area of ​​the target block is determined to be less than or equal to the threshold. The processing circuit reconstructs the target block based on the motion vector prediction list. [Means for solving the problem]

[0037] According to an aspect of the present disclosure, the number of motion vector predictions included in the motion vector prediction list is a first number when it is determined that the area of ​​the target block is less than or equal to a threshold, and the number of motion vector predictions included in the motion vector prediction list is a second number greater than the first number when it is determined that the area of ​​the target block is greater than the threshold.

[0038] According to an aspect of the present disclosure, the prediction information indicates that the target block is further divided into a plurality of smaller blocks. The processing circuit divides the target block into the plurality of smaller blocks. The processing circuit determines whether an area of ​​one of the plurality of smaller blocks is less than or equal to a threshold. The processing circuit constructs a motion vector prediction list based on whether an area of ​​one of the plurality of smaller blocks is determined to be less than or equal to the threshold. The number of motion vector predictions included in the motion vector prediction list is a first number if the area of ​​one of the plurality of smaller blocks is determined to be less than or equal to the threshold, and the number of motion vector predictions included in the motion vector prediction list is a second number if the area of ​​one of the plurality of smaller blocks is determined to be greater than the threshold.

[0039] According to an aspect of the present disclosure, the prediction information indicates a partition size of the current block. The processing circuit determines whether the partition size is less than or equal to a threshold. The processing circuit constructs a motion vector prediction list based on whether the partition size is determined to be less than or equal to the threshold. The number of motion vector predictions included in the motion vector prediction list is a first number if the partition size is determined to be less than or equal to the threshold, and the number of motion vector predictions included in the motion vector prediction list is a second number if the partition size is determined to be greater than the threshold.

[0040] In one embodiment, the threshold is a preset number of luminance samples.

[0041] In one embodiment, the threshold corresponds to the picture resolution of the target picture.

[0042] In one embodiment, the threshold value is signaled in the encoded video sequence.

[0043] In one embodiment, a motion vector prediction list including a first number of motion vector predictions is constructed based on a first set of spatial or temporal motion vector predictions, the first set of spatial or temporal motion vector predictions being smaller than a second set of spatial or temporal motion vector predictions used to construct a motion vector prediction list including a second number of motion vector predictions.

[0044] In one embodiment, the motion vector prediction list that includes the first number of motion vector predictions does not include a spatial or temporal motion vector prediction.

[0045] In one embodiment, a first number of redundancy checks are performed to build a motion vector prediction list including the first number of motion vector predictions, and a second number of redundancy checks are performed to build a motion vector prediction list including the second number of motion vector predictions, the first number of redundancy checks being less than the second number of redundancy checks.

[0046] In one embodiment, the first number of redundancy checks does not include at least one of a comparison between (i) the possible spatial motion vector prediction and a first existing spatial motion vector prediction in the motion vector prediction list, (ii) between the possible temporal motion vector prediction and a first existing temporal motion vector prediction in the motion vector prediction list, (iii) between the possible history-based motion vector prediction and a first or other existing spatial motion vector prediction in the motion vector prediction list, and (iv) between the possible history-based motion vector prediction and a first or other existing temporal motion vector prediction in the motion vector prediction list.

[0047] In one embodiment, the motion vector prediction list including the first number of motion vector predictions is constructed without a redundancy check.

[0048] In one embodiment, at least one of the motion vector predictions uses the current picture as a reference picture such that at least one reference block of the motion vector prediction is within the current picture.

[0049] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any one or combination of methods for video decoding.

[0050] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0051] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 2 is a diagram of an example intra-prediction direction. [Figure 1C] FIG. 2 is a schematic diagram of a target block and its surrounding spatial merging candidates in one example. [Figure 1D] FIG. 1 is a schematic diagram of co-located blocks and temporal merging candidates in an example. [Figure 1E] 1 is a schematic diagram of a current block and its surrounding spatial merging candidates for sub-block-based temporal motion vector prediction (SbTMVP) according to an example. [Figure 1F] 1 is an exemplary process for deriving SbTMVP candidates according to one example. [Figure 1G] 1 is a decoding flow of a history-based motion vector prediction (HMVP) method in one example. [Figure 1H] 1 is an exemplary process for updating a table in HMVP according to one example. [Diagram 2] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system according to one embodiment. [Diagram 3]FIG. 1 is a schematic diagram of a simplified block diagram of a communication system according to one embodiment. [Figure 4] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 6] 4 shows a block diagram of an encoder according to another embodiment; [Figure 7] 4 shows a block diagram of a decoder according to another embodiment; [Figure 8] 1 shows a flowchart outlining an example process according to some embodiments of the present disclosure. [Figure 9] 1 shows another flowchart outlining an example process according to some embodiments of the present disclosure. [Figure 10] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0052] FIG. 2 shows a simplified block diagram of a communication system (200) according to an embodiment of the present disclosure. The communication system (200) includes, for example, a plurality of terminal devices capable of communicating with each other via a network (250). For example, the communication system (200) includes a first pair of terminal devices (210) and (220) interconnected via the network (250). In the example of FIG. 2, the first pair of terminal devices (210) and (220) perform unidirectional transmission of data. For example, the terminal device (210) may encode video data (e.g., a stream of video pictures captured by the terminal device (210)) for transmission to another terminal device (220) via the network (250). The encoded video data may be transmitted in the form of one or more encoded video bitstreams. The terminal device (220) may receive the encoded video data from the network (250), decode the encoded video data to recover the video pictures, and display the video pictures according to the recovered video data. One-way data transmission may be common, such as in media serving applications.

[0053] In another example, the communication system (200) includes a second pair of terminal devices (230) and (240) performing bidirectional transmission of encoded video data, which may occur, for example, during a video conference. In the case of bidirectional transmission of data, in one example, each of the terminal devices (230) and (240) can encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (230) and (240) via the network (250). Each of the terminal devices (230) and (240) can also receive encoded video data transmitted by the other of the terminal devices (230) and (240), decode the encoded video data to recover video pictures, and display the video pictures on an accessible display device according to the recovered video data.

[0054] In the example of FIG. 2, the terminal devices (210), (220), (230), and (240) may be depicted as a server, a personal computer, and a smartphone, although the principles of the present disclosure need not be so limited. The embodiments of the present disclosure find application in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (250) represents any number of networks that convey encoded video data between the terminal devices (210), (220), (230), and (240), including, for example, wireline and / or wireless communication networks. The communication network (250) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, the Internet, and the like. For purposes of this discussion, the architecture and topology of the network (250) may not be important to the operation of the present disclosure unless otherwise described herein below.

[0055] 3 shows an arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0056] The streaming system may include a capture subsystem (313), which may include a video source (301), such as a digital camera, for example, that creates a stream of uncompressed video pictures (302). In one example, the stream of video pictures (302) includes samples captured by a digital camera. The stream of video pictures (302) is shown with a thick line to emphasize the amount of data compared to the encoded video data (304) (or encoded video bitstream) and may be processed by an electronic device (320) that includes a video encoder (303) coupled to the video source (301). The video encoder (303) may include hardware, software, or a combination thereof for enabling or implementing aspects of the disclosed subject matter, as described in more detail below. The encoded video data (304) (or encoded video bitstream (304)) is shown with a thin line to emphasize the amount of data compared to the stream of video pictures (302) and may be stored on a streaming server (305) for future use. One or more streaming client subsystems, such as the client subsystems (306) and (308) of FIG. 3, can access the streaming server (305) to retrieve copies (307) and (309) of the encoded video data (304). The client subsystem (306) may include, for example, a video decoder (310) within an electronic device (330). The video decoder (310) decodes the incoming copy of the encoded video data (307) and creates an output stream (311) of video pictures that may be rendered on a display (312) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (304), (307), and (309) (e.g., a video bitstream) may be encoded according to a number of video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. As an example, a video encoding standard under development is informally known as VVC.The disclosed subject matter may be used in the context of a VVC.

[0057] It should be noted that the electronic devices (320) and (330) may include other components (not shown). For example, the electronic device (320) may include a video decoder (not shown), and the electronic device (330) may also include a video encoder (not shown).

[0058] 4 shows a block diagram of a video decoder (410) according to an embodiment of the present disclosure. The video decoder (410) may be included in an electronic device (430). The electronic device (430) may include a receiver (431) (e.g., a receiving circuit). The video decoder (410) may be used in place of the video decoder (310) in the example of FIG. 3.

[0059] The receiver (431) can receive one or more coded video sequences to be decoded by the video decoder (410), in the same or other embodiments, one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences can be received from a channel (401), which can be a hardware / software link to a storage device that stores the coded video data. The receiver (431) can receive the coded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which can be forwarded to respective usage entities (not shown). The receiver (431) can separate the coded video sequences from the other data. To combat network jitter, a buffer memory (415) can be coupled between the receiver (431) and the entropy decoder / parser (420) (hereinafter "parser (420)"). In some applications, the buffer memory (415) is part of the video decoder (410). In other cases, it may be outside the video decoder (410) (not shown). In still other cases, there may be a buffer memory (not shown) outside the video decoder (410), e.g., to combat network jitter, and yet another buffer memory (415) inside the video decoder (410), e.g., to handle playback timing. When the receiver (431) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from a synchronous network, the buffer memory (415) may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (415) may be needed, may be relatively large, may be advantageously sized, and may be implemented at least in part in an operating system or similar element (not shown) outside the video decoder (410).

[0060] The video decoder (410) may include a parser (420) for reconstructing symbols (421) from the encoded video sequence. These categories of symbols include information used to manage the operation of the video decoder (410) and, in some cases, information for controlling a rendering device (e.g., a display screen), such as the rendering device (412), which may not be an integral part of the electronic device (430) but may be coupled to the electronic device (430) as shown in FIG. 4. The control information for the rendering device may be in the form of a supplemental enhancement information (SEI message) or a video usability information (VUI) parameter set fragment (not shown). The parser (420) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may follow a video encoding technique or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (420) can extract from the encoded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include Group of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (420) can also extract information from the encoded video sequence such as transform coefficients, quantization parameter values, motion vectors, etc.

[0061] The parser (420) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (415) to produce symbols (421).

[0062] The reconstruction of the symbols (421) may involve a number of different units, depending on the type of encoded video picture or part thereof (inter and intra pictures, inter and intra blocks, etc.), and other factors. Which units are involved and how may be controlled by subgroup control information parsed from the encoded video sequence by the parser (420). The flow of such subgroup control information between the parser (420) and the following units is not shown for clarity.

[0063] In addition to the functional blocks already mentioned, the video decoder (410) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0064] The first unit is a scalar / inverse transform unit (451), which receives control information from the parser (420) including the transform to be used, block size, quantization coefficients, quantization scaling matrix, etc., as well as the quantized transform coefficients as symbols (421). The scalar / inverse transform unit (451) may output a block containing sample values ​​that can be input to an aggregator (455).

[0065] In some cases, the output samples of the scaler / inverse transform (451) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (452). In some cases, the intra-picture prediction unit (452) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from a current picture buffer (458). The current picture buffer (458) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (455) adds, in some cases, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (452) to the output sample information provided by the scaler / inverse transform unit (451).

[0066] In other cases, the output samples of the scalar / inverse transform unit (451) may relate to an inter-coded, possibly motion-compensated block. In such cases, the motion compensated prediction unit (453) may access a reference picture memory (457) to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols (421) associated with the block, these samples may be added by an aggregator (455) to the output of the scalar / inverse transform unit (451) to generate output sample information (in this case referred to as residual samples or residual signals). The addresses in the reference picture memory (457) from which the motion compensated prediction unit (453) fetches prediction samples may be controlled by motion vectors available to the motion compensated prediction unit (453) in the form of symbols (421), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (457) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0067] The output samples of the aggregator (455) may be subjected to various loop filtering techniques in the loop filter unit (456). Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video sequence (also called coded video bitstream) and available to the loop filter unit (456) as symbols (421) from the parser (420), but may also be responsive to meta-information obtained during decoding of previous (decoding order) parts of the coded picture or coded video sequence, or may be responsive to previously reconstructed and loop filtered sample values.

[0068] The output of the loop filter unit (456) may be a sample stream that may be output to a rendering device (412) and may be stored in a reference picture memory (457) for use in future inter-picture prediction.

[0069] Some coded pictures, once fully reconstructed, may be used as reference pictures for future predictions. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (420)), the current picture buffer (458) may become part of the reference picture memory (457), and a new current picture buffer may be reallocated before beginning reconstruction of the next coded picture.

[0070] The video decoder (410) may perform decoding operations according to a given video compression technique in a standard such as ITU-T Recommendation H.265. The encoded video sequence may conform to a syntax specified by the video compression technique or standard used in the sense that it conforms to both the syntax of the video compression technique or standard and the profile as a document of the video compression technique or standard. Specifically, a profile may select some tools from all tools available in the video compression technique or standard as the only tools that can be used in that profile. Compliance also requires that the complexity of the encoded video sequence is within a level defined in the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstructed sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited in some cases by the specification of a hypothetical reference decoder (HRD) and metadata for HRD buffer management signaled in the encoded video sequence.

[0071] In one embodiment, the receiver (431) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (410) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0072] 5 shows a block diagram of a video encoder (503) according to an embodiment of the present disclosure. The video encoder (503) is included in an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmission circuit). The video encoder (503) may be used in place of the video encoder (303) in the example of FIG. 3.

[0073] The video encoder (503) can receive video samples from a video source (501) (which is not part of the electronic device (520) in the example of FIG. 5), which can capture video images that are encoded by the video encoder (503). In other examples, the video source (501) is part of the electronic device (520).

[0074] The video source (501) may provide a source video sequence that is encoded by the video encoder (503) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (501) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (501) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that give motion when viewed in sequence. The pictures themselves may be organized as an array of spatial pixels, each of which may contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art will readily appreciate the relationship between pixels and samples. The following description will focus on samples.

[0075] According to one embodiment, the video encoder (503) may encode and compress pictures of a source video sequence into an encoded video sequence (543) in real time or under any other time constraint required by an application. Enforcing an appropriate encoding rate is one function of the controller (550). In some embodiments, the controller (550) controls and is operatively coupled to other functional units as described below. For clarity, coupling is not shown. Parameters set by the controller (550) may include rate control related parameters (e.g., picture skip, quantization, lambda value for rate distortion optimization techniques), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. The controller (550) may be configured to have other appropriate functions related to the optimized video encoder (503) for a certain system design.

[0076] In some embodiments, the video encoder (503) is configured to operate in an encoding loop. As an oversimplified explanation, in one example, the encoding loop may include a source coder (530) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture), and a (local) decoder (533) embedded in the video encoder (503). The decoder (533) reconstructs the symbols to create sample data in a similar manner as the (remote) decoder does (since the compression between the symbols and the encoded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory (534). Since the decoding of the symbol stream results in a bit-perfect result regardless of the location of the decoder (local or remote), the content of the reference picture memory (534) is also bit-perfect between the local and remote encoders. In other words, the prediction part of the encoder "sees" the samples of the reference picture as exactly the same sample values ​​that the decoder "sees" when it uses the prediction during decoding. This basic principle of reference picture synchrony (and the drift that occurs when synchrony cannot be maintained, e.g., due to channel errors) is also used in several related technologies.

[0077] The operation of the "local" decoder (533) may be the same as that of a "remote" decoder, such as the video decoder (410), already described in detail in connection with Figure 4. However, with brief reference also to Figure 4, because symbols are available and the encoding / decoding of symbols into an encoded video sequence by the entropy coder (545) and parser (420) may be lossless, the entropy decoding portion of the video decoder (410), as well as the buffer memory (415) and parser (420), may not be fully implemented in the local decoder (533).

[0078] At this point, it can be noted that any decoder techniques other than analysis / entropy decoding present in a decoder must necessarily be present in the corresponding encoder in substantially identical functional form. For this reason, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder techniques can be omitted, since they are the inverse of the decoder techniques described generically. Only a few areas require more detailed description, which are provided below.

[0079] In operation, in some cases, the source coder (530) may perform motion-compensated predictive encoding, which predictively encodes an input picture with reference to one or more previously encoded pictures from the video sequence designated as “reference pictures.” In this manner, the encoding engine (532) encodes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0080] The local video decoder (533) may decode the encoded video data of pictures that may be designated as reference pictures based on the symbols created by the source coder (530). The operation of the encoding engine (532) may advantageously be a lossy process. When the encoded video data may be decoded in a video decoder (not shown in FIG. 5), the reconstructed video sequence may be a replica of the source video sequence, usually with some errors. The local video decoder (533) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in a reference picture cache (534). In this way, the video encoder (503) may locally store copies of reconstructed reference pictures that have common content (no transmission errors) as reconstructed reference pictures obtained by the far-end video decoder.

[0081] The predictor (535) may perform the prediction search of the coding engine (532). That is, for a new picture to be coded, the predictor (535) may search the reference picture memory (534) for sample data (as candidate reference pixel blocks) or certain metadata of the reference pictures, such as motion vectors, block shapes, etc., which may serve as suitable prediction references for the new picture. The predictor (535) may operate on each pixel block of sample blocks to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (535), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (534).

[0082] The controller (550) may manage the encoding operations of the source coder (530), including, for example, setting parameters and subgroup parameters used to encode the video data.

[0083] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (545), which converts the symbols produced by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0084] The transmitter (540) may buffer the encoded video sequence created by the entropy coder (545) for transmission over a communication channel (560), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (540) may merge the encoded video data from the video coder (503) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0085] The controller (550) may manage the operation of the video encoder (503). During encoding, the controller (550) may assign a coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, in many cases, pictures may be assigned one of the following picture types:

[0086] An intra picture (I-picture) may be one that can be coded and decoded without using other pictures in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, such as, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0087] A predictive picture (P picture) may be one that can be encoded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block.

[0088] Bidirectionally predicted pictures (B-pictures) may be those that can be encoded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0089] A source picture is usually spatially divided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to the respective picture of the block. For example, blocks of an I picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). Pixel blocks of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0090] The video encoder (503) may perform encoding operations according to a given video encoding technique or standard, such as ITU-T Recommendation H.265. In its operations, the video encoder (503) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to a syntax specified in the video encoding technique or standard being used.

[0091] In one embodiment, the transmitter (540) can transmit additional data along with the encoded video. The source coder (530) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0092] A video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits correlations (temporal or otherwise) between pictures. In one example, a particular picture being coded / decoded, called a current picture, is divided into blocks. If a block of the current picture is similar to a reference block of a reference picture in a previously coded and still buffered video, the block of the current picture may be coded by a vector called a motion vector. A motion vector points to a reference block in a reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0093] In some embodiments, a bidirectional prediction technique may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures, such as a first reference picture and a second reference picture, both of which precede a current picture in a video in decoding order (but may be in the past and future, respectively, in display order), are used. A block in the current picture may be coded by a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block may be predicted by a combination of the first and second reference blocks.

[0094] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0095] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs of a picture are of the same size, such as 64×64 pixels, 32×32 pixels, 16×16 pixels, etc. In general, a CTU includes three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels may be partitioned into one CU of 64×64 pixels, four CUs of 32×32 pixels, or sixteen CUs of 16×16 pixels. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is partitioned into one or more prediction units (PUs) depending on the temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using luma prediction blocks as an example of prediction blocks, the prediction blocks include matrices of pixel values ​​(such as luma values) of 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0096] 6 shows a diagram of a video encoder (603) according to another embodiment of the present disclosure. The video encoder (603) is configured to receive a processed block (e.g., a predictive block) of sample values ​​in a target video picture in a sequence of video pictures and encode the processed block into an encoded picture that is part of an encoded video sequence. In one example, the video encoder (603) is used in place of the video encoder (303) in the example of FIG. 3.

[0097] In an HEVC example, the video encoder (603) receives a matrix of sample values ​​for a processing block, such as a predictive block, such as 8×8 samples. The video encoder (603) determines whether the processing block is best coded using intra mode, inter mode, or bidirectional prediction mode, for example, using rate-distortion optimization. If the processing block is coded in intra mode, the video encoder (603) may code the processing block into a coded picture using an intra prediction technique, and if the processing block is coded in inter mode or bidirectional prediction mode, the video encoder (603) may code the processing block into a coded picture using an inter prediction or bidirectional prediction technique, respectively. In some video coding techniques, the merge mode may be an inter-picture prediction submode in which a motion vector is derived from one or more motion vector predictions without benefit of coded motion vector components outside the predictors. In some other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (603) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0098] In the example of FIG. 6, the video encoder (603) includes an inter-encoder (630), an intra-encoder (622), a residual calculator (623), a switch (626), a residual encoder (624), a general controller (621), and an entropy encoder (625) coupled together as shown in FIG. 6.

[0099] The inter-encoder (630) is configured to receive a sample of a current block (e.g., a processing block), compare the block to one or more reference blocks of reference pictures (e.g., blocks of previous and subsequent pictures), generate inter-prediction information (e.g., inter-coding technique, motion vectors, description of redundant information by merge mode information), and calculate an inter-prediction result (e.g., a prediction block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the coded video information.

[0100] The intra encoder (622) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with already encoded blocks in the same picture, generate quantized coefficients after transformation, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (622) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.

[0101] The general controller (621) is configured to determine general control data and control other components of the video encoder (603) based on the general control data. In one example, the general controller (621) determines the mode of the block and provides a control signal to the switch (626) based on the mode. For example, if the mode is an intra mode, the general controller (621) controls the switch (626) to select an intra mode result to be used by the residual calculator (623), controls the entropy encoder (625) to select intra prediction information, and includes the intra prediction information in the bitstream, and if the mode is an inter mode, the general controller (621) controls the switch (626) to select an inter prediction result to be used by the residual calculator (623), controls the entropy encoder (625) to select inter prediction information, and includes the inter prediction information in the bitstream.

[0102] The residual calculator (623) is configured to calculate a difference (residual data) between the received block and a prediction result selected from the intra-encoder (622) or the inter-encoder (630). The residual encoder (624) is configured to operate on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (624) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. Then, a quantization process is performed on the transform coefficients to obtain quantized transform coefficients. In various embodiments, the video encoder (603) also includes a residual decoder (628). The residual decoder (628) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used by the intra-encoder (622) and the inter-encoder (630) as appropriate. For example, the inter-encoder (630) may generate decoded blocks based on the decoded residual data and the inter-prediction information, and the intra-encoder (622) may generate decoded blocks based on the decoded residual data and the intra-prediction information. The decoded blocks are appropriately processed to generate decoded pictures, which may be buffered in a memory circuit (not shown) and used as reference pictures in some examples.

[0103] The entropy encoder (625) is configured to format the bitstream to include the encoded block. The entropy encoder (625) is configured to include various information according to an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (625) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. It should be noted that, according to the disclosed subject matter, there is no residual information when encoding a block in a merged sub-mode of either the inter mode or the bi-prediction mode.

[0104] 7 shows a diagram of a video decoder (710) according to another embodiment of the present disclosure. The video decoder (710) is configured to receive encoded pictures that are part of an encoded video sequence and decode the encoded pictures to generate reconstructed pictures. In one example, the video decoder (710) is used in place of the video decoder (310) in the example of FIG. 3.

[0105] In the example of FIG. 7, the video decoder (710) includes an entropy decoder (771), an inter-decoder (780), a residual decoder (773), a reconstruction module (774), and an intra-decoder (772) coupled together as shown in FIG.

[0106] The entropy decoder (771) may be configured to reconstruct from the coded picture a number of symbols representing syntax elements that make up the coded picture. Such symbols may include, for example, prediction information (e.g., intra- or inter-prediction information) that may identify the mode in which the block is coded (e.g., intra- or inter-prediction mode, inter- or bi-prediction mode in merged or other submodes), certain samples or metadata used for prediction by the intra- or inter-decoder (772) or inter-decoder (780), respectively, residual information, for example in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter- or bi-prediction mode, the inter-prediction information is provided to the inter-decoder (780), and if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra-decoder (772). The residual information may undergo inverse quantization and is provided to the residual decoder (773).

[0107] The inter decoder (780) is configured to receive the inter prediction information and to generate inter prediction results based on the inter prediction information.

[0108] The intra decoder (772) is configured to receive the intra prediction information and to generate a prediction result based on the intra prediction information.

[0109] The residual decoder (773) is configured to perform inverse quantization to extract inverse quantized transform coefficients and process the inverse quantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require some control information (to include quantization parameters (QPs)), which may be provided by the entropy decoder (771) (data path not shown, as it is only low volume control information).

[0110] The reconstruction module (774) is configured to combine, in the spatial domain, the residual output by the residual decoder (773) and a prediction result (possibly output by an inter- or intra-prediction module) to form a reconstructed block, which may be part of a reconstructed picture, which may be part of a reconstructed video. It should be noted that other suitable operations, such as a deblocking operation, may be performed to improve visual quality.

[0111] It should be noted that the video encoders (303), (503), and (603) and the video decoders (310), (410), and (710) may be implemented using any suitable technique. In one embodiment, the video encoders (303), (503), and (603) and the video decoders (310), (410), and (710) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (303), (503), and (603) and the video decoders (310), (410), and (710) may be implemented using one or more processors executing software instructions.

[0112] For blocks coded in non-subblock mode, the motion vector predictor list can be constructed as follows: (1) candidate predictors from spatially and temporally neighboring blocks, (2) candidate predictions from the HMVP buffer, (3) candidate predictions from a pair-wise average of existing motion vector predictions, and (4) zero motion vector predictions using different reference pictures. Examples of these candidate predictors are given above.

[0113] When constructing a motion vector prediction list, a series of operations may be performed, including availability checks and redundancy checks. The availability check checks whether the spatial or temporal neighboring blocks are coded in inter prediction mode. The redundancy check checks whether the new motion vector prediction overlaps with an existing motion vector prediction in the motion vector prediction list. For example, the availability check and redundancy check are performed for spatial and temporal candidates. The redundancy check is performed for existing candidates. In addition, a redundancy check is performed for existing spatial / temporal candidates for candidates from the HMVP buffer. If the motion vector prediction list is long, performing these operations may require several cycles. In the worst case, the picture is divided into many small blocks. The total number of cycles required to complete the construction of the motion vector prediction list for all these small blocks may be more than desired. For example, the merge candidate list for each coding block may not be the same as the neighboring blocks. Therefore, it is desirable to reduce the number of operations and simplify the construction process of the motion vector prediction list, especially for small blocks.

[0114] The present disclosure provides an improved technique (e.g., merge mode, skip mode, or AMVP mode) for simplifying the construction process of the motion vector prediction list when some block size constraints are satisfied. That is, when the target block is considered as a small block, the construction process of the motion vector prediction list of the target block can be simplified to include fewer motion vector predictions and / or perform fewer availability check and / or redundancy check operations.

[0115] Various thresholds may be used to indicate a small block. In one embodiment, if the block area of ​​the target block is less than or equal to the threshold, the target block may be considered a small block. In another embodiment, the target block is further divided into multiple smaller blocks, and if one of the multiple smaller blocks is less than or equal to the threshold, the target block may be considered a small block. The threshold may be a pre-set number of luma samples, such as 32 or 64 luma samples. The threshold may correspond to the picture resolution of the target picture. For example, for a higher resolution picture, the threshold may be larger compared to the threshold for a lower resolution picture. Furthermore, the threshold may be signaled in the bitstream or coded video sequence, such as in a sequence parameter set (SPS), picture parameter set (PPS), slice or tile header, etc.

[0116] According to aspects of the present disclosure, when a target block is considered to be a small block, the process of constructing a motion vector prediction list for the target block can be simplified to include a reduced number of motion vector predictions compared to blocks that are not considered to be small blocks.

[0117] Each of the motion vector prediction categories, such as spatial motion vector prediction, temporal motion vector prediction, and history-based motion vector prediction, may include multiple motion vector predictions. In one embodiment, if the target block is considered to be a small block, the number of motion vector predictions from one or more of these categories may be reduced. For example, for each motion vector prediction category, a subset of each of the multiple motion vector predictions is included in the motion vector prediction list, so that the motion vector prediction list is shorter than including all motion vector predictions in the list. The subset of motion vector predictions may be the first N (e.g., N=1, 2, etc.) candidates from each or some categories, where N is an integer smaller than the number of possible candidates of the respective category. However, other predetermined candidates may be selected in other embodiments. The number N of each or some categories may be the same or different. Furthermore, the selection of the subset of motion vector predictions of each or some categories may be the same or different. In one example, only the first N spatial and / or temporal candidates are allowed in the motion vector prediction list, where the number N is smaller than the number of possible spatial and temporal candidates. In some examples, N may be different for spatial and temporal candidates.

[0118] In one embodiment, for spatial and temporal motion vector prediction categories, a subset of spatial and temporal motion vector predictions are included in the motion vector prediction list, followed by HMVP candidates, etc. In another embodiment, the motion vector prediction list does not include spatial and / or temporal motion vector predictions, i.e., spatial and / or temporal candidates are not allowed and the motion vector prediction list starts with an HMVP candidate.

[0119] According to aspects of the present disclosure, when a current block is considered a small block, the process of building a motion vector prediction list for the current block can be simplified to perform fewer operations, such as a redundancy check operation, which compares a new motion vector prediction with existing motion vector predictions in the motion vector prediction list.

[0120] In one embodiment, one or more redundancy checks between two spatial and / or temporal candidates can be eliminated.For example, the motion vector prediction list does not perform a redundancy check between a possible spatial motion vector prediction and an existing spatial motion vector prediction in the motion vector prediction list.In another example, the motion vector prediction list does not perform a redundancy check between a possible temporal motion vector prediction and an existing temporal motion vector prediction in the motion vector prediction list.

[0121] In one embodiment, one or more redundancy checks between candidates from the HMVP buffer and existing candidates from spatial and / or temporal neighboring positions can be eliminated.For example, the motion vector prediction list does not perform a redundancy check between possible history-based motion vector predictions and existing spatial motion vector predictions in the motion vector prediction list.In another example, the motion vector prediction list does not perform a redundancy check between possible history-based motion vector predictions and existing temporal motion vector predictions in the motion vector prediction list.

[0122] In one embodiment, no redundancy check is performed, i.e., the motion vector prediction list is built without performing a redundancy check.

[0123] 8 shows a flow chart outlining an example process (800) for simplifying the motion vector prediction list construction process according to some embodiments of the present disclosure. In various embodiments, the process (800) is performed by processing circuitry of the terminal devices (210), (220), (230) and (240), processing circuitry performing the functions of the video encoder (303), processing circuitry performing the functions of the video decoder (310), processing circuitry performing the functions of the video decoder (410), processing circuitry performing the functions of the intra prediction module (452), processing circuitry performing the functions of the video encoder (503), processing circuitry performing the functions of the predictor (535), processing circuitry performing the functions of the intra encoder (622), processing circuitry performing the functions of the intra decoder (772), etc. In some embodiments, the process (800) is implemented in software instructions, such that the processing circuitry performs the process (800) when the processing circuitry executes the software instructions.

[0124] The process (800) may generally begin at step (S801), where the process (800) determines whether to further divide the target block. If it is determined not to further divide the target block, the process (800) proceeds to step (S802), otherwise the process (800) proceeds to step (S803).

[0125] In step (S802), the process (800) determines whether the block area of ​​the target block is equal to or less than the threshold value. If it is determined that the block area of ​​the target block is equal to or less than the threshold value, the process (800) proceeds to step (S805); otherwise, the process (800) proceeds to step (S806).

[0126] In step (S803), the process (800) divides the current block into multiple smaller blocks and then proceeds to step (S804).

[0127] In step (S804), the process (800) determines whether the block area of ​​one of the multiple smaller blocks is less than or equal to a threshold. The threshold may be a preset number of luma samples, may correspond to the picture resolution of the target picture, or may be signaled in the encoded video sequence. If it is determined that the block area of ​​one of the multiple smaller blocks is less than or equal to the threshold, the process (800) proceeds to step (S805); otherwise, the process (800) proceeds to step (S806).

[0128] In some embodiments, step (S801) and / or step (S803) are optional and may not be performed. For example, if the decoded prediction information of the target block indicates a partition size of the target block, the partition size indicates that the target block is further divided into multiple smaller blocks. The process (800) determines whether the partition size is equal to or smaller than a threshold. This procedure is the same as step (S804). If it is determined that the partition size is equal to or smaller than the threshold, the process (800) proceeds to step (S805), otherwise, the process proceeds to step (S806). In other embodiments, the comparison of the block with the threshold is performed after the block division has already been performed.

[0129] In step (S805), the process (800) builds a motion vector prediction list including a first number of motion vector predictions, and then proceeds to step (S807).

[0130] In step (S806), the process (800) constructs a motion vector prediction list including a second number of motion vector predictions. The second number is greater than the first number. That is, the motion vector prediction list is simplified by reducing the second number of motion vector predictions to the first number of motion vector predictions, for example in the manner described above. The process (800) then proceeds to step (S807).

[0131] In step (S807), the process (800) reconstructs the current block based on the motion vector prediction list, after which the process (800) ends.

[0132] 9 shows a flow chart outlining an example process (900) according to some embodiments of the present disclosure. In various embodiments, the process (900) is performed by processing circuitry of the terminal devices (210), (220), (230) and (240), processing circuitry performing the functions of the video encoder (303), processing circuitry performing the functions of the video decoder (310), processing circuitry performing the functions of the video decoder (410), processing circuitry performing the functions of the intra prediction module (452), processing circuitry performing the functions of the video encoder (503), processing circuitry performing the functions of the predictor (535), processing circuitry performing the functions of the intra encoder (622), processing circuitry performing the functions of the intra decoder (772), etc. In some embodiments, the process (900) is implemented with software instructions, and thus the processing circuitry performs the process (900) when the processing circuitry executes the software instructions.

[0133] The process (900) may generally begin at step (S901), where the process (900) decodes prediction information for a current block in a current picture that is part of an encoded video sequence. After decoding the prediction information, the process (900) proceeds to step (S902).

[0134] In step (S902), the process (900) determines whether the area of ​​the current block is less than or equal to a threshold. Then, the process (900) proceeds to step (S903). The threshold may be a pre-set number of luma samples, such as 32 or 64 luma samples. The threshold may correspond to a picture resolution of the current picture. The threshold may be signaled in the encoded video sequence.

[0135] In step (S903), the process (900) constructs a motion vector prediction list including a number of motion vector predictions. The number of motion vector predictions is based on whether the area of ​​the target block is determined to be equal to or less than a threshold. In one embodiment, if the area of ​​the target block is determined to be equal to or less than a threshold, the number of motion vector predictions included in the motion vector prediction list is a first number, and if not, the number of motion vector predictions included in the motion vector prediction list is a second number greater than the first number.

[0136] In some embodiments, the prediction information indicates that the target block is further divided into a plurality of smaller blocks, and the process (900) divides the target block into the plurality of smaller blocks and determines whether an area of ​​one of the plurality of smaller blocks is equal to or less than a threshold. The process (900) further builds a motion vector prediction list based on whether an area of ​​one of the plurality of smaller blocks is determined to be equal to or less than a threshold. If an area of ​​one of the plurality of smaller blocks is determined to be equal to or less than a threshold, the number of motion vector predictions included in the motion vector prediction list is a first number, and otherwise, the number of motion vector predictions included in the motion vector prediction list is a second number.

[0137] In some embodiments, the prediction information indicates a partition size of the current block. In such embodiments, the process (900) determines whether the partition size is less than or equal to a threshold. The process (900) further builds a motion vector prediction list based on whether the partition size is determined to be less than or equal to the threshold. If the partition size is determined to be less than or equal to the threshold, the number of motion vector predictions included in the motion vector prediction list is a first number, and if not, the number of motion vector predictions included in the motion vector prediction list is a second number.

[0138] In one embodiment, a motion vector prediction list including a first number of motion vector predictions is constructed based on a first set of spatial or temporal motion vector predictions, the first set of spatial or temporal motion vector predictions being smaller than a second set of spatial or temporal motion vector predictions used to construct a motion vector prediction list including a second number of motion vector predictions.

[0139] In one embodiment, the motion vector prediction list that includes the first number of motion vector predictions does not include a spatial or temporal motion vector prediction.

[0140] In one embodiment, some redundancy checks are reduced to build a motion vector prediction list including a first number of motion vector predictions. For example, a first number of redundancy checks are performed to build a motion vector prediction list including a first number of motion vector predictions, and a second number of redundancy checks are performed to build a motion vector prediction list including a second number of motion vector predictions. The first number of redundancy checks is less than the second number of redundancy checks.

[0141] In one embodiment, the first number of redundancy checks does not include at least one of a comparison between (i) the possible spatial motion vector prediction and a first existing spatial motion vector prediction in the motion vector prediction list, (ii) between the possible temporal motion vector prediction and a first existing temporal motion vector prediction in the motion vector prediction list, (iii) between the possible history-based motion vector prediction and a first or other existing spatial motion vector prediction in the motion vector prediction list, and (iv) between the possible history-based motion vector prediction and a first or other existing temporal motion vector prediction in the motion vector prediction list.

[0142] In one embodiment, the motion vector prediction list including the first number of motion vector predictions is constructed without a redundancy check.

[0143] In one embodiment, at least one of the motion vector predictions uses the current picture as a reference picture such that at least one reference block of the motion vector prediction is within the current picture.

[0144] In step (S904), the process (900) reconstructs the current block based on the motion vector prediction list.

[0145] After reconstructing the subject block, the process (900) ends.

[0146] It should be noted that while the above embodiments are based on reducing the number of redundancy checks in addition to the number of motion vector predictions, in other embodiments the number of redundancy checks may be reduced without reducing the number of motion vector predictions.

[0147] The motion vector prediction method can use a picture different from the target picture as a reference picture. However, the motion vector prediction or block compensation can be performed from a previously reconstructed region in the target picture. Such motion vector prediction or block compensation can be called intra-picture block compensation, target picture reference (CPR), or intra-block copy (IBC). In the IBC prediction mode, the displacement vector indicating the offset between the target block and the reference block in the target picture is called a block vector (BV). The reference block is already reconstructed before the target block. It is noted that the BV can be considered as the MV described in this application.

[0148] The techniques described above can be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, Figure 10 illustrates a computer system (1000) suitable for implementing some embodiments of the disclosed subject matter.

[0149] Computer software can be encoded using any suitable machine code or computer language and can be subject to assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly, or through interpretation, execution of microcode, etc.

[0150] The instructions may be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0151] 10 for the computer system (1000) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having a dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of the computer system (1000).

[0152] The computer system (1000) may include several human interface input devices. Such human interface input devices may be responsive to input by one or more users, such as, for example, tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown), etc. The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as sound (speech, music, environmental sounds, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), and video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0153] The input human interface devices may include a keyboard (1001), a mouse (1002), a trackpad (1003), a touch screen (1010), a data glove (not shown), a joystick (1005), a microphone (1006), a scanner (1007), a camera (1008), etc. (only one of each is shown).

[0154] The computer system (1000) may also include several human interface output devices. Such human interface output devices may stimulate one or more of the user's senses, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (1010), data gloves (not shown), or joystick (1005), although there may also be haptic feedback devices that do not function as input devices), audio output devices (such as speakers (1009), headphones (not shown)), visual output devices (such as screens (1010), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which may output two-dimensional visual output or more than three-dimensional output via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown). These visual output devices, such as a screen (1010), may be connected to the system bus (1048) via a graphics adapter (1050).

[0155] The computer system (1000) may also include human accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1020) along with media such as CDs / DVDs (1021), thumb drives (1022), removable hard drives or solid state drives (1023), legacy magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), etc.

[0156] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not encompass transmission media, carrier waves, or other transitory signals.

[0157] The computer system (1000) may also include a network interface (1054) to one or more communication networks (1055). The one or more communication networks (1055) may be, for example, wireless, wired, optical. Additionally, the one or more communication networks (1055) may be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of the one or more networks (1055) include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable television, satellite television, terrestrial broadcast television, vehicular, industrial including CANBus, etc. Some networks typically require an external network interface adapter connected to some general data port or peripheral bus (1049) (e.g., a USB port on the computer system (1000)), while others are typically integrated into the core of the computer system (1000) by connecting to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1000) can communicate with other entities. Such communication may be unidirectional, receive only (e.g., broadcast TV), unidirectional transmit only (e.g., a CANbus to some CANbus devices), or bidirectional, e.g., communication to other computer systems using local area or wide area digital networks. As noted above, several protocols and protocol stacks may be used with each of these networks and network interfaces.

[0158] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be connected to the core (1040) of the computer system (1000).

[0159] The cores (1040) may include one or more central processing units (CPUs) (1041), graphics processing units (GPUs) (1042), specialized programmable processing units in the form of field programmable gate areas (FPGAs) (1043), hardware accelerators (1044) for some tasks, etc. These devices may be connected via a system bus (1048), along with read-only memory (ROM) (1045), random access memory (1046), internal mass storage devices (1047) such as internal hard drives, SSDs, etc. that are not accessible to the user. In some computer systems, the system bus (1048) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (1048) or via a peripheral bus (1049). Peripheral bus architectures include PCI, USB, etc.

[0160] The CPU (1041), GPU (1042), FPGA (1043), and accelerator (1044) may execute a number of instructions that may combine to constitute the aforementioned computer code. That computer code may be stored in a ROM (1045) or a RAM (1046). Persistent data may be stored, for example, in an internal mass storage device (1047), while transient data may also be stored in the RAM (1046). Rapid storage and retrieval from any memory device may be enabled by the use of a cache memory, which may be closely associated with one or more of the CPUs (1041), GPUs (1042), mass storage devices (1047), ROMs (1045), RAMs (1046), etc.

[0161] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0162] As an example and not by way of limitation, the architecture, particularly the computer system (1000) having the core (1040), can provide functionality as a result of a processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be the user-accessible mass storage devices introduced above, as well as media associated with some storage of the core (1040) that is non-transitory in nature, such as the core internal mass storage (1047) or ROM (1045). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1040). The computer-readable media can include one or more memory devices or chips according to particular needs. The software can cause the core (1040) and particularly the processor therein (including CPU, GPU, FPGA, etc.) to perform certain processes or certain parts of certain processes described herein, including the definition of data structures stored in RAM (1046) and modifying such data structures according to the software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embedded in circuitry (e.g., accelerator (1044)) that may operate in place of or together with software to perform particular processes or particular portions of particular processes described herein. References to software may include logic, and vice versa, as appropriate. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that store software for execution, circuitry embodying logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0163] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that are within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.

[0164] Appendix A: Acronyms AMVP: Advanced Motion Vector Prediction ASIC: Application-Specific Integrated Circuit BMS:Benchmark Set BS:Boundary Strength BV: Block Vector CANBus: Controller Area Network Bus CD: Compact Disc CPR: Current Picture Referencing CPU: Central Processing Units CRT:Cathode Ray Tube CTB: Coding Tree Blocks CTU: Coding Tree Units CU: Coding Unit DPB: Decoder Picture Buffer DVD: Digital Video Disc FPGA: Field Programmable Gate Areas GOP:Groups of Pictures GPU: Graphics Processing Units GSM: Global System for Mobile communications (pan-European digital mobile communications system) HDR: High Dynamic Range HEVC: High Efficiency Video Coding HRD: Hypothetical Reference Decoder IBC: Intra Block Copy IC: Integrated Circuit JEM: Joint Exploration Model LAN: Local Area Network LCD: Liquid-Crystal Display LIC: Local Illumination Compensation LTE: Long-Term Evolution MR-SAD: Mean-Removed Sum of Absolute Difference MR-SATD: Mean-Removed Sum of Absolute (Hadamard-Transformed Difference) MV: Motion Vector OLED: Organic Light-Emitting Diode PB: Prediction Blocks PCI: Peripheral Component Interconnect PLD: Programmable Logic Device PPS: Picture Parameter Set PU: Prediction Units RAM: Random Access Memory ROM: Read-Only Memory SCC: Screen Content Coding SDR: Standard Dynamic Range SEI: Supplementary Enhancement Information SMVP: Spatial Motion Vector Predictor SNR: Signal Noise Ratio SPS: Sequence Parameter Set SSD: Solid-state drive TMVP: Temporal Motion Vector Predictor TU: Transform Units USB: Universal Serial Bus VUI: Video Usability Information VVC:Versatile Video Coding [Explanation of symbols]

[0165] 111 Target Block 121 Blocks in the same location 131 Target Block 141 Target Picture 142 Target Block 144 Sub-CU 149 Movement Shift 151 Same Position Picture 152 Blocks in the same location 154 Sub-CU 200 Communication Systems 210 Terminal Devices 220 Terminal Devices 230 Terminal Devices 240 Terminal Devices 250 Communication Network 301 Video Source 302 Video Picture Stream 303 Video Encoder 304 Encoded Video Data 305 Streaming Server 306 Client Subsystem 307 Copy 308 Client Subsystem 309 Incoming copy, video data 310 Video Decoder 311 Video Picture Output Stream 312 Display 313 Capture Subsystem 320 Electronic Devices 330 Electronic Devices 401 Channel 410 Video Decoder 412 Rendering Device 415 Buffer Memory 420 Parser 421 Symbols 430 Electronic Devices 431 Receiver 451 Scaler / Descaler Unit 452 Intra Picture Prediction Unit, Intra Prediction Module 453 Motion Compensation Prediction Unit 455 Aggregator 456 Loop Filter Unit 457 Reference Picture Memory 458 Target Picture Buffer 501 Video Sources 503 Video Encoder 520 Electronic Devices 530 Source Coder 532 encoding engine 533 Decoder 534 Reference Picture Memory 535 Predictor 540 Transmitter 543 coded video sequence 545 Entropy Coder 550 Controller 560 Communication Channels 603 Video Encoder 621 General Controller 622 Intra Encoder 623 Residual Calculator 624 Residual Encoder 625 Entropy Encoder 626 Switch 628 Residual Decoder 630 InterEncoder 710 Video Decoder 771 Entropy Decoder 772 Intra Decoder 773 Residual Decoder 774 Reconstruction Module 780 Inter Decoder 800 processes 900 Processes 1000 Computer Systems, Architecture 1001 Keyboard 1002 Mouse 1003 Trackpad 1005 Joystick 1006 Mike 1007 Scanner 1008 Camera 1009 Speaker 1010 Touch Screen 1020 CD / DVD ROM / RW 1021 CD / DVD 1022 Thumb Drive 1023 Removable Hard Drive or Solid State Drive 1040 cores 1041 Central Processing Unit (CPU) 1042 Graphics Processing Unit (GPU) 1043 Field Programmable Gate Area (FPGA) 1044 Hardware Accelerator 1045 Read-Only Memory (ROM) 1046 Random Access Memory 1047 Core Internal Mass Storage 1048 System Bus 1049 Surrounding Bus 1050 graphics adapter 1054 Network Interface 1055 Communication Network

Claims

1. 1. A method for video decoding in a decoder, comprising: decoding prediction information for a current block in a current picture that is part of an encoded video sequence; determining whether the area of ​​the target block is less than or equal to a threshold; constructing a motion vector prediction list including a number of motion vector predictions, the number of motion vector predictions being based on whether the area of ​​the current block is determined to be less than or equal to the threshold; reconstructing the current block based on the motion vector prediction list; A method comprising:

2. the number of motion vector predictions included in the motion vector prediction list is a first number if it is determined that the area of ​​the current block is less than or equal to the threshold; if the area of ​​the current block is determined to be greater than the threshold, the number of motion vector predictions included in the motion vector prediction list is a second number greater than the first number. The method of claim 1.

3. When the prediction information indicates a partition size of the target block, the construction includes: constructing the motion vector prediction list based on whether the partition size is determined to be less than or equal to the threshold; the number of motion vector predictions included in the motion vector prediction list is the first number if it is determined that the partition size is less than or equal to the threshold; if the partition size is determined to be greater than the threshold, the number of motion vector predictions included in the motion vector prediction list is the second number. The method of claim 2.

4. The method of claim 1 , wherein the threshold is a preset number of luminance samples.

5. The method of claim 1 , wherein the threshold corresponds to a picture resolution of the target picture.

6. The method of claim 1 , wherein the threshold value is signaled in the encoded video sequence.

7. 3. The method of claim 2, wherein the motion vector prediction list including the first number of motion vector predictions is constructed based on a first set of spatial or temporal motion vector predictions, the first set of spatial or temporal motion vector predictions being smaller than a second set of spatial or temporal motion vector predictions used to construct the motion vector prediction list including the second number of motion vector predictions.

8. The method of claim 2 , wherein the motion vector prediction list including the first number of motion vector predictions does not include a spatial or temporal motion vector prediction.

9. 3. The method of claim 2, further comprising: performing a first number of redundancy checks to build the motion vector prediction list including the first number of motion vector predictions; and performing a second number of redundancy checks to build the motion vector prediction list including the second number of motion vector predictions, the first number of redundancy checks being less than the second number of redundancy checks.

10. 10. The method of claim 9, wherein the first number of redundancy checks does not include at least one of a comparison between (i) a possible spatial motion vector prediction and a first existing spatial motion vector prediction in the motion vector prediction list, (ii) a possible temporal motion vector prediction and a first existing temporal motion vector prediction in the motion vector prediction list, (iii) a possible history-based motion vector prediction and the first or other existing spatial motion vector prediction in the motion vector prediction list, and (iv) a comparison between the possible history-based motion vector prediction and the first or other existing temporal motion vector prediction in the motion vector prediction list.

11. The method of claim 2 , wherein the motion vector prediction list including the first number of motion vector predictions is constructed without a redundancy check.

12. The method of claim 1 , wherein at least one of the motion vector predictions uses the current picture as a reference picture.

13. 1. An apparatus for video decoding, comprising: Decoding prediction information for a current block in a current picture that is part of an encoded video sequence; determining whether the area of ​​the object block is less than or equal to a threshold; constructing a motion vector prediction list including a number of motion vector predictions, the number of motion vector predictions being based on whether the area of ​​the current block is determined to be less than or equal to the threshold; a processing circuit configured to reconstruct the current block based on the motion vector prediction list; 23. An apparatus for video decoding comprising:

14. the number of motion vector predictions included in the motion vector prediction list is a first number if it is determined that the area of ​​the current block is less than or equal to the threshold; if the area of ​​the current block is determined to be greater than the threshold, the number of motion vector predictions included in the motion vector prediction list is a second number greater than the first number.

14. The apparatus of claim 13.

15. The processing circuitry includes: if the prediction information indicates a target partition size, and constructing the motion vector prediction list based on whether the partition size is determined to be less than or equal to the threshold. the number of motion vector predictions included in the motion vector prediction list is the first number if it is determined that the partition size is less than or equal to the threshold; if the partition size is determined to be greater than the threshold, the number of motion vector predictions included in the motion vector prediction list is the second number.

15. The apparatus of claim 14.

16. The apparatus of claim 13 , wherein the threshold is a preset number of luminance samples.

17. The apparatus of claim 13 , wherein the threshold corresponds to a picture resolution of the target picture.

18. The apparatus of claim 13 , wherein the threshold value is signaled in the encoded video sequence.

19. 15. The apparatus of claim 14, wherein the motion vector prediction list including the first number of motion vector predictions is constructed based on a first set of spatial or temporal motion vector predictions, the first set of spatial or temporal motion vector predictions being smaller than a second set of spatial or temporal motion vector predictions used to construct the motion vector prediction list including the second number of motion vector predictions.

20. decoding prediction information for a current block in a current picture that is part of an encoded video sequence; determining whether the area of ​​the target block is less than or equal to a threshold; constructing a motion vector prediction list including a number of motion vector predictions, the number of motion vector predictions being based on whether the area of ​​the current block is determined to be less than or equal to the threshold; reconstructing the current block based on the motion vector prediction list; A non-transitory computer-readable storage medium storing a program executable by at least one processor to perform the method.

Citation Information

Patent Citations

  • VIDEO DECODING METHOD, VIDEO DECODING APPARATUS, COMPUTER PROGRAM, AND VIDEO ENCODING METHOD

    JP7625667B2