Constrained Motion Vector Derivation for Long-Term Reference Images in Video Encoding and Decoding

By determining if reference images are long-term and imposing constraints on inter-mode encoding and decoding tools, the method addresses the lack of constraints in current video coding standards, enhancing motion vector derivation and overall video coding efficiency.

JP7688186B2Active Publication Date: 2025-06-03BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024028537
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-02-20
Filing Date
2024-02-28
Publication Date
2025-06-03
Estimated Expiration
2040-02-19

AI Technical Summary

Technical Problem

Current video coding standards, such as HEVC and VVC, lack well-defined constraints for the scaling process of spatial and temporal motion candidates, particularly when long-term reference images are involved in new inter-mode encoding and decoding tools.

Method used

The method involves determining whether reference images related to inter-mode encoded blocks are long-term reference images and imposing specific constraints on the operation of inter-mode encoding and decoding tools, such as prohibiting certain operations like scaling or using invalid motion vectors.

Benefits of technology

This approach enhances the reliability and accuracy of motion vector derivation for inter-mode encoded blocks, improving the overall efficiency and quality of video coding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007688186000002
    Figure 0007688186000002
  • Figure 0007688186000003
    Figure 0007688186000003
  • Figure 0007688186000004
    Figure 0007688186000004
Patent Text Reader

Abstract

To provide a method of constraining an operation of a specific inter-mode encoding and decoding tool in deriving a motion vector candidate for an inter-mode encoded block employed in a video encoding standard like a current versatile video encoding and decoding (VVC).SOLUTION: In a method executed by a computing device, the computing device determines whether or not one or more of reference images related to an inter-mode encoded block associated with an operation of an inter-mode encoding and decoding tool are a long period reference image, and constrains, on the basis of the determination, the operation of the inter-mode encoding and decoding tool for the inter-mode encoded block.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference to related applications

[0001] This application claims priority to U.S. Provisional Application No. 62 / 808,271, filed on February 20, 2019, and the entire specification of this patent application is incorporated herein by reference. and is cited herein.

Technical Field

[0002] The present disclosure generally relates to video encoding, decoding, and compression. In particular, it relates to a system and method for performing video encoding and decoding using constraints on motion vector derivation for long - term reference images.

Background Art

[0003] Here, background information related to the present disclosure is provided. The information contained herein should not necessarily be construed as prior art.

[0004] Video data can be compressed by any of various video encoding and decoding techniques. Video encoding and decoding can be performed according to one or more video encoding standards. Exemplary video encoding and decoding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM) coding, High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), and Moving Picture Experts Group (MPEG).

[0005] ​​​​​​​​​​In video encoding and decoding, generally, prediction methods (e.g., inter prediction, intra prediction, etc.) based on the redundancy inherent in a video image or sequence are utilized. One of the goals of video encoding and decoding technology is to compress video data into a form with a lower bit rate while avoiding or minimizing the degradation of video quality. The prediction methods used in video encoding and decoding usually perform spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or eliminate the redundancy inherent in video data, and are associated with block-based video encoding and decoding. In block-based video encoding, the input video signal is processed block by block. For each block (also called a coding unit (CU)), spatial prediction and / or temporal prediction can be performed. Spatial prediction (also called "intra prediction") predicts the current block using pixels from samples of already-encoded adjacent blocks (also called reference samples) within the same video image / slice. Spatial prediction reduces the spatial redundancy inherent in the video signal.

[0006] Temporal prediction (also called "inter prediction" or "motion compensation prediction") predicts the current block using reconstructed pixels from already-encoded video images. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a particular CU usually indicates the amount and direction of motion between the current CU and its temporal reference.

[0007]

[0008]

[0009] ​​​​​​​​​​​​​is signaled by a motion vector (MV). Also, when multiple reference images are supported, one reference picture index for identifying from which reference picture in the reference picture memory the temporal prediction signal is, is additionally transmitted.

[0010] After spatial and / or temporal prediction, the mode decision block in the encoder selects the best prediction mode, for example based on a rate-distortion optimization method. Next, the prediction block is subtracted from the current block; the prediction residual is decorrelated by a transform and quantized. The quantized residual coefficients are inverse quantized and inverse transformed to form the reconstruction residual, and this reconstruction residual is then added to the prediction block to form the reconstructed signal of this block.

[0011] After spatial and / or temporal prediction, in-loop filtering is performed. For example, non-blocking filtering, sample adaptive offset (SAO), and adaptive loop filter (ALF) are applied to the reconstructed CU, and then the reconstructed CU is put into the reference picture memory and used for the encoding / decoding of future video blocks. To form the output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit, and further compressed and packed to form the bitstream. During decoding processing, the video bitstream is first entropy decoded by the entropy decoding unit.

[0012] During decoding processing, the video bitstream is first entropy decoded by the entropy decoding unit. ​​​The coding mode and prediction information are sent to the spatial predictor (in the case of intra coding). Or sent to a temporal predictor (in the case of inter-coding) to form a prediction block. The residual transform coefficients are sent to an inverse quantification unit and an inverse transform unit to reconstruct the residual block. Then the prediction block and the residual block are added together. The reconstructed block is , and may be stored in the reference image store after further in-loop filtering. The reconstructed video in the reference image store is then sent to drive a display device. or used to predict future video blocks.

[0013] In video coding standards such as HEVC and VVC, a reference picture set (RPS) is used. The concept of a "picture set" is that previously decoded images are used as references, i.e., sample data prediction and motion. The decoded picture buffer (DPB) is used for vector prediction. In general, the outline of the RPS for reference image management is as follows: The main point is the DPB in each slice (also called a "tile" in the current VVC). The status of the

[0014] The images in the DPB are classified as "used for short-term reference", "used for long-term reference", or "used for short-term reference". Images can be marked as "not used for reference". If a field is marked as "unused," it will no longer be available for prediction and will no longer be needed in the output. When the file is deleted, it can be removed from the DPB.

[0015] In general, long-term reference pictures are usually sorted by display order (i.e., picture order count). far from the current image in terms of (also called r Count or POC) the short-term reference image. The distinction between the long-term reference image and the short-term reference image affects some decoding processes such as motion vector scaling in spatial and temporal MV prediction or implicit weighted prediction and can be given.

[0016] In video coding and decoding standards such as HEVC and VVC, when deriving spatial and / or temporal motion vector candidates, specific constraints are placed on the scaling process that forms part of the derivation of spatial and / or temporal motion vector candidates, based on whether the specific reference image involved in the process is a long-term reference image.

[0017] However, according to current video codec specifications such as the VVC standardization, similar constraints are not placed on such new inter-mode video coding and decoding tools for the motion vector candidate derivation of blocks that are still inter-mode encoded. placed. SUMMARY OF THE INVENTION

[0018] Here, a general overview of the present disclosure is provided, which is not a complete disclosure of its full scope or all of its features.

[0019] According to a first aspect of the present disclosure, a method for video coding and decoding is executed on a computing device having one or more processors and a memory storing a plurality of programs executed by the one or more processors. The method includes partitioning each image in a video stream into a plurality of blocks or coding units (CUs). This method further includes performing inter-mode motion vector derivation for these inter-mode encoded blocks. This method further includes operating specific inter-mode encoding tools while performing inter-mode motion vector derivation for the inter-mode encoded blocks. This method further includes determining whether one or more of the reference images related to the inter-mode encoded blocks involved in the operation of the inter-mode decoding tool are long-term reference images, and based on the

[0020] determination, further restricting the operation of the inter-mode decoding tool for the inter-mode encoded blocks described above. According to a second aspect of the present application, a computing device includes one or more processors, a

[0021] memory, and a plurality of programs stored in the memory. When executed by the one or more processors, the programs cause the computing device to perform the operations as described above. According to a third aspect of the present application, a non-transitory computer-readable storage medium stores a

[0022] plurality of programs executed by a computing device having one or more processors. When executed by the one or more processors, It is possible, and all such modifications can be included within the technical scope of the present disclosure. Without contradiction, the teachings of different embodiments, although not necessarily so, can be combined with each other. When there is no contradiction, the teachings of different embodiments, although not necessarily so, can be combined with each other. When there is no contradiction, the teachings of different embodiments, although not necessarily so, can be combined with each other.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7A

Figure 7B

Embodiments for Carrying out the Invention

[0023] The terms used in this disclosure are for illustrative purposes of specific examples and are not intended to limit this disclosure. As used in this disclosure and the appended claims, the singular forms "a", "one", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The term "and / or" as used herein is to be understood to refer to any and all possible combinations of one or more of the associated listed items.

[0024] Here, terms such as "first", "second", "third", etc. can be used to describe various information, but it should be understood that such information should not be limited by such terms. These terms are only used to distinguish one category of information from another. For example, without departing from the scope of this disclosure, the first information can be called the second information, and similarly, the second information can also be called the first information. As used herein, the terms "if...then", "if...", or "if...and" can, depending on the context, mean "when..." or "in response to...".

[0025] In this specification, references in the singular or plural to "one embodiment", "an embodiment", "another embodiment", or the like mean that one or more specific features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of this disclosure. Thus, the phrases "in one embodiment", "in an example", "in a certain embodiment", and similar expressions that appear in the singular or plural in multiple places throughout this specification do not necessarily all They do not all refer to the same embodiment. Furthermore, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable way.

[0026] Conceptually, many video coding standards are similar and include those described in the background art. For example, almost all video coding standards use block-based processing and share similar video coding block diagrams to achieve video compression.

[0027] FIG. 1 shows a block diagram of an exemplary block-based hybrid video encoder 100 that can be used in combination with many video coding / decoding standards. The encoder 10 0 partitions a video frame into a plurality of video blocks for processing. For each given video block, a prediction is formed based on an inter-prediction approach or an intra-prediction approach. In inter-prediction, one or more predictors are formed by motion estimation and motion compensation based on pixels from a previously reconstructed frame. In intra-prediction, a predictor is formed based on the reconstructed pixels in the current frame. Through mode decision, the best predictor for predicting the current block can be selected. The prediction residual representing the difference between the current video block and its predictor is sent to the transform circuit 102. Then, for entropy reduction, the transform coefficients are sent from the transform circuit 102 to the quantization circuit 104. Next, the quantized coefficients are supplied to the entropy coding circuit 106 to generate a compressed video bitstream. As shown in FIG. 1, for inter-prediction

[0028] The prediction residual representing the difference between the current video block and its predictor is sent to the transform circuit 102. And, for entropy reduction, the transform coefficients are sent from the transform circuit 102 to the quantization circuit 104. Next, the quantized coefficients are supplied to the entropy coding circuit 106 to generate a compressed video bitstream. As shown in FIG. 1, for inter-prediction Video block partitioning information, motion vectors from the loop and / or intra prediction circuit 112, prediction-related information 110 such as reference picture indexes, and intra prediction modes are also supplied via the entropy encoding circuit 106 and stored in the compressed video bitstream 114. In the encoder 100, decoder-related circuits are also necessary to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed via the inverse quantization 116 and inverse transform circuit 118. This reconstructed prediction residual is combined with the block predictor 120 to generate the unfiltered reconstructed pixels of the current video block. In general, in-loop filters are used to improve the coding / decoding efficiency and visual quality. For example, current versions of AVC, HEVC, and VVC provide non-blocking filters. In HEVC, an additional in-loop filter called SAO (Sample Adaptive Offset) is defined to further improve the coding efficiency. In the current version of the VVC standard, another in-loop filter called ALF (Adaptive Loop Filter) is actively being studied and is likely to be included in the final standard. These in-loop filter operations are selectable. Performing these operations improves the coding / decoding efficiency and visual quality. On the other hand, they can be turned off according to the decision of the encoder 100 to save computational complexity.

[0029]

[0030]

[0031]

[0032] Note that when these filter options are turned on by the encoder 100, the intra prediction is usually based on the unfiltered reconstructed pixels, while the inter prediction is based on the filtered reconstructed pixels.

[0033] FIG. 2 is a block diagram showing an exemplary video decoder 200 that can be used in combination with many video encoding / decoding standards. This decoder 200 is similar to the reconstruction-related part present in the encoder 100 of FIG. 1. In the decoder 200 (FIG. 2), the input video bitstream 201 is first decoded through entropy decoding 202 to derive the quantized coefficient levels and prediction-related information. Next, the quantized coefficient levels are processed through inverse quantization 204 and inverse transformation 206 to obtain the reconstructed prediction residual. The intra / inter mode selection unit 212 implements a block prediction mechanism that is configured to perform intra prediction 208 or motion compensation 210 based on the decoded prediction information. The reconstructed prediction residual obtained from the inverse transformation 206 and the prediction output generated by the block predictor mechanism are added by the adder 214 to obtain a set of unfiltered reconstructed pixels. When the in-loop filter is on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video. Next, the reconstructed video in the reference picture memory unit is sent to drive a display device or used to predict future video blocks.

[0034] In video coding and decoding standards such as HEVC, blocks can be partitioned based on a quadtree. This is possible. In new video coding and decoding standards such as the current VVC, more partitioning methods are adopted, and one coding tree unit (CTU) can be partitioned into coding units (CUs) based on a quadtree, a binary tree, or a ternary tree to adapt to various local characteristics. In the current VVC, the distinction between CUs, prediction units (PUs), and transform units (TUs) does not exist in most coding modes, and each CU is always used as the basic unit for both prediction and transform without further partitioning. However, in specific coding modes such as the intra-subpartition coding mode, there are cases where each CU contains multiple TUs. In the multi-type tree structure, one CTU is first partitioned by a quadtree structure. Next, each quadtree leaf node can be further partitioned by binary tree and ternary tree structures.

[0035]

[0036]

[0037] Figure 3 shows the five splitting types adopted in the current VVC, namely, quaternary partitioning 301, horizontal binary partitioning 302, vertical binary partitioning 303, horizontal ternary partitioning 304, and vertical ternary partitioning 305.

[0036]

[0037] In video coding and decoding standards such as HEVC and the current VVC, previously decoded images are managed in a decoded picture buffer (DPB) and used as references in the concept of a reference picture set (RPS). The images in the DPB can be marked as "used for short-term reference", "used for long-term reference", or "not used for reference".If the predetermined target reference image of the current block is different from the reference image of the adjacent block, then the spatially scaled motion vector of the adjacent block can be used as a prediction child of the motion vector of the current block. In the scaling process for spatial motion candidates, the scaling factor is calculated based on the picture order count (POC) distance between the current image and the target reference image, and the POC distance between the current image and the reference image of the adjacent block. In video coding / decoding standards such as HEVC and the current VVC, specific constraints are imposed on the scaling process for spatial

[0038] motion candidates, based on whether a specific reference image involved in the process is a long-term reference image. If one of the two reference images is a long-term reference image and the other is not, the MV of the adjacent block is considered invalid. If both of the two reference images are long-term reference images, the POC distance between these two long-term reference images is usually large, and therefore the scaled MV may be unreliable. Therefore, the MV of the spatial adjacent block is directly used as the MVP of the current block, and the scaling process is prohibited. Similarly, in the scaling process for temporal motion candidates, the scaling factor is calculated based on the POC distance between the current image and the target reference image, and the POC distance between the parallel image and the reference image of the temporal adjacent block (also called the parallel block).

[0039]

[0040] ​​​​​In video encoding and decoding standards such as HEVC and the current VVC, specific constraints are imposed on the scaling process for temporal motion candidates based on whether a specific reference image involved in the process is a long-term reference image. When one of the two reference images is a long-term reference image and the other is not, the MV of adjacent blocks is regarded as invalid. When both of the two reference images are long-term reference images, the POC distance between these two long-term reference images is usually large, and thus the scaled MV may be unreliable. Therefore, temporal the MV of adjacent blocks is directly used as the MVP of the current block, and the scaling process is prohibited.

[0041] In new video encoding and decoding standards such as the current VVC, new inter-mode encoding decoding tools have been introduced. Some examples of the new inter-mode encoding tools are bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), motion mode using MVD (MMVD), symmetric MVD (SMVD), bi-prediction with weighted averaging (BWA: Bi-prediction with Weighted Averaging), pairwise averaging merge candidate derivation, and sub-block based temporal motion vector prediction (SbTMVP).

[0042] Conventional bi-prediction in video encoding and decoding is a simple combination of two temporal prediction blocks obtained from already reconstructed reference images. However, due to the limitation of block-based motion compensation, there is a possibility of observing the remaining small motion between samples of the two prediction blocks, so the efficiency of motion-compensated prediction decreases. To solve this problem, BDOF is applied to the current VVC to reduce the influence of such motion for each sample within a block.

[0043] Figure 4 shows an example of BDOF processing. BDOF is per-sample motion refinement that is performed on block-based motion compensated prediction when dual prediction is used. The motion refinement for each 4×4 sub-block is calculated by minimizing the difference between the reference picture list 0 (L0) prediction sample and the reference picture list 1 (L1) prediction sample after BDOF is applied within one 6×6 window around the sub-block. Based on the motion refinement thus derived, the final dual prediction sample of the CU is calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on the optical flow model. ... ... ... ... ... ...

[0044] DMVR is a dual prediction technique for merge blocks that have two MVs that are first signaled and can be further refined by dual matching prediction. ...

[0045] Figure 5 shows an example of dual matching used in DMVR. Dual matching is used to derive the motion information of the current CU by finding the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures. The cost function used in the matching process is the sum of absolute difference (SAD) of the downsampled rows. After the matching process is completed, the refined MVs are used for motion compensation in the prediction stage, boundary strength calculation with non-block filters, temporal motion vector prediction for subsequent pictures, and cross-CTU spatial motion vector prediction for subsequent CUs. ... ... ... ... ... ... Assuming a continuous motion trajectory, the motion vectors MV0 and MV1 to the two reference blocks are proportional to the temporal distances TD0 and TD1 between the current picture and the two reference pictures. ... It should be. As a special case, if the current image is temporally between two reference images and the temporal distance from the current image to the two reference images is the same, the dual matching becomes the bi-directional MV based on the median.

[0046] The current VVC has introduced MMVD in addition to the existing merge mode. In the existing merge mode the implicitly derived motion information is directly used for generating the prediction samples of the current CU. In the MMVD mode, after a merge candidate is selected, the merge candidate is further refined by the signaled MVD information.

[0047] The MMVD flag is signaled to specify whether the MMVD mode is used for the CU immediately after the skip flag and the merge flag are sent. The MMVD mode information includes a merge candidate flag, a distance index specifying the degree of motion, and a direction index indicating the direction of motion.

[0048] In the MMVD mode, only one of the first two candidates in the merge list is allowed to be selected as the leading MV, and the merge candidate flag is signaled to specify which of the first two candidates is used.

[0049] Figure 6 is an example of the search points used for MMVD. To derive these search points, an offset is added to the horizontal or vertical component of the leading MV. The distance index specifies the information of the degree of motion, indicates a predetermined offset from the starting point, and the direction index represents the direction of the offset with respect to the starting point by a predetermined mapping from the direction index to the offset sign. ​​​​​​​​​​​

[0050] The meaning of the mapped offset sign may vary according to the information of the leading MV. The leading MV is either a single prediction MV or a bi-prediction MV where the referenced reference image points to the same side of the current image (i.e., the POCs of both of the up to two reference images are both larger than the POC of the current image or both are smaller than the POC of the current image), the mapped offset sign specifies the sign of the MV offset added to the leading MV. The leading MV is a bi-prediction MV where the two motion vectors point to different sides of the current image (i.e., the POC of one of the reference images is larger than the POC of the current image and the POC of the other reference image is smaller than the POC of the current image), the mapped offset sign specifies the sign of the MV offset added to the L0 motion vector of the leading MV and the opposite sign of the MV offset added to the L1 motion vector of the leading MV.

[0051] Next, both components of the MV offset are derived from the signaled MMVD distance and sign, and the final MVD is further derived from the MV offset components.

[0052] The current VVC has also introduced the SMVD mode. In the SMVD mode, the motion information including both reference image indexes of L0 and L1 and the MVD of L1 is not signaled but derived. In the encoder, the SMVD motion estimation starts with an initial MV evaluation. The set of initial MV candidates consists of MVs obtained from a single prediction search, MVs obtained from a bi-prediction search, and MVs from the AMVP list. The initial MV candidate with the lowest rate-distortion cost is selected as the initial MV for the SMVD motion search. ​​​

[0053] The current VVC also introduced BWA. In HEVC, the bi-predicted signal is obtained by averaging two predicted signals from two reference images and / or by using two motion vectors. In the current VVC, BWA extends the bi-prediction mode to allow not only simple averaging but also weighted averaging of the two predicted signals.

[0054] In the current VVC, five weights are allowed for BWA. For each bi-predicted CU, the weights are determined in one of two ways. For non-merged CUs, the weight index is signaled after the difference of the motion vectors, while for merged CUs, the weight index is inferred from adjacent blocks based on the merge candidate index. Weighted-average bi-prediction is applied only to CUs with 256 or more luma samples (i.e., the product of the width of the CU and the height of the CU is 256 or more). In the case of an image that does not use backward prediction, all five weights are used. In the case of an image that uses backward prediction, only a predetermined subset of three of the five weights is used.

[0055] The current VVC also introduced the derivation of pairwise-averaged merge candidates. In the derivation of pairwise-averaged merge candidates, pairwise-averaged candidates are generated by averaging a predetermined pair of candidates within the existing merge candidate list. The averaged motion vectors are calculated separately for each reference list. If both motion vectors can be obtained from one list, these two motion vectors are averaged even if they point to different reference images; if there is only one available motion vector, this single motion vector is used directly; if available motion vectors are only one, this one motion vector is directly used; available If there is no motion vector, this list remains invalid. After the pairwise averaging marker candidates are added, if the merge list is not full, zero MVPs are inserted at the end of this merge list until the maximum merge candidate number is reached.

[0056] The current reference software codebase for VVC, known as the VVC Test Model (VTM), has also introduced the SbTMVP mode. Similar to the Temporal Motion Vector Prediction (TMVP) in HEVC, SbTMVP uses the motion fields in the parallel images to improve the motion vector prediction and merge mode for the CUs in the current image. The same parallel images used in TMVP are used in SbTMVP. SbTMVP differs from TMVP in the following two main aspects. First, TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level. Second, TMVP obtains the temporal motion vector from the parallel blocks in the parallel images (this parallel block is the lower-right or central block relative to the current CU), while SbTMVP applies the motion shift obtained from the motion vector from one of the spatial neighboring blocks of the current CU before obtaining the temporal motion information from the parallel images.

[0057] Figures 7A and 7B illustrate the operation of the SbTMVP mode. SbTMVP predicts the motion vectors of the sub-CUs within the current CU in two steps. Figure 7A illustrates the first step where the spatial neighbors are examined in the order of A1, B1, B0, and A0. The first spatial neighboring block with a motion vector using the parallel image as the reference image is identified. When it is determined, this motion vector is selected as the applied motion shift. When such motion cannot be identified from spatial neighbors, the motion shift is set to (0, 0). FIG. 7B illustrates a second step in which the motion shift identified in the first step is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vectors and reference indices) from the parallel images. The example adopted in FIG. 7B shows an example where the motion shift is set to the motion of block A1. Next, for each sub-CU, the motion information of the corresponding block (the smallest motion grid covering the central sample) in the parallel images is used to derive the motion information of this sub-CU. After the motion information of the parallel sub-CUs is identified, using temporal motion scaling similar to the TMVP process of HEVC, the reference images of the temporal motion vectors and the reference images of the temporal motion vectors of the current CU are aligned in such a way that this motion information is converted into the motion vectors and reference indices of the current sub-CU. After the motion information of the parallel sub-CUs is identified, using temporal motion scaling similar to the TMVP process of HEVC, the reference images of the temporal motion vectors and the reference images of the temporal motion vectors of the current CU are aligned in such a way that this motion information is converted into the motion vectors and reference indices of the current sub-CU. In the third version of VTM (VTM3), a combined sub-block based merge list including both SbTMVP candidates and affine merge candidates is used for signaling by the sub-block based merge mode. The SbTMVP mode is enabled or disabled by a sequence parameter set

[0058] (SPS: Sequence Parameter Set) flag. When the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block based merge candidates, followed by the affine merge candidates. The size of the sub-block based merge list is signaled in the SPS and in VTM3 The size of the sub-block based merge list is signaled in the SPS and in VTM3 (SPS: Sequence Parameter Set) flag. When the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block based merge candidates, followed by the affine merge candidates. MVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block based merge candidates, followed by the affine merge candidates. The size of the sub-block based merge list is signaled in the SPS and in VTM3 The size of the sub-block based merge list is signaled in the SPS and in VTM3 The maximum allowable size of the merge list based on this sub-block is fixed at 5. Sb The sub-CU size used in TMVP is fixed at 8×8, and similar to the case of the affine merge mode the SbTMVP mode can only be applied to CUs with both width and height of 8 or more. The encoding logic for additional SbTMVP merge candidates is the same as that of other merge candidates, i.e., for each CU in a P or B slice, an additional RD check is performed to determine whether to use the SbTMVP candidate.

[0059] The current VVC has introduced new inter-mode encoding and decoding tools, but the scaling process for deriving spatial and temporal motion candidates and the constraints regarding the long-term reference images existing in HEVC and the current VVC are not well-defined for some of the new tools. In the present disclosure, some constraints regarding long-term reference images for the new inter-mode encoding and decoding tools are proposed.

[0060] According to the present disclosure, during the operation of an inter-mode encoding and decoding tool for an inter-mode encoded block, it is determined whether one or more of the reference images related to the inter-mode encoded block involved in the operation of the inter-mode encoding and decoding tool are long-term reference images, and then, based on that determination, constraints are imposed on the operation of the inter-mode encoding and decoding tool for the inter-mode encoded block.

[0061] According to an embodiment of the present disclosure, the inter-mode encoding and decoding tool includes the generation of pairwise averaged merge candidates.

[0062] In one example, when an averaging merge candidate involved in generating pairwise averaging merge candidates consists of one reference image that is a long-term reference image and another reference image that is not a long-term reference image, and is generated from a predetermined candidate pair, the averaging merge candidate is regarded as invalid.

[0063] In the same example, during the generation of an averaging merge candidate from a predetermined candidate pair consisting of two reference images that are both long-term reference images, scaling processing is prohibited.

[0064] According to another embodiment of the present disclosure, the inter-mode encoding / decoding tool includes BDOF, and the inter-mode encoded block is a bi-directional prediction block.

[0065] In one example, when it is determined that one reference image of the bi-directional prediction block involved in the operation of BDOF is a long-term reference image and the other reference image of the bi-directional prediction block involved in the operation of BDOF is not a long-term reference image, the execution of BDOF is prohibited.

[0066] According to another embodiment of the present disclosure, the inter-mode encoding / decoding tool includes DMVR, and the inter-mode encoded block is a bi-directional prediction block.

[0067] In one example, when it is determined that one reference image of the bi-directional prediction block involved in the operation of DMVR is a long-term reference image and the other reference image of the bi-directional prediction block involved in the operation of DMVR is not a long-term reference image, the execution of DMVR is prohibited.

[0068] In another example, one reference image of the bi-directional prediction block involved in the operation of DMVR is a long-term reference image, and the other reference image of the bi-directional prediction block involved in the operation of DMVR is a long-term reference If it is determined that the image is not an image, the execution range of DMVR is limited to the execution range of integer pixel DMVR. is limited to the range.

[0069] According to another embodiment of the present disclosure, the inter-mode encoding / decoding tool includes the derivation of MMVD candidates. including.

[0070] In one example, when it is determined that the motion vector candidate involved in the derivation of the MMVD candidate has its own motion vector pointing to a reference image that is a long-term reference image, it is prohibited to use the motion vector candidate as the base motion vector (also called the leading motion vector). as the base motion vector (also called the leading motion vector). is prohibited.

[0071] In a second example, if one of the reference images of the inter-mode encoded block involved in the derivation of the MMVD candidate is a long-term reference image, the other reference image of the inter-mode encoded block involved in the derivation of the MMVD candidate is not a long-term reference image, and furthermore, if the base motion vector is a bidirectional motion vector, for one motion vector that points to the long-term reference image and is also included in the bidirectional base motion vector, it is prohibited to change this one motion vector by the signaled motion vector difference (MVD). involved in the derivation of the MMVD candidate is a long-term reference image, and the other reference image of the inter-mode encoded block involved in the derivation of the MMVD candidate is not a long-term reference image, and furthermore, if the base motion vector is a bidirectional motion vector, involved in the derivation of the MMVD candidate is a long-term reference image, and the other reference image of the inter-mode encoded block involved in the derivation of the MMVD candidate is not a long-term reference image, and furthermore, if the base motion vector is a bidirectional motion vector, points to the long-term reference image and is also included in the bidirectional base motion vector, for one motion vector that points to the long-term reference image and is also included in the bidirectional base motion vector, is prohibited.

[0072] In the same second example, the proposed MVD change process instead becomes as shown in the following frame, and the emphasized part of the text indicates the proposed change from the existing MVD change process in the current VVC.

Table 1

[0073] In a third example, at least one of the inter-mode encoded blocks involved in the derivation of the MMVD candidate At least one reference picture is a long-term reference picture, and furthermore, when the basic motion vector is a bidirectional motion vector, the scaling process in the derivation of the final MMVD candidate is prohibited.

[0074] According to one or more embodiments of the present disclosure, the inter-mode encoding / decoding tool includes deriving an SMVD candidate.

[0075] In one example, when it is determined that a motion vector candidate has its own motion ve ctor pointing to a reference picture that is a long-term reference picture, it is prohibited to use the motion vector candidate as the basic motion vector.

[0076] In one example, when at least one reference picture of an inter-mode encoded block involved in the derivation of an SMVD candidate is a long-term reference picture, and furthermore, when the basic motion vector is a bidirectional motion vector, it is prohibited to change one motion vector by a signal-notified MVD for one motion vector that points to the long-term reference picture and is included in the bidirectional basic motion vector. In another example, when one reference picture of an inter-mode encoded block involved in the derivation of an SMVD candidate is a long-term reference picture, and the other reference picture of the inter-mode encoded block involved in the derivation of the SMVD candidate is not a long-term reference picture, and furthermore, when the basic motion vector is a bidirectional motion vector, it is prohibited to change the one motion vector by the signal-notified MVD for one motion vector that points to the long-term reference picture and is also included in the bidirectional basic motion vector.

[0077] According to another embodiment of the present disclosure, the inter-mode encoding / decoding tool is by weighted averaging​​​​​​​​​​​ A block that includes dual prediction and is inter-mode encoded is a bi-directional prediction block.

[0078] In one example, when it is determined that at least one reference image involved in dual prediction by weighted averaging is a long-term reference image, the use of unequal weighting is prohibited. .

[0079] According to another embodiment of the present disclosure, the inter-mode encoding / decoding tool includes the derivation of motion vector candidates, and the block encoded in the inter-mode is a block encoded by SbTMVP. lock.

[0080] In one example, the same constraints as those for the derivation of motion vector candidates for a conventionally TMVP-encoded block are used for the derivation of motion vector candidates for an SbTMVP-encoded block.

[0081] In one improvement of the above-described example, the constraints for the derivation of motion vector candidates used for both a conventionally TMVP-encoded block and an SbTMVP-encoded block are such that when one of the two reference images consisting of the target reference image and the reference image for the temporally adjacent block is a long-term reference image and the other reference image is not a long-term reference image, the motion vector of the adjacent block is regarded as invalid, and when both the target reference image and the reference image for the adjacent block are long-term reference images, the operation of scaling processing for the motion vector of the adjacent block is prohibited, and the motion vector of the adjacent block is directly used as the motion vector prediction for the current block. temporal temporal

[0082] According to another embodiment of the present disclosure, the inter-mode encoding / decoding tool includes using an affine motion model in the derivation of motion vector compensation. ​​

[0083] In one example, when it is determined that the reference image involved in the use of the affine motion model is a long-term reference image, the use of the affine motion model in the derivation of motion vector candidates is prohibited. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium.

[0084] In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium. In one or more examples, the above-described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, those functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a processing unit in hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to search for instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium.

[0085] Furthermore, the above method may be implemented using an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), or any combination thereof. Furthermore, the above method may be implemented using an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), or any combination thereof. It may be implemented by a device including one or more circuits including a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components. To implement the above method, these circuits can be used in combination with other hardware or software components. Each module, sub-module, unit, or sub-unit disclosed above can be at least partially implemented using one or more circuits. Other embodiments of the present invention will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention following the general principles of the invention, including departures from the present disclosure that come within known or customary practice in the art. Note that this description and embodiments are to be regarded as illustrative only, and the substantial scope and spirit of the invention are indicated by the appended claims. The present invention is not limited to the specific examples illustrated in the above description and the accompanying drawings, and various changes and modifications can be made without departing from its scope. The scope of the present invention is intended to be limited only by the appended patent claims.

[0086]

[0087] ​

Claims

1. determining whether one or more of the reference pictures associated with the inter mode coded blocks involved in the decoding process are long-term reference pictures; constraining operations during a decoding process for the inter mode coded block based on the determination; and Including, If the decoding process includes a derivation process for motion vector difference (MVD) in a merge mode with motion vector difference (MMVD), constraining operations during a decoding process for the inter mode coded block based on the determination includes: prohibiting a scaling operation in the derivation of the final MVD if at least one reference image of the inter mode coded block is a long-term reference image and the base motion vector is a bidirectional motion vector.

13. A method for video decoding, comprising:

2. If the decoding process includes sub-block based temporal motion vector prediction (SbTMVP) and the inter mode coded block is a SbTMVP coded block, constraining operations during a decoding process for the inter mode coded block based on the determination includes: In the derivation process for motion vector candidates for the SbTMVP coded block, the same restrictions are used as for the derivation process for motion vector candidates for temporal motion vector prediction (TMVP) coded blocks. The method of claim 1 , comprising:

3. The constraint on the derivation process for the motion vector candidates used for both the TMVP coded block and the SbTMVP coded block is: Among two reference pictures, the target reference picture and the reference picture for the temporally adjacent block, when one reference picture is a long-term reference picture and the other reference picture is not a long-term reference picture, the motion vector of the temporally adjacent block is regarded as invalid; The method of claim 2 , comprising:

4. The constraint on the derivation process for the motion vector candidates used for both the TMVP coded block and the SbTMVP coded block is: When both the target reference image and the reference image for the temporally adjacent block are long-term reference images, the operation of scaling processing for the motion vector of the temporally adjacent block is prohibited, and the motion vector of the temporally adjacent block is directly used as the motion vector prediction for the current block, The method according to claim 2, comprising:

5. One or more processors, A non-transitory memory connected to the one or more processors, A plurality of programs stored in the non-transitory memory, Including, When the plurality of programs are executed by the one or more processors, the computing device is caused to Determine whether one or more of the reference images related to the inter-mode encoded blocks involved in the decoding process are long-term reference images, Based on the determination, restrict the operations during the decoding process for the inter-mode encoded blocks, Execute a method for video decoding including, When the decoding process includes a derivation process for the motion vector difference (MVD) in the merge mode by motion vector difference (MMVD), Based on the determination, restricting the operations during the decoding process for the inter-mode encoded blocks means When at least one reference image of the inter-mode encoded block is a long-term reference image and the basic motion vector is a bi-directional motion vector, prohibiting the scaling process in the derivation of the final MVD Including, A computing device.

6. When the decoding process includes sub-block based temporal motion vector prediction (SbTMVP) and the inter-mode encoded block is an SbTMVP encoded block, based on the determination, restricting the operations during the decoding process in the inter prediction mode for the inter-mode encoded block means Using the same restrictions as those for the derivation process of the motion vector candidates for the conventionally TMVP encoded blocks in the derivation process of the motion vector candidates for the SbTMVP encoded blocks The computing device according to claim 5, including:

7. The constraint on the derivation process for the motion vector candidates used for both the TMVP coded block and the SbTMVP coded block is: Among two reference pictures consisting of a target reference picture and a reference picture for a temporally adjacent block, when one reference picture is a long-term reference picture and the other reference picture is not a long-term reference picture, the motion vector of the temporally adjacent block is regarded as invalid. The computing device of claim 6 , comprising:

8. The constraint on the derivation process for the motion vector candidates used for both the TMVP coded block and the SbTMVP coded block is: prohibiting the operation of a scaling process on the motion vector of the temporal neighboring block when both the target reference image and the reference image for the temporal neighboring block are long-term reference images, and directly using the motion vector of the temporal neighboring block as a motion vector prediction for the current block; The computing device of claim 6 , comprising:

9. Dividing a video frame into a number of blocks; determining whether one or more of the reference pictures associated with the inter mode coded blocks involved in the decoding process are long-term reference pictures; constraining operations during a decoding process for the inter mode coded block based on the determination; and Including, If the decoding process includes a derivation process for motion vector difference (MVD) in a merge mode with motion vector difference (MMVD), constraining operations during a decoding process for the inter mode coded block based on the determination includes: prohibiting a scaling operation in the derivation of the final MVD if at least one reference image of the inter mode coded block is a long-term reference image and the base motion vector is a bidirectional motion vector.

16. A method for video encoding comprising:

10. If the decoding process includes sub-block based temporal motion vector prediction (SbTMVP) and the inter mode coded block is a SbTMVP coded block, constraining operations during a decoding process for the inter mode coded block based on the determination includes: In the derivation process for motion vector candidates for the SbTMVP coded block, the same restrictions are used as for the derivation process for motion vector candidates for temporal motion vector prediction (TMVP) coded blocks.

10. The method of claim 9, comprising:

11. The constraint on the derivation process for the motion vector candidates used for both the TMVP coded block and the SbTMVP coded block is: Among two reference pictures, the target reference picture and the reference picture for the temporally adjacent block, when one reference picture is a long-term reference picture and the other reference picture is not a long-term reference picture, the motion vector of the temporally adjacent block is regarded as invalid; The method of claim 10, comprising:

12. The constraint on the derivation process for the motion vector candidates used for both the TMVP coded block and the SbTMVP coded block is: prohibiting the operation of a scaling process on the motion vector of the temporal neighboring block when both the target reference image and the reference image for the temporal neighboring block are long-term reference images, and directly using the motion vector of the temporal neighboring block as a motion vector prediction for the current block; The method of claim 10, comprising:

13. one or more processors; a non-transitory memory coupled to the one or more processors; a plurality of programs stored in the non-transitory memory; Including, The programs, when executed by the one or more processors, cause the computing device to: Dividing a video frame into a number of blocks; determining whether one or more of the reference pictures associated with the inter mode coded blocks involved in the decoding process are long-term reference pictures; constraining operations during a decoding process for the inter mode coded block based on the determination; and 4. A method for video encoding comprising: If the decoding process includes a derivation process for motion vector difference (MVD) in a merge mode with motion vector difference (MMVD), constraining operations during a decoding process for the inter mode coded block based on the determination includes: prohibiting a scaling operation in the derivation of the final MVD if at least one reference image of the inter mode coded block is a long-term reference image and the base motion vector is a bidirectional motion vector. Including, Computing device.

14. If the decoding process includes sub-block based temporal motion vector prediction (SbTMVP) and the inter mode coded block is a SbTMVP coded block, constraining an operation during the decoding process in an inter prediction mode for the inter mode coded block based on the determination comprises: In the derivation process for motion vector candidates for the SbTMVP coded block, the same restrictions as those for the derivation process for motion vector candidates for conventional TMVP coded blocks are used.

14. The computing device of claim 13, comprising:

15. The constraint on the derivation process for the motion vector candidates used for both the TMVP coded block and the SbTMVP coded block is: Among two reference pictures consisting of a target reference picture and a reference picture for a temporally adjacent block, when one reference picture is a long-term reference picture and the other reference picture is not a long-term reference picture, the motion vector of the temporally adjacent block is regarded as invalid.

15. The computing device of claim 14, comprising:

16. The constraint on the derivation process for the motion vector candidates used for both the TMVP coded block and the SbTMVP coded block is: prohibiting the operation of a scaling process on the motion vector of the temporal neighboring block when both the target reference image and the reference image for the temporal neighboring block are long-term reference images, and directly using the motion vector of the temporal neighboring block as a motion vector prediction for the current block; 15. The computing device of claim 14, comprising:

17. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon, A non-transitory computer-readable storage medium, the instructions, when executed by one or more computer processors, cause the one or more computer processors to receive a bitstream and perform a method for video decoding of any one of claims 1 to 4 to decode the bitstream, or perform a method for video encoding of any one of claims 9 to 12 to generate a bitstream and transmit the bitstream.

18. Generating a bitstream by executing a method for video encoding according to any one of claims 9 to 12; transmitting said bitstream.

19. A computer program having instructions which, when executed by one or more computer processors, cause the one or more computer processors to perform a method for video decoding as described in any one of claims 1 to 4, or a method for video encoding as described in any one of claims 9 to 12.