Constrained motion vector derivation for long-term reference picture in video coding and decoding

By applying constraints on motion vector derivation based on long-term reference images, the method addresses inefficiencies in existing video coding standards, improving the reliability and efficiency of motion-compensated prediction in video encoding and decoding.

JP2025131632APending Publication Date: 2025-09-09BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025085514
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-02-20
Filing Date
2025-05-22
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing video coding standards like HEVC and VVC lack well-defined constraints for deriving motion vectors, particularly for long-term reference pictures, leading to inefficiencies in motion-compensated prediction.

Method used

Implementing constraints on motion vector derivation for inter-mode coded blocks by determining whether reference images are long-term reference images, and adjusting operations based on this determination to enhance the reliability of motion vector prediction.

Benefits of technology

Improves the efficiency and accuracy of motion-compensated prediction by restricting scaling processes for motion candidates, especially when long-term reference images are involved, thereby enhancing video encoding and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025131632000001_ABST
    Figure 2025131632000001_ABST
Patent Text Reader

Abstract

To provide systems and methods for performing video coding and decoding using constraints on motion vector derivation for long-term reference pictures.SOLUTION: A method of constraining the operations of certain inter-mode coding / decoding tools in derivation of motion vector candidates for inter-mode coded blocks employed in video coding standards, such as the now-current Versatile Video Coding (VVC), in a hybrid video encoder 100 includes: determining whether one or more of reference pictures associated with an inter-mode coded block involved in an operation of an inter-mode coding / decoding tool are long-term reference pictures; and constraining the operation of the inter-mode coding / decoding tool on the inter-mode coded block based on the determination.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is a sequel to U.S. Provisional Application No. 62 / 808,271, filed February 20, 2019. This application claims priority from the same patent application, the entire specification of which is incorporated herein by reference. Quoted in. [Technical Field]

[0002] This disclosure relates generally to video encoding, decoding, and compression. In particular, to a method for encoding, decoding, and compressing long-term reference pictures. System for performing video encoding and decoding using constraints on motion vector derivation and methods. [Background technology]

[0003] This section provides background information related to the present disclosure. The information contained herein is not necessarily conventional. It should not be interpreted as a technique.

[0004] Video data can be compressed by any of a variety of video encoding and decoding techniques. Video encoding and decoding may be performed according to one or more video coding standards. Exemplary video encoding and decoding standards include Versatile Video Coding (VVC) and ile Video Coding), Joint Exploration Test Model (JEM) coding, High Efficiency Video Coding (H.265 / HEVC), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC) and Video Expert Group This includes MPEG (Moving Picture Experts Group).

[0005] In video coding and decoding, prediction is generally performed due to the redundancy inherent in a video image or sequence. Video coding and decoding techniques are used. One of the goals of this technique is to stream video data while avoiding or minimizing video quality degradation. The idea is to compress it into a form with a lower bit rate.

[0006] Prediction methods used in video encoding and decoding are usually spatial (intra-frame) prediction and / or performs temporal (inter-frame) prediction to reduce or remove redundancy inherent in video data and associated with block-based video encoding and decoding.

[0007] In block-based video coding, the input video signal is processed block by block. The block (also called coding unit (CU)) is used for spatial prediction and and / or temporal prediction can be performed.

[0008] Spatial prediction (also called "intra prediction") is the process of predicting the current block from a / Samples of already coded neighboring blocks in a slice (also called reference samples) ) to predict the image. Spatial prediction reduces the spatial redundancy inherent in video signals. do.

[0009] Temporal prediction (also called "inter prediction" or "motion compensated prediction") is the process of The temporal prediction is performed by using pixels reconstructed from already coded video images. The prediction reduces the temporal redundancy inherent in the video signal. is typically one or more CUs indicating the amount and direction of motion between the current CU and its temporal reference. The motion vector (MV) is signaled by the If a picture is supported, the temporal prediction signal is taken from any reference picture in the reference picture store. A reference image index for identifying which of the images is the reference image is additionally transmitted.

[0010] After spatial and / or temporal prediction, the mode decision block in the encoder For example, the best prediction mode is selected based on a rate-distortion optimization method. The prediction residual is decorrelated by a transformation and quantified. The quantified residual coefficients are inversely quantified and inversely transformed to form the residual of the reconstruction; This reconstruction residual is then added to the predicted block to obtain the reconstructed signal of this block. Form.

[0011] After spatial and / or temporal prediction, in-loop filtering is performed, e.g., to eliminate non-blocking Locked filter, Sample Adaptive Offset (SAO) and The adaptive in-loop filter (ALF) is used to reconstruct the CU. After applying the CU, the reconstructed CU is stored in the reference image store, and the Used for encoding and decoding. The encoding mode is used to form the output video bitstream. the code (inter or intra), prediction mode information, motion information, and the quantified residual coefficient All the numbers are sent to the entropy coding section, which then compresses and packs them into a bitstream. A ream is formed.

[0012] During the decoding process, the video bitstream is first decoded by the entropy decoder. The coding mode and prediction information are sent to the spatial predictor (in the case of intra-coding). Or sent to a temporal predictor (in the case of inter-coding) to form a prediction block. The residual transform coefficients are sent to an inverse quantification unit and an inverse transform unit to reconstruct the residual block. Then the predicted block and the residual block are added together. The reconstructed block is , and may be stored in the reference image storage unit after further in-loop filtering. The reconstructed video in the reference image store was then sent to drive a display device. and used to predict future video blocks.

[0013] In video coding standards such as HEVC and VVC, a reference picture set (RPS) is used. The concept of a picture set is that previously decoded pictures are used as references, i.e., sample data prediction and motion. The decoded picture buffer (DPB) is used for vector prediction. In general, the outline of the RPS for reference image management is The key is to consider the DPB in each slice (also called a "tile" in the current VVC). The status of the

[0014] Images in the DPB are classified as "used for short-term reference," "used for long-term reference," or Images can be marked as "not used for reference" or "not used for reference". If a field is marked as "unused," it will no longer be available for prediction and will no longer be needed in the output. When the data is deleted, it can be removed from the DPB.

[0015] In general, the long-term reference images are usually sorted by display order (i.e., Picture Order Count). The distance from the current image compared to the short-term reference image in terms of the distance (called the point of contact count or POC) This distinction between long-term and short-term reference images is important for temporal and spatial MV prediction. It also affects some decoding processes, such as motion vector scaling in implicit weighted prediction. It is possible to give

[0016] Video encoding and decoding standards such as HEVC and VVC allow for spatial and / or temporal When deriving candidate motion vectors, certain constraints may be imposed on the spatial and / or temporal motion vectors. The scaling process that forms part of the derivation of the candidate vectors involves the specific reference images involved in that process. The image is placed based on whether it is a long-term reference image.

[0017] However, similar constraints are imposed according to video codec specifications such as the current VVC standard. , such as for deriving motion vector candidates for inter-mode coded blocks. It is designed for new inter-mode video encoding and decoding tools adopted in the video codec specifications. It is not placed. Summary of the Invention

[0018] This section provides a general overview of the disclosure and does not attempt to provide a comprehensive overview of its full scope or all of its features. This is not a disclosure of the nature of disclosure.

[0019] According to a first aspect of the present disclosure, a method for video encoding / decoding includes: and a plurality of programs executed by the one or more processors. The method is implemented on a computing device having a memory for storing a video stream. This involves partitioning each image in the image frame into multiple blocks or coding units (CUs). This method performs inter-mode motion estimation for these inter-mode coded blocks. The method further includes performing vector derivation. While performing inter-mode motion vector derivation for a lock, The method further includes operating an inter-mode encoding / decoding tool. one of the reference images related to the inter-mode coded block involved in the operation of the coding tool determining whether one or more of the images are long-term reference images; and based on said determination, The inter-mode encoding / decoding tool for the inter-mode encoded block and constraining the operation of the

[0020] In accordance with a second aspect of the present application, a computing device includes one or more processors; The program includes a memory and a plurality of programs stored in the memory. which, when executed by the one or more processors, to perform the operations described above.

[0021] According to a third aspect of the present application, a non-transitory computer-readable storage medium includes one or more A computing device having a plurality of processors and a plurality of programs executed by the computing device. The program, when executed by the one or more processors, The computing device is caused to perform the operations described above. [Brief explanation of the drawings]

[0022] A set of illustrative, non-limiting embodiments of the present disclosure are described below in conjunction with the accompanying drawings. Those skilled in the art will understand that modifications of the structure, method, or function may be made based on the examples presented herein. and all such variations may be included within the scope of the present disclosure. In the absence of a shield, the teachings of different embodiments may be, but need not be, mutually exclusive. They can be combined. [Figure 1] FIG. 1 is a block diagram illustrating an exemplary block-based hybrid video encoder that can be used in combination with many video encoding and decoding standards. [Figure 2] FIG. 2 is a block diagram illustrating an exemplary video decoder that can be used in conjunction with many video encoding and decoding standards. [Figure 3] FIG. 3 is an example of a block partition in a multi-type tree structure that can be used in combination with many video encoding and decoding standards. [Figure 4] FIG. 4 is an example of Bi-Directional Optical Flow (BDOF) processing. [Figure 5] FIG. 5 is an example of bidirectional matching used in decoder-side motion vector refinement (DMVR). [Figure 6] FIG. 6 shows an example of search points used in Merge Mode with Motion Vector Difference (MMVD). [Figure 7A] FIG. 7A is an example of spatially neighboring blocks used in a Subblock-based Temporal Motion Vector Prediction (SbTMVP) mode. [Figure 7B] FIG. 7B is an example of deriving sub-CU level motion information by motion shifts identified from spatially neighboring blocks in SbTMVP mode. DETAILED DESCRIPTION OF THE INVENTION

[0023] The terms used in this disclosure describe particular examples and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "one," "one," and "the" The words "of" and "of" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The term "and / or" as used herein means the inclusion of one or more of the associated listed items. It should be understood to refer to any or all possible combinations.

[0024] Terms such as "first," "second," and "third" are used here to describe different types of information. However, it is understood that this information should not be limited by such terms. These terms are used only to distinguish one type of information from another. For example, without departing from the scope of this disclosure, the first information may be referred to as the second information. Similarly, the second information can be referred to as the first information. As used here, "(if)...tara" or "(if)...ba" or "(if)...to" The term can mean "when" or "depending on" depending on the context. is.

[0025] As used herein, the terms "one embodiment," "an embodiment," "an alternative embodiment," "another ... "embodiment" or similar references refer to one or more particular features, structures, or aspects described in connection with an embodiment. or feature is included in at least one embodiment of the present disclosure. The phrase "in one embodiment" or "in one embodiment" appears in several places throughout this specification in either the singular or plural. The terms "in," "in an example," "in one embodiment," and similar expressions are not necessarily intended to be limiting. Furthermore, specific examples of one or more embodiments may not all refer to the same embodiment. The features, structures, or characteristics may be combined in any suitable manner.

[0026] Conceptually, many video coding standards are similar, including those mentioned above in the background section. For example, almost all video coding standards use block-based processing, and similar video Video compression is realized by sharing the coding block diagram.

[0027] FIG. 1 shows an example of a video encoding / decoding standard that can be used in conjunction with many 1 shows a block diagram of a typical block-based hybrid video encoder 100. In .0, a video frame is divided into multiple video blocks for processing. For each video block obtained, prediction can be performed using either an inter-prediction approach or an intra-prediction approach. In inter prediction, pixels from a previously reconstructed frame are used. Based on the above, one or more predictors are formed by motion estimation and motion compensation. Prediction involves forming a predictor based on reconstructed pixels in the current frame. Through the decision, the best predictor can be selected to predict the current block. .

[0028] A prediction residual, which represents the difference between the current video block and its predictor, is sent to the transform circuit 102. Then, for entropy reduction, the transform coefficients are transferred from the transform circuit 102 to the quantization circuit. The quantized coefficients are then fed to an entropy coding circuit 106. As shown in Figure 1, inter prediction is video block partition information from the circuit and / or intra-prediction circuit 112, motion vectors, Prediction related information 110 such as the image length, reference image index, and intra prediction mode is also included in the is fed through an entropy coding circuit 106 and stored in a compressed video bitstream 114. It will be preserved.

[0029] The encoder 100 also includes decoder-related circuitry for reconstructing pixels for prediction purposes. First, the prediction residual is reconstructed via an inverse quantification 116 and an inverse transform circuit 118. This reconstructed prediction residual is combined with the block predictor 120 to produce the current , generating unfiltered reconstructed pixels for the video block.

[0030] In general, to improve coding / decoding efficiency and visual quality, in-loop filters are used. For example, current versions of AVC, HEVC, and VVC are non-blocking. In order to further improve coding efficiency, HEVC provides a locking filter. An additional in-loop filter called SAO (Sample Adaptive Offset) is defined. In the current version of the VVC standard, it is called the ALF (Adaptive Loop Filter). Alternative in-loop filters are being actively researched and may be included in the final standard. Highly sexual.

[0031] These in-loop filter operations are optional. When these operations are performed, the sign On the other hand, they are used to save computational complexity. For example, the signal may be turned off as determined by the encoder 100.

[0032] Note that if these filter options are turned on by the encoder 100, While the prediction is usually based on the pixels of the unfiltered reconstruction, The target prediction is based on the pixels of the filtered reconstruction.

[0033] FIG. 2 shows an example of a video encoding / decoding standard that can be used in conjunction with many 1 is a block diagram showing a typical video decoder 200. This decoder 200 is similar to the decoder 200 of FIG. The reconstruction-related parts are similar to those present in the coder 100. In the decoder 200 (FIG. 2), The input video bitstream 201 is first subjected to entropy decoding 202 The quantified coefficient levels and prediction related information are then derived. The calculated coefficient levels are processed through inverse quantization 204 and inverse transformation 206 to produce a reconstructed The intra / inter mode selector 212 obtains the predicted residual. The lock prediction mechanism performs intra prediction 208 or It is configured to perform motion compensation 210. Prediction of the reconstruction obtained from the inverse transform 206 The residual and the predicted output generated by the block predictor mechanism are added together by an adder 214. are added together to obtain the set of unfiltered reconstruction pixels. If the in-loop filter is on, the filtering operation will The final reconstructed video is then derived from the reference image memory. The reconstructed video in the section can be sent out to drive a display device or used as a future video broadcast. It is sometimes used to predict locks.

[0034] In video coding and decoding standards such as HEVC, blocks are partitioned based on a quadtree. With the current new video coding and decoding standards such as VVC, more A partitioning scheme is used, which creates a single coding tree based on a quadtree, binary tree, or ternary tree. The coding tree unit (CTU) is divided into CUs to adapt to various local characteristics. In the current VVC, CU, prediction unit (PU), and transform unit (TU) are The unit (TU) distinction does not exist in most coding modes, and each CU is always a further It is used as the basic unit for both prediction and transformation without partitioning. However, In certain coding modes, such as the video coding mode, each CU may contain multiple TUs. In the multi-type tree structure, one CTU is first partitioned by a quadtree structure. Each quadtree leaf node is then further partitioned by binary and ternary tree structures. It is possible.

[0035] Figure 3 shows five partition types currently used in VVC: quaternary partition 301; Horizontal binary partition 302, vertical binary partition 303, horizontal ternary partition 304, and vertical ternary partition 305. It shows 05.

[0036] Video encoding / decoding standards such as HEVC and now VVC require that previously decoded images The images are stored in the decoded image buffer, as referenced in the concept of Reference Picture Set (RPS). Images in the DPB are classified as "images to be used for short-term reference" and "images to be used for long-term reference." can be marked as "Used for Reference" or "Not Used for Reference".

[0037] If the predetermined target reference image of the current block is different from the reference image of the neighboring block, The scaled motion vector of the target neighboring block is used as the motion vector prediction for the current block. In the scaling process for spatial motion candidates, the scale The matching coefficient is the Picture Order Count (POC) distance between the current image and the target reference image, and is calculated based on the POC distance between the current image and the reference image of the neighboring block.

[0038] In video coding and decoding standards such as HEVC and now VVC, certain constraints are imposed on the spatial In the scaling process for motion candidates, the specific reference image involved in the process is the long-term reference image. The two reference images are set based on whether they are long-term reference images. If the other is not a long-term reference image, the MV of the neighboring block is considered invalid. If both of the two reference images are long-term reference images, the difference between the two long-term reference images is The POC distance is usually large, and therefore the scaled MV is unreliable. Therefore, the MV of the spatially adjacent blocks is directly used as the MVP of the current block. and scaling is prohibited.

[0039] Similarly, the scaling process for temporal motion candidates involves the scaling of the current image and the target reference image. The POC distance between parallel images and temporally adjacent blocks (also called parallel blocks) A scaling factor is calculated based on the POC distance between the reference image.

[0040] In video coding and decoding standards such as HEVC and now VVC, certain constraints are imposed on the temporal In the scaling process for motion candidates, the specific reference image involved in the process is the long-term reference image. The two reference images are set based on whether they are long-term reference images. If the other is not a long-term reference image, the MV of the neighboring block is considered invalid. If both of the two reference images are long-term reference images, the difference between the two long-term reference images is The POC distance is usually large, and therefore the scaled MV is unreliable. Therefore, the MV of the spatially adjacent blocks is directly used as the MVP of the current block. and scaling is prohibited.

[0041] New video coding and decoding standards, such as VVC, are currently implementing new inter-mode coding and decoding. Some examples of new inter-mode encoding tools are: Directional Optical Flow (BDOF), Decoder-Side Motion Vector Refinement (DMVR), Marker by MVD Multimode Multi-Modal Difference (MMVD), Symmetric Multi-Modal Difference (SMVD), Bi-Prediction with Weighted Averaging (BWA) Bi-prediction with Weighted Averaging), pairwise averaging merge candidate derivation, and Sub-block based temporal motion vector prediction (SbTMVP).

[0042] Traditional bi-prediction in video coding and decoding involves the use of previously reconstructed reference images. It is a simple combination of two temporal prediction blocks, but block-based motion compensation is not available. The compensation limit allows for the possibility of observing small remaining movements between the samples of the two prediction blocks. This reduces the efficiency of motion-compensated prediction. OF is applied to the current VVC to determine such a Reduce the effects of movement.

[0043] Figure 4 shows an example of BDOF processing. BDOF is a process where the block It is a sample-by-sample motion refinement performed on top of motion compensated prediction based on the block. Block motion refinement is applied within a 6x6 window around the sub-block whose BDOF is After that, the predicted samples of reference picture list 0 (L0) and the predicted samples of reference picture list 1 (L1) are Based on the motion refinement thus derived, The final bi-predicted sample of the CU is then calculated along the motion trajectory based on the optical flow model. It is calculated by interpolating the L1 prediction samples.

[0044] DMVR is initially signaled and can be further refined by bi-matching prediction. It is a bi-prediction technique for merging blocks with two MVs.

[0045] Figure 5 shows an example of bi-matching used in DMVR. Bi-matching is the process of matching two different Find the closest match between two blocks along the motion trajectory of the current CU in the reference image. This is used to derive the motion information of the current CU. The cost function used is the row-subsampled sum of absolute differences (SAD). After the matching process is completed, the refined MV is Motion compensation, boundary strength calculation in deblocking filters, temporal motion vectors for subsequent images It is used for CU prediction, cross-CTU spatial motion vector prediction for subsequent CUs, and cross-CTU spatial motion vector prediction for subsequent CUs. Assuming a continuous motion trajectory, the motion vectors MV0 and MV1 to the two reference blocks are , proportional to the temporal distance between the current image and the two reference images, i.e., TD0 and TD1. As a special case, if the current image is temporally between two reference images, If the temporal distances from the current image to the two reference images are the same, then the bi-matching is It will be an interactive MV based on Ra.

[0046] The current VVC has introduced MMVD in addition to the existing merge mode. In, the implicitly derived motion information is directly used to generate the predicted samples for the current CU. In MMVD mode, after a merge candidate is selected, the MVD information signaled is used. This further refines the merging candidates.

[0047] The MMVD flag is sent immediately after the skip and merge flags are sent. Signaled to specify whether the mode is to be used for the CU. The information includes a merge candidate flag, a distance index that specifies the degree of movement, and a motion The direction index indicates the direction of the

[0048] In MMVD mode, only one of the first two candidates in the merge list is selected as the leading MV. The merge candidate flag indicates which of the first two candidates is The user is signaled to specify which one is to be used.

[0049] Figure 6 shows an example of search points used in MMVD. To derive these search points, An offset is added to the horizontal or vertical component of the first MV. Specifies the degree of fill information, indicates a predetermined offset from the starting point, and provides a direction index. The code is generated by a predetermined mapping from direction index to offset code. Indicates the direction of the offset relative to the starting point.

[0050] The meaning of the mapped offset code can change depending on the information in the first MV. In a uni-predictive MV or a bi-predictive MV where the reference picture points to the same side of the current picture. (i.e., the POC of up to two reference images is both greater than the POC of the current image. , or both are less than the current image's POC), the mapped offset sign specifies the sign of the MV offset added to the first MV. The first MV is the sum of the two motion vectors. A bi-predictive MV where the reference points to different sides of the current picture (i.e., one of the references If the POC of the other reference image is greater than the POC of the current image, C), the mapped offset code is used to map the L0 motion vector of the first MV. The sign of the added MV offset and the MV offset added to the L1 motion vector of the first MV Specifies the opposite sign of the set.

[0051] Both components of the MV offset are then derived from the signaled MMVD distance and code. The final MVD is further derived from the MV offset component.

[0052] The current VVC also introduces the SMVD mode, which allows for the L0 and L1 Although motion information including both reference picture index and L1 MVD is not signaled, At the encoder, SMVD motion estimation begins with an initial MV estimate. The set of candidates includes MVs obtained from a uni-predictive search, MVs obtained from a bi-predictive search, and The initial MV candidate with the lowest rate-distortion cost is , is selected as the initial MV for SMVD motion search.

[0053] The current VVC also introduces BWA. In HEVC, bi-predictive signals are encoded using two reference images. and / or two motion vectors are calculated by averaging the two prediction signals obtained from In the current VVC, BWA allows bi-prediction mode to be generated using a simple This is extended to allow not only averaging but also weighted averaging of the two prediction signals.

[0054] Currently, VVC allows 5 weights for BWA. For each bi-predicted CU, The weight is determined in one of two ways: for a non-merged CU, the weight index is The difference of the motion vectors is signaled after the merged CU, while the weight index is The data is inferred from neighboring blocks based on the merge candidate index. applies only to CUs with 256 or more luma samples (i.e., the width of the CU and the (The product of the height and the pixel size is 256 or more.) For images that do not use backward prediction, five weights are used. For images using backward prediction, three of the five weights are used. Only a predetermined subset of is used.

[0055] The current VVC also introduces the derivation of pairwise averaging merge candidates. In deriving merge candidates, pairwise averaging candidates are selected from the predefined candidates in the existing merge candidate list. The averaged motion vector is generated by averaging the candidate pairs. If both the motion vectors can be obtained from a single list, These two motion vectors are averaged even if they point to different reference images; If there is only one motion vector, this one motion vector is used directly; If there are no valid motion vectors, the list remains invalid. After the merge candidates are added, if the merge list is not full, the maximum number of merge candidates is reached. A zero MVP is inserted at the end of this merge list until

[0056] The current VVC test model, known as the VVC Test Model (VTM), is The current reference software codebase also introduced the SbTMVP mode. Similar to temporal motion vector prediction (TMVP) in parallel images, SbTMVP The motion field is used to perform motion vector prediction and mapping for the CU in the current picture. The same parallel images used in TMVP are used in SbTMVP. SbTMVP differs from TMVP in two main respects. First, TMVP is While TMV predicts motion at the CU level, SbTMVP predicts motion at the sub-CU level. P is obtained from the temporal motion vectors from the parallel blocks in the parallel image (these parallel blocks block is the lower right or center block relative to the current CU), but SbTMVP One of the spatially adjacent blocks of the current CU before obtaining temporal motion information from the sequence image Apply the motion shift obtained from the motion vectors from

[0057] Figures 7A and 7B illustrate the operation of the SbTMVP mode. The motion vectors of sub-CUs within the current CU are predicted in two steps. This illustrates the first step where A1, B1, B0, and A0 are examined in this order. The first spatially adjacent block having a motion vector using the image as a reference image is identified. If a motion vector is selected, this motion vector is selected as the motion shift to be applied. If is not distinguished from its spatial neighbors, the motion shift is set to (0,0). , the motion shift identified in the first step is applied (i.e., the current block position The sub-CU level motion information (motion vectors and reference images) is extracted from the parallel images. The example taken in FIG. 7B illustrates the second step of obtaining the index. In this example, the motion shift is set to the motion of block A1. The motion of the corresponding block (the smallest motion grid covering the central sample) in the parallel image The motion information of this sub-CU is derived using the motion information of the parallel sub-CU. After separation, temporal motion scaling is used to process the video in a similar way to the TMVP processing in HEVC. Align the reference image of the motion vector with the reference image of the temporal motion vector of the current CU In this method, the motion information is converted into a motion vector and a reference index of the current sub-CU. do.

[0058] In the third version of VTM (VTM3), SbTMVP candidates and affine merge candidates are The combined sub-block based merge list containing both SbTMVP mode is used to signal the sequence parameter setting. It is enabled or disabled by the SPS (Sequence Parameter Set) flag. When MVP mode is enabled, the SbTMVP predictor uses sub-block based markers. The affine merge candidate is added as the first entry in the list of merge candidates, followed by the affine merge candidates. The size of the merge list based on the subblock is signaled by the SPS and is set by the VTM3. The maximum allowed size of a merge list based on this subblock is fixed at 5. The sub-CU size used in TMVP is fixed at 8x8, and the affine merge mode As in the previous case, SbTMVP mode is only applicable to CUs with both width and height equal to or greater than 8. The encoding logic of the additional SbTMVP merge candidates is the same as that of the other merge candidates. That is, for each CU in a P or B slice, it is decided whether to use the SbTMVP candidate. Perform additional RD testing to determine if

[0059] The current VVC introduces new inter-mode encoding and decoding tools, but does not support spatial and For HEVC and current VVC scaling processes for derivation of temporal motion candidates The constraints on the long-term reference images that exist are not well defined in some new tools. The presentation presents some of the new inter-mode encoding / decoding tools related to long-term reference images. We propose some constraints.

[0060] According to the present disclosure, inter-mode coding of inter-mode coded blocks During the operation of the decoding tool, the interface involved in the operation of the inter-mode encoding / decoding tool One or more of the reference pictures associated with the pixel-mode coded block are long-term reference pictures. and then, based on that determination, This places constraints on the operation of intermode encoding / decoding tools on the network.

[0061] According to one embodiment of the present disclosure, the inter-mode encoding / decoding tool performs pairwise averaging. Includes generating merge candidates.

[0062] In one example, the averaged merge candidates involved in generating the pairwise averaged merge candidates are A predetermined set of reference images is composed of one reference image that is a long-term reference image and another reference image that is not a long-term reference image. If generated from a candidate pair, the averaged merge candidate is considered invalid.

[0063] In the same example, from a given candidate pair of two reference images, both of which are long-term reference images, Scaling operations are inhibited while generating averaged merge candidates from the .

[0064] According to another embodiment of the present disclosure, an inter mode encoding / decoding tool includes a BDOF, Inter-mode coded blocks are bidirectionally predicted blocks.

[0065] In one example, one reference image of the bidirectionally predicted blocks involved in the BDOF operation is a long-term reference image. The other reference image of the bidirectional prediction block involved in the BDOF operation is a long-term reference image. If it is determined that the image is not an image, the execution of BDOF is prohibited.

[0066] According to another embodiment of the present disclosure, an inter-mode encoding / decoding tool includes a DMVR, Inter-mode coded blocks are bidirectionally predicted blocks.

[0067] In one example, one reference image of the bidirectionally predicted blocks involved in the operation of DMVR is a long-term reference image. The other reference image of the bidirectionally predicted block involved in the DMVR operation is a long-term reference image. If it is determined that the image is not a statue, the execution of the DMVR will be prohibited.

[0068] In another example, one of the reference images of the bidirectionally predicted blocks involved in the DMVR operation is a long-term reference image. The other reference picture of the bidirectionally predicted block involved in the DMVR operation is the long-term reference If it is determined that it is not an image, the scope of DMVR implementation is the scope of integer pixel DMVR implementation. It is limited to the area.

[0069] According to another embodiment of the present disclosure, an inter-mode encoding / decoding tool may Includes derivation.

[0070] In one example, the motion vector candidates involved in deriving the MMVD candidates are referenced to a long-term reference image. If it is determined that it has its own motion vector pointing to the image, it will use the motion vector candidate as a basis. It is prohibited to use this motion vector as a main motion vector (also called a leading motion vector).

[0071] In the second example, one of the inter-mode coded blocks involved in the derivation of the MMVD candidates is One reference picture is a long-term reference picture, and the other is an inter-mode code involved in the derivation of MMVD candidates. The other reference picture of the converted block is not a long-term reference picture, and the basic motion vector is In the case of a bidirectional motion vector, it points to the long-term reference image and the bidirectional basic motion vector. The motion vector signaled by the signal is used for one motion vector that is also included in the vector. It is forbidden to change this one motion vector in the motion vector difference (MVD).

[0072] In the same second example, the proposed MVD modification process would instead be as shown in the box below: The highlighted part of the text is changed from the existing MVD change process in the current VVC. Shows proposed changes to [Table 1]

[0073] In the third example, a small number of inter-mode coded blocks are involved in the derivation of MMVD candidates. At least one reference picture is a long-term reference picture, and the base motion vectors are bidirectional. If it is a vector, scaling is prohibited in the derivation of the final MMVD candidate. will be done.

[0074] According to one or more embodiments of the present disclosure, an inter mode encoding / decoding tool includes: Includes derivation of MVD candidates.

[0075] In one example, a candidate motion vector may have its motion vector pointing to a reference image that is a long-term reference image. If it is determined that the candidate motion vector has a vector, the candidate motion vector is used as the base motion vector. It is prohibited to do so.

[0076] In one example, a small number of inter-mode coded blocks are involved in the derivation of SMVD candidates. At least one reference picture is a long-term reference picture, and the base motion vector is a bidirectional motion vector. If the vector is a long-term reference image, it refers to the long-term reference image and is included in the bidirectional base motion vector. For one motion vector, one motion vector is generated by the MVD signaled by the signal. In other examples, the inputs involved in the derivation of the SMVD candidates are One reference image of the target-mode coded block is the long-term reference image, and the SMVD candidate The other reference image of the inter-mode coded block involved in the derivation of If the base motion vector is a bidirectional motion vector, the long-term reference For one motion vector that points to an image and is also included in the bidirectional basic motion vector, , the change of the one motion vector by the MVD notified by the signal is prohibited.

[0077] According to another embodiment of the present disclosure, the inter-mode encoding / decoding tool performs weighted averaging. Bi-prediction is included, and inter-mode coded blocks are bi-directionally predicted blocks.

[0078] In one example, at least one reference of a bidirectionally predicted block involved in bi-prediction with weighted averaging is If the reference image is determined to be a long-term reference image, the use of unequal weighting is prohibited. .

[0079] According to another embodiment of the present disclosure, an inter-mode encoding / decoding tool is The inter-mode coded block is derivation of the SbTMVP coded block. It is.

[0080] In one example, for deriving motion vector candidates for a conventional TMVP coded block, The same constraints are applied to the motion vector candidates for SbTMVP coded blocks. is used to derive

[0081] One refinement of the previous example is to combine a conventional TMVP-encoded block and SbTMVP Constraints on the derivation of motion vector candidates used for both coded blocks are , of the two reference images consisting of the target reference image and the reference image for the temporally adjacent block. , if one reference image is a long-term reference image and the other reference image is not a long-term reference image. The motion vector of the neighboring block is regarded as invalid, while the target reference image and the neighboring block are regarded as invalid. When both of the reference images for the locking are long-term reference images, the motion of the spatially adjacent blocks is The scaling process for the motion vector of the spatially adjacent block is prohibited. and directly using the vector as a motion vector prediction for the current block.

[0082] According to another embodiment of the present disclosure, an inter mode encoding / decoding tool is This involves using an affine motion model in the derivation of the complement.

[0083] In one example, the reference image involved in using the affine motion model is determined to be a long-term reference image. In this case, the use of affine motion models in deriving motion vector candidates is prohibited.

[0084] In one or more examples, the functionality described above may be implemented in hardware, software, firmware, or other similar devices. If implemented in software, These functions may be implemented as one or more instructions or code on a computer-readable medium. stored in or transmitted through the hardware and executed by the processing unit A computer-readable medium is a computer program that corresponds to a tangible medium such as a data storage medium. computer-readable storage medium or from a location, for example, according to a communication protocol It may include communication media, including any medium that facilitates transfer of a computer program from one location to another. Thus, computer-readable media generally includes: (1) a non-transitory tangible copy; (2) a computer-readable storage medium, or (3) a communication medium such as a signal or carrier wave The data storage medium may include instructions, code, and one or more computers or one or more The medium may be any available medium that can be accessed by multiple processors. A computer program product may include a computer-readable medium.

[0085] Furthermore, the above method can be implemented in an application specific integrated circuit (ASIC), a digital signal processor ( DSP), Digital Signal Processor (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Arrays (FPGAs), Controllers, Microcontrollers One or more circuits containing a controller, microprocessor, or other electronic components To realize the above method, these circuits may be integrated into other hardware. The above-disclosed Each module, submodule, unit, or subunit that is may be implemented at least in part using the circuitry of

[0086] Other embodiments of the invention may be apparent from consideration of the specification and practice of the invention disclosed herein. This application does not include any modifications of the invention that follow the general principles of the invention. This disclosure is intended to cover any modification, use, or application of the present invention, and does not preclude departures from the present disclosure. as within known or customary practice in the art. The embodiments are to be considered as examples only, and the true scope and spirit of the invention is not to be construed as limiting the scope of the invention as defined by the accompanying drawings. As set forth in the claims.

[0087] The present invention is not limited to the embodiments described above and illustrated in the accompanying drawings, and various modifications and variations are possible. The scope of the present invention is defined in the appended patents. It is intended to be limited only by the scope of the claims that follow.

Claims

1. Inter-mode coded blocks involved in the operation of the inter-mode coding / decoding tool determining whether one or more of the reference images associated with the block are long-term reference images; 、 Based on the determination, the inter-mode coding is performed for the inter-mode coded block. Restricting the operation of code encoding / decoding tools; 1. A method for video encoding and decoding, comprising:

2. The inter-mode encoding / decoding tool includes generating pairwise averaging merge candidates. The method of claim 1 .

3. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule The averaged merge candidate is a reference image that is a long-term reference image and another reference image that is not a long-term reference image. and the first reference image is generated from a predetermined candidate pair of the first reference image. Deeming the page candidate invalid; A pair of averaged merge candidates is selected from a given pair of two reference images, both of which are long-term reference images. inhibiting scaling operations while generating the complement; The method of claim 2 , comprising:

4. The inter-mode encoding / decoding tool includes a bidirectional optical flow (BDOF), The method of claim 1 , wherein the mode-coded block is a bidirectionally predicted block.

5. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? One of the reference images of the bidirectionally predicted block is a long-term reference image, If it is determined that the other reference image of the block is not a long-term reference image, the execution of BDOF is prohibited. To The method of claim 4, comprising:

6. The inter-mode encoding / decoding tool implements decoder-side motion vector refinement (DMVR).

10. The method of claim 1, wherein the inter-mode coded block is a bidirectionally predicted block. The method described below.

7. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? One of the reference images of the bidirectionally predicted block is a long-term reference image, If it is determined that the other reference image of the block is not a long-term reference image, the execution of the DMVR is prohibited. To The method of claim 6, comprising:

8. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? One of the reference images of the bidirectionally predicted block is a long-term reference image, If it is determined that the other reference image of the block is not a long-term reference image, the scope of the DMVR execution is Limiting the scope of integer pixel DMVR implementation The method of claim 6, comprising:

9. The inter-mode encoding / decoding tool is a merge mode based on motion vector difference (MMV D) candidate derivation.

10. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule If a candidate motion vector decides to have its motion vector pointing to a reference picture that is a long-term reference picture, When the motion vector candidate is set, the motion vector candidate is used as the base motion vector (also called the leading motion vector). (This will be disclosed) 10. The method of claim 9, comprising:

11. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule At least one reference picture for the inter-mode coded block is a long-term reference picture. and if the base motion vector is a bidirectional motion vector, the long-term reference image is and the bidirectional basic motion vector is included in the signal of one motion vector. Inhibiting changes due to reported motion vector difference (MVD) 10. The method of claim 9, comprising:

12. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule At least one reference picture for the inter-mode coded block is a long-term reference picture. and if the base motion vector is a bidirectional motion vector, the final MMVD Prohibiting scaling in candidate derivation 10. The method of claim 9, comprising:

13. The inter-mode encoding / decoding tool includes a deriving set of symmetric motion vector difference (SMVD) candidates. The method of claim 1 , further comprising:

14. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule A motion vector candidate has its motion vector pointing to a reference image that is a long-term reference image. When it is determined that the motion vector candidate is a basic motion vector, the motion vector candidate is prohibited from being used as a basic motion vector. To stop 14. The method of claim 13, comprising:

15. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule At least one reference picture for the inter-mode coded block is a long-term reference picture. and if the base motion vector is a bidirectional motion vector, the long-term reference image is and the bidirectional basic motion vectors are included in the signal of one motion vector. Prohibit changes by notified MVD 14. The method of claim 13, comprising:

16. The inter mode encoding / decoding tool includes bi-prediction with weighted averaging, and The method of claim 1 , wherein the mode-coded block is a bidirectionally predicted block.

17. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? If it is determined that at least one reference image of the bidirectionally predicted block is a long-term reference image, prohibit the use of unequal weighting in 17. The method of claim 16, comprising:

18. The inter mode encoding / decoding tool includes deriving motion vector candidates, Mode-coded blocks are subjected to sub-block-based temporal motion vector prediction (SbTM) 2. The method of claim 1, wherein the block is a VP (Propagation-Propagation) coded block.

19. Operation of the intermode encoding / decoding tool for SbTMVP encoded blocks Constraining the work In deriving the motion vector candidates for the SbTMVP coded block, Motion vectors for future Temporal Motion Vector Prediction (TMVP) coded blocks Use the same restrictions on candidate derivation 20. The method of claim 18, comprising:

20. The conventional TMVP encoded block and the SbTMVP encoded block The constraints on the derivation of the motion vector candidates used for both One of the two reference images consisting of the target reference image and the reference image for the adjacent block If the reference image is a long-term reference image and the other reference image is not a long-term reference image, Deeming the motion vectors of the adjacent blocks (also called parallel blocks) invalid; If both the target reference picture and the reference picture for the neighboring block are long-term reference pictures, If the motion vector of the spatially adjacent block is larger than the predetermined value, the scaling process is prohibited. and uses the motion vector of the spatially neighboring block as the motion vector for the current block. directly as a rule prediction, 20. The method of claim 19, comprising:

21. The inter mode encoding / decoding tool uses affine motion in deriving motion vector candidates. The method of claim 1 , comprising using a metric model.

22. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule determining that the reference image involved in using the affine motion model is a long-term reference image; prohibiting the use of an affine motion model in deriving the motion vector candidates when 22. The method of claim 21, comprising:

23. one or more processors; a non-transitory memory coupled to the one or more processors; a plurality of programs stored in the non-transitory memory; Including, The plurality of programs, when executed by the one or more processors, To a computing device, Inter-mode coded blocks involved in the operation of the inter-mode coding / decoding tool determining whether one or more of the reference images associated with the block are long-term reference images; Based on the determination, the inter-mode coding is performed for the inter-mode coded block. restricting the operation of code encoding / decoding tools, A computing device that performs operations such as

24. The inter-mode encoding / decoding tool includes generating pairwise averaging merge candidates.

24. The computing device of claim 23.

25. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule The averaged merge candidate is a reference image that is a long-term reference image and another reference image that is not a long-term reference image. and the first reference image is generated from a predetermined candidate pair of the first reference image. Deeming the page candidate invalid; A pair of averaged merge candidates is selected from a given pair of two reference images, both of which are long-term reference images. inhibiting scaling operations while generating the complement; 25. The computing device of claim 24, comprising:

26. The inter-mode encoding / decoding tool includes a BDOF, and the inter-mode encoded 24. The computing device of claim 23, wherein the selected block is a bi-predicted block. 。

27. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? One of the reference images of the bidirectionally predicted block is a long-term reference image, If it is determined that the other reference image of the block is not a long-term reference image, the execution of BDOF is prohibited. To 27. The computing device of claim 26, comprising:

28. The intermode encoding / decoding tool includes a DMVR, and the intermode encoded 24. The computing device of claim 23, wherein the selected block is a bi-predicted block. 。

29. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? One of the reference images of the bidirectionally predicted block is a long-term reference image, If it is determined that the other reference image of the block is not a long-term reference image, the execution of the DMVR is prohibited. To 30. The computing device of claim 28, comprising:

30. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? One of the reference images of the bidirectionally predicted block is a long-term reference image, If it is determined that the other reference image of the block is not a long-term reference image, the scope of the DMVR execution is Limiting the scope of integer pixel DMVR implementation 30. The computing device of claim 28, comprising:

31. 24. The method of claim 23, wherein the inter mode encoding and decoding tool includes deriving MMVD candidates. On-board computing device.

32. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule If a candidate motion vector decides to have its motion vector pointing to a reference picture that is a long-term reference picture, When the motion vector candidate is set, the motion vector candidate is used as the base motion vector (also called the leading motion vector). (This will be disclosed) 32. The computing device of claim 31, comprising:

33. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule At least one reference picture for the inter-mode coded block is a long-term reference picture. and if the base motion vector is a bidirectional motion vector, the long-term reference image is and the bidirectional basic motion vector is included in the signal of one motion vector. Prohibit changes by notified MVD 32. The computing device of claim 31, comprising:

34. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule At least one reference picture for the inter-mode coded block is a long-term reference picture. and if the base motion vector is a bidirectional motion vector, the final MMVD Prohibiting scaling in candidate derivation 32. The computing device of claim 31, comprising:

35. 24. The method of claim 23, wherein the inter mode encoding and decoding tool includes SMVD candidate derivation. On-board computing device.

36. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule A motion vector candidate has its motion vector pointing to a reference image that is a long-term reference image. When it is determined that the motion vector candidate is a basic motion vector, the motion vector candidate is prohibited from being used as a basic motion vector. To stop 36. The computing device of claim 35, comprising:

37. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule At least one reference picture for the inter-mode coded block is a long-term reference picture. and if the base motion vector is a bidirectional motion vector, the long-term reference image is and the bidirectional basic motion vectors are included in the signal of one motion vector. Prohibit changes by known MVD 36. The computing device of claim 35, comprising:

38. The inter mode encoding / decoding tool includes bi-prediction with weighted averaging, and 24. The computer system of claim 23, wherein the mode-coded block is a bidirectionally predicted block. Coating device.

39. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? If it is determined that at least one reference image of the bidirectionally predicted block is a long-term reference image, prohibit the use of unequal weighting in 39. The computing device of claim 38, comprising:

40. The inter mode encoding / decoding tool includes deriving motion vector candidates, 24. The method of claim 23, wherein the mode-encoded block is an SbTMVP-encoded block. A computing device as described herein.

41. Operation of the intermode encoding / decoding tool for SbTMVP encoded blocks Constraining the work In deriving the motion vector candidates for the SbTMVP coded block, Similar to the previous restrictions on the derivation of motion vector candidates for TMVP coded blocks, Use the same thing 41. The computing device of claim 40, comprising:

42. The conventional TMVP encoded block and the SbTMVP encoded block The constraints on the derivation of the motion vector candidates used for both One of the two reference images consisting of the target reference image and the reference image for the adjacent block If the reference image is a long-term reference image and the other reference image is not a long-term reference image, Deeming the motion vectors of the intermediate neighboring blocks invalid; If both the target reference picture and the reference picture for the neighboring block are long-term reference pictures, If the motion vector of the spatially adjacent block is larger than the predetermined value, the scaling process is prohibited. and uses the motion vector of the spatially neighboring block as the motion vector for the current block. directly as a rule prediction, 42. The computing device of claim 41, comprising:

43. The inter mode encoding / decoding tool uses affine motion in deriving motion vector candidates.

24. The computing device of claim 23, further comprising using a metric model.

44. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule determining that the reference image involved in using the affine motion model is a long-term reference image; prohibiting the use of an affine motion model in deriving the motion vector candidates when 44. The computing device of claim 43, comprising:

45. A plurality of methods executed by a computing device having one or more processors A non-transitory computer-readable storage medium storing a program, The plurality of programs, when executed by the one or more processors, To a computing device, Regarding inter-mode coded blocks involved in the operation of the inter-mode coding tool determining whether one or more of the successive reference images are long-term reference images; Based on the determination, the inter-mode coding is performed for the inter-mode coded block. restricting the operation of code encoding / decoding tools, A non-transitory computer-readable storage medium that causes an operation such as

46. The inter-mode encoding / decoding tool includes generating pairwise averaging merge candidates.

46. ​​The non-transitory computer-readable storage medium of claim 45.

47. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule The averaged merge candidate is a reference image that is a long-term reference image and another reference image that is not a long-term reference image. and the first reference image is generated from a predetermined candidate pair of the first reference image. Deeming the page candidate invalid; A pair of averaged merge candidates is selected from a given pair of two reference images, both of which are long-term reference images. inhibiting scaling operations while generating the complement; 46. ​​The non-transitory computer-readable storage medium of claim 45, comprising:

48. The inter-mode encoding / decoding tool includes a BDOF, and the inter-mode encoded 46. ​​The non-transitory computer of claim 45, wherein the selected block is a bi-directionally predicted block. Readable storage medium.

49. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? One of the reference images of the bidirectionally predicted block is a long-term reference image, If it is determined that the other reference image of the block is not a long-term reference image, the execution of BDOF is prohibited. To 49. The non-transitory computer-readable storage medium of claim 48, comprising:

50. The intermode encoding / decoding tool includes a DMVR, and the intermode encoded 46. ​​The non-transitory computer of claim 45, wherein the selected block is a bi-directionally predicted block. Readable storage medium.

51. Constraining the operation of the inter mode coding tool for bidirectionally predicted blocks comprises: One of the reference images of the bidirectionally predicted block is a long-term reference image, If it is determined that the other reference image of the block is not a long-term reference image, the execution of the DMVR is prohibited. To 51. The non-transitory computer-readable storage medium of claim 50, comprising:

52. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? One of the reference images of the bidirectionally predicted block is a long-term reference image, If it is determined that the other reference image of the block is not a long-term reference image, the scope of the DMVR execution is Limiting the scope of integer pixel DMVR implementation 51. The non-transitory computer-readable storage medium of claim 50, comprising:

53. 46. ​​The method of claim 45, wherein the inter mode encoding and decoding tool includes deriving MMVD candidates. A non-transitory computer-readable storage medium.

54. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule If a candidate motion vector decides to have its motion vector pointing to a reference picture that is a long-term reference picture, When the motion vector candidate is set, the motion vector candidate is used as the base motion vector (also called the leading motion vector). (This will be disclosed) 54. The non-transitory computer-readable storage medium of claim 53, comprising:

55. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule At least one reference picture for the inter-mode coded block is a long-term reference picture. and if the base motion vector is a bidirectional motion vector, the long-term reference image is and the bidirectional basic motion vector is included in the signal of one motion vector. Prohibit changes by notified MVD 54. The non-transitory computer-readable storage medium of claim 53, comprising:

56. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule At least one reference picture for the inter-mode coded block is a long-term reference picture. and if the base motion vector is a bidirectional motion vector, the final MMVD Prohibiting scaling in candidate derivation 54. The non-transitory computer-readable storage medium of claim 53, comprising:

57. 46. ​​The method of claim 45, wherein the inter mode encoding and decoding tool includes SMVD candidate derivation. A non-transitory computer-readable storage medium.

58. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule A motion vector candidate has its motion vector pointing to a reference image that is a long-term reference image. When it is determined that the motion vector candidate is a basic motion vector, the motion vector candidate is prohibited from being used as a basic motion vector. To stop 58. The non-transitory computer-readable storage medium of claim 57, comprising:

59. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule At least one reference picture for the inter-mode coded block is a long-term reference picture. and if the base motion vector is a bidirectional motion vector, the long-term reference image is and the bidirectional basic motion vectors are included in the signal of one motion vector. Prohibit changes by notified MVD 58. The non-transitory computer-readable storage medium of claim 57, comprising:

60. The inter mode encoding / decoding tool includes bi-prediction with weighted averaging, and 46. ​​The non-temporal block of claim 45, wherein the mode-coded block is a bidirectionally predicted block. A computer-readable storage medium.

61. Constraining the operation of the inter mode encoding / decoding tool for bidirectionally predicted blocks. What is that? If it is determined that at least one reference image of the bidirectionally predicted block is a long-term reference image, prohibit the use of unequal weighting in 61. The non-transitory computer-readable storage medium of claim 60, comprising:

62. The inter mode encoding / decoding tool includes deriving motion vector candidates, 46. ​​The method of claim 45, wherein the mode-encoded block is an SbTMVP-encoded block.

2. A non-transitory computer-readable storage medium as described herein.

63. Operation of the intermode encoding / decoding tool for SbTMVP encoded blocks Constraining the work In deriving the motion vector candidates for the SbTMVP coded block, Similar to the previous restrictions on the derivation of motion vector candidates for TMVP coded blocks, Use the same thing 63. The non-transitory computer-readable storage medium of claim 62, comprising:

64. The conventional TMVP encoded block and the SbTMVP encoded block The constraints on the derivation of the motion vector candidates used for both One of the two reference images consisting of the target reference image and the reference image for the adjacent block If the reference image is a long-term reference image and the other reference image is not a long-term reference image, Deeming the motion vectors of the intermediate neighboring blocks invalid; If both the target reference picture and the reference picture for the neighboring block are long-term reference pictures, If the motion vector of the spatially adjacent block is larger than the predetermined value, the scaling process is prohibited. and uses the motion vector of the spatially neighboring block as the motion vector for the current block. directly as a rule prediction, 64. The non-transitory computer-readable storage medium of claim 63, comprising:

65. The inter mode encoding / decoding tool uses affine motion in deriving motion vector candidates.

46. ​​The non-transitory computer readable method of claim 45, comprising using a model. storage medium.

66. The inter-mode coding / decoding tool for the inter-mode coded block Constraining the behavior of the rule determining that the reference image involved in using the affine motion model is a long-term reference image; prohibiting the use of an affine motion model in deriving the motion vector candidates when 66. The non-transitory computer-readable storage medium of claim 65, comprising: