Method, system and computer program for deriving temporal motion vector prediction candidates

By increasing the number of temporal MVP candidates based on contextual information, the method enhances video encoding/decoding accuracy and efficiency, addressing the limitations of conventional processes.

JP2025528985APending Publication Date: 2025-09-04TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024548551
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-06-08
Filing Date
2023-06-12
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Conventional video encoding processes limit the number of temporal motion vector predictor candidates, leading to suboptimal encoding/decoding accuracy due to the restricted insertion of temporal MVP candidates.

Method used

Implementing a method to insert up to a predetermined number of temporal MVP candidates based on contextual and previously coded information, enhancing the MVP list generation process.

Benefits of technology

Improves encoding/decoding accuracy by allowing for more optimal motion vector prediction, resulting in better video compression efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528985000001_ABST
    Figure 2025528985000001_ABST
Patent Text Reader

Abstract

Various embodiments described herein include methods and systems for encoding and decoding video. In one aspect, the method includes receiving video data from a video bitstream, the video data including a plurality of blocks including a first block. The method also includes obtaining a first syntax element from the video bitstream, the first syntax element indicating a quantity N of temporal motion vector predictor (TMVP) candidates for a motion vector predictor (MVP) list. The method further includes identifying a set of TMVP candidates, the set of TMVP candidates having a size less than or equal to N, and generating an MVP list using at least the set of TMVP candidates. The method also includes reconstructing the first block using the MVP list.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 403,642, entitled "Improved TMVP Candidates Derivation," filed September 2, 2022, and is a continuation of and claims priority to U.S. Patent Application No. 18 / 207,582, entitled "Systems and Methods for Temporal Motion Vector Prediction Candidate Derivation," filed June 8, 2023, all of which are incorporated herein by reference in their entireties.

[0002] [Technical field] The disclosed embodiments relate generally to video encoding and decoding, including, but not limited to, systems and methods for motion vector prediction candidate derivation. [Background technology]

[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. The electronic devices transmit, receive, or otherwise communicate digital video data over communication networks and / or store the digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding may be used to compress the video data according to one or more video coding standards before the video data is communicated or stored.

[0004] Several video codec standards have been developed. For example, video coding standards include AOMedia Video 1 (AV1), Versatile Video Coding (VVC), Joint Exploration test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Expert Group (MPEG) coding. Video coding generally uses prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. The goal of video coding is to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality.

[0005] HEVC, also known as H.265, is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4). Versatile Video Coding (VVC), also known as H.266, is a video compression standard intended as the successor to HEVC. ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (Version 1) and 2022 (Version 2). AV1 is an open video coding format designed as a replacement for HEVC. Validated version 1.0.0, including errata 1, was released on January 8, 2019. Summary of the Invention

[0006] As described above, a video stream may be encoded into a bitstream with compression and then transmitted to a decoder that can decode / decompress the video stream for viewing or further processing. Compression of the video stream may exploit spatial and temporal correlation of the video signal through spatial and / or motion compensated prediction. Motion compensated prediction may include inter-prediction. Inter-prediction may use one or more motion vectors to generate coded blocks using previously coded and decoded pixels. A decoder receiving the coded signal may recreate the blocks. As used herein, the term block may be interpreted as a predictive block, a coding block, or a coding unit (CU), depending on the context.

[0007] The motion vectors used to encode / decode a video signal can be from spatially adjacent blocks in the same frame as the block being encoded. In addition, the motion vectors can be from temporally adjacent blocks (e.g., from blocks in the previous or subsequent frame). The number of motion vector candidates is limited, for example, for purposes of coding efficiency. Conventionally, spatial motion vectors are preferred (e.g., temporal motion vectors are sometimes limited to 1).

[0008] In some embodiments, a method for encoding video is provided, the method including: (i) obtaining video data including a plurality of blocks including a first block; (ii) obtaining a first syntax element, the first syntax element indicating a quantity N of temporal motion vector predictor (TMVP) candidates for a motion vector predictor (MVP) list; (iii) identifying a set of TMVP candidates, the set of TMVP candidates having a size less than or equal to N; (iv) identifying a set of spatial MVP candidates; (v) generating an MVP list using the set of TMVP candidates and the set of spatial MVP candidates; and (vi) signaling the MVP list in a video bitstream.

[0009] In some embodiments, a method of decoding video is provided, comprising: (i) receiving, from a video bitstream, video data including a plurality of blocks including a first block; (ii) obtaining, from the video bitstream, a first syntax element, the first syntax element indicating a quantity N of temporal motion vector predictor (TMVP) candidates for a motion vector predictor (MVP) list; (iii) identifying a set of TMVP candidates, the set of TMVP candidates having a size less than or equal to N; (iv) identifying a set of spatial MVP candidates; (v) generating an MVP list using the set of TMVP candidates and the set of spatial MVP candidates; and (vi) reconstructing the first block using the MVP list.

[0010] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes control circuitry and a memory that stores one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.

[0011] According to some embodiments, a non-transitory computer-readable storage medium is provided that stores one or more instruction sets for execution by a computing device, the one or more instruction sets including instructions for performing any of the methods described herein.

[0012] Accordingly, devices and systems are disclosed, along with methods for encoding and decoding video, which can complement or replace conventional methods, devices and systems for video encoding / decoding.

[0013] The features and advantages described herein are not necessarily all-inclusive, and some additional features and advantages will become apparent to those skilled in the art, especially in view of the drawings, specification, and claims provided in this disclosure. Furthermore, it should be noted that the language used herein has been chosen primarily for ease of reading and educational purposes, and not necessarily to delineate or limit the subject matter described herein. [Brief explanation of the drawings]

[0014] In order that the present disclosure may be more fully understood, a more particular description may be had by reference to features of various embodiments, some of which are illustrated in the accompanying drawings. However, the accompanying drawings merely illustrate relevant features of the present disclosure, and therefore the description should not be considered necessarily limiting, as those skilled in the art may recognize other useful features as they would understand upon reading the present disclosure.

[0015] [Figure 1] 1 is a block diagram illustrating an exemplary communication system, according to some embodiments.

[0016] [Figure 2A] FIG. 2 is a block diagram illustrating exemplary elements of an encoder component, according to some embodiments.

[0017] [Figure 2B] 3 is a block diagram illustrating exemplary elements of a decoder component, according to some embodiments.

[0018] [Figure 3] FIG. 1 is a block diagram illustrating an exemplary server system, according to some embodiments.

[0019] [Figure 4A] FIG. 2 illustrates an exemplary coding tree structure, according to some embodiments. [Figure 4B] FIG. 2 illustrates an exemplary coding tree structure, according to some embodiments. [Figure 4C] FIG. 2 illustrates an exemplary coding tree structure, according to some embodiments. [Figure 4D] FIG. 2 illustrates an exemplary coding tree structure, according to some embodiments.

[0020] [Figure 5A] FIG. 10 illustrates MMDV search points in two reference frames, according to some embodiments.

[0021] [Figure 5B] FIG. 2 illustrates exemplary spatially neighboring motion candidates for motion vector prediction, according to some embodiments.

[0022] [Figure 5C] 1A-1C are diagrams illustrating examples of exemplary temporally neighboring motion candidates for motion vector prediction, according to some embodiments.

[0023] [Figure 5D] FIG. 10 illustrates exemplary block locations for deriving a temporal motion vector predictor according to some embodiments.

[0024] [Figure 5E] FIG. 10 illustrates an exemplary motion vector candidate generation for a single inter-predicted block, according to some embodiments.

[0025] [Figure 5F] FIG. 10 illustrates an exemplary motion vector candidate generation for a composite prediction block according to some embodiments.

[0026] [Figure 5G] FIG. 1 illustrates an exemplary motion vector candidate bank process, according to some embodiments.

[0027] [Figure 5H] FIG. 10 illustrates an exemplary motion vector predictor list construction order, according to some embodiments.

[0028] [Figure 5I] 4A and 4B are diagrams illustrating exemplary scanning orders of motion vector predictor candidate blocks according to some embodiments.

[0029] [Figure 6A] 1 is a flow diagram illustrating an exemplary method for encoding video, according to some embodiments.

[0030] [Figure 6B] 1 is a flow diagram illustrating an exemplary method for decoding video, according to some embodiments.

[0031] According to common practice, the various features illustrated in the drawings are not necessarily drawn to scale and like reference numerals may be used to refer to like features throughout the specification and drawings. DETAILED DESCRIPTION OF THE INVENTION

[0032] This disclosure describes, among other things, improvements to motion vector predictor derivation. The available slots for MVP candidates are limited in many video encoding processes. Due to the limited slots, many conventional processes limit the insertion of temporal MVP candidates to only one candidate, which may not be optimal (e.g., resulting in lower encoding / decoding accuracy). Embodiments described herein include inserting up to a predetermined number of temporal MVP candidates. The predetermined number may be based on contextual information and / or previously coded information.

[0033] Exemplary Systems and Devices 1 is a block diagram illustrating a communication system 100 according to some embodiments. Communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 through 120-m) that are communicatively coupled to each other via one or more networks. In some embodiments, communication system 100 is a streaming system for use with video-enabled applications, such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0034] Source device 102 includes a video source 104 (e.g., a camera component or media storage) and an encoder component 106. In some embodiments, video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). Encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from video source 104 may be higher in data volume than encoded video bitstream 108 generated by encoder component 106. Because encoded video bitstream 108 has a smaller data volume (less data) than the video stream from the video source, encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared to the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video data to network 110).

[0035] The one or more networks 110 represent any number of networks that convey information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wireline and / or wireless communication networks. The one or more networks 110 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0036] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from source device 102). Server system 112 includes a coder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, coder component 114 includes an encoder component and / or a decoder component. In various embodiments, coder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, coder component 114 is configured to decode encoded video bitstream 108 and re-encode the video data using a different encoding standard and / or methodology to generate encoded video data 116. In some embodiments, server system 112 is configured to generate multiple video formats and / or encodings from encoded video bitstream 108.

[0037] In some embodiments, server system 112 functions as a Media-Aware Network Element (MANE). For example, server system 112 may be configured to prune encoded video bitstream 108 to adapt potentially different bitstreams to one or more of electronic devices 120. In some embodiments, a MANE is provided separate from server system 112.

[0038] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, decoder component 122 is configured to decode encoded video data 116 to generate an output video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of electronic devices 120 does not include a display component (e.g., is communicatively coupled to an external display device and / or includes media storage). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.

[0039] The source device and / or the plurality of electronic devices 120 may also be referred to as “terminal devices” or “user devices.” In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smartphone, tablet, or laptop), a wearable device, a videoconferencing device, and / or other types of electronic devices.

[0040] In an exemplary operation of communication system 100, source device 102 transmits encoded video bitstream 108 to server system 112. For example, source device 102 may code a stream of pictures captured by the source device. Server system 112 may receive encoded video bitstream 108 and decode and / or encode encoded video bitstream 108 using coder component 114. For example, server system 112 may apply encoding to the video data that is more optimal for network transmission and / or storage. Server system 112 may transmit encoded video data 116 (e.g., one or more coded video bitstreams) to one or more of electronic devices 120. Each electronic device 120 may decode encoded video data 116 to recover and optionally display video pictures.

[0041] In some embodiments, the transmission is a one-way data transmission. One-way data transmission may be used in media serving applications, etc. In some embodiments, the transmission is a two-way data transmission. Two-way data transmission may be used in video conferencing applications, etc. In some embodiments, the coded video bitstream 108 and / or the coded video data 116 are encoded and / or decoded according to any of the video coding / compression standards described herein, such as HEVC, VVC, and / or AV1.

[0042] FIG. 2A is a block diagram illustrating exemplary elements of the encoder component 106 according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a different device than the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCB or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores pre-captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, where each pixel may contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can readily understand the relationship between pixels and samples. The following discussion focuses on samples.

[0043] The encoder component 106 is configured to code and / or compress pictures of a source video sequence into a coded video sequence 216 in real time or under any other time constraints required by the application. Enforcing an appropriate coding rate is one function of the controller 204. In some embodiments, the controller 204 controls and is operatively coupled to other functional units, as described below. Parameters set by the controller 204 may include rate control-related parameters (e.g., picture skip, quantizer, and / or lambda values ​​for rate-distortion optimization techniques), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller 204, as they may be relevant to the encoder component 106 being optimized for a particular system design.

[0044] In some embodiments, the encoder component 106 is configured to operate in a coding loop. In a simplified example, the coding loop includes a source coder 202 (responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture, for example) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to a (remote) decoder (when the compression between the symbols and the coded video bitstream is lossless). The reconstructed sample stream (sample data) may be input to a reference picture memory 208. Because decoding of the symbol stream yields bit-exact results independent of the decoder location (local or remote), the contents of the reference picture memory 208 are also bit-exact between the local and remote encoders. In this way, the prediction portion of the encoder interprets the same sample values ​​as reference picture samples that the decoder predicts when using prediction during decoding. This basic principle of reference picture synchronicity (and the resulting drift when synchronicity cannot be maintained due to, for example, channel errors) is known to those skilled in the art.

[0045] The operation of decoder 210 may be the same as a remote decoder, such as decoder component 122, which is described in detail below in connection with Figure 2B. However, with brief reference to Figure 2B, because symbols are available and the encoding / decoding of the symbols into a coded video sequence by entropy coder 214 and parser 254 may be lossless, the entropy decoding portion of decoder component 122, including buffer memory 252 and parser 254, may not be fully implemented in local decoder 210.

[0046] An observation that can be made at this point is that any decoder technology, excluding analysis / entropy decoding, present in the decoder may need to exist in substantially identical functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder technology may be omitted, as it may be the opposite of the decoder technology, which is exhaustively described. Only in certain areas is a more detailed description required and provided below.

[0047] As part of its operation, source coder 202 may perform motion-compensated predictive coding, which predictively codes an input frame with respect to one or more previously coded frames from the video sequence designated as reference frames. In this manner, coding engine 212 codes differences between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as prediction references for the input frame. Controller 204 may manage the coding operations of source coder 202, including, for example, setting parameters and subgroup parameters used to encode the video data.

[0048] The decoder 210 may decode coded video data of frames that may be designated as reference frames based on symbols created by the source coder 202. The operation of the coding engine 212 may advantageously be a lossy process. When the coded video data is decoded in a video decoder (not shown in FIG. 2A ), the reconstructed video sequence may be a replica of the source video sequence with some errors. The decoder 210 may replicate the decoding process that may be performed by a remote video decoder on the reference frames and store the reconstructed reference frames in the reference picture memory 208. In this way, the encoder component 106 locally stores copies of reconstructed reference frames that have common content as the reconstructed reference frames obtained by the remote video decoder (without transmission errors).

[0049] The predictor 206 may perform a predictive search for the coding engine 212. That is, for a new frame to be coded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc., which may serve as appropriate prediction references for the new picture. The predictor 206 may operate on a sample block-by-pixel block basis to find appropriate prediction references. In some cases, as determined by the search results obtained by the predictor 206, the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory 208.

[0050] The outputs of all of the aforementioned functional units may be subject to entropy coding in entropy coder 214. Entropy coder 214 converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing them into symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).

[0051] In some embodiments, the output of the entropy coder 214 is coupled to a transmitter. The transmitter may buffer coded video sequences created by the entropy coder 214 and prepare them for transmission over a communication channel 218, which may be a hardware / software link to a storage device that stores the coded video data. The transmitter may merge the coded video data from the source coder 202 with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown). In some embodiments, the transmitter may transmit additional data along with the coded video. The source coder 202 may include such data as part of the coded video sequence. The additional data may include other forms of redundant data, such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, supplemental enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0052] The controller 204 may manage the operation of the encoder component 106. During coding, the controller 204 may assign a particular coded picture type to each coded picture, which may affect the coding that may be applied to the respective picture. For example, pictures may often be assigned as intra pictures (I pictures), predicted pictures (P pictures), or bidirectionally predicted pictures (B pictures). Intra pictures may be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will be aware of these variations of I pictures and their respective uses and characteristics, and therefore will not be repeated here. Predicted pictures may be coded and decoded using intra prediction or inter prediction, using at most one motion vector and reference index to predict sample values ​​for each block. Bidirectionally predicted pictures can be coded and decoded using intra- or inter-prediction, using up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple-predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0053] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded relative to other (already coded) blocks, as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be nonpredictively coded, or they may be predictively coded relative to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be nonpredictively coded via spatial prediction or via temporal prediction relative to one previously coded reference picture. Blocks of a B-picture may be nonpredictively coded via spatial prediction or via temporal prediction relative to one or two previously coded reference pictures.

[0054] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture being encoded / decoded (called the current picture) is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0055] Encoder component 106 may perform coding operations according to a given video coding technique or standard, such as any of those described herein. In doing so, encoder component 106 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0056] 2B is a block diagram illustrating exemplary elements of a decoder component 122 according to some embodiments. The decoder component 122 of FIG. 2B is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to a loop filter unit 256 and configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0057] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more coded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each coded video sequence is independent of the other coded video sequences. Each coded video sequence may be received from channel 218, which is a hardware / software link to a storage device that stores the coded video data. The receiver may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be transferred respectively using entities (not shown). The receiver may separate the coded video sequence from the other data. In some embodiments, the receiver receives additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.

[0058] In some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra-picture prediction unit 262, a motion compensated prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. In some embodiments, the decoder component 122 is implemented at least partially in software.

[0059] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to combat network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., configured to address playback timing), a separate buffer memory is provided external to decoder component 122 (e.g., to combat network jitter). When receiving data from a sufficient bandwidth and controllable storage / forwarding device or from an isosynchronous network, buffer memory 252 may not be needed or may be small. For use over best-effort packet networks such as the Internet, buffer memory 252 may be needed, may be relatively large, may advantageously be adaptively sized, and may be implemented, at least in part, within an operating system or similar element (not shown) external to decoder component 122.

[0060] Parser 254 is configured to reconstruct symbols 270 from the coded video sequence. The symbols may include, for example, information used to manage the operation of decoder component 122 and / or information to control a rendering device such as display 124. The rendering device control information may be in the form of, for example, a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). Parser 254 parses (entropy decodes) the coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 254 may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. Subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. Parser 254 may also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0061] The reconstruction of symbols 270 may involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter and intra picture, inter and intra block) and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed by parser 254 from the coded video sequence. The flow of such subgroup control information between parser 254 and the following units is not shown for clarity.

[0062] In addition to the functional blocks already mentioned, the decoder component 122 may be conceptually subdivided into multiple functional units, as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the conceptual division into functional units will be maintained hereinafter.

[0063] The scalar / inverse transform unit 258 receives the quantized transform coefficients as well as control information (such as which transform to use, block size, quantization coefficients and / or quantization scaling matrices) as symbols 270 from the parser 254. The scalar / inverse transform unit 258 may output blocks containing sample values ​​that may be input to the aggregator 268.

[0064] In some cases, the output samples of the scaler / inverse transform unit 258 relate to intra-coded blocks; i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 may generate blocks of the same size and shape as the blocks during reconstruction using surrounding already reconstructed information fetched from the current (partially reconstructed) picture from the current picture memory 264. The aggregator 268 may add, on a sample-by-sample basis, the prediction information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258.

[0065] In other cases, the output samples of the scalar / inverse transform unit 258 relate to an inter-coded and potentially motion-compensated block. In such cases, the motion-compensated prediction unit 260 can access the reference picture memory 266 to fetch samples used for prediction. After motion-compensating the fetched samples according to symbols 270 associated with the block, these samples can be added by an aggregator 268 to the output of the scalar / inverse transform unit 258 (in this case, referred to as residual samples or a residual signal) to generate output sample information. The addresses in the reference picture memory 266 from which the motion-compensated prediction unit 260 fetches the prediction samples can be controlled by motion vectors. The motion vectors can be available to the motion-compensated prediction unit 260 in the form of symbols 270, which can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values ​​fetched from the reference picture memory 266 when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0066] The output samples of aggregator 268 may be subjected to various loop filtering techniques in loop filter unit 256. The video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video stream and made available to loop filter unit 256 as symbols 270 from parser 254, but may also be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, as well as to previously reconstructed loop-filtered sample values.

[0067] The output of the loop filter unit 256 may be a sample stream that can be output to a rendering device such as the display 124 and stored in the reference picture memory 266 for use in future inter-picture prediction.

[0068] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once the current coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 254), the current reference picture can become part of reference picture memory 266, and the fresh current picture memory can be reallocated before beginning reconstruction of a subsequent coded picture.

[0069] Decoder component 122 may perform decoding operations according to a given video compression technology, which may be documented in a standard, such as any of the standards described herein. The coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense of adhering to the syntax of the video compression technology or standard, as specified in the video compression technology documents and standards, and particularly in the profile documents therein. Also, for compliance with some video compression technologies and standards, the complexity of the coded video sequence may be within a range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference pixel size, etc. The limits set by the level may be further constrained, in some cases, through a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled with the coded video sequence.

[0070] 3 is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., a CPU, a GPU, and / or a DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).

[0071] The network interface 304 may be configured to interface with one or more communications networks (e.g., wireless, wired, and / or optical networks). The communications networks may be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of communications networks include local area networks such as Ethernet, WLAN, and the like; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wireline or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including CANBus and the like. Such communications may be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., from a CANbus to a specific CANbus device), or bidirectional (e.g., to another computer system using a local or wide-area digital network). Such communications may include communications to one or more cloud computing networks.

[0072] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input devices 310 may include one or more of a keyboard, a mouse, a trackpad, a touchscreen, a data glove, a joystick, a microphone, a scanner, a camera, etc. The output devices 308 may include one or more of an audio output device (e.g., a speaker), a visual output device (e.g., a display or monitor), etc.

[0073] Memory 314 may include high-speed random-access memory (such as DRAM, SRAM, DDR RAM, and / or other random-access solid-state memory devices) and / or non-volatile memory (such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Memory 314 optionally includes one or more storage devices located remotely from control circuitry 302. Memory 314, or alternatively, a non-volatile solid-state memory device within memory 314, comprises a non-transitory computer-readable storage medium. In some embodiments, memory 314, or the non-transitory computer-readable storage medium of memory 314, stores the following programs, modules, instructions, and data structures, or a subset or superset thereof: • an operating system 316 that contains procedures for handling various basic system services and performing hardware-dependent tasks; • A network communications module 318 used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); A coding module 320 for performing various functions related to encoding and / or decoding data, such as video data. In some embodiments, the coding module 320 is an instance of the coder component 114. The coding module 320 may include, but is not limited to, one or more of the following: a decoding module 322 that performs various functions related to decoding of encoded data, such as the functions described above with respect to the decoder component 122; an encoding module 340 for performing various functions related to encoding data, such as the functions described above with respect to the encoder component 106; A picture memory 352 that stores pictures and picture data, e.g., for use with coding module 320. In some embodiments, picture memory 352 includes one or more of reference picture memory 208, buffer memory 252, current picture memory 264, and reference picture memory 266.

[0074] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions described above with respect to the parser 254), a transform module 326 (e.g., configured to perform various functions described above with respect to the scalar / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions described above with respect to the motion compensated prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions described above with respect to the loop filter unit 256).

[0075] In some embodiments, encoding module 340 includes a code module 342 (e.g., configured to perform various functions described above with respect to source coder 202, coding engine 212, and / or entropy coder 214) and a prediction module 344 (e.g., configured to perform various functions described above with respect to predictor 206). In some embodiments, decoding module 322 and / or encoding module 340 include a subset of the modules shown in Figure 3. For example, a shared prediction module is used by both decoding module 322 and encoding module 340.

[0076] Each of the above-identified modules stored in memory 314 corresponds to a set of instructions for performing functions described herein. The above-identified modules (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules; thus, various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, coding module 320 optionally does not include separate decoding and encoding modules, but rather uses the same set of modules to perform both sets of functionality. In some embodiments, memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0077] In some embodiments, the server system 112 includes a web or Hypertext Transfer Protocol (HTTP) server, a File Transfer Protocol (FTP) server, and web pages and applications implemented using Common Gateway Interface (CGI) script, PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), Hypertext Markup Language (HTML), Extensible Markup Language (XML), Java, JavaScript, Asynchronous JavaScript and XML (AJAX), XHP, Javelin, Wireless Universal Resource Files (WURFL), and the like.

[0078] While FIG. 3 illustrates a server system 112 according to some embodiments, FIG. 3 is intended as a functional illustration of various features that may be present in one or more server systems, rather than a structural schematic of the embodiments described herein. In practice, those skilled in the art will recognize that items shown separately can be combined and some items can be separated. For example, some items shown separately in FIG. 3 can be implemented on a single server, and single items can be implemented by more than one server. The actual number of servers used to implement server system 112, and how functionality is allocated among them, may vary from one implementation to another and, optionally, depend in part, on the amount of data traffic the server system handles during peak and average usage periods.

[0079] Example Coding Approach 4A-4D show exemplary coding tree structures according to some embodiments. As shown in the first coding tree structure (400) of FIG. 4A, some coding approaches (e.g., VP9) use a 4-way partition tree from a 64x64 level down to a 4x4 level, with some additional restrictions on 8x8 blocks. In FIG. 4A, the partitions shown as R can be referred to as recursive in that the same partition tree is repeated at lower scales until the lowest 4x4 level is reached.

[0080] As shown in the second coding tree structure (402) of FIG. 4B, some coding approaches (e.g., AV1) expand the partition tree to a 10-way structure and increase the maximum size (e.g., called a superblock in VP9 / AV1 terminology) to start at 128x128. The second coding tree structure includes a 4:1 / 1:4 rectangular partition not found in the first coding tree structure. The partition type with three subpartitions in the second row of FIG. 4B is called a T-shaped partition. The rectangular partitions in this tree structure cannot be further subdivided. In addition to the coding block size, a coding tree depth can be defined to indicate the division depth from the root node. For example, the coding tree depth of the root node, e.g., 128x128, is set to 0, and after one further tree block division, the coding tree depth increases by 1.

[0081] As an example, instead of enforcing a fixed transform unit size as in VP9, ​​AV1 allows partitioning of luma coding blocks into transform units of multiple sizes that can be represented by recursive partitions down to a maximum of two levels. To incorporate AV1's extended coding block partitions, square, 2:1 / 1:2, and 4:1 / 1:4 transform sizes from 4x4 to 64x64 are supported. For chroma blocks, only the largest possible transform units are allowed.

[0082] As an example, a coding tree unit (CTU) may be split into coding units (CUs) by using a quad-tree structure, denoted as a coding tree, to adapt to various local characteristics in HEVC and the like. In some embodiments, the decision of whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU can be further split into one, two, or four PUs depending on the PU split type. The same prediction process is applied within one PU, and related information is transmitted to the decoder on a PU-by-PU basis. After obtaining a residual block by applying a prediction process based on the PU split type, the CU can be partitioned into TUs according to another quad-tree structure, such as the coding tree of the CU. One important feature of the HEVC structure is its multiple partition concepts, including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square, but a PU can be square or rectangular for inter-predicted blocks. In HEVC, one coding block is further split into four square sub-blocks, and a transform is performed on each sub-block (TU). Each TU can be further split recursively (using quad-tree splitting) into smaller TUs, which are called residual quad-trees (RQTs). At picture boundaries such as HEVC, implicit quad-tree splitting may be used so that blocks maintain quad-tree splitting until their size fits the picture boundary.

[0083] In VVC and other technologies, a quad tree with nested multitype trees using binary and ternary split segmentation structures can replace the concept of multiple partition unit types, eliminating the separation of the concepts of CU, PU, ​​and TU, except as needed for CUs with sizes too large for the maximum transform length, and supporting more flexibility for CU partition shapes. In the coding tree structure, CUs can have either square or rectangular shapes. ACTUs are first partitioned using a quadtree (also called a quadtree) structure. The quadtree leaf nodes can be further partitioned using a multitype tree structure. As shown in the third coding tree structure (404) in FIG. 4C, the multitype tree structure includes four split types. For example, the multitype tree structure includes a vertical binary split (SPLIT_BT_VER), a horizontal binary split (SPLIT_BT_HOR), a vertical ternary split (SPLIT_TT_VER), and a horizontal ternary split (SPLIT_TT_HOR). The leaf nodes of the multitype tree are called CUs, and this segmentation is used for prediction and transform processing without further partitioning, unless the CU is too large for the maximum transform length. This means that in most cases, CUs, PUs, and TUs have the same block size in a quad-tree with a nested multitype tree coding block structure. An exception occurs when the supported maximum transform length is smaller than the width or height of the color components of the CU. An example of block partitioning for one CTU (406) is shown in Figure 4D, which illustrates an exemplary quad-tree with a nested multitype tree coding block structure.

[0084] In VVC, etc., the maximum supported luma transform size may be 64x64, and the maximum supported chroma transform size may be 32x32. When the width or height of the CB is larger than the maximum transform width or height, the CB is automatically split horizontally and / or vertically to meet the transform size limit in that direction.

[0085] The coding tree scheme, such as in VTM7, supports the ability for luma and chroma to have separate block tree structures. In some cases, for P and B slices, the luma and chroma CTBs in one CTU share the same coding tree structure. However, for I slices, luma and chroma can have separate block tree structures. When the separate block tree mode is applied, the luma CTB is partitioned into CUs by one coding tree structure, and the chroma CTB is partitioned into chroma CUs by another coding tree structure. This means that a CU in an I slice may include or consist of a coding block for the luma component or a coding block for two chroma components, and a CU in a P or B slice may always include or consist of coding blocks for all three color components unless the video is monochrome.

[0086] To support extended coding block partitions, multiple transform sizes (e.g., ranging from 4 points to 64 points in each dimension) and transform shapes (e.g., square or rectangular with width / height ratios of 2:1 / 1:2 and 4:1 / 1:4) may be used in AV1 and elsewhere.

[0087] In merge mode, implicitly derived motion information can be directly used to generate predicted samples for the current CU. A merge mode with motion vector difference (MMVD) has been introduced in VVC. An MMVD flag can be signaled immediately after sending the skip flag and merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after a merge candidate is selected, it can be further refined by the signaled MVD information. The MVD information may include a merge candidate flag, an index specifying the magnitude of the motion, and an index indicating the direction of the motion. In merge mode, one of the first two merge candidate flags in the merge list can be used as the motion vector (MV) base. A merge candidate flag can be signaled to specify which flag is used.

[0088] The distance index specifies motion magnitude information and indicates a predefined offset from the starting point. Figure 5A shows MMDV search points in two reference frames according to some embodiments. As shown in Figure 5A, the offset can be added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1 below. [Table 1]

[0089] The direction index represents the direction of the MVD relative to the starting point. The direction index may represent one of four directions, as shown in Table 2 below. The meaning of the MVD code may change depending on the information of the starting MV. When the starting MV is a uni-predictive MV or a bi-predictive MV and both lists point to the same side of the current picture (e.g., the POCs of the two references are both greater than the POC of the current picture, or the POCs of the two references are both less than the POC of the current picture), the code in Table 2 specifies the code of the MV offset added to the starting MV. When the starting MV is a bi-predictive MV and the two MVs point to different sides of the current picture (e.g., the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), and the difference in POC in list 0 (L0) is greater than the difference in POC of list 1 (L1), the code in Table 2 specifies the sign of the MV offset added to the L0 MV component of the starting MV, and the sign of the L1 MV has an opposite value. If the difference in POC of L1 is greater than that of L0, the sign in Table 2 specifies the sign of the MV offset added to the L1 MV component of the starting MV, and the sign of the L0 MV has the opposite value.

[0090] In some embodiments, the MVD is scaled according to the difference in POC in each direction. For example, if the difference in POC in both lists is the same, no scaling is required. If the difference in POC in L0 is greater than that of L1, the MVD of L1 is scaled. If the difference in POC in L1 is greater than that of L0, the MVD of L0 is scaled in the same way. If the starting MV is uni-predicted, the MVD is added to the available MV. [Table 2]

[0091] As an example, in VVC, in addition to the MVD signaling of normal unidirectional prediction and bidirectional prediction modes, a symmetric MVD mode for bidirectional MVD signaling may be applied. In the symmetric MVD mode, motion information including both L0 and L1 reference picture indexes and the MVD of L1 may be derived (but not signaled).

[0092] The decoding process for symmetric MVD mode may be as follows: First, at the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived. For example, if mvd_l1_zero_flag is 1, then BiDirPredFlag is set equal to 0. If the nearest reference picture in L0 and the nearest reference picture in L1 form a forward and backward pair of reference pictures or a backward and forward pair of reference pictures, then BiDirPredFlag is set to 1, and both the L0 and L1 reference pictures are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. Next, at the CU level, if a CU is bi-predictively coded and BiDirPredFlag is equal to 1, then a symmetric mode flag is explicitly signaled, indicating whether symmetric mode is used or not. If the symmetric mode flag is true (e.g., equal to 1), only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The L0 and L1 reference indices are set equal to the reference picture pair, respectively. Finally, MVD1 is set equal to (-MVD0).

[0093] In some embodiments, for each coded block in an inter frame, if the mode of the current block is an inter-coding mode rather than a skip mode, a separate flag is signaled to indicate whether a single reference mode or a mixed reference mode is used for the current block. In a single reference mode, a predictive block may be generated by one motion vector. In a mixed reference mode, a predictive block is generated by a weighted average of two predictive blocks derived from two motion vectors. The modes that may be signaled for the single reference case are detailed in Table 3 below. [Table 3] The modes that can be signaled for the mixed reference case are detailed in Table 4 below. [Table 4]

[0094] Some standards, such as AV1, allow for a motion vector precision (or accuracy) of 1 / 8 pixel. Syntax can be used to signal the motion vector difference in reference frame list zero (L0) or list 1 (L1) as follows: For example, the syntax mv_joint specifies which components of the motion vector difference are non-zero. The syntax mv_joint value 0 indicates that there are no non-zero MVDs in either the horizontal or vertical direction, a value 1 indicates that there are non-zero MVDs in only the horizontal direction, a value 2 indicates that there are non-zero MVDs in only the vertical direction, and a value 3 indicates that there are non-zero MVDs in both the horizontal and vertical directions. The syntax mv_sign specifies whether the motion vector difference is positive or negative. The syntax mv_class specifies the class of the motion vector difference. A higher class means that the motion vector difference has a larger magnitude, as shown in Table 5 below. The syntax mv_bit specifies the integer part of the offset between the motion vector difference and the starting magnitude of each MV class. The syntax mv_fr specifies the first two fractional bits of the motion vector difference. The syntax mv_hp specifies the third fractional bit of the motion vector difference. [Table 5]

[0095] For NEW_NEARMV and NEAR_NEWMV modes (shown in Table 4 above), the precision of the MVD depends on the associated class and the magnitude of the MVD. For example, fractional MVD is only allowed if the MVD magnitude is 1 pixel or less. In addition, only one MVD value is allowed when the associated MV class value is MV_CLASS_1 or greater, and the MVD value for each MV class is derived as 4, 8, 16, 32, or 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5). The allowed MVD values ​​for each MV class are shown in Table 6 below. [Table 6]

[0096] In some embodiments, one context is used to signal mv_joint or mv_class if the current block is coded using NEW_NEARMV or NEAR_NEWMV mode, and another context is used to signal mv_joint or mv_class if the current block is not coded using NEW_NEARMV or NEAR_NEWMV mode.

[0097] The inter-coding mode JOINT_NEWMV may be applied to indicate whether the MVDs of two reference lists are jointly signaled. When the inter-prediction mode is equal to JOINT_NEWMV mode, the MVDs of reference L0 and reference L1 are jointly signaled. In this way, only one MVD named joint_mvd is signaled and transmitted to the decoder, and the delta MVs of reference L0 and reference L1 can be derived from joint_mvd. The JOINT_NEWMV mode is signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes.

[0098] When JOINT_NEWMV mode is signaled and the Picture Order Count (POC) distances between two reference frames and the current frame are different, the MVD is scaled for reference L0 or reference L1 based on the POC distance. For example, the distance between reference frame L0 and the current frame is denoted as td0, and the distance between reference frame L1 and the current frame is denoted as td1. If td0 is greater than or equal to td1, joint_mvd is used directly for reference L0, and the mvd of reference L1 is derived from joint_mvd based on the following Equation 1:

number

number

[0099] The inter-coding mode AMVDMV may be added to the single-reference case. When the AMVDMV mode is selected, this indicates that an adaptive MVD solution (AMVD) is applied to the signal MVD. A flag, such as amvd_flag, may be added under the JOINT_NEWMV mode to indicate whether AMVD is applied to the joint MVD coding mode. When the adaptive MVD solution is applied to the joint MVD coding mode, named joint AMVD coding, the MVDs for two reference frames are jointly signaled, and the accuracy of the MVD is implicitly determined by the size of the MVD. The MVDs for two (or more) reference frames are jointly signaled, and MVD coding is applied.

[0100] AMVR, first proposed in CWG-C012, supports a total of seven MV precisions (8, 4, 2, 1, 1 / 2, 1 / 4, and 1 / 8). For each prediction block, the AVM encoder explores all supported precision values ​​and signals the best precision to the decoder. To reduce encoder runtime, two precision sets are supported. Each precision set contains four predefined precisions. The precision set is adaptively selected at the frame level based on the frame's maximum precision value. Similar to AV1, the maximum precision is signaled in the frame header. Table 7 summarizes the supported precision values ​​based on the frame-level maximum precision. [Table 7]

[0101] In the AOM Video Model (AVM), like AV1, there is a frame-level flag that indicates whether the motion vectors of a frame contain sub-pel precision. AMVR is enabled only if the value of the cur_frame_force_integer_mv flag is 0. In AMVR, if the block precision is less than the maximum precision, the motion model and interpolation filter are not signaled. If the block precision is less than the maximum precision, the motion mode is estimated to translation motion and the interpolation filter is estimated to the regular interpolation filter. Similarly, if the block precision is either 4-pel or 8-pel, the inter-intra mode is not signaled and is assumed to be 0.

[0102] Some embodiments include a spatial motion vector predictor (SMVP), a temporal motion vector predictor (TMVP), extra MV candidates, derived MVPs, and / or reference bank MVPs. For example, a fixed-size stack (e.g., MVP list) can be generated at both the encoder and decoder sides to store MVPs.

[0103] The spatial motion vector predictor is derived from spatial neighboring blocks, including adjacent spatial neighboring blocks that are direct neighbors to the top and left of the current block, and non-adjacent spatial neighboring blocks that are close to, but not directly adjacent to, the current block. An example set of spatial neighboring blocks for a luma block is shown in Figure 5B (e.g., each spatial neighboring block is an 8x8 block).

[0104] Spatial neighboring blocks may be examined to find one or more MVs associated with the same reference frame index as the current block. Taking the current block as an example, the search order for spatially neighboring 8x8 luma blocks is as shown by numbers 1-8 in FIG. 5B. In some embodiments, fewer spatial neighboring blocks are scanned (e.g., numbers 5 and 7 are skipped). In this example, first, the upper neighboring row is checked from left to right. Next, the left neighboring column is checked from top to bottom. Third, the upper right neighboring block is checked. Fourth, the neighboring block of the upper left block is checked. Fifth, the next non-adjacent row above is checked from left to right. Sixth, the next non-adjacent column to the left is checked from top to bottom. Seventh, the next non-adjacent row above is checked from left to right. Eighth, the next non-adjacent column to the left is checked from top to bottom.

[0105] In some embodiments, neighboring candidates (e.g., numbers 1-3 in FIG. 5B) are placed first in the MV predictor list, before TMVP candidates. In some embodiments, non-neighboring candidates (e.g., numbers 4-8 in FIG. 5B) are placed in the MV predictor list after TMVP candidates. In this example, all SMVP candidates should have the same reference picture as the current block. If the current block has a single reference picture, MVP candidates with a single reference picture should have the same reference picture. For blocks with mixed reference pictures (e.g., two reference pictures), one of the reference pictures should be the same reference picture as the current block. If the current block has two reference pictures, only MVP candidates with both of the same reference pictures are added to the MVP list.

[0106] In addition to spatially neighboring blocks, MV predictors known as temporal MV predictors can also be derived using collocated blocks in a reference frame. For example, to generate a temporal MV predictor, the MVs of the reference frames are stored together with the reference indexes associated with the respective reference frames. Then, for each 8x8 block in the current frame, the MVs of the reference frames whose trajectories pass through the 8x8 block are identified and stored in a temporal MV buffer together with the reference frame index. For example, in inter prediction using a single reference frame, regardless of whether the reference frame is a forward reference frame or a backward reference frame, the MVs are stored in 8x8 units to perform temporal motion vector prediction of future frames. As another example, in hybrid inter prediction, only forward MVs are stored in 8x8 units to perform temporal motion vector prediction of future frames.

[0107] FIG. 5C illustrates exemplary temporally neighboring motion candidates for motion vector prediction according to some embodiments. In the example of FIG. 5C, the MV of reference frame 1 (R1), MVref, points from R1 to the reference frame of R1. In doing so, MVref passes through an 8x8 block of the current frame. MVref can be stored in a temporal MV buffer associated with this 8x8 block. During the motion projection process for deriving a temporal MV predictor, the reference frames can be scanned in a predefined order, for example, in the order of LAST_FRAME, BWDREF_FRAME, ALTREF_FRAME, ALTREF2_FRAME, and LAST2_FRAME. As an example, an MV from a reference frame indexed higher in the scanning order may not replace a previously identified MV assigned by a reference frame indexed lower in the scanning order.

[0108] Given predefined block coordinates, the associated MV stored in the temporal MV buffer can be identified and projected onto the current block to derive a temporal MV predictor from the current block that points to its reference frame, e.g., MV0 in Figure 5C.

[0109] Figure 5D shows exemplary block positions for deriving a temporal motion vector predictor according to some embodiments. Figure 5D shows predefined block positions for deriving a temporal MV predictor for a 16x16 block. For example, up to seven blocks are checked for valid temporal MV predictors. Temporal MV predictors may be checked after adjacent spatial MV predictors but before non-adjacent spatial MV predictors (as described above with respect to Figure 5B).

[0110] In Figure 5D, B0, B1, B2, and B3 may be referred to as inner TMVP blocks, and B4, B5, and B6 may be referred to as outer TMVP blocks. In some embodiments, the inner TMVP blocks are checked before the outer TMVP blocks. For example, the scan order of the three TMVP locations shown in Figure 5D may be B0 → B1 → B2 → B3 → B4 → B5 → B6.

[0111] For MV predictor derivation, all spatial and temporal MV candidates can be pooled, and each predictor can be assigned a weight determined during scanning of spatial and temporal neighboring blocks. Based on the associated weight, the candidates can be sorted and ranked. As an example, up to four candidates are identified and added to an MV predictor list. This list of MV predictors, sometimes referred to as a dynamic reference list (DRL), can be used in dynamic MV prediction mode. If the MVP list is not full after scanning the spatial and temporal candidates, additional searches can be performed, and additional MVP candidates can be used to fill the MVP list. For example, the additional MVP candidates can include global MVs, zero MVs, combined composite MVs without scaling, etc.

[0112] In some embodiments, adjacent SMVP candidates, TMVP candidates, and / or non-adjacent SMVP candidates added to the MVP list are reordered. For example, the reordering process may be based on a weight assigned to each candidate. Candidate weights may be predefined based on the overlap area between the current block and the candidate block. In some embodiments, the weighting of non-adjacent (external) SMVP and TMVP candidates is not considered during the reordering process (e.g., the reordering process only affects adjacent candidates).

[0113] The derived MVP candidates may include both derived MVPs for a single reference picture and mixed modes. For single inter prediction, if the reference frame of a neighboring block is different from the reference frame of the current block but they are in the same direction, a temporal scaling algorithm can be used to scale the MV to that reference frame to form an MVP for the motion vector of the current block. Figure 5E shows an example motion vector candidate generation for a single inter prediction block according to some embodiments. As shown in Figure 5E, mv1 from neighboring block A is used to derive an MVP for the motion vector mv0 of the current block using temporal scaling.

[0114] In mixed inter prediction, synthesized MVs from different neighboring blocks are used to derive the MVP of the current block, but the reference frame of the synthesized MV must be the same as that of the current block. Figure 5F shows an example of motion vector candidate generation for a mixed prediction block according to some embodiments. As shown in Figure 5F, the synthesized MVs (mv2, mv3) have the same reference frame as the current block but are from different neighboring blocks.

[0115] Some embodiments include a reference motion vector candidate bank as shown in FIG. 5G. For example, each buffer corresponds to a unique reference frame type, corresponding to a single or pair of reference frames covering single inter mode and mixed inter mode, respectively. In some embodiments, all buffers are the same size. In some embodiments, when a new MV is added to a full buffer, existing MVs are evicted to make space for the new MV.

[0116] A coding block can refer to an MV candidate bank to collect reference MV candidates, for example, in addition to those obtained in the reference MV list generation described previously. For example, after coding a superblock, the MV bank is updated with the MVs used by the coding blocks of the superblock. Each tile can have an independent MV reference bank that is utilized by all superblocks within the tile. For example, at the start of encoding each tile, the corresponding bank is emptied. Thereafter, during coding of each superblock within that tile, MVs from the bank can be used as MV reference candidates. At the end of encoding the superblock, the bank is updated.

[0117] 5G shows an example of motion vector candidate bank processing according to some embodiments. As shown in FIG. 5G, the bank update process may be based on a superblock. For example, after a superblock is coded, the first (e.g., up to 64) candidate motion vectors used by each coding block within the superblock are added to the bank. A pruning process may also be included during the update.

[0118] After the reference MV candidate scan is performed as previously described, if there are open slots in the candidate list, the system may refer to the MV candidate bank (e.g., in the buffer with a matching reference frame type) for additional MV candidates. For example, going backward from the end of the buffer to the beginning, MVs in the bank buffer are added to the candidate list if they are not already in the list.

[0119] Figure 5H shows an exemplary motion vector predictor list construction order according to some embodiments. In the example of Figure 5H, the MVP list is constructed in the following order (e.g., with pruning): (i) adjacent SMVP candidates, (ii) a sorting process for existing candidates, (iii) TMVP candidates, (iv) non-adjacent SMVP candidates, (v) derived candidates, (vi) additional MVP candidates, and (vii) candidates from the reference MV candidate bank. In some embodiments, only one TMVP candidate can be added to the MVP candidate list. For example, when one TMVP candidate is added to the MVP list, the remaining TMVP candidate blocks are skipped. In some embodiments, a reverse horizontal scan order is used to check internal TMVP candidates. An example is shown in Figure 5I, where the scan order for internal TMVP candidates is B3 → B2 → B1 → B0. In some embodiments, external TMVP candidates are not scanned. For example, B4, B5, and B6 in Figure 5D are not scanned.

[0120] In some embodiments, the skip mode motion information fetching of spatially neighboring blocks includes both reference picture indexes and motion vectors. The information fetching may include (i) inserting MVs from adjacent spatial neighboring blocks, (ii) inserting temporal motion vector predictors from reference pictures, (iii) inserting MVs from non-adjacent spatial neighboring blocks, (iv) sorting MV candidates from adjacent spatial neighboring blocks, (v) inserting composite MVs when the existing list size is less than a predetermined threshold (e.g., less than 2 or 3), and / or (vi) inserting MVs from a reference MV bank. The temporal motion vectors and MVs from the reference MV bank may use pre-selected reference pictures. For example, the pruning process may take reference picture indexes into account.

[0121] The slots for MVP candidates are limited to 4 in existing processes, e.g., AV2. Furthermore, regardless of the length of the MVP list, at most one TMVP may be inserted in the MVP list, which may lead to precision loss in the encoding / decoding process.

[0122] 6A is a flow diagram illustrating a method 600 for encoding video, according to some embodiments. Method 600 may be performed in a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 600 is performed by executing instructions stored in memory (e.g., memory 314) of the computing system.

[0123] The system obtains (602) video data including a plurality of blocks, including a first block. In some embodiments, the system obtains (604) a first syntax element, the first syntax element indicating a number N of temporal motion vector predictor (TMVP) candidates for a motion vector predictor (MVP) list. In some embodiments, the number N of TMVP candidates is set to a default value. The system identifies (606) a set of TMVP candidates, the set of TMVP candidates having a size less than or equal to N. In some embodiments, the system identifies (608) a set of spatial MVP candidates. The system generates (610) an MVP list using the set of TMVP candidates and the set of spatial MVP candidates. In some embodiments, the system signals (612) the MVP list in the video bitstream.

[0124] 6B is a flow diagram illustrating a method 650 for decoding video according to some embodiments. Method 650 may be performed in a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 650 is performed by executing instructions stored in memory (e.g., memory 314) of the computing system.

[0125] The system receives video data from a video bitstream (652), including a plurality of blocks, including a first block. In some embodiments, the system obtains (654) a first syntax element from the video bitstream, the first syntax element indicating a number N of temporal motion vector predictor (TMVP) candidates for a motion vector predictor (MVP) list. In some embodiments, the number N of TMVP candidates is set to a default value (e.g., rather than being obtained from the video bitstream). The system identifies (656) a set of TMVP candidates, the set of TMVP candidates having a size less than or equal to N. In some embodiments, the system identifies (658) a set of spatial MVP candidates. The system generates (660) an MVP list using the set of TMVP candidates and the set of spatial MVP candidates. The system reconstructs (662) the first block using the MVP list.

[0126] In some embodiments, up to N TMVP candidates are inserted into the skip mode and non-skip mode MVP candidate lists, where N is a positive integer such as 1 or 2. In some embodiments, N is a different value for the skip mode MVP candidate list and the non-skip mode MVP candidate list. For example, for the non-skip mode MVP candidate list, N may be equal to 1, and for the skip mode MVP candidate list, N may be equal to 2. In some embodiments, different values ​​of N are used for different MVP lists, for example, the first MVP list for normal, the second MVP list for inter-interwedge mode, and the third MVP list for warped reference lists (when TMVP is used).

[0127] In some embodiments, the value of N can be updated during encoding and decoding. For example, more or fewer TMVP candidates (as indicated by the value of N) can be used depending on how often TMVP is applied during encoding and decoding. In some embodiments, the value of N is signaled at a higher level syntax, for example at the sequence level, frame level, slice level, or tile level within the bitstream.

[0128] In some embodiments, the value of N depends on the maximum allowed length of the MVP list. In some embodiments, if the maximum allowed MVP list length is less than or equal to a predetermined threshold (e.g., 4 or 5), N is set to a first value (e.g., 1 or 2). Otherwise, N is set to a second value (e.g., 2 or 3).

[0129] In some embodiments, if N is set to a first value (e.g., 1), the outer TMVP candidate blocks are not scanned / checked. Otherwise, if N is greater than the first value (e.g., equal to 2 or 3), the outer TMVP candidate blocks are checked. In some embodiments, the scan order and / or position of the inner TMVP candidates are different for the skip mode MVP candidate list and the non-skip mode MVP candidate list. For example, a reverse horizontal scan order is used to check inner TMVP candidates for the non-skip mode MVP list, and a raster scan order is used to check inner TMVP candidates for the skip mode MVP list.

[0130] In some embodiments, outer TMVP candidate blocks are not scanned for non-skip mode MVP lists, but outer TMVP candidate blocks are scanned for skip mode MVP lists.

[0131] 6A and 6B depict some logical steps in a particular order, steps that are not order-dependent may be rearranged, and other steps may be combined or separated. Some rearrangements or other groupings not specifically mentioned will be apparent to those skilled in the art, and therefore the order and groupings presented herein are not exhaustive. Furthermore, it should be recognized that the various steps may be implemented in hardware, firmware, software, or any combination thereof.

[0132] We now turn to some exemplary embodiments.

[0133] (A1) In one aspect, some embodiments include a method of video coding (e.g., method 600). In some embodiments, the method is implemented in a computing system (e.g., server system 112) having memory and one or more processors. In some embodiments, the method is implemented in a coding module (e.g., coding module 320). In some embodiments, the method is implemented in an entropy coder (e.g., entropy coder 214). The method includes: (i) receiving video data including a plurality of blocks including a first block; (ii) obtaining a first syntax element, the first syntax element indicating a quantity N of temporal motion vector predictor (TMVP) candidates for a motion vector predictor (MVP) list; (iii) identifying a set of TMVP candidates, the set of TMVP candidates having a size less than or equal to N; and (iv) generating an MVP list using at least the set of TMVP candidates. In some embodiments, the quantity of TMVP candidates is based on a frequency of use of the MVP candidates in one or more previously coded blocks. In some embodiments, the number N of TMVP candidates is set to a default value (e.g., rather than being signaled via the first syntax element).

[0134] (A2) In some embodiments of A1, the number of TMVP candidates is based on whether skip mode is enabled for the first block. In some embodiments, the number of TMVP candidates is based on whether translation mode, wedge mode, and / or warp mode is active.

[0135] (A3) In some embodiments of A1 or A2, identifying a set of TMVP candidates includes scanning a plurality of TMVP candidate blocks in a particular scanning order, the particular scanning order being selected based on whether skip mode is enabled for the first block.

[0136] (A4) In some embodiments of any of A1 to A3, the step of identifying the set of TMVP candidates includes: (i) scanning a first set of TMVP blocks in accordance with a determination that skip mode is enabled; and (ii) scanning a second set of TMVP blocks in accordance with a determination that skip mode is disabled, wherein the first set of TMVP blocks is different from the second set of TMVP blocks.

[0137] (A5) In some embodiments of any of A1 to A4, (i) the MVP list is a first MVP list and corresponds to a first mode, and (ii) the method further includes obtaining an indication of a second quantity M of TMVP candidates for a second MVP list, the second MVP list corresponding to a second mode, and N not equal to M.

[0138] (A6) In some embodiments of any of A1 through A5, the quantity of TMVP candidates is selected based on previously decoded information.

[0139] (A7) In some embodiments of A6, the quantity of TMVP candidates is updated based on the frequency of application of TMVP during encoding.

[0140] (A8) In some embodiments of any of A1 through A7, the first syntax element is part of a high-level syntax in a video bitstream. In some embodiments, the quantity of TMVP candidates is signaled in the video bitstream. In some embodiments, the quantity of TMVP candidates for a particular mode (e.g., skip mode) is signaled relative to the quantity of TMVP candidates for a second mode. For example, a quantity N is signaled for non-skip mode, and a relative quantity (e.g., +1) is signaled for skip mode.

[0141] (A9) In some embodiments of any of A1 through A8, the quantity of TMVP candidates is selected based on the maximum length of the MVP list.

[0142] (A10) In some embodiments of A9, the quantity of TMVP candidates is selected based on whether the maximum length of the MVP list meets one or more criteria.

[0143] (A11) In some embodiments of any of A1 to A10, identifying the set of TMVP candidates includes (i) checking one or more outer TMVP candidate blocks in accordance with a determination that N satisfies one or more criteria, and (ii) not checking one or more outer TMVP candidate blocks in accordance with a determination that N does not satisfy one or more criteria.

[0144] (A12) In some embodiments of any of A1 to A11, the method further comprises signaling the MVP list in the video bitstream.

[0145] (A13) In some embodiments of any of A1 to A12, the method further includes identifying a set of spatial MVP candidates, and the MVP list is generated using at least the set of TMVP candidates and the set of spatial MVP candidates.

[0146] (B1) In another aspect, some embodiments include a method of video decoding (e.g., method 650). In some embodiments, the method is implemented in a computing system (e.g., server system 112) having memory and one or more processors. In some embodiments, the method is implemented in a coding module (e.g., coding module 320). In some embodiments, the method is implemented in a parser (e.g., parser 254). The method includes: (i) receiving video data from a video bitstream, the video data including a plurality of blocks including a first block; (ii) obtaining a first syntax element from the video bitstream, the first syntax element indicating a number N of temporal motion vector predictor (TMVP) candidates for a motion vector predictor (MVP) list, where N is an integer greater than 1; (iii) identifying a set of TMVP candidates, the set of TMVP candidates having a size less than or equal to N; (iv) generating an MVP list using at least the set of TMVP candidates; and (v) reconstructing the first block using the MVP list. For example, N is a positive integer such as 1 or 2. In some embodiments, up to N TMVP candidates are inserted into the MVP candidate list for skip mode and non-skip mode. In some embodiments, the quantity N of TMVP candidates is set to a default value (e.g., rather than being signaled via the first syntax element).

[0147] (B2) In some embodiments of B1, the quantity of TMVP candidates is based on whether skip mode is enabled for the first block. In some embodiments, the method further includes obtaining a second syntax element from the bitstream, the second syntax element indicating whether skip mode is enabled for the first block. In some embodiments, the first syntax element indicates the quantity of TMVP candidates for skip mode. In some embodiments, the first syntax element indicates a second quantity of TMVP candidates for non-skip mode. In some embodiments, the third syntax element indicates the second quantity of TMVP candidates for non-skip mode. For example, for a non-skip mode MVP candidate list, N is equal to 1, and for a skip mode MVP candidate list, N is equal to 2.

[0148] (B3) In some embodiments of B1 or B2, identifying the set of TMVP candidates includes scanning a plurality of TMVP candidate blocks in a particular scan order, the particular scan order being selected based on whether skip mode is enabled for the first block. For example, the scan order and / or position of internal TMVP candidates may be different for skip mode MVP candidate lists and non-skip mode MVP candidate lists. In some embodiments, the scan order and / or position of internal TMVP candidates is different for skip mode MVP candidate lists compared to non-skip mode MVP candidate lists. For example, a reverse horizontal scan order is used to check internal TMVP candidates for non-skip mode MVP lists, and a raster scan order is used to check internal TMVP candidates for skip mode MVP lists.

[0149] (B4) In some embodiments of any of B1 through B3, identifying a set of TMVP candidates includes (i) scanning a first set of TMVP blocks in accordance with a determination that skip mode is enabled, and (ii) scanning a second set of TMVP blocks in accordance with a determination that skip mode is disabled, the first set of TMVP blocks being different from the second set of TMVP blocks. For example, outer TMVP candidate blocks are not checked against the non-skip mode MVP list, and outer TMVP candidate blocks are checked against the skip mode MVP list.

[0150] (B5) In some embodiments of any of B1 through B4, (i) the MVP list is a first MVP list and corresponds to a first mode, and (ii) the method further includes obtaining from the bitstream an indication of a second quantity M of TMVP candidates for a second MVP list, the second MVP list corresponding to a second mode, and N not equal to M. Multiple different quantity values ​​can be used for different MVP lists, such as a normal case MVP list, an inter-interwedge mode MVP list, and a warp reference list (when TMVP is used).

[0151] (B6) In some embodiments of any of B1 through B5, the number of TMVP candidates is selected based on previously decoded information. For example, the value of N can be updated during encoding and decoding.

[0152] (B7) In some embodiments of B6, the number of TMVP candidates is updated based on the frequency of TMVP application during encoding. For example, more or fewer TMVP candidates (as indicated by the value of N) can be used depending on how frequently TMVP is applied during encoding and / or decoding.

[0153] (B8) In some embodiments of any of B1 to B7, the first syntax element is part of a high-level syntax in a video bitstream. For example, the high-level syntax corresponds to a sequence level, a frame level, a slice level, or a tile level. In some embodiments, the high-level syntax is higher than the block level. For example, the high-level syntax may include a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a picture header, a tile header, and / or a CTU header.

[0154] (B9) In some embodiments of any of B1 through B8, the number of TMVP candidates is selected based on a maximum length of the MVP list. In some embodiments, the method further includes obtaining the length of the MVP list from the video bitstream. In some embodiments, the number N of TMVP candidates is based on the length of the MVP list. For example, the value of N depends on the maximum allowable length of the MVP list.

[0155] (B10) In some embodiments of B9, the quantity of TMVP candidates is selected based on whether the maximum length of the MVP list satisfies one or more criteria. For example, if the maximum allowable MVP list length is less than or equal to a threshold S1, then N is set to 1. Otherwise, N is set to 2. In one example, S1 is set to 4.

[0156] (B11) In some embodiments of any of B1 through B10, identifying the set of TMVP candidates includes (i) checking one or more outer TMVP candidate blocks pursuant to a determination that N satisfies one or more criteria, and (ii) not checking one or more outer TMVP candidate blocks pursuant to a determination that N does not satisfy one or more criteria. For example, if N is set to 1, the outer TMVP candidate blocks are not checked. Otherwise, if N is greater than 1, the outer TMVP candidate blocks are checked.

[0157] (B12) In some embodiments of any of A1 to A11, the method further includes identifying a set of spatial MVP candidates, and the MVP list is generated using at least the set of TMVP candidates and the set of spatial MVP candidates.

[0158] (B13) In some embodiments of any of A1 to A12, the bitstream corresponds to video coded according to any of A1 to A13.

[0159] The methods described herein may be used separately or combined in any order. Each of the methods may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In some embodiments, the processing circuitry executes a program stored on a non-transitory computer-readable medium.

[0160] In another aspect, some embodiments include a computing system (e.g., server system 112) including a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing one or more instruction sets configured to be executed by the control circuit, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1-A13 and B1-B13 above).

[0161] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more instruction sets for execution by control circuitry of a computing device, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1-A13 and B1-B13 above).

[0162] Although terms such as "first," "second," etc. may be used herein to describe various elements, it will be understood that these elements are not to be limited by these terms; these terms are used only to distinguish one element from another.

[0163] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the claims. When used in the description of the embodiments and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" will be understood to refer to and encompass any and all possible combinations of one or more of the associated listed items. It will be further understood that as used herein, the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0164] As used herein, the term "if" can be interpreted to mean "when" or "if" or "in response to determining that" or "in accordance with a determination that" or "in response to detecting that" the previously-stated condition is true, depending on the context. Similarly, the phrase "if determining that [the previously-stated condition is true]" or "when [the previously-stated condition is true]" can be interpreted to mean "upon determining that" or "in response to determining that" or "in accordance with a determination that" or "upon detecting that" or "in response to detecting that" the previously-stated condition is true, depending on the context.

[0165] The foregoing description has been described with reference to specific embodiments for purposes of explanation. However, the exemplary discussions above are not intended to be exhaustive or to limit the scope of the claims to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described to best explain the principles of operation and practical applications so as to enable others skilled in the art to understand them.

Claims

1. 1. A processor-implemented method of video decoding, comprising: receiving a video bitstream including a plurality of blocks including a first block; obtaining a first syntax element from the video bitstream, the first syntax element indicating a quantity N of temporal motion vector predictor (TMVP) candidates for a motion vector predictor (MVP) list, where N is an integer greater than 1; identifying a set of TMVP candidates, the set of TMVP candidates having a size less than or equal to N; generating the MVP list using at least the set of TMVP candidates; reconstructing the first block using the MVP list; A method comprising:

2. The number of TMVP candidates is based on whether skip mode is enabled for the first block. The method of claim 1.

3. identifying the set of TMVP candidates includes scanning a plurality of TMVP candidate blocks in a particular scanning order, the particular scanning order being selected based on whether skip mode is enabled for the first block; The method of claim 1.

4. The step of identifying a set of TMVP candidates comprises: scanning a first set of TMVP blocks in accordance with a determination that skip mode is enabled; scanning a second set of TMVP blocks in accordance with the determination that skip mode is disabled, wherein the first set of TMVP blocks is different from the second set of TMVP blocks; The method of claim 1 , comprising:

5. the MVP list is a first MVP list and corresponds to a first mode; The method further includes obtaining, from the video bitstream, an indication of a second number M of TMVP candidates for a second MVP list, the second MVP list corresponding to a second mode, and N not equal to M; The method of claim 1.

6. The quantity of TMVP candidates is selected based on previously decoded information. The method of claim 1.

7. The quantity of TMVP candidates is updated based on the frequency of application of TMVP during encoding. The method of claim 6.

8. the first syntax element is part of a high-level syntax in the video bitstream. The method of claim 1.

9. The quantity of TMVP candidates is selected based on the maximum length of the MVP list. The method of claim 1.

10. The quantity of TMVP candidates is selected based on whether the maximum length of the MVP list satisfies one or more criteria.

10. The method of claim 9.

11. The step of identifying a set of TMVP candidates comprises: checking one or more outer TMVP candidate blocks according to a determination that N satisfies one or more criteria; not checking the one or more outer TMVP candidate blocks in accordance with a determination that N does not satisfy the one or more criteria; The method of claim 1 , comprising:

12. further comprising identifying a set of spatial MVP candidates, wherein the MVP list is generated using the set of TMVP candidates and the set of spatial MVP candidates. The method of claim 1.

13. 1. A computing system comprising: a control circuit; Memory and one or more sets of instructions stored in the memory and configured for execution by the control circuitry; 13. A computing system comprising: a processor configured to execute a programmable logic circuit for executing a programmable logic array (100) of a plurality of processors;

14. A computer program which, when executed by a control circuit, causes the control circuit to carry out a method according to any one of claims 1 to 12.