Systems, methods, apparatuses, and computer program products for improved motion vector prediction by chaining
Patent Information
- Application Number
- PCT/EP2026/054235
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-21
- Filing Date
- 2026-02-17
- Publication Date
- 2026-09-24
Smart Images

Figure EP2026054235_24092026_PF_FP_ABST
Abstract
Description
SYSTEMS, METHODS, APPARATUSES, AND COMPUTER PROGRAM PRODUCTS FOR IMPROVED MOTION VECTOR PREDICTION BY CHAININGField
[0001] This disclosure relates generally to systems, methods, apparatuses, and computer program products for video coding and decoding, and more specifically for improving motion vector prediction by chaining.Background
[0002] Hybrid video codecs, such as ITU-T H.263, H.264 / AVC, and HEVC, are technologies for video encoding and decoding, or so-called versatile video coding (WC), enabling efficient compression and transmission of video data. These codecs employ a combination of techniques to achieve high compression ratios while maintaining video quality.
[0003] ITU-T H.263 is a video compression standard designed for low-bit-rate communication, initially developed for videotelephony and videoconferencing. It uses a blockbased hybrid video coding scheme, incorporating motion-compensated prediction, discrete cosine transform (DCT) for prediction differences, and variable-length coding. H.263 has evolved through several versions, adding features like overlapped block motion compensation and variable block-size motion compensation.
[0004] H.264 / AVC (Advanced Video Coding), also known as MPEG-4 Part 10, is a widely used video compression standard that offers significant improvements over H.263. It utilizes a more advanced block-based hybrid video coding approach, including techniques like intra-frame prediction, inter-frame prediction, and context-adaptive binary arithmetic coding (CABAC). H.264 / AVC provides better compression efficiency and video quality, making it suitable for various applications, from streaming to Blu-ray discs.
[0005] HEVC (High Efficiency Video Coding), also known as H.265, represents the next generation of video compression standards. It builds upon the principles of H.264 / AVC but introduces several enhancements to achieve even higher compression efficiency. HEVC uses larger coding units, improved motion compensation, and more sophisticated entropy coding methods.These advancements allow HEVC to deliver high-quality video at lower bit rates, making it ideal for 4K and higher resolution video.
[0006] There is a current and continued need across all different compression standards and technologies, to improve WC to achieve higher coding efficiency and increased bitrate reduction, and to achieve broad versatility that supports efficient compression of a wide range of video content and application.Brief Summary
[0007] Some embodiments and aspects of the present disclosure comprise systems, methods, apparatuses, and / or computer program products for improving motion vector prediction (MVP) by chaining.
[0008] Methods are provided for improving motion vector prediction (MVP) by chaining. One method includes checking a spatial neighbor’s motion information, if that motion information points to a reference picture not currently being searched, recursively tracing motion vectors of this spatial neighbor until the reference picture being searched is found and updating an MVP list based on an accumulation of recursively traced motion vectors. In another method, when chained motion information candidates are being generated for a merge mode, chaining or tracing is started from a position within the source block where the initial motion is fetched in addition or alternative to starting chaining from current block.
[0009] Methods are provided that include tracing motion information of neighboring blocks to a reference picture until it matches the reference picture being searched, checking a spatial neighbor’s motion information, if that motion information points to a reference picture not currently being searched, recursively tracing motion vectors of this spatial neighbor until the reference picture being searched is found, and updating an MVP list based on an accumulation of recursively traced motion vectors. In another method, while searching for MVP candidates that point to a first reference picture, in an instance in which motion information associated with source block(s) within a same picture that are already parsed or decoded indicate motion that is pointing to a second reference picture, recursively tracing motion vectors of these source block(s) to the first reference picture, creating MVP information pointing to the first reference picture that is recursively traced from a location within source block(s) where the initial motion is fetched, adding the MVP information or the recursively traced motion vectors information to a predictor list, and signaling an index of the predictor list.
[0010] According to an embodiment, a method can be provided or carried out that comprises checking a spatial neighbor’s motion information from either of a first prediction direction or a second prediction direction; in an instance in which the motion information points to a reference picture not currently being searched, recursively tracing the motion of this neighbor until the reference picture being searched is found; and creating a new motion vector candidate from the accumulation of the recursively traced motion vectors to be utilized in an advance motion vector prediction (AMVP) list.
[0011] In some embodiments, the method can further comprise adding the traced motion vector to the AMVP list either when the spatial candidate that is the source of the tracing is being searched or after all spatial candidates are searched but before temporal motion vector predictors are checked. In some embodiments, the method can further comprise adding the traced motion vector to the AMVP list as soon as the spatial neighbor position is checked, and the reference picture is found to be different. In some embodiments, the neighbor positions checked depend on a motion vector storage granularity and / or a block size. In some embodiments, the method can further comprise, in an instance in which the block size is larger than the motion vector storage granularity, checking more than one position from the area above and or to the left of the current block. In some embodiments, the method can further comprise starting chain motion vectors from a position within the source block and or from a position within the current block. In some embodiments, the method can further comprise adding the traced motion vector to the AMVP list in the order of neighbor checking position after all these positions are parsed without tracing and before temporal motion vector predictors are checked. In some embodiments, in merge mode, modifying chained motion vector prediction (CMVP) candidates or generating new candidates by tracing motion vectors from a position within the source block they are fetched. In some embodiments, the method can further comprise, in an instance in which the tracing motion comprises tracing motion in a picture that is not the target picture, continuing tracing using a block vector if an intra-block copy (IBC) block, or intra-template matching prediction (intraTMP) block is encountered. In some embodiments, the method can further comprise terminating tracing and skipping the neighbor motion candidate if in a reference picture that is not the target picture, a block that is inter-coded, IBC coded, or intraTMP coded is not encountered, or alternatively no valid motion information or block vector information is found. In some embodiments, the method can further comprise applying the tracing of motion vectors to other types of motion, such as duringthe list generation of affine motion by generating chained motion vector candidates to predict affine motion. In some embodiments, the method can further comprise signaling, to a decoder device, that some pictures are motion references, indicating that the motion buffer of that picture might be utilized in the chain process. In some embodiments, the method can further comprise adding CMVPs to the merge candidate list that are generated by starting to trace from a position within the source block, for example the position from where the motion is fetched. In some embodiments, the method can further comprise checking the chained motion vector candidate traced from a neighbor coded with IBC or intraTMP to be added to the list at a later stage after all spatial candidates are tried.
[0012] According to other embodiments, a method can be provided or carried out that comprises, while searching for motion vector predictor candidates that point to a first reference picture in a group of pictures, in an instance in which motion information associated with one or more source blocks that are within a same picture and are already parsed or decoded indicates motion that is pointing to a second reference picture that is different from the first reference picture currently being searched, recursively tracing the motion information of the one or more of these source blocks to the first reference picture in the group of pictures, creating a motion vector predictor pointing to the first picture that is recursively traced from a location within a source block where the initial motion is fetched, adding the recursively traced motion information to a predictor list, and signaling an index of the predictor list.
[0013] In some embodiments, the method can further comprise selecting, from the predictor list, a best candidate; and determining, for the best candidate selected from the predictor list, a best motion vector prediction mode. In some embodiments, the signaling the index of the predictor list is performed based on the best mode vector prediction mode being a mode used when recursively tracing the motion information of the one or more source blocks to the first reference picture. In some embodiments, the method can further comprise continuing to recursively trace the motion information of the one or more source blocks to the first reference picture until the first reference picture currently being searched is found, and, once the first reference picture in the group of pictures is found based on the recursive tracing of the motion information of the one or more source blocks to the first reference picture, discontinuing recursive tracing of the motion information of the one or more source blocks to the first reference picture.
[0014] In some embodiments, the motion information of the one or more source blocks comprise motion vectors and or block vectors. In some embodiments, the method can further comprise, in the instance in which the motion information associated with the one or more source blocks to the first reference picture indicate motion that is point to the second reference picture that is different from the first reference picture currently being searched, refraining from discarding the motion information of one or more source blocks. In some embodiments, the motion information comprises motion vector weighting information. In some embodiments, the method can further comprise, in the instance in which motion vector weighting information for the first reference picture is not received, inferring the motion vector weighting information for the first reference picture from motion vector weighting information associated with the one or more source blocks to the first reference picture.
[0015] In some embodiments, the method can further comprise setting a motion vector for the first reference picture to an accumulation of the recursively traced motion vectors of the one or more source blocks. In some embodiments, the setting the motion vector for the first reference picture to the accumulation of the recursively traced motion vectors of the one or more source blocks occurs while the motion vectors are being checked before a temporal candidate. In some embodiments, the setting the motion vector for the first reference picture to the accumulation of the recursively traced motion vectors of the one or more source block occurs after all motion vectors obtained from spatially located blocks are checked as spatial candidates but before temporal motion vector predictor checking. In some embodiments, the temporal motion vector predictor checking comprises fetching the motion information from temporally collocated blocks and adding the motion information to a candidate list if the motion information meets one or more criteria.
[0016] In some embodiments, the spatial candidate comprises an advanced motion vector prediction (AMVP) candidate. In some embodiments, the method can further comprise providing or generating an AMVP list for the block being encoded / decoded. In some embodiments, the method can further comprise adding, to an AMVP list, a traced motion vector for the first reference picture. In some embodiments, the motion information comprises at least motion vectors, reference indices and prediction direction associated with regions of a picture. In some embodiments, the recursively tracing the motion vectors of the one or more neighboring blocks generates one ormore motion vector predictors from the one or more neighboring blocks to the first reference picture currently being checked as a new candidate.
[0017] In some embodiments, the candidate comprises a chained motion vector prediction (CMVP) candidate. In some embodiments, CMVP candidates are used in a merge list. In some embodiments, the method can further comprise starting a motion vector chain for the CMVP candidate from a position within one of the one or more source blocks or within a current block being coded / decoded. In some embodiments, the method can further comprise modifying or generating the CMVP candidate to include motion tracing information indicating a position from which one or more positions of one or more positions in the one or more neighboring blocks were fetched for initial tracing of motion vectors for the CMVP candidate. In some embodiments, the method can further comprise generating a number of CMVP candidates, calculating a cost associated with each candidate, and ranking these candidates based on their cost from lowest cost to highest cost. In some embodiments, a predetermined number or an adaptively determined number of the candidates from the ranked list are utilized in the merge list. In some embodiments, the cost associated with a candidate is calculated by generating the prediction for the template area of the current block and calculating the error between the prediction and the reconstructed samples of the template area based on a distance metric. In some embodiments, the distance metric comprises one or more of: a sum of absolute distance (SAD), a sum of squared error (SSE), or a sum of absolute transformed differences (SATD).
[0018] In some embodiments, the method can further comprise, in an instance in which one or more block vectors is identified during the recursive tracing of the motion information for the one or more neighboring blocks to the first reference picture in the group of pictures, continuing the recursive tracing of the motion information using the one or more block vectors for the one or more neighboring blocks. In some embodiments, the method can further comprise, in an instance in which a particular neighboring block of the one or more neighboring blocks is not an interceded block, an intra-block copy (IBC) coded block or an intra-template matching prediction (TMP) coded block, discontinuing the recursive tracing of the motion vectors from the one or more neighboring blocks and skipping that particular neighboring block for the first reference picture. In some embodiments, the method can further comprise selecting one or more positions within the one or more neighboring blocks based on one or more of: a motion vector storage granularity, or a block size. In some embodiments, the method can further comprise receiving an indication thatone or more pictures in the group of pictures are motion reference (i.e. reconstructed samples of the picture is not needed for this purpose but the motion buffer of that picture might be utilized in chain process) as during tracing, motion information of the pictures that are not in current reference picture list might still be needed. This can be checked by the encoder and signaled when needed in picture or slice level. In some embodiments, at least a portion of one or more of the methods described herein can be carried out or performed by an encoder device or a decoder device.
[0019] According to some embodiments, an apparatus can be provided that comprises means for performing at least one portion or element from one or more of the methods described herein. For example, the apparatus can comprise at least one processor and at least one memory comprising instructions stored thereon that, when executed by the at least one processor, cause the apparatus to perform at least one portion or element from one or more of the methods described herein.
[0020] According to other embodiments, a computer program product can be provided that comprises at least one non-transitory computer-readable storage medium comprising instructions stored therein that, when executed by at least one processor of an apparatus, cause the apparatus to perform at least one portion or element from one or more of the methods described herein.
[0021] The above-noted aspects and features may be implemented in systems, apparatuses, methods, articles, and non-transitory computer-readable media depending on the desired configuration. The subject disclosure may be implemented in and used with a number of different types of devices, such as one or more computing devices, one or more codecs, one or more encoders, one or more user equipment, one or more rendering engines, one or more servers, one or more network access nodes, one or more relay stations, one or more display devices, and / or the like. An example device can comprise at least one processor and at least one memory that stores thereon instructions which, when executed by the at least one processor, cause the device to perform some or all of the elements of the above-described method, according to various embodiments. In other examples, a computer program product, such as a non-transitory computer-readable storage medium can be provided that comprises instructions stored thereon that, when executed by at least one processor of an apparatus, cause the apparatus to perform some or all elements of a method such as that descried above, according to some embodiments. In other examples, an apparatus can be provided that comprises means for carrying out a method - such means can include, e.g., a processor and a memory storing computer-executable instructions orcomputer codes thereon that, when executed by the processor, cause the apparatus to perform some or all of a method such as one of the methods described herein.
[0022] The foregoing and other objectives, features, and advantages of the invention will be more readily understood upon consideration of the following detailed description of the invention taken in conjunction with the accompanying drawings.
[0023] This summary is intended to provide a brief overview of some of the aspects and features according to the subject disclosure. Accordingly, it will be appreciated that the abovedescribed features are merely examples and should not be construed to narrow the scope of the subject disclosure in any way. Other features, aspects, and advantages of the subject disclosure will become apparent from the following detailed description, drawings, and claims.Brief Description of the Drawings
[0024] Having thus described the invention in general terms, reference will now be made to the accompanying drawings. The skilled artisan will understand that the drawings primarily are for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the inventive subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).
[0025] In the accompanying drawings:
[0026] FIG. 1 is a block flow diagram illustrating a process for video encoding and decoding, in accordance with various embodiments of the present disclosure;
[0027] FIG. 2 shows schematically an example of a system for video encoding and decoding, in accordance with various embodiments of the present disclosure;
[0028] FIG. 3 shows schematically an example of a decoder-side device configured for carrying out video decoding, in accordance with various embodiments of the present disclosure;
[0029] FIG. 4 shows schematically an example of an encoder-side device configured for carrying out video encoding, in accordance with various embodiments of the present disclosure;
[0030] FIG. 5 illustrates an example of a current block in a current picture, in accordance with various embodiments of the present disclosure;
[0031] FIG. 6 illustrates an example of a motion information chaining approach, in accordance with various embodiments of the present disclosure;
[0032] FIG. 7 illustrates an example of a motion information chaining approach, in accordance with various embodiments of the present disclosure;
[0033] FIG. 8 illustrates a process for AMVP candidate derivation and generating or modifying an AMVP candidate list, in accordance with various embodiments of the present disclosure;
[0034] FIG. 9 is a block flow diagram illustrating a method for improved motion vector prediction using chaining, in accordance with various embodiments of the present disclosure; and
[0035] FIG. 10 is a block flow diagram illustrating a method for improved motion vector prediction using chaining, in accordance with various embodiments of the present disclosure.Acronyms
[0036] ABT Asymmetric binary tree
[0037] AF INTER Affine motion vector derivation inter mode
[0038] AF MERGE Affine motion vector derivation merge mode
[0039] AFP Adaptive frame packing
[0040] AHG Ad-hoc group
[0041] Al All intra (CTC)
[0042] AIF Adaptive interpolation filter
[0043] ALF Adaptive loop filter
[0044] ALTR Adaptive long-term reference
[0045] AMT Adaptive multiple core transform (also EMT, MTS)
[0046] AMVP Advanced motion vector prediction
[0047] AMVR Adaptive motion vector resolution
[0048] AQS Adaptive quantization step size scaling
[0049] ARC Adaptive resolution change (cp. RPR)
[0050] ATMVP Alternative temporal motion vector prediction
[0051] BARC Block adaptive resolution coding
[0052] BCBR Block-composed background reference
[0053] BCW Bi-prediction with CU based weighting
[0054] BDIP Bi-directional intra prediction
[0055] BCPCM Block differential pulse coded modulation
[0056] BDSNR Bjontegaard Delta PSNR
[0057] BD -rate Bjontegaard Delta rate
[0058] BIO Bi-directional optical flow (BDOF)
[0059] BDOF Bi-directional optical flow (formerly known as BIO)
[0060] BLA Broken link access
[0061] CAB AC Context adaptive binary arithmetic coding
[0062] CAVLC Context adaptive variable length coding
[0063] CB Coding block
[0064] CCIP Cross-component intra prediction
[0065] CCLM Cross-component linear model prediction
[0066] CF Combined filter (intra prediction)
[0067] CfE Call for evidence
[0068] CfP Call for proposals
[0069] CIIP Combined inter and intra prediction
[0070] CMP Cube map projection
[0071] CNNLF Convolution neural network loop filter
[0072] CNNSR CNN for super-resolution
[0073] CPMV Control point motion vector
[0074] CPR Current picture referencing
[0075] CR-CNN CNN for compact-resolution
[0076] CRA Clean random access
[0077] CS Constraint set (in CfP)
[0078] CTC Common testing conditions
[0079] CTU Coding tree unit
[0080] CTX CABAC context
[0081] CU Coding unit
[0082] CVS Coded Video Sequence
[0083] CVSG Coded Video Sequence Group
[0084] DCT Discrete cosine transform
[0085] DST Discrete sine transform
[0086] DPB Decoded picture buffer
[0087] DCT Discrete cosine transform
[0088] DCTIF DCT interpolation filter
[0089] DMVD Decoder side motion vector derivation
[0090] DMVR Decoder-side motion vector refinement
[0091] DPB Decoded picture buffer
[0092] DRA Dynamic range adaptation
[0093] DST Discrete sine transform
[0094] EAC Enhanced angular cubemap
[0095] EDSR Enhanced deep residual network for super-resolution
[0096] EMT Explicit multiple core transforms (also AMT)
[0097] EOTF Electro-optical transfer function
[0098] ERP Equirectangular projection
[0099] FRUC Frame rate up conversion
[0100] GALF Geometry transformation-based adaptive loop filter
[0101] GEO Geometric partitioning
[0102] GRL Givens rotation layer
[0103] HAC Hybrid angular cubemap
[0104] HEVC High efficiency video coding
[0105] HM HEVC test model
[0106] HMVP History based motion vector prediction
[0107] HWT Hadamart- Walsh transform
[0108] HyGT Hypercube- Givens transform
[0109] IBC Intra block copy
[0110] IBDI Internal bit-depth increase
[0111] IDCT Inverse discrete cosine transform
[0112] IDR Instantaneous decoder refresh
[0113] IPR Inter prediction refinement
[0114] ISP Intra subblock partitioning
[0115] JCCR Joint coding of chroma residuals
[0116] JCT Joint collaborative team (of ISO and ITU)
[0117] JCT- VC Joint collaborative team on video coding
[0118] JCT-3V Joint collaborative team on 3D video coding extension development
[0119] JEM Joint exploration model (JVET test model)
[0120] JTC Joint technical committee
[0121] JVET Joint video exploration team
[0122] KLT Karhunen-Loeve transform
[0123] KTA Key Technical Areas (H.264 based exploration software of VCEG)
[0124] LAMVR Locally adaptive motion vector resolution / Low-frequency non-separable transform
[0125] LGT Layered Givens transform
[0126] LIC Local illumination compensation
[0127] LIP Linear intra prediction
[0128] LM Linear model prediction
[0129] LMCS Luma mapping with chroma scaling
[0130] LPS Least probably symbol
[0131] MAP Merge assistant prediction
[0132] MBF Multiple boundary filtering
[0133] MCP Modified cubemap projection
[0134] MDCS Mode-dependent coefficient scanning
[0135] MDIS Mode-dependent intra reference sample smoothing
[0136] MDNSST Mode-dependent non-separable secondary transforms
[0137] Merge Merge Mode (MV prediction)
[0138] MFLM Multiple filter linear model prediction
[0139] MIP Multi-line intra prediction
[0140] MIP Multi-combined intra prediction
[0141] MMLM Multi-Model Cross-component linear model prediction
[0142] MMVD Merge with motion vector difference
[0143] MNLM Multiple neighbor linear model
[0144] MPCR Motion predictor candidate refinement
[0145] MPEG Moving picture experts group
[0146] MPM Most probable mode
[0147] MPS Most probably symbol
[0148] MRIP Multi-reference intra prediction
[0149] MRM Motion refinement mode
[0150] MSB Most Significant Bit
[0151] MSE Mean squared error
[0152] MTS Multiple transform selection
[0153] MTT Multi-type tree
[0154] MV Motion vector
[0155] MVD Motion vector difference
[0156] NAL Network abstraction layer
[0157] MB Macroblock (H.264jAVC)
[0158] NAL Network abstraction
[0159] NALU NAL unit
[0160] NLMLF Non-local mean loop filter
[0161] NLSF Non-local structure-based filter
[0162] NSF Noise suppression filter
[0163] NSST Non-separable secondary transforms
[0164] NUH NAL unit header
[0165] NUT NAL unit type
[0166] OBMC Overlapped block motion compensation
[0167] OETF Opto-electrical transfer function
[0168] PAU Parallel-to-axis uniform cubemap projection format
[0169] PB Prediction block
[0170] PDPC Position dependent intra prediction combination for planar mode
[0171] PMMVD Pattern matched motion vector derivation
[0172] PMVD Pattern matched motion vector derivation
[0173] PMVR Pattern-matched motion vector refinement
[0174] POC Picture order count
[0175] PPS Picture parameter set
[0176] PROF Prediction refinement with optical flow
[0177] PSNR Peak signal to noise ratio
[0178] PU Prediction unit
[0179] QP Quantization parameter
[0180] QT Quad-tree
[0181] QTBTT Quad-tree plus binary tree and ternary tree (also QTBTTT)
[0182] RA Random access (CTC)
[0183] RADL Random access decodable leading picture
[0184] RAP Random access point
[0185] RASL Random access skipped leading picture
[0186] RBSP Raw byte sequence payload
[0187] RC-ALF Reduced-complexity adaptive loop filter
[0188] RD Rate-distortion
[0189] RDO Rate-distortion optimization
[0190] RDOQ Rate-distortion optimized quantization
[0191] RPL Reference picture list
[0192] RPR Reference picture resampling (cp. ARC)
[0193] RPS Reference picture set
[0194] RQT Residual quad-tree
[0195] RSP Rotated sphere projection
[0196] RST Reduced secondary transform
[0197] SAG Sample adaptive offset
[0198] SATD Sum of absolute transformed differences
[0199] SBT Subblock transform
[0200] SDP Signal dependent transform
[0201] SEI Supplemental enhancement information
[0202] SODB String of data bits
[0203] SHVC Scalable high efficiency video coding
[0204] SRCC Scan region-based coefficient coding
[0205] STMVP Spatial-temporal motion vector prediction
[0206] STSA Stepwise temporal sub-layer access
[0207] SUCO Split unit coding order
[0208] SVD Singular value decomposition
[0209] SVT Spatial varying transform
[0210] TB Transform block
[0211] TD-RSP Residual signs prediction in transform domain
[0212] TMM Template matched merge mode
[0213] TMVP Temporal motion vector predictor
[0214] TSA Temporal sub-layer access
[0215] TSB Transform sub-block
[0216] TSR Transform syntax reorder
[0217] TSRC Transform skip residual coding
[0218] TU Transform unit
[0219] UHD Ultra High Definition
[0220] UMVE Ultimate motion vector expression
[0221] UWP Unequal weight planar prediction
[0222] VCEG Video coding experts group
[0223] VCL Video coding layer
[0224] VLC Variable length code
[0225] VPDU Virtual pipeline data units
[0226] VPS Video parameter set
[0227] VUI Video usability information
[0228] WCG Wide colour gamut
[0229] WPP Wavefront parallel processing
[0230] XGA Extended Graphics Array
[0231] XYZ XYZ color space, also color format
[0232] YCbCr Color format with luma and two chroma components
[0233] YUV Color format with luma and two chroma components Detailed Description
[0234] The present disclosure more fully describes various embodiments with reference to the accompanying drawings. It should be understood that some, but not all embodiments are shown and described herein. Indeed, the embodiments may take many different forms, and accordingly this disclosure should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like numbers refer to like elements throughout.
[0235] Hybrid video codecs, for example ITU-T H.263, H.264 / AVC and HEVC, may encode the video information in two phases. At first, pixel values in a certain picture (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). In the first phase, predictive coding may be applied, for example, as so-called sample prediction and / or so-called syntax prediction.
[0236] In the sample prediction, pixel or sample values in a certain picture area or "block" are predicted. These pixel or sample values can be predicted, for example, using one or more of motion compensation or intra prediction mechanisms.
[0237] Motion compensation mechanisms (which may also be referred to as inter prediction, temporal prediction or motion-compensated temporal prediction or motion-compensated prediction or MCP) involve finding and indicating an area in one of the previously encoded video frames that corresponds closely to the block being coded. Inter prediction may reduce temporal redundancy.
[0238] Intra prediction, where pixel or sample values can be predicted by spatial mechanisms, involve finding and indicating a spatial region relationship. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0239] In the syntax prediction, which may also be referred to as parameter prediction, syntax elements and / or syntax element values and / or variables derived from syntax elements are predicted from syntax elements (de)coded earlier and / or variables derived earlier. Non-limiting examples of syntax prediction are provided below.
[0240] In motion vector prediction, motion vectors e.g., for inter and / or inter- view prediction may be coded differentially with respect to a block-specific predicted motion vector. In many video codecs, the predicted motion vectors are created in a predefined way, for example by calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions, sometimes referred to as advanced motion vector prediction (AMVP), is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks intemporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or co-located blocks in temporal reference picture. Differential coding of motion vectors is typically disabled across slice boundaries.
[0241] The block partitioning, e.g., from CTU to CUs and down to PUs, may be predicted.
[0242] In filter parameter prediction, the filtering parameters e.g., for sample adaptive offset may be predicted.
[0243] Prediction approaches using image information from a previously coded image can also be called as inter prediction methods which may also be referred to as temporal prediction and motion compensation. Prediction approaches using image information within the same image can also be called as intra prediction methods.
[0244] Secondly, the prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This may be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size of transmission bitrate).
[0245] In many video codecs, including H.264 / AVC and HEVC, motion information is indicated by motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder) or decoded (at the decoder) and the prediction source block in one of the previously coded or decoded images (or pictures). H.264 / AVC and HEVC, as many other video compression standards, a picture is divided into a mesh of rectangles, for each of which a similar block in one of the reference pictures is indicated for inter prediction. The location of the prediction block is coded as a motion vector that indicates the position of the prediction block relative to the block being coded.
[0246] In video coding, inter prediction utilizes previously coded / decoded frames for prediction. Frames that are used by other frames for prediction are referred to as reference pictures. Each inter coded frame, has an at least one associated reference picture. Usually, there are multiplereference pictures available for prediction, and these are collected in at least one list, which is called as reference picture list, such as reference picture list 0 and reference picture list 1. A reference picture within a reference picture list is identified by its index in the list, which is called as reference index. For a given reference list, e.g., reference list 0, reference index within that list and associated motion vectors are part of the motion information of the block. Additionally, motion information can include prediction direction, motion type, bi-prediction with CU-level weight (BCW) for bi-directionally predicted blocks, local illumination compensation flag, alternative half-pel interpolation filter flag, etc.
[0247] In inter prediction, when a block is predicted using one hypothesis, it is called uniprediction. In uni-prediction, prediction direction can be either from reference picture list 0 or from reference picture list 1. Knowing the prediction direction and reference index, reference picture can be identified as it is the reference picture located in the position that the reference index indicates in the reference picture list that the prediction direction identifies. Alternatively, a block can have bi-prediction by having the prediction as a weighted average of the two uni-predictors. In this case, prediction direction is from both list, so information such as motion vectors and reference indices are found for both lists.
[0248] In many codecs, Motion Vector Prediction operates by generating a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or co-located blocks in temporal reference picture.
[0249] In Versatile Video Codec (WC), there are the following new coding tools.• Intra predictiono 67 intra mode with wide angles mode extensiono Block size and mode dependent 4 tap interpolation filtero Position dependent intra prediction combination (PDPC)o Cross component linear model intra prediction (CCLM)o Multi-reference line intra predictiono Intra sub-partitionso Weighted intra prediction with matrix multiplication• Inter-picture predictiono Block motion copy with spatial, temporal, history-based, and pairwise average merging candidateso Affine motion inter predictiono Sub-block based temporal motion vector predictiono Adaptive motion vector resolutiono 8x8 block-based motion compression for temporal motion predictiono High precision (1 / 16 pel) motion vector storage and motion compensation with 8- tap interpolation filter for luma component and 4-tap interpolation filter for chroma componento Combined intra and inter predictiono Merge with MVD (MMVD)o Symmetrical MVD codingo Bi-directional optical flowo Decoder side motion vector refinemento Bi-prediction with CU-level weight• Transform, quantization and coefficients codingo Multiple primary transform selection with DCT2, DST7 and DCT8o Secondary transform for low frequency zoneo Sub-block transform for inter predicted residualo Dependent quantization with max QP increased from 51 to 63o Transform coefficient coding with sign data hidingo Transform skip residual coding• Entropy Codingo Arithmetic coding engine with adaptive double windows probability update o In loop filtero In-loop reshapingo Deblocking filter with strong longer filtero Sample adaptive offseto Adaptive Loop Filter• Screen content codingo Current picture referencing with reference region restrictiono 360-degree video codingo Horizontal wrap-around motion compensationo High-level syntax and parallel processingo Reference picture management with direct reference picture list signaling o Tile groups with rectangular shape tile groups• Inter prediction in WCMerge list may include the following candidateo Spatial MVP from spatial neighbor CUso Temporal MVP from collocated CUso History-based MVP from a FIFO tableo Pairwise average MVP (using the candidates already in the list)o Zero MVs.
[0250] Merged mode width motion vector difference (MMVD) is to signal MVDs and a resolution index after signaling merge candidate.
[0251] In Symmetric MVD, motion information of list- 1 are derived from motion information of list-0 in bi-prediction case.
[0252] In Affine prediction, several motion vectors are indicated / signaled for different corners of a block, which are used to derive the motion vectors of sub-block. In affine merge, affine motion information of a block is generated based on the normal or affine motion information of the neighboring blocks.
[0253] In Sub-block-based temporal motion vector prediction, motion vectors of sub-blocks of the current block are predicted from a proper subblocks in the reference frame which are indicated by the motion vector of a spatial neighboring block (if available).
[0254] In Adaptive motion vector resolution (AMVR), precision of MVD is signaled for each CU.
[0255] In Bi-prediction with CU-level weight, an index indicated the weight values for weighted average of two prediction block.
[0256] Bi-directional optical flow (BDOF) refines the motion vectors in bi-prediction case. BDOF generates two prediction blocks using the signaled motion vectors. Then a motionrefinement is calculated two minimize the error between two prediction blocks using their gradient values. The final prediction blocks are refined using the motion refinement and gradient values.
[0257] Bi-prediction with CU-level weight (BCW) and weighted prediction (WP)
[0258] In HEVC, the bi-prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In WC, the bi-prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals.Pbi-pred
[0259] Five weights are allowed in the weighted averaging bi-prediction, w E {— 2, 3, 4, 5, 10}. For each bi-predicted CU, the weight w is determined in one of two ways: 1) for a non-merge CU, the weight index is signalled after the motion vector difference; 2) for a merge CU, the weight index is inferred from neighbouring blocks based on the merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (wG {3,4,5}) are used.• At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. For further details readers are referred to the VTM software and document JVET- L0646. When combined with AMVR, unequal weights are only conditionally checked for 1-pel and 4-pel motion vector precisions if the current picture is a low-delay picture. • When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode.• When the two reference pictures in bi-prediction are the same, unequal weights are only conditionally checked.• Unequal weights are not searched when certain conditions are met, depending on the POC distance between current picture and its reference pictures, the coding QP, and the temporal level.
[0260] The BCW weight index is coded using one context coded bin followed by bypass coded bins. The first context coded bin indicates if equal weight is used; and if unequal weight is used, additional bins are signaled using bypass coding to indicate which unequal weight is used.
[0261] Weighted prediction (WP) is a coding tool supported by the H.264 / AVC and HEVC standards to efficiently code video content with fading. Support for WP was also added into the WC standard. WP allows weighting parameters (weight and offset) to be signalled for each reference picture in each of the reference picture lists L0 and LI. Then, during motion compensation, the weight(s) and offset(s) of the corresponding reference picture(s) are applied. WP and BCW are designed for different types of video content. In order to avoid interactions between WP and BCW, which will complicate WC decoder design, if a CU uses WP, then the BCW weight index is not signaled, and w is inferred to be 4 (i.e. equal weight is applied). For a merge CU, the weight index is inferred from neighboring blocks based on the merge candidate index. This can be applied to both normal merge mode and inherited affine merge mode. For constructed affine merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index for a CU using the constructed affine merge mode is simply set equal to the BCW index of the first control point MV
[0262] In WC, CIIP and BCW cannot be jointly applied for a CU. When a CU is coded with CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weight.
[0263] In ECM, there is Chained Motion Vector Prediction (CMVP) candidates for regular and BM merge mode, which are generated from other merge candidates in the list and added to the merge list after history-based merge candidates (HMVP) and additionally before zero candidates. CMVP candidates are derived as the accumulation of the recursively traced MVs and BVs based on the pre-derived MVs of the merge list. CMVP candidates can be derived for each merge index and each reference picture list. The traceable reference pictures are restricted to the reference pictures in the reference picture list. When deriving the additional CMVP candidates, several positions within the current block is checked to find traced MVs or BVs. After all merge candidates are derived, ARMC reordering is performed.
[0264] In ECM, CMVP is utilized in merge mode. In case of AMVP mode, where motion information is signaled relative to a predictor with additional delta allowed, when the predictor list is being calculated, motion information of spatial candidates (fetched from close neighbor blocks) is discarded if the reference picture they are pointing is not the one that is being searched (alternatively, these motion vectors could be scaled to point to the target reference picture). Motion information obtained from spatial neighbors are usually good candidates for the current block andfor this reason they are prioritized in the AMVP / merge list construction, therefore it is beneficial to utilize information from these blocks.
[0265] In some embodiments, a method can comprise the use of chaining in a merge mode. For example, CMVP in the merge mode may use a position within the current block (e.g., center of the block or top left corner, etc.) to chain the motion vectors regardless of the source of the initial motion vector that is being chained. If the initial MV is not a good predictor for current block, the chained MV points to a random position in the process. A chained motion vector traced from a position within a source block would potentially trace the motion of that block.
[0266] Described herein are approaches for improving motion vector prediction using chaining. For example, when generating the AMVP list, in accordance with some embodiments, a method can be performed which includes tracing the motion information of the neighboring blocks to the reference picture until it matches the reference picture that is being searched. For this purpose, when a spatial neighbor’s motion information from either prediction direction is checked and it is found that the motion is pointing to a reference picture that is not currently being searched, instead of skipping this neighbor, the motion of this neighbor can be traced recursively until the reference picture being searched is found and the motion vector is set to the accumulation of the recursively traced motion vectors. This step can be done either when the spatial candidate is being searched or after all the spatial candidates are searched but before temporal motion vector predictors are checked. Additionally, chaining can be applied to the other types of candidates such as to the non-adjacent spatial candidates.
[0267] In some embodiments, a left neighbor of a current block is pointing to a block in a first reference picture, however the block in the first reference picture corresponding to this neighbor when motion vector tracing / motion vector prediction is applied, is found to be pointing to a block in a second reference picture that is different from the first reference picture. Said otherwise, the motion vector in the first reference picture may refer to another motion vector in the second reference picture. As such, an accumulated motion vector can be generated or determined that points or ‘traces’ back to the block in the first reference picture and also points or ‘traces’ back to the block in the second reference picture. This accumulated motion vector may be a more viable motion vector candidate from the left neighbor of the current block for the reference picture. If these three blocks in the ‘chain’ are the same object, the block in the second reference picturemight have been coded with higher quality and might be a more suitable reference than the block in the first reference picture.
[0268] According to some embodiments, a method can be performed that comprises recursively tracing reference motion buffers to reach a reference picture being searched in order to obtain a motion vector predictor from a neighbor position being checked. The neighbor position can be a temporal neighbor or a spatial neighbor.
[0269] In some embodiments, if the traced motion vector is found or generated, it may be added to an AMVP list. This can happen as soon as the neighbor position is checked which would be prioritizing the position being checked
[0270] In some embodiments, traced motions are only added to the list after the motion information of all neighbors are tried to be added to the list, which would be prioritizing the initial motion.
[0271] In some embodiments, traced motions are only added to the list before the motion information of all neighbors are tried to be added to the list, which would be prioritizing the traced motion.
[0272] In some embodiments, neighbor positions checked may be more than as initially described, depending on motion vector storage granularity and block size. For example, when the block size is larger than the motion vector storage granularity, more than one position (such as left and right of midpoint) can be checked from the area above the current block, from which only Bl is checked and similarly from left.
[0273] In another embodiment, chain motion vectors can be started from neighbor position and or from a position within the current block, such as center, top left, top right, bottom right, etc.
[0274] In another embodiment, the traced mv can be added to the AMVP list, in the order of neighbor checking position (AMVP spatial neighbors on above and left Ax and Bx) after all these positions are parsed without tracing and before temporal motion vector predictors are checked (prioritizing untraced over traced).
[0275] Similarly, in case of merge mode, CMVP candidates of the merge mode can be modified, or new candidates can be generated so that instead of tracing motion from the current block, motion vectors are traced from the position they are fetched so that initial step of trace motion vector investigates a position that it had used as the predictor of that position.
[0276] In one perspective, the first step of tracing MV candidate can be seen as a temporal motion vector as the motion is being predicted from a temporal location (a position found in a different picture to fetch information) but instead of looking into a fixed position in another picture (collocated position C in collocated picture), the position to look at is found by tracing the predictor of a neighbor block in another picture.
[0277] When tracing the motion, in a picture that is not the target picture, if a position does not have associated motion information (i.e., not coded with inter prediction) but has an associated block vector (i.e., coded with a block copy method such as Intra Block Copy - IBC or intra template matching -, tracing continues using the block vector recursively.
[0278] When tracing the motion, in a picture that is not the target picture, if a block that is not inter or IBC or intraTMP coded (e.g., intra coded), tracing is terminated, and this neighbor motion candidate for this reference picture is skipped.
[0279] This type of tracing of motion vectors can be applied to other types of motion, for example during the list generation of affine motion by generating chained mv candidates to predict affine motion.
[0280] When this method is enabled, the encoder may need to signal to the decoder that some pictures are motion reference (i.e. reconstructed samples of the picture is not needed for this purpose but the motion buffer of that picture might be utilized in chain process) as during tracing, motion information of the pictures that are not in current reference picture list might still be needed. This can be checked by the encoder and signaled when needed in picture or slice level.
[0281] In another embodiment, chained MVPs for the neighbor positions can be added to the list additionally when a neighbor has a motion vector pointing to the reference picture being searched.
[0282] In another embodiment, when a neighbor block is coded with IBC or intraTMP, the chained motion vector candidate traced from this neighbor may be checked to be added to the list at a later stage such as after all spatial candidates are tried (including non-adjacent neighbors).
[0283] FIG. 1 shows a process 100 in which video compression is applied. The process 100 can comprising generating, recording, rendering, receiving, retrieving, or otherwise providing original video data 101. The original video data 101 can be encoded by a video encoder 102, using, e.g., one or more algorithms. Algorithms, such as Discrete Cosine Transform-based video compression algorithms, e.g., MPEG-2, MPEG-4, H.263, and H.264, can be used by the videoencoder 102 to encode the original video data 101. The output from the video encoder 102 is compressed video data 103. Compressed video data 103 is sent to a network 104 that provides the compressed video data 105 to a video decoder 106. The video decoder 106 decodes the compressed video data 105 to generate decoded video data 107, which is approximately equivalent to the original video data 101.
[0284] The video encoder 102 compresses the original video data 101 in such a way that the compressed video data 103 does not exceed an available bandwidth of the network 104 in order for the video decoder 106 to be able to receive and decode the compressed video data 105. However, communication bandwidth may vary depending on the type of the network 104. For example, the available communication bandwidth of an Ethernet is different from that of a wireless local area network (WLAN). The network 104, which may be e.g., a cellular communication network, may have a very narrow bandwidth. Thus, it can be important to generate compressed video data 103 at various bit-rates from the same original video data 101, such as by using scalable video coding. Scalable video coding is a video compression technique that allows video data to provide scalability. Scalability is the ability to generate video sequences at different resolutions, frame rates, and qualities from the same compressed bitstream. In some embodiments, the video encoder 102 can achieve temporal scalability can be provided using, e.g., Motion Compensation Temporal filtering (MCTF), Unconstrained MCTF, Successive Temporal Approximation and Referencing, and / or the like. In some embodiments, the video encoder 102 can achieve Signal -to-Noise Ratio (SNR) scalability or Signal-to-Noise-plus-Interference Ratio (SNIR) scalability using, e.g., Embedded ZeroTrees Wavelet (EZW), Set Partitioning in Hierarchical Trees (SPIHT), Embedded ZeroBlock Coding (EZBC), Embedded Block Coding with Optimized Truncation (EBCOT), etc. In some embodiments, the video encoder 102 can transmit only a portion of a scene, image, or picture need be transmitted as compressed video data 105 to the video decoder 106, which may improve bit-rate efficiency of video compression / coding.
[0285] In some embodiments, the video encoder 102 can achieve spatial scalability by using, e.g., a wavelet transform algorithm or multi-layer coding. For example, in some embodiments the video encoder 102 can use a multi-layer bitstream to transmit different portions of a scene, image, picture, inlay, overlay, background imagery, and / or the like, as separate layers, overlays, textures, alphas, or the like, to the video decoder 106 via the network 104.
[0286] When the video encoder 102 uses a multi-layer bitstream to transmit different portions of a scene, image, picture, inlay, overlay, background imagery, and / or the like, as separate layers, overlays, textures, alphas, or the like, to the video decoder 106 via the network 104, the video encoder 102 must typically provide metadata before, with, or after transmitting a portion or component of the multi-layer bitstream. For example, metadata may include bitstream information, such as attributes of a frame, overlay, layer, texture, alpha, or the like. In some embodiments, the metadata can be provided in a network abstraction layer (NAL) unit, e.g., in accordance with the H.264 / AVC and HEVC video coding standards, the entire disclosures of which are hereby incorporated herein by reference in their entireties for all purposes.
[0287] FIG. 2 illustrates a system 200 within which at least some of the embodiments described in the present disclosure can be utilized. The system 200 comprises multiple communication devices which can communicate through one or more networks. The system 200 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.
[0288] The system 200 may include both wired and wireless communication devices and / or electronic devices suitable for implementing select, various, or all of the embodiments described herein.
[0289] For example, the system 200, as shown in FIG. 2, is illustrated as comprising a mobile network 210 and a representation of the internet 220. The mobile network 210 can be or comprise, e.g., a fourth generation (4G) network, a Long Term Evolution (LTE), a fifth generation (5G) network, a sixth generation (6G) network, and / or the like. Connectivity to the internet 220 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0290] The example communication devices shown in the system 200 may include, but are not limited to, an electronic device or apparatus, such as mobile device 211, user equipment 212, etc. User equipment 212 can be connected to the internet 220 by way of at least an access point 213, which may be or comprise a gNodeB (gNB), an eNodeB (eNB), base station, access network node,radio access network (RAN) node, and / or the like. The mobile device 211 can be connected to the internet 220 by way of at least a cell tower 214, e.g., via radio signaling 215 with the cell tower 214, short messaging service (SMS) with the cell tower 214, and / or the like.
[0291] Additionally or alternatively, the mobile device 211 and / or user device 212 can be connected to the internet 220 by way of a WiFi access point 216 or the like. Access point 213, cell tower 214, and / or WiFi access point 216 can be configured to communicate directly with the internet 220 or with the internet 220 by way of a network server 219, which can comprise, be comprised in, hosted on, or otherwise functionalized via any suitable network-side device. Such network-side devices can include, but are not limited to, a server, a computing device, a centralized processing unit (CPU), a graphics processing unit (GPU), a processor, processing circuitry, a controller, a network element, a virtualized network function, a mobility management entity (MME), a serving gateway (SGW), a packet data network (PDN) gateway (PGW), a home subscriber server (HSS), a public data network (PDN), an access and mobility management function (AMF), a user plane function (UPF), a data network (DN), an authentication server function (AUSF), a session management function (SMF), a network slice selection function (NSSF), a network exposure function (NEF), a network function repository function (NRF), a policy control function (PCF), a unified data management (UDM) function, an application function (AF), or any other suitable network-side device, element, function, hardware, device, etc.
[0292] In the system 200, electronic devices such as, e.g., 211, 212, 217, etc. may be stationary or mobile when carried by an individual who is moving. For example, the user equipment 212 can be or comprise a head-mounted display, a body-worn display, an immersive gaming system, a smartphone, a laptop (e.g., 217), or the like. In other embodiments, electronic devices in the system 200, such as 211, 212, 217, can be mobile by virtue of being located in, mounted on, coupled to, or otherwise supported by a device configured for transportation, including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle, or any similar suitable mode of transport. In other embodiments, the computing device 218 can also be stationary or can be configured to be used while stationary and while mobile, or to be operationally or configurationally switched from a mobile use mode to a stationary use mode. For example, in embodiments in which the user device 212 is or comprises a head mounted display, the userdevice 212 may be effectively used by a user wearing the user device 212 while the user is stationary or while the user is moving.
[0293] Additionally or alternatively, electronic devices in the system 200 can be stationary. For example, the system 200 can comprise a computing device 218, which can be or comprise a gaming console, a desktop computer, a three-dimensional gaming system, a virtual reality display system, an augmented reality display system, an interactive-display system, an image projection system, and / or the like. In other embodiments, the computing device 218 can also be mobile or can be configured to be used while stationary and while mobile, or to be operationally or configurationally switched from a stationary use mode to a mobile use mode.
[0294] In some embodiments, one or more of the electronic devices in the system 200, e.g., one of 211, 212, 217, 218 may be or comprise a set-top box, a digital TV receiver, a device configured to transmit / receive streaming content or audio / video via a bitstream, etc., but which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware or software or combination of the encoder / decoder implementations, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.
[0295] In some embodiments, certain electronic devices (e.g., 211, 212, 217, 218) in the system 200 may be configured to send and receive calls and messages and communicate with service providers through a wireless connection, such as 215, to the cell tower 214 or the access point 213. The cell tower 214 and / or the access point 213 may be connected to the network server 219 that allows communication between the mobile network 210 and the internet 220. The system 200 may include additional communication devices and communication devices of various types, such as electronic devices 221, 222, and 223, which may be outside of the mobile network 210 but nevertheless connected to the internet 220, e.g., by way of a wired or wireless connection 224.
[0296] The various communication devices (e.g., 211, 212, 217, 218, 221, 222, 223) illustrated in the system 200 of FIG. 2 may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (TCP-IP), short messaging service (SMS), multimedia messagingservice (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11 and any similar wireless communication technology. A communications device involved in implementing various embodiments or aspects of the present disclosure may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0297] Among other transmissions between two or more of the communication devices (e.g., 211, 212, 217, 218, 221, 222, 223) illustrated in the system 200 of FIG. 2, video and / or audio transmissions can be carried out. In order to improve the efficiency and / or effectiveness of resource use, reduce bit-rate, reduce and / or improve signaling, and reduce transmission-side (TX-side) and / or receiver-side (RX-side) computational complexity, data to be transmitted can be compressed (i.e., encoded) at the TX-side and decoded at the RX-side. To carry out such video / audio data compression, a coder / decoder device (i.e., codec), or one or more codecs, can be used. In some embodiments, the video encoder 102 and / or video decoder 106 described above with reference to FIG. 1 can comprise at least one codec.
[0298] Real-time Transport Protocol (RTP) is widely used for real-time transport of timed media such as audio and video. RTP may operate on top of the User Datagram Protocol (UDP), which in turn may operate on top of the Internet Protocol (IP). RTP is specified in Internet Engineering Task Force (IETF) Request for Comments (RFC) 3550, available from www.ietf.org / rfc / rfc3550.txt. In RTP transport, media data is encapsulated into RTP packets. Each media type or media coding format may have a dedicated RTP payload format.
[0299] An RTP session is an association among a group of participants communicating with RTP. It is a group communications channel which can potentially carry a number of RTP streams. An RTP stream is a stream of RTP packets comprising media data.
[0300] Communication systems may include any number of media-aware network elements (MANEs). For example, many multipoint audio-visual conferences operate utilizing a centralized unit called Multipoint Control Unit (MCU). An MCU may implement the functionality of an RTP translator or an RTP mixer. An RTP translator may be a media translator that may modify the media inside the RTP stream. A media translator may for example decode and re-encode the media content (i.e., transcode the media content). An RTP mixer is a middlebox that aggregates multiple RTP streams that are part of a session by generating one or more new RTP streams. An RTP mixer may manipulate the media data. One common application for a mixer is to allow a participant to receive a session with a reduced amount of resources compared to receivingindividual RTP streams from all endpoints. A mixer can be viewed as a device terminating the RTP streams received from other endpoints in the same RTP session. Using the media data carried in the received RTP streams, a mixer generates derived RTP streams that are sent to the receiving endpoints. In another example, a MANE is a selective forward unit (SFU) that selectively forwards incoming RTP packets from one or more senders to one or more receivers.
[0301] According to some embodiments, a video codec can consist of an encoder that transforms the input video into a compressed representation suited for storage / transmission and / or a decoder that can uncompress the compressed video representation back into a viewable form. A video encoder (e.g., 102) and / or a video decoder (e.g., 106) may be combined within a singular device or can be separate from each other, i.e., need not form a codec within a singular device. Typically, a video encoder (e.g., 102) discards some information in the original video data 101 in order to represent the original video data 101 in a more compact form (that is, at a lower bit-rate).
[0302] Hybrid video encoders, for example many encoder implementations of ITU-T H.263 and H.264, often may encode the original video data 101 in two or more phases. According to some embodiments, pixel values in a certain picture area (or “block”) are initially predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner), and thereafter a prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This can be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT), a DCT algorithm, or a variant of the same), quantizing the coefficients, and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the video encoder (e.g., 102) can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate), as reflected for example in the compressed video data 103 and / or the compressed video data 105.
[0303] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, reduces temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients canbe predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0304] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy- coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0305] Referring now to FIG. 3, a video coding system or device is illustrated, according to an example embodiment, as a schematic block diagram of an exemplary apparatus, referred to herein as decoding device 300, which is configured to carry out at least a portion of the video compression / encoding / decoding processes and tasks described herein, e.g., 100.
[0306] The decoding device 300 may for example be configured to function as the video decoder 106. In other embodiments, the decoding device 300 can be, comprise, or be comprised within, e.g., mobile device 211, user equipment 212, computing device 217, or computing device 218 in the wireless network 210. In other embodiments, the decoding device 300 can be, comprise, or be comprised within a heads-up display, a head-mounted display, a gaming console, a user’s computer, and / or the like. However, it will be appreciated that various aspects or embodiments of the present disclosure may be implemented within any electronic device or apparatus which may require encoding and decoding, or encoding or decoding, of video images, multi-layer bitstreams, scalable coded video, bitstreams comprising audio and video, bidirectional bitstreams, conversational content bitstreams, non-conversational content bitstreams, and / or the like.
[0307] The decoding device 300 may comprise a controller 301 in operable communication with a memory 302 and a radio interface 303. The decoding device 300 can be further may comprise a display 308, e.g., in the form of a liquid crystal display (LCD), light emitting diode (LED) display, organic LED (OLED) display, plasma display, Active-Matrix OLED (AMOLED) display, Quantum dot LED (QLED) display, micro-LED display, augmented reality display, virtual reality display, projected image display, any combination thereof, and / or the like. In other embodiments, the display 308 may be any other display technology suitable to display an image and / or video. The decoding device 300 may, optionally, further comprise akeypad 309. In other embodiments, any suitable data or user interface mechanism may be employed. For example, a user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0308] The decoding device 300 may comprise a microphone (not shown) or any suitable audio input which may be a digital or analogue signal input. The decoding device 300 may further comprise an audio output device which, in some embodiments, may be any one of: an earpiece, a speaker, or an analogue audio or digital audio output connection. The decoding device 300 may also comprise a battery (not shown). In other embodiments, the decoding device 300 may be powered by any suitable mobile energy device such as a solar cell, a fuel cell, a clockwork generator, etc. The decoding device 300 may, optionally, further comprise a camera 310 capable of recording or capturing images and / or video. The decoding device 300 may further comprise an infrared port (not shown) for short range line of sight communication to other devices. In other embodiments the decoding device 300 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.
[0309] According to some embodiments, the controller 301 can comprise, e.g., a processor or the like configured for controlling at least some aspects, functionalities, equipment, subcomponents, or subsystems of the decoding device 300. The controller 301 may be connected either directly or indirectly to the memory 302 which, in some embodiments, may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 301. The controller 301 may further be connected to codec circuitry 305 and the codec circuitry 305 can be arranged and dimensioned for, operably programmed for, programmatically capable of, or otherwise suitably configured for carrying out coding and / or decoding of audio and / or video data or assisting in coding and decoding carried out by the controller 301.
[0310] The decoding device 300 may further comprise a card reader (not shown) and / or a smart card (not shown), for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network (e.g., 210).
[0311] The decoding device 300 may comprise radio interface circuitry 303 connected to, or otherwise in operable communication with, the controller 301. The radio interface circuitry 303 can be arranged and dimensioned for, operably programmed for, programmatically capable of, orotherwise suitably configured for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network. The decoding device 300 may further comprise an antenna array 304 connected to the radio interface circuitry 303 for transmitting radio frequency signals generated at the radio interface circuitry 303 to other apparatus(es) and for receiving radio frequency signals from other apparatus(es).
[0312] The antenna array 304 can be configured to receive radio signals comprising or representing the compressed / coded audio / video content from, e.g., an encoder-side device. The antenna array 304 can relay the radio signals to the radio interface 303, which can be configured to convert the radio signals to decodable information, which it then passes along to the codec circuitry 305. The codec circuitry 305 then decodes, e.g., with the aid of, and / or upon receiving instructions from, the controller 301. The codec circuitry 305 can then, once the decodable information is decoded, provide decoded information to the controller 301. The controller 301 can interpret the decoded information to synchronize the audio / video content, and otherwise determine how to reconstitute, build, reconstruct, render, overlay, display, emit, broadcast, and / or present the decoded audio, images, and / or video frames to one or more users, either directly on the decoding device 300 (e.g., via the display 308) or by transmitting / providing the decoded audio, images, and / or video frames to another device (e.g., 211, 212, 217, 218) for display thereon.
[0313] The decoding device 300 may, optionally, further comprise a camera 310 capable of recording or detecting individual frames which are then passed to the codec 305 or the controller 301 for processing. The decoding device 300 may receive the video image data for processing from another device prior to transmission and / or storage. The decoding device 300 may also receive either wirelessly or by a wired connection the image for coding / decoding. The incorporation of the camera 310 into / with the decoding device 300 can be helpful in certain circumstances when, e.g., the image and / or video content is or comprises virtual reality content or augmented reality content in which real world objects and imagery located about the decoding device 300 or other device (e.g., 211, 212, 217, 218) displaying thereon the content may need to record and return images and / or video to assist with rendering subsequent images and / or video frames of the content for the user(s).
[0314] Referring now to FIG. 4, a video coding system or device is illustrated, according to an example embodiment, as a schematic block diagram of an exemplary apparatus, referred to hereinas an encoding device 400, which is configured to carry out at least a portion of the video compression / encoding / decoding processes and tasks described herein, e.g., 100.
[0315] The encoding device 400 may for example be configured to function as the video encoder 102. In other embodiments, the encoding device 400 can be, comprise, or be comprised within a network-side or encoder-side device, e.g., 219. However, it will be appreciated that the encoding device 400 can be, comprise, or be comprised within another device within the mobile network (e.g., 210) in which the decoding device 300 is located. In other embodiments, the encoding device 400 can be, comprise, or be comprised within a device operationally and / or physically located outside the system or network (e.g., 210) in which the decoding device 300 is located. For example, the encoding device 400 can be, comprise, or be comprised within, e.g., 221, 222, 223, or the like. However, it will be appreciated that various aspects or embodiments of the present disclosure may be implemented within any electronic device or apparatus which may require encoding and decoding, or encoding or decoding, of video images, multi-layer bitstreams, scalable coded video, bitstreams comprising audio and video, bidirectional bitstreams, conversational content bitstreams, non-conversational content bitstreams, and / or the like.
[0316] The encoding device 400 may comprise a controller 401 in operable communication with a memory 402 and a radio interface 403. The encoding device 400 can be further may comprise codec circuitry 405 in operable communication with one or both of the controller 401 and / or the radio interface 403. The encoding device 400 can further comprise an antenna array 404 in operable communication with the radio interface 403.
[0317] In some embodiments, the encoding device 400 can be configured to capture audio and / or video content to be encoded and transmitted to a decoder-side device (e.g., 300). In such embodiments, the encoding device 400 can, optionally, comprise a camera 407, a microphone (not shown), and / or the like.
[0318] In other embodiments, the encoding device 400 can be configured to generate audio and / or video content to be encoded and transmitted to a decoder-side device (e.g., 300). In such embodiments, the encoding device 400 can, optionally, comprise a graphics generator 409.
[0319] In other embodiments, encoding device 400 can be configured to request, retrieve, or otherwise receive audio and / or video from one or more external devices, subcomponents, systems, etc. For example, the encoding device 400 can be configured to receive audio from an externalmicrophone (not shown) or an external audio generation device (not shown). In some embodiments, the encoding device 400 can be configured to receive video from an external camera 408. Whether video content is captured by the camera 407 or by the external camera 408, these cameras are capable of recording or capturing images and / or video.
[0320] In some embodiments, the encoding device 400 can be configured to receive generated graphics or other rendered content from an external rendering device or graphics generating device, such as the graphics generator 409.
[0321] According to some embodiments, the controller 401 can be or comprise, e.g., a processor, processing circuitry, or the like, that is configured for controlling at least some aspects, functionalities, equipment, subcomponents, or subsystems of the encoding device 400. The controller 401 may be connected either directly or indirectly to the memory 402 which, in some embodiments, may store both data in the form of image and / or audio data and / or may also store instructions for implementation of the same using the controller 401. The controller 401 may further be connected to the codec circuitry 405 and the codec circuitry 405 can be arranged and dimensioned for, operably programmed for, programmatically capable of, or otherwise suitably configured for carrying out coding and / or decoding of audio and / or video data or assisting in coding and decoding carried out by the controller 401.
[0322] The radio interface circuitry 403 of the encoding device 400 can be configured to be connected to, or otherwise in operable communication with, the controller 401. The radio interface circuitry 403 can be arranged and dimensioned for, operably programmed for, programmatically capable of, or otherwise suitably configured for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network. The encoding device 400 may further comprise the antenna 404 connected to the radio interface circuitry 403 for transmitting radio frequency signals generated at the radio interface circuitry 403 to other apparatus(es) and for receiving radio frequency signals from other apparatus(es).
[0323] The codec circuitry 405 of the encoding device 400 may be configured to receive images and / or video (e.g., as bitstream data) from the controller 401. The codec circuitry 405 can be further configured to encode / compress this image and / or video data, and optionally metadata for decoder-side use in decoding / interpreting the encoded / compressed image and / or video data. The codec circuitry 405 can then provide the encoded / compressed image and / or video data to theradio interface 403, which can convert the encoded / compressed image and / or video data to a form that can be provided / transmitted to a decoder-side device (e.g., 300) via radio signalling using the antenna array 404.
[0324] In order to ensure that the encoded / compressed image and / or video data being provided from, e.g., the encoder device 400 to the decoder device 300, is properly and synchronously encoded and decoded, standard means can be defined for how the encoded / compressed image and / or video data is encoded. Likewise, metadata that is associated with the encoded / compressed image and / or video data can likewise be provided between the encoder-side and the decoder-side (e.g., from the encoder device 400 to the decoder device 300). Standard means can also be defined for what the metadata associated with the encoded / compressed image and / or video data comprises, the form in which the metadata associated with the encoded / compressed image and / or video data is provided, how syntax used in the metadata associated with the encoded / compressed image and / or video data is selected and used, and / or how the metadata associated with the encoded / compressed image and / or video data is encoded / decoded.
[0325] The High Efficiency Video Coding (H.265 / HEVC a.k.a. HE VC) standard was originally developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard was published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Version 2 of the H.265 / HEVC standard included scalable, multiview, fidelity range, three-dimensional, and screen content coding extensions which may be abbreviated SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.
[0326] Versatile Video Coding (VVC) (MPEG-I Part 3), a.k.a. ITU-T H.266, is a video compression standard developed by the Joint Video Experts Team (JVET) of the Moving Picture Experts Group (MPEG), (formally ISO / IEC JTC1 SC29 WG11) and Video Coding Experts Group (VCEG) of the International Telecommunication Union (ITU) to be the successor to HEVC / H.265.
[0327] A specification of the AVI bitstream format and decoding process were developed by the Alliance for Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.
[0328] Some key definitions, bitstream and coding structures, and concepts of some video coding standards and specifications are described in this section for providing background for a video encoder, decoder, encoding method, decoding method, and a bitstream structure, wherein the embodiments may be implemented. It is to be understood that embodiments are not limited to the referenced video coding standards or specifications.
[0329] An elementary unit for the input to an encoder and the output of a decoder, respectively, in many cases is a picture. A picture given as an input to an encoder may also be referred to as a source picture, and a picture decoded by a decoded may be referred to as a decoded picture or a reconstructed picture.
[0330] The source and decoded pictures are each comprised of one or more sample arrays. The sample arrays of a picture may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr; regardless of the actual color representation method in use. The actual color representation method in use can be indicated e.g., in a coded bitstream e.g., using the Video Usability Information (VUI) syntax of HEVC or alike. A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma) or the array or a single sample of the array that compose a picture in monochrome format.
[0331] Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays. Chroma formats comprise monochrome format and non-monochrome formats, and these may be summarized as follows:In monochrome sampling there is only one sample array, which may be nominally considered the luma array.In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array.In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array.In 4:4:4 sampling when no separate color planes are in use, each of the two chroma arrays has the same height and width as the luma array.
[0332] Samples of a sample array have a certain bit depth, such as 8 bits per sample or 10 bits per sample. A bit depth implicitly specifies a value range, which may be referred to as the full range. For example, the full range is from 0 to 255, inclusive, for 8 bits per sample, or from 0 to1023, inclusive, for 10 bits per sample. The source video may use allocate a narrower sample value range than the full range. A specific value range, sometimes referred to as the studio range, has been specified in the ITU-T H.273 standard specifying coding-independent code points for video. A source value range may interchangeably be referred to as a source sample value range, and may be defined as the sample value range of the video that is given as input to a video encoder to be encoded.
[0333] A picture may be defined to be either a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, when the source signal is interlaced.
[0334] A bitstream may be defined as a sequence of bits or a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream.
[0335] A syntax element may be defined as an element of data represented in a bitstream. A syntax structure may be defined as zero or more syntax elements present together in a bitstream in a specified order.
[0336] Syntax structures may be specified, for example, using arithmetic, logical, relational, bit- wise, and assignment operators similar to those available in many programming languages. For example, & may indicate a bit-wise ‘AND’ operation. Furthermore, syntax structures may be specified with reference to mathematical functions.
[0337] Syntax structures and semantics may use the values of variables derived from the values of syntax elements. Naming conventions may be defined for variables. For example, variables may be named by a mixture of lower case and upper case letter and without any underscore characters. Variables starting with an upper case letter may be derived for the decoding of the current syntax structure and all depending syntax structures. Variables starting with an upper case letter may, in some cases, be used in the decoding process for later syntax structures without mentioning the originating syntax structure of the variable. Variables starting with a lower case letter may only be used in relation to the syntax structure or function they have been defined for.
[0338] Video coding specifications may define an elementary unit that for the output an of an encoder and / or for the input to a decoder. For example, such an elementary unit may be an open bitstream unit (OBU), as specified e.g., in AVI, or a Network Abstraction Layer (NAL) unit, as specified e.g., in HEVC or WC.
[0339] In some video codecs, an elementary unit for the output of an encoder and the input of a decoder, respectively, may be a Network Abstraction Layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bytestream format has been specified in some video coding standards for transmission or storage environments that do not provide framing structures. The bytestream format separates NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders run a byte-oriented start code emulation prevention algorithm, which adds an emulation prevention byte to the NAL unit payload if a start code would have occurred otherwise. In order to enable straightforward gateway operation between packet- and stream-oriented systems, start code emulation prevention may always be performed regardless of whether the bytestream format is in use or not. ANAL unit may be defined as a syntax structure containing an indication of the type of data to follow and bytes containing that data in the form of an RBSP interspersed as necessary with emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure containing an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits containing syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0340] A bitstream may be defined to logically include a syntax structure, such as a NAL unit, when the syntax structure is transmitted along the bitstream but may be included in the bitstream according to the bitstream format. A bitstream may be defined to natively comprise a syntax structure, when the bitstream includes the syntax structure.
[0341] In some coding formats or standards, a bitstream may be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.
[0342] In some coding formats, such as AVI, a bitstream may comprise a sequence of open bitstream units (OBUs). An OBU comprises a header and a payload, wherein the header identifies a type of the OBU. Furthermore, the header may comprise a size of the payload in bytes.
[0343] In some coding standards, NAL units include a header and payload. In some coding standards, the NAL unit header indicates the type of the NAL unit. In some coding standards, the NAL unit header indicates a scalability layer identifier (e.g., called nuh_layer_id), which may be used, e.g., for indicating spatial or quality layers, views of a multiview video, or auxiliary layers(such as depth maps or alpha planes). In some coding standards, the NAL unit header includes a temporal sublayer identifier, which may be used for indicating temporal subsets of the bitstream, such as a 30-frames-per-second subset of a 60-frames-per-second bitstream.
[0344] Scalable video coding refers to coding structure where one bitstream can contain multiple representations of the content, for example at different bitrates, resolutions, or frame rates. In these cases, the receiver can extract the desired representation depending on its characteristics (e.g., resolution that matches best the display device). Alternatively, a server or a network element can extract the portions of the bitstream to be transmitted to the receiver depending on e.g., the network characteristics or processing capabilities of the receiver. A scalable bitstream may consist of a “base layer” providing the lowest quality video available and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. To improve coding efficiency for the enhancement layers, the coded representation of that layer typically depends on the lower layers. For example, the motion and mode information of the enhancement layer can be predicted from lower layers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer.
[0345] A scalable video codec for quality scalability (also known as Signal-to-Noise or SNR) and / or spatial scalability may be implemented as follows. For a base layer, a conventional non-scalable video encoder and decoder is used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer for an enhancement layer. In H.265 / HEVC and similar codecs using reference picture list(s) for inter prediction, the base layer decoded pictures may be inserted into a reference picture list(s) for coding / decoding of an enhancement layer picture similarly to the decoded reference pictures of the enhancement layer. Consequently, the encoder may choose a base-layer reference picture as inter prediction reference and indicate its use typically with a reference picture index in the coded bitstream. The decoder decodes from the bitstream, for example from a reference picture index, that a base-layer picture is used as inter prediction reference for the enhancement layer. When a decoded base-layer picture is used as prediction reference for an enhancement layer, it is referred to as an inter-layer reference picture.
[0346] While the previous paragraph described a scalable video codec with two scalability layers with an enhancement layer and a base layer, it needs to be understood that the description can be generalized to any two layers in a scalability hierarchy with more than two layers. In this case, a second enhancement layer may depend on a first enhancement layer in encoding and / ordecoding processes, and the first enhancement layer may therefore be regarded as the base layer for the encoding and / or decoding of the second enhancement layer. Furthermore, it needs to be understood that there may be inter-layer reference pictures from more than one layer in a reference picture buffer or reference picture lists of an enhancement layer, and each of these inter-layer reference pictures may be considered to reside in a base layer or a reference layer for the enhancement layer being encoded and / or decoded. Furthermore, it needs to be understood that other types of inter-layer processing than reference- layer picture upsampling may take place instead or additionally. For example, the bit-depth of the samples of the reference- layer picture may be converted to the bit-depth of the enhancement layer and / or the sample values may undergo a mapping from the color space of the reference layer to the color space of the enhancement layer.
[0347] In addition to quality scalability, there are also other scalability modes, such as spatial scalability and bit-depth scalability. In spatial scalability, enhancement layer pictures are coded at a higher resolution than the base layer pictures. In bit-depth scalability, enhancement layer pictures are coded at higher bit-depth (e.g., 10 or 12 bits) than base layer pictures (e.g., 8 bits).
[0348] A multi-layer bitstream is a bitstream comprising multiple layers, which may be, but are not limited to, base and enhancement layers as discussed above for scalable video coding. A multi-layer bitstream may additionally or alternatively comprise independent layers that do not have inter-layer prediction relationship between each other and may even represent different types of content.
[0349] Layers of a multi-layer bitstream may be identified by a layer identifier or layer ID. In some video coding specifications, such as HEVC and WC, the layer ID is represented by the nuh layer id syntax element. In some video coding specifications, such as AVI, the spatial id syntax element of an extension of an OBU header (obu extension header) may be regarded as a layer ID.
[0350] NAL units can be categorized into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units.
[0351] A non-VCL NAL unit may be for example one of the following types: a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end of sequence (EOS) NAL unit, an end of bitstream (EOB) NAL unit, or a filler data NAL unit. Parameter sets may be needed for the reconstruction of decoded pictures, whereasmany of the other non-VCL NAL units may not be necessary for the reconstruction of decoded sample values.
[0352] In some coding formats, picture unit (PU) may be defined as a set of data units, such as NAL units, which are associated with each other, are consecutive in decoding order, and contain exactly one coded picture. For example, certain non-video-coding data units, such as non-VCL NAL units, may be next to coded video data units in decoding order and the respective picture unit may comprise both these non-video-coding data units and the video coding data units of a coded picture.
[0353] In some coding formats, an access unit (AU) may be defined as a set of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and include at most one coded picture at any scalability layer (e.g., with any specific value of nuh layer id in some coding formats, such as HEVC or VVC). In some coding formats, an access unit comprises one or more complete picture units. In some coding formats, in addition to including the VCL NAL units of a coded picture, an access unit may also include non-VCL NAL units associated with the coded picture. Said specified classification rule may, for example, associate pictures with the same output time or picture order count value into the same access unit.
[0354] In some coding formats, a coded video sequence (CVS) may be defined as a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream.
[0355] In some coding formats, such as AVI, a coded video sequence comprises one or more temporal units. A temporal unit consists of a series of OBUs starting from a temporal delimiter, optional sequence headers, optional metadata OBUs, a sequence of one or more frame headers, each followed by zero or more tile group OBUs as well as optional padding OBUs. A temporal unit may be defined to comprise all the OBUs that are associated with a specific, distinct time instant. A temporal unit may comprise a temporal delimiter OBU, and all the OBUs that follow, up to but not including the next temporal delimiter. A temporal delimiter OBU may be defined as an indication that the following OBUs will have a different presentation / decoding time stamp from one or more of the last frames prior to the temporal delimiter.
[0356] A coded layer video sequence (CLVS) may be defined as a sequence of pictures and associated other data within the same scalable layer (e.g., with the same value of nuh layer id) that is decodable independently of other pictures in the same layer.
[0357] In some video coding formats, such as WC, a subpicture may be defined as a rectangular region of one or more slices within a picture, wherein the one or more slices are complete and a slice is a unit (e.g., a NAL unit) that can be decoded independently of other slices of the same coded picture. Thus, a subpicture consists of one or more slices that collectively cover a rectangular region of a picture. Consequently, each subpicture boundary is also always a slice boundary. The slices of a subpicture may be required to be rectangular slices.
[0358] An independent subpicture (also known as an extractable subpicture) may be defined as a subpicture with subpicture boundaries that are treated as picture boundaries. Additionally, it may be required that an independent subpicture has no loop filtering across the subpicture boundaries.
[0359] Output order may be defined as the order in which the decoded pictures are output by a decoder.
[0360] Some coding formats use a concept of picture order count (POC). A value of POC is derived for each picture and is non-decreasing with increasing picture position in output order. In some coding formats, an increasing value of POC indicates the output order of pictures within a single scalability layer and a single CVS. POC may be used in the decoding process for example for implicit scaling of motion vectors and for reference picture list initialization. Furthermore, POC may be used in the verification of output order conformance.
[0361] Some or all of the elements, steps, or components of the approaches described herein can be carried out by a computing device or an apparatus comprising a processor and memory. Examples of such computing devices and apparatuses are described in more detail below. Referring now to both FIG. 3 and FIG. 4, various aspects related to the functionality of the decoder device 300 and / or the encoder device 400 and components / configurations thereof are described. Aspects and / or embodiments of the present disclosure can be implemented as an apparatus or device, such as described above with regard to one or more embodiments of the decoder device 300 and / or one or more embodiments of the encoder device 400. In other embodiments, various aspects or embodiments of the present disclosure can be implemented as a computer program product thatis executable on a computing device - such as by execution of program codes or computer-readable instructions stored on at least one memory device.
[0362] Various aspects or embodiments of the present disclosure may be implemented in various other ways, such as an article of manufacture. One example of an article of manufacture in the context of the present disclosure is a computer program product that includes one or more software components including, for example, software objects, methods, data structures, program codes, computer- readable instructions, application-specific software, and / or the like. A software component may be coded in any of a variety of programming languages. An illustrative programming language may be a lower-level programming language such as an assembly language associated with a particular hardware architecture and / or operating system platform. A software component comprising assembly language instructions may require conversion into executable machine code by an assembler prior to execution by the hardware architecture and / or platform. Another example programming language may be a higher-level programming language that may be portable across multiple architectures. A software component comprising higher-level programming language instructions may require conversion to an intermediate representation by an interpreter or a compiler prior to execution.
[0363] Other examples of programming languages include, but are not limited to, a macro language, a shell or command language, a job control language, a script language, a database query or search language, and / or a report writing language. In one or more example embodiments, a software component comprising instructions in one of the foregoing examples of programming languages may be executed directly by an operating system or other software component without having to be first transformed into another form. A software component may be stored as a file or other data storage construct. Software components of a similar type or functionally related may be stored together such as, for example, in a particular directory, folder, or library. Software components may be static (e.g., pre-established, or fixed) or dynamic (e.g., created or modified at the time of execution).
[0364] A computer program product may include a non-transitory computer-readable storage medium storing applications, programs, program modules, scripts, source code, program code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like (also referred to herein as executable instructions, instructions for execution, computer program products, program code, and / or similar terms used herein interchangeably).Such non-transitory computer-readable storage media include all computer-readable media (including volatile and non-volatile media).
[0365] In one embodiment, a non-volatile computer-readable storage medium may include a floppy disk, flexible disk, hard disk, solid-state storage (SSS) (e.g., a solid-state drive (SSD), solid state card (SSC), solid state module (SSM), enterprise flash drive, magnetic tape, or any other non-transitory magnetic medium, and / or the like. A non-volatile computer-readable storage medium may also include a punch card, paper tape, optical mark sheet (or any other physical medium with patterns of holes or other optically recognizable indicia), compact disc read only memory (CD-ROM), compact disc-rewritable (CD-RW), digital versatile disc (DVD), Blu-ray disc (BD), any other non-transitory optical medium, and / or the like. Such a non-volatile computer- readable storage medium may also include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory (e.g., Serial, NAND, NOR, and / or the like), multimedia memory cards (MMC), secure digital (SD) memory cards, SmartMedia cards, CompactFlash (CF) cards, Memory Sticks, and / or the like. Further, a non-volatile computer-readable storage medium may also include conductive-bridging random access memory (CBRAM), phase-change random access memory (PRAM), ferroelectric random-access memory (FeRAM), non-volatile random-access memory (NVRAM), magnetoresistive randomaccess memory (MRAM), resistive random-access memory (RRAM), Silicon-Oxide-Nitride-Oxide-Silicon memory (SONOS), floating junction gate random access memory (FJG RAM), Millipede memory, racetrack memory, and / or the like.
[0366] In one embodiment, a volatile computer-readable storage medium may include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), fast page mode dynamic random access memory (FPM DRAM), extended data-out dynamic random access memory (EDO DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), double data rate type two synchronous dynamic random access memory (DDR2 SDRAM), double data rate type three synchronous dynamic random access memory (DDR3 SDRAM), Rambus dynamic random access memory (RDRAM), Twin Transistor RAM (TTRAM), Thyristor RAM (T-RAM), Zero-capacitor (Z-RAM), Rambus in-line memory module (RIMM), dual in-line memory module (DIMM), single in-line memory module (SIMM),video random access memory (VRAM), cache memory (including various levels), flash memory, register memory, and / or the like. It will be appreciated that where embodiments are described to use a computer-readable storage medium, other types of computer-readable storage media may be substituted for or used in addition to the computer-readable storage media described above.
[0367] As should be appreciated, various embodiments of the present disclosure may also be implemented as methods, apparatus, systems, computing devices, computing entities, and / or the like. As such, embodiments of the present disclosure may take the form of an apparatus, system, computing device, computing entity, and / or the like executing instructions stored on a computer-readable storage medium to perform certain steps or operations. Thus, embodiments of the present disclosure may also take the form of an entirely hardware embodiment, an entirely computer program product embodiment, and / or an embodiment that comprises combination of computer program products and hardware performing certain steps or operations.
[0368] In some embodiments, the decoding device 300 and / or the encoding device 400 according to one embodiment of the present disclosure. In general, the terms computing device, computing entity, computer, entity, device, system, and / or similar words used herein interchangeably may refer to, for example, one or more computers, computing entities, desktops, mobile phones, tablets, phablets, notebooks, laptops, distributed systems, kiosks, input terminals, servers or server networks, blades, gateways, switches, processing devices, processing entities, set-top boxes, relays, routers, network access points, base stations, the like, and / or any combination of devices or entities adapted to perform the functions, operations, and / or processes described herein. Such functions, operations, and / or processes may include, for example, transmitting, receiving, operating on, processing, displaying, storing, determining, creating / generating, monitoring, evaluating, comparing, and / or similar terms used herein interchangeably. In one embodiment, these functions, operations, and / or processes can be performed on data, content, information, and / or similar terms used herein interchangeably.
[0369] In some embodiments, the decoding device 300 and / or the encoding device 400 may include or be in communication with one or more processing elements (also referred to as processors, processing circuitry, and / or similar terms used herein interchangeably) that communicate with other elements within the decoding device 300 and / or the encoding device 400 via a bus, for example. As will be understood, the processing element of the decoding device 300 and / or the encoding device 400 may be embodied as one or more complex programmable logicdevices (CPLDs), microprocessors, multi-core processors, coprocessing entities, applicationspecific instruction-set processors (ASIPs), microcontrollers, and / or controllers. Further, the processing element may be embodied as one or more other processing devices or circuitry. The term circuitry may refer to an entirely hardware embodiment or a combination of hardware and computer program products. Thus, the processing element may be embodied as integrated circuits, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), hardware accelerators, other circuitry, and / or the like. As will therefore be understood, the processing element may be configured for a particular use or configured to execute instructions stored in volatile or non-volatile media or otherwise accessible to the processing element 402. As such, whether configured by hardware or computer program products, or by a combination thereof, the processing element may be capable of performing steps or operations according to embodiments of the present disclosure when configured accordingly.
[0370] In one embodiment, the decoding device 300 and / or the encoding device 400 may further include or be in communication with non-volatile media (also referred to as non-volatile storage, memory, memory storage, memory circuitry and / or similar terms used herein interchangeably). In one embodiment, the non-volatile storage or memory may include one or more non-volatile storage or memory media, including but not limited to hard disks, ROM, PROM, EPROM, EEPROM, flash memory, MMCs, SD memory cards, Memory Sticks, CBRAM, PRAM, FeRAM, NVRAM, MRAM, RRAM, SONOS, FJGRAM, Millipede memory, racetrack memory, and / or the like. As will be recognized, the non-volatile storage or memory media may store databases, database instances, database management systems, data, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like. The term database, database instance, database management system, and / or similar terms used herein interchangeably may refer to a collection of records or data that is stored in a computer-readable storage medium using one or more database models, such as a hierarchical database model, network model, relational model, entity-relationship model, object model, document model, semantic model, graph model, and / or the like.
[0371] In one embodiment, the decoding device 300 and / or the encoding device 400 may further include or be in communication with volatile media (also referred to as volatile storage, memory, memory storage, memory circuitry and / or similar terms used herein interchangeably). Inone embodiment, the volatile storage or memory may also include one or more volatile storage or memory media 404, including but not limited to RAM, DRAM, SRAM, FPM DRAM, EDO DRAM, SDRAM, DDR SDRAM, DDR2 SDRAM, DDR3 SDRAM, RDRAM, TTRAM, T-RAM, Z-RAM, RIMM, DIMM, SIMM, VRAM, cache memory, register memory, and / or the like. As will be recognized, the volatile storage or memory media may be used to store at least portions of the databases, database instances, database management systems, data, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like being executed by, for example, the processing element. Thus, the databases, database instances, database management systems, data, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like may be used to control certain aspects of the operation of the decoding device 300 and / or the encoding device 400 with the assistance of the processing element and operating system.
[0372] In some embodiments, the decoding device 300 and / or the encoding device 400 may also include one or more network interfaces, such as a transceiver for communicating with various computing entities, such as by communicating data, content, information, and / or similar terms used herein interchangeably that can be transmitted, received, operated on, processed, displayed, stored, and / or the like. Such communication may be executed using a wired data transmission protocol, such as fiber distributed data interface (FDDI), digital subscriber line (DSL), Ethernet, asynchronous transfer mode (ATM), frame relay, data over cable service interface specification (DOCSIS), or any other wired transmission protocol. Similarly, the decoding device 300 and / or the encoding device 400 may be configured to communicate via wireless external communication networks using any of a variety of protocols, such as general packet radio service (GPRS), Universal Mobile Telecommunications System (UMTS), Code Division Multiple Access 2000 (CDMA2000), CDMA2000 IX (IxRTT), Wideband Code Division Multiple Access (WCDMA), Global System for Mobile Communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), Time Division- Synchronous Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), Evolved Universal Terrestrial Radio Access Network (E-UTRAN), Evolution-Data Optimized (EVDO), High Speed Packet Access (HSPA), High-Speed Downlink Packet Access (HSDPA), IEEE 802.11 (Wi-Fi), Wi-Fi Direct, 802.16 (WiMAX), ultra-wideband (UWB), infrared (IR) protocols, near field communication (NFC) protocols, Wibree,Bluetooth protocols, wireless universal serial bus (USB) protocols, and / or any other wireless protocol.
[0373] Although not shown, the decoding device 300 and / or the encoding device 400 may include or be in communication with one or more input elements, such as a keyboard input, a mouse input, a touch screen / display input, motion input, movement input, audio input, pointing device input, joystick input, keypad input, and / or the like. The decoding device 300 and / or the encoding device 400 may also include or be in communication with one or more output elements (not shown), such as audio output, video output, screen / display output, motion output, movement output, and / or the like.
[0374] The signals provided to and received from the decoding device 300 and / or the encoding device 400 may include signaling information / data in accordance with air interface standards of applicable wireless systems. In this regard, the decoding device 300 and / or the encoding device 400 may be capable of operating with one or more air interface standards, communication protocols, modulation types, and access types. More particularly, the decoding device 300 and / or the encoding device 400 may operate in accordance with any of a number of wireless communication standards and protocols, such as those described above. In a particular embodiment, the decoding device 300 and / or the encoding device 400 may operate in accordance with multiple wireless communication standards and protocols, such as UMTS, CDMA2000, IxRTT, WCDMA, GSM, EDGE, TD-SCDMA, LTE, E-UTRAN, EVDO, HSPA, HSDPA, Wi-Fi, Wi-Fi Direct, WiMAX, UWB, IR, NFC, Bluetooth, USB, and / or the like. Similarly, the decoding device 300 and / or the encoding device 400 may operate in accordance with multiple wired communication standards and protocols, such as those described above, via a network interface.
[0375] Via these communication standards and protocols, the decoding device 300 and / or the encoding device 400 can communicate with various other entities using concepts such as Unstructured Supplementary Service Data (USSD), Short Message Service (SMS), Multimedia Messaging Service (MMS), Dual-Tone Multi-Frequency Signaling (DTMF), and / or Subscriber Identity Module Dialer (SIM dialer). The decoding device 300 and / or the encoding device 400 can also download changes, add-ons, and updates, for instance, to its firmware, software (e.g., including executable instructions, applications, program modules), and operating system.
[0376] According to one embodiment, the decoding device 300 and / or the encoding device 400 may include location determining aspects, devices, modules, functionalities, and / orsimilar words used herein interchangeably. For example, the decoding device 300 and / or the encoding device 400 may include outdoor positioning aspects, such as a location module adapted to acquire, for example, latitude, longitude, altitude, geocode, course, direction, heading, speed, universal time (UTC), date, and / or various other information / data. In one embodiment, the location module can acquire data, sometimes known as ephemeris data, by identifying the number of satellites in view and the relative positions of those satellites (e.g., using global positioning systems (GPS)). The satellites may be a variety of different satellites, including Low Earth Orbit (LEO) satellite systems, Department of Defense (DOD) satellite systems, the European Union Galileo positioning systems, the Chinese Compass navigation systems, Indian Regional Navigational satellite systems, and / or the like. This data can be collected using a variety of coordinate systems, such as the Decimal Degrees (DD); Degrees, Minutes, Seconds (DMS); Universal Transverse Mercator (UTM); Universal Polar Stereographic (UPS) coordinate systems; and / or the like.
[0377] Alternatively, the location information / data can be determined by triangulating a position of the decoding device 300 and / or the encoding device 400 in connection with a variety of other systems, including cellular towers, Wi-Fi access points, and / or the like. Similarly, the decoding device 300 and / or the encoding device 400 may include indoor positioning aspects, such as a location module adapted to acquire, for example, latitude, longitude, altitude, geocode, course, direction, heading, speed, time, date, and / or various other information / data. Some of the indoor systems may use various position or location technologies including RFID tags, indoor beacons or transmitters, Wi-Fi access points, cellular towers, nearby computing devices (e.g., smartphones, laptops) and / or the like. For instance, such technologies may include the iBeacons, Gimbal proximity beacons, Bluetooth Low Energy (BLE) transmitters, NFC transmitters, and / or the like. These indoor positioning aspects can be used in a variety of settings to determine the location of someone or something to within inches or centimeters.
[0378] The decoding device 300 and / or the encoding device 400 may also comprise a user interface (that can include a display coupled to the processing element / controller) and / or a user input interface (coupled to the processing element / controller). For example, the user interface may be a user application, browser, user interface, and / or similar words used herein interchangeably executing on and / or accessible via the decoding device 300 and / or the encoding device 400 to interact with and / or cause display of information / data from the decoding device 300 and / or theencoding device 400, as described herein. The user input interface can comprise any of a number of devices or interfaces allowing the decoding device 300 and / or the encoding device 400 to receive data, such as a keypad (hard or soft), a touch display, voice / speech or motion interfaces, or other input device. In embodiments in which the decoding device 300 and / or the encoding device 400 comprises a keypad, the keypad can include (or cause display of) the conventional numeric (0-9) and related keys (#, *), and other keys used for operating the decoding device 300 and / or the encoding device 400 and may include a full set of alphabetic keys or set of keys that may be activated to provide a full set of alphanumeric keys. In addition to providing input, the user input interface can be used, for example, to activate or deactivate certain functions, such as screen savers and / or sleep modes.
[0379] The decoding device 300 and / or the encoding device 400 can also include volatile storage or memory and / or non-volatile storage or memory, which can be embedded and / or may be removable. For example, the non-volatile memory may be ROM, PROM, EPROM, EEPROM, flash memory, MMCs, SD memory cards, Memory Sticks, CBRAM, PRAM, FeRAM, NVRAM, MRAM, RRAM, SONOS, FJG RAM, Millipede memory, racetrack memory, and / or the like. The volatile memory may be RAM, DRAM, SRAM, FPM DRAM, EDO DRAM, SDRAM, DDR SDRAM, DDR2 SDRAM, DDR3 SDRAM, RDRAM, TTRAM, T-RAM, Z-RAM, RIMM, DIMM, SIMM, VRAM, cache memory, register memory, and / or the like. The volatile and nonvolatile storage or memory can store databases, database instances, database management systems, data, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like to implement the functions of the decoding device 300 and / or the encoding device 400. As indicated, this may include a user application that is resident on the entity or accessible through a browser or other user interface for the decoding device 300 to communicate with the encoding device 400 and / or for the encoding device 400 to communication with the decoder device 300.
[0380] In another embodiment, the decoding device 300 and / or the encoding device 400 may include other components or functionalities. As will be recognized, these architectures and descriptions are provided for exemplary purposes only and are not limiting to the various embodiments.
[0381] Referring again to VVC, although the WC standard inherits the framework of blockbased hybrid coding, similar to HEVC, it adopts several highly adaptive and sophisticated codingtools. In general, WC follows a multi-type tree structure (i.e., quadripartite, binary and / or ternary tree) to split a picture into a variety of block shapes (i.e., square, or non-square). Each block is a basic unit for signaling prediction information. Then, intra prediction and / or inter prediction operates on a block-by-block or subblock-by-subblock basis within the basic unit, followed by transform and quantization processes with switchable bases for residual coding, a chain of in-loop filters (i.e., deblocking, sample adaptive offset, adaptive loop filtering) for subjective quality improvement and syntax coding / parsing for transmission.
[0382] In a video signal, high temporal redundancy exists between sequential pictures. Therefore, inter prediction, targeting at reducing the temporal redundancy, makes a major contribution in the video compression capability and plays a key role in the hybrid video coding scheme. In WC, a lot of novel coding tools are developed to further improve inter prediction.
[0383] In general, those tools can be classified into two major groups, depending on whether the whole block share the same set of motion information, i.e., “whole block-based inter prediction” wherein only one set of motion information is utilized and “subblock-based inter prediction” wherein each sub-block could have its own set of motion information.
[0384] Generally, the whole block-based inter-prediction coding tools include the extended adaptive motion vector prediction (AMVP) mode and block merging which are employed in the HEVC inter prediction scheme, and multiple other coding tools introduced in the WC standardization work.
[0385] In HEVC, inter prediction is represented by two modes:a. the AMVP mode and merge mode, wherein reference picture indices and MVDs are signaled in the former mode but not signaled in the latter one. Skip mode is a special merge mode, in which residuals are inferred to be zero and thus not signaled. b. AMVP mode origins from MV competition wherein one of the best MVPs could be selected according to rate distortion cost. For the AMVP mode in HEVC, motion vector predictions are used to exploit spatio-temporal correlations of MVs among prediction units (PUs). The encoder can select the best MVP from an MVP candidate list and transmit the corresponding index together with the reference picture index and MVD. The MVP candidate list for each reference picture list with up to two candidates is constructed. More specifically, the MVPs from spatial neighboring PUs, left and above to the current PU, an MVP derived from temporal motion vector prediction (TMVP)and zero MVPs are added to fulfill the candidate list in order. An TMVP candidate can be derived by scaling an MV in a collocated picture to a target reference picture.
[0386] With the merge mode in HEVC, motion information of the current PU can be directly inherited from spatial or temporal neighboring blocks. A merge candidate list with five candidates is constructed as demonstrated in FIG. 8. Like in the AMVP mode, the encoder selects the best merge candidate from the candidate list and transmits the corresponding index, but without any reference index or MVDs.
[0387] To derive spatial merge candidates, a maximum of four merge candidates are selected among various candidates. After the candidate is added, the addition of the remaining candidates is subject to a redundancy check which ensures that candidates with same motion information are excluded from the list to improve the coding efficiency.
[0388] To reduce computational complexity, only linked pairs are compared, and a candidate is added to the list only if it passes the redundancy check.
[0389] In HEVC, a CU may be partitioned into PUs, which may bring redundancy with the merge mode.
[0390] Beside spatio-temporal merge candidates, there are two additional types of merge candidates: combined bi-predictive merge candidate and zero motion candidate with (0, 0) motion vector. Combined bi-predictive merge candidates are generated by utilizing spatio-temporal merge candidates for B-slice only.
[0391] The combined bi-predictive candidates are generated by combining a first MV referring to reference list 0 of a first merge candidate, and a second MV referring to reference list 1 of a second merge candidate where the first and second merge candidates are selected from available merge candidates in the merge candidate list according to a pre-defined order. The two MVs will form a new bi-predictive candidate. If the merge candidate list is not fulfilled, zero merge candidates will be appended to the list to fill it up.
[0392] In WC, CUP and BCW cannot be jointly applied for a CU. When a CU is coded with CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weight.
[0393] In ECM, there is Chained Motion Vector Prediction (CMVP) candidates for regular and BM merge mode, which are generated from other merge candidates in the list and added to the merge list after history-based merge candidates (HMVP) and additionally before zero candidates. CMVP candidates are derived as the accumulation of the recursively traced MVs and BVs basedon the pre-derived MVs of the merge list. CMVP candidates can be derived for each merge index and each reference picture list. The traceable reference pictures are restricted to the reference pictures in the reference picture list. When deriving the additional CMVP candidates, only the center position of the current block is checked to find traced MVs or BVs. After TMVP and the related CMVP candidates are derived, ARMC reordering is performed.
[0394] However, in ECM, CMVP is only utilized in merge mode. In case of AMVP mode, where motion information is signaled relative to a predictor with additional delta allowed, when the predictor list is being calculated, motion information of spatial candidates (fetched from close neighbor blocks) is discarded if the reference picture they are pointing is not the one that is being searched (alternatively, there is prior art on these motion vectors being scaled to point to the target reference picture). Motion information obtained from spatial neighbors are usually good candidates for the current block and for this reason they are prioritized in the AMVP / merge list construction, therefore it is beneficial to utilize information from these blocks.
[0395] Another potential area to be improved is that CMVP in the merge mode is using the current block (center, top left, etc.) position to chain the motion vectors regardless of the source of the motion vector that is being chained. If the initial MV is not a good predictor for current block, the chained MV points to a random position in the process. For the current block, it is not yet established if the initial MV is suitable, but for the neighbors, it is already chosen by the encoder as the motion of the block. Therefore a chained motion vector traced also starting from the source position might be beneficial.
[0396] Described herein are approaches for improving motion vector prediction using chaining. For example, when generating the AMVP list, in accordance with some embodiments, a method can be performed which includes tracing the motion information of the neighboring blocks to the reference picture until it matches the reference picture that is being searched. For this purpose, when a spatial neighbor’s motion information from either prediction direction is checked and it is found that the motion is pointing to a reference picture that is not currently being searched, instead of skipping this neighbor, the motion of this neighbor can be traced recursively until the reference picture being searched is found and the motion vector is set to the accumulation of the recursively traced motion vectors. This step can be done either when the spatial candidate is being searched or after all the spatial candidates are searched but before temporal motion vector predictors are checked.
[0397] FIG. 5 is a schematic is provided that illustrates spatial neighbor positions and temporal positions checked around a current block in a current picture.
[0398] FIGs. 6 and 7 are schematics that illustrate motion vector chaining processes, in accordance with some embodiments. As illustrated, a left neighbor of the current block is pointing to reference picture Pl, whereas the block in Pl corresponding to this neighbor when the motion vector is applied, is pointing to P0. Therefore, the accumulated motion vector pointing from the current block to the block in P0 would be a viable new motion vector candidate from the left neighbor of the current block for reference picture P0. If these three blocks are the same object, P0 might have been coded with higher quality and might be a more suitable reference.
[0399] In some embodiments, a method can comprise recursively tracing the reference motion buffers to reach to the reference picture being searched, to obtain a motion vector predictor from the neighbor position being checked.
[0400] In one embodiment, the traced motion vector (if found) is added to the AMVP list, as soon as the spatial neighbor position is checked, and reference picture is found to be different (prioritizing position).
[0401] In one embodiment, neighbor positions checked may be more than those illustrated in FIGs. 5-7, depending on motion vector storage granularity and block size. For example, when the block size is larger than the motion vector storage granularity, more than one position (such as left and right of midpoint) can be checked from the area above the current block. In one example, only Bl is checked and similarly from left. In other examples, other blocks (e.g., B0, H, B2, Al, A0) can be additionally or alternatively checked.
[0402] In another embodiment, chain motion vectors can be started from neighbor position and or from a position within the current block such as center, top left, top right, bottom right, etc.
[0403] In another embodiment, the traced mv can be added to the AMVP list, in the order of neighbor checking position (AMVP spatial neighbors on above and left Ax and Bx) after all these positions are parsed without tracing and before temporal motion vector predictors are checked (prioritizing untraced over traced).
[0404] Similarly, in case of merge mode, CMVP candidates of the merge mode can be modified, or new candidates can be generated so that instead of tracing motion from the current block, motion vectors are traced from the position they are fetched so that initial step of trace motion vector investigates a position that it had used as the predictor of that position.
[0405] In one perspective, the first step of tracing MV candidate can be seen as a temporal motion vector as the motion is being predicted from a temporal location (a position found in a different picture to fetch information) but instead of looking into a fixed position in another picture (collocated position C in collocated picture), the position to look at is found by tracing the predictor of a neighbor block in another picture.
[0406] When tracing the motion, in a picture that is not the target picture, if an IBC or intraTMP coded block is encountered for which a block vector is available, tracing continues using the block vector recursively.
[0407] When tracing the motion, in a picture that is not the target picture, if a block that is not inter or IBC or intraTMP coded (e.g., intra coded), tracing is terminated, and this neighbor motion candidate for this reference picture is skipped.
[0408] This type of tracing of motion vectors can be applied to other types of motion, for example during the list generation of affine motion by generating chained mv candidates to predict affine motion.
[0409] When this method is enabled, the encoder may need to signal to the decoder that some pictures are motion reference (i.e. reconstructed samples of the picture is not needed for this purpose but the motion buffer of that picture might be utilized in chain process) as during tracing, motion information of the pictures that are not in current reference picture list might still be needed. This can be checked by the encoder and signaled when needed in picture or slice level.
[0410] In another embodiment, chained MVPs for the neighbor positions can be added to the list additionally when a neighbor has a motion vector pointing to the reference picture being searched.
[0411] In another embodiment, when a neighbor block is coded with IBC or intraTMP, the chained motion vector candidate traced from this neighbor may be checked to be added to the list at a later stage such as after all spatial candidates are tried (including non-adjacent neighbors).
[0412] AMVP mode origins from MV competition wherein one of the best MVPs could be selected according to rate distortion cost. For the AMVP mode in HEVC, motion vector predictions are used to exploit spatio-temporal correlations ofMVs among prediction units (PUs). The encoder can select the best MVP from an MVP candidate list and transmit the corresponding index together with the reference picture index and MVD. The MVP candidate list for each reference picture list with up to two candidates is constructed, following the flow depicted in FIG. 8. More specifically,the MVPs from spatial neighboring PUs, left and above to the current PU, an MVP derived from temporal motion vector prediction (TMVP) and zero MVPs are added to fulfill the candidate list in order. The left spatial neighboring MVP candidate can be derived from one or more first spatial neighbors when at least one of them is inter predicted, and the spatial neighboring MVP candidate can be derived from one or more second spatial neighbors when at least one of them is inter predicted. A scaled MV, according to the temporal distances between the reference picture associated with the left MVP and the current reference picture, will be output as the left MVP if no MV refers to the target reference picture in the one or more first spatial neighbors. Similarly, a scaled MV will be output as the above MVP if no MV refers to the target reference picture in the one or more second spatial neighbors. The spatial neighboring MVP candidate can be discarded if it is identical to the left neighboring MVP candidate. The TMVP candidate is derived by scaling an MV in the collocated picture to the target reference picture. If no MV is available at the first checked location, a second location is checked.
[0413] With the merge mode in HEVC, motion information of the current PU can be directly inherited from spatial or temporal neighboring blocks. A merge candidate list with candidates (e.g., five candidates) can be constructed. Like in the AMVP mode, the encoder selects the best merge candidate from the candidate list and transmits the corresponding index, but without any reference index or MVDs.
[0414] Referring now to FIG. 9, a method 500 according to a particular embodiment of the present disclosure is illustrated. The method 500 can comprise: checking, from a plurality of candidate block positions, for motion vector information associated with one or more motion vectors for respective candidate block positions of the plurality of candidate block positions, at 502. The method 500 can further comprise, in an instance in which the motion vector information associated with one or more candidate block positions from among the plurality of candidate block positions points to a reference picture different than a target reference picture for which motion vector predictor candidates are being derived, recursively tracing, starting from one or more block positions in a current picture, along a plurality of different paths through a plurality of checked blocks over one or more other pictures, the motion vector information until the target reference picture is found, at 504. The method 500 can further comprise creating a new motion vector predictor candidate, from an accumulation of the one or more motion vectors along the plurality of different paths through the plurality of checked blocks over the one or more otherpictures, at 506. The method 500 can, optionally, further comprise adding the new motion vector predictor candidate to a predictor list, at 508. The method 500 can, optionally, further comprise signaling an index of the predictor list, at 510.
[0415] In some embodiments, the candidate is a spatial candidate. In some embodiments, the recursively traced motion vectors are operable for use in generating or modifying an advance motion vector prediction (AMVP) list. In some embodiments, the method 500 can further comprise adding the traced motion vector to the AMVP list either when the spatial candidate that is the source of the tracing is being searched or after all spatial candidates are searched but before temporal motion vector predictors are checked (not shown). In some embodiments, the method 500 can further comprise adding the traced motion vector to the AMVP list as soon as the spatial neighbor position is checked, and the reference picture is found to be different (not shown). In some embodiments, the neighbor positions checked depend on a motion vector storage granularity and / or a block size. In some embodiments, the method 500 can further comprise, in an instance in which the block size is larger than the motion vector storage granularity, checking more than one position from the area above and or to the left of the current block (not shown). In some embodiments, the method 500 can further comprise starting chain motion vectors from a position within the source block and or from a position within the current block (not shown). In some embodiments, the method 500 can further comprise adding the traced motion vector to the AMVP list in the order of neighbor checking position after all these positions are parsed without tracing and before temporal motion vector predictors are checked (not shown). In some embodiments, in merge mode, modifying chained motion vector prediction (CMVP) candidates or generating new candidates by tracing motion vectors from a position within the source block they are fetched. In some embodiments, the method 500 can further comprise, in an instance in which the tracing motion comprises tracing motion in a picture that is not the target picture, continuing tracing using a block vector if an intra-block copy (IBC) block, or intra-template matching prediction (intraTMP) block is encountered (not shown). In some embodiments, the method 500 can further comprise terminating tracing and skipping the neighbor motion candidate if in a reference picture that is not the target picture, a block that is inter-coded, IBC coded, or intraTMP coded is not encountered, or alternatively no valid motion information or block vector information is found (not shown). In some embodiments, the method 500 can further comprise applying the tracing of motion vectors to other types of motion, such as during the list generation of affine motion bygenerating chained motion vector candidates to predict affine motion (not shown). In some embodiments, the method 500 can further comprise signaling, to a decoder device, that some pictures are motion references, indicating that the motion buffer of that picture might be utilized in the chain process (not shown). In some embodiments, the method 500 can further comprise adding CMVPs to the merge candidate list that are generated by starting to trace from a position within the source block, for example the position from where the motion is fetched (not shown). In some embodiments, the method 500 can further comprise checking the chained motion vector candidate traced from a neighbor coded with IBC or intraTMP to be added to the list at a later stage after all spatial candidates are tried (not shown).
[0416] Some or all of the method 500 can be carried out using a system (e.g., 100), by an apparatus, or by a computing device, such as 102, 104, 106, 219, 211, 212, 217, 218, 300, 400, etc. For example, an apparatus can comprise at least one processor and at least one memory storing instructions thereon that, when executed by the at least one processor, cause the apparatus to perform some or all of the method 500. Additionally, a computer program product can be provided that comprises a non-transitory computer-readable storage medium comprising instructions (e.g., one or more programs or applications, program codes, computer- executable program codes, or the like) thereon that, when executed by at least one processor, cause a machine or apparatus to perform some or all of the method 500.
[0417] Referring now to FIG. 10, a method 600 according to a particular embodiment of the present disclosure is illustrated. The method 600 can comprise: while deriving a motion vector predictor candidate list for a reference picture, checking motion information from one or more block positions in a current picture, at 602. The method 600 can further comprise, in an instance in which the motion information of the one or more block positions points to a reference picture that is not a target reference picture for which the motion vector predictor candidate list is being derived, recursively tracing the motion information, starting from the one or more block positions, along a plurality of motion vector paths through a plurality of block positions within a plurality of pictures until the target reference picture is found, at 604. The method 600 can further comprise creating a new motion vector prediction candidate from an accumulation of the motion information recursively traced along the plurality of motion vector paths through the plurality of block positions within the plurality of pictures, at 606. The method 600 can, optionally, further compriseadding the new motion vector predictor candidate to a predictor list, at 608. The method 600 can, optionally, further comprise signaling an index of the predictor list, at 610.
[0418] In some embodiments, the method 600 can further comprise selecting, from the predictor list, a best candidate; and determining, for the best candidate selected from the predictor list, a best motion vector prediction mode (not shown). In some embodiments, the signaling the index of the predictor list is performed based on the best mode vector prediction mode being a mode used when recursively tracing the motion information of the one or more source blocks to the first reference picture. In some embodiments, the method 600 can further comprise continuing to recursively trace the motion information of the one or more source blocks to the first reference picture until the first reference picture currently being searched is found, and, once the first reference picture in the group of pictures is found based on the recursive tracing of the motion information of the one or more source blocks to the first reference picture, discontinuing recursive tracing of the motion information of the one or more source blocks to the first reference picture (not shown).
[0419] In some embodiments, the motion information of the one or more source blocks comprise motion vectors and or block vectors. In some embodiments, the method 600 can further comprise, in the instance in which the motion information associated with the one or more source blocks to the first reference picture indicate motion that is point to the second reference picture that is different from the first reference picture currently being searched, refraining from discarding the motion information of one or more source blocks (not shown). In some embodiments, the motion information comprises motion vector weighting information. In some embodiments, the method 600 can further comprise, in the instance in which motion vector weighting information for the first reference picture is not received, inferring the motion vector weighting information for the first reference picture from motion vector weighting information associated with the one or more source blocks to the first reference picture (not shown).
[0420] In some embodiments, the method 600 can further comprise setting a motion vector for the first reference picture to an accumulation of the recursively traced motion vectors of the one or more source blocks (not shown). In some embodiments, the setting the motion vector for the first reference picture to the accumulation of the recursively traced motion vectors of the one or more source blocks occurs while the motion vectors are being checked before a temporal candidate. In some embodiments, the setting the motion vector for the first reference picture to theaccumulation of the recursively traced motion vectors of the one or more source block occurs after all motion vectors obtained from spatially located blocks are checked as spatial candidates but before temporal motion vector predictor checking. In some embodiments, the temporal motion vector predictor checking comprises fetching the motion information from temporally collocated blocks and adding the motion information to a candidate list if the motion information meets one or more criteria.
[0421] In some embodiments, the spatial candidate comprises an advanced motion vector prediction (AMVP) candidate. In some embodiments, the method 600 can further comprise providing or generating an AMVP list for the block being encoded / decoded (not shown). In some embodiments, the method 600 can further comprise adding, to an AMVP list, a traced motion vector for the first reference picture (not shown). In some embodiments, the motion information comprises at least motion vectors, reference indices and prediction direction associated with regions of a picture. In some embodiments, the recursively tracing the motion vectors of the one or more neighboring blocks generates one or more motion vector predictors from the one or more neighboring blocks to the first reference picture currently being checked as a new candidate.
[0422] In some embodiments, the candidate comprises a chained motion vector prediction (CMVP) candidate. In some embodiments, CMVP candidates are used in a merge list. In some embodiments, the method 600 can further comprise starting a motion vector chain for the CMVP candidate from a position within one of the one or more source blocks or within a current block being coded / decoded (not shown). In some embodiments, the method 600 can further comprise modifying or generating the CMVP candidate to include motion tracing information indicating a position from which one or more positions of one or more positions in the one or more neighboring blocks were fetched for initial tracing of motion vectors for the CMVP candidate (not shown). In some embodiments, the method 600 can further comprise generating a number of CMVP candidates, calculating a cost associated with each candidate, and ranking these candidates based on their cost from lowest cost to highest cost (not shown). In some embodiments, a predetermined number or an adaptively determined number of the candidates from the ranked list are utilized in the merge list. In some embodiments, the cost associated with a candidate is calculated by generating the prediction for the template area of the current block and calculating the error between the prediction and the reconstructed samples of the template area based on a distance metric. In some embodiments, the distance metric comprises one or more of: a sum ofabsolute distance (SAD), a sum of squared error (SSE), or a sum of absolute transformed differences (SATD).
[0423] In some embodiments, the method 600 can further comprise, in an instance in which one or more block vectors is identified during the recursive tracing of the motion information for the one or more neighboring blocks to the first reference picture in the group of pictures, continuing the recursive tracing of the motion information using the one or more block vectors for the one or more neighboring blocks (not shown). In some embodiments, the method 600 can further comprise, in an instance in which a particular neighboring block of the one or more neighboring blocks is not an inter-coded block, an intra-block copy (IBC) coded block or an intra-template matching prediction (TMP) coded block, discontinuing the recursive tracing of the motion vectors from the one or more neighboring blocks and skipping that particular neighboring block for the first reference picture. In some embodiments, the method 600 can further comprise selecting one or more positions within the one or more neighboring blocks based on one or more of: a motion vector storage granularity, or a block size (not shown). In some embodiments, the method 600 can further comprise receiving an indication that one or more pictures in the group of pictures are motion reference (i.e. reconstructed samples of the picture is not needed for this purpose but the motion buffer of that picture might be utilized in chain process) as during tracing, motion information of the pictures that are not in current reference picture list might still be needed (not shown). This can be checked by the encoder and signaled when needed in picture or slice level. In some embodiments, at least a portion of one or more of the methods described herein can be carried out or performed by an encoder device or a decoder device.
[0424] Some or all of the method 600 can be carried out using a system (e.g., 100), by an apparatus, or by a computing device, such as 102, 104, 106, 219, 211, 212, 217, 218, 300, 400, etc. For example, an apparatus can comprise at least one processor and at least one memory storing instructions thereon that, when executed by the at least one processor, cause the apparatus to perform some or all of the method 600. Additionally, a computer program product can be provided that comprises a non-transitory computer-readable storage medium comprising instructions (e.g., one or more programs or applications, program codes, computer- executable program codes, or the like) thereon that, when executed by at least one processor, cause a machine or apparatus to perform some or all of the method 600.
[0425] Where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.
[0426] The above-noted aspects and features may be implemented in systems, apparatuses, methods, articles, and non-transitory computer-readable media depending on the desired configuration. The subject disclosure may be implemented in and used with a number of different types of devices, such as one or more computing devices, one or more codecs, one or more encoders, one or more user equipment, one or more rendering engines, one or more servers, one or more network access nodes, one or more relay stations, one or more display devices, and / or the like.
[0427] An example device can comprise at least one processor and at least one memory that stores thereon instructions which, when executed by the at least one processor, cause the device to perform some or all of the elements of the above-described method, according to various embodiments. In other examples, a computer program product, such as a non-transitory computer-readable storage medium can be provided that comprises instructions stored thereon that, when executed by at least one processor of an apparatus, cause the apparatus to perform some or all elements of a method such as that descried above, according to some embodiments. In other examples, an apparatus can be provided that comprises means for carrying out a method - such means can include, e.g., a processor and a memory storing computer-executable instructions or computer codes thereon that, when executed by the processor, cause the apparatus to perform some or all of a method such as one of the methods described herein.
[0428] As used herein, the terms “instructions,” “file,” “designs,” “data,” “content,” “information,” and similar terms may be used interchangeably, according to some example embodiments of the present disclosure, to refer to data capable of being transmitted, received, operated on, displayed, and / or stored. Thus, use of any such terms should not be taken to limit the spirit and scope of the disclosure. Further, where a computing device is described herein to receive data from another computing device, it will be appreciated that the data may be received directly from the other computing device or may be received indirectly via one or more computing devices,such as, for example, one or more servers, relays, routers, network access points, base stations, and / or the like.
[0429] As used herein, the term “circuitry” refers to all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) to combinations of circuits and computer program product(s) comprising software (and / or firmware instructions stored on one or more computer readable memories), such as (as applicable): (i) to a combination of processor(s) or (ii) to portions of processor(s) / software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions described herein); and (c) to circuits, such as, for example, a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even if the software or firmware is not physically present. This definition of “circuitry” applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term “circuitry” would also cover an implementation of merely a processor (or multiple processors) or portion of a processor and its (or their) accompanying software and / or firmware. The term “circuitry” would also cover, for example and if applicable to the particular claim element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, other network device, and / or other computing device.
[0430] As used herein, the term “computing device” refers to a specialized, centralized device, network, or system, comprising at least a processor and a memory device including computer program code.
[0431] As used herein, the terms “about,” “substantially,” and “approximately” generally mean plus or minus 10% of the value stated, e.g., about 250 pm would include 225 pm to 275 pm, about 1,000 pm would include 900 pm to 1,100 pm. Any provided value, whether or not it is modified by terms such as “about,” “substantially,” or “approximately,” all refer to and hereby disclose associated values or ranges of values thereabout, as described above.
[0432] The described and illustrated embodiments are examples. Although the specification may refer to “an,” “one,” or “some” embodiment(s) in several locations of the text, this does not necessarily mean that each reference is made to the same embodiment(s), or that a particular feature only applies to a single embodiment. Single features of different embodiments may also be combined to provide other embodiments. Further, when a particular feature, structure, orcharacteristic is described in connection of an embodiment, it is within the knowledge of one skilled in the art to apply such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described. It shall be understood that although the terms “first,” “second” and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.
[0433] For the purposes of the present disclosure, the phrases “at least one of A or B,” “at least one of A and B,” and “A and / or B” means (A), (B), or (A and B). For the purposes of the present disclosure, the phrase “A, B, and / or C” means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).
[0434] As used herein, "plurality" means two or more. As used herein, a "set" of items may include one or more of such items. As used herein, whether in the subject disclosure or the claims, the terms "comprising", "including", "carrying", "having", "containing", "involving", and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of' and "consisting essentially of', respectively, are closed or semiclosed transitional phrases with respect to claims. Use of ordinal terms such as "first", "second", "third", etc., in the claims or the subject disclosure to modify an element does not by itself connote any priority, precedence, or order of one element over another or the temporal order in which acts of a method are performed, but are used merely as labels to distinguish one element having a certain name from another element having a same name (but for use of the ordinal term) to distinguish the elements. As used herein, "and / or" and "at least one of' means that the listed items are alternatives, but the alternatives also include any combination of the listed items.
[0435] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, the combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein. It should be appreciated that terminology explicitly employed herein that also may appear in any disclosure incorporated by reference should be accorded a meaning consistent with the particular concepts disclosed herein.
[0436] In some embodiments, one or more of the operations, steps, elements, or processes described herein may be modified or further amplified as described below. Moreover, in someembodiments, additional optional operations may also be included. It should be appreciated that each of the modifications, optional additions, and / or amplifications described herein may be included with the operations previously described herein, either alone or in combination, with any others from among the features described herein.
[0437] The provided method description, illustrations, and process flow diagrams are provided merely as illustrative examples and are not intended to require or imply that the steps of the various embodiments must each or all be performed and / or should be performed in the order presented or described. As will be appreciated by one of skill in the art, the order of steps in some or all of the embodiments described may be performed in any order. Words such as “thereafter,” “then,” “next,” etc. are not intended to limit the order of the steps; these words are simply used to guide the reader through the description of the methods. Further, any reference to claim elements in the singular, for example, using the articles “a,” “an,” or “the” is not to be construed as limiting the element to the singular.
[0438] Many modifications and other embodiments, in addition to those described in the present disclosure, will come to mind to one skilled in the art to which the present disclosure pertains and having the benefit of teachings presented in the foregoing descriptions and the associated drawings. Although the figures only show certain components of the apparatus and systems described herein, it is understood that various other components may be used in conjunction with the system. Therefore, it is to be understood that the scope of the inventions disclosed herein are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, the steps in the method described above may not necessarily occur in the order depicted in the accompanying diagrams, and in some cases one or more of the steps depicted may occur substantially simultaneously, or additional steps may be involved. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation. Specific equipment, apparatuses, systems, computing devices, communications equipment, and / or components described in the examples are for illustration only and not for purposes of limitation. For instance, any and all electronic devices, user equipment, mobile devices, terminal devices, network nodes, access nodes, radio access network nodes, cell towers, base stations, network functions, network elements, servers, structures, and / or the like, having any form factor, scale, dimensions, aesthetic attributes, internal structures or circuitry, external portsor antennas, and / or functional or mechanical properties, which are formed according to any of the disclosed methods, approaches, processes, or variations thereof, using any devices, equipment, apparatuses, systems, or variations thereof, or variations thereof, are all contemplated and covered by the present disclosure. None of the examples provided are intended to, nor should they, limit in any way the scope of the present disclosure.
[0439] Every document cited or referenced herein, including any cross referenced or related patent or application is hereby incorporated herein by reference in its entirety unless expressly excluded or otherwise limited. The citation of any document and / or the mention of methods or apparatuses as being conventional, typical, usual, or the like is not, and should not be taken as an acknowledgement or any form of suggestion that the reference or mentioned method / apparatus is prior art with respect to any invention disclosed or claimed herein or that it alone, or in any combination with any other reference or references, teaches, suggests or discloses any such invention or forms part of the common general knowledge in any country in the world. Further, to the extent that any meaning or definition of a term in this document conflicts with any meaning or definition of the same term in a document incorporated by reference, the meaning or definition assigned to that term in this document shall govern.
[0440] The various portions of the present disclosure, such as the Background, Summary, Brief Description of the Drawings, and Abstract sections, are provided to comply with requirements of the MPEP and are not to be considered an admission of prior art or a suggestion that any portion or part of the disclosure constitutes common general knowledge in any country in the world. The present disclosure is provided as a discussion of the inventor’s own work and improvements based on the inventor’s own work. See, e.g., Riverwood Int ’I Corp. v. R.A. Jones & Co., 324 F.3d 1346, 1354 (Fed. Cir. 2003).
Claims
1. Claims1. An apparatus comprising:at least one processor; andat least one memory comprising instructions stored therein that, when executed by the at least one processor, cause the apparatus to perform at least:checking, from a plurality of candidate block positions, for motion vector information associated with one or more motion vectors for respective candidate block positions of the plurality of candidate block positions;in an instance in which the motion vector information associated with one or more candidate block positions from among the plurality of candidate block positions points to a reference picture different than a target reference picture for which motion vector predictor candidates are being derived, recursively tracing, starting from one or more block positions in a current picture, along a plurality of different paths through a plurality of checked blocks over one or more other pictures, the motion vector information until the target reference picture is found; andcreating a new motion vector predictor candidate, from an accumulation of the one or more motion vectors along the plurality of different paths through the plurality of checked blocks over the one or more other pictures.
2. The apparatus of claim 1, wherein the one or more block positions are selected from among a plurality of spatial positions comprising at least one of: one or more spatially adjacent blocks or one or more spatially non-adjacent blocks.
3. The apparatus of claim 1, wherein the one or more block positions are located in one or more neighboring pictures in either a first prediction direction or a second prediction direction relative to the current picture.
4. The apparatus of claim 3, wherein the candidate block positions are associated with one or more of: a spatial candidate, a non-adjacent candidate, a temporal candidate, a history based candidate, or a chained motion predictor candidate.
695. The apparatus of claim 4, wherein the motion information comprises one or more motion vectors.
6. The apparatus of claim 5, wherein the motion information is recursively traced by at least recursively tracing the one or more motion vectors along the plurality of different paths through the plurality of checked blocks over the one or more other pictures.
7. The apparatus of claim 6, wherein the new motion vector prediction candidate is operable for use in generating or modifying an advance motion vector prediction (AMVP) list.
8. The apparatus of claim 7, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:adding the new motion vector prediction candidate to the AMVP list either during or after the spatial motion vector predictor candidate derivation but before temporal motion vector predictor candidate derivation.
9. The apparatus of claim 7, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:adding the new motion vector prediction candidate to the AMVP list when one or more of the positions checked for spatial candidate derivation is checked and the motion information is found to point to a reference picture that is not current the target reference.
10. The apparatus of claim 7, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:adding the new motion vector prediction candidate to the AMVP list after a redundancy check with the candidates that are already in the list is conducted.
11. The apparatus of claim 2, wherein the one plurality of checked blocks chosen for the plurality of different paths to derive the new motion vector prediction candidate depend on amotion vector storage granularity and a current block size of the current block in the current picture.
12. The apparatus of claim 11, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:in an instance in which the current block size is larger than the motion vector storage granularity, checking more than one block position in the one or more other pictures that correspond to an area above or left of the current block in the current picture.
13. The apparatus of claim 12, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:initiating chaining of motion vectors from a position within a candidate block that is checked and from which a motion vector is obtained or from a position within the current block.
14. The apparatus of claim 13, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:in merge mode, modifying chained motion vector prediction (CMVP) candidate derivation and deriving additional CMVP candidates by allowing the tracing of motion vectors to start from a position within a particular block from which the additional CMVP candidates are derived.
15. The apparatus of any prior claim, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform: recursively tracing motion vectors by recursively tracing motion using at least one motion vector if an inter-coded block is encountered and, in an instance in which the reference picture is not the target picture, continuing recursively tracing using at least one block vector if an intrablock copy (IBC) coded block or an intra-template matching prediction (intraTMP) block is encountered.
16. The apparatus of claim 15, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:terminating the recursively tracing and skipping the position by not creating a candidate if one of an inter-coded block, an IBC coded block, or an intraTMP coded block is not encountered in the reference picture that is not the target picture, or no valid motion information or block vector information is found.
17. The apparatus of claim 16, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:terminating the recursively tracing and skipping the candidate if no valid motion information or block vector information is found in the reference picture that is not the target picture.
18. The apparatus of claim 17, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:applying the recursive tracing of motion vectors to one or more other types of motion, the one or more other types of motion by generating chained motion vector candidates to predict the one or more other types of motion.
19. The apparatus of claim 18, wherein the one or more other types of motion comprise affine motion.
20. The apparatus of claim 14, wherein, in the merge mode, the recursively traced motion vectors are operable for use in generating or modifying a merge candidate list.
21. The apparatus of claim 20, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:adding one or more CMVPs to the merge candidate list that are generated by starting the recursive tracing from a position within the source block.
22. The apparatus of claim 21, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:atempting to add the new motion vector predictor candidate aggregated from the recursively traced blocks along the plurality of different paths through the plurality of checked blocks to the merge candidate list after a redundancy check with all spatial, temporal, nonadj acent, and history-based candidates is conducted.
23. The apparatus of claim 22, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:atempting to add the new motion vector predictor candidate to the merge candidate list before the CMVP candidate is recursively traced starting from the one or more block positions in the current block.
24. An apparatus comprising:at least one processor; andat least one memory comprising instructions stored therein that, when executed by the at least one processor, cause the apparatus to perform at least:while deriving a motion vector predictor candidate list for a reference picture, checking motion information from one or more block positions in a current picture;in an instance in which the motion information of the one or more block positions points to a reference picture that is not a target reference picture for which the motion vector predictor candidate list is being derived, recursively tracing the motion information, starting from the one or more block positions, along a plurality of motion vector paths through a plurality of block positions within a plurality of pictures until the target reference picture is found; and creating a new motion vector prediction candidate from an accumulation of the motion information recursively traced along the plurality of motion vector paths through the plurality of block positions within the plurality of pictures.
25. The apparatus of claim 24, wherein the instructions stored on the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:adding the new motion vector predictor candidate to a predictor list; andsignaling an index of the predictor list.
26. The apparatus of claim 25, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:selecting, from the predictor list, a best candidate; anddetermining, for the best candidate selected from the predictor list, a best motion vector prediction mode,wherein the signaling the index of the predictor list is performed based on the best mode vector prediction mode being a mode used when recursively tracing the motion information of the one or more source blocks to the first reference picture.
27. The apparatus of claim 25 or claim 26, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform: continuing to recursively trace the motion information of the one or more source blocks to the first reference picture until the first reference picture currently being searched is found; andonce the first reference picture in the group of pictures is found based on the recursive tracing of the motion information of the one or more source blocks to the first reference picture, discontinuing recursive tracing of the motion information of the one or more source blocks to the first reference picture.
28. The apparatus of any one of claims 24 to 27, wherein the motion information of the one or more source blocks comprise motion vectors and or block vectors.
29. The apparatus of any one of claims 24 to 28, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:in the instance in which the motion information associated with the one or more source blocks to the first reference picture indicate motion that is point to the second reference picture that is different from the first reference picture currently being searched, refraining from discarding the motion information of one or more source blocks.
30. The apparatus of any one of claims 24 to 29, wherein the motion information comprises motion vector weighting information.
31. The apparatus of any one of claims 24 to 30, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:in the instance in which motion vector weighting information for the first reference picture is not received, inferring the motion vector weighting information for the first reference picture from motion vector weighting information associated with the one or more source blocks to the first reference picture.
32. The apparatus of any one of claims 24 to 31, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:setting a motion vector for the first reference picture to an accumulation of the recursively traced motion vectors of the one or more source blocks.
33. The apparatus of claim 32, wherein the setting the motion vector for the first reference picture to the accumulation of the recursively traced motion vectors of the one or more source blocks occurs while the motion vectors are being checked before a temporal candidate.
34. The apparatus of claim 32 or claim 33, wherein the setting the motion vector for the first reference picture to the accumulation of the recursively traced motion vectors of the one or more source block occurs after all motion vectors obtained from spatially located blocks are checked as spatial candidates but before temporal motion vector predictor checking.
35. The apparatus of claim 34, wherein the checking the temporal motion vector predictor comprises fetching the motion information from temporally collocated blocks and adding the motion information to a candidate list if the motion information meets one or more criteria.
36. The apparatus of any one of claims 24 to 35, wherein the spatial candidate comprises an advanced motion vector prediction (AMVP) candidate.
37. The apparatus of claim 36, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:providing or generating an AMVP list for the AMVP candidate.
38. The apparatus of claim 37, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:adding, to the AMVP list, a traced motion vector for the first reference picture.
39. The apparatus of any one of claims 24 to 38, wherein the motion information comprises at least one of: a motion vector, a reference index, or a prediction direction.
40. The apparatus of claim 39, wherein the recursively tracing the motion information generates motion vector prediction information.
41. The apparatus of any one of claims 24 to 40, wherein the candidate comprises a chained motion vector prediction (CMVP) candidate.
42. The apparatus of claim 41, where CMVP candidates are used in a merge list.
43. The apparatus of claim 41 or claim 42, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform: starting a motion vector chain for the CMVP candidate from a position within one of the one or more source blocks or within a current block.
44. The apparatus of any one of claims 41 to 43, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:modifying or generating the CMVP candidate to include motion tracing information indicating a position from which one or more positions of one or more positions in the one or more neighboring blocks were fetched for initial tracing of motion vectors for the CMVP candidate.
45. The apparatus of any one of claims 43 to 44, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:generating a number of CMVP candidates;calculating a cost associated to each CMVP candidate;ranking the CMVP candidates based on their cost from lowest cost to highest cost; and utilizing, in a merge list, a predetermined number or an adaptively determined number of the CMVP candidates from the ranked list.
46. The apparatus of claim 45, where the cost associated with a candidate is calculated by generating the prediction for the template area of the current block and calculating the error between the prediction and the reconstructed samples of the template area based on a distance metric.
47. The apparatus of claim 46, wherein the distance metric comprises one or more of: a sum of absolute distance (SAD), a sum of squared error (SSE), or a sum of absolute transformed differences (SATD).
48. The apparatus of any one of claims 24 to 47, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:in an instance in which one or more block vectors is identified during the recursive tracing, continuing the recursive tracing using the one or more block vectors.
49. The apparatus of claim 43, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:in an instance in which the current block is not an inter-coded block, an intra-block copy (IBC) coded block or an intra-template matching prediction (TMP) coded block, discontinuing the recursive tracing of the motion vectors from the one or more source blocks and skipping the current block.
50. The apparatus of any one of claims 24 to 49, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:selecting one or more positions within the one or more source blocks based on one or more of: a motion vector storage granularity, or a block size.
51. The apparatus of any one of claims 24 to 50, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:receiving an indication that one or more pictures in the group of pictures contain motion reference information.
52. The apparatus of claim 51, wherein the motion reference information of the one or more pictures outside the predictor list is needed for decoding of the group of pictures.
53. The apparatus of claim 51 or claim 52, wherein the indication is received in a picture level or a slice level.
54. The apparatus of any one of claims 24 to 53, wherein the apparatus is, comprises, or is comprised as a part of an encoder device or a decoder device.
55. An apparatus comprising:at least one processor; andat least one memory comprising instructions stored therein that, when executed by the at least one processor, cause the apparatus to perform at least:searching for motion vector predictor candidates that point to a first reference picture in a group of pictures;in an instance in which motion information associated with one or more source blocks points to a second reference picture that is different from the first reference picture currently being searched, recursively tracing the motion information of the one or more source blocks along a plurality of different paths through a plurality of different checked blocks in a plurality of pictures in the group of pictures until the first reference picture is found; and creating a new motion vector predictor candidate based on a sum of the motion information along a final path through the plurality of different checked blocks in the plurality of different pictures, the new motion vector predictor candidate pointing to the first reference picture.
56. A method comprising:checking, from a plurality of candidate block positions, for motion vector information associated with one or more motion vectors for respective candidate block positions of the plurality of candidate block positions;in an instance in which the motion vector information associated with one or more candidate block positions from among the plurality of candidate block positions points to a reference picture different than a target reference picture for which motion vector predictor candidates are being derived, recursively tracing, starting from one or more block positions in a current picture, along a plurality of different paths through a plurality of checked blocks over one or more other pictures, the motion vector information until the target reference picture is found; andcreating a new motion vector predictor candidate, from an accumulation of the one or more motion vectors along the plurality of different paths through the plurality of checked blocks over the one or more other pictures.
57. The method of claim 56, wherein the one or more block positions are selected from among a plurality of spatial positions comprising at least one of: one or more spatially adjacent blocks or one or more spatially non-adjacent blocks.
58. The method of claim 56, wherein the one or more block positions are located in one or more neighboring pictures in either a first prediction direction or a second prediction direction relative to the current picture.
59. The method of claim 58, wherein the candidate block positions are associated with one or more of: a spatial candidate, a non-adjacent candidate, a temporal candidate, a history based candidate, or a chained motion predictor candidate.
60. The method of claim 59, wherein the motion information comprises one or more motion vectors.
61. The method of claim 60, wherein the motion information is recursively traced by at least recursively tracing the one or more motion vectors along the plurality of different paths through the plurality of checked blocks over the one or more other pictures.
62. The method of claim 61, wherein the new motion vector prediction candidate is operable for use in generating or modifying an advance motion vector prediction (AMVP) list.
63. The method of claim 62, further comprising:adding the new motion vector prediction candidate to the AMVP list either during or after the spatial motion vector predictor candidate derivation but before temporal motion vector predictor candidate derivation.
64. The method of claim 62, further comprising:adding the new motion vector prediction candidate to the AMVP list when one or more of the positions checked for spatial candidate derivation is checked and the motion information is found to point to a reference picture that is not current the target reference.
65. The method of claim 62, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the method to perform:adding the new motion vector prediction candidate to the AMVP list after a redundancy check with the candidates that are already in the list is conducted.
66. The method of claim 57, wherein the one plurality of checked blocks chosen for the plurality of different paths to derive the new motion vector prediction candidate depend on a motion vector storage granularity and a current block size of the current block in the current picture.
67. The method of claim 66, further comprising:in an instance in which the current block size is larger than the motion vector storage granularity, checking more than one block position in the one or more other pictures that correspond to an area above or left of the current block in the current picture.
68. The method of claim 67, further comprising:initiating chaining of motion vectors from a position within a candidate block that is checked and from which a motion vector is obtained or from a position within the current block.
69. The method of claim 68, further comprising:in merge mode, modifying chained motion vector prediction (CMVP) candidate derivation and deriving additional CMVP candidates by allowing the tracing of motion vectors to start from a position within a particular block from which the additional CMVP candidates are derived.
70. The method of any one of claims 56 to 69, further comprising:recursively tracing motion vectors by recursively tracing motion using at least one motion vector if an inter-coded block is encountered and, in an instance in which the reference picture is not the target picture, continuing recursively tracing using at least one block vector if an intrablock copy (IBC) coded block or an intra-template matching prediction (intraTMP) block is encountered.
71. The method of claim 70, further comprising:terminating the recursively tracing and skipping the position by not creating a candidate if one of an inter-coded block, an IBC coded block, or an intraTMP coded block is not encountered in the reference picture that is not the target picture, or no valid motion information or block vector information is found.
72. The method of claim 71, further comprising:terminating the recursively tracing and skipping the candidate if no valid motion information or block vector information is found in the reference picture that is not the target picture.
73. The method of claim 72, further comprising:applying the recursive tracing of motion vectors to one or more other types of motion, the one or more other types of motion by generating chained motion vector candidates to predict the one or more other types of motion.
74. The method of claim 73, wherein the one or more other types of motion comprise affine motion.
75. The method of claim 69, wherein, in the merge mode, the recursively traced motion vectors are operable for use in generating or modifying a merge candidate list.
76. The method of claim 75, further comprising:adding one or more CMVPs to the merge candidate list that are generated by starting the recursive tracing from a position within the source block.
77. The method of claim 76, further comprising:attempting to add the new motion vector predictor candidate aggregated from the recursively traced blocks along the plurality of different paths through the plurality of checked blocks to the merge candidate list after a redundancy check with all spatial, temporal, nonadj acent, and history-based candidates is conducted.
78. The method of claim 77, further comprising:attempting to add the new motion vector predictor candidate to the merge candidate list before the CMVP candidate is recursively traced starting from the one or more block positions in the current block.
79. A method comprising:while deriving a motion vector predictor candidate list for a reference picture, checking motion information from one or more block positions in a current picture;in an instance in which the motion information of the one or more block positions points to a reference picture that is not a target reference picture for which the motion vector predictorcandidate list is being derived, recursively tracing the motion information, starting from the one or more block positions, along a plurality of motion vector paths through a plurality of block positions within a plurality of pictures until the target reference picture is found; and creating a new motion vector prediction candidate from an accumulation of the motion information recursively traced along the plurality of motion vector paths through the plurality of block positions within the plurality of pictures.
80. The method of claim 79, wherein the instructions stored on the at least one memory, when executed by the at least one processor, further cause the method to perform:adding the new motion vector predictor candidate to a predictor list; andsignaling an index of the predictor list.
81. The method of claim 80, further comprising:selecting, from the predictor list, a best candidate; anddetermining, for the best candidate selected from the predictor list, a best motion vector prediction mode,wherein the signaling the index of the predictor list is performed based on the best mode vector prediction mode being a mode used when recursively tracing the motion information of the one or more source blocks to the first reference picture.
82. The method of claim 80 or claim 81, further comprising:continuing to recursively trace the motion information of the one or more source blocks to the first reference picture until the first reference picture currently being searched is found; andonce the first reference picture in the group of pictures is found based on the recursive tracing of the motion information of the one or more source blocks to the first reference picture, discontinuing recursive tracing of the motion information of the one or more source blocks to the first reference picture.
83. The method of any one of claims 79 to 82, wherein the motion information of the one or more source blocks comprise motion vectors and or block vectors.
84. The method of any one of claims 79 to 83, further comprising:in the instance in which the motion information associated with the one or more source blocks to the first reference picture indicate motion that is point to the second reference picture that is different from the first reference picture currently being searched, refraining from discarding the motion information of one or more source blocks.
85. The method of any one of claims 79 to 84, wherein the motion information comprises motion vector weighting information.
86. The method of any one of claims 79 to 85, further comprising:in the instance in which motion vector weighting information for the first reference picture is not received, inferring the motion vector weighting information for the first reference picture from motion vector weighting information associated with the one or more source blocks to the first reference picture.
87. The method of any one of claims 79 to 86, further comprising:setting a motion vector for the first reference picture to an accumulation of the recursively traced motion vectors of the one or more source blocks.
88. The method of claim 87, wherein the setting the motion vector for the first reference picture to the accumulation of the recursively traced motion vectors of the one or more source blocks occurs while the motion vectors are being checked before a temporal candidate.
89. The method of claim 87 or claim 88, wherein the setting the motion vector for the first reference picture to the accumulation of the recursively traced motion vectors of the one or more source block occurs after all motion vectors obtained from spatially located blocks are checked as spatial candidates but before temporal motion vector predictor checking.
90. The method of claim 89, wherein the checking the temporal motion vector predictor comprises fetching the motion information from temporally collocated blocks and adding the motion information to a candidate list if the motion information meets one or more criteria.
91. The method of any one of claims 79 to 90, wherein the spatial candidate comprises an advanced motion vector prediction (AMVP) candidate.
92. The method of claim 91, further comprising:providing or generating an AMVP list for the AMVP candidate.
93. The method of claim 92, further comprising:adding, to the AMVP list, a traced motion vector for the first reference picture.
94. The method of any one of claims 79 to 93, wherein the motion information comprises at least one of: a motion vector, a reference index, or a prediction direction.
95. The method of claim 94, wherein the recursively tracing the motion information generates motion vector prediction information.
96. The method of any one of claims 79 to 95, wherein the candidate comprises a chained motion vector prediction (CMVP) candidate.
97. The method of claim 96, where CMVP candidates are used in a merge list.
98. The method of claim 96 or claim 97, further comprising:starting a motion vector chain for the CMVP candidate from a position within one of the one or more source blocks or within a current block.
99. The method of any one of claims 96 to 98, further comprising:modifying or generating the CMVP candidate to include motion tracing information indicating a position from which one or more positions of one or more positions in the one ormore neighboring blocks were fetched for initial tracing of motion vectors for the CMVP candidate.
100. The method of claim 98 or claim 99, further comprising:generating a number of CMVP candidates;calculating a cost associated to each CMVP candidate;ranking the CMVP candidates based on their cost from lowest cost to highest cost; and utilizing, in a merge list, a predetermined number or an adaptively determined number of the CMVP candidates from the ranked list.
101. The method of claim 100, where the cost associated with a candidate is calculated by generating the prediction for the template area of the current block and calculating the error between the prediction and the reconstructed samples of the template area based on a distance metric.
102. The method of claim 101, wherein the distance metric comprises one or more of: a sum of absolute distance (SAD), a sum of squared error (SSE), or a sum of absolute transformed differences (SATD).
103. The method of any one of claims 79 to 102, further comprising:in an instance in which one or more block vectors is identified during the recursive tracing, continuing the recursive tracing using the one or more block vectors.
104. The method of claim 103, further comprising:in an instance in which the current block is not an inter-coded block, an intra-block copy (IBC) coded block or an intra-template matching prediction (TMP) coded block, discontinuing the recursive tracing of the motion vectors from the one or more source blocks and skipping the current block.
105. The method of any one of claims 79 to 104, further comprising:selecting one or more positions within the one or more source blocks based on one or more of: a motion vector storage granularity, or a block size.
106. The method of any one of claims 79 to 105, further comprising:receiving an indication that one or more pictures in the group of pictures contain motion reference information.
107. The method of claim 106, wherein the motion reference information of the one or more pictures outside the predictor list is needed for decoding of the group of pictures.
108. The method of claim 106 or claim 107, wherein the indication is received in a picture level or a slice level.
109. The method of any one of claims 56 to 108, wherein the method is carried out at least in part by one of: an encoder device or a decoder device.
110. A method comprising:searching for motion vector predictor candidates that point to a first reference picture in a group of pictures;in an instance in which motion information associated with one or more source blocks points to a second reference picture that is different from the first reference picture currently being searched, recursively tracing the motion information of the one or more source blocks along a plurality of different paths through a plurality of different checked blocks in a plurality of pictures in the group of pictures until the first reference picture is found; andcreating a new motion vector predictor candidate based on a sum of the motion information along a final path through the plurality of different checked blocks in the plurality of different pictures, the new motion vector predictor candidate pointing to the first reference picture.87111. An apparatus comprising means for perform a method according to any one of claims 56 to 110.
112. A computer program product comprising at least one non-transitory computer-readable storage media comprising instructions stored therein that, when executed by at least one processor of an apparatus, cause the apparatus to perform a method according to any one of claims 56 to 110.88