Decryption method, apparatus, and encryption method
The SbTMVP method addresses the challenge of predicting motion vectors in video coding by using template matching to derive displacement vectors and motion vector offsets for sub-blocks, resulting in improved compression efficiency and reduced bandwidth requirements.
Patent Information
- Application Number
- JP2024514692
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-09
- Filing Date
- 2022-11-11
- Publication Date
- 2025-06-12
AI Technical Summary
Existing video coding technologies face challenges in efficiently predicting motion vectors, particularly in sub-blocks, which affects compression efficiency and bandwidth utilization.
The proposed solution involves a subblock-based template matching motion vector predictor (SbTMVP) that uses a displacement vector and a motion vector offset derived through template matching to predict motion vectors for sub-blocks within a current block.
This approach enhances compression efficiency by accurately predicting motion vectors at the sub-block level, thereby reducing the bitstream size and improving decoding performance.
Smart Images

Figure 2025517840000001_ABST
Abstract
Description
Technical Field
[0001] [Related Applications] This application claims the benefit of priority of U.S. Patent Application No. 17 / 983,866, "SUBBLOCK-BASED MOTION VECTOR PREDICTOR WITH MV OFFSET DERIVED BY TEMPLATE MATCHING", filed on November 9, 2022, which claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 344,840, "Subblock Based Motion Vector Predictor With MV Offset Derived By Template Matching", filed on May 23, 2022. The disclosures of the foregoing applications are hereby incorporated herein by reference in their entireties.
[0002] [Technical Field] The present disclosure generally describes embodiments related to video coding.
Background Art
[0003] The background description provided herein is for the purpose of generally presenting the background of the present disclosure. The research of the presently named inventors, to the extent that it is not considered prior art at the time of filing in the context of the research described in this background chapter, is not expressly or implicitly admitted as prior art to the present disclosure, in the same manner as aspects of the description that may not be considered prior art at the time of filing.
[0004] Uncompressed digital images and / or videos can include a series of pictures, each picture having a spatial dimension of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate, for example, 60 pictures per second or 60 Hz (also known as the frame rate in short form). Uncompressed images and / or videos have specific bitrate requirements. For example, an 8-bit / sample 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) requires a bandwidth close to 1.5 Gbit / s. One hour of such a video requires more than 600 GByte of storage space.
[0005] One purpose of image and / or video coding and decoding can be to reduce the redundancy in the input image and / or video signal through compression. Compression can, in some cases, help reduce the bandwidth and / or storage space requirements by more than an order of magnitude in size. The description in this specification uses video encoding / decoding as an example for illustration, but the same techniques can be applied to image encoding / decoding in a similar way without departing from the spirit of the present disclosure. Both lossless compression and lossy compression, and combinations thereof, can be utilized. Lossless compression represents techniques where an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal is not identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to generate a useful reconstructed signal for the intended application. In the case of video, lossy compression is widely used. The amount of tolerable distortion depends on the application, and users of a particular consumer streaming application can tolerate higher distortion than users of a television distribution application. The achievable compression ratio can reflect that the higher the acceptable / tolerable distortion, the higher the compression ratio that can be achieved.
[0006] Video encoders and decoders can utilize techniques from several broad classifications, including, for example, motion compensation, transform processing, quantization, and entropy coding.
[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without referring to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in an intra mode, that picture can be an intra picture. Intra pictures, and their derivatives such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session, or as a still image. Samples of an intra block can be transformed, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique that minimizes sample values in the domain before transformation. In some cases, the smaller the DC value after transformation and the smaller the AC coefficients, the fewer bits are required at a given quantization step size to represent the block after entropy coding.
[0008] For example, traditional intra coding used in MPEG-2 generation coding technology does not use intra prediction. However, some new video compression techniques attempt to perform prediction based on, for example, surrounding sample data and / or metadata obtained during encoding and / or decoding of data blocks. Such techniques are hereinafter referred to as "intra prediction" techniques. In at least some cases, intra prediction uses only reference data from the current picture being reconstructed, rather than from a reference picture.
[0009] There can be many different forms of intra prediction. When more than one such technique can be used in a given video coding technique, the particular technique used can be coded as a particular intra prediction mode that uses the particular technique. In certain cases, the intra prediction mode can have sub - modes and / or parameters, and the sub - modes and / or parameters can be coded either individually or included in the mode codeword that defines the prediction mode being used. Which codeword should be used for a given combination of mode, sub - mode, and / or parameter can affect the improvement of coding efficiency through intra prediction, and thus entropy coding techniques can be used to convert the codeword into a bitstream.
[0010] A particular intra prediction mode was introduced by H.264, improved in H.265, and further improved in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). The prediction block can be formed using neighboring sample values of already available samples. The sample values of the neighboring samples are copied into the prediction block according to a direction. The reference to the direction in use can be coded within the bitstream or can itself be predicted.
[0011] Referring to FIG. 1A, a subset of 9 prediction directions can be seen in the lower right, corresponding to 33 of the 35 intra modes defined in H.265. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples at an angle of 45 degrees from horizontal and towards the upper right. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples at an angle of 22.5 degrees from horizontal and towards the lower left of sample (101).
[0012] Referring further to FIG. 1A, in the upper left, a square block (104) of 4×4 samples (shown by thick dashed lines) is shown. The square block (104) contains 16 samples, each sample being labeled with "S", its Y - dimensional position (e.g., row index), and its X - dimensional position (e.g., column index). For example, sample S21 is the second sample from the top in the Y - dimension and the first sample from the left in the X - dimension. Similarly, sample S44 is the fourth sample within block (104) in both the Y and X dimensions. When the block is of size 4×4 samples, S44 is in the lower right. Further, reference samples are shown following a similar numbering scheme. The reference samples are labeled with "R", its Y - position (e.g., row index) and X - position (column index) with respect to block (104). In both H.264 and H.265, the predicted samples are in the neighborhood of the block being reconstructed, and thus negative values need not be used.
[0013] Intra picture prediction can operate by copying the reference sample value from neighboring samples when indicated by the signaled prediction direction. For example, a coded video bitstream includes signaling indicating a prediction direction that matches arrow (102) for this block. That is, the samples are predicted at an angle of 45 degrees from horizontal and to the upper right. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Sample S44 is then predicted from reference sample R08.
[0014] In certain cases, in order to calculate the reference sample, especially when the direction cannot be evenly divided by 45 degrees, the values of multiple reference samples may be combined, for example, through interpolation.
[0015] The number of possible directions has been increasing as video coding technology has evolved. In H.264 (2003), there were potentially 9 different directions presented. That increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and specific techniques in entropy coding are used to represent these likely directions with a small number of bits while accepting a specific penalty for less likely directions. Further, the direction itself may be predicted from neighboring directions in neighboring already decoded blocks.
[0016] FIG. 1B shows a diagram (110) showing 65 intra prediction directions according to JEM to illustrate the increasing number of prediction directions over time.
[0017] The mapping of intra prediction direction bits representing directions within a coded video bitstream may vary depending on the video coding technology. Such mappings can range from simple direct mappings to complex adaptive schemes including codewords, most likely modes, and similar techniques. However, in most cases, there may exist certain directions in video content that statistically occur less frequently than certain other directions. Since the goal of video compression is redundancy reduction, these less likely directions will be represented by more bits than the more likely directions in a well - operating video coding technology.
[0018] Image and / or video coding and decoding can be performed using inter - picture prediction with motion compensation. Motion compensation is a lossy compression technique related to a technique where a block of sample data from a previously reconstructed picture or a portion thereof (reference picture) is spatially shifted in a direction indicated by a motion vector (hereinafter, MV) and then used for prediction of a newly reconstructed picture or picture portion. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or three dimensions where the third dimension is an indication of the reference picture in use (the latter can be indirectly the temporal dimension).
[0019] In some video compression techniques, the motion vectors (MVs) applicable to a particular region of sample data can be predicted from other MVs, for example, from MVs associated with another region of sample data that is spatially adjacent to the region being reconstructed and that precedes the MV in the decoding order. Doing so can, as a result, reduce the amount of data necessary to code the MVs, thereby removing redundancy and improving compression. MV prediction can, for example, be statistically possible when coding an input video signal obtained from a camera (known as natural video) because regions larger than the region to which a single MV is applicable move in a similar direction and, thus, in some cases can be predicted using similar motion vectors derived from the MVs of neighboring regions. This results in an MV found for a given region that is similar or identical to the MV predicted from surrounding MVs. Also, this can be presented in fewer bits than would be used if the MVs were directly coded after entropy coding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MVs) obtained from the original signal (i.e., the sample stream). In other cases, MV prediction itself can result in loss because, for example, errors are rounded when calculating predictors from some surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). From among the many MV prediction mechanisms provided by H.265, the technique called "spatial merge" will be described below with reference to FIG. 2.
[0021] Referring to FIG. 2, the current block (201) includes samples found by the encoder as being predictable from a previous block of the same size that has been spatially shifted during motion search processing. Instead of directly coding the MV, the MV can be derived from metadata associated with one or more reference pictures, for example from the nearest reference picture (in decoding order), using an MV associated with any one of five surrounding samples A0, A1, and B0, B1, B2 (202-206 respectively). In H.265, MV prediction can use predictors from the same reference picture as that used by neighboring blocks.
Summary of the Invention
[0022] Aspects of the disclosure provide methods and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit.
[0023] According to aspects of the present disclosure, a method of video decoding executed in a video decoder is provided. In this method, a coded video bitstream including a current block in a current picture can be received. The current block includes a plurality of sub-blocks and is predicted by a SbTMVP (subblock-based template matching motion vector prediction) mode. Each collocated reference sub-block of each sub-block can be determined based on a combination of a displacement vector (DV) and a motion vector offset (MVO) associated with each sub-block. A motion vector (MV) field in each collocated reference sub-block of each sub-block in the current block can be determined. Each reference template of each sub-block can be derived based on the determined MV field of the collocated reference sub-block. In the SbTMVP mode, a plurality of sub-blocks of the current block can be reconstructed by predicting each sub-block using each reference template.
[0024] To determine each identical - position reference sub - block, a search region located in one of the current picture and the reference pictures of the current picture can be determined. One or more reference blocks of the current block can be determined based on template matching between the template of the current block and the templates of one or more reference blocks in the search region. The template of the current block can include samples adjacent to the current block. The template of each of the one or more reference blocks can include samples adjacent to each of the one or more reference blocks. Each identical - position reference sub - block of each sub - block can be determined as a sub - block that is in the same position as each sub - block in one of the one or more reference blocks.
[0025] The template matching between the template of the current block and the templates of each of the one or more reference blocks can be determined based on one of the sum of absolute differences (SAD), sum of absolute transformed differences (SATD), sum of squared errors (SSE), subsampled SAD, and mean - removed SAD.
[0026] To determine one or more reference blocks, a plurality of candidate reference blocks can be determined within the search region. A plurality of cost values can be determined based on template matching between the template of the current block and the templates of the plurality of candidate reference blocks. One or more reference blocks can be determined as one or more candidate reference blocks among the plurality of candidate reference blocks corresponding to one or more of the lowest cost values among the plurality of cost values.
[0027] In one embodiment, the search region can include one of (i) a region centered at a position that is in the same position as the current block in the reference picture and (ii) a region centered at the current block in the current picture.
[0028] In one example, the search region can be determined based on a displacement vector (DV). The DV can be derived from one of (i) the motion vectors of blocks spatially adjacent to the current block, and (ii) the motion vectors of the merge candidate list of the current block.
[0029] In one example, the search region can be determined as a region centered on the sample indicated by the DV, and the region can be one of a square, a rectangle, and a rhombus.
[0030] In one example, the search region can be determined as a sample group centered on the sample indicated by the DV. The sample group can be located at at least one of 0 degrees, 45 degrees, 90 degrees, or 135 degrees with respect to the sample indicated by the DV.
[0031] To determine one or more reference blocks, a first reference block of the one or more reference blocks can be determined. The first reference block can be indicated by a first displacement vector (DV) from the template of the current block to the template of the first reference block. In one example, the first DV can be derived based on template matching such that the first DV corresponds to a cost value related to the difference between the template of the first reference block and the template of the current block. In one example, the first DV can be signaled.
[0032] The search region can be determined based on a first displacement vector (DV) derived before template matching. The reference block among the one or more reference blocks can be determined based on a second displacement vector (DV) from the template of the current block to the template of the first reference block. The second DV can be derived based on template matching such that the second DV corresponds to a cost value related to the difference between the template of the first reference block and the template of the current block.
[0033] To reconstruct the sub - blocks of the current block, one or more motion vectors (MVs) of the first sub - block among the plurality of sub - blocks in the current block can be determined based on one or more MVs of the sub - block at the same position as the first sub - block in one or more reference blocks. One or more predicted sub - blocks of the first sub - block among the plurality of sub - blocks can be determined based on one or more MVs of the first sub - block. The predicted samples of the first sub - block can be determined based on one or a weighted combination of one or more of the predicted sub - blocks.
[0034] In some embodiments, a plurality of candidate reference blocks of the current block can be determined based on a plurality of displacement vectors (DVs). Each of the plurality of candidate reference blocks can be indicated by a respective one of the plurality of DVs. One or more reference blocks of the current block can be determined from the plurality of candidate reference blocks based on one or more cost values of template matching.
[0035] According to another aspect of the present disclosure, a device is provided. The device includes a processing circuit. The processing circuit can be configured to execute any of the video encoding / decoding methods.
[0036] Aspects of the present disclosure also provide a non - transitory computer - readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to execute any of the methods for video encoding / decoding.
Brief Description of the Drawings
[0037] Further features, characteristics, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
[0038]
Figure 1A
[0039]
Figure 1B
[0040]
Figure 2
[0041]
Figure 3
[0042]
Figure 4
[0043]
Figure 5
[0044]
Figure 6
[0045]
Figure 7
[0046]
Figure 8
[0047]
Figure 9
[0048]
Figure 10
[0049]
Figure 11
[0050]
Figure 12
[0051]
Figure 13
[0052]
Figure 14A
[0053]
Figure 14B
[0054]
Figure 15
[0055]
Figure 16
[0056]
Figure 17
[0057]
Figure 18
[0058]
Figure 19
[0059]
Figure 20
[0060]
Figure 21
DETAILED DESCRIPTION OF THE INVENTION
[0061] FIG. 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) can code video data (a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission may be common in media serving applications and the like.
[0062] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of video data encoded, for example, during a video conference. In the bidirectional transmission of data, the terminal devices (330) and (340) may encode video data (e.g., a stream of video pictures captured by a terminal device) for transmission to the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) may receive the encoded video data transmitted by the other of the terminal devices (330) and (340), may decode the encoded video data to restore the video pictures, and may display the video pictures on an accessible display device according to the restored video data.
[0063] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) are each shown as a server, a personal computer, and a smartphone, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure have applications in laptop computers, tablet computers, media players, and / or dedicated video conferencing facilities. The network (350) represents any number of networks that carry encoded video data among the terminal devices (310), (320), (330), and (340), and includes, for example, wired (wired) and / or wireless communication networks. The communication network (350) may exchange data on circuit-switched and / or packet-switched channels. Representative networks include electronic communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of the discussion of the present invention, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise specifically noted hereinafter.
[0064] FIG. 4 shows a video encoder and a video decoder in a streaming environment as an example of the application of the disclosed subject matter. The disclosed subject matter is equally applicable to, for example, storage of compressed video on digital media including video conferencing, digital TV, streaming services, CDs, DVDs, memory sticks, etc., other video-enabled applications, etc.
[0065] A streaming system may include a capture subsystem (413) that can include, for example, a video source (401) that generates an uncompressed video picture stream (402). In one example, the video picture stream (402) includes samples captured by a digital camera. The video picture stream (402) is shown in bold lines to emphasize its high data capacity when compared to the encoded video data (404) (or coded video bitstream), and can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) includes hardware, software, or a combination thereof, and can enable or implement aspects of the disclosed subject matter as detailed below. The encoded video data (404) (or encoded video bitstream) is shown in thin lines to emphasize its low data capacity when compared to the video picture stream (402), and can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to read copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and generates an output video picture stream (411) that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to specific video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is known informally as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC.
[0066] Note that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can also include a video encoder (not shown).
[0067] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). In the example of FIG. 4, the video decoder (510) can be used instead of the video decoder (410).
[0068] The receiver (531) can receive one or more coded video sequences that are decoded by the video decoder (510). In an embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequence may be received from a channel (501) that may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams that may be transferred to respective using entities (not shown). The receiver (531) may separate the coded video sequence from the other data. To remove network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, “parser (520)”). In certain applications, the buffer memory (515) is part of the video decoder (510). Alternatively, it may be external to the video decoder (510) (not shown). Still alternatively, there may be a buffer memory (not shown) in addition to the buffer memory (515) that may be external to the video decoder (510), for example, to remove network jitter, or internal to the video decoder (510), for example, to process playout timing. When the receiver (531) is receiving data controllably from a storage / transfer device with sufficient bandwidth or from an isosynchronous network, the buffer memory (515) may be unnecessary or can be made small. For use in a best-effort packet network such as the Internet, a buffer memory (515) may be required, may be relatively large, advantageously be of an adaptable size, and may be implemented at least in part external to the operating system or similar elements (not shown) external to the video decoder (510).
[0069] Video decoder (510) may include a parser (520) to reconstruct symbols (521) from the coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510) and, in some cases, information for controlling a rendering device (such as a display screen) like a rendering device (512) that is not an integrated part of the electronic device (530) but can be coupled to the electronic device (530) as shown in FIG. 5. The control information for the rendering device may be in the form of an SEI (Supplemental Enhancement Information) message or a VUI (Video Usability Information) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context dependence, etc. The parser (520) may extract a set of subgroup parameters from the coded video sequence based on at least one parameter corresponding to at least one subgroup of pixels in the video decoder for at least one of the subgroups. Subgroups may include GOP (Groups of Picture), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (520) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.
[0070] The parser (520) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (515) to generate symbols (521).
[0071] The reconstruction of symbol (521) may include a plurality of different units depending on the type of the coded video picture or a portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How which units are included can be controlled by group control information parsed by a parser (520) from the coded video sequence. Such a flow of subgroup control information between the parser (520) and the following plurality of units is not shown for clarity.
[0072] Beyond the function blocks already mentioned, the video decoder (510) can be conceptually subdivided into a number of functional units, as will be described later. In an actual implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0073] The first unit is a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbol (521) from the parser (520), quantized transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block including sample values that can be input to an aggregator (555).
[0074] In some cases, the output samples of the scaler / inverse transform unit (551) may be related to intra-coded blocks. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed, using the surrounding already-reconstructed information fetched from the current picture buffer (558). The current picture buffer (558) buffers, for example, the reconstructed current picture partially and / or the reconstructed current picture completely. The aggregator (555) adds, in some cases, sample-by-sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0075] In other cases, the output samples of the scaler / inverse transform unit (551) may be related to inter-coded, and in some cases motion-compensated, blocks. In such cases, the motion-compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbol (521) related to the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) to generate the output sample information (in this case, called the residual samples or the residual signal). The address in the reference picture memory (557) from which the motion-compensation prediction unit (553) fetches the prediction samples can be controlled by the available motion vectors of the motion-compensation prediction unit (553) in the form of, for example, a symbol (521) having X, Y, and reference picture components. Motion compensation can include interpolation of the sample values fetched from the reference picture memory (557) when an exact sub-sample motion vector is in use, a motion vector prediction mechanism, etc.
[0076] The output samples of the aggregator (555) may undergo various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filter techniques that are controlled by parameters included in a coded video sequence (also referred to as a coded video bitstream) and enable the loop filter unit (556) to use symbols (521) from the parser (520). Video compression can respond not only to previously reconstructed and loop-filtered sample values, but also to meta-information obtained during the decoding of the previous part (in decoding order) of a coded picture or coded video sequence.
[0077] The output of the loop filter unit (556) can be an output to the renderer device (512) and can be a sample stream that can be stored in the reference picture memory (557) for use in future inter-picture prediction.
[0078] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a fresh current picture buffer can be reallocated before starting the reconstruction of subsequent coded pictures.
[0079] The video decoder (510) may perform a decoding operation according to a standard such as ITU-T Rec.H.265 or a predetermined video compression technique. The coded video sequence may conform to the syntax specified by the video compression technique or standard in use in the sense that it conforms to both the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile can select specific tools from all the tools available in the video compression technique or standard as tools that can only be used under the profile. Also, what is necessary for compliance may be that the complexity of the coded video sequence is within the limits defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured, for example, in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted through the HRD (Hypothetical Reference Decoder) specification and the metadata for HRD buffer management signaled in the coded video sequence.
[0080] In an embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0081] FIG. 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used instead of the video encoder (403) in the example of FIG. 4.
[0082] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture a video image to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).
[0083] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media providing system, the video source (601) may be a storage device storing pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may subsequently be provided as a plurality of individual pictures that give movement when viewed in succession. The pictures themselves may be organized as a spatial array of pixels. Each pixel may include one or more samples depending on the sampling structure, color space, etc. in use. A person skilled in the art can immediately understand the relationship between a pixel and a sample. The following description focuses on samples.
[0084] According to an embodiment, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required. Implementing an appropriate coding speed is one function of the control unit (650). In some embodiments, the control unit (650) controls other functional units described below and is functionally coupled to the other functional units. The coupling is not shown for clarity. Parameters set by the control unit (650) may include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, ...), picture size, GOP (group of pictures) layout, maximum motion vector search range, etc. The control unit (650) may be configured to have other appropriate functions related to the video encoder (603) optimized for a specific system design.
[0085] In some embodiments, the video encoder (603) is configured to operate within a coding loop. As a very simplified explanation, in one example, the coding loop may include a source coder (630) (which is responsible for generating symbols such as a symbol stream based on the input picture and reference pictures to be coded), and a (local) decoder (633) built into the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in the same way as a (remote) decoder does. The reconstructed sample stream (sample data) is input into the reference picture memory (634). When the decoding of the symbol stream yields bit-exact results independent of the decoder location (local or remote), the content of the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" exactly the same sample values as the decoder "sees" as reference picture samples when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is similarly used in several related technologies.
[0086] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder such as the video decoder (510) detailed above in connection with FIG. 5. Briefly referring to FIG. 5 for a moment, however, since the symbols are available and the encoding / decoding of the symbols into the encoded video sequence by the entropy encoder (645) and the parser (520) can be lossless, the entropy decoding part of the video decoder (510) including the buffer memory (515), and the parser (520) need not be fully implemented in the local decoder (633).
[0087] In an embodiment, decoder techniques other than parsing / entropy decoding present in a decoder exist in a corresponding encoder in the same or substantially the same functional form. Thus, the described subject matter of the disclosure focuses on decoder operations. The description of encoder techniques can be omitted since they are the reverse of the decoder techniques described comprehensively. In certain areas, more detailed descriptions are provided below.
[0088] During operation, in some examples, the source coder (630) may perform motion-compensated predictive coding. This predictively codes an input picture by referring to one or more previous coded pictures from a video sequence designated as a "reference picture". In this method, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that may be selected as a prediction reference for the input picture.
[0089] The local video decoder (633) may decode the coded video data of a picture that may be designated as a reference picture based on the symbols generated by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data can be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder for the reference picture, resulting in a reconstructed reference picture to be stored in the reference picture memory (634). Thus, the video encoder (603) may store a copy of the reconstructed reference picture having the same content as the reconstructed reference picture obtained by the remote video decoder (if there are no transmission errors).
[0090] The predictor (635) may perform predictive search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (such as a candidate reference pixel block) or specific metadata such as reference picture motion vectors, block shapes, etc. that can function as an appropriate predictive reference for the new picture. The predictor (635) may operate on a sample block - pixel block basis to find an appropriate predictive reference. In some examples, the input picture may have a predictive reference drawn from a plurality of reference pictures stored in the reference picture memory (634) as determined by the search results obtained by the predictor (635).
[0091] The control unit (650) may manage the coding operations of the source coder (630), including, for example, setting parameters and subgroup parameters used for the coding of video data.
[0092] The outputs of all the aforementioned functional units may undergo entropy coding in the entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable - length coding, arithmetic coding, etc.
[0093] The transmitter (640) may buffer the coded video sequence generated by the entropy coder (645) for transmission via a communication channel (660) which may be a hardware / software link to a storage device capable of storing the coded video data. The transmitter (640) may merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (not shown source).
[0094] The control unit (650) may manage the operation of the video encoder (603). During coding, the control unit (650) may assign to each coding picture a specific coding picture type that may affect the coding technique applicable to each picture. For example, a picture may often be assigned as one of the following picture types.
[0095] An intra picture (I picture) may be a picture that can be coded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, for example, IDR (Independent Decoder Refresh) pictures. Those skilled in the art recognize the variations of I pictures, and their individual applications and characteristics.
[0096] A predictive picture (P picture) may be a picture that can be coded and decoded using intra prediction or inter prediction, typically using one motion vector and a reference index to predict the sample values of each block.
[0097] A bi-directionally predictive picture (B picture, Bi-directionally Predictive Picture (B Picture)) may be a picture that can be coded and decoded using intra prediction or inter prediction, typically using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predictive picture can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0098] The source picture may generally be spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and coded block by block. The blocks may be coded predictively by reference to other (already coded) blocks determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively, or they may be coded predictively by reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be coded predictively via spatial prediction or via temporal prediction by reference to one previously coded reference picture. Blocks of a B picture may be coded predictively via spatial prediction or via temporal prediction by reference to one or two previously coded reference pictures.
[0099] The video encoder (603) may perform a coding operation in accordance with a predetermined video coding technology or standard such as ITU-T Rec.H.265. In that operation, the video encoder (603) may perform various compression operations including a predictive coding operation that utilizes the temporal and spatial redundancy in the input video sequence. The coded video data may thus conform to the syntax specified by the video coding technology or standard being used.
[0100] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR extension layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0101] Video may be captured as a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (which may be abbreviated as intra prediction) utilizes the spatial correlation within a given picture, and inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture during encoding / decoding is referred to as the current picture and is partitioned into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture when multiple reference pictures are in use.
[0102] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. According to the bi-prediction technique, two reference pictures such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future in display order respectively) are used. A block within the current picture can be coded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by combining the first reference block and the second reference block.
[0103] Furthermore, to improve coding efficiency, merge mode techniques can be used in inter-picture prediction.
[0104] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed within a unit of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are partitioned into coding tree units (CTUs) for compression. CTUs within a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Usually, a CTU includes three coding tree blocks (CTBs), that is, one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a 64×64 pixel CTU can be partitioned into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter-prediction type or an intra-prediction type. The CU is partitioned into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Usually, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed within a unit of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0105] FIG. 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence and encode the processing block into a coding picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) in the example of FIG. 4.
[0106] In an example of HEVC, a video encoder (703) receives a matrix of sample values of a processing block, such as a prediction block of 8×8 samples. The video encoder (703) determines whether the processing block is optimally coded using an intra mode, an inter mode, or a bi-prediction mode, for example, using rate-distortion optimization. When the processing block is coded in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into a coding picture. When the processing block is coded in the inter mode or the bi-prediction mode, the video encoder (703) may use inter prediction or bi-prediction techniques, respectively, to encode the processing block into a coding picture. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without benefit of coding motion vector components external to the predictor. In certain other video coding techniques, there may be coding motion vector components applicable to the target block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown), to determine the mode of the processing block.
[0107] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general control unit (721), and an entropy encoder (725) that are coupled together as shown in FIG. 7.
[0108] The inter-encoder (730) is configured to receive samples of a current block (e.g., a block being processed), compare the block with one or more reference blocks (e.g., blocks in a previous picture and a subsequent picture) in a reference picture, generate inter-prediction information (e.g., an explanation of redundant information by an inter-coding technique, a motion vector, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the coded video information.
[0109] The intra-encoder (722) is configured to receive samples of a current block (e.g., a block being processed), and in some cases, compare the block with already-coded blocks in a sample picture, generate quantized coefficients after transformation, and in some cases also generate intra-prediction information (e.g., intra-prediction direction information by one or more intra-coding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks in the same picture.
[0110] The general control unit (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general control unit (721) determines a mode of a block and provides a control signal to a switch (726) based on the mode. For example, when the mode is an intra mode, the general control unit (721) controls the switch (726) to select an intra-mode result for use by a residual calculator (723), and controls an entropy encoder (725) to select intra-prediction information and include the intra-prediction information in a bitstream. When the mode is an inter mode, the general control unit (721) controls the switch (726) to select an inter-prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter-prediction information and include the inter-prediction information in the bitstream.
[0111] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the selected prediction result from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data and generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, which in some examples is buffered in a memory circuit (not shown) and can be used as a reference picture.
[0112] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to a suitable standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. Note that there is no residual information when coding a block in either the merge submode of the inter mode or the bi-prediction mode according to the disclosed subject matter.
[0113] FIG. 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a reconstructed picture. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.
[0114] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) that are coupled together as shown in FIG. 8.
[0115] The entropy decoder (871) may be configured to reconstruct from the coded picture certain symbols that represent the generated syntax elements of the coded picture. Such symbols may include, for example, the coded mode of a block (e.g., the latter two of an intra mode, an inter mode, a bi - directional mode, a merge sub - mode or another sub - mode), prediction information (e.g., intra prediction information or inter prediction information) that can identify specific samples or metadata used for prediction by each of the intra decoder (872) or the inter decoder (880). The symbols may also include residual information, for example, in the form of quantized transform coefficients. In one example, when the prediction mode is an inter or bi - directional prediction mode, the inter prediction information is provided to the inter decoder (880), and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information is inverse - quantized and provided to the residual decoder (873).
[0116] The inter decoder (880) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.
[0117] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0118] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients, and process the inverse quantized transform coefficients to convert the residual information from the frequency domain information to the spatial domain. The residual decoder (873) may also require specific control information (for including the quantizer parameter (QP)). This information may be provided by the entropy decoder (871) (since this is only low-capacity control information, the data path is not shown).
[0119] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information as the output by the residual decoder (873) and the prediction result (optionally as the output by the inter or intra prediction module) to form a reconstruction block that can be part of the reconstructed picture and on the other hand can be part of the reconstructed video. Other appropriate operations such as a deblocking operation can be performed to improve the visual quality.
[0120] Note that the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using any appropriate technology. In one embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603), and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.
[0121] The present disclosure includes embodiments related to the derivation of SbTMVP (subblock-based temporal motion vector prediction). SbTMVP can be derived using a displacement motion vector with a motion vector offset.
[0122] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) issued the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, the two standardization organizations formed JVET (Joint Video Exploration Team) together to explore the possibility of developing the next-generation video coding standard after HEVC. In October 2017, the two standardization organizations announced the Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, 22 CfP responses regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding 360 video categories were each submitted. In April 2018, all the received CfP responses were evaluated at the 122nd MPEG / 10th JVET meeting. As a result of this meeting, JVET officially started the standardization process for next-generation video coding beyond HEVC, the new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the VVC video coding standard (version 1).
[0123] In inter prediction, for each inter-predicted coding unit (CU), motion parameters are required for the coding characteristics of VVC, such as those used for generating inter-predicted samples. The motion parameters can include a motion vector, a reference picture index, a reference picture list use index, and / or additional information. The motion parameters can be signaled explicitly or implicitly. If the CU is coded in skip mode, the CU can be associated with one PU, and significant residual coefficients, coded motion vector deltas, and / or reference picture indices may not be required. When the CU is coded in merge mode, the motion parameters of the CU can be obtained from neighboring CUs. The neighboring CUs can include spatial and temporal candidates, and additional schedules (or additional candidates) as introduced in VVC. The merge mode can be applied not only to skip mode but also to inter-predicted CUs. An alternative to the merge mode is the explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index for each reference picture list, the reference picture list use flag, and / or other necessary information can be signaled explicitly for each CU.
[0124] In VVC, the VVC Test model (VTM) reference software can include a number of improved new inter prediction coding tools, including one or more of the following. (1) Extended merge prediction (2) Merge motion vector difference (MMVD) (3) AMVP mode using symmetric MVD signaling (4) Affine motion compensation prediction (5) Subblock-Based Temporal Motion Vector Prediction (SbTMVP) (6) Adaptive motion vector resolution (AMVR) (7) Motion field storage: 1 / 16 luma sample MV storage and 8×8 motion field compression (8) Bi-prediction with CU-level weights (BCW) (9) Bi-directional optical flow (BDOF) (10) Decoder side motion vector refinement (DMVR) (11) Combined Inter-Intra Prediction (CIIP) (12) Geometric partitioning mode (GPM).
[0125] The merge candidate list can be constructed by including five types of candidates such as those in VTM4. The merge candidate list can be constructed in the following order: (1) Spatial MVP from spatially neighboring CUs (2) Temporal MVP from CUs at the same position (3) MVP based on history from the FIFO table (4) Pairwise average MVP (5) Zero MV.
[0126] The size of the merge candidate list can be signaled in the slice header. The maximum allowable size of the merge list is 6 as in VTM4. For each CU coded in the merge mode, the index of the best merge candidate can be coded using, for example, truncated unary binarization. The first bin of the merge index can be coded in context and bypass coding can be used for the other bins.
[0127] For example, in VVC, in spatial candidate derivation, the derivation of spatial merge candidates can be the same as or similar to the derivation of spatial merge candidates in HEVC. For example, from among the candidates located at the positions shown in FIG. 9, the maximum number of merge candidates (e.g., four merge candidates) can be selected. As shown in FIG. 9, the current block (901) is at position A 0 、A 1 、B 0 、B 1 、and B 2 can include neighboring blocks (902) to (906) located at each of them respectively. The derivation order of spatial merge candidates can be B 1 、A 1 、B 0 、A 0 and B 2 . Position B 2 can be considered only when any of the CUs (or blocks) at position A 0 、B 0 、B 1 、or A 1 is not available (e.g., because the CU belongs to another slice or tile), or when it is intra-coded. After the candidate (or block) at position A 1 is added, the addition of the remaining candidates (or blocks) can be subject to redundancy checking. Redundancy checking can ensure that candidates with the same motion information are excluded from the merge list so as to improve coding efficiency. To reduce the computational complexity, redundancy checking may not consider all possible candidate pairs. Instead, only the candidate pairs linked by the arrows in FIG. 10 need to be considered. For example, redundancy checking can be applied to five candidate pairs such as the candidate pair of A1 and B1 and the candidate pair of A1 and A0. A candidate may be added to the merge list only if the corresponding candidates used for redundancy checking do not contain the same motion information. For example, candidate B0 may be added to the merge list only if the corresponding candidate B1 does not contain the same motion information.
[0128] In the derivation of temporal candidates, only one candidate may be added to the merge list. For example, as shown in FIG. 11, in the derivation of the temporal merge candidates for the current CU (1114), scaled motion vectors can be derived based on the same-position CUs (1104) belonging to the same-position reference picture (1112). The reference picture list used to derive the same-position CUs (1104) can be explicitly signaled in the slice header. The scaled motion vectors for the temporal merge candidates can be obtained as shown by the dotted line (1102) in FIG. 11, which is scaled from the motion vectors of the same-position CUs (1104) using the picture order count (POC) distances tb and td. tb can be defined as the POC difference between the reference picture (e.g., Curr_ref) (1106) of the current picture and the current picture (e.g., Curr_pic) (1108). td can be defined as the POC difference between the reference picture (e.g., Col_ref) (1110) of the same-position picture and the same-position picture (e.g., Col_pic) (1112). The reference picture index of the temporal merge candidate can be set to 0.
[0129] As shown in FIG. 12, the position of the temporal merge candidate is candidate C 0 and C 1 can be selected between. For example, if the CU at position C 0 is not available, is intra-coded, or is outside the current row of the CTU, position C 1 can be used. Otherwise, position C 0 can be used for the derivation of the temporal merge candidate.
[0130] Merge with Motion Vector Difference (MMVD) can be used for specific prediction modes, such as either the skip mode or the merge mode using a motion vector representation method. MMVD can reuse merge candidates such as those in VVC. A merge candidate can be selected from among the merge candidates and further extended (or subdivided) by a motion vector representation method. MMVD can provide a new motion vector representation with simplified signaling. The motion vector representation method can include a starting point, a magnitude of motion, and a direction of motion.
[0131] MMVD can utilize a merge candidate list such as that in VVC. Candidates with a default merge type (e.g., MRG_TYPE_DEFAULT_N) can consider MMVD extension. In MMVD, the starting point can be defined with a base candidate index. The base candidate index (IDX) can indicate the best candidate among the candidates in the list, for example, as shown in Table 1. [Table 1] Base candidate IDX [Table 1]
[0132] When the number of base candidates is equal to 1, the base candidate IDX may not be signaled. The distance index can provide information on the magnitude of motion. The distance index can indicate a predefined distance from the starting point. The predefined distance based on the distance index can be provided in Table 2 as follows: [Table 2] Distance IDX [Table 2]
[0133] The direction index can represent the direction of the MVD based on the starting point. The direction index can represent four directions as shown in Table 3. The MMVD flag can be signaled when the skip flag and the merge flag are sent. When the skip flag and the merge flag are true, the MMVD flag can be parsed. When the MMVD flag is 1, the MMVD syntax can be parsed. However, when the MMVD flag is not 1, the AFFINE flag can be parsed. When the AFFINE flag is 1, the AFFINE mode can be applied. However, when the AFFINE flag is not 1, the skip / merge index can be parsed for the skip / merge mode. [Table 3] Direction IDX [Table 3]
[0134] Figure 13 shows an exemplary search process for MMVD. As shown in Figure 13, in Figure 13, the starting point MV can be indicated by (1311) (e.g., according to the direction IDX and the basic candidate IDX), the offset can be indicated by (1312) (e.g., according to the distance IDX and the direction IDX), and the final MV predictor is indicated by (1313). In another example, in Figure 13, the starting point MV can be indicated by (1321) (e.g., according to the direction IDX and the basic candidate IDX), the offset can be indicated by (1322) (e.g., according to the distance IDX and the direction IDX), and the final MV predictor can be indicated by (1323).
[0135] Figures 14A and 14B show exemplary search points of MMVD. As shown in Figure 14A, in the first reference list L0, the starting point MV is indicated at (1411) (e.g., according to the direction IDX and the basic candidate IDX). In the example of Figure 14A, four search directions such as +Y, -Y, +X, and -X are used, and the four search directions are indexed by 0, 1, 2, 3. The distance can be indexed by 0 (distance 0 to the starting point MV), 1 (1s (or 1 sample) to the starting point MV), 2 (2s to the starting point MV), 3 (3s to the starting point MV), etc. Thus, when the direction IDX is 3 and the distance IDX is 2, the final MV predictor is shown as (1415).
[0136] In another example, the search direction and the distance can be combined for indexing. For example, in the second reference list L1, the starting point MV is indicated at (1421) (e.g., according to the direction IDX and the basic candidate IDX). The search direction and the distance are combined to be indexed by 0 to 12 as shown in Figure 14B.
[0137] In candidate reordering based on template matching on MMVD and affine MMVD, the MMVD offset can be extended for the MMVD mode and the affine MMVD mode. In one example, additional subdivision positions along the diagonal angle of k×π / 8 can be added first. Exemplary additional subdivision positions can be shown in Figure 15, and the number of directions can be increased from 4 to 16. Next, based on the SAD cost between the template (e.g., one row above and one column to the left of the current block) and the reference of the template for each subdivision position, all possible MMVD subdivision positions (e.g., 16×6) for each basic candidate can be reordered. Finally, the top 1 / 8 of the subdivision positions with the minimum template SAD cost can be retained as available positions, and as a result, can be retained for MMVD index coding. The MMVD index can be binarized by a code such as a rice code with a parameter equal to 2.
[0138] In another example, on top of the MMVD extension as described above, the affine MMVD rearrangement can also be extended, where additional subdivision positions along the diagonal angle of k×π / 4 can be added. After the rearrangement, the top 1 / 2 of the subdivision positions with the minimum template SAD cost can be retained.
[0139] To improve coding efficiency and reduce the transmission overhead of motion vectors, sub-block level motion vector subdivision can be applied to extend the temporal motion vector prediction (TMVP) at the CU level. Subblock-based TMVP (SbTMVP) can inherit motion information at the sub-block level from the same-position reference picture (or the reference picture at the same position as the current picture). Each sub-block of a large-sized CU can have its own motion information without explicitly transmitting the block partition structure or motion information. SbTMVP can obtain the motion information of each sub-block in three steps. The first step can include the derivation of the displacement vector (DV) of the current CU. In the second step, the availability of SbTMVP candidates can be checked and the central motion can be derived. In the third step, the motion information of the sub-block can be derived from the corresponding sub-block indicated by the DV. Different from the derivation of TMVP candidates that derive the temporal motion vector from the same-position block within the reference frame, SbTMVP can apply the DV derived from the MV of the left neighboring CU of the current CU to find the corresponding sub-block within the same-position picture of each sub-block of the current CU. If the corresponding sub-block is not inter-coded, the motion information of the current sub-block can be set as the central motion.
[0140] SbTMVP can be supported by related coding standards such as VVC. For example, similar to TMVP in HEVC, SbTMVP can use motion fields within the same picture at the same position of the current picture to improve the motion vector prediction (e.g., merge mode) of CUs within the current picture. The same-position picture used by TMVP can also be used by SbTMVP. SbTMVP may be different from TMVP in one or more aspects as follows: (1) TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level. (2) TMVP fetches the temporal motion vector from the same-position block within the same-position picture (e.g., the same-position block can be the block at the lower right or center with respect to the current CU). SbTMVP applies a motion shift before fetching the temporal motion information from the same-position picture. Here, the motion shift can be obtained from the motion vector of one of the spatial neighboring blocks of the current CU.
[0141] Exemplary spatial neighboring blocks applied to SbTMVP can be shown in FIG. 16. As shown in FIG. 16, SbTMVP can predict the motion vectors of sub-CUs (not shown) within the current CU (1602) in two steps. In the first step, the spatial neighbor A1 (1604) in FIG. 16 can be examined. If A1 (1604) has a motion vector using the same-position picture of the current picture as the reference picture, the motion vector of A1 (1604) can be selected as the motion shift (or displacement vector) of SbTMVP, and the corresponding sub-block within the same-position picture can be found for each sub-block of the current CU. If such a motion vector is not identified, the motion shift can be set to (0, 0).
[0142] In the second stage, the motion shift identified in the first stage is applied (e.g., added to the coordinates of the current CU) to obtain motion information at the sub-CU level (e.g., the motion vector and reference index at the sub-CU level) from the same-position picture. As shown in FIG. 17, the current CU (1704) can be included in the current picture (1702). The current CU (1704) can include a plurality of sub-CUs (or sub-blocks) such as the sub-CU (1706). The neighboring block A1 (1708) can be located at the lower left of the current CU (1704). In the example of FIG. 17, a motion shift (or DV) (1710) can be set as the motion vector of the neighboring block A1 (1708). According to the DV (1710), the reference block A1' (1718) of the neighboring block A1 (1708) can be determined. The reference block (1714) adjacent to the reference block A1' (1718) can be determined as the reference block of the current block (1704). For each sub-CU (e.g., (1706)) of the current block (1704), in the same-position picture (1712) of the current picture (1702), using the motion information of the corresponding block (or corresponding sub-CU) (e.g., (1716)) in the reference block (1714), which can be the smallest motion grid covering the central sample of the corresponding block, the motion information of each sub-CU can be derived. After the motion information of the same-position sub-CU (e.g., (1716)) is identified, the motion information can be converted into the motion vector and reference index of the current sub-CU (e.g., (1706)). The motion information can be converted in a manner similar to the TMVP process of HEVC, where temporal motion scaling can be applied to align the temporal motion vector of the reference picture with the temporal motion vector of the current CU.
[0143] A merge list based on a combined sub-block that includes both SbTMVP candidates and affine merge candidates can be used to signal a merge mode (e.g., the SbTMVP mode) based on sub-blocks such as in VVC. The SbTMVP mode can be enabled / disabled by a sequence parameter set (SPS) flag. When the SbTMVP mode is enabled, SbTMVP candidates can be added as the first entry in the list of merge candidates based on sub-blocks, followed by affine merge candidates. The size of the merge list based on sub-blocks can be signaled in the SPS, and the maximum allowable size of the merge list based on sub-blocks such as in VVC can be set to 5.
[0144] The sub-CU size used in the SbTMVP mode can be fixed at 8x8. This can be the same as the sub-CU size in the affine merge mode. For example, the SbTMVP mode can only be applied to CUs where both the width and height are 8 (or 8 pixels) or more. The sub-block size of the SbTMVP mode can be configured to other sizes such as 4x4 in the ECM software model used in the search after VVC.
[0145] In related coding standards such as VVC and ECM, the sub-block-based TMVP (or SbTMVP) can be derived based on the DV derived from the MVs of neighboring CUs of the current CU. However, the derived SbTMVP and the derived DV may not be optimal or have an optimal match with each other.
[0146] To derive the sub-block-based TMVP, an additional motion offset of the DV can be signaled. However, signaling the additional motion offset can be costly because additional bits are required.
[0147] In the present disclosure, SbTMVP based on template matching (TM) can be applied. The template of the current block can indicate an area adjacent to the current block, such as a reconstructed area of a predetermined neighborhood of the current block. In one example, the template can include the upper N rows of the neighboring reconstructed samples above the current block and / or the left M columns of the neighboring reconstructed samples to the left of the current block. Examples of the values of M and N can include, but are not limited to, 1, 2, 3, 4, etc.
[0148] According to the TM-based SbTMVP, the SbTMVP can use template matching to derive the DV instead of using the DV derived from the neighboring CUs of the current CU or instead of further signaling an additional motion offset of the DV. The template of the current coding block (or current block) in the current picture can be compared with each of one or more templates (or reference templates) of a plurality of blocks (or a plurality of reference blocks in the same-position reference picture of the current picture) located at the specified candidate positions. The cost value C can be calculated for each candidate position (or each of the reference blocks) and can be associated with the difference between the template of the current CU and the template of each reference block. The motion information in the sub-blocks of the K blocks associated with the K minimum cost values (for example, the K reference blocks in the same-position picture) can be used to derive the SbTMVP (for example, the motion vectors and reference indices of the sub-blocks within the current block). K can be the number of possible DVs.
[0149] Each template of a plurality of reference blocks can include M×N sub-blocks (or regions) adjacent to each of the reference blocks within the same-position picture. The motion vector of each M×N sub-block adjacent to a block (or a reference block within the same-position picture) and associated with a reference template can be derived from the MV field of the same-position block (or the same-position reference block) of the current block. The same-position block can be derived by one of K displacement motion vectors (DVs) having a given additional motion vector offset (MVO). The derived MV, SbMV(i, j) of the sub-block of the current block Lx can be used to indicate the position of the sub-block template in the reference frame (or the same-position picture).
[0150] FIG. 18 can show an example of SbTMVP based on TM. As shown in FIG. 18, the current block (1802) can be included in the current picture (1804). The current block (1802) can include a template Tc (1806). The template Tc (1806) can include neighboring samples adjacent to the upper side and / or the left side of the current block (1802). In the search region of the reference picture (1822), a plurality of candidate reference blocks can be determined. A plurality of cost values can be determined based on template matching between the template (1806) of the current block (1802) and the templates of the plurality of candidate reference blocks. Each of the plurality of cost values can be based on the difference between the template of the current block and the template of each candidate reference block. One or more reference blocks can be selected from the plurality of candidate reference blocks corresponding to one or more of the lowest cost values among the plurality of cost values. The selected one or more candidate reference blocks can be indicated by one or more DVs. For example, templates T 0 (1820) and template T 1 (1816) can be selected. The candidate reference block (1818) is the template T0 (1820), and template Tc to template T 0 Domestic violence against 0 (1812). The candidate reference block (1814) can be represented by the template T 1 (1816), and template Tc to template T 1 Domestic violence against 1 (1810). Further, motion vectors of the sub-blocks of the candidate reference blocks (1814) and (1818) can be used in TMVP for the sub-blocks of the current block (1802). For example, the motion information of the sub-block (1816) of the candidate reference block (1814) and the motion information of the sub-block (1824) of the candidate reference block (1818) can be transformed into a motion vector and / or a reference index of the sub-block (1808) of the current block (1802), e.g., based on temporal motion scaling. A prediction sample for the sub-block (1808) can further be determined based on the motion vector and / or the reference index.
[0151] In one embodiment, the cost value C can be determined based on the difference between a template of the current block (e.g., (1802)) and a template of each candidate reference block (e.g., (1818)). In one example, the cost value C can be one of the sum of absolute difference (SAD), sum of absolute transformed difference (SATD), sum of squares error (SSE), subsampled SAD, and mean-removed SAD.
[0152] In one embodiment, the selected one or more candidate reference blocks are 1 (1810) or DV 0 (1812)), where k can be, but is not limited to, 1, 2, 3, 4, etc.
[0153] In one embodiment, a search region for finding one or more candidate reference blocks (e.g., candidate reference blocks (1814) and (1818)) can be defined and centered at a position that is at the same position as the current block within the reference picture / frame. For example, the search region can be determined as a region centered at a position that is at the same position as the current block (1802) within the reference picture (1822).
[0154] In one embodiment, the search region can be determined from the coded region of the same picture (e.g., the current picture). For example, the search region can be determined as a region centered on the current block (1802) within the current picture (1804).
[0155] In one embodiment, candidate positions (or candidate reference blocks) can be specified by a search region identified by a displacement vector. Thus, the search region can be identified by a first DV (not shown). Each candidate position (or candidate reference block) within the search region can be identified by a respective second DV. For example, candidate reference block (1814) can be indicated by DV 1 (1810).
[0156] In one embodiment, the DV can be derived from the motion vectors of spatially neighboring coded blocks of the current block within the current picture. For example, the DV can be derived based on the MVP list of the current block, where the neighboring coded blocks can be candidates in the MVP list of the current block.
[0157] In one embodiment, the displacement vector can be derived from the motion vectors of the normal (e.g., no sub-blocks) merge candidate list of the current block within the current picture. Thus, temporal motion vector predictor (TMVP) candidates can be excluded from the normal merge candidate list.
[0158] In one embodiment, the search region can be a specific region range (or specific region) centered on a sample indicated (or directed) by a displacement vector. Examples of the specific region range (or specific region) can include, but are not limited to, a square region, a rectangular region, a rhombus region, and the like.
[0159] In one embodiment, the search region can be a sample group centered on a sample indicated (or directed) by a displacement vector.
[0160] In one example, the sample group can be located at the same position as the position of the MVD search point, such as the search points shown in FIGS. 14A and 14B for MMVD.
[0161] In one example, the sample group can include samples located in the horizontal direction (or 0-degree direction), vertical direction (or 90-degree direction), 45-degree direction, or 135-degree direction with respect to the sample indicated by the displacement vector.
[0162] In one embodiment, the DV derivation based on TM (or TM-based SbTMVP) can be used in combination with another method (e.g., the inter prediction mode or the affine mode). Signaling information such as a flag can be signaled or implicitly derived to indicate whether the DV is derived by template matching or signaled by another method.
[0163] In one embodiment, the DV derivation based on TM can be used in combination with another method (e.g., the inter prediction mode or the affine mode). For example, the first DV can be derived by another method, and the second DV can be derived by template matching in the search region identified by the first DV.
[0164] When a plurality of blocks (e.g., K is greater than 1) are identified, for each sub-block (e.g., (1808)) of the current coding block (e.g., (1802)), a plurality of predicted blocks (or a plurality of predicted sub-blocks) can be derived using the motion vectors associated with the sub-blocks at the same position within the plurality of K blocks, and motion compensation can be performed based on a combination such as a weighted sum of the plurality of predicted blocks. For example, as shown in FIG. 18, candidate reference blocks (1814) and (1818) can be identified for the current block (1802) based on SbTMVP based on TM. The motion vectors of sub-block (1826) within candidate reference block (1814) and sub-block (1824) within candidate reference block (1818) can be applied to derive the first predicted sub-block and the second predicted sub-block of sub-block (1808) within the current block (1802). The predicted samples of sub-block (1808) can be determined based on the weighted sum of the first predicted sub-block and the second predicted sub-block.
[0165] In one embodiment, the derived DVs up to a maximum of S with the lowest template matching cost can be used as DV candidates for signaling. S can be, but is not limited to, 1, 2, 3, 4, etc. Thus, the maximum codeword for signaling the DV candidates can be restricted. For example, the encoder can signal up to a maximum of S DV candidates, and the decoder selects one or more of the signaled SDV candidates for use.
[0166] In one embodiment, the best (or selected) DV having the lowest template matching cost can be used. Thus, no additional signaling is required to signal the best DV. On the decoder side, the decoder can perform the same template matching process as the encoder to derive the best (or selected) DV.
[0167] FIG. 19 shows a flowchart illustrating an overview of an exemplary decoding process (1900) according to some embodiments of the present disclosure. FIG. 20 shows a flowchart illustrating an overview of an exemplary encoding process (2000) according to some embodiments of the present disclosure. The proposed processes may be used separately or combined in any order. Further, each of the processes (or embodiments), encoders, and decoders may be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0168] The operations of the processes (e.g., (1900) and (2000)) can be combined or arranged in any amount or order as needed. In embodiments, two or more of any of the operations of the processes (e.g., (1900) and (2000)) can be executed in parallel.
[0169] The processes (e.g., (1900) and (2000)) can be used in the reconstruction and / or encoding of blocks to generate predicted blocks for the blocks being reconstructed. In various embodiments, the processes (e.g., (1900) and (2000)) are executed by processing circuits such as the processing circuits within the terminal devices (310), (320), (330), and (340), the processing circuit that executes the functions of the video encoder (403), the processing circuit that executes the functions of the video decoder (410), the processing circuit that executes the functions of the video decoder (510), the processing circuit that executes the functions of the video encoder (603), and the like. In some embodiments, the processes (e.g., (1900) and (2000)) are implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the processes (e.g., (1900) and (2000)).
[0170] As shown in FIG. 19, process (1900) starts from (S1901) and can proceed to (S1910). At (S1910), a coded video bitstream including the current block within the current picture is received. The current block includes a plurality of sub-blocks and is predicted by the SbTMVP (subblock-based template matching motion vector prediction) mode.
[0171] At (S1920), the same-position reference sub-blocks of each of the sub-blocks are determined based on a combination of the displacement vector (DV) and the motion vector offset (MVO) associated with each of the sub-blocks.
[0172] At (S1930), the motion vector (MV) field at the same-position reference sub-block of each of the sub-blocks within the current block is determined.
[0173] At (S1940), the reference template of each of the sub-blocks is derived based on the determined MV field of the same-position reference sub-block.
[0174] At (S1950), in the SbTMVP mode, the plurality of sub-blocks of the current block are reconstructed by predicting each of the sub-blocks using each of the reference templates.
[0175] To determine each identical-position reference sub-block, a search region located in one of the current picture and the reference pictures of the current picture is determined. One or more reference blocks of the current block are determined based on template matching between the template of the current block and the templates of each of one or more reference blocks within the search region. The template of the current block includes samples adjacent to the current block. The template of each of the one or more reference blocks includes samples adjacent to each of the one or more reference blocks. Each identical-position reference sub-block of each sub-block can be determined as a sub-block that is in the same position as each sub-block in one of the one or more reference blocks.
[0176] The template matching between the template of the current picture and the templates of each of the one or more reference blocks is determined based on one of SAD, SATD, SSE, sub-sampled SAD, and mean-removed SAD.
[0177] To determine one or more reference blocks, a plurality of candidate reference blocks are determined within the search region. A plurality of cost values are determined based on template matching between the template of the current block and the templates of the plurality of candidate reference blocks. One or more reference blocks are determined as one or more of the candidate reference blocks corresponding to one or more of the lowest cost values among the plurality of cost values.
[0178] In one embodiment, the search region includes one of (i) a region centered at a position that is in the same position as the current block in the reference picture, and (ii) a region centered at the current block in the current picture.
[0179] In one example, the search region can be determined based on DV. DV can be derived from one of (i) the motion vectors of blocks spatially adjacent to the current block, and (ii) the motion vectors of the merge candidate list of the current block.
[0180] In one example, the search area is determined as an area centered on the sample indicated by the DV, and the area is one of a square, a rectangle, and a rhombus.
[0181] In one example, the search area is determined as a sample group centered on the sample indicated by the DV. The sample group can be located at at least one of 0 degrees, 45 degrees, 90 degrees, or 135 degrees with respect to the sample indicated by the DV.
[0182] To determine one or more reference blocks, a first reference block of the one or more reference blocks is determined. The first reference block is indicated by a first DV from the template of the current block to the template of the first reference block. In one example, the first DV is derived based on template matching such that the first DV corresponds to a cost value related to the difference between the template of the first reference block and the template of the current block. In one example, the first DV is signaled.
[0183] The search area is determined based on a first DV derived before template matching. The first reference block among the one or more reference blocks is determined based on a second DV from the template of the current block to the template of the first reference block. The second DV is derived based on template matching such that the second DV corresponds to a cost value related to the difference between the template of the first reference block and the template of the current block.
[0184] To reconstruct sub - blocks of the current block, one or more motion vectors (MVs) of a first sub - block among a plurality of sub - blocks in the current block are determined based on one or more MVs of a sub - block at the same position as the first sub - block in one or more reference blocks. One or more predicted sub - blocks of the first sub - block among the plurality of sub - blocks are determined based on one or more MVs of the first sub - block. Predicted samples of the first sub - block are determined based on one of or a weighted combination of one or more of the predicted sub - blocks.
[0185] In some embodiments, a plurality of candidate reference blocks of the current block are determined based on a plurality of DVs. Each of the plurality of candidate reference blocks is indicated by a DV of each of the plurality of DVs. One or more reference blocks of the current block are determined from the plurality of candidate reference blocks based on one or more cost values of template matching.
[0186] After (S1940), the process proceeds to (S1999) and ends.
[0187] The process (1900) can be appropriately adapted. The steps of the process (1900) can be changed and / or omitted. Additional steps can be added. Any suitable implementation order can be used.
[0188] As shown in FIG. 20, the process (2000) starts from (S2001) and can proceed to (S2010). At (S2010), a search area located in one of the current picture and a reference picture of the current picture is determined.
[0189] At (S2020), one or more reference blocks of the current block are determined based on template matching between a template of the current block and templates of each of one or more reference blocks in the search area. The template of the current block includes samples adjacent to the current block. The template of each of the one or more reference blocks includes samples adjacent to each of the one or more reference blocks.
[0190] (S2030), based on the sub-blocks of the determined one or more reference blocks, a predicted sample of the sub-blocks of the current block is generated.
[0191] Next, the process proceeds to (S2099) and ends.
[0192] The process (2000) can be appropriately adapted. The steps of the process (2000) can be changed and / or omitted. Additional steps can be added. Any appropriate implementation order can be used.
[0193] The above-described technology can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, FIG. 21 shows a computer system (2100) suitable for implementing a particular embodiment of the subject matter of the present disclosure.
[0194] The computer software can be processed by mechanisms such as assembly, compilation, linking, etc. to generate code containing instructions executable directly or through interpretation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. It can be coded using any appropriate machine code or computer language.
[0195] The instructions can be executed on various computers or their components including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.
[0196] The components shown in FIG. 21 of the computer system (2100) are exemplary in nature and do not imply any limitation to the use or functionality of the computer software implementing the embodiments of the present disclosure. Further, the component configuration should not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system (2100).
[0197] The computer system (2100) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users through, for example, sensory input (e.g., keystrokes, swipes, data glove movements), voice input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). The human interface device can also be used to capture certain media that is not necessarily directly related to conscious input by a human, such as voice (e.g., conversation, music, ambient sound), images (e.g., scanned images, photographic images obtained from a digital camera), video (e.g., including 2D video, 3D video, stereoscopic video).
[0198] The input human interface device may include one or more of a keyboard (2101), a mouse (2102), a trackpad (2103), a touch screen (2110), a data glove (not shown), a joystick (2105), a microphone (2106), a scanner (2107), a camera (2108) (only one of which is shown).
[0199] The computer system (2100) may also include a specific human interface output device. Such a human interface output device may stimulate the senses of one or more human users, for example, through sensory output, sound, light, and smell / taste. Such a human interface output device may include a sensory output device (e.g., a touch screen (2110), a data glove (not shown), or a joystick (2105) for sensory feedback, although there may also be sensory feedback devices that do not function as input devices), a voice output device (e.g., a speaker (2109), headphones (not shown), a visual output device (e.g., a screen (2110), a CRT screen, an LCD screen, a plasma screen, an OLED screen, each with or without touch screen input capabilities, each with or without sensory feedback capabilities, some of which may output 2D visual output or output of three or more dimensions through means such as stereoscopic output, virtual reality glasses (not shown), a holographic display, and a smoke agent tank (not shown), and a printer (not shown))).
[0200] The computer system (2100) may also include a human-accessible storage device and related media such as an optical medium like a CD / DVDROM / RW (2120) with a medium (2121) such as a CD / DVD, a thumb drive (2122), a removable hard drive or a solid-state drive (2123), legacy magnetic media such as tapes and floppy disks (not shown), and devices based on dedicated ROM / ASIC / PLD such as a security dongle (not shown).
[0201] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.
[0202] The computer system (2100) may also include an interface (2154) to one or more communication networks (2155). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicle and industrial including CANBus, etc. A particular network generally requires an external network interface attached to a particular general-purpose data port or peripheral device bus (2149) (e.g., the USB port of the computer system (2100)). Others are generally integrated into the core of the computer system (2100) by attachment to a system bus as described later (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using these networks, the computer system (2100) can communicate with other entities. Such communication can be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CANbus to a particular CANbus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Particular protocols and protocol stacks can be used for each of the networks and network interfaces described above.
[0203] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (2140) of the computer system (2100).
[0204] The core (2140) may include one or more central processing units (CPUs) (2141), a graphics processing unit (GPU) (2142), a dedicated programmable processing unit in the form of an FPGA (2143), a hardware accelerator for specific tasks (2144), a graphics adapter (2150), etc. These devices may be connected through a system bus (2148) together with a built-in large-capacity storage device (2147) such as a read-only memory (ROM) (2145), a random access memory (2146), an internal hard drive that is not accessible to users, an SSD, etc. In some computer systems, the system bus (2148) is accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (2148) of the core or through a peripheral device bus (2149). In the example, the screen (2110) can be connected to the graphics adapter (2150). The architecture of the peripheral device bus includes PCI, USB, etc.
[0205] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) can be combined to execute specific instructions capable of generating the aforementioned computer code. The computer code can be stored in the ROM (2145) or the RAM (2146). Temporary data can also be stored in the RAM (2146), while permanent data can be stored, for example, in the built-in large-capacity storage device (2147). Fast storage and reading to any of the memory devices can be enabled through the use of a cache memory that can be closely associated with one or more of the CPU (2141), GPU (2142), large-capacity storage device (2147), ROM (2145), RAM (2146), etc.
[0206] The computer-readable medium may have computer code for executing operations performed by various computers. The medium and the computer code may be specially designed and configured for the purposes of the present disclosure, or may be of the types well known and available to those skilled in the field of computer software.
[0207] By way of example and not limitation, a computer system (2100) having an architecture, and specifically a core (2140), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be associated with a specific storage device of the core (2140) having non-transitory characteristics such as an on-core mass storage device (2147) or ROM (2145), and media associated with a user-accessible mass storage device as described above. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2140). The computer-readable media can include one or more memory devices or chips as required by a particular need. The software can cause the core (2140) and specifically the processor (including a CPU, GPU, FPGA, etc.) therein to execute a specific process or a specific portion of a specific process described herein, including the definition of a data structure stored in a RAM (2146) and the modification of the data structure according to the process defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of an implementation in a logic hardwired or other circuitry (e.g., an accelerator (2144)) operable with or instead of the software to execute a specific process or a specific portion of a specific process described herein. References to software include logic and, where appropriate, vice versa. References to computer-readable media can include, where appropriate, circuitry (such as an integrated circuit (IC)) for storing software for execution, circuitry for implementing logic for execution, or both. The present disclosure includes any suitable combination of hardware and software. Appendix A: Glossary JEM: joint exploration model VVC: versatile video coding BMS: benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOPs: Groups of Pictures TUs: Transform Units, PUs: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Areas SSD: solid-state drive IC: Integrated Circuit CU: Coding Unit
[0208] Although the present disclosure has described several exemplary embodiments, there are alternatives, substitutions, and various equivalents thereof, which are encompassed within the scope of the present disclosure. It will be apparent to those skilled in the art that many systems and methods, although not explicitly shown or described herein, can be devised that implement the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
Claim 1 A method of decoding executed in a decoder, the method comprising: receiving a coded video bitstream including a current block within a current picture, the current block including a plurality of sub-blocks and being predicted by a SbTMVP (subblock-based template matching motion vector prediction) mode; determining, for each sub-block, a co-located reference sub-block at the same position of each sub-block based on a combination of a displacement vector (DV) and a motion vector offset (MVO) associated with each sub-block; determining a motion vector (MV) field within the co-located reference sub-block at the same position of each sub-block within the current block; deriving, for each sub-block, a reference template based on the determined MV field of the co-located reference sub-block; reconstructing the plurality of sub-blocks of the current block by predicting each sub-block using the respective reference template in the SbTMVP mode; A method comprising the above steps. Claim 2 The step of determining each co-located reference sub-block comprises: determining a search region located in one of the current picture and a reference picture of the current picture; determining, based on template matching between a template of the current block and templates of one or more reference blocks within the search region, the one or more reference blocks of the current block, wherein the template of the current block includes samples adjacent to the current block and the template of each of the one or more reference blocks includes samples adjacent to the respective reference block of the one or more reference blocks; determining, for each sub-block, the respective co-located reference sub-block as a sub-block at the same position as the sub-block within one of the one or more reference blocks; The method according to claim 1, further comprising the above steps. Claim 3 The method according to claim 2, wherein the template matching between the template of the current block and the template of each of the one or more reference blocks is determined based on one of sum of absolute differences (SAD), sum of absolute transformed differences (SATD), sum of squared errors (SSE), subsampled SAD, and mean removed SAD.
4. The step of determining the one or more reference blocks comprises: determining a plurality of candidate reference blocks within the search region; determining a plurality of cost values based on template matching between the template of the current block and the templates of the plurality of candidate reference blocks; determining the one or more reference blocks as one or more of the candidate reference blocks corresponding to one or more of the lowest cost values among the plurality of cost values; The method according to claim 3, further comprising.
5. The search region includes one of (i) a region centered at a position that is in the same position as the current block in the reference picture, and (ii) a region centered at the current block in the current picture. The method according to claim 2.
6. The step of determining the search region comprises: determining the search region based on a displacement vector (DV), wherein the DV is derived from one of (i) motion vectors of spatially neighboring blocks of the current block, and (ii) motion vectors of a merge candidate list of the current block; The method according to claim 2, further comprising.
7. The step of determining the search region comprises: determining the search region as a region centered at a sample indicated by the DV, wherein the region is one of a square, a rectangle, and a rhombus; The method according to claim 6, further comprising.
8. The step of determining the search region comprises: determining the search region as a sample group centered at a sample indicated by the DV, wherein the sample group is located at at least one of 0 degrees, 45 degrees, 90 degrees, or 135 degrees with respect to the sample indicated by the DV; The method according to claim 6, further comprising.
9. The step of determining the one or more reference blocks comprises: Determining the first reference block among the one or more reference blocks indicated by a first displacement vector (DV) from the template of the current block to the template of the first reference block, wherein the first DV is (i) derived based on template matching such that the first DV corresponds to a cost value related to the difference between the template of the first reference block and the template of the current block, and (ii) signaled, The method according to claim 2, further comprising.
10. The step of determining the search area comprises Further comprising determining the search area based on a first displacement vector (DV) derived prior to the template matching, The step of determining the one or more reference blocks comprises Determining the first reference block among the one or more reference blocks indicated by a second DV from the template of the current block to the template of the first reference block, wherein the second DV is derived based on template matching such that the second DV corresponds to a cost value related to the difference between the template of the first reference block and the template of the current block, further comprising the step of The method according to claim 2.
11. The step of reconstructing the plurality of sub-blocks of the current block comprises Determining one or more motion vectors (MVs) of a first sub-block among the plurality of sub-blocks in the current block based on one or more MVs of sub-blocks at the same position as the first sub-block in the one or more reference blocks, Determining one or more predicted sub-blocks of the first sub-block among the plurality of sub-blocks based on the one or more MVs of the first sub-block, Determining a predicted sample of the first sub-block based on one or more of the one or more predicted sub-blocks or a weighted combination thereof, The method according to claim 2, further comprising.
12. Determining a plurality of candidate reference blocks for the current block based on a plurality of displacement vectors (DVs), each of the plurality of candidate reference blocks being indicated by a DV of each of the plurality of DVs, determining, based on one or more cost values of the template matching, one or more reference blocks of the current block from the plurality of candidate reference blocks; The method according to claim 2, further comprising.
13. An apparatus, comprising: a processing circuit, the processing circuit: receives a coded video bitstream including a current block in a current picture, the current block includes a plurality of sub-blocks, and is predicted by a SbTMVP (subblock-based template matching motion vector prediction) mode; determines, for each of the sub-blocks, a same-position reference sub-block of each of the sub-blocks based on a combination of a displacement vector (DV) and a motion vector offset (MVO) associated with each sub-block; determines a motion vector (MV) field in the same-position reference sub-block of each of the sub-blocks within the current block; derives, for each of the sub-blocks, a reference template based on the determined MV field of the same-position reference sub-block; reconstructs the plurality of sub-blocks of the current block by predicting each sub-block using each reference template in the SbTMVP mode; An apparatus configured as such.
14. The processing circuit: determines a search region located in one of the current picture and a reference picture of the current picture; determines the one or more reference blocks of the current block based on template matching between a template of the current block and templates of each of the one or more reference blocks in the search region, the template of the current block includes samples adjacent to the current block, and the template of each of the one or more reference blocks includes samples adjacent to the reference block of each of the one or more reference blocks; determines, for each of the sub-blocks, the respective same-position reference sub-blocks of each of the sub-blocks as sub-blocks at the same position as each of the sub-blocks in one of the one or more reference blocks; The apparatus according to claim 13, further configured as such.
15. The apparatus according to claim 14, wherein template matching between the template of the current block and the template of each of the one or more reference blocks is determined based on one of sum of absolute differences (SAD), sum of absolute transformed differences (SATD), sum of squared errors (SSE), subsampled SAD, and mean removed SAD.
16. The processing circuit determines a plurality of candidate reference blocks within the search region, determines a plurality of cost values based on template matching between the template of the current block and the templates of the plurality of candidate reference blocks, and determines the one or more reference blocks as one or more candidate reference blocks among the plurality of candidate reference blocks corresponding to one or more of the lowest cost values among the plurality of cost values. The apparatus according to claim 15, further configured as such.
17. The apparatus according to claim 14, wherein the search region includes one of (i) a region centered at a position that is at the same position as the current block in the reference picture, and (ii) a region centered at the current block in the current picture.
18. The processing circuit determines the search region based on a displacement vector (DV), the DV being derived from one of (i) motion vectors of spatially neighboring blocks of the current block, and (ii) motion vectors of a merge candidate list of the current block. The apparatus according to claim 14, further configured as such.
19. The processing circuit determines the search region as a region centered at a sample indicated by the DV, the region being one of a square, a rectangle, and a rhombus. The apparatus according to claim 18, further configured as such.
20. The processing circuit determines the search region as a sample group centered at a sample indicated by the DV, the sample group being located at at least one of 0 degrees, 45 degrees, 90 degrees, or 135 degrees with respect to the sample indicated by the DV. The apparatus according to claim 18, further configured as such.
21. A method of encoding executed in an encoder, the method comprising: generating a coded video bitstream including a current block within a current picture, the current block including a plurality of sub-blocks. The method includes The plurality of sub-blocks of the current block are determining, for each of the sub-blocks, a corresponding reference sub-block at the same position based on a combination of a displacement vector (DV) and a motion vector offset (MVO) associated with each sub-block; determining a motion vector (MV) field within the corresponding reference sub-block at the same position for each of the sub-blocks within the current block; deriving, for each of the sub-blocks, a reference template based on the determined MV field of the corresponding reference sub-block at the same position; predicting each sub-block using each reference template in SbTMVP (subblock-based template matching motion vector prediction) mode; and thereby reconstructed method. **Claim 22** A method of encoding performed in an encoder, the method comprising: determining a search region located in one of a current picture and a reference picture of the current picture; determining the one or more reference blocks of the current block based on template matching between a template of the current picture and templates of each of the one or more reference blocks within the search region, wherein the template of the current picture within the current picture includes samples adjacent to the current picture, and the template of each of the one or more reference blocks includes samples adjacent to each of the one or more reference blocks; generating predicted samples of the sub-blocks of the current block based on the sub-blocks of the determined one or more reference blocks; and a method including the above.
Citation Information
Patent Citations
Method for encoding and / or decoding images on macroblock level using intra-prediction
US20140205008A1
Motion vector derivation apparatus, video decoding apparatus, and video coding apparatus
US20210067798A1
Skipping refinement based on patch similarity in bilinear interpolation based decoder-side motion vector refinement
US20210084328A1
Motion vector refinement search with integer pixel resolution
US20210195232A1
Enhanced decoder side motion vector refinement
US20210314596A1