Flipping modes for chroma and intra-template matching

By employing flip modes and template matching for adjusting sample locations within video blocks, the method addresses inefficiencies in existing video coding technologies, leading to improved compression and reconstruction of video data.

JP2025536385APending Publication Date: 2025-11-05TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025523073
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-23
Filing Date
2023-10-24
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in compressing video data due to limitations in intra-prediction and inter-prediction methods, particularly in handling spatial and temporal redundancies, leading to suboptimal compression and reconstruction of video blocks.

Method used

The method involves adjusting sample locations within video blocks using flip modes, such as vertical and horizontal flips, based on template matching costs, and determining reference blocks for improved reconstruction, along with inheriting luma coding mode information for chroma components to enhance coding efficiency.

Benefits of technology

This approach enhances video encoding and decoding by optimizing block reconstruction, improving compression efficiency and reducing data redundancy, thereby enhancing the quality of reconstructed video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536385000001_ABST
    Figure 2025536385000001_ABST
Patent Text Reader

Abstract

A video bitstream including coding information for a current block in a current picture is received. The coding information indicates that the current block is coded by a flip mode in which sample locations of the current block are adjusted within the current block. A reference block is determined from a plurality of candidate reference blocks in a reconstructed region of the current picture for the current block based on a template matching (TM) cost. The TM cost indicates a difference between a template of the current block and each template of the plurality of candidate reference blocks. A reconstructed block for the current block is determined based on the determined reference block. The current block is reconstructed by adjusting sample locations of the reconstructed block within the reconstructed block based on the flip mode.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Patent Application No. 18 / 382,922, entitled "FLIPPING MODE FOR CHROMA AND INTRA TEMPLATE MATCHING," filed October 23, 2023, which in turn claims the benefit of priority to U.S. Provisional Application No. 63 / 418,944, entitled "FLIPPING MODE FOR Chroma and Intra Template Matching," filed October 24, 2022. The disclosures of the prior applications are incorporated herein by reference in their entireties.

[0002] This disclosure generally describes embodiments related to video coding. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context for the present disclosure. To the extent described in this background section, the inventors' work, as well as aspects of the description that may not be admitted as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.

[0004] Image / video compression can help transmit image / video data across different devices, storage, and networks while minimizing quality loss. In some examples, video codec technology is capable of compressing video based on spatial and temporal redundancy. In one example, a video codec may use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, a video codec may use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention [Means for solving the problem]

[0005] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit.

[0006] According to one aspect of the present disclosure, a method of video decoding is provided. In the method, a video bitstream including coding information for a current block in a current picture is received. The coding information indicates that the current block is coded by a flip mode in which sample locations of the current block are adjusted within the current block. A reference block is determined from multiple candidate reference blocks within a reconstructed region of the current picture for the current block based on a template matching (TM) cost. The TM cost indicates a difference between a template for the current block and each template of the multiple candidate reference blocks. A reconstructed block for the current block is determined based on the determined reference block. The current block is reconstructed by adjusting sample locations of the reconstructed block within the reconstructed block based on the flip mode.

[0007] In one example, the flip mode includes one of (i) a vertical flip mode configured to adjust the locations of samples of the current block such that the top and bottom of the current block are flipped within the current block, and (ii) a horizontal flip mode configured to adjust the locations of samples of the current block such that the left and right parts of the current block are flipped within the current block.

[0008] In one example, first coding information is received from a received video bitstream. The first coding information indicates whether a flip mode should be applied to the reconstructed block. In response to the first coding information indicating that a flip mode should be applied to the reconstructed block, a type of flip mode is determined based on second coding information in the received video bitstream. The current block is reconstructed by adjusting locations of samples of the reconstructed block based on the determined type of flip mode.

[0009] In one example, a plurality of candidate reference blocks are determined within a search area of ​​a reconstructed region of a current picture. A TM cost between a template of the current block and each template of the plurality of candidate reference blocks is determined. A reference block is determined from the plurality of candidate reference blocks corresponding to the smallest TM cost among the TM costs between the template of the current block and the templates of the plurality of candidate reference blocks.

[0010] In one example, the vertical range of the search area is smaller than the horizontal range of the search area based on the flip mode being horizontal flip. In one example, the vertical range of the search area is larger than the horizontal range of the search area based on the flip mode being vertical flip.

[0011] In one example, the determined reference block is further flipped by adjusting the location of the samples of the determined reference block within the determined reference block based on the flip mode, and the reconstruction block is determined based on the flipped reference block.

[0012] According to another aspect of the present disclosure, a method of video decoding is provided. In the method, a video bitstream including a current chroma block and a co-located luma block of the current chroma block in a current picture is determined. It is determined whether a characteristic value of the co-located luma block is greater than a predefined threshold. The characteristic value is associated with one or more pre-defined coding modes to be applied to the co-located luma block. Based on the characteristic value being greater than the pre-defined threshold, chroma coding mode information of the current chroma block is determined based on the pre-defined luma coding mode information of the luma block. The current chroma block is reconstructed based on the determined chroma coding mode information.

[0013] In one example, the characteristic value indicates one of: (i) the number of luma collocated block positions in a collocated luma block coded by one or more predefined coding modes; (ii) the size of the luma area in a collocated luma block coded by one or more predefined coding modes; and (iii) the ratio of the luma area in a collocated luma block coded by one or more predefined coding modes.

[0014] In one example, the luminance collocated block positions include a center position and four corner positions of the collocated luminance block.

[0015] In one example, the one or more predefined coding modes include at least one of an intra block copy (IBC) flipping mode, an IBC mode indicated by a block vector (BV), an IBC rotation mode, and an IBC geometric partition mode.

[0016] In one example, the one or more predefined coding modes include at least one of an intra-template matching prediction (IntraTMP) flipping mode and an IntraTMP mode indicated by a displacement vector.

[0017] In one example, the chroma coding mode information of the current chroma block is determined based on one of (i) luma coding mode information of a first luma block coded by one of one or more predefined coding modes along a predefined scanning order, (ii) luma coding mode information of a most frequently used mode among the luma co-located block positions coded by one or more predefined coding modes, and (iii) a first luma coding mode information of one or more predefined coding modes used more than N times along the predefined scanning order, where N is a positive integer.

[0018] In one example, the chroma coding mode information of the current chroma block is determined based on luma coding mode information of samples at predefined collocated luma locations, where the samples at the predefined collocated luma locations are coded by one of one or more predefined coding modes, and the predefined collocated luma locations include one of a center location and a top-left location of the collocated luma block.

[0019] In one example, multiple candidate reference chroma blocks are determined within a search range in the current picture indicated by a block vector (BV) of a current chroma block based on one or more predefined coding modes, including an IntraTMP mode. The BV of the current chroma block is determined based on one of multiple associated luma blocks. The multiple associated luma blocks include (i) a first IntraTMP-coded block along a predefined scanning order, (ii) a co-located luma block, and (iii) an IntraTMP-coded block corresponding to the first IntraTMP mode information used more than N times along the predefined scanning order. Multiple flip modes are determined for each of the multiple candidate reference chroma blocks. A TM cost between a template of each of the multiple candidate reference chroma blocks and the template of the current chroma block is determined according to each of the multiple flip modes. The flip mode is determined from multiple flip modes corresponding to the smallest TM cost among the TM costs between the template of each of the multiple candidate reference chroma blocks and the template of the current chroma block according to the multiple flip modes. The flip mode of the current chroma block is determined as a flip mode determined from a plurality of flip modes.

[0020] In one example, the BV of the current chroma block is determined as a scaled BV of one of the multiple associated luma blocks based on a scaling factor, where the scaling factor is determined based on a subsampling ratio associated with the current chroma block.

[0021] In one example, the chroma coding mode information of the current chroma block is determined based on the scaled luma coding mode information of the predefined luma block according to a scaling factor, and the scaling factor is determined based on the sampling format of the current chroma block.

[0022] According to another aspect of the present disclosure, an apparatus is provided, the apparatus including a processing circuit, the processing circuit being configured to perform any of the described methods for video decoding / encoding.

[0023] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding.

[0024] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0025] [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100). [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4A] FIG. 1 illustrates an exemplary reference sample for an intra-block copy (IBC) process. [Figure 4B] FIG. 1 illustrates an exemplary reference sample for an intra-block copy (IBC) process. [Figure 4C] FIG. 1 illustrates an exemplary reference sample for an intra-block copy (IBC) process. [Figure 4D] FIG. 1 illustrates an exemplary reference sample for an intra-block copy (IBC) process. [Figure 5] FIG. 1 is a schematic diagram of an exemplary block vector (BV) adjustment based on a horizontal flip. [Figure 6] FIG. 10 is a schematic diagram of an exemplary BV adjustment based on vertical flip. [Figure 7] 1 is a schematic diagram of intra-template matching prediction (IntraTMP). [Figure 8] 1 is a flowchart outlining a first decoding process according to some embodiments of the present disclosure. [Figure 9] 1 is a flowchart outlining a first encoding process according to some embodiments of the present disclosure. [Figure 10] 10 is a flowchart outlining a second decoding process according to some embodiments of the present disclosure. [Figure 11] 10 is a flowchart outlining a second encoding process according to some embodiments of the present disclosure. [Figure 12] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0026] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is an example application for the disclosed subject matter, a video encoder and video decoder in a streaming environment. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0027] The video processing system (100) includes a capture subsystem (113) that may include a video source (101), such as a digital camera, that generates a stream of uncompressed video pictures (102). In one example, the stream of video pictures (102) includes samples taken by the digital camera. The stream of video pictures (102) is depicted with bold lines to emphasize its large amount of data compared to the encoded video data (104) (or coded video bitstream) that may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or coded video bitstream) is depicted with thin lines to emphasize its small amount of data compared to the stream of video pictures (102) that may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems (106) and (108) of FIG. 1, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) within an electronic device (130). The video decoder (110) decodes an input copy (107) of the encoded video data and creates an output stream (111) of video pictures that can be rendered on a display (112) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265.In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in conjunction with VVC.

[0028] It should be noted that the electronic devices (120) and (130) may include other components (not shown). For example, the electronic device (120) may also include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0029] 2 shows an example block diagram of a video decoder (210). The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used in place of the video decoder (110) in the example of FIG. 1.

[0030] The receiver (231) may receive one or more coded video sequences, e.g., included in a bitstream, to be decoded by the video decoder (210). In one embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of the other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (231) may also receive the encoded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver (231) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter, "parser (220)"). In certain applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be external to the video decoder (210) (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder (210), for example, to combat network jitter, plus another buffer memory (215) internal to the video decoder (210), for example, to handle playout timing. When the receiver (231) receives data from a store-and-forward device with sufficient bandwidth and controllability, or from an asynchronous network, the buffer memory (215) may be unnecessary or may be small.For use with best-effort packet networks such as the Internet, the buffer memory (215) may be necessary and may be relatively large, may advantageously be adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (210).

[0031] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories potentially include information used to manage the operation of the video decoder (210) and information for controlling a rendering device, such as a render device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but may be coupled to the electronic device (230), as shown in FIG. 2. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract from the coded video sequence a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. A subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) may also extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0032] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0033] The reconstruction of the symbols (221) can involve several different units, depending on the type of video picture or portion thereof being coded (inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how may be controlled by subgroup control information parsed from the coded video sequence by the parser (220). The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.

[0034] In addition to the functional blocks already described, the video decoder (210) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0035] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients from the parser (220) as well as control information including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. as symbols (221). The scalar / inverse transform unit (251) can output blocks containing sample values ​​that can be input to an aggregator (255).

[0036] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258), for example, buffers a partially reconstructed and / or fully reconstructed current picture. The aggregator (255) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0037] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (253) may access a reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) associated with the block, these samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (253) in the form of symbols (221), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.

[0038] The output samples of the aggregator (255) can be subjected to various loop filtering techniques in a loop filter unit (256). Video compression techniques can include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called a coded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression can also be responsive to meta-information obtained during decoding of a coded picture or previous portion (in decoding order) of the coded video sequence, and to previously reconstructed, loop-filtered sample values.

[0039] The output of the loop filter unit (256) may be a sample stream that may be output to a render device (212) and stored in a reference picture memory (257) for use in future inter-picture prediction.

[0040] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.

[0041] The video decoder (210) can perform decoding operations according to a given video compression technology or standard, such as ITU-T Recommendation H.265. A coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence adheres to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile may select certain tools from among all tools available in the video compression technology or standard as the only tools available under that profile. Compliance may also require that the complexity of the coded video sequence be within a range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by the specification of a hypothetical reference decoder (HRD) and metadata for HRD buffer management signaled within the coded video sequence.

[0042] In one embodiment, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0043] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmitting circuit). The video encoder (303) may be used in place of the video encoder (103) in the example of FIG. 1.

[0044] The video encoder (303) may receive video samples from a video source (301) (which is not part of the electronic device (320) in the example of FIG. 3) that may capture video images to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0045] The video source (301) may provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (301) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. The following discussion focuses on samples.

[0046] According to one embodiment, the video encoder (303) may code and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other required time constraints. Enforcing an appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is operatively coupled to other functional units described below. Coupling is not shown for clarity. Parameters set by the controller (350) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured with other appropriate functions associated with the video encoder (303) optimized for a particular system design.

[0047] In some embodiments, the video encoder (303) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop can include a source coder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to that used by a (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-exact results regardless of the location of the decoder (local or remote), the contents of the reference picture memory (334) are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the decoder would "see" when using prediction during decoding. This basic principle of reference picture synchronization (the resulting drift when synchronization cannot be maintained due to, for example, channel errors) is also used in several related technologies.

[0048] The operation of the "local" decoder (333) may be the same as the operation of a "remote" decoder, such as the video decoder (210), already described in detail above in conjunction with Figure 2. However, with brief reference also to Figure 2, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (345) and parser (220) may be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333).

[0049] In one embodiment, decoder technology, excluding analysis / entropy decoding, present in a decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operation. Descriptions of encoder technology may be omitted, as they are the reverse of the decoder technology described generically. In certain areas, more detailed descriptions are provided below.

[0050] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0051] The local video decoder (333) may decode the coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence may typically be a copy of the source video sequence, with some errors. The local video decoder (333) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (334). In this way, the video encoder (303) can locally store copies of reconstructed reference pictures that have common content with reconstructed reference pictures obtained by the far-end video decoder (without transmission errors).

[0052] The predictor (335) may perform predictive searches for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., which may serve as appropriate predictive references for the new picture. The predictor (335) may operate on sample blocks, pixel block by pixel block, to find appropriate predictive references. In some cases, as determined by search results obtained by the predictor (335), the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (334).

[0053] The controller (350) may manage the coding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0054] The output of all the aforementioned functional units may be entropy coded by an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.

[0055] The transmitter (340) may buffer the coded video sequence created by the entropy coder (345) and prepare it for transmission over a communication channel (360), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (340) may merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0056] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as one of the following picture types:

[0057] Intra-pictures (I-pictures) can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of intra-pictures, including, for example, independent decoder refresh ("IDR") pictures.

[0058] Predictive pictures (P pictures) may be coded and decoded using intra- or inter-prediction, which uses motion vectors and reference indices to predict the sample values ​​of each block.

[0059] Bidirectionally predicted pictures (B pictures) can be coded and decoded using intra- or inter-prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multi-predicted pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0060] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0061] The video encoder (303) may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. In doing so, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0062] In one embodiment, the transmitter (340) may transmit additional data along with the encoded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0063] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture may be coded by a vector called a motion vector. The motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0064] In some embodiments, a bi-prediction technique may be used in inter-picture prediction. According to the bi-prediction technique, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are before the decoding order of the current picture in the video (but the display order may be past and future, respectively). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0065] Furthermore, merge mode techniques may be used in inter-picture prediction to improve coding efficiency.

[0066] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed block-by-block. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree-divided into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal predictability and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values ​​(e.g., luma values) for pixels of 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0067] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technique. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.

[0068] The present disclosure includes aspects related to inheriting luma intra block copy (IBC) mode information and / or luma IntraTMP mode information for chroma color components and applying flipping modes for IntraTMP coded blocks.

[0069] IBC, or Current Picture Reference (CPR), is a tool that can be used to improve the coding efficiency of screen content material, as adopted in the HEVC extension to Screen Content Coding (SCC). IBC mode can be implemented as a block-level coding mode, so block matching (BM) can be performed in the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector can be used to indicate the displacement from the current block to a reference block, which is already reconstructed within the current picture. The luma block vectors of an IBC-coded CU can be defined with integer precision. The chroma block vectors can also be rounded to integer precision. When combined with Adaptive Motion Vector Resolution (AMVR), IBC mode can switch between 1-pel and 4-pel motion vector precision. IBC-coded CUs can be treated as a third prediction mode other than intra or inter prediction modes. IBC mode can be applicable to CUs whose width and height are both 64 luma samples or less.

[0070] On the encoder side, hash-based motion estimation can be performed for IBC. The encoder can perform a rate-distortion (RD) check for blocks with width or height of 16 luma samples or less. For non-merge mode, a block vector search can be performed first using a hash-based search. If the hash-based search does not return a valid candidate, a block-matching-based local search can be performed.

[0071] In hash-based search, hash key matching (32-bit CRC) between the current block and the reference block can be extended to the allowable block size. Hash key calculation for all locations in the current picture can be based on 4x4 subblocks. For a larger-sized current block, if all hash keys of all 4x4 subblocks match the hash keys of the corresponding reference locations, the hash key can be determined to match the hash key of the reference block. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matched reference block can be calculated, and the matched reference block with the smallest cost can be selected.

[0072] In the block matching search, the search range can be set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode can be signaled by a flag, and the IBC mode can be signaled as IBC Advanced Motion Vector Prediction (AMVP) mode or IBC Skip / Merge mode. Examples of AMVP mode and IBC Skip / Merge mode are as follows: (1) IBC skip / merge mode: A merge candidate index can be used to indicate which of the block vectors in a list from neighboring candidate IBC-coded blocks is used to predict the current block. The merge list can include spatial, history-based MVP (HMVP), and pairwise candidates. (2) IBC AMVP mode: Block vector differentials can be coded in the same way as motion vector differentials. The block vector prediction method can use two candidates as predictors: one from the left neighbor and one from the top neighbor (if IBC coded). If either the left neighbor or the top neighbor is unavailable, a default block vector can be used as the predictor. A flag can be signaled to indicate the block vector predictor index.

[0073] To reduce memory consumption and decoder complexity, IBC can be applied to the reconstructed portion of a predefined area that includes the region of the current CTU and some regions of the left CTU. Figures 4A, 4B, 4C, and 4D show example reference regions for IBC mode, where each block can represent 64x64 luma sample units.

[0074] Depending on the location of the current coding CU location within the current CTU, the reference region for the IBC mode can be defined as follows: (1) As shown in Figure 4A, when a current block enters the upper-left 64x64 block (402) of a current CTU (400A), in addition to the already reconstructed samples in the current CTU, the current block can also use CPR mode to reference samples in the lower-right 64x64 block (406) of the left CTU (400B). The current block can also use CPR mode to reference samples in the lower-left 64x64 block (408) of the left CTU and the upper-right 64x64 block (404) of the left CTU. (2) As shown in Figure 4B, when the current block falls into the upper right 64x64 block (410) of the current CTU, in addition to the already reconstructed samples in the current CTU, if the luminance location (0,64) for the current CTU has not yet been reconstructed, the current block can also refer to reference samples in the lower left 64x64 block (416) and lower right 64x64 block (414) of the left CTU using CPR mode. Otherwise, the current block can also refer to reference samples in the lower right 64x64 block (414) of the left CTU. (3) As shown in Figure 4C, when the current block falls into the lower-left 64x64 block (418) of the current CTU, in addition to the already reconstructed samples in the current CTU, if the luminance location (64,0) for the current CTU has not yet been reconstructed, the current block can also use the CPR mode to reference samples in the upper-right 64x64 block (420) of the current CTU and the lower-right 64x64 block (424) of the left CTU. Otherwise, the current block can also use the CPR mode to reference samples in the lower-right 64x64 block (424) of the left CTU. (4) As shown in Figure 4D, if the current block falls into the bottom right 64x64 block (426) of the current CTU (400C), the current block can only reference already reconstructed samples within the current CTU using CPR mode. This restriction may allow the IBC mode to be implemented using local on-chip memory for hardware implementation.

[0075] For IBC-coded blocks such as those in JVET-AA0070, the reconstructed-reordered IBC (RR-IBC) mode is allowed. When RR-IBC is applied, samples in the reconstructed block can be flipped according to the flip type of the current block. On the encoder side, the original block can be flipped before motion search and residual calculation for the original block (or the current block), but the predicted block can be derived without flipping. On the decoder side, the reconstructed block can be flipped to restore the original block.

[0076] Two flip methods, horizontal flip and vertical flip, may be supported for RR-IBC coded blocks. First, a syntax flag for an IBC AMVP coded block may be signaled, indicating whether the reconstructed block should be flipped. If the syntax flag indicates that the reconstructed block should be flipped, another flag specifying the flip type (e.g., vertical flip or horizontal flip) may be further signaled. In the case of IBC merging, the flip type may be inherited from a neighboring block without syntax signaling. Considering horizontal or vertical symmetry, the current block and the reference block may usually be aligned horizontally or vertically. Therefore, when a horizontal flip is applied, the vertical component of the block vector (BV) may not be signaled and may be inferred to be equal to 0. Similarly, the horizontal component of the BV may not be signaled and may be inferred to be equal to 0 when a vertical flip is applied.

[0077] To better utilize the symmetry, a flip-aware BV adjustment technique can be applied to refine the block vector candidates. Figure 5 shows an example BV adjustment based on a horizontal flip, and Figure 6 shows an example BV adjustment based on a vertical flip. As shown in Figures 5 and 6, (x nbr , y nbr ) and (x cur , y cur ) can represent the coordinates of the center samples of the neighboring blocks (e.g., (504) or (604)) and the current block (e.g., (502) or (602)), respectively. nbr and BV cur can represent the BVs of the neighboring block (504) and the current block (502), respectively. In the horizontal flip shown in Figure 5, instead of directly inheriting the BV from the neighboring block (504), the BV of the current block (502) is used. cur The horizontal component of the horizontal component of the neighboring block (504) is the motion shift of the BV of the neighboring block (504) if the neighboring block (504) is coded with a horizontal flip. nbr The horizontal component of (BVnbr h Therefore, the BV of the current block (502) can be calculated by adding cur The horizontal component of BV cur h =2(x nbr -x cur )+BV nbr h Similarly, in the vertical flip shown in FIG. 6, the BV of the current block (602) can be defined as cur The vertical component of the BV of the neighboring block (604) is used to calculate the motion shift when the neighboring block (604) is coded with a vertical flip. nbr The vertical component (BV nbr v Therefore, the BV of the current block (602) can be calculated by adding cur The vertical component of BV cur v =2(y nbr -y cur )+BV nbr v It can be defined as:

[0078] Intra-template matching prediction (IntraTMP) may be a type of intra-prediction mode for a current block in ECM software, etc. IntraTMP can copy the best predicted block (e.g., the predicted block with the smallest difference from the current block) from the reconstructed portion of the current frame, and the L-shaped template of the reconstructed portion matches the current template of the current block. For a predefined search range, the encoder can search for a template that is most similar to the current template in the reconstructed portion of the current frame and use the corresponding block as the predicted block for the current block. The encoder can then signal the use of IntraTMP mode, and the same prediction operation can be performed at the decoder side.

[0079] The prediction signal can be generated by matching the L-shaped causal neighborhood of the current block with another block within a predefined search area. An exemplary IntraTMP process can be shown in FIG. 7. As shown in FIG. 7, a current block (702) in a current frame (704) can include an L-shaped template (706) adjacent to the current block (702). The search area for determining a reference block for the current block (702) can include: (1) R1: the current CTU; (2) R2: the upper left CTU; (3) R3: the upper CTU; and (4) R4: the left CTU. The search process can use a cost function such as sum of absolute differences (SAD). Within each region (e.g., R1, R2, R3, and R4), the decoder can search for a template (e.g., (710)) with the smallest SAD relative to the current one (e.g., template (706)). The block (e.g., (708)) corresponding to the template (e.g., (710)) can be used as the prediction block. The dimensions of all regions (e.g., SearchRange_w, SearchRange_h) can be set to be proportional to the block dimensions (e.g., BlkW, BlkH) of the current block (402) to have a fixed number of SAD comparisons per pixel. For example, the dimensions of the search regions can be defined by the following equations (1) and (2): SearchRange_w=a*BlkW Equation (1) SearchRange_h=a*BlkH Equation (2) where a may be a constant that controls the gain / complexity tradeoff. In one example, a is equal to 5.

[0080] The intra template matching tool may be enabled for CUs with width and / or height sizes less than or equal to 64. The maximum CU size for intra template matching may be configurable.

[0081] If decoder-side intra mode derivation (DIMD) is not used for the current CU, the intra template matching prediction mode may be signaled at the CU level via a dedicated flag.

[0082] Although the IntraTMP mode allows template matching with L-type templates, as mentioned above, introducing multiple types of templates for template matching can be beneficial to improve the accuracy of template matching and thereby improve coding performance.

[0083] IBC modes may not be applied to chroma color components when the partitioning trees between the luma and chroma components are different, such as in VVC. However, IBC can be applied to chroma components to provide coding gain. When IBC is applied to chroma components, such as on top of ECM, it may be necessary to design how to reuse luma IBC mode information (e.g., IBC flipping modes) for the chroma components.

[0084] IntraTMP mode assumes that there is no flipping between the reference block and the current block, which may limit the coding performance of IntraTMP mode.

[0085] In this disclosure, a reconstruction rearrangement (i.e., flipping) mode can be applied to intra-template matching (i.e., RR-ITM). When RR-ITM is applied to a current block predicted using intra-template matching, samples in the reconstructed block of the current block can be flipped according to the flip type of the current block. For example, on the encoder side, the original block (or the current block) can be flipped before motion search and residual calculation, but the predicted block can be derived without flipping. On the decoder side, the reconstructed block can be flipped to restore the original block.

[0086] In one example, at the encoder side, a flip mode, such as horizontal flip or vertical flip, may be determined for a current block in a current picture. The current block may be flipped (e.g., flipped vertically or flipped horizontally) within the current picture such that sample locations of the current block are flipped within the current block based on the determined flip mode. A predictive block may be determined from multiple candidate predictive blocks within the current picture for the flipped current block based on a template matching (TM) cost. The TM cost may indicate a difference between a template of the flipped current block and each template of the multiple candidate predictive blocks. Furthermore, the current block may be encoded based on the determined predictive block. In one example, a residual block indicating a difference between the flipped current block and the determined predictive block may be encoded. In one example, the determined predictive block may also be flipped based on the flip mode. A residual block indicating a difference between the flipped current block and the flipped predictive block may be encoded.

[0087] In one example, the flip mode can include a vertical flip configured to adjust the locations of samples of a block (e.g., a current block, a reconstructed block for the current block, a predicted block for the current block, or a reference block for the current block) so that the top and bottom of the block are flipped within the block with respect to a horizontal dividing line, where the block is equally divided into top and bottom by the horizontal dividing line. In one example, the flip mode can include a horizontal flip configured to adjust the locations of samples of the block so that the left and right portions of the block are flipped within the block with respect to a vertical dividing line, where the block is equally divided into left and right portions by the vertical dividing line.

[0088] In one example, a decoder side may receive a video bitstream including coding information for a current block in a current picture. The coding information may indicate that the current block is coded by a flip mode in which sample locations of the current block are flipped within the current block when the current block is encoded in an encoder. A reference block may be determined from multiple candidate reference blocks for the current block in a reconstructed region of the current picture based on a template matching (TM) cost. The TM cost may indicate a difference between a template for the current block and each template of the multiple candidate reference blocks. A reconstructed block for the current block may be determined based on the determined reference block. The current block may be reconstructed by adjusting sample locations of the reconstructed block based on the flip mode. In one example, the reconstructed block may be determined as a sum of the determined reference block and a residual block. The residual block may be obtained from the coding information. In one example, the determined reference block may be first flipped according to the flip mode. The reconstructed block may be determined as a sum of the flipped reference block and the residual block.

[0089] In one example, multiple candidate reference blocks can be determined within a search area of ​​a reconstructed region of the current picture. A TM cost between a template of the current block and each template of the multiple candidate reference blocks can be determined. A reference block can be determined from the multiple candidate reference blocks corresponding to a minimum TM cost among the TM costs between the template of the current block and the templates of the multiple candidate reference blocks.

[0090] In one aspect, multiple flip methods may be supported for RR-ITM coded blocks. In one example, the multiple flip methods include two flip methods: horizontal flip and vertical flip. In one aspect, one or more syntax elements may be signaled to indicate whether a reconstructed block is flipped and / or the flip type. In one example, a syntax flag may be first signaled to indicate whether the reconstructed block should be flipped. If the syntax flag indicates that the reconstructed block should be flipped, another flag specifying the flip type (e.g., horizontal flip or vertical flip) may be further signaled.

[0091] In one example, first coding information (e.g., a first syntax flag) of a received video bitstream may indicate whether a flip mode should be applied to the reconstructed block. In response to the first coding information indicating that a flip mode should be applied to the reconstructed block, the type of flip mode is determined based on second coding information (e.g., a second syntax flag) of the received video bitstream. The current block may be reconstructed by adjusting the locations of samples of the reconstructed block based on the determined type of flip mode.

[0092] In one embodiment, one or more different search ranges may be applied for flips. In one example, when applying template matching, different search ranges for the template matching process may be applied for horizontal and vertical flips.

[0093] In one aspect, in the case of horizontal flip mode, the template matching search candidate position (or the template matching candidate search position) may have a vertical range smaller than the horizontal range. For example, the search candidate position in horizontal flip may have the same vertical coordinate as the current block. In one aspect, in the case of vertical flip mode, the template matching search candidate position (or the template matching candidate search position) may have a horizontal range smaller than the vertical range. For example, the search candidate position may have the same horizontal coordinate as the current block.

[0094] In one example, the vertical range of the search area is smaller than the horizontal range of the search area based on the flip mode being horizontal flip. In one example, the vertical range of the search area is larger than the horizontal range of the search area based on the flip mode being vertical flip.

[0095] In one aspect, horizontal flips and / or vertical flips can have the same search candidate numbers as non-flip modes.

[0096] In one aspect, when different coding block partitions are applied to the luma and chroma components of a picture, and when an IBC mode is applied to the luma area, the associated mode information of the luma area can be reused to predict the co-located chroma color component (or chroma block). In addition, the luma area corresponding to a chroma block can cover multiple luma coding blocks, and the boundaries of the luma blocks may not perfectly align with the boundaries of the chroma blocks.

[0097] In one aspect, the mode information reused for chroma may include, but is not limited to, an IBC flipping mode (e.g., horizontal flip or vertical flip), a block vector (BV), an IBC rotation mode (e.g., rotation angle), an IBC geometric partition mode (e.g., division pattern). In one aspect, when luma mode information is applied to chroma samples, the luma mode information, such as BV components, size-related parameters, and distance-related parameters used in the calculation, may be scaled according to the chroma sampling format.

[0098] In one aspect, for certain IBC modes, the chroma block may not inherit IBC mode information from the luma area (or luma block). Exemplary IBC modes may include, but are not limited to, an IBC horizontal flipping mode and / or an IBC vertical flipping mode.

[0099] In the present disclosure, it may be determined whether a characteristic value of a co-located luma block of a current chroma block is greater than a predefined threshold. The characteristic value may be associated with one or more pre-defined coding modes to be applied to the co-located luma block. Based on the characteristic value being greater than the pre-defined threshold, chroma coding mode information of the current chroma block may be determined based on luma coding mode information of the pre-defined luma block. In one example, the characteristic value indicates one of: (i) the number of luma collocated block positions within the co-located luma block coded by one or more pre-defined coding modes; (ii) the size of the luma area within the co-located luma block coded by one or more pre-defined coding modes; and (iii) the ratio of the luma area within the co-located luma block coded by one or more pre-defined coding modes. In one example, the one or more pre-defined coding modes may include at least one of an intra block copy (IBC) flipping mode, an IBC mode indicated by a block vector (BV), an IBC rotation mode, and an IBC geometric partition mode. In one example, the chroma coding mode information of the current chroma block can be determined based on the scaled luma coding mode information of the predefined luma block according to a scaling factor, where the scaling factor is determined based on the sampling format of the current chroma block.

[0100] In one aspect, a set of luma co-located block positions may be pre-defined for the current chroma block. If more than thr1 luma positions are coded by one or more pre-defined IBC modes, the IBC mode information associated with the pre-defined luma blocks may be inherited for the current chroma block. thr1 may be a pre-defined threshold or may be signaled by a high-level syntax, such as at the sequence level, picture level, or slice level. In one example, the set of luma co-located block positions may include, but is not limited to, the center position and four corner positions of the co-located luma blocks of the current chroma block.

[0101] In one aspect, IBC mode information associated with the first IBC coded block along a given scanning order may be fetched and reused for the co-located chroma block. In one aspect, the most frequently used IBC mode information among pre-defined luma co-located block positions may be fetched and reused for the co-located chroma block. In one aspect, the first IBC mode information that is used more than N times along a given scanning order may be fetched and reused for the co-located chroma block. Example values ​​for N may be integers such as 1, 2, 3, or 4.

[0102] In one aspect, if a characteristic value, such as a luma area size covered by one or more predefined IBC modes in the co-located luma block area of ​​the current chroma block, is greater than thr2, the IBC mode information associated with the pre-defined luma block can be inherited for the current chroma block, which can be a pre-defined threshold or can be signaled by a high-level syntax, such as at the sequence level, picture level, or slice level.

[0103] In one aspect, IBC mode information associated with the first IBC-coded block along a given scanning order may be fetched and reused for the co-located chroma block. In one aspect, the most frequently used IBC mode information among pre-defined luma co-located block positions may be fetched and reused for the co-located chroma block. In one aspect, the first IBC mode information that is used more than N times along a given scanning order may be fetched and reused for the co-located chroma block. An example value of N may be a positive integer, such as 1, 2, 3, or 4.

[0104] In one aspect, IBC mode information associated with a predefined luma block can be inherited for a chroma block if a characteristic value, such as a ratio of the co-located luma block area covered by one or more pre-defined IBC modes, in the co-located luma block area is greater than thr3, which can be a pre-defined threshold or can be signaled by a high-level syntax, e.g., at the sequence level, picture level, slice level, etc.

[0105] In one aspect, IBC mode information associated with the first IBC coded block along a given scanning order may be fetched and reused for the co-located chroma block. In one aspect, the most frequently used IBC mode information among pre-defined luma co-located block positions may be fetched and reused for the co-located chroma block. In one aspect, the first IBC mode information that is used more than N times along a given scanning order may be fetched and reused for the co-located chroma block. Example values ​​for N may be integers such as 1, 2, 3, or 4.

[0106] In one example, the chroma coding mode information of the current chroma block may be determined based on one of: (i) luma coding mode information of a first luma block coded by one of one or more predefined coding modes along a predefined scanning order; (ii) luma coding mode information of a most frequently used mode among the luma co-located block positions coded by one or more predefined coding modes; and (iii) a first luma coding mode information of one or more predefined coding modes used more than N times along the predefined scanning order, where N may be a positive integer.

[0107] In one aspect, a specific co-located luma location may be predefined. If the coding mode associated with a sample located at the pre-defined luma location is coded by the IBC mode, the IBC mode information associated with the sample may be inherited for the chroma block. Examples of the specific co-located luma location may include, but are not limited to, the center position of the co-located luma block of the chroma block or the top left position of the co-located luma block.

[0108] In one example, the chroma coding mode information of the current chroma block can be determined based on luma coding mode information of samples at predefined collocated luma locations. The samples at the predefined collocated luma locations can be coded by one of one or more predefined coding modes, and the predefined collocated luma locations can include one of a center location and a top-left location of the collocated luma block.

[0109] In one aspect, when different coding block partitioning is applied to the luma component and the co-located chroma component of a picture, and when ITM mode (or IntraTMP) is applied to the luma component, the associated mode information of the luma component may be reused to predict the co-located chroma color component.

[0110] In one aspect, the mode information reused for the chroma block may include, but is not limited to, the IntraTMP flipping mode (e.g., horizontal flip or vertical flip), the displacement vector used for the IntraTMP.

[0111] In one aspect, for certain IntraTMP modes, the chroma block may not inherit IntraTMP mode information from the luma block. Exemplary IntraTMP modes may include, but are not limited to, an IntraTMP horizontal flipping mode and / or an IntraTMP vertical flipping mode.

[0112] In one aspect, a set of luma-colocated block positions may be predefined for the current chroma block. If more than thr1 luma-colocated block positions are coded by one or more predefined IntraTMP modes, the IntraTMP mode information associated with the predefined luma blocks may be inherited for the current chroma block. thr1 may be a predefined threshold or may be signaled by a high-level syntax, such as at the sequence level, picture level, or slice level. In one example, the set of luma-colocated block positions may include, but is not limited to, the center position and four corner positions of the co-located luma blocks of the current chroma block.

[0113] In one example, the one or more predefined coding modes include at least one of an intra-template matching prediction (IntraTMP) flipping mode and an IntraTMP mode indicated by a displacement vector.

[0114] In one aspect, IntraTMP mode information associated with the first IntraTMP coded block along a given scanning order may be fetched and reused for the co-located chroma block. In one aspect, the most frequently used IntraTMP mode information among pre-defined luma co-located block positions may be fetched and reused for the co-located chroma block. In one aspect, the first IntraTMP mode information used more than N times along a given scanning order may be fetched and reused for the co-located chroma block. An example value of N may be a positive integer, such as 1, 2, 3, or 4.

[0115] In one aspect, if the luma area size covered by one or more predefined IntraTMP modes in a co-located luma block area of ​​a chroma block is larger than thr2, the IntraTMP mode information associated with the predefined luma block can be inherited for the chroma block, where thr2 can be a predefined threshold or can be signaled by a high-level syntax, for example, at the sequence level, picture level, slice level, etc.

[0116] In one aspect, IntraTMP mode information associated with the first IntraTMP coded block along a given scanning order may be fetched and reused for the co-located chroma block. In one aspect, the most frequently used IntraTMP mode information among pre-defined luma co-located block positions may be fetched and reused for the co-located chroma block. In one aspect, the first IntraTMP mode information used more than N times along a given scanning order may be fetched and reused for the co-located chroma block. An example value of N may be a positive integer, such as 1, 2, 3, or 4.

[0117] In one aspect, if the ratio of the co-located luma block area size covered by one or more pre-defined IntraTMP modes in the co-located luma block area of ​​a chroma block is greater than thr2, the IntraTMP mode information associated with the pre-defined luma block can be inherited for the chroma block, where thr2 can be a pre-defined threshold or can be signaled by a high-level syntax, e.g., at the sequence level, picture level, slice level, etc.

[0118] In one aspect, IntraTMP mode information associated with the first IntraTMP coded block along a given scanning order may be fetched and reused for the co-located chroma block. In one aspect, the most frequently used IntraTMP mode information among pre-defined luma co-located block positions may be fetched and reused for the co-located chroma block. In one aspect, the first IntraTMP mode information used more than N times along a given scanning order may be fetched and reused for the co-located chroma block. An example value of N may be a positive integer, such as 1, 2, 3, or 4.

[0119] In one aspect, the search candidates pointed to by the block vectors of the associated luma IntraTMP blocks having the scaling factor of the subsampling ratio can be searched to find the best (or selected) template match with the smallest template matching cost. In one example, the associated luma IntraTMP blocks can include (i) the first IntraTMP coded block along the predefined scanning order, (ii) the collocated luma blocks, and (iii) the IntraTMP coded block corresponding to the first IntraTMP mode information used more than N times.

[0120] In one aspect, the flip mode of each search candidate may be inherited from the flip mode of the associated IntraTMP luma block. In one aspect, template matching may be performed for all possible flip modes (e.g., vertical flip or horizontal flip) for each candidate to find the best (or selected) flip mode with the smallest template matching cost. In one aspect, the flip mode may be signaled for each chroma coding unit.

[0121] In one example, multiple candidate reference chroma blocks may be determined within a search range within a current picture indicated by a block vector (BV) of a current chroma block based on one or more predefined coding modes, including an IntraTMP mode. The BV of the current chroma block may be determined based on one of multiple associated luma blocks. The multiple associated luma blocks may include (i) a first IntraTMP-coded block along a predefined scanning order, (ii) a co-located luma block, and (iii) an IntraTMP-coded block corresponding to first IntraTMP mode information used more than N times along the predefined scanning order. Multiple flip modes may be determined for each of the multiple candidate reference chroma blocks. A TM cost between a template of each of the multiple candidate reference chroma blocks and a template of the current chroma block may be determined according to each of the multiple flip modes. The flip mode may be determined from multiple flip modes corresponding to the smallest TM cost among the TM costs between a template of each of the multiple candidate reference chroma blocks and a template of the current chroma block according to the multiple flip modes. The flip mode of the current chroma block may be determined as a flip mode determined from a plurality of flip modes.

[0122] In one example, the BV of the current chroma block can be determined as a scaled BV of one of multiple associated luma blocks based on a scaling factor, which can be determined based on a subsampling ratio associated with the current chroma block.

[0123] In one aspect, a specific co-located luma location can be predefined. If the coding mode associated with a sample located at the pre-defined luma location is coded by the IntraTMP mode, the IntraTMP mode information associated with the sample can be inherited for the chroma block. In one example, the specific co-located luma location can include, but is not limited to, the center position of the co-located luma area for the current chroma block.

[0124] 8 shows a flowchart outlining a process (800) according to one embodiment of the present disclosure. The process (800) may be used in a video decoder. In various embodiments, the process (800) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), or the like. In some embodiments, the process (800) is implemented by software instructions, such that the processing circuit performs the process (800) when it executes the software instructions. The process begins at (S801) and proceeds to (S810).

[0125] At (S810), a video bitstream including coding information for a current block in a current picture is received, the coding information indicating that the current block is coded by a flip mode in which locations of samples of the current block are adjusted within the current block.

[0126] At (S820), a reference block is determined from a plurality of candidate reference blocks in a reconstructed region of the current picture for the current block based on a template matching (TM) cost, where the TM cost indicates a difference between the template of the current block and each template of the plurality of candidate reference blocks.

[0127] In (S830), a reconstruction block for the current block is determined based on the determined reference block.

[0128] At (S840), the current block is reconstructed by adjusting the location of the samples of the reconstructed block within the reconstructed block based on the flip mode.

[0129] In one example, the flip mode includes one of (i) a vertical flip mode configured to adjust the locations of samples of the current block such that the top and bottom of the current block are flipped within the current block, and (ii) a horizontal flip mode configured to adjust the locations of samples of the current block such that the left and right parts of the current block are flipped within the current block.

[0130] In one example, first coding information is received from a received video bitstream. The first coding information indicates whether a flip mode should be applied to the reconstructed block. In response to the first coding information indicating that a flip mode should be applied to the reconstructed block, a type of flip mode is determined based on second coding information in the received video bitstream. The current block is reconstructed by adjusting locations of samples of the reconstructed block based on the determined type of flip mode.

[0131] In one example, a plurality of candidate reference blocks are determined within a search area of ​​a reconstructed region of a current picture. A TM cost between a template of the current block and each template of the plurality of candidate reference blocks is determined. A reference block is determined from the plurality of candidate reference blocks corresponding to the smallest TM cost among the TM costs between the template of the current block and the templates of the plurality of candidate reference blocks.

[0132] In one example, the vertical range of the search area is smaller than the horizontal range of the search area based on the flip mode being horizontal flip. In one example, the vertical range of the search area is larger than the horizontal range of the search area based on the flip mode being vertical flip.

[0133] In one example, the determined reference block is further flipped by adjusting the location of the samples of the determined reference block within the determined reference block based on the flip mode, and the reconstruction block is determined based on the flipped reference block.

[0134] The process then proceeds to (S899) and ends.

[0135] Process 800 may be adapted as appropriate. Steps of process 800 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0136] 9 shows a flowchart outlining a process (900) according to one embodiment of the present disclosure. The process (900) can be used in a video encoder. In various embodiments, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of the video encoder (103), a processing circuit that performs the functions of the video encoder (303), or the like. In some embodiments, the process (900) is implemented by software instructions, such that the processing circuit performs the process (900) when it executes the software instructions. The process starts at (S901) and proceeds to (S910).

[0137] At (S910), the current block in the current picture is flipped according to a flip mode in which the locations of the samples of the current block are adjusted within the current block.

[0138] At (S920), a reference block is determined from a plurality of candidate reference blocks in a reconstructed region of the current picture for the flipped current block based on a template matching (TM) cost, where the TM cost indicates a difference between a template of the flipped current block and each template of the plurality of candidate reference blocks.

[0139] At (S930), the flipped current block is encoded based on the determined reference block.

[0140] At (S940), coding information indicating that the current block is encoded by flip mode is signaled.

[0141] The process then proceeds to (S999) and ends.

[0142] The process (900) may be adapted as appropriate. Steps of the process (900) may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0143] 10 shows a flowchart outlining a process (1000) according to one embodiment of the present disclosure. The process (1000) can be used in a video decoder. In various embodiments, the process (1000) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), or the like. In some embodiments, the process (1000) is implemented by software instructions, such that the processing circuit performs the process (1000) when it executes the software instructions. The process starts at (S1001) and proceeds to (S1010).

[0144] At (S1010), a video bitstream including a current chroma block and a co-located luma block of the current chroma block in the current picture is determined.

[0145] At (S1020), it is determined whether a characteristic value of the collocated luma block is greater than a predefined threshold, the characteristic value being associated with one or more predefined coding modes to be applied to the collocated luma block.

[0146] In (S1030), based on the characteristic value being greater than the predefined threshold, the chroma coding mode information of the current chroma block is determined based on the predefined luma coding mode information of the luma block.

[0147] At (S1040), the current chroma block is reconstructed based on the determined chroma coding mode information.

[0148] In one example, the characteristic value indicates one of: (i) the number of luma collocated block positions in a collocated luma block coded by one or more predefined coding modes; (ii) the size of the luma area in a collocated luma block coded by one or more predefined coding modes; and (iii) the ratio of the luma area in a collocated luma block coded by one or more predefined coding modes.

[0149] In one example, the luminance collocated block positions include a center position and four corner positions of the collocated luminance block.

[0150] In one example, the one or more predefined coding modes include at least one of an intra block copy (IBC) flipping mode, an IBC mode indicated by a block vector (BV), an IBC rotation mode, and an IBC geometric partition mode.

[0151] In one example, the one or more predefined coding modes include at least one of an intra-template matching prediction (IntraTMP) flipping mode and an IntraTMP mode indicated by a displacement vector.

[0152] In one example, the chroma coding mode information of the current chroma block is determined based on one of (i) luma coding mode information of a first luma block coded by one of one or more predefined coding modes along a predefined scanning order, (ii) luma coding mode information of a most frequently used mode among the luma co-located block positions coded by one or more predefined coding modes, and (iii) a first luma coding mode information of one or more predefined coding modes used more than N times along the predefined scanning order, where N is a positive integer.

[0153] In one example, the chroma coding mode information of the current chroma block is determined based on luma coding mode information of samples at predefined collocated luma locations, where the samples at the predefined collocated luma locations are coded by one of one or more predefined coding modes, and the predefined collocated luma locations include one of a center location and a top-left location of the collocated luma block.

[0154] In one example, multiple candidate reference chroma blocks are determined within a search range in the current picture indicated by a block vector (BV) of a current chroma block based on one or more predefined coding modes, including an IntraTMP mode. The BV of the current chroma block is determined based on one of multiple associated luma blocks. The multiple associated luma blocks include (i) a first IntraTMP-coded block along a predefined scanning order, (ii) a co-located luma block, and (iii) an IntraTMP-coded block corresponding to the first IntraTMP mode information used more than N times along the predefined scanning order. Multiple flip modes are determined for each of the multiple candidate reference chroma blocks. A TM cost between a template of each of the multiple candidate reference chroma blocks and the template of the current chroma block is determined according to each of the multiple flip modes. The flip mode is determined from multiple flip modes corresponding to the smallest TM cost among the TM costs between the template of each of the multiple candidate reference chroma blocks and the template of the current chroma block according to the multiple flip modes. The flip mode of the current chroma block is determined as a flip mode determined from a plurality of flip modes.

[0155] In one example, the BV of the current chroma block is determined as a scaled BV of one of the multiple associated luma blocks based on a scaling factor, where the scaling factor is determined based on a subsampling ratio associated with the current chroma block.

[0156] In one example, the chroma coding mode information of the current chroma block is determined based on the scaled luma coding mode information of the predefined luma block according to a scaling factor, and the scaling factor is determined based on the sampling format of the current chroma block.

[0157] The process then proceeds to (S1099) and ends.

[0158] The process 1000 may be adapted as appropriate. Steps of the process 1000 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0159] 11 shows a flowchart outlining a process (1100) according to one embodiment of the present disclosure. The process (1100) can be used in a video encoder. In various embodiments, the process (1100) is performed by a processing circuit, such as a processing circuit that performs the functions of the video encoder (103), a processing circuit that performs the functions of the video encoder (303), or the like. In some embodiments, the process (1100) is implemented by software instructions, such that the processing circuit performs the process (1100) when it executes the software instructions. The process starts at (S1101) and proceeds to (S1110).

[0160] At (S1110), it is determined whether a characteristic value of a co-located luma block of the current chroma block in the current picture is greater than a predefined threshold, the characteristic value being associated with one or more pre-defined coding modes to be applied to the co-located luma block.

[0161] In (S1120), based on the characteristic value being greater than the predefined threshold, the chroma coding mode information of the current chroma block is determined based on the luma coding mode information of the predefined luma block in the current picture.

[0162] At (S1130), the current chroma block is encoded based on the determined chroma coding mode information.

[0163] The process then proceeds to (S1199) and ends.

[0164] The process 1100 may be adapted as appropriate. Steps of the process 1100 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0165] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 12 illustrates a computer system (1200) suitable for implementing certain embodiments of the disclosed subject matter.

[0166] Computer software may be coded using any suitable machine code or computer language that can undergo mechanisms such as assembly, compilation, linking, etc. to produce code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or that can be executed via interpretation, microcode execution, etc.

[0167] The instructions may be executed on various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.

[0168] 12 for computer system (1200) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having a dependency or requirement related to any one or combination of components illustrated in the exemplary embodiment of computer system (1200).

[0169] The computer system (1200) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0170] The input human interface devices may include one or more (only one of each is shown) of a keyboard (1201), a mouse (1202), a trackpad (1203), a touchscreen (1210), a data glove (not shown), a joystick (1205), a microphone (1206), a scanner (1207), and a camera (1208).

[0171] The computer system (1200) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1210), data gloves (not shown), or joystick (1205), although some haptic feedback devices may not function as input devices), audio output devices (such as speakers (1209), headphones (not shown)), visual output devices (such as screens (1210), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities and with or without haptic feedback capabilities, some of which may output two-dimensional visual output or output in more than three dimensions via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0172] The computer system (1200) may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (1220) along with CD / DVD or similar media (1221), thumb drives (1222), removable hard drives or solid-state drives (1223), legacy magnetic media (not shown) such as tape and floppy disks, and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0173] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0174] The computer system (1200) also includes an interface (1254) to one or more communication networks (1255). The networks can be, for example, wireless, wired, or optical. The networks can further be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet and wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including CAN Bus. Certain networks generally require an external network interface adapter attached to a particular general-purpose data port or peripheral bus (1249) (e.g., a USB port on the computer system (1200)), while other networks are generally integrated into the core of the computer system (1200) by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1200) can communicate with other entities. Such communications can be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a specific CANbus device), or two-way, for example, to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks, as described above, can be used with each of these networks and network interfaces.

[0175] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1240) of the computer system (1200).

[0176] The cores (1240) may include one or more central processing units (CPUs) (1241), graphics processing units (GPUs) (1242), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1243), task-specific hardware accelerators (1244), graphics adapters (1250), etc. These devices may be connected via a system bus (1248), along with read-only memory (ROM) (1245), random access memory (1246), and internal mass storage (1247) such as a non-user-accessible internal hard drive or SSD. In some computer systems, the system bus (1248) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1248) or via a peripheral bus (1249). In one example, a screen (1210) may be connected to the graphics adapter (1250). Architectures for peripheral buses include PCI, USB, and the like.

[0177] The CPU (1241), GPU (1242), FPGA (1243), and accelerator (1244) can execute specific instructions that, in combination, can constitute the aforementioned computer code. That computer code can be stored in ROM (1245) or RAM (1246). Transient data can also be stored in RAM (1246), while permanent data can be stored, for example, in internal mass storage (1247). Rapid storage and retrieval from any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more of the CPU (1241), GPU (1242), mass storage (1247), ROM (1245), RAM (1246), etc.

[0178] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0179] By way of example and not limitation, a computer system (1200) having the architecture, and specifically the core (1240), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be the user-accessible mass storage introduced above, as well as media associated with specific storage of the core (1240) that is non-transitory in nature, such as the core's internal mass storage (1247) or ROM (1245). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1240). The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause the core (1240), and specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM (1246) and modifying such data structures according to the software-defined processes. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1244)), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software may encompass logic, where appropriate, and vice versa. References to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0180] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C, i.e., at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A through C, is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B and one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of the listed elements where applicable, such as when the elements are not mutually exclusive.

[0181] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]

[0182] 100 communication system / video processing system, 101 video source, 102 stream of video pictures, 103 video encoder, 104 encoded video data, 105 streaming server, 106 client subsystem, 107 copy of encoded video data / encoded video data, 108 client subsystem, 109 copy of encoded video data / encoded video data, 110 video decoder, 111 output stream of video pictures, 112 display, 113 capture subsystem, 120 electronic device, 130 electronic device, 201 channel, 210 video decoder, 212 render device, 215 buffer memory, 220 parser, 221 symbols, 230 electronic device, 231 receiver, 251 scaler / inverse transform unit, 252 intra-picture prediction unit, 253 motion compensation prediction unit, 255 aggregator, 256 loop filter unit, 257 Reference picture memory, 258 Current picture buffer, 301 Video source, 303 Video encoder, 320 Electronic device, 330 Source coder, 332 Coding engine, 333 Local video decoder, 334 Reference picture memory, 335 Predictor, 340 Transmitter, 343 Video sequence, 345 Entropy coder, 350 Controller, 360 Communication channel, 400A Current CTU, 400B Left CTU, 400C Current CTU, 402 Top left 64x64 block of current CTU / current block, 404 Top right 64x64 block of left CTU, 406 Bottom right 64x64 block of left CTU, 408 Bottom left 64x64 block of left CTU, 410 Top right 64x64 block of current CTU, 414 Bottom right 64x64 block of left CTU, 416 64x64 block below left of left CTU, 418 64x64 block below left of current CTU, 420 64x64 block above right of current CTU, 424 64x64 block below right of left CTU, 426 64x64 block below right of current CTU, 502 Current block, 504 Nearby block, 602 Current block, 604 Nearby block, 702 Current block, 704current frame, 706 L-shaped template, 708 block, 710 template, 1200 computer system, 1201 keyboard, 1202 mouse, 1203 trackpad, 1205 joystick, 1206 microphone, 1207 scanner, 1208 camera, 1209 speaker, 1210 touchscreen / screen, 1220 CD / DVD ROM / RW, 1221 CD / DVD / similar media, 1222 thumb drive, 1223 removable hard drive / solid state drive, 1240 core, 1241 central processing unit (CPU), 1242 graphics processing unit (GPU), 1243 field programmable gate array (FPGA), 1244 hardware accelerator, 1245 read-only memory (ROM), 1246 random access memory (RAM), 1247 internal mass storage, 1248 system bus, 1249 Peripheral bus, 1250 graphics adapter, 1254 interface, 1255 communication network

Claims

1. 1. A method of video decoding, comprising: receiving a video bitstream including coding information for a current block in a current picture, the coding information indicating that the current block is coded according to a flip mode in which locations of samples of the current block are adjusted within the current block; determining a reference block from a plurality of candidate reference blocks in a reconstructed region of the current picture for the current block based on a template matching (TM) cost, the TM cost indicating a difference between a template of the current block and each template of the plurality of candidate reference blocks; determining a reconstruction block for the current block based on the determined reference block; reconstructing the current block by adjusting locations of samples of the reconstructed block within the reconstructed block based on the flip mode; A method comprising:

2. 2. The method of claim 1, wherein the flip modes include one of: (i) a vertical flip mode configured to adjust the locations of the samples of the current block such that a top and a bottom of the current block are inverted within the current block; and (ii) a horizontal flip mode configured to adjust the locations of the samples of the current block such that a left and a right portion of the current block are inverted within the current block.

3. The step of reconstructing the current block comprises: receiving first coding information from the received video bitstream, the first coding information indicating whether the flip mode should be applied to the reconstructed block; determining a type of the flip mode based on second coding information in the received video bitstream in response to the first coding information indicating that the flip mode should be applied to the reconstructed block; reconstructing the current block by adjusting the locations of the samples of the reconstructed block based on the determined type of the flip mode; 10. The method of claim 1, further comprising:

4. The step of determining the reference block from the plurality of candidate reference blocks comprises: determining the plurality of candidate reference blocks within a search area of ​​the reconstructed area of ​​the current picture; determining the TM cost between the template of the current block and a template of each of the plurality of candidate reference blocks; determining the reference block from the plurality of candidate reference blocks corresponding to the smallest TM cost among the TM costs between the template of the current block and the templates of the plurality of candidate reference blocks; 10. The method of claim 1, further comprising:

5. based on the flip mode being horizontal flip, the vertical extent of the search area is smaller than the horizontal extent of the search area; based on the flip mode being vertical flip, the vertical extent of the search area is greater than the horizontal extent of the search area; The method of claim 4.

6. The step of determining the reconstruction block comprises: flipping the determined reference block by adjusting locations of samples of the determined reference block within the determined reference block based on the flip mode; determining the reconstructed block based on the flipped reference block; 10. The method of claim 1, further comprising:

7. 1. A method of video decoding, comprising: receiving a video bitstream including a current chroma block and a co-located luma block of the current chroma block in a current picture; determining whether a characteristic value of the collocated luma block is greater than a predefined threshold, the characteristic value being associated with one or more predefined coding modes to be applied to the collocated luma block; determining chroma coding mode information of the current chroma block based on luma coding mode information of predefined luma blocks in the current picture based on the characteristic value being greater than the predefined threshold; reconstructing the current chroma block based on the determined chroma coding mode information; A method comprising:

8. 8. The method of claim 7, wherein the characteristic value indicates one of: (i) a number of luminance collocated block positions within the collocated luminance blocks coded by the one or more predefined coding modes; (ii) a size of a luminance area within the collocated luminance blocks coded by the one or more predefined coding modes; and (iii) a ratio of luminance areas within the collocated luminance blocks coded by the one or more predefined coding modes.

9. The method of claim 8 , wherein the luminance collocated block locations include a center location and four corner locations of the collocated luminance block.

10. 8. The method of claim 7, wherein the one or more predefined coding modes include at least one of an intra block copy (IBC) flipping mode, an IBC mode indicated by a block vector (BV), an IBC rotation mode, and an IBC geometric partition mode.

11. The method of claim 7 , wherein the one or more predefined coding modes include at least one of an intra-template matching prediction (IntraTMP) flipping mode and an IntraTMP mode indicated by a displacement vector.

12. The step of determining the chroma coding mode information comprises: determining the chroma coding mode information of the current chroma block based on (i) luma coding mode information of a first luma block coded by one of the one or more predefined coding modes along a predefined scanning order, (ii) luma coding mode information of a most frequently used mode among the luma co-located block positions coded by the one or more predefined coding modes, and (iii) one of first luma coding mode information of the one or more predefined coding modes used more than N times along a predefined scanning order, where N is a positive integer; 9. The method of claim 8, further comprising:

13. The step of determining the chroma coding mode information comprises: determining the chroma coding mode information of the current chroma block based on luma coding mode information of samples at predefined collocated luma locations, the samples at the predefined collocated luma locations being coded by one of the one or more predefined coding modes, the predefined collocated luma locations including one of a center location and a top-left location of the collocated luma block; 8. The method of claim 7, further comprising:

14. based on the one or more predefined coding modes, including an IntraTMP mode; determining a plurality of candidate reference chroma blocks within a search range in the current picture indicated by a block vector (BV) of the current chroma block, the BV of the current chroma block being determined based on one of a plurality of associated luma blocks, the plurality of associated luma blocks including (i) a first IntraTMP coded block along a predefined scanning order, (ii) the co-located luma block, and (iii) an IntraTMP coded block corresponding to a first IntraTMP mode information that is used more than N times along the predefined scanning order; determining a plurality of flip modes for each of the plurality of candidate reference chroma blocks; determining a TM cost between a template of each of the plurality of candidate reference chroma blocks and a template of the current chroma block according to a respective flip mode of the plurality of flip modes; determining a flip mode from the plurality of flip modes corresponding to a smallest TM cost among the TM costs between the template of each of the plurality of candidate reference chroma blocks and the template of the current chroma block according to the plurality of flip modes; determining a flip mode for the current chroma block as the determined flip mode from the plurality of flip modes; 12. The method of claim 11, further comprising:

15. 15. The method of claim 14, wherein the BV of the current chroma block is determined as a scaled BV of the one of the plurality of associated luma blocks based on a scaling factor, the scaling factor being determined based on a subsampling ratio associated with the current chroma block.

16. The step of determining the chroma coding mode information comprises: determining the chroma coding mode information of the current chroma block based on scaled luma coding mode information of the predefined luma block according to a scaling factor, the scaling factor being determined based on a sampling format of the current chroma block; 8. The method of claim 7, further comprising:

17. 1. An apparatus comprising: receiving a video bitstream including coding information for a current block in a current picture, the coding information indicating that the current block is coded using a flip mode in which locations of samples of the current block are adjusted within the current block; determining a reference block from a plurality of candidate reference blocks in a reconstructed region of the current picture for the current block based on a template matching (TM) cost, the TM cost indicating a difference between a template of the current block and each template of the plurality of candidate reference blocks; determining a reconstruction block for the current block based on the determined reference block; reconstructing the current block by adjusting the location of the samples of the reconstructed block within the reconstructed block based on the flip mode; A processing circuit configured as follows: An apparatus comprising:

18. 18. The apparatus of claim 17, wherein the flip modes include one of: (i) a vertical flip mode configured to adjust the locations of the samples of the current block such that a top and a bottom of the current block are inverted within the current block; and (ii) a horizontal flip mode configured to adjust the locations of the samples of the current block such that a left and a right portion of the current block are inverted within the current block.

19. The processing circuitry receiving first coding information from the received video bitstream, the first coding information indicating whether the flip mode should be applied to the reconstructed block; determining a type of the flip mode based on second coding information in the received video bitstream in response to the first coding information indicating that the flip mode should be applied to the reconstructed block; reconstructing the current block by adjusting the locations of the samples of the reconstructed block based on the determined type of the flip mode.

20. The apparatus of claim 17, further configured to:

20. The processing circuitry determining the plurality of candidate reference blocks within a search area of ​​the reconstructed region of the current picture; determining the TM cost between the template of the current block and a template of each of the plurality of candidate reference blocks; determining the reference block from the plurality of candidate reference blocks corresponding to the smallest TM cost among the TM costs between the template of the current block and the templates of the plurality of candidate reference blocks; 20. The apparatus of claim 17, further configured to:

Citation Information

Patent Citations

  • Image encoding apparatus, image decoding apparatus, image encoding method, and image decoding method

    JP2011114572A

  • Method and device for encoding / decoding image, and recording medium in which bitstream is stored

    US20200204799A1