Method for video processing and apparatus for video decoding

By determining chroma block vectors from luma block vectors and performing template matching, the method addresses inefficiencies in encoding and decoding chroma blocks within a chroma split tree, enhancing compression efficiency and reducing bandwidth and storage requirements.

JP2025523730APending Publication Date: 2025-07-25TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024522326
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-08
Filing Date
2022-11-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently encoding and decoding chroma blocks within a chroma split tree, particularly in scenarios where chroma blocks are collocated with luma blocks, leading to inefficiencies in bandwidth and storage requirements.

Method used

The proposed solution involves decoding a syntax element indicating a current picture referencing (CPR) mode for chroma blocks and determining a chroma block vector based on associated luma block vectors, using techniques such as deriving block vector predictors from luma block vectors and performing template matching to reconstruct chroma blocks within the same luma area.

Benefits of technology

This approach enhances video encoding and decoding efficiency by reducing redundancy and improving compression ratios for chroma blocks, thereby optimizing bandwidth and storage needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523730000001_ABST
    Figure 2025523730000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, a processing circuit receives a coded video bitstream including a current picture. The current picture includes chroma blocks within a chroma split tree, and the chroma blocks are collocated within the same luma area as one or more luma blocks. The processing circuit decodes a syntax element indicating a current picture reference (CPR) mode of a chroma block from the coded video bitstream, and in response to the CPR mode, determines a chroma block vector of the chroma block according to one or more luma block vectors associated with the one or more luma blocks. The chroma block vector indicates a reference chroma block within the current picture. The processing circuit reconstructs the chroma block based on the reference chroma block within the current picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure describes embodiments generally related to video coding.

Background Art

[0002] The description of the background given herein is for the purpose of generally presenting the background of the present disclosure. The research of the presently named inventors, to the extent that the research is described in this background section and aspects of the description that might otherwise be eligible as prior art at the time of filing, is not admitted as prior art to the present disclosure, either expressly or implicitly.

[0003] Uncompressed digital images and / or videos can include a sequence of pictures, each picture having, for example, spatial dimensions of 1920×1080 luminance samples and associated chrominance samples. The sequence of pictures can have, for example, a fixed or variable picture rate of 60 pictures per second, i.e., 60 Hz (commonly also known as the frame rate). Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p60 4:2:0 video at 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires more than 600 Gbytes of storage space.

[0004] One purpose of the encoding and decoding of images and / or videos can be the reduction of redundancy of the input image and / or video signal by compression. Compression can, in some cases, help reduce the above bandwidth and / or storage space requirements by more than two orders of magnitude. Although the description in this specification uses video encoding / decoding as an example, the same techniques can be applied to image encoding / decoding in the same way without departing from the spirit of the present disclosure. Both lossless compression and lossy compression as well as combinations thereof can be used. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. In the case of video, lossy compression is widely used. The amount of acceptable distortion depends on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that higher acceptable / tolerable distortion can result in a higher compression ratio.

[0005] Video encoders and decoders can utilize techniques from several broad categories including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0006] Video coding technology can include techniques known as intracoding. In intracoding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some video coders, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in an intra mode, that picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used as the first pictures of a coded video bitstream and a video session to reset the decoder state, or as still pictures, since they can be used to reset the decoder state. Samples of an intra block can be subject to a transform, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique for minimizing sample values in a pre-transform region. In some cases, the smaller the DC value after transformation and the smaller the AC coefficients, the fewer bits are required at a given quantization step size to represent the block after entropy coding.

[0007] For example, conventional intracoding used in MPEG-2 generation coding technology does not use intra prediction. However, some newer video compression technologies include techniques that attempt to perform prediction, for example, based on surrounding sample data and / or metadata obtained during encoding and / or decoding of data. Such techniques are hereinafter referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and does not use reference pictures.

[0008] There can be a variety of forms of intra prediction. When more than one such technique can be used with a given video coding technique, the particular technique in use can be coded as a particular intra prediction mode that uses that particular technique. In some cases, the intra prediction mode can have sub - modes and / or parameters, and the sub - modes and / or parameters can be coded independently or can be included in the mode codeword. This defines the prediction mode being used. Which codeword should be used for a given mode, sub - mode, and / or parameter combination can affect the coding efficiency gain through intra prediction, so entropy coding techniques can be used to convert the codeword into a bitstream.

[0009] Intra prediction for a particular mode was introduced by H.264, refined in H.265, and further refined in more recent coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). The predictor block can be formed using adjacent sample values of already available samples. The sample values of the adjacent samples are copied into the predictor block according to the direction. The reference to the direction in use can be coded in the bitstream or can itself be predicted.

[0010] Referring to FIG. 1A, in the lower right, a subset of 9 predictor directions known from 33 possible predictor directions of H.265 (corresponding to 33 of the 35 intra - modes' angular modes) is represented. The point (101) where the arrows converge corresponds to the sample being predicted. The arrows represent the direction in which the sample is being predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples at a 45 - degree angle from the horizontal and to the upper right. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples at a 22.5 - degree angle from the horizontal and to the lower left of sample (101).

[0011] Still referring to FIG. 1A, in the upper left, a square block (104) of 4×4 samples (indicated by the thick dashed line) is shown. The square block (104) contains 16 samples, and each sample is labeled using "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample within the block (104) in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is in the lower right. Further, reference samples following a similar numbering scheme are shown. The reference samples are labeled for the block (104) using "R", its Y position (e.g., row index), and its X position (column index). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, and thus negative values need not be used.

[0012] Intra-picture prediction can work by copying the reference sample value from adjacent samples according to the prediction direction signaled by the signal. For example, assume that the coded video bitstream contains signaling indicating that for this block, the prediction direction coincides with arrow (102), i.e., the sample is predicted from one or more predicted samples at a 45-degree angle from the horizontal and in the upper right. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0013] In some cases, the values of multiple reference samples may be combined, particularly when the direction is not evenly divisible by 45 degrees, for example, through interpolation, to calculate the reference sample.

[0014] The number of possible directions has been increasing as video coding technology develops. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those likely directions with fewer bits while accepting some penalty for the less likely directions. Furthermore, the direction itself can sometimes be predicted from the neighboring directions used in adjacent, already decoded blocks.

[0015] FIG. 1B shows a schematic diagram (110) representing 65 intra prediction directions by JEM to illustrate the number of prediction directions increasing over time.

[0016] The mapping of intra prediction direction bits representing the direction in the coded video bitstream can be different for each video coding technology. Such mapping can range from a simple direct mapping to complex adaptive schemes including the most probable mode, and similar techniques up to codewords. In most cases, however, there can be certain directions that occur less statistically likely in video content than certain other directions. Since the goal of video compression is to reduce redundancy, those less likely directions will be represented by more bits than the more likely directions in a well - functioning video coding technology.

[0017] Image and / or video encoding and decoding can be performed using inter-picture prediction with motion compensation. Motion compensation can be an irreversible compression technique, and blocks of sample data from a previously reconstructed picture or a portion thereof (reference picture) are spatially shifted in the direction indicated by a motion vector (hereinafter MV) and then used for prediction of a newly reconstructed picture or picture portion. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or three dimensions, where the third dimension is an indication of the reference picture in use (the latter can indirectly be the temporal dimension).

[0018] In some video compression techniques, the MV applicable to a particular area of sample data can be predicted from other MVs, for example, from other areas of sample data spatially adjacent to the area being reconstructed and from those preceding it in the decoding order. By doing so, the amount of data required to code the MV can be significantly reduced, thereby removing redundancy and increasing compression. For example, when coding an input video signal obtained from a camera (known as natural video), there is a statistical likelihood that areas larger than the area to which a single MV is applicable move in a similar direction, and thus, in some cases, MV prediction can work effectively by predicting with a similar motion vector derived from the MVs of adjacent areas. As a result, the MV required for a given area can be similar or the same as the MV predicted from surrounding MVs and can be represented in fewer bits than the number of bits that would be used if the MV were directly coded after entropy coding. In some cases, MV prediction can be an example of reversible compression of a signal (i.e., an MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be irreversible, for example, due to rounding errors when calculating predictors from some surrounding MVs.

[0019] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms proposed by H.265, in this specification, the technique hereinafter referred to as "spatial merge" will be described with reference to FIG. 2.

[0020] Referring to FIG. 2, the current block (201) has samples recognized by the encoder during the motion search process as being predictable from a previous block of the same size that has been spatially shifted. Instead of directly coding the MV, the MV can be derived from metadata associated with one or more reference pictures, for example, from the most recent reference picture (in decoding order) using an MV associated with any one of five surrounding samples represented as A0, A1, and B0, B1, B2 (202 to 206 respectively). In H.265, MV prediction can use predictors from the same reference picture that adjacent blocks are using. SUMMARY OF THE INVENTION

[0021] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video decoding includes a receiving circuit and a processing circuit. The receiving circuit receives a coded video bitstream including a current picture. The current picture includes chroma blocks within a chroma separation tree, and the chroma blocks are collocated within the same luma area as one or more luma blocks. The processing circuit decodes a syntax element indicating a current picture referencing (CPR) mode of a chroma block from the coded video bitstream, and in response to the CPR mode, determines a chroma block vector of the chroma block according to one or more luma block vectors associated with one or more luma blocks. The chroma block vector indicates a reference chroma block within the current picture. The processing circuit reconstructs the chroma block based on the reference chroma block of the current block within the current picture.

[0022] In some examples, the processing circuit determines a block vector predictor according to one or more luma block vectors, and decodes a block vector difference from a coded video bitstream. The processing circuit determines a chroma block vector based on the block vector predictor and the block vector difference.

[0023] In some examples, the processing circuit derives a block vector predictor from at least one of an average value of one or more luma block vectors and a weighted average value of one or more luma block vectors.

[0024] In some examples, the processing circuit determines a first luma block from one or more luma blocks, where the first luma block includes a sample point corresponding to a specific sample position of a chroma block. The processing circuit derives a block vector predictor from a luma block vector associated with the first luma block. In an example, the specific sample position is a center sample position of the chroma block. In other examples, the specific sample position is an upper left sample position of the chroma block.

[0025] In some examples, the processing circuit decodes an index indicating a first luma block vector from one or more luma block vectors from a coded video bitstream, and derives a block vector predictor from the first luma block vector.

[0026] In some examples, the processing circuit determines a first accuracy of a chroma block vector from candidates coarser than a second accuracy of one or more luma block vectors.

[0027] In some examples, the processing circuit derives the chroma vector vector from at least one of an average value of one or more luma block vectors and a weighted average value of one or more luma block vectors.

[0028] In some examples, the processing circuit determines a first luma block from one or more luma blocks, and the first luma block includes a sample point corresponding to a specific sample position of a chroma block. The processing circuit derives a chroma block vector from a luma block vector associated with the first luma block. In an example, the specific sample position is the center sample position of the chroma block. In other examples, the specific sample position is the upper left sample position of the chroma block.

[0029] In some examples, the processing circuit decodes an index indicating a first luma block vector from one or more luma block vectors, and derives a chroma block vector from the first luma block vector.

[0030] In some examples, the luma area corresponding to a chroma block includes a plurality of luma blocks each having a luma block vector, and the processing circuit derives a candidate chroma block vector from the luma block vector, determines a candidate reference template corresponding to the current template of the chroma block according to the candidate chroma block vector, calculates a template matching cost associated with each candidate chroma block vector according to the distortion between the current template and the candidate reference template, selects a chroma block vector from the candidate chroma block vectors based on the template matching cost, and the chroma block vector has the minimum template matching cost among the candidate chroma block vectors.

[0031] In some examples, the processing circuit orders candidate chroma block vectors according to the template matching cost into a list of ordered candidate chroma block vectors, decodes an index indicating a chroma block vector from the bit stream from the list of ordered candidate chroma block vectors, and selects a chroma block vector from the list of ordered candidate chroma block vectors according to the index.

[0032] In some examples, the processing circuit derives an initial chroma block vector according to one or more luma block vectors and performs a template matching search starting from the initial chroma block vector to determine the chroma block vector.

[0033] In some examples, the processing circuit derives the initial chroma block vector from at least one of the average value of one or more luma block vectors and the weighted average value of one or more luma block vectors.

[0034] In some examples, the processing circuit determines a first luma block having sample points corresponding to specific sample positions of the chroma block from one or more luma blocks and derives the initial chroma block vector from the luma block vector associated with the first luma block. In the example, the specific sample position is the center sample position of the chroma block. In other examples, the specific sample position is the upper left sample position of the chroma block.

[0035] In the example, the processing circuit decodes an index indicating the first luma block vector from one or more luma block vectors and derives the initial chroma block vector from the first luma block vector.

[0036] In some examples, performing a template matching search involves, for an intermediate chroma block vector, the processing circuit determining an intermediate chroma reference template corresponding to the current chroma template of the chroma block according to the intermediate chroma block vector, determining an intermediate luma reference template collocated with the intermediate chroma reference template, calculating a first template matching cost according to the distortion between the current chroma template and the intermediate chroma reference template, calculating a second template matching cost according to the distortion between the current luma template collocated with the current chroma template and the intermediate luma reference template, and calculating a combined template matching cost associated with the intermediate chroma block vector by combining the first template matching cost with the second template matching cost.

[0037] The disclosed aspects also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a method of video decoding.

[0038] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0039]

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

DETAILED DESCRIPTION OF THE INVENTION

[0040] FIG. 3 represents an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) may encode video data (e.g., a stream of video data captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to recover a video picture, and display the video picture according to the recovered video data. Unidirectional data transmission can be common in media serving applications and the like.

[0041] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, for example, during a video conference. For the bidirectional transmission of data, in the example, each of the terminal devices (330) and (340) may encode video data (e.g., a stream of video pictures captured by that terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) may also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), may decode the encoded video data to recover the video pictures, and may display the video pictures on an accessible display device according to the recovered video data.

[0042] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) are each represented as a server, a personal computer, and a smartphone, but the principles of the present disclosure cannot be so limited. Embodiments of the present disclosure find use in laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) corresponds to any number of networks that transmit encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wireline (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (350) can be irrelevant to the operation of the present disclosure unless otherwise described hereinafter.

[0043] FIG. 4 depicts a video encoder and a video decoder in a streaming environment as an application example of the disclosed subject matter. The disclosed subject matter can be similarly applicable to other video-related applications including, for example, storage of compressed video on digital media including video conferencing, digital TV, streaming services, CDs, DVDs, memory sticks, etc.

[0044] A streaming system may include a capture subsystem (413) that includes, for example, a video source (401) that generates a stream (402) of uncompressed video pictures, such as a digital camera. In the example, the stream (402) of video pictures includes samples taken by a digital camera. The stream (402) of video pictures is represented by a thick line to emphasize its high data volume compared to the encoded video data (404) (or coded video bitstream), and can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement the disclosed aspects of the subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream) is represented by a thin line to emphasize its lower data volume compared to the stream (402) of video pictures, and can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to read copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include a video decoder (410), for example, in an electronic device (430). The video decoder (410) decodes an incoming copy (407) of the encoded video data and generates an outgoing stream (411) of video pictures that can be rendered on a display (412) (such as a display screen) or another rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (such as a video bitstream) can be encoded according to a particular video coding / compression standard. An example of such a standard is ITU-T Recommendation H.265.In an example, the video coding standard under development is commonly known as Versatile Video Coding (VVC). The subject matter disclosed may be used in relation to VVC.

[0045] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can similarly include a video encoder (not shown).

[0046] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used instead of the video decoder (410) in the example of FIG. 4.

[0047] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In an embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data together with other data, such as coded audio data and / or auxiliary data streams, which may be transferred to their respective using entities (not shown). The receiver (531) may separate the coded video sequence from other data. To counter network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter “parser (520)”). For certain applications, the buffer memory (515) is part of the video decoder (510). Otherwise, it can be outside the video decoder (510) (not shown). Still otherwise, for example, there can be a buffer memory outside the video decoder (not shown) for countering network jitter, and in addition, another buffer memory (515) within the video decoder (510) for, for example, manipulating the playback timing. When the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be needed, or may be small. For use in a best effort packet network such as the Internet, the buffer memory (515) may be needed, may be relatively large, and advantageously may be of an adaptable size and may be implemented at least in part in an operating system or similar element outside the video decoder (510) (not shown).

[0048] Video decoder (510) may include a parser (520) for reconstructing symbols (521) from the coded video sequence. The categories of those symbols include information used to manage the operation of the video decoder (510) and, potentially, information for controlling a rendering device such as a rendering device (512) (e.g., a display screen) that may be coupled to the electronic device (530) although not an essential part of the electronic device (530) as shown in FIG. 5. The control information for the rendering device may take the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context dependence, etc. The parser (520) may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to that group. The subgroups can include group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (520) may also extract quantization parameter values, motion vectors, etc. from the coded video sequence information such as transform coefficients.

[0049] The parser (520) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (515) to generate the symbols (521).

[0050] The reconstruction of symbol (521) can have a number of different units depending on the type of the coded video picture or a portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How the units are included can be controlled by subgroup control information parsed by parser (520) from the coded video sequence. The flow of such subgroup control information between parser (520) and the following multiple units is not shown for clarity.

[0051] Beyond the function blocks already described, video decoder (510) can conceptually be subdivided into a number of functional units described below. In an actual implementation operating under commercial constraints, many of those units interact closely with each other and can be at least partially incorporated into each other. However, for the purpose of explaining the disclosed subject matter, the conceptual subdivision into functional units below is appropriate.

[0052] The first unit is a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbols (521) from parser (520), the quantized transform coefficients together with control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block including sample values that can be input to aggregator (555).

[0053] In some cases, the output samples of the scaler / inverse converter (551) can be related to the intra-coded blocks. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture and can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be supplied by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information fetched from the current picture buffer (558). The current picture buffer (558) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (555) adds, in some cases, for each sample, the prediction information generated by the intra prediction unit (552) to the output sample information supplied by the scaler / inverse conversion unit (551).

[0054] In other cases, the output samples of the scaler / inverse transform unit (551) can relate to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. According to the symbols (521) related to the block, after motion-compensating the fetched samples, those samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called the residual samples or residual signal) to generate output sample information. The address in the reference picture memory (557) where the motion compensation prediction unit (553) fetches the prediction samples can be controlled by the motion vectors that the motion compensation prediction unit (553) can utilize in the form of symbols (521) that can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (557) when exact sub-sample motion vectors are used, a motion vector prediction mechanism, and the like.

[0055] The output samples of the aggregator (555) can undergo various loop filtering techniques in the loop filter unit (556). Video compression techniques can include in-loop filter techniques. This technique is included in the coded video sequence (also called the coded video bitstream) and is controlled by the parameters made available to the loop filter unit (556) as symbols (521) from the parser (520). Video compression can also respond to the meta information obtained during the decoding of the previous part of the coded picture or coded video sequence (in the decoding order), and further, can also respond to the previously configured loop filter processed sample values.

[0056] The output of the loop filter unit (556) can be a sample stream that is output to the rendering device (512) and further stored in the reference picture memory (557) for use in future inter-picture prediction.

[0057] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and the unused current picture buffer can be reallocated before starting the reconstruction of subsequent coded pictures.

[0058] The video decoder (510) may perform a decoding operation according to a predetermined video compression technology or standard specification such as ITU-T Recommendation H.265. The coded video sequence may conform to the syntax defined by the video compression technology or standard specification in use, in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard specification and the profile documented in the video compression technology or standard specification. Specifically, the profile can select specific tools as the only tools available for use under that profile from all the tools available in the video compression technology or standard specification. Also, the complexity of the coded video sequence needs to be within the bounds defined by the level of the video compression technology or standard specification for compliance. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted through the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the coded video sequence.

[0059] In an embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may also be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can take forms such as, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0060] FIG. 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used in place of the video encoder (403) of the example of FIG. 4.

[0061] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that may capture a video image to be coded by the video encoder (603). In other examples, the video source (601) is part of the electronic device (620).

[0062] The video source (601) may supply a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device storing pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be supplied as a plurality of individual pictures that impart motion when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, and each pixel may have one or more samples depending on the sampling structure, color space, etc. in use. One of ordinary skill in the art can readily understand the relationship between pixels and samples. This specification will hereinafter focus on samples.

[0063] According to an embodiment, the video encoder (603) can code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other required time constraints. Enforcing an appropriate coding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to the other functional units. The couplings are not shown for clarity. Parameters set by the controller (650) can include parameters related to rate control (picture skip, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.

[0064] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified description, in the example, the coding loop may involve a source coder (630) (e.g., generating symbols such as a symbol stream based on an input picture to be coded and a reference picture), and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in the same way as a (remote) decoder would also generate. The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in a bit-exact result independent of the location of the decoder (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the reference picture samples that the decoder would "see" when using prediction during decoding. This basic principle of the synchrony of the reference picture (and the resulting drift if the synchrony cannot be maintained, e.g., due to channel errors) is also used in some related technologies.

[0065] The operation of the "local" decoder (633) can be the same as that of a "remote" decoder such as the video decoder (510), which has already been described in detail previously with reference to FIG. 5. However, referring temporarily also to FIG. 5, since the symbols are available and the encoding / decoding of the symbols to the coded video sequence by the entropy encoder (645) and the parser (520) can be reversible, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).

[0066] In an embodiment, decoder techniques, excluding parsing / entropy decoding that exists in a decoder, exist in a corresponding encoder in the same or substantially the same functional form. Accordingly, the disclosed subject matter focuses on the operation of the decoder. The description of encoder techniques may be omitted since they are the reverse of the decoder techniques described comprehensively. In certain ranges, more detailed descriptions are given below.

[0067] During operation, in some examples, the source coder (630) may perform motion-compensated predictive coding. This predictively codes an input picture by referring to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (632) codes the difference between a pixel block of a reference picture that can be selected as a prediction reference for the input picture and a pixel block of the input picture.

[0068] The local video decoder (633) may decode the coded video data of a picture that can be designated as a reference picture based on the symbols generated by the source coder (630). The operation of the coding engine (632) may advantageously be an irreversible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence is usually a reproduction of the source video sequence with some errors. The local video decoder (633) may reproduce the decoding process that can be performed by the video decoder for the reference picture and store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture having the same content as the reconstructed reference picture that would be obtained by a remote video decoder (without transmission errors).

[0069] Predictor (635) can perform predictive search for the coding engine (632). That is, for a new picture to be coded, Predictor (635) can search the reference picture memory (634) for specific metadata such as reference picture motion vectors, block shapes, etc. that can serve as appropriate prediction criteria for that new picture, or sample data (as candidate reference pixel blocks). Predictor (635) may operate on a sample block-by-pixel block basis to find appropriate prediction criteria. In some cases, the input picture may have prediction criteria drawn from a plurality of reference pictures stored in the reference picture memory (634), as determined by the search results obtained by Predictor (635).

[0070] Controller (650) may manage the coding operations of source coder (630), including, for example, setting parameters and subgroup parameters used for encoding video data.

[0071] The outputs of all the above functional units can undergo entropy coding in entropy coder (645). Entropy coder (645) converts the symbols generated by various functional units into a coded video sequence by reversibly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0072] The transmitter (640) may buffer the coded video sequence generated by the entropy coder (645) for transmission via the communication channel (660). The communication channel (660) may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).

[0073] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign to each coded picture a particular coded picture type that may affect the coding technique applicable to each picture. For example, a picture may often be assigned as one of the following picture types.

[0074] An Intra Picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow various types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will be aware of such variations of I pictures and their respective applications and characteristics.

[0075] A Predictive Picture (P picture) may be a picture that can be encoded and decoded by intra prediction or inter prediction using at most one motion vector and a reference index to predict the sample values of each block.

[0076] A bi-directionally predictive picture (B picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple-predictive picture(s) can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0077] A source picture may generally be spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each), and each block may be coded. The blocks may be coded predictively with reference to other (already coded) blocks determined by the coding assignment applied to each of the pictures of the block. For example, blocks of an I picture may be coded non-predictively, or they may be coded predictively with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be coded predictively by spatial prediction or temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be coded predictively by spatial prediction or temporal prediction with reference to one or two previously coded reference pictures.

[0078] Video encoder (603) may perform coding operations according to a given video coding technology or standard specification such as ITU-T Recommendation H.265. During the operation, video encoder (603) may perform various compression operations including predictive coding operations that utilize temporal and spatial redundancies in the input video sequence. Accordingly, the coded video data may conform to the syntax defined by the video coding technology or standard specification being used.

[0079] In an embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may have temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0080] Video may be captured as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes spatial correlation within a given picture, and inter-picture prediction utilizes correlation (temporal or otherwise) between pictures. In an example, a particular picture being encoded / decoded, referred to as the current picture, is partitioned into blocks. If a block within the current picture is similar to a reference block within a reference picture that was previously coded and is still buffered within the video, that block within the current picture may be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture if multiple reference pictures are being used.

[0081] In some embodiments, dual prediction techniques may be used in inter-picture prediction. According to the dual prediction technique, two reference pictures, e.g., a first reference picture and a second reference picture that both precede the current picture in decoding order within the video (however, in display order, they may be in the past and future respectively), are used. A block within the current picture may be coded by a first motion vector that points to a first reference block within the first reference picture and a second motion vector that points to a second reference block within the second reference picture. The block is predictable by a combination of the first reference block and the second reference block.

[0082] Furthermore, the merge mode technique can be used in inter-picture prediction to improve coding efficiency.

[0083] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard specification, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs) which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In the example, each CU is analyzed to determine a prediction type for the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In an embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0084] FIG. 7 shows an example of a video encoder (703). The video encoder (703) receives a processing block (e.g., a prediction block) of sample values within a current video picture included in a sequence of video pictures, and is configured to encode the processing block into a coded picture that is part of a coded video sequence. In the example, the video encoder (703) is used instead of the video encoder (403) in the example of FIG. 4.

[0085] In an example of HEVC, the video encoder (703) receives a matrix of sample values of a processing block such as a prediction block of 8×8 samples. The video encoder (703) determines, for example using rate distortion optimization, whether the processing block is best coded in an intra mode, an inter mode, or a bi-prediction mode. If the processing block is to be coded in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into the coded picture, and if the processing block is to be coded in the inter mode or the bi-prediction mode, the video encoder (703) may use inter prediction or bi-prediction techniques respectively to encode the processing block into the coded picture. In certain video coding techniques, the merge mode can be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictor. In certain other video coding techniques, there may be coded motion vector components applicable to the target block. In the example, the video encoder (703) includes other components such as a mode decision module (not shown) that determines the mode of the processing block.

[0086] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculation unit (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725) that are coupled as shown in FIG. 7.

[0087] The inter-encoder (730) receives samples of the current block (e.g., a processing block), compares the block with one or more reference blocks (e.g., blocks in the previous and subsequent pictures) in the reference picture, generates inter-prediction information (e.g., a description of redundant information according to an inter-coding technique, a motion vector, merge mode information), and is configured to calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using some suitable technique. In some examples, the reference picture is a decoded reference picture that has been decoded based on the encoded video information.

[0088] The intra-encoder (722) receives samples of the current block (e.g., a processing block), and in some cases, compares the block with blocks that have already been coded within the same picture, and also generates quantized coefficients after transformation, and in some cases, also generates intra-prediction information (e.g., intra-prediction direction information according to one or more intra-coding techniques). In an example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks within the same picture.

[0089] The general-purpose controller (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In an example, the general-purpose controller (721) determines the mode of a block and supplies a control signal to the switch (726) based on the mode. For example, when the mode is the intra mode, the general-purpose controller (721) controls the switch (726) to select the intra mode result for use by the residual calculation unit (723), and selects the intra prediction information and controls the entropy encoder (725) to include the intra prediction information in the bitstream. When the mode is the inter mode, the general-purpose controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculation unit (723), and selects the inter prediction information and controls the entropy encoder (725) to include the inter prediction information in the bitstream.

[0090] The residual calculation unit (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data so as to generate transform coefficients. In an example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. Next, the transform coefficients are subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown.) and can be used as a reference picture in some examples.

[0091] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to an appropriate standard such as the HEVC standard. In an example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. Note that there is no residual information when coding a block in either the merge submode of the inter mode or the bi-prediction mode according to the disclosed subject matter.

[0092] FIG. 8 shows an example of a video decoder (810). The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence, decode the coded picture, and generate a reconstructed picture. In the example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.

[0093] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) that are coupled as shown in FIG. 8.

[0094] The entropy decoder (871) may be configured to reconstruct specific symbols representing syntax elements from the coded picture, from which the coded picture is composed. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, or inter mode or bi-prediction mode in a merge sub-mode or other sub-mode), and prediction information (e.g., intra prediction information or inter prediction information) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), respectively. The symbols can also include residual information in the form of, for example, quantized transform coefficients. In the example, when the prediction mode is inter or bi-prediction mode, the inter prediction information is supplied to the inter decoder (880), and when the prediction type is intra prediction type, the intra prediction information is supplied to the intra decoder (872). The residual information can undergo inverse quantization and is supplied to the residual decoder (873).

[0095] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0096] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0097] The residual decoder (873) performs inverse quantization to extract the inverse quantized transform coefficients, processes the inverse quantized transform coefficients, and is configured to convert the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also request specific control information (for including quantization parameter (QP)), and that information may be supplied by the entropy decoder (871) (this is only low-capacity control information and the data path is not shown).

[0098] The reconstruction module (874) is configured to combine, in the spatial domain, the residual output by the residual decoder (873) and the prediction result (optionally output by an inter or intra prediction module) to form a reconstructed block. The reconstructed block may be a part of the reconstructed picture, and then the reconstructed picture may be a part of the reconstructed video. Note that other appropriate operations such as a deblocking operation may be performed to improve visual quality.

[0099] Note that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented by any appropriate technology. In an embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In other embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.

[0100] Aspects of the present disclosure provide techniques for enabling the current picture reference (e.g., IBC, IntraBC, IntraTMP) mode for chroma components when a separate intraluma / chrominance coding tree structure (intra dual tree) is used.

[0101] The IBC mode has applications in various video codecs such as HEVC, VVC, AOMedia Video 1 (AV1), etc. In some examples, techniques can derive a chroma BV from a luma block vector (BV).

[0102] Some IBC coding tools are used in the HEVC screen content coding (SCC) extension as a current picture reference (CPR). The IBC mode can use coding techniques that are used for inter prediction where the current picture is used as a reference picture in the IBC mode. The advantage of using the IBC mode is the reference structure of the IBC mode where 2 - dimensional (2D) spatial vectors can be used as an expression of the addressing mechanism to reference samples. The advantage of the IBC mode architecture is that the integration of IBC requires relatively few changes to the specification and can reduce the implementation burden when manufacturers already have a specific inter - prediction technology such as HEVC version 1 implemented. CPR in the HEVC SCC extension can be a special inter - prediction model that results in a syntax structure identical to that of the inter - prediction mode and a decoding process similar to the decoding process of the inter - prediction mode.

[0103] The IBC mode can be integrated into the inter prediction process. In some examples, the IBC mode (or CPR) is an inter prediction mode, and the intra-only prediction slice should be a prediction slice that allows the use of the IBC mode. When the IBC mode is applicable, the coder can extend the reference picture list with only one entry for the pointer to indicate the current picture. For example, the current picture uses a buffer of one picture size in the shared Decoded Picture Buffer (DPB). The IBC mode signaling can be implicit. For example, if the selected reference picture indicates the current picture, the CU() can use the IBC mode. In various embodiments, the reference samples used in the IBC process are not filtered, which is different from normal inter prediction. The corresponding reference picture used in the IBC process is a long-term reference. To minimize memory requirements, the coder can release the buffer after reconstructing the current picture. For example, the coder can release the buffer immediately after reconstructing the current picture. The filtered version of the reconstructed picture can be returned to the DPB by the coder as a short-term reference if the reconstructed picture is a reference picture.

[0104] In block vector (BV) coding, referring to the reconstructed area can be performed by a 2D BV similar to that in inter prediction. The prediction and coding of BV can reuse the prediction and coding of the MV in the inter prediction process. In some examples, the luma BV is at integer resolution rather than 1 / 4 the accuracy of the MV used for a normally inter-coded CTU.

[0105] FIG. 9 shows the BV associated with the current CU (901) according to an embodiment of the present disclosure. Each square (900) can represent a CTU. The shaded area represents an area that has already been coded (e.g., an area that has already been encoded), and the area without white shading represents an area to be coded (e.g., an area to be encoded). The current CTU (900(4)) being reconstructed includes the current CU (901), the coded area (902), and the area scheduled for coding (903). In the example, the area (903) is coded after the coding of the current CU (901).

[0106] In the example, for example, in HEVC, the shaded area excluding the two CTUs ((900(1) to 900(2)) above the upper right of the current CTU (900(4)) can be used as a reference area in the IBC mode to allow wavefront parallel processing (WPP). The BV allowed in HEVC can indicate a block within the reference area (the shaded area excluding the two CTUs (900(1) to 900(2))). For example, the BV allowed in HEVC indicates the reference block (911).

[0107] In the example, for example, in VVC, in addition to the current CTU (900(4)), the left adjacent CTU (900(3)) to the left of the current CTU (900(4)) is allowed as a reference area in the IBC mode. In the example, the reference area used in the IBC mode in VVC is within the dashed line area (915) and includes the coded samples. For example, the BV allowed in VVC indicates the reference block (912).

[0108] In some examples, the decoded motion vector difference (Motion Vector Difference, MVD) of the BV (also referred to as BV difference (BVD)) can be left-shifted by 2 before being added to the corresponding BV predictor to reconstruct the final BV.

[0109] In some embodiments, special handling of the IBC mode may be required for implementation and execution, and the IBC mode and the inter prediction mode (e.g., the normal inter prediction mode) may be different, as described below. In an example, the reference samples used in the IBC mode are not filtered (e.g., the reconstructed samples before the in-loop filtering process such as DBF and sample adaptive offset (SAO) filters are applied). Other inter prediction modes of HEVC (e.g., the normal inter prediction mode) can use filtered samples, e.g., reference samples filtered by the in-loop filtering process.

[0110] In some examples, luma sample interpolation is not performed in the IBC mode. Chroma sample interpolation can be performed in the IBC mode. In some examples, chroma sample interpolation is only necessary when the chroma BV is non-integer when the chroma BV is derived from the corresponding luma BV. In some examples, luma sample interpolation and chroma sample interpolation can be performed in the normal inter prediction mode.

[0111] In the IBC mode, in special cases, it may occur when the chroma BV is a non-integer BV and the reference block is close to the boundary of the available area (e.g., the reference area). For example, the surrounding reconstructed samples may be outside the boundary for performing chroma interpolation. In an example, the BV indicating a single line adjacent to the boundary may result in the surrounding reconstructed samples being outside the boundary.

[0112] According to aspects of the present disclosure, the IBC architecture in VVC has unique features.

[0113] For the IBC mode in the HEVC SCC extension, the valid reference area can include the entire already reconstructed area of the current picture with some exceptions for parallel processing, as described in FIG. 9. A drawback of the reference area used in HEVC is the need for additional memory in the DPB, for which hardware implementations may use external memory. Additional accesses to external memory can increase the memory bandwidth and may reduce the attractiveness of using the DPB. In some embodiments, on-chip fixed memory (e.g., memory with a fixed size) that can be implemented for the IBC mode can be used in VVC. The on-chip fixed memory in the IBC mode can significantly reduce the complexity of implementing the IBC mode in the hardware architecture. In an example, the on-chip fixed memory in the IBC mode can reduce latency. In some examples, the changes address signaling concepts that deviate from the integration within the inter-prediction process as seen in the HEVC SCC extension.

[0114] In the examples shown in FIGS. 10A - 10D, fixed memory can be allocated to store the reference area used in the IBC mode. The fixed memory may be referred to as Reference Sample Memory (RSM). A portion of the RSM can be updated at different intermediate points during the coding process (encoding process or reconstruction process). FIGS. 10A - 10D show the RSM update process at various intermediate points during the coding process (e.g., encoding process or reconstruction process) according to embodiments of the present disclosure. FIGS. 10A - 10D show the reference area for the IBC mode in VVC and the configuration in VVC.

[0115] Referring to FIGS. 10A to 10D, currently, CTU (1020) is adjacent to a CTU (for example, the left adjacent CTU) (1010) on the left side of the current CTU (1020). In some examples, the current CTU (1020) includes four areas (1021) to (1024). The left adjacent CTU (1010) can include four areas (1011) to (1014) corresponding to the areas (1021) to (1024), respectively. The positions of the areas (1011) to (1014) are shifted to the left by the width of the CTU (1020) from the positions of the areas (1021) to (1024), respectively. The RSM can include the position of the current CTU (1020) and / or the position of the left adjacent CTU (1010). In the example shown in FIGS. 10A to 10D, the size of the RSM is equal to the size of the CTU. The brightly shaded area can include the reference samples of the current CTU (1020), and the area without white shading can represent the area to be coded (the area scheduled for coding).

[0116] Referring to FIG. 10A, at the first intermediate point of the coding process, which is the beginning of the coding process of the current CTU (1020), the RSM includes the entire left adjacent CTU (1010), and the entire left adjacent CTU (1010) can be the reference area in the IBC mode at the start of the coding process of the current CTU (1020). The RSM at the start of the coding process of the current CTU (1020) does not include any of the areas (1021) to (1024).

[0117] Referring to FIG. 10B, area (1021) includes sub-areas (1031) to (1033). Sub-area (1031) has already been coded (e.g., encoded or reconstructed), sub-area (1032) is the current CU being coded (e.g., during encoding or reconstruction), and sub-area (1033) is to be coded later. At the second intermediate point of the coding process of the current CTU (1020) where the sub-area (1032) of the current CTU (1020) is being coded, the RSM is updated to include a part of the left adjacent CTU (1010) and a part of the current CTU (1020). For example, the RSM includes areas (1012) to (1014) of the left adjacent CTU (1010) and sub-area (1031) of the current CTU (1020). The reference area at the second intermediate point includes areas (1012) to (1014) of the left adjacent CTU (1010) and sub-area (1031) of the current CTU (1020).

[0118] Referring to FIG. 10C, area (1022) includes sub-areas (1041) to (1043). Sub-area (1041) (shaded dark gray) has already been coded (e.g., encoded or reconstructed), sub-area (1042) is the current CU being coded (e.g., during encoding or reconstruction), and sub-area (1043) (white) is to be encoded later. At the third intermediate point of the coding process of the current CTU (1020) where the sub-area (1042) of the current CTU (1020) is being coded, the RSM is updated to include (i) areas (1013) to (1014) of the left adjacent CTU (1010) and (ii) area (1021) and sub-area (1041) of the current CTU (1020). Within the RSM, area (1012) is replaced by sub-area (1041). The reference area at the third intermediate point can include (i) areas (1013) to (1014) of the left adjacent CTU (1010) and (ii) area (1021) and sub-area (1041) of the current CTU (1020).

[0119] Referring to FIG. 10D, area (1024) includes sub-areas (1051) to (1053). Sub-area (1051) (shaded dark gray) has already been coded (e.g., encoded or reconstructed), sub-area (1052) is the current CU being coded (e.g., during encoding or reconstruction), and sub-area (1053) (white) is to be coded later. At the fourth intermediate point in the coding process of the current CTU (1020) where the sub-area (1052) of the current CTU (1020) is being coded, the RSM is updated to include areas (1021) to (1023) and sub-area (1051) of the current CTU (1020). The RSM at the fourth intermediate point does not include the area within the left adjacent CTU (1010). The reference area at the fourth intermediate point can include areas (1021) to (1023) and sub-area (1051) of the current CTU (1020).

[0120] According to an aspect of the present disclosure, VVC has a specific syntax and semantics for the IBC mode.

[0121] The IBC architecture in VVC can form a dedicated coding mode. In addition to the intra prediction mode and the inter prediction mode (e.g., the normal inter prediction mode), the IBC mode is the third prediction mode. The bitstream can include, for example, an IBC syntax element indicating the IBC mode for a CU when the size of the CU is equal to or smaller than 64×64. In some examples, the maximum CU size for which the IBC mode can be used is 64×64 to implement a continuous memory update mechanism of RSM as described with reference to FIGS. 10A - 10D. In an example, the reference sample addressing mechanism is the same as that used in the HEVC SCC extension by expressing a 2D offset and reusing the vector (e.g., MV) coding process of the inter prediction mode. In an example, when the Chroma Separate Tree (CST) is active, the coder cannot derive the chroma BV from the corresponding luma BV, and as a result, the IBC mode is used only for the luma CB.

[0122] The IBC design in VVC can use a fixed memory size (e.g., 128×128) for each color component to store reference samples. As described above, the fixed memory size can enable on - chip placement of the memory (e.g., RSM) in hardware implementation. In an example, for example, in VVC, the maximum CTU size and the fixed memory size for the IBC mode are 128×128. In an example, when the maximum CTU size is equal to the fixed memory size (e.g., 128×128) for the IBC mode, the RSM includes samples of a single CTU.

[0123] As depicted in FIGS. 10A to 10C, the characteristic of RSM is a continuous update mechanism that replaces the reconstructed samples of the left adjacent CTU with the reconstructed samples of the current CTU. FIGS. 10A to 10C show a simple example of RSM for the update mechanism at four intermediate time points during the coding process (e.g., the reconstruction process). The brightly shaded areas in FIGS. 10A to 10C can include the reference samples of the left adjacent CTU (1010), and the darkly shaded areas in FIGS. 10B to 10D can include the reference samples of the current CTU (1020). Referring to FIG. 10A, at the first intermediate time point representing the start of the coding (e.g., encoding or reconstruction) of the current CTU (1020), RSM consists only of the reference samples of the left adjacent CTU (1010). At the other three intermediate time points shown in FIGS. 10B to 10D, the coding process (e.g., the encoding process or the reconstruction process) replaces the samples of the left adjacent CTU (1010) with the samples within the current CTU (1020).

[0124] In some examples, the RSM is implicitly divided into four areas, such as four non - overlapping areas of 64×64. Reset of the areas within the RSM can occur when the coder processes the first CU within the corresponding area in the current CTU, which eases the hardware implementation effort. For example, the RSM is mapped to areas within the CTU (e.g., the left - adjacent CTU and the current CTU). FIG. 11 spatially shows the continuous update process (1100) of the RSM. The left - adjacent CTU (1010) and the current CTU (1020) are described in FIGS. 10A - 10D. The left - adjacent CTU (1010) can include areas (1011)-(1014). The current CTU (1020) can include areas (1021)-(1024). Area (1023) within the current CTU (1020) includes the current CU (1152) being coded, the sub - area (1151) that has already been coded, and the sub - area (1153) that is going to be coded. The gray - shaded areas can include the samples stored in the RSM, and the non - shaded white areas can include the replaced samples or the samples that have not been coded (e.g., the samples that have not been reconstructed).

[0125] At the coding time (e.g., reconstruction time) shown in FIG. 11, the RSM update process replaces the samples covered by the non - shaded white areas (e.g., areas (1011)-(1031)) in the left - adjacent CTU (1010) with the gray - shaded areas (e.g., areas (1021)-(1022) and sub - area (1051)) of the current CTU (1020). In FIG. 11, the RSM can include (i) area (1014) in the left - adjacent CTU (1010) and (ii) areas (1021)-(1022) and sub - area (1051) of the current CTU (1020).

[0126] In some examples, when the maximum CU size is smaller than the RSM size (e.g., 128×128), the RSM may include more than a single left-adjacent CTU, and multiple adjacent CTUs may be used as a reference area in the IBC mode. For example, when the maximum CTU size is 32×32, an RSM having a size of 128×128 can include samples of 15 adjacent CTUs.

[0127] In VVC, BV coding in the IBC mode can use a process specific to inter prediction (e.g., normal inter prediction). BV coding can use rules that are simpler than the rules used in inter prediction (e.g., normal inter prediction) to construct a candidate list.

[0128] For example, the candidate list for inter prediction includes 5 spatial candidates, 1 temporal candidate, and candidates based on 6 histories. Multiple candidate comparisons can be used for candidates based on history to avoid duplicate entries in the final candidate list for inter prediction. The candidate list for inter prediction may include pairwise averaged candidates.

[0129] The candidate list for the IBC mode can include two BVs from each spatial neighborhood and a BV based on 5 histories (HBVP). In the example, the candidate list for the IBC mode is limited to two BVs from each spatial neighborhood and a BV based on 5 histories (HBVP). In an embodiment, in the IBC mode, only the first HBVP is compared with the spatial candidates when the first HBVP is added to the candidate list.

[0130] The normal inter prediction mode can use two different candidate lists. For example, one is the candidate list for the merge mode, and the other is the candidate list for the normal mode (e.g., the inter prediction mode that is not the merge mode). The candidate list in the IBC mode can be the same for both IBC modes (e.g., the merge IBC mode and the normal IBC mode). In the IBC mode, the merge mode may use up to the first 6 candidates in the candidate list, and the normal mode uses only the first 2 candidates in the candidate list.

[0131] Block vector difference (BVD) coding can use the MVD process used in the normal inter prediction mode, and the final BV can have any size. The determined BV (e.g., the reconstructed BV) may point to an area outside the reference sample area. In an example, the correction for the absolute offset in each direction can be applied using a modulo operation based on the width and / or height of the RSM.

[0132] According to an aspect of the present disclosure, the block vector of a chroma block can be derived from the block vector of a luma block in some examples.

[0133] In some examples, when the current coding tree type is SINGLE_TREE, the chroma block always has a corresponding luma block. In the IBC mode, the BV of the chroma block can be derived from the BV of the corresponding luma block using appropriate scaling according to the chroma sampling format (e.g., 4:2:0, 4:2:2) and the chroma BV precision.

[0134] In some examples, a derivation process is used to derive the BV of a chroma block from the BV of the corresponding luma block. The input to the derivation process includes a luma block vector with 1 / 16 fractional sample accuracy (bvL represents the luma block vector, bvL[0] represents the x component, and bvL[1] represents the y component), and the output of the derivation process includes a chroma block vector with 1 / 32 fractional sample accuracy (bvC represents the chroma block vector, bvC[0] represents the x component, and bvC[1] represents the y component).

[0135] In some examples, the chroma block vector is derived from the corresponding luma block vector according to equations (1) and (2):

Equation

Table 1

[0136] For example, when sps_chroma_format_idc is equal to 0, the chroma format is monochrome sampling format, and there is only one sample array which is usually regarded as the luma array. When sps_chroma_format_idc is equal to 1, the chroma format is 4:2:0 sampling format, and each of the two chroma arrays has half the height and half the width of the luma array. When sps_chroma_format_idc is equal to 2, the chroma format is 4:2:2 sampling format, and each of the two chroma arrays has the same height and half the width of the luma array. When sps_chroma_format_idc is equal to 3, the chroma format is 4:4:4 sampling format, and each of the two chroma arrays has the same height and width as the luma array.

[0137] In some examples, the number of bits required for the representation of each sample in the luma and chroma arrays in a video sequence is in the range from 8 to 16.

[0138] In some examples, for example in AV1, the IBC mode is called the IntraBC mode, and the BV is used to find the prediction block within the same picture of the current block. The BV can be signaled in the bitstream, and the accuracy of the signaled BV is at integer points. The prediction process in the IBC mode can be similar to the prediction process in the inter prediction mode (e.g., inter picture prediction). The differences between the IBC mode and the inter picture prediction are described as follows. In the IBC mode, the predictor block can be formed from the reconstructed samples of the current picture (e.g., before applying loop filtering). The IBC mode is regarded as "motion compensation" within the current picture using the BV as the MV.

[0139] In AV1, a flag indicating whether the IBC mode is valid for the current block can be transmitted in the bitstream. If the IBC mode is valid for the current block, the BV difference can be derived by subtracting the predicted BV from the current BV, and the BV difference can be classified into four types according to the horizontal and vertical components of the BV difference value. The type information can be signaled in the bitstream, and the BV difference values of the two components (e.g., the horizontal and vertical components) can be signaled according to the type information.

[0140] In AV1, the IBC mode can be effective in coding screen content. The IBC mode may introduce challenges to hardware design. To facilitate the hardware design, some changes can be adopted in the IBC mode, e.g., in AV1.

[0141] In an example of the first change, when the IBC mode is permitted, the loop filter can be disabled. The loop filter can include a deblocking filter, a Constrained Directional Enhancement Filter (CDEF), and a Loop Restoration (LR) filter. By disabling the loop filter, a dedicated second picture buffer can be avoided to enable the IBC mode.

[0142] In a second example of change, in order to facilitate parallel decoding, the prediction cannot cross the restricted area. The coordinates of the upper left position of the superblock are (x0, y0). For the superblock, the prediction at position (x, y) can be accessed by the IBC mode if the vertical coordinate is smaller than y0 and the horizontal coordinate is smaller than (x0 + 2(y0 - y)). In the example, the prediction at position (x, y) can be accessed by the IBC mode only if the vertical coordinate is smaller than y0 and the horizontal coordinate is smaller than (x0 + 2(y0 - y)). In the example, the prediction at position (x, y) can be accessed by the IBC mode only if the vertical coordinate is smaller than or equal to y0 and the horizontal coordinate is smaller than (x0 + 2(y0 - y)).

[0143] In a third example of a change, in order to allow for hardware write-back delays, the intermediate reconfigured area cannot be accessed by the IBC mode. The restricted intermediate reconfigured area can contain from 1 to N superblocks, where N is a positive integer. In addition to the second change described above, if the coordinates of the upper left position of the superblock (1210) being reconfigured are (x0, y0), the prediction at position (x, y) can be accessed by the IBC mode if the vertical coordinate is less than or equal to y0 and the horizontal coordinate is less than (x0 + 2(y0 - y) - D). D can indicate the size of the intermediate reconfigured area that is restricted for the IBC mode. FIG. 12 shows an example of the restricted intermediate reconfigured area. The shaded gray area includes the allowable search area that can be accessed in the IBC mode for each current superblock (1210) being reconfigured. The shaded black area includes the non-allowable search area that cannot be accessed in the IBC mode for each current superblock (1210). The unshaded white area includes the superblocks that are to be coded (e.g., reconfigured). In the case of the current superblock (1210(1)), the intermediate reconfigured area includes two superblocks (1221) to (1222) that are to the left of the current superblock (1210(1)). The superblocks (1221) to (1222) are not accessible for the current superblock (1210(1)). Area (1230) is accessible for the current superblock (1210(1)).

[0144] According to aspects of the present disclosure, AV1 can use a local reference range defined in the IBC mode. For example, a part of the on-chip memory (e.g., memory fabricated on the same chip as the processor) with a size of M×M (e.g., 128×128) can be allocated to store reference samples used in the IBC mode, and that part of the on-chip memory is called the RSM. The RSM can store reconstructed samples that can be used as reference samples. The reconstructed samples stored in the RSM are updated according to an update process, and the range of available reference samples in the RSM can be called the local reference range. In an embodiment, the size of the RSM is equal to the size of the superblock. The memory reuse mechanism can be applied to the RSM in units of L×L (e.g., 64×64). The RSM can be divided into I RSM units, and I is equal to the ratio of M×M to L×L. For example, when M×M is 128×128 and L×L is 64×64, I is 4 (128×128 / 64×64). Some changes can be made to the IBC mode due to the local reference range.

[0145] In an example of the first change, the maximum block size in the IBC mode is limited to L×L (e.g., 64×64).

[0146] In an example of the second change, the reference block and the corresponding current block within the current superblock (SB) can be in the same SB row. In the example, the reference block is located only in the current SB or the left adjacent SB on the left side of the current SB.

[0147] In an example of the third change, when a unit having an RSM unit of size L×L (e.g., 64×64) starts to be updated by the reconstructed samples of the current SB, the previously stored reference samples (e.g., the reference samples of the left adjacent SB) within the entire L×L unit can be marked as unavailable for generating the predicted samples used in the IBC mode.

[0148] FIG. 13 shows an exemplary memory reuse mechanism (1300) in which a memory (e.g., RSM (1310)) is updated during the coding (e.g., encoding or decoding) of the current SB (1301) in the current picture according to an embodiment of the present disclosure. The top block shows the RSM (1310) in state (0). The upper row shows the RSM (1310) in states (1)-(4). The lower row shows the current SB (1301) and the left adjacent SB (1302) during coding in the current picture in states (0)-(4). The left adjacent SB (1302) can be on the left side of the current SB (1301). In the example of FIG. 13, quadtree splitting is used at the SB root and the SB can include four regions. In the example, the size of each of the four regions is 64×64. In the example, the current SB (1301) includes four regions 4-7 and the left adjacent SB (1302) includes four regions 0-3.

[0149] In state (0), which is the first to code each SB such as the current SB (1301), the RSM (1310) can store samples of a previously coded SB (e.g., the left adjacent SB (1302)). When the current block is located in one of the four regions (e.g., four 64×64 regions) within the current SB (1301), the corresponding region within the RSM (1310) is emptied and can be used to store samples of the current coding region (e.g., the current 64×64 coding region). The samples within the RSM (1310) can be gradually updated by the samples within the current SB (1301).

[0150] Referring to state (1), the current block (1311) is located in region 4 within the current SB (1301), the corresponding region (e.g., the upper left region) within the RSM (1310) is emptied, and can be used to store samples of region 4 which is the currently coded region. Referring to the lower row, the BV (e.g., the coded BV or the decoded BV) (1321) can point from the current block (1311) to the reference block (1331) within the search range (1341) for the current block (1311) (the boundaries of the search range (1341) are marked by dashed lines). Referring to the upper row, the corresponding offset (1351) within the RSM (1310) can point from the current block (1311) to the reference block (1331) within the RSM (1310). Referring to state (1), the search range (1341) includes regions 1 - 3 within the left adjacent SB (1302) and the coded sub-region (1361) within region 4. The search range (1341) does not include region 0 within the left adjacent SB (1302).

[0151] Referring to state (2), the current block (1312) is located in region 5 within the current SB (1301), the corresponding region (e.g., the upper right region) within the RSM (1310) is emptied, and can be used to store samples of region 5 which is the currently coded region. The BV (e.g., the coded BV or the decoded BV) (1322) can point from the current block (1312) to the reference block (1332) within the search range (1342) for the current block (1312) (the boundaries of the search range (1342) are marked by dashed lines). The corresponding offset (1352) within the RSM (1310) can point from the current block (1312) to the reference block (1332) within the RSM (1310). Referring to state (2), the search range (1342) includes regions 2 - 3 within the left adjacent SB (1302) and the coded sub-region (1362) within region 5 within the current SB (1301). The search range (1342) does not include regions 0 - 1 within the left adjacent SB (1302).

[0152] Referring to state (3), the current block (1313) is located in region 6 within the current SB (1301), the corresponding region (e.g., the lower left region) within the RSM (1310) is emptied, and can be used to store samples of region 6 which is the currently coded region. The BV (e.g., the coded BV or the decoded BV) (1323) can point from the current block (1313) to a reference block (1333) within the search range (1343) (the boundaries of the search range (1342) are marked by dashed lines) for the current block (1313). The corresponding offset (1353) within the RSM (1310) can point within the RSM (1310) from the current block (1313) to the reference block (1333). Referring to state (3), the search range (1343) includes (i) region 3 within the left adjacent SB (1302), and (ii) the coded sub-regions (1363) within regions 4 - 5 and region 6 within the current SB (1301). The search range (1343) does not include regions 0 - 2 within the left adjacent SB (1302).

[0153] Referring to state (4), the current block (1314) is located in region 7 within the current SB (1301), the corresponding region (e.g., the lower right region) within the RSM (1310) is emptied, and can be used to store samples of region 7 which is the currently coded region. The BV (e.g., the coded BV or the decoded BV) (1324) can point from the current block (1314) to a reference block (1334) within the search range (1344) (the boundaries of the search range (1344) are marked by dashed lines) for the current block (1314). The corresponding offset (1354) within the RSM (1310) can point within the RSM (1310) from the current block (1314) to the reference block (1334). Referring to state (4), the search range (1344) includes the coded sub-regions (1364) within regions 4 - 6 and region 7 within the current SB (1301). The search range (1344) does not include regions 0 - 3 within the left adjacent SB (1302).

[0154] When the current SB (1301) is fully coded, the entire RSM (1310) can be filled with all samples of the current SB (1301).

[0155] In the example shown in FIG. 13, the current SB (1301) is partitioned using a quadtree split. The coding order of the four regions within the current SB (1301) can be the upper left region (e.g., region 4), the upper right region (e.g., region 5), the lower left region (e.g., region 6), and the lower right region (e.g., region 7). In other split decisions such as those shown in FIGS. 14A - 14B, the RSM update process can be the same as that shown in FIG. 13, for example, by replacing each region within the RSM using the reconstructed samples within the current SB.

[0156] FIGS. 14A - 14B show an exemplary update process within the RSM during the coding (e.g., encoding or decoding) of the current SB (1401). In FIGS. 14A - 14B, the left - adjacent SB (1402) is to the left of the current SB (1401) that is in coding (e.g., encoding or decoding). In the example, the size of each of the current SB (1401) and the left - adjacent SB (1402) is 128×128. Each of the current SB (1401) and the left - adjacent SB (1402) can include four regions (e.g., four blocks) of size 64×64. The current SB (1401) can include blocks 4 - 7, and the left - adjacent SB (1402) can include blocks 0 - 3.

[0157] In FIG. 14A, a horizontal split is performed at the SB root, followed by a vertical split. The SB (e.g., the current SB (1401)) can include four blocks, namely, the upper left block (e.g., block 4), the lower left block (e.g., block 6), the upper right block (e.g., block 5), and the lower right block (e.g., block 7). The coding order of the current SB (1401) can be the upper left block (state 1), the upper right block (state 2), the lower left block (state 3), and the lower right block (state 4).

[0158] In FIG. 14B, vertical splitting with SB included is performed, followed by horizontal splitting. The coding order of the current SB (1401) can be the upper left block (state 1), the lower left block (state 2), the upper right block (state 3), and the lower right block (state 4).

[0159] Depending on the position of the current block (e.g., (1431)) with respect to the current SB (1401), the following can be applied.

[0160] (i) Referring to state (1) in FIGS. 14A - 14B, the current block (1431) is within the upper left block (e.g., block 4) of the current SB (1401), and the RSM can include reference samples within the lower right block (e.g., block 3), the lower left block (e.g., block 2), and the upper right block (e.g., block 1) of the left adjacent SB (1402) in addition to the already reconstructed samples within block 4.

[0161] (ii) Referring to state (2) in FIG. 14A or state (3) in FIG. 14B, the current block (1432) is within the upper right block (e.g., block 5) of the current SB (1401).

[0162] As shown in state (2) of FIG. 14A, if the luma sample located at the upper left corner of block 6 ((0, 64) with respect to the current SB (1401)) has not been reconstructed yet, in addition to the already reconstructed samples in blocks (1462) within blocks 4 and 5, the current block (1432) can refer to the reference samples within the lower left block (e.g., block 2) and the lower right block (e.g., block 3) of the left adjacent SB (1402). The corresponding RSM can include the reference samples within the lower left block (e.g., block 2) and the lower right block (e.g., block 3) of the left adjacent SB (1402) in addition to the blocks (1462) within blocks 4 and 5.

[0163] Instead, as shown in state (3) of FIG. 14B, when the lumen sample located at the upper left corner of block 6 (currently (0, 64) with respect to SB(1401)) is being reconstructed, the current block (1432) can refer to the reference samples within the lower right block of the left adjacent SB(1402) (e.g., block 3). The corresponding RSM can include the reference samples within the lower right block of the left adjacent SB(1402) (e.g., block 3) in addition to the already reconstructed samples in blocks 4 and 5 within block (1462).

[0164] (iii) Referring to state (3) of FIG. 14A or state (2) of FIG. 14B, the current block (1433) is within the lower left block of the current SB(1401) (e.g., block 6).

[0165] As shown in state (2) of FIG. 14B, when the lumen sample located at the upper left corner of block 5 (currently (64, 0) with respect to SB(1401)) has not yet been reconstructed, in addition to the already reconstructed samples in block 4 and block (1463) within the current SB(1401), the current block (1433) can refer to the reference samples within the upper right block (e.g., block 1) and the lower right block (e.g., block 3) of the left adjacent SB(1402). The corresponding RSM can include the reference samples within the upper right block (e.g., block 1) and the lower right block (e.g., block 3) of the left adjacent SB(1402) in addition to block 4 and block (1463) within the current SB(1401).

[0166] Instead, as shown in state (3) of FIG. 14A, when the lumen sample located at the upper left corner of block 5 (currently (64, 0) with respect to SB(1401)) has been reconstructed, the current block (1433) can refer to the reference sample within the lower right block of the left adjacent SB(1402) (for example, block 3). The corresponding RSM can include the reference sample within the lower right block of the left adjacent SB(1402) (for example, block 3) in addition to the already reconstructed samples in blocks 4-5 and the block (1463) within the current SB(1401).

[0167] (iv) Referring to state (4) of FIGS. 14A - 14B, the current block (1434) is within the lower right block of the current SB(1401) (for example, block 7). The current block (1434) can refer to the already reconstructed samples within the current SB(1401), such as the already reconstructed samples in blocks 4 - 6 and block (1464). The corresponding RSM can include the reference samples within blocks 4 - 6 and block (1464). In the example, when the current block (1434) fits within the lower right block of the current SB(1401), the current block can only refer to the already reconstructed samples within the current SB(1401).

[0168] In accordance with aspects of the present disclosure, in some examples (e.g., ECM software), prediction techniques based on template matching can be used for intra prediction. Intra prediction based on template matching is referred to as Intra Template Matching Prediction (IntraTMP). In the IntraTMP mode, the best prediction block for the current block is determined from the reconstructed portion of the current picture based on the comparison between the L-shaped reference template of the best prediction block and the current template of the current block. In some examples, the encoder searches for a block having a template that is most similar to the current template of the current block within a predefined search range in the reconstructed portion of the current frame, and uses that block as the prediction block for the current block. The encoder then signals the use of IntraTMP for the prediction of the current block, and the same prediction operation can be performed on the decoder side.

[0169] FIG. 15 shows a diagram representing a search area for intra template matching prediction in some examples. In FIG. 15, the current picture (1500) is partitioned into CTUs as shown by the horizontal CTU boundary (1501) and the vertical CTU boundary (1502) in FIG. 15. The current block (1501) is the current CTU. The adjacent samples of the current block form the current template (1515) that is L-shaped. FIG. 15 shows a predefined search area for intra template matching prediction that includes four regions R1 to R4. R1 is within the current CTU, R2 is within the upper left CTU, R3 is within the upper CTU, and R4 is within the left CTU.

[0170] In an example, within each region, for each potential matching block (1520), the L-shaped adjacent samples of the potential matching block form a potential template (1525) that is L-shaped. The sum of absolute differences (SAD) between the potential template and the current template is calculated as the template matching cost of the potential matching block (1520).

[0171] The encoder or decoder can search a predefined area to determine a matching block having the lowest template matching cost, and the matching block is used as a prediction block for the current block (1510).

[0172] In some examples, the determination of predefined areas such as R1 to R4 is defined in proportion to the size (dimensions) of the current block because it has a certain number of comparisons per pixel. In the example, the area R2 can have a size of (SearchRange_w, SearchRange_h), where SearchRange_w is the width of the area R2 and SearchRange_h is the height of the area R2. The current block (1510) can have a size of (BlkW, BlkH), where BlkW is the width of the current block (1510) and BlkH is the height of the current block (1510). The width and height of the area R2 can be set according to Equations (3) and (4): [Number] Here, "a" is a constant that can control the gain / complexity trade-off. In the example, "a" is equal to 5.

[0173] In some examples, the IntraTMP tool is enabled for CUs having a size with a width and height of 64 or less. In some examples, the maximum CU size for IntraTMP is configurable.

[0174] In some examples, the IntraTMP mode is signaled at the CU level by a dedicated flag when the decoder-side intra mode derivation (DIMD) is not used for the current CU.

[0175] In some related examples (e.g., VVC), when the coding tree type is a dual tree type, IBC is only applied to the luma coding block, the chroma intra coding block can only use other intra prediction modes, and the coding efficiency of the intra chroma coding block is limited in the case of a dual tree chroma.

[0176] However, in the following description, the template of a block can refer to any suitable portion of the neighboring samples of the block, such as the samples neighboring the block above, to the left, to the right, and below the block.

[0177] FIG. 16 shows an example of a current block and a template of the current block. The template (indicated by the gray area) includes the reconstructed samples neighboring above and to the left. Note that while the template in FIG. 16 is L-shaped, the template can be defined to have other suitable patterns.

[0178] Aspects of the present disclosure provide techniques for enabling current picture reference techniques, such as IBC mode, IntraBC mode, IntraTMP mode, etc., for use in chroma blocks (in some examples, also referred to as chroma coding blocks) in a chroma separation tree, for example when a coding tree is of a dual-tree type. The current reconstruction of chroma blocks is based on reference chroma blocks within the reconstructed portion of the current picture. In some embodiments, a block vector (also referred to as a chroma block vector) is determined for a chroma block, the block vector indicating a reference chroma block within the same picture as the chroma block, and the reference chroma block can be copied to be the current chroma block. A block vector predictor for a chroma block vector (BV) or chroma BV can be derived from a luma block (also referred to as a corresponding luma area) from the same region, the luma block being from a luma separation tree. It can be said that the luma block and the chroma block are collocated in the same luma area. The chroma block vector or the block vector predictor for the chroma block can be generated by various techniques described in the present disclosure. In some examples, when IntraTMP is used to determine a chroma BV, an initial chroma block vector for a template matching search process can be derived from a luma block collocated with the chroma block within the same luma area.

[0179] In some embodiments, the chroma BV of a chroma block is signaled in a manner similar to luma BV signaling techniques. In some examples, chroma BV signaling for a chroma block can include BV predictor (BVP) signaling by a BVP index, BV difference (BVD) signaling, and BV precision index signaling.

[0180] In some examples, the luma blocks within the area corresponding to the current chroma block (the same area) are predicted by applying a current picture reference mode such as the IBC mode, IntraBC mode, or IntraTMP mode, and the associated luma BV value of the luma block is derived. In some examples, the luma BV value of the luma block can be used to derive BVP candidates for the chroma block, and one of the BVP candidates can be selected as the BV predictor for the chroma block.

[0181] FIG. 17 shows an example of the corresponding luma area of a chroma block in the case of a chroma sampling format 4:2:0. In FIG. 17, the chroma block (1701) is generated according to the chroma sampling format 4:2:0. The corresponding luma area (1702) includes one or more luma blocks such as five luma blocks (1721) to (1725). In some examples, the luma blocks (1721) to (1725) can each have a respective luma BV, as indicated by BV0, BV1, BV2, BV3, and BV4 in FIG. 17. In some examples, when a luma BV is selected to derive a BV predictor or BV predictor candidate for a chroma block, the luma BV is appropriately converted to a chroma BV according to, for example, Equation (1), Equation (2), and Table 1.

[0182] In some examples, the chroma BV predictor candidate is derived from the average value of the luma BVs within the corresponding luma area. For example, in FIG. 17, BV C represents the chroma BV value of the chroma BV predictor candidate and is derived from the average of BV0, BV1, BV2, BV3, and BV4 using, for example, Equation (1), Equation (2), and Table 1.

[0183] In some examples, the chroma BV predictor candidate is derived from the luma BV applied at the luma sample position corresponding to a particular sample position of the chroma block. For example, when a particular sample position of the chroma block is the center sample position of the chroma block, in that case, the luma BV applied at the luma sample position corresponding to the center sample position of the chroma block is used to derive the chroma BV predictor candidate. In other examples, when a particular sample position of the chroma block is the top-left sample position of the chroma block, in that case, the luma BV applied at the luma sample position corresponding to the top-left sample position of the chroma block is used to derive the chroma BV predictor candidate.

[0184] In some examples, the chroma BV predictor candidate is derived from one of the luma BVs within the corresponding luma area. In the example, the chroma BV predictor candidate is derived from BV0. In other examples, the chroma BV predictor candidate is derived from BV1. In other examples, the chroma BV predictor candidate is derived from BV2. In other examples, the chroma BV predictor candidate is derived from BV3. In other examples, the chroma BV predictor candidate is derived from BV4.

[0185] In some examples, the BV prediction of the chroma block may be selected from an option that is coarser than the luma BV accuracy. In the example, the luma BV accuracy is 1-pel, and the chroma BV accuracy can be 2-pel, 4-pel, or 8-pel.

[0186] In some examples, chroma BV predictor candidates can be generated according to the above techniques and placed in a candidate list. A BVP index indicating a selected chroma BV predictor from the chroma BV predictor candidates can be signaled from the encoder side to the decoder side. Further, a BV difference (BVD) may be signaled from the encoder side to the decoder side, and a BV accuracy index may be signaled from the encoder side to the decoder side. In an example, a chroma BV difference value of a chroma block can be determined based on the signaled BV difference and BV accuracy index. The BV difference value and the selected chroma BV predictor can be combined to determine a final chroma BV that indicates a reference chroma block for copying as the current chroma block in the example.

[0187] In some embodiments, the chroma BV of a chroma block can be derived from a collocated luma block without signaling.

[0188] In some examples, the chroma BV of the current chroma block is derived from the average value of the luma BV within the corresponding luma area. For example, in FIG. 17, BV C represents the chroma BV and is derived from the average of BV0, BV1, BV2, BV3, and BV4. In the example, the chroma BV is derived from the average according to Equation (1), Equation (2), and Table 1. The chroma BV indicates a reference chroma block for copying as the current chroma block.

[0189] In some examples, the chroma BV of the current chroma block is derived from the luma BV applied at the luma sample position corresponding to a particular sample position of the chroma block. For example, when a particular sample position of the chroma block is the center sample position of the chroma block, in that case, the luma BV applied at the luma sample position corresponding to the center sample position of the chroma block is used to derive the chroma BV. In other examples, when a particular sample position of the chroma block is the upper left sample position of the chroma block, in that case, the luma BV applied at the luma sample position corresponding to the upper left sample position of the chroma block is used to derive the chroma BV. The chroma BV indicates a reference chroma block for copying as the current chroma block. However, Equation (1), Equation (2), and Table 1 are used in the derivation of the vector of the chroma block in some examples.

[0190] In some examples, the chroma BV of the current chroma block is derived from one of the luma BVs within the corresponding luma area. In an example, the chroma BV is derived from BV0. In other examples, the chroma BV is derived from BV1. In other examples, the chroma BV is derived from BV2. In other examples, the chroma BV is derived from BV3. In other examples, the chroma BV is derived from BV4. The chroma BV indicates a reference chroma block for copying as the current chroma block. However, Equation (1), Equation (2), and Table 1 are used in the derivation in some examples.

[0191] In an embodiment, an index indicating one of the plurality of luma BVs within the corresponding luma area may be signaled from the encoder side to the decoder side to indicate which luma BV should be used to derive the chroma BV.

[0192] In some embodiments, a technique based on template matching is used to determine which of a plurality of lumen BVs within a corresponding lumen area of a chroma block should be used to derive a chroma BV. For example, according to each lumen BV of the plurality of lumen BVs, a corresponding candidate chroma BV of the current chroma block is derived. The corresponding candidate chroma BV (of the lumen BV) indicates a candidate reference block related to the lumen BV for the current chroma block. By applying the corresponding candidate chroma BV (of the lumen BV) to the current template of the current chroma block, a candidate reference template related to the lumen BV is determined. Then, a template matching cost can be calculated between the candidate reference template and the current template of the current chroma block, and the template matching cost is associated with the candidate chroma BV. The template matching cost is calculated based on a distortion measurement between the current template of the current chroma block and the candidate reference template associated with the candidate chroma BV, for example, based on SAD.

[0193] FIG. 18 shows an example of template matching in some examples. In the example, the candidate chroma BV is derived according to the lumen BV of the lumen block collocated with the current chroma block (e.g., according to Equation (1), Equation (2), and Table 1). The candidate chroma BV (associated with the lumen BV) indicates a candidate chroma reference block of the current chroma BV. By applying the candidate chroma BV (associated with the lumen BV) to the current template of the current chroma block, a candidate reference template associated with the candidate chroma BV is determined. Then, a template matching cost can be calculated between the candidate reference template and the current template of the current chroma block, and the template matching cost is associated with the candidate chroma BV. The template matching cost is calculated based on a distortion measurement between the current template of the current chroma block and the candidate reference template associated with the candidate chroma BV, for example, based on SAD.

[0194] In some examples, one of the candidate chroma BVs having the minimum template matching cost is used as the chroma BV of the chroma block for prediction in the current picture reference mode.

[0195] In some examples, the candidate chroma BVs are reordered into a reordered list, for example having an ascending order, based on the template matching cost. In an example, the BV candidate indexes of the reordered list are signaled from the encoder side to the decoder side to indicate the chroma BVs in the reordered list for prediction in the current picture reference mode.

[0196] In some embodiments, the current chroma block may be coded by the IntraTMP mode, and the collocated luma block is coded by IntraBC or IntraTMP. In some examples, the chroma BV derived according to the luma BV of the collocated luma block is used as a starting point for the template matching search.

[0197] In some examples, in the case of dual-tree coding (the coding tree type is the dual-tree type), multiple collocated luma blocks exist in the corresponding luma area for the current chroma block in the IntraTMP mode, and the multiple collocated luma blocks may be coded by IntraBC or IntraTMP. In an example, the initial chroma BV derived from the luma BV of a certain luma block among the multiple collocated luma blocks is used as a starting point for the template matching search to determine the final chroma BV for prediction of the current chroma block in the current picture reference. In an example, the luma blocks among the multiple collocated luma blocks are selected according to a predefined order.

[0198] In some examples, the current chroma block and its corresponding luma block are used together as a template for template matching search. In an example, for the current chroma block, a current chroma template is determined, and a current luma template collocated with the current chroma template is determined. During the template matching search, for a chroma BV, the corresponding luma BV can be determined, for example, according to Equation (1), Equation (2), and Table 1. The chroma BV is applied to the current chroma template to determine a chroma reference template. The corresponding luma BV is applied to the current luma template to determine a luma reference template. In some examples, the luma reference template can be determined based on the chroma reference template. For example, the luma reference template includes luma samples within the same luma area as the chroma reference template. In an example, a first template matching cost is calculated based on, for example, the SAD between the current luma template and the chroma reference template, and a second template matching cost is calculated based on, for example, the SAD between the current luma template and the luma reference template. The first template matching cost and the second template matching cost are combined to determine a combined template matching cost for the chroma BV in the template matching search. The template matching search can determine, in an example, the chroma BV with the minimum template matching cost.

[0199] FIG. 19 shows a flowchart illustrating a process (1900) according to an embodiment of the present disclosure. The process (1900) can be used in a video encoder. In various embodiments, the process (1900) is performed by processing circuits in terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of video encoder (403), a processing circuit that executes the functions of video encoder (603), a processing circuit that executes the functions of video encoder (703), and the like. In some embodiments, since the process (1900) is implemented by software instructions, when the processing circuit executes the software instructions, the processing circuit executes the process (1900). The process (1900) starts at (S1901) and proceeds to (S1910).

[0200] (S1910), the use of the current picture reference (CPR) mode for chroma blocks in the chroma separation tree is determined. The CPR mode can code chroma blocks in the chroma separation tree according to the reconstructed portions within the current picture containing the chroma blocks. The CPR mode can be, for example, the IBC mode, the IntraBC mode, the IntraTMP mode, etc.

[0201] (S1920), the chroma block vector of the chroma block is determined according to one or more luma block vectors associated with one or more luma blocks located in the luma area corresponding to the chroma block, and the chroma block vector indicates a reference chroma block within the current picture.

[0202] (S1930), a signal indicating the use of the current picture reference mode for coding the chroma block is encoded in the bitstream carrying the video.

[0203] In some examples, a block vector predictor is derived from a first luma block vector selected from one or more luma block vectors to predict a chroma block vector, and a block vector difference between the block vector predictor and the chroma block vector is determined. Signals indicating the first luma block vector and the block vector difference are encoded in a bitstream.

[0204] In some examples, the block vector predictor is derived from an average value of one or more luma block vectors.

[0205] In some examples, a first luma block is determined from one or more luma blocks, and the first luma block includes sample points corresponding to specific sample positions of a chroma block. The block vector predictor is derived from a luma block vector associated with the first luma block. In an example, the specific sample position is a center sample position of the chroma block. In other examples, the specific sample position is an upper left sample position of the chroma block.

[0206] In some examples, a first accuracy of a chroma block vector is determined from candidates that are coarser than a second accuracy of one or more luma block vectors. An accuracy index indicating the first accuracy is encoded within the bitstream.

[0207] In some examples, the chroma block vector is derived from an average value of one or more luma block vectors

[0208] In some examples, a first luma block is determined from one or more luma blocks, and the first luma block includes sample points corresponding to specific sample positions of a chroma block. The chroma block vector is derived from a luma block vector associated with the first luma block. In an example, the specific sample position is a center sample position of the chroma block. In other examples, the specific sample position is an upper left sample position of the chroma block.

[0209] In some examples, the chroma block vector is derived from a first luma block vector among one or more luma block vectors. An index indicating the first luma block vector selected from among the one or more luma block vectors is encoded in the bitstream.

[0210] In some examples, the luma area corresponding to a chroma block includes a plurality of luma blocks each having a luma block vector. In some examples, a candidate chroma block vector is derived from the luma block vector. A candidate reference template corresponding to the current template of the chroma block is determined respectively according to the candidate chroma block vector. The template matching cost respectively associated with the candidate chroma block vector is calculated according to the distortion between the current template and the candidate reference template. The chroma block vector is selected from the candidate chroma block vectors based on the template matching cost. For example, the chroma block vector has the minimum template matching cost among the candidate chroma block vectors.

[0211] In some examples, the candidate chroma block vectors are ordered according to the template matching cost into a list of ordered candidate chroma block vectors. The chroma block vector is selected from the ordered list according to other suitable techniques such as rate distortion measurement. An index indicating the chroma block vector selected from the list of ordered candidate chroma block vectors is encoded in the bitstream.

[0212] In some examples, an initial chroma block vector is derived according to one or more luma block vectors. A template matching search is performed starting from the initial chroma block vector to determine the chroma block vector.

[0213] In some examples, the initial chroma block vector is derived from the average value of one or more luma block vectors.

[0214]

[0214] In some examples, the first lumablock is determined from one or more lumablocks, and the first lumablock includes sample points corresponding to specific sample positions of the chromablock. The initial chromablock vector is derived from the lumablock vector associated with the first lumablock. In an example, the specific sample position is the central sample position of the chromablock. In other examples, the specific sample position is the upper left sample position of the chromablock.

[0215] In some examples, the first lumablock vector is selected from one or more lumablock vectors to derive the initial chromablock. An index indicating the first lumablock vector selected from one or more lumablock vectors is encoded in the bitstream.

[0216] In some examples, the chromablock and the corresponding lumablock are used together as a template for template matching search. For example, in the case of an intermediate chromablock vector, an intermediate chroma reference template corresponding to the current chroma template of the chromablock is determined according to the intermediate chromablock vector. Then, an intermediate luma reference template collocated with the intermediate chroma reference template is determined. The first template matching cost is calculated according to the distortion between the current chroma template and the intermediate chroma reference template, and the second template matching cost is calculated according to the distortion between the current chroma template and the intermediate luma reference template collocated with the current luma template. The combined template matching cost associated with the intermediate chromablock vector is calculated by combining the first template matching cost and the second template matching cost. The combined template matching cost is used as a cost metric for the intermediate chromablock vector.

[0217] Then, the process proceeds to (S1999) and ends.

[0218] The process (1900) can be suitably adapted. The steps of the process (1900) can be changed and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0219] FIG. 20 shows a flowchart illustrating a process (2000) according to an embodiment of the present disclosure. The process (2000) can be used in a video decoder. In various embodiments, the process (2000) is executed by processing circuits such as in terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of video decoder (410), a processing circuit that executes the functions of video decoder (510), and the like. In some embodiments, since the process (2000) is implemented by software instructions, when the processing circuit executes the software instructions, the processing circuit executes the process (2000). The process (2000) starts from (S2001) and proceeds to (S2010).

[0220] In (S2010), a signal (also referred to as a syntax element) is decoded from a bitstream carrying video (also referred to as a coded video bitstream), and the signal indicates the use of a current picture reference (CPR) mode for coding chroma blocks in a reconstructed portion within the current picture that includes chroma blocks. In some examples, a coded video bitstream including the current picture is received. The current picture includes chroma blocks, and the chroma blocks are collocated within the same luma area as one or more luma blocks. A syntax element indicating the use of the current picture reference (CPR) mode for chroma blocks is decoded from the coded video bitstream. The CPR mode can be, for example, the IBC mode, the IntraBC mode, the IntraTMP mode, etc.

[0221] In (S2020), the chroma block vector of the chroma block is determined according to one or more luma block vectors associated with one or more luma blocks, and the one or more luma blocks are located within the luma area corresponding to the chroma block (e.g., collocated with the chroma block), and the chroma block vector indicates a reference chroma block in the current picture.

[0222] In (S2030), the chroma block is reconstructed based on the reference chroma block.

[0223] In some embodiments, a block vector predictor is determined according to one or more luma block vectors, and a block vector difference is decoded from the bit stream. The chroma block vector is determined based on the block vector predictor and the block vector difference. In some examples, the block vector predictor is derived from the average value of one or more luma block vectors using, for example, Equation (1), Equation (2), and Table 1.

[0224] In some examples, a first luma block is determined from one or more luma blocks, and the first luma block includes a sample point corresponding to a specific sample position of the chroma block. The block vector predictor is derived from the luma block vector associated with the first luma block. In the example, the specific sample position is the central sample value of the chroma block. In other examples, the specific sample position is the upper left sample position of the chroma block.

[0225] In some examples, an index indicating the first luma block vector from one or more luma block vectors is decoded from the bit stream. The block vector predictor is derived from the first luma block vector.

[0226] In some examples, a first precision of the chroma block vector is determined from candidates that are coarser than a second precision of one or more luma block vectors based on, for example, a precision index decoded from the bit stream.

[0227] In some examples, the chroma block vector is derived from the average value of one or more luma block vectors, for example, using Equation (1), Equation (2), and Table 1.

[0228] In some examples, the first luma block is determined from one or more luma blocks, and the first luma block includes a sample point corresponding to a specific sample position of the chroma block. The chroma block vector is derived from the luma block vector associated with the first luma block. In an example, the specific sample position is the central sample value of the chroma block. In other examples, the specific sample position is the upper left sample position of the chroma block.

[0229] In some examples, the index indicating the first luma block vector from one or more luma block vectors is decoded from the bitstream. The chroma block vector is derived from the first luma block vector.

[0230] In some examples, the luma area corresponding to the chroma block includes a plurality of luma blocks each having a luma block vector. Candidate chroma block vectors are respectively derived from the luma block vectors. The candidate reference template corresponding to the current template of the chroma block is determined according to the candidate chroma block vectors. The template matching cost respectively associated with the candidate chroma block vectors is calculated according to the distortion between the current template and the candidate reference template, for example, using SAD. The chroma block vector is selected from the candidate chroma block vectors based on the template matching cost. For example, the chroma block vector has the minimum template matching cost among the candidate chroma block vectors.

[0231] In some examples, candidate chroma block vectors are ordered into a list of sorted candidate chroma block vectors according to a template matching cost. An index indicating a chroma block vector from the list of sorted candidate chroma block vectors is then decoded from the bitstream. The chroma block vector is selected from the sorted list according to the index.

[0232] In some examples, an initial chroma block vector is determined according to one or more luma block vectors. A template matching search is performed starting from the initial chroma block vector to determine the chroma block vector. In some examples, the initial chroma block vector is derived from an average value of one or more luma block vectors, for example, using Equation (1), Equation (2), and Table 1.

[0233] In some examples, a first luma block is determined from one or more luma blocks, and the first luma block includes sample points corresponding to specific sample positions of the chroma block. The initial chroma block vector is derived from the luma block vector associated with the first luma block. In an example, the specific sample position is the center sample position of the chroma block. In other examples, the specific sample position is the upper left sample position of the chroma block.

[0234] In some examples, an index indicating a first luma block vector from one or more luma block vectors is decoded from the bitstream. The initial chroma block vector is derived from the first luma block vector.

[0235] In some examples, to perform template matching search, a chroma block and a corresponding luma block are used together as a template for the template matching search. For example, in the case of an intermediate chroma block vector, an intermediate chroma reference template corresponding to the current chroma template of the chroma block is determined according to the intermediate chroma block vector. Then, an intermediate luma reference template collocated (e.g., within the same luma area) with the intermediate chroma reference template is determined. A first template matching cost is calculated according to the distortion between the current chroma template and the intermediate chroma reference template, and a second template matching cost is calculated according to the distortion between the current chroma template and the intermediate luma reference template and the current luma template collocated with the current chroma template. A combined template matching cost associated with the intermediate chroma block vector is calculated by combining the first template matching cost and the second template matching cost. The combined template matching cost is used as a cost metric for the intermediate chroma block vector.

[0236] Then, the process proceeds to (S2099) and ends.

[0237] The process (2000) can be appropriately adapted. The steps of the process (2000) can be changed and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0238] The above technology can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 21 shows a computer system (2100) suitable for implementing a particular embodiment of the disclosed subject matter.

[0239] Computer software can be coded in any suitable machine code or computer language that can generate code containing instructions that can be executed directly or through interpretation, microcode execution, etc. by one or more central processing units (CPUs), graphics processing units (GPUs), etc., according to mechanisms such as assembly, compilation, and linking.

[0240] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, devices for the Internet of Things, etc.

[0241] The components shown in FIG. 21 with respect to the computer system (2100) are essentially illustrative and are not intended to suggest any limitation regarding the use or function scope of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be interpreted as having any dependence or requirement regarding any one or combination of the components described in the exemplary embodiments of the computer system (2100).

[0242] The computer system (2100) may include a specific human interface input device. Such a human interface input device can respond to input by one or more users through, for example, tactile input (e.g., keyboard, swipe, data glove operation), voice input (e.g., voice, clapping), visual input (e.g., gesture), olfactory input (not shown). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious human input, such as voice (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still camera), videos (e.g., two-dimensional videos, three-dimensional videos including stereoscopic videos).

[0243] The input human interface device may include one or more of a keyboard (2101), a mouse (2102), a trackpad (2103), a touch screen (2110), a data glove (not shown), a joystick (2105), a microphone (2106), a scanner (2107), a camera (2108) (only one of each is shown).

[0244] The computer system (2100) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (2110), a data glove (not shown), or a joystick (2105), but there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (2109), headphones (not shown)), visual output devices (e.g., a CRT screen, an LCD screen, a plasma screen, an OLED screen, regardless of the presence or absence of a touch screen input function and regardless of the presence or absence of a tactile feedback function, and some of them can output two-dimensional visual output or output of more than three dimensions by means such as stereoscopic output, virtual reality glasses (not shown), holographic display, and smoke tank (not shown), a screen (2110)), and a printer (not shown).

[0245] The computer system (2100) can also include memory devices and their associated media that are accessible to humans, such as a CD / DVD ROM / RW (2120) including a CD / DVD or similar media (2121), a thumb drive (2122), a removable hard disk or solid state drive (2123), legacy magnetic media such as tapes and floppy (registered trademark) disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0246] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0247] The computer system (2100) can also include an interface (2154) to one or more communication networks (2155). The network can be, for example, wireless, wireline, optical. The network can further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet®, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wireline or wireless wide area digital networks including cable TV, satellite TV, and over-the-air broadcast TV, vehicular and factory networks including CAN bus, etc. A particular network generally requires an external network interface adapter attached to a particular general-purpose digital port or peripheral bus (2149) (e.g., a USB port of the computer system (2100)). Others are generally incorporated into the core of the computer system (2100) by attachment to a system bus as described later (e.g., an Ethernet network to a PC computer system, or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2100) can communicate with other entities. Such communication can be unidirectional receive-only (e.g., broadcast TV) or unidirectional transmit-only (e.g., CAN bus to a particular CAN bus device), or bidirectional to other computer systems using, for example, a local or wide area digital network. A particular protocol or protocol stack can be used with each of the networks and network interfaces described above.

[0248] The above human interface device, the accessible memory device by a person, and the network interface can be attached to the core (2140) of the computer system (2100).

[0249] The core (2140) can include one or more central processing units (CPUs) (2141), a graphics processing unit (GPU) (2142), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2143), a hardware accelerator (2144) for a specific task, a graphics adapter (2150), and the like. These devices may be connected through a system bus (2148) together with a read-only memory (ROM) (2145), a random access memory (RAM) (2146), a built-in large-capacity storage device such as an internal hard drive inaccessible to the user, an SSD, etc. (2147). In some computer systems, the system bus (2148) can be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (2148) of the core or through a peripheral bus (2149). In the example, the display (2110) can be connected to the graphics adapter (2150). Architectures for peripheral buses include PCI, USB, and the like.

[0250] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) are capable of executing specific instructions that, in combination, can constitute the above computer code. The computer code can be stored in the ROM (2145) or RAM (2146). Temporary data can also be stored in the RAM (2146), while persistent data can be stored, for example, in the built-in mass storage device (2147). Fast storage and retrieval to / from any of the memory devices can be enabled by the use of cache memory. The cache memory can be closely related to one or more CPUs (2141), GPUs (2142), mass storage devices (2147), ROMs (2145), RAMs (2146), etc.

[0251] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and configured for the purposes of this disclosure, or they can be of the kind well known and available to those skilled in the art of computer software technology.

[0252] As an example, and not by way of limitation, a computer system having an architecture (2100), specifically a core (2140), can provide functionality as a result of software embodied in one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be media associated with the user-accessible mass storage device introduced above, in addition to specific storage devices of the core (2140) having a non-transitory nature, such as an on-core mass storage device (2147) or a ROM (2145). The software implementing various embodiments of the present disclosure is stored in such a device and is executable by the core (2140). The computer-readable media can include one or more memory devices or chips depending on specific needs. The software can cause the core (2140), and specifically the processors (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in a RAM (2146) and change such data structures according to processes defined by the software, so as to execute specific processes or specific parts of specific processes described herein. Additionally, or alternatively, the computer system can provide functionality as a result of logic (e.g., an accelerator (2144)) hardwired or otherwise embodied in a circuit that operates instead of or in conjunction with software to execute specific processes or specific parts of specific processes described herein. References to software can, if necessary, include logic, and vice versa. References to computer-readable media can, if necessary, include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0253] Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Group of Picture(s) TU: Transform Unit(s) PU: Prediction Unit(s) CTU: Coding Tree Unit(s) CTB: Coding Tree Block(s) PB: Prediction Block(s) HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPU: Central Processing Unit(s) GPU: Graphics Processing Unit(s) CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area(s) SSD: Solid-State Drive IC: Integrated Circuit CU: Coding Unit

[0254] Although the present disclosure has described several exemplary embodiments, there are alternative, interchangeable, and various substitution equivalents within the scope of the present disclosure. Thus, as is apparent, those skilled in the art can embody the principles of the present disclosure, and thus can conceive of numerous systems and methods within its spirit and scope, even if not explicitly illustrated or described herein.

[0255] [Incorporation by Reference] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 388,597, filed on July 12, 2022, under the title "IBC Chroma Block Vector Derivation from Luma Block Vectors" and U.S. Patent Application No. 17 / 983,353, filed on November 8, 2022, under the title "IBC CHROMA BLOCK VECTOR DERIVATION FROM LUMA BLOCK VECTORS". The disclosures of these prior applications are hereby incorporated by reference in their entirety into this application.

Claims

1. 1. A method of video processing performed by a decoder, comprising: receiving a coded video bitstream including a current picture, the current picture including a chroma block in a chroma separation tree, the chroma block being co-located within the same luma area as one or more luma blocks; decoding, from the coded video bitstream, a syntax element indicating a current picture reference (CPR) mode for the chroma block; determining, in response to the CPR mode, a chroma block vector for the chroma block according to one or more luma block vectors associated with the one or more luma blocks, the chroma block vector indicating a reference chroma block in the current picture; reconstructing the chroma block based on the reference chroma block; The method according to claim 1,

2. The CPR mode is an intra block copy (IBC) mode, and the step of determining the chroma block vector comprises: determining a block vector predictor according to the one or more luma block vectors; decoding block vector differentials from the coded video bitstream; determining the chroma block vector based on the block vector predictor and the block vector differential; Further comprising The method of claim 1.

3. The step of determining a block vector predictor comprises: deriving the block vector predictor from at least one of an average of the one or more luma block vectors and a weighted average of the one or more luma block vectors. The method of claim 2.

4. The step of determining a block vector predictor comprises: determining a first luma block from the one or more luma blocks having a sample point corresponding to a central sample position of the chroma block; deriving the block vector predictor from a luma block vector associated with the first luma block; Further comprising The method of claim 2.

5. The step of determining a block vector predictor comprises: Determining a first luma block having a sample point corresponding to the upper left sample position of the chroma block from the one or more luma blocks; Deriving the block vector predictor from the luma block vector associated with the first luma block and further comprising The method according to claim 2.

6. The step of determining the block vector predictor comprises Decoding an index indicating a first luma block vector from the one or more luma block vectors; Deriving the block vector predictor from the first luma block vector and further comprising The method according to claim 2.

7. Further comprising determining a first accuracy of the chroma block vector from candidates coarser than a second accuracy of the one or more luma block vectors, The method according to claim 2.

8. The CPR mode is an intra-block copy (IBC) mode, and the step of determining the chroma block vector comprises Deriving the chroma block vector from at least one of an average value of the one or more luma block vectors and a weighted average value of the one or more luma block vectors, The method according to claim 1.

9. The CPR mode is an intra-block copy (IBC) mode, and the step of determining the chroma block vector comprises Determining a first luma block having a sample point corresponding to the center sample position of the chroma block from the one or more luma blocks; Deriving the chroma block vector from the luma block vector associated with the first luma block and further comprising The method according to claim 1.

10. The CPR mode is an intra-block copy (IBC) mode, and the step of determining the chroma block vector comprises Determining a first luma block having a sample point corresponding to the upper left sample position of the chroma block from the one or more luma blocks; Deriving the chroma block vector from the luma block vector associated with the first luma block and further comprising The method according to claim 1.

11. The CPR mode is an intra-block copy (IBC) mode, and the step of determining the chroma block vector comprises decoding an index indicating a first luma block vector from the one or more luma block vectors; deriving the chroma block vector from the first luma block vector further comprising The method according to claim 1.

12. The CPR mode is an intra-block copy (IBC) mode, and the luma area corresponding to the chroma block includes luma blocks each having a luma block vector, and the method comprises: deriving candidate chroma block vectors from the luma block vectors; determining candidate reference templates corresponding to the current template of the chroma block according to the candidate chroma block vectors; calculating template matching costs respectively associated with the candidate chroma block vectors according to a distortion between the current template and the candidate reference templates; selecting, from the candidate chroma block vectors based on the template matching costs, the chroma block vector having the minimum template matching cost among the candidate chroma block vectors comprising The method according to claim 1.

13. The CPR mode is an intra-block copy (IBC) mode, and the luma area corresponding to the chroma block includes luma blocks each having a luma block vector, and the method comprises: deriving candidate chroma block vectors from the luma block vectors; determining candidate reference templates corresponding to the current template of the chroma block according to the candidate chroma block vectors; calculating template matching costs respectively associated with the candidate chroma block vectors according to a distortion between the current template and the candidate reference templates; ordering the candidate chroma block vectors according to the template matching costs into a list of ordered candidate chroma block vectors; decoding an index indicating the chroma block vector from the ordered list of candidate chroma block vectors from the coded video bitstream selecting the chroma block vector from the list of the sorted candidate chroma block vectors according to the index having The method according to claim 1.

14. The CPR mode is an Intra Template Matching Prediction (IntraTmp) mode, and the step of determining the chroma block vector includes deriving an initial chroma block vector according to the one or more luma block vectors, and performing a template matching search starting from the initial chroma block vector to determine the chroma block vector further comprising The method according to claim 1.

15. The step of deriving the initial chroma block vector includes further deriving the initial chroma block vector from at least one of an average value of the one or more luma block vectors and a weighted average value of the one or more luma block vectors The method according to claim 14.

16. The step of deriving the initial chroma block vector includes determining a first luma block having a sample point corresponding to a central sample position of the chroma block from the one or more luma blocks, and deriving the initial chroma block vector from a luma block vector associated with the first luma block further comprising The method according to claim 14.

17. The step of deriving the initial chroma block vector includes determining a first luma block having a sample point corresponding to an upper left sample position of the chroma block from the one or more luma blocks, and deriving the initial chroma block vector from a luma block vector associated with the first luma block further comprising The method according to claim 14.

18. The step of deriving the initial chroma block vector includes decoding an index indicating a first luma block vector from the one or more luma block vectors, and deriving the initial chroma block vector from the first luma block vector further comprising The method according to claim 14.

19. The step of performing the template matching search is for an intermediate chroma block vector Determining an intermediate chroma reference template corresponding to the current chroma template of the chroma block according to the intermediate chroma block vector; Determining an intermediate luma reference template collocated with the intermediate chroma reference template; Calculating a first template matching cost according to the distortion between the current chroma template and the intermediate chroma reference template; Calculating a second template matching cost according to the distortion between the current luma template collocated with the current chroma template and the intermediate luma reference template; Calculating a combined template matching cost associated with the intermediate chroma block vector by combining the first template matching cost and the second template matching cost; further comprising; The method according to claim 14.

20. An apparatus for video decoding, comprising a processing circuit configured to execute the method according to any one of claims 1 to 19.

21. A method for video processing executed by an encoder, comprising: determining the use of a current picture reference (CPR) mode for a chroma block in a chroma separation tree, wherein the CPR mode can code the chroma block in the chroma separation tree according to a reconstructed portion within a current picture containing the chroma block; determining the chroma block vector of the chroma block according to one or more luma block vectors associated with one or more luma blocks located in a luma area corresponding to the chroma block, wherein the chroma block vector indicates a reference chroma block within the current picture; encoding a signal indicating the use of the CPR mode for coding the chroma block in a bitstream carrying the video including the current picture. having a method. ​