Non-residual coding based on block vector coding
By introducing format rules and skip flags in video encoding, the complexity problem of residual-related syntax elements writing and skipping in intra prediction in the prior art is solved, and a more efficient encoding process and efficiency are achieved.
Patent Information
- Application Number
- CN202480004444.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-13
- Filing Date
- 2024-05-10
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video encoding techniques are difficult to effectively handle the writing and skipping of residual-related syntax elements in intra prediction, resulting in increased encoding efficiency and complexity.
By introducing format rules in video encoding, a code stream of visual media data is processed, which contains the skip flag of the current block. This flag indicates whether the residual-related syntax element is written to the current block in IntraTMP mode. If the skip flag indicates that it is not written, the current block is processed by directly copying the sample value in the predicted block indicated by the block vector.
The intra prediction encoding process is simplified, the encoding complexity and signaling overhead are reduced, and the encoding efficiency is improved.
Smart Images

Figure CN120077638A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 526,668, titled "Non - Residual Coding On IntraPrediction Coding", filed on July 13, 2023. The entire content of the provisional application is incorporated herein by reference. Technical Field
[0002] This disclosure describes aspects generally related to video coding. Background Art
[0003] The background description provided herein is for the purpose of generally presenting the context of the disclosure. The extent of the work of the presently - named inventors, described in this background art section and in various aspects of this specification, does not indicate that it was qualified as prior art at the time of filing and has never been expressly or implicitly admitted as prior art to this disclosure.
[0004] Image / video compression helps to transmit image / video data between different devices, memories, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial and temporal redundancies. In one example, a video codec can use a technique called intra - prediction, which can compress an image based on spatial redundancy. For example, intra - prediction can use reference data from a reconstructed current picture for sample prediction. In another example, a video codec can use a technique called inter - prediction, which can compress an image based on temporal redundancy. For example, inter - prediction can predict samples in the current picture according to a previously reconstructed picture using motion compensation. Motion compensation can be represented by a motion vector (MV). Summary of the Invention
[0005] Aspects of the present disclosure include methods and apparatuses for video encoding / decoding.
[0006] In one aspect, a method of processing visual media data includes processing a bitstream of the visual media data according to formatting rules. The bitstream includes a skip flag for a current block, and the skip flag for the current block indicates whether at least one residual - related syntax element for the current block is written when encoding the current block in the current picture in an intratemplate matching prediction (IntraTMP) mode. The formatting rules specify that when the skip flag indicates that the residual - related syntax element is not written, the current block is processed by directly copying the values of the processed samples in a predicted block indicated by a block vector (BV) of the current block in the current picture.
[0007] In one example, the formatting rules specify that the current block is a chrominance block of the chrominance component in an intra slice, and the chrominance component and the luminance component in the intra slice are separately partitioned.
[0008] In one aspect, a method for video coding includes determining whether to apply residual coding to a current block in a current picture encoded in one of an intra block copy (IBC) mode and an intra template matching prediction (IntraTMP) mode. When it is determined that residual coding is not applied to the current block, the method for video coding includes encoding, in a bitstream, a syntax element that indicates that no residual-related syntax elements are written for the current block, where the current block is predicted by directly copying values of samples in a predicted block indicated by BV in the current picture.
[0009] In one example, the syntax element is a skip flag of the current block.
[0010] In one example, the current block is encoded in the IntraTMP mode and the current block is a luminance block. The method for video coding includes determining a context of a skip flag of the current block based on skip flags of neighboring blocks of the current block, where the neighboring blocks are encoded in one of the IBC mode and the IntraTMP mode, and encoding the skip flag of the current block using the determined context.
[0011] In one example, the current block is a chrominance block of a first chrominance component in an intra slice, and the luminance component and the first chrominance component in the intra slice are separately partitioned.
[0012] According to one aspect of the present invention, an apparatus for video decoding includes a processing circuit. The processing circuit is configured to receive a bitstream including a syntax element. The syntax element indicates whether at least one residual-related syntax element for a current block in a current picture is written in the bitstream when the current block is encoded according to a predicted block in the current picture and an offset between the current block and the predicted block is indicated by BV. When the syntax element indicates that the residual-related syntax elements are not written, the processing circuit is configured to reconstruct the current block by directly copying values of reconstructed samples in the predicted block.
[0013] In one aspect, the syntax element is a skip flag of the current block.
[0014] In one aspect, the current block is encoded in one of the IBC mode and the IntraTMP mode.
[0015] In one example, the current block is encoded in IntraTMP mode and the current block is a luma block. The processing circuitry is further configured to determine the context of the skip flag of the current block based on the skip flags of the neighboring blocks of the current block and use the determined context to determine the value of the skip flag. The neighboring blocks are encoded in one of IBC mode and IntraTMP mode.
[0016] In one aspect, the current block is a chroma block of the first chroma component in an intra slice, and the luma component and the first chroma component in the intra slice are separately partitioned. In one example, the processing circuitry is further configured to determine the context of the skip flag of the current block based on the skip flag of the co-located luma block of the current block and use the determined context to determine the value of the skip flag of the current block. The co-located luma block is encoded in one of IBC mode and IntraTMP mode. In one example, the current block is encoded in direct block vector (DBV) mode, and the processing circuitry is configured to determine the BV of the current block according to the BV of the co-located luma block from the luma component.
[0017] Aspects of the present disclosure also provide an apparatus for video coding. The apparatus for video coding includes processing circuitry. The processing circuitry is configured to implement any of the methods for video coding described.
[0018] Aspects of the present disclosure also provide a method for video decoding. The method includes any of the methods implemented by an apparatus for video decoding.
[0019] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods for video decoding / encoding described. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0021] Figure 1 is a schematic diagram of an example of a block diagram of a communication system (100).
[0022] Figure 2 is a schematic diagram of an example of a block diagram of a decoder.
[0023] Figure 3 is a schematic diagram of an example of a block diagram of an encoder.
[0024] Figure 4AShows an example of intra picture block compensation such as the intra block copy (IBC) mode according to one aspect of the present disclosure.
[0025] Figure 4B Shows an example of intra picture block compensation according to one aspect of the present disclosure, which has a coding tree unit (CTU) size search range, and in some examples, a memory area is reused to search for a part of the left CTU.
[0026] Figure 5 Shows an example of the Intra Template Matching Prediction (IntraTMP) mode according to one aspect of the present disclosure.
[0027] Figure 6 Shows an example of the separate partitioning structure of the luma coding tree block (CTB) and chroma CTB in a CTU according to one aspect of the present disclosure.
[0028] Figure 7 Shows an example of different template types for determining the current block of a prediction block in an intra prediction method such as the IntraTMP mode according to one aspect of the present disclosure.
[0029] Figure 8 Shows an example of a method for deriving the BV of the current chroma block in the chroma component from the co-located luma region in the luma component, such as the Direct Block Vector (DBV) mode.
[0030] Figure 9 Shows a flowchart outlining the decoding process according to some aspects of the present disclosure.
[0031] Figure 10 Shows a flowchart outlining the encoding process according to some aspects of the present disclosure.
[0032] Figure 11 Is a schematic diagram of a computer system according to one aspect. Detailed Description
[0033] Figure 1 Shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is an application example of the disclosed subject matter, namely a video encoder and a video decoder located in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-supported applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0034] The video processing system (100) includes an acquisition subsystem (113), and the acquisition subsystem may include a video source (101) such as a digital camera. The video source creates an uncompressed video picture stream (102), for example. In an example, the video picture stream (102) includes samples taken by the digital camera. Compared with the encoded video data (104) (or the encoded video bitstream), the video picture stream (102), depicted as a thick line to emphasize the high data volume, can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the video picture stream (102), the encoded video data (104) (or the encoded video bitstream), depicted as a thin line to emphasize the lower data volume, can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 the client subsystem (106) and the client subsystem (108) in, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes an incoming copy (107) of the encoded video data and creates an output video picture stream (111) that can be presented on a display (112) (such as a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), the video data (107), and the video data (109) (such as the video bitstream) may be encoded according to certain video coding / compression standards. Examples of these standards include the ITU-T H.265 recommendation. In one example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the disclosed subject matter can be used in the context of VVC.
[0035] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).
[0036] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used to replace Figure 1 the video decoder (110) in the example.
[0037] A receiver (231) may receive, for example, one or more encoded video sequences included in a bitstream to be decoded by a video decoder (210). In one aspect, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not shown). The receiver (231) may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and an entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be provided external to the video decoder (210) (not shown). And in other cases, a buffer memory (not shown) may be provided external to the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be configured inside the video decoder (210) to, for example, handle presentation timing. And when the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may also be possible not to configure the buffer memory (215), or the buffer memory (215) may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (215) may also be required, which may be relatively large and may have an adaptive size, and may be implemented at least partially in an operating system or a similar element (not shown) external to the video decoder (210).
[0038] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequences. The categories of these symbols include information for managing the operation of the video decoder (210), and potential information for controlling a display device (212) (e.g., a display screen), etc., which is not part of the electronic device (230), but may be coupled to the electronic device (230), such as Figure 2As shown. The control information for the display device may be a parameter set segment (not labeled) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be carried out according to video coding techniques or standards and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (220) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0039] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0040] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. Which units are involved and the way they are involved may be controlled by the subgroup control information parsed by the parser (220) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (220) and multiple units below are not described.
[0041] In addition to the functional blocks already mentioned, the video decoder (210) may be conceptually divided into several functional units as described below. In practical implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide them into the functional units below.
[0042] The first unit is a scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) may output a block including sample values, and the sample values may be input into the aggregator (255).
[0043] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. The intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a surrounding block having the same size and shape as the block being reconstructed using the reconstructed information extracted from the current picture buffer (258). For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.
[0044] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-coded and potentially motion-compensated block. In this case, the motion compensation prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After motion compensation of the extracted samples according to the symbol (221) belonging to the block, these samples may be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (referred to as residual samples or a residual signal in this case), thereby generating output sample information. The motion compensation prediction unit (253) obtaining the prediction samples from an address within the reference picture memory (257) may be controlled by a motion vector, and the motion vector may be in the form of the symbol (221) for use by the motion compensation prediction unit (253), and the symbol (221) includes, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, a motion vector prediction mechanism, and so on.
[0045] The output samples of the aggregator (255) can be adopted by various loop filtering techniques in the loop filter unit (256). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream), and the parameters can be used in the loop filter unit (256) as symbols (221) from the parser (220). Video compression can also respond to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0046] The output of the loop filter unit (256) can be a sample stream that can be output to the display device (212) and stored in the reference picture memory (257) for subsequent inter-picture prediction.
[0047] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.
[0048] The video decoder (210) can perform decoding operations according to a predetermined video compression technique or standard such as ITU-T H.265. In the sense that the encoded video sequence conforms to the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence can comply with the syntax specified by the video compression technique or standard used. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0049] In one aspect, a receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by a video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.
[0050] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is disposed in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) may be used to replace Figure 1 the video encoder (103) in the example.
[0051] The video encoder (303) may receive video samples from a video source (301) (which is not Figure 3 part of the electronic device (320) in the example), and the video source may capture video images to be encoded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).
[0052] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device that stores previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following focuses on describing the samples.
[0053] According to one aspect, a video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other desired time constraint. Implementing an appropriate encoding speed is a function of a controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to these units. For simplicity, couplings are not labeled in the figures. Parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be used for other suitable functions that relate to the video encoder (303) optimized for a certain system design.
[0054] In some aspects, the video encoder (303) is configured to operate in an encoding loop. As a simple description, in an example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data. The reconstructed sample stream (sample data) is input into a reference picture memory (334). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples “seen” by the prediction part of the encoder are exactly the same as the sample values that the decoder will “see” when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related technologies.
[0055] The operation of the “local” decoder (333) may be the same as that of the “remote” decoder of the video decoder (210) described in detail above, for example, in combination with Figure 2 However, briefly referring additionally to Figure 2 , when symbols are available and the entropy encoder (345) and parser (220) can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).
[0056] In some aspects, decoder techniques other than parsing / entropy decoding present in the decoder exist in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on decoder operations. The description of encoder techniques can be simplified because encoder techniques are inverse to the decoder techniques described in detail. A more detailed description in certain areas is provided below.
[0057] During operation, in some examples, the source encoder (330) may perform motion-compensated predictive coding. With reference to one or more previously encoded pictures designated as "reference pictures" in a video sequence, the motion-compensated predictive coding performs predictive coding on an input picture. In this way, the coding engine (332) encodes the difference between a pixel block of the input picture and a pixel block of the reference picture, which can be selected as the prediction reference for the input picture.
[0058] The local video decoder (333) may decode the encoded video data of a picture that can be designated as a reference picture based on the symbols created by the source encoder (330). The operation of the coding engine (332) can advantageously be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 3 not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that can be performed by the video decoder on the reference picture and can store the reconstructed reference picture in the reference picture memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference picture that has the same content (in the absence of transmission errors) as the reconstructed reference picture to be obtained by a remote video decoder.
[0059] The predictor (335) may perform a prediction search for the coding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor (335) may operate on a per-pixel block basis of sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it can be determined that the input picture may have a prediction reference taken from multiple reference pictures stored in the reference picture memory (334).
[0060] The controller (350) may manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0061] The outputs of all the above functional units can be entropy - encoded in the entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable - length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0062] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) to prepare for transmission over the communication channel (360), which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter (340) can merge the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0063] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can generally be assigned to any of the following picture types:
[0064] Intra picture (I - picture), which can be a picture that is encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0065] Predictive picture (P - picture), which can be a picture that is encoded and decoded using intra - prediction or inter - prediction, where the intra - prediction or inter - prediction uses one motion vector and a reference index to predict the sample values of each block.
[0066] Bi - directional predictive picture (B - picture), which can be a picture that is encoded and decoded using intra - prediction or inter - prediction, where the intra - prediction or inter - prediction uses two motion vectors and a reference index to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0067] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the coding assignment of the corresponding pictures applied to the blocks. For example, blocks of I pictures can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of P pictures can be prediction-encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of B pictures can be prediction-encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0068] Video encoder (303) can perform encoding operations according to a predetermined video coding technology or standard such as ITU-T Recommendation H.265. In operation, video encoder (303) can perform various compression operations, including prediction coding operations that utilize temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technology or standard used.
[0069] In one aspect, transmitter (340) can transmit additional data when transmitting the encoded video. Source encoder (330) can include such data as part of what can be an encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0070] The captured video can be multiple source pictures (video pictures) in a time series. Intra-picture prediction (often simplified to intra-frame prediction) utilizes the spatial correlation in a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is called the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0071] In some aspects, bidirectional prediction techniques can be used in inter - picture prediction. According to bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.
[0072] In addition, merge mode techniques can be used in inter - picture prediction to improve coding efficiency.
[0073] According to some aspects disclosed in the present application, predictions such as inter - picture prediction and intra - picture prediction are performed on a block - by - block basis. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression. CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be recursively divided into one or more coding units (CUs) in a quadtree. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, four 32×32 - pixel CUs, or sixteen 16×16 - pixel CUs. In one example, each CU is analyzed to determine the prediction type for the CU, such as an inter - frame prediction type or an intra - frame prediction type. In addition, depending on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In one aspect, prediction operations in encoding (encoding / decoding) are performed on a prediction - block basis. Taking the luminance prediction block as an example of the prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0074] Note that any suitable technology may be used to implement video encoder (103) and video encoder (303), as well as video decoder (110) and video decoder (210). In one aspect, one or more integrated circuits may be used to implement video encoder (103) and video encoder (303), as well as video decoder (110) and video decoder (210). In another aspect, one or more processors executing software instructions may be used to implement video encoder (103), video encoder (303), as well as video decoder (110) and video decoder (210).
[0075] Video coding has been widely used in many applications, such as broadcasting, video recording, video streaming, etc. Many emerging video coding standards, such as H.264, H.265 / HEVC, H.266 / VVC, and AV1 have been released and widely adopted in video applications. In one aspect, a hybrid video codec may include the following coding modules, such as intra prediction, inter prediction, transform coding, quantization, entropy coding, post-processing loop filtering, etc.
[0076] In various examples, the current picture (which may also be interchangeably referred to as the current frame) may be used as a reference region for block-based compensation (such as in intra block copy (IBC), intra template matching prediction (IntraTMP) mode, etc.).
[0077] In one aspect, block-based compensation from different pictures may be referred to as motion compensation. Similarly, block compensation may be performed based on a previously reconstructed region within the same picture, and the block compensation may include intra picture block compensation (also referred to as current picture referencing (CPR) or IBC mode). Figure 4A An example of intra picture block compensation such as the IBC mode according to one aspect of the present disclosure is shown. A displacement vector indicating the offset between the current block (430) and the reference block (440) may be referred to as a block vector (BV) (450). The current block (430) and the reference block (440) are in the current picture (400).
[0078] Different from the MV in motion compensation (the MV can be any value (positive or negative, in the x or y direction)), the BV may have some constraints such that the pointed-to reference block is available and has been reconstructed. In the example, the reference Figure 4A , the current picture (400) may include a region to be decoded (420) and a reconstructed region (410). In one example, the BV may be constrained to point to a reference block in the reconstructed region (410). In some examples, for parallel processing considerations, some reference regions that are tile boundaries or wavefront trapezoid boundaries may be excluded.
[0079] Block matching (BM) can be performed at the encoder to find the best BV for each CU. On the encoder side, hash-based motion estimation can be performed for the IBC mode. The encoder can perform rate-distortion (RD) checks on blocks with a width or height not greater than 16 luma samples. For non-merge modes, BV search can be first performed using hash-based search. If the hash-based search does not return a valid candidate, block-matching based local search can be performed.
[0080] In the hash-based search, the hash key matching (32-bit cyclic redundancy check (CRC)) between the current block and the reference block can be extended to all allowed block sizes. The hash key calculation for each position in the current picture can be based on 4×4 sub-blocks. For a current block of a larger size, when all the hash keys of the corresponding 4×4 sub-blocks match the hash keys in the corresponding reference positions, it can be determined that the hash key of the current block matches the hash key of the reference block. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the BV cost for each matching reference can be calculated, and the reference with the minimum cost can be selected.
[0081] In the block-matching search, the search range can be set to include both the previous CTU and the current CTU.
[0082] The encoding of the BV can be explicit or implicit. In the explicit mode, the difference between the BV and the BV prediction can be written. In the implicit mode, in a manner similar to the MV in the merge mode, the BV can be recovered only based on the BV prediction value. In some embodiments, the resolution of the BV can be limited to integer positions; in other systems, the resolution of the BV can be allowed to point to fractional positions.
[0083] In one example, intra block copy is regarded as an inter mode (also interchangeably referred to as an inter prediction mode). By regarding intra block copy as an inter mode, block vector prediction in the implicit mode and the explicit mode can be analogous to the merge mode and the AMVP mode respectively. In one example, intra block copy is regarded as a third mode different from the intra prediction mode or the inter prediction mode. By regarding intra block copy as a third mode, block vector prediction in the implicit mode and the explicit mode can be separated from the conventional inter modes. In one example, the above-mentioned explicit mode can be referred to as the IBC AMVP mode, and the above-mentioned implicit mode can be referred to as the IBC merge mode. For example, a separate merge candidate list is defined for the IBC mode (e.g., the IBC merge mode), where, for example, all entries in the list are BVs. Similarly, in one example, the block vector prediction list in the IBC AMVP mode consists only of BVs. In some examples, the general rules applied to both lists include: in terms of the candidate derivation process, both lists can follow the same logic as the inter merge candidate list used in the inter mode or the AMVP prediction list used in the inter mode. For example, five spatially adjacent positions in an inter merge mode such as the HEVC or VVC inter merge mode can be accessed for the IBC mode to derive its own merge candidate list.
[0084] A block-level flag (referred to as the IBC flag) can be used to write the use of intra block copy at the block level. In one aspect, the IBC flag is written when the current block is not encoded in the merge mode. In one example, the IBC flag can be written by a reference index method (e.g., by regarding the current decoded picture as a reference picture). In one example, such as in HEVC SCC, such a reference picture (e.g., the current decoded picture) is placed at the last position in a list (e.g., the reference picture list). The special reference picture (e.g., the current decoded picture) can be managed together with other temporal reference pictures in the decoded picture buffer (DPB).
[0085] In one example, at the CU level, the IBC mode can be written with a flag and can be indicated as the IBC AMVP mode or the IBC skip / merge mode as follows: - IBC skip / merge mode: The merge candidate index is used to indicate which BVs from adjacent candidate IBC-encoded blocks in the merge list are used to predict the current block. The merge list can include spatial candidates, HMVP candidates, and paired candidates, or the merge list consists of spatial candidates, HMVP candidates, and paired candidates. -IBC AMVP mode: The BV difference is encoded in the same way as the MV difference. The BV prediction method can use two candidates as the predicted value (e.g., the BV predicted value): one from the left adjacent block (if the left adjacent block is encoded by IBC), and the other from the upper adjacent block (if the upper adjacent block is encoded by IBC). When either adjacent block is not available, the default BV can be used as the predicted value (e.g., the BV predicted value). A write flag is used to indicate the BV predicted value index.
[0086] Figure 4B An example of intra-picture block compensation according to an aspect of the present disclosure is shown, which has a CTU size search range and, in some examples, reuses memory for searching a portion of the left CTU.
[0087] In some examples, such as in VVC, the search range of the IBC mode is constrained within the current CTU. In one example, the effective memory area requirement for storing the reference samples of the IBC mode is one CTU size of samples. Considering that the existing reference sample memory area is used to store the reconstructed samples in the current 64×64 region, three additional 64×64 sized reference sample memory areas can be used. Therefore, a method can be used to extend the effective search range of the IBC mode to a portion of the left CTU while the total memory area requirement for storing the reference pixels can remain unchanged. For example, the total memory area requirement is one CTU size, such as a total of four 64×64 reference sample memory areas. Figure 4B An example of this memory area reuse mechanism is shown. Each vertically striped block is the current coding region (Curr), and the samples in each gray area are the encoded samples. The crossed-out areas (marked with "X") are not available for reference because the crossed-out areas can be replaced by the coding region in the current CTU in the reference sample memory area.
[0088] Figure 5 An example of the IntraTMP mode according to an aspect of the present disclosure is shown. In one example, the IntraTMP mode is a special intra prediction mode. In one example, the IntraTMP mode is different from the intra prediction mode. Refer to Figure 5, in an example of the IntraTMP mode, a prediction block (521) such as the best prediction block from the reconstructed part of the current frame can be copied. A template (520) such as an L-shaped template of the prediction block (521) can match the current template (530) of the current block (531). For a predefined search range, the encoder can search for the template most similar to the current template (530) in the reconstructed part of the current frame and use the corresponding block (521) as the prediction block. In one example, the encoder then writes the use of the IntraTMP mode and performs the same prediction operation on the decoder side.
[0089] Reference Figure 5 , a prediction signal can be generated by matching the L-shaped causal neighbor of the current block (531) with another block in a predefined search area. In one example, the predefined search area includes or consists of: R1, R2, R3, and R4. R1 is the current CTU; R2 is the upper left CTU of the current CTU; R3 is the upper CTU of the current CTU, and R4 is the left CTU of the current CTU. In one example, the sum of absolute differences (SAD) is used as the cost function.
[0090] Within each region, the decoder can search for the template with the minimum cost (e.g., minimum SAD) relative to the current template and can use the block corresponding to the minimum cost as the prediction block.
[0091] The sizes (SearchRange_w, SearchRange_h) of all regions can be set to be proportional to the block sizes (BlkW, BlkH) so that there is a fixed number of SAD comparisons per pixel. In one example, SearchRange_w = a × BlkW and SearchRange_h = a × BlkH, where "a" is a constant that controls the trade-off between gain and complexity. In the example, "a" is equal to 5.
[0092] To accelerate the template matching process, in some examples, the search ranges of all search regions are subsampled by a factor of 2, for example, so that the template matching search is reduced to 1 / 4. After finding the best match, a refinement process can be performed. The refinement process can be performed via a second template matching search around the best match with a reduced range. In one example, the reduced range is defined as min(BlkW, BlkH) / 2.
[0093] For CUs with both width and height dimensions less than or equal to 64, the intra-template matching tool can be enabled. The maximum CU size for IntraTMP mode can be configurable. For example, when decoder-side intra mode derivation (DIMD) is not used for the current CU, the IntraTMP mode can be written at the CU level through a dedicated flag.
[0094] Various partitions can be applied in video and / or image coding, such as in VVC or HEVC. A picture can be partitioned into coding tree units (CTUs). For example, a picture is partitioned into a sequence of CTUs. In some examples, for a picture with three sample arrays, a CTU can include one N×N block of luma samples and two corresponding chroma sample blocks.
[0095] In one example, the maximum allowed size of the luma block in a CTU is specified as 128×128. In one example, the maximum size of a luma transform block is 64×64.
[0096] A tree structure can be used to partition a CTU. In some examples, such as in HEVC, a CTU is split into multiple CUs using a quaternary-tree (QT) structure represented as an encoding tree to accommodate various local characteristics. A decision can be made at the leaf CU level on whether to use inter-picture (temporal) prediction or intra-picture (spatial) prediction to encode a picture region. Depending on the PU split type, each leaf CU can be further split into one, two, or four PUs. Within a PU, the same prediction process can be applied, and relevant information can be sent to the decoder based on the PU. After obtaining the residual block by applying the prediction process based on the PU split type, the leaf CU can be partitioned into transform units (TUs) according to another quaternary-tree structure similar to the encoding tree for the CU. In one example, the HEVC structure has multiple partitioning concepts including CUs, PUs, and TUs.
[0097] In some examples, such as in VVC, a quaternary tree with a multi-type tree (MTT) having a split structure using binary and ternary splits can replace the concept of multiple partitioning unit types. For example, the separation of CU, PU, and TU concepts can be removed, for example, except for CUs whose size is too large for the maximum transform length, and more flexible CU partitioning shapes are supported. In the encoding tree structure, a CU can have any suitable shape, such as a square or rectangular shape. A CTU can first be partitioned by a quaternary-tree (quad-tree) structure. Then the leaf nodes of the quaternary tree can be further partitioned by a multi-type tree structure.
[0098] The coding tree scheme can support the ability for the luma and chroma to have the same block tree structure or independent block tree structures. For P and B slices, in some examples, the luma CTB and chroma CTB in a CTU can share the same coding tree structure. The coding tree type can be a dual-tree type (also known as chroma-independent tree, such as DUAL_TREE_CHORMA) used in VVC, and the luma CTB and chroma CTB in a CTU can have independent block tree structures. For I slices, in some examples, the luma CTB and chroma CTB in a CTU can have independent block tree structures. When applying the independent block tree mode, one luma CTB can be divided into multiple CUs (e.g., luma CUs) through one coding tree structure, and the chroma CTB can be divided into multiple chroma CUs (e.g., chroma CUs) through another coding tree structure. Therefore, unless the video is monochrome, a CU in an I slice can include a coded block of one luma component or two chroma components or be composed of a coded block of one luma component or two chroma components, and a CU in a P or B slice includes coded blocks of all three color components or is composed of coded blocks of all three color components.
[0099] Figure 6 An example of the independent partitioning structure of the luma CTB (1020) and chroma CTB (610) in a CTU (602) according to one aspect of the present disclosure is shown. The CTU (602) is in a picture (601).
[0100] The chroma CTB (610) can be co-located with the luma CTB (620). For example, the chroma CTB (610) and the luma CTB (620) correspond to the same physical region (602) in the picture (601). The sizes (e.g., width and / or height) of the chroma CTB (610) and the luma CTB (620) can be indicated by a chroma subsampling format (also known as sampling format) such as 4:2:0, 4:2:2, etc. In Figure 6 the example shown, the chroma format sampling structure is 4:2:0, and the width and height of the chroma CTB (610) are each 1 / 2 of the width and height of the luma CTB (620).
[0101] For the chroma-independent tree, two independent coding tree structures can be used to partition the luma CTB (620) and chroma CTB (610) in the CTU (602). In one example, the picture (601) is an intra picture (I picture). Independent intra-luma / chroma coding tree structures can be used to partition the luma CTB (620) and chroma CTB (610).
[0102] Any suitable coding tree structure can be used to partition the luminance CTB (620). Any suitable coding tree structure can be used to partition the chrominance CTB (610). In Figure 6 the example shown, the chrominance CTB (610) is partitioned into chrominance blocks (611)-(612) by, for example, a binary tree. In Figure 6 the example shown, the luminance CTB (620) is partitioned into luminance blocks (621)-(629) by, for example, quadtree splitting, binary tree, etc. For example, the luminance CTB (620) is partitioned into 4 blocks using quadtree splitting. The upper left block of the luminance CTB (620) is further partitioned into luminance blocks (621)-(622) using a binary tree. The upper right block of the luminance CTB is further partitioned into luminance blocks (627)-(628) using a binary tree. The lower left block of the luminance CTB (620) is further partitioned into luminance blocks (623)-(626) using quadtree splitting. The lower right block of the luminance CTB (620) is not partitioned and is the luminance block (629).
[0103] When using a chroma separate tree (CST), the current chrominance block (611) can be co-located with one or more luminance blocks in the luminance CTB (620). The number of luminance blocks co-located with the current chrominance block (611) can depend on the coding tree structure used to partition the CTU (602). Referring to Figure 6 , the current chrominance block (611) can be co-located with the luminance blocks (621)-(626).
[0104] According to one aspect of the present disclosure, a first method (e.g., the first method of intra prediction) can be designed to determine (e.g., find) a prediction block based on the reconstructed region in the current picture (also referred to as the current frame). The current picture will be encoded or decoded. In one example, the prediction block can be determined by using at least one written (e.g., such as in the IBC AMVP mode) or inherited (e.g., IBC merge / skip mode) block vector (BV). The written BV or the inherited BV can be used to indicate the displacement from the current block to the reconstructed prediction block within the current picture. The current block can be encoded or decoded based on the prediction block. An example of the first method of intra prediction is a variant of the intra block copy mode or IBC mode described above (such as in Figures 4A - 4B ).
[0105] The intra block copy mode can be considered a specific type of prediction mode or an independent prediction mode. In one example, the IBC mode can be considered a method of intra prediction because the prediction block is in the current picture where the current block is located. However, if the IBC mode is considered a method of inter prediction or a third mode different from intra prediction and inter prediction, the methods described in the present disclosure can also be applicable or can be appropriately adapted.
[0106] The second method of intra prediction can be designed to determine (e.g., find) a predicted block (e.g., the best predicted block) of a current block in a current picture based on the reconstructed region in the current picture. In the second method of intra prediction, a template can be used and the predicted block can be determined without BV signaling. Any suitable template shape and / or any suitable template size can be used. The template of the current block can be referred to as the current template and can be adjacent to the current block. In one aspect, the current template can include reconstructed samples in the current picture.
[0107] Figure 7 Examples of different template types of a current block for determining a predicted block in the second method of intra prediction according to one aspect of the present disclosure are shown. The template type of the current block can include a plurality of templates, and the plurality of templates include different combinations of samples above and / or to the left of the current block. For example, the template type includes a left template T l (734), an above template (also referred to as a top template) T a (733), a template T a+l (732), an L-shaped template T L (731), etc. The left template T l can include reconstructed samples to the left of the current block (730). The above template T a can include reconstructed samples above (e.g., directly above) the current block (730). The template T a+l (732) can include a combination of the left template T l and the top template T a . The L-shaped template T L (731) can be L-shaped and include the left template T l , the top template T a , and the upper left corner between the left template T l and the top template T a .
[0108] In the second method of intra prediction, for a predefined search range, the encoder can search the reconstructed part of the current picture for a template that is most similar to the current template (e.g., T L , T a+l , T a or T l ), and can use the corresponding block as the predicted block. In one example, the encoder then writes the use of the mode (e.g., the second method of intra prediction), and the same prediction operation can be performed on the decoder side. It can be done by using the current template (e.g., Figure 7 the T shown in L , T a+l , T a or Tl ) generating a prediction signal by matching with a reference template of another block in a predefined search region. In one example, the current template has been reconstructed and is a causal neighbor of the current block. An example of the second method of intra prediction is the IntraTMP mode described in Figure 5 . Examples of templates such as Figure 7 shown in L , T a+l , T a or T l and the like can be used for the IntraTMP mode.
[0109] For example, when the encoded slice is an intra slice, a third method can be applied to separately partition the luminance component and the chrominance component. For a superblock (SB) or CTU including luminance and chrominance components, for example, separate partitioning can be applied in the intra slice. In this case, the CTU can include a luminance CTU and a chrominance CTU, or the SB can include a luminance SB and a chrominance SB. The luminance CTU (or luminance SB) forms a partition (coding tree). The chrominance CTU (or chrominance SB) can include two chrominance channels (also referred to as two chrominance components), such as two chrominance CTBs, and can form a chrominance partition. The two chrominance channels can use the same chrominance partition. Examples of the third method are such as Figure 6 shown in
[0110] Figure 8 shows an example of a fourth method for deriving the BV of the current chrominance CU (e.g., current chrominance block) (811) in the chrominance component from the co-located luminance region (812) in the luminance component according to an aspect of the present disclosure. To improve coding efficiency, the fourth method can be used to directly reuse the BV (BV) of the co-located luminance region (812) in the luminance component determined using the first method or the second method. For example, when the encoded slice is an intra slice, the third method is applied to the encoded slice, and the color component of the current block (e.g., current chrominance block (811)) is the chrominance component, and the co-located luminance region (812) in the luminance component is encoded using the first method or the second method.
[0111] Refer to Figure 8, the CTUs in the current picture may include a luminance CTB (801) and a chrominance CTU (802), and the chrominance CTU (802) may be co-located with the luminance CTB (801). The chrominance CTU (802) may include one or more chrominance CTBs. The current picture is in an intra slice. A third method, such as a dual-tree partitioning (e.g., chrominance-independent tree), is applied to partition the CTU. Thus, two independent coding tree structures can be used to partition the luminance CTB (801) and the chrominance CTU (802). Any suitable coding tree structure can be used to partition the luminance CTB (801). Any suitable coding tree structure can be used to partition the chrominance CTU (802). In one example, the chrominance CTU (802) includes two chrominance CTBs corresponding to the chrominance components Cr and Cb. The two chrominance CTBs may share the same coding tree structure.
[0112] In Figure 8 the example shown, the chrominance CTU (802) is partitioned into two chrominance CUs (811) and (813), and the luminance CTB (801) is partitioned into luminance blocks (831) to (843). The chrominance CU (811) is co-located with a luminance region (812). The luminance region (812) may overlap with one or more luminance blocks in the luminance CTB (801). In Figure 8 the example shown, the luminance region (812) includes luminance blocks (831)-(840). In one example, one or more predefined positions in the luminance region (812) may be checked to determine whether a BV (e.g., represented by bvL) is used to encode a luminance block associated with one of the one or more predefined positions. In one example, the one or more predefined positions (also referred to as one or more predefined luminance positions) in the luminance region (812) include Figure 8 the center position C, top-left position (TL), top-right position (TR), bottom-left (BL) position, and bottom-right (BR) position indicated in
[0113] An example of a fourth method is the direct block vector (DBV) mode. For example, when the chrominance dual-tree is activated in an intra slice, for a chrominance CU (e.g., (811)) encoded using the DBV mode, if the five positions ( Figure 8If one of the luminance blocks (C, TL, TR, BL, and BR shown in the figure) is encoded using the IBC mode or the IntraTMP mode, the bvL of one of the luminance blocks can be used to derive the chrominance block vector bvC of the chrominance CU (811). In Figure 8 In the example shown, the luminance blocks (836), (831), (833), (838), and (840) are associated with positions C, TL, TR, BL, and BR, respectively.
[0114] Reference Figure 8 , when checking the predefined positions, position C is determined to be associated with the BV (821) used to encode the luminance block (836). The bvL or BV (821) can be used to determine the bvC, which is the BV (822) of the current chrominance CU (811). In one example, the bvC (822) is the bvL (821) scaled according to the chrominance subsampling format. In one example, when the chrominance subsampling format is 4:2:0, bvC = bvL / 2.
[0115] Aspects of the present disclosure provide techniques, apparatuses, and methods related to intra prediction coding tools without using residual coding, such as non-residual coding that can use a prediction block in a current picture to predict the prediction mode of a current block in the current picture. The prediction mode can include a first method, a second method, a fourth method, etc. The first method can include the IBC mode. The first method and the IBC mode can be considered intra prediction coding tools in the present disclosure because the prediction block and the current block are in the same current picture. The second method can include the IntraTMP mode. The fourth method can include the DBV mode. The term "IBC mode" can refer to the IBC mode or its variants, such as the IBC skip mode, the IBC merge mode, or the IBC AMVP mode. The term "IntraTMP mode" can refer to the IntraTMP mode or its variants. The term DBV mode can refer to the DBV mode or its variants.
[0116] According to one aspect of the present disclosure, the current block in the current picture is encoded based on a prediction block in the current picture, and the offset between the current block and the prediction block is indicated by the BV. The current block can be encoded using the first method (e.g., the IBC mode), the second method (e.g., the IntraTMP mode), the fourth method (e.g., the DBV mode), etc. On the decoder side, a bitstream can be received. The bitstream can include a syntax element that indicates whether at least one residual-related syntax element is written into the bitstream.
[0117] When a syntax element indicates that residual-related syntax elements are not written, the current block can be reconstructed by directly copying the values of the reconstructed samples in the prediction block. In one aspect, the current block can be encoded using a skip mode. On the encoder side, the skip transformation is used, so the residual, such as the difference between the prediction block and the current block, is not transformed and thus not encoded into the bitstream. Additionally, the residual-related syntax elements are not encoded into the bitstream. On the decoder side, when the syntax element indicates that the residual-related syntax elements are not written, the current block is reconstructed by directly copying the values of the reconstructed samples in the prediction block.
[0118] In an example, the syntax element is the skip flag of the current block. The skip flag can be written to indicate whether to skip one or more residual-related signaling syntax elements (e.g., at least one residual-related syntax element). If the skip flag is true, there is no residual-related signaling syntax after the skip flag, thus reducing the signaling overhead. In one example, the residual-related signaling syntax elements include the code block pattern (CBP).
[0119] In the following description, the skip flag refers to the skip flag used to determine whether the current block has residual data. Thus, if the skip flag is true, the skip mode is applied to the current block, the current block does not have residual data (e.g., the residual data of the current block is not transformed and not encoded into the bitstream), and there is no residual-related signaling syntax after the skip flag.
[0120] In one aspect, the current block can be encoded using one of a first method (e.g., IBC mode) and a second method (e.g., IntraTMP mode). The skip flag can be written to indicate whether to skip the residual-related signaling syntax elements when encoding the current block using the first method (e.g., IBC mode) or the second method (e.g., IntraTMP mode). When the skip flag is true, the residual-related syntax elements are not written, and the reconstructed samples of the current block are directly copied from the samples of the prediction block in the current picture.
[0121] In one aspect, the current block (e.g., the current luma block) is encoded using a second method (e.g., IntraTMP mode), and the current block is a luma block in the luma component. For example, when encoding the current block in the luma component using the second method, a skip flag is written for the current block. In one example, the context of the skip flag for the current block (e.g., the current luma block) can be determined based on the skip flags of the neighboring blocks of the current block (e.g., neighboring coded blocks), and the determined context can be used to determine the value of the skip flag for the current luma block. One of the first method (e.g., IBC mode) and the second method (e.g., IntraTMP mode) can be used to encode the neighboring blocks. For example, the context of the skip flag for the current luma block is determined based on the state (e.g., true or false) of the skip flag of the neighboring coded block when encoding the neighboring coded block using the second method. In one example, the context of the skip flag for the current luma block is determined based on the state of the skip flag of the neighboring coded block when encoding the neighboring coded block using the first method or the second method.
[0122] In one aspect, the current block (e.g., the current chroma block) is a chroma block of the first chroma component in an intra slice, and the first chroma component and the luma component in the intra slice are separately partitioned, such as using a third method (e.g., Figure 6 the dual-tree partitioning described in). In one example, for instance, when encoding the current chroma block using the first method or the second method and applying the third method to the intra slice, a skip flag is written for the current block (e.g., the current chroma block) in the chroma component (e.g., the first chroma component) of the intra slice.
[0123] In one aspect, the context of the skip flag for the current chroma block can be determined based on the skip flag of the co-located luma block of the current chroma block. One of the first method (e.g., IBC mode) and the second method (e.g., IntraTMP mode) can be used to encode the co-located luma block. The determined context of the skip flag for the current chroma block can be used to determine the value of the skip flag for the current chroma block. In one example, the context of the skip flag for the current chroma block in the first chroma component is determined based on the state (e.g., true or false) of the skip flag of the co-located block (i.e., the co-located luma block) in the luma component and whether the co-located block is encoded using the first method or the second method. In one example, the context of the skip flag for the current chroma block in the first chroma component is determined based on the state of the skip flag of the co-located block in the luma component and whether the co-located block is encoded using the second method. In one example, the context of the skip flag for the current chroma block in the first chroma component is determined based on the state of the skip flag of the co-located block in the luma component and whether the co-located block is encoded using the first method.
[0124] In one aspect, the context of the skip flag of the current chrominance block can be determined based on the skip flag of a co-located block of the current chrominance block, where the co-located block can be in the luma component or in the second chrominance component. The determined context of the skip flag of the current chrominance block can be used to determine the value of the skip flag of the current chrominance block. In one example, the context of the skip flag of the current chrominance block in the first chrominance component is determined based on the status of the skip flag of a co-located block (e.g., co-located luma block) in the luma component. If the co-located block in the luma component does not have a skip flag, the status of the skip flag of the current chrominance block can be inferred as false. In one example, the context of the skip flag of the current chrominance block in the first chrominance component (e.g., Cr) is determined based on the status of the skip flag of a co-located block (e.g., co-located chrominance block) in the second chrominance component (e.g., Cb).
[0125] In one aspect, the current block is encoded using a fourth method (e.g., DBV mode), and as Figure 8 shown, the BV of the current block (e.g., bvC) can be determined based on the BV of a co-located luma block from the luma component (e.g., bvL). In one example, the current block has a first chrominance component. A skip flag can be written to indicate whether to skip residual-related signaling syntax elements when encoding the current block using the fourth method.
[0126] In one aspect, when encoding the current block using the fourth method, the context of the skip flag of the current block is determined based on the status of the skip flag of a co-located block in an alternative color component (e.g., the luma component or a second chrominance component different from the first chrominance component) and whether the co-located block in the alternative color component is encoded using the first method or the second method. For example, the alternative color component is the luma component, the co-located block in the alternative color component is a co-located luma block, and the co-located luma block is encoded using one of the first method (e.g., IBC mode) and the second method (e.g., IntraTMP mode), so the context of the skip flag of the current block can be determined based on the skip flag of the co-located luma block. The determined context can be used to determine the value of the skip flag of the current block.
[0127] In one aspect, when encoding the current block using the fourth method, the context of the skip flag of the current block is determined based on the status of the skip flag of a co-located block in an alternative color component and whether the co-located block in the alternative color component is encoded using the second method.
[0128] In one aspect, when encoding the current block using the fourth method, the context of the skip flag of the current block is determined based on the status of the skip flag of a co-located block in an alternative color component and whether the co-located block in the alternative color component is encoded using the first method.
[0129] In one aspect, the context of the skip flag of the current block encoded by the fourth method is determined based on the state of the skip flag of the co-located block in the alternative color component. If the co-located block in the alternative color component does not have a skip flag, the state of the skip flag of the co-located block can be inferred as false. For example, when the co-located block in the alternative color component does not have a skip flag, the skip flag of the current block encoded by the fourth method can be inferred as false.
[0130] In one aspect, a skip flag associated with the corresponding luminance position of the co-located luminance region can be obtained. Such as Figure 8 shown, the luminance positions can include C, TL, TR, BL, and BR. The context of the skip flag of the current block encoded by the fourth method can be determined based on the skip flag associated with the corresponding luminance position. The determined context can be used to determine the value of the skip flag of the current block. For example, the context of the skip flag of the current block can be determined based on the state(s) of the skip flag(s) associated with the luminance position(s) (e.g., Figure 8 all the luminance positions C, TL, TR, BL, and BR shown).
[0131] In one example, the scan of the luminance positions can determine the number of skip flags associated with the luminance positions in the luminance component. The context of the skip flag of the current block encoded by the fourth method can be determined based on the number of skip flags associated with the luminance positions in the luminance component.
[0132] In one example, when the skip flag is true in a predefined scan order (e.g., C, TL, TR, BL, and BR, where position C is scanned first and position BR is scanned last), the scan can determine the skip flag of the first position.
[0133] In one example, if the luminance skip flags are false for all scanned positions, the skip flag of the current block encoded by the fourth method is inferred as false. Referring to Figure 8 , if the skip flags associated with the luminance positions C, TL, TR, BL, and BR are false or do not exist, the skip flag of the current block encoded by the fourth method is inferred as false.
[0134] Figure 9FIG. 0 shows a flowchart outlining a process (900) according to one aspect of the present disclosure. The process (900) can be used in a device such as a video decoder. In various aspects, the process (900) is executed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some aspects, the process (900) is implemented as software instructions, such that when the processing circuit executes the software instructions, the processing circuit performs the process (900). The process starts at (S901) and proceeds to (S910).
[0135] At (S910), a bitstream including syntax elements is received. The syntax elements indicate whether at least one residual-related syntax element for a current block in the current picture is written to the bitstream when the current block is encoded based on a predicted block in the current picture and the offset between the current block and the predicted block is indicated by a block vector (BV).
[0136] In one example, the syntax element is a skip flag for the current block.
[0137] In one example, the current block is encoded using one of an intra block copy (IBC) mode and an intra template matching prediction (IntraTMP) mode.
[0138] In one example, the current block is encoded in the IntraTMP mode and the current block is a luminance block. The context of the skip flag for the current block (e.g., the current luminance block) is determined based on the skip flags of adjacent blocks of the current luminance block, where the adjacent blocks are encoded using one of the IBC mode and the IntraTMP mode. The value of the skip flag for the current luminance block is determined using the determined context.
[0139] In one example, the current block (e.g., the current chrominance block) is a chrominance block of a first chrominance component in an intra slice, and the first chrominance component and the luminance component in the intra slice are separately partitioned, for example, using a third method (e.g., dual-tree partitioning). In one example, the context of the skip flag for the current chrominance block is determined based on the skip flag of the co-located luminance block of the current chrominance block, where the co-located luminance block is encoded using one of the IBC mode and the IntraTMP mode; and the value of the skip flag for the current chrominance block is determined using the determined context. In one example, the context of the skip flag for the current chrominance block is determined based on the skip flag of the co-located block of the current chrominance block, where the co-located block belongs to the luminance component or a second chrominance component; and the value of the skip flag for the current chrominance block is determined using the determined context.
[0140] In one example, the current block is encoded in direct block vector (DBV) mode, and the BV (e.g., bvL) of the current chroma block is determined based on the BV (e.g., bvC) of the co-located luma block from the luma component.
[0141] In one example, the co-located luma block is encoded in one of the IBC mode and the IntraTMP mode, the context of the skip flag of the current chroma block is determined based on the skip flag of the co-located luma block, and the value of the skip flag of the current chroma block is determined using the determined context.
[0142] In one example, the skip flag associated with the corresponding luma position of the co-located luma block is obtained, and the luma positions include the center position C, the top-left position (TL), the top-right position (TR), the bottom-left (BL) position, and the bottom-right (BR) position. The context of the skip flag of the current chroma block is determined based on the skip flag associated with the corresponding luma position, and the value of the skip flag of the current chroma block is determined using the determined context.
[0143] At (S920), when the syntax element indicates that the residual-related syntax element is not written, the current block is reconstructed by directly copying the values of the reconstructed samples in the prediction block.
[0144] Then, the process proceeds to (S999) and ends.
[0145] The process (900) can be adjusted appropriately. The steps in the process (900) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.
[0146] Figure 10 A flowchart outlining a process (1000) according to one aspect of the present disclosure is shown. The process (1000) can be used in a video encoder. In various aspects, the process (1000) is executed by a processing circuit, such as a processing circuit that executes the functions of the video encoder (103), a processing circuit that executes the functions of the video encoder (303), etc. In some aspects, the process (1000) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (1000). The process starts at (S1001) and proceeds to (S1010).
[0147] At (S1010), it is determined whether residual coding is applied to the current block in the current picture encoded in one of the intra block copy (IBC) mode and the intra template matching prediction (IntraTMP) mode.
[0148] At (S1020), when it is determined that residual coding is not applied to the current block, a syntax element is coded in the bitstream that indicates that no residual-related syntax elements are written for the current block. The current block is predicted by directly copying the values of the samples in the predicted block in the current picture indicated by the block vector (BV).
[0149] In one example, the syntax element is the skip flag of the current block.
[0150] Then, the process proceeds to (S1099) and ends.
[0151] The process (1000) can be adjusted appropriately. The steps in the process (1000) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.
[0152] In one example, the current block is coded in IntraTMP mode and the current block is a luma block. In one example, the context of the skip flag of the current block is determined based on the skip flags of the neighboring blocks of the current block. The neighboring blocks are coded in one of IBC mode and IntraTMP mode, and the skip flag of the current block is coded using the determined context.
[0153] In one example, the current block is a chroma block of the first chroma component in an intra slice, and the first chroma component and the luma component in the intra slice are separately partitioned.
[0154] In one aspect, a method of processing visual media data includes processing a bitstream of visual media data according to format rules. For example, the bitstream can be a bitstream decoded / encoded by any decoding and / or encoding method described herein. The format rules can specify one or more constraints of the bitstream and / or one or more processes performed by a decoder and / or an encoder.
[0155] The bitstream includes a first syntax element (e.g., the skip flag of the current block) that indicates whether at least one residual-related syntax element for the current block is written when the current block in the current picture is coded in intra template matching prediction (IntraTMP) mode. The format rules specify that when the first syntax element indicates that the residual-related syntax elements are not indicated, the current block is processed by directly copying the values of the processed samples in the predicted block in the current picture indicated by the block vector (BV) of the current block.
[0156] In one example, the format rules specify that the current block is a chroma block of a chroma component in an intra slice, and the chroma component and the luma component in the intra slice are separately partitioned.
[0157] The methods, aspects, and examples described in this disclosure may be used alone or in any combination. For example, some aspects and / or examples performed by a decoder may be performed by an encoder, and vice versa. Each of the method (or aspect), encoder, and decoder may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0158] The above techniques may be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media. For example, Figure 11 FIG. shows a computer system (1100) suitable for implementing certain aspects of the disclosed subject matter.
[0159] The computer software may be encoded using any suitable machine code or computer language, and any suitable machine code or computer language may be subject to mechanisms such as assembly, compilation, linking, or the like to create code that includes instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretive microcode, etc.
[0160] The instructions may be executed on various types of computers or their components, such as personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0161] Figure 11 The components of the illustrated computer system (1100) are exemplary and are not intended to impose any limitation on the scope of use or functionality of the computer software for implementing aspects of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary aspects of the computer system (1100).
[0162] The computer system (1100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, the following: tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, captured images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.
[0163] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (1101), mouse (1102), touchpad (1103), touch screen (1110), data glove (not shown), joystick (1105), microphone (1106), scanner (1107), camera (1108).
[0164] The computer system (1100) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users, for example, through haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (such as haptic feedback of the touch screen (1110), data glove (not shown), or joystick (1105), but may also be haptic feedback devices that are not input devices), audio output devices (such as: speakers (1109), headphones (not shown)), visual output devices (such as screens (1110) including CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input function, each screen having or not having haptic feedback function - some of which can output two-dimensional visual output or output beyond three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted).
[0165] The computer system (1100) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1120) with media such as CD / DVD (1121), thumb drives (1122), removable hard disk drives or solid state drives (1123), traditional magnetic media such as tapes and floppy disks (not shown), devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown), etc.
[0166] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the currently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0167] The computer system (1100) may also include an interface (1154) to one or more communication networks (1155). The network can be, for example, a wireless network, a wired network, or an optical network. The network can further be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CANBus, and so on. Some networks typically require an external network interface adapter connected to certain common data ports or peripheral buses (1149) (e.g., the USB port of the computer system (1100); as described below, other network interfaces are typically integrated into the kernel of the computer system (1100) by connecting to the system bus (e.g., connecting the Ethernet interface in a PC computer system or connecting the cellular network interface in a smartphone computer system). The computer system (1100) can use any of these networks to communicate with other entities. Such communication can be only one-way receiving (e.g., broadcast television), only one-way sending (e.g., CANbus connected to certain CANbus devices), or two-way, for example, using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.
[0168] The above-mentioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel (1140) of the computer system (1100).
[0169] The kernel (1140) may include one or more central processing units (CPUs) (1141), a graphics processing unit (GPU) (1142), a dedicated programmable processing unit in the form of a Field Programmable Gate Area (FPGA) (1143), a hardware accelerator (1144) for certain tasks, a graphics adapter (1150), etc. These devices, as well as a read-only memory (ROM) (1145), a random access memory (1146), an internal mass storage such as an internal hard disk drive, SSD, etc. that is not user-accessible (1147), may be connected via a system bus (1148). In some computer systems, the system bus (1148) may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the system bus (1148) of the kernel or connected to the system bus (1148) of the kernel via a peripheral bus (1149). In one example, a screen (1110) may be connected to the graphics adapter (1150). The architecture of the peripheral bus includes PCI, USB, etc.
[0170] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) may execute certain instructions, which may be combined to form the aforementioned computer code. The computer code may also be stored in the ROM (1145) or the RAM (1146). Transitional data may also be stored in the RAM (1146), while permanent data may be stored, for example, in the internal mass storage (1147). Fast storage and retrieval of any storage device may be performed by using a cache, which may be closely associated with one or more CPUs (1141), GPUs (1142), mass storage (1147), ROM (1145), RAM (1146), etc.
[0171] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the medium and the computer code may be of the type well-known and available to those skilled in the field of computer software.
[0172] By way of example, and not limitation, a computer system having an architecture (1100), particularly a core (1140), can provide functionality due to software executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage described above, as well as certain non-transitory memories of the core (1140), such as the on-core mass storage (1147) or ROM (1145). Software implementing aspects of the present disclosure can be stored in such devices and executed by the core (1140). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core (1140), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in the RAM (1146) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system can provide functionality due to logic hardwired or otherwise embodied in a circuit (e.g., an accelerator (1144)), which can replace the software or operate in conjunction with the software to execute specific processes or specific portions of specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0173] As used in this disclosure, "at least one" or "one of" is intended to include any one or combination of the recited elements. For example, reference to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A through C is intended to include only A, only B, only C, or any combination thereof. Reference to one of A or B and one of A and B is intended to include A or B or (A and B). When applicable, e.g., when the elements are not mutually exclusive, use of "one of" does not exclude any combination of the recited elements.
[0174] Although the present disclosure has described examples of multiple aspects, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.
Claims
1. A method for processing visual media data, comprising: Processing the code stream of the visual media data according to the format rules, wherein: The code stream includes a skip flag of the current block; the skip flag of the current block indicates whether at least one residual-related syntax element for the current block is written when the current block in the current picture is encoded in an intra template matching prediction (IntraTMP) mode, The format rule specifies that when the skip flag indicates that the residual-related syntax element is not written, the current block is processed by directly copying the values of processed samples in a prediction block indicated by a block vector (BV) of the current block in the current picture.
2. The method according to claim 1, wherein: The format rule specifies that the current block is a chroma block of a chroma component in an intra slice, and the chroma component and luma component in the intra slice are separately divided.
3. A device for video decoding, comprising: The processing circuit is configured to: Receiving a codestream including a syntax element, the syntax element indicating whether at least one residual-related syntax element for the current block in the current picture is written into the codestream when a current block is encoded according to a prediction block in a current picture and an offset between the current block and the prediction block is indicated by a block vector (BV); as well as When the syntax element indicates that the residual-related syntax element is not written, the current block is reconstructed by directly copying the values of the reconstructed samples in the prediction block.
4. The device according to claim 3, wherein: The syntax element is a skip flag for the current block.
5. The device according to claim 4, wherein: The current block is encoded in one of an intra block copy (IBC) mode and an intra template matching prediction (IntraTMP) mode.
6. The device according to claim 5, wherein: The current block is encoded in the IntraTMP mode, and the current block is a luminance block.
7. The device according to claim 6, wherein: The processing circuit is further configured to: Determining a context of a skip flag of the current block based on a skip flag of a neighboring block of the current block, wherein the neighboring block is encoded in one of the IBC mode and the IntraTMP mode; and The value of the skip flag is determined using the determined context.
8. The device according to claim 5, wherein: The current block is a chroma block of a first chroma component in an intra slice, and the first chroma component and a luma component in the intra slice are separately divided.
9. The device according to claim 8, wherein: The processing circuit is further configured to: determining a context of a skip flag of the current block based on a skip flag of a co-located luma block of the current block, wherein the co-located luma block is encoded in one of the IBC mode and the IntraTMP mode; and A value of a skip flag for the current block is determined using the determined context.
10. The device according to claim 8, wherein: The current block is encoded in direct block vector mode (DBV), and The processing circuitry is configured to determine a BV of the current block based on a BV of a co-located luma block from the luma component.
11. A method for video encoding, comprising: determining whether to apply residual coding to a current block in a current picture encoded in one of an intra block copy (IBC) mode and an intra template matching prediction (IntraTMP) mode; as well as When it is determined that the residual coding is not applied to the current block, a syntax element is encoded in the codestream, the syntax element indicating that the residual-related syntax element is not written for the current block, and the current block is predicted by directly copying the values of the samples in the prediction block indicated by the block vector (BV) in the current picture.
12. The method according to claim 11, wherein: The syntax element is a skip flag for the current block.
13. The method according to claim 12, wherein: The current block is encoded in the IntraTMP mode, and the current block is a luminance block.
14. The method according to claim 13, further comprising: Determining a context of a skip flag of the current block based on a skip flag of a neighboring block of the current block, the neighboring block being encoded in one of the IBC mode and the IntraTMP mode; as well as A skip flag of the current block is encoded using the determined context.
15. The method according to claim 11 or 12, wherein: The current block is a chroma block of a first chroma component in an intra slice, and the first chroma component and a luma component in the intra slice are separately divided.