Video Decoding Methods

By determining intra-prediction modes and transform sets based on reconstructed samples and content type, the processing circuit optimizes transform selection for intra-prediction, addressing inefficiencies in existing video coding technologies and improving compression efficiency.

JP2025536301APending Publication Date: 2025-11-05TENCENT AMERICA LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025522104
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-18
Filing Date
2023-10-19
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently determining the appropriate transform set for intra-prediction modes, particularly in intra block copy (IBC) and intra template matching (IntraTMP) modes, leading to suboptimal compression efficiency.

Method used

A processing circuit determines an intra-prediction mode based on reconstructed samples and content type, calculates edge directions and template matching costs to select the most suitable transform set for the current block, and uses a lookup table to map intra-prediction modes to transform sets, including low-frequency non-separable transforms.

Benefits of technology

Improves video decoding efficiency by optimizing transform selection for intra-prediction, enhancing compression performance and reducing data volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536301000001_ABST
    Figure 2025536301000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure include methods and apparatuses for video coding. One of the apparatuses includes a processing circuit that receives an encoded video bitstream including a current picture having a current block coded in one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode. The processing circuit determines one of (i) an intra prediction mode based on reconstructed samples of the current picture and (ii) a content type of the reconstructed samples of the current picture. The processing circuit determines a transform set for the current block coded in the one of the IBC mode and the IntraTMP mode. The transform set is determined as being associated with the one of (i) the determined intra prediction mode and (ii) the determined content type. The processing circuit performs an inverse transform on the current block according to the determined transform set.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 417,937, entitled "Transform Selection Through Block Matching," filed October 20, 2022, which claims the benefit of priority to U.S. Patent Application No. 18 / 381,618, entitled "TRANSFORM SELECTION THROUGH BLOCK MATCHING," filed October 18, 2023. The disclosures of these prior applications are incorporated herein by reference in their entireties.

[0002] This disclosure describes aspects generally relating to video coding. [Background technology]

[0003] The background discussion provided herein is intended to provide a general overview of the context for the disclosure. To the extent described in this background section, the work of the named inventors, and aspects of the disclosure that may not otherwise qualify as prior art at the time of filing, are not admitted, explicitly or implicitly, as prior art to the present disclosure.

[0004] Image / video compression can help transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0005] Aspects of the present disclosure include methods and apparatuses for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit receives an encoded video bitstream including a current picture having a current block coded in one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode. The processing circuit determines one of (i) an intra prediction mode based on reconstructed samples of the current picture and (ii) a content type of the reconstructed samples of the current picture, and determines a transform set for the current block coded in the one of the IBC mode and the IntraTMP mode, the transform set being determined as associated with the one of (i) the determined intra prediction mode and (ii) the determined content type. The processing circuit performs an inverse transform on the current block according to the determined transform set.

[0006] In one example, the reconstructed samples include neighboring reconstructed samples of the current block.

[0007] In one example, the processing circuit calculates the frequency of edge directions of the current block using neighboring reconstructed samples of the current block, and determines an intra-prediction mode associated with the most frequently used edge direction among the edge directions. In one example, the processing circuit determines a transform set for the current block associated with the determined intra-prediction mode.

[0008] In one example, the processing circuit calculates template matching costs between a current template including neighboring reconstructed samples of the current block and each of the current templates indicated by the candidate intra-prediction modes, and selects the candidate intra-prediction mode associated with the smallest template matching cost as the determined intra-prediction mode. The processing circuit determines a transform set for the current block associated with the determined intra-prediction mode.

[0009] In one example, the reconstructed samples in the current picture include reconstructed samples of a reference block indicated by a block vector of the current block. The processing circuit calculates the frequency of directions of the reconstructed samples in the reference block and determines an intra-prediction mode associated with the most frequently used direction among the directions. In one example, the processing circuit determines a transform set for the current block associated with the determined intra-prediction mode.

[0010] In one example, the processing circuit determines a content type of one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block. The reference block is indicated by a block vector of the current block. The reconstructed samples in the current picture include the one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block. In one example, the processing circuit determines a transform set for the current block according to whether the determined content type is screen content.

[0011] In one example, the processing circuit determines the transform set to be a first transform set based on the determined content type being screen content, and determines the transform set to be a second transform set based on the determined content type being non-screen content, the second transform set being different from the first transform set.

[0012] In one example, the processing circuit determines a brightness number of one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block, and in response to the brightness number being less than a brightness threshold, the processing circuit determines the content type as screen content.

[0013] In one example, the luminosity of said one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block is associated with a particular color component.

[0014] In one example, the luminosity of the one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block is associated with multiple color components.

[0015] In one example, the steps of determining one of (i) an intra prediction mode and (ii) a content type, and determining a transform set, are performed only in response to the current block being within an intra slice of the current picture.

[0016] In one example, the processing circuitry determines the transform set for the current block according to the determined intra-prediction mode using a lookup table that maps intra-prediction modes to transform sets.

[0017] In one example, the transform set for the current block is a secondary transform set, and the mapping between the intra prediction mode indicated by the mode number (IntraPredMode) and the transform set including four low-frequency non-separable transform (LFNST) sets indicated by LFNST set indices 0 to 3 is shown in the following lookup table: [Table 1]

[0018] In one aspect, a processing circuit receives an encoded video bitstream including a current picture having a current block coded in one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode. The processing circuit determines an intra prediction mode associated with a reference block of the current block. The reference block is indicated by a block vector (BV) of the current block. The processing circuit determines a transform set for the current block coded in one of the IBC mode and the IntraTMP mode according to the intra prediction mode associated with the reference block, and performs an inverse transform on the current block according to the determined transform set.

[0019] In one example, the processing circuit obtains prediction mode information associated with at least one sub-block in the reference block by checking the at least one sub-block in a predetermined order, and determines the intra-prediction mode associated with the reference block according to the prediction mode information associated with the at least one sub-block.

[0020] In one example, the processing circuit obtains prediction mode information associated with at least one sub-block in a reference block located at one or more predetermined sub-block positions, and determines an intra-prediction mode associated with the reference block according to the prediction mode information associated with the at least one sub-block.

[0021] In one aspect, a processing circuit receives an encoded video bitstream including a current picture having a current block coded in an inter-prediction mode. The processing circuit determines an intra-prediction mode associated with a reference block of the current block indicated by a motion vector (MV) of the current block. At least a portion of the reference block is intra-coded. The processing circuit determines a transform set for the current block according to the intra-prediction mode associated with the reference block, and performs an inverse transform on the current block according to the determined transform set.

[0022] In one example, the processing circuit obtains prediction mode information associated with at least one sub-block in the reference block by checking the sub-blocks in a predetermined order, and determines an intra-prediction mode associated with the reference block according to the prediction mode information associated with the at least one sub-block.

[0023] In one example, samples of sub-blocks within a reference block are coded with different intra-prediction modes, and the processing circuitry applies filtering to the sub-blocks to determine the intra-prediction mode.

[0024] In one example, the processing circuit obtains prediction mode information associated with at least one sub-block in a reference block located at one or more predetermined sub-block positions, and determines an intra-prediction mode associated with the reference block according to the prediction mode information associated with the at least one sub-block.

[0025] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding / encoding. [Brief explanation of the drawings]

[0026] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100). [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4] 1 illustrates an example of an intra-block copy (IBC) mode according to an example of the present disclosure. [Figure 5] 1 shows the reference area for IBC mode in some examples. [Figure 6] 1 illustrates an example of an intra-template matching prediction (IntraTMP) mode according to one aspect of the present disclosure. [Figure 7] 1 shows an example of template-based intra-mode derivation (TIMD). [Figure 8] 1 shows an example of decoder-side intra-mode derivation (DIMD). [Figure 9A] 1 shows an example of a transform coding process using a 16x64 transform. [Figure 9B] 1 illustrates an example mapping from intra-prediction modes to transform sets according to one aspect of the present disclosure. [Figure 10] 1 shows a flowchart outlining a process according to some aspects of the present disclosure. [Figure 11] 1 shows a flowchart outlining a process according to some aspects of the present disclosure. [Figure 12] 1 shows a flowchart outlining a process according to some aspects of the present disclosure. [Figure 13] 1 shows a flowchart outlining a process according to some aspects of the present disclosure. [Figure 14] 1 shows a flowchart outlining a process according to some aspects of the present disclosure. [Figure 15] 1 shows a flowchart outlining a process according to some aspects of the present disclosure. [Figure 16] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0027] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is an example of an application of the disclosed subject matter, which is a video encoder and video decoder in a streaming environment. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, and storage of compressed video on digital media, including CDs, DVDs, memory sticks, and the like.

[0028] The video processing system 100 may include a capture subsystem 113, which may include a video source 101, such as a digital camera, that produces a stream of uncompressed video pictures 102. In one example, the stream of video pictures 102 includes samples captured by the digital camera. The stream of video pictures 102 is depicted as a thick line to emphasize its high data volume compared to the encoded video data 104 (or encoded video bitstream) and may be processed by an electronics device 120 that includes a video encoder 103 coupled to the video source 101. The video encoder 103 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data 104 (or encoded video bitstream) is depicted as a thin line to emphasize its low data volume compared to the stream of video pictures 102 and may be stored on a streaming server 105 for later use. One or more streaming client subsystems, such as the client subsystems 106 and 108 of FIG. 1, can access the streaming server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 can include a video decoder 110, for example, within an electronics device 130. The video decoder 110 can decode the incoming copy of the encoded video data 107 and produce an outgoing stream of video pictures 111, which can be rendered on a display 112 (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data 104, 107, and 109 (e.g., a video bitstream) can be encoded according to a particular video coding / compression standard.Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The subject matter disclosed herein may be used in the context of VVC.

[0029] It should be noted that electronics devices 120 and 130 may include other components (not shown). For example, electronics device 120 may include a video decoder (not shown), and electronics device 130 may also include a video encoder (not shown).

[0030] 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) may be included in an electronics device (230). The electronics device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used in place of the video decoder (110) in the example of FIG. 1.

[0031] The receiver (231) can receive one or more coded video sequences, e.g., included in a bitstream, to be decoded by the video decoder (210). In one aspect, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences can be received from a channel (201), which can be a hardware / software link to a storage device that stores the coded video data. The receiver (231) can also receive the coded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which can be forwarded to their respective using entities (not shown). The receiver (231) can separate the coded video sequences from other data. To combat network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser 520 (hereinafter, "parser (220)"). In certain applications, the buffer memory (215) is part of the video decoder (210). In others, it may be external to the video decoder 210 (not shown). In still others, there may be a buffer memory (not shown) external to the video decoder 210, e.g., to combat network jitter, and there may be another buffer memory 215 internal to the video decoder 210, e.g., to handle playback timing. When the receiver 231 is receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory 215 may not be required or may be small. For use over best-effort packet networks, such as the Internet, the buffer memory 215 may be required and may be relatively large and advantageously sized adaptively, and may be implemented, at least in part, in an operating system or similar element (not shown) external to the video decoder 210.

[0032] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and possibly information for controlling a rendering device, such as a render device (212) (e.g., a display screen) that is not an integral part of the electronics device (230) but can be coupled to the electronics device (230), as shown in FIG. 2. The control information for the rendering device(s) may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be according to a video coding technique or standard and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) may also extract information from the coded video sequence information, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0033] The parser (220) may perform an entropy decoding / parsing process on the video sequence received from the buffer memory (215) to produce symbols (221).

[0034] The reconstruction of the symbols (221) may involve several different units, depending on the type of coded video picture or portion thereof and other factors (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.). Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (220). The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.

[0035] Beyond the functional blocks already described, the video decoder (210) may be conceptually subdivided into a number of functional units, as described below. In practical implementations operating within commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the subject matter of this disclosure, the following conceptual division into functional units is appropriate:

[0036] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as symbol(s) (221) from the parser (220), along with control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. The scalar / inverse transform unit (251) can output blocks of sample values ​​that can be input to an aggregator (255).

[0037] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258), for example, buffers partially reconstructed and / or fully reconstructed current pictures. In some cases, the aggregator (255) adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0038] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to a block that may be inter-coded and motion-compensated. In such cases, the motion-compensated prediction unit (253) may access the reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) related to the block, these samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches prediction samples may be controlled by a motion vector and are available to the motion-compensated prediction unit (253) in the form of symbols (221), which may have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values ​​fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.

[0039] The output samples of the aggregator (255) may be subjected to various loop filtering techniques in a loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the coded video sequence (also referred to as a coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression may also respond to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, as well as to previously reconstructed and loop-filtered sample values.

[0040] The output of the loop filter unit (256) can be a sample stream that can be output to a render device (212), which can also be stored in a reference picture memory (257) for use in future inter-picture prediction.

[0041] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before beginning reconstruction of the next coded picture.

[0042] The video decoder (210) may perform decoding according to a given video compression technology or standard, such as ITU-T Recommendation H.265. The coded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense of adhering to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile may select specific tools from all tools available in the video compression technology or standard, such that only those tools are available for use under that profile. Compliance also requires that the complexity of the coded video sequence be within a range specified by the level of the video compression technology or standard. In some cases, the level constrains the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may optionally be further constrained through a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0043] In one aspect, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of one or more encoded video sequences. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0044] 3 shows an exemplary block diagram of a video encoder (303). The video encoder (303) is included in an electronics device (320). For example, the electronics device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.

[0045] The video encoder (303) may receive video samples from a video source (301) (not part of the electronics device (320) in the example of FIG. 6) that may capture video image(s) to be encoded by the encoder (303). In another example, the video source (301) is part of the electronics device (320).

[0046] The video source (301) may provide a source video sequence to be encoded by the video encoder (303) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service provision system, the video source (301) may be a storage device containing pre-prepared video. In a video conferencing system, the video source (301) may be a camera capturing local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may have one or more samples, depending on the sampling structure, color space, etc. used. The following discussion focuses on samples.

[0047] According to one aspect, the video encoder (303) may code and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under other required time constraints. Enforcing an appropriate coding rate is one function of the controller (350). In some aspects, the controller (350) controls and is operatively coupled to other functional units, such as those described below, which are not shown for clarity. Parameters set by the controller (350) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured with other suitable functions associated with the video encoder (303) that are optimized for a particular system design.

[0048] In some aspects, the video encoder (303) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop can include a source coder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and one or more reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data, in a manner similar to that used by a (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-accurate results independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronism (and the resulting drift when synchronism cannot be maintained, for example due to channel errors) is also used in some related art.

[0049] The operation of the "local" decoder (333) may be the same as that of a "remote" decoder, such as the video decoder (210), which has already been described in detail above in connection with Figure 2. However, briefly referring also to Figure 2, because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (345) and parser (220) may be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333).

[0050] In one aspect, decoder technology, excluding parsing / entropy decoding, present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the subject matter of the disclosure focuses on decoder operation. A description of the encoder technology can be omitted, as it is the reverse of the decoder technology that has been thoroughly described. In certain areas, more detailed descriptions are provided below.

[0051] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture relative to one or more previously coded pictures from a video sequence designated as “reference pictures.” Thus, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of one or more reference pictures that may be selected as prediction reference(s) for the input picture.

[0052] The local video decoder (333) may decode coded video data for pictures that may be designated as reference pictures based on symbols created by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence may typically be a replica of the source video sequence, with some error. The local video decoder (333) may replicate the decoding process that may be performed by a video decoder on the reference pictures, causing the reconstructed reference pictures to be stored in the reference picture memory (334). In this way, the video encoder (303) may locally store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures that will be obtained by a far-end video decoder.

[0053] The predictor (335) may perform a predictive search for the coding engine (332). That is, for a new picture to be coded, the predictor (336) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors or block shapes, that can serve as an appropriate prediction reference for the new picture. The predictor (335) may operate pixel block by pixel block to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).

[0054] The controller (350) may manage the coding process of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0055] The outputs of all the aforementioned functional units may be subjected to entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.

[0056] A transmitter (340) may buffer the coded video sequence(s) produced by the entropy coder (345) and prepare them for transmission over a communication channel (360), which may be a hardware or software link to a storage device that stores the coded video data. The transmitter (340) may merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0057] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign each coded picture a particular coded picture type, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned one of the following picture types:

[0058] Intra-pictures (I-pictures) can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow several different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) pictures.

[0059] Predictive pictures (P pictures) can be encoded and decoded using intra- or inter-prediction, using motion vectors and reference indices to predict the sample values ​​of each block.

[0060] Bidirectionally predicted pictures (B pictures) can be coded and decoded using intra- or inter-prediction, using two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multi-predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0061] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to their respective pictures. For example, blocks of an I-picture may be coded non-predictively, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be coded non-predictively or via spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be coded non-predictively or via spatial or temporal prediction with reference to one or two previously coded reference pictures.

[0062] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. In operation, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. The coded video data may therefore conform to a syntax defined by the video coding technique or standard being used.

[0063] In one aspect, the transmitter (340) may transmit additional data along with the coded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0064] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture being coded / decoded, called the current picture, is divided into multiple blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. A motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0065] In some aspects, bi-prediction techniques can be used in inter-picture prediction. Bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, both of which are prior to the current picture in decoding order (but may be prior and future, respectively, in display order) in a video. A block in the current picture can be coded with a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first and second reference blocks.

[0066] Furthermore, merge mode techniques can be used to improve coding efficiency in inter-picture prediction.

[0067] According to some aspects of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed on a block-by-block basis. For example, according to the HEVC standard, a picture in a sequence of video pictures is divided into multiple coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the prediction type of that CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, prediction operations during coding (encoding / decoding) are performed in units of prediction blocks. Taking a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, and the like.

[0068] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In one aspect, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another aspect, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.

[0069] An example of an intra block copy mode (also called IntraBC mode or IBC mode), as used in, for example, HEVC and VVC, is described below.

[0070] FIG. 4 illustrates an example of an IBC mode according to an example of the present disclosure. The reference block used to predict the current CU (401) may be indicated by the block vector (BV) associated with the current CU (401). Each square (400) may represent a CTU. Gray-shaded areas represent areas or regions that have already been coded, while white unshaded areas represent areas or regions to be coded. The current CTU (400(4)) being reconstructed includes the current CU (401), a coded area (402), and an area to be coded (403). In one example, the area (403) will be coded after coding the current CU (401).

[0071] In one example, such as in HEVC, the gray-shaded area excluding the two CTUs (400(1)-400(2)) to the upper right of the current CTU (400(4)) can be used as a reference area in IBC mode to enable wavefront parallel processing (WPP). A BV allowed in HEVC can refer to a block within the reference area (e.g., the gray-shaded area excluding the two CTUs (400(1)-400(2))). For example, a BV (405) allowed in HEVC refers to the reference block (411).

[0072] In one example, such as in VVC, only the current CTU (400(4)) and the left neighboring CTU (400(3)) to the left of the current CTU (400(4)) are allowed as reference areas in IBC mode. In one example, the reference area used in IBC mode in VVC is within the dotted area (415) and contains coded samples. For example, the BV (406) allowed in VVC refers to the reference block (412).

[0073] In IBC mode BV coding, referencing the reconstructed area can be done by 2D BVs similar to the MVs used in inter prediction. Prediction and coding of BVs can reuse MV prediction and coding in the inter prediction process. In some examples, the luma BVs are in integer resolution, rather than 1 / 4 (or 1 / 4 pel) precision MVs as used in normal inter-coding CTUs.

[0074] In one example, the decoded motion vector differential (MVD) or block vector differential (BVD) of the BV is left-shifted (e.g., by 2) before being added to the BV predictor to determine the final BV.

[0075] The effective reference area for the IBC mode in some examples, such as the HEVC SCC extension, is almost the entire already reconstructed area of ​​the current picture, with some exceptions for parallel processing purposes. Figure 5 shows the reference area for the IBC mode in some examples, such as the configurations in HEVC and VVC, where only the CTU to the left of the current CTU serves as the reference sample area at the start of the reconstruction process of the current CTU. A drawback of the IBC concept in some implementations, such as in HEVC, is the additional memory requirement in the decoded picture buffer (DPB), for which hardware implementations may use external memory. The additional access to external memory involves increased memory bandwidth, making the concept of using the DPB less attractive. Some implementations, such as VVC, may use fixed memory, which can realize the IBC mode by using on-chip memory, significantly reducing memory bandwidth requirements and hardware complexity. In one embodiment, a significant change addresses the signaling concept that deviates from integration within the inter prediction process, such as in the HEVC SCC extension.

[0076] 5 illustrates a reference sample memory (RSM) (510) update process at four intermediate points (501)-(504) during the reconstruction process according to one embodiment of the present disclosure. The light gray shaded area may represent the reference sample of the left adjacent CTU. The dark gray shaded area may represent the reference sample of the current CTU. The white unshaded area may represent the area to be coded (e.g., the next coding area).

[0077] Referring to Figure 5, at the first intermediate point (501), which represents the start of reconstruction of the current CTU, in one example, the RSM (510) contains only the reference samples of the left adjacent CTU. At the other three intermediate points (502)-(504), the reconstruction process replaces the samples of the left adjacent CTU with variants of the current CTU. An implicit partitioning of the RSM (510) can be applied, dividing the RSM (510) into four disjoint 64x64 areas (511)-(514). Area resetting can be performed when the coder processes the first coding unit within the corresponding area when mapping the RSM to a CTU, reducing hardware implementation effort.

[0078] In the example shown in Figure 5, a fixed memory (e.g., RSM (510)) may be allocated to store reference areas used in IBC mode. At different intermediate points (e.g., (501)-(504)) during the coding process (e.g., encoding process or reconstruction process), parts of the RSM may be updated. Figure 5 shows the reference areas for IBC mode in VVC and their configuration in VVC.

[0079] 5, the RSM (510) can include a portion of the current CTU and / or a portion of the left adjacent CTU. In the example shown in FIG. 5, the size of the RSM is equal to the size of the CTU. The RSM (510) can include portions (511)-(514).

[0080] At a first intermediate point (501) in the coding process, the RSM (510) contains the entire left adjacent CTU, which can serve as a reference area in IBC mode at the start of the coding process for the current CTU, and does not currently contain the CTU, and portions (511)-(514) contain reconstructed samples of the left adjacent CTU.

[0081] At a second intermediate point (502) in the coding process for the current CTU, a sub-area (531) in the upper left region of the current CTU has already been coded (e.g., encoded or reconstructed), a sub-area (532) in the upper left region of the current CTU is the current CU being coded, and a sub-area (533) in the upper left region of the current CTU will be coded later. The RSM (510) is updated to include a portion of the left adjacent CTU and a portion of the current CTU. For example, portions (512)-(514) in the RSM (510) store the reconstructed samples in the same left adjacent CTU as at the first intermediate point (501), while portion (511) is updated to store the sub-area (531) of the current CTU. The reference area at the second intermediate point (502) may include reconstructed samples of the left adjacent CTU stored in portions (512)-(514) and reconstructed samples of the subarea (531) of the current CTU stored in portion (511).

[0082] At the third intermediate point (503) in the coding process for the current CTU, the upper left region of the current CTU has already been reconstructed. The upper right region of the current CTU includes subareas (541)-(543). Subarea (541) (shaded dark gray) has already been coded (e.g., encoded or reconstructed), subarea (542) is the current CU being coded (e.g., being coded or reconstructed), and subarea (543) (not shaded white) will be coded later. Portions (513)-(514) in the RSM (510) store reconstructed samples in the same left-neighboring CTU as at the first intermediate point (501), while portions (511)-(512) are updated so that portion (511) stores reconstructed samples for the upper left region of the current CTU and portion (512) stores subarea (541) of the current CTU. The reference area at the third intermediate point (503) may include (i) reconstructed samples of the left adjacent CTU stored in portions (513)-(514), and (ii) reconstructed samples of the upper left region of the current CTU stored in portion (511) and the sub-area (541) of the current CTU stored in portion (512).

[0083] At a fourth intermediate point (504) in the coding process for the current CTU, the upper-left, upper-right, and lower-left regions of the current CTU have already been reconstructed. The lower-right region of the current CTU includes sub-areas (551)-(553). Sub-area (551) (shaded dark gray) has already been coded (e.g., encoded or reconstructed), sub-area (552) is the current CU being coded (e.g., being coded or reconstructed), and sub-area (553) (not shaded white) will be coded later. Portion (511) stores reconstructed samples of the upper-left region within the current CTU, as at the third intermediate point (503), while portions (512)-(514) are updated so that portions (512)-(513) store reconstructed samples of the upper-right and lower-left regions of the current CTU, respectively, and portion (514) stores sub-area (551) of the current CTU. The reference area at the fourth intermediate time point (504) can include the reconstructed samples of the current CTU stored in portions (511)-(514). The RSM (510) at the fourth intermediate time point (504) does not include the area in the left adjacent CTU.

[0084] BV coding in IBC mode can adopt the concept of a merge list used for inter prediction. The IBC list construction process can consider two spatial neighbor BVs and five history-based BVs (HBVPs). In one example, only the first HBVP is compared with a spatial candidate when added to a candidate list. While normal inter prediction uses two different candidate lists, one for merge mode and the other for normal mode, the candidate list in IBC mode is used for both cases (e.g., IBC merge mode and IBC normal mode). IBC mode can include multiple different modes, such as IBC merge mode and IBC normal mode. A merge mode (e.g., IBC merge mode) can use up to six candidates in the list, while a normal mode (e.g., IBC normal mode) uses only the first two candidates. Block vector differential (BVD) coding employs a motion vector differential (MVD) process and can result in a final BV of any size. The reconstructed BV may point to an area outside the reference sample area and in some cases may need to be corrected by removing the absolute offset in each direction using modulo arithmetic with the width and height of the RSM.

[0085] The syntax and semantics of the IBC mode in some examples, such as in VVC, are described below. The IBC architecture in VVC can form a dedicated coding mode, where the IBC mode is a third prediction mode in addition to the intra prediction mode and the inter prediction mode. In one example, the bitstream carries an IBC syntax element indicating the IBC mode for a coding unit when the block size is 64x64 or smaller. As a result, the maximum CU size that can utilize the IBC mode can be 64x64 to realize the continuous memory update mechanism of the RSM. In one example, the reference sample addressing mechanism represents a two-dimensional offset and is the same as in the HEVC SCC extension by reusing the vector coding process of inter prediction. In one example, when a chroma separation tree (CST) is active, the coder cannot derive chroma BVs from luma BVs, and the use of the IBC mode is only for luma coding blocks.

[0086] 6 illustrates an example of an intra-template matching prediction (IntraTMP) mode according to one aspect of the present disclosure. In one aspect, such as in enhanced compression model (ECM) software, IntraTMP is a special intra-prediction mode that can copy a best predicted block (e.g., a matching block (621)) from a reconstructed portion of a current frame (or current picture) and match the template (e.g., an L-shaped template) (620) of the best predicted block with the current template (610) of a current block (611) (e.g., a current PU or current CU). For a given search range, the encoder can search for a template most similar to the current template within the reconstructed portion of the current frame and use the corresponding block as the predicted block. The encoder can signal the use of IntraTMP mode, and the same prediction operation can be performed at the decoder side.

[0087] The prediction signal can be generated by matching a current template (610), such as an L-shaped causal neighbor of the current block (611), with a template of another block within a predetermined search area. The exemplary search area shown in Figure 6 may include multiple CTUs (or superblocks). Referring to Figure 6, the search area may include the current CTU R1 (e.g., a portion of the current CTU R1), the CTU R2 above and to the left, the CTU R3 above, and the CTU R4 to the left. The cost function may include any suitable cost function, such as the sum of absolute differences (SAD).

[0088] Within each region, the decoder can search for the template with the smallest cost (e.g., smallest SAD) relative to the current template and use the block associated with the template with the smallest cost as the predicted block.

[0089] The dimensions of the region indicated by (SearchRange_w, SearchRange_h) can be set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. SearchRange_w=a×BlkW Equation (1) SearchRange_h=a×BlkH Equation (2) is.

[0090] The parameter 'a' may be a constant that controls the trade-off between gain and complexity. In one example, 'a' is 5.

[0091] The intra template matching tool may be enabled for CUs with a specific size, for example, width and height sizes less than or equal to 64. The maximum CU size for IntraTMP mode may be configurable. IntraTMP mode may be signaled at the CU level, for example, through a dedicated flag when decoder-side intra mode derivation (DIMD) is not currently used for the CU.

[0092] Template-based intra mode derivation (TIMD) can use a reference sample of the current CU as a template to select an intra mode from a set of candidate intra prediction modes associated with the TIMD. The selected intra mode can be determined as the best intra mode based on, for example, a cost function. As shown in FIG. 7, neighboring reconstructed samples of the current CU (702) can be used as a template (704). Reconstructed samples in the template (704) can be compared with predicted samples of the template (704). The predicted samples can be generated using a reference sample (706) of the template (704). The reference sample (706) can be neighboring reconstructed samples around the template (704). A cost function can be used to calculate a cost (or distortion) between the predicted samples and the reconstructed samples in the template (704) based on each intra prediction mode in the set of candidate intra prediction modes. The intra-prediction mode with the lowest cost (or distortion) may be selected as the intra-prediction mode (eg, the best intra-prediction mode) for intra-predicting the current CU (702).

[0093] When decoder-side intra-mode derivation (DIMD) is applied, N intra-modes can be derived from reconstructed neighbor samples around the current block (801), and N predictors obtained using these N intra-modes can be combined with the planar mode predictor using corresponding weights. The weights can be derived from gradients, such as a histogram of gradients (HoG) calculation. FIG. 8 shows an example of DIMD. The HoG calculation can be performed by applying a filter (e.g., horizontal and vertical Sobel filters) to pixels in a template (802) around the current block (801). The template (802) can include reconstructed neighbor samples around the current block (801). In one example, the template has a width of 3. In one example, pixels in the middle line of the template (802) (marked in gray) can be involved in the HoG calculation. Referring to FIG. 8, a window (803) around a pixel (805) can be used to determine the gradient associated with the pixel (805). The window (803) may have a size of 3x3 with the pixel (805) at the center of the window (803). Horizontal and vertical gradients may be obtained, for example, using horizontal and vertical Sobel filters, respectively. A direction or orientation may be obtained from the horizontal and vertical gradients. An intra-prediction mode (IPM) associated with the direction may be determined. Subsequently, a histogram (also referred to as HoG) (810) of the IPMs may be obtained. The IPM corresponding to the N highest histogram bars may be selected for the current block (801).

[0094] A transform may be applied to the block, such as a linear transform, a secondary transform, etc. In one example, the transform includes a combination of a linear transform and a secondary transform. In one example, the transform includes a non-separable transform. In one example, the transform includes a separable transform.

[0095] A secondary transform may be performed, such as in VVC. In some examples, such as in VVC, a low-frequency non-separable transform (LFNST) may be applied between the forward primary transform and quantization at the encoder side, and between dequantization and inverse primary transform at the decoder side, as shown in Figure 9A. In the LFNST, a reduced secondary transform (RST) method may be used.

[0096] The application of a non-separable transform, such as may be used in an LFNST, may be described as follows, using a 4×4 input block (or input matrix) X (shown in Equation (3)) as an example: To apply a 4×4 non-separable transform (e.g., an LFNST), the 4×4 input block X may be represented by a vector X, as shown in Equations 3-4:

number

[0097] The non-separable transformation is

number

[0098] FIG. 9A shows an example of a transform coding process (900) using a 16×64 transform (or a 64×16 transform, depending on whether the transform is a forward or inverse quadratic transform). Referring to FIG. 9A, in the process (900), at the encoder side, a forward linear transform (910) may first be performed on a block (e.g., a residual block) to obtain a coefficient block (913). A forward quadratic transform (or forward LFNST) (912) may then be applied to the coefficient block (913). In the forward quadratic transform (912), the 64 coefficients of a 4×4 sub-block AD in the upper left corner of the coefficient block (913) may be represented by a 64-length vector, which may be multiplied by a 64×16 (i.e., 64 widths and 16 heights) transform matrix to produce a 16-length vector. The elements of the 16-length vector are backfilled (913) into the upper left 4×4 sub-block A of the coefficient block. The coefficients of the sub-block BD may be zero, and the coefficients obtained after the forward binary transform (912) are then quantized in a quantization step (914) and entropy coded to generate bits that are coded into a bitstream (916).

[0099] The coded bits may be received at the decoder side and entropy decoded, followed by a dequantization step (924) to generate a coefficient block (923). An inverse quadratic transform (or inverse LFNST) (922), such as an inverse RST 8x8, may be performed to obtain 64 coefficients, for example, from the 16 coefficients in the top-left 4x4 sub-block E. These 64 coefficients may then be backfilled into the 4x4 sub-block EH. The coefficients in the coefficient block (923) after the inverse quadratic transform (922) may then be processed with an inverse linear transform (920) to obtain a reconstructed residual block.

[0100] In one example, according to the block size of the block, a 4×4 non-separable transform (e.g., 4×4 LFNST) or an 8×8 non-separable transform (e.g., 8×8 LFNST) is applied. The block size of the block can include the width, height, or the like. For example, the 4×4 LFNST is applied to a block where the smaller of the horizontal and vertical is less than a threshold value such as 8 (e.g., min(width, height) < 8). For example, the 8×8 LFNST is applied to a block where the smaller of the horizontal and vertical is greater than a threshold value such as 4 (e.g., min(width, height) > 4).

[0101] The non-separable transform (e.g., LFNST) can be based on a direct matrix multiplication approach and thus can be implemented in a single pass without repetition. In order to reduce the dimensions of the non-separable transform matrix and minimize the memory space for calculating complexity and storing transform coefficients, in LFNST, a reduced non-separable transform method (or RST) can be used. Therefore, in the reduced non-separable transform, an N-dimensional vector (e.g., N is 64 in an 8×8 non-separable second-order transform (NSST)) can be mapped to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Therefore, instead of an N×N matrix, the RST matrix is an R×N matrix as described in Equation (5):

Equation

[0102] In Equation (5), the R rows of the R×N transform matrix are the R bases in the N-dimensional space. The inverse transform matrix can be the transpose of the transform matrix (e.g., T RxN ) used in the forward transform. In 8×8 LFNST, a reduction factor of 4 can be applied, and the 64×64 direct matrix used in the 8×8 non-separable transform can be reduced to a 16×64 direct matrix as shown in FIG. 9A. Alternatively, a reduction factor greater than 4 may be applied, and the 64×64 direct matrix used in the 8×8 non-separable transform can be reduced to a 16×48 direct matrix. Therefore, a 48×16 inverse RST matrix can be used on the decoder side to generate the core (primary) transform coefficients in the 8×8 upper left region.

[0103] If a 16x48 matrix is ​​applied instead of a 16x64 matrix with the same transform set configuration, the input to the 16x48 matrix will contain 48 input data from three 4x4 blocks A, B, and C in the upper left 8x8 block, excluding the lower right 4x4 block D. As the dimensions are reduced, the memory usage for storing the LFNST matrix can be reduced, for example, from 10KB to 8KB, with minimal performance degradation.

[0104] To reduce complexity, the LFNST may be constrained to be applicable when coefficients outside the first coefficient subgroup are insignificant. In one example, the LFNST may be constrained to be applicable only when all coefficients outside the first coefficient subgroup are insignificant. Referring to Figure 9A, the first coefficient subgroup corresponds to the top-left block E, and therefore, the coefficients outside block E are insignificant.

[0105] In one example, when LFNST is applied, the primary-only transform coefficients are insignificant (e.g., zero). In one example, when LFNST is applied, all primary-only transform coefficients are zero. Primary-only transform coefficients may refer to transform coefficients obtained from a primary transform without a secondary transform.

[0106] Therefore, LFNST index signaling is conditional on the last significant position, and therefore, an extra coefficient scan in LFNST can be avoided. In some examples, the extra coefficient scan is used to check for significant transform coefficients at specific positions. In one example, LFNST's worst-case handling, e.g., with respect to per-pixel multiplication, constrains non-separable transforms for 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In the above cases, when LFNST is applied, the last significant scan position may be less than 8. For other sizes, when LFNST is applied, the last significant scan position may be less than 16. For 4xN and Nx4 blocks, where N is greater than 8, the constraint can mean that LFNST is applied to the top-left 4x4 region within the block. In one example, the constraint means that LFNST is applied only once to the top-left 4x4 region within the block. In one example, when LFNST is applied, all first-order coefficients are insignificant (e.g., zero), reducing the number of operations for the linear transform. From the encoder's perspective, quantization of transform coefficients can be significantly simplified when the LFNST transform is tested. Rate-distortion optimized quantization can be performed on, for example, up to the first 16 coefficients in scan order, and the remaining coefficients can be set to zero.

[0107] The LFNST transforms (also referred to as transform kernels, transform cores, or transform matrices) may be selected as described below. In one embodiment, multiple transform sets may be used, and each of the multiple transform sets in the LFNST may include one or more non-separable transform matrices (or kernels). A transform set may be selected from multiple transform sets, and a non-separable transform matrix may be selected from one or more non-separable transform matrices in a transform set.

[0108] Table 1 shows an example mapping from intra-prediction modes to multiple transform sets according to one embodiment of the present disclosure. The mapping indicates the relationship between the intra-prediction modes and multiple transform sets. The relationship, such as that shown in Table 1, can be predetermined and stored in the encoder and decoder. [Table 2]

[0109] Referring to Table 1, the plurality of transform sets includes four transform sets, for example, transform sets 0 to 3, represented by transform set indices (e.g., Tr.set indices) of 0 to 3, respectively. The index (e.g., intra-prediction mode index or IntraPredMode) may indicate an intra-prediction mode, and the transform set index may be obtained based on the index and Table 1. Thus, the transform set may be determined based on the intra-prediction mode. In one example, if one of three cross-component linear model (CCLM) modes (e.g., INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (e.g., 81≦IntraPredMode≦83), transform set 0 is selected for the current block.

[0110] As described above, each transform set may include one or more non-separable transform matrices. One of the one or more non-separable transform matrices may be selected, for example, by an explicitly signaled LFNST index. The LFNST index may be signaled in the bitstream once for each intra-coded CU, for example, after signaling the transform coefficients. For each transform set, the selected non-separable secondary transform candidate may be specified by the explicitly signaled LFNST index.

[0111] In one embodiment, LFNST is constrained to be applicable only when all coefficients outside the first coefficient subgroup are non-significant, and the coding of the LFNST index may depend on the position of the last significant coefficient. The LFNST index may be context coded. In one example, the context coding of the LFNST index does not depend on the intra prediction mode, and only the first bin is context coded. LFNST can be applied to intra-coded CUs within intra slices or inter slices for both the luma and chroma components. If dual trees are enabled, LFNST indexes can be signaled separately for the luma and chroma components. In inter slices (e.g., dual trees are disabled), a single LFNST index can be signaled and used for both the luma and chroma components.

[0112] Considering that large CUs larger than 64x64 are implicitly split (TU tiling) due to existing maximum transform size constraints (e.g., 64x64), LFNST index search may increase data buffering by four times for a certain number of decoding pipeline stages. Therefore, in some examples, the maximum size allowed for LFNST is constrained to 64x64. In one example, LFNST is enabled only for DCT2. In one example, LFNST index signaling is placed before MTS index signaling.

[0113] In one example, the use of a scaling matrix for perceptual quantization does not reveal that a scaling matrix specified for a primary matrix may be useful for the LFNST coefficients. Therefore, in some examples, the use of a scaling matrix for the LFNST coefficients is not permitted. In one example, in single-tree partition mode, chroma LFNST is not applied.

[0114] In one embodiment, such as in ECM, the LFNST design in VVC is extended as follows: The number of LFNST sets (S) and candidates (C) is expanded to S=35 and C=3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: If predModeIntra<0, lfnstTrSetIdx is equal to 2 For predModeIntra in [0,34], lfnstTrSetIdx=predModeIntra For predModeIntra in [35,66], lfnstTrSetIdx=68-predModeIntra Three different kernels, LFNST4, LFNST8, and LFNST16, are defined to denote the LFNST kernel sets that apply to 4xN / Nx4 (N≧4), 8xN / Nx8 (N≧8), and MxN (M, N≧16), respectively.

[0115] 9B shows an example of a mapping from intra prediction modes to these sets according to one aspect of the present disclosure. As shown in FIG. 9B, a table such as Table 2 is used to show the mapping from intra prediction modes (denoted as Intra pred. mode in Table 2) to secondary transform sets, such as LFNST sets indicated by respective LFNST set indices (denoted as lfnstTrSetIdx in Table 2).

[0116] In related art, when applying transforms to IBC-coded blocks (e.g., IntraBC-coded blocks) or IntraTMP-coded blocks, the same primary and secondary transform sets as those applied to planar intra-prediction modes are applied. However, IntraBC-coded blocks or IntraTMP-coded blocks may exhibit specific orientations in texture, and therefore sharing the same transform set as that used for planar intra-prediction modes may be suboptimal. According to one aspect of the present disclosure, transform selection may be performed through block matching.

[0117] A current picture is being coded (e.g., encoded or reconstructed). A current block in the current picture may be coded in one of an IBC mode and an IntraTMP mode. The current block may be referred to as an IBC-coded block or an IntraTMP-coded block. According to one aspect of the present disclosure, a transform set for the current block may be determined (e.g., selected) based on reconstructed samples (also referred to as reconstructed samples) of the current picture. In one example, the reconstructed samples of the current picture include neighboring reconstructed samples of the current block. In one example, the reconstructed samples of the current picture include reconstructed samples of a reference block indicated by the BV of the current block. The reference block of the current block may also be referred to as a predicted block of the current block. The reconstructed samples of the reference block may also be referred to as reconstructed samples of a predictive block. For example, a transform set is selected for a current block (e.g., an IntraBC-coded block or an IntraTMP-coded block) using reconstructed samples of a predictive block (also referred to as a reference block) or neighboring reconstructed samples of the current block.

[0118] In one example, the transform set for a current block (e.g., an IntraBC-coded block or an IntraTMP-coded block) may refer to a primary transform set or a secondary transform set. The transform set for the current block may include a primary transform set or a secondary transform set.

[0119] In one example, one of (i) an intra-prediction mode based on a reconstructed sample of a current picture and (ii) a content type of the reconstructed sample of the current picture is determined. A transform set associated with the one of (i) the determined intra-prediction mode and (ii) the determined content type may be determined as a transform set for a current block. In one aspect, a transform set for a current block coded in one of an IBC mode and an IntraTMP mode is determined as being associated with the one of (i) the determined intra-prediction mode and (ii) the determined content type. A transform may be performed on the current block according to the determined transform set. For example, at the encoder side, an inverse transform is performed on the current block according to the determined transform set. In one example, at the decoder side, a forward transform is performed on the current block according to the determined transform set.

[0120] In one aspect, when neighboring reconstructed samples are used to determine a transform set, a DIMD method, such as the method used in the DIMD mode, is used to derive an intra-prediction mode for selecting the transform set, and the intra-prediction mode associated with the largest histogram amplitude value is used to determine the transform set. For example, in the DIMD mode, the frequency of edge directions of the current block is calculated using neighboring reconstructed samples of the current block as described in FIG. 8. The edge direction associated with one of the neighboring reconstructed samples of the current block (e.g., (805)) may be calculated using a gradient window (e.g., (803)). In one example, the frequency of edge directions of the current block (e.g., the number of times each edge direction occurs) may be plotted in a histogram similar to histogram (810). The frequency of the edge directions corresponds to each histogram amplitude value in the histogram. The intra-prediction mode associated with the most frequently used edge direction may be determined. The most frequently used edge direction corresponds to the largest histogram amplitude value. The transform set for the current block can be determined as the transform set associated with the determined intra prediction mode, for example, the transform set associated with the determined intra prediction mode (e.g., the intra prediction mode associated with the most frequently used edge direction) can be used as the transform set for the current block. In one example, one or more intra prediction modes are associated with a transform set, for example, as shown in Table 1 or Table 2. If a block is predicted using one of the one or more intra prediction modes associated with a transform set, a transform in the transform set can be used to transform the block.

[0121] In one example, the intra prediction mode determined using the above-described DIMD method (eg, the intra prediction mode associated with the most frequently used edge direction among the edge directions) is an angular mode or a directional intra prediction mode.

[0122] In one aspect, when neighboring reconstructed samples are used to determine a transform set, a TIMD method, such as the template matching method used in the TIMD mode shown in Figure 7, can be used to derive an intra-prediction mode for selecting a transform set, and the transform set can be determined using the intra-prediction mode associated with the smallest template matching cost. For example, as shown in Figure 7, template matching costs between a current template including neighboring reconstructed samples of the current block and each template of the current template indicated by the candidate intra-prediction modes are calculated. Among the candidate intra-prediction modes, the candidate intra-prediction mode associated with the smallest template matching cost can be selected as the determined intra-prediction mode. The transform set for the current block is determined as associated with the determined intra-prediction mode.

[0123] In one aspect, when reconstructed samples of a predictive block (e.g., a reference block) are used to determine a transform set, a directional histogram (also referred to as an edge directional histogram) is derived based on the reconstructed samples of the predictive block using a DIMD method, such as the method used in the DIMD mode described above and in FIG. 8. The direction (referred to as the edge direction) with the largest histogram amplitude value is derived, and the associated intra-prediction mode is used to determine the transform set. In one example, reconstructed samples of the entire reference block are used to determine the directional histogram. In one example, reconstructed samples at the boundary of the reference block are used to determine the directional histogram, and the directional histogram is referred to as an edge directional histogram. In one example, deriving the directional histogram includes calculating the frequency of the directions of the reconstructed samples within the reference block. The intra-prediction mode associated with the most frequently used direction among those directions is determined, and the transform set for the current block is determined as being associated with the determined intra-prediction mode. As described above, the direction with the largest histogram amplitude value is the most frequently used direction.

[0124] Intra-prediction modes may be associated with transform sets. In one aspect, one or more intra-prediction modes may be associated (e.g., mapped) to distinct transform sets, for example, as shown in Tables 1 and 2.

[0125] According to an aspect of the present disclosure, a transform set for the current block may be determined according to the determined intra-prediction mode using a lookup table (e.g., Table 1 or Table 2) that maps intra-prediction modes to transform sets.

[0126] In one aspect, the method of transform set selection based on intra prediction mode may be the same as that in the VVC standard, for example, as described in Table 1. For example, the transform set for the current block is the secondary transform set. The mapping between the intra prediction mode indicated by the mode number (IntraPredMode) and the transform set including four LFNST sets indicated by LFNST set indexes 0 to 3 is shown in a lookup table, for example, Table 1.

[0127] In one aspect, the method of transform set selection based on intra prediction modes for primary and secondary transforms may be the same as in ECM. Referring back to Figure 9B, intra prediction modes may be mapped to secondary transform sets (e.g., LFNST transform sets) based on the mapping relationship between intra prediction modes (indicated by intra prediction mode numbers between -14 and 80) and LFNST transform sets (indicated by LFNST set indices between 0 and 34) shown in Table 2.

[0128] In one aspect, content type detection is performed on neighboring reconstructed samples of a current block or reconstructed samples of a predictive block to determine whether the current block is a screen content or a non-screen content. Based on this content type determination, different transform sets may be applied. For example, the content type of the reconstructed samples of a reference block or neighboring reconstructed samples of the current block is determined. Depending on whether the determined content type is a screen content, a transform set for the current block may be determined. For example, two different transform sets may be applied to screen content and non-screen content, respectively. In one example, there is a one-to-one correspondence between transform sets and content types.

[0129] In one example, the transform set is determined to be a first transform set based on the determined content type being screen content. In one example, the transform set is determined to be a second transform set based on the determined content type being non-screen content. The second transform set is different from the first transform set.

[0130] In one example, the content type detection process involves checking how many different intensities there are in neighboring reconstructed samples of a current block or in reconstructed samples of a reference block. If there are less than a threshold of intensities (e.g., a given threshold of intensities), the content type of the neighboring reconstructed samples of the current block or in the reconstructed samples of the reference block is screen content. For example, the number of intensities of the reconstructed samples of the reference block or the neighboring reconstructed samples of the current block is determined. If the number of intensities is less than the threshold of intensities, the content type is determined as screen content.

[0131] The brightness of a reconstructed sample of a reference block or a neighboring reconstructed sample of a current block is associated with a specific color component. In one example, brightness refers to the value of one specific color component, such as the luma component. In one example, brightness is an integer ranging from 0 to 255.

[0132] The brightness of a reconstructed sample of a reference block or a neighboring reconstructed sample of a current block may be associated with multiple color components. In one example, brightness refers to a combination of values ​​of multiple color components, such as a combination of Y, Cb, and Cr, or a combination of R, G, and B. If the value of the same color component changes, a first brightness is different from a second brightness. In one example, brightness is a combination of Y, Cb, and Cr, and a first brightness of [100,10,5] for Y, Cb, and Cr is different from a second brightness of [100,5,5] for Y, Cb, and Cr.

[0133] In one aspect, the above-described methods may be applied only to blocks within an intra slice, such as when the current block is within an intra slice. For example, the methods for determining an intra-prediction mode or content type and the methods for determining a transform set are performed only when the current block is within an intra slice of the current picture.

[0134] According to one aspect of the present disclosure, for an IntraBC-coded block or an IntraTMP-coded block, such as a current block coded using an IBC mode or an IntraTMP mode, a transform set can be determined using an intra-prediction mode associated with a block (e.g., a reference block) identified by a block vector (BV) used in IntraBC / IntraTMP. In one example, at least a portion of the reference block is intra-coded. For example, an intra-prediction mode associated with the reference block of the current block is determined. A transform set for the current block can be determined according to the intra-prediction mode associated with the reference block.

[0135] In one example, to identify an intra-prediction mode, a block identified by a block vector (e.g., a reference block) may be examined sub-block by sub-block, for example, according to a predetermined order. The reference block may include the sub-blocks. A prediction mode of a first sub-block may be different from a prediction mode of a second sub-block. Prediction mode information associated with at least one sub-block in the reference block may be obtained by examining the at least one sub-block, for example, in a predetermined order. The prediction mode information associated with a sub-block may indicate a prediction mode of the sub-block. An intra-prediction mode associated with the reference block may be determined according to the prediction mode information associated with the at least one sub-block. In one example, the intra-prediction mode associated with the reference block is a first intra-prediction mode identified according to a predetermined order. In one example, the at least one sub-block includes multiple sub-blocks coded in respective prediction modes. The prediction mode may include multiple intra-prediction modes used to code the sub-block (e.g., the intra-coded sub-block). In one example, the intra-prediction mode associated with the reference block is determined as the one most frequently used to code the intra-coded sub-block among multiple intra-prediction modes.

[0136] In another aspect, the reference block identified by the block vector may be examined at one or more predetermined sub-block positions. For example, prediction mode information associated with at least one sub-block in the reference block at the one or more predetermined sub-block positions is obtained. An intra-prediction mode associated with the reference block may be determined according to the prediction mode information associated with the at least one sub-block at the one or more predetermined sub-block positions.

[0137] According to one aspect of the present disclosure, for multiple inter-coded blocks including an inter-coded block, if a motion vector (MV) points to a reference block in which at least partial samples in the reference block (or at least a portion of the reference block) are intra-coded, the transform set may be determined using an intra-prediction mode associated with the reference block. The reference block of an inter-coded block and the inter-coded block may be located in different pictures. In one aspect, the intra-prediction mode associated with the reference block of the inter-coded block is determined. At least a portion of the reference block may be intra-coded. A transform set for a current block may be determined according to the intra-prediction mode associated with the reference block. A transform may be performed on the current block according to the determined transform set.

[0138] In one aspect, to identify an intra-prediction mode, a reference block may be examined in units of sub-blocks according to a predetermined order. The reference block may include the sub-blocks. The prediction mode of a first sub-block may be different from the prediction mode of a second sub-block. Prediction mode information associated with at least one sub-block in the reference block may be obtained by examining the at least one sub-block, for example, in a predetermined order. The prediction mode information associated with a sub-block may indicate the prediction mode of the sub-block. The intra-prediction mode associated with the reference block may be determined according to the prediction mode information associated with the at least one sub-block. In one example, the intra-prediction mode associated with the reference block is a first intra-prediction mode identified according to a predetermined order. In one example, the at least one sub-block includes multiple sub-blocks coded in respective prediction modes. The prediction modes may include multiple intra-prediction modes used to code sub-blocks (e.g., intra-coded sub-blocks). In one example, the intra-prediction mode associated with the reference block is determined as the one most frequently used to code intra-coded sub-blocks among multiple intra-prediction modes.

[0139] In another aspect, filtering may be applied to units of sub-blocks within a reference block of an inter-coded block to derive an intra-prediction mode when samples within the reference block are coded in multiple different intra-prediction modes.

[0140] In another aspect, the reference block identified by the MV may be inspected at one or more predetermined sub-block positions. For example, prediction mode information associated with at least one sub-block in the reference block at one or more predetermined sub-block positions is obtained. An intra-prediction mode associated with the reference block may be determined according to the prediction mode information associated with the at least one sub-block at one or more predetermined sub-block positions.

[0141] 10 shows a flowchart outlining a process (1000) according to one aspect of the present disclosure. The process (1000) can be used in a video decoder. In various aspects, the process (1000) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), and the like. In some aspects, the process (1000) is implemented in software instructions, and thus the processing circuit performs the process (1000) when the processing circuit executes the software instructions. The process begins at (S1001) and proceeds to (S1010).

[0142] At (S1010), a coded video bitstream is received that includes a current picture having a current block coded in one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode.

[0143] At (S1020), one of (i) an intra prediction mode based on the reconstructed samples of the current picture and (ii) a content type of the reconstructed samples of the current picture may be determined.

[0144] At (S1030), a transform set for a current block coded in one of IBC mode and IntraTMP mode may be determined, where the transform set may be determined as associated with one of (i) the determined intra-prediction mode and (ii) the determined content type.

[0145] At (S1040), an inverse transform may be performed on the current block according to the determined transform set.

[0146] Then, the process proceeds to (S1099) and ends.

[0147] The process 1000 may be adapted as desired. One or more steps of the process 1000 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.

[0148] In one example, the reconstructed samples include neighboring reconstructed samples of the current block.

[0149] In one example, determining one of (i) the intra-prediction mode and (ii) the content type includes calculating a frequency of edge directions of the current block using neighboring reconstructed samples of the current block, for example, as described in Figure 8 using a method used in DIMD mode, and determining an intra-prediction mode associated with a most frequently used edge direction among the edge directions. Determining a transform set includes determining a transform set for the current block associated with the determined intra-prediction mode.

[0150] In one example, determining one of (i) the intra-prediction mode and (ii) the content type includes calculating a template matching cost between a current template including neighboring reconstructed samples of the current block and each of the current templates indicated by the candidate intra-prediction modes, and selecting, from the candidate intra-prediction modes, a candidate intra-prediction mode associated with a minimum template matching cost as the determined intra-prediction mode. Determining a transform set includes determining a transform set for the current block associated with the determined intra-prediction mode.

[0151] In one example, the reconstructed samples in the current picture include reconstructed samples of a reference block indicated by a block vector of the current block. Determining one of (i) an intra-prediction mode and (ii) a content type includes calculating a frequency of directions of the reconstructed samples in the reference block and determining an intra-prediction mode associated with a most frequently used direction among the directions. Determining a transform set includes determining a transform set for the current block associated with the determined intra-prediction mode.

[0152] In one example, determining one of (i) an intra-prediction mode and (ii) a content type includes determining a content type of one of reconstructed samples of a reference block and neighboring reconstructed samples of the current block. The reference block is indicated by a block vector of the current block. The reconstructed samples in the current picture include the one of the reconstructed samples of the reference block and neighboring reconstructed samples of the current block. Determining a transform set includes determining a transform set for the current block according to whether the determined content type is screen content.

[0153] In one example, the transform set is determined to be a first transform set based on the determined content type being screen content, and in one example, the transform set is determined to be a second transform set based on the determined content type being non-screen content, the second transform set being different from the first transform set.

[0154] In one example, a brightness number of the one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block is determined. If the brightness number is less than a brightness threshold, the content type is determined as screen content. In one example, the brightness of the one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block is associated with a specific color component. In one example, the brightness of the one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block is associated with multiple color components.

[0155] In one aspect, determining one of (i) an intra prediction mode and (ii) a content type, and determining a transform set, is performed only in response to the current block being within an intra slice of the current picture.

[0156] In one example, the transform set for the current block is determined according to the determined intra-prediction mode using a lookup table that maps intra-prediction modes to transform sets. The method of claim 1, wherein the transform set for the current block is a secondary transform set. A mapping between an intra-prediction mode indicated by a mode number (IntraPredMode) and a transform set including a low-frequency non-separable transform (LFNST) set indicated by an LFNST set index (e.g., 0 to 3 in Table 1, 0 to 34 in Table 2) is indicated in a look-up table, e.g., Table 1 or Table 2.

[0157] 11 shows a flowchart outlining a process (1100) according to one aspect of the present disclosure. The process (1100) can be used in a video encoder. In various aspects, the process (1100) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), and the like. In some aspects, the process (1100) is implemented in software instructions, and thus the processing circuit performs the process (1100) when the processing circuit executes the software instructions. The process begins at (S1101) and proceeds to (S1110).

[0158] At (S1110), (i) an intra-prediction mode may be determined based on the reconstructed samples of the current picture, or (ii) a content type of the reconstructed samples of the current picture may be determined.

[0159] At (S1120), a transform set may be determined for a current block coded in one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode. The transform set may be determined as associated with one of (i) the determined intra prediction mode and (ii) the determined content type, such as a transform set associated with one of (i) the determined intra prediction mode and (ii) the determined content type.

[0160] At (S1130), a forward transform may be performed on the current block according to the determined transform set.

[0161] Then, the process proceeds to (S1199) and ends.

[0162] The process 1100 may be adapted as desired. One or more steps of the process 1100 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.

[0163] 12 shows a flowchart outlining a process (1200) according to one aspect of the present disclosure. The process (1200) can be used in a video decoder. In various aspects, the process (1200) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), and the like. In some aspects, the process (1200) is implemented in software instructions, and thus the processing circuit performs the process (1200) when the processing circuit executes the software instructions. The process begins at (S1201) and proceeds to (S1210).

[0164] At (S1210), a coded video bitstream is received that includes a current picture having a current block coded in one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode.

[0165] At (S1220), an intra-prediction mode associated with a reference block of the current block may be determined. The reference block may be indicated by a block vector (BV) of the current block.

[0166] In one example, prediction mode information associated with at least one sub-block in a reference block is obtained by checking the at least one sub-block in a predetermined order, and an intra-prediction mode associated with the reference block is determined according to the prediction mode information associated with the at least one sub-block.

[0167] In one example, prediction mode information associated with at least one sub-block in a reference block located at one or more predetermined sub-block positions is obtained, and an intra-prediction mode associated with the reference block is determined according to the prediction mode information associated with the at least one sub-block.

[0168] At (S1230), a transform set for a current block coded in one of IBC mode and IntraTMP mode may be determined according to the intra prediction mode associated with the reference block.

[0169] At (S1240), an inverse transform may be performed on the current block according to the determined transform set.

[0170] Then, the process proceeds to (S1299) and ends.

[0171] Process 1200 may be adapted as desired. Step(s) of process 1200 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.

[0172] 13 shows a flowchart outlining a process (1300) according to one aspect of the present disclosure. The process (1300) can be used in a video encoder. In various aspects, the process (1300) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), and the like. In some aspects, the process (1300) is implemented in software instructions, and thus the processing circuit performs the process (1300) when the processing circuit executes the software instructions. The process begins at (S1301) and proceeds to (S1310).

[0173] At (S1310), an intra prediction mode associated with a reference block of the current block that is coded in one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode may be determined.

[0174] At (S1320), a transform set for a current block coded in one of IBC mode and IntraTMP mode may be determined according to the intra prediction mode associated with the reference block.

[0175] At (S1330), a forward transform may be performed on the current block according to the determined transform set.

[0176] Then, the process proceeds to (S1399) and ends.

[0177] Process 1300 may be adapted as desired. Step(s) of process 1300 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.

[0178] 14 shows a flowchart outlining a process (1400) according to one aspect of the present disclosure. The process (1400) can be used in a video decoder. In various aspects, the process (1400) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), and the like. In some aspects, the process (1400) is implemented in software instructions, and thus the processing circuit performs the process (1400) when the processing circuit executes the software instructions. The process begins at (S1401) and proceeds to (S1410).

[0179] At (S1410), a coded video bitstream including a current picture having a current block coded in an inter-prediction mode is received.

[0180] At (S1420), an intra-prediction mode associated with a reference block of the current block indicated by the motion vector (MV) of the current block may be determined, at least a portion of the reference block being intra-coded.

[0181] In one example, prediction mode information associated with at least one sub-block in the reference block is obtained by checking the sub-blocks in a predetermined order, and an intra-prediction mode associated with the reference block is determined according to the prediction mode information associated with the at least one sub-block.

[0182] In one example, samples of sub-blocks within a reference block are coded in different intra-prediction modes, and filtering may be applied to the sub-blocks to determine the intra-prediction mode.

[0183] In one example, prediction mode information associated with at least one sub-block in a reference block located at one or more predetermined sub-block positions is obtained, and an intra-prediction mode associated with the reference block is determined according to the prediction mode information associated with the at least one sub-block.

[0184] At (S1430), a transform set for the current block may be determined according to the intra-prediction mode associated with the reference block.

[0185] At (S1440), an inverse transform may be performed on the current block according to the determined transform set.

[0186] Then, the process proceeds to (S1499) and ends.

[0187] Process 1400 may be adapted as desired. Step(s) of process 1400 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.

[0188] 15 shows a flowchart outlining a process (1500) according to one aspect of the present disclosure. The process (1500) can be used in a video encoder. In various aspects, the process (1500) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), and the like. In some aspects, the process (1500) is implemented in software instructions, and thus the processing circuit performs the process (1500) when the processing circuit executes the software instructions. The process begins at (S1501) and proceeds to (S1510).

[0189] At (S1510), an intra-prediction mode associated with a reference block of the current block indicated by a motion vector (MV) of the current block may be determined, where at least a portion of the reference block is intra-coded. The current block may be coded in an inter-prediction mode.

[0190] At (S1520), a transform set for the current block may be determined according to the intra-prediction mode associated with the reference block.

[0191] At (S1530), a forward transform may be performed on the current block according to the determined transform set.

[0192] Then, the process proceeds to (S1599) and ends.

[0193] Process 1500 may be adapted as desired. Step(s) of process 1500 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.

[0194] The aspects and / or examples in this disclosure may be used separately or in combination in any order. Each of these methods (or aspects), encoders, and decoders may be implemented by processing circuitry (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0195] The techniques described above can be implemented as computer software with computer-readable instructions physically stored on one or more computer-readable media. For example, Figure 16 illustrates a computer system (1600) suitable for implementing certain aspects of the disclosed subject matter.

[0196] Computer software may be coded using any suitable machine code or computer language that can be assembled, compiled, linked, or similarly subjected to mechanisms to produce code having instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), and the like, either directly or via interpretation, microcode execution, and the like.

[0197] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0198] 16 with respect to computer system (1600) are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing aspects of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement related to any one or combination of components shown in this exemplary embodiment of computer system (1600).

[0199] The computer system (1600) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, moving a data glove, etc.), audio input (e.g., voice, clapping, etc.), visual input (e.g., gestures, etc.), or olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds, etc.), images (e.g., scanned images, photographic images obtained from a still camera, etc.), or video (e.g., two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0200] The input human interface devices may include one or more of a keyboard (1601), a mouse (1602), a trackpad (1603), a touchscreen (1610), a data glove (not shown), a joystick (1605), a microphone (1606), a scanner (1607), and a camera (1608) (only one of each is shown).

[0201] The computer system (1600) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1610), data gloves (not shown), or joystick (1605), although some haptic feedback devices may not function as input devices), audio output devices (e.g., speakers (1609), headphones (not shown), etc.), visual output devices (e.g., screens (1610) including CRT screens, LCD screens, plasma screens, and OLED screens (each with or without touchscreen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or four or more dimensional output through means such as stereoscopic output), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), etc.), and printers (not shown).

[0202] The computer system (1600) may also include human-accessible storage devices and their associated media, such as optical media including, for example, a CD / DVD ROM / RW (1620) with CD / DVD or similar media (1621), a thumb drive (1622), a removable hard drive or solid state drive (1623), legacy magnetic media such as tape and floppy disks (registered trademark, not shown), specialized ROM / ASIC / PLD-based devices (not shown) such as security dongles, and the like.

[0203] Those skilled in the art will also appreciate that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other transitory signals.

[0204] The computer system 1600 may also include an interface 1654 to one or more communications networks 1655. Networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, latency-tolerant, and the like. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, and the like, TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and vehicular and industrial networks including CANbus. Certain networks typically require an external network interface adapter that attaches to a particular general-purpose data port or peripheral bus 1649 (e.g., a USB port on the computer system 1600), while others are typically integrated into the core of the computer system 1600 by attachment to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, computer system 1600 can communicate with other entities. Such communication may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a specific CANbus device), or two-way, for example, to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks may be used on each network and network interface, as described above.

[0205] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1640) of the computer system (1600).

[0206] The core (1640) may include one or more central processing units (CPUs) (1641), graphics processing units (GPUs) (1642), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (1643), task-specific hardware accelerators (1644), graphics adapters (1650), etc. These devices may be connected via a system bus (1648), along with read-only memory (ROM) (1645), random access memory (1646), and internal mass storage (1647), such as internal non-user-accessible hard drives, SSDs, and the like. In some computer systems, the system bus (1648) may be made accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, and the like. Peripheral devices may be attached either directly to the core's system bus (1648) or via a peripheral bus (1649). In one example, a screen 1610 can be connected to a graphics adapter 1650. Peripheral bus architectures include PCI, USB, and the like.

[0207] The CPU (1641), GPU (1642), FPGA (1643), and accelerator (1644) may execute specific instructions that, in combination, may constitute the aforementioned computer code. The computer code may be stored in ROM (1645) or RAM (1646). Transient data may also be stored in RAM (1646), while permanent data may be stored, for example, in internal mass storage (1647). Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be associated with one or more of the CPU (1641), GPU (1642), mass storage (1647), ROM (1645), RAM (1646), and the like.

[0208] The computer-readable media may have computer code thereon for performing various computer-implemented processes. The media and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0209] By way of example, and not limitation, a computer system having the architecture (1600), and in particular the core (1640), can provide functionality as a result of the execution by one or more processors (including CPUs, GPUs, FPGAs, accelerators, and the like) of software embodied in one or more tangible computer-readable media. Such computer-readable media can be specific storage of the core (1640) that is non-transitory in nature, such as the core's internal mass storage (1647) or ROM (1645), and media associated with user-accessible mass storage as introduced above. Software implementing various aspects of the present disclosure can be stored in such devices and executed by the core (1640). The computer-readable media can include one or more memory devices or chips, depending on specific needs. Software may cause the core (1640) and particularly the processors therein (including CPUs, GPUs, FPGAs, and the like) to perform particular processes or portions of particular processes described herein, including by defining data structures stored in RAM (1646) and modifying such data structures according to processes defined by the software. Additionally, or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerators (1644)) that can operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both, where appropriate. The present disclosure includes any suitable combination of hardware and software.

[0210] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the described elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B and one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of the described elements where applicable, such as when the elements are not mutually exclusive.

[0211] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various equivalent alternatives that fall within the scope of the disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the disclosure and are therefore within its spirit and scope.

Claims

1. 1. A method of video decoding executed by one or more processors, comprising: receiving a coded video bitstream including a current picture including a current block coded in one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode; determining one of (i) an intra-prediction mode based on reconstructed samples of the current picture, and (ii) a content type of the reconstructed samples of the current picture; determining a transform set for the current block coded in the one of the IBC mode and the IntraTMP mode, the transform set being determined as associated with the one of (i) the determined intra-prediction mode and (ii) the determined content type; performing an inverse transform on the current block according to the determined set of transforms; A method having the following.

2. The method of claim 1 , wherein the reconstructed samples include neighboring reconstructed samples of the current block.

3. The step of determining one of (i) the intra-prediction mode and (ii) the content type comprises: calculating a frequency of edge directions of the current block using the neighboring reconstructed samples of the current block; determining an intra-prediction mode associated with a most frequently used edge direction among the edge directions; Including, determining the transform set includes determining a transform set for the current block associated with the determined intra-prediction mode. The method of claim 2.

4. The determining of one of (i) the intra-prediction mode and (ii) the content type is based on template-based intra-mode derivation (TIMD), i.e.: calculating a template matching cost between a current template including the neighboring reconstructed samples of the current block and each template of the current template indicated by a candidate intra-prediction mode; selecting, from the candidate intra-prediction modes, a candidate intra-prediction mode associated with a minimum template matching cost among the template matching costs as the determined intra-prediction mode; determining the one of (i) the intra-prediction mode and (ii) the content type based on determining the transform set includes determining a transform set for the current block associated with the determined intra-prediction mode. The method of claim 2.

5. the reconstructed samples in the current picture include reconstructed samples of a reference block indicated by a block vector of the current block; The determining of one of (i) the intra prediction mode and (ii) the content type is based on decoder-side intra mode derivation (DIMD), i.e.: calculating the frequency of orientations of the reconstructed samples within the reference block; determining an intra-prediction mode associated with a most frequently used one of the directions; determining the one of (i) the intra-prediction mode and (ii) the content type based on determining the transform set includes determining a transform set for the current block associated with the determined intra-prediction mode. The method of claim 1.

6. the determining of one of (i) the intra prediction mode and (ii) the content type includes determining a content type of one of reconstructed samples of a reference block and neighboring reconstructed samples of the current block, the reference block being indicated by a block vector of the current block, and the reconstructed samples in the current picture including the one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block; the step of determining the transformation set includes determining the transformation set for the current block according to whether the determined content type is a screen content; The method of claim 1.

7. The determining of the transformation set comprises: determining the transform set to be a first transform set based on the determined content type being the screen content; determining, based on the determined content type being non-screen content, that the transform set is a second transform set different from the first transform set; 7. The method of claim 6, comprising:

8. determining the content type determining a number of intensities of the one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block; determining the content type as the screen content in response to the number of luminosity values ​​being less than a luminosity threshold; 7. The method of claim 6, comprising:

9. The method of claim 8 , wherein the lightness of the one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block is associated with a particular color component.

10. The method of claim 8 , wherein the lightness of the one of the reconstructed samples of the reference block and the neighboring reconstructed samples of the current block is associated with multiple color components.

11. 2. The method of claim 1, wherein the steps of determining one of (i) the intra prediction mode and (ii) the content type and determining the transform set are performed only in response to the current block being within an intra slice of the current picture.

12. The step of determining the transformation set comprises: determining the transform set for the current block according to the determined intra-prediction mode using a lookup table that maps intra-prediction modes to transform sets; 2. The method of claim 1, comprising:

13. the transform set for the current block is a secondary transform set; The mapping between the intra prediction modes, indicated by a mode number (IntraPredMode), and the transform set, which includes four low frequency non-separable transform (LFNST) sets, indicated by LFNST set indices 0 to 3, is shown in the following lookup table: Table 1 The method of claim 12.

14. 1. A method of video decoding executed by one or more processors, comprising: receiving a coded video bitstream including a current picture including a current block coded in one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode; determining an intra prediction mode associated with a reference block of the current block, the reference block being indicated by a block vector (BV) of the current block used in the one of the IBC mode and the IntraTMP mode; determining a transform set for the current block coded in the one of the IBC mode and the IntraTMP mode according to the intra prediction mode associated with the reference block; performing an inverse transform on the current block according to the determined set of transforms; A method having the following.

15. The step of determining the intra prediction mode includes: Obtaining prediction mode information associated with at least one sub-block in the reference block by checking the at least one sub-block in a predetermined order; determining the intra prediction mode associated with the reference block according to the prediction mode information associated with the at least one sub-block; 15. The method of claim 14, comprising:

16. The step of determining the intra prediction mode includes: obtaining prediction mode information associated with at least one sub-block within the reference block located at one or more predetermined sub-block positions; determining the intra prediction mode associated with the reference block according to the prediction mode information associated with the at least one sub-block; 15. The method of claim 14, comprising:

17. 1. A method of video decoding executed by one or more processors, comprising: receiving a coded video bitstream including a current picture including a current block coded in an inter prediction mode; determining an intra-prediction mode associated with a reference block of the current block indicated by a motion vector (MV) of the current block, wherein at least a portion of the reference block is coded for intra-prediction; determining a transform set for the current block according to the intra-prediction mode associated with the reference block; performing an inverse transform on the current block according to the determined set of transforms; A method having the following.

18. The step of determining the intra prediction mode includes: Obtaining prediction mode information associated with at least one sub-block in the reference block by checking the sub-block in a predetermined order; determining the intra prediction mode associated with the reference block according to the prediction mode information associated with the at least one sub-block; 18. The method of claim 17, comprising:

19. Samples of sub-blocks in the reference block are coded in different intra prediction modes; the determining the intra-prediction mode includes applying filtering to the sub-block to determine the intra-prediction mode.

18. The method of claim 17.

20. The step of determining the intra prediction mode includes: obtaining prediction mode information associated with at least one sub-block within the reference block located at one or more predetermined sub-block positions; determining the intra prediction mode associated with the reference block according to the prediction mode information associated with the at least one sub-block; 18. The method of claim 17, comprising:

Citation Information

Patent Citations

  • Transform selection in a video encoder and / or video decoder

    US20200186815A1

  • Transform-based image coding method and device for same

    US20220086449A1

  • Methods and apparatuses for cross-component prediction

    US20220239897A1

  • Intra-mode dependent multiple transform selection for video coding

    US20220329800A1

  • A method, an apparatus and a computer program product for video encoding and video decoding

    WO2021244935A1