Signaling indication and filtering of templates and flexible partitioning split

CN122804398APending Publication Date: 2026-09-22TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580017492.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-05-22
Filing Date
2025-05-23
Publication Date
2026-09-22

Smart Images

  • Figure CN122804398A_ABST
    Figure CN122804398A_ABST
Patent Text Reader

Abstract

Methods and apparatus, including computer programs encoded on a computer-readable medium, for video decoding and video encoding, and for processing visual media data are provided. One of the methods for video decoding includes receiving coded information for a current block in a current picture. It is determined whether to apply a filter to a neighboring reconstructed area of the current block. The neighboring reconstructed area is adjacent to the current block and includes reconstructed samples in the current picture. The filter is applied to the neighboring reconstructed area of the current block. The current block is reconstructed from the filtered neighboring reconstructed area.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications This application claims priority to U.S. Patent Application No. 19 / 216,658, filed May 22, 2025, entitled “Signaling and Filtering of Template and Flexible Partition Split,” which claims priority to U.S. Provisional Application No. 63 / 651,932, filed May 24, 2024, entitled “Flexible Partition Split,” U.S. Provisional Application No. 63 / 702,436, filed October 2, 2024, entitled “Template Filtering for Template-Based Interpretation,” and U.S. Provisional Application No. 63 / 702,436, filed October 2, 2024, entitled “Method and Apparatus for Signaling Based on Reference Sample Features.” Priority is claimed in U.S. Provisional Application No. 63 / 702,463, entitled “Reference Pixel or Template Enhancement for Intra Prediction Mode,” filed October 4, 2024, and in whole or in part in U.S. Provisional Application No. 63 / 703,796, entitled “Enhancement of Reference Pixel or Template for Intra Prediction Mode.” The entire disclosure of these prior applications is incorporated herein by reference. Technical Field

[0002] This invention describes various aspects generally related to video encoding and decoding. Background Technology

[0003] The background description provided in this application is intended to generally present the context of the invention. To the extent that the work described in this background section is intended, the work of the currently named inventors and aspects of the description that may not have otherwise qualified as prior art at the time of filing are neither explicitly nor implicitly considered as prior art to the invention.

[0004] Image / video compression facilitates the transfer of image / video data across different devices, storage devices, and networks with minimal quality loss. In some examples, video codecs can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which compresses images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image being reconstructed to predict samples. In another example, a video codec can use a technique called inter-frame prediction, which compresses images based on temporal redundancy. For example, inter-frame prediction can utilize motion compensation to predict samples in the current image based on previously reconstructed images. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0005] Various aspects of the present invention include methods and apparatus for video encoding / decoding.

[0006] Various aspects of the present invention provide a method for video decoding, wherein encoded information of a current block in a current image is received. It is determined whether to apply a filter to adjacent reconstructed regions of the current block. The adjacent reconstructed regions are adjacent to the current block and include reconstructed samples from the current image. The filter is applied to the adjacent reconstructed regions of the current block. The current block is reconstructed based on the filtered adjacent reconstructed regions.

[0007] The present invention also provides a method for video coding, wherein, based on one of the partitioning information and prediction information of adjacent reconstructed regions of a current block in a current image, it is determined whether to indicate one of the following in the video bitstream by signaling: (i) a control flag for predicting a prediction mode of the current block, and (ii) the type of prediction mode. The current block is predicted based on reconstructed samples in adjacent reconstructed regions adjacent to the current block. The current block is encoded according to the prediction mode and the reconstructed samples in the adjacent reconstructed regions. When it is determined that one of the control flag and the type of prediction mode should be indicated by signaling, one of the control flag and the type of prediction mode is encoded into the video bitstream.

[0008] The present invention also provides a method for video decoding, wherein encoded information of a current region to be segmented in a current image is received. A segmentation structure of the current region is determined based on one of (i) the size information of the current region and (ii) segmentation information of another region. The current region is reconstructed based on the determined segmentation structure of the current region.

[0009] The present invention also provides an apparatus for video decoding. The apparatus for video decoding includes processing circuitry configured to implement any of the described methods for video decoding.

[0010] Aspects of the present invention also provide an apparatus for video encoding. The apparatus for video encoding includes processing circuitry configured to implement any of the described methods for video encoding.

[0011] Aspects of the present invention also provide a non-transitory computer-readable storage medium storing instructions that, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding. Attached Figure Description

[0012] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which: Figure 1 An example block diagram of a communication system (100) is shown schematically.

[0013] Figure 2 An example block diagram of a decoder is shown schematically.

[0014] Figure 3 An example block diagram of an encoder is shown schematically.

[0015] Figure 4 An example of a template in inter-frame prediction according to one aspect of the invention is shown.

[0016] Figure 5 An example of a current template and a reference template according to one aspect of the invention is shown, the current template including neighboring samples of the current block, and the reference template including neighboring samples of the corresponding reference block for deriving local illumination compensation (LIC) parameters.

[0017] Figure 6 An example of a deblocking filter according to one aspect of the invention is shown.

[0018] Figure 7 An example of a loop filter according to one aspect of the invention is shown.

[0019] Figure 8 An example is shown of a current block in a current image according to one aspect of the invention, and the adjacent reconstructed regions of that current block.

[0020] Figure 9 An example of adjacent reconstructed regions of the current block in the current image is shown according to one aspect of the invention.

[0021] Figure 10 An example is shown of a block according to one aspect of the invention and an adjacent reconstructed region formed by three adjacent blocks of that block.

[0022] Figure 11-12An example of a block and a current template according to one aspect of the invention is shown.

[0023] Figure 13 An example of a block partitioning structure according to one aspect of the present invention is shown.

[0024] Figure 14A An example of vertical center-side ternary tree partitioning according to one aspect of the invention is shown.

[0025] Figure 14B An example of horizontal center-side ternary tree partitioning according to one aspect of the invention is shown.

[0026] Figure 15 Examples of five partitioning methods according to one aspect of the invention are shown.

[0027] Figure 16 An example of a flexible partitioning based on one aspect of the invention is shown, whereby the flexible partitioning is predicted based on the area above or to the left of the current area, wherein the partitioning depth is reduced by 1.

[0028] Figures 17A-17B An example of a flexible partitioning scheme predicted using the directionality of partitioning information in adjacent regions of the current region, according to one aspect of the invention, is shown.

[0029] Figure 18A-18G An example of a flexible partitioning method according to one aspect of the invention is shown, which predicts partitioning information by extrapolating partitioning information in neighboring regions of the current region.

[0030] Figures 19A-19C An example is shown of using prior information to modify a derived flexible partitioning according to one aspect of the invention.

[0031] Figures 20A-20B An example is shown of the scan order of the current block (2001) according to one aspect of the invention.

[0032] Figure 21 A flowchart illustrating an overview process (2100) according to one aspect of the present invention is shown.

[0033] Figure 22 A flowchart illustrating an overview process (2200) according to one aspect of the present invention is shown.

[0034] Figure 23 A flowchart illustrating an overview process (2300) according to one aspect of the present invention is shown.

[0035] Figure 24 A computer system is illustrated schematically according to one aspect. Detailed Implementation

[0036] Figure 1 Block diagrams of some example video processing systems (100) are shown. The video processing system (100) is an example application of the disclosed subject matter, namely, video encoders and video decoders in a streaming media environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming media services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0037] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101), such as a digital camera, for creating, for example, an uncompressed video image stream (102). In one example, the video image stream (102) includes samples captured by a digital camera. The video image stream (102) is depicted as a thick line compared to encoded video data (104) (or encoded video bitstream) to emphasize the high data volume of the video image stream, which may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to enable or implement aspects of the disclosed subject matter, which are described in more detail below. Compared to the video image stream (102), the encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize the low data volume of the encoded video data, which can be stored on a streaming media server (105) for future use. At least one streaming media client subsystem, such as... Figure 1 Client subsystems (106) and (108) can access a streaming media server (105) to retrieve copies (107) and (109) of encoded video data (104). Client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an outgoing video image stream (111) that can be presented on a display (112) (e.g., a screen) or other display device (not depicted). In some streaming media systems, the encoded video data (104), (107), and (109) (e.g., a video stream) may be encoded according to certain video coding / compression standards. For example, those standards include the ITU-T H.265 Recommendation. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed topics are applicable in the context of VVC.

[0038] It should be noted that electronic devices (120) and (130) may include other components (not shown). For example, electronic device (120) may also include a video decoder (not shown), while electronic device (130) may also include a video encoder (not shown).

[0039] Figure 2 An example block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., receiving circuitry). The video decoder (210) may replace... Figure 1 The video decoder (110) in the example is used.

[0040] The receiver (231) can receive, for example, at least one encoded video sequence included in the bitstream, which is decoded by the video decoder (210). In one aspect, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences can be received from a channel (201), which can be a hardware / software link connected to a storage device storing the encoded video data. The receiver (231) can receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams that can be forwarded to their respective user entities (not depicted). The receiver (231) can separate the encoded video sequences from other data. To combat / mitigate network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be external to the video decoder (210) (not depicted). In other applications, a buffer memory (not depicted) may exist external to the video decoder (210), for example, to combat / mitigate network jitter, and another buffer memory (215) may exist internally to the video decoder (210), for example, to handle playback timing. The buffer memory (215) may not be necessary or may be made smaller when the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isosynchronous network. For use on best-effort packet networks such as the Internet, a buffer memory (215) may be required. The buffer memory (215) may be relatively large and may advantageously have an adaptive size, and may be implemented in part in the operating system or in a similar component (not depicted) external to the video decoder (210).

[0041] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the encoded video sequence. Categories of those symbols include information for managing the operation of the video decoder (210) and may include information for controlling a display device, such as a component of a non-electronic device (230) but which may be coupled to the electronic device (230) (e.g., a display screen). Figure 2 As shown in the diagram. Control information for one or more display devices may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (220) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set from the encoded video sequence for at least one pixel subgroup / subgroup in the video decoder based on at least one parameter corresponding to a group. Subgroups may include Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0042] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) in order to create symbols (221).

[0043] Depending on the type of the encoded video image or its portions (e.g., inter-frame and intra-frame images, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (220). For clarity, the flow of such subgroup control information between the parser (220) and the multiple units described below is not depicted.

[0044] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In practical implementations under commercial constraints, many of these units interact closely with each other and may be partially integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0045] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives from the parser (220) quantization transform coefficients as one or more symbols (221) and control information, including the transform to be used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output blocks containing sample values, which can be input into the aggregator (255).

[0046] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks do not use predictive information from previously reconstructed images, but instead can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-image prediction unit (252). In some cases, the intra-image prediction unit (252) uses surrounding reconstructed information obtained from the current image buffer (258) to generate blocks of the same size and shape as the blocks in the reconstruction. For example, the current image buffer (258) is used to buffer partially reconstructed current images and / or fully reconstructed current images. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0047] In other cases, the output samples of the scaler / inverse transform unit (251) may involve inter-frame coded and possibly motion-compensated blocks. In this case, the motion compensation prediction unit (253) can access the reference image memory (257) to obtain samples for prediction. After motion compensation of the obtained samples according to the block-related symbols (221), these samples can be added to the output of the scaler / inverse transform unit (251) (referred to in this case as residual samples or residual signals) by the aggregator (255) to generate output sample information. The address of the predicted samples obtained by the motion compensation prediction unit (253) in the reference image memory (257) can be controlled by motion vectors used by the motion compensation prediction unit (253) in the form of symbols (221), which may have, for example, X, Y, and reference image components. Motion compensation may also include interpolation of sample values ​​obtained from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0048] The output samples of the aggregator (255) can be employed by various loop filtering techniques in the loop filter unit (256). Video compression techniques may include loop filtering techniques controlled by parameters included in the encoded video sequence (also known as the encoded video stream) and available as symbols (221) from the parser (220) for use in the loop filter unit (256). Video compression may also respond to metadata obtained during the decoding of a previous (in decoding order) portion of the encoded image or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0049] The output of the loop filter unit (256) can be a sample stream, which can be output to the display device (212) and stored in the reference image memory (257) for future inter-frame image prediction.

[0050] Some encoded images can be used as reference images for future predictions after they have been fully reconstructed. For example, after the encoded image corresponding to the current image has been fully reconstructed and the encoded image has been identified as a reference image (by, for example, a parser (220)), the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.

[0051] The video decoder (210) can perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax of the video compression technology or standard and the configuration file documented therein if the encoded video sequence follows both the syntax of the video compression technology or standard and the configuration file documented therein. Specifically, the configuration file may select certain tools from all available tools in the video compression technology or standard as the only tools available under that configuration file. For compliance, the complexity of the encoded video sequence must also be within the limits defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer indicated by signaling in the encoded video sequence.

[0052] In one aspect, the receiver (231) may receive supplemental (redundant) data along with the encoded video. This supplemental data may be part of one or more encoded video sequences. The supplemental data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The supplemental data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0053] Figure 3 An example block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) may replace Figure 1 The video encoder (103) in the example is used.

[0054] The video encoder (303) can obtain data from the video source (301) (non- Figure 3 In one example, an electronic device (320) receives video samples, and a video source (301) can capture one or more video images to be encoded by a video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0055] A video source (301) can provide a sequence of source video samples in the form of a digital video sample stream to be encoded by a video encoder (303). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit…), any color space (e.g., BT.601 YCrCb, RGB…), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual images that, when viewed sequentially, produce an animated effect. The images themselves can be organized as spatial pixel arrays, where each pixel can include at least one sample, depending on the sampling structure, color space, etc., used. Samples are described in detail below.

[0056] According to one aspect, the video encoder (303) can encode and compress images of a source video sequence into an encoded video sequence (343) in real time or as needed, according to any other time constraints. Implementing an appropriate encoding rate is a function of the controller (350). In some aspects, the controller (350) controls and is functionally coupled to other functional units described below. For clarity, the coupling is not depicted. Parameters set by the controller (350) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, GOP layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions related to the video encoder (303) optimized for a specific system design.

[0057] In some respects, the video encoder (303) is configured to operate in an encoding loop. Simply put, in one example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and one or more reference images), and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to that created by the (remote) decoder. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since the decoding of the symbol stream produces a bit-by-bit accurate result unaffected by the decoder's location (local or remote), the contents of the reference image memory (334) are also bit-by-bit accurate between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same as the sample values ​​"seen" by the decoder during decoding when using the prediction. This fundamental principle of reference image synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related fields.

[0058] The operation of the “local” decoder (333) can be the same as that of a “remote” decoder such as a video decoder (210), which has already been described above. Figure 2 A detailed description has been provided. However, a brief additional reference is also available. Figure 2 When symbols are available and the entropy encoder (345) and parser (220) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (210), including the buffer (215) and parser (220), may not be fully implemented in the local decoder (333).

[0059] In one respect, decoder techniques other than parsing / entropy decoding, which exist in the decoder, exist in the corresponding encoder with the same or substantially the same functional form. Therefore, the disclosed subject focuses on decoder operation. The description of encoder techniques can be simplified, as encoder techniques are the reverse operation of the comprehensively described decoder techniques. More detailed descriptions are only required in certain areas, and these descriptions are provided below.

[0060] In some examples, during operation, the source encoder (330) may perform motion-compensated predictive coding, i.e., predictively coding the input image against at least one previously encoded image designated as a “reference image” in the reference video sequence. In this way, the coding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of one or more reference images, which may be selected as one or more predictive references for the input image.

[0061] The local video decoder (333) can decode encoded video data of an image that can be designated as a reference image, based on symbols created by the source encoder (330). Advantageously, the operation of the encoding engine (332) can be a lossy process. When the encoded video data can be decoded by the video decoder (333), Figure 3 During decoding at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process, which can be performed by the video decoder on the reference image, and can store the reconstructed reference image in the reference image memory (334). In this way, the video encoder (303) can store a copy of the reconstructed reference image locally, which shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0062] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can be used as suitable prediction references for the new image. The predictor (335) can operate on a pixel-by-pixel basis based on sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (334).

[0063] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for video data encoding.

[0064] The outputs of all the aforementioned functional units can be entropy encoded in the entropy encoder (345). By applying lossless compression to the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding, the entropy encoder (345) converts these symbols into an encoded video sequence.

[0065] The transmitter (340) may buffer one or more encoded video sequences created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which may be a hardware / software link connected to a storage device that will store the encoded video data. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0066] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a specific encoded image type to each encoded image, which may affect the encoding techniques applicable to the corresponding image. For example, one of the following image types can typically be assigned to an image.

[0067] Intraframe images (I-images) are images that are encoded and decoded without using any other images in the sequence as prediction sources. Some video codecs allow different types of intraframe images, including, for example, Independent Decoder Refresh (IDR) images.

[0068] Predictive images (P-images) can be images that are encoded and decoded by predicting sample values ​​for each block via intra-frame prediction or via inter-frame prediction using motion vectors and reference indices.

[0069] A bidirectional predictive image (B-image) can be an image encoded and decoded by predicting sample values ​​for each block via intra-frame prediction or via inter-frame prediction using two motion vectors and a reference index. Similarly, a multi-predictive image can reconstruct a single block using more than two reference images and associated metadata.

[0070] The source image can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the coding assignment of the corresponding images applied to the blocks. For example, blocks of an I-image can be unpredictably encoded, or these blocks can be predictively encoded with reference to already encoded blocks of the same image (spatial prediction or intra-frame prediction). Pixel blocks of a P-image can be predictively encoded via spatial prediction or via temporal prediction with reference to a previously encoded reference image. Blocks of a B-image can be predictively encoded via spatial prediction or via temporal prediction with reference to one or two previously encoded reference images.

[0071] The video encoder (303) can perform encoding operations according to predetermined video coding techniques or standards, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0072] In one respect, the transmitter (340) may transmit additional data along with the encoded video. The source encoder (330) may transmit such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant images and slices, SEI messages, VUI parameter set fragments, etc.

[0073] Video can be captured as multiple source images (video images) in a time-series manner. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In one example, the specific image being encoded / decoded is called the current image, which is divided into blocks. Where a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image and, in the case of multiple reference images, may have a third dimension identifying the reference images.

[0074] In some aspects, inter-frame image prediction can utilize bidirectional prediction techniques. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly representing past and future images in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. A block can be predicted using a combination of the first and second reference blocks.

[0075] Furthermore, merging mode techniques can be used to improve coding efficiency in inter-frame image prediction.

[0076] According to some aspects of the invention, prediction is performed on a block-by-block basis, such as inter-frame image prediction and intra-frame image prediction. For example, according to the HEVC standard, images in a video image sequence are divided into coding tree units (CTUs) for compression. Coding tree units in an image have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a coding tree unit comprises three coding tree blocks (CTBs): one luma coding tree block and two chroma coding tree blocks. Each coding tree unit can be recursively quadtree-based into one or more coding units. For example, a 64×64 pixel coding tree unit can be split into one 64×64 pixel coding unit, four 32×32 pixel coding units, or sixteen 16×16 pixel coding units. In one example, each coding unit is analyzed to determine its prediction type, such as inter-frame prediction or intra-frame prediction. Coding units are split into at least one prediction unit based on temporal and / or spatial predictability. Typically, each prediction unit comprises one luma prediction block (PB) and two chroma prediction blocks. In one aspect, prediction operations in encoding / decoding are performed on a block-by-block basis. Using the luma prediction block as an example, a prediction block comprises a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0077] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using at least one integrated circuit (IC). In another aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using at least one processor that executes software instructions.

[0078] Video coding is widely used in many applications such as broadcasting, video recording, and video streaming. Various emerging video coding standards are employed in video applications, such as H.264, H.265 / HEVC, H.266 / VVC, and AV1. Hybrid video codecs may include coding modules, intra-frame prediction, inter-frame prediction, transform coding, quantization, entropy coding, loop filtering (or loop filters), etc. A set of methods for video compression is disclosed, including methods related to template matching prediction.

[0079] In one example, a template is used in video encoding. In one aspect, the template can include a predefined region adjacent to the current block, such as... Figure 4 As shown. Because the template is close to the current block, the template and the current block may be highly correlated. When a reference template (e.g., the best reference template) is found in the reference frame, the reference block corresponding to that reference template may be optimal (e.g., the best). The reference template may be adjacent to the reference block. To find this template (e.g., the best template), the template cost needs to be evaluated. The template cost can be the sum of absolute differences (SAD), the sum of transformed differences (SATD), or other metrics, such as the mean removal SAD (MRSAD) between the reference template and the current template.

[0080] Using templates allows codecs to evaluate encoded information and reliably predict block content to a certain extent. Figure 4 An example of a template in inter-frame prediction according to one aspect of the invention is shown. In one example, such as Figure 4 As shown, a predefined search range (440) is used to search for an MV (e.g., the optimal MV) starting from the initial MV. The advantages of the template method include: it relies on already encoded regions, allowing the encoder and decoder to easily use the template without additional data transfer. Therefore, the template can efficiently guide block prediction and improve overall coding efficiency.

[0081] exist Figure 4In the current image (410), the current block (401) is being reconstructed. The current template (421) of the current block (401) may include the upper template (422) and the left template (423). The initial MV (402) points to the reference block (403) in the reference image (411). The MV can be searched within the search range (440). The template cost can be calculated based on the current template (421) and the reference template in the reference image (411). In one example, the template cost includes the template cost between the current template (421) and the reference template (425) corresponding to the initial MV (402). The reference template (425) includes the upper template (426) and the left template (427).

[0082] Templates are not only used in the template matching process, for example Figure 4 The motion vector refinement shown is also used in other video coding methods, such as LIC. LIC can be an inter-frame prediction technique used to model local illumination variations between the current block and its predicted block. Local illumination variations can be modeled as a function of local illumination variations between the current block template and the reference block template.

[0083] For example, when determining the motion vector, local illumination compensation is modeled using the following formula: p′[x] = α•p[x] + β. p[x] is the original pixel at position x in the reference template, p′[x] is the compensated pixel at the corresponding position in the current template, and α and β are compensation parameters. The parameters α and β can be determined using linear regression to ensure an optimal match between the compensated pixel value in the current template and the corresponding pixel value in the reference template. The derived pattern is then applied to the reference sample to determine the predicted value for the current coded block.

[0084] Figure 5 Examples of the current template (501) and the reference template (511) are shown. The current template (501) includes the neighboring samples of the current block (500), and the reference template (511) includes the neighboring samples of the corresponding reference block (510) used to derive the LIC parameters α and β.

[0085] In one example, when the coding unit encodes in merge mode, the LIC flag can be copied from a neighboring block in a manner similar to motion information copying in merge mode. In another example, the LIC flag can be signaled to the coding unit to indicate whether LIC is applied.

[0086] In one aspect, template-based inter-frame prediction methods can refer to inter-frame prediction methods that predict the current block based on reconstructed samples in the reconstructed neighboring region (e.g., the current template). The reconstructed neighboring region can be the current template of the current block, for example... Figure 4 The current template (421) or Figure 5The current template (501) is used. In some examples, misalignment occurs between the current template and the reference template when applying template-based inter-frame prediction methods. In some examples, this misalignment occurs because the current template is constructed using reconstructed samples before loop filtering, while the reference template is derived from samples that have already undergone loop filtering. For example, the reconstructed samples in the current template have not undergone loop filtering, while the reconstructed samples in the reference template have. In some examples, this mismatch between the two templates may reduce the accuracy of the matching process or model construction, resulting in poor coding efficiency.

[0087] In one aspect, reconstructed samples from neighboring coded blocks of the current block are used as reference samples for intra-prediction modes (including angular intra-prediction modes and non-angular intra-prediction modes) to generate prediction samples for the current block.

[0088] Reconstructed samples from spatially adjacent coded blocks can be used as templates for the current block (called the current template). This template can be used for template search to find a prediction block with at least one minimum template matching cost, or it can be used to derive intra-frame prediction modes using template features.

[0089] Template-based intra-frame prediction methods refer to intra-frame prediction methods that predict the current block based on reconstructed samples in the reconstructed neighboring regions (e.g., the current template). In one aspect, template-based prediction methods refer to prediction methods that predict the current block based on reconstructed samples in the reconstructed neighboring regions (e.g., the current template). Template-based prediction methods can refer to template-based inter-frame prediction methods, template-based intra-frame prediction methods, etc. In some examples, template-based intra-frame prediction methods include block vector (BV) based methods, such as intra-template matching prediction (IntraTMP) mode. Template-based prediction methods can refer to BV based methods, such as IntraTMP mode, intra-block copy (IBC) mode, etc.

[0090] In some examples, the loop filter includes one or more of the following: (i) a deblocking filter, (ii) a sample adaptive offset (SAO) filter, (iii) an adaptive loop filter (ALF), etc.

[0091] Figure 6Figures (610), (620), (630), and (640) illustrate examples of deblocking filter applications. For example, in Figure (610), a deblocking filter is applied to perform horizontal filtering on a vertical edge (611). Samples on both sides of the vertical edge (611) participate in the horizontal filtering, and these samples on both sides of the vertical edge (611) are modified by the horizontal filtering. The filter length can indicate the number of filtered samples along the horizontal axis, such as the samples modified by the deblocking filter. The filter length can be varied based on the filtering parameters. For example, in Figure (620), a similar deblocking filter is applied to perform horizontal filtering on a vertical edge (621), and the number of modified samples along the horizontal axis in Figure (620) is less than in the example in Figure (610). The filter lengths of the deblocking filters shown in (610) and (620) are 3 and 2, respectively. The deblocking filters shown in (610) and (620) are also referred to as horizontal deblocking filters.

[0092] In Figure (630), a deblocking filter is applied to perform vertical filtering on the horizontal edge (631). Samples on both sides of the horizontal edge (631) participate in the vertical filtering, and these samples on both sides of the horizontal edge (631) can be modified by the vertical filtering. The number of filtered samples (samples modified by the deblocking filter) can be changed based on the filtering parameters. For example, in Figure (640), a similar deblocking filter is applied to perform vertical filtering on the horizontal edge (641), and the number of modified samples in Figure (640) is less than the example in Figure (630). The filter lengths of the deblocking filters shown in (630) and (640) are 3 and 2, respectively. The deblocking filters shown in (630) and (640) are also called vertical deblocking filters.

[0093] In one example, a sub-level of a deblocking filter refers to a deblocking filter that filters along a single direction (e.g., horizontal or vertical). Each of the deblocking filters shown in (610), (620), (630), and (640) can be referred to as a sub-level of a deblocking filter.

[0094] Figure 7 Various filters, including loop filters, are shown according to one aspect of the invention. Figure 7 The loop filter for the luminance component is shown. Figure 7 Loop filters for chrominance components are shown, such as cross-component filters (e.g., CC-ALF) used to generate chrominance components. In some examples, Figure 7The filtering process for the first chromaticity component, the second chromaticity component, and the luminance component is illustrated. The luminance component can be filtered by the SAO filter (710) to generate the SAO-filtered luminance component (741). The SAO-filtered luminance component (741) can be further filtered by the ALF luminance filter (716) to become a filtered luminance coding block (CB) (761) (e.g., 'Y').

[0095] The first chromaticity component can be filtered by a SAO filter (712) and an ALF chromaticity filter (718) to generate a first intermediate component (752). Furthermore, the SAO-filtered luminance component (741) can be filtered by a cross-component filter (e.g., CC-ALF) (721) of the first chromaticity component to generate a second intermediate component (742). Subsequently, a filtered first chromaticity component (762) (e.g., 'Cb') can be generated based on at least one of the second intermediate component (742) and the first intermediate component (752). In one example, the filtered first chromaticity component (762) (e.g., 'Cb') is generated by combining the second intermediate component (742) and the first intermediate component (752) using an adder (722). The cross-component adaptive loop filtering process for the first chromaticity component may include steps performed by the CC-ALF (721) and steps performed by, for example, the adder (722).

[0096] The above description applies to the second chromaticity component. The second chromaticity component can be filtered by a SAO filter (714) and an ALF chromaticity filter (718) to generate a third intermediate component (753). Furthermore, the SAO-filtered luminance component (741) can be filtered by a cross-component filter (e.g., CC-ALF) (731) of the second chromaticity component to generate a fourth intermediate component (743). Subsequently, a filtered second chromaticity component (763) (e.g., 'Cr') can be generated based on at least one of the fourth intermediate component (743) and the third intermediate component (753). In one example, the filtered second chromaticity component (763) (e.g., 'Cr') can be generated by combining the fourth intermediate component (743) and the third intermediate component (753) using an adder (732). In one example, the cross-component adaptive loop filtering process for the second chromaticity component may include steps performed by a CC-ALF (731) and steps performed by, for example, an adder (732).

[0097] Cross-component filters (e.g., CC-ALF (721), CC-ALF (731)) can be operated by applying a linear filter with any suitable filter shape to the luminance component (or luminance channel) to refine each chrominance component (e.g., the first chrominance component, the second chrominance component).

[0098] According to one aspect of the invention, in video coding, template filtering based on template prediction (e.g., template-based inter-frame prediction, template-based intra-frame prediction, etc.) can be used. In one example, template filtering for inter-frame prediction is applied, wherein template filtering is performed on the template using one or more loop filters. In one aspect, it can be determined whether to apply filters to adjacent reconstructed regions of the current block. Figure 8 An example is shown of the current block (801) and its adjacent reconstructed regions (810) in the current image (800). In one example, the current block (801) is undergoing encoding / decoding (e.g., encoding or decoding). The adjacent reconstructed regions (810) may be adjacent to the current block (801). The adjacent reconstructed regions (810) may include reconstructed samples from the current image (800). In one aspect, the reconstructed samples in the adjacent reconstructed regions (810) are unfiltered, e.g., not filtered using a loop filter. The adjacent reconstructed regions (810) may include the current template of the current block. Figure 8 In the example shown, the current template (810) includes (i) an upper template (811) located directly above the current block (801) and (ii) a left template (812) located directly to the left of the current block (801). Filters can be applied to the adjacent reconstructed regions (810) of the current block (801). The current block (801) can be reconstructed based on the filtered adjacent reconstructed regions (810).

[0099] In one aspect, inter-frame prediction is used to predict the current block (801), and the current template (810) used for inter-frame prediction can be filtered, for example, before a template matching or model building process between the current template (810) and a reference template. In one example, inter-frame prediction includes one of the template matching and model building processes based on the current template (810) and the reference template. In one example, the reference template is located in a reference image different from the current image (800).

[0100] In one respect, the current block is encoded using one of (i) intra-frame prediction and (ii) BV-based prediction modes (e.g., IntraTMP mode, IBC mode, etc.).

[0101] In one example, the filter includes one or more loop filters. The one or more loop filters may include one or more of the following: (i) a deblocking filter, (ii) a SAO filter, (iii) a bilateral filter, (iv) an ALF, etc. For example, the applied filter includes at least one or more loop filters, such as a deblocking filter, a SAO filter, a bilateral filter, and an ALF.

[0102] In one example, the applied filter is, but is not limited to, a deblocking filter or a variant of a deblocking filter.

[0103] In one aspect, the applied filter is adaptively triggered based on the coding information in the adjacent reconstructed regions (810). In one example, whether to apply a filter is based on the partitioning information (also called partitioning information) of the adjacent reconstructed regions (or the current template) (810). For example, whether to apply a filter is determined based on the coding information of the adjacent reconstructed regions (810). In one example, the coding information of the adjacent reconstructed regions (810) includes the partitioning information of the current template (810).

[0104] In one aspect, when an edge is detected in the current template (810), a deblocking filter is applied to reduce blocking artifacts. When an edge is detected in an adjacent reconstructed region (810), a filter such as a deblocking filter is determined to be applied. In one example, the edge detection region is not limited to the current template (810). The edge detection region can be larger than the current template (810). For example, the current template (810) is... N ×1 template, including N ×1 reconstructed samples, of which N The width of the template above the current block (801) 。 Edge detection area (or N (×4 zone) can be used for N ×1 template for edge detection. N The ×4 region may include (i) N ×1 template and (ii) N ×3 area, this N The ×3 area is the vertical extension of the upper template (811).

[0105] In one aspect, sub-levels of the deblocking filter (e.g., (i) a horizontal deblocking filter for vertical edges or (ii) a vertical deblocking filter for horizontal edges) are adaptively applied to the current template (810), such as Figure 8As shown. In one example, the upper template (811) includes reconstructed samples from adjacent blocks (841)-(842). The adjacent blocks (841)-(842) share a vertical edge (831). In this case, a sub-level of the deblocking filter (e.g., a horizontal deblocking filter for the vertical edge (831)) can be applied to the upper template (811), and no vertical deblocking filter is applied to the vertical edge (831). In one example, the left template (812) includes reconstructed samples from adjacent blocks (843)-(844). The adjacent blocks (843)-(844) share a horizontal edge (832). In this case, a sub-level of the deblocking filter (e.g., a vertical deblocking filter for the horizontal edge (832)) can be applied to the left template (812), and no horizontal deblocking filter is applied to the horizontal edge (832).

[0106] In one example, an edge is detected in an adjacent reconstructed region. When the edge is a vertical edge, for example, when the edge is a vertical edge (831), the deblocking filter is a horizontal deblocking filter. When the edge is a horizontal edge, for example, when the edge is a horizontal edge (832), the deblocking filter is a vertical deblocking filter.

[0107] In one example, when the upper template (811) includes one or more internal vertical edges, samples in the upper template (811) along one or more vertical edges can be filtered using a horizontal deblocking filter targeting those one or more vertical edges. Reference Figure 8 The upper template (811) includes a vertical edge (831), and samples along the vertical edge (831) in the upper template (811) can be filtered using a horizontal deblocking filter.

[0108] In one example, when the left template (812) includes one or more internal horizontal edges, samples in the left template (812) along those one or more horizontal edges can be filtered using a vertical deblocking filter. Reference Figure 8 The left template (812) includes a horizontal edge (832), and samples along the horizontal edge (832) in the left template (812) can be filtered using a vertical deblocking filter.

[0109] like Figure 6 The above describes the application of a horizontal deblocking filter to vertical edges and a vertical deblocking filter to horizontal edges. Figure 6 The left column corresponds to a filter length of 3 (e.g., filtering up to 3 samples on each side of the edge), and filters up to 3 samples perpendicular to the edge. The right column has a filter length of 2.

[0110] In one example, refer to Figure 8The samples along the two vertical edges (851)-(852) of the template (811) are filtered. In this case, only the samples on one side of the corresponding edge (851) or (852) are filtered, for example, only the samples within the template (811) are filtered. For example, only the samples to the right of edge (851) are filtered, and only the samples to the left of edge (852) are filtered.

[0111] In one example, refer to Figure 8 The samples along the two horizontal edges (853)-(854) in the left template (812) are filtered. In this case, only the samples on one side of the corresponding edge (853) or (854) are filtered, for example, only the samples in the left template (812) are filtered. For example, only the samples below the edge (853) are filtered, and only the samples above the edge (854) are filtered.

[0112] The minimum of the width and height of adjacent reconstructed regions (810) can be at least four samples. In one aspect, the template size is at least four samples. The template size can refer to the height of the upper template (811) or the width of the left template (812). For example, the height of the upper template (811) and the width of the left template (812) are four samples. In one example, the minimum of the width and height of adjacent reconstructed regions is at least four samples.

[0113] In one respect, the above discussion applies to any suitable filter, including loop filters. The above discussion can be similarly applied to other loop filters besides deblocking filters.

[0114] In one respect, the template area is not limited to the area directly above and to the left, for example Figure 8 The top template (811) and left template (812) are shown. The template area may include the top left template, the top right template, and the bottom left template, such as... Figure 9 As shown. In Figure 9 In the example, the template size is 3. Figure 9An example of a template region or adjacent reconstructed region of a current block (901) in a current image (900) according to one aspect of the invention is shown. The adjacent reconstructed regions of the current block (901) may include one or more of the following adjacent to the current block (901): (i) the upper template (also called the upper adjacent reconstructed region) (911), (ii) the left template (also called the left adjacent reconstructed region) (912), (iii) the upper left template (also called the upper left adjacent reconstructed region) (913), (iv) the upper right template (also called the upper right adjacent reconstructed region) (914), and (v) the lower left template (also called the lower left adjacent reconstructed region) (915). In one example, each of (i) the upper template (911), (ii) the left template (912), (iii) the upper left template (913), (iv) the upper right template (914), and (v) the lower left template (915) includes a reconstructed sample in the current image (900). Figure 9 As shown, the adjacent reconstruction regions of the current block (901) include one or more of the following: the upper adjacent reconstruction region (911) located directly above the current block (901), the left adjacent reconstruction region (912) located directly to the left of the current block (901), the upper left adjacent reconstruction region (913), the upper right adjacent reconstruction region (914), and the lower left adjacent reconstruction region (915).

[0115] In one respect, the above discussion applies to either the luminance component or the chrominance component. In one example, the above discussion similarly applies to the chrominance component using a cross-component loop filter (e.g., a cross-component SAO (CCSAO) filter, CC-ALF, etc.).

[0116] In one aspect, the filtering strength and / or filtering conditions of the chroma component are derived from information in the luminance component, which directly corresponds to the chroma component. Filtering conditions may include the presence of edges, the type of filter to be applied, etc. Filtering strength may include the degree or intensity of the filter. An example of filtering strength is boundary strength (Bs), which indicates the degree or intensity of the deblocking filtering process at edges. For example, the filtering strength and / or filtering conditions of the chroma component may inherit from or reuse the filtering strength and / or filtering conditions of the luminance component, respectively.

[0117] One approach involves using enhancement or augmentation methods based on reference pixels (also known as reference samples), or using templates from intra-frame prediction modes. Figure 8The current block (801) in the current image (800) is not encoded or decoded using inter-frame prediction. For example, the current block (801) is encoded or decoded using reference samples (e.g., reconstructed samples) in the current image (800), such as using intra-frame prediction, BV-based prediction modes, etc. Examples of BV-based prediction modes include IBC mode, IntraTMP mode, etc.

[0118] A filter can be applied to a reference sample or a neighboring reconstructed region (or current template) (810) of the current block (801). The reference sample may include reconstructed samples in the neighboring reconstructed region (or current template) (810).

[0119] In one aspect, filters are adaptively applied based on partitioning and / or encoding information within adjacent reconstructed regions (810) (e.g., one or more reference sample rows / columns). The adjacent reconstructed regions (810) may be referred to as template regions or the current template, and may include one or more reference sample rows and / or one or more reference sample columns.

[0120] In one example, the encoding information within adjacent reconstructed regions includes partitioning information. This partitioning information indicates that the adjacent reconstructed regions include reconstructed samples from multiple coded blocks in the current image. These multiple coded blocks form at least one coded block boundary within the adjacent reconstructed regions. Each of these at least one coded block boundary can be located between two adjacent coded blocks. The filter is applied only to reconstructed samples adjacent to at least one coded block boundary.

[0121] In one aspect, the filter is applied only to samples within the reference sample row / column or template along the boundaries of coded blocks. Adjacent reconstructed regions can be formed by multiple coded blocks, and adjacent coded blocks within multiple blocks can form coded block boundaries that separate adjacent coded blocks. Figure 10 An example is shown of a block (or the current block) (1001) and the current template (or adjacent reconstructed region) (1021) formed by the three adjacent blocks (1011)-(1013) of block (1001). In one example, blocks (1011)-(1013) have been encoded, for example, reconstructed and include a reconstructed sample or reference sample. The encoded block boundary (also called the edge) (1031) separates blocks (1011)-(1012). The encoded block boundary (1032) separates blocks (1011) and (1013).

[0122] In one example, the filter is applied to samples along the vertical edge of the template or reference line above, such as... Figure 10As shown. The upper template (1015) includes blocks (1011)-(1012) separated by a vertical edge (1031). A filter is applied to samples in the upper template (1015) along the vertical edge (1031). In one example, the current template (1021) includes a vertical edge (1032). A filter can be applied to samples in the current template (1021) along the vertical edge (1032). Figure 10 The black dots in the diagram indicate the filter length (2 in this case) perpendicular to the edges (1031)-(1032). The filter applied to the vertical edge (1031) can be the same as or different from the filter applied to the vertical edge (1032).

[0123] In one aspect, the filter is applied row-wise and / or column-wise to each adjacent reconstruction region (1021) of the current block (1001) (e.g., one or more reference sample rows and / or one or more reference columns or the current template). M One sample. M It can be an integer, such as 1, 2, 4, or 8. In one example, the filter is applied row by row to the upper template (1015), for example, filtering samples in every M columns of the upper template (1015). The current template (1021) includes the upper template (1015) and the left template (1013).

[0124] In one respect, the filter type decision and / or filter strength in the deblocking filter are applied to filtering the reference sample or the current template (1021).

[0125] In one respect, refer again Figure 8 The current block (801) does not use inter-frame prediction for prediction, and instead uses intra-frame prediction or a BV-based prediction mode. When an edge is detected in the current template (810), a filter can be applied to samples along that edge. In one example, as referenced above... Figure 8 The edge detection area is not limited to the current template (810). The edge detection area can be larger than the current template (810). For example, the current template (1021) can include... N ×1 reconstructed samples N ×1 template 。 Edge detection area (or N (×4 zone) can be used for N ×1 template for edge detection. N The ×4 region may include (i) N ×1 template and (ii) N ×3 area, this N The ×3 region is the vertical extension of the template (811) described above.

[0126] In one respect, a reference sample or current template (e.g., Figure 8 (810) or Figure 10 (1021) in the middle) and the current block (e.g., Figure 8 (801) or Figure 10 The intra-frame prediction of (1001) is associated with this.

[0127] In one aspect, during the template matching process, for example in IntraTMP mode, when the reference image refers to the reconstructed region of the current coded frame (also known as the current image), a filter is applied to the template of the reference block. Figure 10 The current image (1000) includes the reconstructed area, the current block (1001) being reconstructed, and the area to be reconstructed (after the current block (1001) is being reconstructed). The template of the reference block is located in the reconstructed area of ​​the current image (1000) and includes the reconstructed sample.

[0128] In one aspect, when the reference image is the reconstructed region of the current coded frame, during the template matching process, the filter is applied not only to the current template of the current block but also to the template of the reference block. Figure 10 The current image (1000) includes the reconstructed region, the current block (1001) being reconstructed, and the region to be reconstructed (after the current block (1001) is reconstructed). The template of the reference block is located in the reconstructed region of the current image (1000) and includes reconstructed samples. During the template matching process, such as during IntraTMP, filters are applied to the current template (1021) of the current block (1001) and the template of the reference block.

[0129] According to one aspect of the invention, a signaling indication based on reference sample characteristics is described, including an efficient method for control information signaling indication of an encoding tool utilizing reference sample information from the decoding end, for example, to improve compression efficiency.

[0130] Related video and image codecs utilize reference sample information available to the decoder in their workflows. Some examples of such methods include intra-frame prediction (e.g., regular intra-frame prediction), matrix-based intra-frame prediction (MIP), LIC, etc. In some examples, one such method has control flags indicating the use of that method, as well as certain syntax indicating a specific template type, such as the template being a left-side template, an upper-side template, an L-shaped template, etc. Methods for improving the efficiency of both control flag signaling indication and template type signaling indication based on available information (e.g., partitioning of adjacent blocks and / or prediction information) are disclosed.

[0131] In one aspect, a template-based prediction method (also known as a template-based approach or template-based prediction mode) refers to a prediction method that predicts the current block (also known as the current coding block) in the current image based on reconstructed samples in the current image. The reconstructed samples in the current image may be referred to as decoder-available reconstructed samples or decoder-available reference samples. In one aspect, the reconstructed samples in the current image used in the template-based prediction method are located in the reconstructed neighbor region (e.g., the current template) of the current block. In this invention, without loss of generality, any intra-frame or inter-frame coding method that can utilize information from the decoder-available reconstructed samples of the current coding block may be called a "template-based method," and the decoder-available reference sample region may be called a "template" (e.g., the current template) or a "template region." For example, the reconstructed neighbor region or the current template of the current block is used for pattern derivation of the current block, prediction of the current block, etc.

[0132] Examples of template-based prediction methods include MIP, LIC, intra-prediction modes (e.g., angle intra-prediction modes), IntraTMP, and template-matching-based inter-prediction.

[0133] In one aspect, based on at least one of the partitioning information and prediction information corresponding to the template region, signaling is used to indicate the control flag, template type, and template-based method type.

[0134] According to one aspect of the invention, based on either segmentation information or prediction information of reconstructed samples (or reconstructed samples available to the decoder) in the current image, it can be determined whether to indicate one of the following in the video stream via signaling: (i) a control flag for predicting a prediction mode (e.g., a template-based prediction mode) for the current block in the current image, and (ii) the type of prediction mode (e.g., the type of template-based method). In one example, based on either segmentation information or prediction information of adjacent reconstructed regions of the current block in the current image, it is determined whether to indicate one of the following in the video stream via signaling: (i) a control flag for predicting a prediction mode for the current block, and (ii) the type of prediction mode. The current block can be predicted based on reconstructed samples in adjacent reconstructed regions adjacent to the current block. The current block can be encoded according to the prediction mode and the reconstructed samples in adjacent reconstructed regions. When it is determined that one of the control flag and the type of prediction mode should be indicated via signaling, this one of the control flag and the type of prediction mode is encoded into the video stream.

[0135] In one aspect, signaling can be used to indicate control flags for template-based methods based on partitioning information corresponding to the template region. In one example, the partitioning information corresponding to the template region refers to the partitioning information of that template region.

[0136] In one aspect, signaling control flags are used only when the template (current template) meets specific conditions determined based on neighboring block partitioning information, which is the partitioning information of neighboring blocks in the adjacent reconstructed region (e.g., the current template). For example, whether to use signaling control flags to indicate the prediction mode in the video stream is determined based on the partitioning information of the adjacent reconstructed regions of the current block.

[0137] In one example, only if the template region belongs to no more than Signaling is used to indicate control flags only when there are L different coding blocks. Otherwise, signaling is not used to indicate control flags, and the value of the control flags is inferred to be zero (e.g., disabled). For example, the control flags for indicating the prediction mode in the video stream are determined only when the partitioning information indicates that the reconstructed samples in adjacent reconstructed regions come from no more than L different coding blocks.

[0138] Figure 11 An example of a block (e.g., the current block) (1101) and a template (e.g., the current template) (1121) according to one aspect of the invention is shown. The current template (1121) is formed by two adjacent blocks (1111)-(1112). Figure 11 In the example shown, the current block (1101) is a square block, and the current template (1121) is an L-shaped template region formed by two adjacent blocks (1111)-(1112). In the value Set as In this example, signaling is used to indicate control flags for the template-based method.

[0139] Figure 12 An example of a block (e.g., the current block) (1201) and a template (e.g., the current template) (1221) according to one aspect of the invention is shown. The current template (1221) is formed by three adjacent blocks (1211)-(1213). Figure 12 In the example shown, the current block (1201) is a square block, and the current template (1221) is an L-shaped template region formed by three adjacent blocks (1211)-(1213). In the value In this example, where the value is set to 2, no signaling is used to indicate the control flags for the template-based method, and it can be inferred that they are 0.

[0140] In one example, if the actual template region (e.g., the current template) to be used by the template-based method cannot be determined before control flag signaling indication and / or reading, (i) a subset of all possible template regions or (ii) the maximum possible template region can be used to determine whether to indicate and / or read control flags with signaling. In one example, the maximum possible template region includes all possible template regions.

[0141] In one example, when it is determined that a control flag should be indicated by signaling, the entropy context of the control flag is determined based on the partitioning information in adjacent reconstructed regions (e.g., the current template).

[0142] In one aspect, the entropy context of the control flag is determined based on neighbor block partitioning information, which is the partitioning information of neighboring blocks in adjacent reconstructed regions (e.g., the current template). In one example, if the template region belongs to no more than If there are multiple distinct coded blocks, then a context (e.g., the first context) is selected; otherwise, if the template region belongs to no more than... If there is one block, then another context (e.g., a second context) is selected, where In one example, if the template region belongs to more than If there is a block, then the third context is selected.

[0143] In one example, if the template region is formed by two adjacent blocks, the context of the control flag is set to 0, indicating the first context; otherwise, if the template region is formed by three adjacent blocks, the context of the control flag is set to 1, indicating the second context; otherwise, the context of the control flag is set to 2, indicating the third context, and so on.

[0144] In another example, if the template region is formed by two adjacent blocks, the context of the control flag is set to 0, indicating that the context is the first context; otherwise, the context of the control flag is set to 1, indicating that the context is the second context.

[0145] In another example, if the template region is formed by the minimum possible number of adjacent blocks, the context of the control flag is set to 0, indicating the first context; otherwise, the context of the control flag is set to 1, indicating the second context. For example, if the minimum possible number is 2 and the template region includes two adjacent blocks, the context of the control flag is the first context; otherwise, the context of the control flag is the second context.

[0146] In another example, if the template region is formed by two adjacent blocks, the previously determined context number of the control flag is multiplied by 1; otherwise, if the template region is formed by three adjacent blocks, the context number of the control flag is multiplied by 2; otherwise, the context number of the control flag is multiplied by 3, and so on. In one example, the previously determined context number of the control flag is based on either the decoding order or the encoding order. In one example, the decoding order is the same as the encoding order.

[0147] In another example, if the template region is formed by two adjacent blocks, the previously determined context number of the control flag is multiplied by 1; otherwise, it is multiplied by 2.

[0148] The context selection method in this invention can be combined with any other context selection method, such as analyzing the similarity of control flags of adjacent blocks (e.g., the left block and the top block).

[0149] In one aspect, the template type of the template-based method can be indicated by signaling based on the partitioning information corresponding to the template region. In one example, the adjacent reconstruction regions (or the current template) used to predict the current block are determined, and the adjacent reconstruction regions (or the current template) are indicated by signaling based on the partitioning information of the adjacent reconstruction regions.

[0150] In one respect, a specific template type can only be indicated by signaling if the template meets specific conditions determined based on the partitioning information of the template region (e.g., the partitioning information of adjacent blocks in the template region).

[0151] In one example, if the template region consists of more than If a block is formed, the template cannot be indicated by signaling, and the template cannot be used. For example, the reconstructed samples in the template region (e.g., adjacent reconstructed regions) are constrained to no more than N different coded blocks. Figure 12 An example of the current square block (1201) and the current template (1221) is shown. The current template (1221) may include an upper template and a left template, the left template being block (1212). The upper template is formed by two coded blocks (1211) and (1213). If If set to 1, the upper template cannot be signaled and cannot be used, for example, because the upper template is formed by more than N (e.g., 1) blocks.

[0152] In one aspect, signaling indications for a specific template type can be determined based on the partitioning information of the template region (or the current template) (e.g., a binarization scheme).

[0153] In one aspect, template candidates (e.g., all possible templates) can be formed into an ordered list (also called a template list) based on partitioning information corresponding to each template candidate (e.g., the corresponding one among the possible templates), and can be indicated by signaling in different ways based on the different positions of the template candidates in the template list. Template candidates can refer to multiple adjacent reconstructed regions adjacent to the current block.

[0154] In one example, a list (e.g., a template list) including multiple adjacent reconstructed regions (e.g., multiple template candidates) adjacent to the current block is constructed based on the corresponding partitioning information of the multiple adjacent reconstructed regions. The multiple adjacent reconstructed regions include this adjacent reconstructed region. The adjacent reconstructed regions can be indicated by signaling based on their position in the list.

[0155] In one example, if the template-based prediction method can indicate the left template, the top template, or the L-shaped template with signaling, then the template list is first constructed by checking the number of partitions belonging to each template (e.g., the left template, the top template, or the L-shaped template), and the template list can be sorted in ascending order such that the first position of the template list is set to the template with fewer partitions (e.g., the fewest partitions), and the smaller number in the template list corresponds to the shorter codeword in the bitstream.

[0156] exist Figure 12 In the example shown, the template candidates include a left template (1212), an upper template (1210) comprising blocks (1211) and (1213), and an L-shaped template (1221). An ordered list of possible partitions (i.e., a template list) can be formed as 1: "left template (1212)", 2: "upper template (1210)", and 3: "L-shaped template (1221)", because the left template (1212) is formed by only one block (e.g., the left template (1212) does not involve template partitioning), the upper template (1210) is formed by two blocks, and the L-shaped template (1221) is formed by three blocks (1211)-(1213). In one example, the binarized signaling indication is implemented as follows: "0" indicates the left template (1212), "01" indicates the upper template (1210), and "11" indicates the L-shaped template (1221). In this example, a bit "0" is used to signal the left template (1212), two bits "01" are used to signal the top template (1210), and two bits "11" are used to signal the L-shaped template (1221).

[0157] On the other hand, the entropy context of the template type is determined based on the partitioning information of the current template.

[0158] In one aspect, based on partitioning information and / or prediction information corresponding to the template region, a signaling method type is used to indicate the template-based method type. The template-based method type can refer to the type of template-based method. Multiple template-based methods can have the same template-based method type. An example of a template-based method type is intra-frame prediction, and an example of a template-based method is a specific-angle intra-frame prediction mode.

[0159] In one aspect, only partitioning information from reference data (e.g., reconstructed samples) from the template region is used for signaling indication and / or entropy context modeling, and the reference data can be used to generate predicted samples for the current block. For example, when predicting predicted samples for the current block based solely on the upper template, only partitioning information within the upper template region is used for signaling indication and / or entropy context modeling.

[0160] An image can be divided into multiple coding units using any suitable method. For example, according to the HEVC standard, an image can be split into multiple coding tree units. Furthermore, a quad-tree (QT) structure, called a coding tree, can be used to further divide coding tree units into coding units to accommodate various local characteristics of the image. The decision on whether to use inter-frame image prediction (also known as temporal prediction or inter-prediction type), intra-frame image prediction (also known as spatial prediction or intra-prediction type), etc., to encode image regions is made at the coding unit level. Each coding unit can be further divided into one, two, or four prediction units based on the prediction unit splitting type. Within a prediction unit, the same prediction process is applied, and the same prediction information is transmitted to the decoder on a unit-by-unit basis. After obtaining residual data or residual information by applying the prediction process based on the prediction unit splitting type, the coding unit can be divided into transform units according to another quad-tree structure similar to the coding tree of the coding unit. In one example, a transform is applied to each transform unit with the same transform information. The HEVC structure has multiple partitioning units, including coding units, prediction units, and transform units. Samples in a coding unit can have the same prediction type, samples in a prediction unit can have the same prediction information, and samples in a transform unit can have the same transform information. For inter-frame prediction blocks, the coding unit or transform unit has a square shape, while the prediction unit can have a rectangular shape; in some embodiments, the rectangular shape includes a square shape. In some examples, such as in the JEM standard, prediction units with a rectangular shape can be used for intra-frame prediction.

[0161] According to the HEVC standard, implicit QT splitting is applied to the coding tree units located at the image boundary to recursively split the coding tree units into multiple coding units, so that each coding unit is located inside the image boundary.

[0162] In various embodiments, such as in the HEVC standard, coded tree blocks, coded blocks, prediction blocks, and transform blocks (TBs) can be used to specify, for example, a two-dimensional sample array of a color component associated with a corresponding coded tree unit, coded unit, prediction unit, and transform unit, respectively. Thus, a coded tree unit may include one or more coded tree blocks, such as one luma coded tree block and two chroma coded tree blocks. Similarly, a coded unit may include one or more coded blocks, such as one luma coded block and two chroma coded blocks.

[0163] In addition to the block divisions mentioned above Figure 13 An example of a block partitioning structure according to one aspect of the invention is also shown. The block partitioning structure uses a QT plus a binary tree (BT) and may be referred to as a QTBT structure or QTBT partitioning. Compared to the QT structure described above, the QTBT structure eliminates the separation of encoding units, prediction units, and transform units, and supports more flexible encoding unit partitioning shapes. In the QTBT structure, the encoding tree unit is split into multiple encoding units using the QTBT structure, and the encoding units can have a rectangular shape; in some embodiments, this rectangular shape includes a square shape. In various embodiments, the encoding unit serves as both a prediction and transform unit; therefore, samples within the encoding unit can have the same prediction type, can be encoded using the same prediction process, and can have the same prediction information and the same transform information.

[0164] Figure 13 (Left side) shows an example of block partitioning using QTBT. Figure 13 (Right side) shows the corresponding QTBT tree representation (1315). Solid lines indicate QT splits, and dashed lines indicate BT splits. In each split (i.e., non-leaf) node of the binary tree, a signaling flag is used to indicate the type of split used (i.e., symmetric horizontal split or symmetric vertical split). For example, "0" indicates a symmetric horizontal split, and "1" indicates a symmetric vertical split. For quadtree splits, no signaling is used to indicate the split type because quadtree splits non-leaf nodes simultaneously in both the horizontal and vertical directions, resulting in four smaller nodes of equal size.

[0165] refer to Figure 13First, the coding tree unit (1310) is divided (or split) into nodes (1301)-(1304) using a quadtree structure. Nodes (1301)-(1302) are further divided using a binary tree structure. As mentioned above, BT splitting includes two types of splitting: symmetrical horizontal splitting and symmetrical vertical splitting. The quadtree nodes (1303) are further divided using a combination of the BT and QT structures. Nodes (1304) are not further divided. Therefore, the un-split binary tree leaf nodes (1311)-(1320) and quadtree leaf nodes (1304)-(1306) are coding units used for prediction and transformation processing. In this example, the coding unit, prediction unit, and transformation unit are identical in the QTBT structure. For example, samples in the coding unit have the same prediction type, the same prediction information, and the same transformation information. In QTBT partitioning, a coding unit may include coding blocks for different color components. For example, for P and B slices in 4:2:0 chroma format, a coding unit includes one luma coding block and two chroma coding blocks. In some examples, a coding unit may include coding blocks for a single component; for example, for I slices, a coding unit includes one luma coding block or two chroma coding blocks.

[0166] The following parameters are defined for QTBT partitioning. The encoding tree unit size refers to the size of the root node of the quadtree. For example, Figure 13 In the example, the root node or encoding tree unit is (1310). MinQTSize refers to the minimum allowed size of a quadtree leaf node. MaxBTSize refers to the maximum allowed size of a binary tree root node. For example, Figure 13 In the example, node (1301) is the root node of the binary tree. MaxBTDepth refers to the maximum allowed depth of the binary tree. MinBTSize refers to the minimum allowed size of the leaf nodes in the binary tree.

[0167] In one example of QTBT partitioning, the coding tree unit size is set to 128×128 luma samples, corresponding to two 64×64 chroma sample blocks. MinQTSize is set to 16×16, MaxBTSize to 64×64, MinBTSize (the width and height of the binary tree leaf node) is set to 4×4, and MaxBTDepth is set to 4. First, quadtree partitioning is applied to the coding tree unit to generate quadtree leaf nodes. The size of the quadtree leaf node can range from 16×16 (i.e., MinQTSize) to 128×128 (i.e., the coding tree unit size). If the size of the quadtree leaf node is 128×128, the binary tree will not further partition the quadtree leaf node because 128×128 exceeds MaxBTSize (i.e., 64×64). Otherwise, the binary tree can further partition the quadtree leaf nodes. Therefore, the quadtree leaf node can serve as the root node of the binary tree, and its binary tree depth is 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further splitting is performed. When the width of a binary tree node equals MinBTSize (i.e., 4), no further horizontal splitting is performed. Similarly, when the height of a binary tree node equals MinBTSize, no further vertical splitting is performed. Leaf nodes of the binary tree undergo further processing or encoding through prediction and transformation, without any further partitioning. In the JEM standard, in some examples, the maximum coding tree unit size is 256×256 luminance samples.

[0168] In some examples, such as for P-slices and B-slices, the luma coding tree block and chroma coding tree block within a coding tree unit share the same QTBT structure. On the other hand, QTBT partitioning supports luma and chroma having independent QTBT structures. For example, for an I-slice, the luma coding tree block is partitioned into luma coding units using one QTBT structure, while the chroma coding tree block is partitioned into chroma coding units using another QTBT structure. Therefore, a coding unit in an I-slice can include a coding block for one luma component or two chroma components, while a coding unit in a P-slice or B-slice can include coding blocks for all three color components.

[0169] In some examples, such as the HEVC standard, inter-frame prediction for small blocks is limited to reduce memory access for motion compensation; therefore, 4×8 blocks and 8×4 blocks do not support bidirectional prediction, and 4×4 blocks do not support inter-frame prediction. In some embodiments, such as QTBT implemented in the JEM standard, the above limitations are removed.

[0170] Multi-type tree (MTT) structures can be flexible tree structures. In an MTT, horizontal and vertical center-side ternary trees (TT) can be used for partitioning or splitting, such as... Figures 14A-14B As shown. Ternary tree partitioning can also be called three-level tree partitioning. Figure 14A An example of a vertical center-side ternary tree partition is shown. For example, region (1420) is vertically split into three subregions (1421)-(1423), where subregion (1422) is located in the middle of region (1420). Figure 14B An example of a horizontally centered ternary tree partition is shown. For example, region (1430) is horizontally partitioned into three smaller subregions (1431)-(1433), with subregion (1432) located in the middle of region (1430). In various examples, regions (1420) and (1430) can be coding tree units or coding units, or nodes that can be further partitioned, such as node (1301). One or more of the subregions (1421)-(1423) and (1431)-(1433) can be coding units that are no longer further partitioned, or nodes that can be subsequently partitioned.

[0171] According to one aspect of the present invention, a set of video compression methods is described, including partitioning methods. The partitioning methods include flexible partitioning and splitting methods in video coding.

[0172] When compressing video frames, block-based video codecs divide pixels or samples in a video frame into multiple rectangular regions called blocks, and can apply different compression strategies to each block. Pixel regions can be recursively divided into blocks. At each division depth, a signaling indicator flag can be used to indicate whether the current region should be further subdivided. If no further subdivision is needed, a block is formed in the current region. If further subdivision is needed, a signaling indicator type can be used to indicate how the current region will be subdivided into multiple sub-regions. This process can continue recursively. In some video coding standards, each possible partitioning type is typically fixed. For example, the H.266 video coding standard provides five partitioning types to divide a region into smaller sub-regions. In various examples, each of the five partitioning types divides the current region into sub-regions in a fixed manner. Figure 15 Examples of possible partitioning (e.g., five fixed partitionings) according to one aspect of the invention are shown. Figure 15The partitions (1501)-(1505) shown can be used in the H.266 video coding standard. Partition (1501) refers to, for example, a symmetrical horizontal partition of BT. Partition (1502) refers to, for example, a symmetrical vertical partition of BT. Partition (1503) refers to QT. Partition (1504) refers to, for example, a horizontal center-side ternary tree partition of TT. Partition (1505) refers to, for example, a vertical center-side ternary tree partition of TT.

[0173] In one aspect, a type of partitioning called flexible partitioning is described. Similar to, for example, reference... Figure 13-15 The fixed partitioning method allows for the selection of a flexible partitioning type at each partitioning depth. Unlike fixed partitioning, which divides the current region into sub-regions in a specific way, the flexible partitioning mechanism (also known as the partitioning mechanism or partitioning structure) is derived based on neighborhood information or information indicated by signaling. The partitioning mechanism refers to the way sub-regions are generated or the way the current region is divided. When using flexible partitioning, the partitioning mechanism can be varied.

[0174] According to one aspect of the invention, encoded information of a current region to be segmented in the current image can be received. The segmentation mechanism or structure of the current region can be determined based on one of (i) the size information of the current region and (ii) the segmentation information of another region. The current region can be reconstructed based on the determined segmentation structure of the current region.

[0175] In one respect, flexible partitioning reuses the same partitioning mechanism used in another region of the same size. In one example, within a frame, the other region is located in the same image (e.g., the current image) as the current region. Between frames, the other region may be located in (i) an image different from the current image, or (ii) within the current image.

[0176] In one aspect, flexible partitioning reuses the partitioning mechanism of regions of the same size located at any position in an encoded frame for the current region. In another aspect, the partitioning is determined during the parsing stage, and another region has already been parsed. For example, if another region has the same size as the current region and is located at any position in an encoded frame, the partitioning mechanism of the other region is reused as the partitioning mechanism of the current region. In some examples, the encoded frame may refer to the current frame being encoded, and the partitioning information available in the encoded frame is the partitioning information of the parsed region. Therefore, if another region is located in the current image (or the current frame), then the other region has been parsed. In some examples, the other region may be located in a parsed image (different from the current image).

[0177] In one example, the flexible partitioning reuses the partitioning mechanism of the same-sized region located to the left of the current region.

[0178] In one example, flexible partitioning reuses the partitioning mechanism of a region of the same size located above the current region (e.g., at the top of the current region).

[0179] In one example, the flexible partitioning reuses the partitioning mechanism of the same-sized region located to the upper left of the current region.

[0180] In one example, flexible partitioning reuses the partitioning mechanism in the bitstream that uses regions of the same size indicated by displacement vectors relative to the current region.

[0181] In one example, the flexible partitioning reuses a mechanism that infers partitioning regions of the same size based on the neighborhood partitioning information of the current region. The neighborhood partitioning information of the current region indicates the partitioning information of the adjacent regions. In one example, the adjacent regions have already been resolved.

[0182] In one aspect, flexible partitioning is used to predict the partitioning regions of the current region based on the size information, adjacency information, and / or information indicated by signaling (e.g., partitioning structure).

[0183] In one aspect, flexible partitioning is predicted based on the size information of the current region and / or neighborhood partitioning information (e.g., partitioning information of neighboring regions of the current region). In one example, neighboring regions have already been resolved.

[0184] In such Figure 16 In the example shown, flexible partitioning is predicted by using partitioning information from regions of the same size above or to the left of the current region, by reducing the partitioning depth by a constant factor (e.g., 1). Figure 16 An example of flexible partitioning based on predictions of the upper region (1601) or the left region (1602) of the current region (1605) is shown, where the partitioning depth is reduced by 1. The upper region (1601) and the left region (1602) have the same size as the current region (1605). The partitioning depth of the upper region (1601) and the left region (1602) is 3.

[0185] The partitioning depth of partitioning pattern (1611) is 2, which is the partitioning depth of the upper region (1601) minus 1. For example, the upper region (1601) is first partitioned into partitioning pattern (1611), then BT is used to further partition block (1621) into blocks (1631)-(1632), and TT is used to further partition block (1622) into blocks (1641)-(1643). When partitioning the current region (1605) using the partitioning information of the upper region (1601), partitioning pattern (1611) is used to partition the current region (1605).

[0186] The partitioning depth of partitioning pattern (1612) is 2, which is the partitioning depth of the left region (1602) minus 1. For example, the left region (1602) is first partitioned into partitioning pattern (1612), then BT is used to further partition block (1651) into blocks (1661)-(1662), and TT is used to further partition block (1652) into blocks (1671)-(1673). When partitioning the current region (1605) using the partitioning information of the left region (1602), partitioning pattern (1612) is used to partition the current region (1605).

[0187] In one example, the directionality of partitioning information in the neighborhood of the current region is used to predict flexible partitioning. Figures 17A-17B This illustrates a flexible partitioning scheme predicted using the directionality of partitioning information in the neighborhood (or adjacent regions) of the current region (1701). The predicted flexible partitioning scheme is as follows: Figure 17B As shown.

[0188] In one aspect, neighboring regions of the current region (1701) are analyzed to identify emerging partitions (e.g., small partitions whose partition depth meets a condition (e.g., greater than or equal to a threshold)). These emerging partitions can be grouped according to their potential orientation, and the orientation information can then be used to predict partitions within the current region (1701). In one example, the orientation information can be extrapolated to the current region (1701) to generate a prediction of flexible partitioning within the current region (1701). Reference Figures 17A-17B The adjacent regions of the current region (1701) include a first group region (1711) and a second group region (1712). The first group region (1711) is arranged along a first direction (1721). A partition (1731) in the current region (1701) can be generated by extrapolating the first group region (1711) into the current region (1701) along the first direction (1721). The second group region (1712) is arranged along a second direction (1722). A partition (1732) in the current region (1701) can be generated by extrapolating the second group region (1712) into the current region (1701) along the second direction (1722).

[0189] In one example, flexible partitioning is predicted by extrapolating partitioning information from the neighborhood of the current region. Figure 18A-18G This demonstrates a flexible partitioning predicted by extrapolating partitioning information in the neighborhood (also known as adjacent regions) of the current region (1801). Figure 18A The current region (1801) and the adjacent regions (1802)-(1804) are shown.

[0190] exist Figure 18BIn this context, the partitioning pattern of the current region (1801) is predicted based on the partitioning pattern of the adjacent region (1804). For example, the partitioning pattern of the adjacent region (1804) is directly used as the prediction of the partitioning pattern of the current region (1801).

[0191] exist Figure 18C In this context, the partitioning pattern of the current region (1801) is predicted based on the partitioning pattern of the adjacent region (1802). For example, the partitioning pattern of the adjacent region (1802) is directly used as the prediction of the partitioning pattern of the current region (1801).

[0192] exist Figure 18D In this context, the current region (1801) is divided based on the division patterns of adjacent regions such as (1802) and (1804).

[0193] exist Figure 18E In this context, the partitioning pattern of the current region (1801) is predicted based on the partitioning pattern of the adjacent region (1804), and the partitioning pattern of the current region (1801) is different from that of the adjacent region (1804).

[0194] exist Figure 18F-18G In this context, the current region (1801) is divided based on the division patterns of adjacent regions such as (1802) and (1804).

[0195] In one example, a learning-based approach is used to predict flexible partitions using the size information of the current region and / or the neighborhood partitioning information of the current region as one or more inputs. The neighborhood partitioning information of the current region indicates the partitioning information of the neighboring regions of the current region.

[0196] In one example, a learning-based approach is used, taking the size information of the current region and / or the neighborhood partitioning information of the current region as one or more inputs, to predict multiple flexible partitioning possibilities. Additional flags can be indicated using signaling to determine which flexible partitioning possibility to use.

[0197] In one respect, flexible partitioning can be predicted based on partitioning information in another region (also known as a reference region). In one example, the other region (or reference region) has been resolved, and partitioning information is available.

[0198] In one example, the reference region is indicated by the displacement vector relative to the current region.

[0199] In one example, the displacement vector is derived based on the pixel similarity between a template of a reference region (e.g., a reference template) and a template of the current region (e.g., a current template).

[0200] In one example, signaling is used to indicate the number of encoded / predicted blocks to indicate how many associated encoded / predicted blocks to parse during the bitstream parsing phase when flexible partitioning cannot be further subdivided.

[0201] In one example, a signaling indicator is used to indicate the end of a continuous encoded / predicted block to indicate whether the end of the flexible partition has been reached.

[0202] In one example, signaling is used to indicate the number of coding units in a flexible partition.

[0203] In one example, signaling is used to indicate the width and / or height of the coding unit in a flexible partition.

[0204] In one example, the horizontal and vertical components of the displacement vector (which can be derived or indicated by signaling) are powers of 2.

[0205] In one example, the magnitudes of the horizontal and vertical components of the displacement vector are N times the width and height of the horizontal and vertical components of the displacement vector, respectively, where N is a non-zero positive integer value.

[0206] In one example, pixel similarity is calculated between the reference region and its adjacent above, left, and upper left regions. The partitioning decision of the block with the highest pixel similarity is reused for the current region.

[0207] In one respect, after deriving the flexible partitioning, the derived flexible partitioning can be further modified (e.g., corrected) using the previous partitioning mechanism.

[0208] Figures 19A-19C An example of a flexible partitioning derivation is shown, which uses prior information to correct the derivation. Figures 19A-19B Description and Figures 17A-17B The description is the same. Figure 19A The partitioning pattern of the current region (1701) is predicted using the directionality of the partitioning information of the neighboring regions of the current region (1701), such as... Figures 17A-17B As shown. Figure 19B The splitting mechanism of the lower right region shows a flexible division (1911) derived from the initial derivation from the adjacent regions of the current region (1701).

[0209] Figure 19B-19C The diagram shows modifications to the initially derived flexible partition (1911) to generate the final derived flexible partition (1921). Figure 19B As shown, the initial derived flexible partition (1911) of the current region (1701) includes the partition pattern (1931) of the lower right region in the current region (1701), and as Figure 19CAs shown, the partitioning pattern (1931) is replaced with no partitioning. Therefore, the final derived flexible partition (1921) is obtained by modifying the initially derived flexible partition (1911) using prior information, thereby improving flexibility.

[0210] In one respect, when deriving flexible partitioning for the current partition, multiple scanning and / or encoding orders can be implemented, and the actual order can be determined by the syntax indicated by signaling, or implicitly (e.g., without signaling) by adjacent block information.

[0211] Figures 20A-20B An example of a flexible partitioning of the current block (2001) with multiple scan orders is shown according to one aspect of the invention. Figure 20A The first scan order of the current block (2001) is shown. The first scan order is the upper left region, the upper right region, the lower left region, and the lower right region, with the upper left region being scanned first and the lower right region being scanned last. Figure 20B The second scan order of the current block (2001) is shown. The second scan order is the upper left region, the lower left region, the upper right region, and the lower right region, with the upper left region being scanned first and the lower right region being scanned last.

[0212] Figure 21 A flowchart illustrating an overview process (2100) according to one aspect of the present invention is shown. The process (2100) can be used in devices such as video decoders. In various aspects, the process (2100) is executed by processing circuitry, such as processing circuitry that performs the functions of the video decoder (110), processing circuitry that performs the functions of the video decoder (210), etc. In some aspects, the process (2100) is implemented in the form of software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (2100). The process begins at (S2101) and continues to (S2110).

[0213] In (S2110), the encoded information of the current block in the current image can be received.

[0214] In one example, the current block is encoded using inter-frame prediction, and the adjacent reconstructed regions are the current templates for the current block.

[0215] In one example, the current block is encoded using one of (i) intra-frame prediction and (ii) BV-based prediction modes.

[0216] In (S2120), it is determined whether to apply a filter to the adjacent reconstructed regions of the current block. The adjacent reconstructed regions are adjacent to the current block and include reconstructed samples from the current image.

[0217] In one example, the filter includes one or more loop filters.

[0218] In one example, one or more loop filters include one or more of the following: (i) a deblocking filter, (ii) a SAO filter, (iii) a bilateral filter, and (iv) an ALF.

[0219] In one example, when an edge is detected in an adjacent reconstructed region, a deblocking filter is determined to be applied. The filter includes a deblocking filter.

[0220] In one example, an edge is detected in an adjacent reconstructed region. When the edge is a vertical edge, the deblocking filter is a horizontal deblocking filter, and when the edge is a horizontal edge, the deblocking filter is a vertical deblocking filter.

[0221] In one example, the decision to apply a filter is based on the encoding information of adjacent reconstructed regions.

[0222] In one example, the encoding information includes partitioning information. The partitioning information indicates that adjacent reconstructed regions comprise reconstructed samples from multiple coded blocks in the current image, forming at least one coded block boundary within the adjacent reconstructed regions, and each of the at least one coded block boundary lies between two adjacent coded blocks. The determination filter is applied only to reconstructed samples adjacent to at least one coded block boundary.

[0223] In one example, the minimum value of the width and height of adjacent reconstructed regions is at least 4 samples.

[0224] In one example, the adjacent reconstruction region includes one or more of the following: the upper adjacent reconstruction region directly above the current block, the left adjacent reconstruction region directly to the left of the current block, the upper left adjacent reconstruction region, the upper right adjacent reconstruction region, and the lower left adjacent reconstruction region.

[0225] In (S2130), a filter is applied to the adjacent reconstructed regions of the current block.

[0226] In (S2140), the current block is reconstructed based on the filtered adjacent reconstruction regions.

[0227] The process then continues to (S2199) and terminates.

[0228] The process (2100) can be adjusted as appropriate. One or more steps in the process (2100) can be modified and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used.

[0229] In one aspect, a method for video coding includes determining whether to apply a filter to neighboring reconstructed regions of a current block in a current image. The neighboring reconstructed regions are adjacent to the current block and include reconstructed samples from the current image. The filter is applied to the neighboring reconstructed regions of the current block, and the current block is encoded based on the filtered neighboring reconstructed regions. In one example, the encoded information of the current block is encoded in the video bitstream.

[0230] The following will describe, for example, references Figure 8-10 The template filtering process described in section 21 offers benefits in template-based prediction modes. For template-based inter-frame prediction modes, a mismatch can occur between the unfiltered current template (e.g., not filtered using one or more loop filters) and the filtered reference template (e.g., filtered using one or more loop filters). In some examples, this mismatch between the two templates reduces the accuracy of template-based prediction modes based on a comparison of the two templates, resulting in poor coding efficiency. The template filtering process for template-based inter-frame prediction modes, for example, filters the current template using one or more filters (e.g., including one or more loop filters), thereby improving the accuracy of template-based prediction modes based on a comparison of the two templates, and thus improving coding efficiency.

[0231] In one aspect, template-based intra-prediction modes or BV-based prediction modes use the current template to predict the current block, for example: (i) generating prediction samples for the current block using reconstructed samples from the current template, (ii) performing template search using reconstructed samples from the current template to find a prediction block with at least one minimum template matching cost, (iii) deriving the intra-prediction mode using reconstructed samples from the current template, and so on. Therefore, by utilizing one or more filters (e.g., one or more loop filters) to filter the current template of the current block, the current block can be predicted more accurately.

[0232] Figure 22 A flowchart illustrating an overview process (2200) according to one aspect of the present invention is shown. Process (2200) can be used in devices such as video encoders. In various aspects, process (2200) is executed by processing circuitry, such as processing circuitry executing the functions of a video encoder (103), processing circuitry executing the functions of a video encoder (303), etc. In some aspects, process (2200) is implemented in the form of software instructions, so that when processing circuitry executes software instructions, processing circuitry executes process (2200). The process begins at (S2201) and continues to (S2210).

[0233] At (S2210), based on one of the segmentation information and prediction information of the adjacent reconstructed regions of the current block in the current image, it is determined whether to indicate one of the following in the video bitstream via signaling: (i) a control flag for the prediction mode used to predict the current block, and (ii) the type of prediction mode. The current block is predicted based on reconstructed samples in the adjacent reconstructed regions adjacent to the current block.

[0234] In one example, based on the partitioning information of the adjacent reconstructed regions of the current block, it is determined whether to use signaling to indicate the control flags of the prediction mode in the video stream.

[0235] In one example, a control flag indicating the prediction mode in the video stream is determined only if the partitioning information indicates that the reconstructed samples in adjacent reconstructed regions come from no more than N different coding blocks.

[0236] In one example, when it is determined that a control flag should be indicated by signaling, the entropy context of the control flag is determined based on partitioning information.

[0237] In (S2220), the current block is encoded based on the prediction pattern and reconstruction samples in adjacent reconstruction regions.

[0238] In (S2230), when it is determined that one of the control flags and the type of prediction mode should be indicated by signaling, that one of the control flags and the type of prediction mode is encoded into the video stream.

[0239] The process then continues to (S2299) and terminates.

[0240] The process (2200) can be adjusted as appropriate. One or more steps in the process (2200) can be modified and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used.

[0241] In one example, it is determined that adjacent reconstruction regions will be used to predict the current block. Based on the partitioning information of adjacent reconstruction regions, signaling is used to indicate the adjacent reconstruction regions. In one example, the reconstructed samples in the adjacent reconstruction regions are constrained to no more than N distinct coded blocks.

[0242] In one example, a list (e.g., a template list) is constructed based on the corresponding partitioning information of multiple adjacent reconstructed regions. This list includes multiple adjacent reconstructed regions adjacent to the current block. The multiple adjacent reconstructed regions include this adjacent reconstructed region. Based on the position of the adjacent reconstructed region in the list, signaling is used to indicate the adjacent reconstructed regions.

[0243] In one aspect, a method for video decoding includes: determining, based on one of the partitioning information and prediction information of adjacent reconstructed regions of a current block in a current image, whether to indicate in the video bitstream one of the following: (i) a control flag for predicting a prediction mode for the current block, and (ii) the type of prediction mode. The current block is predicted based on reconstructed samples in adjacent reconstructed regions adjacent to the current block. The current block is reconstructed based on the prediction mode and the reconstructed samples in the adjacent reconstructed regions. In one example, when one of the control flag and the type of prediction mode is indicated in the video bitstream via signaling, that one of the control flag and the type of prediction mode is decoded.

[0244] The following will describe, for example, references Figure 11-12 The control information signaling indications described in section 22 offer advantages in utilizing encoding tools available to reconstruct sample information from the decoder. For example, whether to signal control flags based on template patterns, whether to signal template types, and / or whether to signal template-based method types is determined based on prediction information (e.g., partitioning information) of the current template for the current block. This differs from related techniques, which signal control flags (e.g., regardless of the partitioning information of the current block) and also signal some syntax representing a specific template type. Therefore, control information signaling indications that depend on prediction information (e.g., partitioning information) of the current template for the current block may be more efficient, for example, under certain conditions, eliminating the need to signal control flags, template types, and / or template-based method types, thereby reducing encoding overhead.

[0245] Figure 23 A flowchart illustrating an overview process (2300) according to one aspect of the present invention is shown. The process (2300) can be used in devices such as video decoders. In various aspects, the process (2300) is executed by processing circuitry, such as processing circuitry executing the functions of video decoder (110), processing circuitry executing the functions of video decoder (210), etc. In some aspects, the process (2300) is implemented in the form of software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (2300). The process begins at (S2301) and continues to (S2310).

[0246] In (S2310), the encoded information of the current region to be divided in the current image is received.

[0247] In (S2320), the partitioning structure of the current region is determined based on one of (i) the size information of the current region and (ii) the partitioning information of another region.

[0248] In one example, the current region has the same size as another region, and the partitioning structure of the current region is determined to be the partitioning structure of the other region.

[0249] In (S2330), the current region is reconstructed based on the established partitioning structure of the current region.

[0250] The process then continues to (S2399) and terminates.

[0251] The process (2300) can be adjusted as appropriate. One or more steps in the process (2300) can be modified and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used.

[0252] In one aspect, a method for video coding includes determining the partitioning structure of the current region based on one of (i) the size information of the current region and (ii) the partitioning information of another region. The current region can then be encoded based on its determined partitioning structure.

[0253] The following will describe, for example, references Figure 16-2 The benefits of flexible partitioning as described in 0 and 23: Unlike fixed partitioning that divides the current region in a specific way, the partitioning mechanism of flexible partitioning (e.g., the way sub-regions are generated) can be derived based on neighborhood information. Therefore, when using flexible partitioning, the partitioning mechanism can vary, making the partitioning mechanism more flexible. Furthermore, since it may not be necessary to indicate the partitioning mode with signaling, signaling overhead can be reduced.

[0254] In one aspect, a method for processing visual media data includes processing a video bitstream of the visual media data according to format rules. For example, the bitstream may be a bitstream decoded / encoded using any of the decoding and / or encoding methods described in this application. The format rules may specify at least one constraint on the bitstream and / or at least one process to be performed by the decoder and / or encoder.

[0255] In one aspect, the video stream includes encoded information for the current block in the current image. Format rules specify whether to apply filters to adjacent reconstructed regions of the current block. Adjacent reconstructed regions are adjacent to the current block and include reconstructed samples from the current image. Format rules specify that filters are applied to adjacent reconstructed regions of the current block, and the current block is reconstructed based on the filtered adjacent reconstructed regions.

[0256] In one aspect, the video stream includes encoded information about the current block in the current image. The format rules specify that, based on one of the partitioning information and prediction information of the adjacent reconstructed regions of the current block in the current image, it should be determined whether to indicate one of the following in the video stream via signaling: (i) a control flag for predicting the prediction mode of the current block, and (ii) the type of prediction mode, whereby the current block is predicted based on reconstructed samples from adjacent reconstructed regions. The format rules specify that the current block is reconstructed based on the prediction mode and reconstructed samples from adjacent reconstructed regions. The format rules specify that when one of the control flag and the type of prediction mode is indicated via signaling, that one of the control flag and the type of prediction mode is decoded.

[0257] In one aspect, the video stream includes encoded information about the current region to be segmented in the current image. The format rules specify that the segmentation structure of the current region is determined based on one of (i) the size information of the current region and (ii) the segmentation information of another region, and the current region is reconstructed based on the determined segmentation structure of the current region.

[0258] The methods, aspects, and / or examples in this invention may be used individually or in any combination in any order. For example, some aspects and / or examples executed by the decoder may be executed by the encoder, and vice versa. Each method (or aspect), encoder, and decoder may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, at least one processor executes a program stored in a non-transitory computer-readable storage medium. The disclosed methods can be used with various codecs (e.g., video codecs), such as those described in this invention.

[0259] The process (or method) described in this invention can be implemented in an image and / or video decoding process or an image and / or video encoding process. The decoding / encoding process can be used in a video decoding device. The decoding / encoding process can be used in a video encoding device. In some examples, the process is executed by processing circuitry, such as processing circuitry that performs the function of a video decoder, processing circuitry that performs the function of a video encoder, etc. In some examples, the process is implemented as software instructions, so the processing circuitry executes the process when it executes the software instructions. In some examples, the process can be implemented on a chip as a hardware process, so the processing circuitry executes the process when it executes the hardware instructions. The process can be appropriately adjusted. Steps in the process described in this invention can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.

[0260] The above-described techniques can be implemented in the form of computer software using computer-readable instructions and physically stored in at least one computer-readable medium. For example, Figure 24 A computer system (2400) suitable for implementing certain aspects of the disclosed subject matter is shown.

[0261] Computer software can be written using any suitable machine code or computer language. This machine code or computer language may be processed by mechanisms such as assembly, compilation, and linking to generate code containing instructions that can be executed directly by at least one computer central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, or other methods.

[0262] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things (IoT) devices.

[0263] Figure 24 The components of the computer system (2400) shown are merely examples and are not intended to imply any limitation on the scope or functionality of computer software used to implement aspects of the present invention. The configuration of the components should also not be construed as having any dependency or requirement on any component or combination of components shown in the example aspects of the computer system (2400).

[0264] The computer system (2400) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from at least one human user via, for example, tactile input (e.g., key presses, swipes, data glove movements), audio input (e.g., voice, clapping sounds), visual input (e.g., gestures), or olfactory input (not depicted). The human-machine interface devices may also be used to acquire media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0265] Human-machine interface input devices may include at least one of the following (only one of each is depicted): keyboard (2401), mouse (2402), touchpad (2403), touch screen (2410), data glove (not shown), joystick (2405), microphone (2406), scanner (2407), and camera (2408).

[0266] The computer system (2400) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback provided by a touchscreen (2410), a data glove (not shown), or a joystick 2405, but there may also be tactile feedback devices that do not act as input devices), audio output devices (e.g., speakers (2409), headphones (not depicted)), visual output devices (e.g., screens (2410), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input functionality, each with or without tactile feedback functionality—some of which may be able to output two-dimensional or more three-dimensional visual outputs through stereoscopic output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), etc.), and printers (not depicted).

[0267] The computer system (2400) may also include user-accessible storage devices and their associated media, such as optical media or similar media (2421) including CD / DVD ROM / RW (2420) equipped with CD / DVD, flash drives (2422), removable hard disk drives or solid-state drives (2423), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD devices such as security dongles (not depicted), etc.

[0268] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the present disclosure does not cover transmission media, carrier waves, or other transient signals. "Computer-readable storage medium" as used herein refers to any tangible medium capable of storing data, including but not limited to random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical disks, magnetic disks, etc., but excluding the transiently propagated signal itself.

[0269] The computer system (2400) may also include an interface (2454) connected to at least one communication network (2455). The network may be, for example, a wireless network, a wired network, or an optical network. The network may also be a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), an in-vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include: local area networks such as Ethernet and wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; in-vehicle and industrial networks including CAN buses, etc. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (2400)) attached to some general-purpose data port or peripheral bus (2449); other networks are typically integrated into the core of the computer system (2400) via a system bus described below (e.g., integrated into a PC computer system via an Ethernet interface, or integrated into a smartphone computer system via a cellular network interface). Using any of these networks, the computer system (2400) can communicate with other entities. This communication can be one-way receiving (e.g., broadcasting TV), one-way transmitting (e.g., CAN bus to certain CAN bus devices), or bidirectional, such as connecting to other computer systems using a local area or wide area digital network. Certain protocols and protocol stacks can be used on each of the networks and network interfaces described above.

[0270] The aforementioned human-machine interface device, user-accessible storage device, and network interface can be attached to the core (2440) of the computer system (2400).

[0271] The core (2440) may include at least one CPU (2441), GPUs (2442), a dedicated programmable processing unit in the form of a Field Programmable Gate Array (FPGA) (2443), a hardware accelerator (2444) for certain tasks, a graphics adapter (2450), etc. These devices, along with read-only memory (ROM) (2445), random-access memory (RAM) (2446), and internal mass storage devices (2447) such as internal non-user-accessible hard disk drives, SSDs, etc., can be connected via a system bus (2448). In some computer systems, the system bus (2448) may be accessed in the form of at least one physical connector, allowing for expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (2448) or connected to the core's system bus (2448) via a peripheral bus (2449). In one example, a screen (2410) may be connected to a graphics adapter (2450). Peripheral bus architectures include PCI, USB, etc.

[0272] CPUs (2441), GPUs (2442), FPGAs (2443), and accelerators (2444) can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (2445) or RAM (2446). Temporary data can also be stored in RAM (2446), while permanent data can be stored, for example, in an internal mass storage device (2447). Fast storage and retrieval of any of the storage devices can be achieved by using a cache memory, which can be closely associated with at least one CPU (2441), GPU (2442), mass storage device (2447), ROM (2445), RAM (2446), etc.

[0273] The computer-readable medium may contain computer code for performing various computer-implemented operations. This medium and computer code may be designed and constructed specifically for the purposes of this invention, or may belong to a class well-known and available to those skilled in the art of computer software.

[0274] For example, but not as a limitation, a computer system having an architecture (2400), particularly a core (2440), can provide functionality generated when one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in at least one tangible computer-readable medium. Such a computer-readable medium can be a medium associated with certain storage devices (e.g., internal core mass storage (2447) or ROM (2445)) that are user-accessible mass storage devices as described above and with the non-volatile nature of the core (2440). Software implementing various aspects of the invention can be stored in such devices and executed by the core (2440). Depending on specific needs, the computer-readable medium may include at least one storage device or chip. The software can cause the core (2440), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute a particular process or a particular portion of a particular process described in this application, including defining data structures stored in RAM (2446) and modifying such data structures according to the process defined by the software. Alternatively or as an alternative, the computer system may provide functionality generated by hard-wired logic or otherwise embodied in circuitry (e.g., an accelerator (2444)) that operates in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may cover logic, and vice versa. Where appropriate, references to computer-readable media may cover circuitry (e.g., integrated circuits) storing software for execution, circuitry embodying logic for execution, or both. This invention covers any suitable combination of hardware and software.

[0275] The terms “at least one of…” or “one of…” as used in this invention are intended to include any one or a combination of the listed elements. For example, mentioning at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A to C are intended to include only A, only B, only C, or any combination thereof. Mentioning one of A or B and one of A and B are intended to include either A or B or (A and B). Where applicable, such as when the elements are not mutually exclusive, the use of “one of…” does not exclude any combination of the listed elements.

[0276] Although several examples of various aspects have been described in this invention, modifications, substitutions, and various alternative equivalents exist that fall within the scope of this invention. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, while not explicitly stated or described herein, embody the principles of the invention and are therefore within its spirit and scope.

[0277] The aforementioned disclosure also covers the following features. These features can be combined in various ways, and are not limited to the following combinations.

[0278] (1) A method for video decoding, the method comprising: receiving encoded information of a current block in a current image; determining whether to apply a filter to an adjacent reconstructed region of the current block, the adjacent reconstructed region being adjacent to the current block and including reconstructed samples in the current image; applying the filter to the adjacent reconstructed region of the current block; and reconstructing the current block based on the filtered adjacent reconstructed region.

[0279] (2) According to the method of feature (1), wherein the filter includes one or more loop filters.

[0280] (3) The method according to feature (2), wherein the one or more loop filters include one or more of the following: (i) a deblocking filter, (ii) a SAO filter, (iii) a bilateral filter and (iv) an ALF.

[0281] (4) A method according to any one of features (1) to (3), wherein determining whether to apply the filter includes: when an edge is detected in the adjacent reconstructed region, determining to apply a deblocking filter, the filter including the deblocking filter.

[0282] (5) According to the method of feature (4), wherein the edge is detected in the adjacent reconstruction region, the deblocking filter is a horizontal deblocking filter when the edge is a vertical edge, and the deblocking filter is a vertical deblocking filter when the edge is a horizontal edge.

[0283] (6) A method based on any one of features (1) to (3), wherein determining whether to apply the filter includes: determining whether to apply the filter based on the coding information of the adjacent reconstructed region.

[0284] (7) The method according to feature (6), wherein the encoding information includes segmentation information; the segmentation information indicates that the adjacent reconstruction region includes reconstruction samples from a plurality of encoded blocks in the current image, the plurality of encoded blocks forming at least one encoded block boundary in the adjacent reconstruction region, each of the at least one encoded block boundary being located between two adjacent encoded blocks in the plurality of encoded blocks; and determining whether to apply the filter includes determining that the filter is applied only to the reconstruction sample adjacent to the at least one encoded block boundary.

[0285] (8) According to any one of the features (1) to (7), wherein the minimum of the width and height of the adjacent reconstructed region is at least 4 samples.

[0286] (9) The method according to any one of features (1) to (8), wherein the adjacent reconstruction region includes one or more of the following: the upper adjacent reconstruction region located directly above the current block, the left adjacent reconstruction region located directly to the left of the current block, the upper left adjacent reconstruction region, the upper right adjacent reconstruction region, and the lower left adjacent reconstruction region.

[0287] (10) The method according to any one of features (1) to (9), wherein the current block is encoded using inter-frame prediction and the adjacent reconstructed region is the current template of the current block.

[0288] (11) The method according to any one of features (1) to (9), wherein the current block is encoded using one of (i) intra-frame prediction and (ii) BV-based prediction modes.

[0289] (12) A method for video coding, the method comprising: determining, based on one of partitioning information and prediction information of adjacent reconstructed regions of a current block in a current image, whether to signal in a video bitstream one of: (i) a control flag for predicting a prediction mode of the current block, and (ii) a type of the prediction mode, the current block being predicted based on reconstructed samples in the adjacent reconstructed regions adjacent to the current block; encoding the current block according to the prediction mode and the reconstructed samples in the adjacent reconstructed regions; and when it is determined that the one of the control flag and the type of the prediction mode should be signaled, encoding the one of the control flag and the type of the prediction mode into the video bitstream.

[0290] (13) According to the method of feature (12), wherein determining includes: determining whether to indicate the control flag of the prediction mode in the video stream by signaling based on the partition information of the adjacent reconstructed region of the current block.

[0291] (14) According to the method of feature (13), wherein determining includes: determining the control flag to indicate the prediction mode in the video stream by signaling only when the partition information indicates that the reconstructed sample in the adjacent reconstructed region comes from no more than N different coding blocks.

[0292] (15) The method according to feature (13) or (14), wherein the method further includes determining the entropy context of the control flag based on the partition information when it is determined that the control flag should be indicated by signaling.

[0293] (16) The method according to feature (12), wherein the method includes determining to use the adjacent reconstruction region to predict the current block; and indicating the adjacent reconstruction region by signaling based on the partition information of the adjacent reconstruction region.

[0294] (17) According to the method of feature (16), the reconstructed sample in the adjacent reconstructed region is constrained to no more than N different coding blocks.

[0295] (18) The method according to feature (16), wherein the method includes constructing a list based on corresponding partition information of a plurality of adjacent reconstruction regions, the list including the plurality of adjacent reconstruction regions adjacent to the current block, the plurality of adjacent reconstruction regions including the adjacent reconstruction regions; and signaling indicating the adjacent reconstruction regions based on the position of the adjacent reconstruction regions in the list.

[0296] (19) A method for video decoding, the method comprising: receiving encoded information of a current region to be segmented in a current image; determining a segmentation structure of the current region based on one of (i) size information of the current region and (ii) segmentation information of another region; and reconstructing the current region based on the determined segmentation structure of the current region.

[0297] (20) According to the method of feature (19), wherein the current region and the other region have the same size; and determining includes determining the partitioning structure of the current region as the partitioning structure of the other region.

[0298] (21) An apparatus for decoding, comprising a processing circuit configured to perform a method according to any one of features (1) to (11).

[0299] (22) An apparatus for encoding, comprising a processing circuit configured to perform a method according to any one of features (12) to (18).

[0300] (23) An apparatus for decoding, comprising a processing circuit configured to perform a method according to any one of features (19) to (20).

[0301] (24) A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform a method according to any one of features (1) to (20).

Claims

1. A method for video decoding, characterized in that, The method includes: Receive the encoded information of the current block in the current image; Determine whether to apply a filter to the adjacent reconstructed regions of the current block, the adjacent reconstructed regions being adjacent to the current block and including reconstructed samples in the current image; The filter is applied to the adjacent reconstructed regions of the current block; and The current block is reconstructed based on the filtered adjacent reconstruction regions.

2. The method according to claim 1, characterized in that, The filter includes one or more loop filters.

3. The method according to claim 2, characterized in that, The one or more loop filters include one or more of the following: (i) a deblocking filter, (ii) a SAO filter, (iii) a bilateral filter, and (iv) an ALF.

4. The method according to any one of claims 1 to 2, characterized in that, Determining whether to apply the filter includes: When an edge is detected in the adjacent reconstructed region, it is determined that a deblocking filter, including the deblocking filter, will be applied.

5. The method according to claim 4, characterized in that, The edge was detected in the adjacent reconstructed region. When the edge is a vertical edge, the deblocking filter is a horizontal deblocking filter, and When the edge is a horizontal edge, the deblocking filter is a vertical deblocking filter.

6. The method according to any one of claims 1 to 3, characterized in that, Determining whether to apply the filter includes: Whether to apply the filter is determined based on the encoding information of the adjacent reconstructed regions.

7. The method according to claim 6, characterized in that, The encoding information includes partitioning information; The partitioning information indicates that the adjacent reconstruction region includes reconstructed samples from multiple coded blocks in the current image, the multiple coded blocks forming at least one coded block boundary in the adjacent reconstruction region, and each of the at least one coded block boundary being located between two adjacent coded blocks among the multiple coded blocks; and Determining whether to apply the filter includes determining that the filter is applied only to the reconstructed samples adjacent to the boundary of the at least one coded block.

8. The method according to any one of claims 1 to 7, characterized in that, The current block is encoded using inter-frame prediction, and the adjacent reconstructed regions are the current templates for the current block.

9. The method according to any one of claims 1 to 7, characterized in that, The current block is encoded using one of (i) intra-frame prediction and (ii) BV-based prediction modes.

10. A method for video encoding, characterized in that, The method includes: Based on one of the partitioning information and prediction information of the adjacent reconstructed regions of the current block in the current image, determine whether to indicate one of the following in the video stream by signaling: (i) a control flag for predicting the prediction mode of the current block, and (ii) the type of the prediction mode, wherein the current block is predicted based on reconstructed samples in the adjacent reconstructed regions adjacent to the current block; The current block is encoded based on the prediction pattern and the reconstructed samples in the adjacent reconstructed regions; and When it is determined that signaling is needed to indicate one of the control flags and the type of the prediction mode, the control flags and the type of the prediction mode are encoded into the video stream.

11. The method according to claim 10, characterized in that, The determination includes: Based on the partitioning information of the adjacent reconstructed regions of the current block, determine whether to indicate the control flag of the prediction mode in the video stream using signaling.

12. The method according to claim 11, characterized in that, The determination includes: The control flag indicating the prediction mode in the video stream is determined only when the partitioning information indicates that the reconstructed samples in the adjacent reconstructed regions come from no more than N different coding blocks.

13. The method according to claim 11 or 12, characterized in that, Also includes: When it is determined that the control flag needs to be indicated by signaling, the entropy context of the control flag is determined based on the partitioning information.

14. The method according to claim 10, characterized in that, Also includes: Determine whether to use the adjacent reconstruction region to predict the current block; as well as Based on the division information of the adjacent reconstruction areas, the adjacent reconstruction areas are indicated by signaling.

15. A method for video decoding, characterized in that, The method includes: Receive the encoded information of the current region to be divided in the current image; Based on one of (i) the size information of the current region and (ii) the partitioning information of another region, determine the partitioning structure of the current region; and The current region is reconstructed based on its established partitioning structure.