Video decoding method, apparatus and storage medium
By employing an adaptive hybrid processing approach that combines geometric partitioning and template matching techniques in video encoding, the problem of low efficiency in video block encoding is solved, resulting in more efficient video encoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2026-03-10
AI Technical Summary
Existing video coding technologies struggle to effectively utilize geometric partitioning patterns for efficient sample prediction and reconstruction when processing video blocks, resulting in low coding efficiency.
Adaptive blending of video blocks is performed using Geometric Partitioning (GPM). The width candidate list is reordered using template matching technology, a suitable width candidate is selected and the blending region is determined, and the samples in the current block are reconstructed using adaptive blending technology.
It improves the efficiency and quality of video encoding, reduces the amount of data, and enhances the encoding effect.
Smart Images

Figure CN117280687B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of priority to U.S. Patent Application No. 18 / 136,091, filed April 18, 2023, entitled “ADAPTIVE BLENDING FOR GEOMETRIC PARTITION MODE (GPM),” which in turn claims the benefit of priority to U.S. Provisional Application No. 63 / 332,741, filed April 20, 2022, entitled “ADAPTIVE BLENDING FOR GEOMETRIC PARTITION MODE (GPM),” the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] This disclosure describes embodiments that are generally related to video coding. Background Technology
[0004] The background description provided herein is for the purpose of presenting the general content of this disclosure. The extent of the work of the currently named inventors described in the background section and various aspects of this specification does not indicate that it was prior art at the time of filing of this application, nor is it expressly or implied that it was acknowledged as prior art to this disclosure.
[0005] Image / video compression enables the transfer of image / video files with minimal quality degradation between different devices, storage devices, and networks. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress images based on spatial redundancy. For example, intra-frame prediction can use reference data from the reconstructed current image to predict samples. In another example, a video codec can use a technique called inter-frame prediction, which can compress images based on temporal redundancy. For example, inter-frame prediction can predict samples in the current image from a previously reconstructed image with motion compensation. Motion compensation is typically indicated by motion vectors (MV). Summary of the Invention
[0006] This disclosure provides methods and apparatus for video encoding or decoding. In some examples, the apparatus for video decoding includes processing circuitry. The processing circuitry receives a video bitstream comprising a current block in a current image. It decodes the value of a syntax element associated with the current block in the current image, the syntax element indicating whether the current block is encoded using Geometric Partition Mode (GPM) along partition edges intersecting the current block. In response to the syntax element value indicating that the current block is encoded using GPM and satisfies adaptive blending conditions, it reorders width candidates in a width candidate list using template matching (TM) based on a current template of the current block and a reference template corresponding to the respective width candidate. The processing circuitry can select a width candidate from the reordered width candidate list and determine the width of a blending region based on the selected width candidate. The blending region surrounds a partition edge and is defined by boundaries on either side of the partition edge. The boundaries are parallel to the partition edge, and the width of the blending region is measured perpendicular to the partition edge. The processing circuitry can determine the blending region based on its width and reconstruct samples within the blending region in the current block by applying adaptive blending using the determined blending region.
[0007] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a method for video decoding. Attached Figure Description
[0008] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0009] Figure 1 This is a schematic diagram of an exemplary block diagram of a communication system (100).
[0010] Figure 2 This is a schematic diagram of an exemplary block diagram of a decoder.
[0011] Figure 3 This is a schematic diagram of an exemplary block diagram of an encoder.
[0012] Figure 4A A predetermined number of angles from 0° to 360° are shown in the geometric partition pattern (GPM) according to an embodiment of the present disclosure.
[0013] Figure 4B Multiple partition edges corresponding to the angle of GPM are shown according to embodiments of the present disclosure.
[0014] Figure 4CAn exemplary GPM mixing process applied to the current block according to an embodiment of this disclosure is shown.
[0015] Figure 5A The partition edges of the GPM segmentation mode applied to the current block according to an embodiment of this disclosure are shown.
[0016] Figure 5B The GPM mixing process applied to obtain a reference template according to an embodiment of the present disclosure is illustrated.
[0017] Figure 6 The weights are shown under different mixing region sizes according to the embodiments. and displacement The relationship between them.
[0018] Figures 7A-7B An exemplary adaptive blending process using template matching (TM) is shown according to an embodiment of this disclosure.
[0019] Figures 8A-8B Exemplary mixed areas not fully covered within the template used for template matching, according to embodiments of this disclosure.
[0020] Figure 9 An example of obtaining the sample value difference across partition edges according to an embodiment of the present disclosure is shown.
[0021] Figure 10 A flowchart outlining the encoding process according to an embodiment of this disclosure is shown.
[0022] Figure 11 A flowchart outlining the decoding process according to an embodiment of the present disclosure is shown.
[0023] Figure 12 A flowchart outlining the encoding process according to an embodiment of this disclosure is shown.
[0024] Figure 13 A flowchart outlining the decoding process according to an embodiment of the present disclosure is shown.
[0025] Figure 14 This is a schematic diagram of a computer system according to an embodiment. Detailed Implementation
[0026] Figure 1 Block diagrams of some examples of video processing systems (100) are shown. The video processing system (100) is an example of an application of the disclosed subject matter, video encoder, and video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0027] The video processing system (100) includes an acquisition subsystem (113) that may include a video source (101) such as a digital camera, which creates, for example, an uncompressed video image stream (102). In one example, the video image stream (102) includes samples taken by a digital camera. The video image stream (102), depicted as a thick line to emphasize its high data volume, may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or encoded video stream), depicted as a thin line to emphasize its lower data volume, may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 Client subsystems (106) and (108) can access a streaming server (105) to retrieve copies (107) and (109) of encoded video data (104). Client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and produces an output video picture stream (111) that can be displayed on a display (112) (e.g., a screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video stream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Next Generation Video Coding (VVC). The disclosed topics can be used in the context of VVC.
[0028] It should be noted that electronic devices (120) and (130) may include other components (not shown). For example, electronic device (120) may include a video decoder (not shown), and electronic device (130) may also include a video encoder (not shown).
[0029] Figure 2 An exemplary block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., receiving circuitry). The video decoder (210) may be used in place of... Figure 1 The example video decoder (110).
[0030] The receiver (231) may receive one or more encoded video sequences to be decoded by the video decoder (210). In one embodiment, one encoded video sequence may be received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequences may be received from a channel (201), which may be a hardware / software link leading to a storage device storing the encoded video data. The receiver (231) may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective user entities (not depicted). The receiver (231) may separate the encoded video sequences from other data. To prevent network jitter, a buffer (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer (215) is part of the video decoder (210). In other applications, the buffer (215) may be located outside the video decoder (210) (not depicted). In other applications, a buffer memory (not depicted) may be provided externally to the video decoder (210) to, for example, prevent network jitter, and another additional buffer memory (215) may be provided internally to the video decoder (210) to, for example, handle broadcast timing. When the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer memory (215) may not be necessary, or it may be made smaller. For use on packet-switched networks such as the Internet, a buffer memory (215) may be required, which may be relatively large, advantageously having an adaptive size, and may be implemented at least partially in the operating system or in a similar component (not depicted) external to the video decoder (210).
[0031] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (210), and potential information for controlling a presentation device such as a presentation device (212) (e.g., a display screen), which is not integral to the electronic device (230) but may be coupled to it. Figure 2As shown. The control information used for the presentation device can be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (220) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be based on video coding techniques or standards and can follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract a subgroup parameter set of at least one subgroup of pixels in the subgroups of pixels for use in the video decoder based on at least one parameter corresponding to a group. Subgroups may include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (220) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0032] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0033] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by the parser (220) through subgroup control information parsed from the encoded video sequence. For clarity, the flow of such subgroup control information between the parser (220) and the various units described below is not depicted.
[0034] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into multiple functional units as described below.
[0035] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantization transform coefficients as symbols (221) from the parser (220) and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block containing sample values, which can be input into the aggregator (255).
[0036] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use prediction information from a previously reconstructed image, but can use prediction information from a previously reconstructed portion of the current image. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from the current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some cases, the aggregator (255) adds the prediction information generated by the intra-picture prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.
[0037] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to blocks of inter-frame coding and potential motion compensation. In this case, the motion compensation prediction unit (253) may access the reference image memory (257) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (221) belonging to the block, these samples may be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The extraction of prediction samples by the motion compensation prediction unit (253) from the address in the reference image memory (257) may be controlled by motion vectors, which may be provided to the motion compensation prediction unit (253) in the form of symbols (221), which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values extracted from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0038] The output samples of the aggregator (255) can be subjected to various loop filtering techniques in the loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video stream) and used as symbols (221) from the parser (220) for the loop filter unit (256). Video compression techniques may also respond to metadata obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0039] The output of the loop filter unit (256) can be a sample stream that can be output to the presentation device (212) and stored in the reference image memory (257) for future inter-frame image prediction.
[0040] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded images corresponding to the current image have been fully reconstructed and the encoded images (by, for example, the parser (220)) are identified as reference images, the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.
[0041] The video decoder (210) can perform decoding operations according to a predetermined video compression technique as specified in a standard such as ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technique or standard as the only tools available under that configuration file. For compliance, it may also be necessary that the complexity of the encoded video sequence be within the limits defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.
[0042] In one embodiment, the receiver (231) may receive additional (redundant) data when receiving encoded video. This additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0043] Figure 3 A block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) can be used in place of... Figure 1 The example video encoder (103).
[0044] The video encoder (303) can obtain data from the video source (301) (not...). Figure 3In one example, an electronic device (320) receives video samples, and the video source (301) can capture video images that will be encoded by a video encoder (303). In another example, the video source (301) is part of the electronic device (320).
[0045] A video source (301) can provide a sequence of source video samples to be encoded by a video encoder (303) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601YCrCB, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera that captures local image information as a video sequence. The video data can be provided as multiple individual images, which are given motion when viewed in sequence. The images themselves can be reconstructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc., used. The relationship between pixels and samples will be readily understood by those skilled in the art. The following focuses on describing samples.
[0046] According to one embodiment, the video encoder (303) can encode and compress images of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of the controller (350). In some embodiments, the controller (350) controls and is functionally coupled to other functional units described below. For clarity, the coupling is not depicted in the figures. Parameters set by the controller (350) may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of images (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured to have other suitable functions related to the video encoder (303) optimized for a particular system design.
[0047] In some embodiments, the video encoder (303) is configured to operate within an encoding loop. As an oversimplification, in one example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (333) embedded within the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder can also create sample data. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since decoding of the symbol stream produces bit-accurate results independent of the decoder's location (local or remote), the contents of the reference image memory (334) are also bit-accurately corresponding between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related techniques.
[0048] The operation of the "local" decoder (333) can be combined with, for example, those already mentioned above. Figure 2 The video decoder (210) described in detail is the same as the "remote" decoder. However, a further brief reference is provided. Figure 2 Since symbols are available and the entropy encoder (345) and parser (220) can encode / decode symbols into encoded video sequences without loss, the entropy decoding portion of the video decoder (210), which includes the buffer (215) and parser (220), may not be fully implemented in the local decoder (333).
[0049] In one embodiment, any decoder technique other than parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operation. The description of encoder techniques can be simplified because encoder techniques are inverses of the fully described decoder techniques. Only in certain areas is a more detailed description necessary and provided below.
[0050] During operation, in some examples, the source encoder (330) may perform motion-compensated predictive coding, which predictively encodes the input image by referencing one or more previously encoded images from the video sequence designated as "reference images." In this way, the encoding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.
[0051] The local video decoder (333) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (330). The operation of the encoding engine (332) can advantageously be lossy processing. When the encoded video data can be decoded by the video decoder (333), Figure 3 When decoded in (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process, which can be performed by the video decoder on the reference picture, and allows the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference picture that shares the same content (no transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder.
[0052] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image. The predictor (335) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (334).
[0053] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0054] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by the various functional units using techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0055] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which may be a hardware / software link to a storage device capable of storing the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0056] The controller (350) manages the operation of the video encoder (303). During encoding, the controller (350) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types:
[0057] An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (“IDR”) pictures. Those skilled in the art will understand these variations of I-pictures and their corresponding applications and characteristics.
[0058] A predictive picture (P-picture) can be a picture that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and reference index to predict sample values for each block.
[0059] A bidirectional predictive picture (B-picture) can be a picture that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values for each block. Similarly, multiple predictive pictures can be used to reconstruct a single block using more than two reference pictures and associated metadata.
[0060] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignment of the corresponding images applied to the block. For example, blocks of an I-image can be non-predictively coded, or blocks of an I-image can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. Blocks of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.
[0061] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard such as ITU-T H.265 Recommendation. In operation, the video encoder (303) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0062] In one embodiment, the transmitter (340) may transmit additional data while transmitting encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0063] The captured video can be presented as multiple source images (video images) in a time-series format. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In one example, a specific image being encoded / decoded is divided into blocks, referred to as the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference image.
[0064] In some embodiments, bidirectional prediction techniques can be used for inter-frame image prediction. According to this bidirectional prediction technique, two reference images are used, such as a first reference image and a second reference image that precede the current image in the video in decoding order (but may be past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. The block can be predicted using a combination of the first and second reference blocks.
[0065] In addition, merging mode techniques can be used for inter-frame image prediction to improve coding efficiency.
[0066] According to some embodiments of this disclosure, predictions such as inter-frame picture prediction and intra-frame picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a video picture sequence are divided into Coding Tree Units (CTUs) for compression, with CTUs in the pictures having the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU comprises three Coding Tree Blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively partitioned into one or more Coding Units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be partitioned into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. Based on temporal and / or spatial predictability, the CU is divided into one or more Prediction Units (PUs). Typically, each PU comprises a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in the encoding (encoding / decoding) is performed on a per-prediction-block basis. Using a luminance prediction block as an example, the prediction block comprises a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0067] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more processors that execute software instructions.
[0068] Various inter-frame prediction modes can be used in VVC. For inter-frame prediction CUs, motion parameters may include MV(s), one or more reference picture indices, reference picture list usage indexes, and additional information about certain coded features used to generate inter-frame prediction samples. Motion parameters can be signaled explicitly or implicitly. When a CU is encoded in skip mode, the CU can be associated with a PU and may not have significant residual coefficients, no coded motion vector increments, MV differences (e.g., MVD), or reference picture indices. A merge mode can be specified in which the motion parameters of the current CU are obtained from neighboring CUs, including spatial and / or temporal candidates, and optionally, additional information introduced in VVC. This merge mode can be applied to inter-frame prediction CUs, not just skip mode. In one example, an alternative to the merge mode is explicit transmission of motion parameters, in which MV(s), the corresponding reference picture index and reference picture list usage flag for each reference picture list, and other information are explicitly signaled to each CU.
[0069] In one embodiment (e.g., in VVC), the VVC Test Model (VTM) reference software includes one or more refined inter-frame prediction coding tools, including: extended merge prediction, merged motion vector difference (MMVD) mode, adaptive motion vector prediction mode with symmetric MVD signaling, affine motion compensation prediction, sub-block-based temporal motion vector prediction (SbTMVP), combined inter-frame and intra-frame prediction (CIIP), geometric partitioning mode (GPM), etc. Inter-frame prediction and related methods will be described in detail below.
[0070] Geometric Partitioning Mode (GPM) can be used for intra-frame prediction, such as in VVC. In one example, GPM is applied only to CUs of 8×8 or larger. A CU-level flag can be used to signal GPM as a type of merge mode, while other merge modes include regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. When using GPM, the CU can be partitioned into two geometrically shaped partitions (also called geometric partitions or partitions) using a predetermined number (e.g., 64) of different partitioning schemes or one of the GPM partitioning modes. A geometric partitioning index (or GPM partitioning mode index) can be used to indicate the partitioning scheme or GPM partitioning mode, such as one of 64 different partitioning schemes. In one example, the CU is uniformly partitioned into two geometrically shaped partitions. Different partitioning schemes can be distinguished by a predetermined number (e.g., 24) of angles (e.g., non-uniform quantization between 0 and 360°) and up to a predetermined number (e.g., 4) of partition edges relative to the center of the CU, such as... Figure 4A and Figure 4B As shown.
[0071] Figure 4A The illustration shows a predetermined number (or multiple angles) distributed from 0° to 360° according to embodiments of the present disclosure, such as the 24 supported angles in a VVC. These multiple angles can be indicated by corresponding angle indices 0-23. Figure 4B The illustration shows, according to an embodiment of the present disclosure, multiple partition edges corresponding to one of a plurality of angles (e.g., four partition edges indicated by corresponding indices (idx) 0-3), for example, possible partition edges supported for angle index 3. A set of GPM segmentation patterns that can be used to encode blocks can be based on Figure 4A and Figure 4B The appropriate combination of multiple angles and associated partition edges is shown.
[0072] In one embodiment, each geometric partition in the CU is predicted inter-frame using the corresponding motion information. In one example, only unidirectional prediction is allowed for each partition; for example, each partition has a corresponding MV and a corresponding reference index. Unidirectional prediction motion constraints can be applied to ensure that only two motion-compensated predictions are used for each CU, which is the same as when applying bidirectional prediction to the entire CU. In some examples, bidirectional prediction is applied to partitions in the CU.
[0073] In one embodiment, inter-frame prediction is used for geometric partitioning in the CU. In another embodiment, inter-frame prediction and another prediction (e.g., intra-frame prediction) are used for geometric partitioning in the CU.
[0074] If GPM is used for CU, signaling can be sent to inform information indicating the geometric partition index of the CU and the prediction information for two geometric partitions in the CU. In one example, if each geometric partition in the CU is inter-predicted, the prediction information includes two merge indices. If inter-prediction and intra-prediction are used for the geometric partitions in the CU, the prediction information may include a merge index and an index indicating the intra-prediction mode.
[0075] The number of maximum GPM candidate sizes can be explicitly signaled, for example, at the slice level, and a syntax binary can be specified for the GPM merge index.
[0076] After predicting each of the two geometric partitions, a blending process with adaptive weights (or GPM blending) can be used to adjust the sample values in the blended region (including samples along the partition edges). The size of the blended region can be determined by the blending intensity or, for example,... Figure 4C The width of the blending region, θ (also referred to as τ in some examples), is used to indicate this. The width of the blending region, θ, can be the width of the blending region measured perpendicular to the partition edge (452).
[0077] In one embodiment, the blending region width θ is fixed for CUs with different content, such as natural content, screen content, a mixture of natural content and screen content, etc. In another embodiment, the blending region width θ is adaptive, for example, by selecting the blending region width θ from a predefined list of width candidates. Figure 6 As shown.
[0078] Figure 4C An exemplary GPM mixing process applied to a CU (450) according to an embodiment of the present disclosure is illustrated. The CU (450) may be divided into geometric partitions (461)-(462) by partition edges (452). In one example, a first prediction mode and a second prediction mode are applied to predict samples in the CU (450) as P0 and P1, respectively. The first prediction mode and the second prediction mode may include suitable prediction modes, such as inter-frame prediction mode, intra-frame prediction mode, intra-block copy (IBC) mode, etc.
[0079] The partition edge (452) can be indexed at the corresponding angle (e.g., Figure 4A The angular orientation is 10 or 22 in the angular index. The partition edge (452) and the angular index can correspond to the geometric partition index of CU (450).
[0080] A blending region or blending area (451) may include samples along the partition edge (452). A blending region (451) may include samples within a distance (or blending region width) θ from the partition edge (452). The boundaries (471)-(472) of the blending region (451) are parallel to the partition edge (452) and at a distance θ from the partition edge (452). In one example, the blending region (451) includes a first blending region (465) and a second blending region (466) separated by the partition edge (452). The first blending region (465) may be within the partition (461), and the second blending region (466) may be within the partition (462).
[0081] The samples in CU(450) can be determined based on a mixture process, such as a weighted sum P.
[0082] P = (1-W)×P1 + W×P0 Equation 1
[0083] P0 and P1 represent the predicted values of the samples based on the first and second prediction modes, respectively. The values can be based on the displacement d(x) of the samples from the partition edge (452) to the CU (450). c y c Determine the weights (W) of the samples in CU(450).
[0084] A blending mask can be applied to CU(450), and the weights (or weight values) ω in the blending mask xc,yc This can be given by the ramp function below. In one example, W is ω. xc,yc / 8.
[0085]
[0086] When sample (x) c y c Located in partition (462) and outside the second mixing region (466), displacement d(x) c y c ) less than or equal to -θ, ω xc,yc W is 0. Therefore, samples located outside the mixed region (451) in partition (462) can be predicted as P1 based on the second prediction mode.
[0087] When sample (x) c y c Located in partition (461) and outside the first mixing region (465), displacement d(x) c y c ) greater than or equal to θ, ω xc,yc The value is 8, and W is 1. Therefore, samples in partition (461) located outside the mixing region (451) can be predicted as P0 based on the first prediction mode. In one example, mixing is not used for samples outside the mixing region (451).
[0088] When sample (x) c y c Located in the mixed region (451), displacement d(x) c y c ) lies between -θ and θ, ω xc,yc Based on displacement d(x) c y c The sample in the mixed region (451) can be predicted as a weighted sum of P0 and P1 as described in Equation 1.
[0089] In one example, θ is fixed at 2 pixels (pel), for example, in the current VVC design, the ramp function ω xc,yc It can be quantized as ω m,n .
[0090] ω m,n =Clip3(0, 8, (d(m,n)+32+4)>>3) Equation 3
[0091] In one example, d(m, n) is 16 × d(x c y c ).
[0092] The resulting mixture P (e.g., predicted sample values of CU(450)) can include the predicted signal of CU(450) (e.g., the entire CU(450)). Transformation and quantization processing can be applied to CU(450) as in other prediction modes. The motion field of CU(450) predicted using GPM can be stored.
[0093] GPM can be used in addition to VVC, for example. For instance, template matching (TM) can be used to refine motion on the decoder side to improve compression efficiency. In TM mode, motion is refined by constructing templates from the left-side adjacent reconstructed samples and the top-side adjacent reconstructed samples, and finding the closest match between the template in the current image and the template in the reference image (or reference frame).
[0094] In one example, the TM pattern can be applied to the GPM. When encoding the CU in the GPM, each motion of the geometric partition can determine whether to use TM for refinement. When TM is selected, a template is constructed using the left and top neighbor samples, and the motion is then refined by finding the best match between the current template and a reference region in the reference frame that has the same template pattern. The refined motion can be used to perform motion compensation on the geometric partition, and the refined motion is stored in the motion field.
[0095] Template matching (TM)-based reordering can be applied to GPM partition patterns. For example, the signaling cost of the GPM partition pattern index can be reduced by using TM costs to reorder the index. A reordering method for GPM partition patterns may include a two-step process performed after generating corresponding reference templates for the two GPM partitions in the CU, as described below.
[0096] Reference templates for two GPM partitions corresponding to a given GPM partitioning pattern can be mixed using appropriate weights (e.g., W is 0 or 1) to obtain a reference template (or a mixed reference template) corresponding to the given GPM partitioning pattern. A predetermined number (e.g., 64) of mixed reference templates can be determined for a given GPM partitioning pattern, and the TM generation value (or TM cost) of the corresponding mixed reference template can be determined (e.g., calculated).
[0097] GPM partitioning patterns can be reordered in ascending order based on their respective TM generation values. In one example, the TM generation values of the corresponding M1 (e.g., 64) GPM partitioning patterns are calculated, and the best M2 (e.g., 32) GPM partitioning patterns (e.g., the 32 GPM partitioning patterns with the minimum TM cost values) are marked as available GPM partitioning patterns, from which GPM partitioning patterns can be selected for the CU.
[0098] After reordering the GPM partitioning patterns based on the TM cost, a signaling instruction can be sent to the index to indicate the GPM partitioning pattern selected from the reordered GPM partitioning patterns (e.g., the 32 GPM partitioning patterns with the lowest TM cost values) for predicting the CU. In one example, the index is encoded using a Golomb-Rice code (division by 4).
[0099] In the hybrid processing, the corresponding weights used to obtain the hybrid reference template can be calculated using, for example, a similar GPM weight derivation process described in Equations 1-3, except for, for example, Figure 5A and Figure 5B The weight ω depends on the displacement from the corresponding reference sample to the partition edge. xc,yc Aside from the case where it maps to two values, such as 0 and 8 (corresponding to W being 0 or 1).
[0100] Figure 5A A partition edge (502) of a GPM partitioning mode applied to the current block (501) according to an embodiment of the present disclosure is shown. The partition edge (502) can divide the current block (501) into partitions or GPM partitions (511) and (512). The partition edge (502) can extend into the current template (521) of the current block (501). The current template (521) can include an upper template (522) and a left template (523). The partition edge (502) can divide the current template (521) into a first template (561) located on one side (e.g., the right side) of the partition edge (502) and a second template (562) located on the other side (e.g., the left side) of the partition edge (502).
[0101] Figure 5B The embodiments of the present disclosure are shown for obtaining and Figure 5A The blending process is applied to the reference template (or blended reference template) (531) corresponding to the current template (521). In one example, Figure 5B Mixed processing and Figure 4C The mixing process is the same, except... Figure 5BThe width θ of the mixing region is outside 0. The reference template (531) may include a first reference template (541) corresponding to the first template (561) and a second reference template (542) corresponding to the second template (562). The boundaries (571)-(572) of the first reference template (541) and the second reference template (542) are parallel to the partition edge (502). In one example, the first reference template (541) may be determined based on the first template (561) using the first prediction mode, and the second reference template (542) may be determined based on the second template (562) using the second prediction mode. In one example, the sample in the first reference template (541) is predicted as (1-W)×P1+W×P0, where W is 1 (e.g., ω xc,yc Mapped to 8), the samples in the second reference template (542) are predicted as (1-W)×P1+W×P0, where W is 0 (e.g., ω xc,yc (Mapped to 0). Figure 5B In the example shown, the first prediction mode and the second prediction mode are inter-frame predictions with two different prediction modes (e.g., the first MV and the second MV). In one example, the first MV is associated with a first reference image, and the second MV is associated with a second reference image. The first reference image and the second reference image can be the same or different. If an intra-frame prediction or IBC mode is used as one of the first and second prediction modes, a reference template (531) can be predicted similarly. After determining the reference template (531), the TM cost corresponding to the GPM segmentation mode is determined. The GPM segmentation modes can be reordered based on the determined TM cost. The GPM segmentation mode to be applied to the current block (501) can be selected based on the reordered GPM segmentation modes.
[0102] GPM can be applied in conjunction with TM. GPM can be applied in conjunction with MMVD. GPM can be applied in conjunction with inter-frame prediction and intra-frame prediction. In some examples, a fixed blending region with a fixed θ is used to blend two different partitions, such as... Figure 4C And the mixed weights described in Equations 1-3 (e.g., ω) xc,yc Alternatively, the derivation (W) may not be optimal for different types of content (e.g., natural content and screen content). In one example, screen content video includes strong textures and sharp edges, while narrow blending regions (e.g., with relatively small θ) can preferably preserve edge information in screen content video.
[0103] Adaptive blending can be applied to GPM. Two different adaptive blending methods can be used. The blending region width θ for signaling notification can differ between these two methods. A series of blending region sizes (or predefined blending region width candidates) can be used to select the GPM blending region from the two different adaptive blending methods. This series of blending region sizes can include multiple blending region sizes, such as {θ1, θ2, θ3}. In one example, θ1 = θ / 4, θ2 = θ, and θ3 = 4θ.
[0104] Figure 6 The weights under different mixing area sizes according to the embodiments are shown. and displacement The relationship between them. In one example, Equation 2 describes the weights under different mixing region sizes. and displacement The relationship between them.
[0105] refer to Figure 4C and Figure 6 The CU (e.g., (450)) can be divided into partition A (e.g., (462)) and partition B (e.g., (461)). The blending region size can be chosen from θ / 4, θ, and 4θ. When the blending region size is θ, the GPM blending region can be between -θ and +θ. Weights and displacement The relationship or ramp function can be indicated by curve (612). When the mixing region size is θ / 4, the GPM mixing region can be between -θ / 4 and +θ / 4. Weights and displacement The relationship between them can be indicated by curve (611). When the mixing region size is 4θ, the GPM mixing region can be between -4θ and +4θ. Weights and displacement The relationship between them can be indicated by curve (613). Curve (611)-curve (613) can represent the ramp function under different mixing region sizes.
[0106] refer to Figure 6 And Equation 2, the maximum value of the weight (ω) max The value is 8. The slope of the curve (611)-curve (613) within the corresponding GPM mixing region can depend on the maximum value ω of the weight. maxFor example, the slope of curve (611)-curve (613) within the corresponding GPM mixing region increases (e.g., linearly) as the maximum value of the weight increases. The slope of curve (611)-curve (613) within the corresponding GPM mixing region can be 16 / θ, 4 / θ, and 1 / θ corresponding to mixing region sizes of θ / 4, θ, and 4θ, respectively. In one example, θ = 2, then the slope of curve (611)-curve (613) within the corresponding GPM mixing region can be 8, 2, and 1 / 2 corresponding to mixing region sizes of 1 / 2, 2, and 4, respectively.
[0107] In adaptive blending methods, the selectable width or blending region size can be limited based on the current block size to maximize coding performance. In one example, when the shorter side of the current block is equal to or less than 16 samples, only two narrower widths can be chosen, such as (θ, θ / 4) (or (τ, τ / 4)). Otherwise, when the shorter side of the current block is greater than 16 samples, only two wider widths can be chosen, such as (θ, 4θ) (or (τ, 4τ)).
[0108] In another adaptive blending method, up to N (e.g., 5) different blending weight regions can be selected, such as {θ1, θ2, θ3, θ4, θ5}. In one example, the choice of blending weights is unrestricted. In one example, the blending weight regions include {1 / 2, 1, 2, 4, 8} or {0, 1, 2, 4, 8}. To accommodate the increased width of the GPM blending region (e.g., the blending region width θ increases from 2 to 8), the maximum value of the weights can be changed from 8 in Equation 2 to 32 in Equation 3. Thus, the ramp function in Equation 2 can be transformed into the ramp function described in Equation 4.
[0109]
[0110] In one example, the ramp function ω in Equation 4 xc,yc It can be quantified as ω described in Equation 5 m,n .
[0111]
[0112] In one example, d(m, n) is 16 × d(x c, y c A weight index (e.g., weightIdx) can be defined as follows.
[0113]
[0114] The size (e.g., width) θ of the blending region can be selected from a set of predefined values, such as {1 / 2,1,2,4,8} or {0,1,2,4,8}.
[0115] As described above, predefined values (e.g., predefined blending region width candidates or predefined width candidates), such as {1 / 2,1,2,4,8} or {0,1,2,4,8}, can be used for adaptive blending regions. For example, signaling can be sent at the CU level to the index to indicate which blending function or which blending region θ (e.g., which predefined width candidate was selected) was chosen. Signaling the index may not be effective when the number of predefined width candidates is relatively large or when the index is relatively large (e.g., 5).
[0116] According to embodiments of this disclosure, an adaptive blending process with TM can be applied to reorder predefined width candidates (e.g., {θ1, θ2, θ3, θ4, θ5}) in the width candidate list, and the width of the blending region of the current block encoded with GPM can be determined based on the reordered width candidate list. For example Figure 4C The sample within the mixed region in the current block can be reconstructed based on the width of the mixed region.
[0117] In one embodiment, the current block is encoded using GPM along the partition edges that divide the current block. Samples within the blended regions surrounding the partition edges or within the GPM blended regions can be reconstructed based on blending processes, such as... Figure 4C As described in [the document]. The blending region can be defined by the boundaries on both sides of the partition edge, and the boundaries are parallel to the partition edge. The width of the blending region is perpendicular to the partition edge and can be based on a width candidate list, which includes predefined width candidates, such as {θ1, θ2, θ3, θ4, θ5}. According to embodiments of this disclosure, when the adaptive blending condition is met, an adaptive blending process using TM can be applied to reorder the width candidates in the width candidate list (e.g., {θ1, θ2, θ3, θ4, θ5}), and the width of the blending region can be determined based on the reordered width candidate list to reconstruct the current block. For example, the width of the blending region is determined by selecting a width candidate from the reordered width candidate list. In the adaptive blending process with TM, the width candidates in the width candidate list can be reordered based on the current template of the current block and a reference template corresponding to the respective width candidate. Samples within the blending region in the current block can be reconstructed by applying the blending process using the selected width candidate.
[0118] In one embodiment, the width of the mixing region is determined based on N1 width candidates in a reordered width candidate list (e.g., the N1 width candidates with the minimum TM cost). N1 can be less than the number of width candidates in the width candidate list. The TM cost corresponding to the N1 width candidates can be less than or equal to one or more TM costs corresponding to one or more remaining width candidates in the reordered width candidate list.
[0119] The index of the adaptive hybrid region can be reordered in ascending order based on the TM cost. In one example, the best N1 (e.g., 2 or 3) candidates after only TM reordering are selected, and a signaling message is sent to the CU to notify the best N1 (e.g., 2 or 3) candidates after only TM reordering.
[0120] In another embodiment, the selected width candidate is the one with the lowest TM cost from the reordered list of width candidates. For example, when TM-based reordering is effective for the CU, only the width candidate with the lowest TM cost is used in the CU without signaling the index of the adaptive blending region. When multiple width candidates have the same TM cost, the width candidate with the lowest or highest index value can be selected.
[0121] When the adaptive blending condition is not met, the width of the blending region can be determined based on the width candidate list without reordering the width candidates in the width candidate list.
[0122] In one embodiment, the encoding information of the current block indicates whether an adaptive blending condition is met. In one example, a signaling flag is sent at the CU level to indicate whether a width candidate with the lowest TM cost can be used. If the flag is true, the selected blending region can be determined based on the TM cost, and the selected blending region corresponding to the lowest TM cost can be used for the specified GPM segmentation pattern. In one example, multiple width candidates may have the same TM cost (e.g., the lowest TM cost), and a width candidate can be selected as one of multiple width candidates with (i) the lowest index value or (ii) the highest index value. Otherwise, if the flag is false, no reordering of the width candidates in the width candidate list is performed, and other adaptive blending methods such as those described in this disclosure can be used.
[0123] In one embodiment, the adaptive blending condition is not met if the upper template above the current template or the left template to the left of the current template is unavailable. Reordering is ineffective when either the upper or left template is unavailable. Therefore, no reordering of width candidates in the width candidate list is performed; instead, other adaptive blending methods, such as those described in this disclosure (e.g., an adaptive blending weight index (unreordered) method), can be used to signal the selected adaptive weight index.
[0124] Figures 7A-7B An exemplary adaptive blending process using TM is illustrated according to an embodiment of the present disclosure. The current block (701) in the current image (711) is encoded using GPM along the partition edge (703). A first prediction mode and a second prediction mode can be used in the GPM. The partition edge (703) can correspond to a GPM partitioning mode. The width candidate list can include predefined width candidates, such as {θ1, θ2, θ3, θ4, θ5}. Figures 7A-7B An example of a width candidate θ1 is shown. The current template (702) of the current block (701) may include reconstructed samples from neighboring blocks of the current block (701). In one example, the current template (702) includes an upper template (712) above the current block (701) and a left template (713) to the left of the current block (701).
[0125] The partition edge (703) can extend from the current block (701) to the current template (702). The extension portion (715) of the partition edge (703) can be within the current template (702). The template blending region (717) including the template sample in the current template (702) can be located between the boundary (704) and the boundary (705), which are separated from the extension portion (715) by a distance (or blending region width) θ1. The template blending region (717) does not include the sample in the current block (701).
[0126] For a width candidate θ1 in the width candidate list, the reference template or hybrid reference template (732) corresponding to the corresponding width candidate θ1 can be determined based on GPM and the current template (702). Figure 4CThe blending process described herein can be applied to determine a reference template (732) based on a width candidate θ1 and a current template (702) segmented by partition edges (703). In one example, a first MV of a first prediction mode indicates a first reference block in a first reference image, and a first reference template of the first reference block is determined based on the current template (702). In one example, a second MV of a second prediction mode indicates a second reference block in a second reference image, and a second reference template of the second reference block is determined based on the current template (702). In one example, the first and second reference templates have the same shape and size as the current template. In one example, the blended reference template (732) is associated with the first and second reference images. The blended reference template (732) can be determined by a blending process, which is as follows: Figure 4C Equations 1-5 and / or similar processes described therein for measuring displacement from a sample in the current template (702) to the partition edge (703).
[0127] refer to Figure 7B The extension (731) of the partition edge corresponding to the partition edge (703) may intersect with the reference template (732). The extension (731) is parallel to the extension (715). In one example, the reference template sample in the reference template (732) may include a first reference template sample in region (721) whose displacement from the extension (731) is greater than or equal to +θ1, a second reference template sample in region (724) whose displacement from the extension (731) is less than or equal to -θ1, and a third reference template sample in the template blending region (722) located between the boundary (732) and the boundary (733). Each boundary in the boundary (732) and the boundary (733) is separated from the extension (731) by a distance θ1 (the width of the blending region). The template blending region (722) in the reference template (732) corresponds to the template blending region (717) in the current template (702).
[0128] The reference template sample in the reference template (732) can be based on, for example Figure 4CThe weight W of the reference template sample is determined by the mixed processing described in Equations 2-3 and 4-5. P0 and P1 can represent the predicted values of the reference template samples of the first and second prediction modes based on GPM, respectively. The weight W of the reference template sample can be determined based on the displacement of the reference template sample in the reference template (732) to the extension portion (731) or based on the displacement of the template sample in the current template (702) to the partition edge (703). For example, the first reference template sample in region (721) is determined as P0, the second reference template sample in region (724) is determined as P1, and the third reference template sample in the template mixing region (722) is determined as a weighted sum of P0 and P1 using Equation 1.
[0129] The TM cost corresponding to width candidate θ1 can be determined based on the current template (702) and the reference template (732). The above description can be applied to other width candidates in the width candidate list, such as θ2, θ3, θ4, and θ5, to obtain other TM costs (e.g., four other TM costs). The width candidates in the width candidate list can be reordered based on the determined TM costs, for example, based on the five TM costs corresponding to θ1, θ2, θ3, θ4, and θ5 respectively.
[0130] As described above, the TM cost of the current template (702) and the TM costs of the corresponding reference templates for possible blending regions (e.g., all possible blending regions), such as {θ1, θ2, θ3, θ4, θ5}, can be calculated to reorder the indices indicating the adaptive blending regions. The adaptive blending regions or their corresponding indices can be reordered in ascending order based on the TM cost.
[0131] In one example, the current block (701) is divided into two GPM partitions (761)-(762) by partition edge (703). The GPM partitions (761)-(762) can be predicted using a first prediction mode and a second prediction mode, for example, as... Figure 4C Or as described in Equations 1-3. The first prediction mode and the second prediction mode are inter-frame prediction modes having first motion information (e.g., first MV) and second motion information (e.g., second MV), respectively. Each TM cost corresponding to the width of the mixing region (e.g., θ1) is calculated using the current template (702) and a reference template (e.g., the reference template (732) corresponding to θ1). It can be based on a mixing method (e.g., for samples covered by the extended mixing region in the template region) Figure 4CAlternatively, a hybrid method described in Equations 1-3 can be used to derive a reference template (e.g., (732)) from first motion information (e.g., first MV) and second motion information (e.g., second MV). In one example, the sum of absolute differences (SAD) between the current template (702) and the reference template (e.g., (732)) is used as the TM cost.
[0132] Figures 8A-8B An exemplary blending area not fully covered within a template used for template matching is shown according to an embodiment of this disclosure. Figure 8A and Figure 8B In this code, the current block (801) is encoded using GPM. A partition edge (803) can divide the current block (801) into two partitions. The current template (802) can include an upper template (811) and a left template (812). Boundaries (804) and (805) can be located on opposite sides of the partition edge (803) and can be offset by a distance l1 from the partition edge (803) based on θ1 and the direction of the partition edge (803). Boundaries (804) and (805) are parallel to the partition edge (803).
[0133] exist Figure 8A In this process, the partition edge (803) can extend into the current template (802). The partition edge (803) can intersect with the current template (802). For a width candidate such as θ1, only a portion of the blending region (e.g., a portion of the template blending region) is covered within the current template (802). Figure 8A The template blending region (821) can include the region between the boundary (804) and the boundary (805). The template blending region (821) does not include samples in the current block (801).
[0134] refer to Figure 8A The template blending region (821) is partially covered by the current template (802). For example, the template blending region (821) may include a first region (822) (in gray) that overlaps with the current template (802) between the boundary (804) and the partition edge (803), a second region (823) (in gray) that overlaps with the current template (802) between the boundary (805) and the partition edge (803), and a third region (824) (in white) to the right of the current template (802). The first region (822) and the second region (823) in the template blending region (821) are covered by the current template (802), while the third region (824) is not covered by the current template (802).
[0135] refer to Figure 8BThe partition edge (803) does not intersect with the current template (802). For example, the extended partition edge (803) is outside the current template (802). The template blending region (831) may include the region between the boundary (804) and the boundary (805). The template blending region (831) does not include samples in the current block (801). The template blending region (831) is completely outside the current template (802).
[0136] refer to Figure 7A For each width candidate in the width candidate list {θ1, θ2, θ3, θ4, θ5}, the adaptive blending condition is satisfied when the template blending region (e.g., (717)) is completely covered by the current template (e.g., (702)). For the largest width candidate (e.g., θ5) in the width candidate list {θ1, θ2, θ3, θ4, θ5}, the adaptive blending condition is satisfied when the template blending region (e.g., (717)) is completely covered by the current template (e.g., (702)).
[0137] refer to Figure 8A and Figure 8B For width candidates in the width candidate list {θ1, θ2, θ3, θ4, θ5}, the adaptive blending condition is not satisfied when the template blending region (e.g., (821) or (831)) is not completely covered by the current template (e.g., (802)). In Figure 8A, the partition edge (803) intersects with the current template (802), but the adaptive blending condition is not satisfied when the template blending region (821) is not completely covered by the current template (802).
[0138] exist Figure 8B In the example, the partition edge (803) does not intersect with the current template (802), thus failing to meet the adaptive blending condition. In another example, the template blending region (831) is not covered by the current template (802), thus failing to meet the adaptive blending condition.
[0139] For width candidates in the width candidate list, when the template blending region is not completely covered by the current template, for example... Figure 8A and Figure 8B As shown, the template blending region is represented as invalid for TM-based reordering (e.g., NOT_VALID) for the GPM segmentation mode, and no reordering of width candidates in the width candidate list is performed for the GPM segmentation mode. An index indicating the size of the blending region for the GPM segmentation mode can be determined based on the unordered width candidate list.
[0140] In one embodiment, reordering is ineffective when potential blending regions (e.g., all possible blending regions) cannot be fully covered within the current template. In this case, alternative signaling methods for the adaptive blending weight index can be used to signal the selected adaptive weight index.
[0141] In one embodiment, when the blended region cannot completely cover the current template, the blended region cannot be effectively used for reordering GPM segmentation patterns. This invalid blended region cannot be used for signaling for the specified GPM segmentation pattern.
[0142] In one embodiment, the current block is encoded using GPM along the partition edge. The blending region surrounding the partition edge can be defined by the boundaries on either side of the partition edge. These boundaries can be parallel to the partition edge. The width of the blending region measured perpendicular to the partition edge can be determined based on a predefined list of width candidates, including width candidates. According to embodiments of this disclosure, a subset of width candidates or a subset of blending regions can be adaptively selected at the CU level or slice level, for example, based on signaling information at the corresponding level (e.g., CU level or slice level). Signaling can be sent to notify only the selected subset of width candidates or the selected subset of blending regions to reduce signaling costs.
[0143] A subset of width candidates can be determined based on a predefined list of width candidates, and the number of width candidates in the subset can be less than the number of width candidates in the predefined list. The subset of width candidates can come from the predefined list. The width of the blended region to be used for the current block can be determined by selecting a width candidate from the subset. In one example, a signaling notification indicating the index of the blended region's width can be sent. Samples within the blended region in the current block can be reconstructed by applying blending processing using the selected width candidate.
[0144] In one embodiment, if screen content encoding is enabled for multiple blocks including the current block, N² minimum width candidates can be selected as a subset of the width candidates from a predefined width candidate list. For example, when screen content encoding is enabled, for instance, at the sequence parameter set (SPS), picture parameter set (PPS), or slice level, only the first N² (e.g., 2 or 3) narrowest blending regions are signaled at the slice level or CU level. The screen content encoding tool can include tools for encoding screen content, such as IBC mode, palette encoding, etc. In one example, when screen content encoding is applied to the current block, N² is 1, and the width of the blending region is the minimum width candidate in the predefined width candidate list, for example, 0 or 1 / 2. A blending region width of 0 is equivalent to disabling GPM blending for the CU; for example, in Equation 1, the weight W is 0 or 1.
[0145] In one embodiment, a subset of width candidates can be determined based on the difference (or motion difference) between the two MVs of the first and second partitions of the current block divided by the partition edges. For example, a subset of only mixed regions can be selected based on the motion difference between the two resolved MVs between merge indices 0 and 1 associated with the first and second partitions, respectively. Two motion difference thresholds can be predefined as a first threshold (denoted by thr1) and a second threshold (denoted by thr2), where thr1 is less than thr2. In one example, the number of width candidates in each subset of the width candidates is the same.
[0146] For example, five width candidates (e.g., five possible mixed regions), such as {θ1, θ2, θ3, θ4, θ5}, can be used with GPM. For natural content with SCC encoding tools disabled for the current block, a first subset of width candidates (e.g., {θ1, θ2, θ3}) can be used as a subset of width candidates when the motion difference is less than a first threshold (thr1). When the motion difference is between the first threshold (thr1) and the second threshold (thr2) (e.g., the motion difference is greater than or equal to the first threshold and less than or equal to the second threshold), a second subset of width candidates (e.g., {θ2, θ3, θ4}), different from the first subset, can be used as a subset of width candidates for the current block using the parsed GPM segmentation pattern. When the motion difference is greater than the second threshold (thr2), a third subset of width candidates (e.g., {θ3, θ4, θ5}), different from the first and second subsets, can be used as a subset of width candidates for the current block. In a bidirectional prediction block using GPM encoding, the motion difference can be the maximum of the motion differences from reference list 0 and reference list 1. In one example, the widths, ordered in ascending order, include θ1, θ2, θ3, θ4, and θ5. The first subset of width candidates includes the three narrowest blending regions {θ1, θ2, θ3}, the second subset includes the three middle blending regions {θ2, θ3, θ4}, and the third subset includes the three largest blending regions {θ3, θ4, θ5}.
[0147] In one embodiment, the partition edge is extended into the current template, the extension portion is within the current template, and a subset of width candidates can be determined based on the sample value difference of the extension portion across the partition edge. The current template can include samples from the left-hand adjacent reconstructed block and / or the upper-hand adjacent reconstructed block of the current block. The samples used to calculate the sample value difference can be within the current template. In some examples, the sample value difference is referred to as the gradient value of the samples across the extension portion of the partition edge.
[0148] In one example, signaling is sent at the CU level to a subset of only mixed regions or a subset of width candidates based on the difference between sample values (e.g., gradient values of predicted samples) across the extended portion of partition edges or GPM segmentation boundaries. Two sample difference thresholds (e.g., two gradient thresholds) can be predefined as a first gradient threshold (denoted by gthr1) and a second gradient threshold (denoted by gthr2), where gthr1 is less than gthr2. For example, five width candidates (e.g., five possible mixed regions), such as {θ1, θ2, θ3, θ4, θ5}, can be used with GPM. When the gradient value is less than the first gradient threshold (gthr1), a fourth subset of width candidates (e.g., {θ3, θ4, θ5}) can be used as a subset of width candidates for the current block. When the gradient value is between the first gradient threshold (gthr1) and the second gradient threshold (gthr2), a fifth subset of width candidates (e.g., {θ2, θ3, θ4}), different from the fourth subset, can be used as a subset of width candidates for the current block using the resolved GPM segmentation pattern. When the gradient value is greater than the second gradient threshold (gthr2), a sixth subset (e.g., {θ1, θ2, θ3}), which is different from the fourth and fifth subsets of width candidates, can be used as a subset of width candidates. In one example, the widths in ascending order include θ1, θ2, θ3, θ4, and θ5. The sixth subset of width candidates includes the three narrowest blending regions {θ1, θ2, θ3}, the fifth subset of width candidates includes the three middle blending regions {θ2, θ3, θ4}, and the fourth subset of width candidates includes the three largest blending regions {θ3, θ4, θ5}.
[0149] In one example, {θ1, θ2, θ3, θ4, θ5} is {0, 1, 2, 4, 8}. In another example, {θ1, θ2, θ3, θ4, θ5} is {1 / 2, 1, 2, 4, 8}.
[0150] Figure 9 An example of obtaining the sample value difference (also called gradient value) across partition edges according to an embodiment of this disclosure is shown. The current block (901) is encoded using the GPM partitioning mode. In the GPM partitioning mode, partition edges (903) can divide the current block (901) into two partitions. A predefined list of width candidates can include width candidates, such as {θ1, θ2, θ3, θ4, θ5}. For example, as shown... Figure 9 The sample value difference can be determined for one of the width candidates (e.g., θi, i = 1, 2, 3, 4 or 5).
[0151] refer to Figure 9The first boundary (904) and the second boundary (905) can be located on opposite sides of the partition edge (903), and can be offset from the partition edge (903) by a distance li based on θi and the orientation of the partition edge (903). The partition edge (903), the first boundary (904), and the second boundary (905) can extend into the current template (921). As described above, the current template (921) can include an upper template (922) and a left template (923).
[0152] The region (915) between the first boundary (904) and the second boundary (905) within the current template (921) can be used to determine the sample value difference. The extension (910) of the partition edge (903) within the current template (921) can divide the current template (921) into two templates (911)-(912) located on different sides of the extension (910). Figure 9 In the example, template (912) includes the left portion of the left template (923) and the left portion of the upper template (922), and template (911) includes the right portion of the upper template (922). The sample value difference can be determined based on the sample difference across the extension portion (910), which is the boundary between the two templates.
[0153] In one embodiment, a first sample average is obtained based on a first template sample that surrounds and / or intersects with the extension portion (910), and a second sample average is obtained based on a second template sample that surrounds or intersects with the first boundary (904) or the second boundary (905). The sample value difference may be the absolute difference between the first sample average and the second sample average. The first template sample and the second template sample are located in region (915).
[0154] In another embodiment, the sample difference can be the average of the absolute sample differences between the first template sample and the corresponding second template sample.
[0155] In one example, referring to Equations 4-6, the second sample mean is derived from the average of the sample values with a weight index (e.g., weightIdx) of 0 (e.g., indicating that the sample falls on the second boundary (905) in the current template (921). The first sample mean is derived from the average of the sample values with a weight index (e.g., weightIdx) of 16×θi (e.g., indicating that the sample falls on the extension (910)). The weight index (e.g., weightIdx) can be determined based on the template sample (x c ,y c The weight index (e.g., weightIdx) is calculated using the distance between θ and the center position of the current block. i Calculate equation 6. θi It can be one of the width candidates from a predefined list of width candidates. In one example, θ i It is 8. In one example, θ i It is 2 or another value. d(m, n) is As described in Equation 5.
[0156] weightIdx is 16×θ i The first and second sample means when 0 can be expressed as follows: Figure 9 It is obtained as described above. The sample value difference can be the absolute difference between the first sample mean and the second sample mean.
[0157] In one embodiment, when using intra-frame and inter-frame modes in the GPM to predict the current block (901), the three largest mixing regions, such as {2, 4, 8}, are used as a subset of width candidates. In one example, the first GPM partition is predicted from the intra-frame mode, and the second GPM partition is predicted from the inter-frame mode.
[0158] Figure 10 A flowchart of an overview process (1000) according to an embodiment of the present disclosure is shown. The process (1000) can be used with a video encoder. In various embodiments, the process (1000) is executed by processing circuitry, such as processing circuitry that performs the functions of a video encoder (103), processing circuitry that performs the functions of a video encoder (303), etc. In some embodiments, the process (1000) is implemented as software instructions, so that the processing circuitry executes the processing (1000) when the software instructions are executed. The process (1000) begins at (S1001) and continues to (S1010).
[0159] At (S1010), it can be determined whether the current block satisfies the adaptive blending condition. The current block will be encoded using the Geometric Partitioning Pattern (GPM) along the partition edges intersecting with the current block. The blending region surrounding the partition edges can be defined by the boundaries on both sides of the partition edges (e.g., a first boundary and a second boundary). The boundaries can be parallel to the partition edges. The width of the blending region measured perpendicular to the partition edges can be based on a width candidate list including width candidates. If it is determined that the adaptive blending condition is satisfied, the process at (1000) can continue to (S1020).
[0160] If the adaptive blending condition is determined not to be met, the width of the blending region can be determined based on the width candidate list without reordering the width candidates in the list, and samples within the blending region of the current block can be reconstructed by applying blending processing, for example in... Figure 4C As described in [the document]. Processing (1000) can continue to (S1099) and then terminate.
[0161] For a width candidate in the width candidate list (e.g., the maximum width candidate), a region (e.g., a template blending region (821) or a template blending region (831)) can be defined by a first boundary and a second boundary (e.g., (804)-(805)) parallel to the partition edge. This region does not include samples within the current block. The first and second boundaries can be located on opposite sides of the partition edge. In one embodiment, if the region is only partially covered by the current template, for example... Figure 8A As shown, or the area is completely outside the current template, for example, as Figure 8B As shown, the adaptive mixing condition is not satisfied.
[0162] In one example, the distances from the first and second boundaries to the partition edge are the largest width candidate in the width candidate list, and the extension of the partition edge intersects with the current template. The region between the first and second boundaries includes both a first region and a second region. This region may include the extension of the partition edge. The first region is the area between the first boundary and the partition edge that overlaps with the current template. The second region is the area between the second boundary and the partition edge that overlaps with the current template. Based on the first region being equal to the second region, the adaptive blending condition can be determined. (Reference) Figure 7A If the first region in region (717) is equal to the second region in region (717), then the adaptive mixing condition is satisfied. If the first region is not equal to the second region, then the adaptive mixing condition is not satisfied. (See reference...) Figure 8A If the first region (e.g. (822)) is not equal to the second region (e.g. (823)), it is determined that the adaptive mixing condition is not satisfied.
[0163] In one example, the extended portion of the partition edge (e.g.) Figure 8B (803) in the template does not intersect with the current template (e.g., Figure 8B (802) can be used to determine that the adaptive mixing condition is not satisfied.
[0164] In one example, if the top template above the current block or the left template to the left of the current block is unavailable, it is determined that the adaptive blending condition is not met.
[0165] At (S1020), the width candidates in the width candidate list can be reordered using template matching (TM) based on the current template of the current block and the reference template corresponding to the width candidate.
[0166] In one embodiment, the partition edge can extend from the current block into the current template, and the extended portion of the partition edge can be within the current template. For each width candidate in the width candidate list, a reference template corresponding to the corresponding width candidate can be determined based on the GPM and the current template. Reference samples within the template blending region surrounding the extended portion of the partition edge can be determined based on a blending process. The template blending region can be within the reference template, and the width of the template blending region can be based on the corresponding width candidate. The TM cost corresponding to the corresponding width candidate can be determined based on the current template and the reference template, and the width candidate list can be reordered based on the determined TM cost.
[0167] In (S1030), the width of the blending region can be determined by selecting a width candidate from the reordered list of width candidates.
[0168] In one embodiment, the width of the blending region is based on N1 width candidates from a reordered list of width candidates. N1 may be less than the number of width candidates in the list. The TM cost corresponding to the N1 width candidates may be less than or equal to one or more TM costs corresponding to one or more remaining width candidates in the reordered list. For example, the width of the blending region is selected from the N1 width candidates with the lowest TM cost from the reordered list of width candidates.
[0169] In one example, the selected width candidate is the one with the lowest TM cost from the reordered list of width candidates.
[0170] At (S1040), samples within the mixed region of the current block can be encoded by applying the mixed processing using the selected width candidate.
[0171] In one example, the encoding information (e.g., index) of the current block indicating the width candidate selected in the reordered list of width candidates can be encoded and included in the bitstream to be sent to the decoder.
[0172] In one example, the encoding information for the current block indicates whether the adaptive mixing condition is met.
[0173] Then, processing (1000) proceeds to (S1099) and terminates.
[0174] Process (1000) can be appropriately adapted to various scenarios, and the steps in process (1000) can be adjusted accordingly. One or more steps in process (1000) can be adjusted, omitted, repeated, and / or combined. Process (1000) can be implemented using any suitable order. Additional steps can be added.
[0175] Figure 11 A flowchart of an overview process (1100) according to an embodiment of the present disclosure is shown. Process (1100) can be used with a video decoder. In various embodiments, process (1100) is executed by processing circuitry, such as processing circuitry that performs the functions of a video decoder (110), processing circuitry that performs the functions of a video decoder (210), etc. In some embodiments, process (1100) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes process (1100). Process (1100) begins at (S1101) and continues to (S1110).
[0176] In (S1110), the encoding information of the current block in the current image can be decoded. The encoding information may indicate that the current block is encoded using the geometric partitioning pattern (GPM) along a partition edge (e.g., (452)) intersecting with the current block. The blending region surrounding the partition edge may be defined by the boundaries on both sides of the partition edge (e.g., (471)-(472)). The boundaries may be parallel to the partition edge. The width of the blending region measured perpendicular to the partition edge may be based on a width candidate list including width candidates.
[0177] At (S1120), it can be determined whether the adaptive mixing condition is satisfied. If it is determined that the adaptive mixing condition is satisfied, then process (1100) can proceed to (S1130).
[0178] In one example, the encoding information for the current block indicates whether the adaptive mixing condition is met.
[0179] If it is determined that the adaptive blending condition is not met, the width of the blending region can be determined based on the width candidate list without reordering the width candidates in the width candidate list, and the samples in the blending region in the current block can be reconstructed by applying the blending process. Then, the process (1100) can proceed to (S1199) and terminate.
[0180] For a width candidate in the width candidate list (e.g., the maximum width candidate), a region (e.g., a template blending region (821) or a template blending region (831)) can be defined by a first boundary and a second boundary (e.g., (804)-(805)) parallel to the partition edge. This region does not include samples within the current block. The first and second boundaries can be located on opposite sides of the partition edge. In one embodiment, if the region is only partially covered by the current template, for example... Figure 8A As shown, or the area is completely outside the current template, for example, as Figure 8B As shown, the adaptive mixing condition is not satisfied.
[0181] In one example, the distances from the first and second boundaries to the partition edge are the largest width candidate in the width candidate list, and the extension of the partition edge intersects with the current template. The region between the first and second boundaries includes both a first region and a second region. This region may include the extension of the partition edge. The first region is the area between the first boundary and the partition edge that overlaps with the current template. The second region is the area between the second boundary and the partition edge that overlaps with the current template. Based on the first region being equal to the second region, the adaptive blending condition can be determined. (Reference) Figure 7A If the first region in region (717) is equal to the second region in region (717), then the adaptive mixing condition is satisfied. If the first region is not equal to the second region, then the adaptive mixing condition is not satisfied. (See reference...) Figure 8A If the first region (e.g. (822)) is not equal to the second region (e.g. (823)), it is determined that the adaptive mixing condition is not satisfied.
[0182] In one example, the extended portion of the partition edge (e.g.) Figure 8B (803) in the template does not match the current template (e.g., Figure 8B The intersection of (802) in the equation indicates that the adaptive mixing condition is not satisfied.
[0183] In one example, if the top template above the current block or the left template to the left of the current block is unavailable, it is determined that the adaptive blending condition is not met.
[0184] At (S1130), the width candidates in the width candidate list can be reordered using template matching (TM) based on the current template of the current block and the reference template corresponding to the width candidate.
[0185] In one embodiment, the partition edge can extend from the current block into the current template, and the extended portion of the partition edge can be within the current template. For each width candidate in the width candidate list, a reference template corresponding to the corresponding width candidate can be determined based on the GPM and the current template. Reference samples within the template blending region surrounding the extended portion of the partition edge can be determined based on a blending process. The template blending region can be within the reference template, and the width of the template blending region can be based on the corresponding width candidate. The TM cost corresponding to the corresponding width candidate can be determined based on the current template and the reference template, and the width candidate list can be reordered based on the determined TM cost.
[0186] In (S1140), the width of the blending region can be determined by selecting a width candidate from the reordered list of width candidates.
[0187] In one embodiment, the width of the blending region is based on N1 width candidates from a reordered list of width candidates. N1 may be less than the number of width candidates in the list. The TM cost corresponding to the N1 width candidates may be less than or equal to one or more TM costs corresponding to one or more remaining width candidates in the reordered list. For example, the width of the blending region is selected from the N1 width candidates with the lowest TM cost from the reordered list of width candidates.
[0188] In one example, the selected width candidate is the one with the lowest TM cost from the reordered list of width candidates.
[0189] At (S1150), samples within the blended region of the current block can be reconstructed by applying blending processing using the selected width candidate.
[0190] Processing (1100) proceeds to (S1199) and terminates.
[0191] Process (1000) can be appropriately adapted to various scenarios, and the steps in process (1000) can be adjusted accordingly. One or more steps in process (1000) can be adjusted, omitted, repeated, and / or combined. Process (1000) can be implemented using any suitable order. Additional steps can be added.
[0192] In one embodiment, a video bitstream including the current block in the current image is received. The value of the syntax element associated with the current block in the current image is decoded. The syntax element indicates whether the current block is encoded using Geometric Partitioning (GPM) along partition edges intersecting the current block. If the value of the syntax element indicates that the current block is encoded using GPM and satisfies adaptive blending conditions, a TM can be used to reorder width candidates in the width candidate list, based on the current template of the current block and a reference template corresponding to the respective width candidate. Width candidates can be selected from the reordered width candidate list. The width of the blending region can be determined based on the selected width candidate. The blending region surrounds the partition edge and is defined by the boundaries on both sides of the partition edge. The boundaries are parallel to the partition edge, and the width of the blending region is measured perpendicular to the partition edge. The blending region can be determined based on the width of the blending region. Samples within the blending region in the current block can be reconstructed by applying adaptive blending using the determined blending region.
[0193] Figure 12A flowchart of an overview process (1200) according to an embodiment of the present disclosure is shown. Process (1200) can be used with a video encoder. In various embodiments, process (1200) is executed by processing circuitry, such as processing circuitry that performs the functions of a video encoder (103), processing circuitry that performs the functions of a video encoder (303), etc. In some embodiments, process (1200) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes process (1200). Process (1200) begins at (S1201) and continues to (S1210).
[0194] At (S1210), a subset of width candidates for the current block to be encoded using Geometric Partitioning Pattern (GPM) along the partition edge intersecting the current block can be determined based on a width candidate list (e.g., a predefined width candidate list). The blending region surrounding the partition edge can be defined by the boundaries on either side of the partition edge. The boundaries can be parallel to the partition edge. The width of the blending region measured perpendicular to the partition edge can be based on a predefined width candidate list that includes width candidates. In one example, the number of width candidates in the subset is less than the number of width candidates in the predefined width candidate list.
[0195] In one embodiment, if screen content encoding tools are enabled for multiple blocks including the current block, then N2 minimum width candidates are selected as a subset of the width candidates from a predefined list of width candidates.
[0196] In one embodiment, a subset of width candidates is determined to reconstruct the current block based on the difference between two motion vectors of a first prediction mode and a second prediction mode used in GPM. In one example, screen content encoding tools are disabled for the current block, and a first threshold is less than a second threshold. If the difference between the two motion vectors is less than the first threshold, a subset of width candidates is determined as a first subset of width candidates. If the difference between the two motion vectors is greater than or equal to the first threshold and less than or equal to the second threshold, a subset of width candidates is determined as a second subset of width candidates. If the difference between the two motion vectors is greater than the second threshold, a subset of width candidates is determined as a third subset of width candidates.
[0197] In one embodiment, the partition edge is extended into the current template, wherein the extended portion is within the current template. The current template may include samples from the left adjacent reconstructed block and / or the upper adjacent reconstructed block of the current block. A subset of width candidates may be determined based on the sample value difference of the extended portion across the partition edge. The samples are within the current template. In one example, a first sample difference threshold is less than a second sample difference threshold. If the absolute value of the sample value difference is less than the first sample difference threshold, a subset of width candidates is determined as a fourth subset of width candidates. If the absolute value of the sample value difference is greater than or equal to the first threshold and less than or equal to the second threshold, a subset of width candidates is determined as a fifth subset of width candidates. If the absolute value of the sample value difference is greater than the second sample difference threshold, a subset of width candidates is determined as a sixth subset of width candidates.
[0198] In one example, the current block partition is predicted using intra-frame prediction. A subset of the width candidates can be selected from a predefined list of maximum width candidates.
[0199] In (S1220), the width of the blending region can be determined by selecting a width candidate from a subset of the width candidates.
[0200] At (S1230), samples within the mixed region of the current block can be encoded by applying the mixed processing using the selected width candidate.
[0201] In one example, encoding information (e.g., an index) of the current block, indicating a selected width candidate in a subset of width candidates, is encoded and included in the bitstream to be sent to the decoder.
[0202] Then, processing (1200) proceeds to (S1299) and terminates.
[0203] Process (1200) can be appropriately adapted to various scenarios, and the steps in process (1200) can be adjusted accordingly. One or more steps in process (1200) can be adjusted, omitted, repeated, and / or combined. Process (1200) can be implemented using any suitable order. Additional steps can be added.
[0204] Figure 13A flowchart of an overview process (1300) according to an embodiment of the present disclosure is shown. This process (1300) can be used with a video decoder. In various embodiments, the process (1300) is executed by processing circuitry, such as processing circuitry that performs the functions of a video decoder (110), processing circuitry that performs the functions of a video decoder (210), etc. In some embodiments, the process (1300) is implemented as software instructions, so that the processing circuitry executes the processing (1300) when the software instructions are executed. The process (1300) begins at (S1301) and continues to (S1310).
[0205] In (S1310), the encoding information of the current block in the current image can be decoded. The encoding information can indicate that the current block is encoded using a geometric partitioning pattern (GPM) along a partition edge (e.g., (452)) that intersects with the current block. The blending region surrounding the partition edge can be defined by the boundaries on both sides of the partition edge (e.g., (471)-(472)). The boundaries can be parallel to the partition edge. The width of the blending region measured perpendicular to the partition edge can be based on a width candidate list including width candidates (e.g., a predefined width candidate list).
[0206] In (S1320), a subset of width candidates can be determined based on a predefined list of width candidates. In one example, the number of width candidates in the subset is less than the number of width candidates in the predefined list of width candidates.
[0207] In one embodiment, if screen content encoding tools are enabled for multiple blocks including the current block, N2 minimum width candidates are selected as a subset of the width candidates from a predefined list of width candidates.
[0208] In one embodiment, a subset of width candidates is determined to reconstruct the current block based on the difference between two motion vectors of a first prediction mode and a second prediction mode used in GPM. In one example, screen content encoding tools are disabled for the current block, and a first threshold is less than a second threshold. If the difference between the two motion vectors is less than the first threshold, a subset of width candidates is determined as a first subset of width candidates. If the difference between the two motion vectors is greater than or equal to the first threshold and less than or equal to the second threshold, a subset of width candidates is determined as a second subset of width candidates. If the difference between the two motion vectors is greater than the second threshold, a subset of width candidates is determined as a third subset of width candidates.
[0209] In one embodiment, the partition edge is extended into the current template, wherein the extended portion is within the current template. The current template may include samples from the left adjacent reconstructed block and / or the upper adjacent reconstructed block of the current block. A subset of width candidates may be determined based on the sample value difference of the extended portion across the partition edge. The samples are within the current template. In one example, a first sample difference threshold is less than a second sample difference threshold. If the absolute value of the sample value difference is less than the first sample difference threshold, a subset of width candidates is determined as a fourth subset of width candidates. If the absolute value of the sample value difference is greater than or equal to the first threshold and less than or equal to the second threshold, a subset of width candidates is determined as a fifth subset of width candidates. If the absolute value of the sample value difference is greater than the second sample difference threshold, a subset of width candidates is determined as a sixth subset of width candidates.
[0210] In one example, the current block partition is predicted using intra-frame prediction. A subset of the width candidates can be selected from a predefined list of maximum width candidates.
[0211] In (S1330), the width of the blending region can be determined by selecting a width candidate from a subset of the width candidates.
[0212] At (S1340), samples within the blended region of the current block can be reconstructed by applying blending processing using the selected width candidate.
[0213] Processing (1300) proceeds to (S1399) and terminates.
[0214] The process (1300) can be appropriately adapted to various scenarios, and the steps in the process (1300) can be adjusted accordingly. One or more steps in the process (1300) can be adjusted, omitted, repeated, and / or combined. The process (1300) can be implemented using any suitable order. Additional steps can be added.
[0215] The embodiments in this disclosure can be used individually or in combination in any order. Furthermore, each of the methods (or embodiments), encoders, and decoders can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0216] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 14 A computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0217] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or through interpretation, microcode execution, etc.
[0218] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0219] Figure 14 The components of the computer system (1400) shown are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependencies or requirements relating to any one or a combination of components shown in the exemplary embodiments of the computer system (1400).
[0220] The computer system (1400) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographs acquired from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0221] Human-machine interface input devices may include one or more of the following (only one of each is shown): keyboard (1401), mouse (1402), touchpad (1403), touch screen (1410), data glove (not shown), joystick (1405), microphone (1406), scanner (1407), camera (1408).
[0222] The computer system (1400) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of a touchscreen (1410), data gloves (not shown), or joysticks (1405), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (1409), headphones (not depicted)), and visual output devices (e.g., screens (1410) including CRT screens, LCD screens, plasma screens, OLED screens, each screen may or may not have touchscreen input functionality, each screen may or may not have tactile feedback functionality, some of which are capable of outputting two-dimensional or more three-dimensional visual outputs via devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).
[0223] The computer system (1400) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1420) having media such as CD / DVD (1421), finger drives (1422), removable hard disk drives or solid-state drives (1423), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0224] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0225] The computer system (1400) may also include an interface (1454) leading to one or more communication networks (1455). The network may be, for example, a wireless network, a wired network, or an optical network. The network may further be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (1400)) attached to some general-purpose data port or peripheral bus (1449); other network interfaces are typically integrated into the core of the computer system (1400) by being attached to a system bus (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system (1400) can use any of these networks to communicate with other entities. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., a CANBus connected to certain CANBus devices), or bidirectional, such as connecting to other computer systems using a local area network (LAN) or wide area network (WAN) digital network. As mentioned above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.
[0226] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel (1440) of the computer system (1400).
[0227] The core (1440) may include one or more central processing units (CPU) (1441), graphics processing units (GPUs) (1442), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1443), hardware accelerators (1444) for certain tasks, graphics adapters (1450), etc. These devices, along with read-only memory (ROM) (1445), random access memory (1446), and internal mass storage (1447) such as internal non-user-accessible hard disk drives, SSDs, etc., may be connected via a system bus (1448). In some computer systems, the system bus (1448) may be accessed in the form of one or more physical connectors to allow for extension via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1448) or attached to the core's system bus (1448) via a peripheral bus (1449). In one example, a screen (1410) may be connected to a graphics adapter (1450). Peripheral bus architectures include PCI, USB, etc.
[0228] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute certain instructions that can be combined to form the aforementioned computer code. This computer code can be stored in ROM (1445) or RAM (1446). Transient data can be stored in RAM (1446), while permanent data can be stored, for example, in internal mass storage (1447). Fast storage and retrieval to any storage device can be achieved using a cache, which can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage (1447), ROM (1445), RAM (1446), etc.
[0229] Computer-readable media may have computer code thereon that performs various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.
[0230] By way of example, and not limitation, software contained in one or more tangible computer-readable media can be executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) to enable a computer system having an architecture (1400), particularly a kernel (1440), to provide functionality. Such computer-readable media can be media associated with user-accessible mass storage as described above, and memory of certain non-transitory kernels (1440), such as internal kernel mass storage (1447) or ROM (1445). Software implementing various embodiments of this disclosure can be stored in such devices and executed by the kernel (1440). Depending on specific needs, the computer-readable media may include one or more storage devices or chips. The software can cause the kernel (1440), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1446) and modifying such data structures according to software-defined processing. Additionally or alternatively, the computer system may be made functional by logic hard-wired or otherwise embodied in circuitry (e.g., accelerator (1444)), which may replace or operate with the software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., integrated circuit (IC)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0231] While several exemplary embodiments have been described in this disclosure, modifications, substitutions, and various equivalent alternatives fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.
Claims
1. A method for video decoding in a video encoder, the method comprising: The method comprises: receiving a video bitstream, the video bitstream comprising a current block in a current picture; decoding a value of a syntax element associated with the current block in the current picture, the syntax element indicating whether the current block is coded with a geometric partition mode (GPM) along a partition edge intersecting the current block; and in response to the value of the syntax element indicating that the current block is coded with the GPM and that an adaptive blending condition is satisfied, reordering width candidates in a width candidate list using template matching (TM) based on a current template of the current block and reference templates corresponding to the respective width candidates; selecting a width candidate from the reordered width candidate list; determining a width of a blending region based on the selected width candidate, the blending region surrounding the partition edge and being defined by boundaries on both sides of the partition edge, the boundaries being parallel to the partition edge, the width of the blending region being measured perpendicular to the partition edge; determining the blending region based on the width of the blending region; and reconstructing samples within the blending region in the current block by applying adaptive blending using the determined blending region.
2. The method of claim 1, wherein, The reordering of the width candidates comprises: extending the partition edge from the current block into the current template, the extended portion of the partition edge being in the current template; for each width candidate in the width candidate list, determining a reference template corresponding to the respective width candidate based on the GPM and the current template, determining reference samples within a template blending region based on blending processing, the template blending region surrounding the extended portion of the partition edge, the template blending region being in the reference template, the width of the template blending region being based on the respective width candidate; and determining a TM cost corresponding to the respective width candidate based on the current template and the reference template; and reordering the width candidate list based on the determined TM costs.
3. The method according to claim 1 or 2, characterized in that, The method further comprises: in response to the adaptive blending condition not being satisfied, determining the width of the blending region based on the width candidate list without reordering the width candidates in the width candidate list; and reconstructing the samples within the blending region in the current block by applying blending processing.
4. The method of claim 3, wherein the coding information of the current block indicates whether the adaptive blending condition is satisfied.
5. The method of claim 3, wherein in response to the extended portion of the partition edge intersecting the current template, an area between a first boundary and a second boundary parallel to the partition edge comprises a first area and a second area, the area including the extended portion of the partition edge, the first boundary and the second boundary being located on opposite sides of the partition edge, and a distance of the first boundary and the second boundary to the partition edge being a maximum width candidate in the width candidate list, the first area is an area between the first boundary and the partition edge that overlaps the current template, the second area is an area between the second boundary and the partition edge that overlaps the current template. the second region is a region between the second boundary and the partition edge that overlaps the current template, the adaptive blending condition is satisfied based on the first region being equal to the second region; and the adaptive blending condition is not satisfied based on the first region not being equal to the second region. and the adaptive blending condition is not satisfied in response to the extended portion of the partition edge not intersecting the current template.
6. The method of claim 3, wherein the adaptive blending condition is not satisfied in response to an upper template above the current block or a left template to the left of the current block being unavailable.
7. The method of claim 2, wherein a width of the blending region is based on Nl width candidates of the reordered width candidate list, Nl being less than a number of width candidates in the width candidate list, and a TM cost corresponding to the Nl width candidates being less than or equal to one or more TM costs corresponding to one or more remaining width candidates in the reordered width candidate list.
8. The method of claim 2, wherein the selected width candidate is a width candidate in the reordered width candidate list having a minimum TM cost.
9. A method of video decoding for use in a video decoder, the method comprising: The method comprises: decoding coding information of a current block in a current picture, the coding information indicating that the current block is coded with a geometric partition mode (GPM) along a partition edge, a blending region around the partition edge being defined by boundaries on both sides of the partition edge, the boundaries being parallel to the partition edge, a width of the blending region measured perpendicular to the partition edge being based on a pre-defined width candidate list including width candidates; determining a subset of width candidates based on the pre-defined width candidate list, a number of width candidates in the subset of width candidates being less than a number of width candidates in the pre-defined width candidate list; determining the width of the blending region by selecting a width candidate from the subset of width candidates; and reconstructing samples within the blending region in the current block by applying a blending process using the selected width candidate.
10. The method of claim 9, wherein, The determining the subset of width candidates comprises: selecting N2 smallest width candidates from the pre-defined width candidate list as the subset of width candidates in response to a screen content coding tool being enabled for a plurality of blocks including the current block.
11. The method of claim 9, wherein, The determining the subset of width candidates comprises: determining the subset of width candidates to reconstruct the current block based on a difference of two motion vectors of a first prediction mode and a second prediction mode used in the GPM.
12. The method of claim 11, wherein a screen content coding tool is disabled for the current block, a first threshold is less than a second threshold, and The determining the subset of width candidates comprises: determining the subset of width candidates as a first subset of width candidates in response to the difference of the two motion vectors being less than the first threshold; in response to the difference between the two motion vectors being greater than or equal to the first threshold value and less than or equal to the second threshold value, determining the subset of width candidates as a second subset of width candidates, the second subset of width candidates being different from the first subset; and in response to the difference between the two motion vectors being greater than the second threshold value, determining the subset of width candidates as a third subset of width candidates, the third subset of width candidates being different from the first subset and the second subset.
13. The method of claim 9, wherein extending the partition edge into a current template, wherein the extended portion is in the current template, the current template including samples in a left neighboring reconstructed block or an above neighboring reconstructed block of the current block, and the determining the subset of width candidates includes determining the subset of width candidates based on a sample value difference across the extended portion of the partition edge, the samples being in the current template.
14. The method of claim 13, wherein a first sample difference threshold value is less than a second sample difference threshold value, and the determining the subset of width candidates includes: in response to an absolute value of the sample value difference being less than the first sample difference threshold value, determining the subset of width candidates as a fourth subset of width candidates; in response to the absolute value of the sample value difference being greater than or equal to the first sample difference threshold value and less than or equal to the second sample difference threshold value, determining the subset of width candidates as a fifth subset of width candidates, the fifth subset of width candidates being different from the fourth subset; and in response to the absolute value of the sample value difference being greater than the second sample difference threshold value, determining the subset of width candidates as a sixth subset of width candidates, the sixth subset of width candidates being different from the fourth subset and the fifth subset.
15. The method of claim 9, wherein the current block is partitioned by intra prediction, and the determining the subset of width candidates includes selecting three largest width candidates from the predefined list of width candidates as the subset of width candidates.
16. An apparatus for video decoding, comprising processing circuitry, characterized in that, the processing circuitry is configured to perform the method for video decoding in a video encoder of any of claims 1 to 8 or the method for video decoding in a video decoder of any of claims 9 to 15.
17. A computer readable storage medium storing program instructions, wherein, the program instructions, when run on a computer, cause the computer to perform the method for video decoding in a video encoder of any of claims 1 to 8 or the method for video decoding in a video decoder of any of claims 9 to 15.
18. A method of storing a bitstream, the method comprising: performing the method for video decoding in a video encoder of any of claims 1 to 8 or the method for video decoding in a video decoder of any of claims 9 to 15 generates a bitstream; and storing the bitstream.
19. A method of transmitting a bitstream, the method comprising: performing the method for video decoding in a video encoder of any of claims 1 to 8 or the method for video decoding in a video decoder of any of claims 9 to 15 generates a bitstream; and transmitting the bitstream.
Citation Information
Patent Citations
Geometric partitioning mode in video coding
CN113796083A
Geometric segmentation mode with coordinated motion field storage and motion compensation
CN114342373A