Method, apparatus, and computer program for boundary filtering for intraBC and intraTMP modes

Boundary filtering for IBC and intraTMP modes addresses block discontinuities, improving video compression efficiency by enabling adaptive filtering based on syntax elements in the encoded video stream.

JP2025523724APending Publication Date: 2025-07-25TENCENT AMERICA LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024515304
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-03
Filing Date
2022-11-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Current coding standards for intra-block copy (IBC) and intra-template matching prediction (intraTMP) modes do not adequately address block discontinuities at prediction block boundaries, leading to inefficiencies in video compression.

Method used

Implement boundary filtering for blocks located at the boundaries of current pictures in IBC and intraTMP modes, enabling or disabling filtering based on syntax elements in the encoded video stream to generate filtered samples for decoding.

Benefits of technology

Improves video compression efficiency by effectively handling block discontinuities, enhancing the decoding process for IBC and intraTMP modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523724000001
    Figure 2025523724000001
  • Figure 2025523724000002
    Figure 2025523724000002
  • Figure 2025523724000003
    Figure 2025523724000003
Patent Text Reader

Abstract

A method performed by a video decoder includes receiving an encoded video bitstream including a current picture having at least one block located at a boundary of the current picture. The method includes determining, based on syntax elements within the received encoded video stream, whether boundary filtering is enabled for the at least one block. The method further includes filtering one or more boundary samples corresponding to the at least one block based on a determination that boundary filtering is enabled to generate one or more filtered samples, and decoding the at least one block based on the one or more generated filtered samples. The method further includes decoding the at least one block without filtering the one or more boundary samples based on a determination that boundary filtering is not enabled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority based on U.S. Provisional Patent Application No. 63 / 388,983, filed on July 13, 2022, and U.S. Patent Application No. 17 / 980,316, filed on November 3, 2022, and the entire disclosures of these are incorporated herein by reference.

[0002] Technical Field The present disclosure generally relates to communication systems, and more particularly, to methods and apparatuses for boundary filtering for intra - block copy (IBC) and intra - template - matching prediction (intraTMP) modes.

Background Art

[0003] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) announced the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, these two standardization organizations jointly formed the JVET (Joint Video Exploration Team) to explore the possibility of developing the next video coding standard beyond HEVC. In October 2017, they announced a Call for Proposals (CfP) for video compression with capabilities beyond HEVC. By February 15, 2018, a total of 22 CfP responses regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding the 360 video category were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / JVET meeting. As a result of this meeting, JVET officially started the standardization process for the next-generation video coding beyond HEVC, and the new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) announced the VVC video coding standard (version 1). ECM (Enhanced Compression Model) software is developed under the collaborative exploration research by the Joint Video Exploration Team (JVET) of ITU-T VCEG and ISO / IEC MPEG as a potential enhanced video coding technology beyond the capabilities of VVC. The current coding standards for intra-block copy (IBC) and the intra-template matching prediction mode do not fully consider the block discontinuity at the prediction block boundary. Summary of the Invention

Problems to be Solved by the Invention

[0004] The following presents a simplified summary of one or more embodiments of the present disclosure to provide a basic understanding of such embodiments. This summary is not an exhaustive overview of all contemplated embodiments, nor is it intended to identify key or critical elements of all embodiments or to delineate the scope of any or all embodiments. Its sole purpose is to present some concepts of one or more embodiments of the present disclosure in a simplified form as a prelude to the more detailed description that follows.

[0005] A method, apparatus, and non-transitory computer-readable medium for boundary filtering for intra-block copy (IBC) and intra-template matching prediction (intraTMP) modes are disclosed by the present disclosure.

Means for Solving the Problems

[0006] According to one exemplary embodiment, a method executed by at least one processor in a video decoder includes receiving an encoded video bitstream including a current picture having at least one block located at a boundary of the current picture, encoded according to one of (i) an intra block copy (IBC) mode, (ii) an intra template matching (intraTMP) mode, and (iii) an IBC merge mode. The method further includes determining, based on syntax elements in the received encoded video stream, whether boundary filtering is enabled for the at least one block. The method further includes, based on a determination that boundary filtering is enabled: filtering one or more boundary samples corresponding to the at least one block to generate one or more filtered samples, and decoding the at least one block based on the one or more generated filtered samples. The method further includes, based on a determination that boundary filtering is not enabled, decoding the at least one block without filtering the one or more boundary samples.

[0007] According to an exemplary embodiment, a video decoder includes at least one memory configured to store computer program code, and at least one processor configured to access the computer program code and operate as instructed by the computer program code. The computer program code includes reception code configured to cause the at least one processor to receive an encoded video bitstream including a current picture having at least one block located at a boundary of the current picture, encoded according to one of (i) an intra block copy (IBC) mode, (ii) an intra template matching (intraTMP) mode, and (iii) an IBC merge mode. The computer program code includes determination code configured to cause the at least one processor to determine whether boundary filtering is to be enabled for the at least one block based on syntax elements in the received encoded video stream. The computer program code further includes filtering code and determination code. Based on a determination that boundary filtering is to be enabled, the filtering code is configured to cause the at least one processor to filter one or more boundary samples corresponding to the at least one block to generate one or more filtered samples, and the decode code is configured to cause the at least one processor to decode the at least one block based on the one or more generated filtered samples. Based on a determination that the boundary filtering is not to be enabled, the decode code is configured to cause the at least one processor to decode the at least one block without filtering the one or more boundary samples.

[0008] According to an exemplary embodiment, a non-transitory computer-readable medium storing instructions, which when executed by a processor in a video decoder, cause the processor to execute a method, the method including: receiving an encoded video bitstream including a current picture having at least one block located at a boundary of the current picture, the at least one block being encoded according to one of (i) an intra-block copy (IBC) mode, (ii) an intra-template matching (intraTMP) mode, and (iii) an IBC merge mode; further determining, based on syntax elements in the received encoded video stream, whether boundary filtering is to be enabled for the at least one block; further, based on a determination that boundary filtering is to be enabled: filtering one or more boundary samples corresponding to the at least one block to generate one or more filtered samples, and decoding the at least one block based on the one or more generated filtered samples; and further, based on a determination that boundary filtering is not to be enabled, decoding the at least one block without filtering the one or more boundary samples.

[0009] Additional embodiments are described in the following description, will be apparent in part from the description, and / or may be learned by practice of the presented embodiments of the disclosure.

Brief Description of the Drawings

[0010] The above and other aspects, features, and aspects of the embodiments of the present disclosure will become apparent from the following description related to the accompanying drawings.

[0011]

Figure 1

[0012]

Figure 2

[0013]

Figure 3

[0014]

Figure 4

[0015]

Figure 5

[0016]

Figure 6

[0017]

Figure 7

[0018]

Figure 8

[0019]

Figure 9

[0020]

Figure 10

[0021]

Figure 11

[0022]

Figure 12

[0023]

Figure 13

[0024]

Figure 14

Best Mode for Carrying Out the Invention

[0025] The following detailed description of the exemplary embodiments refers to the accompanying drawings. The same reference numbers in different drawings can identify the same or similar elements.

[0026] The foregoing disclosure provides examples and explanations, but is not intended to be exhaustive or to limit implementation to the exact forms disclosed. Modifications and variations are possible in light of the above disclosure or may be obtained from practice of the implementation. Further, one or more features or components of an embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Further, in the flowcharts and operation descriptions provided below, one or more operations may be omitted, one or more operations may be added, one or more operations may be (at least partially) executed simultaneously, and the order of one or more operations may be interchanged.

[0027] It is apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting. Thus, the operations and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0028] Certain combinations of features are recited in the claims and / or disclosed herein, but these combinations are not intended to limit the disclosure of possible implementations. Indeed, many of these features can be combined in ways that are not specifically recited in the claims and / or disclosed herein. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible implementations includes combinations of each dependent claim with each of the other claims in the claim set.

[0029] Elements, operations, or instructions used in this specification should not be construed as important or essential unless explicitly stated as such. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." When only one item is intended, the term "one" or similar language is used. Also, as used herein, terms such as "have," "possess," "include," "comprise," etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise specified. Additionally, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0030] Throughout this specification, references to "one embodiment," "an embodiment," or similar terms mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the solution of the present invention. Thus, the phrases "in one embodiment," "in an embodiment," and similar language throughout this specification may refer to the same embodiment, but do not necessarily all refer to the same embodiment.

[0031] Furthermore, the described features, advantages, and characteristics of the present disclosure can be combined in any suitable manner in one or more embodiments. Those skilled in the art will recognize, in light of the description herein, that the present disclosure can be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages that are not necessarily present in all embodiments of the present disclosure may be recognized in certain embodiments.

[0032] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) can include at least two terminals (110, 120) interconnected via a network (150). For one-way data transmission, the first terminal (110) can encode video data at a local location for transmission to another terminal (120) via the network (150). The second terminal (120) can receive the encoded video data of the counterpart terminal from the network (150), decode the encoded data, and display the recovered video data. One-way data transmission is common in media providing applications and the like.

[0033] FIG. 1 shows a second pair of terminals (130, 140) provided to support two-way transmission of encoded video that can occur, for example, during a video conference. For two-way data transmission, each terminal (130, 140) may encode video data captured at a local location for transmission to the counterpart terminal via the network (150). Each terminal (130, 140) may also receive the encoded video data transmitted by the counterpart terminal, decode the encoded data, and display the recovered video data on a local display device.

[0034] In FIG. 1, the terminals (110 - 140) may be shown as servers, personal computers, smartphones, and / or other types of terminals. For example, the terminals (110 - 140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing facilities. The network (150) represents any number of networks that transmit encoded video data among the terminals (110 - 140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data over circuit - switched channels and / or packet - switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless otherwise described hereinafter in this specification.

[0035] FIG. 2 shows the arrangement of video encoders and decoders in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video - enabled applications, including, for example, storage of compressed video on digital media including video conferencing, digital TV, CD, DVD, memory sticks, etc.

[0036] As shown in FIG. 2, the streaming system (200) can include a capture subsystem (213) that can include a video source (201) and an encoder (203). The video source (201) can be, for example, a digital camera and can be configured to generate an uncompressed video sample stream (202). The uncompressed video sample stream (202) can provide a high data volume as compared to an encoded video bitstream and can be processed by an encoder (203) coupled to the camera (201). The encoder (203) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in detail below. The encoded video bitstream (204) can include a lower data volume as compared to the sample stream and can be stored in a streaming server (205) for future use. One or more streaming clients (206) can access the streaming server (205) to retrieve a video bitstream (209) that can be a copy of the encoded video bitstream (204).

[0037] In embodiments, the streaming server (205) can function as a Media-Aware Network Element (MANE). For example, the streaming server (205) can be configured to clip the encoded video bitstream (204) to adapt potentially different bitstreams to one or more of the streaming clients (206). In embodiments, the MANE can be provided separately from the streaming server (205) within the streaming system (200).

[0038] The streaming client (206) can include a video decoder (210) and a display (212). The video decoder (210) can decode, for example, a video bit stream (209) that is an incoming copy of the encoded video bit stream (204) and generate an outgoing video sample stream (211) that can be rendered on the display (212) or another rendering device (not shown). In some streaming systems, the video bit streams (204, 209) can be encoded according to certain video encoding / compression standards. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. A video encoding standard under development is informally known as Versatile Video Coding (VVC). Embodiments of the present disclosure may be used in the context of VVC.

[0039] Figure 3 shows an exemplary functional block diagram of a video decoder (210) attached to a display (212) according to an embodiment of the present disclosure. The video decoder (210) can include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (358). In at least one embodiment, the video decoder (210) can include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder (210) may be partially or fully embodied as software executed on one or more CPUs with associated memory.

[0040] In this and other embodiments, a receiver (310) can receive one or more encoded video sequences, decoded by a decoder (210), one at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel (312) that can be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, and those data may be transferred to respective usage entities (not shown). The receiver (310) can separate the encoded video sequences from other data. To handle network jitter, a buffer memory (315) may be coupled between the receiver (310) and an entropy decoder / parser (320) (hereinafter, "parser"). If the receiver (310) receives data from a store / forward device having sufficient bandwidth and controllability, or from an isochronous network, the buffer (315) may not be used or may be small. For use in a best effort packet network, such as the Internet, a buffer (315) may be necessary, may be relatively large, and may be of an adaptable size.

[0041] The video decoder (210) can include a parser (320) that reconstructs symbols (321) from an entropy-encoded video sequence. The categories of these symbols can include, for example, information used to manage the operation of the decoder (210) and potential information for controlling a rendering device such as a display (212) that can be coupled to the decoder as shown in FIG. 2. The control information for the rendering device can be in the form of, for example, a supplementary enhancement information (SEI) message or a video user utility information (VUI) parameter set fragment (not shown). The parser (320) can parse / entropy-decode the received encoded video sequence. The encoding of the encoded video sequence can be in accordance with a video encoding technique or standard and can follow principles well known to those skilled in the art, including variable-length encoding, Huffman encoding, arithmetic encoding with or without context sensitivity, etc. The parser (320) can extract a set of subgroup parameters from the encoded video sequence based on at least one parameter corresponding to a group for at least one subgroup of pixels within the video decoder. The subgroups can include, for example, picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (320) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0042] The parser (320) can perform an entropy decoding / parsing operation on the video sequence received from the buffer (315) to generate symbols (321). The reconstruction of the symbols (321) can involve multiple different units depending on the type of the encoded video picture or a part thereof (inter and intra pictures, inter and intra blocks, etc.) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (320). Such a flow of subgroup control information between the parser (320) and the following multiple units is not shown for clarity.

[0043] In addition to the function blocks already described, the decoder (210) can be conceptually subdivided into several functional units as described below. In a practical implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate.

[0044] One unit may be a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive, as symbols (321) from the parser (320), control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc., together with the quantized transform coefficients. The scaler / inverse transform unit (351) may output a block including sample values, and the sample values may be input to the aggregator (355).

[0045] In some cases, the output samples of the scaler / inverse transform unit (351) may be related to intra-coded blocks. That is, a block that does not use prediction information from a previously reconstructed picture but may use prediction information from a previously reconstructed part of the current picture. Such prediction information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses the surrounding already reconstructed information fetched from the current (partially reconstructed) picture in the current picture memory (358) to generate a block of the same size and shape as the block being reconstructed. The aggregator (355) may, in some cases, add, sample by sample, the prediction information generated by the intra-coded block (352) to the output sample information provided by the scaler / inverse transform unit (351).

[0046] In other cases, the output samples of the scaler / inverse transform unit (351) may be related to inter-coded, potentially motion-compensated blocks. In such cases, the motion-compensation prediction unit (353) may access the reference picture memory (357) to fetch the samples used for prediction. After motion-compensating the fetched samples according to the symbols (321) regarding the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (in this case, called the residual samples or residual signal) to generate the output sample information. The address in the reference picture memory (357) from which the motion-compensation prediction unit (353) fetches the prediction samples can be controlled by the motion vector. The motion vector may be available to the motion-compensation prediction unit (353) in the form of, for example, a symbol (321) that can have X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (357) when an exact sub-sample motion vector is used, a motion vector prediction mechanism, etc.

[0047] The output samples of the aggregator (355) can follow various loop filtering techniques within the loop filter unit (356). The video compression technique can include in-loop filter techniques that are included in the encoded video bitstream and controlled by parameters made available to the loop filter unit (356) as symbols (321) from the parser (320), but may also respond to meta information obtained during the decoding of the previous part (in decode order) of an encoded picture or an encoded video sequence, or to previously reconstructed and loop filtered sample values.

[0048] The sample stream of the output of the loop filter unit (356) is not only output to a rendering device such as the display (212), but may also be stored in the reference picture memory (357) for use in future inter-picture prediction.

[0049] Once fully reconstructed, certain encoded pictures may be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and identified as a reference picture (e.g., by the parser (320)), the current reference picture becomes part of the reference picture memory (357), and a fresh current picture memory may be reallocated before starting the reconstruction of the next encoded picture.

[0050] The video decoder (210) can perform a decoding operation according to a predetermined video compression technique that may be described in a standard such as ITU-T Rec. H.265. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used. This means conforming to the syntax of the video compression technique or standard, as specified in the document of that video compression technique or standard, particularly in its profile document. Also, in order to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level can limit, for example, the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured in units of megasamples per second), the maximum reference picture size, etc. The limitations set by the level can, in some cases, be further restricted through the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.

[0051] In one embodiment, the receiver (310) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0052] FIG. 4 shows an exemplary functional block diagram of a video encoder (203) associated with a video source (201) according to an embodiment of the present disclosure. The video encoder (203) may include, for example, an encoder which is a source encoder (430), an encoding engine (432), a (local) decoder [decoder] (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy encoder (445), a controller (450), and a channel (460).

[0053] The encoder (203) can receive video samples from a video source (201) (which is not part of the encoder) that can capture the video images to be encoded by the encoder (203). The video source (201) can provide a source video sequence to be encoded by the encoder (203) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media providing system [media serving system], the video source (201) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that give motion when viewed in sequence. Each picture itself may be configured as a spatial array of pixels, and each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0054] According to an embodiment, the encoder (203) may encode and compress pictures of the source video sequence into the encoded video sequence (443) in real time or under any other arbitrary time constraints required by the application. Enforcing an appropriate encoding speed is one function of the controller (450). The controller (450) can also control other functional units as described hereinafter and may be functionally coupled to these units. The coupling is not shown for clarity. The parameters set by the controller (450) can include rate control related parameters (picture skip, quantizer, lambda value of rate-distortion optimization techniques, …), picture size, picture group (GOP) layout, maximum motion vector search range, etc. A person skilled in the art can easily identify other functions of the controller (450) related to a video encoder (203) optimized for a certain system design.

[0055] Some video encoders operate in what those skilled in the art would readily recognize as an "encoding loop". As a drastically simplified explanation, the encoding loop may consist of an encoding portion of a source encoder (430) (responsible for creating symbols based on an input picture and reference picture(s) to be encoded) and a (local) decoder (433) embedded in the encoder (203). The (local) decoder embedded in the encoder reconstructs the symbols and generates sample data that a (remote) decoder would also generate in some video compression techniques where the compression between the symbols and the encoded video bitstream is reversible. This reconstructed sample stream may be input into a reference picture memory (434). Since the decoding of the symbol stream yields a bit-exact result independent of the location of the decoder (local or remote), the content of the reference picture memory is also bit-exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder "sees" the exact same sample values as reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift, if synchronization cannot be maintained, e.g., due to channel errors) is known to those skilled in the art.

[0056] The operation of the "local" decoder (433) may be the same as that of the "remote" decoder (210) already described in detail in connection with FIG. 3. However, since the symbols are available and the encoding / decoding of the symbols into the encoded video sequence by the entropy encoder (445) and the parser (320) can be lossless, the entropy decoding portion of the decoder (210) including the channel (312), the receiver (310), the buffer (315), and the parser (320) may not be fully implemented in the local decoder (433).

[0057] What can be observed at this point is that any decoder technology other than the parse / entropy decoding existing in the decoder may need to exist in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the operation of the decoder. The description of the encoder technology may be omitted since it may be the reverse of the decoder technology described comprehensively. More detailed descriptions are necessary only in certain areas and are provided below.

[0058] As part of its operation, the source coder (430) may perform motion-compensated predictive coding. This is to predictively code an input frame by referring to one or more previously encoded frames from a video sequence designated as a "reference frame". In this way, the coding engine (432) codes the difference between a pixel block of the input frame and pixel blocks of the reference frame(s) that can be selected as a predictive reference (s) for the input frame.

[0059] The local video decoder (433) may decode the encoded video data of a frame that can be designated as a reference frame based on the symbols generated by the source coder (430). It may be advantageous for the operation of the coding engine (432) to be an irreversible process. If the encoded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (433) may replicate the decoding process that can be performed by the video decoder for the reference frame and store the reconstructed reference frame in the reference picture memory (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame having the same content as the reconstructed reference frame obtained by the remote video decoder (in the absence of transmission errors).

[0060] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) can search the reference picture memory (434) for sample data (as candidate reference pixel blocks) that can serve as appropriate prediction references for the new picture, or certain metadata such as motion vectors and block shapes of the reference pictures. The predictor (435) can operate on a pixel block basis for each sample block to find an appropriate prediction reference. In some cases, the input picture can have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (434) as determined by the search results obtained by the predictor (435).

[0061] The controller (450) can manage the encoding operation of the video encoder (430), including setting parameters and subgroup parameters used, for example, to encode video data. The outputs of all the functional units described above can be entropy encoded in the entropy encoder (445). The entropy encoder converts the symbols generated by the various functional units into an encoded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0062] The transmitter (440) can buffer the encoded video sequence generated by the entropy encoder (445) and prepare it for transmission via a communication channel (460) which can be a hardware / software link to a storage device that stores the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (430) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown). The controller (450) can manage the operation of the encoder (203). During encoding, the controller (450) can assign a certain type of encoded picture type to each encoded picture, which can affect the encoding technique applicable to each picture. For example, a picture may often be assigned as an intra picture (I picture), a predicted picture (P picture), or a bi-directionally predicted picture (B picture).

[0063] An intra picture (I picture) may be one that can be encoded and decoded without using other frames in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including for example, an Independent Decoder Refresh (IDR) picture. Those skilled in the art are aware of these variations of I pictures, and their respective uses and characteristics.

[0064] A predicted picture (P picture) may be one that can be encoded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0065] Bidirectional prediction pictures (B pictures) may be encoded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple prediction pictures can use more than three reference pictures and associated metadata for the reconstruction of a single block.

[0066] The source picture is generally spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be encoded block by block. The blocks can be predictively encoded by referring to other (already encoded) blocks as determined by the encoding assignment applied to each picture of the block. For example, blocks of an I picture can be encoded non-predictively or predictively by referring to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be encoded non-predictively, via spatial prediction, or via temporal prediction that refers to one previously encoded reference picture. Blocks of a B picture can be encoded non-predictively, via spatial prediction, or via temporal prediction that refers to one or two previously encoded reference pictures.

[0067] The video encoder (203) can perform an encoding operation according to a predetermined video encoding technology or standard such as ITU-T Rec. H.265. In that operation, the video encoder (203) can perform various compression operations including a predictive encoding operation that utilizes temporal and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technology or standard being used.

[0068] In one embodiment, the transmitter (440) can transmit additional data along with the encoded video. The video encoder (430) can include such data as part of the encoded video sequence. The additional data can include, for example, temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual user information (VUI) parameter set fragments, and the like.

[0069] Before describing some aspects of certain embodiments of the present disclosure in more detail, some terms that are referred to in the remainder of this description are introduced below.

[0070] "Sub-picture" hereinafter refers to samples, blocks, macroblocks, coding units, or similar entities that are arranged in a rectangular pattern, grouped semantically, and can be independently coded at a modified resolution. One or more sub-pictures can form a picture. One or more coded sub-pictures can form a coded picture. One or more sub-pictures can be grouped together to form a picture, and one or more sub-pictures can be extracted from a picture. In certain environments, one or more coded sub-pictures can be grouped in the compressed domain to form a coded picture without transcoding at the sample level, and in the same or certain other cases, one or more coded sub-pictures can be extracted from the coded picture in the compressed domain.

[0071] "Adaptive Resolution Change" (ARC) hereinafter refers to a mechanism that allows for changing the resolution of a picture or sub-picture within an encoded video sequence, for example, by resampling a reference picture. "ARC parameter" hereinafter refers to the control information required to perform adaptive resolution change and can include, for example, filter parameters, scaling factors, the resolution of the output and / or reference pictures, various control flags, and the like.

[0072] For each inter-predicted CU, motion parameters consisting of a motion vector, a reference picture index, and a reference picture list use index, as well as additional information required for the new coding functions of VVC, are used for inter-prediction sample generation. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded in skip mode, the CU is associated with one PU, there are no significant residual coefficients, and there is no coded motion vector delta or reference picture index. When the merge mode is specified, the motion parameters for the current CU are obtained from neighboring CUs including spatial and temporal candidates, and additional schedules introduced in VVC. The merge mode can be applied not only to skip mode but also to any inter-predicted CU. An alternative to the merge mode is the explicit transmission of motion parameters. In that case, the motion vector, the corresponding reference picture index for each reference picture list, and the reference picture list use flag as well as other required information are explicitly signaled for each CU.

[0073] In VVC, the VTM reference software includes several new refined inter-prediction coding tools as follows. · Extended merge prediction · Merge motion vector difference (MMVD) · AMVP mode using symmetric MVD signaling · Affine motion compensation prediction · Subblock-based temporal motion vector prediction (SbTMVP) · Adaptive motion vector resolution (AMVR) · Motion Field Storage: 1 / 16 Luma Sample MV Memory and 8×8 Motion Field Compression · Bi-prediction with CU-level weights (BCW) · Bi-directional optical flow (BDOF) · Decoder side motion vector refinement (DMVR) · Combined inter and intra prediction (CIIP) · Geometric partitioning mode (GPM)

[0074] The following provides details about inter prediction and related methods.

[0075] In VTM4, the merge candidate list is constructed by sequentially including the following five types of candidates: · Spatial MVP from spatial neighboring CUs · Temporal MVP from co-located CUs · History-based MVP from the FIFO table · Pairwise average MVP · Zero MV

[0076] The size of the merge list may be signaled in the slice header. The maximum allowable size of the merge list is 6 in VTM4. For each CU encoded in merge mode, the index of the best merge candidate can be encoded using truncated unary binarization (TU). The first bin of the merge index is encoded using context, and bypass coding is used for the other bins. The generation process for each category of merge candidates is provided in this session.

[0077] The derivation of spatial merge candidates in VVC can be the same as that in HEVC. Up to four merge candidates can be selected from among the candidates located at the positions shown in Fig. 5. The order of derivation is B1, A1, B0, A0, and B2. The position B2 can be considered only when any of the CUs at positions A0, B0, B1, A1 are not available (e.g., because it belongs to another slice or tile), or when it is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check. The redundancy check ensures that candidates with the same motion information are excluded from the list so that the coding efficiency is improved. To reduce the computational complexity, it is not necessary to consider all possible candidate pairs in the above redundancy check. Instead, only the pairs connected by the arrows in Fig. 6 are considered. A candidate is added to the list only if the corresponding candidate used for the redundancy check does not contain the same motion information.

[0078] In this operation, only one candidate is added to the list. In particular, in the derivation of this temporal merge candidate, a scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list used for the derivation of the co-located CU is explicitly signaled in the slice header. The scaled motion vector for the temporal merge candidate is obtained as shown by the dotted line in Fig. 7. This is scaled from the motion vector of the co-located CU using the POC distances tb and td. Here, tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set equal to 0.

[0079] The position for the temporal candidate is selected between candidates C0 and C1 as shown in Fig. 8. Position C1 is used when the CU at position C0 is not available, is intra-coded, or is outside the current row of the CTU. Otherwise, position C0 is used in the derivation of the temporal merge candidate.

[0080] The IBC mode was previously incorporated into the HEVC standard. However, there was a need to reduce the implementation cost due to the entire already-reconstructed area of the picture. The drawback of the IBC mode in HEVC is the need for additional memory in the DPB, and for this reason, hardware implementations typically use external memory. Additional external memory access is accompanied by increased memory bandwidth. VVC realizes the IBC mode using fixed memory by using on-chip memory to significantly reduce the memory bandwidth requirements and hardware complexity. The reference sample memory (RSM) can hold samples of a single CTU. A special feature of the RSM is the continuous update mechanism that replaces the reconstructed samples of the left neighboring CTU with the reconstructed samples of the current CTU. Furthermore, the block vector (BV) coding of the IBC mode uses the concept of the merge list for inter prediction. The construction process of the IBC list considers two spatially neighboring BVs and five history-based BVs (HBVPs). Here, when added to the candidate list, only the first HBVP can be compared with the spatial candidates. Normal inter prediction uses two different candidate lists for the merge mode and the normal mode, but the candidate list in IBC is for both cases. However, the merge mode can use up to six candidates in the list, while the normal mode can only use the first two candidates. The block vector difference (BVD) coding uses the motion vector difference (MVD) process, which generates the final BV of any size. There is also the fact that the reconstructed BV may point to an area outside the reference sample area, and correction is required by removing the absolute offset in each direction using modulo operations with the width and height of the RSM.

[0081] In ECM5, Intra-Template Matching Prediction (IntraTMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, and its L-shaped template matches the current template. For the pre-defined search range, the encoder searches for the template that is most similar to the current template in the reconstructed part of the current frame and uses the corresponding block as the prediction block. Then, the encoder signals the use of this mode, and the same prediction operation is executed on the decoder side.

[0082] The prediction signal can be generated by matching the L-shaped causal neighborhood of the current block with another block within the pre-defined search area of FIG. 9 consisting of: R1: The current CTU R2: The upper left CTU R3: The upper CTU R4: The left CTU

[0083] SAD can be used as the cost function. Within each region, the decoder searches for the template with the minimum SAD with respect to the current one and uses the corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) can be set proportional to the block dimensions (BlkW, BlkH) so as to have a fixed number of SAD comparisons per pixel. That is, Equation (1) SearchRange_w = a × BlkW Equation (2) SearchRange_h = a × BlkH where "a" is a constant that controls the trade-off between gain and complexity. In practice, "a" may be equal to 5.

[0084] The Intra-Template Matching tool may be enabled for CUs with a size of width and height of 64 or less. This maximum CU size for Intra-Template Matching may be configurable. The Intra-Template Matching prediction mode may be signaled at the CU level through a dedicated flag when DIMD is not used for the current CU.

[0085] In ECM5, IntraTMP may need to access 320 up-samples and 320 left-samples for the 64×64 blocks it supports. Additional memory may sometimes help improve the coding efficiency of the IBC mode. The reference area for the IBC mode may be extended up to the top 2 rows of the CTU. Figure 10 shows a picture (1000) having a reference area for encoding CTU(m,n). Specifically, to encode CTU(m,n), the reference area can include CTUs with indices (m-2,n-2)…(W,n-2), (0,n-1)...(W,n-1), (0,n)…(m,n). Here, W represents the maximum horizontal index within the current tile, slice, or picture. This setting ensures that when the CTU size is 128, IBC does not require extra memory in the current ETM platform. The per-sample block-vector search (or local search as it is called) range can be restricted to [-(C<<1),C>>2] horizontally and [-C,C>>2] vertically to conform to the extension of the reference area. Here, C indicates the CTU size.

[0086] To reduce complexity and signaling overhead, in JVET-W0110, possible intra prediction mode deformations are also studied. In JVET-W0110, two configurations are performed to study the influence from the said deformations of possible intra prediction modes for GPM using inter and intra prediction. The first configuration tries only the intra direction modes parallel and perpendicular to the geometric partition line. In the second configuration, in addition to the intra angle modes parallel and perpendicular to the geometric partition line, the planar mode is also tested. There are two or three possible intra prediction modes tested for geometric partitioning in GPM using inter and intra prediction.

[0087] In VVC, the results of intra prediction for DC, planar, and some angle modes are further refined by the Position-Dependent Intra Prediction Combination (PDPC) method. PDPC is an intra prediction method that calls for a combination of HEVC-style intra prediction with boundary reference samples and filtered boundary reference samples. PDPC may be applied to the following intra modes without signaling: planar, DC, intra angles below horizontal, and intra angles above vertical and below 80. If the current block is in the Bdpcm mode, or if the MRL index is greater than 0, PDPC is not applied. The predicted sample pred(x',y') is predicted according to the following formula using a linear combination of the intra prediction mode (DC, planar, angle) and the reference samples. Equation (1) pred(x',y') = Clip(0, (1 << BitDepth)-1, (wL × R -1,y' + wT × R x',-1 + (64 - wL - wT) × pred(x',y') + 32) >> 6), where R x,-1 and R -1,y represent the reference samples located at the upper and left boundaries of the current sample (x,y), respectively. Equation (2) w L = 32 >> ((x << 1) >> s Equation (3) w T = 32 >> ((y << 1) >> s Here, s is a parameter that controls the attenuation rate of the upper and left reference sample weightings from top to bottom and from left to right, respectively.

[0088] When PDPC is applied to DC, planar, horizontal, and vertical intra - modes, no additional boundary filter is required as necessary for the boundary filter of the HEVC DC mode or the edge filter of the horizontal / vertical modes. The PDPC processes for the DC mode and the planar mode may be the same. For the angular modes, when the current angular mode is HOR_IDX or VER_IDX, the left or upper reference samples may not be used. The weights and scale factors of PDPC may depend on the prediction mode and the block size. PDPC can be applied to blocks where both the width and height are 4 or more.

[0089] (A) - (D) of FIG. 11 show the definitions of the reference samples (R x,-1 and R -1,y ) for PDPC applied to various prediction modes. The predicted sample pred(x', y') can be located at (x', y') within the prediction block. As an example, the coordinate x of the reference sample R x,-1 is given by x = x'+y'+1, and the coordinate y of the reference sample R -1,y for the diagonal mode can similarly be given by y = x'+y'+1. For other angular modes, the reference samples R x,-1 and R -1,y can be located at fractional sample positions. In this case, the sample value at the nearest integer sample position can be used.

[0090] In IntraBC and IntraTMP, the prediction block may be indicated by a block vector that points to another block within the same picture. Thus, there may be block discontinuities at the prediction block boundaries compared to the neighboring reconstructed samples. Such discontinuities at the prediction block boundaries limit the coding efficiency of IntraBC and IntraTMP, especially in content captured by a camera.

[0091] Embodiments of the present disclosure may be directed to applying boundary filtering to prediction blocks of IntraBC and IntraTMP. Boundary filtering applies an adjustment to prediction samples at a block boundary using nearby reconstruction samples from neighboring encoded blocks. Embodiments of the present disclosure may be used separately or combined in any order. Further, embodiments of the present disclosure may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0092] In some embodiments, boundary filtering is applied to prediction blocks of IntraBC and IntraTMP. Boundary filtering applies one or more adjustments to prediction samples at a block boundary using nearby reconstruction samples from a previously encoded area. In some embodiments, boundary filtering is the same as PDPC applied to other intra prediction modes (e.g., DC and planar modes).

[0093] In some embodiments, boundary filtering is PDPC with some adjustments compared to PDPC applied to other intra prediction modes (e.g., DC and planar modes). In one example, different values of parameter s may be used compared to PDPC used for other intra prediction modes.

[0094] In some embodiments, for left (upper) boundary prediction samples, boundary filtering is a weighted average of the left (upper) neighboring reconstruction samples and the left (upper) boundary prediction samples. An example of boundary filtering for block (1200) is shown in FIG. 12 using a two-tap filter for the boundary prediction samples of the upper row and the left column.

[0095] In some embodiments, the number of lines above and columns to the left of the predicted samples filtered using the boundary filter may depend on the block size. In some embodiments, a block-level and / or HLS-level flag may be signaled to indicate whether PDPC is applied to IntraBC and / or IntraTMP, and the HLS may be a flag within a VPS, PPS, SPS, APS, slice header, frame header, tile header, or CTU header.

[0096] In some embodiments, the template matching (TM) cost (similar to that used for IntraTMP as in 4.3) of the current IBC or IntraTMP block may be used to determine whether and how to apply the boundary filter. For example, for a block encoded in the IBC model, the template matching cost may be calculated based on the template area pointed to by the BV of the current block. If the TM cost is below a threshold T1, the boundary filtering may be disabled. For example, T1 may be equal to 0.

[0097] In another example, for a block encoded in the IntraTMP mode, the template matching cost of the IntraTMP mode may be used to check against a threshold T2. If the TM cost is below the threshold T2, the proposed boundary filtering may be disabled. In some embodiments, the values of T1 and T2 may be different. In some embodiments, one or more PDPC parameters, such as s, may depend on the template matching cost.

[0098] In some embodiments, when the current coding block is coded in IBC merge mode, the boundary filter is not applied. In other embodiments, when the current block is coded in IBC merge mode and the BVP is derived from a neighboring upper or upper-right spatial candidate, only the reconstructed samples of the left neighborhood may be used for boundary filtering. In other embodiments, when the current block is coded in IBC merge mode and the BVP is derived from a neighboring left or lower-left spatial candidate, only the reconstructed samples of the upper neighborhood are used for boundary filtering.

[0099] In some embodiments, the residuals of the blocks coded by IntraTMP and IntraBC may be used to determine whether to apply the boundary filter and how to apply it. For example, if the energy of the residual is greater than a threshold T1', boundary filtering may be applied. In another example, if the energy of the residual is less than a threshold T2', boundary filtering may not be applied. In another embodiment, the energy may be measured by SAD, SSE, SATD, MSE of the residual block. In another embodiment, the values of the thresholds T1' and T2' may be different for the blocks coded by IntraTMP and IntraBC. In one example, PDPC parameters such as s may depend on the energy of the residual.

[0100] FIG. 13 shows an embodiment of a process (1300) for performing boundary filtering. The process (1300) may be performed by a decoder such as decoder (210). The process may start at operation (1302), where the coded video bitstream includes a current picture having at least one block located at the boundary of the current picture, and the at least one block is coded according to one of (i) IBC mode, (ii) intraTMP mode, and (iii) IBC merge mode.

[0101] The process proceeds to an operation (1304) where it is determined whether boundary filtering is enabled. The determination of whether boundary filtering is enabled can be made based on whether a predetermined condition is satisfied. For example, the predetermined condition can be satisfied when the at least one block is encoded in one of the IBC mode and the intraTMP mode. As another example, the predetermined condition is satisfied in response to a determination that the template matching cost is below a threshold when the at least one block is encoded in one of the IBC mode and the intraTMP mode. The predetermined condition may be based on syntax elements included in the received encoded video bitstream.

[0102] If it is determined that a predetermined condition enabling boundary filtering is satisfied, the process proceeds to operation (1306), where one or more boundary samples are filtered to generate boundary samples. The process proceeds to operation (1308), where the at least one block is decoded based on the filtered samples. If it is determined that the predetermined condition enabling boundary filtering is not satisfied, the process proceeds from operation (1304) to operation (1310), where the at least one block is decoded without filtering.

[0103] The components shown in FIG. 14 for the computer system (1400) are exemplary in nature and are not intended to imply any limitation as to the use or functionality of the computer software implementing embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement related to any one or combination of the components shown in the exemplary embodiment of the computer system (1400).

[0104] The computer system (1400) can include certain types of human interface input devices. Such human interface input devices can respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), voice input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). Also, the human interface device can be used to capture certain types of media that are not necessarily directly related to conscious input by a human, such as voice (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., 2D video, 3D video including stereoscopic video).

[0105] The input human interface device may include one or more of a keyboard (1401), a mouse (1402), a trackpad (1403), a touch screen (1410), a data glove, a joystick (1405), a microphone (1406), a scanner (1407), and a camera (1408) (only one of each is shown).

[0106] The computer system (1400) may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback by a touch screen (1410), data glove, or joystick (1405); however, there may also be tactile feedback devices that do not function as input devices). For example, such devices may be audio output devices (e.g., speakers (1409), headphones (not shown)), visual output devices (e.g., a screen including a CRT screen, LCD screen, plasma screen, OLED screen (1410); each may or may not have a touch screen input function, each may or may not have a tactile feedback function, and some of them can output higher than three-dimensional output through means such as two-dimensional visual output or stereoscopic output; virtual reality glasses (not shown), holographic display, and smoke tank (not shown)), and printers (not shown).

[0107] The computer system (1400) may also include a human-accessible memory device and associated media, such as an optical medium including a CD / DVD ROM / RW (1420) together with a CD / DVD or similar media (1421), a thumb drive (1422), a removable hard drive or solid state drive (1423), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0108] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0109] The computer system (1400) may also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, in-vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include Ethernet®, wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., cellular networks including cable TV, satellite TV, terrestrial broadcast TV, wired or wireless wide area digital networks including TV, in-vehicle and industrial including CANBus, etc. Certain types of networks typically require an external network interface adapter attached to a certain type of general-purpose data port or peripheral bus (1449) (such as a USB port of the computer system (1400)). Others are typically integrated into the core of the computer system (1400) by attachment to a system bus as described later (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be unidirectional, receive-only (such as broadcast TV), dedicated unidirectional transmission (such as CANbus to certain types of CANbus devices), or bidirectional to other computer systems using, for example, local or wide area digital networks. Such communication may include communication to a cloud computing environment (1455). For each of the networks and network interfaces as described above, certain protocols and protocol stacks can be used.

[0110] The aforementioned human interface device, human-accessible storage device, and network interface (1454) may be attached to the core (1440) of the computer system (1400).

[0111] The core (1440) can include one or more central processing units (CPUs) (1441), a graphics processing unit (GPU) (1442), a specialized programmable processing device in the form of a field programmable gate array (FPGA) (1443), a hardware accelerator (1444) for certain tasks, etc. These devices can be connected through a system bus (1448) together with internal mass storage devices (1447) such as read-only memory (ROM) (1445), random access memory (1446), an internal hard drive that is not user-accessible, a solid state drive (SSD). In some computer systems, the system bus (1448) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core's system bus (1448) or through a peripheral bus (1449). Architectures for peripheral buses include PCI, USB, etc. A graphics adapter (1450) may be included in the core (1440).

[0112] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute certain instructions that can together constitute the above-described computer code. The computer code can be stored in the ROM (1445) or RAM (1446). Temporary data can also be stored in the RAM (1446), while persistent data can be stored, for example, in the internal mass storage device (1447). By using cache memory that can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage devices (1447), ROM (1445), RAM (1446), etc., fast storage and retrieval to any of the memory devices can be enabled.

[0113] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well known and available to those having skill in the computer software arts.

[0114] By way of example and not limitation, a computer system having an architecture (1400), specifically a core (1440), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be associated with user-accessible mass storage as introduced above, as well as certain types of storage of the core (1440) having a non-transitory nature, such as a mass storage device (1447) inside the core or a ROM (1445). The software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (1440). The computer-readable media can include one or more memory devices or chips depending on specific needs. The software can include defining a data structure stored in a RAM (1446) and modifying such a data structure according to a process defined by the software, and causing a specific process or a specific part of a specific process described herein to be executed by the core (1440) and specifically by a processor (including a CPU, GPU, FPGA, etc.) therein. Additionally or alternatively, the computer system can provide functionality as a result of logic wired within a circuit (e.g., an accelerator (1444)) or otherwise embodied, which can operate instead of or in conjunction with software for executing a specific process or a specific part of a specific process described herein. References to software include logic and, where appropriate, vice versa. References to computer-readable media can, where appropriate, include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0115] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementation to the exact forms disclosed. Modifications and variations are possible in light of the above disclosure or can be obtained from practice of the implementation.

[0116] It is understood that the specific order or hierarchical structure of the blocks in the processes / flowcharts disclosed herein are illustrations of exemplary approaches. It is understood that, based on design preferences, the specific order or hierarchical structure of the blocks in a process / flowchart may be rearranged. Additionally, some blocks may be combined or omitted. The presented methods present the elements of the various blocks in an order as an example, and are not meant to be limited to the specific order or hierarchical structure presented.

[0117] Some embodiments can relate to systems, methods, and / or computer-readable media in any possible integration of technical detail levels. Further, one or more of the above-described components may be implemented as instructions stored on a computer-readable medium and executable by at least one processor (and / or may include at least one processor). The computer-readable medium can include a computer-readable non-transitory storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform operations.

[0118] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as raised structures within grooves in which instructions are recorded, and any suitable combination of the foregoing. A computer-readable storage medium as used herein should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0119] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.

[0120] The computer-readable program code / instructions for performing the operations may be in source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for performing aspects or operations.

[0121] These computer-readable program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to generate a machine that causes the instructions executed via the processor of the computer or other programmable data processing apparatus to create means for implementing the functions / steps specified in the flowchart and / or block of the block diagram. These computer-readable program instructions may be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, so that the computer-readable storage medium having the instructions stored therein includes a manufactured article that includes instructions for implementing aspects of the functions / steps specified in the flowchart and / or block of the block diagram.

[0122] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to generate a computer-implemented process, and the instructions executed on the computer, other programmable apparatus, or other device implement the functions / steps specified in the flowchart and / or block of the block diagram.

[0123] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible embodiments of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of one or more executable instructions for implementing the specified logical function. The methods, computer systems, and computer-readable media can include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those shown in the figures. In some alternative implementations, the functions shown in the blocks may not be in the order shown in the figures. For example, two blocks shown in succession may actually be executed concurrently or substantially concurrently, or the blocks may sometimes be executed in the reverse order depending on the related functions. It should also be noted that each block of the illustrations of the block diagrams and / or flowcharts, and combinations of blocks in the illustrations of the block diagrams and / or flowcharts, may be implemented by a special-purpose hardware-based system that performs the specified functions or steps, or by a combination of special-purpose hardware and computer instructions.

[0124] It is obvious that the systems and / or methods described herein may be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0125] Abbreviations HEVC: High Efficiency Video Coding (High-Efficiency Video Encoding) HDR: High dynamic range (High Dynamic Range) SDR: Standard dynamic range (Standard Dynamic Range) VVC: Versatile Video Coding (Versatile Video Encoding) JVET: Joint Video Exploration Team (Joint Video Exploration Team) GPM: Geometric partition mode (Geometric Partition Mode) IBC, IntraBC: Intra block copy (Intra Block Copy) IntraTMP: Intra template matching (Intra Template Matching) PDPC: Position-Dependent Predictor Combinations (Position-Dependent Predictor Combinations) HLS: High-Level Syntax (High-Level Syntax)

[0126] The above disclosure also includes the embodiments listed below.

[0127] (1) A method executed by at least one processor in a video decoder, comprising: receiving an encoded video bitstream including a current picture having at least one block located at a boundary of the current picture, encoded according to one of (i) an intra block copy (IBC) mode, (ii) an intra template matching (intraTMP) mode, and (iii) an IBC merge mode; determining whether boundary filtering is enabled for the at least one block based on syntax elements in the received encoded video stream; based on a determination that boundary filtering is enabled: filtering one or more boundary samples corresponding to the at least one block to generate one or more filtered samples, and decoding the at least one block based on the generated one or more filtered samples; and based on a determination that boundary filtering is not enabled, decoding the at least one block without filtering the one or more boundary samples. (2) The method according to item (1), wherein the syntax elements specify that boundary filtering is enabled based on a determination that the at least one block is encoded in one of the IBC mode and the intraTMP mode. (3) The method according to item (1) or (2), wherein the boundary filtering includes a position-dependent predictor combination (PDPC) filter. (4) The method according to any one of items (1) to (3), wherein the boundary filtering includes a weighted average of neighboring reconstructed samples and adjacent boundary prediction samples. (5) The method according to item (4), wherein the neighboring reconstructed samples include one or more columns located to the left of the at least one block. (6) The method according to item (5), wherein the number of columns located to the left of the at least one block used for the boundary filtering depends on the block size of the at least one block. (7) The method according to item (4), wherein the reconstructed sample in the vicinity includes one or more rows located on the at least one block. (8) The method according to item (7), wherein the number of rows located on the at least one block used in the boundary filtering depends on the block size of the at least one block. (9) The method according to any one of items (1) to (8), wherein the at least one block is encoded in the IBC mode, and the boundary filtering is specified to be disabled based on a determination that a template matching cost calculated based on a template area pointed to by a block vector of the at least one block is less than or equal to a threshold value for the syntax element. (10) The method according to item (9), wherein the boundary filtering includes a position-dependent predictor combination (PDPC) filter, and an S parameter of the PDPC filter depends on the template matching cost. (11) The method according to any one of items (1) to (8) and (10), wherein the at least one block is encoded in the intraTMP mode, and the boundary filtering is specified to be disabled based on a determination that a template matching cost of a template used in the intraTMP mode is less than or equal to a threshold value for the syntax element. (12) The method according to item (11), wherein the boundary filtering includes a position-dependent predictor combination (PDPC) filter, and an S parameter of the PDPC filter depends on the template matching cost. (13) The method according to any one of items (1) to (12), wherein the syntax element specifies that the boundary filtering is disabled based on a determination that the at least one block is encoded in the IBC merge mode. (14) The syntax element specifies that the boundary filtering is enabled based on a determination that the at least one block is encoded in the IBC merge mode. The boundary filtering uses the reconstructed neighboring samples located to the left of the at least one block, and the block vector of the at least one block is derived from the spatial candidates located above the at least one block or diagonally above and to the right of the at least one block. The method according to any one of items (1) to (13). (15) The syntax element specifies that the boundary filtering is enabled based on a determination that the at least one block is encoded in the IBC merge mode. The boundary filtering uses the reconstructed neighboring samples located above the at least one block, and the block vector of the at least one block is derived from the spatial candidates located to the left of the at least one block or diagonally below and to the left of the at least one block. The method according to any one of items (1) to (13). (16) At least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate as instructed by the computer program code, a video decoder comprising: the computer program code causes the at least one processor to: (i) receive an encoded video bitstream including a current picture having at least one block located at a boundary of the current picture encoded according to one of an intra-block copy (IBC) mode, (ii) an intra-template matching (intraTMP) mode, and (iii) an IBC merge mode; a determination code configured to cause the at least one processor to determine whether boundary filtering is to be enabled for the at least one block based on syntax elements in the received encoded video stream; filtering code; and determination code, based on a determination that boundary filtering is to be enabled: the filtering code is configured to cause the at least one processor to filter one or more boundary samples corresponding to the at least one block to generate one or more filtered samples, and the decoding code is configured to cause the at least one processor to decode the at least one block based on the one or more generated filtered samples, and based on a determination that boundary filtering is not to be enabled, the decoding code is configured to cause the at least one processor to decode the at least one block without filtering the one or more boundary samples, a video decoder. (17) The video device according to item (16), wherein the syntax elements specify that the boundary filtering is to be enabled based on a determination that the at least one block is encoded in one of the IBC mode and the intraTMP mode. (18) The video device according to item (16) or (17), wherein the boundary filtering includes a position-dependent predictor combination (PDPC) filter. (19) The video device according to any one of (16) to (18), wherein the boundary filtering includes a weighted average of neighboring reconstruction samples and adjacent boundary prediction samples. (20) A non-transitory computer-readable medium storing instructions, which, when executed by a processor in a video decoder, cause the processor to execute a method, the method comprising: receiving an encoded video bitstream including a current picture having at least one block located at a boundary of the current picture, the at least one block being encoded according to one of (i) an intra block copy (IBC) mode, (ii) an intra template matching (intraTMP) mode, and (iii) an IBC merge mode; determining whether boundary filtering is to be enabled for the at least one block based on syntax elements in the received encoded video stream; based on a determination that boundary filtering is to be enabled: filtering one or more boundary samples corresponding to the at least one block to generate one or more filtered samples, and decoding the at least one block based on the generated one or more filtered samples; and based on a determination that boundary filtering is not to be enabled, decoding the at least one block without filtering the one or more boundary samples.

Claims

1. A method performed by at least one processor in a video decoder, comprising: receiving an encoded video bitstream including a current picture having at least one block located at a boundary of the current picture, encoded according to one of (i) an intra block copy (IBC) mode, (ii) an intra template matching (intraTMP) mode, and (iii) an IBC merge mode; determining, based on syntax elements in the received encoded video stream, whether boundary filtering is enabled for the at least one block; based on a determination that boundary filtering is enabled: filtering one or more boundary samples corresponding to the at least one block to generate one or more filtered samples; decoding the at least one block based on the one or more generated filtered samples; based on a determination that boundary filtering is not enabled, decoding the at least one block without filtering the one or more boundary samples. A method.

2. The method of claim 1, wherein the syntax elements specify that boundary filtering is enabled based on a determination that the at least one block is encoded in one of the IBC mode and the intraTMP mode.

3. The method of claim 1, wherein the boundary filtering includes a position dependent predictor combination (PDPC) filter.

4. The method of claim 1, wherein the boundary filtering includes a weighted average of neighboring reconstructed samples and adjacent boundary prediction samples.

5. The method of claim 4, wherein the neighboring reconstructed samples include one or more columns located to the left of the at least one block.

6. The method of claim 5, wherein the number of columns located to the left of the at least one block used for the boundary filtering depends on the block size of the at least one block.

7. The method of claim 4, wherein the neighboring reconstructed samples include one or more rows located above the at least one block.

8. The method according to claim 7, wherein the number of lines located on top of the at least one block used in the boundary filtering depends on the block size of the at least one block.

9. The method according to claim 1, wherein the at least one block is encoded in IBC mode, and the boundary filtering is specified to be disabled based on a determination that a template matching cost calculated based on a template area pointed to by a block vector of the at least one block is below a threshold value.

10. The method according to claim 9, wherein the boundary filtering includes a position-dependent predictor combination (PDPC) filter, and an S parameter of the PDPC filter depends on the template matching cost.

11. The method according to claim 1, wherein the at least one block is encoded in intraTMP mode, and the boundary filtering is specified to be disabled based on a determination that a template matching cost of a template used in the intraTMP mode is below a threshold value.

12. The method according to claim 11, wherein the boundary filtering includes a position-dependent predictor combination (PDPC) filter, and an S parameter of the PDPC filter depends on the template matching cost.

13. The method according to claim 1, wherein the syntax element specifies that the boundary filtering is disabled based on a determination that the at least one block is encoded in IBC merge mode.

14. The method according to claim 1, wherein the syntax element specifies that the boundary filtering is enabled based on a determination that the at least one block is encoded in IBC merge mode, the boundary filtering uses a reconstructed neighboring sample located to the left of the at least one block, and a block vector of the at least one block is derived from a spatial candidate located above or in the upper right of the at least one block.

15. The syntax element specifies that the boundary filtering is enabled based on a determination that the at least one block is encoded in an IBC merge mode, the boundary filtering uses a reconstructed neighboring sample located above the at least one block, and a block vector of the at least one block is derived from a spatial candidate located to the left of the at least one block or to the lower left of the at least one block, the method according to claim 1.

16. at least one memory configured to store computer program code; at least one processor accessing the computer program code and configured to operate as commanded by the computer program code, a video decoder comprising: the computer program code causes the at least one processor to execute the method according to any one of claims 1 to 15; video decoder.

17. A computer program for causing a processor to execute the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Usage of templates for decoder-side intra mode derivation

    US11388421B1

  • Nter PDPC mode

    US20200296421A1

  • Motion compensation boundary filtering

    US20210360285A1

  • High level control of PDPC and intra reference sample filtering of video coding

    US20210368170A1

  • Image encoding device, image decoding device, and programs for same

    WO2017030200A1