Method, apparatus, and computer program product for video encoding and decoding

By using directional intra prediction or intra prediction methods in video encoding to generate extended reference frames, the filling problem when the motion vector points to the area outside the boundary of the reference frame is solved, and the motion compensation efficiency and video encoding performance are improved.

CN119999210APending Publication Date: 2025-05-13NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069616.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-06-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During the motion compensation process of existing video encoding technology, when the motion vector points to an area outside the reference frame, it is difficult to effectively fill and process the area outside the boundary, resulting in poor motion compensation performance.

Method used

The extended reference frame is generated by using the boundary samples of the reference frame for directional intra prediction, or using the decoded samples of the current frame for intra prediction, thereby filling and processing samples pointing to areas outside the motion vector.

Benefits of technology

The motion compensation efficiency is improved, especially when the motion vector points to the area outside the reference frame boundary, the video block can be predicted and encoded more accurately, thereby improving the overall performance of video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119999210A_ABST
    Figure CN119999210A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a coding method and technical equipment for implementing the method. The method comprises: determining (1610) a motion vector and a reference frame of a current block; determining (1620) when the motion vector points to an area outside the reference frame; an extended reference frame is generated (1630) by filling a region outside the reference frame with one or more of: directional intra prediction using boundary samples of the reference frame according to the direction of the motion vector; an intra prediction method using a decoded sample within the current frame; an intra prediction sample of the current block; predicting (1640) the current block by using motion compensation from the extended reference frame; and encoding (1650) the current block into a bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present solution generally relates to video encoding and video decoding. In particular, the present solution relates to motion compensation in video encoding and decoding. Background Art

[0002] This section is intended to provide a background or context to the inventions described in the claims. The description herein may include concepts that could be pursued, but not necessarily concepts that have been previously conceived or pursued. Therefore, unless otherwise indicated herein, what is described in this section is not prior art to the description and claims in this application and is not admitted to be prior art by inclusion in this section.

[0003] A video coding system may include an encoder that transforms an input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form, e.g., to enable storage / transmission of the video information at a lower bit rate than might otherwise be required. Summary of the invention

[0004] The independent claims define the scope of protection of various embodiments of the present invention. Embodiments and features described in this specification that do not belong to the scope of the independent claims (if any) are to be construed as examples that help understand the various embodiments of the present invention.

[0005] Various aspects include a method, an apparatus, and a computer readable medium comprising a computer program stored therein, which are characterized by what is stated in the independent claims. Various embodiments are disclosed in the dependent claims.

[0006] According to a first aspect, there is provided an apparatus comprising: means for determining a motion vector and a reference frame for a current block; means for determining when a motion vector points to an area outside the reference frame; means for generating an extended reference frame by filling the area outside the reference frame with one or more of the following:

[0007] ○ Oriented frame using boundary samples of reference frame according to the direction of motion vector

[0008] Internal prediction;

[0009] ○ Intra-frame prediction method using decoded samples within the current frame;

[0010] ○ Intra-frame prediction samples of the current block;

[0011] means for predicting a current block by using motion compensation from an extended reference frame; and means for encoding the current block into a bitstream.

[0012] According to a second aspect, a method is provided, the method comprising: determining a motion vector and a reference frame of a current block; determining when the motion vector points to an area outside the reference frame; generating an extended reference frame by filling the area outside the reference frame with one or more of the following:

[0013] ○ Oriented frame using boundary samples of reference frame according to the direction of motion vector

[0014] Internal prediction;

[0015] ○ Intra-frame prediction method using decoded samples within the current frame;

[0016] ○ Intra-frame prediction samples of the current block;

[0017] Predicting a current block by using motion compensation from an extended reference frame; and encoding the current block into a bitstream.

[0018] According to a third aspect, there is provided an apparatus comprising at least one processor, a memory comprising computer program code, the memory and the computer program code being configured to, with the at least one processor, cause the apparatus to perform at least the following: determining a motion vector and a reference frame for a current block; determining when a motion vector points to an area outside the reference frame; generating an extended reference frame by filling the area outside the reference frame with one or more of the following:

[0019] ○ Oriented frame using boundary samples of reference frame according to the direction of motion vector

[0020] Internal prediction;

[0021] ○ Intra-frame prediction method using decoded samples within the current frame;

[0022] ○ Intra-frame prediction samples of the current block;

[0023] Predicting a current block by using motion compensation from an extended reference frame; and encoding the current block into a bitstream.

[0024] According to a fourth aspect, there is provided a computer program product comprising computer program code configured to, when executed on at least one processor, cause an apparatus or system to: determine a motion vector and a reference frame for a current block; determine when a motion vector points to an area outside the reference frame; generate an extended reference frame by filling the area outside the reference frame with one or more of the following:

[0025] ○ Oriented frame using boundary samples of reference frame according to the direction of motion vector

[0026] Internal prediction;

[0027] ○ Intra-frame prediction method using decoded samples within the current frame;

[0028] ○ Intra-frame prediction samples of the current block;

[0029] Predicting a current block by using motion compensation from an extended reference frame; and encoding the current block into a bitstream.

[0030] According to one embodiment, the intra prediction method is determined by using a texture analysis method.

[0031] According to one embodiment, the texture analysis method is a decoder-side intra mode derivation method.

[0032] According to one embodiment, the intra prediction method is determined by using a template-based intra mode derivation method.

[0033] According to one embodiment, samples of an area outside a picture are predicted by cross-component prediction, wherein samples of an area outside a picture in a reference channel are used to predict samples of an area outside a picture in a current channel.

[0034] According to one embodiment, boundary samples between intra prediction and motion compensated prediction are filtered in the final prediction. According to one embodiment, the computer program product is embodied on a non-transitory computer readable medium. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Hereinafter, various embodiments will be described in more detail with reference to the accompanying drawings, in which

[0036] Figure 1 An example of the positions of the left and above samples of the current block involved in CCLM mode is shown.

[0037] Figure 2a An example of deriving the chroma prediction mode from the luma mode when CCLM is enabled is shown.

[0038] Figure 2b An example of a unified binarization table for chroma prediction mode is shown.

[0039] Figure 3a Two luma-to-chroma models obtained for a luma Y threshold of 17 are shown.

[0040] Figure 3b An example of the correspondence of each luminance-to-chrominance model to the spatial segmentation of the content is shown.

[0041] Figure 4 An example of the location of samples used to derive the CCCM filter is shown.

[0042] Figure 5Examples of various filter kernels are shown;

[0043] Figure 6 An example of four reference lines adjacent to a prediction block is shown.

[0044] Figure 7 An example of a matrix-weighted intra prediction process is shown.

[0045] Figure 8 An example of HoG calculation from a template of 3 pixels in width is shown.

[0046] Fig. 9 An example of a low frequency non-separable transform (LFNST) process is shown.

[0047] Fig.10 An example of using intra prediction to generate OOB region samples is shown.

[0048] Fig.11 An example of predicting only a portion of the OOB region samples is shown.

[0049] Fig.12 An example of a motion vector pointing to a block of an OOB region in a reference frame is shown.

[0050] Fig.13 An example of combining motion compensated prediction and intra prediction from neighboring reference samples when the motion vector points to an OOB region in a reference frame is shown.

[0051] Fig.14 An example is shown where motion compensated samples are used together with reconstructed reference samples.

[0052] Fig.15 An example of directionally filling OOB region samples in a reference picture is shown.

[0053] Fig.16 is a flow chart illustrating a method according to an embodiment.

[0054] Fig.17 An apparatus according to an embodiment is shown.

[0055] Fig.18 An encoding process according to an embodiment is shown.

[0056] Fig.19 A decoding process according to an embodiment is shown. DETAILED DESCRIPTION

[0057] The following description and drawings are illustrative and should not be construed as unnecessary limitations. Specific details are provided for a thorough understanding of the present disclosure. However, in some cases, in order to avoid obscuring the description, well-known or conventional details are not described. References to an embodiment or embodiments in the present disclosure may be, but are not necessarily, references to the same embodiment, and such references represent at least one embodiment.

[0058] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure.

[0059] In the following, several embodiments will be described in the context of a video coding arrangement. However, it should be noted that the present embodiments are not necessarily limited to this particular arrangement. The embodiments relate to boundary samples in motion compensation and methods for processing the same.

[0060] The Advanced Video Coding standard (which may be abbreviated as AVC or H.264 / AVC) was developed by the Video Coding Experts Group (VCEG) of the Telecommunication Standardization Sector of the International Telecommunication Union (ITU-T) and the Joint Video Team (JVT) of the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is issued by two parent standardization organizations and is known as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There are multiple versions of the H.264 / AVC standard, each of which integrates new extensions or features into the specification. These extensions include scalable video coding (SVC) and multi-view video coding (MVC).

[0061] The High Efficiency Video Coding standard (which may be abbreviated as HEVC or H.265 / HEVC) is developed by the Joint Collaboration Team on Video Coding (JCT-VC) of VCEG and MPEG. The standard is issued by two superior standardization organizations and is referred to as ITU-TH.265 Recommendation and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Extensions to H.265 / HEVC include scalable, multi-view, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV-HEVC, 3D-HEVC, and REXT, respectively. Unless otherwise stated, references to H.265 / HEVC, SHVC, MV-HEVC, 3D-HEVC, and REXT in this specification for understanding the definitions, structures, or concepts of these standard specifications should be understood as references to the latest versions of these standards available before the date of this application.

[0062] Versatile Video Coding (which may be abbreviated as VVC, H.266, or H.26 / VVC) is a video compression standard developed as a successor to HEVC. VVC is specified in the ITU-T H.266 recommendation and equivalently specified in ISO / IEC 23090-3, which is also known as MPEG-I Part 3.

[0063] The Alliance for Open Media (AOM) developed the specifications for the AV1 bitstream format and decoding process. The AV1 specification was released in 2018. AOM is reportedly working on the AV2 specification.

[0064] This section describes some key definitions, bitstreams, and coding structures and concepts of H.264 / AVC, HEVC, VVC, and / or AV1 and some extensions thereof, taking video encoders, decoders, encoding methods, decoding methods, and bitstream structures as examples, in which embodiments may be implemented. Aspects of the various embodiments are not limited to H.264 / AVC, HEVC, VVC, and / or AV1 or their extensions, but are described for a possible basis on which the embodiments may be partially or fully implemented.

[0065] A video codec may include an encoder that transforms an input video into a compressed representation suitable for storage / transmission, and a decoder that can decompress the compressed video representation back into a visual form. The compressed representation may be referred to as a bitstream or a video bitstream. The video encoder and / or video decoder may also be separate from each other, i.e., need not form a codec. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bit rate). The notation "encoder (decoder)" refers to an encoder and / or a decoder.

[0066] Hybrid video codecs (e.g., ITU-T H.263, H.264 / AVC, and HEVC) can encode video information in two stages. First, the pixel values ​​in a certain picture area (or "block") are predicted, for example, by motion compensation means (finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being encoded) or by spatial means (using the pixel values ​​around the block to be encoded in a specified way). In the first stage, predictive coding can be applied, for example, as so-called sample prediction and / or so-called grammatical prediction.

[0067] In sample prediction, the pixel or sample values ​​in a certain picture area or "block" are predicted. For example, these pixel or sample values ​​may be predicted using one or more motion compensation or intra prediction mechanisms.

[0068] The motion compensation mechanism (which may also be referred to as inter-frame prediction, temporal prediction, motion compensated temporal prediction, or motion compensated prediction, or MCP) involves finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being encoded. Inter-frame prediction can reduce temporal redundancy.

[0069] Intra-frame prediction (where pixel sample values ​​can be predicted by a spatial mechanism) involves finding and indicating spatial region relationships. Intra-frame prediction exploits the fact that neighboring pixels within the same picture can be correlated. Intra-frame prediction can be performed in the spatial domain or in the transform domain, i.e., sample values ​​or transform coefficients can be predicted. Intra-frame prediction can be used for intra-frame coding, where inter-frame prediction is not applied.

[0070] In syntax prediction (which may also be referred to as parameter prediction), syntax elements and / or syntax element values ​​and / or variables derived from syntax elements are predicted from earlier encoded (decoded) syntax elements and / or earlier derived variables. A non-limiting example of syntax prediction is provided below.

[0071] In motion vector prediction, motion vectors for inter-frame and / or inter-view prediction, for example, can be differentially encoded relative to block-specific predicted motion vectors. In many video codecs, the predicted motion vector is created in a predefined manner, for example, by calculating the median of the encoded or decoded motion vectors of neighboring blocks. Another method for creating motion vector predictions, sometimes called advanced motion vector prediction (AMVP), is to generate a candidate prediction list from neighboring blocks and / or collocated blocks in a temporal reference picture, and signal the selected candidate as a motion vector predictor. In addition to predicting motion vector values, reference indexes of previously encoded / decoded pictures can also be predicted. The reference index can be predicted based on neighboring blocks and / or collocated blocks in a temporal reference picture. Differential encoding of motion vectors can be disabled across slice boundaries. Block partitions can be predicted, for example from coding tree units (CTUs) to coding units (CUs) to prediction units (PUs).

[0072] In filter parameter prediction, filter parameters, for example, used for sample adaptive offset, may be predicted.

[0073] Prediction methods using picture information from previously coded pictures may also be referred to as inter-frame prediction methods, which may also be referred to as temporal prediction and motion compensation.

[0074] Prediction methods using image information within the same image may also be referred to as intra-frame prediction methods.

[0075] Second, the prediction error (i.e., the difference between the predicted pixel block and the original pixel block) is encoded. This can be achieved by transforming the difference in pixel values ​​using a specified transform (e.g., discrete cosine transform (DCT) or its variants), quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and the size of the final encoded video representation (file size for transmission bitrate).

[0076] In most cases, the corresponding basic unit for the input of the encoder and the output of the decoder is a picture. The picture as the input of the encoder can also be called a source picture, and the picture decoded by the decoder can be called a decoded picture or a reconstructed picture.

[0077] The source picture and the decoded picture each consist of one or more sample arrays, such as one of the following sets of sample arrays:

[0078] -Luma(Y) only (monochrome).

[0079] - Luma and two chrominances (YCbCr or YCgCo).

[0080] Green, Blue, and Red (GBR, also called RGB).

[0081] - arrays representing other unspecified monochromatic or tristimulus color sampling (e.g., YZX,

[0082] Also known as XYZ).

[0083] In the following, these arrays may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr; regardless of the actual color representation method used. The actual color representation method in use may be indicated, for example, in the coded bitstream, such as using the Video Usability Information (VUI) syntax of HEVC, etc. A component may be defined as an array or a single sample from one of the three sample arrays (luminance and two chroma), or an array or a single sample of an array that makes up a picture in monochrome format.

[0084] A picture can be defined as a frame or a field. A frame consists of a matrix of luma samples and possibly corresponding chroma samples. A field is a set of alternating sample rows of a frame and can be used as encoder input when the source signal is interlaced. The chroma sample array may not be present (and therefore monochrome sampling may be used), or the chroma sample array may be subsampled compared to the luma sample array.

[0085] The decoder reconstructs the output video by applying a prediction component similar to the encoder to form a predicted representation of the pixel block (using the motion or spatial information created by the encoder and stored in the compressed representation) and applying prediction error decoding (the inverse operation of prediction error encoding, which recovers the quantized prediction error signal in the spatial pixel domain). After applying the prediction and prediction error decoding components, the decoder adds the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) may also apply additional filtering components to improve the quality of the output video before passing it on for display and / or storing it as a prediction reference for upcoming frames in the video sequence.

[0086] Motion information can be indicated by a motion vector associated with each motion compensated image block in the video codec. Each of these motion vectors represents the displacement of an image block in a picture to be encoded (on the encoder side) or decoded (on the decoder side) and a prediction source block in one of the previously encoded or decoded images (or pictures). Like many other video compression standards, H.264 / AVC and HEVC divide the picture into a rectangular grid, and for each rectangular grid, a similar block in one of the reference pictures is indicated for inter-frame prediction. The position of the prediction block is encoded as a motion vector that indicates the position of the prediction block relative to the block being encoded.

[0087] A bitstream can be defined as a sequence of bits or a sequence of syntactic structures. A bitstream format can constrain the order of syntactic structures in a bitstream.

[0088] A syntax element may be defined as a data element represented in a bitstream. A syntax structure may be defined as zero or more syntax elements that appear together in a specified order in a bitstream.

[0089] In some coding formats or standards, the bitstream may be in the form of a stream of Network Abstraction Layer (NAL) units or a byte stream that forms a representation of coded pictures and associated data to form one or more coded video sequences.

[0090] A NAL unit may be defined as a syntax structure containing an indication of the type of data to follow, and bytes containing that data in the form of an RBSP, interspersed with start code emulation prevention bytes as necessary. A Raw Byte Sequence Payload (RBSP) may be defined as a syntax structure containing an integer number of bytes encapsulated in a NAL unit. An RBSP is either empty or in the form of a data bit string containing the syntax elements, followed by an RBSP stop bit, followed by zero or more subsequent bits equal to 0.

[0091] The NAL unit includes a header and a payload. The NAL unit header indicates the type of the NAL unit, etc.

[0092] In some coding formats (such as AV1), a bitstream may include a sequence of open bitstream units (OBUs). The OBU includes a header and a payload, where the header identifies the type of the OBU. In addition, the header may include the size of the payload in bytes.

[0093] In the claims and described embodiments, the phrase with the bitstream (e.g., with a bitstream indication) or with a coding unit of a bitstream (e.g., with a coding block indication) may be used to refer to transmission, signaling, or storage in a manner that "out-of-band" data is associated with the bitstream or coding unit, respectively, but is not included therein. The phrases decode with the bitstream or with a coding unit of a bitstream, etc. may refer to decoding the referenced out-of-band data associated with the bitstream or coding unit, respectively (which may be obtained from the out-of-band transmission, signaling, or storage). For example, the phrase with the bitstream may be used when the bitstream is contained in a container file, such as a file conforming to the ISO base media file format, and certain file metadata is stored in the file in a manner that associates the metadata with the bitstream, such as a box in a sample entry of a track containing the bitstream, a sample group of a track containing the bitstream, or a timed metadata track associated with a track containing the bitstream.

[0094] Below, the partitioning of pictures into sub-pictures, slices and tiles according to H.266 / VVC will be described in more detail. Similar concepts can also be applied to other video coding specifications.

[0095] A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of coding tree units (CTUs) covering a rectangular area of ​​a picture. The CTUs in a tile are scanned in raster scan order within the tile.

[0096] A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture. Therefore, every vertical slice boundary is always also a vertical tile boundary. The horizontal boundaries of a slice can be not tile boundaries, but instead consist of horizontal CTU boundaries within a tile; this occurs when a tile is split into multiple rectangular slices, where each slice consists of an integer number of consecutive complete CTU rows within the tile.

[0097] Two slicing modes are supported, namely raster scan slicing mode and rectangular slicing mode. In raster scan slicing mode, a slice contains a sequence of complete tiles in a tile raster scan of a picture. In rectangular slicing mode, a slice contains multiple complete tiles that together form a rectangular area of ​​a picture, or multiple consecutive complete CTU rows of a tile that together form a rectangular area of ​​a picture. Tiles within a rectangular slice are scanned in tile raster scan order within the rectangular area corresponding to the slice.

[0098] A sub-picture can be defined as a rectangular area of ​​one or more slices within a picture, where one or more slices are complete. Thus, a sub-picture consists of one or more slices that together cover a rectangular area of ​​the picture. Thus, every sub-picture boundary is also always a slice boundary, and every vertical sub-picture boundary is also a vertical tile boundary. Slices of a sub-picture may need to be rectangular slices.

[0099] For each sub-picture and tile, one or both of the following conditions may be satisfied: i) All CTUs in a sub-picture belong to the same tile. ii) All CTUs in a tile belong to the same sub-picture.

[0100] Below, the partitioning of pictures into tiles and tile groups according to AV1 will be described in more detail. Similar concepts can also be applied to other video coding specifications.

[0101] A tile consists of an integer number of complete superblocks that together form a complete rectangular area of ​​the picture. Intra-picture prediction is disabled across tile boundaries. The minimum tile size is one superblock, and the maximum tile size in the currently specified level is 4096×2304 in terms of luma sample count. A picture is partitioned into a tile grid of one or more tile rows and one or more tile columns. Tile grids with uniform tile sizes or non-uniform tile sizes can be signaled in the picture header, in the latter case the tile row height and tile column width are signaled. Superblocks in a tile are scanned in raster scan order within that tile.

[0102] A tile group OBU carries one or more complete tiles. The first tile and the last tile in a tile group OBU may be indicated in the tile group OBU before the encoded tile data. Tiles within a tile group OBU may appear in a tile raster scan of a picture.

[0103] The features and coding tools included in VVC include the following:

[0104] - Intra-frame prediction

[0105] ○ 67 intra-frame modes, with wide-angle mode expansion

[0106] ○ Block size and mode dependent 4-tap interpolation filter

[0107] ○Position-dependent intra prediction combination (PDPC)

[0108] ○Cross-Component Linear Model Intra Prediction (CCLM)

[0109] ○Multi-reference intra prediction

[0110] ○ Intra-frame sub-partition

[0111] ○ Weighted intra prediction with matrix multiplication

[0112] -Inter-picture prediction

[0113] ○ Block motion replication using spatial, temporal, history-based, and pairwise average merge candidates

[0114] ○Affine motion inter-frame prediction

[0115] ○ Sub-block based temporal motion vector prediction

[0116] ○Adaptive motion vector resolution

[0117] ○ 8×8 block-based motion compression for temporal motion prediction

[0118] ○ High-precision (1 / 16 pixel) motion vector storage and motion compensation, using an 8-tap interpolation filter for the luminance component and a 4-tap interpolation filter for the chrominance component

[0119] ○Triangular partition

[0120] ○ Combined intra and inter prediction

[0121] ○Combined with MVD (MMVD)

[0122] ○Symmetrical MVD encoding

[0123] ○ Bidirectional optical flow

[0124] ○Decoder side motion vector refinement

[0125] ○Use CU-level weights for bidirectional prediction

[0126] - Transformation, quantization, and coefficient coding

[0127] ○ Multiple main transform options of DCT2, DST7, and DCT8

[0128] ○ Secondary transformation in low frequency region

[0129] ○ Sub-block transform of inter-frame prediction residual

[0130] ○ Maximum QP dependent quantization increased from 51 to 63

[0131] ○ Transform coefficient coding with sign data hiding

[0132] ○ Transform skip residual coding

[0133] - Entropy coding

[0134] ○ Arithmetic coding engine with adaptive dual-window probability update

[0135] - In-loop filter

[0136] ○Intra-ring plastic surgery

[0137] ○ Deblocking filter with robust long filter

[0138] ○ Sample Adaptive Offset

[0139] ○Adaptive loop filter

[0140] -Screen content encoding

[0141] ○The current image reference has reference area restrictions

[0142] -360-degree video encoding

[0143] ○Horizontal surround motion compensation

[0144] - Advanced syntax and parallel processing

[0145] ○ Reference picture management via direct reference picture list signaling

[0146] ○ Tile groups with rectangular tile groups

[0147] In H.26 / VVC, the following block partitioning is applied. Pictures are partitioned into CTUs. Pictures can also be divided into slices, tiles, bricks, and sub-pictures. CTUs can be split into smaller CUs using a quadtree structure. Each CU can be partitioned using quadtrees and nested multi-type trees, including ternary and binary splits.

[0148] There are specific rules to infer partitions in the image boundaries.

[0149] Redundant split patterns are not allowed in nested build type partitions.

[0150] To reduce cross-component redundancy, VVC uses a cross-component linear model (CCLM) prediction mode, in which chrominance samples are predicted based on the reconstructed luma samples of the same CU by using the following linear model:

[0151] pred C (i, j) = α·rec L ′(i, j)+β

[0152] where pred C (i,j) represents the predicted chroma sample in CU, rec L ′(i,j) represents the downsampled reconstructed luma sample of the same CU.

[0153] The CCLM parameters (α and β) are derived from at most four adjacent chroma samples and their corresponding downsampled luma samples. Assuming the current chroma block size is W×H, W' and H' are set to

[0154] - When LM mode is applied, W'=W, H'=H;

[0155] - When LM-A mode is applied, W'=W+H;

[0156] - When LM-L mode is applied, H'=H+W.

[0157] The upper adjacent positions are denoted as S[0, -1], ..., S[W'-1, -1], and the left adjacent positions are denoted as S[-1, 0], ..., S[-1, H'-1]. Then, four samples are selected as follows:

[0158] - When LM mode is applied and both upper and left neighboring samples are available, S[W' / 4, -1], S[3*W' / 4, -1], S[-1, H' / 4], S[-1, 3*H' / 4];

[0159] - When LM-A mode is applied or only upper neighboring samples are available, S[W' / 8, -1], S[3*W' / 8, -1], S[5*W' / 8, -1], S[7*W' / 8, -1];

[0160] - When LM-L mode is applied or only left neighboring samples are available, S[-1, H' / 8], S[-1, 3*H' / 8], S[-1, 3*H' / 8], S[-1, 7*H' / 8].

[0161] The four adjacent luma samples at the selected position are downsampled and compared four times to find two smaller values: x0A and x1A, and two larger values: x0B and x1B. Their corresponding chroma sample values ​​are denoted as y0A, y1A, y0B, and y1B. Then, xA, xB, yA, and yB are derived as follows

[0162] X a =(x 0 A +x 1 A +1)>>1;

[0163] X b =(x 0 B +x 1 B +1)>>1;

[0164] Y a =(y 0 A +y 1 A +1)>>1;

[0165] Y b =(y0 B +y 1 B +1)>>1

[0166] Finally, the linear model parameters are obtained according to the following equation:

[0167]

[0168] β=Y b -α·X b

[0169] Figure 1 An example of the positions of the left and upper samples and the samples of the current block involved in the CCLM mode is shown. The division operation for calculating the parameter α is implemented by a lookup table. In order to reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented in exponential notation. For example, diff is approximated with a 4-bit significant part and an exponent. Therefore, the table of 1 / diff is reduced to 16 elements of 16 significant values, as shown below:

[0170] DivTable[ ]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}

[0171] This will help reduce the complexity of the calculations and the amount of memory required to store the required tables.

[0172] In addition to the upper template and the left and side templates being used together to calculate the linear model coefficients, they can also be used alternately in other 2LM modes, called LM_A and LM_L modes.

[0173] In LM_A mode, only the upper template is used to calculate the linear model coefficients. In order to obtain more samples, the upper template is expanded to (W+H). In LM_L mode, only the left template is used to calculate the linear model coefficients. In order to obtain more samples, the left template is expanded to (H+W).

[0174] For non-square blocks, the upper template is expanded to W+W and the left template is expanded to H+H.

[0175] To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to the luma samples to achieve a 2 to 1 downsampling ratio in both the horizontal and vertical directions. The choice of downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "type 0" and "type 2" content, respectively.

[0176]

[0177] It should be understood that when the upper reference line is located at a CTU boundary, only one luma line (common line buffer in intra prediction) is used to make the downsampled luma samples.

[0178] This parameter calculation is performed as part of the decoding process, not just as an encoder search operation. Therefore, no syntax is used to communicate the α and β values ​​to the decoder.

[0179] For chroma intra mode coding, a total of eight intra modes are allowed for chroma intra mode coding. These modes include five traditional intra modes and three cross-component linear model modes (CCLM, LM_A, and LM_L). The chroma mode signaling and derivation process is as follows Figure 2a As shown in Table 1 in . Chroma mode coding directly depends on the intra prediction of the corresponding luminance block. Since the separate block partition structure of luminance and chrominance components is enabled in I slices, a chrominance block can correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chrominance block is directly inherited.

[0180] Use a single binarization table regardless of the value of sps_cclm_enabled_flag, such as Figure 2b As shown in Table 2 in . In Table 2, the first interval indicates whether it is normal (0) or LM mode (1). If it is LM mode, the next interval indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next 1 interval indicates whether it is LM_L (0) or LM_A (1). For this case, when sps_cclm_enabled_flag is 0, the first interval of the binarization table corresponding to intra_chroma_pred_mode can be discarded before entropy coding. Or, in other words, the first interval is inferred to be 0 and therefore not encoded. This single binarization table is used for both cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two intervals in Tables 3-4 are context encoded with their own context models, and the remaining intervals are bypassed for encoding.

[0181] Additionally, to reduce luma-chroma latency in dual trees, when a 64×64 luma CT node is partitioned using unsplit (and ISP is not used for 64×64 CUs) or QT, chroma CUs in 32×32 / 32×16 chroma CT nodes are allowed to use CCLM as follows:

[0182] - If a 32×32 chroma node is not split or partitioned QT split, all chroma CUs in the 32×32 node can use CCLM

[0183] - If the 32×32 chroma node is partitioned using horizontal BT and the 32×17 child nodes are not split or use vertical BT split, all chroma CUs in the 32×16 chroma node can use CCLM.

[0184] Under all other luma and chroma coding tree split conditions, CCLM is not allowed for chroma CUs.

[0185] The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, the reconstructed neighboring samples are divided into two categories using a threshold value, which is the average value of the brightness reconstructed neighboring samples. The linear model for each category is derived using the least mean square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. Figure 3a The figure shows two luminance-to-chrominance models obtained when the luminance (Y) threshold is 17. Each luminance-to-chrominance model has its own linear model parameters α and β. Figure 3b As shown, each luminance-to-chrominance model corresponds to a spatial segmentation of the content (ie, they correspond to different objects or textures in the scene).

[0186] An improved version of cross-component prediction, called convolutional cross-component model (CCCM), uses a 2D filter kernel to derive a luma-to-chroma model. The filter coefficients are derived at the decoder side using the reconstructed input data and the chroma sample set. For filter coefficient derivation, Figure 4 As shown in , collocated reference sample regions (consisting of reconstructed luma and chroma samples) are defined for luma and chroma, but any number of reference lines (which can be implemented by the encoder and decoder) can be used. In general, the reference samples can contain any chroma and luma samples reconstructed by both the encoder and the decoder. After the reference samples have been determined, different types of linear regression tools (such as ordinary least squares estimation, orthogonal matching pursuit, optimized orthogonal matching pursuit, ridge regression, or least absolute shrinkage and selection operators) can be used to derive the filter coefficients.

[0187] The size of the filter kernel can be, for example, 1×3 (1D vertical), 3×1 (1D horizontal), 3×3, 7×7, or any size, and can be shaped (by selecting only a subset of all possible kernel positions) into a cross, a diamond, or any given shape. When referring to samples in the filter kernel, the following notation is used: North (up), East (right), South (down), West (left), and Center, as in Figure 5 The letters N, E, S, W, and C are used. Figure 5 A 3-tap vertical kernel 501 , a 3-tap horizontal kernel 502 , a 5-tap cross kernel 503 , and a 25-tap diamond kernel 504 are illustrated.

[0188] The overall approach of reconstructing chrominance samples using convolution between the filter kernel obtained at the decoder side and the input data set is referred to herein as the convolutional cross-component model (CCCM). The following steps can be used to perform the CCCM operation:

[0189] - defining collocated reference regions on luma and chroma components;

[0190] - Downsample luma samples to match the chroma grid (optional);

[0191] - Scanning the luminance and chrominance samples of the reference region and collecting available statistics (such as autocorrelation matrix and cross-correlation vector) based on the filter shape;

[0192] - Solving for the filter coefficients by minimizing the squared error (or any other metric) based on available statistics (such as the autocorrelation matrix and the cross-correlation vector);

[0193] -The predicted chroma block is computed by convolving the downsampled luma samples with the filter kernel.

[0194] In the following, the (possibly downsampled) luma samples are defined as a 2D array Y(x,y) indexed using a horizontal x-coordinate and a vertical y-coordinate. Furthermore, the collocated chroma samples are defined as a 2D array C(x,y), and the filter kernel (i.e., coefficients) are defined as a 3x3 array F(i,j). At the sample level, the convolution between Y and F is defined as

[0195]

[0196] When other data terms are used (such as nonlinear square root terms), the additional convolution becomes

[0197]

[0198] in are the filter coefficients that are outside the 2D filter kernel, but have been obtained as part of the linear system of equations used to solve the 2D filter coefficients in step 4 above. Similarly, a bias term can be added to the convolution

[0199]

[0200] Angular intra prediction (also called directional intra prediction) can be performed by extrapolating sample values ​​from reconstructed reference samples with a given directionality. Reference samples may include the immediately above and to the right of the current block (when available) and the immediately to the left of the current block (when available), where availability may require an earlier decoding order than the current block and exist in the same image segment, such as in the same tile. To simplify the process, all sample positions within one prediction block may be projected to a single reference row or column, depending on the directionality of the selected prediction mode. The prediction samples within the block being encoded / decoded may be obtained by the following steps:

[0201] - Project the position of the prediction sample to a position within the reference row or column by applying the selected prediction direction. The position within the reference row or column may have fractional sampling precision, such as 1 / 32 pixel precision.

[0202] -Interpolate the values ​​of the sample positions on the reference row or column from the reference samples at the reference row / column.

[0203] Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Figure 6 In Figure 1, an example of 4 reference lines is depicted, where the samples of segments A and F are not extracted from the reconstructed neighboring samples, but are filled with the closest samples in segments B and E, respectively. HEVC intra-picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used.

[0204] The index of the selected reference line (mrl_idx) is signaled and used to generate the intra predictor. For reference line idx greater than 0, only other reference line modes are included in the MPM list, and only the mpm index is signaled without sending the remaining modes. The reference line index is signaled before the intra prediction mode, and in case a non-zero reference line index is signaled, the planar mode is excluded from the intra prediction mode.

[0205] MRL is disabled for the first row of blocks within a CTU to prevent the use of extended reference samples outside the current CTU line. In addition, PDPC is disabled when additional lines are used. For MRL mode, the derivation of DC values ​​in DC intra prediction mode for non-zero reference line index is consistent with the derivation of reference line index 0. MRL requires 3 adjacent luma reference lines stored using the CTU to generate the prediction. The Cross-Component Linear Model (CCLM) tool also requires 3 adjacent luma reference lines for its own downsampling filter. The definition of MLR using the same 3 lines is consistent with CCLM to reduce the storage requirements of the decoder.

[0206] Intra sub-partitioning (ISP) divides the luma intra prediction block vertically or horizontally into 2 or 4 sub-partitions depending on the block size. For example, the minimum block size of ISP is 4×8 (or 8×4). If the block size is larger than 4×8 (or 8×4), the corresponding block is divided into 4 sub-partitions. It has been noted that M×12 (M≤64) and 128×N (N≤64) ISP blocks may cause potential problems for 64×64 VDPU. For example, in the single-tree case, an M×128 CU has an M×128 luma TB and two corresponding Chroma TB. If the CU uses ISP, the luma TB will be split into four M×32 TBs (only horizontal splitting is possible), each TB is smaller than 64×64 blocks. However, in the current ISP design, the chroma blocks are not split. Therefore, the size of both chroma components is larger than 32×32 blocks.

[0207] Similarly, a 128×N CU using ISP may also produce similar situations. Therefore, both situations are a problem for the 64×64 decoder pipeline. Therefore, the CU size that can use ISP is limited to a maximum of 64×64. All sub-partitions meet this condition of having at least 16 samples.

[0208] The matrix weighted intra prediction (MIP) method is an intra prediction technique in VVC. To predict samples of a rectangular block of width W and height H, the matrix weighted intra prediction (MIP) takes a row of H reconstructed neighboring boundary samples on the left side of the block and a row of W reconstructed neighboring boundary samples above the block as input. If the reconstructed samples are not available, they are generated in the same way as traditional intra prediction. The generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication, and linear interpolation, as shown in Figure 7 as shown in .

[0209] When decoder-side intra mode derivation (DIMD) is applied, two intra modes are derived from reconstructed neighboring samples and these predictors are combined with the planar mode predictor with weights derived from the gradients. The division operations in the weight derivation are performed using the same lookup table (LUT) based integerization scheme used by CCLM. For example, the division operation in the orientation calculation

[0210] Orient=G y / G x

[0211] This is calculated using the following LUT-based scheme:

[0212] x=Floor(Log2(Gx))

[0213] normDiff=((Gx<<4)>>x)&15

[0214] x+=(3+(normDiff!=0)?1:0)

[0215] Orient=(Gy*(DivSigTable[normDiff]|8)+(1<<(x-1)))>>x

[0216] in

[0217] DivSigTable

[16] ={0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}.

[0218] The derived intra modes are included in the master list of intra most probable modes (MPMs), so the DIMD process is performed before constructing the MPM list. The master derived intra modes of a DIMD block are stored with the block and are used for MPM list construction of neighboring blocks. Figure 8 An example of calculating HoG (Histogram of Oriented Gradients) from a template with a width of 3 pixels is illustrated.

[0219] For each intra prediction mode in the MPM, the sum of absolute transform differences (SATD) between the prediction and reconstructed samples of the template is calculated. The first two intra prediction modes with the smallest SATD are selected as TIMD modes. The two TIMD modes are fused with weights after applying the PDPC process, and the current CU is encoded using this weighted intra prediction. Position-dependent intra prediction combination (PDPC) is included in the derivation of the TIMD mode. The costs of the two selected modes are compared with a threshold, and in the test with a cost factor of 2, the following is applied:

[0220] costMode2<2*costMode1

[0221] If this condition is true, then fusion is applied, otherwise only mode 1 is used.

[0222] The weight of a mode is calculated based on its SATD cost as follows:

[0223] weight1=costMode2 / (costMode1+costMode2)

[0224] weight2=1-weight1

[0225] The division operation is performed using the same lookup table (LUT) based integerization scheme used by CCLM.

[0226] In VVC, a low-frequency non-separable transform (LFNST) is applied between the forward main transform and quantization (at the encoder) and between dequantization and the inverse main transform (at the decoder side), as Fig. 9 As shown in Fig. 9 . In the LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform is applied according to the block size. For example, the 4×4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and the 8×8 LFNST is applied to large blocks (i.e., min(width, height) > 4).

[0227] The application of the non-separable transform used in the LFNST is described below by taking an input as an example. To apply the 4×4 LFNST, the 4×4 input block X

[0228]

[0229] is first represented as a vector

[0230]

[0231] The non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16×16 transform matrix. The 16×1 coefficient vector is then reorganized into a 4×4 block using the scan order (horizontal, vertical, or diagonal) of the block. The coefficients with smaller indices will be placed in the 4×4 coefficient block together with smaller scan indices.

[0232] The LFNST applies the non-separable transform based on the direct matrix multiplication method, thus being implemented in a single iteration without multiple iterations. However, it is necessary to reduce the dimension of the non-separable transform matrix to minimize the computational complexity and the storage space for storing the transform coefficients. Therefore, the reduced non-separable transform (or RST) method is used in the LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (for the 8×8 NSST, N is usually equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Therefore, the RST matrix becomes an R×N matrix instead of an N×N matrix, as follows:

[0233]

[0234] Where the R rows of the transform are the R basis of the N-dimensional space. The inverse transform matrix of RT is the transpose of its forward transform. For 8×8 LFNST, a reduction factor of 4 is applied to reduce the 64×64 direct matrix of the size of the traditional 8×8 non-separable transform matrix to a 16×48 direct matrix. Therefore, the core (main) transform coefficients are generated in the 8×8 upper left corner area using a 48×16 inverse RST matrix on the decoder side. When a 16×48 matrix is ​​applied instead of a 16×64 matrix with the same transform set configuration, each matrix takes 48 input data from three 4×4 blocks in the upper left 8×8 block (excluding the lower right 4×4 block). With the help of dimensionality reduction, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, and the performance degradation is reasonable. To reduce complexity, LFNST is only applicable when all coefficients outside the first coefficient subgroup are insignificant. Therefore, when LFNST is applied, all main-only transform coefficients must be zero. This allows the LFNST index signaling to be adjusted at the last significant position, thus avoiding the extra coefficient scan in the current LFNST design, which is only used to check the significant coefficients at a specific position. The worst-case processing of LFNST (in terms of multiplications per pixel) restricts the inseparable transforms for 4×4 and 8×8 blocks to 8×16 and 8×48 transforms, respectively. In these cases, when LFNST is applied, the last significant scan position must be less than 8 for other sizes less than 16. For blocks of shape 4×N, N×4 and N>8, the proposed restriction means that LFNST is now applied only once and only to the 4×4 region in the top left corner. Since all the main-only coefficients are zero when LFNST is applied, the number of operations required for the main transform is reduced in this case. From the encoder perspective, the quantization of the coefficients is significantly simplified when testing the LFNST transform. For the first 16 coefficients (in scan order), rate-distortion optimized quantization must be performed to the maximum extent possible, and the remaining coefficients are forced to zero.

[0235] There are 4 transform sets in LFNST, and 2 inseparable transform matrices (kernels) are used for each transform set. The mapping from intra prediction modes to transform sets is predefined, as shown in the following table. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81<=predModeIntra<=83), transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable secondary transform candidate is further specified by an explicitly signaled LFNST index. After the transform coefficients, the index is signaled once in the bitstream for each intra CU.

[0236] IntraPredMode Transform Set Index IntraPredMode<0 1 0<=IntraPredMode<=1 0 2<=IntraPredMode<=12 1 13<=IntraPredMode<=23 2 24<=IntraPredMode<=44 3 45<=IntraPredMode<=55 2 56<=IntraPredMode<=80 1 81<=IntraPredMode<=83 0

[0237] Since LFNST is only applicable when all coefficients outside the first coefficient subgroup are unimportant, the LFNST index encoding depends on the position of the last significant coefficient. In addition, the LFNST index is context-coded, but does not depend on the intra prediction mode, and only the first bit is context-coded. In addition, LFNST is applicable to intra CUs both within and between slices and for both luma and chroma. If dual tree is enabled, the LFNST indexes for luma and chroma will be signaled separately. For inter slices (dual tree is disabled), a single LFNST index is signaled and used for both luma and chroma.

[0238] Considering that due to the existing maximum transform size limit (64×64), large CUs larger than 64×64 are implicitly split (TU tiles), and LFNST index search can increase the data buffer by four times in a certain number of decoding pipeline stages. Therefore, the maximum size allowed by LFNST is limited to 64×64. It is worth noting that LFNST is only enabled under DCT2. LFNST index signaling is placed before MTS index signaling.

[0239] It is not obvious that using a scaling matrix for perceptual quantization, the scaling matrix specified for the master matrix can be useful for LFNST coefficients. Therefore, using a scaling matrix for LFNST coefficients is not allowed. For single-tree partitioning mode, chroma LFNST is not applied.

[0240] Scalable video coding refers to a coding architecture in which a bitstream can contain multiple representations of the content, for example, at different bitrates, resolutions, or frame rates. In these cases, the receiver can extract the desired representation based on its characteristics (e.g., the resolution that best matches the display device). Alternatively, a server or network element can extract the portion of the bitstream to be sent to the receiver based on, for example, the receiver's network characteristics or processing capabilities.

[0241] Scalable video coding can be achieved through multi-layer coding. Multi-layer coding is a concept in which an uncoded visual representation of a scene is mapped to multiple dependent or independent representations (called layers) through processes such as transformation and filtering. One or more encoders are used to encode the layered visual representation. When a layer contains redundancy, it can be encoded with a significant gain in coding efficiency using a single encoder by using inter-layer prediction techniques. Layered video coding is typically used to provide some form of scalability in a service, such as quality scalability, spatial scalability, temporal scalability, and view scalability.

[0242] Temporal scalability can be treated differently than other types of scalability. A sublayer or temporal sublayer can be defined as a temporal scalable layer (or temporal layer TL) of a temporal scalable bitstream. Each picture of a temporal scalable bitstream can be assigned a temporal identifier, for example, which can be assigned to a variable TemporalId. For example, the temporal identifier can be indicated in a NAL unit header or an OBU extension header. TemporalId equal to 0 corresponds to the lowest temporal level. The bitstream created by excluding all coded pictures with TemporalId greater than or equal to the selected value and including all other coded pictures remains consistent. Therefore, a picture with TemporalId equal to tid_value does not use any picture with TemporalId greater than tid_value as a prediction reference.

[0243] Inter prediction may involve sample positions outside the reference picture boundary (also referred to as out-of-bounds (OOB)), i.e., the motion vector in the inter prediction may point outside the picture boundary, which is due to at least (but not necessarily limited to) the following two reasons. First, the position of the prediction block corresponding to the motion vector used in the inter prediction may be partially or completely outside the picture boundary. Second, the prediction block corresponding to the motion vector used in the inter prediction may include non-integer sample positions, which are located within the picture boundary, but whose sample values ​​are interpolated using filtering that obtains input samples from positions outside the picture boundary.

[0244] To allow motion vectors to point outside the picture boundaries, existing codecs (such as HEVC and VVC) obtain samples outside the picture boundaries by efficiently copying samples on the picture boundaries. Samples on the left and right picture boundaries are horizontally copied to the OOB samples on the left and right sides of the picture, respectively. Samples on the top and bottom picture boundaries are vertically copied to the OOB samples above and below the picture, respectively. This duplication can be called "padding". This extension of the reference picture allows the motion compensation process to point to the OOB area, thereby improving prediction efficiency. However, since the extended area is effectively filled with padding samples, which do not represent the actual texture behavior of the video, the motion compensation performance is still not optimal when using such samples for prediction.

[0245] During VVC decoding, independent VVC sub-pictures are treated as pictures. Motion vectors pointing outside the boundaries of independent sub-pictures will cause the duplication of boundary samples of the sub-picture. In addition, for independent VVC sub-pictures, it may also be necessary to disable loop filtering across the boundaries of independent VVC sub-pictures. When sps_subpic_treated_as_pic_flag[i] of a sub-picture is equal to 1, the boundaries of the sub-picture are treated as picture boundaries during VVC decoding. When sps_Loop_filter_across_subpic_enabled_pic_flag[i] is equal to 0, loop filtering across the boundaries of sub-pictures is disabled during VVC decoding.

[0246] The mechanism for efficiently copying boundary samples to OOB samples during inter prediction can be implemented in a variety of ways. One way is to allocate a sample array that is larger than the size of the decoded picture, i.e., there are margins on the top, bottom, right, and left of the image. In addition to or instead of using such margins, the position of the sample used for prediction (whether as input to fractional sample interpolation of the prediction block, or as samples in the prediction block itself) can be saturated so that the position does not exceed the picture boundary (or margin if margins are used). Some video coding standards describe support for motion vectors on picture boundaries in this way.

[0247] In some projection formats of 360-degree panoramic video and omnidirectional video, when samples that are horizontally beyond the picture boundary are needed during inter-frame prediction, the sample values ​​on the other side of the picture can be used instead of using duplicate boundary samples. This prediction mode can be called surround motion compensation.

[0248] The present embodiment relates to various examples of processing out of bounds (OOB) regions with samples when a motion vector points to a region outside a picture during motion compensation.

[0249] In some embodiments, the out-of-bounds area is filled with samples copied according to the direction of the motion vector(s), rather than the traditional horizontal or vertical copying.

[0250] In some embodiments, samples in the outer region of the boundary are predicted or extrapolated by intra-frame prediction methods using decoded samples within the picture. In such embodiments, the OOB region can be divided into M×N blocks, where M is the number of samples in the extended OOB region and N is the size of the other dimensions of the block.

[0251] In some embodiments, intra prediction is used from decoded samples of the current block, and the intra predicted samples are used as replacements for regions in the motion compensated block that fall within the OOB region.

[0252] Fig.10An example process is illustrated in which samples in the OOB region 1010 are predicted by intra prediction using decoded samples in the picture 1000 .

[0253] According to one embodiment, the desired intra prediction mode(s) for predicting samples in the OOB region may be determined using a texture analysis method. For example, a texture analysis mechanism such as DIMD may be applied to decoded samples in a picture to determine the intra prediction mode(s).

[0254] According to one embodiment, the (multiple) desired intra-frame prediction modes for predicting samples in the OOB region can be determined using a template-based intra-frame derivation method. For example, a template-based method (such as TIMD) can be applied to decoded samples in a picture to determine (multiple) intra-frame prediction modes.

[0255] According to one embodiment, the final prediction of samples in the OOB region may be accomplished by combining two or more different intra predictions using certain weights. One or more of the following methods may be available and / or may be used for encoding and / or decoding:

[0256] The weight value for each prediction can be predefined.

[0257] • The weight value may be selected among predefined values, and the encoder may signal the relevant index in or along with the bitstream, and correspondingly, the decoder may decode the relevant index from or along with the bitstream.

[0258] • The weight values ​​may be signaled in or along with the bitstream by an encoder, and correspondingly, the weight values ​​may be decoded from or along with the bitstream by a decoder.

[0259] The weight values ​​can be derived at both the encoder and decoder side.

[0260] according to Fig.11 In the embodiment shown in , intra prediction can be used to predict samples of only a portion 1110 of the OOB region, and intra prediction is performed using decoded samples in the picture 1100. The remaining portion 1120 can be filled horizontally or vertically with the last predicted samples, similar to how picture boundary samples are used to fill outside the picture boundaries in published coding standards (such as VVC). According to one embodiment, when samples in the remaining portion (hereinafter referred to as "corner samples") 1120 have vertical and horizontal neighboring samples in the intra-predicted OOB region 1110, the sample value of the corner sample is derived as a weighted sum of the sample values ​​of the neighboring samples, where the weight may be inversely proportional to the spatial distance from the corner sample to the neighboring samples.

[0261] According to one embodiment, a cross-component prediction method such as CCLM or CCCM may be used to predict samples in the OOB region. For example, model parameters may be derived using decoded samples in the current frame and the reference channel, and then, intra-predicted samples in the OOB region of the reference channel may be used to predict samples in the OOB region of the current channel using the derived parameters.

[0262] According to one embodiment, in a motion compensation (MC) process, a region of a block falling within an OOB region may be predicted using intra prediction of reference samples from a current block 1225 in a current frame 1220 . Fig.12 An example when a motion vector of a block points to an OOB region in the reference frame 1210 is illustrated. Fig.13 An example is illustrated of using intra prediction samples 1320 from neighboring reference samples for block areas in the reference frame that fall in the OOB, and the remainder of the block is filled with motion compensated prediction 1310 .

[0263] According to one embodiment, the intra prediction 1320X and the MC prediction 1310X in the final prediction may be Fig.13 ) to generate a smoother block prediction. For example, a hybrid operation of intra-frame and inter-frame prediction samples can be applied.

[0264] According to one embodiment, intra prediction may use motion compensated samples as well as reconstructed reference samples from the neighborhood of the block. Fig.14 An example is illustrated where motion compensated samples are used together with reconstructed reference samples from the neighborhood for bi-intra prediction of regions falling in the OOB region during the motion compensation process.

[0265] In another embodiment, intra prediction may be obtained using only reconstructed reference samples from the neighborhood of the current block, and motion compensated samples may be used to filter the intra prediction samples.

[0266] According to one embodiment, the intra-frame prediction used to predict the area of ​​the block falling in the OOB region can be a fixed mode, or can be determined using rate-distortion optimization associated with signaling the best execution mode for each block, or alternatively, can be determined using a texture analysis method (e.g., DIMD) or a template-based method (e.g., TIMD).

[0267] According to one embodiment, the intra prediction of the described region may be obtained by combining two or more intra prediction methods.

[0268] According to one embodiment, when bi-predictive motion compensation is used and only one motion vector points to the OOB region, the second motion compensated predicted samples can also be used for the intra prediction process.

[0269] According to one embodiment, when encoding or decoding a subsequent picture, the OOB area samples of the reference picture may be modified. When encoding or decoding a subsequent picture, the motion vector direction of the boundary block of the subsequent picture is used to derive the prediction direction to perform direction filling of the reference picture using intra-frame prediction or the like.

[0270] According to one embodiment, the modification of the OOB region samples of the reference picture is invoked by the subsequent picture, provided that the reference picture meets certain requirements, which may include but are not necessarily limited to one or more of the following:

[0271] - The reference picture has a certain predefined picture type. For example, the reference picture is an intra random access point (IRAP) picture defined in VVC or a key frame defined in AV1.

[0272] - The reference picture has a specific picture type indicated in or with the bitstream or decoded from or with the bitstream.

[0273] - The reference picture has a certain predefined coding type. For example, the reference picture is an intra frame, or all slices of the reference picture are intra slices.

[0274] - The reference picture has a specific coding type indicated in or along the bitstream or decoded from or along the bitstream.

[0275] According to one embodiment, when encoding or decoding a single specific subsequent picture (hereinafter referred to as the current picture), the OOB area samples of the reference picture may be modified, provided that the current picture satisfies certain requirements, which may include but are not necessarily limited to one or more of the following:

[0276] - The current picture is the first subsequent picture in decoding order in the lowest temporal sub-layer (denoted as TL0) after the reference picture. In this case, the reference picture is also in the lowest temporal sub-layer.

[0277] - The current picture is the first subsequent picture following the reference picture in decoding order and is located in the same temporal sub-layer as the reference picture.

[0278] -The current picture uses the reference picture as a reference for inter-frame prediction.

[0279] According to one embodiment, when encoding or decoding a subsequent picture, the modification of the OOB area samples of the reference picture may be repeatedly performed, provided that the subsequent picture meets certain requirements, which may include but are not necessarily limited to one or more of the following items:

[0280] - A subsequent picture follows the reference picture in decoding order and has a lower temporal identifier value than any other picture that immediately follows the reference picture in decoding order before the subsequent picture.

[0281] - Subsequent pictures use the reference picture as a reference for inter prediction.

[0282] The above embodiments of modifying the OOB area samples of the reference picture based on a limited reference picture or a subsequent picture can also be combined. In other words, when both the reference picture and the subsequent picture meet certain requirements described in the above embodiments, the OOB area samples of the reference picture will be modified.

[0283] According to one embodiment, in the case where the boundary block is encoded in intra-frame mode, the (multiple) motion vectors of the blocks near the intra-frame coded block can be used for directional filling. It can be a direct neighboring block of the intra-frame coded boundary block, or it can be the first inter-frame coded block at a specific distance and direction of the intra-frame coded boundary block. The nearby area can be defined in the coding standard; for example, it can be one or more CTUs, tiles, or slices. In another example, if the boundary block is encoded in intra-frame mode, the directional filling can follow the intra-frame prediction direction of the intra-frame coded block.

[0284] The width or height of the directional fill region can be selected based on the motion vector size or can be hard-coded (e.g., to 64).

[0285] In some examples, samples that have undergone directional filling may undergo another directional filling. According to one embodiment, such samples are overwritten by the new directional filling. According to one embodiment, the results of the directional filling can be mixed to form sample values ​​for samples within the OOB region.

[0286] Fig.15 An example embodiment of directionally filling the OOB region with samples in a reference picture 1510 is illustrated. When encoding a block at a picture boundary, the following steps are performed for each sub-block (such as a 4×4 block), where for each sub-block, a motion vector is derived that references samples in the OOB region of the reference picture 1510. In this process, for each block (such as a 4×4 block) that references samples outside the picture boundary, the motion vector is copied. The motion vector is used to directionally fill the boundary samples in the reference picture. Motion compensation is performed from the filled reference picture.

[0287] It is important to note in encoding that this process can be applied multiple times for the candidate motion vector being selected by the encoder. Similarly, when decoding a block at a picture boundary, the following steps are performed for each sub-block, where for each sub-block, a motion vector is derived that refers to a sample in the OOB region of the reference picture 1510

[0288] 1. Motion vectors are used to derive directional intra prediction modes, etc.

[0289] 2. The derived intra prediction mode is used to fill the boundary samples of the reference picture into the OOB area of ​​the reference picture.

[0290] 3. The prediction block used for motion compensation of the sub-block being encoded or decoded is formed using a reference picture with an OOB region that has directionally filled samples.

[0291] According to one embodiment, when intra prediction is used for OOB region samples, the direction of the motion vector of the boundary blocks may be considered so that an angular intra mode with the same or close direction is used to predict the motion vector of the sample.

[0292] According to one embodiment, one or more of the described embodiments may be used alone or in combination to handle out-of-boundary samples in motion compensation.

[0293] According to one embodiment, the solution described in any other (multiple) embodiments may be applied to the boundaries of tiles, sub-pictures, and / or slices when their corresponding boundaries are configured in an independently decodable manner while motion vectors are allowed to point to the OOB region. For example, when sps_subpic_treated_as_pic_flag[i] is indicated or inferred to be equal to 1 in VVC, the sub-picture boundary is treated as a picture boundary according to any of the above embodiments.

[0294] According to one embodiment, the number of padding samples (i.e., the width and height of the OOB region) is predefined. According to one embodiment, the width and / or height of the OOB region is indicated in the bitstream or with the bitstream, for example, at the channel, picture, or sequence level. According to one embodiment, the encoder indicates the width and / or height of the OOB region in the bitstream or with the bitstream, for example, in a picture parameter set or a sequence parameter set. The width and / or height can be indicated separately or jointly for all boundaries (i.e., a single value applicable to both width and height), separately for horizontal and vertical boundaries, or separately for top, bottom, left and right boundaries. The width and / or height can be indicated jointly for all independent regions applicable to OOB sample derivation, or can be indicated separately on the basis of independent regions. Similarly, in one embodiment, the decoder decodes (multiple) width and / or height values ​​of (multiple) OOB regions from the bitstream or with the bitstream (such as from a picture parameter set or a sequence parameter), and applies them accordingly in decoding. Controlling the width and / or height of the OOB region may be used to limit memory usage, which may be beneficial, for example, when embodiments are applied to independent regions, such as independent sub-pictures.

[0295] According to one embodiment, different methods for filling the OOB region are determined to be used for different boundaries. The filling method for a particular boundary can be signaled in the bitstream. For example, it can be signaled that the outfill based on texture analysis is used for the left and right boundaries, while the filling based on the closest available sample is used for the upper and lower boundaries.

[0296] According to one embodiment, the use of horizontal wraparound motion compensation is enabled and indicated in the bitstream by the encoder (e.g., in VVC, pps_ref_wraparound_enabled_flag is equal to 1, etc.), or decoded from the bitstream by the decoder. When horizontal wraparound motion compensation is used, horizontal sample padding of the OOB area is turned off, and the sample padding described in this embodiment is only applied to the OOB area above or below the upper and lower boundaries. This can reduce memory usage in the implementation.

[0297] According to one embodiment, if the motion vector points to an area outside the first reference image, the motion vector is inverted by changing the signs of its horizontal and vertical motion vector components, and the second reference image is selected from a different temporal direction than the first reference image. In addition to inverting the motion vector, the first and second reference images may be scaled according to their temporal distance to the current image. The decision to invert the motion vector may be configured to depend on the size of the area outside the first reference image required for the prediction block.

[0298] According to the method of the embodiment Fig.16 The method generally includes determining 1610 a motion vector and a reference frame of a current block; determining 1620 when a motion vector points to an area outside the reference frame; and generating 1630 an extended reference frame by filling the area outside the reference frame with one or more of the following:

[0299] ○ Directional intra prediction using boundary samples of the reference frame according to the direction of the motion vector;

[0300] ○ Intra-frame prediction method using decoded samples within the current frame;

[0301] ○ Intra-frame prediction samples of the current block;

[0302] The current block is predicted 1640 by using motion compensation from the extended reference frame; and the current block is encoded 1650 into a bitstream. Each step can be implemented by a corresponding module of a computer system.

[0303] An apparatus according to an embodiment includes: means for determining a motion vector and a reference frame for a current block; means for determining when a motion vector points to an area outside the reference frame; means for generating an extended reference frame by filling the area outside the reference frame with one or more of the following:

[0304] ○ Directional intra prediction using boundary samples of the reference frame according to the direction of the motion vector;

[0305] ○ Intra-frame prediction method using decoded samples within the current frame;

[0306] ○ Intra-frame prediction samples of the current block;

[0307] A component for predicting a current block by using motion compensation from an extended reference frame; and a component for encoding the current block into a bitstream. The component includes at least one processor and a memory including computer program code, wherein the processor may also include a processor circuit system. The memory and the computer program code are configured to work with the at least one processor so that the apparatus performs according to various embodiments. Fig.16 method.

[0308] Fig.17 An example of a data processing system of an apparatus is illustrated. Several functions may be performed by a single physical device, for example, all computing processes may be performed in a single processor if desired. The data processing system includes a main processing unit 100, a memory 102, a storage device 104, an input device 106, an output device 108, and a graphics subsystem 110, all of which are connected to each other via a data bus 112.

[0309] Main processing unit 100 is a conventional processing unit arranged to process data within a data processing system. Main processing unit 100 may include or be implemented as one or more processors or processor circuitry. Memory 102, storage device 104, input device 106, and output device 108 may include conventional components recognized by those skilled in the art. Memory 102 and storage device 104 store data in data processing system 100.

[0310] Computer program code resides in the memory 102 for implementing, for example, Fig.16 The method is shown in the flowchart of . Input device 106 inputs data into the system, while output device 108 receives data from the data processing system and forwards the data to, for example, a display. Data bus 112 is a conventional data bus, and although shown as a single line, it can be any combination of the following items: processor bus, PCI bus, graphics bus, ISA bus. Therefore, it is easy for a person skilled in the art to recognize that the device can be any data processing device, such as a computer device, a personal computer, a server computer, a mobile phone, a smart phone or an Internet access device, such as an Internet tablet.

[0311] Fig.18 An example of a video encoder is shown in FIG. n: The image to be encoded; P' n : Prediction representation of image block; D n : prediction error signal; D' n : Reconstructed prediction error signal; I' n : Preliminary reconstructed image; R' n : Final reconstructed image; T, T -1 : Transformation and inverse transformation; Q, Q -1 : quantization and inverse quantization; E: entropy coding; RFM: reference frame memory; P inter : Inter-frame prediction; P intra : Intra-frame prediction; MS: Mode selection; F: Filtering. Fig.19 The block diagram of a video decoder is shown, where P' n : Prediction representation of image block; D' n : Reconstructed prediction error signal; I' n : Preliminary reconstructed image; R' n : Final reconstructed image; T -1 : Inverse transform; Q -1 : Inverse quantization; E -1 : Entropy decoding; RFM: Reference frame memory; P: Prediction (intra or intra); F: Filtering. The apparatus according to the embodiment may include only an encoder or a decoder or both.

[0312] Various embodiments may be implemented with the aid of a computer program code resident in a memory and causing an associated device to perform the method. For example, a device may include circuitry and electronics for processing, receiving, and sending data, computer program code in a memory, and a processor that causes the device to perform the features of the embodiments when the computer program code is run. In addition, a network device such as a server may include circuitry and electronics for processing, receiving, and sending data, computer program code in a memory, and a processor that causes the network device to perform the features of the various embodiments when the computer program code is run.

[0313] The present solution provides advantages. For example, the present embodiment improves the efficiency of motion compensation when the motion vectors point to areas outside the picture boundaries in the reference frame(s).

[0314] If desired, the different functions discussed herein may be performed in different orders and / or concurrently with other functions. In addition, if desired, one or more of the above functions and embodiments may be optional or may be combined.

[0315] Although various aspects of the embodiment are listed in the independent claim, further aspects include other combinations of features of the described embodiment and / or dependent claims with features of the independent claim and not just the combinations explicitly listed in the claim.

[0316] It should also be noted herein that while the above describes example embodiments, these descriptions should not be viewed in a limiting sense. Rather, several variations and modifications may be made without departing from the scope of the present disclosure as defined in the appended claims.

Claims

1. A device for encoding a video frame, the device comprising: -Components for determining the motion vector and reference frame of the current block; - means for determining when said motion vector points to an area outside said reference frame; - means for generating an extended reference frame by filling an area outside said reference frame with one or more of: ○ Directional intra-frame prediction using boundary samples of the reference frame according to the direction of the motion vector; ○ Intra-frame prediction method using decoded samples within the current frame; ○Intra-frame prediction samples of the current block; - means for predicting said current block by using motion compensation from said extended reference frame; and - means for encoding said current block into a bitstream.

2. The apparatus according to claim 1, further comprising: means for determining said intra prediction method by using a texture analysis method.

3. The apparatus according to claim 2, wherein the texture analysis method is a decoder-side intra-frame mode derivation method.

4. The device according to any one of claims 1 to 3, further comprising: Means for determining said intra prediction method by using a template based intra mode derivation method.

5. The apparatus according to claim 1, further comprising: Means for predicting samples of an area outside a picture by cross-component prediction, wherein samples in the area outside the picture in a reference channel are used to predict samples of the area outside the picture in a current channel.

6. The device according to any one of claims 1 to 5, further comprising: Means for filtering boundary samples between the intra prediction and the motion compensated prediction in the final prediction.

7. A method for encoding a video frame, the method comprising: - Determine the motion vector and reference frame of the current block; - determining when the motion vector points to an area outside the reference frame; - generating an extended reference frame by filling the area outside the reference frame with one or more of the following: ○ Directional intra-frame prediction using boundary samples of the reference frame according to the direction of the motion vector; ○ Intra-frame prediction method using decoded samples within the current frame; ○Intra-frame prediction samples of the current block; - predicting the current block by using motion compensation from the extended reference frame; as well as - encoding the current block into a bitstream.

8. The method according to claim 7, further comprising: The intra prediction method is determined by using a texture analysis method.

9. The method according to claim 8, wherein the texture analysis method is a decoder-side intra mode derivation method.

10. The method according to any one of claims 7 to 9, further comprising: The intra prediction method is determined by using a template-based intra mode derivation method.

11. The method according to claim 7, further comprising: Samples of a region outside a picture are predicted by cross-component prediction, wherein samples in the region outside the picture in a reference channel are used to predict samples of the region outside the picture in a current channel.

12. The method according to any one of claims 7 to 11, further comprising: The boundary samples between the intra prediction and the motion compensated prediction are filtered in the final prediction.

13. An apparatus comprising at least one processor, a memory comprising computer program code, the memory and the computer program code being configured to, together with the at least one processor, cause the apparatus to at least perform the following: - Determine the motion vector and reference frame of the current block; - determining when the motion vector points to an area outside the reference frame; - generating an extended reference frame by filling the area outside the reference frame with one or more of the following: ○ Directional intra-frame prediction using boundary samples of the reference frame according to the direction of the motion vector; ○ Intra-frame prediction method using decoded samples within the current frame; ○Intra-frame prediction samples of the current block; - predicting the current block by using motion compensation from the extended reference frame; as well as - encoding the current block into a bitstream.