Video decoding method, device, electronic device and computer-readable storage medium

By segmenting the current block into L-shaped partitions and using adjacent reconstruction samples for non-directional intra prediction, the problem of intra prediction schemes in the prior art lacks flexibility, and the encoding efficiency of video encoding is improved.

CN115004700BActive Publication Date: 2025-07-25TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180006891.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-08
Filing Date
2021-09-22
Publication Date
2025-07-25
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

In the existing video encoding technology, the intra prediction scheme lacks flexibility, especially the intra prediction of non-directional intra prediction cannot be effectively utilized on the right and bottom sides, resulting in insufficient encoding efficiency.

Method used

The reconstruction of the non-directional intra prediction mode is employed by segmenting the current block into multiple partitions, including at least one L-shaped partition, and the reconstruction of the non-directional intra prediction mode is performed based on the adjacent reconstruction sample or the adjacent reconstruction sample of the current block.

Benefits of technology

Improves the flexibility of non-directional intra prediction and enhances the encoding efficiency of video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115004700B_ABST
    Figure CN115004700B_ABST
Patent Text Reader

Abstract

The present application discloses a video decoding method, apparatus, electronic device, and non-transitory computer-readable storage medium. The method includes: decoding prediction information of a current block in a current picture that is part of an encoded video bitstream, the prediction information indicating a non-directional intra prediction mode for the current block; dividing the current block into a plurality of partitions, the plurality of partitions including at least one L-shaped partition; and reconstructing one of the plurality of partitions based on at least one of the following: (i) neighboring reconstruction samples of one of the plurality of partitions; or (ii) neighboring reconstruction samples of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Patent Application No. 17 / 469,500, filed on September 8, 2021, "METHOD AND APPARATUS FOR VIDEO CODING", which claims the benefit of priority of U.S. Provisional Application No. 63 / 084,460, filed on September 28, 2020, "NON - DIRECTIONAL INTRA PREDICTION FOR L - SHAPE PARTITION", the entire contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] This application relates to the field of video coding and decoding, and particularly to a video decoding method, apparatus, electronic device, and computer - readable storage medium. Background Art

[0004] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work described in this background section is not otherwise qualified as prior art at the time of filing, the work of the currently named inventors and aspects that are not otherwise described as prior art may not be explicitly or implicitly recognized as prior art against the present disclosure.

[0005] Video coding and decoding can be performed using inter - picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having spatial dimensions such as 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (also informally referred to as the frame rate), such as 60 pictures per second or 60 Hz. Uncompressed video has significant bit - rate requirements. For example, a 1080P60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 gigabytes of storage space.

[0006] One purpose of video encoding and decoding can be to reduce redundancy in an input video signal through compression. In some cases, compression can help reduce the foregoing bandwidth or storage space requirements by two orders of magnitude or more. Lossless compression, lossy compression, and combinations thereof can be employed. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed signal. When lossy compression is used, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be useful for the intended application. In the case of video, lossy compression is widely adopted. The amount of distortion tolerated depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that higher allowable / tolerable distortion can result in a higher compression ratio.

[0007] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy coding.

[0008] Video codec techniques can include techniques referred to as intra coding. In intra coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are encoded in the intra mode, the picture can be an intra picture. Intra pictures and their derivatives such as independent decoder refresh pictures can be used to reset the decoder state and can thus be used as the first picture in an encoded video bitstream and a video session or as a still image. Samples of intra blocks can be subjected to transformation, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique that minimizes the sample values in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after transformation, the fewer bits are required to represent the block after entropy coding for a given quantization step.

[0009] Conventional intra coding, such as known from, for example, MPEG-2 generation coding techniques, does not use intra prediction. However, some newer video compression techniques include techniques that attempt to predict based on, for example, metadata and / or surrounding sample data obtained during the encoding and / or decoding of spatially neighboring and earlier in the decoding order data blocks. Such techniques are hereinafter referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses only data from the current picture being reconstructed and not reference data from reference pictures.

[0010] There can be many different forms of intra prediction. When more than one such technique can be used in a given video coding technique, the technique used can be encoded in an intra prediction mode. In some cases, the mode can have sub - modes and / or parameters, and these sub - modes and / or parameters can be encoded separately or included in the mode codeword. Which codeword is used for a given mode, sub - mode, and / or parameter combination can affect the coding efficiency gain through intra prediction and, thus, can affect the entropy coding technique used to convert the codeword into a bitstream.

[0011] Certain modes of intra prediction were introduced in H.264, refined in H.265, and further refined in more recent coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). The neighboring sample values belonging to the available samples can be used to form a prediction block. The sample values of the neighboring samples are copied into the predictor block according to the direction. The reference to the direction used can be encoded in the bitstream or can itself be predicted.

[0012] The intra prediction schemes in the prior art typically use the top and left reference samples to perform non - directional intra prediction and cannot be compatible with using the neighboring reconstructed samples from the right and / or bottom side, which makes the intra prediction scheme lack flexibility. Summary of the Invention

[0013] Aspects of the present disclosure provide a method for video decoding, including: decoding prediction information of a current block in a current picture that is part of an encoded video bitstream, the prediction information indicating a non - directional intra prediction mode for the current block; dividing the current block into a plurality of partitions, the plurality of partitions including at least one L - shaped partition; and reconstructing one of the plurality of partitions based on at least one of: (i) neighboring reconstructed samples of one of the plurality of partitions; or (ii) neighboring reconstructed samples of the current block.

[0014] In one embodiment, at least one of the neighboring reconstructed samples is adjacent to the right or bottom side of one of the plurality of partitions.

[0015] In one embodiment, one of the plurality of partitions is an L - shaped partition, and the number of neighboring reconstructed samples depends on the size of the L - shaped partition. In one example, the number of neighboring reconstructed samples is the sum of the width and height of the L - shaped partition. In another example, the number of neighboring reconstructed samples is the sum of the shorter width and shorter height of the L - shaped partition. In another example, the number of neighboring reconstructed samples is the maximum value between the width and height of the L - shaped partition. In another example, the number of neighboring reconstructed samples is the minimum value between the width and height of the L - shaped partition.

[0016] In one embodiment, at least one of the neighboring reconstructed samples is located in another one of the plurality of partitions that is reconstructed before one of the plurality of partitions. In an example, the another one of the plurality of partitions is an L-shaped partition, and at least one of the neighboring reconstructed samples is adjacent in position to one of the right side or the bottom side of one of the plurality of partitions.

[0017] In one embodiment, a plurality of neighboring reference samples for one of the plurality of partitions are determined based on at least one of the following: (i) neighboring reconstructed samples of one of the plurality of partitions; or (ii) neighboring reconstructed samples of the current block. One of the plurality of partitions is reconstructed based on the plurality of neighboring reference samples.

[0018] In an example, the neighboring reconstructed samples include neighboring reconstructed samples of the left column and the right column of one of the plurality of partitions. The processing circuitry determines neighboring reference samples for the bottom row of one of the plurality of partitions based on the neighboring reconstructed samples of the left column and the right column of one of the plurality of partitions. One of the plurality of partitions is reconstructed based on the neighboring reference samples for the bottom row of one of the plurality of partitions.

[0019] In an example, the neighboring reconstructed samples include neighboring reconstructed samples of the top row and the bottom row of one of the plurality of partitions. The processing circuitry determines neighboring reference samples for the right column of one of the plurality of partitions based on the neighboring reconstructed samples of the top row and the bottom row of one of the plurality of partitions. One of the plurality of partitions is reconstructed based on the neighboring reference samples for the right column of one of the plurality of partitions.

[0020] In one embodiment, one of the plurality of partitions is an L-shaped partition, and one of the plurality of partitions is reconstructed based on the neighboring reconstructed samples of the left column and the top row of the current block.

[0021] In one embodiment, based on one of the plurality of partitions being an L-shaped partition, a plurality of neighboring reference samples are determined for each sample of the L-shaped partition based on the position of the sample. Each sample of the L-shaped partition is reconstructed based on the plurality of neighboring reference samples of the sample.

[0022] In one embodiment, the plurality of neighboring reference samples for each sample include one of the reconstructed neighboring samples and a neighboring sample to be reconstructed based on the reconstructed neighboring sample.

[0023] Aspects of the present disclosure provide a method for video encoding / decoding. In the method, prediction information in a current picture of a coded video bitstream is decoded, the prediction information indicating a non-directional intra prediction mode for a current block; the current block is segmented into a plurality of partitions, the plurality of partitions including at least one L-shaped partition; one of the plurality of partitions is reconstructed based on at least one of the following: (i) neighboring reconstructed samples of one of the plurality of partitions; or (ii) neighboring reconstructed samples of the current block.

[0024] Aspects of the present disclosure also provide an electronic device, including a memory and a processor. The memory stores computer instructions, and when the computer instructions are run by the processor, the electronic device is caused to execute any one of the above video decoding methods.

[0025] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions, which when executed by at least one processor, cause the at least one processor to execute any one of the above video decoding methods.

[0026] Thus, in the video decoding method provided by the present application, by using the L-shaped partition, neighboring reconstructed samples of the current block can be obtained from the right side and / or the bottom side of the current block. The present application further solves the problem that the neighboring reconstructed samples on the right side and / or the bottom side are not compatible with the conventional scheme of performing non-directional intra prediction using the top and left reference samples, and improves the flexibility of non-directional intra prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] According to the following detailed description and the drawings, other features, properties, and various advantages of the disclosed subject matter will become more apparent, in the drawings:

[0028] FIG. 1A is a schematic illustration of an exemplary subset of intra prediction modes;

[0029] FIG. 1B is an illustration of exemplary intra prediction directions;

[0030] FIG. 1C is a schematic illustration of a current block and its surrounding spatial merge candidates in an example of the present application;

[0031] Figure 2 is a schematic illustration of a simplified block diagram of a communication system according to an embodiment of the present application;

[0032] Figure 3 is a schematic illustration of a simplified block diagram of a communication system according to another embodiment of the present application;

[0033] Figure 4 is a schematic illustration of a simplified block diagram of a decoder according to an embodiment of the present application;

[0034] Figure 5 is a schematic illustration of a simplified block diagram of an encoder according to an embodiment of the present application;

[0035] Figure 6 shows a block diagram of an encoder according to another embodiment of the present application;

[0036] Figure 7 shows a block diagram of a decoder according to another embodiment of the present application;

[0037] Figure 8 Shows an exemplary block partition according to some embodiments of the present application;

[0038] Figure 9 Shows an exemplary quadtree with a nested binary tree structure according to an embodiment of the present application;

[0039] Figure 10 Shows an exemplary block partition in a multi-type tree structure according to some embodiments of the present application;

[0040] Figure 11 Shows an exemplary L-shaped partition according to an embodiment of the present application;

[0041] Figure 12 Shows an exemplary block partition using L-shaped splitting according to some embodiments of the present application;

[0042] Figure 13 Shows an exemplary nominal angle according to an embodiment of the present application;

[0043] Figure 14 Shows the positions of the top, left, and top-left samples of a pixel in the current block according to an embodiment of the present application;

[0044] Figure 15 Shows an exemplary recursive filter intra mode according to an embodiment of the present application;

[0045] Figure 16 Shows an exemplary multi-line intra prediction using four reference lines adjacent to the coding block unit according to an embodiment of the present application;

[0046] Figures 17A to 17F Shows six exemplary Reference Sample Chains (RSCs) according to some embodiments of the present application;

[0047] Figure 18 Shows an exemplary RSC according to an embodiment of the present application;

[0048] Figures 19A to 19B Shows two exemplary RSCs according to some embodiments of the present application;

[0049] Figures 20A to 20D Shows four exemplary RSCs according to other embodiments of the present application;

[0050] Figure 21 Shows an exemplary RSC according to an embodiment of the present application;

[0051] Figures 22A to 22B Shows an exemplary RSC according to some embodiments of the present application;

[0052] Figures 23A to 23B Illustrates an exemplary RSC according to some other embodiments of the present application;

[0053] Figure 24 Illustrates an exemplary RSC according to another embodiment of the present application;

[0054] Figures 25A to 25B Illustrates an exemplary RSC according to yet some other embodiments of the present application;

[0055] Figure 26 Illustrates an exemplary flowchart according to an embodiment of the present application; and

[0056] Figure 27 Is a schematic illustration of a computer system according to an embodiment of the present application. Detailed Description

[0057] I. Video Decoder and Encoder System

[0058] Referring to FIG. 1A, depicted in the lower right is a subset of nine predictor directions known from the 33 possible predictor directions of H.265 (corresponding to the 33 angular modes of the 35 intra modes). The point (101) where the arrows converge represents the sample being predicted. The arrows indicate the direction of predicting the sample. For example, arrow (102) indicates that sample (101) is predicted based on one or more samples at a 45-degree angle to the horizontal in the upper right. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more samples at a 22.5-degree angle to the horizontal in the lower left of sample (101).

[0059] Still referring to FIG. 1A, depicted in the upper left is a square block (104) of 4×4 samples (indicated by the bold dashed line). The square block (104) includes 16 samples, each labeled with “S”, its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in both the Y dimension and the X dimension of block (104). Since the size of the block is 4×4 samples, S44 is in the lower right. Also shown are reference samples following a similar numbering scheme. The reference samples are labeled with “R”, its Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, the predicted sample is adjacent to the block in reconstruction; thus, negative values are not required.

[0060] Intra picture prediction can work by copying reference sample values from neighboring samples depending on the predicted direction signaled. For example, assume that an encoded video bitstream includes signaling that indicates, for a block, a prediction direction consistent with arrow (102) - i.e., samples are predicted based on one or more predicted samples at a 45-degree angle to the horizontal in the upper right. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Then, sample S44 is predicted based on reference sample R08.

[0061] In some cases, values of multiple reference samples can be combined, for example, by interpolation, to compute a reference sample; especially when the direction is not divisible by 45 degrees.

[0062] As video coding technology has evolved, the number of possible directions has increased. In H.264 (in 2003), nine different directions could be represented. This increased to 33 in H.265 (in 2013), and JEM / VVC / BMS can support up to 65 directions when disclosed. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those likely directions with a small number of bits, thereby accepting some penalty for less likely directions. Additionally, sometimes the direction itself can be predicted based on neighboring directions used in neighboring already decoded blocks.

[0063] Figure 1B shows a schematic diagram (105) depicting 65 intra prediction directions according to JEM to show the increasing number of prediction directions over time.

[0064] The mapping of intra prediction direction bits representing the direction in an encoded video bitstream can vary with different video coding technologies; and can range, for example, from a simple direct mapping of the prediction direction to an intra prediction mode, to codewords, to complex adaptive schemes involving the most likely modes and the like. However, in all cases, there can be certain directions that are statistically less likely to occur in video content than some other directions. Since the goal of video compression is to reduce redundancy, those less likely directions will be represented by a larger number of bits compared to the more likely directions in a well-working video coding technology.

[0065] Motion compensation can be a lossy compression technique and can involve the following techniques: where blocks of sample data from a previously reconstructed picture or a part thereof (reference picture) are spatially shifted in the direction indicated by a motion vector (hereinafter referred to as MV (Motion Vector, MV)) and then used to predict a newly reconstructed picture or picture part. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or can have three dimensions, the third dimension being an indication of the reference picture in use (the latter can be indirectly a temporal dimension).

[0066] In some video compression techniques, the MV applicable to a specific region of sample data can be predicted based on other MVs, for example, based on an MV that is related to another region of sample data that is spatially adjacent to the region being reconstructed and that is before the MV in decoding order. Doing so can significantly reduce the amount of data required to encode the MV, thereby eliminating redundancy and improving compression. MV prediction can work effectively, for example, because when encoding an input video signal (referred to as natural video) obtained from a video camera device, there is a statistical likelihood that there are regions moving in a similar direction that are larger than the region to which a single MV is applied, and thus in some cases, a similar MV obtained from a neighboring region can be used for prediction. This results in the MV found for a given region being similar or identical to the MV predicted based on surrounding MVs and can in turn be represented with fewer bits than the number of bits used in the case of directly encoding the MV. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) obtained from an original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy, for example, because of rounding errors when calculating a predicted value based on several surrounding MVs.

[0067] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, “High Efficiency Video Coding”, December 2016). Among the various MV prediction mechanisms provided by H.265, the technique hereinafter referred to as “spatial merge” is described herein.

[0068] Referring to FIG. 1C, the current block (111) may include samples that have been found by the encoder during motion search processing and that can be predicted from a previous block of the same size that has been spatially shifted. Instead of directly encoding the MV, an MV associated with any one of five surrounding samples represented by A0, A1 and B0, B1, B2 (112 to 116 respectively) can be used, and the MV can be derived from metadata associated with one or more reference pictures, for example, from the most recent (in decoding order) reference picture. In H.265, MV prediction can use a predictor from the same reference picture that neighboring blocks are using.

[0069] Figure 2 FIG. 4 shows a simplified block diagram of a communication system (200) according to an embodiment of the present disclosure. The communication system (200) includes a plurality of terminal devices that can communicate with each other via, for example, a network (250). For example, the communication system (200) includes a first pair of terminal devices (210) and (220) interconnected via a network (250). In Figure 2 an example, the first pair of terminal devices (210) and (220) perform unidirectional data transmission. For example, the terminal device (210) may encode video data (e.g., a video picture stream captured by the terminal device (210)) for transmission via the network (250) to another terminal device (220). The encoded video data may be transmitted in the form of one or more encoded video bitstreams. The terminal device (220) may receive the encoded video data from the network (250), decode the encoded video data to recover the video pictures, and display the video pictures based on the recovered video data. Unidirectional data transmission may be common in media service applications and the like.

[0070] In another example, the communication system (200) includes a second pair of terminal devices (230) and (240) that perform two-way transmission of encoded video data, which may occur, for example, during a video conference. For two-way data transmission, in an example, each of the terminal devices (230) and (240) may encode video data (e.g., a video picture stream captured by the terminal device) for transmission via the network (250) to the other of the terminal devices (230) and (240). Each of the terminal devices (230) and (240) may also receive the encoded video data transmitted by the other of the terminal devices (230) and (240), and may decode the encoded video data to recover the video pictures, and may display the video pictures at an accessible display device based on the recovered video data.

[0071] In Figure 2In the example, the terminal devices (210), (220), (230), and (240) may be shown as servers, personal computers, and smart phones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (250) represents any number of networks that convey encoded video data between the terminal devices (210), (220), (230), and (240), such as including wired (wired) and / or wireless communication networks. The communication network (250) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise stated herein, the structure and topology of the network (250) may be immaterial to the operation of the present disclosure.

[0072] As an example of an application for the disclosed subject matter, Figure 3 the placement of a video encoder and a video decoder in a streaming environment is shown. The disclosed subject matter may equally apply to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0073] A streaming system may include a capture subsystem (313), which may include a video source (301) that creates, for example, an uncompressed video picture stream (302), such as a digital imaging device. In the example, the video picture stream (302) includes samples taken by the digital imaging device. The video picture stream (302), depicted as a thick line to emphasize the high data volume compared to the encoded video data (304) (or encoded video bitstream), may be processed by an electronic device (320) that includes a video encoder (303) coupled to the video source (301). The video encoder (303) may include hardware, software, or a combination thereof to implement or perform aspects of the disclosed subject matter as described in more detail below. The encoded video data (304) (or encoded video bitstream (304)), depicted as a thin line to emphasize the lower data volume compared to the video picture stream (302), may be stored on a streaming server (305) for future use. One or more streaming client subsystems such as Figure 3The client subsystems (306) and (308) can access the streaming server (305) to retrieve copies (307) and (309) of the encoded video data (304). The client subsystem (306) can include, for example, a video decoder (310) in an electronic device (330). The video decoder (310) decodes the incoming copy (307) of the encoded video data and creates an output video picture stream (311) that can be presented on a display (312) (e.g., a display screen) or other rendering device (not depicted). In some streaming systems, the encoded video data (304), (307), and (309) (e.g., video bitstreams) can be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In an example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.

[0074] Note that the electronic devices (320) and (330) can include other components (not shown). For example, the electronic device (320) can include a video decoder (not shown), and the electronic device (330) can also include a video encoder (not shown).

[0075] Figure 4 A block diagram of a video decoder (410) according to an embodiment of the present disclosure is shown. The video decoder (410) can be included in an electronic device (430). The electronic device (430) can include a receiver (431) (e.g., receiving circuitry). The video decoder (410) can be used instead of Figure 3 the video decoder (310) in the example.

[0076] A receiver (431) may receive one or more encoded video sequences to be decoded by a video decoder (410); in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel (401), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (431) may receive the encoded video data along with other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not depicted). The receiver (431) may separate the encoded video sequences from the other data. To counteract network jitter, a buffer memory (415) may be coupled between the receiver (431) and the entropy decoder / parser (420) (hereinafter referred to as "parser (420)"). In some applications, the buffer memory (415) is part of the video decoder (410). In other applications, the buffer memory (415) may be external to the video decoder (410) (not depicted). In still other applications, a buffer memory (not depicted) may exist external to the video decoder (410) to, for example, counteract network jitter, and additionally, another buffer memory (415) may exist internal to the video decoder (410) to, for example, handle playback timing. When the receiver (431) is receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (415) may not be required or the buffer memory (415) may be small. For use on a best-effort packet network such as the Internet, a buffer memory (415) may be required, which may be relatively large and may advantageously have an adaptive size and may be implemented at least partially in an operating system or a similar element (not depicted) external to the video decoder (410).

[0077] The video decoder (410) may include a parser (420) to reconstruct symbols (421) from the encoded video sequences. The categories of these symbols include information for managing the operation of the video decoder (410), and potentially information for controlling a rendering device such as a renderer device (412) (e.g., a display screen), which is not part of the electronic device (430) but may be coupled to the electronic device (430), as Figure 4As shown. The control information for presenting the device can be in the form of Supplemental Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (420) can perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be carried out according to video coding techniques or standards and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (420) can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroups can include Groups of Picture (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The parser (420) can also extract information such as transform coefficients, quantizer parameter values, MVs, etc. from the encoded video sequence.

[0078] The parser (420) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (415) to create symbols (421).

[0079] Depending on the type of the encoded video picture or a part thereof (e.g., inter - picture and intra - picture, inter - block and intra - block) and other factors, the reconstruction of the symbols (421) can involve multiple different units. Which units are involved and the way they are involved can be controlled by subgroup control information parsed by the parser (420) from the encoded video sequence. For clarity, such subgroup control information flows between the parser (420) and the following multiple units are not depicted.

[0080] In addition to the functional blocks already mentioned, the video decoder (410) can be conceptually subdivided into multiple functional units as described below. In the actual implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.

[0081] The first unit is a scaler / inverse transform unit (451). The scaler / inverse transform unit (451) receives the quantized transform coefficients as symbols (421) and control information from the parser (420), including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (451) may output a block including sample values, and the sample values may be input into the aggregator (455).

[0082] In some cases, the output samples of the scaler / inverse transform unit (451) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed part of the current picture. Such predictive information may be provided by the intra-picture prediction unit (452). In some cases, the intra-picture prediction unit (452) uses the already reconstructed surrounding information obtained from the current picture buffer (458) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (458) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (455) adds the predictive information already generated by the intra-prediction unit (452) to the output sample information provided by the scaler / inverse transform unit (451) on a per-sample basis.

[0083] In other cases, the output samples of the scaler / inverse transform unit (451) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (453) may access the reference picture memory (457) to obtain samples for prediction. After motion-compensating the obtained samples according to the symbols (421) belonging to the block, these samples may be added by the aggregator (455) to the output of the scaler / inverse transform unit (451) (in this case, referred to as residual samples or a residual signal) to generate output sample information. The address from which the motion compensation prediction unit (453) in the reference picture memory (457) obtains the prediction samples may be controlled by the MV, and the MV is available to the motion compensation prediction unit (453) in the form of symbols (421). The MV may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values obtained from the reference picture memory (457) when using sub-sample accurate MV, an MV prediction mechanism, etc.

[0084] The output samples of the aggregator (455) can undergo various loop filtering techniques in the loop filter unit (456). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream) and are available to the loop filter unit (456) as symbols (421) from the parser (420). However, video compression techniques can also respond to meta-information obtained during the decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0085] The output of the loop filter unit (456) can be a sample stream that can be output to a renderer device (412) and stored in a reference picture memory (457) for future inter-picture prediction.

[0086] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (420)), the current picture buffer (458) can become part of the reference picture memory (457), and a new current picture buffer can be reallocated before starting the reconstruction of subsequent encoded pictures.

[0087] The video decoder (410) can perform decoding operations according to predetermined video compression techniques in standards such as ITU-T Recommendation H.265. In the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence can conform to the syntax specified by the video compression technique or standard being used. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as tools that can only be used under that profile. For compliance, it may also be required that the complexity of the encoded video sequence be within the range defined by the video compression technique or standard level. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0088] In an embodiment, a receiver (431) may receive additional (redundant) data and an encoded video. The additional data may be included as part of an encoded video sequence. The additional data may be used by a video decoder (410) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, a redundant slice, a redundant picture, a forward error correction code, and the like.

[0089] Figure 5 A block diagram of a video encoder (503) according to an embodiment of the present disclosure is shown. The video encoder (503) is included in an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmission circuit system). The video encoder (503) may be used instead of Figure 3 the video encoder (303) in the example.

[0090] The video encoder (503) may receive video samples from a video source (501) (which is not Figure 5 part of the electronic device (520) in the example), and the video source (501) may capture video images to be encoded by the video encoder (503). In another example, the video source (501) is part of the electronic device (520).

[0091] The video source (501) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (503), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 Y CrCb, RGB,...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (501) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (501) may be a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that impart motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, and the like. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0092] According to an embodiment, the video encoder (503) may encode pictures of a source video sequence in real time or under any constraints required by the application and compress them into an encoded video sequence (543). Enforcing an appropriate encoding speed is a function of the controller (550). In some embodiments, the controller (550) controls other functional units as described below and is functionally coupled to other functional units. For clarity, the coupling is not depicted. Parameters set by the controller (550) may include rate control related parameters (picture skip, quantizer, λ value of rate-distortion optimization technique,...), picture size, Group of Picture (GOP) layout, maximum MV allowed reference area, etc. The controller (550) may be configured to have other suitable functions, which belong to the video encoder (503) optimized for a specific system design.

[0093] In some embodiments, the video encoder (503) is configured to operate in an encoding loop. As an overly simplified description, in an example, the encoding loop may include a source encoder (530) (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and reference pictures) and a (local) decoder (533) embedded in the video encoder (503). The decoder (533) reconstructs the symbols in a manner similar to how a (remote) decoder would also create sample data to create sample data (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory (534). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (534) is also bit-exact between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related fields.

[0094] The operation of the "local" decoder (533) may be the same as the operation of a "remote" decoder such as the video decoder (410), which has been described in detail above in connection with Figure 4 the video decoder (410). However, also briefly referring to Figure 4 , when symbols are available and the entropy encoder (545) and the parser (420) can encode / decode the symbols into the encoded video sequence losslessly, the entropy decoding part of the video decoder (410) including the buffer memory (415) and the parser (420) may not be fully implemented in the local decoder (533).

[0095] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also necessarily exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described in detail. A more detailed description is needed only in certain areas and is provided below.

[0096] In some examples, during operation, the source encoder (530) may perform motion-compensated predictive coding that predictively encodes an input picture by referring to one or more previously encoded pictures designated as "reference pictures" from a video sequence. In this way, the encoding engine (532) encodes the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture, which can be selected as the prediction reference for the input picture.

[0097] The local video decoder (533) may decode the encoded video data of a picture that can be designated as a reference picture based on the symbols created by the source encoder (530). The operation of the encoding engine (532) can advantageously be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 5 (not shown)), the reconstructed video sequence can generally be a replica of the source video sequence with some errors. The local video decoder (533) replicates the decoding process that can be performed by the video decoder on the reference picture and can store the reconstructed reference picture in the reference picture memory (534). In this way, the video encoder (503) can locally store a copy of the reconstructed reference picture, which has the same content (in the absence of transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder.

[0098] The predictor (535) may perform a prediction search for the encoding engine (532). That is, for a new picture to be encoded, the predictor (535) may search the reference picture memory (534) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture MVs, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor (535) may operate on a per-pixel block basis of the sample blocks to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (535), the input picture may have a prediction reference taken from multiple reference pictures stored in the reference picture memory (534).

[0099] The controller (550) may manage the encoding operations of the source encoder (530), including, for example, setting parameters and subgroup parameters for encoding video data.

[0100] The outputs of all the foregoing functional units may be entropy encoded in the entropy encoder (545). The entropy encoder (545) converts symbols generated by various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0101] The transmitter (540) may buffer the encoded video sequence created by the entropy encoder (545) in preparation for transmission via a communication channel (560), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (540) may merge the encoded video data from the video encoder (503) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0102] The controller (550) may manage the operations of the video encoder (503). During encoding, the controller (550) may assign a specific encoded picture type to each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures may generally be assigned to one of the following picture types:

[0103] An intra picture (I picture), which may be a picture that is encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, an Independent Decoder Refresh (IDR) picture. Those skilled in the art are aware of those variations of I pictures and their corresponding applications and characteristics.

[0104] A predictive picture (P picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction that predicts the sample values of each block using at most one motion vector (MV) and a reference index.

[0105] A bi-predictive picture (B picture), which may be a picture that is encoded and decoded using intra prediction or inter prediction that predicts the sample values of each block using at most two MVs and reference indices. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0106] The source picture may typically be spatially subdivided into a plurality of blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples, respectively) and encoded block by block. A block may be predictively encoded with reference to other (already encoded) blocks as determined by the coding assignment applied to the block's corresponding picture. For example, a block of an I picture may be non-predictively encoded, or a block of an I picture may be predictively encoded (spatial prediction or intra-frame prediction) with reference to already encoded blocks of the same picture. A pixel block of a P picture may be predictively encoded via spatial prediction or via temporal prediction with reference to one previously encoded reference picture. A block of a B picture may be predictively encoded via spatial prediction or via temporal prediction with reference to one or two previously encoded reference pictures.

[0107] The video encoder (503) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In its operation, the video encoder (503) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.

[0108] In an embodiment, the transmitter (540) may transmit additional data with the encoded video. The source encoder (530) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0109] Video may be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an example, a particular picture in encoding / decoding, referred to as the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously encoded and still buffered reference picture in the video, the block in the current picture may be encoded by a vector referred to as an MV. The MV points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the MV may have a third dimension that identifies the reference picture.

[0110] In some embodiments, dual prediction techniques can be used for inter - picture prediction. According to the dual prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in coding order (but may be past and future respectively in display order). A block in the current picture can be encoded by a first motion vector (MV) pointing to a first reference block in the first reference picture and a second MV pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0111] In addition, the merge mode technique can be used in inter - picture prediction to improve the coding efficiency.

[0112] According to some embodiments of the present disclosure, predictions such as inter - picture prediction and intra - picture prediction are performed on a block - by - block basis. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression, and CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) in a quadtree manner. For example, a 64×64 - pixel CTU can be divided into a 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In an example, each CU is analyzed to determine the prediction type of the CU, such as an inter - prediction type or an intra - prediction type. Depending on temporal and / or spatial prediction, the CU is divided into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in decoding (encoding / decoding) is performed on a prediction - block basis. Using a luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0113] Figure 6 A diagram of a video encoder (603) according to another embodiment of the present disclosure is shown. The video encoder (603) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence and encode the processing block into an encoded picture that is part of an encoded video sequence. In an example, the video encoder (603) is used instead of Figure 3 the video encoder (303) in the example.

[0114] In the HEVC example, a video encoder (603) receives a matrix of sample values for processing blocks such as prediction blocks of 8×8 samples. The video encoder (603) uses, for example, rate-distortion optimization to determine whether to use an intra mode, an inter mode, or a bi-prediction mode to optimally encode the processing block. When the processing block is to be encoded in the intra mode, the video encoder (603) may use an intra prediction technique to encode the processing block into an encoded picture; and when the processing block is to be encoded in the inter mode or the bi-prediction mode, the video encoder (503) may use an inter prediction or a bi-prediction technique to encode the processing block into an encoded picture, respectively. In some video coding techniques, a merge mode may be an inter-picture prediction sub-mode, where an MV is derived from one or more MV predictors without resorting to an encoded MV component external to the predictor. In some other video coding techniques, there may be an MV component applicable to a subject block. In the example, the video encoder (603) includes other components such as a mode decision module (not shown) for determining the mode of the processing block.

[0115] In Figure 6 the example, the video encoder (603) includes as Figure 6 shown an inter encoder (630), an intra encoder (622), a residual calculator (623), a switch (626), a residual encoder (624), a general controller (621), and an entropy encoder (625) coupled together.

[0116] The inter encoder (630) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter prediction information (e.g., a description of redundant information according to an inter coding technique, an MV, merge mode information), and calculate an inter prediction result (e.g., a prediction block) based on the inter prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information.

[0117] The intra encoder (622) is configured to receive samples of a current block (e.g., a processing block), compare the block with already encoded blocks in the same picture in some cases, generate quantized coefficients after transformation, and also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques) in some cases. In the example, the intra encoder (622) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.

[0118] The general controller (621) is configured to determine general control data and control other components of the video encoder (603) based on the general control data. In an example, the general controller (621) determines the mode of a block and provides a control signal to the switch (626) based on the mode. For example, when the mode is the intra mode, the general controller (621) controls the switch (626) to select the intra mode result used by the residual calculator (623), and controls the entropy encoder (625) to select the intra prediction information and include the intra prediction information in the bitstream; and when the mode is the inter mode, the general controller (621) controls the switch (626) to select the inter prediction result used by the residual calculator (623), and controls the entropy encoder (625) to select the inter prediction information and include the inter prediction information in the bitstream.

[0119] The residual calculator (623) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (622) or the inter encoder (630). The residual encoder (624) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In an example, the residual encoder (624) is configured to transform the residual data from the spatial domain to the frequency domain and generate transform coefficients. Then, the transform coefficients undergo quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (603) further includes a residual decoder (628). The residual decoder (628) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (622) and the inter encoder (630). For example, the inter encoder (630) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (622) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and in some examples, the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.

[0120] The entropy encoder (625) is configured to format the bitstream to include the encoded block. The entropy encoder (625) is configured to include various information according to a suitable standard such as HEVC. In an example, the entropy encoder (625) is configured to include general control data, the selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. Note that according to the disclosed subject matter, there is no residual information when encoding a block in the inter mode or the merge submode of the bi-prediction mode.

[0121] Figure 7FIG. showing a video decoder (710) according to another embodiment of the present disclosure. The video decoder (710) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (710) is used instead of Figure 3 the video decoder (310) in the example.

[0122] In Figure 7 the example, the video decoder (710) includes an entropy decoder (771), an inter-frame decoder (780), a residual decoder (773), a reconstruction module (774), and an intra-frame decoder (772) coupled together as Figure 7 shown.

[0123] The entropy decoder (771) may be configured to reconstruct certain symbols from the encoded picture, where the symbols represent syntax elements that make up the encoded picture. Such symbols may include, for example, the mode for encoding a block (e.g., intra mode, inter mode, bi-prediction mode, a merge sub-mode of the latter two, or another sub-mode), prediction information that can respectively identify certain samples or metadata for prediction by the intra-frame decoder (772) or the inter-frame decoder (780) (e.g., intra prediction information or inter prediction information), residual information in the form of, for example, quantized transform coefficients, etc. In an example, when the prediction mode is inter mode or bi-prediction mode, the inter prediction information is provided to the inter-frame decoder (780); and when the prediction type is intra prediction type, the intra prediction information is provided to the intra-frame decoder (772). The residual information may be inverse quantized and provided to the residual decoder (773).

[0124] The inter-frame decoder (780) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.

[0125] The intra-frame decoder (772) is configured to receive the intra prediction information and generate a prediction result based on the intra prediction information.

[0126] The residual decoder (773) is configured to perform inverse quantization to extract the dequantized transform coefficients and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require certain control information (to include Quantizer Parameter (QP)), and this information may be provided by the entropy decoder (771) (the data path is not depicted as this may be only low-volume control information).

[0127] The reconstruction module (774) is configured to combine in the spatial domain the residual output by the residual decoder (773) with the prediction result (optionally output by an inter-frame prediction module or an intra-frame prediction module) to form a reconstructed block, which can be part of a reconstructed picture, and the reconstructed picture can in turn be part of a reconstructed video. Note that other suitable operations such as deblocking operations can be performed to improve the visual quality.

[0128] Note that any suitable technique can be used to implement the video encoders (303), (503) and

[0129] (603) and the video decoders (310), (410) and (710). In an embodiment, one or more integrated circuits can be used to implement the video encoders (303), (503) and (603) and the video decoders (310), (410) and (710). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (303), (503) and

[0130] (603) and the video decoders (310), (410) and (710).

[0131] II. Block Partitioning

[0132] Figure 8 An exemplary block partitioning according to some embodiments of the present disclosure is shown.

[0133] In some related examples such as VP9 proposed by the Alliance for Open Media (AOMedia), a 4-way partitioning tree can be used, which goes from the 64×64 level down to the 4×4 level, with some additional restrictions on blocks of 8×8 and below, as Figure 8 shown. Note that the partitioning designated as R can be referred to as recursive partitioning. That is, the same partitioning tree is repeated at a lower scale until the lowest 4×4 level is reached.

[0134] In some related examples such as AV1 proposed by AOMedia, the partitioning tree can be extended to a 10-way structure as Figure 8 shown, and the maximum coding block size (referred to as a superblock in VP9 / AV1 terminology) is increased to start from 128×128. Note that 4:1 / 1:4 rectangular partitioning is included in AV1 but not in VP9. None of the rectangular partitions can be further subdivided. Additionally, AV1 can support greater flexibility when using partitions below the 8×8 level, as in some examples inter-frame prediction can be performed on 2×2 chrominance blocks.

[0135] In some related examples such as HEVC, a CTU can be partitioned into CUs by using a quadtree structure represented as an encoding tree to adapt to various local features. A decision can be made at the CU level on whether to use inter-picture (temporal) prediction or intra-picture (spatial) prediction to encode a picture region. Each CU can be further partitioned into one, two, or four PUs according to the PU partition type. Inside a PU, the same prediction process can be applied, and relevant information can be transmitted to the decoder based on the PU. After obtaining a residual block by applying the prediction process based on the PU partition type, the CU can be split into TUs according to another quadtree structure such as the encoding tree of the CU. A key feature of the HEVC structure is that it has multiple partitioning concepts including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square-shaped, while a PU can be square or rectangular for an inter-prediction block. In HEVC, an encoding block can be further partitioned into four square sub-blocks, and a transform process can be performed on each sub-block, i.e., TU. Each TU can be further recursively partitioned (e.g., using quadtree partitioning) into smaller TUs. Quadtree partitioning can be referred to as a Residual Quadtree (RQT).

[0136] At the picture boundary, HEVC adopts implicit quadtree partitioning so that the block can continue to perform quadtree partitioning until the size of the block fits the picture boundary.

[0137] Figure 9 An exemplary quadtree with a nested binary tree structure according to an embodiment of the present disclosure is shown.

[0138] Figure 10 An exemplary block partition in a multi-type tree structure according to some embodiments of the present disclosure is shown.

[0139] In some related examples such as VVC, a Multi-Type-Tree (MTT) structure can be used. The MTT structure is a combination of a Quadtree (QT) nested with a Binary Tree (BT) and a Triple (Ternary) Tree (TT). First, a Coding Tree Unit (CTU) or Coding Unit (CU) can be recursively divided into square blocks by the QT. Then, each QT leaf can be further divided by the BT or TT, where the BT and TT divisions can be recursively applied and interleaved, but no additional QT division can be applied. In some examples, the TT divides a rectangular block vertically or horizontally into three blocks using a 1:2:1 ratio to avoid non-power-of-two widths and heights. To prevent partitioning competition, additional partitioning constraints are usually imposed on the MTT to avoid duplicate partitioning (e.g., prohibiting vertical / horizontal binary partitioning on intermediate partitions resulting from vertical / horizontal ternary partitioning). Additional limits are set for the maximum depth of the BT and TT divisions.

[0140] Figure 11 An exemplary L-shaped partitioning according to an embodiment of the present disclosure is shown. Instead of using rectangular block partitioning, the L-shaped split can divide a block into one or more L-shaped partitions and one or more rectangular partitions. As Figure 11 shown, an L-shaped (or L-type) partition can have a width, a height, a shorter width, and a shorter height. In the present disclosure, a rotated L-shaped partition can also be considered an L-shaped partition.

[0141] Figure 12 An exemplary block partitioning using the L-shaped split according to some embodiments of the present disclosure is shown. Based on the L-shaped partition, a block can be split into two partitions, including an L-shaped partition (Partition 1) and a rectangular partition (Partition 0).

[0142] III. Intra Prediction

[0143] In some related examples such as VP9, 8 direction modes are supported, and the 8 direction modes correspond to angles from 45 to 207 degrees. To utilize more types of spatial redundancy in the directional texture, in some related examples such as AV1, the directional intra mode is extended to a set of angles with finer granularity. The original 8 angles are slightly changed and are called nominal angles, and these 8 nominal angles are named V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED.

[0144] Figure 13An exemplary nominal angle according to an embodiment of the present disclosure is shown. Each nominal angle may be associated with seven finer angles, so in some related examples such as AV1, there may be a total of 56 directional angles. The predicted angle may be represented by adding an angle increment to the nominal intra-frame angle. The angle increment may be equal to a coefficient multiplied by a step size of 3 degrees. The coefficient may be in the range of -3 to 3. To implement the directional prediction mode in AV1 in a general way, all 56 directional intra-frame prediction angles in AV1 may be implemented with a unified directional predictor that projects each pixel to a reference sub-pixel position and interpolates the reference sub-pixels with a 2-tap bilinear filter.

[0145] In some related examples such as AV1, there are five non-directional smooth intra-frame prediction modes, which are DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. For DC prediction, the average of the left and upper neighboring samples is used as the predictor for the block to be predicted. For

[0146] PAETH prediction, first the top, left, and upper-left reference samples are obtained, and then the value closest to (top + left - upper-left) is set as the predictor for the pixel to be predicted.

[0147] Figure 14 The positions of the top, left, and upper-left samples of a pixel in the current block according to an embodiment of the present disclosure are shown. For the SMOOTH mode, SMOOTH_V mode, and SMOOTH_H mode, quadratic interpolation in the vertical or horizontal direction or the average of the two directions is used to predict the block.

[0148] Figure 15 An exemplary recursive filter intra-frame mode according to an embodiment of the present disclosure is shown.

[0149] To capture the decaying spatial correlation regarding edges, the FILTER INTRA mode is designed for luminance blocks. Five filter intra-frame modes are defined in AV1, and each filter intra-frame mode is represented by a set of eight 7-tap filters that reflect the correlation between the pixels in a 4×2 tile and the seven neighboring pixels adjacent to the tile. For example, the weighting factors of the 7-tap filters are position-dependent. As Figure 15As shown, the 8×8 block is divided into eight 4×2 tiles indicated by B0, B1, B2, B3, B4, B5, B6, and B7. For each tile, its 7 neighboring tiles indicated by R0 to R7 are used to predict the pixels in the corresponding tile. For tile B0, all neighboring tiles have been reconstructed. However, for other tiles, not all neighboring tiles are reconstructed, and then the prediction values of the closest neighbors are used as reference values. For example, all neighboring tiles of tile B7 have not been reconstructed, so the prediction samples of the neighboring tiles of tile B7 (i.e., B5 and B6) are used alternatively.

[0150] For the chrominance component, the chroma-only intra prediction mode, called Chroma from Luma (CfL) mode, models the chroma pixels as a linear function of the coincidentally reconstructed luma pixels. The CfL prediction can be expressed as follows:

[0151] CfL(α) = α × L AC + DC Equation (1)

[0152] where L AC represents the AC contribution of the luma component, α represents the parameter of the linear model, and DC represents the DC contribution of the chroma component. In the example, the reconstructed luma pixels are subsampled to the chroma resolution and then the average value is subtracted to form the AC contribution. To approximate the chroma AC component from the AC contribution, as in some related examples, the CfL mode in AC1 determines the parameter α based on the original chroma pixels and signals them in the bitstream, rather than requiring the decoder to calculate the scaling parameter. This reduces the decoder complexity and produces a more accurate prediction. Regarding the DC contribution of the chroma component, it is calculated using the intra DC mode, which is sufficient for most chroma content and has a mature fast implementation.

[0153] Figure 16 Illustrates an exemplary multi-line intra prediction using four reference lines adjacent to the coding block unit according to an embodiment of the present disclosure. For multi-line intra prediction, the encoder determines and signals which reference lines are used to generate the intra predictor. The reference line index is signaled before the intra prediction mode, and only the most likely mode is allowed in the case of signaling a non-zero reference line index. In Figure 16 an example of 4 reference lines is depicted, where each reference line consists of four segments (i.e., segments A to D) together with the top-left reference sample. Additionally, in Figure 16 the reconstructed samples in different reference lines are filled with different patterns. The multi-line intra prediction mode can also be referred to as the multiple reference line prediction (MRLP) mode.

[0154] IV. Non-directional intra prediction for L-shaped partitions

[0155] With an L-shaped partition, neighboring reconstruction samples of the current block can be obtained from the right side and / or the bottom side of the current block. However, the available neighboring reconstruction samples from the right side and / or the bottom side are not fully compatible with some related intra prediction schemes that use top and left reference samples to perform non-directional intra prediction.

[0156] The present disclosure includes methods for non-directional intra prediction modes for L-shaped partitions. The proposed methods can be used alone or in any order of combination. In the present disclosure, an L-shaped (or L-type) partition can be defined as shown, and a rotated L-shaped partition can also be regarded as an L-shaped partition. Figure 11 as shown, and a rotated L-shaped partition can also be regarded as an L-shaped partition.

[0157] Intra prediction modes can include different types of intra prediction modes, such as angular intra prediction modes or directional intra prediction modes and non-angular intra prediction modes or non-directional intra prediction modes. For example, if the predicted samples of a mode can be generated according to a given prediction direction, the mode can be called an angular intra prediction mode or a directional intra prediction mode. Otherwise, the mode can be called a non-angular intra prediction mode or a non-directional intra prediction mode. Examples of non-angular intra prediction modes include but are not limited to DC mode, Planar mode, Plane mode (defined in H.264 / AVC), SMOOTH mode, SMOOTH_H mode, SMOOTH_V mode, Paeth mode, recursive filtering mode, and / or Matrix-based Intra Prediction (MIP) mode. In some embodiments, modes that are not smooth modes can be regarded as angular intra prediction modes or directional intra prediction modes.

[0158] In related intra prediction schemes, top and / or left neighboring reference samples are used to perform non-directional intra prediction modes. However, for an L-shaped partition, additional neighboring samples may be available and reconstructed. For example, right and / or bottom neighboring samples may be available and reconstructed, and thus can be used to predict the L-shaped partition.

[0159] According to aspects of the present disclosure, when a block is divided into at least one L-shaped partition (LP) and at least one rectangular partition (RP), the reference samples for performing the intra prediction mode of the L-shaped partition can come from neighboring reconstruction samples of another LP or RP or other blocks. In some embodiments, the neighboring reconstruction samples can form a continuous chain of any shape, rather than a horizontal line and / or a vertical line.

[0160] In the present disclosure, the reference samples together may be referred to as a Reference Sample Chain (RSC). All samples or subsets of samples in the RSC may be used for non-directional intra prediction modes. The RSC may include reference samples of more than one horizontal line or vertical line.

[0161] Figures 17A to 17F Six exemplary RSCs according to some embodiments of the present disclosure are shown. Figures 17A to 17F Each block in has a size of 8×8 and is divided into two partitions: an LP and an RP. The RP has a size of 4×4 and is located at the upper left corner of each block. The LP has a height of 8 and a width of 8. Figures 17A to 17F Each RSC in includes reference samples of two horizontal lines and two vertical lines of reference samples.

[0162] According to some embodiments of the present disclosure, the total number of reference samples included in the RSC may be a power of 2. One or more samples in the RSC may be excluded from the reference samples such that the total does not exceed the total number of reference samples. For example, in Figures 17A to 17F , the total number of reference samples included in each RSC is 16. For Figures 17A to 17C each RSC in, one corner sample of the corresponding RSC is excluded from the reference samples such that the total number of reference samples included in the corresponding RSC is 16. For Figure 17D and Figure 17E each RSC in, one sample at the head or tail of the corresponding RSC is excluded from the reference samples such that the total number of reference samples included in the corresponding RSC is 16. For Figure 17F the RSC in, two corner samples are excluded from the reference samples, and one intermediate corner sample is used twice in the intra prediction mode (e.g., DC mode) such that the total number of reference samples included in the RSC is 16.

[0163] In one embodiment, only a subset of the reference samples in the RSC may be used for the intra prediction mode (e.g., DC mode).

[0164] In some embodiments, a block may be divided into two partitions: an LP and an RP. The RP is located at the upper left corner of the block, and the height and width of the LP may be equal or unequal. When performing a non-directional intra prediction mode (e.g., DC mode) on the LP, in some embodiments, the total number of reference samples used in the non-directional intra prediction mode may be the sum of the width and height of the LP (e.g., width + height), such as Figures 17A to 17F shown. In an embodiment, the total number of reference samples used in the non-directional intra prediction mode may be the sum of the shorter width and shorter height of the LP (e.g., shorter width + shorter height). Figure 18An example is shown in which the reference samples used are marked in gray.

[0165] In some embodiments, a block can be divided into two partitions: an LP and an RP. The RP is located at the upper left corner of the block, and the height and width of the LP are not equal. When performing a non-directional intra prediction mode (e.g., DC mode) on the LP, the total number of reference samples used in the non-directional intra prediction mode is the maximum or minimum of the width and height of the LP (e.g., max(width, height) or min(width, height)). For example, in Figure 19A , the width of the LP is greater than the height of the LP, so the value of the width is selected as the total number of reference samples used in the non-directional intra prediction mode (e.g., DC mode). In Figure 19B , the height of the LP is greater than the width of the LP, so the value of the height is selected as the total number of reference samples used in the non-directional intra prediction mode (e.g., DC mode). In Figure 19A and Figure 19B the total number of reference samples used in the prediction process is 16.

[0166] In some embodiments, a block can be divided into two partitions: an LP and an RP. When the height and width of the LP are not equal or the RP is not located at the lower right corner of the block, only the reference samples along the vertical or horizontal side of the block are used in the non-directional intra prediction mode (e.g., DC mode). The total number of reference samples used in the non-directional intra prediction mode is the maximum or minimum of the width and height of the LP (e.g., max(width, height) or min(width, height)). Figures 20A to 20D Some exemplary reference samples for the LP according to some embodiments of the present disclosure are shown.

[0167] In one embodiment, a block can be divided into two partitions: an LP and an RP. The LP can be Figure 12 one of the four L-shaped types in Figure 21 When performing a non-directional intra prediction mode (e.g., DC mode) on the LP, the total number of reference samples used in the non-directional intra prediction mode is the sum of the width and height of the LP (e.g., width + height), and all reference samples are outside the LP and RP partitions.

[0168] According to some embodiments of the present disclosure, a block can be divided into multiple partitions. For the current partition, when reconstructing the right and / or bottom neighboring samples from different partitions (LP or RP) before reconstructing the samples of the current partition, the right and / or bottom neighboring samples can form an RSC and be used to perform a non-directional intra prediction mode (e.g., DC mode) on the current partition. AsFigures 22A to 22B As shown, LP (Partition 1) is reconstructed before RP (Partition 0). Thus, samples of LP can form an RSC and be used for non-directional intra prediction modes (e.g., DC mode) of RP. In Figures 22A to 22B , the reference samples in the upper row of RP are marked in dark gray, the reference samples in the left column of RP are marked in gray, and the reference samples in the right column or bottom row of RP are marked in white.

[0169] In one embodiment, only the neighboring samples in one of the upper row, left column, right column, and bottom row of the RP block can be used for non-directional intra prediction modes (e.g., DC mode) of RP.

[0170] In one embodiment, only the neighboring samples in the left column and upper row of RP can be used for non-directional intra prediction modes (e.g., DC mode) of RP.

[0171] In one embodiment, as Figure 22A shown, when RP is located at the lower left corner of the block, only the neighboring samples in the left column and right column of RP can be used for non-directional intra prediction modes (e.g., DC mode) of RP.

[0172] In one embodiment, as Figure 22B shown, when RP is located at the upper right corner of the block, only the neighboring samples in the upper row and bottom row of RP can be used for non-directional intra prediction modes (e.g., DC mode) of RP.

[0173] According to aspects of the present disclosure, when performing one of certain non-directional intra prediction modes (e.g., the planar mode defined in HEVC and VVC, the SMOOTH, SMOOTH-H, or SMOOTH-V mode defined in AV1) and reconstructing right or bottom neighboring samples, the reconstructed neighboring samples can be directly used for 4-tap interpolation in the non-directional intra prediction mode, rather than extrapolating the right and / or bottom neighboring samples from the top and left reconstructed neighboring samples.

[0174] In one embodiment, when performing one of certain non-directional intra prediction modes (e.g., the planar mode defined in HEVC and VVC, the SMOOTH, SMOOTH-H, or SMOOTH-V mode defined in AV1) and the bottom row neighboring samples are not available, the bottom row neighboring samples can be linearly extrapolated from the left column and right column neighboring samples. As Figure 23AAs shown, if the bottom-left neighboring sample (labeled BL) is not available, the BL neighboring sample can be used directly or obtained by copying from the nearest neighbor in the left column, and the bottom-right neighboring sample (labeled BR) can be obtained by copying from the nearest neighbor in the right column. The remaining bottom-row neighboring samples between the BL and BR neighboring samples can be extrapolated by using, for example, linear interpolation.

[0175] In one embodiment, when performing one of certain non-directional intra prediction modes (e.g., the planar mode defined in HEVC and VVC, the SMOOTH, SMOOTH-H, or SMOOTH-V modes defined in AV1) and the right-column neighboring samples are not available, the right-column neighboring samples can be linearly extrapolated from the top-row and bottom-row neighboring samples. As Figure 23B shown, if the top-right neighboring sample (labeled TR) is not available, the TR neighboring sample can be used directly or obtained by copying from the nearest neighbor in the top row, and the bottom-right neighboring sample (labeled BR) can be obtained by copying from the nearest neighbor in the bottom row. The remaining right-column neighboring samples between the TR and BR neighboring samples can be extrapolated by using, for example, linear interpolation.

[0176] In one embodiment, when performing one of certain non-directional intra prediction modes (e.g., the planar mode defined in HEVC and VVC, the SMOOTH, SMOOTH-H, or SMOOTH-V modes defined in AV1), only the top and left neighboring samples outside the RP and LP blocks can be used as reference samples, and the right and bottom neighboring samples can be obtained by copying or extrapolating from the top and left neighboring samples. Figure 24 An example showing how to select the left and upper neighboring samples for LP in such an embodiment.

[0177] In one embodiment, when performing one of certain non-directional intra prediction modes (e.g., the planar mode defined in HEVC and VVC, the SMOOTH, SMOOTH-H, or SMOOTH-V modes defined in AV1), for samples located at different positions in the LP, the left, right, top, and bottom neighboring reference samples can be from different rows, and the right and bottom neighboring reference samples (marked with diagonal textures in Figures 25A to 25B can be obtained by copying or extrapolating from the top and left neighboring reference samples.

[0178] Figures 25A to 25B Two examples showing how to select the left, right, top, and bottom neighboring reference samples for LP.

[0179] In Figure 25AIn it, block (2501) is divided into LP (labeled 1) and RP (labeled 0). RP is located at the lower left corner of block (2501). For sample (2510) in LP, the top neighboring reference sample (2511) comes from the top reference row of block (2501), the left neighboring reference sample (2512) comes from the left reference row of block (2501), the bottom neighboring reference sample (2513) comes from the top row of RP, and the right neighboring reference sample (2514) comes from the right reference row of block (2501). Note that the reference sample in the right reference row of block (2501), such as the right neighboring reference sample (2514), can be obtained by copying or extrapolating from the reference sample in the top reference row of block (2501).

[0180] For sample (2520) in LP, the top neighboring reference sample (2521) comes from the top reference row of block (2501), the left neighboring reference sample (2522) comes from the right row of RP, the bottom neighboring reference sample (2523) comes from the bottom reference row of block (2501), and the right neighboring reference sample (2524) comes from the right reference row of block (2501). Note that the reference sample in the bottom reference row of block (2501), such as the bottom neighboring reference sample (2523), can be obtained by copying or extrapolating from the reference sample in the left reference row of block (2501).

[0181] In Figure 25B it, block (2502) is divided into LP (labeled 1) and RP (labeled 0). RP is located at the upper left corner of block (2502). For sample (2530) in LP, the top neighboring reference sample (2531) comes from the top reference row of block (2502), the left neighboring reference sample (2532) comes from the right row of RP, the bottom neighboring reference sample (2533) comes from the bottom reference row of block (2502), and the right neighboring reference sample (2534) comes from the right reference row of block (2502). Note that the reference sample in the right reference row of block (2502), such as the right neighboring reference sample (2534), can be obtained by copying or extrapolating from the reference sample in the top reference row of block (2502).

[0182] For sample (2540) in LP, the top neighboring reference sample (2541) comes from the bottom row of RP, the left neighboring reference sample (2542) comes from the left reference row of block (2502), the bottom neighboring reference sample (2543) comes from the bottom reference row of block (2502), and the right neighboring reference sample (2544) comes from the right reference row of block (2502). Note that the reference sample in the bottom reference row of the block, such as the bottom neighboring reference sample (2543), can be obtained by copying or extrapolating from the reference sample in the left reference row of block (2502).

[0183] V. Flowchart

[0184] Figure 26 A flowchart is shown that outlines an exemplary process (2600) according to embodiments of the present disclosure. In various embodiments, the process (2600) is performed by processing circuitry, such as the processing circuitry in terminal devices (210), (220), (230), and (240), the processing circuitry that performs the functions of video encoder (303), the processing circuitry that performs the functions of video decoder (310), the processing circuitry that performs the functions of video decoder (410), the processing circuitry that performs the functions of intra prediction module (452), the processing circuitry that performs the functions of video encoder (503), the processing circuitry that performs the functions of predictor (535), the processing circuitry that performs the functions of intra encoder (622), the processing circuitry that performs the functions of intra decoder (772), and so on. In some embodiments, the process (2600) is implemented as software instructions, and thus when the processing circuitry executes the software instructions, the processing circuitry performs the process (2600).

[0185] The process (2600) generally may begin with step (S2610), in which the process (2600) decodes prediction information for a current block in a current picture that is part of an encoded video bitstream. The prediction information indicates a non-directional intra prediction mode for the current block. Then, the process (2600) proceeds to step (S2620).

[0186] At step (S2620), the process (2600) divides the current block into a plurality of partitions. The plurality of partitions includes at least one L-shaped partition. Then, the process (2600) proceeds to step (S2630).

[0187] At step (S2630), the process (2600) reconstructs one of the plurality of partitions based on at least one of the following: (i) neighboring reconstruction samples of one of the plurality of partitions; or (ii) neighboring reconstruction samples of the current block. Then, the process (2600) terminates.

[0188] In one embodiment, at least one of the neighboring reconstruction samples is adjacent to one of the right or bottom sides of one of the plurality of partitions.

[0189] In one embodiment, one of the plurality of partitions is an L-shaped partition, and the number of neighboring reconstruction samples depends on the size of the L-shaped partition. In one example, the number of neighboring reconstruction samples is the sum of the width and height of the L-shaped partition. In another example, the number of neighboring reconstruction samples is the sum of the shorter width and the shorter height of the L-shaped partition. In another example, the number of neighboring reconstruction samples is the maximum between the width and height of the L-shaped partition. In another example, the number of neighboring reconstruction samples is the minimum between the width and height of the L-shaped partition.

[0190] In one embodiment, at least one of the neighboring reconstruction samples is located in another partition among the plurality of partitions reconstructed before one of the plurality of partitions. In the example, another partition among the plurality of partitions is an L-shaped partition, and at least one of the neighboring reconstruction samples is adjacent to one of the right side or the bottom side of one of the plurality of partitions.

[0191] In one embodiment, the processing (2600) determines a plurality of neighboring reference samples for one of the plurality of partitions based on at least one of the following: (i) the neighboring reconstruction samples of one of the plurality of partitions; or

[0192] (ii) the neighboring reconstruction samples of the current block. The processing (2600) reconstructs one of the plurality of partitions based on the plurality of neighboring reference samples.

[0193] In one example, the neighboring reconstruction samples include the neighboring reconstruction samples of the left column and the right column of one of the plurality of partitions. The processing (2600) determines the neighboring reference samples of the bottom row of one of the plurality of partitions based on the neighboring reconstruction samples of the left column and the right column of one of the plurality of partitions. The processing (2600) reconstructs one of the plurality of partitions based on the neighboring reference samples of the bottom row of one of the plurality of partitions.

[0194] In one example, the neighboring reconstruction samples include the neighboring reconstruction samples of the top row and the bottom row of one of the plurality of partitions. The processing (2600) determines the neighboring reference samples of the right column of one of the plurality of partitions based on the neighboring reconstruction samples of the top row and the bottom row of one of the plurality of partitions. The processing (2600) reconstructs one of the plurality of partitions based on the neighboring reference samples of the right column of one of the plurality of partitions.

[0195] In one embodiment, one of the plurality of partitions is an L-shaped partition, and the processing (2600) reconstructs one of the plurality of partitions based on the neighboring reconstruction samples of the left column and the top row of the current block.

[0196] In one embodiment, based on one of the plurality of partitions being an L-shaped partition, the processing (2600) determines a plurality of neighboring reference samples for each sample of the L-shaped partition based on the position of the sample. The processing (2600) reconstructs each sample of the L-shaped partition based on the plurality of neighboring reference samples of the sample.

[0197] In one embodiment, the plurality of neighboring reference samples for each sample includes one of the reconstructed neighboring samples and the neighboring samples to be reconstructed based on the reconstructed neighboring samples.

[0198] VI. Computer System

[0199] The above techniques can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 27 A computer system (2700) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0200] The computer software can be encoded using any suitable machine code or computer language, which can be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. or executed through interpretation, microcode execution, etc.

[0201] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0202] Figure 27 The components shown for the computer system (2700) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system (2700).

[0203] The computer system (2700) can include certain human-machine interface input devices. Such human-machine interface input devices can respond to inputs made by one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, taps), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices can also be used to capture certain media that are not necessarily directly related to conscious inputs of humans, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0204] The input human-machine interface device may include one or more of the following (only one of each type is depicted): keyboard (2701), mouse (2702), touchpad (2703), touch screen (2710), data glove (not shown), joystick (2705), microphone (2706), scanner (2707), and camera device (2708).

[0205] The computer system (2700) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., through the touch screen (2710), data glove (not shown), or joystick (2705), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (2709), headphones (not depicted)), visual output devices (e.g., screen (2710), which includes CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input capabilities, each screen having or not having tactile feedback capabilities - some of these screens may be capable of outputting two-dimensional visual output or more than three-dimensional output in ways such as stereoscopic graphics output; virtual reality glasses (not depicted), holographic displays, and smell cans (not depicted)), and printers (not depicted). These visual output devices (e.g., screen (2710)) may be connected to the system bus (2748) through a graphics adapter (2750).

[0206] The computer system (2700) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2720) with CD / DVD or similar media (2721), thumb drives (2722), removable hard disk drives or solid-state drives (2723), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0207] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the currently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0208] The computer system (2700) may also include a network interface (2754) to one or more communication networks (2755). The one or more communication networks (2755) may be, for example, wireless, wired, optical. The one or more communication networks (2755) may also be local, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of the one or more communication networks (2755) include: local area networks such as Ethernet; wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television; vehicular and industrial networks including CANBus, etc. Certain networks typically require an external network interface adapter that is attached to certain general-purpose data ports or peripheral buses (2749) (such as, for example, the USB port of the computer system (2700)); other networks are typically integrated into the core of the computer system (2700) by attaching to a system bus as described below (for example, an Ethernet interface to a PC computer system or a cellular network interface to a smart phone computer system). The computer system (2700) may use any of these networks to communicate with other entities. Such communication may be only one-way reception (e.g., broadcast television), only one-way transmission (e.g., CANBus to certain CANBus devices), or two-way (e.g., using a local digital network or a wide area digital network to other computer systems). Certain protocols and protocol stacks may be used on each of these networks and network interfaces as described above.

[0209] The foregoing human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core (2740) of the computer system (2700).

[0210] The core (2740) may include one or more central processing units (CPUs) (2741), a graphics processing unit (GPU) (2742), a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) (2743), a hardware accelerator (2744) for certain tasks, a graphics adapter (2750), etc. These devices, together with a read-only memory (ROM) (2745), a random access memory (2746), and an internal mass storage device (2747) such as an internal non-user-accessible hard disk drive, SSD, etc., can be connected via a system bus (2748). In some computer systems, the system bus (2748) can be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the system bus (2748) of the core or can be attached to the system bus (2748) via a peripheral bus (2749). In an example, a screen (2710) can be connected to the graphics adapter (2750). The architecture of the peripheral bus includes PCI, USB, etc.

[0211] The CPU (2741), GPU (2742), FPGA (2743), and accelerator (2744) can execute certain instructions, and the instructions can be combined to form the aforementioned computer code. The computer code can be stored in the ROM (2745) or the RAM (2746). Interim data can also be stored in the RAM

[0212] (2746), while permanent data can be stored in, for example, the internal mass storage device (2747). Fast storage and retrieval of any storage device in the storage device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (2741), GPUs (2742), mass storage devices (2747), ROMs (2745), RAMs (2746), etc.

[0213] Computer-readable media can have computer code for performing various computer-implemented operations. The media and the computer code can be media and computer code that are specifically designed and constructed for the purposes of this disclosure, or the media and the computer code can be of the type that is well known and available to those skilled in the field of computer software.

[0214] By way of example and not limitation, a computer system (2700) having an architecture and in particular a core (2740) can provide functionality due to software contained in one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such a computer-readable medium can be a medium associated with a user-accessible mass storage device as introduced above and certain storage devices of a non-transitory nature in the core (2740), such as an on-core mass storage device

[0215] (2747) or a ROM (2745). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2740). Depending on specific needs, the computer-readable medium can include one or more memory devices or chips. The software can cause the core (2740) and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to perform specific processes or specific parts of specific processes described herein, including defining data structures stored in a RAM (2746) and modifying such data structures according to the processes defined by the software. Additionally or alternatively, the computer system can provide functionality due to logic embodied in a hardwired or other manner in a circuit (e.g., an accelerator (2744)), which can operate in place of or in conjunction with the software to perform specific processes or specific parts of specific processes described herein. In appropriate cases, references to software can include logic, and references to logic can also include software. In appropriate cases, references to a computer-readable medium can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit implementing logic for execution, or both of the above. The present disclosure encompasses any suitable combination of hardware and software.

[0216] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that although not explicitly shown or described herein, those skilled in the art will be able to envision many systems and methods that embody the principles of the present disclosure and are thus within its spirit and scope.

[0217] Appendix A: Acronyms

[0218] ALF: Adaptive Loop Filter

[0219] AMVP: Advanced Motion Vector Prediction

[0220] APS: Adaptive Parameter Set

[0221] ASIC: Application Specific Integrated Circuit

[0222] ATMVP: Alternate / Advanced Temporal Motion Vector Prediction

[0223] AV1: Aomedia Video 1

[0224] AV2: Aomedia Video 2

[0225] BMS: Baseline Set

[0226] BV: Block Vector

[0227] CANBus: Controller Area Network Bus

[0228] CB: Coding Block

[0229] CC-ALF: Cross-Component Adaptive Loop Filter

[0230] CD: Compact Disc

[0231] CDEF: Constrained Direction Enhancement Filter

[0232] CPR: Current Picture Reference

[0233] CPU: Central Processing Unit

[0234] CRT: Cathode Ray Tube

[0235] CTB: Coding Tree Block

[0236] CTU: Coding Tree Unit

[0237] CU: Coding Unit

[0238] DPB: Decoder Picture Buffer

[0239] DPS: Decoding Parameter Set

[0240] DVD: Digital Video Disc

[0241] FPGA: Field Programmable Gate Array

[0242] JCCR: Joint CbCr Residual Coding

[0243] JVET: Joint Video Exploration Team

[0244] GOP: Group of Pictures

[0245] GPU: Graphics Processing Unit

[0246] GSM: Global System for Mobile Communications

[0247] HDR: High Dynamic Range

[0248] HEVC: High Efficiency Video Coding

[0249] HRD: Hypothetical Reference Decoder

[0250] IBC: Intra Block Copy

[0251] IC: Integrated Circuit

[0252] ISP: Intra Sub - Partition

[0253] JEM: Joint Exploration Model

[0254] LAN: Local Area Network

[0255] LCD: Liquid Crystal Display

[0256] LR: Loop Restorer

[0257] LTE: Long Term Evolution

[0258] MPM: Most Probable Mode

[0259] MV: Motion Vector

[0260] OLED: Organic Light - Emitting Diode

[0261] PB: Prediction Block

[0262] PCI: Peripheral Component Interconnect

[0263] PDPC: Position - Dependent Prediction Combination

[0264] PLD: Programmable Logic Device

[0265] PPS: Picture Parameter Set

[0266] PU: Prediction Unit

[0267] RAM: Random Access Memory

[0268] ROM: Read - Only Memory

[0269] SAO: Sample Adaptive Offset

[0270] SCC: Screen Content Coding

[0271] SDR: Standard Dynamic Range

[0272] SEI: Supplemental Enhancement Information

[0273] SNR: Signal - to - Noise Ratio

[0274] SPS: Sequence Parameter Set

[0275] SSD: Solid State Drive

[0276] TU: Transform Unit

[0277] USB: Universal Serial Bus

[0278] VPS: Video Parameter Set

[0279] VUI: Video Usability Information

[0280] VVC: Versatile Video Coding

[0281] WAIP: Wide Angle Intra Prediction

Claims

1. A video decoding method, characterized in that, Comprising: Decoding prediction information of a current block in a current picture that is part of an encoded video bitstream, the prediction information indicating a non-directional intra prediction mode for the current block; Dividing the current block into a plurality of partitions, the plurality of partitions including at least one L-shaped partition; And Reconstructing one of the plurality of partitions based on at least one of the following: (i) neighboring reconstructed samples of one of the plurality of partitions; or (ii) neighboring reconstructed samples of the current block, the neighboring reconstructed samples including a plurality of reference samples of a horizontal line or a vertical line, and the number of the neighboring reconstructed samples being one of the following: (i) the sum of the width and height of the L-shaped partition; (ii) the sum of the shorter width and the shorter height of the L-shaped partition; (iii) the maximum value of the width and height of the L-shaped partition; and (iv) the minimum value of the width and height of the L-shaped partition.

2. The method according to claim 1, wherein At least one of the neighboring reconstructed samples is located in another partition among the plurality of partitions that is reconstructed before one of the plurality of partitions.

3. The method according to claim 2, wherein, Another partition among the plurality of partitions is an L-shaped partition, and at least one of the neighboring reconstructed samples is adjacent to the right side or the bottom side of one of the plurality of partitions.

4. The method according to any one of claims 1 to 3, wherein The neighboring reconstructed samples include neighboring reconstructed samples of the left column and the right column of one of the plurality of partitions, and reconstructing the plurality of partitions includes: Determining neighboring reference samples of the bottom row of one of the plurality of partitions based on the neighboring reconstructed samples of the left column and the right column of one of the plurality of partitions; and Reconstructing one of the plurality of partitions based on the neighboring reference samples of the bottom row of one of the plurality of partitions.

5. The method according to any one of claims 1 to 3, wherein The neighboring reconstructed samples include neighboring reconstructed samples of the top row and the bottom row of one of the plurality of partitions, and reconstructing the plurality of partitions includes: Determining neighboring reference samples of the right column of one of the plurality of partitions based on the neighboring reconstructed samples of the top row and the bottom row of one of the plurality of partitions; and Reconstructing one of the plurality of partitions based on the neighboring reference samples of the right column of one of the plurality of partitions.

6. The method according to claim 1, wherein One of the plurality of partitions is an L-shaped partition, and reconstructing the plurality of partitions includes: Reconstructing one of the plurality of partitions based on the neighboring reconstructed samples of the left column and the top row of the current block.

7. The method according to claim 1, wherein, Based on one of the plurality of partitions being an L-shaped partition, reconstructing the plurality of partitions includes: For each sample of the L-shaped partition, determining a plurality of neighboring reference samples based on the position of the sample; and Reconstructing each sample of the L-shaped partition based on the plurality of neighboring reference samples of the sample.

8. The method according to claim 7, wherein The plurality of neighboring reference samples of each sample include one of the neighboring reconstructed samples and neighboring samples to be reconstructed based on the neighboring reconstructed samples.

9. The method according to any one of claims 1 to 3, wherein Reconstructing the plurality of partitions includes: Determining a plurality of neighboring reference samples of one of the plurality of partitions based on at least one of the following: (i) neighboring reconstructed samples of one of the plurality of partitions; or (ii) neighboring reconstructed samples of the current block; Reconstructing one of the plurality of partitions based on the plurality of neighboring reference samples.

10. An apparatus comprising processing circuitry, characterized in that, The processing circuit system is configured to: Decode prediction information for a current block in a current picture that is part of an encoded video bitstream, the prediction information indicating a non-directional intra prediction mode for the current block; Partition the current block into a plurality of partitions, the plurality of partitions including at least one L-shaped partition; and Reconstruct one of the plurality of partitions based on at least one of the following: (i) neighboring reconstructed samples of one of the plurality of partitions; or (ii) neighboring reconstructed samples of the current block, the neighboring reconstructed samples including a plurality of reference samples of a horizontal line or a vertical line, the number of the neighboring reconstructed samples being one of the following: (i) the sum of the width and height of the L-shaped partition; (ii) the sum of the shorter width and the shorter height of the L-shaped partition; (iii) the maximum of the width and height of the L-shaped partition; and (iv) the minimum of the width and height of the L-shaped partition.

11. The apparatus according to claim 10, wherein, At least one of the neighboring reconstructed samples is located in another partition among the plurality of partitions that is reconstructed before the one of the plurality of partitions.

12. The apparatus according to claim 11, wherein another partition among the plurality of partitions is an L-shaped partition, and at least one of the neighboring reconstructed samples is adjacent to the right side or the bottom side of one of the plurality of partitions.

13. The apparatus according to any one of claims 10 to 12, wherein, The neighboring reconstructed samples include neighboring reconstructed samples of the left column and the right column of one of the plurality of partitions, and the processing circuitry is further configured to: Determine neighboring reference samples of the bottom row of one of the plurality of partitions based on the neighboring reconstructed samples of the left column and the right column of one of the plurality of partitions; and Reconstruct one of the plurality of partitions based on the neighboring reference samples of the bottom row of one of the plurality of partitions.

14. The device according to any one of claims 10 to 12, wherein, The neighboring reconstructed samples include neighboring reconstructed samples of the top row and the bottom row of one of the plurality of partitions, and the processing circuitry is further configured to: Determine neighboring reference samples of the right column of one of the plurality of partitions based on the neighboring reconstructed samples of the top row and the bottom row of one of the plurality of partitions; and Reconstruct one of the plurality of partitions based on the neighboring reference samples of the right column of one of the plurality of partitions.

15. The apparatus according to claim 10, wherein One of the plurality of partitions is an L-shaped partition, and the processing circuitry is further configured to: Reconstruct one of the plurality of partitions based on the neighboring reconstructed samples of the left column and the top row of the current block.

16. The device according to claim 10, wherein, Based on that one of the plurality of partitions is an L-shaped partition, the processing circuitry is further configured to: For each sample of the L-shaped partition, determine a plurality of neighboring reference samples based on the position of the sample; and Reconstruct each sample of the L-shaped partition based on the plurality of neighboring reference samples of the sample.

17. An electronic device, comprising a memory and a processor, characterized in that, The memory stores computer instructions, and when the computer instructions are run by the processor, the electronic device is caused to execute the method according to any one of claims 1 to 9.

18. A non-transitory computer-readable storage medium storing instructions, characterized in that, The instructions, when executed by at least one processor, cause the at least one processor to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image encoder, image decoder, image encoding method, and image decoding method

    WO2019039324A1