Method, apparatus and readable storage medium for video decoding

By using a multi-reference row intra-frame prediction method, the problem of low direction prediction efficiency in intra-frame prediction is solved, thereby improving video coding efficiency, especially the coding performance in high-resolution and high-frame-rate videos.

CN115152208BActive Publication Date: 2026-02-10TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180015174.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-28
Filing Date
2021-06-29
Publication Date
2026-02-10
Estimated Expiration
2041-06-29

AI Technical Summary

Technical Problem

Existing video coding techniques have low efficiency in direction prediction during intra-frame prediction, resulting in low coding efficiency. This is especially true in high-resolution and high-frame-rate video compression, where more bits are needed to represent unlikely directions, increasing coding complexity and storage requirements.

Method used

An intra-frame prediction method with multiple reference lines is adopted. By determining the first subset of multiple reference lines, the current block is reconstructed based on the intra-frame prediction direction, reducing bit representation for unlikely directions and improving coding efficiency.

Benefits of technology

It improves the efficiency of video encoding, reduces the bit requirements for unlikely directions, lowers storage and bandwidth requirements, and is suitable for efficient compression of high-resolution and high-frame-rate video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115152208B_ABST
    Figure CN115152208B_ABST
Patent Text Reader

Abstract

Aspects of the disclosure include methods, devices, and readable storage media of video decoding. A device includes processing circuitry that decodes prediction information of a current block in a current picture, the current picture being part of a coded video sequence. The prediction information indicates one of a plurality of intra-prediction directions for the current block. The processing circuitry determines a first subset of a plurality of reference lines based on the one of the plurality of intra-prediction directions. The processing circuitry performs an intra-prediction of the current block based on the determined first subset of the plurality of reference lines. The processing circuitry reconstructs the current block based on the intra-prediction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] References merged

[0002] This application claims priority to U.S. Patent Application No. 17 / 360,803, filed June 28, 2021, entitled “Method and Apparatus for Video Coding,” and U.S. Provisional Application No. 63 / 082,806, filed September 24, 2020, entitled “Interpolation-Free Directional Intra Prediction.” The entire disclosure of these earlier applications is incorporated herein by reference. Technical Field

[0003] This application describes embodiments of video decoding in general. Background Technology

[0004] The background description provided herein is intended to present the overall context of this application. The extent of the work of the currently named inventors described in the background section and various aspects of this specification does not imply that it was prior art at the time of filing of this application, nor is it expressly or implied that it was acknowledged as prior art to this application.

[0005] Video encoding and decoding can be performed using inter-frame picture prediction with motion compensation. Uncompressed digital video can comprise a series of pictures, each with spatial dimensions, for example, 1920×1080 luminance samples and correlated chrominance samples. The series of pictures has a fixed or variable picture rate (also informally referred to as the frame rate), for example, 60 pictures per second or 60Hz. Uncompressed video has very high bitrate requirements. For example, a 1080p60 4:2:0 video with 8 bits per sample (1920x1080 luminance sample resolution, 60Hz frame rate) requires close to 1.5 Gbit / s of bandwidth. One hour of such video would require over 600 GB of storage space.

[0006] One goal of video encoding and decoding is to reduce redundant information in the input video signal through compression. Video compression can help reduce bandwidth or storage requirements by two or more orders of magnitude in some cases. Lossless and lossy compression, as well as combinations of both, can be used. Lossless compression refers to the technique of reconstructing an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original and reconstructed signals is small enough that the reconstructed signal can be used for the intended application. Lossy compression is widely used in video. The amount of distortion tolerated depends on the application. For example, users of some consumer streaming applications may tolerate higher distortion than users of television applications. The achievable compression ratio reflects that higher allowable / tolerable distortion results in a higher compression ratio.

[0007] Video encoders and decoders can utilize several major categories of techniques, including motion compensation, transform, quantization, and entropy coding.

[0008] Video codec techniques may include known intra-frame coding techniques. In intra-frame coding, sample values ​​are represented without reference to samples or other data from a previously reconstructed reference image. In some video codecs, an image is spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, the image can be an intra-frame image. Intra-frame images and their derivatives (e.g., independent decoder refresh images) can be used to reset the decoder state and are therefore used as the first image in the encoded video bitstream and video session, or as still images. Samples from an intra-frame block can be used for transform, and the transform coefficients can be quantized before entropy coding. Intra-frame prediction can be a technique that minimizes the sample values ​​in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are needed to represent the entropy-coded block for a given quantization step size.

[0009] As is known from technologies such as MPEG-2, traditional intra-frame coding does not use intra-frame prediction. However, some newer video compression techniques include those that attempt to derive data blocks from, for example, surrounding sample data and / or metadata, which are obtained during spatially adjacent encoding and / or decoding, and prior to the decoding sequence. This technique is later referred to as "intra-frame prediction." It is important to note that, at least in some cases, intra-frame prediction uses only reference data from the current frame being reconstructed, and not reference data from a reference frame.

[0010] There can be many different forms of intra-prediction. When more than one such technique can be used in a given video coding technique, the techniques used can be coded in intra-prediction modes. In some cases, a mode may have sub-modes and / or parameters, and these modes may be encoded individually or contained in mode codewords. Which codeword is used for a given mode, the combination of sub-modes and / or parameters will affect the coding efficiency gain through intra-prediction, and this also applies to entropy coding techniques used to convert codewords into bitstreams.

[0011] H.264 introduced an intra-frame prediction mode, which was improved in H.265 and further refined in newer coding techniques such as Joint Exploration Model (JEM), Universal Video Coding (VVC), and Baseline Set (BMS). Prediction blocks are formed by using neighboring sample values ​​belonging to already available samples. The sample values ​​of neighboring samples are copied into the prediction block in a specific direction. References to the direction used can be encoded in the bitstream or can be predicted themselves.

[0012] Referring to Figure 1A, the lower right corner depicts a subset of nine known prediction directions from the 33 possible prediction directions of H.265 (corresponding to 33 angular patterns out of 35 internal patterns). The point (101) where the arrows converge represents the sample being predicted. The arrow indicates the direction in which the sample is being predicted. For example, arrow (102) indicates that sample (101) is predicted based on one or more samples at a 45-degree angle to the horizontal direction from the upper right. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more samples at a 22.5-degree angle to the horizontal direction from the lower left.

[0013] Referring again to Figure 1A, a square block (104) comprising 4×4 samples is shown in the upper left (represented by a thick dashed line). The square block (104) consists of 16 samples, each labeled with "S" and its position in the Y dimension (e.g., row index) and X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from top to bottom) and the first sample in the X dimension (from left to right). Similarly, sample S44 is the fourth sample in block (104) in both the Y and X dimensions. Since the block is a 4×4 size sample, S44 is located in the lower right corner. Reference samples following a similar numbering scheme are also shown. Reference samples are labeled with "R" and their Y position (e.g., row index) and X position (e.g., column index) relative to block (104). In H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, so negative values ​​are not required.

[0014] Intra-frame image prediction can be performed by copying reference sample values ​​from adjacent samples occupied by the prediction direction indicated by the signal. For example, suppose the encoded video bitstream includes signaling that, for this block, the signaling indicates a prediction direction consistent with arrow (102), i.e., predicting samples based on one or more prediction samples at a 45-degree angle to the horizontal direction from the upper right. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Then, sample S44 is predicted based on reference sample R08.

[0015] In some cases, such as through interpolation, the values ​​of multiple reference samples can be combined to compute a reference sample, especially when the direction is not divisible by 45 degrees.

[0016] With the development of video coding technology, the number of possible directions has increased. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013) and JEM / VVC / BMS, and at the time of this application, up to 65 directions could be supported. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding have been used to represent those possible directions using a small number of bits, while accepting some cost for less likely directions. Furthermore, the direction itself can sometimes be predicted based on the adjacent directions used in adjacent, already decoded blocks.

[0017] Figure 1B is a schematic diagram (105) illustrating 65 intra-frame prediction directions based on JEM to show how the number of prediction directions increases over time.

[0018] The mapping of intra-predicted direction bits in a coded video bitstream can vary depending on the video coding technique, and can range from a simple, direct mapping of intra-predicted modes to predicted directions in codewords to complex adaptive schemes that incorporate the most probable modes and similar techniques. However, in all cases, there may be certain directions in the video content that are statistically less likely to occur than others. Since the purpose of video compression is to reduce redundancy, in well-functioning video coding techniques, less probable directions will be represented using a greater number of bits compared to the more probable directions.

[0019] Motion compensation can be a lossy compression technique and can involve using sample data blocks from a previously reconstructed image or a portion thereof (the reference image), spatially shifted in a direction indicated by a motion vector (MV), for prediction of the newly reconstructed image or a portion thereof. In some cases, the reference image can be the same as the image currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference image in use (indirectly, the latter could be a temporal dimension).

[0020] In some video compression techniques, a video feature (MV) applicable to a region of sample data can be predicted from other MVs, such as those MVs preceding the one being reconstructed, which are spatially adjacent to the region being reconstructed. This significantly reduces the amount of data required to encode the MV, eliminating redundancy and increasing compression. MV prediction can work effectively, for example, because when encoding an input video signal derived from a camera (called natural video), regions larger than the area applicable to a single MV are statistically likely to move in similar directions. Therefore, in some cases, these regions (regions larger than the area applicable to a single MV) can be predicted using similar MVs derived from neighboring regions. This results in MVs found for a given region being similar to or identical to MVs predicted from surrounding MVs, and, conversely, after entropy encoding, the MV can be represented with fewer bits than when directly encoding the MV. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself may be lossy, for example due to rounding errors when calculating predictions from several surrounding MVs.

[0021] H.265 / HEVC (ITU-T H.265 Recommendation, “High Efficiency Video Coding”, December 2016) describes various MV prediction mechanisms. Among the various MV prediction mechanisms provided by H.265, this application describes the technique hereinafter referred to as “spatial combining”.

[0022] Referring to Figure 1C, the current block (111) includes samples discovered by the encoder during the motion search process, which can be predicted based on previous blocks of the same size that have generated spatial offsets. Alternatively, the MV can be derived from metadata associated with one or more reference images, rather than being directly encoded. For example, using the MV associated with any of the five surrounding samples A0, A1 and B0, B1, B2 (corresponding to 112 to 116 respectively), the MV is derived from the metadata of the nearest reference image (in decoding order). In H.265, MV prediction can use predictions from the same reference image that is also being used in adjacent blocks. Summary of the Invention

[0023] Various aspects of this disclosure provide video encoding / decoding apparatus. One apparatus includes processing circuitry that decodes prediction information for a current block in a current image, the current image being part of an encoded video sequence. The prediction information indicates one of a plurality of intra-prediction directions for the current block. Based on the one intra-prediction direction, the processing circuitry determines a first subset of a plurality of reference rows. Based on the determined first subset of the plurality of reference rows, the processing circuitry performs intra-prediction on the current block; based on the intra-prediction, the processing circuitry reconstructs the current block.

[0024] In one embodiment, the number of reference rows in the first subset of the determined plurality of reference rows is greater than 1.

[0025] In one embodiment, among the plurality of intra-prediction directions, the intra-prediction direction associated with a first reference line among the plurality of reference lines is different from the intra-prediction direction associated with a second reference line among the plurality of reference lines.

[0026] In one embodiment, the plurality of intra-prediction directions are associated with a first reference line among the plurality of reference lines, or a second subset of the plurality of intra-prediction directions is associated with the first reference line among the plurality of reference lines, and a third subset of the plurality of intra-prediction directions is associated with the second reference line among the plurality of reference lines.

[0027] In one embodiment, the processing circuitry determines a reference row in the first subset for a corresponding sample based on the intra-frame prediction direction and the position of each sample in the current block.

[0028] In one embodiment, the prediction information includes syntax elements that indicate whether to perform the intra-frame prediction on the current block based on the plurality of reference lines.

[0029] In one embodiment, the current block is not located near the top boundary of the coding tree unit that includes the current block.

[0030] In one embodiment, one of the tangent and cotangent values ​​of the prediction angle associated with the intra-frame prediction direction is an integer.

[0031] In one embodiment, the processing circuitry determines a reference row index for a reference row in the first subset based on the tangent of the prediction angle associated with the intra-frame prediction direction and the row number of each row of the samples in the current block.

[0032] Various aspects of this disclosure provide methods for video encoding / decoding. In this method, prediction information for a current block in a current image, which is part of an encoded video sequence, is decoded. The prediction information indicates one of a plurality of intra-prediction directions for the current block. Based on the one intra-prediction direction, a first subset of a plurality of reference rows is determined. Based on the determined first subset of the plurality of reference rows, intra-prediction of the current block is performed. Based on the intra-prediction, the current block is reconstructed.

[0033] Various aspects of this disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform any one or more combinations of video decoding methods. Attached Figure Description

[0034] Other features, properties, and various advantages of the disclosed subject matter will become further apparent from the following detailed description and accompanying drawings, wherein:

[0035] Figure 1A is a schematic diagram of an example subset of intra-frame prediction modes;

[0036] Figure 1B is an illustration of an exemplary intra-frame prediction direction;

[0037] Figure 1C is a schematic diagram of the current block and its surrounding space merging candidates in an example;

[0038] Figure 2 This is a simplified block diagram of a communication system according to an embodiment;

[0039] Figure 3 This is a simplified block diagram of a communication system according to an embodiment;

[0040] Figure 4 This is a simplified block diagram of the decoder according to an embodiment;

[0041] Figure 5 This is a simplified block diagram of the encoder according to an embodiment;

[0042] Figure 6 A block diagram of an encoder according to an embodiment is shown;

[0043] Figure 7 A block diagram of a decoder according to an embodiment is shown;

[0044] Figure 8 Exemplary block partitions according to some embodiments of this disclosure are shown;

[0045] Figure 9 Exemplary block partitions according to some embodiments of this disclosure are shown;

[0046] Figure 10 Exemplary block partitions according to some embodiments of this disclosure are shown;

[0047] Figure 11 An exemplary quadtree according to an embodiment of the present disclosure is shown, the quadtree having a nested multi-type tree coding block structure;

[0048] Figure 12 Exemplary nominal angles according to embodiments of this disclosure are shown;

[0049] Figure 13 The positions of the top sample, left sample, and top-left sample of a pixel in the current block according to an embodiment of the present disclosure are shown;

[0050] Figure 14 An exemplary bilinear interpolation for deriving predicted samples at fractional positions is shown according to an embodiment of this disclosure;

[0051] Figure 15 An exemplary multi-line intra-frame prediction according to an embodiment of the present disclosure is shown, which uses four reference lines adjacent to a coding block unit;

[0052] Figure 16 An exemplary angle for intra-frame prediction direction according to an embodiment of this disclosure is shown;

[0053] Figure 17 Exemplary prediction angles according to some embodiments of this disclosure are shown;

[0054] Figure 18 An exemplary intra-frame prediction using two reference rows is illustrated according to an embodiment of this disclosure;

[0055] Figure 19 An exemplary intra-frame prediction using three reference rows is illustrated according to an embodiment of this disclosure;

[0056] Figure 20 An exemplary flowchart according to an embodiment of this disclosure is shown; and

[0057] Figure 21 This is a schematic diagram of a computer system according to an embodiment. Detailed Implementation

[0058] I. Video Decoder and Encoder Systems

[0059] Figure 2 This is a simplified block diagram of a communication system (200) according to an embodiment disclosed in this application. The communication system (200) includes a plurality of terminal devices that can communicate with each other via, for example, a network (250). For example, the communication system (200) includes a first terminal device (210) and a second terminal device (220) interconnected via a network (250). Figure 2 In this embodiment, the first terminal device (210) and the second terminal device (220) perform unidirectional data transmission. For example, the first terminal device (210) may encode video data (e.g., a video image stream captured by the first terminal device (210)) for transmission over a network (250) to the second terminal device (220). The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device (220) may receive the encoded video data from the network (250), decode the encoded video data to recover the video data, and display video images based on the recovered video data. Unidirectional data transmission is common in applications such as media services.

[0060] In another embodiment, the communication system (200) includes a third terminal device (230) and a fourth terminal device (240) that perform bidirectional transmission of encoded video data, which may occur, for example, during a video conference. For bidirectional data transmission, each of the third terminal device (230) and the fourth terminal device (240) may encode video data (e.g., a stream of video images captured by the terminal device) for transmission over a network (250) to the other terminal device. Each of the third terminal device (230) and the fourth terminal device (240) may also receive encoded video data transmitted by the other terminal device and may decode the encoded video data to recover the video data, and may display the video images on an accessible display device based on the recovered video data.

[0061] exist Figure 2 In the embodiments disclosed herein, the first terminal device (210), the second terminal device (220), the third terminal device (230), and the fourth terminal device (240) may be servers, personal computers, and smartphones, but the principles disclosed herein are not limited thereto. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (250) refers to any number of networks that transmit encoded video data between the first terminal device (210), the second terminal device (220), the third terminal device (230), and the fourth terminal device (240), including, for example, wired (connected) and / or wireless communication networks. The communication network (250) may exchange data in circuit-switched and / or packet-switched channels. The network may include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of the network (250) may be irrelevant to the operation of this application.

[0062] As an example, Figure 3 The diagram illustrates the placement of a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0063] The streaming system may include an acquisition subsystem (313) that may include a video source (301) such as a digital camera, which creates an uncompressed video image stream (302). In an embodiment, the video image stream (302) includes samples captured by a digital camera. The video image stream (302) is depicted as a thick line to emphasize the high data volume of the video image stream compared to encoded video data (304) (or encoded video bitstream). The video image stream (302) may be processed by an electronic device (320) that includes a video encoder (303) coupled to the video source (301). The video encoder (303) may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter as described in more detail below. Compared to the video image stream (302), the encoded video data (304) (or the encoded video bitstream (304)) is depicted as a thin line to emphasize the lower data volume of the encoded video data (304) (or the encoded video bitstream (304)), which can be stored on a streaming server (305) for future use. One or more streaming client subsystems, such as Figure 3 Client subsystems (306) and (308) can access a streaming server (305) to retrieve copies (307) and (309) of encoded video data (304). Client subsystem (306) may include, for example, a video decoder (310) in an electronic device (330). The video decoder (310) decodes the incoming copy (307) of the encoded video data and produces an output video picture stream (311) that can be displayed on a display (312) (e.g., a screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (304), video data (307), and video data (309) (e.g., video streams) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In embodiments, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application can be used in the context of the VVC standard.

[0064] It should be noted that the electronic devices (320) and (330) may include other components (not shown). For example, the electronic device (320) may include a video decoder (not shown), and the electronic device (330) may also include a video encoder (not shown).

[0065] Figure 4 This is a block diagram of a video decoder (410) according to an embodiment disclosed in this application. The video decoder (410) may be disposed in an electronic device (430). The electronic device (430) may include a receiver (431) (e.g., receiving circuitry). The video decoder (410) may be used in place of... Figure 3 The video decoder (510) in the embodiment.

[0066] The receiver (431) may receive one or more encoded video sequences to be decoded by the video decoder (410); in the same embodiment or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequences may be received from a channel (401), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (431) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not indicated). The receiver (431) may separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (415) may be coupled between the receiver (431) and the entropy decoder / parser (420) (hereinafter referred to as "parser (420)"). In some applications, the buffer memory (415) is part of the video decoder (410). In other cases, the buffer memory (415) may be located external to the video decoder (410) (not indicated). In other cases, an external buffer (not shown) may be provided for the video decoder (410) to prevent network jitter, for example, and another buffer (415) may be configured internally for, for example, handling broadcast timing. When the receiver (431) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer (415) may not be necessary, or it may be made smaller. Of course, for use on packet networks such as the Internet, a buffer (415) may be required; this buffer may be relatively large and adaptive in size, and may be at least partially implemented in the operating system or a similar component (not shown) external to the video decoder (410).

[0067] The video decoder (410) may include a parser (420) to reconstruct symbols (421) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (410) and potential information for controlling a display device (412) (e.g., a display screen), which is not part of the electronic device (430) but may be coupled to it, such as... Figure 4 As shown in the figure. The control information for the display device may be a parameter set fragment (not shown) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (420) may parse / decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (420) may extract a subgroup parameter set of at least one subgroup of pixels in the subgroup of pixels in the encoded video sequence for use in the video decoder based on at least one parameter corresponding to a group. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (420) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, MV, etc.

[0068] The parser (420) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (415) to create symbols (421).

[0069] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (421) may involve multiple different units. Which units are involved and how they are involved can be controlled by the subgroup control information parsed by the parser (420) from the encoded video sequence. For brevity, the flow of such subgroup control information between the parser (420) and the various units described below is not described.

[0070] In addition to the functional blocks already mentioned, the video decoder (410) can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0071] The first unit is the scaler / inverse transform unit (451). The scaler / inverse transform unit (451) receives quantization transform coefficients as symbols (421) and control information from the parser (420), including the transform method used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (451) can output a block containing sample values, which can be input into the aggregator (455).

[0072] In some cases, the output samples of the scaler / inverse transform unit (451) may belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images but can use predictive information from previously reconstructed portions of the current image. Such predictive information may be provided by the intra-picture prediction unit (452). In some cases, the intra-picture prediction unit (452) uses reconstructed information extracted from the current picture buffer (458) to generate surrounding blocks of the same size and shape as the block being reconstructed. For example, the current picture buffer (458) buffers partially reconstructed and / or fully reconstructed current images. In some cases, the aggregator (455) adds the predictive information generated by the intra-picture prediction unit (452) to the output sample information provided by the scaler / inverse transform unit (451) based on each sample.

[0073] In other cases, the output samples of the scaler / inverse transform unit (451) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (453) can access the reference image memory (457) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (421), these samples can be added by the aggregator (455) to the output of the scaler / inverse transform unit (451) (referred to as residual samples or residual signals in this case) to generate output sample information. The motion compensation prediction unit (453) can obtain the prediction samples from the address in the reference image memory (457) under motion vector control, and the motion vector is available to the motion compensation prediction unit (453) in the form of the symbols (421), which, for example, include X, Y and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory (457) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0074] The output samples of the aggregator (455) can be employed by various loop filtering techniques in the loop filter unit (454). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video stream), and these parameters can be used as symbols (421) from the parser (420) in the loop filter unit (456). However, in other embodiments, the video compression techniques may also respond to metadata obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0075] The output of the loop filter unit (456) can be a sample stream, which can be output to a display device (412) and stored in a reference image memory (457) for subsequent inter-frame image prediction.

[0076] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded image corresponding to the current image has been fully reconstructed and the encoded image (by, for example, the parser (420)) is identified as the reference image, the current image buffer (458) can become part of the reference image memory (457), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.

[0077] The video decoder (410) can perform decoding operations according to a predetermined video compression technique, such as that specified in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technique or standard as the only tools available under said configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference picture size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.

[0078] In this embodiment, the receiver (431) may receive additional (redundant) data along with the encoded video. The additional data may be a portion of the encoded video sequence. The additional data may be used by the video decoder (410) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0079] Figure 5 This is a block diagram of a video encoder (503) according to an embodiment disclosed in this application. The video encoder (503) is disposed in an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmission circuit). The video encoder (503) can be used to replace... Figure 3 The video encoder (303) in the embodiment.

[0080] The video encoder (503) can obtain data from the video source (501) (not) Figure 5 In one embodiment, a portion of the electronic device (520) receives video samples, the video source being capable of capturing video images to be encoded by a video encoder (503). In another embodiment, the video source (501) is a portion of the electronic device (520).

[0081] A video source (501) can provide a sequence of source video samples encoded by a video encoder (503) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (501) can be a storage device storing previously prepared video. In a video conferencing system, the video source (501) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc., used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0082] According to an embodiment, the video encoder (503) can encode and compress images of a source video sequence into an encoded video sequence (543) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of the controller (550). In some embodiments, the controller (550) controls and is functionally coupled to other functional units described below. For simplicity, coupling is not shown in the figures. Parameters set by the controller (550) may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum allowed motion vector reference area, etc. The controller (550) may be used with other suitable functions related to the video encoder (503) optimized for a particular system design.

[0083] In some embodiments, the video encoder (503) operates within an encoding loop. As a simplified description, in an embodiment, the encoding loop may include a source encoder (530) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (533) embedded within the video encoder (503). The decoder (533) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (because in the video compression techniques considered in this application, any compression between the symbols and the encoded video stream is lossless). The reconstructed sample stream (sample data) is input to a reference image memory (534). Since decoding of the symbol stream produces bit-precise results independent of the decoder's location (local or remote), the contents of the reference image memory (534) are also bit-precisely corresponding between the local encoder and the remote encoder. In other words, the reference image samples "seen" by the encoder's prediction portion are exactly the same sample values ​​that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.

[0084] The operation of the “local” decoder (533) can be combined with, for example, the above-described method. Figure 4 The video decoder (410) is described in detail as the same as the "remote" decoder. However, a further brief reference is provided. Figure 4 When symbols are available and the entropy encoder (545) and parser (420) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (410), including the buffer (415) and parser (420), may not be fully implemented in the local decoder (533).

[0085] It can be observed that any decoder technique other than parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in essentially the same functional form. For this reason, this application focuses on decoder operation. The description of encoder techniques can be simplified because encoder techniques are inverses of the fully described decoder techniques. More detailed descriptions are only required in certain areas, and are provided below.

[0086] During operation, in some embodiments, the source encoder (530) may perform motion-compensated predictive coding. The motion-compensated predictive coding predictively encodes the input image, referencing one or more previously encoded images from the video sequence designated as "reference images." In this manner, the encoding engine (532) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.

[0087] The local video decoder (533) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (530). The operation of the encoding engine (532) can be a lossy process. When the encoded video data can be decoded by the video decoder (533), Figure 5 When the source video sequence (not shown) is decoded, the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (533) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in a reference image cache (534). In this way, the video encoder (503) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0088] The predictor (535) can perform a prediction search against the encoding engine (532). That is, for a new image to be encoded, the predictor (535) can search in the reference image memory (534) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (535) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, based on the search results obtained by the predictor (535), it can be determined that the input image may have prediction references obtained from multiple reference images stored in the reference image memory (534).

[0089] The controller (550) can manage the encoding operations of the source encoder (530), including, for example, setting parameters and subgroup parameters for encoding video data.

[0090] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (545). The entropy encoder (545) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding, thereby converting the symbols into an encoded video sequence.

[0091] The transmitter (540) can buffer the encoded video sequence created by the entropy encoder (545) in preparation for transmission via a communication channel (560), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (540) can combine the encoded video data from the video encoder (503) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0092] The controller (550) manages the operation of the video encoder (503). During encoding, the controller (550) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types:

[0093] An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will understand variations of I-pictures and their corresponding applications and characteristics.

[0094] A predictive image (P-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and a reference index to predict sample values ​​for each block.

[0095] A bidirectional predictive image (B-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values ​​for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.

[0096] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, determined based on the coding assignments of the corresponding images applied to the blocks. For example, blocks of an I-image can be non-predictively coded, or the blocks can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. Blocks of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.

[0097] The video encoder (503) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (503) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0098] In this embodiment, the transmitter (540) may transmit additional data while transmitting encoded video. The source encoder (530) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0099] The acquired video can serve as multiple source images (video images) presented in a time series. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In an embodiment, a specific image being encoded / decoded is segmented into blocks, referred to as the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when multiple reference images are used, the motion vector may have a third dimension that identifies the reference image.

[0100] In some embodiments, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. Specifically, the block can be predicted using a combination of the first and second reference blocks.

[0101] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0102] According to some embodiments disclosed in this application, predictions such as inter-frame image prediction and intra-frame image prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a video image sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the images have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Furthermore, each CTU can be further subdivided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be subdivided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In embodiments, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. Furthermore, depending on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In embodiments, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Taking a luma prediction block as an example, a prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0103] Figure 6 This is a diagram of a video encoder (603) according to another embodiment disclosed in this application. The video encoder (603) is used to receive processing blocks (e.g., prediction blocks) of sample values ​​within a current video image in a video image sequence, and to encode the processing blocks into an encoded image that is part of an encoded video sequence. In this embodiment, the video encoder (603) is used instead of Figure 3The video encoder (303) in the embodiment.

[0104] In the HEVC embodiment, the video encoder (603) receives a matrix of sample values ​​for a processing block, such as an 8×8 sample prediction block. The video encoder (603) uses, for example, rate-distortion optimization to determine whether to use intra-frame mode, inter-frame mode, or bidirectional prediction mode to encode the processing block. When encoding the processing block in intra-frame mode, the video encoder (603) can use intra-frame prediction techniques to encode the processing block into an encoded picture; and when encoding the processing block in inter-frame mode or bidirectional prediction mode, the video encoder (603) can use inter-frame prediction or bidirectional prediction techniques to encode the processing block into an encoded picture, respectively. In some video coding techniques, the merging mode can be an inter-frame picture prediction sub-mode, in which motion vectors are derived from one or more motion vector prediction values ​​without relying on encoded motion vector components outside the prediction values. In some other video coding techniques, motion vector components applicable to the subject block may exist. In the embodiment, the video encoder (603) includes other components, such as a mode decision module (not shown) for determining the processing block mode.

[0105] exist Figure 6 In one embodiment, the video encoder (603) includes, as shown below: Figure 6 The inter-frame encoder (630), intra-frame encoder (622), residual calculator (623), switch (626), residual encoder (624), general controller (621) and entropy encoder (625) are shown coupled together.

[0106] An inter-frame encoder (630) is configured to receive samples of the current block (e.g., the processing block), compare the block with one or more reference blocks in a reference image (e.g., blocks in previous and later images), generate inter-frame prediction information (e.g., redundancy information description, motion vectors, merging mode information based on inter-frame coding techniques), and calculate inter-frame prediction results (e.g., prediction blocks) based on the inter-frame prediction information using any suitable technique. In some embodiments, the reference image is a decoded reference image based on encoded video information.

[0107] The intra encoder (622) is used to receive samples of the current block (e.g., the processing block), in some cases compare the block with previously encoded blocks in the same image, generate quantization coefficients after transformation, and in some cases also (e.g., based on intra prediction direction information of one or more intra coding techniques) generate intra prediction information. In an embodiment, the intra encoder (622) also calculates intra prediction results (e.g., prediction blocks) based on the intra prediction information and reference blocks in the same image.

[0108] A general-purpose controller (621) determines general-purpose control data and controls other components of the video encoder (603) based on the general-purpose control data. In an embodiment, the general-purpose controller (621) determines the mode of a block and provides control signals to a switch (626) based on the mode. For example, when the mode is an intra-frame mode, the general-purpose controller (621) controls the switch (626) to select an intra-frame mode result for use by the residual calculator (623) and controls the entropy encoder (625) to select intra-frame prediction information and add the intra-frame prediction information to the bitstream; and when the mode is an inter-frame mode, the general-purpose controller (621) controls the switch (626) to select an inter-frame prediction result for use by the residual calculator (623) and controls the entropy encoder (625) to select inter-frame prediction information and add the inter-frame prediction information to the bitstream.

[0109] A residual calculator (623) is used to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (622) or the inter encoder (630). A residual encoder (624) is used to operate on the residual data to encode the residual data to generate transform coefficients. In an embodiment, the residual encoder (624) is used to transform the residual data from the time domain to the frequency domain and generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients. In various embodiments, the video encoder (603) also includes a residual decoder (628). The residual decoder (628) is used to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra encoder (622) and the inter encoder (630). For example, the inter encoder (630) can generate a decoded block based on the decoded residual data and inter-frame prediction information, and the intra encoder (622) can generate a decoded block based on the decoded residual data and intra-frame prediction information. The decoded blocks are processed appropriately to generate a decoded image, and in some embodiments, the decoded image may be buffered in a memory circuit (not shown) and used as a reference image.

[0110] An entropy encoder (625) is used to format the bitstream to produce encoded blocks. The entropy encoder (625) generates various information according to a suitable standard such as HEVC. In an embodiment, the entropy encoder (625) is used to obtain general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. It should be noted that, according to the disclosed subject matter, residual information is not present when blocks are encoded in a merged sub-mode of inter-frame mode or bidirectional prediction mode.

[0111] Figure 7This is a diagram of a video decoder (710) according to another embodiment disclosed in this application. The video decoder (710) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In an embodiment, the video decoder (710) is used instead of Figure 3 The video decoder (310) in the embodiment.

[0112] exist Figure 7 In this embodiment, the video decoder (710) includes, as follows: Figure 7 The entropy decoder (771), inter-frame decoder (780), residual decoder (773), reconstruction module (774), and intra-frame decoder (772) are shown coupled together.

[0113] An entropy decoder (771) can be used to reconstruct certain symbols from an encoded image, these symbols representing the syntax elements constituting the encoded image. Such symbols may include, for example, a mode for encoding the block (e.g., intra-frame mode, inter-frame mode, bidirectional prediction mode, a merged sub-mode of the latter two, or another sub-mode), prediction information (e.g., intra-frame prediction information or inter-frame prediction information) that can respectively identify certain samples or metadata used by the intra-frame decoder (772) or the inter-frame decoder (780) for prediction, residual information in the form of, for example, quantized transform coefficients, and so on. In an embodiment, when the prediction mode is inter-frame or bidirectional prediction mode, inter-frame prediction information is provided to the inter-frame decoder (780); and when the prediction type is intra-frame prediction type, intra-frame prediction information is provided to the intra-frame decoder (772). Residual information may be provided to the residual decoder (773) via inverse quantization.

[0114] The inter-frame decoder (780) is used to receive inter-frame prediction information and generate inter-frame prediction results based on the inter-frame prediction information.

[0115] The intra-frame decoder (772) is used to receive intra-frame prediction information and generate prediction results based on the intra-frame prediction information.

[0116] The residual decoder (773) performs inverse quantization to extract the dequantized transform coefficients and processes the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require some control information (to obtain the quantizer parameters QP), and this information may be provided by the entropy decoder (771) (the data path is not indicated because this is only low-level control information).

[0117] The reconstruction module (774) is used to combine the residual output by the residual decoder (773) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, which may be a part of a reconstructed image, which in turn may be a part of a reconstructed video. It should be noted that other suitable operations, such as deblocking, may be performed to improve visual quality.

[0118] It should be noted that any suitable technology can be used to implement the video encoder (303), video encoder (503), and video encoder (603), as well as the video decoder (310), video decoder (410), and video decoder (710). In one embodiment, one or more integrated circuits can be used to implement the video encoder (303), video encoder (503), and video encoder (603), as well as the video decoder (310), video decoder (410), and video decoder (710). In another embodiment, one or more processors executing software instructions can be used to implement the video encoder (303), video encoder (503), and video encoder (603), as well as the video decoder (310), video decoder (410), and video decoder (710).

[0119] II. Block Partitioning

[0120] Figure 8 Exemplary block partitioning according to some embodiments of the present disclosure is illustrated. In the embodiments, Figure 8 The exemplary block partitioning in the example can be used for VP9 proposed by the Open Media Consortium (AOMedia). For example... Figure 8 As shown, four partitioning trees can be used, ranging from a 64×64 level down to a 4×4 level, with some additional restrictions for 8×8 blocks. It should be noted that a partition designated as R can be called a recursive partition. That is, the same partitioning tree is repeated at lower scales until the lowest 4×4 level is reached.

[0121] Figure 9 Exemplary block partitioning according to some embodiments of the present disclosure is illustrated. In the embodiments, Figure 9 The exemplary block partitioning in the example can be used for AV1 proposed by AOMedia. For example... Figure 9 As shown, the partitioning tree can be extended to 10 structures, and the maximum coded block size (referred to as the superblock in VP9 / AV1 terminology) is increased to start from 128×128. It should be noted that... Figure 9 The 4:1 / 1:4 rectangular partition in the first row does not exist in VP9. Figure 9The partition type with 3 sub-partitions in the second row is called a T-partition. No rectangular partition can be further subdivided. In addition to the coded block size, the coded tree depth can be defined to indicate the partition depth from the root node. In an embodiment, the coded tree depth of the root node, for example, 128×128, can be set to 0. The coded tree depth increases by 1 after each further partition of the coded block.

[0122] Instead of forcing the use of the fixed transform unit size in VP9, ​​AV1 allows luma coding blocks to be partitioned into transform units of multiple sizes, which can be represented by recursive partitions down to level 2. To merge extended coding block partitions in AV1, transform sizes from 4×4 to 64×64 are supported, including square, 2:1 / 1:2, and 4:1 / 1:4. For chroma blocks, only the largest possible transform unit is allowed.

[0123] In some relevant examples such as HEVC, the CTU is partitioned into CUs using a quadtree structure represented as a coding tree to accommodate various local characteristics. A decision can be made at the CU level regarding whether to use inter-frame (temporal) or intra-frame (spatial) prediction to encode a picture region. Each CU can be further partitioned into one, two, or four PUs based on the PU partitioning type. Within a PU, the same prediction process can be applied, and relevant information can be sent to the decoder based on the PU. After obtaining residual blocks by applying a prediction process based on the PU partitioning type, the CU can be partitioned into TUs according to another quadtree structure (such as the coding tree of the CU). One of the key features of the HEVC structure is that it has multiple partitioning concepts including CUs, PUs, and TUs. In HEVC, CUs or TUs are simply square shapes, while PUs can be square or rectangular shapes used for inter-frame prediction blocks. In HEVC, a coding block can be further partitioned into four square sub-blocks, and a transformation process is performed on each sub-block (i.e., TU). Each TU can be further recursively partitioned (e.g., using quadtree partitioning) into smaller TUs. Quadtree partitioning can be called residual quadtree (RQT).

[0124] At the image boundaries, HEVC uses implicit quadtree segmentation, allowing the block to continue quadtree segmentation until the block size fits the image boundaries.

[0125] In some relevant examples such as VVC, a quadtree with nested multi-type trees using binary and ternary segmentation structures can replace the concept of multiple partitioning unit types. That is, the separation of CU, PU, ​​and TU concepts is removed, unless the size of the CU is too large relative to the maximum transform length. Therefore, more flexibility in the shape of the CU partition can be supported in these examples. In the coding tree structure of VVC, the CU can have a square or rectangular shape. The CTU can first be partitioned using a quadtree (or quadtree) structure. Then, the leaf nodes of the quadtree can be further partitioned using a multi-type tree structure.

[0126] Figure 10 Exemplary block partitioning for multi-type tree segmentation patterns is illustrated according to some embodiments of the present disclosure. In the embodiments, Figure 10 The exemplary block partitioning in the example can be used in VVC. For example... Figure 10 As shown, there are four partitioning types in the multi-type tree structure: vertical binary partition (SPLIT_BT_VER), horizontal binary partition (SPLIT_BT_HOR), vertical ternary partition (SPLIT_TT_VER), and horizontal ternary partition (SPLIT_TT_HOR). The leaf nodes of the multi-type tree are called CUs. Unless the CU is too large for the maximum transform length, the multi-type tree structure is used for both the prediction and transform processes without any further partitioning. This means that in most cases, the CU, PU, ​​and TU have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is less than the width or height of the color component of the CU.

[0127] Figure 11 An exemplary quadtree with a nested multi-type tree coding block structure is shown according to an embodiment of the present disclosure.

[0128] In some relevant examples such as VVC, the maximum supported luminance transformation size is 64×64, and the maximum supported chrominance transformation size is 32×32. When the width or height of the CB is greater than the maximum transformation width or height, the CB can be automatically divided along the horizontal and / or vertical directions to meet the transformation size limit in that direction.

[0129] In some relevant examples such as VTM7, the coding tree scheme can support separate block tree structures for the luma and chroma CTBs within a CTU. For example, for P and B slices, the luma and chroma CTBs within a CTU share the same coding tree structure. However, for I slices, the luma and chroma CTBs within a CTU can have separate block tree structures. When the separate block tree mode is applied, the luma CTB is partitioned into CUs using one coding tree structure, and the chroma CTB is partitioned into chroma CUs using another coding tree structure. This means that a CU in an I slice can include either a coding block for the luma component or coding blocks for both chroma components, and a CU in a P or B slice always includes coding blocks for all three color components, unless the video is monochrome.

[0130] III. Intra-frame prediction

[0131] In some related examples such as VP9, ​​eight orientation modes are supported, corresponding to angles from 45 degrees to 207 degrees. To utilize more spatial redundancy in orientation textures, in some related examples such as AV1, the orientation intra-frame modes are extended to a more fine-grained set of angles. The original eight angles are slightly modified, and these original eight angles are called nominal angles, and these eight nominal angles are named V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED.

[0132] Figure 12 An exemplary nominal angle according to an embodiment of this disclosure is shown. Each nominal angle can be associated with seven finer angles, thus a total of 56 orientation angles can exist, as in AV1. The predicted angle is represented by a nominal intra-angle plus an angle increment, which is a step size of -3 to 3 multiplied by 3 degrees. To implement the orientation prediction mode in AV1 in a general manner, all 56 orientation intra-frame predicted angles in AV1 can be implemented using a unified orientation predictor that projects each pixel to a reference sub-pixel location and interpolates the reference sub-pixel through a 2-tap bilinear filter.

[0133] In some relevant examples such as AV1, there are five non-directional smooth intra-prediction modes: DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. For DC prediction, the average of the left and top neighboring samples is used as the predictor for the block to be predicted. For PAETH prediction, the top, left, and top-left reference samples are retrieved first, and then the closest value (top + left - top-left) is set as the predictor for the pixel to be predicted.

[0134] Figure 13 The positions of the top, left, and top-left samples of a pixel in the current block according to an embodiment of this disclosure are shown. For SMOOTH, SMOOTH_V, and SMOOTH_H modes, quadratic interpolation is used to predict the block in the vertical, horizontal, or average of these two directions.

[0135] For the chroma component, in addition to the 56 directional modes and 5 non-directional modes, there is also a chroma-only intra-frame prediction mode, which can be called the Chroma from Luma (CfL) mode. The chroma-only intra-frame prediction mode models chroma pixels as a linear function of overlapping reconstructed luma pixels. CfL prediction can be represented as follows:

[0136] CfL(α)=α×L AC +DC Equation (1)

[0137] Where L AC The AC contribution of the luminance component is represented by α, the parameter of the linear model is represented by α, and the DC contribution of the chrominance component is represented by DC. In the example, the reconstructed luminance pixels are subsampled to chrominance resolution, and then the average value is subtracted to form the AC contribution. To approximate the chrominance AC component from the AC contribution, the decoder does not need to compute scaling parameters as in some related examples. In AV1, the CfL mode determines the parameter α based on the original chrominance pixels and represents the parameter α as a signal in the bitstream. This reduces decoder complexity and produces more accurate predictions. For the DC contribution of the chrominance component, an intra-frame DC mode is used to compute the DC contribution of the chrominance component, which is sufficient for most chrominance content and has a well-established and fast implementation.

[0138] To represent chroma intra-prediction modes using signals, we can first represent the eight nominal directional modes, five non-directional modes, and the CfL mode using signals. The context used to represent these modes using signals can depend on the corresponding luma mode at the top-left position of the current block. Then, if the current chroma mode is a directional mode, an additional flag can be represented using signals to indicate the incremental angle relative to the nominal angle.

[0139] Screen content video encoding and decoding is becoming increasingly important in various applications such as desktop sharing, video conferencing, and distance education. Typically, screen content exhibits different characteristics compared to content captured by natural cameras, such as sharp edges. For conventional directional intra-prediction modes (such as the directional intra-prediction mode introduced above), interpolation operations (e.g., 2-tap bilinear interpolation, 4-tap cubic interpolation) are required to generate predicted sample values ​​at fractional sample locations. Interpolation operations inevitably smooth out sharp edges and produce high frequencies in the expensive residual blocks used for encoding.

[0140] To preserve sharp edges in intra-frame prediction, instead of applying interpolation, Nearest-Neighbor (NN) interpolation can be used. Two alternatives are described below. In the first alternative, a pixel-based implicit method is used, where both the encoder and decoder determine whether to perform NN interpolation based on the predicted pixels. In the second alternative, the encoder performs a rate-distortion search at the block level and explicitly signals a flag to the decoder indicating when to use NN interpolation.

[0141] Figure 14 An exemplary bilinear interpolation for deriving predicted samples at score locations is illustrated according to embodiments of this disclosure. Instead of using a weighted sum of multiple reference samples, the NN interpolation essentially selects one of the reference samples along the prediction direction. For example, in Figure 14 In this context, a bilinear interpolation filter is used to derive a predictor for sample C using two reference samples A and B. With bilinear interpolation, the predicted sample value is calculated as (A*b+B*a) / (a+b). With NN interpolation, the predicted sample value is derived as (a>b)?B:A.

[0142] Figure 15 An exemplary multi-line intra-prediction using four reference lines adjacent to a coding block unit is illustrated according to an embodiment of this disclosure. For multi-line intra-prediction, the encoder determines and signals which reference line is used to generate the intra-predictor. The reference line index is signaled before the intra-prediction mode, and only the most probable mode is allowed when a non-zero reference line index is signaled. Figure 15 The example depicts four reference rows, each consisting of six segments (segments A through F) and a top-left reference sample. Segments A and F are filled with the nearest samples from segments B and E, respectively.

[0143] IV. Interpolation-free directional intra-frame prediction

[0144] In some relevant examples such as AV1, there are multiple incremental angles (e.g., 7) for each directional nominal mode, and it is not optimal to represent and resolve all incremental angles with signals regardless of the orientation of adjacent nominal modes.

[0145] This disclosure includes a method for interpolation-free directional intra-frame prediction.

[0146] In this disclosure, when one directional intra prediction mode is close to another directional intra prediction mode, it means that the absolute difference in the prediction angles between the two modes is within a given threshold T. In one example, T is set to 1 or 2.

[0147] Figure 16An exemplary angle for intra-frame prediction direction according to an embodiment of this disclosure is shown. Figure 16 In the diagram, α is the prediction angle, and the solid arrow indicates the prediction direction. The tangent of the prediction angle is tan(α) = y / x.

[0148] According to various aspects of this disclosure, for each sample of the current block to be predicted, given one of a plurality of intra-frame prediction directions, a sample from one of a plurality of reference rows is selected as the prediction sample, and the selected prediction sample is located at an integer sample position in one of the plurality of reference rows.

[0149] In one embodiment, the number of reference lines is less than a threshold. The threshold is, for example, N+1, where N is a positive integer. For example, up to N reference lines are used for intra-frame prediction of the current block. Example values ​​for N include, but are not limited to, 2, 3, 4, 5, 6, 7, and 8.

[0150] In one embodiment, the tangent value of the prediction angle associated with multiple intra-frame prediction directions includes ±N or ±1 / N, where N is an integer, and example values ​​of N are 1, 2, 3, 4, 5, 6, 7, and 8.

[0151] According to some embodiments of this disclosure, for multiple reference rows having a reference row index m (m can be 0, 1, 2, ... and N-1, such as...) Figure 15 A reference sample in a reference row (as shown) with index m can be used only with a subset of multiple intra-prediction directions. In some embodiments, one or more reference rows can be used only with a subset of multiple intra-prediction directions. For different embodiments, the subsets of multiple intra-prediction directions of one or more reference rows can be different, overlapping, or the same. The tangent values ​​of the prediction angles associated with the subsets of multiple intra-prediction directions are ±(m+1) and ±1 / (m+1).

[0152] Figure 17 Exemplary prediction angles according to some embodiments of this disclosure are shown.

[0153] In some embodiments, different reference lines can be associated with different intra-prediction directions. For example, a solid line can indicate a direction relative to... Figure 15 The reference line 0, along with the reference line 0, is used to determine the intra-prediction direction for intra-prediction. Solid lines represent three diagonal directions (tangent value ±1), a horizontal direction (tangent value 0), and a vertical direction (tangent value ∞). Dashed lines can indicate the direction relative to the reference line. Figure 15 The reference line 1, along with the reference line 1, is used to perform intra-prediction directions. The dashed lines represent four prediction directions (tangent values ​​of ±1 / 2 and ±2). Dotted lines can indicate the direction relative to the reference line. Figure 15Reference line 2 in the table is used together with the intra-prediction direction for performing intra-prediction. The dotted line includes four prediction directions (tangent values ​​of ±1 / 3 and ±3). Dashed and dotted lines can indicate the direction relative to the reference line. Figure 15 The reference line 3 in the table is used together with the intra-prediction directions for performing intra-prediction. The dashed and dotted lines include four prediction directions (with tangent values ​​of ±1 / 4 and ±4).

[0154] In some embodiments, different reference rows may be associated with different subsets of intra-prediction directions. Different subsets of intra-prediction directions associated with certain reference rows may overlap, such as by sharing the same intra-prediction direction. For example, a solid line may indicate an intra-prediction direction used with reference row 0 to perform intra-prediction. A solid line includes three diagonal directions (tangent value ±1), a horizontal direction (tangent value 0), and a vertical direction (tangent value ∞). A dashed line may indicate an intra-prediction direction used with reference rows 0 and / or 1 to perform intra-prediction. A dashed line includes four prediction directions (tangent values ​​±1 / 2 and ±2). A dotted line may indicate an intra-prediction direction used with reference rows 0, 1, and / or 2 to perform intra-prediction. A dotted line includes four prediction directions (tangent values ​​±1 / 3 and ±3). A dashed-dotted line may indicate an intra-prediction direction used with reference rows 0, 1, 2, and / or 3 to perform intra-prediction. A dashed-dotted line includes four prediction directions (tangent values ​​±1 / 4 and ±4).

[0155] In one embodiment, for a reference row index m (m can be 0, 1, 2, ..., N-1, such as...) Figure 15 One of multiple reference rows (as shown), the reference sample in the reference row with index m can be used only with a subset of multiple intra-prediction directions. The tangent values ​​of the prediction angles associated with the subsets of multiple intra-prediction directions are ±(m+1) and ±1 / (m+1). When one of the prediction angles points to a fractional sample position in a given reference row index, the sample at the nearest integer position can be used as the reference sample.

[0156] According to aspects of this disclosure, when intra-prediction is performed at a given intra-prediction angle, the predicted samples of different rows of pixels in the current block can come from different reference rows of the current block. For example, which reference row is used for intra-prediction can vary for one or more rows of pixels in the current block.

[0157] In some embodiments, when performing intra-frame prediction, for a prediction angle with a tangent value of ±m or ±1 / m, the prediction sample of the nth row pixel can come from a reference row with line indices of (m-1)-(n%m), where % is the modulo operation.

[0158] Figure 18 An exemplary intra-frame prediction using two reference rows is illustrated according to an embodiment of this disclosure. Figure 18In the diagram, solid circles indicate reference (or predicted) samples; dashed circles indicate samples to be predicted; and solid lines indicate the prediction direction. The predicted sample for the nth row can come from the reference row with row indices (m-1)-(n%m). In this example, m=2. Therefore, the predicted samples for even rows (rows 0, 2, 4, ...) come from reference row 1, and the predicted samples for odd rows (rows 1, 3, 5, ...) come from reference row 0.

[0159] Figure 19 An exemplary intra-frame prediction using three reference rows is illustrated according to an embodiment of this disclosure. Figure 19 In the diagram, solid circles indicate reference (or predicted) samples, dashed circles indicate samples to be predicted, and solid lines indicate the prediction direction. The predicted sample for the nth row can come from the reference row with row indices (m-1)-(n%m). In this example, m=3. Therefore, the predicted samples for the first plurality of rows (rows 0, 3, 6, ...) come from reference row 2, the predicted samples for the second plurality of rows (rows 1, 4, 7, ...) come from reference row 1, and the predicted samples for the third plurality of rows (rows 2, 5, 8, ...) come from reference row 0.

[0160] According to aspects of this disclosure, the intra-prediction mode described above can be referred to as the interpolation-free intra-prediction mode, and is signaled as an alternative to the conventional intra-prediction mode used to perform intra-prediction. It can be determined which of the interpolation-free intra-prediction mode and the conventional intra-prediction mode is used. For example, for a block, a flag can be signaled to indicate whether the conventional intra-prediction mode (e.g., mode set #0 with interpolation) or the intra-prediction mode described above (e.g., mode set #1, interpolation-free orientation mode) is applied.

[0161] In one embodiment, different intra-frame prediction mode schemes can be applied to mode set #0 (with interpolation) and mode set #1 (without interpolation).

[0162] In one embodiment, the predicted angles in pattern set #1 are a subset of the predicted angles in pattern set #0.

[0163] In one embodiment, pattern set #1 does not include one or more of the vertical, horizontal, and 45-degree angles.

[0164] In one embodiment, to represent the directional prediction modes in mode set #1 (without interpolation) with signals, a flag is first represented with signals to indicate whether the most probable mode (MPM) is applied. If the MPM is not applied, then a fixed-length code can be used to encode one of the remaining intra-prediction modes.

[0165] In one embodiment, for mode set #1, in addition to the non-interpolation directional modes described above, other non-directional modes can also be represented by signals, which may include, but are not limited to, DC mode, planar mode, SMOOTH mode, SMOOTH_H mode, SMOOTH_V mode, Paeth mode, recursive filtering mode and / or matrix-based intra-prediction mode (MIP).

[0166] In one embodiment, when mode set #1 is selected, the reference line index is not represented or parsed in the bitstream using signals.

[0167] In one embodiment, when mode set #1 is selected, the reference line index is not represented or parsed in the bitstream, but reference lines with non-zero indices can still be used for intra-frame prediction.

[0168] In one embodiment, the interpolation-free intra-prediction mode described above may be applied only to certain block locations, such as when the block is not located at the top boundary of the CTU that includes the block.

[0169] V. Flowchart

[0170] Figure 20 A flowchart of an exemplary method (2000) according to an embodiment of this disclosure is shown. In various embodiments, the method (2000) is executed by processing circuitry, such as processing circuitry in a first terminal device (210), a second terminal device (220), a third terminal device (230), and a fourth terminal device (240), processing circuitry performing the functions of a video encoder (303), processing circuitry performing the functions of a video decoder (310), processing circuitry performing the functions of a video decoder (410), processing circuitry performing the functions of an intra-frame image prediction unit (452), processing circuitry performing the functions of a video encoder (503), processing circuitry performing the functions of a predictor (535), processing circuitry performing the functions of an intra-frame encoder (622), processing circuitry performing the functions of an intra-frame decoder (772), and so on. In some embodiments, the method (2000) is implemented in software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the method (2000).

[0171] Method (2000) may typically begin at step (S2010), where method (2000) decodes prediction information for the current block in the current image, which is part of an encoded video sequence. The prediction information indicates one of a plurality of intra-prediction directions for the current block. Then, method (2000) proceeds to step (S2020).

[0172] At step (S2020), the method (2000) determines a first subset of multiple reference lines based on the intra-frame prediction direction. Then, the method (2000) proceeds to step (S2030).

[0173] At step (S2030), the method (2000) performs intra-frame prediction of the current block based on a first subset of the determined plurality of reference rows. Then, the method (2000) proceeds to step (S2040).

[0174] At step (S2040), method (2000) reconstructs the current block based on the intra-frame prediction. Then, method (2000) ends.

[0175] In one embodiment, the number of reference rows in the first subset of the determined plurality of reference rows is greater than 1.

[0176] In one embodiment, among the plurality of intra-prediction directions, the intra-prediction direction associated with a first reference line among the plurality of reference lines is different from the intra-prediction direction associated with a second reference line among the plurality of reference lines.

[0177] In one embodiment, the plurality of intra-prediction directions are associated with a first reference row among the plurality of reference rows, or a second subset of the plurality of intra-prediction directions is associated with the first reference row among the plurality of reference rows, and a third subset of the plurality of intra-prediction directions is associated with the second reference row among the plurality of reference rows. For example, the first reference row may be associated with a set of intra-prediction directions, while the remaining reference rows are associated with a subset of the set of intra-prediction directions.

[0178] In one embodiment, the method (2000) determines a reference row in the first subset for the corresponding sample based on the intra-frame prediction direction and the position of each sample in the current block.

[0179] In one embodiment, the prediction information includes syntax elements that indicate whether to perform the intra-frame prediction on the current block based on the plurality of reference lines.

[0180] In one embodiment, the current block is not located near the top boundary of the encoding tree unit that includes the current block.

[0181] In one embodiment, one of the tangent and cotangent values ​​of the prediction angle associated with the intra-frame prediction direction is an integer.

[0182] In one embodiment, the method (2000) determines a reference row index for a reference row in the first subset based on the tangent of the prediction angle associated with the intra-frame prediction direction and the row number of each row of the samples in the current block.

[0183] VI. Computer System

[0184] The above-described technology can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 21 A computer system (2100) is shown, which is adapted to implement certain embodiments of the disclosed subject matter.

[0185] The computer software can be encoded using any suitable machine code or computer language, and code including instructions can be created through mechanisms such as assembly, compilation, and linking. These instructions can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through decoding, microcode, etc.

[0186] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0187] Figure 21 The components shown for the computer system (2100) are exemplary in nature and are not intended to limit the scope or functionality of the computer software used to implement the embodiments of this application. Nor should the configuration of the components be construed as having any dependency or requirement on any component or combination thereof shown in the exemplary embodiments of the computer system (2100).

[0188] The computer system (2100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through tactile input (e.g., keyboard input, swiping, data glove movement), audio input (e.g., sound, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-machine interface device may also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from still cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0189] Human-machine interface input devices may include one or more of the following (only one is shown): keyboard (2101), mouse (2102), touchpad (2103), touch screen (2110), data glove (not shown), joystick (2105), microphone (2106), scanner (2107), camera (2108).

[0190] The computer system (2100) may also include certain human-machine interface (HMI) output devices. Such HMI output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such HMI output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (2110), data gloves (not shown), or joystick (2105), but may also include tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (2109), headphones (not shown)), visual output devices (e.g., screens (2110) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light-emitting diode screens, each with or without touchscreen input functionality, each with or without tactile feedback functionality—some of which may output two-dimensional or more three-dimensional visual outputs by means such as stereoscopic image output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown). These visual output devices (e.g., screens (2110)) may be connected to the system bus (2148) via a graphics adapter (2150).

[0191] The computer system (2100) may also include human-accessible storage devices and related media, such as optical media including high-density read-only / rewritable optical discs (CD / DVD ROM / RW) (2120) or similar media (2121), thumb drives (2122), removable hard disk drives or solid-state drives (2123), conventional magnetic media such as magnetic tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.

[0192] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0193] The computer system (2100) may also include a network interface (2154) leading to one or more communication networks (2155). The one or more communication networks (2155) may be wireless, wired, or optical. The one or more communication networks (2155) may also be local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), vehicular and industrial networks, real-time networks, latency-tolerant networks, etc. Examples of the one or more communication networks (2155) include Ethernet, wireless LANs, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANbus), etc. Some networks typically require external network interface adapters for connection to certain general-purpose data ports or peripheral buses (2149) (e.g., a USB port on the computer system (2100)); other systems are typically integrated into the core of the computer system (2100) via a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). By using any of these networks, the computer system (2100) can communicate with other entities. This communication can be unidirectional, for receiving only (e.g., wireless television), unidirectional, for sending only (e.g., CAN bus to certain CAN bus devices), or bidirectional, such as via a local area or wide area digital network to other computer systems. Each of the aforementioned networks and network interfaces can use certain protocols and protocol stacks.

[0194] The aforementioned human-computer interface device, human-accessible storage device, and network interface can be connected to the core (2140) of the computer system (2100).

[0195] The core (2140) may include one or more central processing units (CPU) (2141), graphics processing units (GPUs) (2142), dedicated programmable processing units in the form of field-programmable gate arrays (FPGAs) (2143), task-specific hardware accelerators (2144), etc. These devices, along with read-only memory (ROM) (2145), random access memory (2146), internal mass storage (e.g., internal non-user-accessible hard disk drives, solid-state drives, etc.) (2147), etc., can be connected via a system bus (2148). In some computer systems, the system bus (2148) may be accessed as one or more physical connectors to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (2148) or connected via a peripheral bus (2149). In one example, a screen (2110) may be connected to a graphics adapter (2150). Peripheral bus architectures include external controller interfaces (PCI), universal serial buses (USB), etc.

[0196] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (2145) or RAM (2146). Transient data can also be stored in RAM (2146), while permanent data can be stored, for example, in internal mass storage (2147). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (2141), GPUs (2142), mass storage (2147), ROM (2145), RAM (2146), etc.

[0197] The computer-readable medium may contain computer code for performing various computer-implemented operations. The medium and computer code may be specifically designed and constructed for the purposes of this application, or they may be media and code well-known and usable by those skilled in the art of computer software.

[0198] By way of example and not limitation, a computer system having an architecture (2100), particularly a core (2140), can provide functionality as a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the aforementioned user-accessible mass storage, as well as specific memory of the non-volatile core (2140), such as internal mass storage (2147) or ROM (2145). Software implementing various embodiments of this application can be stored in such a device and executed by the core (2140). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the core (2140), particularly the processor therein (including a CPU, GPU, FPGA, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (2146) and modifying such data structures according to software-defined processes. Alternatively or as an alternative, the computer system may provide logic hardwired or otherwise incorporated into circuitry (e.g., an accelerator (2144)) that may replace or operate with the software to perform the specific process or a specific portion of the specific process described herein. References to software may include logic, and vice versa, where appropriate. References to computer-readable media may include, where appropriate, circuitry storing the execution of software (such as an integrated circuit (IC)), circuitry containing execution logic, or both. This application includes any suitable combination of hardware and software.

[0199] While this application has described several exemplary embodiments, various modifications, arrangements, and equivalent substitutions of the embodiments are all within the scope of this application. Therefore, it should be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of this application and are thus within its spirit and scope.

[0200] Appendix A: Acronyms

[0201] ALF: Adaptive Loop Filter

[0202] AMVP: Advanced Motion Vector Prediction

[0203] APS: Adaptive Parameter Set

[0204] ASIC: Application-Specific Integrated Circuit

[0205] ATMVP: Alternative / Advanced Temporal Motion Vector Prediction

[0206] AV1: Open Media Alliance Video 1 (AOMedia Video 1)

[0207] AV2: Open Media Alliance Video 2 (AOMedia Video 2)

[0208] BMS: Benchmark set

[0209] BV: Block Vector

[0210] CANBus: Controller Area Network Bus

[0211] CB: Coding Block

[0212] CC-ALF: Cross-Component Adaptive Loop Filter

[0213] CD: Compact Disc

[0214] CDEF: Constrained Directional Enhancement Filter

[0215] CPR: Current Picture Referencing

[0216] CPUs: Central Processing Units

[0217] CRT: Cathode Ray Tube

[0218] CTBs: Coding Tree Blocks

[0219] CTUs: Coding Tree Units

[0220] CU: Coding Unit

[0221] DPB: Decoder Picture Buffer

[0222] DPS: Decoding Parameter Set

[0223] DVD: Digital Video Disc

[0224] FPGA: Field Programmable Gate Array

[0225] JCCR: Joint CbCr Residual Coding

[0226] JVET: Joint Video Exploration Team

[0227] GOPs: Groups of Pictures

[0228] GPUs: Graphics Processing Units

[0229] GSM: Global System for Mobile Communications

[0230] HDR: High Dynamic Range

[0231] HEVC: High Efficiency Video Coding

[0232] HRD: Hypothetical Reference Decoder

[0233] IBC: Intra Block Copy

[0234] IC: Integrated Circuit

[0235] ISP: Intra Sub-Partitions

[0236] JEM: Joint Exploration Model

[0237] LAN: Local Area Network

[0238] LCD: Liquid Crystal Display

[0239] LR: Loop Restoration Filter

[0240] LTE: Long-Term Evolution

[0241] MPM: Most Probable Mode

[0242] MV: Motion Vector

[0243] OLED: Organic Light-Emitting Diode

[0244] PBs: Prediction Blocks

[0245] PCI: Peripheral Component Interconnect

[0246] PDPC: Position-Dependent Prediction Combination

[0247] PLD: Programmable Logic Device

[0248] PPS: Picture Parameter Set

[0249] PUs: Prediction Units

[0250] RAM: Random Access Memory

[0251] ROM: Read-Only Memory

[0252] SAO: Sample Adaptive Offset

[0253] SCC: Screen Content Coding

[0254] SDR: Standard Dynamic Range

[0255] SEI: Supplementary Enhancement Information

[0256] SNR: Signal-to-Noise Ratio

[0257] SPS: Sequence Parameter Set

[0258] SSD: Solid-state drive

[0259] TUs: Transform Units

[0260] USB: Universal Serial Bus

[0261] VPS: Video Parameter Set

[0262] VUI: Video Usability Information

[0263] VVC: Versatile Video Coding

[0264] WAIP: Wide-Angle Intra Prediction.

Claims

1. A method for video decoding, characterized in that, The method includes: Decode the prediction information of the current block in the current image, which is part of an encoded video sequence, and the prediction information indicates one of multiple intra prediction directions for the current block; A first subset of multiple reference rows is determined based on the tangent of a prediction angle associated with the intra-frame prediction direction; Based on a first subset of the determined plurality of reference rows, perform intra-frame prediction for the current block; and The current block is reconstructed based on the intra-frame prediction; The determination of a first subset of multiple reference rows includes: determining a reference row index for a reference row in the first subset based on the tangent of the prediction angle associated with the intra-frame prediction direction and the row number of each row sample in the current block.

2. The method according to claim 1, wherein, The number of reference rows in the first subset of the identified plurality of reference rows is greater than 1.

3. The method according to claim 1, wherein, Among the plurality of intra-prediction directions, the intra-prediction direction associated with the first reference line among the plurality of reference lines is different from the intra-prediction direction associated with the second reference line among the plurality of reference lines.

4. The method according to claim 1, wherein, The plurality of intra-frame prediction directions are associated with a first reference line among the plurality of reference lines, or a second subset of the plurality of intra-frame prediction directions is associated with a first reference line among the plurality of reference lines, and a third subset of the plurality of intra-frame prediction directions is associated with a second reference line among the plurality of reference lines.

5. The method according to any one of claims 1-4, wherein, The determination of a first subset of multiple reference rows includes: determining a reference row in the first subset for a corresponding sample based on the intra-frame prediction direction and the position of each sample in the current block.

6. The method according to any one of claims 1-4, wherein, The prediction information includes syntax elements that indicate whether to perform the intra-frame prediction on the current block based on the plurality of reference lines.

7. The method according to any one of claims 1-4, wherein, The current block is not located near the top boundary of the coding tree unit that includes the current block.

8. The method according to any one of claims 1-4, wherein, One of the tangent and cotangent values ​​of the prediction angle associated with the intra-frame prediction direction is an integer.

9. The method according to claim 1, wherein, The step of determining a reference row index for a reference row in the first subset based on the tangent of the prediction angle associated with the intra-frame prediction direction and the row number of each row sample in the current block includes: For a predicted angle with a tangent of ±m or ±1 / m, the reference row index for the nth row pixel is determined as (m-1) - (n%m), where % is the modulo operation and m is a non-negative integer.

10. The method according to claim 1, wherein, The step of performing intra-frame prediction for the current block based on a first subset of the determined plurality of reference rows includes: When the prediction angle corresponding to the intra-frame prediction direction points to the fractional sample position in the reference row index of the reference row, the sample in the reference row index that is closest to the fractional sample position is used as the reference sample. Based on the reference sample, perform intra-frame prediction for the current block.

11. The method according to any one of claims 1-4, wherein, When the prediction information indicates one of the multiple intra-prediction directions for the current block, the reference line index of the multiple reference lines is not parsed from the bitstream of the encoded video sequence.

12. A video encoding method, characterized in that, The method includes: Determine prediction information for the current block in the current image, which is part of a video sequence, wherein the prediction information indicates one of multiple intra-prediction directions for the current block; A first subset of multiple reference rows is determined based on the tangent of a prediction angle associated with the intra-frame prediction direction; Based on a first subset of the determined plurality of reference rows, perform intra-frame prediction for the current block; and Based on the intra-frame prediction, the current block is encoded in the already encoded video bitstream; The determination of a first subset of multiple reference rows includes: determining a reference row index for a reference row in the first subset based on the tangent of the prediction angle associated with the intra-frame prediction direction and the row number of each row sample in the current block.

13. A method for storing video streams, characterized in that, The video encoding method of claim 12 is used to generate a video stream and to store the video stream.

14. A method for transmitting a video stream, characterized in that, The video encoding method of claim 12 is used to generate a video stream and transmit the video stream.

15. An electronic device, comprising a processing circuit, characterized in that, The processing circuit is configured to perform the method as described in any one of claims 1-14.

16. A non-transitory computer-readable storage medium storing a computer program or instructions that, when executed by at least one processor, cause at least one processor to perform the steps of the method of claim 12 to generate a video stream.

Citation Information

Patent Citations

  • Method and apparatus for video decoding using multiple line intra prediction

    US10419754B1

  • Position dependent intra prediction combination extended with angular modes

    US20190306513A1