Harmonized design for offset-based fine-tuning and multiple reference line selection

Offset refinement and multiple reference line selection enhance intra-prediction and motion compensation in video coding, addressing inefficiencies and improving compression efficiency and bitrate management.

KR102994062B1Active Publication Date: 2026-07-21TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2022-01-27
Publication Date
2026-07-21

Smart Images

  • Figure 112023027710264-PCT00023_ABST
    Figure 112023027710264-PCT00023_ABST
Patent Text Reader

Abstract

A method, apparatus, and computer-readable storage medium for offset fine-tuning for intra prediction and multi-reference line intra prediction in video decoding. The method comprises the step of receiving a coded video bitstream for a block by the apparatus. The apparatus comprises a memory for storing instructions and a processor communicating with the memory. The method further comprises the step of determining, by the apparatus, based on mode information of the block, whether offset fine-tuning for intra prediction is applied to the block, wherein the mode information of the block comprises at least one of the following: a reference line index of the block, an intra prediction mode of the block, and a size of the block; and, in response to the determination that offset fine-tuning for intra prediction is applied to the block, the apparatus comprises the step of performing offset fine-tuning to generate an intra predictor for intra prediction of the block.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Related applications

[0002] This application claims the benefit of priority based on U.S. Provisional Application No. 63 / 217,061 filed on June 30, 2021 and U.S. General Application No. 17 / 573,823 filed on January 12, 2022, the entire contents of which are incorporated herein by reference.

[0003] The present disclosure relates to video coding and / or decoding technology, and in particular to an improved design and signaling of offset-based refinement and multiple reference line selection methods. Background Technology

[0004] The background description provided in this specification is intended to provide the context of the present disclosure in general. To the extent that the work is described in this background section, the works of the currently named inventors, as well as aspects of the description that may not be recognized as prior art at the time of filing this application, are not recognized as prior art of the present disclosure, either expressly or impliedly.

[0005] Image and / or video coding and decoding can be performed using inter-picture prediction along with motion compensation. Uncompressed digital video may contain a series of pictures, each having, for example, a spatial dimension of 1920×1080 luminance samples (also called luminance samples) and associated full or subsampled chrominance samples (also called chroma samples). The series of pictures may have, for example, 60 pictures per second or a fixed or variable picture rate (alternatively called the frame rate) of 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, and 4:2:0 chroma subsampling at 8 bits per pixel per chroma channel requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires more than 600GB of storage space.

[0006] One objective of video coding and decoding may be to reduce the redundancy of the uncompressed input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage requirements by more than double digits in some cases. Both lossless compression and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to a technique that allows an exact copy of the original signal to be reconstructed from the compressed original signal through the decoding process. Lossy compression refers to a coding / decoding process where the original video information is not fully preserved during coding and cannot be fully recovered during decoding. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough that the reconstructed signal remains useful for the intended application despite some loss of information. For video, lossy compression is widely adopted in many application programs. The amount of acceptable distortion varies depending on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of movie or television distribution applications. The compression ratios achievable with specific coding algorithms can be selected or adjusted to reflect varying distortion tolerances: higher acceptable distortion generally allows for coding algorithms that provide higher loss and higher compression ratios.

[0007] Video encoders and decoders can utilize techniques from various broad categories and stages, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.

[0008] Video codec technology may include a technique known as intra coding. In intra coding, sample values ​​are represented without reference to samples of a previously reconstructed reference picture or other data. In some video codecs, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture may be referred to as an intra picture. Derivations such as the intra picture and the independent decoder refresh picture can be used to reset the decoder state and thus serve as the first picture of the coded video bitstream and video session, or as a still image. After intra prediction, samples in the block can be transformed into the frequency domain, and the resulting transformation coefficients can be quantized before entropy coding. Intra prediction refers to a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after conversion and the smaller the AC coefficient, the fewer bits are required at a given quantization step size to represent the block after entropy coding.

[0009] For example, traditional intra-coding, such as that known from MPEG-2 generative coding techniques, does not use intra-prediction. However, some newer video compression techniques include methods that attempt to code / decode blocks based on surrounding sample data and / or metadata acquired, for example, during the encoding and / or decoding of spatially neighboring blocks and preceding the block of data being intra-coded or decoded in the decoding order. These techniques will be referred to as "intra-prediction" techniques. Note that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, rather than from other reference pictures.

[0010] There may be various forms of intra-prediction. If one or more of these techniques are available in a given video coding technique, the technique in use may be referred to as an intra-prediction mode. One or more intra-prediction modes may be provided in a specific codec. In some cases, a mode may have submodes and / or may be associated with various parameters, and the mode / submode information and intra-coding parameters for a video block may be coded individually or collectively included in a mode codeword. Which codeword is used for a given combination of mode, submode, and / or parameters can affect the coding efficiency gain through intra-prediction, and thus may also affect the entropy coding technique used to convert the codeword into a bitstream.

[0011] Intra prediction in specific modes was introduced with H.264, improved in H.265, and further enhanced in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). Generally, for intra prediction, predictor blocks can be formed using available neighbor sample values. For example, available values ​​from a specific set of neighbor samples along a specific direction and / or line can be copied into the predictor block. References to the direction in use can be coded into the bitstream or predicted themselves.

[0012] Referring to FIG. 1a, a subset of nine specified predictor directions among the 33 possible intra predictor directions of H.265 (corresponding to the 33 angle modes of the 35 intra modes specified in H.265) is shown in the lower right. The point where the arrows converge (101) indicates the predicted sample. The arrows indicate the direction in which the sample is predicted. The arrows indicate the direction in which neighboring samples are used to predict the sample (101). For example, arrow (102) indicates that the sample (101) is predicted to the upper right at a 45-degree angle in the horizontal direction from neighboring samples or samples. Similarly, arrow (103) indicates that the sample (101) is predicted to the lower left of the sample (101) at a 22.5-degree angle in the horizontal direction from neighboring samples or samples.

[0013] Referring still to FIG. 1a, a square block (104) of 4×4 samples (indicated by a bold dashed line) is shown in the upper left corner. The square block (104) contains 16 samples, each labeled "S", a position in the Y dimension (e.g., row index) and a position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Likewise, sample S44 is the fourth sample of the block (104) in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is located in the lower right corner. An exemplary reference sample following a similar numbering scheme is additionally illustrated. The reference sample is labeled R, a Y position (e.g., row index) and an X position (column index) relative to the block (104). In both H.264 and H.265, prediction samples adjacent to the block being rebuilt are used.

[0014] Intra-picture prediction of block (104) can be initiated by copying a reference sample value from a neighbor sample along the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating the prediction direction of the arrow (102) for this block (104)—that is, the samples are predicted from one prediction sample or samples toward the upper right at a 45-degree angle from the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0015] In some cases, to calculate reference samples, especially when the direction cannot be divided equally into 45 degrees, the values ​​of multiple reference samples may be combined through interpolation, for example.

[0016] The number of possible directions has increased as video coding technology has continued to advance. In H.264 (2003), for example, nine different directions are available for intra prediction. In H.265 (2013), this increased to 33, and at the time of this disclosure, JEM / VVC / BMS can support up to 65 directions. Experimental studies have been conducted to help identify the most suitable intra prediction direction, and specific techniques in entropy coding can be used to encode such most suitable directions with a small number of bits, accepting a certain bit penalty for the direction. Additionally, the direction itself can sometimes be predicted from the neighboring direction used for intra prediction of the decoded neighboring block.

[0017] FIG. 1b illustrates a schematic diagram (180) showing 65 intra-predicted directions according to JEM to explain the increasing number of predicted directions in various encoding technologies developed over time.

[0018] The method of mapping bits representing the intra-prediction direction to the prediction direction of a coded video bitstream may vary depending on the video coding technique; for example, it ranges from simple direct mapping of the prediction direction to complex adaptive schemes involving intra-prediction modes, codewords, and the most likely mode, as well as similar techniques. However, in all cases, there may be specific directions for intro prediction that are statistically less likely to occur than other specific directions within the video content. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technique, less likely directions may be represented by a larger number of bits than more likely directions.

[0019] Inter-picture prediction or inter-prediction may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or part thereof (reference picture) may be spatially shifted in the direction indicated by a motion vector (hereinafter MV) and then used for prediction of a newly reconstructed picture or part of a picture (e.g., a block). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may be 2-dimensional with X and Y, or 3-dimensional with a third dimension (similar to a time dimension) that is an indication of the reference picture being used.

[0020] In some video compression techniques, the current MV applicable to a specific area of ​​sample data can be predicted from other MVs, for example, from other MVs associated with other areas of sample data that are spatially adjacent to the area being reconstructed and precede the current MV in the decoding order. This can significantly reduce the total amount of data required to code the MV by relying on the deduplication of correlated MVs, thereby increasing compression efficiency. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction can operate effectively because there is a statistical probability that an area larger than the area where a single MV can be applied moves in a similar direction within the video sequence; thus, in some cases, it can be predicted using similar motion vectors derived from MVs of neighboring areas. As a result, the actual MV for a given area is similar or identical to the MV predicted from surrounding MVs. After entropy coding, such MVs can be represented with fewer bits than would have been used if the MV had been coded directly rather than predicted from neighboring MV(s). In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself may be lost, for example, because rounding error occurs when calculating the predictor from several surrounding MVs.

[0021] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms specified by H.265, the technique described here is hereinafter referred to as "spatial merge".

[0022] Referring to FIG. 2, the current block (201) may include samples discovered by an encoder during the motion search process so as to be predictable from a previous block of the same size that is spatially shifted. Instead of directly coding the MV, the MV may be derived from metadata associated with one or more reference pictures, for example from the most recent (in decoding order) reference picture, using an MV associated with one of five neighboring samples denoted as A0, A1 and B0, B1, B2 (202–206, respectively). In H.265, the MV prediction may use a predictor from the same reference picture used by the neighboring block.

[0023] The present disclosure describes various embodiments of a method, apparatus, and computer-readable storage medium for video encoding and / or decoding.

[0024] According to one aspect, an embodiment of the present disclosure provides a method for offset refinement for intra prediction and multiple reference line intra prediction in video decoding, performed by a device. The method comprises the step of the device receiving a coded video bitstream for a block. The method further comprises the step of the device determining, based on mode information of the block, whether offset refinement for intra prediction is applied to the block—the mode information of the block comprises at least one of a reference line index of the block, an intra prediction mode of the block, and a size of the block—and the step of performing the offset refinement to generate an intra predictor for intra prediction of the block in response to the determination that offset refinement for intra prediction is applied to the block.

[0025] According to another aspect, one embodiment of the present invention provides an apparatus for video encoding and / or decoding. The apparatus includes a memory for storing instructions; and a processor that communicates with said memory. When said processor executes instructions, said processor is configured to cause said apparatus to perform said method for video decoding and / or encoding.

[0026] In another aspect, one embodiment of the present invention provides a computer-readable non-transient medium storing instructions, said instructions causing the computer to perform said method for video decoding and / or encoding when the instructions are executed by the computer for video decoding and / or encoding.

[0027] The above-mentioned aspects and other aspects and their implementations are described in more detail in the drawings, description, and claims. Brief explanation of the drawing

[0028] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the attached drawings. Figure 1a illustrates a schematic diagram of an exemplary subset of intra-predicted direction modes. Figure 1b illustrates an example of an exemplary intra-predicted direction. Figure 2 illustrates a schematic diagram of the current block and surrounding space merge candidates for motion vector prediction in one example. FIG. 3 illustrates a schematic diagram of a simplified block diagram of a communication system (300) according to an exemplary embodiment. FIG. 4 illustrates a schematic diagram of a simplified block diagram of a communication system (400) according to an exemplary embodiment. FIG. 5 illustrates a schematic diagram of a simplified block diagram of a video decoder according to an exemplary embodiment. FIG. 6 illustrates a schematic diagram of a simplified block diagram of a video encoder according to an exemplary embodiment. FIG. 7 illustrates a block diagram of a video encoder according to another exemplary embodiment. FIG. 8 illustrates a block diagram of a video decoder according to another exemplary embodiment. FIG. 9 illustrates a directional intra-prediction mode according to an exemplary embodiment of the present disclosure. FIG. 10 illustrates a non-directional intra-prediction mode according to an exemplary embodiment of the present disclosure. FIG. 11 illustrates a recursive intra-prediction mode according to an exemplary embodiment of the present disclosure. FIG. 12 illustrates an intra-prediction method based on various reference lines according to an exemplary embodiment of the present disclosure. FIG. 13 illustrates offset-based fine-tuning for intra-prediction according to an exemplary embodiment of the present disclosure. FIG. 14a illustrates another diagram of offset-based fine-tuning for intra-prediction according to an exemplary embodiment of the present disclosure. FIG. 14b illustrates another illustration of offset-based fine-tuning for intra-prediction according to an exemplary embodiment of the present disclosure. FIG. 15 illustrates a flowchart of a method according to an exemplary embodiment of the present disclosure. FIG. 16 illustrates a schematic diagram of a computer system according to an exemplary embodiment of the present disclosure. Specific details for implementing the invention

[0029] The present invention will be described in detail below with reference to the accompanying drawings, which form part of the invention and exemplarily illustrate specific examples of embodiments. However, it should be noted that the present invention may be embodied in various different forms, and therefore the subject matter covered or claimed should not be interpreted as being limited to any of the embodiments described below. It should also be noted that the present invention may be embodied in a method, device, component, or system. Accordingly, embodiments of the present invention may take the form of, for example, hardware, software, firmware, or a combination thereof.

[0030] Throughout the specification and claims, terms may have nuances implied or suggested by the context beyond their explicitly stated meanings. The phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, nor do the phrases "in another embodiment" or "in another embodiment" as used herein refer to different embodiments. Likewise, the phrases "in one implementation" or "in some implementations" as used herein do not necessarily refer to the same implementation, nor do the phrases "in another implementation" or "in another implementation" as used herein refer to different implementations. For example, the claimed subject matter is intended to include a combination of exemplary embodiments / implements, wholly or partially.

[0031] Generally, terms may be understood at least in part from their usage in context. For example, terms such as “and,” “or,” or “and / or” as used herein may include various meanings that depend at least in part on the context in which such terms are used. Generally, when “or” is used to associate a list such as A, B, or C, it is intended to mean A, B, and C as used herein in an inclusive sense, as well as A, B, or C as used herein in an exclusive sense. Additionally, terms such as “one or more” or “at least one” as used herein may be used to describe any feature, structure, or characteristic in a singular sense, or to describe a combination of features, structures, or characteristics in a plural sense, depending at least in part on the context. Similarly, terms such as “one (a, an)” or “the,” which correspond to English articles, may be understood to convey a singular usage or a plural usage, at least in part, depending on the context. Furthermore, the terms "based on" or "determined by" may be understood as not necessarily intended to convey an exclusive set of factors, but instead may allow for the existence of additional factors that are not necessarily explicitly described, at least in part, depending on the context.

[0032] FIG. 3 illustrates a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, through a network (350). For example, the communication system (300) includes a first pair of terminal devices (310, 320) interconnected through the network (350). In the example of FIG. 3, the first pair of terminal devices (310, 320) can perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) through the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) receives coded video data from the network (350), decodes the coded video data to restore a video picture, and can display the video picture according to the restored video data. Unidirectional data transmission can be implemented in a media serving application, etc.

[0033] In another example, the communication system (300) includes a second pair of terminal devices (330, 340) that perform bidirectional transmission of coded video data, which can be implemented, for example, in a video conferencing application. For bidirectional transmission of data, in one example, each terminal device of the pair of terminal devices (330, 340) can code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to another terminal device of the pair of terminal devices (330, 340) via a network (350). Each terminal device of the pair of terminal devices (330, 340) can also receive coded video data transmitted by another terminal device of the pair of terminal devices (330, 340), decode the coded video data to restore the video, and display the video picture on an accessible display device according to the restored video data.

[0034] In the example of FIG. 3, the terminal devices (310, 320, 330, 340) may be implemented as servers, personal computers, and smartphones, but the applicability of the basic principles of the present disclosure may not be so limited. Embodiments of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, etc. Network (350) represents any number or type of network that transmits coded video data between terminal devices (310, 320, 330, 340), for example, including wired and / or wireless communication networks. The communication network (350) may exchange data via circuit-switched, packet-switched, and / or other types of channels. Representative networks include communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this disclosure, the architecture and topology of the network (350) may not be important to the operation of this disclosure unless explicitly described herein.

[0035] FIG. 4 illustrates the arrangement of a video encoder and a video decoder in a video streaming environment as an example of an application to the disclosed subject. The disclosed subject may be equally applicable to other video applications, such as, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, and the storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0036] A video streaming system may include a video capture subsystem (413) which may include a video source (401), such as a digital camera, for generating a stream (402) of uncompressed video pictures or images. In one example, the stream (402) of video pictures includes samples recorded by the digital camera of the video source (401). The stream (402) of video pictures, indicated by a bold line to emphasize a high data volume compared to encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) comprising a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof capable of enabling or implementing aspects of the disclosed subject matter as described in more detail below. Encoded video data (404) (or encoded video bitstream (404)), indicated by a thin line to emphasize a lower data volume compared to a stream (402) of uncompressed video pictures, may be stored in a streaming server (405) for later use or may be stored directly in a downstream video device (not shown). One or more streaming client subsystems, such as the client subsystems (406, 408) of FIG. 4, may access the streaming server (405) to retrieve copies (407, 409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410) in the electronic device (430), for example. A video decoder (410) decodes an incoming copy (407) of encoded video data and generates an outgoing stream of a video picture (411) that can be rendered on a display (412) (e.g., a display screen) or another rendering device (not shown) without compression.The video decoder (410) may be configured to perform some or all of the various functions described in this disclosure. In some streaming systems, encoded video data (404, 407, 409) (e.g., video bitstream) may be encoded according to a specific video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC and other video coding standards.

[0037] Note that the electronic device (420, 430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0038] FIG. 5 illustrates a block diagram of a video decoder (510) according to any embodiment of the present disclosure below. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used instead of the video decoder (410) in the example of FIG. 4.

[0039] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In the same embodiment or another embodiment, one coded video sequence may be decoded at a time, wherein the decoding of each coded video sequence is independent of other coded video sequences. Each video sequence may be associated with multiple video frames or images. The coded video sequence may be received from a channel (501) which may be a hardware / software link to a storage device that stores the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective processing circuits (not shown). The receiver (531) may separate the coded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be placed between the receiver (531) and the entropy decoder / parser (520) (hereinafter "parser (520)"). In certain applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be separate from the video decoder (510) and located outside (not shown). In yet another application, for example, to prevent network jitter, the buffer memory (not shown) may be located outside the video decoder (510), and for example, to handle playback timing, there may be an additional buffer memory (515) inside the video decoder (510).When the receiver (531) receives data from a storage / forwarding device or an isosynchronous network having sufficient bandwidth and controllability, the buffer memory (515) may not be needed or may be small. For use in a best-effort packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and its size may be relatively large. Such buffer memory may be implemented in an adaptive size and may be implemented at least partially in a similar element (not shown) outside of the operating system or the video decoder (510).

[0040] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from a coded video sequence. Categories of these symbols include information used to manage the operation of the video decoder (510) and potential information for controlling a rendering device, such as a display (512) (e.g., a display screen), which may or may not be an integrated part of the electronic device (530) as illustrated in FIG. 5, but can be coupled to the electronic device (530). The control information for the rendering device(s) may be in the form of Supplemental Enhancement Information (SEI messages) or Video Usability Information (VUI) parameter set fragments (not shown). The parser (520) may parse / entropy decode the coded video sequence received by the parser (520). The coding of an entropy-coded video sequence may follow video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, and arithmetic coding with or without context sensitivity. The parser (520) may extract a set of subgroup parameters for at least one of a subgroup of pixels in a video decoder from the coded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Picture (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (520) may also extract from the coded video sequence information, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc.

[0041] The parser (520) can generate symbols (521) by performing entropy decoding / parsing operations on a video sequence received from the buffer memory (515).

[0042] The reconstruction of the symbol (521) may be associated with a number of different processing or functional units depending on the type of the coded video picture or part thereof (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors. The associated units and how they are associated may be controlled by subgroup control information parsed from the video sequence coded by the parser (520). The flow of such subgroup control information between the parser (520) and the number of processing or functional units below it is not shown for simplicity.

[0043] Beyond the previously mentioned functional blocks, the video decoder (510) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these functional units interact closely with one another and may be integrated with one another, at least partially. However, to clearly explain the various functions of the disclosed subject, the conceptual subdivision into functional units is adopted in the disclosure below.

[0044] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive quantized transform coefficients as well as control information including information indicating the type of inverse transform to be used, block size, quantization factor / parameter, and quantization scaling matrix, and is in the form of symbol(s) from the parser (520). The scaler / inverse transform unit (551) may output a block containing sample values ​​that can be input to an aggregator (555).

[0045] In some cases, the output sample of the scaler / inverse transformation unit (551) may be related to an intra-coded block. That is, a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases,

[0046] The intra-picture prediction unit (552) can generate a block of the same size and shape as the block being reconstructed by using surrounding block information that has already been reconstructed and is stored in the current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) can add the prediction information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse conversion unit (551) on a sample basis.

[0047] In other cases, the output samples of the scaler / inverse unit (551) may be intercoded and potentially associated with a motion-compensated block. In such cases, the motion compensation prediction unit (553) may access the reference picture memory (557) to retrieve samples used for inter-picture prediction. After motion-compensating the retrieved samples according to the symbol (521) associated with the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse unit (551) (the output of the unit (551) may be referred to as a residual sample or residual signal) to generate output sample information. The address in the reference picture memory (557) from which the motion compensation prediction unit (553) retrieves the prediction samples may be controlled by a motion vector and may be used by the motion compensation prediction unit (553) in the form of a symbol (521) that may have, for example, X and Y components (shift) and a reference picture component (time). Motion compensation may also include interpolation of sample values ​​drawn from a reference picture memory (557) when an accurate motion vector of a subsample is used, and may be associated with a motion vector prediction mechanism, etc.

[0048] The output sample of the aggregator (555) may be subject to various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filter techniques that are controlled by parameters contained in the coded video sequence (also referred to as the coded video bit stream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but may also respond to meta-information acquired while decoding a previous (in decoding order) portion of the coded picture or coded video sequence, as well as respond to previously reconstructed and loop-filtered sample values. As described in more detail below, various types of loop filters may be included as part of the loop filter unit (556) in various order.

[0049] The output of the loop filter unit (556) may be a sample stream that can be output to the rendering device (512) and stored in the reference picture memory (557) for use in future inter-picture prediction.

[0050] Once a specific coded picture has been completely reconstructed, it can be used as a reference picture for future inter-picture predictions. For example, when a coded picture corresponding to a current picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and the new current picture buffer can be reallocated before starting the reconstruction of the next coded picture.

[0051] A video decoder (510) can perform decoding operations according to a predetermined video compression technique in a standard, such as ITU-T Rec. H.265. The coded video sequence may follow the syntax specified by the video compression technique or standard used in that it conforms to both the syntax of the video compression technique or standard and the profile documented in the video compression technique. Specifically, the profile may select a specific tool as the only tool available in that profile among all tools available in the video compression technique or standard. To conform to the standard, the complexity of the coded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further limited by HRD specifications and metadata for managing the virtual reference decoder (HRD) buffer signaled in the coded video sequence.

[0052] In some exemplary embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, a time layer, a spatial layer or SNR enhancement layer, a redundant slice, a redundant picture, a forward error correction code, etc.

[0053] FIG. 6 illustrates a block diagram of a video encoder (603) according to one embodiment of the present disclosure. The video encoder (603) may be included in the electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) may be used instead of the video encoder (403) in the example of FIG. 4.

[0054] The video encoder (603) can receive video samples from a video source (601) (in the example of FIG. 6, not part of the electronic device (620)) capable of capturing video image(s) to be encoded by the video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620).

[0055] A video source (601) may provide a source video sequence to be coded by a video encoder (603) in the form of a digital video sample stream, which may have any appropriate bit depth (e.g., 8-bit, 10-bit, 12-bit,…), any color space (e.g., BT.601 YCrCb, RGB, XYZ…), and any appropriate sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device capable of storing a pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures or images that impart motion when viewed sequentially. The picture itself may consist of a spatial array of pixels, wherein each pixel may contain one or more samples depending on the sampling structure, color space, etc. being used. Those skilled in the art will easily understand the relationship between pixels and samples. The explanation below focuses on samples.

[0056] According to some exemplary embodiments, the video encoder (603) can code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraint required by the application. Enforcing an appropriate coding rate constitutes one function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to another function unit to control as described below. For simplicity, the coupling is not shown. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, …), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other appropriate functions related to the video encoder (603) optimized for a specific system design.

[0057] In some exemplary embodiments, the video encoder (603) may be configured to operate in a coding loop. For an oversimplified description, in one example, the coding loop may include a source coder (630) (e.g., responsible for generating symbol and reference picture(s), such as a symbol stream, based on an input picture to be coded), and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the sample to generate sample data in a manner similar to how the (remote) decoder would have generated the video stream coded by the source coder (630) even if the embedded decoder (633) processes the video stream coded by the source coder (630) without entropy coding (where any compression between the symbol and the coded video bitstream may be lossless in the video compression techniques considered in the subject matter).

[0058] The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream leads to a bit-exact result regardless of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "recognizes" the sample value as a reference picture sample that is exactly the same as what the decoder "sees" when using prediction during decoding. This fundamental principle of reference picture synchronicity (and resulting drift, for example, when synchronicity cannot be maintained due to channel error) is used to improve coding quality.

[0059] The operation of the “local” decoder (633) may be the same as the operation of the “remote” decoder, such as the video decoder (510), which has already been described in detail in relation to FIG. 5. Referring briefly back to FIG. 5, since the symbols are available and the encoding / decoding of the symbols into the video sequence coded by the entropy coder (645) and the parser (420) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and the parser (520), may not be entirely implemented in the local decoder (633) within the encoder.

[0060] What can be observed at this point is that, with the exception of parsing / entropy decoding which may exist only in the decoder, all decoder techniques may necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on the operation of the decoder in relation to the decoding portion of the encoder. Descriptions of encoder techniques may be omitted as they are the opposite of the comprehensively described decoder techniques. More detailed descriptions of the encoder are provided below only in specific areas or aspects.

[0061] During operation, in some exemplary implementations, the source coder (630) may perform motion-compensated predictive coding to predictively code an input picture by referencing one or more previously coded pictures from a video sequence designated as a "reference picture." In this way, the coding engine (632) codes the difference (or residual) in color channels between a pixel block of the input picture and a pixel block of the reference picture(s) that can be selected as predictive reference(s) for the input picture.

[0062] The local video decoder (633) can decode the coded video data of a picture that can be designated as a reference picture based on the symbols generated by the source coder (630). The operation of the coding engine (632) can advantageously be a lossy process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may generally be a replica of the source video sequence with some errors. The local video decoder (633) can replicate the decoding process that can be performed by the video decoder on the reference picture and allow the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture to be acquired by the far-end (remote) video decoder (no transmission errors).

[0063] The predictor (635) can perform a predictive search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for sample data (which is a candidate reference pixel block) that can serve as a suitable predictive reference for the new picture, or for specific metadata such as a reference picture motion vector, block shape, etc. The predictor (635) can operate on a sample block-by-pixel block basis to find a suitable predictive reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have a predictive reference drawn from a number of reference pictures stored in the reference picture memory (634).

[0064] The controller (650) can manage the coding operation of the source coder (630), including the settings of parameters and subgroup parameters used to encode video data, for example.

[0065] The output of all the aforementioned function units may be subject to entropy coding in the entropy coder (645). The entropy coder (645) converts the symbols generated by the various function units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.

[0066] The transmitter (640) may buffer the coded video sequence(s) generated by the entropy coder (645) and prepare for transmission through a communication channel (660), which may be a hardware / software link to a storage device for storing the encoded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, e.g., coded audio data and / or an auxiliary data stream (source not shown).

[0067] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each coded picture, which can affect the coding technique that can be applied to each picture. For example, a picture can often be designated as one of the following picture types:

[0068] An Intra Picture (I Picture) may be one that can be coded and decoded without using any other picture within the sequence as a prediction source. Some video codecs allow different types of Intra Pictures, including, for example, Independent Decoder Refresh Picture ("IDR"). Those skilled in the art are aware of these variations of I Pictures and their respective applications and characteristics.

[0069] The predictive picture (P picture) may be coded and decoded using intra-prediction or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0070] A bidirectionally predictive picture (B-picture) may be coded and decoded using intra-prediction or inter-prediction, which utilize up to two motion vectors and reference indices to predict sample values ​​for each block. Similarly, a multiple-predictive picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0071] A source picture can generally be spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks, respectively) and coded on a block-by-block basis. Blocks can be predictively coded by referencing other (already coded) blocks as determined by the coding assignment applied to each picture in the block. For example, blocks in picture I can be non-predictively coded or predictively coded by referencing already coded blocks (spatial prediction or intra prediction) of the same picture. Pixel blocks in picture P can be predictively coded via spatial prediction or temporal prediction by referencing a previously coded reference picture. Blocks in picture B can be predictively coded via spatial prediction or temporal prediction by referencing one or two previously coded reference pictures. A source picture or an intermediately processed picture can be subdivided into different types of blocks for different purposes. The division of coding blocks and other types of blocks may or may not follow the same method, as explained in more detail below.

[0072] The video encoder (603) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In such operations, the video encoder (603) may perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Thus, the coded video data may follow the syntax specified by the video coding technique or standard used.

[0073] In some exemplary embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include this data as part of the encoded video sequence. The additional data may include other forms of redundant data such as time / space / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0074] Video can be captured as a temporal sequence of multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as intra prediction) utilizes spatial correlations within a given picture, while inter-picture prediction utilizes temporal or other correlations between pictures. For example, a specific picture currently being encoded / decoded, referred to as the current picture, can be partitioned into blocks. If a block within the current picture is similar to a reference block in a reference picture that was previously encoded and is still buffered in the video, it can be encoded by a vector referred to as a motion vector. The motion vector points to a reference block within the reference picture and, if multiple reference pictures are in use, may have three dimensions that identify the reference picture.

[0075] In some exemplary embodiments, a bi-prediction technique may be used for inter-picture prediction. According to this bi-prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, both of which precede the current picture in the decoding order in the video (but which may be past or future in the display order, respectively). A block of the current picture may be coded by a first motion vector pointing to a first reference block of the first reference picture and a second motion vector pointing to a second reference block of the second reference picture. A block may be predicted together by a combination of the first reference block and the second reference block.

[0076] In addition, coding efficiency can be improved by using merge mode technology in inter-picture prediction.

[0077] According to some exemplary embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in blocks. For example, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and CTUs within a picture may have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU may include three parallel coding tree blocks (CTBs), namely one luminance CTB and two chroma CTBs. Each CTU may be repeatedly quadtree-divided into one or more coding units (CUs). For example, a 64×64 pixel CTU may be divided into one 64×64 pixel CU, four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. Each of one or more 32×32 blocks may be further subdivided into four CUs of 16×16 pixels. In some exemplary embodiments, each CU may be analyzed during encoding to determine the prediction type for the CU among various prediction types, such as inter-prediction type or intra-prediction type. The CU may be subdivided into one or more prediction units (PUs) based on temporal and / or spatial predictability. Generally, each PU includes a luminance prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation during coding (encoding / decoding) is performed on a unit of prediction blocks. Subdividing the CU into PUs (or PBs of color channels) may be performed in various spatial patterns. The luminance or chroma PBs may include a matrix of values ​​(e.g., luminance values) for samples, such as, for example, 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0078] FIG. 7 illustrates a drawing of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values ​​within the current video picture in a sequence of video pictures and to encode the processing block into a coded picture that is part of a coded video sequence. An exemplary video encoder (703) may be used instead of the video encoder (403) in the example of FIG. 4.

[0079] For example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) then determines whether the processing block is best coded using an intra mode, an inter mode, or a positive-prediction mode, for example, using rate-distortion optimization (RDO). If it is determined that the processing block is coded in intra mode, the video encoder (703) may use an intra prediction technique to encode the processing block into a coded picture; if it is determined that the processing block is coded in inter mode or positive-prediction mode, the video encoder (703) may use an inter prediction or positive-prediction technique, respectively, to encode the processing block into a coded picture. In some exemplary embodiments, a merge mode may be used as a submode of inter-picture prediction where the motion vector is derived from one or more motion vector predictors without the benefit of the coded motion vector component outside the predictor. In some other exemplary embodiments, there may be motion vector components applicable to the subject block. Accordingly, the video encoder (703) may include components not explicitly shown in FIG. 7, such as a mode determination module, to determine the prediction mode of the processing block.

[0080] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) combined together as shown in the exemplary arrangement of FIG. 7.

[0081] The inter-encoder (730) is configured to receive a sample of the current block (e.g., processing block), compare the block with one or more reference blocks within the reference picture (e.g., blocks within the previous and subsequent pictures in display order), generate inter-prediction information (e.g., description of redundancy information according to the inter-encoding technique, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on encoded video information using a decoding unit (633) embedded in the exemplary encoder (620) of FIG. 6 (illustrated as the residual decoder (728) of FIG. 7, as described in more detail below).

[0082] The intra-encoder (722) is configured to receive a sample of the current block (e.g., processing block), compare the block with a block already coded in the same picture, and generate quantized coefficients after conversion, and, in some cases, also intra-prediction information (e.g., intra-prediction direction information according to one or more intra-encoding techniques). The intra-encoder (722) calculates an intra-prediction result (e.g., predicted block) based on the intra-prediction information and reference block of the same picture.

[0083] A general controller (721) may be configured to determine general control data and to control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, if the prediction mode is an intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use in the residual calculator (723) and controls the entropy encoder (725) to select intra prediction information and include the intra prediction information in the bit stream; if the prediction mode for the block is an inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use in the residual calculator (723) and controls the entropy encoder (725) to select inter prediction information and include the inter prediction information in the bit stream.

[0084] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and the prediction result for the block selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transformation coefficients. For example, the residual encoder (724) may be configured to convert the residual data from the spatial domain to the frequency domain to generate transformation coefficients. The transformation coefficients then undergo quantization processing to obtain quantized transformation coefficients. In various exemplary embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transformation and generate decoded residual data. The decoded residual data may be appropriately used by the intra-encoder (722) and the inter-encoder (730). For example, the inter encoder (730) can generate a decoded block based on decoded residual data and inter prediction information, and the intra encoder (722) can generate a decoded block based on decoded residual data and intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.

[0085] The entropy encoder (725) may be configured to format the bitstream to include the encoded block and to perform entropy coding. The entropy encoder (725) may be configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. When coding the block in the inter mode or the merged submode of the two-prediction mode, there may be no residual information.

[0086] FIG. 8 illustrates an exemplary video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and to decode the coded picture to produce a reconstructed picture. In one example, the video decoder (810) may be used instead of the video decoder (410) in the example of FIG. 4.

[0087] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) combined together as shown in the exemplary arrangement of FIG. 8.

[0088] The entropy decoder (871) may be configured to reconstruct, from the coded picture, specific symbols representing the syntax elements constituting the coded picture. These symbols may include, for example, the mode in which the block is coded (e.g., intra mode, inter mode, positive-prediction mode, the latter two being merge submodes or other submodes), prediction information (e.g., intra prediction information or inter prediction information) capable of identifying specific samples or metadata used for prediction by the intra decoder (872) or inter decoder (880), residual information in the form of, for example, quantized transformation coefficients. In one example, if the prediction mode is an inter mode or positive-prediction mode, the inter prediction information is provided to the inter decoder (880); if the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information may be dequantized and provided to the residual decoder (873).

[0089] The inter decoder (880) can be configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0090] The intra decoder (872) can be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0091] The residual decoder (873) may be configured to perform inverse quantization to extract inverse quantized transform coefficients and process the inverse quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize specific control information (to include quantizer parameters (QP)) that may be provided by the entropy decoder (871) (the data path is not shown as it may only contain a small amount of control information).

[0092] The reconstruction module (874) may be configured to combine the residual output by the residual decoder (873) and the prediction result (output by the inter-prediction module or intra-prediction module in some cases) in the spatial domain to form a reconstructed block that forms part of the reconstructed picture as part of the reconstructed video. Note that other appropriate actions, such as deblocking, may be performed to improve visual quality.

[0093] It should be noted that the video encoder (403, 603, 703) and video decoder (410, 510, 810) may be implemented using any suitable technology. In some exemplary embodiments, the video encoder (403, 603, 703) and video decoder (410, 510, 810) may be implemented using one or more integrated circuits. In other embodiments, the video encoder (403, 603, 703) and video decoder (410, 510, 810) may be implemented using one or more processors that execute software instructions.

[0094] Returning to the intra-prediction process, samples within a block (e.g., a luminal or chroma prediction block, or a coding block if not further subdivided into prediction blocks) are predicted by neighbors, next-neighbors, or other lines or lines, or combinations thereof, to generate a prediction block. The residual between the actual block being coded and the prediction block can be processed through a transformation following quantization. Various intra-prediction modes may be available, and parameters related to intra-mode selection and other parameters may be signaled in the bitstream. For example, various intra-prediction modes may relate to line positions or locations for predicting samples, the direction in which prediction samples are selected from the prediction lines or lines, and other special intra-prediction modes.

[0095] For example, a set of intra prediction modes (also interchangeably referred to as "intra modes") may include a predefined number of directional intra prediction modes. As previously described in relation to the exemplary implementation of FIG. 1, these intra prediction modes may correspond to a predefined number of directions in which an out-of-block sample is selected as a prediction for a sample predicted in a specific block. In another specific exemplary implementation, eighteen (8) major directional modes corresponding to angles from 45 degrees to 207 degrees with respect to the horizontal axis may be supported and predefined.

[0096] In some other implementations of intra prediction, to further leverage more diverse spatial redundancy in directional textures, directional intra modes can be extended to a set of angles with finer granularity. For example, the above 8-angle implementation can be configured to provide eight nominal angles designated as V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, as illustrated in FIG. 9, and for each nominal angle, a predefined number (e.g., 7) of finer angles can be added. With such an extension, a total of more directional angles (e.g., 56 in this example) can be used for intra prediction corresponding to the same number of predefined directional intra modes. The predicted angle can be represented as the nominal intra angle plus an angle delta. In the case of the specific example above, where there are 7 finer angle directions for each nominal angle, the angle delta can be -3 to 3, multiplied by a step size of 3 degrees.

[0097] The above-described directional intra-prediction may also be referred to as a unidirectional intra-prediction, distinct from the bidirectional intra-prediction (also referred to as intra-bidirectional prediction) described later in the present disclosure.

[0098] In some implementations, as an alternative or in addition to the directional intra modes above, a predefined number of non-directional intra prediction modes may be predefined and made available. For example, five non-directional intra modes referred to as smooth intra prediction modes may be specified. These non-directional intra prediction modes may specifically be referred to as DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H intra modes. Predictions of samples of a specific block under these exemplary non-directional modes are illustrated in FIG. 10. As an example, FIG. 10 illustrates a 4×4 block (1002) predicted by samples from the top neighbor line and / or the left neighbor line. A specific sample (1010) of block (1002) may directly correspond to the upper sample (1004) of the sample (1010) in the upper neighbor line of block (1002), and the upper-left sample (1006) of the sample (1010) may correspond to the sample (1008) immediately to the left of the sample (1010) in the left neighbor line of block (1002) as the intersection of the upper and left neighbor lines. In the case of an exemplary DC intra prediction mode, the average of the left and upper neighbor samples (1008 and 1004) may be used as the predictor for the sample (1010). For example, in PAETH intra prediction mode, the upper, left, and upper-left reference samples (1004, 1008, and 1006) are drawn up, and then the value closest to (upper + left - upper-left) among these three reference samples may be set as the predictor for the sample (1010). In the case of an exemplary SMOOTH_V intra prediction mode, a sample (1010) can be predicted by quadratic interpolation in the vertical direction of the upper-left neighbor sample (1006) and the left neighbor sample (1008).In an exemplary SMOOTH_H intra prediction mode, a sample (1010) can be predicted by second-order interpolation in the horizontal direction of the top-left neighbor sample (1006) and the top-neighbor sample (1004). For an exemplary SMOOTH intra prediction mode, a sample (1010) can be predicted by the average of second-order interpolation in the vertical and horizontal directions. The above non-directional intra mode implementations are described merely as examples and not as limitations. Other non-directional choices of other neighbor lines and samples, and methods of combining prediction samples to predict a specific sample in a prediction block are also considered.

[0099] At various coding levels (picture, slice, block, unit, etc.), the selection of a specific intra-prediction mode by the encoder from the above directional or non-directional modes can be signaled in the bitstream. In some exemplary implementations, eight exemplary nominal directional modes may be signaled first, along with five non-angular smoothing modes (a total of 13 options). Then, if the signaled mode is one of the eight nominal angular intra-modes, an index is additionally signaled to indicate the selected angular delta for the corresponding signal nominal angle. In some other exemplary implementations, all intra-prediction modes may be indexed together for signaling (e.g., adding 56 directional modes and 5 non-directional modes yields 61 intra-prediction modes).

[0100] In some exemplary implementations, 56 or other exemplary directional intra prediction modes may be implemented as a unified directional predictor that projects each sample of the block to a reference subsample location and interpolates the reference sample with a 2-tap bilinear filter.

[0101] In some implementations, additional filter modes called FILTER INTRA modes may be designed to capture references on edges and attenuating spatial correlations. For these modes, in addition to samples outside the block, prediction samples within the block may be used as intra-prediction reference samples for some patches within the block. For example, these modes may be predefined and may be used for intra-prediction for at least a luminance block (or only a luminance block). A predefined number (e.g., 5) of filter intra modes may be pre-designed, each represented as a set of n-tap filters (e.g., 7-tap filters) that reflect the correlation between samples, for example, between a 4×2 patch and n adjacent neighbors. In other words, the weighting coefficients for the n-tap filters may vary depending on the location. Taking an 8×8 block, 4×2 patch, and 7-tap filtering as an example, the 8×8 block (1102) may be divided into eight 4×2 patches, as illustrated in FIG. 11. These patches are denoted as B0, B1, B3, B4, B5, B6, and B7 in FIG. 11. For each patch, seven neighbors, denoted as R0 through R7 in FIG. 11, can be used to predict samples in the current patch. In the case of patch B0, all neighbors may have already been reconstructed. However, for other patches, some neighbors may not have been reconstructed because they are in the current block, so the predicted values ​​of the immediate neighbors are used as a reference. For example, as shown in FIG. 11, since not all neighbors of patch B7 are reconstructed, predicted samples of neighbors, e.g., parts of B4, B5, and / or B6, are used instead.

[0102] In some implementations of intra-prediction, a single color component can be predicted using one or more other color components. The color components can be any one of the components of the YCrCb, RGB, or XYZ color spaces. For example, the prediction of a chroma component (e.g., a chroma block) from a lumina component (e.g., a lumina reference sample) (referred to as Chroma from Luma, or CfL) can be implemented. In some exemplary implementations, cross-color prediction is primarily allowed from lumina to chroma. For example, the chroma samples of a chroma block can be modeled as a linear function of the matched reconstructed lumina samples. CfL prediction can be implemented as follows:

[0103] (1)

[0104] Here represents the AC contribution of the luminance component, α represents the parameter of the linear model, and DC represents the DC contribution of the chroma component. For example, the AC component is acquired for each sample of the block, while the DC component is acquired for the entire block. Specifically, the reconstructed luminance sample can be subsampled to the chroma resolution, and the average luminance value (DC of luminance) is subtracted from each luminance value to form the AC contribution of the luminance. The AC contribution of the luminance is then used in the linear mode of Equation (1) to predict the AC value of the chroma component. To approximate or predict the chroma AC component from the luminance AC contribution, instead of requiring the decoder to calculate the scaling parameter, an exemplary CfL implementation can determine the parameter α based on the original chroma sample and signal it in the bitstream. This reduces decoder complexity and provides a more accurate prediction. Regarding the DC contribution of the chroma component, in some exemplary implementations, it can be calculated using an intra-DC mode within the chroma component.

[0105] Returning to intra prediction, in some exemplary implementations, the prediction of a sample in a coding block or prediction block may be based on one of a set of reference lines. In other words, instead of always using the nearest neighbor line (e.g., the immediate top neighbor line or the immediate left neighbor line of the prediction block as illustrated in Figure 1 above), multiple reference lines may be provided as options for the selection of intra prediction. Such an intra prediction implementation may be referred to as Multiple Reference Line Selection (MRLS). In such an implementation, the encoder determines and signals which of the multiple reference lines is used to generate the intra predictor. On the decoder side, after parsing the reference line index, the intra prediction of the current intra prediction block may be generated by identifying the reconstructed reference sample by finding the specified reference line according to the intra prediction mode (e.g., directional, non-directional, and other intra prediction modes). In some implementations, the reference line index may be signaled at the coding block level, and only one of the multiple reference lines may be selected and used for the intra prediction of a single coding block. In some examples, multiple reference lines may be selected together for intra-prediction. For example, the multiple reference lines may be combined, averaged, interpolated, or any other method may be performed with or without weights to generate a prediction. In some exemplary implementations, MRLS may be applied only to the luminance component and not to the chroma component(s).

[0106] An example of a four-reference line MRLS is illustrated in FIG. 12. As illustrated in the example of FIG. 12, an intra-coding block (1202) can be predicted based on one of four horizontal reference lines (1204, 1206, 1208, 1210) and four vertical reference lines (1212, 1214, 1216, 1218). Among these reference lines, 1210 and 1218 are immediate neighbor reference lines. Reference lines can be indexed according to their distance from the coding block. For example, reference lines 1210 and 1218 can be referred to as zero reference lines, while other reference lines can be referred to as non-zero reference lines. Specifically, reference lines 1208 and 1216 can be referred to as first reference lines; reference lines 1206 and 1214 can be referred to as second reference lines; Reference lines 1204 and 1212 may be referenced as third reference lines.

[0107] In some embodiments, to improve video encoding / decoding performance, offset-based refinement for intra prediction (ORIP) may be used after generating intra prediction samples. When ORIP is applied, the prediction samples are refined by adding an offset value.

[0108] As illustrated in FIG. 13, intra prediction (1330) is performed based on reference samples. Reference samples may include samples from one or more left reference lines (1312) and / or one or more top reference lines (1310). Offset-based fine-tuning (ORIP) (1350) for intra prediction may generate offset values ​​using neighbor reference samples. In some implementations, the neighbor reference samples for ORIP may be the same set as the reference samples for intra prediction. In some other implementations, the neighbor reference samples for ORIP may be a different set from the reference samples for intra prediction.

[0109] In some implementations, referring to FIG. 14a and FIG. 14b, ORIP may be performed at the 4×4 subblock level. For each 4×4 subblock (1471, 1472, 1473 and / or 1474), an offset is generated from neighbor samples. For example, for the first subblock (1471), the offset is generated from top neighbor samples (P1, P2, P3 and P4 of 1420), left neighbor samples (P5, P6, P7 and P8 of 1410), and / or left top neighbor samples (P0) (1401). In some implementations, the top neighbor samples may include both top neighbor samples (P1, P2, P3 and P4 of 1420) and left top neighbor samples (P0) (1401). In some other implementations, the left neighbor samples may include both the left neighbor samples (P5, P6, P7, and P8 of 1410) and the left top neighbor sample (P0) (1401).

[0110] The first subblock (1471) includes 4×4 pixels, and each pixel of the 4×4 pixels corresponds to predN, which is the Nth neighbor prediction sample before fine-tuning. For example, pred0, pred1, pred2, … pred16.

[0111] In various embodiments, the offset value of each pixel of a given subblock may be calculated based on neighbor samples according to a formula. The formula may be a predefined formula or a formula indicated by parameters coded in a coded bitstream.

[0112] In some implementations referring to FIG. 14b, the offset value (offset(k)) of the k-th position of a given subblock can be generated as follows.

[0113] (2)

[0114] (3)

[0115] W kn is a predefined weight for offset calculation. P nis the value of a neighboring sample (e.g., P0, P1, P2, …, P8). pred k is the predicted value for the pixel after applying intra-prediction or other predictions (e.g., inter-prediction). pred_refined k is the fine-tuned value for the pixel after applying ORIP. clip3() is the clip3 mathematical function. n is an integer from 0 to 8 (inclusive). k is an integer from 0 to 15 (inclusive).

[0116] In some implementations, W kn It can be predefined and obtained according to Table 1.

[0117] Table 1: Predefined weights for offset calculation

[0118]

[0119] In some other implementations, subblock-based ORIP may only apply to a predefined set of intra prediction modes, and / or may differ for luminance and chroma depending on the intra prediction mode. Table 2 shows one implementation of subblock-based ORIP according to various intra prediction modes and the luminance or chroma channel. Taking the luminance channel as an example: if the prediction mode is DC or SMOOTH, ORIP is always ON and no additional signaling is required; if the prediction mode is HOR / VER and angle_delta is 0, block-level signaling is required to enable / disable ORIP; and / or if the intra prediction mode is any other mode, ORIP is always OFF and no additional signaling is required.

[0120] Table 2: Mode-dependent ON / OFF of the proposed method

[0121]

[0122] Referring again to the second 4×4 subblock (1473), due to the relative position with respect to the first 4×4 subblock (1471), the top neighbor sample of the second subblock may be some pixels of the first subblock: P1 of the second subblock may be pred12 of the first block, P2 of the second subblock may be pred13 of the first subblock, P3 of the second subblock may be pred14 of the first subblock, and P4 of the second subblock may be pred15 of the first subblock. The left top neighbor sample (P0) of the second subblock may be the left neighbor sample (P8) of the first subblock.

[0123] There may be some issues / problems associated with certain ORIP implementations. For example, if the intra prediction mode of the block's luminance component is vertical or horizontal intra prediction mode, there may be three different options: 1) select an adjacent reference line and apply ORIP to the block; 2) select an adjacent reference line and do not apply ORIP to the block; and / or 3) select a non-adjacent reference line and do not apply ORIP to the block. Since reference samples from adjacent and non-adjacent reference lines can generally be similar, there may be some overlap between option 2) and option 3), which can degrade the performance of video encoding / decoding.

[0124] The present disclosure describes various embodiments for offset fine-tuning for intra prediction and multi-reference line intra prediction in video coding and / or decoding, and addresses at least one of the issues / problems discussed above.

[0125] In various embodiments, referring to FIG. 15, in an offset fine-tuning method (1500) for intra prediction and multi-reference line intra prediction in video decoding, the method (1500) may include some or all of the following steps: step 1510, a step in which a device including a memory storing instructions and a processor communicating with the memory receives a coded video bitstream for a block; step 1520, a step in which the device determines whether offset fine-tuning for intra prediction is applied to the block based on mode information of the block, wherein the mode information of the block includes at least one of a reference line index of the block, an intra prediction mode of the block, and a size of the block; and / or step 1530, a step in which, in response to the determination that offset fine-tuning for intra prediction is applied to the block, the device performs offset fine-tuning to generate an intra predictor for intra prediction of the block. In some implementations, step 1520 may include the device determining whether offset fine-tuning for intra-prediction is applied to the block based on mode information of the block, and the mode information of the block includes at least one of the following: a reference line index of the block, an intra-prediction mode of the block, or a size of the block.

[0126] In various embodiments of the present disclosure, the size of a block (e.g., a coding block, a prediction block, or a transformation block, but not limited thereto) may refer to the width or height of the block. The width or height of the block may be an integer in pixels.

[0127] In various embodiments of the present disclosure, the size of a block (e.g., a coding block, a prediction block, or a transformation block, but not limited thereto) may refer to the area size of the block. The area size of the block may be an integer obtained by multiplying the width of the block and the height of the block in pixels.

[0128] In some various embodiments of the present disclosure, the size of a block (e.g., a coding block, a prediction block, or a transformation block, but not limited thereto) may mean a maximum value of width or height, a minimum value of width or height of a block, or an aspect ratio of a block. The aspect ratio of a block may be calculated by dividing the width of a block by the height of a block, or by dividing the height of a block by the width of a block.

[0129] In the present disclosure, a reference line index indicates a reference line among a plurality of reference lines. In various embodiments, a reference line index of 0 for a block may indicate an adjacent reference line for the block, which is also the reference line closest to the block. For example, referring to the block (1202) of FIG. 12, the top reference line (1210) is the top adjacent reference line for the block (1202), which is also the top closest reference line for the block; the left reference line (1218) is the left adjacent reference line for the block (1202), which is also the left reference line closest to the block. A reference line index greater than 0 for a block indicates a non-adjacent reference line for the block, which is also the reference line not closest to the block. For example, referring to the block (1202) of FIG. 12, a reference line index of 1 may indicate the top reference line (1208) and / or the left reference line (1216); A reference line index of 2 can indicate an upper reference line (1206) and / or a left reference line (1214); and / or a reference line index of 3 can indicate an upper reference line (1204) and / or a left reference line (1212).

[0130] In various embodiments for video coding and / or decoding, when a directional intra prediction mode is selected for a block (e.g., a coding block or a coded block), it may be determined whether an offset fine-tuning for intra prediction (ORIP) is applied to the block to generate intra predictors from samples at non-adjacent reference lines. This determination may rely on mode information of the block, which may include, but is not limited to, a reference line index, one or more intra prediction angles, and / or the size of the block (or) the block size. If this determination is made based on the mode information of the block, no additional signaling is required to indicate whether ORIP is applied, thereby reducing the overhead of video encoding / decoding and / or improving the performance of video encoding / decoding. In various embodiments, offset fine-tuning for intra prediction (ORIP) may mean offset-based fine-tuning for intra prediction, which is a type of intra prediction in which the prediction result is further fine-tuned (or modified) according to an offset value generated based on one or more neighboring reference samples.

[0131] Referring to step 1510, the device may be the electronic device (530) of FIG. 5 or the video decoder (810) of FIG. 8. In some implementations, the device may be the decoder (633) within the encoder (620) of FIG. 6. In other implementations, the device may be part of the electronic device (530) of FIG. 5, part of the video decoder (810) of FIG. 8, or part of the decoder (633) within the encoder (620) of FIG. 6. The coded video bitstream may be the coded video sequence of FIG. 8 or the intermediate encoded data of FIG. 6 or FIG. 7. The block may mean a coding block or a coded block.

[0132] Referring to step 1520, the device can determine whether offset fine-tuning for intra-prediction is applied to the block based on the block's mode information, and the block's mode information may include at least one of the block's reference line index, the block's intra-prediction mode, and the block's size.

[0133] In various embodiments, step 1520 may include: determining that offset fine-tuning for intra-prediction is applied to the block in response to the block's reference line index being smaller than a predefined threshold and the block's intra-prediction mode belonging to a directional intra-prediction mode; and / or determining that offset fine-tuning for intra-prediction is not applied to the block in response to the block's reference line index not being smaller than a predefined threshold and the block's intra-prediction mode belonging to a directional intra-prediction mode.

[0134] In some implementations, for a set of directional intra-prediction modes, ORIP is always applied if the value of the reference line index is less than a predefined value; and / or ORIP is not applied if the value of the reference line index is greater than or equal to a predefined value. In one example, this set of directional intra-prediction modes includes vertical mode and horizontal mode. The set of directional intra-prediction modes includes vertical mode, horizontal mode, 45-degree directional mode, and 135-degree directional mode. In another example, the predefined value is N and is set to 1. In another example, N is set to 2. In yet another example, N is set to 0 or 3.

[0135] In various embodiments, step 1520 may include: determining that an offset fine-tuning for intra-prediction is applied to the block in response to the block's reference line index being even and the block's intra-prediction mode belonging to a directional intra-prediction mode; and / or determining that an offset fine-tuning for intra-prediction is not applied to the block in response to the block's reference line index being odd and the block's intra-prediction mode belonging to a directional intra-prediction mode.

[0136] In some implementations, for a set of directional intra-prediction modes, the determination of whether ORIP applies depends on whether the value of the reference line index is even or odd. In one example, for a set of directional intra-prediction modes, ORIP is applied if the value of the reference line index is even; ORIP is not applied if the value of the reference line index is odd. In another example, for a set of directional intra-prediction modes, ORIP is applied if the value of the reference line index is odd; ORIP is not applied if the value of the reference line index is even. In another example, this set of directional intra-prediction modes includes vertical mode and horizontal mode. In another example, the set of directional intra-prediction modes includes vertical mode, horizontal mode, 45-degree directional mode, and 135-degree directional mode.

[0137] In various embodiments, step 1520 comprises: determining that an offset fine-tuning for intra-prediction is applied to the block in response to the block's intra-prediction mode belonging to a first selected set of intra-prediction modes; determining that an offset fine-tuning for intra-prediction is applied to the block in response to the block's intra-prediction mode belonging to a second selected set of intra-prediction modes and a reference line index indicating an adjacent reference line; and / or determining that an offset fine-tuning for intra-prediction is not applied to the block in response to the block's intra-prediction mode belonging to a second selected set of intra-prediction modes and a reference line index indicating a non-adjacent reference line. In some other implementations, the first selected set of intra-prediction modes may include the second selected set of intra-prediction modes.

[0138] In some implementations, for an intra prediction mode of a selected set, ORIP applies to both adjacent and non-adjacent reference lines, whereas for an intra prediction mode of the remaining set, ORIP applies only to adjacent reference lines. In one example, if intra prediction is applied using only integer samples, e.g., for a diagonal (45 / 225 degree) intra prediction mode, ORIP may also apply to non-adjacent reference lines; and / or if intra prediction does not use integer samples, ORIP applies only to adjacent reference lines.

[0139] In various embodiments, step 1520 may include the step of determining that an offset fine-tuning for intra prediction is applied to the block in response to the fact that the intra prediction mode of the block belongs to a directional intra prediction mode and the reference line index of the block indicates a non-adjacent reference line.

[0140] In various embodiments, step 1520 may further include the step of using mode information of the block as a context for parameters for offset fine-tuning for intra-prediction during entropy decoding.

[0141] In some other implementations, when a directional intra prediction mode is selected for a block, an ORIP may still be applicable to generate an intra predictor from samples at a non-adjacent reference line, and / or an ORIP may still be signaled by a parameter (e.g., ORIP flag or ORIP index) in the coded bitstream regardless of the block's mode information value.

[0142] In some other implementations during video encoding, the context used for entropy coding of ORIP parameters (e.g., flags / indexes) indicating whether and how ORIP is applied depends on the mode information of the current block. In some implementations, the context for entropy coding of syntax values ​​indicating whether and how other modes of the current block are applied depends on the ORIP flags / indexes indicating whether and how ORIP is applied to the block. Examples of syntax for other modes of the current block include, but are not limited to, reference line indices, one or more intra-predicted angles, and / or block sizes.

[0143] Referring to step 1530, in response to the determination that offset fine-tuning for intra-prediction is applied to a block, the device performs offset fine-tuning for intra-prediction on the block. In some implementations, the block is a 4×4 block or comprises a plurality of 4×4 blocks, and offset fine-tuning for intra-prediction is performed on the block according to a predefined method, for example, using Equations (2) and (3) as described above.

[0144] In various embodiments, step 1530 may include determining a predictor from a non-adjacent reference line indicated by the reference line index of the block; determining a predicted value for the block based on an intra prediction based on the predictor; and / or modifying the predicted value based on an offset fine-tuning for an intra prediction based on a sample from an adjacent reference line. In some implementations, the non-adjacent reference line and the adjacent reference line come from both sides; the sides include a top side and a left side.

[0145] In some implementations, when ORIP is applied, to generate an intra-predictor from samples from non-adjacent reference lines, a predictor (e.g., predictor A) is generated from samples from non-adjacent reference lines selected by MRLS, and then ORIP further modifies the predicted values ​​from predictor A using samples from adjacent reference lines.

[0146] In some other implementations, the reference samples used to generate predictor A and to further modify predictor values ​​from (or based on) predictor A come from different sides of neighboring reference samples. Examples of the said sides of neighboring reference samples include, but are not limited to, top reference samples and left reference samples. In one example, one or more top reference samples are used to generate predictor A, and one or more left reference samples are used to modify predictor values ​​from (or based on) predictor A. In another example, one or more left reference samples are used to generate predictor A, and one or more top reference samples are used to modify predictor values ​​from predictor A.

[0147] In various embodiments, step 1530 may include: determining a predictor from a non-adjacent reference line indicated by a reference line index of the block; determining a predicted value for the block based on an intra prediction based on the predictor; and / or modifying the predicted value based on an offset fine-tuning for an intra prediction based on samples from multiple reference lines.

[0148] In some implementations, step 1530 may further include the step of obtaining samples based on samples and corresponding weights using a linear weighted average, wherein the corresponding weights are predefined according to the relative positions of the samples.

[0149] In some other implementations, where ORIP is applied, to generate intra predictors from samples of non-adjacent reference lines,

[0150] Predicted samples are fine-tuned by ORIP using samples from multiple (more than one) reference lines. In one example, ORIP is generated using a linear weighted sum of samples from multiple (multiple) reference lines. The weights can be predefined based on the relative positions of the samples.

[0151] Embodiments of the present disclosure may be used individually or combined in any order. Additionally, each method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a computer-readable medium. Embodiments of the present disclosure may be applied to a luminance block or a chroma block; in a chroma block, the embodiment may be applied individually to a plurality of color components or together to a plurality of color components.

[0152] In the present disclosure, any step or operation of various embodiments may be combined in any amount or any order as desired. In the present disclosure, two or more steps or operations may be performed in parallel in various embodiments.

[0153] The aforementioned technology may be implemented as computer software that uses computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 16 illustrates a computer system (2600) suitable for implementing a specific embodiment of the disclosed subject matter.

[0154] Computer software may be coded using any suitable machine code or computer language capable of generating code containing instructions that can be executed directly, or through interpretation, micro-code execution, etc., by a computer central processing unit (CPU), graphics processing unit (GPU), etc., via assembly, compilation, linking, or similar mechanisms.

[0155] The instruction can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0156] The components of the computer system (2600) illustrated in FIG. 16 are essentially exemplary and are not intended to imply any limitation to the scope of use or function of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be interpreted as having any dependency or requirement related to any one or combination of the components shown in the exemplary embodiments of the computer system (2600).

[0157] The computer system (2600) may include a specific human interface input device. This human interface input device may respond to input by one or more human users, for example, tactile input (e.g., keystroke, swip, data glove movement), audio input (e.g., voice, clap), visual input (e.g., gesture), and olfactory input (not shown). The human interface device may also be used to capture specific media that are not necessarily directly related to conscious input by a person, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0158] The input human interface device may include one or more of a keyboard (2601), a mouse (2602), a trackpad (2603), a touchscreen (2610), a data glove (not shown), a joystick (2605), a microphone (2606), a scanner (2607), and a camera (2608) (each only one is shown).

[0159] The computer system (2600) may include a specific human interface output device. This human interface output device may stimulate the senses of one or more human users, for example, through tactile output, sound, light and smell / taste. These human interface output devices may include a tactile output device (e.g., a touch screen (2610), a data glove (not shown), or tactile feedback via a joystick (2605), but there may also be a tactile feedback device that does not serve as an input device), an audio output device (e.g., a speaker (2609), headphones (not shown)), a visual output device (e.g., a screen (2610) including a CRT screen, an LCD screen, a plasma screen, and an OLED screen, each with or without a touch screen input function and with or without a tactile feedback function—some of which may produce two-dimensional visual output or three-dimensional or higher output through means such as stereographic output, virtual-reality glasses (not shown), a holographic display and a smoke tank (not shown)—), and a printer (not shown).

[0160] The computer system (2600) may also include human-accessible storage devices and associated media, such as optical media including a CD / DVD ROM RW (2620) having a CD / DVD media (2621), a thumb drive (2622), a removable hard drive or solid-state drive (2623), legacy magnetic media such as tape and floppy disk (not shown), and specialized ROM / ASIC / PLD-based devices such as a security dongle (not shown).

[0161] Those skilled in the art should also understand that the term “computer-readable medium” as used in connection with the subject matter disclosed herein does not include a transmitting medium, a carrier wave, or other transient signals.

[0162] The computer system (2600) may also include an interface (2654) to one or more communication networks (2655). The network may be, for example, a wireless, wired, optical network. The network may also be a local, wide-area, metropolitan, automotive and industrial, real-time, latency-tolerant network. Examples of networks include cellular networks including Ethernet, wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, automotive and industrial networks including a CAN bus, etc. A specific network generally requires an external network interface adapter attached to a specific general-purpose data port or peripheral bus (2649) (e.g., a USB port of the computer system (2600)); others are generally integrated into the core of the computer system (2600) by attaching to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2600) can communicate with another network. This communication may be unidirectional, receive-only (e.g., TV broadcasting), unidirectional transmit-only (e.g., from a CANbus to a specific CANbus device), or bidirectional (e.g., to another computer system using a local or wide-area digital network). Specific protocols and protocol stacks may be used for each of the networks and network interfaces as described above.

[0163] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (2640) of the computer system (2600).

[0164] The core (2640) may include one or more central processing units (CPUs) (2641), graphics processing units (GPUs) (2642), specialized programmable processing units in the form of a field programmable gate array (FPGA) (2643), hardware accelerators (2644) for specific tasks, graphics adapters (2650), etc. Along with read-only memory (ROM) (2645), random access memory (2646), and internal mass storage devices (2647) such as internal hard drives or SSDs that are not accessible to the user, these devices may be connected via a system bus (2648). In some computer systems, the system bus (2648) may be accessible in the form of one or more physical plugs that enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (2648) or connected via a peripheral bus (2649). For example, a screen (2610) can be connected to a graphics adapter (2650). Architectures for peripheral buses include PCI, USB, etc.

[0165] The CPU (2641), GPU (2642), FPGA (2643), and accelerator (2644) can be combined to execute specific instructions that can construct the aforementioned computer code. The computer code may be stored in ROM (2645) or RAM (2646). Transitional data may also be stored in RAM (2646), while persistent data may be stored, for example, in an internal mass storage device (2647). Fast storage and retrieval of any of the memory elements may be made possible through the use of a cache memory that may be closely associated with one or more CPUs (2641), GPUs (2642), mass storage devices (2647), ROM (2645), RAM (2646), etc.

[0166] A computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of a type well known and available to those skilled in the art of computer software.

[0167] As an example, not limited to, the computer system having an architecture (2600), specifically a core (2640), may provide functionality as a result of processor(s) (including CPU, GPU, FPGA, accelerator, etc.) executing software implemented on one or more types of computer-readable media. Such computer-readable media may be a medium associated with a mass storage device accessible to a user as described above, as well as a specific storage device of the core (2640) of a non-transient nature, such as a mass storage device (2647) or ROM (2645) inside the core. Software implementing various embodiments of the present disclosure may be stored on such devices and executed by the core (2640). The computer-readable media may include one or more memory elements or chips as needed. Software may enable the core (2640) and, in particular, the internal processor (including a CPU, GPU, FPGA, etc.) to execute a specific process or a specific part of a specific process described herein, including defining data structures stored in RAM (2646) and modifying such data structures according to a process defined by the software. Additionally or alternatively, the computer system may provide a function that is otherwise implemented in a circuit (e.g., accelerator (2644)) as a result of logic hardwired, which can operate instead of or with the software to execute a specific process or a specific part of a specific process described herein. References to software may include logic, and where appropriate, vice versa. References to computer-readable media may include circuits that store software for execution (e.g., integrated circuits (ICs)), circuits that implement logic for execution, or both, where appropriate. The present disclosure includes any suitable combination of hardware and software.

[0168] Although specific inventions have been described with reference to exemplary embodiments, this description is not intended to be limiting. Various modifications of the exemplary embodiments and additional embodiments of the invention will be apparent to those skilled in the art from this description. Those skilled in the art will readily recognize that these and various other modifications may be made to the exemplary embodiments shown and described herein without departing from the spirit and scope of the invention. Accordingly, the appended claims are intended to include any such modifications and alternative embodiments. Certain proportions in the drawings may be exaggerated, and other proportions may be minimized. Accordingly, the disclosure and drawings should be regarded as illustrative rather than limiting.

Claims

Claim 1 A video decoding method comprising: receiving a coded video bitstream for a block by a device including a memory for storing instructions and a processor communicating with said memory; determining by said device whether an offset-based intra prediction refinement (ORIP) mode is enabled for said block based on a reference line index of said block - the step of determining whether the ORIP mode is enabled based on a reference line index of said block includes: determining that the ORIP mode is enabled for said block when said reference line index is smaller than a predefined value; and determining that the ORIP mode is disabled for said block when said reference line index is larger than the predefined value -; and, when the ORIP mode is enabled for said block, by said device performing offset refinement to generate an intra predictor for intra prediction of said block. Claim 2 A video decoding method according to claim 1, wherein the step of determining whether the ORIP mode is enabled for the block is additionally based on whether the intra prediction mode of the block belongs to a set of directional intra prediction modes. Claim 3 A video decoding method according to claim 1, further comprising the step of performing intra prediction of the block, wherein the intra prediction of the block is performed using a vertical intra prediction mode or a horizontal intra prediction mode. Claim 4 A video decoding method according to claim 1, further comprising the step of performing intra prediction of the block, wherein the intra prediction of the block is performed using one of a vertical mode, a horizontal mode, a 45-degree direction mode, and a 135-degree direction mode. Claim 5 A video decoding method according to claim 1, wherein the predefined value includes one of 1 or 2. Claim 6 A video decoding method according to claim 1, wherein the step of determining whether the ORIP mode is enabled for the block comprises: determining that the ORIP mode is enabled for the block when the intra prediction mode of the block belongs to the intra prediction mode of a first selected set; determining that the ORIP mode is enabled for the block when the intra prediction mode of the block belongs to the intra prediction mode of a second selected set and the reference line index indicates an adjacent reference line; and determining that the ORIP mode is not enabled for the block when the intra prediction mode of the block belongs to the intra prediction mode of the second selected set and the reference line index indicates a non-adjacent reference line, wherein the intra prediction mode of the first selected set does not overlap with the intra prediction mode of the second selected set. Claim 7 In claim 6, the intra prediction mode of the first selected set includes a diagonal intra prediction mode, a video decoding method. Claim 8 A video decoding method according to claim 1, wherein the step of generating an intra predictor for intra prediction of the block by performing offset fine-tuning comprises: determining a predictor from a non-adjacent reference line indicated by a reference line index of the block; determining a predicted value for the block according to intra prediction based on the predictor; and modifying the predicted value according to offset fine-tuning for intra prediction based on samples from adjacent reference lines. Claim 9 In paragraph 8, the above non-adjacent reference line and the above adjacent reference line come from both sides; said both sides include the upper side and the left side, a video decoding method. Claim 10 A video decoding method according to claim 1, wherein the step of generating an intra predictor for intra prediction of the block by performing the offset fine-tuning comprises: determining a predictor from a non-adjacent reference line indicated by a reference line index of the block; determining a predicted value for the block according to an intra prediction based on the predictor; and modifying the predicted value according to the offset fine-tuning for intra prediction based on samples from multiple reference lines. Claim 11 A video decoding method according to claim 10, further comprising the step of acquiring samples based on the samples and corresponding weights using a linear weighted average, wherein the corresponding weights are predetermined according to the relative positions of the samples. Claim 12 A video decoding method according to claim 1, wherein the step of determining whether the ORIP mode is enabled for the block comprises the step of determining that the ORIP mode is enabled for the block when the intra prediction mode of the block belongs to a set of directional intra prediction modes and the reference line index of the block indicates a non-adjacent reference line. Claim 13 A video decoding method according to claim 12, further comprising the step of using mode information of the block as a context for parameters for offset fine-tuning for intra-prediction during entropy decoding. Claim 14 An offset fine-tuning device for intra prediction and multi-reference line intra prediction in video decoding, comprising: a memory for storing instructions; and a processor communicating with said memory, wherein when said processor executes said instructions, said processor is configured to cause said offset fine-tuning device to perform the method of any one of claims 1 to 13. Claim 15 A computer-readable non-transient storage medium for storing instructions, wherein when said instructions are executed by a processor, said instructions are configured to cause said processor to perform the method of any one of claims 1 to 13. Claim 16 delete