Method and computer program for video encoding / decoding
MRL prediction with weighted combinations for intra prediction fusion addresses inefficiencies in existing video coding technologies, enhancing coding efficiency and compression ratios by optimizing intra prediction.
Patent Information
- Application Number
- JP2024515587
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-10
- Filing Date
- 2022-11-11
- Publication Date
- 2025-07-25
AI Technical Summary
Existing video coding technologies face inefficiencies in intra prediction due to the use of single reference lines, which can lead to increased bit usage for less likely prediction directions, and motion vector prediction mechanisms that do not fully exploit spatial and temporal redundancy.
Implementing multi-reference line (MRL) prediction that utilizes multiple reference lines with weighted combinations for intra prediction fusion, determining gradient costs to select optimal weight candidates, and applying these to predict samples in the current block.
Enhances coding efficiency by reducing bit usage for less likely prediction directions and improving compression ratios through more accurate intra prediction, while maintaining image quality.
Smart Images

Figure 2025523725000001_ABST
Abstract
Description
Technical Field
[0001] This application claims priority to U.S. Provisional Application No. 63 / 388,905, filed on July 13, 2022, entitled "ON THE WEIGHT DERIVATION OF MULTIPLE REFERENCE LINE FOR INTRA PREDICTION FUSION", and claims the benefit of priority to U.S. Patent Application No. 17 / 984,926, filed on November 10, 2022, entitled "WEIGHT DERIVATION OF MULTIPLE REFERENCE LINE FOR INTRA PREDICTION FUSION". The disclosure of the prior application is hereby incorporated by reference in its entirety.
[0002] This disclosure generally describes embodiments related to video coding.
Background Art
[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the presently named inventors, to the extent that the research is not otherwise considered prior art at the time of filing and is not explicitly or implicitly admitted as prior art to the present disclosure, is presented in a manner that is not typically considered prior art within the scope described in this background art.
[0004] Uncompressed digital images and / or videos can include a series of pictures, each picture having, for example, spatial dimensions of 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have, for example, a fixed or variable picture rate of 60 pictures per second or 60 Hz (informally also known as the frame rate). Uncompressed images and / or videos have specific bitrate requirements. For example, an 8-bit-per-sample 1080p60 4:2:0 video (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage area of over 600 GB.
[0005] One purpose of image and / or video coding and decoding may be the reduction of redundancy in the input image and / or video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage requirements, sometimes by more than an order of magnitude. The description here uses video encoding / decoding as an illustrative example, but the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of the present disclosure. Both reversible compression and irreversible compression, as well as combinations thereof, can be employed. Reversible compression refers to techniques that can reconstruct an exact copy of the original signal from the compressed signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. In the case of video, irreversible compression is widely used. The amount of distortion tolerated depends on the application. For example, a user of a particular consumer streaming application may tolerate higher distortion than a user of a television distribution application. The achievable compression ratio may reflect the fact that higher allowable / tolerable distortion can result in a higher compression ratio.
[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.
[0007] Video coding technology can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video coders, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in an intra mode, that picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. The samples of an intra block can be transformed, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique that minimizes sample values in a pre-transform domain. In some cases, the smaller the DC value and AC coefficients after transformation, the fewer bits are required at a given quantization step size to represent the block after entropy coding.
[0008] For example, traditional intra coding used in MPEG-2 generation coding technology does not use intra prediction. However, some newer video compression technologies include techniques that attempt to perform prediction based on surrounding sample data and / or metadata obtained during the encoding and / or decoding of blocks of data. Such techniques are referred to below as "intra prediction" techniques. It should be noted that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and not from reference pictures.
[0009] There can be many different forms of intra prediction. In a given video coding technology, when two or more of such technologies can be used, the specific technology in use can be coded as a specific intra prediction mode that uses the specific technology. In some cases, the intra prediction mode can have sub - modes and / or parameters, where the sub - modes and / or parameters can be coded individually or included in a mode codeword that defines the prediction mode being used. Which codeword to use for a given mode / sub - mode and / or parameter combination can affect the coding efficiency gain through intra prediction, and the same can be true for the entropy coding technology used to convert the codeword into a bitstream.
[0010] Specific modes of intra prediction were introduced in H.264, scrutinized in H.265, and further scrutinized in newer coding technologies such as the joint exploration model (JEM), versatile video coding (VVC), and the benchmark set (BMS). A predictor block can be formed using adjacent sample values of already available samples. The sample values of the adjacent samples are copied into the predictor block according to a direction. The reference to the direction in use can be coded in the bitstream or can itself be predicted.
[0011] Referring to FIG. 1A, shown at the lower right is a subset of nine predictor directions, out of the 33 possible predictor directions (corresponding to the 33 angular modes of the 35 intra modes) defined in H.265. The point (101) where the arrows converge represents the predicted sample. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples at an angle of 45 degrees from horizontal, to the upper right. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples at an angle of 22.5 degrees from horizontal, to the lower left of sample (101).
[0012] Continuing to refer to FIG. 1A, shown at the upper left is a square block (104) of 4×4 samples (shown by the thick dashed line). The square block (104) contains 16 samples, each labeled with an "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample of block (104) in both the Y and X dimensions. Since the block size is 4×4 samples, S44 is at the lower right. Further, reference samples following a similar numbering scheme are shown. The reference samples are labeled with an R for the block (104), its Y position (e.g., row index), and X position (column index). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, and thus negative values need not be used.
[0013] Intra picture prediction can function by copying the reference sample value from adjacent samples as indicated by the predicted direction signaled. For example, assume that the coded video bitstream includes signaling indicating a prediction direction that coincides with arrow (102) for this block, i.e., the samples are predicted from the predicted sample at a 45-degree angle from horizontal going up and to the right. In this case, samples S41, S32, S23, S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.
[0014] In some cases, especially when the direction is not evenly divisible by 45 degrees, the values of multiple reference samples can be combined, for example, through interpolation, to calculate the reference sample.
[0015] As video coding technology has evolved, the number of possible directions has increased. In H.264 (2003), nine different directions can be represented. This has increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments are conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those most likely directions with a small number of bits, accepting a specific penalty for less likely directions. Additionally, sometimes the direction itself can be predicted from the adjacent directions used in adjacent already decoded blocks.
[0016] Figure 1B shows a schematic (110) indicating 65 intra prediction directions according to JEM, exemplifying the increase in the number of prediction directions over time.
[0017] The mapping of intra prediction direction bits representing directions within a coded video bitstream may vary for each video coding technology. Such mapping can range from, for example, simple direct mapping to codewords, to complex adaptive schemes including the most probable mode and similar techniques. However, in most cases, there may be certain directions that are statistically less likely to occur within video content than certain other directions. Since the goal of video compression is redundancy reduction, those less likely directions are represented by a larger number of bits in well-functioning video coding technologies than the more likely directions.
[0018] Image and / or video coding and decoding can be performed using inter-picture prediction with motion compensation. Motion compensation can be a non-reversible compression technique, and a block of sample data from a previously reconstructed picture or a part thereof (reference picture) may be spatially shifted in the direction indicated by a motion vector (hereinafter, MV) and then used for prediction of a newly reconstructed picture or a part of the picture. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y or three dimensions, and the third dimension is an indication of the reference picture in use (the latter may indirectly be the temporal dimension).
[0019] In some video compression techniques, the motion vectors (MVs) applicable to an area of sample data can be predicted from other MVs, for example, from an MV related to another area of sample data that is spatially adjacent to the area being reconstructed and that precedes that MV in decoding order. Doing so can substantially reduce the amount of data required to code the MVs, thereby removing redundancy and increasing compression. MV prediction can function effectively because, for example, when coding an input video signal (known as natural video) derived from a camera, there is a statistical likelihood that areas larger than the area to which a single MV is applicable will move in a similar direction, and thus, in some cases, prediction can be made using a similar motion vector derived from the MVs of adjacent areas. As a result, the MVs found for a given area are similar or identical to the MVs predicted from surrounding MVs, and this can then be represented in a smaller number of bits than would be used if the MVs were coded directly after entropy coding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when calculating the predictor from several surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms proposed by H.265, the one described in relation to FIG. 2 is a technique called "spatial merge".
[0021] Referring to FIG. 2, the current block (201) contains samples found to be predictable from the previous block of the same size that has been spatially shifted during the motion search process by the encoder. Instead of directly coding the MV, the MV can be derived using an MV associated with any one of five surrounding samples shown as A0, A1 and B0, B1, B2 (202 to 206 respectively) from the metadata associated with one or more reference pictures, for example from the latest reference picture (in decoding order). In H.265, MV prediction can use predictors from the same reference picture as that used by adjacent blocks. Summary of the Invention
[0022] Aspects of the present disclosure provide methods and apparatuses for video / image encoding and / or decoding. In some embodiments, an apparatus for video / image decoding includes a processing circuit. The processing circuit decodes prediction information of a current block in a current picture from a coded bitstream. The prediction information indicates that an intra prediction mode using multiple reference line (MRL) prediction is applied to the current block. The current block is predicted based on a first reference line and a second reference line. Each piece of weight information among a plurality of pieces of weight information indicates a respective first weight candidate of the first reference line and a respective second weight candidate of the second reference line. For each piece of weight information, the processing circuit predicts a subset of samples in the current block using intra prediction fusion based on the first reference line, the second reference line, and the respective piece of weight information. The subset of samples includes one of (i) upper samples in the uppermost row (also referred to as the top row) in the current block and (ii) left samples in the leftmost column in the current block. The processing circuit determines a gradient cost based on the predicted subset of samples in the current block and reconstructed samples outside the current block. The reconstructed samples outside the current block include samples adjacent to the predicted subset of samples in the current block. The processing circuit selects weight information from among the plurality of pieces of weight information based on the determined gradient cost corresponding to each of the plurality of pieces of weight information. The selected weight information indicates a first weight for the first reference line and a second weight for the second reference line.
[0023] In one example, the first reference line includes first reference samples that are N1 rows or N1 columns away from the current block. The second reference line includes second reference samples that are N2 rows or N2 columns away from the current block. N1 and N2 are different integers greater than or equal to 0.
[0024] In one example, for each of a plurality of weight information and one sample of a subset of samples, the processing circuit uses an intra prediction mode to determine a first predicted value based on one or more first reference samples of a first reference line. The processing circuit uses an intra prediction mode to determine a second predicted value based on one or more second reference samples of a second reference line. The processing circuit predicts one sample of the subset of samples based on the first predicted value corrected by a first weight, the second predicted value corrected by a second weight, and the residual of one sample of the subset of samples.
[0025] In one example, a subset of samples within the current block includes the upper samples of the uppermost row within the current block and the left samples of the leftmost column within the current block.
[0026] In one example, the reconstructed samples outside the current block include the reconstructed samples that are not adjacent to the predicted subset of samples.
[0027] The processing circuit can determine the weight information such that, among the plurality of weight information, the weight information corresponds to the minimum gradient cost among the determined gradient costs.
[0028] In one example, the processing circuit rearranges the plurality of weight information based on the determined gradient costs, and determines the weight information based on an index and the rearranged plurality of weight information.
[0029] The index can be signaled in high-level syntax.
[0030] In one embodiment, the processing circuit receives a coded bitstream including a current block within a current picture. The processing circuit obtains and decodes prediction information from the coded bitstream indicating whether an intra prediction mode using multi-reference line (MRL) prediction is applied to the current block. The current block is predicted based on a first reference line and a second reference line, a first weight candidate is applied to the first reference line, and a second weight candidate is applied to the second reference line. The processing circuit obtains a plurality of weight candidate combinations from the coded bitstream. Each weight candidate combination includes a respective first weight candidate and a respective second weight candidate. For each weight candidate combination, the processing circuit uses intra prediction fusion based on the first reference line weighted by the respective first weight candidate and the second reference line weighted by the respective second weight candidate to predict a subset of samples within the current block, the subset of samples including (i) upper samples in the uppermost row within the current block and (ii) left samples in the leftmost column within the current block. The processing circuit determines a gradient cost based on the predicted subset of samples within the current block and reconstructed samples outside the current block. The reconstructed samples outside the current block include samples adjacent to the predicted subset of samples within the current block. The processing circuit selects a weight candidate combination based on the determined gradient cost corresponding to each weight candidate combination.
[0031] In one embodiment, the processing circuit receives an MRL index i from the coded bitstream indicating that the i-th entry in the MRL list corresponds to the first reference line. The processing circuit determines that the (i + 1)-th entry in the MRL list corresponds to the second reference line and reconstructs samples within the current block using intra prediction fusion based on a plurality of reference lines of the current block. In one example, the first reference line and the second reference line among the plurality of reference lines are not spatially adjacent.
[0032] In one example, the processing circuit uses an intra prediction mode to determine a first predicted value based on one or more first reference samples of a first reference line. The processing circuit uses an intra prediction mode to determine a second predicted value based on one or more second reference samples of a second reference line. The processing circuit predicts a sample based on a weighted average of the first predicted value and the second predicted value.
[0033] In one example, the processing circuit reconstructs a sample based on the predicted sample and the residual of the sample.
[0034] In one example, the MRL list is {1, 3, 5, 7, 12} corresponding to reference lines 1, 3, 5, 7, and 12 respectively. Each of reference lines 1, 3, 5, 7, and 12 is 1, 3, 5, 7, and 12 rows and / or columns away from the current block respectively. The MRL index i which is 0, 2, 3, or 4 corresponds to reference line 1, 5, 7, or 12 respectively. When the MRL index i is 0, the first reference line and the second reference line are reference lines 1 and 3. When the MRL index i is 2, the first reference line and the second reference line are reference lines 5 and 7. When the MRL index i is 3, the first reference line and the second reference line are reference lines 7 and 12. When the MRL index i is 4, the first reference line and the second reference line are reference lines 12 and 1.
[0035] In one example, the first weight associated with the first reference line and the second weight associated with the second reference line depend on the first reference line and the second reference line. Based on the first reference line, the second reference line, the first weight, and the second weight, the samples within the current block can be reconstructed.
[0036] In one example, the first weight associated with the first reference line and the second weight associated with the second reference line depend on the distance between the first reference line and the second reference line. The distance can be the number of rows and / or columns between the first reference line and the second reference line.
[0037] In one example, the first weight associated with the first reference line depends on the distance between the first reference line and the reference line 0 adjacent to the current block. The distance is proportional to the MRL index i.
[0038] Aspects of the present disclosure also provide a non - transitory computer - readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to execute a method for video decoding.
Brief Description of the Drawings
[0039] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
[0040]
Figure 1A
[0041]
Figure 1B
[0042]
Figure 2
[0043]
Figure 3
[0044]
Figure 4
[0045]
Figure 5
[0046]
Figure 6
[0047]
Figure 7
[0048]
Figure 8
[0049]
Figure 9
[0050]
Figure 10
[0051]
Figure 11
[0052]
Figure 12
[0053]
Figure 13A
[0054]
Figure 13B
[0055]
Figure 14
[0056]
Figure 15
[0057]
Figure 16
[0058]
Figure 17
[0059]
Figure 18A
[0060]
Figure 18B
[0061]
Figure 19
[0062]
Figure 20
[0063]
Figure 21
DETAILED DESCRIPTION OF THE INVENTION
[0064] FIG. 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission can be common in media delivery applications and the like.
[0065] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of coded video data, for example, during a video conference. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by that terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to restore the video pictures, and display the video pictures on an accessible display device according to the restored video data.
[0066] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) can each be exemplified as a server, a personal computer, and a smartphone, but the principles of the present disclosure cannot be so limited. Embodiments of the present disclosure find application in laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) represents any number of networks that transmit coded video data among the terminal devices (310), (320), (330), and (340), including, for example, wireline (wired) and / or wireless communication networks. The communication network (350) can exchange data in circuit-switch and / or packet-switch channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described herein below.
[0067] FIG. 4 shows a video encoder and a video decoder in a streaming environment as an example of the application of the disclosed subject matter. The disclosed subject matter may be similarly applicable to other video-enabled applications, including, for example, video conferencing, storage of compressed video in digital media such as digital TV and streaming services, CDs, DVDs, memory sticks, etc.
[0068] A streaming system may include a video source (401) that creates a stream of video pictures (402), for example, uncompressed video pictures, and may include a video capture subsystem (413), which may include, for example, a digital camera. In one example, the stream of video pictures (402) includes samples taken by a digital camera. The stream of video pictures (402) is shown as a thick line to emphasize the high data volume when compared to the encoded video data (404) (or coded video bitstream) and can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or encoded video bitstream) is shown as a thin line to emphasize the low data volume when compared to the stream of video pictures (402) and can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and creates an output stream of video pictures (411) that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, the video coding standard under development is informally known as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC.
[0069] Note that electronic devices (420) and (430) can include other components (not shown). For example, electronic device (420) can include a video decoder (not shown), and electronic device (430) can also include a video encoder (not shown).
[0070] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). Instead of the video decoder (410) of the example of FIG. 4, the video decoder (510) can be used.
[0071] Receiver (531) may receive one or more coded video sequences to be decoded by video decoder (510). In one embodiment, one coded video sequence is received at a time, in which case the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequence may be received from channel (501), which may be a hardware / software link to a storage device storing the encoded video data. Receiver (531) may receive the encoded video data together with other data, such as coded audio data and / or auxiliary data streams, and these data may be transferred to their respective usage entities (not shown). Receiver (531) may separate the coded video sequence from other data. To eliminate network jitter, buffer memory (515) may be coupled between receiver (531) and entropy decoder / parser (520) (hereinafter, "parser (520)"). In certain applications, buffer memory (515) is part of video decoder (510). Otherwise, buffer memory (515) may be external to video decoder (510) (not shown). Still otherwise, there may be a buffer memory (not shown) external to video decoder (510), for example, to eliminate network jitter, in addition to another buffer memory (515) inside video decoder (510) for processing playback timing. When receiver (531) is receiving data from a store / forward device with sufficient bandwidth and controllability or from an isosynchronous network, buffer memory (515) may not be necessary or may be made small.For use in a best effort packet network such as the Internet, buffer memory (515) may be required, the size of which can be relatively large, advantageously an adaptable size, and may be implemented at least in part in an operating system external to the video decoder (510) or a similar element (not shown).
[0072] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510) and, potentially, information for controlling a rendering device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but can be coupled to the electronic device (530) as shown in FIG. 5. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) may syntax analyze / entropy decode the received, coded video sequence. The coding of the coded video sequence can follow a video coding technology or standard and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) can extract a set of subgroup parameters for at least one subgroup of a subgroup of pixels within the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups can include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (520) may also extract from coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.
[0073] The parser (520) may perform an entropy decoding / syntax analysis operation on the video sequence received from the buffer memory (515) so as to generate symbols (521).
[0074] The reconstruction of the symbols (521) can involve multiple different units depending on the type of the coded video picture or a part thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How each unit is involved can be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. Such a flow of subgroup control information between the parser (520) and the multiple units below is not shown for clarity.
[0075] In addition to the functional blocks already described, the video decoder (510) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, a conceptual subdivision into functional units is appropriate below.
[0076] The first unit is a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives not only control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. as symbols (521) from the parser (520), but also quantized transform coefficients. The scaler / inverse transform unit (551) can output a block including sample values that can be input to an aggregator (555).
[0077] In some cases, the output samples of the scaler / inverse transform unit (551) may be related to the intra-coded blocks. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the surrounding already reconstructed information fetched from the current picture buffer (558) to generate blocks of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (555) may, in some cases, add, for each sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0078] In other cases, the output samples of the scaler / inverse transform unit (551) may be related to blocks that are inter-coded and potentially motion-compensated. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. According to the symbols (521) related to the block, after motion-compensating the fetched samples, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (in this case, called the residual samples or residual signal) to generate output sample information. The address in the reference picture memory (557) where the motion compensation prediction unit (553) fetches the prediction samples can be controlled by the motion vectors available to the motion compensation prediction unit (553) in the form of symbols (521) that can have, for example, X, Y and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (557) when exact sub-sample motion vectors are used, a motion vector prediction mechanism, etc.
[0079] The output samples of the aggregator (555) can be subject to various loop filtering techniques within the loop filter unit (556). The video compression technology is controlled by the parameters included in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), and can include in-loop filter techniques. Video compression can also respond to meta information obtained during the decoding of the previous part of the coded picture or coded video sequence (in decoding order), and can respond to previously reconstructed and loop-filtered sample values.
[0080] The output of the loop filter unit (556) can be output to the render device (512) and can be stored in the reference picture memory (557) for use in future inter-picture prediction, and can be a sample stream.
[0081] Once a particular coded picture is completely reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and the fresh current picture buffer can be reallocated before starting the reconstruction of subsequent coded pictures.
[0082] The video decoder (510) may perform a decoding operation according to a predetermined video compression technique or a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used in the sense that the coded video sequence adheres to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile can select specific tools as the only tools available for use under that profile. Also, it may be necessary for compliance that the complexity of the coded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level restricts, for example, the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured, for example, in megasamples per second), the maximum reference pixel size, etc. The limits set by the level may, in some cases, be further restricted through the HRD specifications and metadata for buffer management of a Hypothetical Reference Decoder (HRD) signaled in the coded video sequence.
[0083] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0084] FIG. 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). Instead of the video encoder (403) of the example of FIG. 4, the video encoder (603) can be used.
[0085] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture a video image to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).
[0086] The video source (601) can provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 YCrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media supply system, the video source (601) can be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures that convey motion when viewed in sequence. The picture itself can be organized as a spatial array of pixels, in which case each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0087] According to one embodiment, a video encoder (603) can code and compress pictures of a source video sequence in real time or under any other required time constraints to produce a coded video sequence (643). Implementing an appropriate coding speed is one function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units as described below. This coupling is not shown for clarity. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, ...), picture size, layout of group of pictures (GOP), maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) that are optimized for a particular system design.
[0088] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As a grossly simplified explanation, in one example, the coding loop can include a source coder (630) (which is responsible for creating symbols such as a symbol stream based on the input picture and reference pictures to be coded) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in the same way as a (remote) decoder also does. The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream yields bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as reference picture samples that the decoder "sees" when using the prediction during decoding. This basic principle of the synchronicity of the reference picture (and the resulting drift if the synchronicity cannot be maintained, for example due to channel errors) is also used in several related technologies.
[0089] The operation of the "local" decoder (633) can be the same as that of a "remote" decoder such as the video decoder (510), which has already been described above in relation to FIG. 5. However, referring briefly to FIG. 5 as well, since the symbols are available and the encoding / decoding of the symbols to the coded video sequence by the entropy encoder (645) and the parser (520) can be reversible, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).
[0090] In one embodiment, any decoder technology other than parsing / entropy decoding present in the decoder exists in the corresponding encoder in the same or substantially the same functional form. Accordingly, the disclosed subject matter focuses on the operation of the decoder. Since the description of encoder technology is the opposite of the decoder technology described comprehensively, it may be omitted. In certain areas, more detailed descriptions are provided below.
[0091] In some examples, during operation, the source coder (630) may perform motion compensation predictive coding, which predictively codes an input picture in relation to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.
[0092] The local video decoder (633) may decode the coded video data of a picture that can be designated as a reference picture based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a non-invertible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that can be performed by the video decoder for the reference picture and store the reconstructed reference picture in the reference picture memory (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture obtained by the remote video decoder (without transmission errors).
[0093] Predictor (635) may perform predictive search for the coding engine (632). That is, for a new picture to be coded, predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc., which may function as appropriate prediction references for the new picture. Predictor (635) may operate on a sample block-by-pixel block basis to find an appropriate prediction reference. In some cases, the input picture may have prediction references drawn from a plurality of reference pictures stored in the reference picture memory (634), as determined by the search results obtained by predictor (635).
[0094] Controller (650) may manage the coding operations of source coder (630), including the setting of parameters and subgroup parameters used, for example, to encode video data.
[0095] Outputs of all of the foregoing functional units may be subject to entropy coding in entropy coder (645). Entropy coder (645) converts symbols generated by various functional units into a coded video sequence by applying reversible compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0096] The transmitter (640) may buffer the coded video sequence created by the entropy coder (645) to prepare for transmission via the communication channel (660), which may be a hardware / software link to a storage device storing the encoded video data. The transmitter (640) may merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).
[0097] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a specific coded picture type to each coded picture, which may affect the coding applicable to each picture. For example, a picture may often be assigned as one of the following picture types:
[0098] An intra picture (I picture) may be coded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example independent decoder refresh (IDR) pictures. Those skilled in the art are aware of these variations of I pictures and their respective uses and characteristics.
[0099] A predicted picture (P picture) may be coded and decoded using intra prediction or inter prediction, using at most one motion vector and a reference index, to predict the sample values of each block.
[0100] A bi - directional predicted picture (B picture) can be coded and decoded using intra - prediction or inter - prediction with up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple - predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0101] A source picture is typically subdivided spatially into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block - by - block. Blocks can be coded predictively in relation to other (already - coded) blocks as determined by the coding assignment applied to each block of the picture. For example, blocks of an I picture may be coded non - predictively, or they may be coded predictively in relation to already - coded blocks of the same picture (spatial prediction or intra - prediction). Pixel blocks of a P picture can be coded predictively via spatial prediction or temporal prediction in relation to one previously - coded reference picture. Blocks of a B picture can be coded predictively via spatial prediction or temporal prediction in relation to one or two previously - coded reference pictures.
[0102] A video encoder (603) can perform coding operations according to a given video coding technology or standard such as ITU - T Rec. H.265. In its operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. The coded video data can thus conform to the syntax specified by the video coding technology or standard being used.
[0103] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0104] Video can be captured as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra prediction) uses spatial correlations within a given picture, and inter-picture prediction uses (temporal or other) correlations between pictures. In one example, a particular picture during encoding / decoding is called the current picture and is partitioned into blocks. When a block within the current picture is similar to a reference block within a reference picture that has been previously coded and is still buffered within the video, the block within the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture in cases where multiple reference pictures are being used.
[0105] In some embodiments, dual prediction techniques can be used in inter-picture prediction. According to the dual prediction technique, two reference pictures, such as a first reference picture and a second reference picture, both of which precede the current picture within the video in decoding order (although in display order they can be in the past and future respectively), are used. A block within the current picture can be coded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.
[0106] Furthermore, in order to improve coding efficiency, a merge mode technique can be used in inter-picture prediction.
[0107] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree split into one or more coding units (CUs). For example, a 64×64 pixel CTU can be split into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is split into one or more prediction units (PUs) depending on the temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block may include a matrix of values (e.g., luma values) for pixels, such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0108] FIG. 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) receives a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures, and is configured to encode the processing block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used instead of the video encoder (403) of the example of FIG. 4.
[0109] In an example of HEVC, the video encoder (703) receives a matrix of sample values of a processing block such as a prediction block of 8×8 samples. The video encoder (703) determines whether the processing block is best coded using an intra mode, an inter mode, or a bi-prediction mode, for example, using rate-distortion optimization. When the processing block is to be coded in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into a coded picture, and when the processing block is to be coded in the inter mode or the bi-prediction mode, the video encoder (703) may use inter prediction techniques or bi-prediction techniques, respectively, to encode the processing block into a coded picture. In certain video coding techniques, the merge mode can be an inter-picture prediction sub-mode when the motion vector is derived from one or more motion vector predictors without the benefit of the coded motion vector components outside the predictor. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components such as a mode decision module (not shown) to determine the mode of the processing block.
[0110] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), which are coupled together as shown in FIG. 7.
[0111] The inter-encoder (730) receives samples of the current block (e.g., a processing block), compares the block with one or more reference blocks (e.g., blocks in the previous picture and the subsequent picture) in a reference picture, generates inter-prediction information (e.g., a description of redundant information according to an inter-coding technique, a motion vector, merge mode information), and is configured to calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.
[0112] The intra-encoder (722) receives samples of the current block (e.g., a processing block), optionally compares the block with already-coded blocks in the same picture, generates quantized coefficients after transformation, and is optionally also configured to generate intra-prediction information (e.g., intra-prediction direction information according to one or more intra-coding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and a reference block in the same picture.
[0113] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the mode of a block and provides a control signal to the switch (726) based on the mode. For example, when the mode is the intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream, and when the mode is the inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.
[0114] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result, selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) operates based on the residual data and is configured to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients are then subject to quantization processing to obtain the quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation to generate the decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.
[0115] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to an appropriate standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. Note that there is no residual information when coding a block in either the merge submode of the inter mode or the bi-prediction mode according to the disclosed subject matter.
[0116] FIG. 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a reconstructed picture. In one example, the video decoder (810) is used instead of the video decoder (410) of the example of FIG. 4.
[0117] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) that are coupled together as shown in FIG. 8.
[0118] The entropy decoder (871) can be configured to reconstruct from the coded picture specific symbols that represent syntax elements from which the coded picture is composed. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi-prediction mode, two latter merge sub-modes, or another sub-mode, etc.), and prediction information (e.g., intra prediction information or inter prediction information, etc.) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), respectively. The symbols can also include, for example, residual information in the form of quantized transform coefficients. In one example, when the prediction mode is inter mode or bi-prediction mode, the inter prediction information is provided to the inter decoder (880), and when the prediction mode is intra prediction mode, the intra prediction information is provided to the intra decoder (872). The residual information can be subject to inverse quantization and is provided to the residual decoder (873).
[0119] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.
[0120] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0121] The residual decoder (873) may be configured to perform inverse quantization to extract de-quantized transform coefficients, and process the de-quantized transform coefficients to convert residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including quantizer parameters (QP)), which may be provided by the entropy decoder (871) (since this may be only a small amount of control information, the data path is not shown).
[0122] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information as the output by the residual decoder (873) and the prediction result (optionally, as the output by the inter or intra prediction module) to form a reconstruction block, which may be part of a reconstruction picture that may be part of the reconstructed video. Note that other appropriate operations, such as a deblocking operation, can be performed to improve visual quality.
[0123] Note that the video encoders (403), (603) and (703), and the video decoders (410), (510) and (810) can be implemented using any appropriate technology. In one embodiment, the video encoders (403), (603) and (703), and the video decoders (410), (510) and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603) and (703), and the video decoders (410), (510) and (810) can be implemented using one or more processors that execute software instructions.
[0124] In VVC, various inter-prediction modes can be used. For an inter-prediction CU, the motion parameters can include an MV, one or more reference picture indexes, a reference picture list utilization index, and additional information on specific coding features used for inter-prediction sample generation. The motion parameters can be signaled either explicitly or implicitly. When a CU is coded in skip mode, the CU may be associated with a PU and may not have significant residual coefficients, a coded motion vector delta or MV difference (e.g., MVD) or a reference picture index. The motion parameters for the current CU can specify a merge mode obtained from adjacent CUs that include spatial and / or temporal candidates and optionally additional information as introduced in VVC. The merge mode can be applied not only to skip mode but also to inter-predicted CUs. In one example, an alternative to the merge mode is the explicit transmission of motion parameters, in which case the corresponding reference picture indexes and other information of the MV, each reference picture list, and the reference picture list use flag are signaled explicitly for each CU.
[0125] In embodiments such as VVC, the VVC Test Model (VTM) reference software includes one or more improved inter prediction coding tools, including extended merge prediction, merge with motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8×8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. Inter prediction and related methods are described in detail below.
[0126] In some examples, extended merge prediction can be used. In an example such as VTM4, the merge candidate list includes the following five types of candidates: namely, the spatial motion vector predictor (MVP) from adjacent spatial CUs, the temporal MVP from collocated CUs, the history-based MVP (HMVP) from the first-in first-out (FIFO) table, the pairwise-average MVP, and the zero MV, in this order.
[0127] The size of the merge candidate list can be signaled in the slice header. In one example, the maximum allowable size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate (e.g., the merge index) can be coded using truncation-unary binarization (TU). The first bin of the merge index can be coded in a context (e.g., using context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.
[0128] Some examples of the generation process for each category of merge candidates are provided below. In one embodiment, the spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC can be the same as that in HEVC. In one example, up to four merge candidates are selected from the candidates located at the positions shown in FIG. 9. FIG. 9 shows the positions of the spatial merge candidates according to an embodiment of the present invention. Referring to FIG. 9, the order of derivation is B1, A1, B0, A0, and then B2. The position B2 is considered only when any of the CUs at positions A0, B0, B1, and A1 are not available (e.g., because the CU belongs to another slice or another tile), or when it is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.
[0129] To reduce the computational complexity, not all possible candidate pairs are necessarily considered in the aforementioned redundancy check. Instead, only the pairs linked by the arrows in FIG. 10 are considered, and a candidate is added to the candidate list only if the corresponding candidate used for the redundancy check does not have the same motion information. FIG. 10 shows the candidate pairs considered for the redundancy check of the spatial merge candidates according to an embodiment of the present disclosure. Referring to FIG. 10, the pairs linked by each arrow include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, the candidates at positions B1, A0, and / or B2 can be compared with the candidates at position A1, and the candidates at positions B0 and / or B2 can be compared with the candidates at position B1.
[0130] In one embodiment, the temporal candidates are derived as follows. In one example, only one temporal merge candidate is added to the candidate list. FIG. 11 shows an exemplary motion vector scaling for the temporal merge candidates. To derive the temporal merge candidates for the current CU (1111) in the current picture (1101), the scaled MV (1121) (e.g., shown by the dotted line in FIG. 11) can be derived based on the collocated CU (1112) belonging to the collocated reference picture (1104). The reference picture list used to derive the collocated CU (1112) can be explicitly signaled in the slice header. As shown by the dotted line in FIG. 11, the scaled MV (1121) of the temporal merge candidate can be obtained. The scaled MV (1121) can be scaled from the MV of the collocated CU (1112) using the picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (1102) of the current picture (1101) and the current picture (1101). The POC distance td can be defined as the POC difference between the collocated reference picture (1104) of the collocated picture (1103) and the collocated picture (1103). The reference picture index of the temporal merge candidate can be set to zero.
[0131] FIG. 12 shows exemplary candidate positions (e.g., C0 and C1) for the temporal merge candidates of the current CU. The position of the temporal merge candidate can be selected between the candidate positions C0 and C1. The candidate position C0 is located at the lower right corner of the collocated CU (1210) of the current CU. The candidate position C1 is located at the center of the collocated CU (1210) of the current CU. If the CU at the candidate position C0 is not available, is intra-coded, or is outside the current row of the CTU, the candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at the candidate position C0 is available, is intra-coded, is in the current row of the CTU, the candidate position C0 is used to derive the temporal merge candidate.
[0132] In some examples such as VVC, a quadtree with nested multi-type tree (QTMTT) structure having nested multi-type trees is used to enable further flexibility in CTU splitting. In the QTMTT structure, a CTU can be split by a recursive quadtree (QT) and a multi-type tree (MTT) splitting structure. The QT structure can use the same concept as that used in, for example, HEVC, where the current CU can be divided into four smaller secondary CUs. The MTT structure can split a CU (e.g., a rectangular-shaped CU) by a binary tree (BT) and a ternary tree (TT). The BT split can divide the current CU into two symmetric compartments. The TT split can divide the current CU into three compartments, which can include a central compartment having 1 / 2 and 1 / 4 of the original CU size and two side compartments, respectively. In an intra-slice such as VVC, a dual-tree split is possible where the luminance and chrominance components within a CTU can be split separately. In one example as in VVC, intra prediction of luminance samples can be performed with rectangular CU sizes in the range from 4×4 to 64×64. To improve coding efficiency, various intra prediction coding tools are introduced, including but not limited to, angular intra prediction using 65 angles and 4-tp interpolation filters, wide-angle intra prediction (WAIP), position dependent prediction combination (PDPC), multiple reference line (MRL) prediction, intra subpartition (ISP) mode, matrix-based intra prediction (MIP), cross component linear model (CCLM), intra mode coding using six most probable modes (MPM), etc.
[0133] For MRL intra prediction, multiple reference lines can be used for intra prediction. FIG. 13A shows an example of MRL intra prediction. In FIG. 13A, four reference lines 0 to 3 of the current block (1301) are shown. The reference line i can include reference samples that are i lines away from the current block (1301), for example, reference samples that are i lines away from the boundary of the current block (1301) (e.g., reference samples that are i lines away from the upper boundary and / or reference samples that are i columns away from the left boundary), where i is 0, 1, 2, or 3. For example, the reference line i can include reference samples that are i rows above the upper boundary of the current block (1301) and / or reference samples that are i columns to the left of the left boundary of the current block (1301). In one example, the reference line 0 includes reference samples adjacent to the current block (1301), such as adjacent reconstructed samples that include the upper adjacent sample above the current block (1301) and the left adjacent sample to the left of the block (1301). In one example, the reference line 0 can include the reconstructed adjacent sample in the upper left corner.
[0134] The reference lines 0 to 3 can include a plurality of segments such as segments A to F. In one example, the samples of segments A and F are not fetched from the reconstructed adjacent samples. The samples of segments A and F can be padded (or filled) with the samples closest to them from segments B and E, respectively.
[0135] In an example such as HEVC, the closest reference line (i.e., reference line 0) is used for intra prediction (or intra picture prediction). In MRL intra prediction, multiple reference lines can be used. In one example, two additional lines (e.g., reference line 1 and reference line 3) are used.
[0136] An index (e.g., a reference line index denoted as mrl_idx) used to select a reference line can be signaled, and the selected reference line can be used to generate an intra predictor for the current block (1301). For a reference line index greater than 0, only additional reference line modes can be included in the MPM list, and only MPM indexes that do not include the remaining modes (e.g., intra prediction modes not included in the MPM list) can be signaled. An index can be signaled before the intra prediction mode. In one example, when a non-zero reference line index is signaled, certain intra prediction modes (e.g., the planar mode and / or the DC mode) are excluded from the intra prediction modes.
[0137] MRL intra prediction is extended to include more reference lines for intra prediction, for example, in the Enhanced Compression Model 5 (ECM5). FIG. 13B shows an example of multi-reference lines used in MRL intra prediction. For the current block (1301), reference lines 0 to 7 out of reference lines 0 to 12 are shown. Reference lines 0 to 3 in FIG. 13B are described in FIG. 13A. The description of reference lines 0 to 3 can be adapted to reference lines 4 to 12. In this case, the reference sample of reference line i is i lines away from the current block (1400) as described in FIG. 13A, where i can be 4 to 12. MRL intra prediction can use fewer or more reference lines than reference lines 0 to 12. MRL intra prediction can use spatially adjacent reference lines (e.g., reference lines 1 and 2) and / or non-spatially adjacent reference lines (e.g., reference lines 1 and 3).
[0138] An MRL list, such as an extended reference line list, can include indices of reference lines (e.g., mrl_idx). In one example, the extended reference line list includes one or more reference lines different from reference lines 0 to 3 shown in FIG. 13B. In one example, the MRL list is {1, 3, 5, 7, 12}, where the indices 1, 3, 5, 7, and 12 indicate reference lines 1, 3, 5, 7, and 12, respectively. The MRL list of {1, 3, 5, 7, 12} is an extended reference line list.
[0139] Template-based intra mode derivation (TIMD) can use the reference samples of the current CU as a template and select the best intra mode from a set of candidate intra prediction modes associated with TIMD. In TIMD, instead of a complete MRL list (e.g., {1, 3, 5, 7, 12}), a subset of the complete MRL list is used. In one example, the first two reference line candidates, e.g., {1, 3}, are used.
[0140] In intra prediction or an intra prediction mode, the sample values of a coding block can be predicted from already reconstructed samples (referred to as reference samples). The samples can exist within one or more reference lines.
[0141] An example of intra prediction is the planar mode that can use bilinear interpolation. In the planar mode, one or more positions within the current block can be predicted using the reference samples within the reference lines. Other positions within the current block can be predicted as a linear combination of the samples at the one or more positions and the reference samples. The weights can be determined according to the position of the current sample within the current block.
[0142] An example of intra prediction is the DC mode. To predict the samples within a block in the DC mode, the average of the samples within the reference line can be used as a predictor.
[0143] An example of intra prediction is angular intra prediction. In angular intra prediction, a current sample within a current block can be predicted using a reference sample (e.g., a predicted sample), or for example, an interpolated reference sample within a reference line.
[0144] FIG. 14 shows an example of intra prediction using a reference line. The reference line can be a reference line 0 adjacent to a block (1400), or a reference line not adjacent to the block (e.g., reference line 1). A part of reference lines 0 to 1 is shown. A sample (1401) within the block (1400) is predicted using one or more samples of reference line 1 having an intra predictor direction (or angular direction) (1412). The sample (1401) is projected onto reference line 1 along the angular direction (1412).
[0145] In the example shown in FIG. 14, the projected position (1413) of the sample (1401) is located between two reference samples (1411) to (1412) of reference line 1 and is called the projected fractional position. The sample (1401) can be predicted using the two reference samples (1411) to (1412) of reference line 1. An interpolation filter such as a 2-tap linear interpolation filter can be used to generate the prediction of the sample (1401). The filter coefficients can be inversely proportional to the two distances between the projected fractional position (1413) and the two adjacent integer positions (indicated by black dots) of the two reference samples (1411) to (1412), respectively.
[0146] In one example, the projected position of the sample (1401) is located at the integer position of the reference sample of reference line 1. Therefore, the reference sample of reference line 1 is used as the prediction of the sample (1401) and no interpolation is required.
[0147] Intra prediction fusion can be used to determine the intra-predicted samples of the samples within a block using two reference lines (e.g., a first reference line and a second reference line), as shown in Equation 1. p fusion = w0p line + w1p line+1 (Equation 1)
[0148] Parameter p line represents the first prediction based on the first reference line, and parameter p line+1 represents the second prediction based on the second reference line. Each of the first prediction and the second prediction can be predicted using an intra prediction mode and its respective reference line, such as the angular intra prediction shown in FIG. 14. In the example shown in Equation 1, the two reference lines (e.g., a first reference line and a second reference line) are spatially adjacent to each other, such as reference lines 1 to 2. The two intra predictions obtained from the respective reference lines are weighted using weights w0 and w1 for the first reference line and the second reference line, respectively.
[0149] In one example, the first reference line is the default reference line, and the second reference line is the reference line above the default reference line. In one example, the weights w0 and w1 are set to 3 / 4 and 1 / 4, respectively.
[0150] Intra prediction fusion can be applied to luma blocks. In one example, when the angular intra mode has a non-integer gradient and the block size (e.g., the number of samples within a luma block) is greater than 16, intra prediction fusion is applied to the luma block. In one example, intra prediction fusion is used together with MRL intra prediction. In one example, intra prediction fusion is not applied to ISP-coded blocks.
[0151] Fixed weights are used in intra prediction fusion using the multi-reference lines as described above. In some examples, the fixed weights may not have good coding efficiency. In the example shown in Equation 1, the reference samples of two spatially adjacent reference lines are stored in a buffer. For example, the extra line buffer requirement for MRL for intra prediction fusion is twice the size of the MRL in some embodiments such as ECM5. For example, if the MRL list is {1, 3, 5, 7, 12}, then the samples of reference lines 1, 3, 5, 7, and 12 are stored. To use the intra prediction fusion of Equation 1, each adjacent line including 2, 4, 6, and / or 8 may be memorized, thus increasing the buffer requirement of the MRL.
[0152] Embodiments of the present disclosure describe embodiments related to weight combination derivation for MRL for intra prediction fusion and a multi-reference line selection method for intra prediction fusion.
[0153] Referring to FIG. 15, the current block (1501) is predicted in an intra prediction mode (e.g., an angular intra prediction mode as described in FIGS. 1A-1B and FIG. 14). According to embodiments of the present disclosure, a plurality of reference lines can be used to predict samples within the current block (1501) using intra prediction fusion with adaptive weights. Based on the adaptive weights, a weighted average of individual predictions based on the plurality of reference lines can be used to predict the samples. Each individual prediction can be based on an intra prediction mode using a reference line. The weights can be adapted based on (i) the gradient or difference between the intra-predicted samples within the current block and (ii) the reconstructed samples outside the current block. In one example, the intra-predicted samples within the current block and the reconstructed samples outside the current block are near the boundary of the current block. In one example, two of the plurality of reference lines are not spatially adjacent.
[0154] In one embodiment, two reference lines (e.g., a first reference line and a second reference line) are used in intra prediction fusion. Multiple weight information is available for the current block. In one example, the multiple weight information includes multiple weight candidate combinations. Each weight information among the multiple weight information can indicate a first weight (also called a first weight candidate) (e.g., w0) and a second weight (also called a second weight candidate) (e.g., w1) of the first reference line and the second reference line of the current block, respectively. Each weight information can indicate a weight candidate combination such as each first weight candidate and each second weight candidate. In one example, the first weight and the second weight are related. For example, the sum of the first weight and the second weight is 1. Therefore, when one of the first weight and the second weight is known, the other of the first weight and the second weight is also known.
[0155] For each weight information among the multiple weight information, a subset (1511) of samples within the current block (1501) can be predicted, for example, using intra prediction fusion based on the first reference line, the second reference line, and the respective weight information, as described in, for example, Equation 2. p fusion = w0p1 + w1p2 (Equation 2)
[0156] The parameter p1 represents a first prediction based on the first reference line, and the parameter 2 represents a second prediction based on the second reference line. As an example, the first reference line (e.g., reference line 1) and the second reference line (e.g., reference line 3) are not adjacent and are different from Equation 1. Each of the first prediction and the second prediction can be predicted using an intra prediction mode and the respective reference lines, such as the angular intra prediction shown in FIG. 14. The weights w0 and w1 for the first reference line and the second reference line, respectively, are described with reference to Equation 1.
[0157] In one example, residual data of a subset (1511) of samples is added to a predicted subset (1511) of samples to determine a reconstructed subset (1511) of samples. The subset (1511) of samples can include (i) an upper sample (1512) of the top row (also referred to as the top-most row) of the current block (1501), and / or (ii) a left sample (1513) of the left-most column of the current block (1501).
[0158] Based on the reconstructed subset (1511) of samples within the current block and an adjacent reconstructed sample (1531) of the current block (1501) corresponding to the reconstructed subset (1511) of samples, a gradient cost can be determined. The adjacent reconstructed sample (1531) can include (i) an upper adjacent sample (1532) and / or (ii) a left adjacent sample (1533). In one example, the adjacent reconstructed sample (1531) is a reference sample of reference line 0.
[0159] In one example, the gradient cost is determined using Equation 3.
Equation
[0160] Parameter r m,-1 represents the reference sample value of the upper adjacent sample R m,-1 (1532), and parameter p m,0は , the reconstructed sample value of the upper reconstructed sample P m,0 (1512), where the integer m ranges from 0 to W - 1. W can be the width of the current block (1501), and for example, in the example shown in FIG. 15, it can be 8. The upper adjacent sample R m,-1 corresponds to the upper reconstructed sample P m,0 . For example, the upper adjacent sample R m,-1 is directly above the upper reconstructed sample P m,0 .
[0161] Parameter r -1,j represents the reference sample value of the left adjacent sample R -1,j (1533), and parameter p 0,jは represents the reconstructed sample value of the left reconstructed sample P 0,j (1513), where the integer j ranges from 0 to H - 1. H can be the height of the current block (1501), for example, 4 in the example shown in FIG. 15. The left adjacent sample R -1,j corresponds to the left reconstructed sample P 0,j . For example, the left adjacent sample R -1,j is in the left neighborhood of the left reconstructed sample P 0,j .
[0162] The gradient cost of Equation 3 can include a first sum over the upper reconstructed samples (1512) of the top row of the current block (1501) and a second sum over the left reconstructed samples (1513) of the leftmost column of the current block (1501).
[0163] In some examples, the gradient cost does not include the second sum and includes only the first sum. In some examples, the gradient cost does not include the first sum and includes only the second sum. In some examples, the gradient cost includes a sum over selected samples within a subset (1511) of samples.
[0164] Based on the determined gradient cost corresponding to multiple weight information, a piece of weight information can be determined. The determined weight information can be used to reconstruct the current block (1501). For example, when the two reference lines are reference lines 1 and 3, the determined weight information indicates 4 / 5 and 1 / 5 for reference lines 1 and 3 respectively. Each sample within the current block (1501) can be predicted based on Equation 2. The first prediction p1 and the second prediction p2 can be determined using the intra prediction mode.
[0165] In one embodiment, the best combination of the two weights of the two reference lines is determined by calculating the gradient cost (e.g., along the top and left of the CU boundary) between the adjacent reconstruction sample (1531) and the reconstruction subset of the sample (1511). For example, the predefined weight combinations of the two reference lines are predefined in a weight list. The predefined weight combinations can correspond to multiple weight information. Using each predefined weight combination (e.g., 3 / 4 and 1 / 4), the corresponding intra predictor of the subset of the sample (1511) as described above can be generated. The residual corresponding to the intra predictor (e.g., the decoded residual) can be added to the intra predictor to form the reconstruction sample of each weight combination. For each weight combination, as shown in Equation 3, all the gradient values between the adjacent reconstruction sample (1531) along the top and left CU boundaries and the adjacent reconstruction sample (1511) are added together as the gradient cost.
[0166] FIG. 16 shows another example of samples used to determine the gradient cost of the current block (1501). The subset of samples (1511), the top sample (1512), and the left sample (1513) in FIG. 16 are described in FIG. 15.
[0167] Based on the reconstruction subset of the samples (1511) within the current block and the reconstruction samples (1631) outside the current block corresponding to the reconstruction subset of the samples (1511), the gradient cost can be determined. The reconstruction samples (1631) can include the adjacent reconstruction samples (1531) (e.g., (1532) and (1533)) described in FIG. 15 and additional reconstruction samples (1632) and (1633) that are not adjacent to the current block (1501). The reconstruction sample (1632) is adjacent to and above the upper adjacent sample (1532). The reconstruction sample (1633) is adjacent to the left adjacent sample (1533). In the example of FIG. 16, the reconstruction samples (1632) to (1633) are the reference samples of reference line 1. In one example, the gradient cost is determined using Equation 4.
Equation
[0168] Parameter r m,-2 represents the reference sample value of the reconstruction sample R m,-2 (1632), where the integer m is from 0 to W-1. Parameter r -2,j represents the reference sample value of the left reconstruction sample R -2,j (1633), where the integer j is from 0 to H-1.
[0169] The gradient cost corresponding to each weight information (e.g., a predefined weight combination) can be determined based on the reconstruction samples within the current block and the reconstruction samples outside the current block, where the reconstruction samples within the current block are determined based on the respective weight information.
[0170] In one embodiment, the weight information (e.g., weight combination) having the minimum gradient cost is selected as the best weight combination, and the selected weight combination is not signaled in the bitstream.
[0171] In one embodiment, template matching is disabled, for example, by high-level syntax, or template matching is not supported, and thus, the multiple weight information cannot be sorted based on the determined gradient cost. After the weight information is selected from the multiple weight information (e.g., by an encoder), an index indicating the selected weight combination may be signaled.
[0172] In one embodiment, template matching is enabled. The multiple weight information can be sorted based on the determined gradient cost, for example, in ascending order of the determined gradient cost. An index can be signaled at the CU level to indicate which weight information (e.g., which weight combination) among the sorted multiple weight information is used.
[0173] In one embodiment, an index indicating weight information (e.g., a weight combination) may be signaled at a high level, such as a level higher than the CU level. For example, the index may be signaled at the sequence level, picture level, slice level, etc.
[0174] Two or more selected reference lines in the MRL buffer can be used for intra prediction fusion. The MRL buffer can store the reference sample values of the reference lines. The two or more selected reference lines may not be spatially adjacent to each other. For example, referring to FIG. 13B, the two or more selected reference lines include non-adjacent reference line 1 and reference line 5.
[0175] Referring to FIG. 13B, in one example, the reference line in the MRL buffer having index i and the reference line in the MRL buffer having index (i + 1) are utilized in intra prediction fusion. In one example, i is an integer greater than or equal to 0. Index i and index (i + 1) are reference line indices that point to adjacent entries in an MRL list, such as the MRL list of {1, 3, 5, 7, 12}. For example, index i being 0, 1, 2, 3, or 4 corresponds to reference line 1, 3, 5, 7, or 12, respectively. When index i being 0 is signaled, entries 0 and 1 in the MRL list of {1, 3, 5, 7, 12} are used, and thus, reference line 1 and 3 are selected to be used for intra prediction fusion.
[0176] In one example, the MRL buffer stores the reference lines in the MRL list including reference line 1 and 3, and thus, no extra memory is required to perform intra prediction fusion. This is different from the example shown in Equation 1. For example, in Equation 1, two spatially adjacent reference lines (e.g., reference lines 1 - 2) are used, and thus, extra memory is required to store reference line 2 which is not in the MRL list.
[0177] In one example, three or more reference lines are used for intra prediction fusion. A single index i can be signaled, and reference lines having indices i, (i + 1), (i + 2), etc. can be used. When three reference lines are used and an index i of 0 is signaled, reference lines 1, 3, and 5 corresponding to indices 0 to 2 are selected for intra prediction fusion. In addition, each weight information indicates a weight combination of reference lines 1, 3, and 5 such as 4 / 7, 2 / 7, and 1 / 7, respectively.
[0178] In one example, reference line 1 and reference line 3 are used for intra prediction fusion. For example, an index i of 0 is signaled, indicating reference lines 1 and 3.
[0179] In another example, reference line 5 and reference line 7 are used for intra prediction fusion. For example, an index i of 1 is signaled, indicating reference lines 5 and 7.
[0180] In another example, reference line 7 and reference line 12 are used for intra prediction fusion.
[0181] In another example, reference line 12 and reference line 1 are used for intra prediction fusion. For example, an index i of 4 is signaled, indicating reference lines 12 and 1. In one example, the last entry (e.g., 12), the first entry (e.g., 1), and the MRL list (e.g., {1, 3, 5, 7, 12}) are treated as adjacent entries.
[0182] In one embodiment, the weights applied to the selected reference lines depend on which reference lines are selected. For example, the weights applied to the combination of reference line 1 and reference line 3 are different from the weights applied to the combination of reference line 1 and reference line 5.
[0183] In one embodiment, the weight applied to the selected reference line depends on the relative distance of the selected reference line. The relative distance between two reference lines can be the number of lines (e.g., the number of rows and / or columns) between the two reference lines. For example, the relative distance between reference line i1 (e.g., 1) and reference line i2 (e.g., 3) is |i1 - i2| (e.g., 2).
[0184] In one embodiment, the weight applied to the selected reference line depends on the relative distance (e.g., the number of lines) between the selected reference line and a reference line (e.g., reference line 0) adjacent to the current block. For example, the relative distance between reference line i (e.g., 1) and reference line 0 is i (e.g., 1).
[0185] FIG. 17 shows an exemplary flowchart outlining an encoding process (1700) according to an embodiment of the present disclosure. The process (1700) can be used in a video / image encoder. The process (1700) can be executed by an apparatus for video / image coding that can include a processing circuit. In various embodiments, the process (1700) is executed by a processing circuit such as a processing circuit in terminal devices (310), (320), (330), and (340), a processing circuit that performs the functions of a video encoder (e.g., (403), (603), (703)). In some embodiments, the process (1700) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1700). The process starts at (S1701) and proceeds to (S1710).
[0186] In (S1710), for each of the plurality of weight information, a subset of samples within the current block can be predicted using intra prediction fusion based on a first reference line, a second reference line, and the respective weight information. The current block can be encoded in an intra prediction mode using multi-reference line (MRL) prediction. Each weight information can indicate respective first weight candidates for the first reference line and respective second weight candidates for the second reference line.
[0187] In one example, the subset of samples includes at least one of (i) the upper samples of the top row of the current block or (ii) the left samples of the leftmost column of the current block.
[0188] For each weight information, a gradient cost can be determined based on the predicted subset of samples within the current block and the reconstructed samples outside the current block. In one example, the reconstructed samples outside the current block include samples adjacent to the predicted subset of samples within the current block.
[0189] In (S1720), weight information can be selected from among the plurality of weight information based on the determined gradient costs corresponding to each of the plurality of weight information. The selected weight information can indicate a first weight for the first reference line and a second weight for the second reference line. The current block can be encoded using intra prediction fusion based on the first weight and the second weight.
[0190] The process (1700) then proceeds to (S1799) and ends.
[0191] The process (1700) can be appropriately adapted to various scenarios, and accordingly, the steps of the process (1700) can be adjusted. One or more of the steps of the process (1700) can be adapted, omitted, repeated, and / or combined. The process (1700) can be implemented using any appropriate order. Additional steps can be added.
[0192] In one example, the first reference line includes a first reference sample that is N1 rows or N1 columns away from the current block. The second reference line includes a second reference sample that is N2 rows or N2 columns away from the current block. N1 and N2 are different integers greater than or equal to 0.
[0193] In one example, for each piece of weight information among a plurality of pieces of weight information and one sample of a subset of samples, an intra prediction mode is used to determine a first predicted value based on one or more first reference samples of the first reference line. An intra prediction mode is used to determine a second predicted value based on one or more second reference samples of the second reference line. One sample of the subset of samples is predicted based on the first predicted value corrected by the first weight, the second predicted value corrected by the second weight, and the residual of the one sample of the subset of samples.
[0194] In one example, the subset of samples within the current block includes the upper sample of the top row within the current block and the left sample of the leftmost column within the current block.
[0195] In one example, the reconstructed samples outside the current block include reconstructed samples that are not adjacent to the predicted subset of samples.
[0196] In one example, the weight information is determined to be the weight information corresponding to the minimum gradient cost among the determined gradient costs of the plurality of pieces of weight information.
[0197] In one example, the plurality of weight information is sorted based on the determined gradient cost. The weight information can be determined based on the sorted plurality of weight information.
[0198] In one example, an index is encoded and included in a bitstream. The index can be signaled in high-level syntax.
[0199] FIG. 18A shows an exemplary flowchart outlining a decoding process (1800A) according to an embodiment of the present disclosure. The process (1800A) can be used in a video / image decoder. The process (1800A) can be executed by an apparatus for video / image coding that can include a receiving circuit and a processing circuit. In various embodiments, the process (1800A) is executed by a processing circuit such as a processing circuit in terminal devices (310), (320), (330), and (340), a processing circuit that executes the function of video encoder (403), a processing circuit that executes the function of video decoder (410), a processing circuit that executes the function of video decoder (510), a processing circuit that executes the function of video encoder (603), and the like. In some embodiments, the process (1800A) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1800A). The process starts at (S1801) and proceeds to (S1810).
[0200] (S1810), prediction information of the current block in the current picture can be decoded from the coded bitstream (e.g., the coded video bitstream). The prediction information indicates an intra prediction mode using multi-reference line (MRL) prediction applied to the current block. Based on the first reference line and the second reference line, the current block can be predicted. Each weight information of the plurality of weight information indicates a respective first weight candidate for the first reference line and a respective second weight candidate for the second reference line.
[0201] In (S1820), for each piece of weight information, a subset of samples within the current block can be predicted using intra prediction fusion based on the first reference line, the second reference line, and the respective weight information. The subset of samples can include at least one of (i) the upper samples in the top row of the current block or (ii) the left samples in the leftmost column of the current block. Based on the predicted subset of samples within the current block and the reconstructed samples outside the current block, a gradient cost can be determined. The reconstructed samples outside the current block can include samples adjacent to the predicted subset of samples within the current block.
[0202] In (S1830), based on the determined gradient costs corresponding to each of the plurality of pieces of weight information, weight information can be selected from among the plurality of pieces of weight information. The selected weight information indicates a first weight for the first reference line and a second weight for the second reference line.
[0203] Process (1800A) proceeds to (S1899) and ends.
[0204] Process (1800A) can be appropriately adapted to various scenarios, and accordingly, the steps of process (1800A) can be adjusted. One or more of the steps of process (1800A) can be adapted, omitted, repeated, and / or combined. Process (1800A) can be implemented using any appropriate order. Additional steps can be added.
[0205] In one example, the first reference line includes first reference samples that are N1 rows or N1 columns away from the current block. The second reference line includes second reference samples that are N2 rows or N2 columns away from the current block. N1 and N2 are different integers greater than or equal to 0.
[0206] In one example, for each of a plurality of weight information and one sample among subsets of samples, a first predicted value is determined based on one or more first reference samples of a first reference line using an intra prediction mode. A second predicted value is determined based on one or more second reference samples of a second reference line using the intra prediction mode. One sample among the subsets of samples is predicted based on the first predicted value corrected by a first weight, the second predicted value corrected by a second weight, and a residual of one sample among the subsets of samples.
[0207] In one example, a subset of samples within a current block includes upper samples of the top row within the current block and left samples of the leftmost column within the current block.
[0208] In one example, reconstructed samples outside a current block include reconstructed samples that are not adjacent to a predicted subset of samples.
[0209] In one example, the weight information is determined to be the weight information corresponding to the minimum gradient cost among the determined gradient costs among the plurality of weight information.
[0210] In one example, the plurality of weight information is sorted based on the determined gradient costs. The weight information can be determined based on an index and the sorted plurality of weight information.
[0211] In one example, the index is signaled in high-level syntax.
[0212] Figure 18B shows an exemplary flowchart outlining the decoding process (1800B) according to an embodiment of the present disclosure. The process (1800B) can be used in a video / image decoder. The process (1800B) can be executed by an apparatus for video / image coding that can include a receiving circuit and a processing circuit. In various embodiments, the process (1800B) is executed by a processing circuit such as the processing circuits within terminal devices (310), (320), (330), and (340), the processing circuit that executes the functions of video encoder (403), the processing circuit that executes the functions of video decoder (410), the processing circuit that executes the functions of video decoder (510), the processing circuit that executes the functions of video encoder (603), and the like. In some embodiments, the process (1800B) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1800B). The process starts at (S1802) and proceeds to (S1812).
[0213] At (S1812), receive the coded bitstream including the current block within the current picture.
[0214] At (S1822), obtain prediction information indicating whether an intra prediction mode using multi-reference line (MRL) prediction is applied to the current block from the coded bitstream. The current block is predicted based on a first reference line and a second reference line, where a first weight candidate is applied to the first reference line and a second weight candidate is applied to the second reference line.
[0215] At (S1832), obtain a plurality of weight candidate combinations from the coded bitstream. Each weight candidate combination includes a respective first weight candidate and a respective second weight candidate.
[0216] In (S1842), for each weight candidate combination, use intra-prediction fusion based on a first reference line weighted by each first weight candidate and a second reference line weighted by each second weight candidate to predict a subset of samples within the current block, where the subset of samples includes (i) the upper samples above the uppermost row within the current block and (ii) the left samples in the leftmost column within the current block. For each weight candidate combination, a gradient cost is determined based on the predicted subset of samples within the current block and the reconstructed samples outside the current block. The reconstructed samples outside the current block include samples adjacent to the predicted subset of samples within the current block.
[0217] In (S1852), a weight candidate combination is selected based on the determined gradient cost corresponding to each weight candidate combination.
[0218] Process (1800B) proceeds to (S1892) and ends.
[0219] Process (1800B) can be appropriately adapted to various scenarios and, accordingly, the steps of process (1800B) can be adjusted. One or more of the steps of process (1800B) can be adapted, omitted, repeated, and / or combined. Process (1800B) can be implemented using any suitable order. Additional steps can be added.
[0220] FIG. 19 shows an exemplary flowchart outlining an encoding process (1900) according to an embodiment of the present disclosure. The process (1900) can be used in a video / image encoder. The process (1900) can be executed by an apparatus for video / image coding that can include a processing circuit. In various embodiments, the process (1900) is executed by a processing circuit such as a processing circuit within terminal devices (310), (320), (330), and (340), a processing circuit that performs the functions of a video encoder (e.g., (403), (603), (703)). In some embodiments, the process (1900) is implemented with software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1900). The process starts at (S1901) and proceeds to (S1910).
[0221] (S1910), an intra prediction mode is applied to the current block using multi-reference line (MRL) prediction. The i-th entry and the (i + 1)-th entry in the MRL list correspond to the first reference line and the second reference line, respectively. For example, when i is 1, the first entry and the second entry in the MRL list correspond to the first reference line and the second reference line, respectively. In one example, the first reference line and the second reference line are not spatially adjacent.
[0222] (S1920), samples within the current block are encoded using intra prediction fusion based on multiple reference lines of the current block. The multiple reference lines include a first reference line and a second reference line.
[0223] The process (1900) then proceeds to (S1999) and ends.
[0224] The process (1900) can be appropriately adapted to various scenarios, and accordingly, the steps of the process (1900) can be adjusted. One or more of the steps of the process (1900) can be adapted, omitted, repeated, and / or combined. The process (1900) can be implemented using any appropriate order. Additional steps can be added.
[0225] In one example, using an intra prediction mode, a first predicted value is determined based on one or more first reference samples at a first reference line. Using the intra prediction mode, a second predicted value is determined based on one or more second reference samples at a second reference line. A sample can be predicted based on a weighted average of the first predicted value and the second predicted value.
[0226] In one example, a residual of a sample is determined based on the value of the sample and the predicted value of the sample. The residual of the sample can be encoded and included in a bitstream.
[0227] In one example, the MRL list is {1, 3, 5, 7, 12}, corresponding to reference lines 1, 3, 5, 7, and 12 respectively. Each of reference lines 1, 3, 5, 7, and 12 is 1, 3, 5, 7, and 12 rows and / or columns away from the current block respectively. The MRL index i which is 0, 2, 3, or 4 corresponds to reference lines 1, 5, 7, or 12 respectively. When the MRL index i is 0, the first reference line and the second reference line are reference lines 1 and 3. When the MRL index i is 2, the first reference line and the second reference line are reference lines 5 and 7. When the MRL index i is 3, the first reference line and the second reference line are reference lines 7 and 12. When the MRL index i is 4, the first reference line and the second reference line are reference lines 12 and 1.
[0228] In one example, the first weight associated with the first reference line and the second weight associated with the second reference line depend on the first reference line and the second reference line. Samples within the current block can be predicted based on the first reference line, the second reference line, the first weight, and the second weight.
[0229] In one example, the first weight associated with the first reference line and the second weight associated with the second reference line depend on the distance between the first reference line and the second reference line. The distance can be the number of rows and / or columns between the first reference line and the second reference line.
[0230] In one example, the first weight associated with the first reference line depends on the distance between the first reference line and reference line 0 adjacent to the current block. The distance is proportional to the MRL index i.
[0231] FIG. 20 shows an exemplary flowchart outlining a decoding process (2000) according to an embodiment of the present disclosure. The process (2000) can be used in a video / image decoder. The process (2000) can be executed by an apparatus for video / image coding that can include a receiving circuit and a processing circuit. In various embodiments, the process (2000) is executed by a processing circuit such as the processing circuits within terminal devices (310), (320), (330), and (340), the processing circuit that executes the function of video encoder (403), the processing circuit that executes the function of video decoder (410), the processing circuit that executes the function of video decoder (510), the processing circuit that executes the function of video encoder (603), and the like. In some embodiments, the process (2000) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2000). The process starts at (S2001) and proceeds to (S2010).
[0232] In (S2010), prediction information of a current block within a current picture can be decoded from a coded bitstream. The prediction information indicates that an intra prediction mode is applied to the current block using multi-reference line (MRL) prediction.
[0233] In (S2020), an MRL index i indicating that the i-th entry in the MRL list corresponds to a first reference line is received from the coded bitstream.
[0234] In (S2030), it can be determined that the (i + 1)-th entry in the MRL list corresponds to a second reference line.
[0235] In (S2040), samples of the current block can be reconstructed using intra prediction fusion based on a plurality of reference lines of the current block, and the first reference line and the second reference line among the plurality of reference lines are not spatially adjacent.
[0236] Process (2000) proceeds to (S2099) and ends.
[0237] Process (2000) can be appropriately adapted to various scenarios, and accordingly, steps of process (2000) can be adjusted. One or more steps of process (2000) can be adapted, omitted, repeated, and / or combined. Process (2000) can be implemented using any suitable order. Additional steps can be added.
[0238] In one example, using the intra prediction mode, a first predicted value is determined based on one or more first reference samples at the first reference line. Using the intra prediction mode, a second predicted value is determined based on one or more second reference samples at the second reference line. Samples can be predicted based on a weighted average of the first predicted value and the second predicted value.
[0239] In one embodiment, the sample is reconstructed based on the predicted sample and the residual of the sample.
[0240] In one example, the MRL list is {1, 3, 5, 7, 12}, corresponding to reference lines 1, 3, 5, 7, and 12 respectively. Each of the reference lines 1, 3, 5, 7, and 12 is 1, 3, 5, 7, and 12 rows and / or columns away from the current block respectively. The MRL index i, which is 0, 2, 3, or 4, corresponds to reference lines 1, 5, 7, or 12 respectively. When the MRL index i is 0, the first reference line and the second reference line are reference lines 1 and 3. When the MRL index i is 2, the first reference line and the second reference line are reference lines 5 and 7. When the MRL index i is 3, the first reference line and the second reference line are reference lines 7 and 12. When the MRL index i is 4, the first reference line and the second reference line are reference lines 12 and 1.
[0241] In one example, the first weight associated with the first reference line and the second weight associated with the second reference line depend on the first reference line and the second reference line. The samples within the current block can be reconstructed based on the first reference line, the second reference line, the first weight, and the second weight.
[0242] In one example, the first weight associated with the first reference line and the second weight associated with the second reference line depend on the distance between the first reference line and the second reference line. The distance can be the number of rows and / or columns between the first reference line and the second reference line.
[0243] In one example, the first weight associated with the first reference line depends on the distance between the first reference line and the reference line 0 adjacent to the current block. The distance is proportional to the MRL index i.
[0244] Embodiments in the present disclosure may be used separately or combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0245] The above-described technology can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, FIG. 21 shows a computer system (2100) suitable for implementing certain embodiments of the disclosed subject matter.
[0246] The computer software can be coded using any suitable machine code or computer language that can be the subject of assembly, compilation, linking, or similar mechanisms, and can create code that includes instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.
[0247] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.
[0248] The components shown in FIG. 21 for the computer system (2100) are essentially illustrative and are not intended to imply any limitations regarding the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependencies or requirements regarding any one or combination of the components shown in the exemplary embodiment of the computer system (2100).
[0249] The computer system (2100) may include a specific human interface input device. Such a human interface input device can respond to input by one or more human users, for example, through tactile input (keystrokes, swipes, movement of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). Also, the human interface input device can be used to capture a specific medium, such as audio (voice, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), and video (2D video, 3D video including stereoscopic video, etc.), which is not necessarily directly related to conscious human input.
[0250] The human interface input device may include one or more of a keyboard (2101), a mouse (2102), a trackpad (2103), a touch screen (2110), a data glove (not shown), a joystick (2105), a microphone (2106), a scanner (2107), and a camera (2108) (only one of each is shown).
[0251] The computer system (2100) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, acoustics, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (2110), a data glove (not shown), or a joystick (2105), although there may be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (2109), headphones (not shown), etc.), visual output devices (each regardless of the presence or absence of a touch screen input function and also regardless of the presence or absence of a tactile feedback function, some of which can output two-dimensional visual output or output beyond three dimensions through means such as stereoscopic image output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), including screens (2110) such as CRT screens, LCD screens, plasma screens, and OLED screens), and printers (not shown).
[0252] The computer system (2100) may also include human-accessible storage devices and their associated media, such as optical media or similar media (2121) including CD / DVD ROM / RW (2120) having CD / DVDs, thumb drives (2122), removable hard drives or solid state drives (2123), legacy magnetic media such as tapes and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0253] One of ordinary skill in the art should also understand that the term "computer-readable medium," when used in connection with the presently disclosed subject matter, does not include transmission media, carrier waves, or other transient signals.
[0254] The computer system (2100) can also include an interface (2154) to one or more communication networks (2155). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet (registered trademark), Wi-Fi, cellular networks including GSM (registered trademark), 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and vehicle and industrial networks including CANBus, etc. A particular network generally requires attachment to a particular general-purpose data port or peripheral bus (2149) and an external network interface adapter (such as a USB port of the computer system (2100)), while others are generally integrated into the core of the computer system (2100) by attachment to the system bus described later (such as an Ethernet (registered trademark) interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2100) can communicate with other entities. Such communication can be, for example, one-way receive-only (such as broadcast TV), one-way transmit-only (such as from a particular CANbus to a particular CANbus device), or two-way using a local or wide area digital network. As described above, specific protocols and protocol stacks can be used in each of these networks and network interfaces.
[0255] The foregoing human interface device, human-accessible storage device, and network interface can be attached to the core (2140) of the computer system (2100).
[0256] The core (2140) can include one or more central processing units (CPUs) (2141), a graphics processing unit (GPU) (2142), a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) (2143), a hardware accelerator (2144) for specific tasks, a graphics adapter (2150), etc. These devices can be connected through a system bus (2148) together with a read-only memory (ROM) (2145), a random access memory (RAM) (2146), an internal large-capacity storage such as an internal non-user-accessible hard drive, SSD (2147). In some computer systems, the system bus (2148) can be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (2148) of the core or via a peripheral bus (2149). In one example, a screen (2110) can be connected to a graphics adapter (2150). The architecture of the peripheral bus includes PCI, USB, etc.
[0257] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) can execute specific instructions that can be combined to form the above-mentioned computer code. That computer code can be stored in the ROM (2145) or RAM (2146). Also, temporary data can be stored in the RAM (2146), while permanent data can be stored, for example, in the internal large-capacity storage (2147). Through the use of cache memory that can be closely associated with one or more CPUs (2141), GPUs (2142), large-capacity storage (2147), ROM (2145), RAM (2146), etc., fast storage and retrieval for any of the memory devices can be enabled.
[0258] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well-known and available to those having skill in the computer software arts.
[0259] By way of example and not limitation, a computer system (2100) having an architecture and in particular a core (2140) can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be associated with user-accessible mass storage as introduced above, as well as specific storage of the core (2140) of a non-transitory nature such as internal core mass storage (2147) or ROM (2145). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (2140). The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause the core (2140) and in particular the processor (including a CPU, GPU, FPGA, etc.) therein to define data structures stored in the RAM (2146) and modify such data structures according to processes defined by the software, thereby causing the core to execute specific processes or specific portions of specific processes described herein. Additionally or alternatively, the computer system can provide functionality as a result of being embodied in logic hardware or otherwise in a circuit (such as an accelerator (2144)), which can operate instead of or together with the software to execute specific processes or specific portions of specific processes described herein. References to software include logic and, where appropriate, vice versa. References to computer-readable media can include circuits (such as integrated circuits (ICs), etc.) that store software for execution, circuits that embody logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software. Appendix A: Acronyms JEM: joint exploration model VVC: versatile video coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOPs: Groups of Pictures TUs: Transform Units PUs: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: solid-state drive IC: Integrated Circuit CU: Coding Unit SDR: Standard dynamic range JVET: Joint Video Exploration Team AMVR: Adaptive Motion Vector Resolution POC: Picture Order Count SbTMVP: Subblock-based Temporal Motion Vector Predictor
[0260] Although the present disclosure describes some exemplary embodiments, there are changes, substitutions, and various alternative equivalents that are within the scope of the present disclosure. Thus, it will be understood by those skilled in the art that, although not explicitly shown or described herein, various systems and methods that embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure can be devised.
Claims
1. A method for video decoding in a decoder, comprising: receiving a coded bitstream including a current block in a current picture; obtaining prediction information indicating whether an intra prediction mode using multi-reference lines (MRL) prediction is applied to the current block from the coded bitstream, wherein the current block is predicted based on a first reference line and a second reference line, a first weight candidate is applied to the first reference line, and a second weight candidate is applied to the second reference line; obtaining a plurality of weight candidate combinations from the coded bitstream, each weight candidate combination including a respective first weight candidate and a respective second weight candidate; for each weight candidate combination, predicting a subset of samples in the current block using intra prediction fusion based on the first reference line weighted by the respective first weight candidate and the second reference line weighted by the respective second weight candidate, wherein the subset of samples includes (i) upper samples in the uppermost row in the current block and (ii) left samples in the leftmost column in the current block; determining a gradient cost based on the predicted subset of samples in the current block and reconstructed samples outside the current block, wherein the reconstructed samples outside the current block include samples adjacent to the predicted subset of samples in the current block; selecting a weight candidate combination based on the determined gradient cost corresponding to each weight candidate combination; A method comprising the above steps.
2. The first reference line includes first reference samples that are N1 rows or N1 columns away from the current block, The second reference line includes second reference samples that are N2 rows or N2 columns away from the current block, where N1 and N2 are different integers greater than or equal to 0. The method according to claim 1.
3. The step of predicting the subset of samples comprises: for each weight candidate combination among the plurality of weight candidate combinations and one sample among the subset of samples, Determining a first prediction value based on one or more first reference samples of the first reference line using the intra prediction mode; Determining a second prediction value based on one or more second reference samples of the second reference line using the intra prediction mode; Predicting one sample of the subset of the samples based on the first prediction value corrected by the first weight candidate, the second prediction value corrected by the second weight candidate, and the residual of the one sample of the subset of the samples; The method according to claim 1, comprising:
4. The reconstructed samples outside the current block include reconstructed samples that are not adjacent to the predicted subset of the samples. The method according to claim 3.
5. The step of selecting the weight candidate combination: The step of selecting the weight candidate combination includes selecting the weight candidate combination corresponding to the minimum gradient cost among the determined gradient costs among the plurality of weight candidate combinations. The method according to claim 1.
6. The step of selecting the weight candidate combination: Sorting the plurality of weight candidate combinations based on the determined gradient cost; Determining the weight candidate combination based on an index and the sorted plurality of weight candidate combinations; The method according to claim 1, comprising:
7. The index is signaled in high-level syntax. The method according to claim 6.
8. A method for video decoding in a decoder, comprising: Decoding prediction information of a current block in a current picture from a coded bitstream, the prediction information indicating an intra prediction mode using a multi-reference line (MRL) prediction applied to the current block; Receiving an MRL index i indicating that the i-th entry in the MRL list corresponds to a first reference line from the coded bitstream; Determining that the (i + 1)-th entry in the MRL list corresponds to a second reference line; Reconstructing samples in the current block using intra prediction fusion based on a plurality of reference lines of the current block, wherein the first reference line and the second reference line in the plurality of reference lines are not spatially adjacent; A method comprising
9. The step of reconstructing the sample comprises Determining a first predicted value based on one or more first reference samples of the first reference line using the intra prediction mode; Determining a second predicted value based on one or more second reference samples of the second reference line using the intra prediction mode; Predicting the sample based on a weighted average of the first predicted value and the second predicted value; The method according to claim 8, comprising
10. The step of reconstructing the sample comprises Reconstructing the sample based on the predicted sample and a residual of the sample. The method according to claim 9.
11. The MRL list is {1, 3, 5, 7, 12} corresponding to reference lines 1, 3, 5, 7, and 12 respectively, and each of the reference lines 1, 3, 5, 7, and 12 is 1, 3, 5, 7, and 12 rows and / or columns away from the current block respectively; The MRL index i, which is 0, 2, 3, or 4, corresponds to the reference lines 1, 5, 7, or 12 respectively; In response to the MRL index i being 0, the first reference line and the second reference line are the reference lines 1 and 3; In response to the MRL index i being 2, the first reference line and the second reference line are the reference lines 5 and 7; In response to the MRL index i being 3, the first reference line and the second reference line are the reference lines 7 and 12; In response to the MRL index i being 4, the first reference line and the second reference line are the reference lines 12 and 1. The method according to claim 9.
12. The first weight associated with the first reference line and the second weight associated with the second reference line depend on the first reference line and the second reference line; The step of reconstructing the sample comprises reconstructing the sample within the current block based on the first reference line, the second reference line, the first weight, and the second weight. The method according to claim 8.
13. The first weight associated with the first reference line and the second weight associated with the second reference line depend on the distance between the first reference line and the second reference line, and the distance is the number of rows and / or columns between the first reference line and the second reference line. The method according to claim 8.
14. The first weight associated with the first reference line depends on the distance between the first reference line and the reference line 0 adjacent to the current block, and the distance is proportional to the MRL index i. The method according to claim 8.
15. An apparatus for video decoding in a decoder, receiving a coded bitstream including a current block in a current picture, obtaining, from the coded bitstream, prediction information indicating whether an intra prediction mode using multi-reference line (MRL) prediction is applied to the current block, the current block being predicted based on a first reference line and a second reference line, a first weight candidate being applied to the first reference line, and a second weight candidate being applied to the second reference line, obtaining, from the coded bitstream, a plurality of weight candidate combinations, each weight candidate combination including a respective first weight candidate and a respective second weight candidate, for each weight candidate combination, predicting a subset of samples in the current block using intra prediction fusion based on the first reference line weighted by the respective first weight candidate and the second reference line weighted by the respective second weight candidate, the subset of samples including (i) the upper samples in the uppermost row in the current block and (ii) the left samples in the leftmost column in the current block, determining a gradient cost based on the predicted subset of samples in the current block and the reconstructed samples outside the current block, the reconstructed samples outside the current block including samples adjacent to the predicted subset of samples in the current block, selecting a weight candidate combination based on the determined gradient cost corresponding to each weight candidate combination, An apparatus comprising a processing circuit configured as such.
16. The first reference line includes first reference samples that are N1 rows or N1 columns away from the current block, The second reference line includes second reference samples that are N2 rows or N2 columns away from the current block, N1 and N2 are different integers greater than or equal to 0. The apparatus according to claim 15.
17. The processing circuit, for each weight candidate combination among the plurality of weight candidate combinations and for one sample among the subset of samples, Determine a first prediction value based on one or more first reference samples of the first reference line using the intra prediction mode; Determine a second prediction value based on one or more second reference samples of the second reference line using the intra prediction mode; Predict the one sample of the subset of the samples based on the first prediction value corrected by the first weight candidate, the second prediction value corrected by the second weight candidate, and the residual of the one sample of the subset of the samples; The apparatus according to claim 15, configured as such.
18. The reconstructed samples outside the current block include reconstructed samples that are not adjacent to the predicted subset of the samples. The apparatus according to claim 17.
19. The processing circuit selects the weight candidate combination to be the weight candidate combination corresponding to the minimum gradient cost among the determined gradient costs among the plurality of weight candidate combinations; The apparatus according to claim 15, configured as such.
20. The processing circuit rearranges the plurality of weight candidate combinations based on the determined gradient cost, and determines the weight candidate combination based on an index and the rearranged plurality of weight candidate combinations; The apparatus according to claim 15, configured as such.
21. A computer program that, when executed by a processing circuit, causes the processing circuit to execute the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Method and apparatus for video coding using decoder side intra prediction derivation
US20190215521A1
Image encoding / decoding method and device, and recording medium having bitstream stored therein
US20200244956A1
Image encoding / decoding method and device, and recording medium stored with bitstream
US20200413069A1
Decoder side intra mode derivation for most probable mode list construction in video coding
WO2022140718A1
Intra-frame prediction fusion method, video coding method and apparatus, video decoding method and apparatus, and system
WO2024007366A1