Modified Intra Prediction Fusion

By predicting video blocks using multiple intra-prediction modes with weighted sums based on neighboring block frequencies and spatial relationships, the method addresses inefficiencies in intra-prediction, enhancing coding efficiency and data reduction in video transmission.

JP2025530287AActive Publication Date: 2025-09-11TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025514693
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-18
Filing Date
2023-10-19
Publication Date
2025-09-11
Estimated Expiration
2043-10-19

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in intra-prediction methods, particularly in deriving weights for combining multiple intra-prediction modes, leading to suboptimal coding efficiency.

Method used

A method for video encoding/decoding that involves predicting a current block using multiple candidate intra-prediction modes, deriving weights for each mode based on the frequency and spatial relationships of neighboring blocks, and applying a weighted sum to improve coding efficiency.

Benefits of technology

Enhances coding efficiency by accurately modeling local textures through intra-prediction fusion, improving the prediction accuracy and reducing data volume in video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025530287000001_ABST
    Figure 2025530287000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure include methods and devices for video coding. One of the devices includes a processing circuit that receives a current block in a bitstream. The current block is predicted using intra-prediction fusion including multiple candidate intra-prediction modes. The processing circuit determines, for each of the multiple candidate intra-prediction modes, a candidate predicted value for each of samples in the current block. The processing circuit derives weights for each of the multiple candidate intra-prediction modes based on the intra-prediction modes used to code neighboring blocks of the current block. The processing circuit predicts samples in the current block using a weighted sum of the candidate predicted values ​​associated with the multiple candidate intra-prediction modes according to the derived weights.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure describes embodiments that relate generally to video coding. [Background technology]

[0002] [Incorporated by reference] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. patent application Ser. No. 18 / 381,415, entitled "MODIFIED INTRA PREDICTION FUSION," filed Oct. 18, 2023, which claims the benefit of priority to U.S. Provisional Application No. 63 / 417,650, entitled "Modified Intra Prediction Fusion," filed Oct. 19, 2022, the entire contents of which are incorporated herein by reference. [Background technology] The background art discussion provided herein is intended to generally present the context for the present disclosure. The inventors' work, to the extent described in this background art section, as well as aspects of the discussion that are not admitted as prior art at the time of filing, are not admitted, expressly or impliedly, as prior art to the present disclosure.

[0003] Image / video compression can help transmit image / video data across different devices, storage, and networks with minimal image quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from a current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0004] Aspects of the present disclosure include methods and apparatuses for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit receives a current block in a bitstream. The current block is predicted using multiple candidate intra-prediction modes. For example, the current block is predicted using intra-prediction fusion including multiple candidate intra-prediction modes. The processing circuit determines a candidate prediction value for each of the multiple candidate intra-prediction modes for samples in the current block and derives weights for each of the multiple candidate intra-prediction modes based on the intra-prediction modes used to code neighboring blocks of the current block. The processing circuit predicts samples in the current block by a weighted sum of the candidate prediction values ​​associated with the multiple candidate intra-prediction modes according to the derived weights.

[0005] In one example, the processing circuit determines a frequency at which multiple candidate intra-prediction modes are applied to code the neighboring block based on the intra-prediction modes used to code the neighboring block, and derives a weight for one of the multiple candidate intra-prediction modes based on the determined frequency.

[0006] In one example, the processing circuit derives a horizontal weight for one of a plurality of candidate intra-prediction modes, derives a vertical weight for one of a plurality of candidate intra-prediction modes, and derives a weight for one of the plurality of candidate intra-prediction modes based on the derived horizontal weight and the derived vertical weight.

[0007] In one example, the processing circuit determines a frequency at which one of a plurality of candidate intra-prediction modes is applied to code a left-neighboring block in the neighboring blocks based on one or more of the intra-prediction modes used to code the left-neighboring block, and derives a horizontal weight for one of the plurality of candidate intra-prediction modes based on the determined frequency.

[0008] In one example, the processing circuit determines a frequency at which one of a plurality of candidate intra prediction modes is applied to code an upper adjacent block in the adjacent block based on one or more of the intra prediction modes used to code the upper adjacent block, and derives a vertical weight for one of the plurality of candidate intra prediction modes based on the determined frequency.

[0009] In one example, the processing circuit derives a first weight for a horizontal weight and a second weight for a vertical weight based on the relative coordinates of the sample with respect to the top-left coordinate within the current block, and uses the first weight and the second weight, respectively, to derive a weight for one of a plurality of candidate intra-prediction modes based on a weighted sum of the derived horizontal weight and the derived vertical weight.

[0010] In one example, the processing circuit derives a weight for one of a plurality of candidate intra-prediction modes based on bilinear interpolation of a horizontal weight, a vertical weight, a default horizontal weight, and a default vertical weight.

[0011] In one example, the multiple candidate intra prediction modes comprise one or more of a DC mode, a planar mode, an intra directional prediction mode, a decoder-side intra mode derivation (DIMD) mode, a template based intra mode derivation (TIMD) mode, a cross-component linear model (CCLM), a convolutional cross-component linear model (CCLM), and a multi-model linear mode (MMLM).

[0012] In one example, the processing circuit derives weights for each of a plurality of candidate intra-prediction modes based on neighboring reconstructed samples of the current block, the neighboring reconstructed samples including reconstructed samples within at least one line of the current block.

[0013] In one example, the processing circuit calculates a histogram of edge directions for the current block based on neighboring reconstructed samples of the current block. The histogram of edge directions indicates the frequency of edge directions. The processing circuit derives a weight for one of the plurality of candidate intra-prediction modes based on the frequency of one of the edge directions. One of the plurality of candidate intra-prediction modes is associated with one of the edge directions.

[0014] In one example, the processing circuit calculates a template matching cost between a current template including neighboring reconstructed samples of the current block and each of the current templates indicated by the plurality of candidate intra-prediction modes, and derives weights for the plurality of candidate intra-prediction modes based on the respective template matching costs.

[0015] In one example, the processing circuit calculates a histogram of left edge directions of the current block based on left-neighboring reconstructed samples in neighboring reconstructed samples of the current block. The histogram of left edge directions indicates a frequency of left edge directions. The processing circuit derives the horizontal weight of the one of the plurality of candidate intra-prediction modes based on a frequency of one of the left edge directions. One of the plurality of candidate intra-prediction modes is associated with one of the left edge directions.

[0016] In one example, the processing circuit calculates a histogram of top edge directions of the current block based on the reconstructed samples of neighboring upper adjacent reconstructed samples of the current block. The histogram of top edge directions indicates the frequency of the top edge directions. The processing circuit derives a vertical weight for one of a plurality of candidate intra-prediction modes based on the frequency of one of the top edge directions. One of the plurality of candidate intra-prediction modes is associated with one of the top edge directions.

[0017] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding / encoding. [Brief explanation of the drawings]

[0018] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100). [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4] 1 illustrates intra-prediction modes according to an embodiment of the present disclosure. [Figure 5] 1 illustrates intra prediction modes according to one embodiment of the present invention. [Figure 6] 1 illustrates an example of four reference lines 0 to 3 adjacent to a coding block unit according to an embodiment of the present invention. [Figure 7] 1 shows an example of template-based intra-mode derivation (TIMD). [Figure 8] 1 shows an example of decoder-side intra mode derivation (DIMD). [Figure 9] An example of determining weights based on horizontal and vertical weights will be shown. [Figure 10] 1 shows a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 11] 1 shows a flowchart outlining another process according to some embodiments of the present disclosure. [Figure 12] 1 shows a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 13] 1 shows a flowchart outlining another process according to some embodiments of the present disclosure. [Figure 14] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0019] 1 illustrates a block diagram of a video processing system 100 in accordance with some examples. The video processing system 100 is an example of an application of the disclosed subject matter, a video encoder and a video decoder, in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, and storage of compressed video on digital media, including CDs, DVDs, memory sticks, etc.

[0020] The video processing system (100) includes a capture subsystem (113) that may include a video source (101), such as a digital camera, that creates a stream of uncompressed video pictures (102). In one example, the stream of video pictures (102) includes samples captured by the digital camera. The stream of video pictures (102), shown with a thick line to emphasize its high data volume compared to the encoded video data (104) (or coded video bitstream), may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or coded video bitstream), shown with a thin line to emphasize its lower data volume compared to the stream of video pictures (102), may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems (106) and (108) of FIG. 1, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include a video decoder (110), for example, within an electronic device (130). The video decoder (110) decodes an input copy (107) of the encoded video data and creates an output stream (111) of video pictures that can be rendered on a display (112) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) can be encoded according to several video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, a developing video coding standard is informally known as Universal Video Coding (VVC).The disclosed subject matter can be used in the context of a VVC.

[0021] It should be noted that the electronic devices (120) and (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may include a video encoder (not shown).

[0022] 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used in place of the video decoder (110) in the example of FIG. 1.

[0023] The receiver (231) can receive one or more coded video sequences contained in a bitstream, for example, to be decoded by the video decoder (210). In one embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of the other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (231) can receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to respective using entities (not shown). The receiver (231) can separate the coded video sequences from other data. To address network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter, "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other embodiments, it may be external to the video decoder (210) (not shown). In still other embodiments, there may be a buffer memory (not shown) external to the video decoder (210), for example, to deal with network jitter, and there may be another buffer memory (215) internal to the video decoder (210), for example, to handle playback timing. When the receiver (231) is receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory (215) may not be needed or may be small. For use over best-effort packet networks such as the Internet, the buffer memory (215) may be required and may be relatively large, advantageously adaptively sized, and implemented at least in part within an operating system or similar element (not shown) external to the video decoder (210).

[0024] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and, potentially, information for controlling a rendering device, such as a rendering device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but may be coupled to the electronic device (230), as shown in FIG. 2. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups can include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (220) can also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0025] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0026] The reconstruction of the symbols (221) can involve several different units, depending on the type of coded video picture or portion thereof (interpicture and intrapicture, interblock and intrablock, etc.), as well as other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (220). The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.

[0027] In addition to the functional blocks already mentioned, the video decoder (210) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:

[0028] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as well as control information from the parser (220) as symbols (221), including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. The scalar / inverse transform unit (251) can output blocks containing sample values, which can be input to an aggregator (255).

[0029] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258), for example, buffers partially reconstructed and / or fully reconstructed current pictures. The aggregator (255) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0030] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (253) may access a reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) associated with the block, these samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (253) in the form of symbols (221), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0031] The output samples of the aggregator (255) can be subjected to various loop filtering techniques in a loop filter unit (256). Video compression techniques can include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called a coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression can be responsive to meta-information obtained during decoding of a previous portion (in decoding order) of the coded picture or coded video sequence, or it can be responsive to previously reconstructed, loop-filtered sample values.

[0032] The output of the loop filter unit (256) may be a sample stream that may be output to the rendering device (212) and stored in a reference picture memory (257) for use in future inter-picture prediction.

[0033] Once a coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before beginning reconstruction of the subsequent coded picture.

[0034] Video decoder 210 can perform decoding operations according to a given video compression technology or standard (e.g., ITU-T Rec. H.265). A coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile may select some tools from all tools available in the video compression technology or standard as the only tools available for use under that profile. Also, a requirement for compliance may be that the complexity of the coded video sequence be within boundaries defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained through a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0035] In one embodiment, the receiver (231) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0036] 3 shows an exemplary block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.

[0037] The video encoder (303) can receive video samples from a video source (301) (not part of the electronic device (320) in the example of FIG. 3) that can capture video images to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0038] The video source (301) can provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCb, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media delivery system, the video source (301) can be a storage device that stores previously prepared video. In a videoconferencing system, the video source (301) can be a camera that captures local image information as a video sequence. The video data can be provided as multiple individual pictures that, when viewed sequentially, impart motion. The pictures themselves can be organized as a spatial array of pixels, each of which can comprise one or more samples depending on the sampling structure, color space, etc., in use. The following discussion focuses on samples.

[0039] According to one embodiment, the video encoder (303) can code and compress pictures of a source video sequence into a coded video sequence (343) in real time, or under any other time constraints as needed. Enforcing the appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is operatively coupled to other functional units, as described below. This coupling is not shown for clarity. Parameters set by the controller (350) can include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other appropriate functionality associated with the video encoder (303) optimized for a particular system design.

[0040] In some embodiments, the video encoder (303) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop can include a source coder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and one or more reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to that of a (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-accurate results independent of the location (local or remote) of the decoder, the contents in the reference picture memory (334) are also bit-accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the decoder would "see" when using prediction during decoding. This basic principle of reference picture synchronism (and the resulting drift if synchronism cannot be maintained, eg due to channel errors) is also used in several related techniques.

[0041] The operation of the "local" decoder (333) may be the same as a "remote" decoder, such as the video decoder (210) already described in detail above in connection with Figure 2. However, briefly referring also to Figure 2, because symbols are available and the encoding / decoding of the symbols into an encoded video sequence by the entropy coder (345) and parser (220) may be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333).

[0042] In one embodiment, decoder technology, with the exception of parsing / entropy decoding, present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operation. A description of the encoder technology may be omitted, as it is the reverse of the decoder technology described generically. In certain areas, more detailed descriptions are provided below.

[0043] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0044] The local video decoder (333) can decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded in a video decoder (not shown in FIG. 3), the reconstructed video sequence may generally be a replica of the source video sequence with some errors. The local video decoder (333) can replicate the decoding process that may be performed on the reference pictures by the video decoder and store the reconstructed reference pictures in the reference picture memory (334). In this way, the video encoder (303) can locally store copies of reconstructed reference pictures that have common content with reconstructed reference pictures obtained by the far-end video decoder (in the absence of transmission errors).

[0045] The predictor (335) may perform the prediction search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor (335) may operate on a sample block-by-sample block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).

[0046] The controller (350) can manage the coding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0047] The output of all the aforementioned functional units is entropy coded in entropy coder 345. The entropy coder (345) converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.

[0048] The transmitter (340) may buffer the coded video sequence created by the entropy coder (345) for transmission over a communication channel (360), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (340) may merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0049] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign each coded picture a particular coded picture type, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as one of the following picture types:

[0050] Intra-pictures (I-pictures) can be coded and decoded without using other pictures in the sequence as a source of prediction. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh (“IDR”) pictures.

[0051] Predictive pictures (P pictures) may be coded and decoded using intra- or inter-prediction, which uses motion vectors and reference indices to predict the sample values ​​of each block.

[0052] Bidirectionally predicted pictures (B pictures) can be coded and decoded using intra- or inter-prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multi-predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0053] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0054] The video encoder (303) may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0055] In one embodiment, the transmitter (340) can transmit additional data along with the encoded video. The source coder (330) can include such data as part of the coded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0056] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture may be coded by a vector called a motion vector. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0057] In some embodiments, bi-prediction techniques may be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are earlier in decoding order than a current picture in a video (but may be in the past and future, respectively, in display order). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0058] Also, coding efficiency can be improved by using a merge mode technique in inter-picture prediction.

[0059] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU may be partitioned into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations during coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values) of 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0060] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technique. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.

[0061] Various intra prediction modes for intra prediction may be used in video coding such as HEVC and VVC. Figure 4 illustrates intra prediction modes (e.g., 35 intra prediction modes as used in HEVC) according to an embodiment of the present disclosure. In one example such as HEVC, there are 35 intra prediction modes (e.g., a total of 35 intra prediction modes). Referring to Figure 4, of the 35 intra prediction modes, mode 10 is the horizontal mode, mode 26 is the vertical mode, and mode 2, mode 18, and mode 34 are diagonal modes. The intra prediction modes may be signaled by three most probable modes (MPMs) and 32 remaining modes.

[0062] FIG. 5 illustrates intra-prediction modes, such as those defined in VVC Draft 2, according to an embodiment of the present disclosure. Referring to FIG. 5, in an example VVC, there are 95 intra-prediction modes (e.g., a total of 95 intra-prediction modes). In one example, the 95 intra-prediction modes are represented by modes −14 through 80. For example, mode 18 is a horizontal mode, mode 50 is a vertical mode, and modes 2, 34, and 66 are diagonal modes. Modes −1 through −14 and modes 67 through 80 may be referred to as Wide-Angle Intra Prediction (WAIP) modes.

[0063] Multi-line intra prediction may be applied to video coding. In multi-line intra prediction, more reference lines (e.g., more reference lines neighboring the current block to be coded) may be used for intra prediction. The encoder may determine and signal which reference line is used to generate the intra predictor. The reference line index may be signaled before the intra prediction mode. In one example, if a non-zero reference line index is signaled, only the most probable mode is allowed. Figure 6 shows an example of four reference lines 0 to 3 neighboring a coding block unit (e.g., block) (601) according to an embodiment of the present disclosure. In one embodiment, each reference line includes six segments, i.e., segments A to F, with a top-left reference sample. In one example, segments A and F are padded with the nearest samples from segments B and E, respectively.

[0064] Template-based intra mode derivation (TIMD) can use a reference sample of the current CU as a template to select an intra mode from a set of candidate intra-prediction modes associated with the TIMD. The selected intra mode can be determined as the best intra mode based on a cost function, for example. As shown in FIG. 7, neighboring reconstructed samples of the current CU (702) can be used as a template (704). The reconstructed samples in the template (704) can be compared with predicted samples of the template (704). The predicted samples can be generated using a reference sample (706) of the template (704). The reference sample (706) can be neighboring reconstructed samples around the template (704). The cost function can be used to calculate the cost (or distortion) between the predicted samples in the template (704) and the reconstructed samples based on each of the set of candidate intra-prediction modes. The intra-prediction mode with the lowest cost (or distortion) may be selected as the intra-prediction mode (eg, the best intra-prediction mode) for intra-predicting the current CU (702).

[0065] A spatial geometric partitioning mode (SGPM) may be used in intra prediction. SGPM may be an intra mode similar to the inter-coding tool of GPM. In SGPM, a block may be partitioned into two partitions along a partition boundary according to a partition mode. In one example, the partition mode is one of the partition modes indicated by a partition mode index. Two respective predicted portions may be generated from an intra-predicted process. In one embodiment, SGPM uses blending (e.g., adaptive blending) to determine a predicted value for a sample within a predetermined distance from the partition boundary using a weighted average of first and second predicted values ​​generated from the two intra-prediction modes, respectively, according to the weights of the two intra-prediction modes.

[0066] When decoder-side intra-mode derivation (DIMD) is applied, N intra-modes may be derived from reconstructed neighboring samples around the current block (801), and N predictors obtained using the N intra-modes may be combined with a planar mode predictor with corresponding weights. The weights may be derived from gradients, such as a histogram of gradients (HoG) calculation. FIG. 8 shows an example of DIMD. The HoG calculation may be performed by applying a filter (e.g., horizontal and vertical Sobel filters) to pixels in a template (802) around the current block (801). The template (802) may include reconstructed neighboring samples around the current block (801). In one example, the template has a width of 3. In one example, pixels in the middle line of the template (802) (marked in gray) may be involved in the HoG calculation. Referring to FIG. 8, a window (803) around a pixel (805) may be used to determine the gradient associated with the pixel (805). The window (803) may have a size of 3x3 with a pixel (805) at the center of the window (803). Horizontal and vertical gradients may be obtained, for example, using horizontal and vertical Sobel filters, respectively. A direction or orientation may be obtained from the horizontal and vertical gradients. An intra-prediction mode (IPM) associated with the direction may be determined. Subsequently, a histogram (also called HoG) (810) of the IPMs may be obtained. The IPM corresponding to the N highest histogram bars may be selected for the current block (801).

[0067] In some examples, a specific intra-prediction mode may be inefficient in accurately modeling local texture, and fusion of multiple intra-prediction modes can improve prediction accuracy. However, inefficient design for deriving weights when combining multiple intra-prediction samples may result in lower coding efficiency. This disclosure includes a set of advanced video coding techniques, such as a method for deriving weights (or weights) for multiple intra-prediction modes to improve coding efficiency of fusion of multiple intra-prediction modes. More specifically, a method for deriving weights for multiple intra-prediction modes when applying intra-prediction fusion is described.

[0068] In this disclosure, the term "intra prediction fusion" can refer to any case where multiple intra prediction modes are used to generate predictions for a block and derive one residual block. Examples include, but are not limited to, fusion between DIMD mode and normal intra prediction mode, fusion between TIMD mode and normal intra prediction mode, SGPM (or spatial GPM where two intra prediction modes are signaled for one coding block), fusion of two normal intra prediction modes, and fusion of cross-component linear model (CCLM) / convolutional cross-component model (CCCM) / multi-model linear mode (MMLM) and normal intra prediction mode.

[0069] In this disclosure, the terms "intra prediction fusion" or "fusion of multiple intra prediction modes" may refer to when multiple intra prediction modes are used to predict a block, such as to generate a predictive block of the block. A residual block may be derived using the multiple intra prediction modes.

[0070] The regular intra prediction modes may include the intra prediction modes described with reference to Figures 4-5, such as DC mode, planar mode, and intra directional prediction mode (also referred to as angular prediction mode or angular mode). The regular intra prediction modes may be used with reference line 0 in Figure 6 or another reference line in Figure 6 different from reference line 0. In one example, the regular intra prediction modes do not include DIMD mode, TIMD mode, SGPM, CCLM, CCCM, and MMLM.

[0071] When two or more intra-prediction modes (e.g., multiple intra-prediction modes) are used to generate the final intra-predicted block, i.e., when intra-prediction fusion is applied, for each sample, a weighting for each candidate prediction value is derived based on the coded intra-prediction modes from neighboring blocks. Each intra-prediction mode involved in intra-prediction fusion is referred to as a candidate intra-prediction mode. Therefore, two or more intra-prediction modes or multiple intra-prediction modes used to predict the current block may also be referred to as multiple candidate intra-prediction modes.

[0072] For example, when a current block is predicted using multiple candidate intra-prediction modes, such as using intra-prediction fusion or fusion of multiple intra-prediction modes, a candidate predicted value for each of the samples in the current block may be determined for each of the multiple candidate intra-prediction modes. A weight for each of the multiple candidate intra-prediction modes (or a weighting for each candidate predicted value) may be derived based on the intra-prediction modes used to code neighboring blocks of the current block. Samples in the current block may be predicted according to the derived weight by a weighted sum of the candidate predicted values ​​associated with the multiple candidate intra-prediction modes. As described above, each intra-prediction mode of the multiple candidate intra-prediction modes is referred to as a candidate intra-prediction mode. Neighboring blocks of the current block may include neighboring blocks that are spatially adjacent to the current block.

[0073] The multiple candidate intra prediction modes may include one or more of DC mode, planar mode, intra directional prediction mode (also called angular prediction mode or angular mode), DIMD mode, TIMD mode, CCLM, CCCM, and MMLM.

[0074] In one example, a DIMD mode and an angular mode (e.g., a normal intra prediction mode) are used in the fusion of multiple intra prediction modes, where the multiple intra prediction modes include the DIMD mode and the angular mode. A first predictor of a current block (e.g., a first predicted block) is obtained using the DIMD mode, and a second predictor of the current block (e.g., a second predicted block) is obtained using the angular mode. A predicted block (e.g., a final intra prediction block) may be obtained by a weighted average of the first predictor and the second predictor according to the respective weights of the DIMD mode and the angular mode. In one example, the weights of the DIMD mode and the angular mode are specific to each sample or each sample location. For example, the first prediction includes a first candidate predicted value of a sample in the current block, and the second prediction includes a second candidate predicted value of a sample in the current block. The weights of the samples include a first weight of the DIMD mode and a second weight of the angular mode. A predicted sample of the sample is determined based on the sum of (first candidate predicted value×first weight) and (second candidate predicted value×second weight).

[0075] In one aspect, neighboring blocks are scanned in units of a certain block size, such as 4×4, to collect how frequently each of the candidate intra-prediction modes is applied, and then a weighting for each candidate intra-prediction mode is derived based on the frequency of use of the candidate intra-prediction mode. The neighboring blocks may include one or more units (e.g., 4×4 units), and prediction mode information (e.g., prediction mode) associated with each unit (e.g., 4×4 unit) may be stored for each unit.

[0076] As described above, the frequency at which multiple candidate intra-prediction modes are applied to code neighboring blocks of the current block may be determined by checking prediction mode information of neighboring blocks of the current block, such as prediction mode information associated with each unit in each neighboring block of the current block. The frequency at which multiple candidate intra-prediction modes are applied to code neighboring blocks of the current block may be determined by checking prediction modes used to code the neighboring blocks, such as prediction mode(s) stored for one or more units (e.g., 4x4 units) in each neighboring block. For example, the frequency of a candidate intra-prediction mode may be determined based on the number of times the candidate intra-prediction mode is stored for unit(s) in each neighboring block. A weight for one of the multiple candidate intra-prediction modes may be derived based on the determined frequency. Weights for the multiple candidate intra-prediction modes may be derived based on the determined frequency.

[0077] In another aspect, for each sample, a horizontal weighting (or horizontal weight) for each candidate intra-prediction mode is derived, a vertical weighting (vertical weight) for each candidate intra-prediction mode is also derived, and a final weighting is derived as a weighted sum of the horizontal weighting and the vertical weighting. For example, a horizontal weight for one of the multiple candidate intra-prediction modes is derived, and a vertical weight for one of the multiple candidate intra-prediction modes is derived. The weight (or final weighting) of one of the multiple candidate intra-prediction modes may be derived based on the derived horizontal weight and the derived vertical weight.

[0078] In one aspect, the weights for horizontal weighting and vertical weighting (e.g., a first weight for horizontal weighting and a second weight for vertical weighting, as described below) are derived based on the relative coordinates of the sample with respect to the top-left coordinate in the current block (e.g., (X, Y) in FIG. 9).

[0079] The current block (801) of Figure 8 is redrawn in Figure 9. For clarity, the template (802) is not shown in Figure 9. Referring to Figure 9, the current block (801) includes a sample (820). The sample (820) has relative coordinates (X, Y) with respect to the top-left coordinate (0, 0) within the current block (801). Based on the relative coordinates (X, Y) of the sample (820), a first weight for the horizontal weight and a second weight for the vertical weight are derived. A weight for one of a plurality of candidate intra-prediction modes may be derived based on a weighted sum of the derived horizontal weight and the derived vertical weight using the first weight and the second weight, respectively.

[0080] In another embodiment, the weights for the horizontal and vertical weights are derived using bilinear interpolation between the horizontal weight, the vertical weight, the default horizontal weight, and the default vertical weight, as shown in FIG. Hor and the vertical weighting is W Ver If , the final weighting (eg, Weight in equation (1)) is derived as follows:

[0081]

number

[0082] As described in equation (1), the weight of one of the multiple candidate intra prediction modes is the horizontal weight W Hor , vertical weight W Ver , default horizontal weight W DefaultH and the default vertical weight W DefaultVcan be derived based on bilinear interpolation of

[0083] In one aspect, the horizontal weighting for the candidate intra-prediction mode is derived based on how frequently the candidate intra-prediction mode is used to code the left-neighboring block. The more frequently the candidate intra-prediction mode is used to code the left-neighboring block, the higher the horizontal weighting. For example, the frequency with which one of the multiple candidate intra-prediction modes is applied to code the left-neighboring block in the neighboring block is determined, for example, by checking the prediction mode used to code the left-neighboring block, and the horizontal weighting of one of the multiple candidate intra-prediction modes is derived based on the determined frequency.

[0084] In one aspect, the vertical weighting for the candidate intra-prediction mode is derived based on how frequently the candidate intra-prediction mode is used to code the upper neighboring block. The more frequently the candidate intra-prediction mode is used to code the upper neighboring block, the higher the vertical weighting. For example, the frequency at which one of the multiple candidate intra-prediction modes is applied to code the upper neighboring block in the neighboring block is determined, for example, by checking the prediction mode used to code the upper neighboring block, and the vertical weighting of one of the multiple candidate intra-prediction modes is derived based on the determined frequency.

[0085] When two or more intra-prediction modes are used to generate the final intra-predicted block, i.e., when intra-prediction fusion is applied, for each sample, a weighting for each candidate prediction value is derived based on neighboring reconstructed samples. Each intra-prediction mode involved in intra-prediction fusion may be referred to as a candidate intra-prediction mode.

[0086] As described above, when a current block is predicted using multiple candidate intra-prediction modes, such as using intra-prediction fusion or fusion of multiple intra-prediction modes, a candidate predicted value for each of the samples in the current block may be determined for each of the multiple candidate intra-prediction modes. A weight for each of the multiple candidate intra-prediction modes (or a weighting for each candidate predicted value) may be derived based on neighboring reconstructed samples of the current block. Samples in the current block may be predicted according to the derived weights by a weighted sum of the candidate predicted values ​​associated with the multiple candidate intra-prediction modes. The neighboring reconstructed samples may include reconstructed samples within at least one line of the current block.

[0087] In one aspect, a histogram for edge directions is calculated based on neighboring reconstructed samples in a manner similar to the DIMD or TIMD method, and then weights for candidate intra-prediction modes associated with the edge directions are derived based on the frequency of each edge direction in the histogram. In one example, a histogram of edge directions for a current block is calculated based on neighboring reconstructed samples of the current block, similar to the calculation of HoG (810) used in DIMD as described in FIG. 8. The histogram of edge directions can indicate the frequency of each edge direction. A weight for one of the multiple candidate intra-prediction modes can be derived based on the frequency of one of the edge directions, where one of the multiple candidate intra-prediction modes is associated with one of the edge directions. In one example, similar to TIMD, a template matching cost is calculated between a current template including neighboring reconstructed samples of the current block and each template of the current template indicated by the multiple candidate intra-prediction modes. The weights for the multiple candidate intra-prediction modes can be derived based on their respective template matching costs.

[0088] The same method described with reference to FIG. 9 can be applied to deriving horizontal and vertical weights, for example, when neighboring reconstructed samples of the current block are used to derive weights for each of the multiple candidate intra-prediction modes. Instead of scanning how frequently each candidate intra-prediction mode is used in neighboring blocks, the frequency of the candidate intra-prediction directions reflected by histograms of edge directions using left and top neighboring reconstructed samples, respectively, is used. In one example, the histogram of the left edge directions of the current block is calculated based on the left-neighboring reconstructed samples among the neighboring reconstructed samples of the current block, and the histogram of the left edge directions indicates the frequency of each of the left edge directions. The horizontal weight of one of the multiple candidate intra-prediction modes is derived based on the frequency of one of the left edge directions, where one of the multiple candidate intra-prediction modes is associated with one of the left edge directions. In one example, the histogram of the top edge directions of the current block is calculated based on the top-neighboring reconstructed samples among the neighboring reconstructed samples of the current block, and the histogram of the top edge directions indicates the frequency of each of the top edge directions. The vertical weight of one of the plurality of candidate intra prediction modes may be derived based on the frequency of one of the top edge directions that one of the plurality of candidate intra prediction modes is associated with one of the top edge directions.

[0089] 10 shows a flowchart outlining a process (1000) according to an embodiment of the present disclosure. The process (1000) may be used in a video decoder. In various embodiments, the process (1000) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), or the like. In some embodiments, the process (1000) is implemented with software instructions, and thus, the processing circuit performs the process (1000) when it executes the software instructions. The process begins at (S1001) and proceeds to (S1010).

[0090] At (S1010), a current block in a bitstream is received. The current block is predicted using a plurality of candidate intra-prediction modes.

[0091] At (S1020), for each of a plurality of candidate intra-prediction modes, a respective candidate prediction value for the samples in the current block is determined.

[0092] In step S1030, weights for each of a plurality of candidate intra-prediction modes can be derived based on the intra-prediction modes used to code neighboring blocks of the current block.

[0093] In one example, a frequency at which multiple candidate intra-prediction modes are applied to code neighboring blocks is determined, and a weight for one of the multiple candidate intra-prediction modes is derived based on the determined frequency.

[0094] In one example, a horizontal weight for one of the plurality of candidate intra-prediction modes is derived, and a vertical weight for one of the plurality of candidate intra-prediction modes is derived. The weight for one of the plurality of candidate intra-prediction modes may be derived based on the derived horizontal weight and the derived vertical weight.

[0095] At (S1040), samples in the current block may be predicted by a weighted sum of candidate prediction values ​​associated with multiple candidate intra-prediction modes according to the derived weights.

[0096] Then, the process proceeds to (S1099) and ends.

[0097] The process 1000 may be adapted as appropriate. Steps of the process 1000 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0098] 11 shows a flowchart outlining a process (1100) according to one embodiment of the present disclosure. The process (1100) may be used in a video encoder. In various embodiments, the process (1100) is performed by a processing circuit, such as a processing circuit that performs the functions of the video encoder (103), a processing circuit that performs the functions of the video encoder (303), or the like. In some embodiments, the process (1100) is implemented with software instructions, and thus, the processing circuit performs the process (1100) when it executes the software instructions. The process begins at (S1101) and proceeds to (S1110).

[0099] At (S1110), if the current block is predicted with multiple candidate intra-prediction modes, a candidate predicted value for each of the samples in the current block is determined for each of the multiple candidate intra-prediction modes.

[0100] In step S1120, weights for each of a plurality of candidate intra-prediction modes can be derived based on the intra-prediction modes used to code neighboring blocks of the current block.

[0101] At (S1130), samples in the current block may be predicted by a weighted sum of candidate prediction values ​​associated with multiple candidate intra-prediction modes according to the derived weights.

[0102] Then, the process proceeds to (S1199) and ends.

[0103] The process 1100 may be adapted as appropriate. Steps of the process 1100 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0104] 12 shows a flowchart outlining a process (1200) according to one embodiment of the present disclosure. The process (1200) may be used in a video decoder. In various embodiments, the process (1200) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), or the like. In some embodiments, the process (1200) is implemented with software instructions, and thus, the processing circuit performs the process (1200) when it executes the software instructions. The process begins at (S1201) and proceeds to (S1210).

[0105] At (S1210), a current block in a bitstream is received. The current block is predicted using a plurality of candidate intra-prediction modes.

[0106] At (S1220), a candidate predicted value for each of the samples in the current block for each of a plurality of candidate intra-prediction modes is determined.

[0107] At (S1230), weights for each of the plurality of candidate intra-prediction modes can be derived based on neighboring reconstructed samples of the current block, where the neighboring reconstructed samples include reconstructed samples in at least one line of the current block.

[0108] A histogram of edge directions of the current block may be calculated based on neighboring reconstructed samples of the current block, the histogram of edge directions indicating a frequency of the edge directions, and a weight of one of the plurality of candidate intra-prediction modes based on a frequency of one of the edge directions, wherein one of the plurality of candidate intra-prediction modes is associated with the one of the edge directions.

[0109] In one example, a template matching cost is calculated between a current template including neighboring reconstructed samples of the current block and each of the current templates indicated by a plurality of candidate intra-prediction modes, and weights of the plurality of candidate intra-prediction modes may be derived based on the respective template matching costs.

[0110] In one example, a horizontal weight for one of the plurality of candidate intra-prediction modes is derived, and a vertical weight for one of the plurality of candidate intra-prediction modes is derived. The weight for one of the plurality of candidate intra-prediction modes may be derived based on the derived horizontal weight and the derived vertical weight.

[0111] In one embodiment, a histogram of left edge directions of the current block is calculated based on left-neighboring reconstructed samples of the neighboring reconstructed samples of the current block. The histogram of left edge directions indicates a frequency of the left edge directions. A horizontal weight of one of the plurality of candidate intra prediction modes is derived based on the frequency of one of the left edge directions. One of the plurality of candidate intra prediction modes is associated with one of the left edge directions.

[0112] In one example, a histogram of top edge directions of the current block is calculated based on the top neighboring reconstructed samples among the neighboring reconstructed samples of the current block. The histogram of top edge directions indicates the frequency of the top edge directions. A vertical weight of one of the plurality of candidate intra-prediction modes is derived based on the frequency of one of the top edge directions. One of the plurality of candidate intra-prediction modes is associated with one of the top edge directions.

[0113] In one embodiment, a first weight for a horizontal weight and a second weight for a vertical weight are derived based on the relative coordinates of the samples with respect to the top-left coordinate of the current block, and a weight for one of the plurality of candidate intra-prediction modes is derived based on a weighted sum of the derived horizontal weight and the derived vertical weight using the first weight and the second weight, respectively.

[0114] In one example, a weight for one of the multiple candidate intra-prediction modes is derived based on bilinear interpolation between a horizontal weight, a vertical weight, a default horizontal weight, and a default vertical weight.

[0115] At (S1240), samples in the current block may be predicted by a weighted sum of candidate prediction values ​​associated with multiple candidate intra-prediction modes according to the derived weights.

[0116] Then, the process proceeds to (S1299) and ends.

[0117] Process 1200 may be adapted as appropriate. Steps of process 1200 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0118] 13 shows a flowchart outlining a process (1300) according to one embodiment of the present disclosure. The process (1300) may be used in a video encoder. In various embodiments, the process (1300) is performed by a processing circuit, such as a processing circuit that performs the functions of the video encoder (103), a processing circuit that performs the functions of the video encoder (303), or the like. In some embodiments, the process (1300) is implemented with software instructions, and thus, the processing circuit performs the process (1300) when it executes the software instructions. The process starts at (S1301) and proceeds to (S1310).

[0119] At (S1310), if the current block is predicted with multiple candidate intra-prediction modes, a candidate predicted value for each of the samples in the current block is determined for each of the multiple candidate intra-prediction modes.

[0120] At (S1320), weights for each of the plurality of candidate intra-prediction modes can be derived based on neighboring reconstructed samples of the current block, where the neighboring reconstructed samples include reconstructed samples in at least one line of the current block.

[0121] At (S1330), samples in the current block may be predicted by a weighted sum of candidate prediction values ​​associated with multiple candidate intra-prediction modes according to the derived weights.

[0122] Then, the process proceeds to (S1399) and ends.

[0123] Process 1300 may be adapted as appropriate. Steps of process 1300 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0124] The aspects, embodiments, and / or examples in this disclosure may be used separately or combined in any order. Each of the methods (or aspects), encoders, and decoders may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0125] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 14 illustrates a computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter.

[0126] Computer software can be coded using any suitable machine code or computer language, which can be assembled, compiled, linking, or similar mechanisms to produce code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly, or via interpretation, microcode execution, etc.

[0127] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0128] 14 for computer system 1400 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system 1400.

[0129] The computer system 1400 may include certain human interface input devices that may respond to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).

[0130] The input human interface devices may include one or more (only one of each is shown) of a keyboard (1401), a mouse (1402), a trackpad (1403), a touchscreen (1410), a data glove (not shown), a joystick (1405), a microphone (1406), a scanner (1407), and a camera (1408).

[0131] The computer system (1400) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses through, for example, haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1410), data gloves (not shown), or joystick (1405), although some haptic feedback devices may not function as input devices), audio output devices (such as speakers (1409), headphones (not shown)), visual output devices (such as screens (1410), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capability and each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown).

[0132] The computer system (1400) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1420) with media (1421) such as CD / DVD, thumb drives (1422), removable hard drives or solid state drives (1423), legacy magnetic media such as tape and floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0133] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not encompass transmission media, carrier waves, or other transitory signals.

[0134] The computer system (1400) may also include an interface (1454) to one or more communication networks (1455). The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, WLAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide-area digital networks including cable television, satellite television, and terrestrial broadcast television, and vehicular and industrial networks including CANbus. Some networks generally require an external network interface adapter attached to some general-purpose data port or peripheral bus (1449) (e.g., a USB port on the computer system (1400)), while other networks are generally integrated into the core of the computer system (1400) by attachment to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system), as described below. Using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., from the CANbus to a specific CANbus device), or two-way, for example, to other computer systems using local or wide-area digital networks. Certain protocols and protocol stacks can be used over each of these networks and network interfaces, as described above.

[0135] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1440) of the computer system (1400).

[0136] The core (1440) may include one or more central processing units (CPUs) (1441), graphics processing units (GPUs) (1442), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1443), hardware accelerators for specific tasks (1444), graphics adapters (1450), etc. These devices may be connected via a system bus (1448), along with read-only memory (ROM) (1445), random access memory (1446), and internal mass storage devices (1447), such as internal non-user-accessible hard drives or SSDs. In some computer systems, the system bus (1448) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (1448) or via a peripheral bus (1449). In one example, a screen (1410) may be connected to the graphics adapter (1450). Peripheral bus architectures include PCI, USB, etc.

[0137] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute certain instructions, which in combination can constitute the aforementioned computer code. The computer code can be stored in ROM (1445) or RAM (1446). Also, transient data can be stored in RAM (1446), while persistent data can be stored in, for example, internal mass storage device (1447). Rapid storage and retrieval from any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more of the CPU (1441), GPU (1442), mass storage device (1447), ROM (1445), RAM (1446), etc.

[0138] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0139] By way of example and not limitation, the architecture (1400), and in particular a computer system having a core (1440), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage devices, as introduced above, as well as media associated with specific storage devices of the core (1440) that are non-transitory in nature, such as the core's internal mass storage device (1447) or ROM (1445). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1440). The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause the core (1440), and in particular the processor (including a CPU, GPU, FPGA, etc.) therein, to perform specific processes or portions of specific processes described herein, including defining data structures stored in RAM (1446) and modifying such data structures in accordance with the software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1444)), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software may encompass logic, where appropriate, and vice versa. References to computer-readable media may encompass, where appropriate, circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any appropriate combination of hardware and software.

[0140] The use of "at least one" or "one" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of the listed elements, where applicable, such as when the elements are not mutually exclusive.

[0141] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore are within the spirit and scope of the present disclosure.

Claims

1. 1. A method of video decoding, comprising: receiving a current block in a bitstream, the current block being predicted using intra prediction fusion including a plurality of candidate intra prediction modes; determining a candidate predicted value for each of the samples in the current block in each of the plurality of candidate intra-prediction modes; deriving a weight for each of the plurality of candidate intra-prediction modes based on intra-prediction modes used to code neighboring blocks of the current block; predicting the samples in the current block by a weighted sum of the candidate prediction values ​​associated with the plurality of candidate intra-prediction modes according to the derived weights; A method comprising:

2. The derivation is determining a frequency at which the plurality of candidate intra-prediction modes are applied to code the neighboring blocks; and and deriving a weight for one of the plurality of candidate intra-prediction modes based on the determined frequency.

3. The derivation is deriving a horizontal weight for one of the plurality of candidate intra-prediction modes; deriving a vertical weight for the one of the plurality of candidate intra-prediction modes; deriving a weight for the one of the plurality of candidate intra-prediction modes based on the derived horizontal weight and the derived vertical weight; and The method of claim 1 , comprising:

4. Deriving the horizontal weights comprises: determining a frequency with which the one of the plurality of candidate intra-prediction modes is applied to code a left-neighboring block among the neighboring blocks; deriving the horizontal weight of the one of the plurality of candidate intra-prediction modes based on the determined frequency; and The method of claim 3, comprising:

5. Deriving the vertical weights comprises: determining a frequency at which the one of the plurality of candidate intra-prediction modes is applied to code an upper-neighboring block among the neighboring blocks; deriving the vertical weight of the one of the plurality of candidate intra-prediction modes based on the determined frequency; and The method of claim 3, comprising:

6. Deriving the weights may include: deriving a first weight for the horizontal weight and a second weight for the vertical weight based on a relative coordinate of the sample with respect to a top-left coordinate within the current block; deriving the weight for the one of the plurality of candidate intra-prediction modes based on a weighted sum of the derived horizontal weight and the derived vertical weight using the first weight and the second weight, respectively; The method of claim 3, comprising:

7. Deriving the weights deriving the weight for the one of the plurality of candidate intra-prediction modes based on bilinear interpolation of the horizontal weight, the vertical weight, a default horizontal weight, and a default vertical weight. The method of claim 3.

8. the plurality of candidate intra-prediction modes include one or more of a DC mode, a planar mode, an intra-directional prediction mode, a decoder-side intra-mode derivation (DIMD) mode, a template-based intra-mode derivation (TIMD) mode, a component-to-component linear model (CCLM), a convolutional component-to-component model (CCCM), and a multi-model linear mode (MMLM); The method of claim 1.

9. 1. A method of video decoding, comprising: receiving a current block in a bitstream including coding information indicating that the current block is predicted using intra prediction fusion including a plurality of candidate intra prediction modes; determining a candidate predicted value for each of the plurality of candidate intra-prediction modes for each of the samples in the current block; deriving weights for each of the plurality of candidate intra-prediction modes based on neighboring reconstructed samples of the current block, the neighboring reconstructed samples including reconstructed samples within at least one line of the current block; and predicting the samples in the current block by a weighted sum of the candidate prediction values ​​associated with the plurality of candidate intra-prediction modes according to the derived weights.

10. The derivation is calculating a histogram of edge directions of the current block based on the neighboring reconstructed samples of the current block, the histogram of edge directions indicating a frequency of each of the edge directions; deriving a weight for one of the plurality of candidate intra-prediction modes based on a frequency of one of the edge directions, the one of the plurality of candidate intra-prediction modes being associated with the one of the edge directions; 10. The method of claim 9.

11. The above derivation calculating a template matching cost between a current template including the neighboring reconstructed samples of the current block and each of the current templates indicated by the plurality of candidate intra-prediction modes; deriving the weights for the plurality of candidate intra-prediction modes based on the template matching cost.

10. The method of claim 9.

12. The derivation is deriving a horizontal weight for one of the plurality of candidate intra-prediction modes; deriving a vertical weight for the one of the plurality of candidate intra-prediction modes; deriving a weight for the one of the plurality of candidate intra-prediction modes based on the derived horizontal weight and the derived vertical weight; and 10. The method of claim 9, comprising:

13. Deriving the horizontal weights includes: calculating a histogram of left edge directions of the current block based on left-side adjacent reconstructed samples among the adjacent reconstructed samples of the current block, wherein the histogram of left edge directions indicates a frequency of each of the left edge directions; deriving the horizontal weight for the one of the plurality of candidate intra-prediction modes based on a frequency of one of the left edge directions, the one of the plurality of candidate intra-prediction modes being associated with the one of the left edge directions.

13. The method of claim 12, comprising:

14. Deriving the vertical weights comprises: Calculating a histogram of top edge directions of the current block based on top-neighboring reconstructed samples among the neighboring reconstructed samples of the current block, wherein the histogram of top edge directions indicates a frequency of each of the top edge directions; deriving the vertical weight for the one of the plurality of candidate intra prediction modes based on a frequency of one of the top edge directions, wherein the one of the plurality of candidate intra prediction modes is associated with the one of the top edge directions. The method of claim 12.

15. Deriving the weights deriving a first weight for the horizontal weight and a second weight for the vertical weight based on a relative coordinate of the sample with respect to a top-left coordinate in the current block; deriving the weight for the one of the plurality of candidate intra-prediction modes based on a weighted sum of the derived horizontal weight and the derived vertical weight using the first weight and the second weight, respectively; 13. The method of claim 12, comprising:

16. Deriving the weights may include: deriving the weight for the one of the plurality of candidate intra-prediction modes based on bilinear interpolation between the horizontal weight, the vertical weight, a default horizontal weight, and a default vertical weight. The method of claim 12.

17. 1. An apparatus for video decoding, comprising: a processing circuit, receiving a current block in a bitstream, the current block being predicted using intra prediction fusion including a plurality of candidate intra prediction modes; determining a candidate predicted value for each of the plurality of candidate intra-prediction modes for each of the samples in the current block; deriving a weight for each of the plurality of candidate intra prediction modes based on an intra prediction mode used to code a neighboring block of the current block; predicting the samples in the current block by a weighted sum of the candidate prediction values ​​associated with the plurality of candidate intra-prediction modes according to the derived weights; The apparatus is configured to:

18. The processing circuitry determining a frequency at which the plurality of candidate intra-prediction modes are applied to code the neighboring blocks; configured to derive a weight for one of the plurality of candidate intra-prediction modes based on the determined frequency.

18. The apparatus of claim 17.

19. The processing circuitry deriving a horizontal weight for one of the plurality of candidate intra-prediction modes; deriving a vertical weight for the one of the plurality of candidate intra-prediction modes; deriving the weight for the one of the plurality of candidate intra-prediction modes based on the derived horizontal weight and the derived vertical weight.

18. The apparatus of claim 17, configured to:

20. the plurality of candidate intra-prediction modes include one or more of a DC mode, a planar mode, an intra-directional prediction mode, a decoder-side intra-mode derivation (DIMD) mode, a template-based intra-mode derivation (TIMD) mode, a component-to-component linear model (CCLM), a convolutional component-to-component model (CCCM), and a multi-model linear mode (MMLM); 18. The apparatus of claim 17.

Citation Information

Patent Citations

  • Video signal encoding / decoding method and device for said method

    JP2022505874A