Method, apparatus, and computer program for video decoding
By selecting intra-interpolation filters based on neighboring reconstructed samples and angular intra-prediction modes, the video coding process is enhanced, addressing inefficiencies in existing technologies and improving compression and decoding performance.
Patent Information
- Application Number
- JP2025539636
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-27
- Filing Date
- 2023-10-30
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2043-10-30
AI Technical Summary
Existing video coding technologies face challenges in efficiently utilizing intra-prediction modes, particularly angular intra-prediction, due to limitations in selecting optimal interpolation filters, leading to suboptimal video compression and decoding performance.
The implementation of a processing circuit that selects an intra-interpolation filter from a predetermined set based on neighboring reconstructed samples, considering prediction errors and angular intra-prediction modes, to enhance the prediction accuracy and efficiency of video encoding and decoding.
Improves video compression efficiency by optimizing the selection of intra-interpolation filters, resulting in better prediction accuracy and reduced data volume in video bitstreams.
Smart Images

Figure 2026501685000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 444,877, entitled "Self-guided Intra interpolation filter," filed February 10, 2023, which claims the benefit of priority to U.S. Patent Application No. 18 / 384,759, entitled "SELF-GUIDED INTRA INTERPOLATION FILTER," filed October 27, 2023. The disclosures of these prior applications are incorporated herein by reference in their entireties.
[0002] This disclosure describes aspects generally related to video coding. [Background technology]
[0003] The background discussion provided herein is intended to provide a general overview of the context for the disclosure. To the extent described in this background section, the work of the named inventors, and aspects of the disclosure that may not otherwise qualify as prior art at the time of filing, are not admitted, explicitly or implicitly, as prior art to the present disclosure.
[0004] Image / video compression can help transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from a current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention
[0005] Aspects of the present disclosure include methods and apparatuses for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. In one example, the processing circuit receives, from a bitstream having a current block in a picture, coding information in the bitstream indicating that the current block is coded in an angular intra-prediction mode using intra-interpolation filters. The processing circuit applies each of a predetermined set of intra-interpolation filters to neighboring reconstructed samples within N neighboring lines from a boundary of the current block, where N is a positive integer. The processing circuit selects one intra-interpolation filter from the predetermined set of intra-interpolation filters based on a prediction error associated with each of the intra-interpolation filters, and predicts samples in the current block using the angular intra-prediction mode using the selected intra-interpolation filter.
[0006] In one example, a processing circuit receives a bitstream of a current block in a picture. Coding information in the bitstream indicates that the current block is coded in an angular intra-prediction mode using an intra-interpolation filter. The processing circuit selects an intra-interpolation filter from a predetermined set of intra-interpolation filters based on neighboring reconstructed samples of the current block. The neighboring reconstructed samples include reconstructed samples within N lines from one or more boundaries of the current block. The processing circuit applies the selected intra-interpolation filter to reference samples in reference lines in the picture to predict samples in the current block using the angular intra-prediction mode and reconstructs the current block based on the predicted samples.
[0007] In one example, the processing circuitry selects one of (i) the type of intra-interpolation filter or (ii) the number of taps in the intra-interpolation filter from a predetermined set of intra-interpolation filters based on neighboring reconstructed samples of the current block.
[0008] In one example, the predetermined set of intra-interpolation filters includes different types of intra-interpolation filters, and the processing circuit selects a type of intra-interpolation filter from the different types of intra-interpolation filters based on neighboring reconstructed samples of the current block, and the selected intra-interpolation filter is one of a bilinear interpolation filter, a cubic interpolation filter, a spline interpolation filter, a DCT-based interpolation filter, or a DST-based interpolation filter.
[0009] In one example, the predetermined set of intra-interpolation filters includes different tap numbers, and the processing circuit selects the tap number of the intra-interpolation filter based on neighboring reconstructed samples of the current block, where the tap number of the intra-interpolation filter is one of 2 taps, 4 taps, 6 taps, or 8 taps.
[0010] In one example, N is greater than 1, and the neighboring reconstructed samples include neighboring reconstructed samples of multiple lines. For each combination of an intra-interpolation filter from the predetermined set of intra-interpolation filters and neighboring reconstructed samples of a line from the multiple lines of neighboring reconstructed samples, the processing circuit predicts the neighboring reconstructed samples of the respective line based on one or more remaining lines from the multiple lines of neighboring reconstructed samples using the respective intra-interpolation filter to obtain at least one prediction error, and selects the intra-interpolation filter corresponding to the smallest prediction error from the obtained prediction errors.
[0011] In one example, whether the neighboring reconstructed samples include (i) neighboring reconstructed samples of the left line to the left of the current block, (ii) neighboring reconstructed samples of the line above the current block, or (iii) neighboring reconstructed samples of the left line to the left of the current block and neighboring reconstructed samples of the line above the current block depends on the intra prediction direction of the angular intra prediction mode.
[0012] In one example, the value of N depends on the block size.
[0013] In one example, the processing circuit predicts neighboring reconstructed samples of each line based on the intra prediction direction of the angular intra prediction mode.
[0014] In one example, the processing circuit predicts neighboring reconstructed samples of each line based on a direction opposite to the intra-prediction direction of the angular intra-prediction mode.
[0015] In one example, the processing circuit selects an intra-interpolation filter from a predetermined set of intra-interpolation filters and an angular intra-prediction mode from a plurality of angular intra-prediction modes based on neighboring reconstructed samples of the current block.
[0016] In one example, the processing circuit derives N1 angular intra prediction modes from the plurality of angular intra prediction modes associated with the N1 lowest cost values from the cost values of the respective angular intra prediction modes. Each cost value is determined based on neighboring reconstructed samples of the current block, the respective angular intra prediction modes, and a default intra interpolation filter from a predetermined set of intra interpolation filters. For each of the N1 angular intra prediction modes, the processing circuit selects M1 intra interpolation filters from the predetermined set of intra interpolation filters associated with the M1 lowest cost values from the cost values of the predetermined set of intra interpolation filters. Each cost value is determined based on neighboring reconstructed samples of the current block, the respective intra interpolation filters, and a respective angular intra prediction mode from the N1 angular intra prediction modes. The processing circuit selects one intra interpolation filter and angular intra prediction mode from the N1×M1 combinations. Each combination of the N1×M1 combinations includes an intra-interpolation filter from among the M1 intra-interpolation filters and an angular intra-prediction mode from among the N1 angular intra-prediction modes.
[0017] In one example, the processing circuit selects N2 angular intra prediction modes from a first plurality of modes of the plurality of angular intra prediction modes based on neighboring reconstructed samples of the current block and the intra interpolation filter. The processing circuit selects a first updated intra interpolation filter based on neighboring reconstructed samples of the current block and one of the N2 angular intra prediction modes. The processing circuit selects N3 angular intra prediction modes from a second plurality of modes of the plurality of angular intra prediction modes based on neighboring reconstructed samples of the current block and the first updated intra interpolation filter. The processing circuit selects a second updated intra interpolation filter based on neighboring reconstructed samples of the current block and one of the N3 angular intra prediction modes. The processing circuit selects the second updated intra-interpolation filter as the intra-interpolation filter to be selected and selects one of the N3 angular intra-prediction modes as the angular intra-prediction mode to be selected based on (i) a difference between a first prediction error associated with the first updated intra-interpolation filter and one of the N2 directional intra-prediction modes and a second prediction error associated with the second updated intra-interpolation filter and one of the N3 angular intra-prediction modes, or (ii) the second prediction error.
[0018] In one example, for each combination of an angular intra prediction mode from among the plurality of angular intra prediction modes and an intra interpolation filter from a predetermined set of intra interpolation filters, the processing circuit determines a cost value for each combination based on neighboring reconstructed samples of the current block, selects K combinations based on the determined cost values, and selects an intra interpolation filter and an angular intra prediction mode from the K combinations based on index information signaled in the bitstream.
[0019] In one example, the processing circuit calculates feature values based on neighboring reconstructed samples of the current block, and selects one intra-interpolation filter based on the calculated feature values.
[0020] In one example, the processing circuit calculates the feature value as one of (i) an absolute gradient value associated with neighboring reconstructed samples of the current block, or (ii) a difference between the minimum value among the neighboring reconstructed samples of the current block and the maximum value among the neighboring reconstructed samples of the current block.
[0021] In one example, the processing circuitry obtains an output from a neural network, the inputs of which include neighboring reconstructed samples of the current block and the output of which indicates an intra-interpolation filter to select, and the inputs of the neural network further include an intra-prediction mode index indicating an angular intra-prediction mode.
[0022] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video encoding / decoding. [Brief explanation of the drawings]
[0023] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100). [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4] 1 illustrates nine predictor directions from the possible predictor directions according to one aspect of the present disclosure. [Figure 5] 1 illustrates intra-prediction directions according to one aspect of the present disclosure. [Figure 6] 1 shows eight nominal angles according to one embodiment of the present disclosure: V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED. [Figure 7]1 illustrates an example of a non-directional intra predictor according to one aspect of the present disclosure. [Figure 8] 1 illustrates an example of a recursive intra-filtering mode according to an aspect of the present disclosure. [Figure 9] 1 shows an example of multi-reference line (MRL) intra prediction. [Figure 10] 10 illustrates an example angular intra prediction using reference line 0 adjacent to the block. [Figure 11] 1 illustrates an example of a current block coded using an angular intra prediction mode and neighboring reconstructed samples of the current block according to one aspect of the present disclosure. [Figure 12] 1 shows a flowchart outlining a decoding process according to some aspects of the present disclosure. [Figure 13] 1 shows a flowchart outlining an encoding process according to some aspects of the present disclosure. [Figure 14] 1 shows a flowchart outlining a decoding process according to some aspects of the present disclosure. [Figure 15] FIG. 1 is a schematic diagram of a computer system according to one aspect. DETAILED DESCRIPTION OF THE INVENTION
[0024] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is an example of an application of the disclosed subject matter, which is a video encoder and video decoder in a streaming environment. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, and storage of compressed video on digital media, including CDs, DVDs, memory sticks, and the like.
[0025] The video processing system 100 may include a capture subsystem 113, which may include a video source 101, such as a digital camera, that produces a stream of uncompressed video pictures 102. In one example, the stream of video pictures 102 includes samples captured by the digital camera. The stream of video pictures 102 is depicted as a thick line to emphasize its high data volume compared to the encoded video data 104 (or coded video bitstream), which may be processed by electronics 120 including a video encoder 103 coupled to the video source 101. The video encoder 103 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data 104 (or coded video bitstream), depicted as a thin line to emphasize its low data volume compared to the stream of video pictures 102, may be stored on a streaming server 105 for later use. One or more streaming client subsystems, such as the client subsystems 106 and 108 of FIG. 1, can access the streaming server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 can include a video decoder 110, for example, within an electronic device 130. The video decoder 110 can decode the incoming copy of the encoded video data 107 and produce an outgoing stream of video pictures 111, which can be rendered on a display 112 (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data 104, 107, and 109 (e.g., a video bitstream) can be encoded according to a particular video coding / compression standard.Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The subject matter disclosed herein may be used in the context of VVC.
[0026] It should be noted that the electronics 120 and 130 may include other components (not shown). For example, the electronics 120 may include a video decoder (not shown), and the electronics 130 may also include a video encoder (not shown).
[0027] 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used in place of the video decoder (110) in the example of FIG. 1.
[0028] The receiver (231) can receive one or more coded video sequences, e.g., in a bitstream, to be decoded by the video decoder (210). In one aspect, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences can be received from a channel (201), which can be a hardware / software link to a storage device that stores the coded video data. The receiver (231) can also receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which can be forwarded to their respective using entities (not shown). The receiver (231) can separate the coded video sequences from other data. To combat network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser 520 (hereinafter, "parser (220)"). In certain applications, the buffer memory (215) is part of the video decoder (210). In others, it may be external to the video decoder 210 (not shown). In still others, there may be a buffer memory (not shown) external to the video decoder 210, e.g., to combat network jitter, and there may be another buffer memory 215 internal to the video decoder 210, e.g., to handle playback timing. When the receiver 231 is receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an equally synchronous network, the buffer memory 215 may not be required or may be small. For use over best-effort packet networks, such as the Internet, the buffer memory 215 may be required and may be relatively large and advantageously sized adaptively, and may be implemented, at least in part, in an operating system or similar element (not shown) external to the video decoder 210.
[0029] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and possibly information for controlling a rendering device, such as a render device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but can be coupled to the electronic device (230), as shown in FIG. 2. The control information for the rendering device(s) may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be according to a video coding technique or standard and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) may also extract information from the coded video sequence information, such as transform coefficients, quantization parameter values, motion vectors, etc.
[0030] The parser (220) may perform an entropy decoding / parsing process on the video sequence received from the buffer memory (215) to produce symbols (221).
[0031] The reconstruction of the symbols (221) may involve several different units, depending on the type of video picture or portion thereof being coded and other factors (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.). Which units are involved and how they are involved can be controlled by subgroup control information parsed by the parser (220) from the coded video sequence. The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.
[0032] Beyond the functional blocks already described, the video decoder (210) may be conceptually subdivided into a number of functional units, as described below. In practical implementations operating within commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the subject matter of this disclosure, the following conceptual division into functional units is appropriate:
[0033] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as symbol(s) (221) from the parser (220), along with control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. The scalar / inverse transform unit (251) can output blocks of sample values that can be input to an aggregator (255).
[0034] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258), for example, buffers partially reconstructed and / or fully reconstructed current pictures. In some cases, the aggregator (255) adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).
[0035] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to a block that may be inter-coded and motion-compensated. In such cases, the motion-compensated prediction unit (253) may access the reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) related to the block, these samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches prediction samples may be controlled by a motion vector and are available to the motion-compensated prediction unit (253) in the form of symbols (221), which may have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.
[0036] The output samples of the aggregator (255) may be subjected to various loop filtering techniques in a loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the coded video sequence (also referred to as a coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression may also be responsive to meta-information obtained during decoding of a previous portion (in decoding order) of the coded picture or coded video sequence, as well as to previously reconstructed and loop-filtered sample values.
[0037] The output of the loop filter unit (256) can be a sample stream that can be output to a render device (212), which can also be stored in a reference picture memory (257) for use in future inter-picture prediction.
[0038] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and that coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before beginning reconstruction of the next coded picture.
[0039] The video decoder (210) may perform decoding according to a given video compression technology or standard, such as ITU-T Recommendation H.265. The coded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense of adhering to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile may select specific tools from all tools available in the video compression technology or standard, such that only those tools are available for use under that profile. Compliance also requires that the complexity of the coded video sequence be within a range specified by the level of the video compression technology or standard. In some cases, the level constrains the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may optionally be further constrained through a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0040] In one aspect, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of one or more coded video sequences. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0041] 3 shows an exemplary block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). For example, the electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.
[0042] The video encoder (303) may receive video samples from a video source (301) (not part of the electronics (320) in the example of FIG. 6) that may capture video image(s) to be coded by the video encoder (303). In another example, the video source (301) is part of the electronics (320).
[0043] The video source (301) may provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service provision system, the video source (301) may be a storage device containing pre-prepared video. In a video conferencing system, the video source (301) may be a camera capturing local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may have one or more samples, depending on the sampling structure, color space, etc. used. The following discussion focuses on samples.
[0044] According to one aspect, the video encoder (303) may code and compress pictures of a source video sequence into a coded video sequence (343) in real time or under other required time constraints. Enforcing an appropriate coding rate is one function of the controller (350). In some aspects, the controller (350) controls and is operatively coupled to other functional units, such as those described below, which are not shown for clarity. Parameters set by the controller (350) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured with other suitable functions associated with the video encoder (303) that are optimized for a particular system design.
[0045] In some aspects, the video encoder (303) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop can include a source coder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and one or more reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data, in a manner similar to that used by a (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-accurate results independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronism (and the resulting drift when synchronism cannot be maintained, for example due to channel errors) is also used in some related art.
[0046] The operation of the "local" decoder (333) may be the same as that of a "remote" decoder, such as the video decoder (210), which has already been described in detail above in connection with Figure 2. However, briefly referring also to Figure 2, because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (345) and parser (220) may be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333).
[0047] In one aspect, decoder technology, excluding parsing / entropy decoding, present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the subject matter of the disclosure focuses on decoder operation. A description of the encoder technology can be omitted, as it is the reverse of the decoder technology that has been thoroughly described. In certain areas, more detailed descriptions are provided below.
[0048] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture relative to one or more previously coded pictures from a video sequence designated as “reference pictures.” Thus, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of one or more reference pictures that may be selected as prediction reference(s) for the input picture.
[0049] The local video decoder (333) may decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence may typically be a replica of the source video sequence, with some error. The local video decoder (333) may replicate the decoding process that may be performed by a video decoder on the reference pictures, causing the reconstructed reference pictures to be stored in the reference picture memory (334). In this way, the video encoder (303) may locally store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures that will be obtained by a far-end video decoder.
[0050] The predictor (335) may perform a predictive search for the coding engine (332). That is, for a new picture to be coded, the predictor (336) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors or block shapes, that can serve as an appropriate prediction reference for the new picture. The predictor (335) may operate pixel block by pixel block to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).
[0051] The controller (350) may manage the coding process of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0052] The outputs of all the aforementioned functional units may be subjected to entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.
[0053] A transmitter (340) may buffer the coded video sequence(s) produced by the entropy coder (345) and prepare them for transmission over a communication channel (360), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (340) may merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0054] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign each coded picture a particular coded picture type, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned one of the following picture types:
[0055] Intra-pictures (I-pictures) can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow several different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0056] Predictive pictures (P pictures) can be coded and decoded using intra- or inter-prediction, using motion vectors and reference indices to predict the sample values of each block.
[0057] Bidirectionally predicted pictures (B pictures) can be coded and decoded using intra- or inter-prediction, using two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0058] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to their respective pictures. For example, blocks of an I-picture may be coded non-predictively, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be coded non-predictively or via spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be coded non-predictively or via spatial or temporal prediction with reference to one or two previously coded reference pictures.
[0059] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. In operation, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. The coded video data may therefore conform to a syntax defined by the video coding technique or standard being used.
[0060] In one aspect, the transmitter (340) may transmit additional data along with the encoded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0061] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture being coded / decoded, called the current picture, is divided into multiple blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. A motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0062] In some aspects, bi-prediction techniques can be used in inter-picture prediction. Bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, both of which are prior to the current picture in decoding order (but may be prior and future, respectively, in display order) in a video. A block in the current picture can be coded with a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first and second reference blocks.
[0063] Furthermore, merge mode techniques can be used to improve coding efficiency in inter-picture prediction.
[0064] According to some aspects of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed on a block-by-block basis. For example, according to the HEVC standard, a picture in a sequence of video pictures is divided into multiple coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the prediction type of that CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, prediction operations during coding (encoding / decoding) are performed in units of prediction blocks. Taking a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, and the like.
[0065] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In one aspect, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another aspect, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.
[0066] Video codec techniques can include intra-coding, in which sample values can be represented without reference to samples or other data from a previously reconstructed reference picture. In some video codecs, a picture can be spatially divided into multiple blocks of samples. If all the samples of a block are coded in intra-mode, the picture can be an intra-picture (e.g., an I-picture).
[0067] In some aspects, prediction may be performed based on surrounding sample data and / or metadata obtained during encoding and / or decoding of a block of data. Such techniques are referred to as “intra-prediction” techniques. In one aspect, intra-prediction (e.g., intra-picture prediction) may use only reference data from the current picture being reconstructed, not from reference pictures.
[0068] In intra prediction, a predictor block may be formed using neighboring sample values of already available samples. The sample values of the neighboring samples may be copied to the predictor block according to a certain direction. The reference to the direction to use may be coded in the bitstream or may itself be predicted.
[0069] Various intra-prediction coding tools may be used, such as angular intra-prediction as shown in Figures 4-6 and 10, multiple reference line (MRL) prediction as shown in Figure 9, and / or the like. Non-directional intra-prediction modes may also be used.
[0070] 4 illustrates nine predictor directions from among the 33 possible predictor directions (e.g., corresponding to the 33 angular modes of the 35 intra modes as defined in H.265) according to one embodiment. The dot (401) represents the sample being predicted. The arrows represent directions from which the sample (401) can be predicted. For example, arrow (402) indicates that the sample (401) is predicted from one or more samples to the upper right and at a 45° angle from the horizontal (e.g., X-dimension). Arrow (403) indicates that the sample (401) is predicted from one or more samples to the lower left and below the sample (401), at a 22.5° angle from the horizontal.
[0071] FIG. 4 also shows a block (e.g., a square block of 4×4 samples indicated by a dashed thick line) (404). The block (404) can include 16 samples, each labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. FIG. 4 also shows reference samples with a similar numbering scheme. The reference samples are labeled with R and their Y position (e.g., row index) and X position (column index) relative to the block (404). In some aspects (e.g., in H.264 and H.265), predicted samples may be adjacent to the block being reconstructed. In some aspects, multiple reference lines may be used, as shown in FIG. 9.
[0072] Intra prediction can work by copying reference sample values from adjacent samples indicated by a prediction direction (e.g., a signaled prediction direction). For example, if a prediction direction indicated by an arrow (402) is signaled in the coded video bitstream, samples are predicted from the upper right sample at a 45° angle from horizontal. For example, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Sample S44 can be predicted from reference sample R08.
[0073] In certain cases, the values of multiple reference samples may be combined, for example, through interpolation (e.g., an intra-interpolation filter), to calculate a reference sample, as illustrated in Figure 10. For example, when the direction is not divisible by 45°, the values of multiple reference samples are combined through interpolation.
[0074] Any suitable number of possible directions may be used in intra prediction. In the example of Figure 4 (e.g., in H.264), 9 different directions are used. In an example such as in H.265, 33 different directions are used.
[0075] In some examples, such as finer-granularity angular prediction in VVC, 65 different directions are used. The 65 angular prediction directions can be used for certain block sizes, and the set of angles can depend on the block size. For square blocks, in one example, the 65 angular prediction directions are defined in a clockwise direction from 45° to −135° for a square-shaped coding block. Figure 5 shows a schematic diagram (510) illustrating 65 intra-prediction directions according to one aspect of the present disclosure (e.g., in JEM). In addition to these 65 intra-prediction directions (corresponding to the 65 intra-prediction modes 2-66), the intra-prediction modes can include planar mode (“0”) and DC mode (“1”).
[0076] In some cases, such as in VVC, wide-angle intra prediction (WAIP) may be used, in which, for non-square blocks, the 14 angles using prediction from the short side of the block may be replaced by more extreme angles using prediction from the long side, bringing the total number of angles supported in WAIP to 93, and leaving the number of angle modes that can be signaled for a particular block size at 65.
[0077] In an example such as VP9, eight directional modes are supported, corresponding to angles from 45° to 207°. An example of directional intra prediction as used in Alliance for Open Media (AOMedia) Video 1 (AV1) is shown in FIG. 6. To take advantage of more diverse spatial redundancies in directional textures, for example, in AV1, the directional intra modes can be extended to a finer-grained set of angles. The original eight angles can be slightly modified to become nominal angles. FIG. 6 shows eight nominal angles: V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED. Each nominal angle can have seven finer angles, so for example, AV1 uses 56 directional angles. The prediction angle may be represented by a nominal intra angle (e.g., V_PRED) plus an angle delta, which is a 3° step size multiplied by −3 to 3. In one aspect, to implement directional prediction modes in a general manner in AV1, the directional intra prediction modes in AV1 (e.g., all 56 directional intra prediction modes) are implemented using a unified directional predictor, which projects each pixel to a reference sub-pixel position and interpolates the reference pixel with an intra interpolation filter (e.g., a 2-tap bilinear filter), as shown in FIG.
[0078] Figure 7 shows an example of a non-directional intra predictor as used in AV1 according to one aspect of the present disclosure. Figure 7 shows the locations of the above, left, and above-left samples relative to one pixel (701) in the current block (710) to be predicted. In this example, five non-directional smooth intra prediction modes are used, including DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. In DC prediction, an average of neighboring samples to the left and above may be used as a predictor for the pixel (701) in the current block (710). In the PAETH predictor, reference samples above, left, and above-left may first be fetched, and then the value closest to (above + left - above-left) may be set as the predictor for the pixel (701). The SMOOTH, SMOOTH_V, and SMOOTH_H modes can predict the block (710) using quadratic interpolation in the vertical or horizontal direction, or an average in both directions.
[0079] An intra predictor based on recursive filtering is described using Figure 8. Figure 8 shows an example of a recursive intra filtering mode according to an aspect of the present disclosure. A FILTER INTRA mode may be used for a block (e.g., a luma block) to capture decaying spatial correlation with a reference on an edge. In AV1 and other similar applications, five filter intra modes may be used. Each may be represented by a set of eight 7-tap filters that reflect the correlation between a pixel within a 4x2 patch and its seven neighbors. The weighting coefficients of the 7-tap filters may be position-dependent. Figure 8 shows an example of an 8x8 block (810). The block (810) may be divided into eight 4x2 patches: B0, B1, B2, B3, B4, B5, B6, and B7. Each patch's seven neighbors (e.g., indicated by R0-R7) may be used to predict pixels within the current patch. For patch B0, all neighbors (e.g., indicated by R0-R7 of patch B0) have already been reconstructed. For other patches, not all neighbors have been reconstructed, in which case the predicted values of the immediate neighboring patches can be used as a reference. For example, the neighbors (e.g., all neighbors) of patch B7 have not been reconstructed, and the predicted samples of the neighboring patches (e.g., patches B5 and B6) are used.
[0080] MRL intra prediction can use multiple reference lines for intra prediction. Figure 9 shows an example of MRL intra prediction. Four reference lines 0-3 of a current block (901) are shown in Figure 9. Where i is 0, 1, 2, or 3, reference line i can include reference samples that are i lines away from the current block (901), for example, i lines away from the boundary of the current block (901) (e.g., i rows away from the top boundary and / or i columns away from the left boundary). For example, reference line i can include reference samples on the i-th row of the top boundary of the current block (901) and / or the i-th column to the left of the left boundary of the current block (901). In one example, reference line 0 includes reference samples that are adjacent to the current block (901), such as reconstructed neighboring samples that include an upper neighboring sample above the current block (901) and a left neighboring sample to the left of the block (901). In one example, reference line 0 can include the reconstructed neighboring sample on the upper left.
[0081] Reference lines 0-3 may include multiple segments, such as segments A-F. In one example, samples of segments A and F are not fetched from reconstructed neighboring samples. Samples of segments A and F may be padded (or filled) with the nearest samples from segments B and E, respectively.
[0082] In an example such as in HEVC, the closest reference line (i.e., reference line 0) is used in intra prediction (or intra-picture prediction). In one example, a reference line other than the closest reference line (e.g., reference line 1) is used in intra prediction. In one example, multiple reference lines may be used in MRL intra prediction.
[0083] MRL intra prediction can be extended to include more reference lines for intra prediction, for example, in Enhanced Compression Model 5 (ECM5).
[0084] In angular intra prediction, a current sample in a current block may be predicted using, for example, a reference sample (eg, a predicted sample) in a reference line or an interpolated reference sample.
[0085] FIG. 10 shows an example of angular intra prediction using reference line 0 adjacent to block (1000). This description may also be applied to using a reference line (e.g., reference line 1) that is not adjacent to block (1000). A portion of reference line 0 is shown. A sample (1001) in block (1000) may be predicted using one or more samples of reference line 0 in the intra predictor direction (or angular direction) (1022). Sample (1001) may be projected to reference line 0 along the angular direction (1022).
[0086] In the example shown in FIG. 10 , the projected position (1013) of sample (1001) is between two reference samples (1011)-(1012) on reference line 0 and is referred to as a projected fractional position or fractional sample position. Reference samples around the projected position (1013) may be used to predict sample (1001), for example, by using an interpolation filter (e.g., an intra-interpolation filter). In one aspect, an interpolation filter is applied when the projected position (1013) of sample (1001) is a fractional sample position between two adjacent reference samples. The intra-interpolation filter may be any suitable filter, such as a bilinear filter, a cubic filter, a spline interpolation filter, a DCT-based interpolation filter, a DST-based interpolation filter, or the like. In one aspect, a sample (e.g., sample (1001)) that is intra-predicted using an angular intra-prediction mode is projected to a fractional position (e.g., projection position (1013)) that is between two adjacent reference samples (e.g., (1011)-(1012)), and a filter used to generate a predicted sample value for the sample is referred to as an intra-interpolation filter. In one example, the predicted sample value is based on interpolation of the sample values of the two adjacent reference samples (e.g., (1011)-(1012)), e.g., the predicted sample value is a weighted average of the sample values of the two adjacent reference samples using weights that correspond to the filter coefficients of the intra-prediction filter. In one example, these weights are the respective filter coefficients of the intra-prediction filter. In one example, the predicted sample value is based on interpolation of the sample values of the two adjacent reference samples (e.g., (1011)-(1012)) and an additional reference sample.
[0087] The two reference samples (1011)-(1012) of reference line 0 can be used to predict sample (1001) using a two-tap intra-interpolation filter (e.g., a bilinear filter). In one example, the filter coefficients in the two-tap intra-interpolation filter can be based on two distances between the projected position (1013) and two adjacent integer positions (indicated by black dots) of the two reference samples (1011)-(1012), respectively. For example, if the distance between the projected fractional position (1013) and the reference sample (1011) is smaller than the distance between the projected fractional position (1013) and the reference sample (1012), the filter coefficient for the reference sample (1011) is larger than the filter coefficient for the reference sample (1012).
[0088] In one example, four reference samples (1011), (1012), (1014), and (1015) of reference line 0 can be used to predict sample (1001) using a 4-tap intra-interpolation filter (e.g., a 4-tap linear interpolation filter), such as a DCT-based interpolation filter (DCTIF).
[0089] Intra-interpolation filters may differ by filter type and / or the number of filter coefficients (also referred to as filter taps or taps). Filter types may include different types of intra-interpolation filters, such as bilinear filters, cubic filters, spline filters, DCT-based interpolation filters, DST-based interpolation filters, and the like. In one example, filter types include intra-interpolation filters associated with different directions, such as horizontal filters (e.g., applied to samples in rows, as shown in FIG. 10) and vertical filters (e.g., applied to samples in columns). Referring to FIG. 10, the two-tap intra-interpolation filter applied to two reference samples (1011)-(1012) may be a horizontal filter.
[0090] Intra-interpolation filters can differ by the number of filter taps, for example, different intra-interpolation filters can include intra-interpolation filters with different taps, such as 2-tap, 4-tap, 6-tap, and 8-tap filters.
[0091] Intra-interpolation filters may differ in filter type and filter taps.
[0092] An intra-interpolation filter known in the related art is a fixed interpolation filter for a block (e.g., each block). In one example, an intra-interpolation filter applied to a block is considered fixed when the filter type and number of taps of the intra-interpolation filter are fixed. When the filter type and number of taps of the intra-interpolation filter are fixed, the filter coefficients of the intra-interpolation filter may vary based on the samples to be predicted within the block and the respective intra-prediction modes used to predict those samples.
[0093] The optimal intra-interpolation filter may depend on the statistics of the samples within an image or video, e.g., within a picture or frame, and therefore using a fixed interpolation filter (e.g., a fixed intra-interpolation filter) may be suboptimal.
[0094] A current block in a current picture may be coded using intra prediction using a directional intra prediction mode. The directional intra prediction mode may also be referred to as an angular intra prediction mode, an angular mode, a directional mode, or a directional intra mode. The directional intra prediction mode or angular intra prediction mode is an intra prediction mode that can predict samples in the current block along a prediction direction associated with the angular intra prediction mode, as described in Figures 4 and 10.
[0095] According to one aspect of the present disclosure, an intra-interpolation filter (e.g., as described in FIG. 10 ) used in an angular intra-prediction mode may be determined (e.g., derived or selected) based on neighboring reconstructed samples (also referred to as neighboring reconstructed samples) of a current block. Determining the intra-interpolation filter may include determining one or more characteristics of the intra-interpolation filter. For example, determining the intra-interpolation filter may include determining the type of intra-interpolation filter and / or determining the number of taps used in the intra-interpolation filter. The determined intra-interpolation filter may depend on the neighboring reconstructed samples of the current block and thus may be considered self-guided, e.g., the determination of the intra-interpolation filter is guided by the neighboring reconstructed samples of the current block. The determined intra-interpolation filter may be referred to as a self-guided intra-interpolation filter and may be applied to predict samples within the current block.
[0096] In one aspect, an intra-interpolation filter is determined for a current block based on neighboring reconstructed samples of the current block and the angular intra-prediction mode used to predict the current block. In one example, a first portion of the neighboring reconstructed samples (e.g., reference line 0 in FIG. 11 ) is predicted using a predetermined set of intra-interpolation filters from a second portion of the neighboring reconstructed samples (e.g., reference line 1 in FIG. 11 ) using an angular intra-prediction mode, thereby generating a predicted first portion associated with the predetermined set of intra-interpolation filters. One of the predetermined set of intra-interpolation filters may be selected based on a comparison between the predicted first portion and the first portion. For example, a prediction error may be generated based on the predicted first portion and the first portion, and one of the predetermined set of intra-interpolation filters may be selected as the intra-interpolation filter associated with the smallest prediction error. In one example, the predetermined set of intra-interpolation filters includes a first filter and a second filter. The predicted first portion includes a predicted first portion I predicted using the first filter and a predicted first portion II predicted using the second filter. The prediction errors include a first prediction error between the prediction first portion I and the first portion, and a second prediction error between the prediction first portion II and the second portion. If the first prediction error is the smallest prediction error (e.g., the first prediction error is smaller than the second prediction error), the first filter may be selected as the intra-prediction filter for the current block. The above description may also be suitably applied when multiple portions (e.g., reference lines 0-3) are used to determine the intra-prediction filter (instead of only the first portion described above).
[0097] In one aspect, the intra-interpolation filter and angular intra-prediction mode used to predict a current block are determined for the current block based on neighboring reconstructed samples of the current block.
[0098] Using a self-guided intra-interpolation filter can improve coding efficiency and therefore can be more efficient than using a fixed interpolation filter in related art. In one example, the fixed interpolation filter is predetermined and does not depend on neighboring reconstructed samples of the current block. In one example, the current block is a luma block. In one example, the current block is a chroma block. The current block can be predicted using MRL intra-prediction, which predicts samples within the current block using a reference line other than reference line 0 (e.g., reference line 1 in FIG. 9).
[0099] In this disclosure, the term "block" may be interpreted as a predictive block, a coding block, a coding unit (CU), etc. The term "current block" may be interpreted as a predictive block being coded (e.g., being reconstructed), a coding block being coded (e.g., being reconstructed), or a coding unit (CU) being coded (e.g., being reconstructed), etc. The methods described in this disclosure may be applicable to multiple different video coding standards, including, but not limited to, AV1, AOMedia Video Model (AVM), AOMedia Video 2 (AV2), Versatile Video Coding (VVC), Enhanced Compression Model (ECM), and / or H.267, etc.
[0100] An intra-interpolation filter used in a directional intra-prediction mode (or angular intra-prediction mode) may be derived or selected by neighboring reconstructed samples of a current block. For example, a current block in a picture is coded in an angular intra-prediction mode using an intra-interpolation filter. The intra-interpolation filter used in the angular intra-prediction mode may be determined (e.g., derived or selected) from a predetermined set of intra-interpolation filters based on neighboring samples of the current block. The neighboring samples may include, for example, neighboring reconstructed samples within N lines (e.g., N neighboring lines, N>1) from one or more boundaries of the current block. The determined intra-interpolation filter (e.g., selected or derived intra-interpolation filter) may be applied to reference samples in reference lines in the picture to predict samples in the current block using the angular intra-prediction mode.
[0101] In one aspect, the value of N depends on the block size. N indicates the number of lines and / or columns in the plurality of reference lines. In one example, how many lines and / or columns are used depends on the block size. Whether neighboring reconstructed samples are used to derive an interpolation filter (e.g., an intra-interpolation filter) may depend on the block area. For example, when the block area is smaller than a threshold, neighboring pixels (e.g., neighboring reconstructed samples) are used to derive the cost (e.g., prediction error) of each filter (e.g., each intra-interpolation filter of a given set of intra-interpolation filters) to save (e.g., reduce) signaling overhead.
[0102] Figure 11 illustrates an example of a current block (1101) coded using an angular intra prediction mode and neighboring reconstructed samples (1110) of the current block (1101) according to one aspect of the present disclosure. In one example, the current block (1101) is being reconstructed. The prediction direction of the angular intra prediction mode may be referred to as the intra prediction direction. In the example illustrated in Figure 11, arrows (1121)-(1123) indicate the same direction, e.g., the prediction direction of the angular intra prediction mode. Arrows (1124)-(1126) may indicate the same direction, e.g., the opposite direction to the prediction direction of the angular intra prediction mode. The opposite direction (e.g., downward, as indicated by arrows (1124)-(1126)) may be directly opposite to the prediction direction of the angular intra prediction mode (e.g., upward, as indicated by arrows (1121)-(1123)).
[0103] Neighboring reconstructed samples (1110) of the current block (1101) may be referred to as neighboring reconstructed samples of the current block (1101). The neighboring reconstructed samples (1110) of the current block (1101) may include samples in multiple reference lines (e.g., two or more reference lines). In the example shown in FIG. 11, the neighboring reconstructed samples (1110) of the current block (1101) include neighboring reconstructed samples in reference lines 0-(N-1) (N is 4). The method described with reference to FIG. 11 may be applied to any suitable N greater than 1. As shown in FIG. 11, where i is 0, 1, 2, or 3, reference line i may include reference samples that are i lines away from the current block (1101), for example, i lines away from the boundary of the current block (1101) (e.g., i rows away from the top boundary and / or i columns away from the left boundary). For example, reference line i includes reference samples on row i of the top boundary of the current block (1101) and / or column i to the left of the left boundary of the current block (1101). In one example, reference line 0 includes reference samples adjacent to the current block (1101), such as reconstructed neighboring samples including top neighboring samples above the current block (1101) and left neighboring samples to the left of the block (1101).
[0104] In one example, the neighboring reconstructed samples (1110) of the current block (1101) include a top reference line (1111) that is a reference line above the top boundary of the current block (1101). In one example, the neighboring reconstructed samples (1110) of the current block (1101) include a left reference line (1112) that is a reference line to the left of the left boundary of the current block (1101). In one example, the neighboring reconstructed samples (1110) of the current block (1101) include the top reference line (1111) and the left reference line (1112).
[0105] In one aspect, N is greater than 1 and the angular intra-prediction mode has already been determined (e.g., associated with the prediction direction indicated by arrows (1121)-(1123)). The neighboring reconstructed samples may include multiple lines of neighboring reconstructed samples (e.g., reference lines 0-3 in FIG. 11).
[0106] Each intra-interpolation filter of the predetermined set of intra-interpolation filters may be applied to the neighboring reconstructed samples of each line of the multiple lines of neighboring reconstructed samples to predict the neighboring reconstructed samples of the respective line using an angular intra-prediction mode, and a respective prediction error may be obtained. In one example, for each combination of an intra-interpolation filter of the predetermined set of intra-interpolation filters and a neighboring reconstructed sample of a line of the multiple lines of neighboring reconstructed samples, the neighboring reconstructed sample of the respective line may be predicted based on one or more remaining lines of the multiple lines of neighboring reconstructed samples using the respective intra-interpolation filter and the angular intra-prediction mode, and at least one prediction error may be obtained.
[0107] In one example, referring to FIG. 11 , N is 4, the neighboring reconstructed samples of the multiple lines include reference lines 0-3, one of the predetermined set of intra-interpolation filters is a bilinear filter, and the neighboring reconstructed sample of the certain line of the neighboring reconstructed samples of the multiple lines is reference line 0. The one or more remaining lines of the neighboring reconstructed samples of the multiple lines include reference lines 1-3. The neighboring reconstructed sample of the certain line (e.g., reference line 0) can be predicted based on reference lines 1-3, and at least one prediction error includes three prediction errors respectively associated with reference lines 1-3. For example, a first prediction error is based on a reconstructed sample in reference line 0 and a predicted sample of reference line 0 based on reference line 1 and an angular intra-prediction mode associated with the prediction direction indicated by arrow (1121). Similarly, reference lines 1-3 can be predicted, and corresponding prediction errors are obtained for the bilinear filter. For example, reference line 1 is predicted based on reference lines 0, 2, and 3, respectively, using bilinear filters and angular intra prediction modes associated with the prediction directions indicated by arrows (1122) and (1124), and three prediction errors are obtained.
[0108] The above process for one of the predetermined set of intra-interpolation filters (e.g., a bilinear filter) may be performed or repeated for other intra-interpolation filters (e.g., spline filters and / or cubic filters, etc.) of the predetermined set of intra-interpolation filters. The intra-interpolation filter corresponding to the smallest prediction error among the obtained prediction errors may be selected. For example, if the smallest prediction error is the prediction error obtained using a particular intra-interpolation filter (e.g., a bilinear filter), the particular intra-interpolation filter is selected as the intra-interpolation filter to be applied to the current block (1101) for that angular intra-prediction mode (e.g., associated with the prediction direction indicated by arrows (1121)-(1123)).
[0109] In one aspect, given an angular intra-prediction mode (also referred to as a directional intra-prediction mode) and neighboring reconstructed samples of multiple lines (e.g., neighboring reconstructed samples in reference lines 0-3 in FIG. 11), a predetermined set of intra-interpolation filters are applied to the reference lines (e.g., reference lines 0-3 in FIG. 11) to predict each other along the intra-prediction direction (e.g., upward, the prediction direction indicated by arrows (1121)-(1123)) and / or the opposite direction (e.g., downward, the opposite direction indicated by arrows (1124)-(1126)), and the candidate intra-interpolation filter (also referred to as a candidate interpolation filter) having the smallest prediction error is selected as the intra-interpolation filter for the current block (1101). Referring to FIG. 11, reference line 0 is predicted by other reference lines, such as reference lines 1, 2, 3, ..., and (N-1) (indicated by arrow (1121)). Reference line 1 can be predicted by other reference lines such as reference lines 0, 2, 3, ..., and (N-1) (indicated by arrows 1122 and 1124). Reference line 2 can be predicted by other reference lines such as reference lines 0, 1, 3, ..., and (N-1) (indicated by arrows 1123 and 1125). Reference line 3 can be predicted by other reference lines such as reference lines 0, 1, 2, ..., and (N-1) (indicated by arrow 1126). Each of the candidate interpolation filters is applied to predict the reference lines between each other as described above, and the one with the smallest prediction error is used as the prediction filter (e.g., the determined intra-interpolation filter).
[0110] In one aspect, one of (i) the type of intra-interpolation filter or (ii) the number of taps in the intra-interpolation filter is selected from a predetermined set of intra-interpolation filters based on neighboring reconstructed samples of the current block.
[0111] In one example, the predetermined set of intra-interpolation filters includes different types of intra-interpolation filters, such as one of a bilinear interpolation filter, a cubic interpolation filter, a spline interpolation filter, a DCT-based interpolation filter, or a DST-based interpolation filter. The type of the intra-interpolation filter can be selected from the different types of intra-interpolation filters based on neighboring reconstructed samples of the current block. The selected intra-interpolation filter can be one of a bilinear interpolation filter, a cubic interpolation filter, a spline interpolation filter, a DCT-based interpolation filter, or a DST-based interpolation filter.
[0112] For example, the predetermined set of intra-interpolation filters (e.g., candidate intra-interpolation filters) may include, but are not limited to, bilinear filters, cubic filters, spline interpolation filters, and DCT / DST-based interpolation filters (e.g., DCT or DST-based interpolation filters).
[0113] In one example, the predetermined set of intra-interpolation filters includes different tap numbers, such as one or more of a 2-tap interpolation filter, a 4-tap interpolation filter, a 6-tap interpolation filter, or an 8-tap interpolation filter. The tap number of the intra-interpolation filter can be selected based on neighboring reconstructed samples of the current block. The tap number of the intra-interpolation filter can be one of 2-tap, 4-tap, 6-tap, or 8-tap.
[0114] In one example, a given set of intra-interpolation filters (e.g., candidate intra-interpolation filters) may have different numbers of filter taps. For example, the candidate intra-interpolation filters may be a group of 2-tap, 4-tap, 6-tap, and 8-tap filters.
[0115] In one example, whether the multiple lines of neighboring reconstructed samples (e.g., reference lines 0-(N-1)) include (i) the left reference line (1112) (e.g., neighboring reconstructed samples on the left line to the left of the current block (1101)), (ii) the above reference line (1111) (e.g., neighboring reconstructed samples on the line above the current block (1101)), or (iii) the left reference line (1112) (e.g., neighboring reconstructed samples on the line above the current block (1101)) and the above reference line (1111) (e.g., neighboring reconstructed samples on the line above the current block (1101)) depends on the intra prediction direction of the angular intra prediction mode. In the example shown in Figure 11, the intra prediction direction of the angular intra prediction mode is upward, as indicated by arrows (1121)-(1123). For example, whether a left reference line (e.g., left reference line (1112)), an upper reference line (also referred to as an upper reference line, such as upper reference line (1111)), or both left and upper reference lines are used to derive an intra-interpolation filter may depend on a given intra-prediction direction, for example, as indicated by arrows (1121)-(1123) in Figure 11. In one example, if the prediction direction is horizontal, the left reference line is used. In one example, if the prediction direction is vertical, the upper reference line is used.
[0116] In one aspect, only the intra-prediction direction is used to derive the selectable filter. For example, a neighboring reconstructed sample of each line of a plurality of lines (e.g., reference line 0) is predicted based on the intra-prediction direction of the angular intra-prediction mode, such as the prediction directions indicated by arrows (1121)-(1123). Referring to Figure 11, in this case, reference line 0 can be predicted based on the prediction directions indicated by arrows (1121) from reference lines 1-3, respectively, reference line 1 can be predicted based on the prediction directions indicated by arrows (1122) from reference lines 2-3, respectively, reference line 2 can be predicted based on the prediction direction indicated by arrow (1123) from reference line 3, and no prediction is performed on reference line 3.
[0117] In one aspect, only the opposite direction of the intra-prediction direction is used to derive the selectable filter. For example, a neighboring reconstructed sample of each line of a plurality of lines (e.g., reference line 1) is predicted based on the opposite direction of the intra-prediction direction of the angular intra-prediction mode, such as the opposite direction indicated by arrows (1124)-(1126). Referring to Figure 11, in this case, reference line 1 can be predicted based on the opposite direction indicated by arrow (1124) from reference line 0, and reference line 2 can be predicted based on the opposite direction indicated by arrow (1125) from reference lines 0-1. Reference line 3 can be predicted based on the opposite direction indicated by arrow (1126) from reference lines 0-2, respectively, and no prediction is performed for reference line 0.
[0118] According to one aspect of the present disclosure, the selection of an intra-interpolation filter from a predetermined set of intra-interpolation filters and the selection of an angular intra-prediction mode from a predetermined set of angular intra-prediction modes (e.g., angular intra-prediction modes) may be performed based on neighboring reconstructed samples of the current block (1101). Both the intra-interpolation filter and the angular intra-prediction mode may be selected based on neighboring reconstructed samples of the current block (1101).
[0119] In one example, the selection of both the directional intra-prediction mode (or angular intra-prediction mode) and the intra-interpolation filter can be evaluated by a cost value calculated as the prediction error of the prediction reference line between each other, and the best L combinations of intra-modes (e.g., one or more intra-modes from a predetermined set of angular intra-prediction modes) and filters (e.g., one or more filters from a predetermined set of intra-prediction filters) are selected as candidates. From the best L combinations of intra-modes and filters, the actual angular intra-prediction mode (e.g., associated with the intra-prediction angle) and interpolation filter (e.g., intra-interpolation filter) used in the intra-prediction process can be obtained.
[0120] In one example, the selection of an intra-interpolation filter and an angular intra-prediction mode may include: deriving N1 angular intra-prediction modes from among a plurality of angular intra-prediction modes (e.g., a predetermined set of angular intra-prediction modes), where the N1 angular intra-prediction modes are associated with the N1 lowest cost values from among the cost values of the respective angular intra-prediction modes; determining each cost value (e.g., the cost value of each angular intra-prediction mode) based on neighboring reconstructed samples of the current block (1101), the respective angular intra-prediction modes, and a default intra-interpolation filter (e.g., a filter from the predetermined set of intra-interpolation filters); selecting M1 intra-interpolation filters from the predetermined set of intra-interpolation filters that are associated with the M1 lowest cost values from among the cost values of the respective intra-interpolation filters; Each cost value (e.g., a cost value for each filter in a predetermined set of intra-interpolation filters) may be determined based on neighboring reconstructed samples of the current block, the respective intra-interpolation filters, and a respective angular intra-prediction mode among the N1 angular intra-prediction modes. The intra-interpolation filters and angular intra-prediction modes may be selected from N1×M1 combinations, where each combination of the N1×M1 combinations may include an intra-interpolation filter among the M1 intra-interpolation filters and an angular intra-prediction mode among the N1 angular intra-prediction modes.
[0121] For example, a default intra-interpolation filter (which may be predetermined) is applied to derive the best N1 intra-prediction modes (e.g., the above-mentioned N1 angular intra-prediction modes) using neighboring reconstructed samples. After the best N1 intra-prediction modes are selected, the best M1 interpolation filters may be further selected, for example, from a predetermined set of intra-interpolation filters. Example values of M1 and N1 may include, but are not limited to, 1, 2, 3, or 4. M1 and N1 are positive integers.
[0122] In one aspect, the intra-prediction mode (e.g., angular intra-prediction mode) and the intra-interpolation filter may be determined iteratively by neighboring reconstructed samples. That is, an intra-prediction filter may first be selected using the above-described method, for example, an intra-prediction filter may be selected as one of the M1 intra-interpolation filters. Then, the best L1 intra-prediction modes are further selected by trying each intra-prediction mode in the neighboring reconstructed samples and selecting the L1 intra-prediction modes associated with the smallest L1 prediction errors for that intra-prediction filter. Thereafter, the intra-interpolation filter may be further refined using a similar approach. The improved intra-interpolation filter may be used to further refine the intra-prediction mode.
[0123] In one example, this iterative search of intra-prediction modes (e.g., angular intra-prediction modes) and intra-interpolation filters may be terminated when the prediction error is no longer reduced, when the prediction error is within a predetermined threshold (e.g., a first predetermined threshold), or when the reduction in prediction error during the search process is within a predetermined threshold (e.g., a second predetermined threshold) (e.g., the difference in error between successive iterations is within a second predetermined threshold).
[0124] An example of the last two iterations (e.g., the first iteration and the second iteration) in the iterative process is described below. In the first iteration, N2 angular intra prediction modes may be selected from a first plurality of modes among the plurality of angular intra prediction modes based on neighboring reconstructed samples of the current block and an intra interpolation filter. As described in FIG. 11, a first updated intra interpolation filter may be selected based on neighboring reconstructed samples of the current block and one of the N2 directional intra prediction modes. In the second iteration, N3 angular intra prediction modes may be selected from a second plurality of modes among the plurality of angular intra prediction modes based on neighboring reconstructed samples of the current block and the first updated intra interpolation filter. A second updated intra interpolation filter may be selected based on neighboring reconstructed samples of the current block and one of the N3 angular intra prediction modes. The iterative process may be terminated, and an intra-interpolation filter and an angular intra-prediction mode may be selected based on one of (i) the first prediction error of the first iteration and the second prediction error of the second iteration, and (ii) the second prediction error of the second iteration. The first prediction error may be associated with a first updated intra-interpolation filter and one of the N2 directional intra-prediction modes. The second prediction error may be associated with a second updated intra-interpolation filter and one of the N3 angular intra-prediction modes. In one example, based on a difference between the first prediction error and the second prediction error, for example, an absolute difference between the first prediction error and the second prediction error being within a second predetermined threshold, the second updated intra-interpolation filter is selected as the selected intra-interpolation filter, and one of the N3 angular intra-prediction modes is selected as the selected angular intra-prediction mode. In one example, based on the second prediction error, for example, the second prediction error being within a first predetermined threshold, the second updated intra-interpolation filter is selected as the selected intra-interpolation filter, and the one of the N3 angular intra-prediction modes is selected as the selected angular intra-prediction mode.
[0125] In one example, the first and second modes are the same as the plurality of angular intra prediction modes. In one example, the first and second modes are a subset of the plurality of angular intra prediction modes. This subset of the plurality of angular intra prediction modes may be obtained from a previous iteration. For example, the second modes are the N2 angular intra prediction modes.
[0126] In one aspect, the angular intra-prediction mode and intra-interpolation filter may be determined based on neighboring reconstructed samples in another iterative manner, i.e., an angular intra-prediction mode is first selected, then a best intra-prediction filter is obtained, followed by refining the angular intra-prediction mode using a similar approach, e.g., based on one of the best intra-prediction filters. The refined angular intra-prediction mode may be used to further refine the intra-prediction filter.
[0127] According to one aspect of the present disclosure, for each combination of an angular intra prediction mode among a plurality of angular intra prediction modes and an intra interpolation filter from a predetermined set of intra interpolation filters, a cost value for each combination is determined based on neighboring reconstructed samples of the current block, for example, using the method described in Figure 11. In one example, the cost value is determined based on a reconstructed sample in a reference line (e.g., reference line 0) and a predicted sample of the reference line (e.g., predicted by reference line 1 using the angular intra prediction mode and intra interpolation filter in the combination), for example, as shown in Figure 11. In one example, the cost value is determined based on a reconstructed sample in the reference line (e.g., reference line 0) and a predicted sample of the reference line using each reference line (e.g., predicted by reference lines 1 to 3 using the angular intra prediction mode and intra interpolation filter in the combination), and the cost value is one of the cost values (e.g., the smallest cost value) as shown in Figure 11. K combinations may be selected based on the determined cost values. Based on index information signaled in the bitstream, an intra-interpolation filter and an angular intra-prediction mode can be selected from K combinations. In one example, the combinations of angular intra-prediction modes and intra-interpolation filters are ranked, for example, in ascending order of cost values, and the K combinations are associated with the K smallest cost values.
[0128] In one example, the selection of a directional intra-prediction mode and an intra-interpolation filter by using neighboring reconstructed samples can be performed at both the encoder side and the decoder side. A combination of an intra-prediction mode (e.g., an angular intra-prediction mode) and an intra-interpolation filter can be selected. In one example, the order of the combinations of the directional intra-prediction mode and the intra-interpolation filter is determined based on the calculated cost value of each combination. Figure 11 shows an example of calculating the cost value of each combination. For example, a syntax element can be signaled in the bitstream to indicate which combination should be used for the current block. In one example, the combinations of the intra-prediction mode (e.g., an angular intra-prediction mode) and the intra-interpolation filter can be ranked based on the calculated cost value of the combination.
[0129] According to one aspect of the present disclosure, a feature value may be calculated based on neighboring reconstructed samples of a current block, and an intra-interpolation filter may be selected from a predetermined set of intra-interpolation filters based on the calculated feature value. The feature value may be calculated as an absolute gradient value associated with the neighboring reconstructed samples of the current block, or a difference between the minimum value among the neighboring reconstructed samples of the current block and the maximum value among the neighboring reconstructed samples of the current block. For example, given the neighboring reconstructed samples, feature values such as absolute gradient values and a difference between the minimum and maximum values (e.g., the difference between the minimum and maximum values) are calculated. Then, based on the feature value, an intra-interpolation filter (e.g., an intra-interpolation filter to be used in an angular intra-prediction mode to predict the current block) may be selected from a predetermined group of intra-interpolation filters (e.g., a predetermined set of intra-interpolation filters). In one example, the intra-interpolation filter with the largest or smallest gradient value is selected. In one example, the absolute gradient value(s) are calculated based on one or more predetermined directions.
[0130] According to one aspect of the present disclosure, an output may be obtained from a neural network (e.g., a convolutional neural network (CNN)). The input of the neural network may include neighboring reconstructed samples of a current block, and the output (e.g., an index) indicates a selected intra-interpolation filter, which is one of a predetermined set of intra-interpolation filters. In one example, the input of the neural network further includes an intra-prediction mode index indicating an angular intra-prediction mode.
[0131] In one example, neighboring reconstructed samples are provided to a convolutional neural network (CNN), and the output includes an index of an intra-interpolation filter selected from a predetermined group of intra-interpolation filters (e.g., a predetermined set of intra-interpolation filters). The network parameters (e.g., parameters of the CNN) can be pre-trained and evaluated at both the encoder and decoder sides, or can be trained online using reconstructed blocks. In one example, the intra-prediction mode index, along with neighboring reconstructed samples, is provided to the CNN for deriving the intra-interpolation filter.
[0132] The explanation for determining (e.g., selecting) (i) an intra-interpolation filter, or (ii) determining (e.g., selecting) an intra-interpolation filter and an angular intra-prediction mode for the current block (1101), may be based on any suitable neighboring reconstructed samples of the current block (1101). The size of the neighboring reconstructed samples of the current block (1101) may depend on (e.g., increase with) the block size of the current block (1101). Referring to FIG. 11 , the top reference line (1111) and the left reference line (1112) may include a region (1113) that is an upper-left neighbor of the current block (1101). In some examples, the top reference line (1111) and / or the left reference line (1112) may include a portion of the region (1113). In some examples, the top reference line (1111) and / or the left reference line (1112) may not include the region (1113).
[0133] 12 shows a flowchart outlining a process (1200) according to one aspect of the present disclosure. The process (1200) can be used in a video decoder. In various aspects, the process (1200) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), and the like. In some aspects, the process (1200) is implemented in software instructions, and thus the processing circuit performs the process (1200) when the processing circuit executes the software instructions. The process (1200) begins at (S1201) and proceeds to (S1210).
[0134] At (S1210), a bitstream having a current block in a picture is received. Coding information of the bitstream may indicate that the current block is coded in an angular intra prediction mode using an intra interpolation filter.
[0135] At (S1220), an intra-interpolation filter may be determined (e.g., selected) from a predetermined set of intra-interpolation filters based on neighboring reconstructed samples of the current block. The neighboring reconstructed samples may include reconstructed samples within N lines from one or more boundaries of the current block, for example, as shown in FIG. 11 .
[0136] In one aspect, one of (i) the type of intra-interpolation filter or (ii) the number of taps in the intra-interpolation filter is selected from a predetermined set of intra-interpolation filters based on neighboring reconstructed samples of the current block.
[0137] In one example, the predetermined set of intra-interpolation filters includes different types of intra-interpolation filters. The type of the intra-interpolation filter can be selected from the different types of intra-interpolation filters based on neighboring reconstructed samples of the current block. The selected intra-interpolation filter can be one of a bilinear interpolation filter, a cubic interpolation filter, a spline interpolation filter, a DCT-based interpolation filter, or a DST-based interpolation filter.
[0138] In one example, the predetermined set of intra-interpolation filters includes different tap numbers. The tap number of the intra-interpolation filter can be selected based on neighboring reconstructed samples of the current block. The tap number of the intra-interpolation filter can be one of 2 taps, 4 taps, 6 taps, or 8 taps.
[0139] In one example, whether the neighboring reconstructed samples include (i) neighboring reconstructed samples of the left line to the left of the current block, (ii) neighboring reconstructed samples of the line above the current block, or (iii) neighboring reconstructed samples of the left line to the left of the current block and neighboring reconstructed samples of the line above the current block depends on the intra prediction direction of the angular intra prediction mode.
[0140] In one example, the value of N depends on the block size.
[0141] In one aspect, N is greater than 1, and the neighboring reconstructed samples include neighboring reconstructed samples of multiple lines. For each combination of an intra-interpolation filter from the predetermined set of intra-interpolation filters and neighboring reconstructed samples of a line from the multiple lines of neighboring reconstructed samples, the respective intra-interpolation filter can be used to predict the neighboring reconstructed samples of the respective line based on one or more remaining lines from the multiple lines of neighboring reconstructed samples, and at least one prediction error can be obtained. An intra-interpolation filter corresponding to the smallest prediction error among the obtained prediction errors can be selected.
[0142] In one example, neighboring reconstructed samples of each line are predicted based on the intra-prediction direction of the angular intra-prediction mode. In one example, neighboring reconstructed samples of each line are predicted based on the direction opposite to the intra-prediction direction of the angular intra-prediction mode.
[0143] In one aspect, an intra-interpolation filter is selected from a predetermined set of intra-interpolation filters and an angular intra-prediction mode is selected from a plurality of angular intra-prediction modes based on neighboring reconstructed samples of the current block.
[0144] In one example, N1 angular intra prediction modes from among the plurality of angular intra prediction modes associated with the N1 lowest cost values from among the cost values of the respective angular intra prediction modes are derived. Each cost value may be determined based on neighboring reconstructed samples of the current block, the respective angular intra prediction modes, and a default intra interpolation filter from a predetermined set of intra interpolation filters. For each of the N1 angular intra prediction modes, M1 intra interpolation filters associated with the M1 lowest cost values from among the cost values of the predetermined set of intra interpolation filters may be selected from the predetermined set of intra interpolation filters. Each cost value may be determined based on neighboring reconstructed samples of the current block, the respective intra interpolation filters, and the respective angular intra prediction mode from among the N1 angular intra prediction modes. One intra interpolation filter and angular intra prediction mode may be selected from N1×M1 combinations. Each combination of the N1×M1 combinations may include an intra interpolation filter from among the M1 intra interpolation filters and an angular intra prediction mode from among the N1 angular intra prediction modes.
[0145] In one example, the selection of angular intra prediction modes and intra interpolation filters is performed iteratively in an iterative process, which includes selecting N2 angular intra prediction modes from a first plurality of modes of the plurality of angular intra prediction modes based on neighboring reconstructed samples of the current block and the intra interpolation filters, selecting a first updated intra interpolation filter based on neighboring reconstructed samples of the current block and one of the N2 directional intra prediction modes, selecting N3 angular intra prediction modes from a second plurality of modes of the plurality of angular intra prediction modes based on neighboring reconstructed samples of the current block and the first updated intra interpolation filter, and selecting N4 angular intra prediction modes from a second plurality of modes of the plurality of angular intra prediction modes based on neighboring reconstructed samples of the current block and the first updated intra interpolation filter. The method may include selecting a second updated intra-interpolation filter based on one of the N2 directional intra-prediction modes, and selecting the second updated intra-interpolation filter as the selected intra-interpolation filter and one of the N3 angular intra-prediction modes based on (i) a difference between a first prediction error associated with the first updated intra-interpolation filter and one of the N2 directional intra-prediction modes and a second prediction error associated with the second updated intra-interpolation filter and one of the N3 angular intra-prediction modes, or (ii) the second prediction error. In one example, the first plurality of modes and the second plurality of modes are the same as the plurality of angular intra-prediction modes. In one example, the first plurality of modes and the second plurality of modes are a subset of the plurality of angular intra-prediction modes. The subset of angular intra-prediction modes may be obtained from a previous iteration. For example, the second plurality of modes are the N2 angular intra-prediction modes.
[0146] In one example, for each combination of an angular intra-prediction mode from among a plurality of angular intra-prediction modes and an intra-interpolation filter from a predetermined set of intra-interpolation filters, a cost value for each combination is determined based on neighboring reconstructed samples of the current block. K combinations may be selected based on the determined cost values. The intra-interpolation filter and the angular intra-prediction mode may be selected from the K combinations based on index information signaled in the bitstream.
[0147] In one aspect, feature values are calculated based on neighboring reconstructed samples of the current block, and one intra-interpolation filter can be selected based on the calculated feature values.
[0148] In one aspect, a neural network (e.g., a CNN) may be used to determine an intra-interpolation filter for a current block. The input of the neural network includes neighboring reconstructed samples of the current block, and the output is obtained from the neural network based on the input. The output may indicate the selected intra-interpolation filter. In one example, the input of the neural network further includes an intra-prediction mode index indicating an angular intra-prediction mode.
[0149] At (S1230), the selected intra interpolation filter can be applied to reference samples in a reference line in the picture to predict samples in the current block using an angular intra prediction mode, for example as shown in FIG.
[0150] At (S1240), the current block can be reconstructed based on the predicted samples.
[0151] Then, the process proceeds to (S1299) and ends.
[0152] Process 1200 may be adapted as desired. Step(s) of process 1200 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of execution may be used. Figure 14 shows variations on process 1200.
[0153] 13 shows a flowchart outlining a process (1300) according to one aspect of the present disclosure. The process (1300) can be used in a video encoder. In various aspects, the process (1300) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), and the like. In some aspects, the process (1300) is implemented in software instructions, and thus the processing circuit performs the process (1300) when the processing circuit executes the software instructions. The process begins at (S1301) and proceeds to (S1310).
[0154] At (S1310), an intra-interpolation filter may be selected from a predetermined set of intra-interpolation filters based on neighboring reconstructed samples of the current block. The neighboring reconstructed samples may include reconstructed samples within N lines from one or more boundaries of the current block. The current block is coded in an angular intra-prediction mode using the selected intra-interpolation filter.
[0155] At (S1320), the selected intra-interpolation filter can be applied to reference samples in a reference line in the picture to predict samples in the current block using an angular intra-prediction mode.
[0156] At (S1330), the current block is coded based on the predicted samples.
[0157] Then, the process proceeds to (S1399) and ends.
[0158] Process 1300 may be adapted as desired. Step(s) of process 1300 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.
[0159] 14 shows a flowchart outlining a process (1400) according to one aspect of the present disclosure. The process (1400) can be used in a video decoder. In various aspects, the process (1400) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), and the like. In some aspects, the process (1400) is implemented in software instructions, and thus the processing circuit performs the process (1400) when the processing circuit executes the software instructions. The process (1400) begins at (S1401) and proceeds to (S1410).
[0160] At (S1410), coding information of the bitstream indicating that the current block is coded in an angular intra prediction mode using an intra interpolation filter is received from a bitstream having a current block in a picture.
[0161] At (S1420), each of the predetermined set of intra-interpolation filters is applied to adjacent reconstructed samples in N adjacent lines from the boundary of the current block, where N is a positive integer, for example as described in FIG.
[0162] At (S1430), an intra-interpolation filter is selected from a predetermined set of intra-interpolation filters based on a prediction error associated with each of the intra-interpolation filters in the predetermined set, for example as described in FIGS.
[0163] At (S1440), samples in the current block are predicted in angular intra prediction mode using the selected intra interpolation filter, for example, as described with reference to FIG.
[0164] At (S1450), the current block is reconstructed based on the predicted samples.
[0165] Then, the process proceeds to (S1499) and ends.
[0166] Process 1400 may be adapted as desired. Step(s) of process 1400 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.
[0167] The aspects and examples in this disclosure may be used separately or in combination in any order. Also, each of these methods (or aspects, examples), encoders, and decoders may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.
[0168] The techniques described above can be implemented as computer software with computer-readable instructions physically stored on one or more computer-readable media. For example, Figure 15 illustrates a computer system (1500) suitable for implementing certain aspects of the disclosed subject matter.
[0169] Computer software may be coded using any suitable machine code or computer language that can be assembled, compiled, linked, or similarly subjected to mechanisms to produce code having instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), and the like, either directly or via interpretation, microcode execution, and the like.
[0170] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0171] 15 with respect to computer system (1500) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing aspects of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in this exemplary embodiment of computer system (1500).
[0172] The computer system (1500) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, moving a data glove, etc.), audio input (e.g., voice, clapping, etc.), visual input (e.g., gestures, etc.), or olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds, etc.), images (e.g., scanned images, photographic images obtained from a still camera, etc.), or video (e.g., two-dimensional video, three-dimensional video including stereoscopic video, etc.).
[0173] The input human interface devices may include one or more of a keyboard (1501), a mouse (1502), a trackpad (1503), a touchscreen (1510), a data glove (not shown), a joystick (1505), a microphone (1506), a scanner (1507), and a camera (1508) (only one of each is shown).
[0174] The computer system (1500) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1510), data gloves (not shown), or joystick (1505), although some haptic feedback devices may not function as input devices), audio output devices (e.g., speakers (1509), headphones (not shown), etc.), visual output devices (e.g., screens (1510) including CRT screens, LCD screens, plasma screens, and OLED screens (each with or without touchscreen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or four or more dimensional output through means such as stereoscopic output), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), etc.), and printers (not shown).
[0175] The computer system (1500) may also include human-accessible storage devices and their associated media, such as optical media including, for example, a CD / DVD ROM / RW (1520) with a CD / DVD or similar media (1521), a thumb drive (1522), a removable hard drive or solid state drive (1523), legacy magnetic media such as tape and floppy disks (registered trademark, not shown), specialized ROM / ASIC / PLD-based devices (not shown) such as security dongles, and the like.
[0176] Those skilled in the art will also appreciate that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other transitory signals.
[0177] The computer system 1500 may also include interfaces 1554 to one or more communications networks 1555. Networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, latency-tolerant, and the like. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, and the like, TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and vehicular and industrial networks including CANbus. Certain networks typically require an external network interface adapter that attaches to a particular general-purpose data port or peripheral bus 1549 (e.g., a USB port on the computer system 1500), while others are typically integrated into the core of the computer system 1500 by attachment to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, computer system 1500 can communicate with other entities. Such communication may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a particular CANbus device), or two-way, for example, to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks may be used on each network and network interface, as described above.
[0178] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1540) of the computer system (1500).
[0179] The core (1540) may include one or more central processing units (CPUs) (1541), graphics processing units (GPUs) (1542), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (1543), task-specific hardware accelerators (1544), graphics adapters (1550), etc. These devices may be connected via a system bus (1548), along with read-only memory (ROM) (1545), random access memory (1546), and internal mass storage (1547), such as internal non-user-accessible hard drives, SSDs, and the like. In some computer systems, the system bus (1548) may be made accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, and the like. Peripheral devices may be attached either directly to the core's system bus (1548) or via a peripheral bus (1549). In one example, a screen 1510 can be connected to a graphics adapter 1550. Peripheral bus architectures include PCI, USB, and the like.
[0180] The CPU (1541), GPU (1542), FPGA (1543), and accelerator (1544) may execute specific instructions that, in combination, may constitute the aforementioned computer code. The computer code may be stored in ROM (1545) or RAM (1546). Transient data may also be stored in RAM (1546), while permanent data may be stored, for example, in internal mass storage (1547). Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be associated with one or more of the CPU (1541), GPU (1542), mass storage (1547), ROM (1545), RAM (1546), and the like.
[0181] The computer-readable media may have computer code thereon for performing various computer-implemented processes. The media and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0182] By way of example, and not limitation, a computer system having the architecture (1500), and in particular the core (1540), can provide functionality as a result of the execution by one or more processors (including CPUs, GPUs, FPGAs, accelerators, and the like) of software embodied in one or more tangible computer-readable media. Such computer-readable media can be specific storage of the core (1540) that is non-transitory in nature, such as the core's internal mass storage (1547) or ROM (1545), and media associated with user-accessible mass storage as introduced above. Software implementing various aspects of the present disclosure can be stored in such devices and executed by the core (1540). The computer-readable media can include one or more memory devices or chips, depending on specific needs. Software may cause the core (1540) and particularly the processors therein (including CPUs, GPUs, FPGAs, and the like) to perform particular processes or portions of particular processes described herein, including by defining data structures stored in RAM (1546) and modifying such data structures according to processes defined by the software. Additionally, or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1544)) that can operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both, where appropriate. The present disclosure includes any suitable combination of hardware and software.
[0183] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the described elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B and one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of the described elements where applicable, such as when the elements are not mutually exclusive.
[0184] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various equivalent alternatives that fall within the scope of the disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the disclosure and are therefore within its spirit and scope.
Claims
1. 1. A method for video decoding executed by one or more processors, comprising: receiving, from a bitstream having a current block in a picture, coding information of the bitstream indicating that the current block is coded in an angular intra prediction mode using an intra interpolation filter; applying each of a predetermined set of intra-interpolation filters to adjacent reconstructed samples in N adjacent lines from a boundary of the current block, where N is a positive integer; selecting an intra-interpolation filter from the set of intra-interpolation filters based on a prediction error associated with each of the set of intra-interpolation filters; predicting samples in the current block using the selected intra-interpolation filter using the angular intra-prediction mode; reconstructing the current block based on the predicted samples; A method having the following.
2. The step of selecting the one intra-interpolation filter includes: selecting, from the predetermined set of intra-interpolation filters based on the neighboring reconstructed samples of the current block, one of (i) the type of the intra-interpolation filter or (ii) the number of taps in the one intra-interpolation filter; 2. The method of claim 1, comprising:
3. the predetermined set of intra-interpolation filters includes different types of intra-interpolation filters; the selecting the one intra-interpolation filter includes selecting the type of the one intra-interpolation filter from the different types of intra-interpolation filters based on the neighboring reconstructed samples of the current block, and the selected one intra-interpolation filter is one of a bilinear interpolation filter, a cubic interpolation filter, a spline interpolation filter, a DCT-based interpolation filter, or a DST-based interpolation filter. The method of claim 2.
4. the predetermined set of intra-interpolation filters include different numbers of taps; the selecting the one intra-interpolation filter includes selecting a number of taps of the one intra-interpolation filter based on the neighboring reconstructed samples of the current block, the number of taps of the one intra-interpolation filter being one of 2 taps, 4 taps, 6 taps, or 8 taps. The method of claim 2.
5. N is greater than 1, and the adjacent reconstructed samples include multiple lines of adjacent reconstructed samples; The step of selecting the one intra-interpolation filter includes, for each combination of an intra-interpolation filter from the set of intra-interpolation filters and adjacent reconstructed samples of a line from the plurality of lines of adjacent reconstructed samples: predicting, using a respective intra-interpolation filter, neighboring reconstructed samples of each line based on one or more remaining lines of neighboring reconstructed samples of the plurality of lines to obtain at least one prediction error; selecting the one intra-interpolation filter corresponding to the smallest prediction error among the obtained prediction errors; having The method of claim 1.
6. 2. The method of claim 1, wherein whether the neighboring reconstructed samples include (i) neighboring reconstructed samples of a left line to the left of the current block, (ii) neighboring reconstructed samples of a line above the current block, or (iii) neighboring reconstructed samples of the left line to the left of the current block and neighboring reconstructed samples of the line above the current block depends on the intra prediction direction of the angular intra prediction mode.
7. The method of claim 1 , wherein the value of N depends on the block size.
8. The method of claim 5 , wherein predicting neighboring reconstructed samples of each line comprises predicting neighboring reconstructed samples of each line based on an intra-prediction direction of the angular intra-prediction mode.
9. 6. The method of claim 5, wherein predicting neighboring reconstructed samples of each line comprises predicting neighboring reconstructed samples of each line based on a direction opposite to an intra-prediction direction of the angular intra-prediction mode.
10. 2. The method of claim 1 , wherein the selecting the one intra-interpolation filter comprises selecting the one intra-interpolation filter from the predetermined set of intra-interpolation filters and selecting the angular intra-prediction mode from a plurality of angular intra-prediction modes based on the neighboring reconstructed samples of the current block.
11. The selecting of the one intra-interpolation filter and the angular intra-prediction mode includes: deriving N1 angular intra prediction modes from the plurality of angular intra prediction modes associated with N1 lowest cost values from among the cost values of the respective angular intra prediction modes, wherein each cost value is determined based on the neighboring reconstructed samples of the current block, the respective angular intra prediction modes, and a default intra interpolation filter from the predetermined set of intra interpolation filters; selecting, for each of the N1 angular intra prediction modes, M1 intra interpolation filters from the predetermined set of intra interpolation filters associated with M1 lowest cost values among cost values of the predetermined set of intra interpolation filters, each cost value being determined based on the neighboring reconstructed samples of the current block, the respective intra interpolation filter, and a respective angular intra prediction mode among the N1 angular intra prediction modes; selecting the one intra-interpolation filter and the angular intra-prediction mode from N1×M1 combinations, each of the N1×M1 combinations including an intra-interpolation filter from the M1 intra-interpolation filters and an angular intra-prediction mode from the N1 angular intra-prediction modes; 11. The method of claim 10, comprising:
12. The selecting of the one intra-interpolation filter and the angular intra-prediction mode includes: selecting N2 angular intra prediction modes from a first plurality of modes of the plurality of angular intra prediction modes based on the neighboring reconstructed samples of the current block and an intra-interpolation filter; selecting a first updated intra-interpolation filter based on the neighboring reconstructed samples of the current block and one of the N2 angular intra-prediction modes; selecting N3 angular intra prediction modes from a second plurality of modes of the plurality of angular intra prediction modes based on the neighboring reconstructed samples of the current block and the first updated intra interpolation filter; selecting a second updated intra-interpolation filter based on the neighboring reconstructed samples of the current block and one of the N3 angular intra-prediction modes; selecting the second updated intra-interpolation filter as the selected one of the intra-interpolation filters and the one of the N3 angular intra-prediction modes based on (i) a difference between a first prediction error associated with the first updated intra-interpolation filter and the one of the N2 angular intra-prediction modes and a second prediction error associated with the second updated intra-interpolation filter and the one of the N3 angular intra-prediction modes, or (ii) the second prediction error; 11. The method of claim 10, comprising:
13. The selecting of the one intra-interpolation filter and the angular intra-prediction mode includes: For each combination of an angular intra-prediction mode from the plurality of angular intra-prediction modes and an intra-interpolation filter from the predetermined set of intra-interpolation filters, determining a cost value for each combination based on the neighboring reconstructed samples of the current block; selecting K combinations based on the determined cost values; selecting the one intra-interpolation filter and the angular intra-prediction mode from the K combinations based on index information signaled in the bitstream; 11. The method of claim 10, comprising:
14. The step of selecting the one intra-interpolation filter includes: calculating feature values based on the neighboring reconstructed samples of the current block; selecting the one intra-interpolation filter based on the calculated feature values; 2. The method of claim 1, comprising:
15. The calculating of the feature value comprises: calculating the feature value as one of (i) an absolute gradient value associated with the neighboring reconstructed samples of the current block, or (ii) a difference between a minimum value of the neighboring reconstructed samples of the current block and a maximum value of the neighboring reconstructed samples of the current block; 15. The method of claim 14, comprising:
16. The step of selecting the one intra-interpolation filter includes: obtaining an output from a neural network, the input of which includes the neighboring reconstructed samples of the current block and the output of which indicates the one intra-interpolation filter to select; 2. The method of claim 1, comprising:
17. The method of claim 16 , wherein the inputs of the neural network further include an intra-prediction mode index indicating the angular intra-prediction mode.
18. one or more processors; one or more memories storing a computer program; and The computer program causes the one or more processors to perform the method of any one of claims 1 to 17. Device.
19. A computer program causing a computer to carry out the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Image decoding device, image encoding device, image processing system, and program
JP2020053724A
Intra-prediction mode concepts for block-based image coding.
JP2021519542A
DEVICE AND METHOD FOR INTRA PREDICTION - Patent application
JP2021528919A
Intra video coding using multiple reference filters - Patents.com
JP2022530240A
Interpolation filtering method and apparatus for intra prediction, computer program and electronic device
JP2022546774A