Method, apparatus, and computer program for selecting an intra interpolation filter for multi-line intra prediction.

By selectively applying edge-preserving and edge-smoothing filters based on reference line indices and block characteristics, the method addresses inefficiencies in multi-line intra-prediction, improving video encoding efficiency and compression performance.

JP2026063081APending Publication Date: 2026-04-10TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2026-01-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing video encoding technologies, such as HEVC and VVC, face challenges in effectively utilizing intra-prediction modes and interpolation filters for multi-line intra-prediction, particularly in handling edge regions and smooth image areas, leading to suboptimal compression efficiency.

Method used

A method for selecting intra interpolation filters based on reference line indices, applying edge-preserving filters like cubic and Gaussian filters differently for zero and non-zero reference lines, and adjusting filter taps and coefficients based on block size and intra-prediction modes to enhance prediction accuracy.

Benefits of technology

Improves video encoding efficiency by optimizing filter selection for different image regions, preserving edges and reducing noise, thereby enhancing compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063081000001_ABST
    Figure 2026063081000001_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for selecting an intra interpolation filter for multi-line intra prediction based on a reference line index for decoding a video sequence. [Solution] The method includes: identifying a set of reference lines associated with an encoding unit; applying a first type of interpolation filter to reference samples contained in a first reference line adjacent to the encoding unit, based on the fact that a first reference line in the set of reference lines is associated with a first reference line index, to generate a first set of prediction samples; and applying a second type of interpolation filter to reference samples contained in a second reference line not adjacent to the encoding unit, based on the fact that a second reference line in the set of reference lines is associated with a second reference line index, to generate a second set of prediction samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - reference to Related Applications) This application claims priority under 35 U.S.C. 119 to U.S. Provisional Application No. 62 / 729,395, filed on September 10, 2018, and U.S. Patent Application No. 16 / 202,902, filed on November 28, 2018, with the United States Patent and Trademark Office, the disclosures of which are hereby incorporated by reference in their entireties.

[0002] This disclosure relates to a set of advanced video encoding techniques. More specifically, this disclosure provides a modified intra - interpolation filter scheme for multi - line intra - prediction.

Background Art

[0003] ITU - T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), and provided updates in 2014 (version 2), 2015 (version 3), and 2016 (version 4). Since then, ITU has been researching the potential need for standardization of future video encoding technologies with compression capabilities significantly beyond the HEVC standard (including extensions).

[0004] In October 2017, ITU issued a Call for Proposals (CfP) on video compression with capabilities beyond HEVC. By February 15, 2018, a total of 22 CfP responses regarding Standard Dynamic Range (SDR), 12 CfP responses regarding High Dynamic Range (HDR), and 12 CfP responses regarding the 360 - video category were each submitted.

[0005] In April 2018, 122 MPEG / 10 thAt the JVET (Joint Video Exploration Team - Joint Video Experts Team) meeting, all received CfP responses were evaluated. Based on this careful evaluation, JVET has formally begun the standardization of a next-generation video encoding other than HEVC, so-called Versatile Video Coding (VVC).

[0006] HEVC has a total of 35 intra-predictive modes: mode 10 is the horizontal mode, mode 26 is the vertical mode, and modes 2, 18, and 34 are the diagonal modes. The intra-predictive modes are signaled by three most probable modes (MPMs) and the remaining 32 modes.

[0007] In the current development of VVC, there are a total of 87 intra-prediction modes, with mode 18 being the horizontal mode, mode 50 being the vertical mode, and modes 2, 34, and 66 being the diagonal modes. Modes 1 through 10 and modes 67 through 76 are designated as wide-angle intra-prediction (WAIP) modes.

[0008] To encode an intra-mode, a list of three most likely modes (MPMs) is generated based on the intra-modes of adjacent blocks. This MPM list is called the MPM list or primary MPM list. If an intra-mode is not included in the MPM list, a flag is signaled to indicate whether the intra-mode belongs to the selected mode.

[0009] In the development of VVC, methods for implementing primary and secondary MPM lists have been proposed. Modes in the secondary MPM list are not included in the primary MPM list. The number of modes in the MPM list can be 3, 4, 5, 6, 7, 8, etc., and the number of modes in the secondary MPM list can be 8, 16, 32, etc.

[0010] In VVC, for luminance components, adjacent samples used to generate intra-prediction samples are filtered (i.e., intra-smoothing) before the generation process. Filtering is controlled by a given intra-prediction mode and transformation block size. If the intra-prediction mode is DC or the transformation block size is 4x4, adjacent samples are not filtered. Also, if the distance between a given intra-prediction mode and the vertical mode (or horizontal mode) is greater than a predetermined threshold, filtering is initiated. [1,2,1] filters and bilinear filters are used for filtering adjacent samples. For example, article 8.2.4.2.4 and Table 8-4 of VVC Draft 2 describe the intra-smoothing process proposed in VVC.

[0011] Multiline intra-prediction was proposed to use additional reference lines for intra-prediction. The encoder can determine and signal the reference lines used to generate the intra-predictor. Prior to the intra-prediction mode, the reference line index is signaled, and if a non-zero reference line index is signaled, the planar mode / DC mode is excluded from the intra-prediction mode.

[0012] Wide-angle prediction modes have been proposed that exceed the range of prediction directions covered by conventional intra-prediction modes, and are called wide-angle intra-prediction modes. These wide angles apply only to the following non-square blocks: When the width of the block is greater than the height of the block, the angle is greater than 45 degrees in the upper right direction (HEVC intra-prediction mode 34). When the height of the block is greater than the width of the block, the angle is greater than 45 degrees in the lower left direction (HEVC intra-prediction mode 2).

[0013] Using the original method, the replaced modes are signaled and remapped to the wide-angle mode index after analysis. The total number of intra-predictive modes remains unchanged, i.e., 35 for VTM-1.0 and 67 for BMS-1.0, and the encoding of intra-modes remains unchanged.

[0014] A bilateral filter is a nonlinear, edge-preserving, noise-reducing smoothing filter used for images. It replaces the luminance of each pixel with a weighted average of luminance values ​​from adjacent pixels. These weights can be based on a Gaussian distribution. The weights depend on the Euclidean distance between pixels and, further, on differences in radiation (e.g., differences in range such as color intensity or depth distance). Bilateral filters are useful for preserving sharp edges. Given an original (unfiltered) reference sample I(x) in an intrablock, the bilateral filter function is defined as follows: JPEG2026063081000002.jpg29151

[0015] To generate prediction samples used for directional intra-prediction, 4-tap and 6-tap intra-interpolation filters were proposed. Two types of 4-tap interpolation filters are used: a cubic interpolation filter and a Gaussian interpolation filter. The cubic filter is suitable for preserving image edges, while the Gaussian interpolation filter is suitable for removing image noise. The cubic interpolation filter is implemented when the intra-prediction mode is diagonal mode (i.e., mode 34) or higher and the block width is 8 or less, and when the intra-prediction mode is diagonal mode (i.e., mode 34) or lower and the block height is 8 or less.

[0016] A Gaussian filter is implemented when the intra-prediction mode is greater than or equal to the diagonal mode (i.e., mode 34) and the block width is greater than 8, and when the intra-prediction mode is less than or equal to the diagonal mode (i.e., mode 34) and the block height is greater than 8. [Overview of the project]

[0017] A method for selecting an intra interpolation filter for multiline intra prediction based on a reference line index for decoding a video sequence, comprising: identifying a set of reference lines associated with an encoding unit; applying a first type of interpolation filter to reference samples contained in the first reference line that is adjacent to the encoding unit, based on the fact that a first reference line in the set of reference lines is associated with a first reference line index, to generate a first set of prediction samples; and applying a second type of interpolation filter to reference samples contained in the second reference line that is not adjacent to the encoding unit, based on the fact that a second reference line in the set of reference lines is associated with a second reference line index, to generate a second set of prediction samples.

[0018] A device for selecting an intra interpolation filter for multiline intra prediction based on a reference line index for decoding a video sequence, comprising: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code, wherein the program code comprises: recognition code configured to cause the at least one processor to identify a set of reference lines associated with an encoding unit; first application code configured to cause the at least one processor to apply a first type of interpolation filter to reference samples contained in the first reference line adjacent to the encoding unit, based on the fact that a first reference line in the set of reference lines is associated with a first reference line index, thereby generating a first set of prediction samples; and second application code configured to cause the at least one processor to apply a second type of interpolation filter to reference samples contained in the second reference line not adjacent to the encoding unit, based on the fact that a second reference line in the set of reference lines is associated with a second reference line index, thereby generating a second set of prediction samples.

[0019] A non-temporary computer-readable medium for storing instructions, the instructions comprising one or more instructions, the instructions causing the one or more processors of a device that selects an intra interpolation filter for multi-line intra prediction based on a reference line index for decoding a video sequence, to perform the following operations: identify a set of reference lines associated with an encoding unit; apply a first type of interpolation filter to a reference sample in the first reference line adjacent to the encoding unit, based on the fact that a first reference line in the set of reference lines is associated with a first reference line index, to generate a first set of prediction samples; and apply a second type of interpolation filter to a reference sample in the second reference line not adjacent to the encoding unit, based on the fact that a second reference line in the set of reference lines is associated with a second reference line index, to generate a second set of prediction samples. [Brief explanation of the drawing]

[0020] Other features, characteristics, and various advantages of the disclosed subject matter will become clearer from the following detailed description and accompanying drawings. In the drawings, [Figure 1] This is an illustrative flowchart of the process for selecting an intra interpolation filter for multi-line intra prediction based on a reference line index for decoding a video sequence. [Figure 2] This is a simplified block diagram of a communication system according to an embodiment of the disclosure. [Figure 3] This is a diagram showing the arrangement of video encoders and decoders in a streaming environment. [Figure 4] This is a functional block diagram of a video decoder according to the embodiments of this disclosure. [Figure 5] This is a functional block diagram of a video encoder according to an embodiment of the present disclosure. [Figure 6] This is a diagram of a computer system according to an embodiment.

Best Mode for Carrying Out the Invention

[0021] The Gaussian filter is suitable for smooth image regions, and the cubic filter is suitable for edge regions. The addition of additional reference lines used in multi-line intra prediction is useful for edge regions. However, it is not a desirable design to apply both the same cubic filter and Gaussian filter to all reference lines.

[0022] FIG. 1 is a flowchart of an exemplary process 100 for selecting an intra interpolation filter for multi-line intra prediction based on a reference line index for decoding a video sequence. In some implementations, one or more process blocks of FIG. 1 can be executed by a decoder. In some embodiments, one or more process blocks of FIG. 1 can be executed by a device or group of devices (e.g., an encoder) separate from or including the decoder.

[0023] As shown in FIG. 1, process 100 can include identifying a set of reference lines associated with an encoding unit (block 110).

[0024] In some embodiments, the line index of the closest reference line is 0 (zero reference line). Also, the maximum number of reference lines signaled is represented as N. Hereinafter, the intra interpolation filter described is an interpolation filter for generating a predicted value indicating the fractional position of a reference sample.

[0025] In some implementations, the selection of an intra interpolation filter depends on the signaled reference line index. Types of intra interpolation filters include edge-preserving filters and edge-smoothing filters and / or similar filters. Edge-smoothing filters include linear interpolation filters with positive or zero filter coefficients. Edge-preserving filters include linear or nonlinear interpolation filters with at least one or two negative filter coefficients, such as bilateral filters.

[0026] As further shown in Figure 1, the process 100 may include applying a first type of interpolation filter to reference samples contained in the first reference line that is adjacent to the coding unit, based on the fact that a first reference line in the set of reference lines is associated with a first reference line index, to generate a first set of prediction samples (block 120), and applying a second type of interpolation filter to reference samples contained in the second reference line that is not adjacent to the coding unit, based on the fact that a second reference line in the set of reference lines is associated with a second reference line index, to generate a second set of prediction samples (block 130).

[0027] According to the embodiment, both an edge smoothing filter and an edge preservation filter are applied to zero reference lines, while only an edge preservation filter is applied to non-zero reference lines. The edge smoothing filter may include a bilinear filter or a Gaussian filter.

[0028] According to the embodiment, the edge-preserving filter may be a cubic interpolation filter, a DCT-based interpolation filter (DCT-IF), a 4 / 6 tap polynomial-based interpolation filter, a bilateral filter, a Hermitian interpolation filter, and / or a similar filter.

[0029] According to the embodiment, the edge-preserving filters used for different lines may differ. For example, the taps of the edge-preserving filter may differ for zero-reference lines and non-zero-reference lines. Furthermore, or alternatively, an M-tapped edge-preserving filter may be used for zero-reference lines, and an N-tapped edge-preserving filter may be used for non-zero-reference lines (for example, M and N are positive integers, and M is not equal to N). Furthermore, or alternatively, the filter coefficients of the edge-preserving filters used for different reference lines may differ.

[0030] In other embodiments, the coefficients of the edge preservation filter and the edge smoothing filter are fixed and independent of the sample value.

[0031] In some implementations, the selection of the intra interpolation filter depends on the reference line index and other encoded information, or on any information available to both the encoder and decoder, such as the intra prediction mode and block size.

[0032] According to the embodiment, an edge preservation filter is used based on conditions that are met. For example, the conditions may be met if the intra-prediction mode is greater than or equal to the diagonal mode (i.e., mode 34) and the block width is less than or equal to S. Furthermore, or alternatively, the conditions may be met if the intra-prediction mode is less than or equal to the diagonal mode (i.e., mode 34) and the block height is less than or equal to S. For example, S may represent a threshold for the block width or block height. S may differ depending on the line. For example, according to the embodiment, for a zero reference line, S is 8, and for a non-zero reference line, S is 16 or 32.

[0033] According to one embodiment, the edge smoothing filter is used for wide angles on the zero reference line, and the edge preservation filter is used for wide angles on the non-zero reference line. Alternatively, the edge preservation filter is used for wide angles on the zero reference line, and the edge smoothing filter is used for wide angles on the non-zero reference line.

[0034] According to one embodiment, both the edge-preserving filter and the edge-smoothing filter are applied to the zero reference line, and / or the edge-preserving filter or the edge-smoothing filter is applied to at least one reference line. In this case, it is possible that neither filter is applied. According to another embodiment, if there are three reference lines, the line indices may be {0,1,2} or {0,1,3}, and both the edge-preserving filter and the edge-smoothing filter are applied to line 0, the edge-preserving filter or the edge-smoothing filter only is applied to line 1, and the edge-smoothing filter or the edge-preserving filter only is applied to line 2 or line 3.

[0035] In the embodiment, the edge preservation filter and the edge smoothing filter are applied to all reference lines, but the number of taps in the intra interpolation filter may differ for each line. For example, in the embodiment, the edge preservation filter used for zero reference lines has M taps, and the edge preservation filter used for non-zero reference lines has N taps (for example, M and N are positive integers, and M is not equal to N, e.g., M=6 and N=4). Alternatively, the edge smoothing filter used for zero reference lines has M taps, and the edge smoothing filter used for non-zero reference lines has N taps (for example, M and N are positive integers, and M is not equal to N, e.g., M=6 and N=4).

[0036] According to the embodiment, for some reference lines, wide angles are prohibited or differently limited. For example, wide angles are prohibited for non-zero reference lines. Furthermore, or instead, at least one reference line does not use wide angles. Furthermore, or instead, the wide angles used for different reference lines are different. For example, the number of conventional angular intra-prediction directions replaced by wide-angle intra-prediction directions depends on the reference line index.

[0037] According to the embodiment, the intra-smoothing filter may differ for each line. For example, a bilateral filter is used for intra-smoothing of non-zero reference lines. Furthermore, or alternatively, if a reference line index is signaled, the intra-smoothing filter is prohibited for non-zero reference lines. Furthermore, or alternatively, at least one reference line is intra-smoothed using a bilateral filter.

[0038] According to the embodiment, the number of filter taps in the intra-smoothing filter differs for each line. For example, the number of filter taps in the intra-smoothing filter used for zero-reference lines is M, and the number of filter taps in the intra-smoothing filter used for non-zero-reference lines is N (for example, M and N are positive integers, and M is not equal to N, for example, M=6 and N=4).

[0039] According to the embodiment, the intra-prediction mode to which intra-smoothing is applied differs for each line. For example, the threshold T defines which intra-prediction mode applies intra-smoothing when the intra-prediction mode Mode satisfies the following conditions: min(abs(Mode-Hor),abs(Mode-Ver)) <Tであり、 Here, "Hor" represents the intra-predictive mode index for horizontal mode, and "Ver" represents the intra-predictive mode index for vertical mode. The value of T depends on the reference line index and other encoded information, or any information known to both the encoder and decoder, including, but not limited to, block area size, block width, block height, and block aspect ratio.

[0040] According to one embodiment, for non-zero reference lines, only MPM mode is permitted, including the primary and secondary MPM lists. In one embodiment, if the reference line index of the current block is greater than 0, one bin is signaled to indicate whether the intra-prediction mode of the current block belongs to the primary or secondary MPM list, and is called "primary_mpm_flag". If "primary_mpm_flag" is "true", the primary MPM index is signaled. Otherwise, the secondary MPM index is signaled. Non-MPM mode is not used for non-zero reference lines, and no flag is signaled to indicate whether to use secondary or non-MPM.

[0041] According to this embodiment, the planar mode and DC mode are excluded from the primary MPM list and the secondary MPM list.

[0042] According to the embodiment, several implementations encode the first bin of the reference line index using multiple contexts. For example, the selection of context depends on the reference line index of the adjacent block. As a specific example, if both the reference line index of the block to the left and the reference line index of the block above are equal to 0, context 0 is selected; and if both the reference line index of the block to the left and the reference line index of the block above are not equal to 0, context 1 is selected; otherwise, context 2 is selected.

[0043] According to the embodiment, context selection depends on the CBF (Encoded Block Flag) of the adjacent block. The CBF is an indicator of whether the current block contains non-zero coefficients. If the CBF is equal to 0, the current block does not contain non-zero coefficients. For example, if both the CBF of the block to the left and the CBF of the block above are equal to 0, context 0 is selected. Also, if both the CBF of the block to the left and the CBF of the block above are not equal to 0, context 1 is selected. Otherwise, context 2 is selected.

[0044] According to the embodiment, the context for encoding the transformation selection information according to the reference line index value includes, but is not limited to, the MTS flag, the MTS index, and the NSST index.

[0045] Figure 1 shows an exemplary block of process 100, but in some implementations, process 100 may include other blocks, fewer blocks, different blocks, or blocks arranged differently compared to these blocks shown in Figure 1. Furthermore, or instead, two or more blocks of process 100 may be executed in parallel.

[0046] Figure 2 shows a schematic block diagram of a communication system (200) according to an embodiment of the present disclosure. The communication system (200) may include at least two terminals (210-220) interconnected via a network (250). In the case of one-way data transmission, the first terminal (210) can encode video data at its local location and transmit it to the other terminal (220) via the network (250). The second terminal (220) can receive the encoded video data from the other terminal via the network (250), decode the encoded data, and display the restored video data. One-way data transmission is common in media service applications and the like.

[0047] Figure 2 shows a second pair of terminals (230, 240) provided to support the bidirectional transmission of encoded video, for example, during a video conference. In the case of bidirectional data transmission, each terminal (230, 240) can encode video data captured at its local location and transmit it to the other terminal over the network (250). Each terminal (230, 240) may also receive the encoded video data transmitted by the other terminal, decode the encoded data, and display the restored video data on a local display device.

[0048] In Figure 2, terminals (210-240) are illustrated as servers, personal computers, and smartphones, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure apply to laptop computers, tablets, media players, and / or dedicated video conferencing equipment. Network (250) refers to any number of networks that transmit encoded video data between terminals (210-240), including, for example, wired and / or wireless communication networks. Communication network (250) can exchange data over circuit-switched and / or packet-switched channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of network (250) may not be important to the operation of this disclosure unless described below herein.

[0049] As an example of the application of the disclosed subject matter, Figure 3 shows a configuration of video encoders and decoders in a streaming environment, and the disclosed subject matter can be equivalently used in other video-supporting applications, such as video conferencing, digital TV, and storage of compressed video on digital media such as CDs, DVDs, and memory sticks.

[0050] The streaming system may include a capture subsystem (313), which may include a video source (301) (e.g., a digital camera) to create, for example, an uncompressed video sample stream (302). The video sample stream (302) is drawn as a thick line to highlight a larger amount of data compared to the encoded video bitstream. The sample stream (302) can be processed by an encoder (303) coupled to the camera (301). The encoder (303) may include hardware, software, or a combination thereof to realize or implement each aspect of the disclosed subject matter, which will be described in more detail below. The encoded video bitstream (304) is drawn as a thin line to highlight a smaller amount of data compared to the sample stream. The encoded video bitstream (304) can be stored in a streaming server (305) for future use. One or more streaming clients (306, 308) can access a streaming server (305) to retrieve copies (307, 309) of an encoded video bitstream (304). A client (306) may include a video decoder (310) that decodes an incoming copy of the encoded video bitstream (307) to create an outgoing video sample stream (311) that can be rendered on a display (312) or other rendering device (not shown). Some streaming systems may encode the video bitstream (304, 307, 309) according to specific video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. Video encoding standards, informally called multipurpose video encodings, are under development. The disclosed subject matter can be used in the context of VVC.

[0051] Figure 4 may be a functional block diagram of a video decoder (310) according to an embodiment of the present disclosure.

[0052] The receiver (410) can receive one or more codec video sequences to be decoded by the decoder (310). In the same or different embodiments, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. Encoded video sequences can be received from a channel (412), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (410) may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams that can be transferred to their respective user entities (not shown). The receiver (410) can isolate the encoded video sequence from other data. To prevent network jitter, a buffer memory (415) may be coupled between the receiver (410) and the entropy decoder / parser (420) (hereinafter referred to as the “parser”). When the receiver (410) receives data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous real-time network, the buffer (415) may not be necessary or may be small. For use in best-effort packet networks such as the Internet, the buffer (415) may be necessary, and the buffer (415) may be relatively large and may have a self-adaptive size to its advantage.

[0053] The video decoder (310) may include a parser (420) to reconstruct symbols (421) based on the entropy-encoded video sequence. These symbol categories include information for managing the operation of the decoder (310) and latent information for controlling rendering devices, such as a display (312), which are not components of the decoder but may be coupled to it. As shown in Figure 4, the control information for one or more rendering devices may be in the form of Supplementary Enhancement Information (SEI messages) or Video Usability Information (VUI) parameter set fragments (not shown). The parser (420) can analyze / entropy-decode the received encoded video sequence. The encoding of the encoded video sequence may conform to video coding techniques or standards and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, and arithmetic coding with or without context. The parser (420) can extract a set of subgroup parameters for at least one of the pixel subgroups in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups may include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The entropy decoder / parser may also extract from the encoded video sequence information, for example, transform coefficients, quantizer parameter (QP) values, motion vectors, etc.

[0054] The parser (420) can create symbols (421) by performing an entropy decoding / analysis operation on the video sequence received from the buffer (415). The parser (420) may receive encoded data and selectively decode specific symbols (421). The parser (420) may also determine whether a particular symbol (421) is provided to a motion compensation prediction unit (453), a scaler / inverse transform unit (451), an intra prediction unit (452), or a loop filter (456).

[0055] The reconstruction of the symbol (421) may involve multiple different units, depending on the encoded video picture or some type of encoded video picture (e.g., interpicture and intrapicture, interblock and intrablock) and other factors. The units involved and the manner of involvement may be controlled by subgroup control information analyzed by the parser (420) from the encoded video sequence. For brevity, the flow of such subgroup control information between the parser (420) and the following multiple units will not be described.

[0056] In addition to the functional blocks already mentioned, the decoder (310) can conceptually be subdivided into several functional units, which are described below. In actual implementations carried out under commercial constraints, many of these units can interact closely with each other and integrate with each other at least partially. However, for the purpose of illustrating the disclosed subject, it is appropriate to conceptually subdivide it into the following functional units.

[0057] The first unit is the scaler / inverse unit (451). The scaler / inverse unit (451) receives quantization conversion coefficients and control information as symbols (421) (one or more) from the parser (420), which include the conversion method to use, block size, quantization coefficients, quantization scaling matrix, etc. The scaler / inverse unit (451) can output a block containing sample values, which can be input to the aggregator (455).

[0058] In some cases, the output samples of the scaler / inverse unit (451) may belong to intra-encoded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but do use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (452). In some cases, the intra-picture prediction unit (452) generates a block of the same size and shape as the block being reconstructed, using information extracted from the current (partially reconstructed) picture (456) that has already been reconstructed. In some cases, the aggregator (455) adds the prediction information generated by the intra-prediction unit (452) to the output sample information provided by the scaler / inverse unit (451), based on each sample.

[0059] In other cases, the output samples of the scaler / inverse unit (451) may belong to an inter-coding and latent motion compensation block. In such cases, the motion compensation prediction unit (453) may access the reference picture memory (457) to extract samples for prediction. After motion compensation is performed on the extracted samples based on the symbols (421) belonging to the block, these samples can be added by the aggregator (455) to the output of the scaler / inverse unit (in this case, called residual samples or residual signals) to generate output sample information. The address in the reference picture memory from which the motion compensation unit extracts prediction samples may be controlled by a motion vector, which can be used by the motion compensation unit in the form of a symbol (421), which may have, for example, X, Y, and reference picture components. Motion compensation may include interpolation of sample values ​​extracted from the reference picture memory, a motion vector prediction mechanism, etc., when the exact motion vector of the subsample is used.

[0060] The output samples of the aggregator (455) may be processed by various loop filtering techniques in the loop filter unit (456). The video compression technique may include an in-loop filtering technique, which is controlled by parameters included in the encoded video bitstream and made available to the loop filter unit (456) as symbols (421) from the parser (420). However, the video compression technique may respond to metadata obtained during the period of decoding earlier portions (in decoding order) of the encoded picture or encoded video sequence, or to previously reconstructed and loop-filtered sample values.

[0061] The output of the loop filter unit (456) may be a sample stream, which may be output to the rendering device (312) and stored in the reference picture memory (456) for use in future inter-picture prediction.

[0062] Once fully reconstructed, some encoded pictures can be used as reference pictures for future predictions. When an encoded picture is fully reconstructed and identified as a reference picture (e.g., by the parser (420)), the current reference picture (656) can become part of the reference picture buffer (457) and can then reallocate new current picture memory before starting the subsequent reconstruction of the encoded picture.

[0063] The video decoder (310) may perform the decoding operation in accordance with a specified video compression technique, for example, as recorded in the ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence conforms to the syntax of the video compression technique or standard (as specified in the video compression technique documentation or standard, and in particular in the document files therein). Compliance also requires that the complexity of the encoded video sequence is within the limits set by the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the specification of the Hypothetical Reference Decoder (HRD) and metadata for HRD buffer management signaled in the encoded video sequence.

[0064] In an embodiment, the receiver (410) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of one or more encoded video sequences. The additional data may be used by the video decoder (310) to accurately decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, a time, space, or signal-to-noise ratio (SNR) extension layer, redundant slices, redundant pictures, or forward error correction codes.

[0065] Figure 5 may be a functional block diagram of a video encoder (303) according to an embodiment of the present disclosure.

[0066] The encoder (303) may receive video samples from a video source (301) (which is not part of the encoder), and the video source may capture (one or more) video images that are encoded by the encoder (303).

[0067] The video source (301) may provide a source video sequence in the form of a digital video sample stream encoded by an encoder (303), the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 Y CrCb, RGB, etc.), and any suitable sampling configuration (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (503) may be a camera that captures local image information as a video sequence. The video data may be provided as a series of separate pictures that are given motion when viewed sequentially. These pictures themselves may be organized as a spatial pixel array, and each pixel may contain one or more samples depending on the sampling configuration, color space, etc. used. Those skilled in the art will readily understand the relationship between pixels and samples. The following description will focus on samples.

[0068] According to the embodiment, the encoder (303) may encode pictures of a source video sequence in real time or under any other time constraints required by the application and compress them into an encoded video sequence (543). Implementing an appropriate encoding rate is one of the functions of the controller (550). The controller controls and is functionally coupled to other functional units described below. For brevity, the coupling is not depicted. Parameters set by the controller may include rate control-related parameters (picture skip, quantizer, λ value of rate distortion optimization technique, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc. Those skilled in the art will readily recognize the other functions of the controller (550), which relate to the video encoder (303) optimized for a particular system design.

[0069] Some video encoders operate in a manner readily recognizable to those skilled in the art as an "encoding loop." In a very simplified description, the encoding loop may include an encoding portion of an encoder (530) (hereinafter referred to as the "source encoder") (responsible for creating symbols based on the input picture to be encoded and one or more reference pictures), and a (local) decoder (533) embedded in the encoder (303), which reconstructs the symbols to create sample data that is also created by a (remote) decoder (since compression between the symbols and the encoded video bitstream is lossless in the video compression techniques considered in the disclosed subject). The reconstructed sample stream is input to a reference picture memory (534). Decoding of the symbol stream results in bit-accurate results regardless of the decoder's location (local or remote), so that the contents of the reference picture buffer are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture sample that the predictive part of the encoder "sees" is exactly the same as the sample value that the decoder "sees" when using the prediction during decoding. The basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.

[0070] The operation of the “local” decoder (533) may be the same as the operation of the “remote” decoder (310) as described above with reference to Figure 4. However, also referring briefly to Figure 5, if symbols are available and the entropy encoder (545) and parser (420) can losslessly encode / decode the symbols into the encoded video sequence, the entropy decoding portion of the decoder (310), including the channel (412), receiver (410), buffer (415), and parser (420), may not be fully realized by the local decoder (533).

[0071] In this case, in addition to the parsing / entropy decoding present in the decoder, it becomes clear that any decoder technique also inevitably exists in the corresponding encoder in essentially the same functional form. Since encoder techniques and fully described decoder techniques are inverses of each other, the explanation of encoder techniques can be simplified. A more detailed explanation is necessary only in specific areas and is provided below.

[0072] As part of the operation of the source encoder (530), the source encoder (530) may perform motion-compensated predictive coding, which predictively encodes the input frame by referencing one or more previously encoded frames designated as “reference frames” from a video sequence. In this way, the encoding engine (532) may encode the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame that may be selected as the predictive reference for the input frame.

[0073] A local video decoder (533) may decode the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source encoder (530). The operation of the encoding engine (532) may, advantageously, be lossy. If the encoded video data may be decoded by a video decoder (not shown in Figure 5), the reconstructed video sequence may typically be a copy of the source video sequence with some error. The local video decoder (533) may copy the decoding process that may be performed by the video decoder on the reference frame and store the reconstructed reference frame in a reference picture buffer (534). In this way, the encoder (303) can locally store a copy of the reconstructed reference frame which has the same content (no transmission error) as the reconstructed reference frame obtained by the remote video decoder.

[0074] The predictor (535) can perform a predictive search on the encoding engine (532). That is, for a new frame to be encoded, the predictor (535) may search the reference picture memory (534) for sample data that can be used as a suitable predictive reference for the new picture (as a candidate reference pixel block), or for specific metadata such as the motion vector or block shape of a reference picture. The predictor (535) can operate pixel block by pixel based on the sample blocks to find a suitable predictive reference. In some cases, the input picture may have predictive references extracted from multiple reference pictures stored in the reference picture memory (534), for example, as determined by the search results obtained by the predictor (535).

[0075] The controller (550) can manage the encoding operation of the source encoder (530), including, for example, setting parameters and subgroup parameters for encoding video data.

[0076] In the entropy encoder (545), entropy coding may be performed on the outputs of all the functional units mentioned above. The entropy encoder converts the symbols generated by each functional unit into an encoded video sequence by performing lossless compression on the symbols according to techniques well known to those skilled in the art, such as Huffman coding, variable-length coding, and arithmetic coding.

[0077] The transmitter (540) may buffer the encoded video sequence created by the entropy encoder (545) to prepare it for transmission over a communication channel (560), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (540) may merge the encoded video data from the video encoder (530) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0078] The controller (550) can manage the operation of the encoder (303). During encoding, the controller (550) can assign a specific encoded picture type to each encoded picture, which may affect the encoding technique applicable to the corresponding picture. For example, a picture may typically be assigned as one of the following frame types:

[0079] An intra-picture (I-picture) can be a picture that is encoded and decoded without using any other frame in the sequence as a source for prediction. Some video codecs allow different types of intra-pictures, including, for example, independent decoder refresh pictures. Those skilled in the art will know these variations of I-pictures and their corresponding uses and characteristics.

[0080] A predictive picture (P-picture) may be a picture that is encoded and decoded using intra-prediction or inter-prediction, which can predict the sample value of each block using at most one motion vector and reference index.

[0081] A bidirectional predictive picture (B-picture) can be a picture encoded and decoded using intra-prediction or inter-prediction, which predicts the sample values ​​of each block using at most two motion vectors and a reference index. Similarly, a multi-predictive picture can use two or more reference images and associated metadata to reconstruct a single block.

[0082] A source picture is generally subdivided spatially into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block may be encoded. These blocks may be predictively encoded by referencing other (encoded) blocks determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I-picture may be non-predictively encoded, or these I-picture blocks may be predictively encoded by referencing encoded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively encoded via spatial or temporal prediction by referencing one previously encoded reference picture. Blocks of a B-picture may be predictively encoded via spatial or temporal prediction by referencing one or two previously encoded reference pictures.

[0083] The video encoder (303) can perform encoding operations based on a specified video encoding technique or standard, such as ITU-T Recommendation H.265. During its operation, the video encoder (303) can perform various compression operations, including predictive encoding operations based on temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.

[0084] In an embodiment, the transmitter (540) may transmit additional data along with the encoded video. The video encoder (530) may include such data as part of the encoded video sequence. The additional data may include a temporal / spatial / SNR extension layer, other forms of redundant data such as redundant pictures and slices, supplementary extension information (SEI) messages, visual usability information (VUI) parameter set fragments, and the like.

[0085] Furthermore, the proposed method can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors perform one or more of the proposed methods by executing a program stored in a non-temporary computer-readable medium.

[0086] The above techniques may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, Figure 6 shows a computer system 600 suitable for implementing some embodiments of the disclosed subject matter.

[0087] Computer software can be coded using any suitable machine code or computer language, and by applying mechanisms such as assembly, compilation, and linking to any suitable machine code or computer language, code can be created that contains instructions that are executed directly by a computer's central processing unit (CPU), graphics processing unit (GPU), etc., or by interpretation, microcode, etc.

[0088] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, and Internet of Things devices.

[0089] The components used in the computer system 600 shown in Figure 6 are illustrative in nature and are not intended to limit the scope or functionality of the computer software used to implement embodiments of this disclosure. The arrangement of the components should not be interpreted as having any dependency or requirement on any of the components or combinations thereof shown in the exemplary embodiment of computer system 600.

[0090] The computer system 600 may include several human-machine interface input devices. Such human-machine interface input devices can respond to one or more inputs from a human user, such as tactile input (e.g., keystrokes, slides, data glove movements), audio input (e.g., voices, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voices, music, ambient sounds), images (e.g., scanned images, photographic images acquired from a still image capture device), and video (e.g., 2D video, 3D video including stereo video).

[0091] The input human-machine interface device may have one or more of the following: keyboard 601, mouse 602, touchpad 603, touch panel 610, data glove, joystick 605, microphone 606, scanner 607, and imaging device (camera) 608 (only one of each is shown in the illustration).

[0092] The computer system 600 may further have human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human-machine interface output devices include tactile output devices (e.g., touch panel 610, tactile feedback via data glove or joystick 605, although tactile feedback devices that are not used as input devices also exist), audio output devices (e.g., speaker 609, headphones (not shown)), visual output devices (e.g., screen 610 including CRT screen, LCD screen, plasma screen, OLED screen, virtual reality glasses (not shown), holographic display and smoke tank (not shown), where each screen may or may not have touch panel input capability and tactile feedback capability, and some of them output two-dimensional visual output or three-dimensional or more output by means such as stereoscopic image output), and printers (not shown).

[0093] The computer system 600 may further have human-accessible storage devices and associated media, including, for example, optical media such as a CD / DVD ROM / RW 620 having a medium 621 such as a CD / DVD, a thumb drive 622, a removable hard drive or solid-state drive 623, conventional magnetic media such as magnetic tape and floppy disks (not shown), and devices based on dedicated ROM / ASIC / PLD (e.g., dongles (not shown)).

[0094] Those skilled in the art will understand that the term “computer-readable medium” as used in relation to the subject matter disclosed in this application does not include a transmission medium, carrier wave, or other instantaneous signal.

[0095] The computer system 600 may further include one or more interfaces to one or more communication networks. These networks may be, for example, wireless networks, wired networks, or optical networks. The networks may also be local networks, wide-area networks, metropolitan networks, automotive and industrial networks, real-time networks, or latency-tolerant networks. Examples of networks include local area networks (e.g., Ethernet, wireless LAN), cellular networks (including Global Mobile Communication Network (GSM), 3G, 4G, 5G, Long-Term Evolution (LTE), etc.), wired or wireless wide-area digital television networks (including wired television, satellite television, and terrestrial television), and automotive and industrial networks (including CANBus). Some networks generally require an adapter connected to an external network interface on a general-purpose data port or peripheral bus 649 (e.g., a USB port on computer system 600), while other networks are generally integrated into the core of computer system 600 by connecting to the system bus via, for example, an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system, as described below. Computer system 600 can communicate with other entities via any of these networks. Such communication may be one-way receive only (e.g., broadcast television), one-way transmit only (e.g., a CAN bus to a CAN bus device), or bidirectional (e.g., reaching other computer systems using a local area or wide area digital network). Specific protocols and protocol stacks may be available for each of these networks and network interfaces described above.

[0096] The aforementioned human-machine interface devices, human-accessible storage devices, and network interfaces may be connected to the core 640 of the computer system 600.

[0097] The Core 640 includes one or more central processing units (CPUs) 641, dedicated programmable processing units in the form of graphics processing units (GPUs) 642, field-programmable gate arrays (FPGAs) 643, and hardware accelerators 644 for specific tasks. These devices, along with read-only memory (ROM) 645, random access memory 646, and internal mass memory 647 (such as internal, user-inaccessible hard disk drives and solid-state drives (SSDs)), are connected via the system bus 648. In a computer system, expansion with other CPUs, GPUs, etc., is possible by accessing the system bus 648 in the form of one or more physical plugs. Peripheral devices are connected to the core's system bus 648 directly or via peripheral buses 649. Peripheral bus architectures include peripheral component interconnects (PCI), USB, etc.

[0098] The CPU 641, GPU 642, FPGA 643, and accelerator 644 can execute several instructions, and these instructions can be combined to form the computer code described above. This computer code is stored in ROM 645 or RAM 646. Temporary data may be stored in RAM 646, and permanent data may be stored, for example, in internal mass memory 647. Fast storage and retrieval to any of the memory devices can be achieved by using cache memory, and this cache memory can be closely associated with one or more CPUs 641, GPUs 642, mass memory 647, ROM 645, RAM 646, etc.

[0099] A computer-readable medium may contain computer code for performing various operations that are performed by a computer. The medium and computer code may be a medium and computer code specifically designed and constructed for the purposes of this disclosure, or they may be of a type that is well known and available to those skilled in the computer software field.

[0100] To the extent of exemplary, not limiting, a computer system having architecture 600, particularly a core 640, can provide functionality by having one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media may be media relating to user-accessible mass storage devices as described above, and non-transient storage devices of the core 640, such as the core internal mass memory 647 or ROM 645. Software for implementing each embodiment of the present disclosure may be stored in such devices and executed by the core 640. Depending on the specific requirements, the computer-readable media may include one or more storage devices or chips. The software may cause the core 640, particularly its processors (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific parts of specific processes as described herein, which include limiting data configurations to be stored in RAM 646 and modifying such data configurations based on processes limited by the software. Alternatively, a computer system may provide functionality by being implemented in logic hardwired or otherwise in circuitry (e.g., accelerator 644), which, in place of or in conjunction with software, can perform specific processes or specific parts of specific processes described herein. Where appropriate, the software referred to may include logic, and conversely, the logic referred to may include software. Where appropriate, the computer-readable medium referred to may include circuitry (e.g., integrated circuit (IC)) in which the software to be executed is stored, circuitry embodying the logic to be executed, or both. This disclosure includes any appropriate combination of hardware and software.

[0101] While this disclosure has described several exemplary embodiments, there are various modifications, substitutions, and equivalent alternatives that fall within the scope of this disclosure. Therefore, it should be understood that a person skilled in the art could conceive of many systems and methods that embody the principles of this disclosure and thus fall within its spirit and scope, although these are not expressly shown or described herein.

Claims

[Claim 1] A method for selecting an intra interpolation filter for multiline intra prediction based on a reference line index for decoding a video sequence, Steps include identifying the reference line set associated with the coding unit, Based on the fact that a first reference line in the set of reference lines is associated with a first reference line index, the step of applying an edge-preserving filter and an edge-smoothing filter having only positive or zero filter coefficients to the reference samples included in the first reference line in order to generate a first set of prediction samples, wherein the first reference line is adjacent to the coding unit, A method comprising the steps of applying only an edge-preserving filter, rather than an edge-smoothing filter, to the reference samples contained in the second reference line in order to generate a second set of prediction samples, based on the fact that a second reference line in the set of reference lines is associated with a second reference line index, wherein the second reference line is not adjacent to the coding unit.