Intra prediction using multiple reference lines
By using a subset of intra-prediction modes with optimized reference line access and reduced codeword size, the inefficiencies in video compression due to multiple reference lines are addressed, enhancing encoding efficiency and reducing file size and bandwidth.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-25
AI Technical Summary
Existing video compression techniques using multiple reference lines increase signal transmission overhead, leading to decreased encoding efficiency due to the need to identify and signal multiple reference lines, which can outweigh the compression gain.
Implementing a subset of intra-prediction modes that have access to alternative reference lines, limiting modes excluded from the subset to primary reference lines, and encoding the reference line index only when necessary, along with mechanisms to reduce codeword size and storage requirements.
Reduces signal transmission overhead and increases compression efficiency by optimizing the use of reference lines, resulting in smaller file sizes and reduced bandwidth requirements.
Smart Images

Figure 2026053588000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This patent application claims priority to U.S. Non - Provisional Patent Application No. 15 / 972,870, filed on May 7, 2018, entitled "Intra Prediction Using Multiple Reference Lines", which claims priority to U.S. Provisional Patent Application No. 62 / 503,884, filed on May 9, 2017, by Shan Liu et al., entitled "Methods and Apparatus for Intra Prediction Using Multiple Reference Lines" and U.S. Provisional Patent Application No. 62 / 511,757, filed on May 26, 2017, by Xiang Ma et al., entitled "Methods and Apparatus for Intra Prediction Using Multiple Reference Lines", the teachings and disclosures of which are hereby incorporated in their entirety by reference thereto.
[0002] Statement Regarding Federally Sponsored Research or Development Not applicable.
[0003] Reference to Microfiche Appendix Not applicable.
Background Art
[0004] Even the amount of video data required to depict a relatively short video can be considerable, which can cause difficulties when the data is streamed or otherwise transmitted over communication networks with limited bandwidth. Therefore, video data is generally compressed before being transmitted over modern telecommunications networks. Video size can also be a problem when the video is stored in storage, as storage resources may be limited. Video compressors often use software and / or hardware at the source to encode the video data prior to transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompressor that decodes the video data. Given limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression techniques that improve the compression ratio with little to no sacrifice of image quality are desirable. [Overview of the project]
[0005] In one embodiment, the Disclosure includes: a receiver configured to receive a bitstream; a processor coupled to the receiver, configured to perform the steps of: determining an intra-prediction mode subset, wherein the intra-prediction mode subset includes intra-prediction modes correlated to a plurality of reference lines for the current image block, and excluding intra-prediction modes correlated to a primary reference line for the current image block; decoding the first intra-prediction mode by an alternate intra-prediction mode index when the first intra-prediction mode is included in the intra-prediction mode subset; and decoding the first intra-prediction mode by an intra-prediction mode index when the first intra-prediction mode is not included in the intra-prediction mode subset; and a display coupled to the processor, which presents video data including the image block decoded based on the first intra-prediction mode.
[0006] Optionally, in any aspect described above, as provided by another implementation of the aspect, the processor is further configured to: decode the reference line index when the first intra-prediction mode is included in the intra-prediction mode subset, the reference line index indicating the first reference line from the plurality of reference lines for the first intra-prediction mode; and not decode the reference line index when the first intra-prediction mode is not included in the intra-prediction mode subset.
[0007] Optionally, in any of the above aspects, as provided by another implementation of the aspect, the reference line index is located after the first intra-prediction mode in the bitstream.
[0008] Optionally, in any aspect described above, as provided by another implementation of the aspect, the intra-prediction mode subset includes a starting directional intra-prediction mode (DirS), an ending directional intra-prediction mode (DirE), and N directional intra-prediction modes between DirS and DirE, where N is a predetermined integer.
[0009] Optionally, in any of the above aspects, as provided by another implementation of said aspect, the intra-prediction mode subset further includes a planar prediction mode and a direct current (DC) prediction mode.
[0010] Optionally, in any of the above aspects, as provided by another embodiment of the aspect, the intra-prediction mode subset includes a starting directional intra-prediction mode (DirS), an ending directional intra-prediction mode (DirE), a middle directional intra-prediction mode (DirD), a horizontal directional intra-prediction mode (DirH), a vertical directional intra-prediction mode (DirV), and effective directional intra-prediction modes in positive or negative N directions of DirS, DirE, DirD, DirH, and DirV, where N is a predetermined integer value.
[0011] Optionally, in any of the above aspects, as provided by another implementation of said aspect, the intra-prediction mode subset further includes a planar prediction mode and a direct current (DC) prediction mode.
[0012] Optionally, in any of the above aspects, as provided by another implementation of the aspect, the intra-prediction mode subset includes an intra-prediction mode selected for a decoded adjacent block, the decoded adjacent block being located in a predetermined adjacent state to the current image block.
[0013] Optionally, in any of the aspects described above, as provided by another implementation of the aspect, the intra-predictive mode subset includes modes associated with the most probable mode (MPM) list for the current image block.
[0014] In one embodiment, the present disclosure includes a method comprising: storing a bitstream in memory containing an image block encoded as a prediction block; obtaining a current prediction block encoded by a direct current (DC) intra-prediction mode by a processor coupled to the memory; determining a DC prediction value that approximates the current image block corresponding to the current prediction block by determining the average of all reference samples in at least two of a plurality of reference lines associated with the current prediction block; reconstructing the current image block based on the DC prediction value by the processor; and displaying a video frame containing the current image block on a display.
[0015] Optionally, in any of the above aspects, as provided by another implementation of the said aspect, determining the DC prediction value involves determining the average of all reference samples in N adjacent reference lines to the current prediction block, where N is a predetermined integer.
[0016] Optionally, in any of the aspects described above, as provided by another implementation of the aspect, determining the DC prediction value includes determining the mean of all reference samples within the selected reference line and the corresponding reference line.
[0017] Optionally, in any of the aspects described above, as provided by another implementation of the aspect, determining the DC prediction value includes determining the mean of all reference samples within adjacent reference lines and selected reference lines.
[0018] In one embodiment, the present disclosure includes a non-temporary computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-temporary computer-readable medium, the instructions causing the video coding device, when executed by a processor, to perform the following steps: receiving a bitstream via a receiver; receiving an intra-prediction mode from the bitstream by the processor, the intra-prediction mode indicating a relationship between a current block and a selected reference line, the current block being associated with a plurality of reference lines, including the selected reference line; decoding the selected reference line by the processor based on a selected codeword indicating the selected reference line, the selected codeword including a length based on the selection probability of the selected reference line; and presenting video data on a display, including the image block decoded based on the intra-prediction mode and the selected reference line.
[0019] Optionally, in any of the above aspects, as provided by another implementation of the aspect, the multiple reference lines are represented by multiple codewords, where the reference line furthest from the current block is represented by a codeword having the second shortest length.
[0020] Optionally, in any of the above aspects, as provided by another implementation of the aspect, the multiple reference lines are represented by multiple codewords, where the second furthest reference line from the current block is represented by a codeword having the second shortest length.
[0021] Optionally, in any of the above aspects, as provided by another implementation of the aspect, the plurality of reference lines are represented by a plurality of codewords, and predefined reference lines other than adjacent reference lines are represented by codewords of the second shortest length.
[0022] Optionally, in any aspect described above, as provided by another implementation of that aspect, the plurality of reference lines are represented by a plurality of codewords, the plurality of codewords are classified into Class A groups and Class B groups, the Class A groups include codewords having a length shorter than the length of the codewords in the Class B groups.
[0023] Optionally, in any of the above aspects, as provided by another implementation of the said aspect, the plurality of reference lines include reference rows and reference columns, and the number of reference rows stored for the current block is half the number of reference columns stored for the current block.
[0024] Optionally, in any of the above aspects, as provided by another implementation of the said aspect, the plurality of reference lines include reference rows and reference columns, and the number of reference rows stored for the current block is equal to the number of reference columns stored for the current block minus 1.
[0025] Optionally, in any of the above aspects, as provided by another implementation of the aspect, the plurality of reference lines include reference rows, and the number of reference rows stored for the current block is selected based on the number of reference rows used by the deblocking filter operation.
[0026] For clarity, any of the above embodiments can be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present disclosure.
[0027] These and other features will be more clearly understood from the following detailed description, which is to be read in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0028] To better understand the present disclosure, reference is made to the following brief description, which is to be read in conjunction with the accompanying drawings and the detailed description. Here, like reference numerals represent like parts.
[0029] [Figure 1] It is a flowchart of an exemplary method for encoding a video signal.
[0030] [Figure 2] It is a schematic diagram of an exemplary encoding and decoding (codec) system for video encoding.
[0031] [Figure 3] It is a block diagram showing an exemplary video encoder capable of implementing intra prediction.
[0032] [Figure 4] It is a block diagram showing an exemplary video decoder capable of implementing intra prediction.
[0033] [Figure 5] It is a schematic diagram showing an exemplary intra prediction mode used in video encoding.
[0034] [Figure 6] It is a schematic diagram showing an example of the directional relationship of blocks in video encoding.
[0035] [Figure 7]This is a schematic diagram showing an example of a primary reference line scheme for encoding blocks with intra-prediction.
[0036] [Figure 8] This is a schematic diagram illustrating an example of an alternative reference line scheme for encoding blocks with intra-prediction.
[0037] [Figure 9] This is a schematic diagram showing an exemplary subset of intra-predictive modes.
[0038] [Figure 10] This is a schematic diagram illustrating an exemplary conditional signal transmission representation for alternate reference lines in a video encoded bitstream.
[0039] [Figure 11] This is a schematic diagram illustrating an exemplary conditional signal transmission representation for the main reference line in a video encoded bitstream.
[0040] [Figure 12] This is a schematic diagram illustrating an exemplary mechanism for DC-mode intra-prediction using an alternative reference line.
[0041] [Figure 13] This is a schematic diagram illustrating an exemplary mechanism for encoding alternative reference lines with codewords.
[0042] [Figure 14] This is a schematic diagram illustrating an exemplary mechanism for encoding alternative reference lines with different numbers of rows and columns.
[0043] [Figure 15] This is a schematic diagram illustrating another exemplary mechanism for encoding alternative reference lines with different numbers of rows and columns.
[0044] [Figure 16] This is a schematic diagram of an exemplary video encoding device.
[0045] [Figure 17] This is a flowchart illustrating an exemplary method of video coding using an intra-predictive mode subset with an alternative reference line.
[0046] [Figure 18] This is a flowchart illustrating an exemplary method of video coding using DC-mode intra-prediction with an alternate reference line.
[0047] [Figure 19] This is a flowchart illustrating an exemplary method of video coding using reference lines encoded by codewords based on selection probability. [Modes for carrying out the invention]
[0048] Firstly, while exemplary implementations of one or more embodiments are provided below, it should be understood that the disclosed systems and / or methods may be carried out using any number of techniques, whether currently known or existing. This disclosure should not be limited in any way to the exemplary implementations, drawings and techniques shown below, including the exemplary designs and implementations shown and described herein, and may be modified within the entire scope of the accompanying claims and their equivalents.
[0049] Many methods are used tandem to compress video data during the video encoding process. For example, a video sequence is divided into image frames. The image frames are then divided into image blocks. The image blocks can then be compressed by inter-prediction (correlation between blocks in different frames) or intra-prediction (correlation between blocks within the same frame). In intra-prediction, the current image block is predicted from a reference line of samples. The reference line contains samples from adjacent image blocks, also called neighbor blocks. Samples from the current block are matched with samples from the reference line with the nearest lumen (brightness) or chromen (color) value. The current block is encoded as a prediction mode indicating the matching samples. Prediction modes include angular prediction mode, direct current (DC) mode, and planar mode. The difference between the values predicted by these prediction modes and the actual values is encoded as residual values in the residual block. Matching may be improved by using multiple reference lines. Improved matching results in reduced residual values and thus improved compression. However, increasing the number of reference lines can increase the number of bins (binary values) required to uniquely identify a matching reference line. When multiple reference lines are not needed to determine the best match, the increased signal transmission overhead associated with identifying multiple reference lines can generally outweigh the compression gain associated with multiple reference lines, thus increasing the overall bitstream file size. As a result, encoding efficiency decreases in such cases.
[0050] This application discloses mechanisms to support a video coding process that reduce the signal transmission overhead associated with intra-prediction based on multiple reference lines, thereby increasing compression in a video coding system. For example, a mechanism for creating a subset of intra-prediction modes is disclosed. Allowing all intra-prediction modes to have access to the complete set of multiple reference lines can increase signal transmission overhead, which can result in a large file size to be encoded. Therefore, the intra-prediction mode subset includes the subset of intra-prediction modes that have access to alternative reference lines, and modes excluded from the intra-prediction mode subset are limited to those that have access to primary reference lines. As used herein, the primary reference line is the reference line located closest to the current block (e.g., immediately adjacent). Alternative reference lines include both the primary reference line and a set of reference lines located further from the current block than the primary reference line. The encoder may use modes from the intra-prediction mode subset if alternative reference lines are beneficial, and modes excluded from the intra-prediction mode subset if the primary reference lines are sufficient. Since modes outside the intra-prediction mode subset are limited to accessing the primary reference lines, the reference line index may be omitted for intra-prediction modes not included in the intra-prediction mode subset. This may be achieved by encoding the reference line index after the intra-prediction mode information. If intra-prediction mode information indicates that an intra-prediction mode is not included in the intra-prediction mode subset, the decoder can contextually recognize that a reference line index is not included. The intra-prediction modes included in the intra-prediction mode subset may be predetermined (e.g., stored in a table and / or hardcoded) and / or inferred contextually based on the intra-prediction mode subset of an adjacent block, etc. In some cases, the intra-prediction mode subset may be selected to include modes related to the most likely mode (MPM) list. This allows the most commonly selected mode to have access to alternative reference lines. Furthermore, intra-prediction modes in the intra-prediction mode subset can be signaled based on the intra-prediction mode subset index. Since the intra-prediction mode subset contains fewer prediction modes than the complete intra-prediction mode set, the intra-prediction mode subset index can generally be signaled with fewer bins. Furthermore, an extension to DC intra-prediction modes is disclosed. The disclosed DC intra-prediction modes may allow the DC prediction values to be determined based on alternative reference lines. Furthermore, a mechanism for condensing the codeword size for specific reference lines is disclosed. In this mechanism, reference lines are indexed based on selection probability rather than distance from the current image sample. For example, the reference line index most likely to be selected receives the shortest index, and the reference line least likely to be selected receives the longest index. As a result, the encoding of reference line indices is small in most cases. The codewords for reference line indices may be predetermined and used consistently (e.g., stored in a table and / or hardcoded). In addition, certain hardware designs require more storage to store rows of reference lines than columns of reference lines. Therefore, a mechanism is disclosed to support storing fewer reference rows than reference columns when encoding image samples, thus supporting a reduction in storage requirements during encoding.
[0051] Figure 1 is a flowchart of an exemplary method 100 for encoding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal using various mechanisms to reduce the video file size. The smaller file size allows the compressed video file to be sent to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is generally a mirror image of the encoding process to allow the decoder to reconstruct the video signal in a consistent manner.
[0052] In step 101, a video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. In another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may contain both audio and video components. The video component includes a series of image frames that give the impression of visual motion when viewed in sequence. The frames include pixels represented using brightness, which is referred to herein as the lumen component, and color, which is referred to herein as the chroma component. In some examples, the frames may also include depth values to support three-dimensional viewing.
[0053] In step 103, the video is divided into blocks. This division involves subdividing the pixels within each frame into square and / or rectangular blocks for compression. For example, a coding tree can be used to divide the blocks, and then the blocks can be recursively subdivided until a configuration supporting further encoding is achieved. Thus, a block is sometimes referred to as a coding tree unit in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2). For example, the lumens component of a frame may be subdivided until each individual block contains relatively uniform illumination values. Furthermore, the chromens component of a frame may be subdivided until each individual block contains relatively uniform color values. Therefore, the division mechanism varies depending on the content of the video frame.
[0054] In step 105, various compression mechanisms are used to compress the image blocks divided in step 103. For example, interpretation and / or intrapretation may be used. Interpretation is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Therefore, a block that depicts an object in a reference frame does not need to be described repeatedly in subsequent frames. Specifically, an object such as a table may remain in the same position across multiple frames. Therefore, a table is described once, and subsequent frames can refer to the reference frame. A pattern matching mechanism can be used to match objects across multiple frames. Furthermore, moving objects may be represented across multiple frames, for example, due to the movement of an object or the movement of the camera. As a specific example, a video may show a car moving across the screen across multiple frames. Motion vectors can be used to describe such movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in the frame to the coordinates of an object in the reference frame. Thus, interpretation can encode the image blocks in the current frame as a set of motion vectors indicating offsets from the corresponding block in the reference frame.
[0055] Intra-prediction encodes blocks within a common frame. It leverages the fact that lumens and chroma components tend to cluster within a frame. For example, green patches in a tree tend to be adjacent to similar green patches. Intra-prediction uses multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and DC mode. Directional mode indicates that the current block is similar to / identical to samples of adjacent blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the edges of the row. Planar mode effectively shows smooth brightness / color transitions across rows / columns by using a relatively constant slope when changing values. DC mode is used for boundary smoothing, indicating that a block is similar to / identical to the mean value related to samples of all adjacent blocks in the angular direction of the directional prediction mode. Thus, intra-predicted blocks can represent image blocks as various relational prediction mode values instead of actual values. Furthermore, intra-predicted blocks can represent image blocks as motion vector values instead of actual values. In either case, the predicted block may not accurately represent the image block in some cases. All differences are stored in the residual block. Transformations may be applied to the residual block to further compress the file.
[0056] In step 107, various filtering techniques can be applied. In HEVC, filters are applied according to an in-loop filtering scheme. Block-based prediction, as discussed above, can lead to the generation of blocky images in the decoder. Furthermore, block-based prediction schemes may encode blocks and then reconstruct the encoded blocks for later use as reference blocks. In-loop filtering schemes sequentially apply noise suppression filters, deblocking filters, adaptive loop filters, and sample-adaptive offset (SAO) filters to blocks / frames. These filters mitigate such block artifacts so that the encoded file can be accurately reconstructed. In addition, these filters mitigate artifacts in the reconstructed reference blocks, thereby reducing the likelihood that artifacts will cause additional artifacts in subsequent blocks encoded based on the reconstructed reference blocks.
[0057] Once the video signal has been split, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes the data discussed above, plus any signal transmission data desired to support proper video signal reconstruction by the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide encoding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. Bitstream generation is a sequential iterative process. Therefore, steps 101, 103, 105, 107, and 109 can occur sequentially and / or simultaneously across many frames and blocks. The order shown in Figure 1 is presented for clarity and simplicity of discussion and is not intended to restrict the video encoding process to a specific order.
[0058] The decoder receives the bitstream and begins the decoding process in step 111. Specifically, the decoder uses an entropy decoding scheme that converts the bitstream into corresponding syntax and video data. Using the syntax data from the bitstream, the decoder determines the partitions for the frames in step 111. The partitioning should match the result of the block partitioning in step 103. The entropy coding / decoding used in step 111 is described here. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible options based on the spatial positioning of values in the input image (one or more). Signaling a strict choice may involve using a number of bins. As used herein, a bin is a binary value (e.g., a context-dependent bit value) treated as a variable. Entropy coding allows the encoder to discard any option that is obviously impractical in particular, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of acceptable options (for example, one bin for two options, two bins for three or four options, etc.). The encoder then encodes a codeword for the selected options. This method reduces the size of the codeword so that it is only large enough to uniquely represent a selection from a small subset of acceptable options, rather than a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of acceptable options in a similar manner to the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.
[0059] In step 113, the decoder performs block decoding. Specifically, the decoder generates residual blocks using the inverse transform. The decoder then uses the residual blocks and corresponding prediction blocks to reconstruct the image blocks according to the partitioning. The prediction blocks may include both intra-prediction blocks and inter-prediction blocks, as generated by the encoder in step 105. The reconstructed image blocks are then positioned within the frame of the reconstructed video signal according to the partitioning data determined in step 111. The syntax for step 113 may also be transmitted in the bitstream by entropy coding, as discussed above.
[0060] In step 115, filtering is performed on the frames of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters can be applied to the frames to remove blocking artifacts. Once the frames have been filtered, the video signal can be output to a display in step 117 for viewing by the end user.
[0061] Figure 2 is a schematic diagram of an exemplary encoding and decoding (codec) system 200 for video encoding. Specifically, the codec system 200 provides functionality to support the implementation of Method 100. The codec system 200 is generalized to depict components used in both the encoder and the decoder. The codec system 200 receives and partitions the video signal, as discussed with respect to steps 101 and 103 of Method 100, thereby obtaining a partitioned video signal 201. The codec system 200 then compresses the partitioned video signal 201 into an encoded bitstream, as discussed with respect to steps 105, 107, and 109 of Method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as discussed with respect to steps 111, 113, 115, and 117 of Method 100. The codec system 200 includes a general encoder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context adaptive binary arithmetic coding (CABAC) component 231. Such components are combined as shown in the figure. In Figure 2, black lines indicate the movement of data to be encoded / decoded, and dashed lines indicate the movement of control data that controls the operation of other components. In an encoder, all components of the codec system 200 may be present. A decoder may contain a subset of the components of the codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are described below.
[0062] The partitioned video signal 201 is a captured video stream that has been partitioned into blocks of pixels by a coding tree. The coding tree partitions blocks of pixels into smaller blocks of pixels using various partitioning modes. These blocks can then be further subdivided into even smaller blocks. Blocks may also be referred to as nodes of the coding tree. Larger parent nodes are partitioned into smaller child nodes. The number of times a node is subdivided is referred to as the node / coding tree depth. Partitioned blocks may sometimes be referred to as coding units (CUs). Partitioning modes may include binary trees (BT), triple trees (TT), and quad trees (QT), which are used to partition a node into two, three, or four child nodes of different shapes, depending on the partitioning mode used. The divided video signal 201 is then transferred for compression to the general encoder control component 211, the transformation scaling and quantization component 213, the intra-picture estimation component 215, the filter control analysis component 227, and the motion estimation component 221.
[0063] The general encoder control component 211 is configured to make decisions related to encoding the images of a video sequence into a bitstream, according to application constraints. For example, the general encoder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions may be based on memory space / bandwidth availability and image resolution requirements. The general encoder control component 211 also manages buffer utilization in relation to the transmission rate to mitigate buffer underrun and overrun issues. To manage these issues, the general encoder control component 211 manages partitioning, prediction, and filtering by other components. For example, the general encoder control component 211 can dynamically increase the complexity of compression to increase resolution and bandwidth usage, or decrease the complexity of compression to decrease resolution and bandwidth usage. Thus, the general encoder control component 211 controls other components of the codec system 200 to balance bitrate concerns with video signal reconstruction quality. The general encoder control component 211 generates control data that controls the operation of other components. The control data is also transferred to the header formatting and CABAC component 231, where it is encoded in a bitstream to signal parameters for decoding by the decoder.
[0064] The divided video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for interprediction. The frames or slices of the divided video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform interpredictive coding of the received video blocks for one or more blocks within one or more reference frames in order to provide temporal prediction. The codec system 200 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.
[0065] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are illustrated separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process that generates motion vectors, which estimate the motion about a video block. The motion vectors may, for example, represent the displacement of the prediction unit (PU) of a video block relative to a prediction block in a reference frame (or other encoding unit) with respect to the currently encoded block in the current frame (or other encoding unit). A prediction block is a block that is found to closely match the block to be encoded in terms of pixel difference, which may be determined by the sum of absolute difference (SAD), the sum of square difference (SSD), or other difference metrics. In some examples, the codec system 200 can calculate values for pixel positions of the reference picture stored in the decoded picture buffer 223 that are less than an integer. For example, the video codec system 200 may interpolate values for quarter-pixel, eighth-pixel, or other fractional pixel positions of the reference picture. Therefore, the motion estimation component 221 can perform motion search for full pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision. The motion estimation component 221 calculates motion vectors for the PUs of video blocks in the inter-encoded slice by comparing the PU positions with the predicted block positions in the reference image. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header formatting and CABAC component 231 for encoding and outputs the motion to the motion compensation component 219.
[0066] The motion compensation performed by the motion compensation component 219 may involve fetching or generating a predicted block based on the motion vector determined by the motion estimation component 221. Again, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving the motion vector for the PU of the current video block, the motion compensation component 219 may locate the predicted block that the motion vector points to in the reference picture list. The residual video block is then formed by subtracting the pixel values of the predicted block from the pixel values of the encoded current video block to form a pixel difference value. Generally, the motion estimation component 221 performs motion estimation for the lumen component, and the motion compensation component 219 uses the motion vector calculated based on the lumen component for both the chromen and lumen components. The predicted and residual blocks are then transferred to the transformation scaling and quantization component 213.
[0067] The split video signal 201 is also sent to the intra-picture estimation component 215 and the intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated, but are illustrated separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block for the block in the current frame, instead of the inter-prediction performed by the inter-frame motion estimation component 221 and the motion compensation component 219 described above. In particular, the intra-picture estimation component 215 determines the intra-prediction mode to use for encoding the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode from several tested intra-picture prediction modes to encode the current block. The selected intra-prediction mode is then forwarded to the header formatting and CABAC component 231 for encoding.
[0068] For example, the intra-picture estimation component 215 calculates rate-distortion values for various tested intra-prediction modes using rate-distortion analysis and selects the intra-prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block encoded to produce the encoded block, and the bitrate (e.g., number of bits) used to produce the encoded block. The intra-picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks and determines which intra-prediction mode exhibits the best rate-distortion value for that block. In addition, the intra-picture estimation component 215 may be configured to encode depth blocks of a depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).
[0069] When implemented on an encoder, the intra-picture prediction component 217 generates residual blocks from prediction blocks based on a selected intra-prediction mode determined by the intra-picture estimation component 215; when implemented on a decoder, it can read residual blocks from a bitstream. The residual blocks contain the difference in values between the prediction blocks and the original blocks, represented as a matrix. The residual blocks are then transferred to the transformation scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may act on both the lumen and chroma components.
[0070] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms may also be used. The transform can convert the residual information from the pixel value domain to the transform domain, for example, the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which can affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameters. In some examples, the transformation scaling and quantization component 213 may then perform a scan of a matrix containing the quantized transformation coefficients. The quantized transformation coefficients are then transferred to the header formatting and CABAC component 231 and encoded in the bitstream.
[0071] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 reconstructs the residual block of the pixel region by applying inverse scaling, transform, and / or quantization. This is for use as a reference block, for example, which may later become the predicted block of another current block. The motion estimation component 221 and / or motion compensation component 219 can compute the reference block by adding the residual block to the corresponding predicted block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to mitigate artifacts generated during scaling, quantization, and transform. Otherwise, such artifacts could cause inaccurate predictions (and further artifacts) when subsequent blocks are predicted.
[0072] The filter-controlled analysis component 227 and the in-loop filter component 225 apply the filter to the residual block and / or reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 can be combined with the corresponding predictive block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter can then be applied to the reconstructed image block. In some examples, the filter may be applied to the residual block instead. Like the other components in Figure 2, the filter-controlled analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are depicted separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes several parameters to adjust how such a filter is applied. The filter-controlled analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred to the header formatting and CABAC component 231 as filter-controlled data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. These filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., for a reconstructed pixel block) or in the frequency domain, depending on the example.
[0073] When operating as an encoder, filtered and reconstructed image blocks, residual blocks, and / or predicted blocks are stored in the decoded picture buffer 223 for later use in motion estimation as discussed above. When operating as a decoder, the decoded picture buffer 223 stores the reconstructed and filtered blocks and transfers them to the display as part of the output video signal. The decoded picture buffer 223 may be any memory device capable of storing predicted blocks, residual blocks, and / or reconstructed image blocks.
[0074] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into an encoded bitstream for transmission to the decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data such as general control data and filter control data. Furthermore, prediction data, including intra-prediction and motion data, as well as residual data in the form of quantized transformation coefficient data, are all encoded in the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of encoding contexts for various blocks, indications of the most likely intra-prediction mode, and indications of partition information. Such data may be encoded using entropy coding. For example, the information may be encoded using context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. Following entropy coding, the encoded bitstream may be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.
[0075] Figure 3 is a block diagram showing an exemplary video encoder 300 capable of implementing intra-prediction. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of method 100. The video encoder 300 partitions an input video signal, producing a partitioned video signal 301 which is substantially similar to the partitioned video signal 201. The partitioned video signal 301 is then compressed by the components of the encoder 300 and encoded into a bitstream.
[0076] Specifically, the divided video signal 301 is transferred to an intra-picture prediction component 317 for inter-prediction. The intra-picture prediction component 317 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. The divided video signal 301 is also transferred to a motion compensation component 321 for inter-prediction based on a reference block in the decoded picture buffer 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to a transformation and quantization component 313 for transformation and quantization of the residual block. The transformation and quantization component 313 may be substantially the same as the transformation scaling and quantization component 213. The transformed and quantized residual block and the corresponding prediction block (along with the associated control data) are transferred to an entropy coding component 331 for encoding in the bitstream. The entropy coding component 331 may be substantially the same as the header formatting and CABAC component 231.
[0077] The transformed and quantized residual blocks and / or corresponding prediction blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329, for reconfiguration into reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially the same as the scaling and inverse transform component 229. The in-loop filters in the in-loop filter component 325 are also applied to the residual blocks and / or reconfigured reference blocks, depending on the example. The in-loop filter component 325 may be substantially the same as the filter-controlled analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters such as noise suppression filters, deblocking filters, SAO filters, and / or adaptive loop filters. The filtered blocks are then stored in the decoded picture buffer 323 for use by the motion compensation component 321 in the reference blocks. The decoded picture buffer 323 may be substantially the same as the decoded picture buffer 223.
[0078] As will be discussed later, the intra-picture prediction component 317 may perform intra-prediction by selecting an intra-prediction mode that uses alternative reference lines related to adjacent blocks. To reduce signal transmission overhead, the intra-picture prediction component 317 may determine an intra-prediction mode subset that includes a subset of intra-prediction modes that have access to alternative reference lines. Modes excluded from the intra-prediction mode subset have access to primary reference lines. The intra-picture prediction component 317 then has the option of selecting an intra-prediction mode that uses alternative reference lines to obtain better matching, or an intra-prediction mode that uses primary reference lines to support lower signal transmission overhead. As will be discussed later, when using alternative reference lines, the intra-picture prediction component 317 may also use various other mechanisms to support increased coding efficiency.
[0079] Figure 4 is a block diagram showing an exemplary video decoder 400 that can implement intra-prediction. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of method 100. The decoder 400 receives a bitstream, for example, from the encoder 300 and generates a reconstructed output video signal based on the bitstream for display to the end user.
[0080] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 performs the inverse function of the entropy coding component 331. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 can use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transformation coefficients from residual blocks. The quantized transformation coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be substantially the same as the inverse transform and quantization component 329.
[0081] The reconstructed residual blocks and / or predicted blocks are transferred to the intra-picture prediction component 417 for reconstruction into image blocks based on intra-predictive operation. The intra-picture prediction component 417 may be substantially similar to the intra-picture prediction component 317, but in reverse operation. Specifically, the intra-picture prediction component 417 uses prediction mode to locate reference blocks in the frame and applies the residual blocks to the result to reconstruct the intra-predicted image blocks. The reconstructed intra-predicted image blocks and / or residual blocks and the corresponding intra-predictive data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425. These may be substantially similar to the in-loop filter component 325 and the decoded picture buffer component 323, respectively. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks, and / or predicted blocks, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are transferred to the motion compensation component 421 for interpretation. The motion compensation component 421 may be substantially the same as, or inversely, the motion compensation component 321. Specifically, the motion compensation component 421 generates a prediction block using the motion vector from the reference block, and then applies the residual block to the result to reconstruct the image block. The resulting reconstructed block may be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks. These can be reconstructed into frames via partition information. Such frames may be arranged in a sequence. This sequence is output to the display as a reconstructed output video signal.
[0082] Similar to the intra-picture prediction component 317, the intra-picture prediction component 417 may perform intra-prediction based on an intra-prediction mode using alternate reference lines. Specifically, the intra-picture prediction component 417 recognizes the modes assigned to an intra-prediction mode subset. For example, the intra-prediction mode subset can correspond to a determined MPM list, can be predefined in memory, and / or can be determined based on the intra-prediction mode of an adjacent block. Thus, if the intra-prediction mode for the current block is in the intra-prediction mode subset, the intra-picture prediction component 417 can obtain the reference line index from the bitstream. Otherwise, the intra-picture prediction component 417 can infer that the primary reference line is intended by the encoder. The intra-picture prediction component 417 may use various other mechanisms to support increased coding efficiency when using alternate reference lines, as will be discussed later.
[0083] Figure 5 is a schematic diagram showing an exemplary intra-prediction mode 500 used in video coding. For example, the intra-prediction mode 500 may be used by steps 105 and 113 of Method 100, the intra-picture estimation component 215 and intra-picture prediction component 217 of the codec system 200, the intra-picture prediction component 317 of the encoder 300, and / or the intra-picture prediction component 417 of the decoder 400. Specifically, the intra-prediction mode 500 can be used to compress an image block into a prediction block containing the selected prediction mode and the remaining residual block.
[0084] As described above, intra-prediction involves matching the current image block to the corresponding sample(s) of one or more adjacent blocks. The current image block can then be represented as a selected prediction mode index and residual block, which is far smaller than all the lumens / chroma values contained in the current image block. Intra-prediction can be used when no reference frame is available or when inter-predictive coding is not used for the current block or frame. Reference samples for intra-prediction may be derived from previously coded (or reconstructed) adjacent blocks within the same frame. Advanced Video Coding (AVC), also known as H.264 and H.265 / HEVC, uses reference lines of boundary samples from adjacent blocks as reference samples for intra-prediction. Many different intra-prediction modes are used to cover different textures or structural characteristics. H.265 / HEVC supports a total of 35 different intra-prediction modes, 500 of which spatially correlate the current block to one or more reference samples. Specifically, the intra-prediction mode 500 includes 33 directional prediction modes indexed as modes 2 to 34, a DC mode indexed as mode 1, and a planar mode indexed as mode 0.
[0085] During encoding, the encoder matches the lumen / chroma values of the current block with the lumen / chroma values of the corresponding reference samples in reference lines that pass through the edges of adjacent blocks. When the best match with one of the reference lines is found, the encoder selects one of the directional intra-prediction modes 500 that points to the best-matching reference line. For clarity of discussion, acronyms are used below to refer to specific directional intra-prediction modes 500. DirS represents the starting directional intra-prediction mode (e.g., mode 2 in HEVC) when counting clockwise from the bottom left. DirE represents the ending directional intra-prediction mode (e.g., mode 34 in HEVC) when counting clockwise from the bottom left. DirD represents the intermediate neutral intra-encoding mode (e.g., mode 18 in HEVC) when counting clockwise from the bottom left. DirH represents the horizontal directional intra-prediction mode (e.g., mode 10 in HEVC). DirV represents the vertical intra-prediction mode (e.g., mode 26 in HEVC).
[0086] As discussed above, the DC mode acts as a smoothing function, deriving the predicted value of the current block as the average of all reference samples within the reference line spanning adjacent blocks. Also, as discussed above, the planar mode returns predicted values that show a smooth transition (e.g., a constant gradient of values) between samples below and to the upper left or upper left and to the upper right of the reference line of the reference sample.
[0087] For prediction modes in the plane, DC, and DirH to DirV, samples in both the row above the reference line and the column to the left of the reference line are used as reference samples. For prediction modes with prediction directions from DirS to DirH (including DirS and DirH), reference samples from previously encoded and reconstructed adjacent blocks in the column to the left of the reference line are used as reference samples. For prediction modes with prediction directions from DirV to DirE (including DirV and DirE), reference samples from previously encoded and reconstructed adjacent blocks in the row above the reference line are used as reference samples.
[0088] While there are many intra-prediction modes 500, not all intra-prediction modes 500 are selected with equal probability during video coding. Furthermore, the intra-prediction mode 500 selected by an adjacent block is statistically highly correlated with the intra-prediction mode 500 selected for the current block. Therefore, in some examples, an MPM list may be used. The MPM list is a list containing a subset of the intra-prediction modes 500 that are most likely to be selected. If the intra-prediction mode for the current block is included in the MPM list, the selected mode can be signaled in the bitstream by an MPM list index, which can use a codeword with fewer bins than the number of bins used to uniquely identify all intra-prediction modes 500.
[0089] An MPM list may be constructed with intra-prediction modes 500 of several neighboring decoded blocks and several default intra-prediction modes that generally have a high selection probability. For example, in H.265 / HEVC, an MPM list of length 3 is constructed with intra-prediction modes of two adjacent blocks (one above and one to the left of the current block). In the case of overlapping modes, the MPM list is assigned by default as planar mode, DC mode, or DirV mode, in that order. If the MPM list has a longer length, it may also include intra-prediction modes 500 of further adjacent blocks and / or further default intra-prediction modes 500. The length and construction method of the MPM list may be predefined.
[0090] Figure 6 is a schematic diagram showing an example of the directional relationship of blocks 600 in video coding. For example, blocks 600 may be used when selecting intra-prediction modes 500. Thus, blocks 600 may be used by steps 105 and 113 of method 100, the intra-picture estimation component 215 and intra-picture prediction component 217 of codec system 200, the intra-picture prediction component 317 of encoder 300, and / or the intra-picture prediction component 417 of decoder 400. In video coding, blocks 600 are divided based on video content and can therefore include many rectangles and squares of various shapes and sizes. Blocks 600 are depicted as squares for illustrative purposes and are therefore simplified from actual video coding blocks to support clarity of discussion.
[0091] Block 600 currently contains block 601 and adjacent block 610. Current block 610 is any block being encoded at a given time. Adjacent block 610 is any block immediately adjacent to the left or top edge of current block 601. Video encoding generally proceeds from the top left to the bottom right. Therefore, adjacent block 610 may be encoded and reconstructed prior to the encoding of current block 601. When encoding current block 601, the encoder matches the luma / chroma values of current block 601 with reference samples (one or more) from a reference line following the edge of adjacent block 610. The match is then used to select an intra-prediction mode, for example, from intra-prediction mode 500, because it points to the matched sample (or, if DC or planar mode is selected, the samples). The selected intra-prediction mode indicates that the luma / chroma values of current block 601 are substantially similar to the reference samples corresponding to the selected intra-prediction mode. Any difference can be retained in the residual block. Next, the selected intra-prediction mode is encoded in the bitstream. In the decoder, block 601 can now be reconstructed using the lumen / chroma values of the reference samples in the selected reference line within the adjacent block 610 corresponding to the selected intra-prediction mode (along with any residual information from the residual block).
[0092] Figure 7 is a schematic diagram showing an example of a primary reference line scheme 700 for encoding blocks with intra-prediction. The primary reference line scheme 700 may be used when selecting the intra-prediction mode 500. Thus, the primary reference line scheme 700 may be used by steps 105 and 113 of Method 100, the intra-picture estimation component 215 and the intra-picture prediction component 217 of the codec system 200, the intra-picture prediction component 317 of the encoder 300, and / or the intra-picture prediction component 417 of the decoder 400.
[0093] The primary reference line scheme 700 uses the primary reference line 711 of reference sample 712. The primary reference line 711 includes an edge sample (e.g., a pixel) of an adjacent block as reference sample 712 for the current block 701. As used herein, reference sample 712 is a value such as the chroma or luma value of a pixel or a sub-part of it. The current block 701 is substantially the same as the current block 601. For the purposes of discussion, the primary reference line 711 includes a reference row 713 containing reference sample 712 above the current block 701. The primary reference line 711 also includes a reference column 714 containing reference sample 712 to the left of the current block 701. In the primary reference line scheme 700, the primary reference line 711 is used, and the primary reference line 711 is the reference line immediately adjacent to the current block 701. Thus, the current block 701 is matched as close as possible to the reference sample 712 contained in the primary reference line 711 during intra-prediction.
[0094] Figure 8 is a schematic diagram showing an example of an alternative reference line scheme 800 for encoding blocks with intra-prediction. The alternative reference line scheme 800 may be used when selecting the intra-prediction mode 500. Thus, the alternative reference line scheme 800 may be used by steps 105 and 113 of Method 100, the intra-picture estimation component 215 and the intra-picture prediction component 217 of the codec system 200, the intra-picture prediction component 317 of the encoder 300, and / or the intra-picture prediction component 417 of the decoder 400.
[0095] Alternative reference line scheme 800 uses current block 801, which is substantially the same as current block 701. Multiple reference lines 811 extend from current block 801. Reference line 811 contains reference sample 812 similar to reference sample 712. Reference line 811 is substantially the same as the primary reference line 711, but extends further away from current block 801. By using alternative reference lines 811, the matching algorithm can have access to more reference samples 812. The presence of more reference samples may, in some cases, result in a better match for current block 801, which in turn may lead to fewer residual samples after the prediction mode is selected. The alternative reference line scheme 800 may also be called multiple lines intra prediction (MLIP) and is discussed in detail in Joint Video Experts Team (JVET) documents JVET-C0043, JVET-C0071, JVET-D0099, JVET-D0131, and JVET-D0149. The reference lines 811 may be numbered from 0 to M, where M is an arbitrary predetermined constant value. The reference lines 811 may include reference row 813 and reference column 814 as shown in the figure, which are similar to reference row 713 and reference column 714, respectively.
[0096] During encoding, the encoder can select the best match from the reference lines 811 based on RDO. Specifically, M+1 reference lines are used from the nearest reference line (RefLine0) to the furthest reference line (RefLineM), where M is greater than zero. The encoder selects the reference line with the optimal rate-distortion cost. The index of the selected reference line is signaled to the decoder in the bitstream. Here, RefLine0 may be called the original reference line (e.g., primary reference line 711), and RefLine1 to RefLineM may be called additional reference lines or alternative reference lines. Additional reference lines may be selected for use by any mode of the intra-prediction mode 500. In some cases, these additional reference lines may be selected for use by a subset of the intra-prediction mode 500 (e.g., DirS to DirE may use additional reference lines).
[0097] Once a reference line 812 is selected, the index of the selected reference line from among the reference lines 811 is signaled to the decoder in the bitstream. By using various reference lines, a more accurate prediction signal can be derived, which in some cases can increase coding efficiency by reducing residual samples. However, using an alternate reference line 811 increases the signal transmission overhead because the selected reference line is signaled to the decoder. As the number of reference lines 811 used increases, more bins are used during signal transmission to uniquely identify the selected reference line. Thus, using an alternate reference line 811 can actually decrease coding efficiency when the matching sample is currently in a reference line immediately adjacent to block 801.
[0098] Despite the compression advantages supported by MLIP, several areas can be improved to achieve higher coding gain. For example, an intra-prediction mode subset may be employed to achieve increased compression. Specifically, some intra-prediction modes may be selected to use alternate reference line 811 to support increased matching accuracy. Other intra-prediction modes may be selected to use primary reference line 711 and eliminate the signal transfer overhead associated with using alternate reference line 811. Intra-prediction modes using alternate reference line 811 may be included in the intra-prediction mode subset.
[0099] Figure 9 is a schematic diagram 900 showing an exemplary intra-predictive mode subset 930 for use in video coding of intra-predictive modes in an encoder or decoder. The alternate reference line scheme 800 and the primary reference line scheme 700 can be combined / modified to use both the intra-predictive mode list 920 and the intra-predictive mode subset 930. The intra-predictive mode list 920 contains all intra-predictive modes 923 (e.g., intra-predictive mode 500) that can be used in an intra-predictive scheme. Each such intra-predictive mode 923 may be indexed by a corresponding intra-predictive mode index 921. In this example, some of the intra-predictive modes have access to alternate reference lines. Intra-predictive modes with access to alternate reference lines are stored in the intra-predictive mode subset 930 as alternate reference line predictive modes 933. The alternate reference line predictive modes 933 in the intra-predictive mode subset 930 may be indexed by an alternate intra-predictive mode index 931. The alternate intra-prediction mode index 931 is an index value used to number and indicate the alternate reference line prediction mode 933 in the intra-prediction mode subset 930. Since fewer intra-prediction modes are included in the intra-prediction mode subset 930, the alternate intra-prediction mode index 931 can contain fewer bins than the intra-prediction mode index 921. Therefore, if an intra-prediction mode is included in the intra-prediction mode subset 930, that intra-prediction mode can be matched to an alternate reference line. If an intra-prediction mode is not included in the intra-prediction mode subset 930, that intra-prediction mode can be matched to a primary reference line. This method takes advantage of the fact that the intra-prediction mode list 920 includes many intra-prediction modes 923 to cover texture and structural characteristics in detail. However, the likelihood of using alternate reference lines is relatively low. Therefore, in the case of alternate reference lines, a rough direction is sufficient.Therefore, the intra-prediction modes 923 can be sampled in order to construct an intra-prediction mode subset 930 using a method that will be discussed later.
[0100] The use of both the intra-prediction mode list 920 and the intra-prediction mode subset 930 allows for a significant increase in coding efficiency. Furthermore, the selection of which intra-prediction modes should be included as alternate reference line prediction modes 933 affects coding efficiency. For example, larger mode ranges use more bits to represent alternate intra-prediction mode indices 931, while smaller mode ranges use fewer bits. In one embodiment, the intra-prediction mode subset 930 includes every second intra-prediction mode in [DirS, DirE], where [A, B] represents a set containing integer elements x, and B ≥ x ≥ A. Specifically, the intra-prediction mode subset 930 may be associated with prediction modes {DirS, DirS+2, DirS+4, DirS+6, …, DirE}, where {A, B, C, D} represents a set containing all elements listed between the curly braces. Furthermore, the number of bits representing the selected intra-prediction mode within the mode range of the alternate intra-prediction mode index 931 is reduced. In such cases, the alternate intra-prediction mode index 931 can be derived by dividing the intra-prediction mode index 921 by 2.
[0101] In another embodiment, the intra-prediction mode subset 930 includes intra-prediction modes in the MPM list 935. As described above, the MPM list 935 includes a subset of the intra-prediction modes that are most likely to be selected. In this case, the intra-prediction mode subset 930 may be configured to include intra-prediction modes of the currently decoded and / or reconstructed adjacent blocks of the block. A flag may be used in the bitstream to indicate whether the selected intra-prediction mode is included in the MPM list 935. If the selected intra-prediction mode is in the MPM list 935, the MPM list 935 index is signaled. Otherwise, an alternate intra-prediction mode index 931 is signaled. If the selected intra-prediction mode is a primary reference line mode and is not included in the MPM list 935, the intra-prediction mode index 921 may be signaled. In some examples, a binary representation mechanism can be used in which the alternate intra-prediction mode index 931 and / or the MPM list 935 index are of fixed length. In some examples, the alternative intra-prediction mode index 931 and / or MPM list 935 index are encoded via context-based adaptive binary arithmetic coding (CABAC). Other binary representations and / or entropy coding mechanisms may be used. The binary representation / entropy coding mechanism used is predefined. The mode range of the intra-prediction mode subset 930 includes fewer intra-prediction modes (e.g., coarse modes) than the intra-prediction mode list 920. Therefore, if an intra-prediction mode in MPM list 935 (e.g., a mode in an adjacent block) is not included in the mode range of the intra-prediction mode subset 930, the intra-prediction mode index 921 can be rounded to the nearest alternative intra-prediction mode index 931 value by dividing it (e.g., by 2) and adding or subtracting 1, for example. The rounding mechanism may be predefined.The relative positions of adjacent blocks, the scan order, and / or the size of the MPM list 935 may also be predefined.
[0102] As a specific example, an encoder operating according to H.265 may use a mode range of size 33 [DirS,DirE]. If every second mode is used in the intra-prediction mode subset 930, the size of the mode range for the alternate intra-prediction mode index 931 is 17. The size of the MPM list 935 may be 1, and the intra-prediction modes of the decoded block above are used to construct the MPM list 935. If the mode of the adjacent block is not in the mode range of the alternate intra-prediction mode index 931 (for example, mode 3), the mode of the adjacent block is rounded to mode 4 (or 2). As a specific example, if the selected intra-prediction mode is mode 16 according to the intra-prediction mode index 921, and the selected mode is not in the MPM list 935, the selected intra-prediction mode will have an alternate intra-prediction mode index 931 of 8 (16 / 2=8). When a fixed-length binary representation is used, in this example, 1000 can be used to indicate the selected intra-prediction mode.
[0103] The above example / embodiment assumes that the intra-prediction mode subset 930 contains each odd (or even) intra-prediction mode 923. However, other mechanisms may be used to populate the intra-prediction mode subset 930 while using the mechanism described above. In one example, the intra-prediction mode subset 930 contains every N intra-prediction mode in [DirS, DirE], where N is an integer equal to 0, 1, 2, 3, 4, etc. This example can also be written as {DirS, DirS+N, DirS+2N, ..., DirE}. In another example, the intra-prediction mode subset 930 may contain every N intra-prediction mode in [DirS, DirE], as well as planar and DC intra-prediction modes.
[0104] In another example, the intra-predictive mode subset 930 includes intra-predictive modes with high general selection probabilities. Thus, the intra-predictive mode signal transmission cost is reduced, and coding efficiency is improved. In some examples, intra-predictive modes with a principal direction are generally selected with a higher general probability (e.g., perfectly vertical, horizontal, etc.). Thus, the intra-predictive mode subset 930 may include such modes. For example, the intra-predictive mode subset 930 may include principal directional intra-predictive modes such as DirS, DirE, DirD, DirV, and DirH. Neighbor modes to principal modes may also be included in some examples. Specifically, the intra-predictive modes DirS, DirE, DirD, DirV, and DirH are included in the intra-predictive mode subset 930 along with neighboring intra-predictive modes with positive or negative N indices. This may be expressed as [DirS,DirS+N], [DirE-N,DirE], [DirD-N,DirD+N], [DirH-N,DirH+N], [DirV-N,DirV+N]. In another example, the intra-prediction mode subset 930 includes the intra-prediction modes DirS, DirE, DirD, DirV, DirH, as well as adjacent intra-prediction modes and DC and planar modes with positive or negative N indices.
[0105] In another example, the intra-prediction mode subset 930 is adaptively constructed and therefore not defined using predefined intra-prediction modes. For example, the intra-prediction mode subset 930 may include the intra-prediction modes of neighboring decoded blocks (e.g., adjacent block 610) of the current block (e.g., current block 601). Adjacent blocks may be located to the left and above the current block. Additional adjacent blocks may also be used. The size and configuration of the intra-prediction mode subset 930 are predefined. In another example, the intra-prediction mode subset 930 includes the intra-prediction modes of adjacent blocks and some predetermined default intra-prediction modes, such as DC and planar modes. In yet another example, the intra-prediction mode subset 930 includes intra-prediction modes within the MPM list 935.
[0106] Upon receiving a reference line index in the bitstream, the decoder determines that the intra-prediction mode list 920 is implied if the reference line index points to a primary reference line. In such a case, the decoder can read the subsequent intra-prediction mode index 921 to determine the corresponding intra-prediction mode 923. If the reference line index points to an additional reference line, the intra-mode subset 930 is implied. In such a case, the decoder can read the subsequent alternate intra-prediction mode index 931 to determine the corresponding alternate reference line prediction mode 933. In some examples, a flag may be used to indicate when the indicated alternate reference line prediction mode 933 is in the MPM list 935. In such cases, the alternate reference line prediction mode 933 may be signaled according to the index used by the MPM list 935.
[0107] As described above, plot 900 can be used when encoding intra-prediction modes with alternate or primary reference lines. Below, we discuss a signaling scheme for encoding such data. Specifically, the reference line index can be signaled after the index of the corresponding intra-prediction mode. Furthermore, the reference line index can be signaled conditionally. For example, the reference line index can be signaled if the intra-prediction mode is related to an alternate reference line, and can be omitted if the intra-prediction mode is related to a primary reference line. This allows for the omission of the reference line index whenever the intra-prediction mode is not included in the intra-prediction mode subset 930, because it is not necessary to show the reference line index in the case of a primary reference line. This technique significantly improves encoding efficiency by reducing signaling overhead. The conditional signaling scheme discussed below is an exemplary implementation of this concept.
[0108] Figure 10 is a schematic diagram illustrating an exemplary conditional signaling representation 1000 for alternate reference lines in a video encoded bitstream. For example, the conditional signaling representation 1000 may be used in a bitstream when an intra-prediction mode subset 930 is used as part of an alternate reference line scheme 800 during video encoding that uses an intra-prediction mode, such as an intra-prediction mode 500 in an encoder or decoder. Specifically, the conditional signaling representation 1000 is used when the encoding device encodes or decodes the bitstream with intra-prediction data containing selected intra-prediction modes associated with alternate reference lines. Alternatively, when the encoding device encodes or decodes the bitstream with intra-prediction data containing selected intra-prediction modes associated with primary reference lines, the conditional signaling representation 1100 is used instead, as will be discussed later.
[0109] The conditional signaling representation 1000 may include fields of an encoding unit 1041 containing relevant partitioning information. Such partitioning information indicates block boundaries to the decoder, allowing the decoder to fill the decoded image blocks to produce a frame. The fields of the encoding unit 1041 are included for context and may or may not be adjacent to other fields discussed in relation to the conditional signaling representation 1000. The conditional signaling representation 1000 further includes a flag 1042. The flag 1042 indicates to the decoder that the subsequent information is for an intra-prediction mode related to an alternate reference line. For example, the flag 1042 may indicate whether the subsequent information is encoded as an alternate intra-prediction mode index 931 or as an MPM list index 935. Following the flag 1042, the intra-prediction mode subset index 1043 field is used, for example, to encode the alternate intra-prediction mode index 931. In some examples, the intra-prediction mode subset index 1043 field is used as a substitute for the MPM list index field to hold the MPM list index 935, as discussed above. In any case, the decoder can decode the selected intra-prediction mode for the associated coding unit 1041 based on the flag 1042 and the index. Since the intra-prediction mode subset index 1043 field indicates that the selected intra-prediction mode relates to an alternate reference line, a reference line index 1044 is also included to indicate to the decoder which reference line contains the matching sample. Thus, the intra-prediction mode subset index 1043 gives the direction of the matching sample in the adjacent block, and the reference line index 1044 gives the distance to the matching sample. Based on this information, the decoder can determine the matching sample, generate a prediction block using the matching sample, and optionally reconstruct a block of pixels by combining the prediction block with residual blocks encoded elsewhere in the bitstream.
[0110] Figure 11 is a schematic diagram showing an exemplary conditional signaling representation 1100 for a primary reference line in a video encoded bitstream. For example, the conditional signaling representation 1100 may be used in a bitstream when an intra-predictive mode list 920 is used as part of a primary reference line scheme 700 during video encoding that uses an intra-predictive mode, such as an intra-predictive mode 500 in an encoder or decoder. The conditional signaling representation 1100 may include fields of encoding unit 1141 having partition information in a similar manner to the fields of encoding unit 1041. The conditional signaling representation 1100 also includes a flag 1142 substantially similar to flag 1042, in which case the flag indicates that the selected intra-predictive mode is not associated with an intra-predictive mode subset 930. Based on the information of flag 1142, an intra-predictive mode index 1145 can then be interpreted as an index of mode ranges for an intra-predictive mode index 921. Furthermore, since the selected intra-prediction mode is determined to be associated with a major reference line, the reference line index is omitted.
[0111] Therefore, conditional signaling representations 1000 and 1100 can be used to signal either an intra-prediction mode using an alternate reference line or an intra-prediction mode using a primary reference line, respectively. These schemes can be used in conjunction with any of the examples / embodiments discussed above. In one example, an intra-prediction mode in [DirS,DirE] can use an alternate reference line. In such a case, if the selected intra-prediction mode is not included in [DirS,DirE] (e.g., DC or planar mode), it is not necessary to signal the reference line index. Thus, the reference index is presumed to be equal to 0, where 0 indicates the primary reference line immediately adjacent to the current block. In such a case, the intra-prediction mode subset 930 may not be used, as all directional references use alternate reference lines in such cases. In summary, whenever the intra-prediction mode is not DC or planar mode, the reference line index is signaled in such an example. Furthermore, in such cases, whenever the intra-prediction mode is DC or planar mode, the reference line index is not signaled. In another example, whenever an intra-prediction mode is in the intra-prediction mode subset 930, the reference line index is signaled. Therefore, whenever an intra-prediction mode is not included in the intra-prediction mode subset 930, the reference line index is not signaled.
[0112] Table 1 below is an exemplary syntax table describing when the reference index is omitted when the selected mode is DC or planar mode. [Table 1] As shown in Table 1 above, when the intra-prediction mode is not planar and not DC, the reference line index is signaled along with the x and y positions of the current block in this example.
[0113] Table 2 below is an exemplary syntax table describing cases where a reference index is signaled if the current intra-prediction mode is included in the intra-prediction mode subset, and omitted otherwise. The intra-prediction mode subset can be determined according to one of the mechanisms discussed with respect to Figure 9. [Table 2] As shown in Table 2 above, if the intra-prediction mode is in the intra-prediction subset, the reference line index is signaled.
[0114] Further modifications can be made to the MLIP scheme, either independently or in combination with the examples / embodiments discussed above. For example, in many MLIP schemes, the DC intra-prediction mode is limited to a primary reference line. In such cases, the DC intra-prediction mode generates a predicted value that is the average of all reference samples in the reference line. This results in a smoothing effect between blocks. As will be discussed later, the DC intra-prediction mode can be extended in an MLIP scheme to generate a predicted value that is the average of all reference samples in multiple reference lines.
[0115] Figure 12 is a schematic diagram showing an exemplary mechanism for DC-mode intra-prediction 1200 using an alternative reference line 1211. DC-mode intra-prediction 1200 can be used in an encoder or decoder, such as encoder 300 or decoder 400, when performing DC intra-prediction according to intra-prediction mode 500, for example. The DC-mode intra-prediction 1200 scheme can be used in conjunction with the intra-prediction mode subset 930, conditional signal transmission representations 1000 and 1100, or can be used independently of such examples.
[0116] The DC mode intra-prediction 1200 uses the current block with alternative reference lines 1211 labeled 0-M. Any number of reference lines 1211 can be used as such. Such reference lines 1211 are substantially similar to reference line 811 and include reference sample 1212 which is substantially similar to reference sample 812. Then the DC prediction value 1201 is calculated as the average of reference samples 1212 in multiple reference lines 1211. In one example, the DC prediction value 1201 is calculated as the average of all reference samples 1212 in all reference lines 1211 0-M, regardless of which reference line 1211 is selected during intra-prediction mode selection. This provides a very robust DC prediction value 1201 for the current block. For example, if there are four reference lines 1211, the average of all reference samples 1212 in reference lines 1211 indexed as 0 to 3 is determined and set as the DC prediction value 1201 for the current block. In such cases, the reference line index may not, in some cases, be transmitted in the bitstream.
[0117] In another example, the DC prediction value 1201 is determined based on the selected reference line and its corresponding reference line. This allows the DC prediction value 1201 to be determined based on two correlated reference lines 1211, rather than just all or index 0 reference lines. For example, if four reference lines 1211 are used, the reference lines 1211 may be indexed from 0 to 4. In such a case, Table 3 below shows examples of the corresponding reference lines 1211 used when the reference lines are selected. [Table 3] As shown in the figure, the DC predicted value 1201 can be determined using correlated reference lines 1211. In such a case, the decoder can determine both reference lines 1211 when the selected reference line indices are transmitted in the bitstream.
[0118] In another example, the DC mode intra prediction 1200 can always use the reference line at index 0 (the primary reference line), but may also select an additional reference line 1211. This is when such a selection more accurately matches the DC prediction value 1201 to the predicted pixel block value. Table 4 below shows such an example. In such a case, the decoder may determine both reference lines 1211 once the selected reference line indices are signaled in the bitstream. [Table 4]
[0119] Further modifications can be made to the MLIP scheme, either independently or in combination with the examples / embodiments discussed above. For example, in many MLIP schemes, reference lines are indexed based on their distance from the current block. For instance, in JVET-C0071, codewords 0, 10, 110, and 111 are used to represent RefLine0, RefLine1, RefLine2, and RefLine3, respectively. This technique can be improved by assigning shorter codewords to reference lines with a higher statistical probability of being selected. Since many reference lines are signaled in the bitstream, the average codeword length for reference line signaling decreases when using this mechanism. As a result, coding efficiency improves. Exemplary schemes for alternative reference line signaling using codewords are discussed below.
[0120] Figure 13 is a schematic diagram showing an exemplary mechanism 1300 for encoding alternative reference lines using codewords. Mechanism 1300 can be used in an encoder or decoder, such as encoder 300 or decoder 400, when performing intra-prediction according to intra-prediction mode 500, for example. Mechanism 1300 can be used in conjunction with the intra-prediction mode subset 930, conditional signal transmission representations 1000 and 1100, DC mode intra-prediction 1200, or can be used independently of such examples.
[0121] The mechanism 1300 uses a current block 1301 encoded by a corresponding reference line 1311 containing an intra-prediction and a reference sample. The current block 1301 and the reference line 1311 may be substantially the same as the current block 801 and the reference line 811, respectively. The current block 1301 is matched with one or more reference samples in the reference line 1311. Based on the match, the intra-prediction mode and the reference line 1311 are selected to indicate the matching reference sample. The selected intra-prediction mode is encoded in the bitstream as discussed above. Furthermore, the index of the selected reference line can be encoded in the bitstream by using a reference line codeword 1315. The reference line codeword 1315 is a binary value indicating the index of the reference line. Furthermore, the reference line codeword 1315 omits leading zeros to further compress the data. Thus, the reference line 1311 can, in some cases, be signaled using a single binary value.
[0122] In mechanism 1300, the length of the codeword 1315 for the reference line 1311 index does not have to be related to the distance from the reference line 1311 to the current block 1301. Instead, the codeword 1315 is assigned to the reference line 1311 based on the length of the codeword 1315 and the probability that the corresponding reference line 1311 is selected as the reference line 1311 to be used for the current block 1301. In some cases, a reference line 1311 that is further away from the current block 1301 may be given a shorter codeword than a reference line 1311 that is immediately adjacent to the current block 1301. For example, the codewords 0, 10, 110, and 111 of codeword 1315 can be used to represent RefLine0, RefLine3, RefLine1, and RefLine2, respectively. In addition, the first or first two bins of codeword 1315 may be context-coded, and the last bin(s) may be bypass-coded without context. The context may be derived from spatially adjacent blocks (e.g., above and to the left). Table 5 shows an exemplary encoding scheme for codeword 1315 when four reference lines 1311 are used. [Table 5] As shown in Table 5, reference lines 1311 may be indexed as 0 to 3. In one exemplary MLIP scheme, such indices are represented by codewords 1315 based on their distance from the current block 1301, with the smallest codeword assigned to the closest reference line 1311 and the largest codeword 1315 assigned to the furthest reference line 1311. In exemplary mechanisms 1 and 2, the most likely reference line 1311 is index 0, which receives the shortest codeword 1315. The second most likely reference line 1311 is the furthest reference line (index 3), which receives the second shortest codeword 1315. The remaining codewords 1315 are assigned to indices 1 and 2 in descending or ascending order, depending on the example.
[0123] In another example, the reference line 1311 at index 0 may remain the most likely reference line 1311. Furthermore, the second most likely reference line 1311 may be the reference line 1311 at index 2. The result is shown in Table 6. [Table 6] As shown in Table 6, the most likely reference line 1311 receives index 0, which receives the shortest codeword 1315. The second most likely reference line 1311 is the second furthest reference line (index 2), which receives the second shortest codeword 1315. The remaining codewords 1315 are assigned to indices 1 and 3 in descending or ascending order, depending on the example.
[0124] The above example can be extended when more reference lines are used. For example, when five reference lines are used, the example in Table 5 may be extended as shown in Table 7. [Table 7] As shown in Table 7, the most likely reference line 1311 is at index 0, which receives the shortest codeword 1315. The second most likely reference line 1311 is the one associated with index 2, which receives the second shortest codeword 1315. The remaining codewords 1315 are assigned to indices 1, 2, 4, and 5 in descending or ascending order, depending on the example.
[0125] In another example, the index of reference line 1311 can be sorted into classes, and each class can be assigned to a different group of codewords 1315, which may be arranged in various ways. For example, class A group may receive the shortest codeword 1315, and class B group may receive the longest codeword 1315. Class A codewords 1315 and class B codewords 1315 may be assigned in ascending or descending order, or any combination thereof. The class construction scheme may be predefined. To clarify this scheme, the following example is given. The following example uses six reference lines 1311 indexed from 0 to 5, as shown in Table 8 below. Reference lines 1311 can be assigned to any class as desired. For example, reference lines 1311 indexed as 1, 3, and 5 may be assigned to class A, and reference lines 1311 indexed as 2 and 4 may be assigned to class B. The reference line 1311, indexed as 0, may always be the choice with the highest probability, and therefore may always have the shortest codeword 1315. Class A may be assigned the shortest codeword 1315 other than the codeword 1315 for index 0. Then, Class B may be assigned the longest codeword 1315. The results of such a scheme are shown in Table 8. [Table 8] In the first example in Table 8, the codewords 1315 for class A and class B are both incremented (descending) independently. In the second example in Table 8, the class A codewords 1315 are incremented (descending), while the class B codewords 1315 are decremented (ascending). In the third example in Table 8, the class A codewords 1315 are decremented (ascending), while the class B codewords 1315 are incremented (descending). In the fourth example in Table 9, the codewords 1315 for class A and class B are both decremented (ascending) independently. Any number of reference lines and any number of classes may be used in this scheme to generate codewords 1315 as desired.
[0126] Further modifications can be made to the MLIP scheme, either independently or in combination with the examples / embodiments discussed above. For example, the MLIP scheme uses alternative reference lines that have both rows and columns. Many MLIP schemes use the same number of rows and columns. However, when the MLIP scheme is implemented in hardware, the entire set of rows is stored in memory, for example, in a line buffer on the central processing unit (CPU) chip, during encoding. Columns, on the other hand, are stored in a cache and can be pulled onto the CPU buffer as desired. Thus, using reference line rows is computationally more expensive in terms of system resources than using reference line columns. Since it may be desirable to reduce line buffer memory usage, the following MLIP schemes use different numbers of reference rows and reference columns. Specifically, the following MLIP schemes use fewer reference rows than reference columns.
[0127] Figure 14 is a schematic diagram showing an exemplary mechanism 1400 for encoding alternative reference lines with different numbers of rows and columns. Mechanism 1400 can be used in an encoder or decoder, such as encoder 300 or decoder 400, when performing DC intra-prediction according to intra-prediction mode 500. Mechanism 1400 can be used in conjunction with the intra-prediction mode subset 930, conditional signal transmission representations 1000 and 1100, DC mode intra-prediction 1200, mechanism 1300, or independently of such examples.
[0128] The mechanism 1400 uses the current block 1401 with alternative reference lines 1411 designated as 0-M. Any number of reference lines 1411 can be used as such. Such reference lines 1411 are substantially similar to reference lines 811 and / or 1211 and include reference samples substantially similar to reference samples 812 and / or 1212. Reference line 1411 includes reference line 1413 containing reference samples located above the current block 1401. Reference line 1411 also contains reference column 1414, which contains the reference sample currently located to the left of block 1401. As described above, the MLIP scheme stores all reference rows 1413 in on-chip memory during intra-prediction, which is resource-intensive. Therefore, to reduce memory usage and / or to obtain a certain trade-off between coding efficiency and memory usage, the number of reference rows 1413 may be less than the number of reference columns 1414. As shown in Figure 14, the number of reference rows 1413 may be half the number of reference columns 1414. For example, this can be achieved by removing half of the reference rows 1413 between the first reference row 1413 and the last reference row 1413 (e.g., retaining RefLine0 and RefLine3 but removing RefLine1 and RefLine2). In another example, odd rows 1413 or even rows 1413 can be removed (e.g., retaining RefLine0 and RefLine2 but removing RefLine1 and RefLine3, or vice versa). A mechanism 1400 for removing row 1413 can be predefined.
[0129] Figure 15 is a schematic diagram showing another exemplary mechanism 1500 for encoding alternative reference lines having a different number of rows and columns. Mechanism 1500 can be used in an encoder or decoder, such as encoder 300 or decoder 400, when performing DC intra-prediction according to intra-prediction mode 500, for example. Mechanism 1500 can be used in conjunction with the intra-prediction mode subset 930, conditional signal transmission representations 1000 and 1100, DC mode intra-prediction 1200, mechanism 1300, or independently of such examples.
[0130] Mechanism 1500 is substantially similar to mechanism 1400, but uses a different number of reference rows 1513. Specifically, mechanism 1500 uses current block 1501 with alternative reference lines 1511 designated as 0-M. Any number of reference lines 1511 can be used as such. Such reference lines 1511 are substantially similar to reference lines 811 and / or 1211 and contain reference samples substantially similar to reference samples 812 and / or 1212. Reference line 1511 contains reference row 1513 containing reference samples located above current block 1501. Reference line 1511 also contains reference column 1514 containing reference samples located to the left of current block 1501.
[0131] In this case, the number of reference rows 1513 is the number of reference columns 1514 minus K, where K is a positive integer less than the number of reference columns 1514. K may be predefined or signaled in the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), or slice header in the bitstream. In the example shown, the number of reference rows 1513 is the number of reference columns 1514 minus 1. However, any value for K may be used. Reference rows 1513 are removed using any predefined mechanism. In Figure 15, reference column 1513 associated with reference line 1511 indexed as 1 is removed. In other examples, reference column 1513 associated with reference line 1511 indexed as 2, 3, etc., is removed.
[0132] When using mechanisms 1400 and / or 1500, the reference index of a referenced row may differ from the reference index of a referenced column. Several mechanisms for indicating or deriving reference row indices when the number of referenced rows and referenced columns differ as discussed above are presented below. In some mechanisms, the referenced column index and the left row index are signaled separately. For example, one syntax element is used to signal the referenced row index between 0 and NumTop, where NumTop is the number of referenced rows (including 0 and NumTop). Another syntax element is used to signal the referenced column index between 0 and NumLeft, where NumLeft is the number of referenced columns (including 0 and NumLeft).
[0133] In another mechanism, a single syntax element is used to signal the reference line index for both the reference row and the reference column. In most examples, NumLeft is greater than NumTop. However, the mechanism described below also applies to situations where NumTop is greater than NumLeft. When NumLeft > NumTop, in one mechanism, the index of the reference column is signaled using the mechanism described above. In such cases, the reference line index is between 0 and NumLeft (including 0 and NumLeft). If a reference line or reference pixel from the reference row is used and the reference line index is greater than NumTop, the reference line associated with the index of NumTop is signaled.
[0134] In another mechanism, if NumLeft > NumTop, the index of the reference row is signaled using the mechanism described above. In such a case, the reference index is between 0 and NumTop (including 0 and NumTop). When a reference line or pixel from the reference row is used, the signaled (or parsed) reference line index indicates the selected reference line. Since NumLeft is greater than NumTop, when a reference sample is selected from a reference column, an index mapping table is used to map the signaled reference line index to the selected reference column. Exemplary mapping tables 9-11 are shown below. [Table 9] [Table 10] [Table 11] In some examples, the table mappings shown above may be replaced by computations, depending on the implementation.
[0135] Furthermore, some intra-prediction modes use reference columns as reference lines, while others use reference rows as reference lines. Conversely, some intra-prediction modes use both reference rows and reference columns as reference lines. For example, the index range definition may depend on the intra-prediction mode, because the phase process and phase result are determined by the index range.
[0136] As a concrete example, an intra-predictive mode using reference columns (e.g., [DirS,DirH]) has an index range [0,NumLeft]. An intra-predictive mode using reference rows ([DirV,DirE]) has an index range [0,NumTop]. There are two possible intra-predictive modes that use both reference columns and reference rows. In the first case, the index range is between [0,NumTop] (denoted as Modi1). In this case, NumLeft > NumTop, and the index range is [1,NumTop]. Therefore, several reference columns are selected from NumLeft reference columns, where the number of selected reference columns is less than or equal to NumTop. The selection mechanism is predefined in the encoder and decoder. In the second case, the index range is between [0,NumLeft] (denoted as Modi2). In this case, a mapping mechanism is used to determine the index of the reference column and the index reference row, for example, based on the signal-transmitted reference line index. The mapping mechanism is predefined. For example, any of the mapping mechanisms listed in Tables 9 to 11 above can be used.
[0137] As another specific example, when the number of reference lines is four, the reference rows for reference lines RefLine1 and RefLine3 may be removed to reduce on-chip memory usage. In this case, four reference columns and two reference rows are used. When an intra-prediction mode uses both reference columns and reference rows, the index range is redefined to uniquely signal the selected reference sample(s). In this case, the intra-prediction mode [DirS,DirH] has an index range of [0,3]. For the intra-prediction mode [DirV,DirE], the index range is [0,1].
[0138] For intra-prediction modes that use both reference columns and reference rows (e.g., intra-prediction modes in a plane, DC, or (DirH,DirV)), the index range is either [0,1] (Modi1) or [0,3] (Modi2). When the index range is [0,1] (denoted as Modi1), two of the reference columns are used (e.g., the left portions of RefLine0 and RefLine2 are used). In this case, both the number of reference rows and reference columns are 2. The remaining two reference columns (e.g., RefLine1 and RefLine3) may also be selected for use. The selection mechanism is predefined. When the index range is [0,3] (denoted as Modi2), the number of reference columns is 4, the number of reference rows is 2, and the index range is [0,3]. In this case, an index mapping mechanism is employed (e.g., the reference column index in [0,1] corresponds to reference row index 0, and the reference column index in [2,3] corresponds to reference row index 1). Other mapping mechanisms can also be used (for example, column indices in [0,2] correspond to row index 0, column 3 corresponds to row index 1, etc.). However, the mapping mechanism used is predefined.
[0139] By combining a reference index coding mechanism and a reference line construction mechanism, coding efficiency can be improved and on-chip memory usage can be reduced. The following is an example combination of such a mechanism. In this case, the shortest codeword is assigned to the index of the furthest reference line, and NumTop is half of NumLeft. Furthermore, the number of reference lines is 4, and Ex1 in Table 8 is used. For the index coding mechanism, if NumTop is less than NumLeft, Modi1 is used, and the columns RefLine0 and RefLine3 are retained. As a result, a reference line signaling mechanism as shown in Table 12 below is obtained. [Table 12]
[0140] In this example, when the intra-prediction mode is between [DirS,DirH], the index range is [0,3], and the mechanism for representing the reference line index is shown in Table 12. Reference columns can be selected as reference lines according to the reference line index. Furthermore, when the intra-prediction mode is between [DirV,DirE], the index range is [0,1], and the mechanism for representing the reference line index is defined in Table 12. Reference rows can be selected as reference lines according to the reference line index. In this case, RefLine0 and RefLine3 are taken into consideration, but RefLine1 and RefLine2 are not. When the intra-prediction mode is between [Planar,DC] and (DirH,DirV), the index range is [0,1], and the mechanism for representing the reference line index is defined in Table 12. Then, reference columns and reference rows are selected as reference lines according to the reference line index. In this case, RefLine0 and RefLine3 are taken into consideration, but RefLine1 and RefLine2 are not. Similarly, the binary representation table discussed above may be applied to further improve encoding efficiency.
[0141] Figure 16 is a schematic diagram of a video coding device 1600 according to one embodiment of the present disclosure. The video coding device 1600 is suitable for implementing the disclosed embodiments as described herein. The video coding device 1600 has a downstream port 1620, an upstream port 1650, and / or a transceiver unit (Tx / Rx) 1610 for communicating data upstream and / or downstream over a network. The video coding device 1600 also includes a processor 1630, which includes a logic unit and / or a central processing unit (CPU) for processing data, and a memory 1632 for storing data. The video coding device 1600 may also have optical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components coupled to the upstream port 1650 and / or the downstream port 1620 for communicating data over an optical or wireless communication network. The video encoding device 1600 may also include an input and / or output (I / O) device 1660 for communicating data with the user. The I / O device 1660 may include output devices such as a display for showing video data and speakers for outputting audio data. The I / O device 1660 may also include a corresponding interface for interacting with input devices such as a keyboard, mouse, or trackball and / or such output devices.
[0142] The processor 1630 is implemented by hardware and software. The processor 1630 may be implemented as one or more CPU chips, cores (e.g., a multi-core processor), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 1630 communicates with the downstream port 1620, the Tx / Rx 1610 upstream port 1650, and the memory 1632. The processor 1630 has an encoding module 1614. The encoding module 1614 implements the embodiments disclosed above, such as methods 1700, 1800, 1900, and / or any other mechanisms described above. Furthermore, the encoding module 1614 implements a codec system 200, an encoder 300, and a decoder 400, performs intra-prediction using an intra-prediction mode 500 with blocks 600, uses a primary reference line scheme 700, uses an alternate reference line scheme 800, uses an intra-prediction mode subset 930, uses representations 1000 and / or 1100, uses DC-mode intra-prediction 1200, and may use mechanisms 1300, 1400 and / or 1500, as well as any combination thereof. Thus, by including the encoding module 1614, the functionality of the video encoding device 1600 is greatly improved, and conversion of the video encoding device 1600 to different states is performed. Alternatively, the encoding module 1614 can also be implemented as instructions stored in memory 1632 and executed by processor 1630 (for example, as a computer program product stored on a non-temporary medium).
[0143] Memory 1632 includes one or more memory types such as disks, tape drives, solid-state drives, read-only memory (ROM), random access memory (RAM), flash memory, ternary content-addressable memory (TCAM), and static random access memory (SRAM). Memory 1632 may also be used as an overflow data storage device to store a program when such a program is selected for execution and to store instructions and data that are read during program execution.
[0144] Figure 17 is a flowchart of an exemplary method 1700 for video coding using an intra-predictive mode subset with alternate reference lines. For example, method 1700 may operate on a video coding device 1600 configured to function as a decoder 400. Method 1700 may also use an intra-predictive mode 500 in relation to an alternate reference line scheme 800 and an intra-predictive mode subset 930.
[0145] In step 1701, a bitstream is received. The bitstream contains encoded compressed video data compressed by the decoder. In step 1703, an intra-predictive mode subset is determined, such as the intra-predictive mode subset 930. The intra-predictive mode subset includes intra-predictive modes correlated to multiple reference lines for the current image block, and excludes intra-predictive modes correlated to the primary reference line for the current image block. The intra-predictive mode subset may contain various groups of intra-predictive modes, as discussed above. In some examples, the intra-predictive mode subset includes directional intra-predictive modes for every Nth DirS, DirE, and between DirS and DirE, where N is a given integer value. In other examples, the intra-predictive mode subset may further include planar predictive modes and DC predictive modes. In a further example, the intra-prediction mode subset includes valid directional intra-prediction modes in the positive or negative N directions of DirS, DirE, DirD, DirH, DirV, and DirV, where N is a predetermined integer value. Such an intra-prediction mode subset may further include planar prediction modes and DC prediction modes. In yet another example, the intra-prediction mode subset includes intra-prediction modes selected for a decoded adjacent block located in a predetermined adjacency state to the current image block. In yet another example, the intra-prediction mode subset includes modes associated with the MPM list for that block. Furthermore, the intra-prediction mode subset may include any other combination of the intra-prediction modes discussed above.
[0146] In the optional step 1705, if the selected intra-prediction mode is included in the intra-prediction mode subset, the selected intra-prediction mode is decoded by the intra-prediction mode subset index. In the optional step 1707, if the selected intra-prediction mode is not included in the intra-prediction mode subset, the selected intra-prediction mode is decoded by the intra-prediction mode index. In some cases, the selected intra-prediction mode may be decoded based on the MPM index, as discussed with respect to Figure 9. Flags may be used to provide context for determining the indexing scheme, as discussed with respect to Figures 10-11.
[0147] In step 1709, if the selected intra-prediction mode is included in the intra-prediction mode subset, the reference line index is decoded from the encoded form. The reference line index indicates the selected reference line from the multiple reference lines for the selected intra-prediction mode. As discussed above, if the selected intra-prediction mode is not included in the intra-prediction mode subset, the reference line index is not decoded in order to reduce extra bits in the bitstream. To further support such conditional signaling in the bitstream, the reference line index may be encoded after the selected intra-prediction mode in the encoding (for example, as discussed with respect to Figures 10-11).
[0148] In step 1711, video data is presented to the user via a display. The video data includes image blocks decoded based on the selected intra-prediction mode and the corresponding reference line(s).
[0149] Figure 18 is a flowchart of an exemplary method 1800 for video coding using DC-mode intra-prediction with alternate reference lines. For example, method 1800 may operate on a video coding device 1600 configured to act as a decoder 400. Method 1800 may use an intra-prediction mode 500 in relation to an intra-prediction mode subset 930 with an alternate reference line scheme 800 and DC-mode intra-prediction 1200.
[0150] In step 1801, the bitstream is stored in memory. The bitstream contains compressed video data. The compressed video data contains image blocks encoded as prediction blocks according to a certain intra-prediction scheme. In step 1803, a current prediction block encoded by the DC intra-prediction mode is obtained. In step 1805, a DC prediction value is determined for the current prediction block. The DC prediction value approximates the current image block corresponding to the current prediction block by determining the average of all reference samples in at least two of a plurality of reference lines associated with the current prediction block. Thus, step 1805 extends the DC prediction mode to an alternate reference line context. In some examples, determining the DC prediction value may involve determining the average of all reference samples in N reference lines adjacent to the current prediction block, where N is a predetermined integer. In some examples, determining the DC prediction value involves determining the average of all reference samples in a selected reference line and its corresponding (e.g., predefined) reference line. In yet another example, determining the DC prediction value involves determining the average of all reference samples in an adjacent reference line (e.g., a reference line with index 0) and a selected reference line that has been signaled in the bitstream. Furthermore, the DC prediction values may be determined using any combination of schemes as discussed above with respect to Figure 12. In step 1807, the current image block is reconstructed based on the DC prediction values. In block 1809, the frame containing the current image block is displayed to the user.
[0151] Figure 19 is a flowchart of an exemplary method 1900 for video coding using reference lines encoded by codewords based on selection probability. For example, method 1900 may operate on a video coding device 1600 configured to act as a decoder 400. Method 1900 may also use an intra-prediction mode 500 in relation to an alternative reference line scheme 800 and an intra-prediction mode subset 930 using mechanism 1300. Furthermore, different numbers of reference rows and reference columns may be employed by method 1900, as discussed with respect to Figures 14 and 15.
[0152] In step 1901, a bitstream containing the encoded data is received. The encoded data contained video data compressed by the encoder. In step 1903, an intra-prediction mode is decoded from the encoded data. The intra-prediction mode shows the relationship between the current block and the reference samples in the selected reference lines. Furthermore, the current block is associated with a plurality of reference lines, including the selected reference lines (e.g., an alternate reference line scheme). In step 1905, the selected reference lines are decoded based on a selected codeword that indicates the selected reference line. The selected codeword includes a length based on the selection probability of the selected reference line, as discussed with respect to Figure 13. For example, the plurality of reference lines may be indicated by a plurality of codewords. Furthermore, the reference line furthest from the current block may be indicated by a codeword with the second shortest length. In another example, the second furthest reference line from the current block may be indicated by a codeword with the second shortest length. In yet another example, predefined reference lines other than adjacent reference lines may be indicated by codewords with the second shortest length. In yet another example, multiple codewords may be sorted into Class A and Class B groups. Class A groups may contain codewords shorter in length than those in Class B groups. Furthermore, Class A and Class B groups can be incremented and decremented independently of each other. In yet another example, a reference line may contain a different number of reference rows and reference columns, as discussed with respect to Figures 14 and 15. For example, multiple reference lines may contain reference rows and reference columns. Furthermore, the number of reference rows stored for the current block may be half the number of reference columns stored for the current block. In yet another example, the number of reference rows stored for the current block may be equal to the number of reference columns stored for the current block minus one. In yet another example, the number of reference rows stored for the current block may be selected based on the number of reference rows used by the deblocking filter operation. In step 1907, video data is presented to the user via the display.The video data includes image blocks decoded based on the intra-prediction mode and selected reference lines. Thus, methods 1700, 1800, and 1900 can be applied individually or in any combination to improve the effectiveness of alternative reference line schemes when encoding video via intra-prediction.
[0153] A video encoding apparatus comprising: receiving means for receiving a bitstream; processing means for determining an intra-prediction mode subset, wherein the intra-prediction mode subset includes intra-prediction modes correlated with a plurality of reference lines for the current image block, excluding intra-prediction modes correlated with a primary reference line for the current image block; processing means configured to perform, if a first intra-prediction mode is included in the intra-prediction mode subset, decoding the first intra-prediction mode by an alternate intra-prediction mode index; and if the first intra-prediction mode is not included in the intra-prediction mode subset, decoding the first intra-prediction mode by an intra-prediction mode index; and display means for presenting video data including an image block decoded based on the first intra-prediction mode.
[0154] A method comprising: storing a bitstream containing image blocks encoded as prediction blocks in memory means; obtaining a current prediction block encoded by a direct current (DC) intra-prediction mode by processing means; determining a DC prediction value for approximating the current image block corresponding to the current prediction block by determining the average of all reference samples in at least two of a plurality of reference lines associated with the current prediction block; reconstructing the current image block based on the DC prediction value by the processor; and displaying a video frame containing the current image block on display means.
[0155] A video coding apparatus comprising: receiving means for receiving a bitstream; processing means configured to perform the steps of: decoding an intra-prediction mode from the bitstream, wherein the intra-prediction mode indicates a relationship between a current block and a selected reference line, wherein the current block is associated with a plurality of reference lines, including the selected reference line; decoding the selected reference line based on a selected codeword indicating the selected reference line, wherein the selected codeword includes a length based on the selection probability of the selected reference line; and display means for presenting video data including an image block decoded based on the intra-prediction mode and the selected reference line.
[0156] A first component is directly coupled to a second component if there are no intervening components between them other than lines, traces, or other media. A first component is indirectly coupled to a second component if there are intervening components other than lines, traces, or other media between them. The term "coupled" and its variations include both direct and indirect coupling. The use of the term "about" means a range including ±10% of the following number unless otherwise specified.
[0157] While several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The examples of this application are illustrative and not limiting, and their intent is not limited to the details given herein. For example, different elements or components can be combined or integrated into another system, or certain features may be omitted or not implemented.
[0158] Furthermore, the technologies, systems, subsystems, and methods described and illustrated discretely or separately in various embodiments may be combined with or integrated with other systems, components, technologies, or methods without departing from the scope of this disclosure. Other examples of modifications, substitutions, and alterations may be identified by those skilled in the art and made without departing from the spirit and scope disclosed herein.
Claims
1. This is a video coding method: A step of obtaining a reference line index of a coding unit, wherein the coding unit is associated with a plurality of reference lines, the plurality of reference lines including a primary reference line and one or more additional reference lines located further away from the coding unit than the primary reference line; If the reference line index indicates a primary reference line, the step is to determine a first intra-prediction mode for the coding unit from the intra-prediction mode list; If the reference line index indicates one of the additional reference lines, the step of determining a second intra-prediction mode for the coding unit is to be taken only from a subset of the intra-prediction mode list. Methods that include...
2. The method according to claim 1, wherein the subset of the intra-predictive mode list includes only the modes associated with the most probable mode (MPM) list.
3. The method according to claim 1 or 2, wherein the second intra-prediction mode is indicated by an MPM list index.
4. The method according to claim 2 or 3, wherein the candidate intra-prediction modes in the MPM list include intra-prediction modes used by the neighboring coding units of the coding unit.
5. The method according to any one of claims 1 to 4, further comprising the step of coding the coding unit by obtaining block prediction values based on the first intra prediction mode or the second intra prediction mode and the plurality of reference lines.
6. The method according to any one of claims 1 to 5, wherein the value of the reference line index is presumed to be equal to zero if it does not exist.
7. The method according to any one of claims 1 to 6, wherein the subset of the intra-prediction mode list includes planar prediction modes and direct current (DC) prediction modes.
8. A step of obtaining a reference line index of a coding unit, wherein the coding unit is associated with a plurality of reference lines, the plurality of reference lines including a primary reference line and one or more additional reference lines located further away from the coding unit than the primary reference line; If the reference line index indicates a primary reference line, the step is to determine a first intra-prediction mode for the coding unit from the intra-prediction mode list; If the reference line index indicates one of the additional reference lines, the step of determining a second intra-prediction mode for the coding unit is to be taken only from a subset of the intra-prediction mode list. A video coding device having at least one processor configured to perform the following:
9. The video coding apparatus according to claim 8, wherein the subset of the intra-predictive mode list includes only the modes associated with the most probable mode (MPM) list.
10. The video coding apparatus according to claim 8 or 9, wherein the second intra-prediction mode is indicated by an MPM list index.
11. The video coding apparatus according to claim 9 or 10, wherein the candidate intra-prediction modes in the MPM list include intra-prediction modes used by neighboring coding units of the coding unit.
12. The video coding apparatus according to any one of claims 8 to 11, wherein one of the processors is further configured to code the coding unit by obtaining block prediction values based on the first intra prediction mode or the second intra prediction mode and the plurality of reference lines.
13. The video coding apparatus according to any one of claims 8 to 12, wherein the value of the reference line index is presumed to be equal to zero if it does not exist.
14. The video coding apparatus according to any one of claims 8 to 12, wherein the subset of the intra prediction mode list includes a planar prediction mode and a direct current (DC) prediction mode.
15. Memory containing instructions; One or more processors that communicate with the aforementioned memory A video decoding device having, The one or more processors execute the instruction to perform the method according to any one of claims 1 to 7. Video decoding device.
16. Memory containing instructions; One or more processors that communicate with the aforementioned memory A video encoding device having, The one or more processors execute the instruction to perform the method according to any one of claims 1 to 7. Video encoding device.
17. A non-temporary storage medium comprising a bitstream encoded by the method described in any one of claims 1 to 7.