Weighted planar and DC modes for intra prediction
By using weighted average vertical and horizontal predictors in video coding and adjusting the weights according to the block shape, the problem of insufficient intra-frame prediction efficiency in existing technologies is solved, and more efficient video compression is achieved.
Patent Information
- Application Number
- CN202480024536.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-06
- Filing Date
- 2024-03-25
- Publication Date
- 2025-11-11
AI Technical Summary
Existing video coding techniques fail to fully utilize block shape information in intra-frame prediction, resulting in insufficient compression efficiency.
A weighted average vertical and horizontal predictor is used, with the weights adjusted according to the aspect ratio of the target block to form a more accurate predictor for sample prediction of intra-frame blocks.
It improves the accuracy of intra-frame prediction and compression efficiency, especially in video coding performance under different target block shapes.
Smart Images

Figure CN120937364A_ABST
Abstract
Description
Technical Field
[0001] This embodiment generally relates to methods and apparatus for intra-frame prediction in video encoding and decoding. Background Technology
[0002] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transform to utilize spatial and temporal redundancy in the video content. Generally, intra-frame or inter-frame prediction is used to leverage intra-frame or inter-frame image correlations. Then, the differences between the original and predicted blocks (often denoted as prediction error or prediction residual) are transformed, quantized, and entropy-coded. To reconstruct the video, the compressed data is decoded through the inverse process corresponding to entropy coding, quantization, transform, and prediction. Summary of the Invention
[0003] According to one embodiment, a method for video decoding is presented, comprising: obtaining a vertical predictor and a horizontal predictor for samples in an intra-frame block; obtaining a weighted average of the vertical predictor and the horizontal predictor to form a predictor for the samples, wherein, in forming the weighted average, a first weight for the vertical predictor is different from a second weight for the horizontal predictor; and predictively decoding the samples based on the predictor for the samples.
[0004] According to another embodiment, a video coding method is presented, comprising: obtaining a vertical predictor and a horizontal predictor for samples in an intra-frame block; obtaining a weighted average of the vertical predictor and the horizontal predictor to form a predictor for the samples, wherein a first weight for the vertical predictor is different from a second weight for the horizontal predictor when forming the weighted average; and predictively coding the samples based on the predictor for the samples.
[0005] According to another embodiment, an apparatus is presented comprising: at least one memory; and one or more processors coupled to the memory, wherein the one or more processors are configured to: obtain a vertical predictor and a horizontal predictor for samples in an intra-frame block; obtain a weighted average of the vertical predictor and the horizontal predictor to form a predictor for the samples, wherein a first weight for the vertical predictor differs from a second weight for the horizontal predictor when forming the weighted average; and decode the samples based on the predictor for the samples.
[0006] According to another embodiment, an apparatus is provided, comprising: at least one memory; and one or more processors coupled to the memory, wherein the one or more processors are configured to: obtain a vertical predictor and a horizontal predictor for samples in an intra-frame block; obtain a weighted average of the vertical predictor and the horizontal predictor to form a predictor for the samples, wherein a first weight for the vertical predictor is different from a second weight for the horizontal predictor when forming the weighted average; and encode the samples based on the predictor for the samples.
[0007] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding or decoding method according to any of the embodiments described herein. One or more of these embodiments also provide a computer-readable storage medium having instructions stored thereon for video encoding or decoding according to the methods described herein.
[0008] One or more embodiments also provide a computer-readable storage medium storing video data generated according to the methods described above. One or more embodiments also provide methods and apparatus for transmitting or receiving video data generated according to the methods described herein. Attached Figure Description
[0009] Figure 1 A block diagram of a system in which aspects of this embodiment can be implemented is shown.
[0010] Figure 2 A block diagram illustrating an embodiment of a video encoder is shown.
[0011] Figure 3 A block diagram illustrating an embodiment of a video decoder is shown.
[0012] Figure 4A and Figure 4B The diagrams illustrate the horizontal and vertical interpolation in the PLANA prediction within VVC.
[0013] Figure 5 The illustration shows the DC prediction in VVC and ECM for different target block shapes.
[0014] Figure 6 The illustration depicts a method for intra-frame prediction with weighted PLANA or DC modes according to an embodiment. Detailed Implementation
[0015] Figure 1The diagram illustrates an example of a system in which various aspects and embodiments can be implemented. System 100 may be embodied as a device including the various components described below and configured to perform one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 100 may be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or to other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in this application.
[0016] System 100 includes: at least one processor 110 configured to execute instructions loaded therein for implementing various aspects, such as those described in this application. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes: a storage device 140 which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. Storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices, as non-limiting examples.
[0017] System 100 includes an encoder / decoder module 130 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 130 may be implemented as a separate element of system 100, or may be incorporated into processor 110 as a combination of hardware and software as known to those skilled in the art.
[0018] Program code to be loaded onto processor 110 or encoder / decoder 130 to execute the various aspects described herein may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more items of various kinds during execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing of equations, formulas, operations, and operational logic.
[0019] In several embodiments, memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 110 or encoder / decoder module 130) is used for one or more of these functions. External memory may be memory 120 and / or storage device 140, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, fast external volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2, HEVC, or VVC.
[0020] Input to the components of system 100 can be provided by various input devices as indicated in box 105. Such input devices include, but are not limited to: (i) an RF section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) a composite input terminal; (iii) a USB input terminal; and / or (iv) an HDMI input terminal.
[0021] In various embodiments, the input device of block 105 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal or limiting a signal to a band); (ii) down-converting the selected signal; (iii) further limiting the band to a narrower band to select, for example, a signal band that may be referred to as a channel in some embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select the desired stream of data packets. The RF section in various embodiments includes one or more elements to perform these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to the desired frequency band. Various embodiments rearrange the order of the components described above (and others), remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0022] Additionally, the USB and / or HDMI endpoints may include corresponding interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or, if necessary, within processor 110. Similarly, aspects of USB or HDMI interface processing may be implemented, either within a separate interface IC or, if necessary, within processor 110. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on an output device.
[0023] Various components of the system 100 can be provided within an integrated housing, in which the various components can be interconnected and data can be transmitted therebetween using a suitable connection arrangement 115, such as an internal bus as known in the art, including an I2C bus, wiring, and printed circuit board.
[0024] System 100 includes a communication interface 150, which enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 190 may be implemented over, for example, wired and / or wireless media.
[0025] In various embodiments, a Wi-Fi network such as IEEE 802.11 is used to stream data to system 100. The Wi-Fi signals in these embodiments are received on a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that delivers data over an HDMI connection in input box 105 to provide streaming data to system 100. Still other embodiments use an RF connection in input box 105 to provide streaming data to system 100.
[0026] System 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various examples of embodiments, other peripheral devices 185 include one or more of a standalone DVR, disc player, stereo system, lighting system, and other devices that provide output functionality based on system 100. In various embodiments, signaling such as AV is used to transmit control signals between system 100 and the display 165, speaker 175, or other peripheral devices 185. Device-to-device control links, CEC, or other communication protocols are implemented with or without user intervention. Output devices can be communicatively coupled to system 100 via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 via communication interface 150 using communication channel 190. The display 165 and speaker 175 can be integrated into a single unit within an electronic device (e.g., a television). In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0027] Display 165 and speaker 175 can alternatively be separated from one or more other components, for example, if the RF section of input 105 is part of a separate set-top box. In various embodiments where display 165 and speaker 175 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0028] Figure 2 The illustration shows an example video encoder 200, such as a VVC (Various Video Coding) encoder. Figure 2 The diagram can also show encoders that improve upon the VVC standard or encoders that use technologies similar to VVC.
[0029] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, as may the terms "encoded" or "coded," and the terms "image," "picture," and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side and "decoded" is used on the decoder side.
[0030] Before being encoded, the video sequence may undergo pre-encoding processing (201), such as applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a more resilient signal distribution for compression (e.g., using histogram equalization with one of the color components). Metadata may be associated with preprocessing and attached to the bitstream.
[0031] In encoder 200, the image is encoded by encoder elements, as described below. The image to be encoded is partitioned (202) and processed in units such as CUs. Each unit is encoded using, for example, an intra-frame or inter-frame mode. When a unit is encoded in an intra-frame mode, it performs intra-frame prediction (260). In an inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use for encoding the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. Prediction enhancement (285) is applied to the prediction block before prediction. The prediction residual is calculated, for example, by subtracting (210) the prediction block from the original image block.
[0032] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (245) to output a bit stream. The encoder can skip the transform and apply the quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., encode the residual directly without applying the transform or quantization process.
[0033] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are combined (255) to reconstruct the image blocks. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (280).
[0034] Figure 3 A block diagram of an example video decoder 300 is shown. In decoder 300, the bitstream is decoded by decoder elements, as described below. Video decoder 300 generally performs operations similar to... Figure 2 The encoding passes described herein are mutually decoded passes. Encoder 200 generally also performs video decoding as part of the encoding of video data.
[0035] Specifically, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. First, entropy decoding (330) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Image partitioning information indicates how the image should be partitioned. Therefore, the decoder can partition (335) the image based on the decoded image partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (355) to reconstruct the image block. The predicted block (370) can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). After prediction, prediction enhancement (390) is applied to the predicted block. A loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference image buffer (380).
[0036] The decoded image can undergo further post-decoding processing (385), such as inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4), or inverse remapping, which is the inverse operation of the remapping process performed in pre-encoding processing (201). Post-decoding processing can utilize metadata derived in pre-encoding processing and signaled in the bitstream.
[0037] In VVC and ECM, the encoding of video sequence frames is based on a quadtree (QT) / binary tree (BT) / tritree (TT) block structure. Frames are divided into non-overlapping square coding tree units (CTUs), which undergo QT / BT / TT-based splitting into multiple coding units (CUs) based on a rate distortion criterion. In intra-frame prediction, different prediction modes are used to spatially predict CUs from causal neighbor CUs (i.e., the decoded CUs at the top and left). From 95 defined modes, one is a planar mode (indexed as mode 0), one is a DC mode (indexed as mode 1), and the remaining 93 (indexed as modes -14, ..., -1, 2, ..., 80) are corner modes. From the 93 corner modes, depending on their shape, only 65 neighboring modes are selected for any target CU. The planar and DC modes are designed to model regions of slowly changing intensity, while the directional modes are designed to model the directional structure within the frame. It is generally observed that non-angular patterns (especially planar patterns) have the highest probability within intra-frame sequences.
[0038] With some minor modifications, two non-angular modes have been implemented from the HEVC standard to adapt to possible rectangular target blocks in VVC and ECM. Planar mode predictions are obtained by averaging the horizontal and vertical interpolations, regardless of the target block shape. DC predictions are obtained by averaging reference samples on the longer side of the block. In this document, we propose a simple and implementable modification using weighted combinations. Before describing the proposed method, we briefly present the planar mode and DC mode predictions as used in VVC and ECM 7.0 below. For ease of reference, we will use the terms "CU" and "block" interchangeably throughout the text.
[0039] Planar pattern prediction in ECM Figure 4A and Figure 4B The diagrams illustrate the horizontal and vertical interpolations in the PLANA prediction within the VVC. The average of the two interpolations gives the first prediction in the planar pattern. The horizontal and vertical interpolations also constitute the horizontal and vertical PLANA in the ECM.
[0040] Planar prediction is constructed using the average of horizontal and vertical interpolations, such as... Figure 4A and Figure 4B As shown. In such Figure 4A In the horizontal interpolation shown, the pixel values in row y are linearly interpolated using the left reference sample R(-1,y) and the upper-right reference sample R(W,-1) at (W,-1), where W denotes the target block width. Similarly, in... Figure 4BIn the vertical interpolation shown, the pixel values on column x are linearly interpolated using the upper reference sample R(x,-1) and the lower left reference sample R(-1,H) at (-1,H), where H indicates the target block height. The average of these two interpolations makes the first prediction in the planar pattern. In this document, the calculation is based on an integer implementation. It should be noted that other forms of integer implementation besides the form used (e.g., at other precisions) are also possible. Vertical Predictor P v ( x,y (Vertical interpolation), Horizontal predictor P h ( x,y (Horizontal interpolation) and plane predictor P ( x,y ) can be expressed as: .
[0041] Alternatively, the planar predictor can be computed as: in .
[0042] Subsequently, the predicted values are processed using PDPC as shown in equation (3) to eliminate discontinuities with respect to the upper and left reference arrays: in w T = 32>>((y<<1)>>expansion / contraction ratio), w L = 32>>((x<<1)>>Scaling ratio), Expansion / contraction ratio = ( log 2( W ) + log 2( H ) – 2)>>2 .
[0043] Due to weight w T and w L These are decreasing functions of y and x, respectively, and therefore depend on the scaling factor. Only the top few rows and the left few columns of the block are affected by the PDPC process. The remaining values remain unchanged as in the first prediction step.
[0044] In ECM, horizontal and vertical interpolation have also been adopted as the horizontal and vertical planar modes, respectively. Normalizing the interpolation values given above, they are derived as follows: For the vertical plane (5) For the horizontal plane (6).
[0045] As in the normal plane, the first prediction value is then processed using PDPC in the same manner as given above.
[0046] And DC mode prediction in ECM Figure 5 Illustrates DC prediction in VVC and ECM for different target block shapes. The initial prediction value at any pixel is the average of the colored reference pixels.
[0047] Depending on the block shape, as Figure 5 shown, the DC prediction is constructed using the upper or left reference sample or the average of both. If the block is square-shaped, the average of the upper and left reference samples excluding the top-left pixel at (-1, -1) is used as the first prediction for all target pixels. If the block is flat, i.e., W > H, the average of the upper reference samples with coordinates from (0, -1) to (W–1, -1) is used as the first predictor. Similarly, if the block is tall, i.e., W < H, the average of the left reference samples with coordinates from (-1, 0) to (-1, H–1) is used as the first predictor. Note that the top-right reference sample (on the upper reference array) and the bottom-left reference sample (on the left reference array) are excluded from the DC value calculation. Mathematically, the first prediction value can be expressed as: .
[0048] Subsequently, the first prediction value is processed using PDPC in the same manner as done in the case of the planar mode to eliminate the discontinuities with respect to the upper and left reference arrays.
[0049] Weighted planar mode prediction As described above, in VVC, planar mode prediction is the average of two interpolations. Equivalently, a weight of 0.5 is given to each interpolation in the combination to derive the final prediction. In ECM, in addition, there are horizontal PLANAR and vertical PLANAR modes, which can be understood as giving a weight of 1 to one interpolation and a weight of 0 to the other interpolation. In this document, we propose a weighted planar mode, where in one example, the weights are derived based on the target block shape.
[0050] In the first method, the vertical and horizontal interpolations are obtained as in the vertical and horizontal PLANAR modes: .
[0051] They are then combined as follows: The weighting parameter wP is determined as follows: .
[0052] In the above combinations, horizontal interpolation is given a higher weight than vertical interpolation for flat rectangular blocks and vice versa for tall rectangular blocks. Alternatively, they can be combined as follows: The weighting is done in the reverse manner.
[0053] In another variation, horizontal interpolation can be combined with a strictly vertical pattern ( R ( x (, -1) is a predictor) can be combined, and vertical interpolation can be combined with a strictly horizontal mode ( R (-1, y (The predictor) is combined as follows: Alternatively, as: .
[0054] In special cases, by using equal weighting, two different versions of the PLANA pattern can be obtained by averaging the horizontal interpolation and strictly vertical patterns, or by averaging the vertical interpolation and strictly horizontal patterns.
[0055] Note that it is possible to combine normalization with scaling in these equations, resulting in only one rounding. For square target blocks, at wP = 32, this method becomes equivalent to the normal PLANA mode.
[0056] The weighting parameter wP is a function of the aspect ratio of the target block. In one variant, it can be determined based on the mean absolute difference (MAD) score calculated at the top and left reference arrays, after linear interpolation between the top-left and top-right reference pixels and between the top-left and bottom-left reference pixels, respectively.
[0057] Furthermore, the weighting parameter wP is a fixed value for the entire target block. In a variant, it can be chosen to vary depending on the (x,y) coordinates of the target pixels.
[0058] In another variation, the weighting parameter wP can be selected from a list of values using a template as done in TIMD, depending on the aspect ratio of the target block. The planar prediction for the template with each wP is subtracted from the template pixel values, and the SATD of the residuals is calculated. The wP value that gives the minimum SATD value is selected from the weighted planar predictions for the current block.
[0059] In another variation, if the left neighbor block is unavailable but the upper neighbor block is available, then a weight of 0 can be assigned to the horizontal interpolation and a weight of 1 can be assigned to the vertical interpolation. In this case, the prediction is equivalent to the vertical plane pattern. Similarly, if the upper neighbor block is unavailable but the left neighbor block is available, then a weight of 0 can be assigned to the vertical interpolation and a weight of 1 can be assigned to the horizontal interpolation. In this case, the prediction is equivalent to the horizontal plane pattern.
[0060] More generally, the weighting parameter wP can be adapted to the block width and height, and / or to the decoded samples at the left and top of the block.
[0061] As in a normal plane, the initial prediction can then be followed by the PDPC.
[0062] Weighted DC model prediction Regardless of the block shape, we first derive the DC values of the upper and left reference samples, which are then used as the average of the upper and left reference samples, respectively. Horizontal Predictor horDC (DC value of the upper reference sample) and vertical predictor verDC (The DC value of the left reference sample) can be calculated as: .
[0063] Obtain the final DC value for the prediction, as a weighted combination of the two DC values: The weighting parameter wP is derived, for example, as follows: .
[0064] For square blocks, wP = 32, and this becomes equivalent to the existing DC prediction. As in the case of weighted PLANA, it is possible to combine two rounding operations into a single operation.
[0065] The weighting parameter wP is a function of the aspect ratio of the target block. In one variant, it can be determined based on the mean absolute difference (MAD) score calculated at the top and left reference arrays, after linear interpolation between the top-left and top-right reference pixels and between the top-left and bottom-left reference pixels, respectively.
[0066] In another variation, fixed weights can be assigned to two DC values. For example, when W > 2H or H > 2W, we can set wP = 16. In the first case, the upper DC receives 3 / 4 of the weight and the left DC receives 1 / 4 of the weight, and in the second case, the reverse is true.
[0067] In another variation, the weighting parameter wP can be selected from a list of values using a template as done in TIMD, depending on the aspect ratio of the target block. The DC prediction for the template with each wP is subtracted from the template pixel values, and the SATD of the residuals is calculated. The wP value that gives the minimum SATD value is selected based on the weighted DC prediction for the current block.
[0068] In another variation, if the left neighbor block is unavailable but the upper neighbor block is available, then a horizontal DC value can be assigned. horDC Assign a weight of 1 and can give a vertical DC value verDC Assign a weight of 0. Similarly, if the upper neighbor block is unavailable but the left neighbor block is available, then a weight of 1 can be assigned to the vertical DC value and a weight of 0 can be assigned to the horizontal DC value.
[0069] More generally, the weighting parameter wP can be adapted to the block width and height, and / or to the decoded samples at the left and top of the block.
[0070] The initial prediction can then be followed by PDPC to arrive at the final prediction.
[0071] Figure 6 A method for intra-frame prediction with weighted PLANAR or DC modes according to one embodiment is illustrated. This method can be used at both the encoder and decoder. Specifically, the encoder or decoder computes (620) a horizontal predictor for the PLANAR or DC mode as in equation (9) or (13), and / or computes (630) a vertical predictor for the PLANAR or DC mode as in equation (8) or (14). Then, a weighted planar or DC predictor is computed (640) based on the vertical and / or horizontal predictors. A PDPC can be applied (650) to obtain the final prediction.
[0072] In video codecs, such as those based on VVC, ECM, etc., the encoder and decoder can follow the weighted predictions above for planar and DC modes. In a variant, the encoder can use classic planar and DC predictions, and the decoder, which includes the encoder, can use weighted predictions, or vice versa.
[0073] Hereinafter, we assume that the video codecs include those with block-based CU partitioning and various angular and non-angular prediction modes, such as codecs based on HEVC, VVC, or ECM standards.
[0074] In one embodiment, we replace existing PLANAR and / or DC intra-frame prediction modes with a proposed weighted planar / DC pattern. Horizontal and vertical interpolation in PLANAR or horizontal and vertical DC in DC prediction are combined with variable weights, where the weights are derived based on the target block width and height.
[0075] In another embodiment, in addition to the existing PLANAR / DC pattern, the proposed weighted PLANAR / DC pattern is also included as a new prediction pattern. The horizontal and vertical interpolations in PLANAR or the horizontal and vertical DC in DC prediction are combined with variable weights, where the weights are derived based on the target block width and height. Explicit signaling of the new pattern is then implemented.
[0076] In another embodiment, the signal notification for the new mode is not explicitly performed. Both the encoder and decoder use templates, such as those used in TIMD, to make decisions between normal PLANA / DC and weighted PLANA / DC.
[0077] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first," "second," etc., may be used in various embodiments to modify elements, components, steps, operations, etc., e.g., "first decoding" and "second decoding." The use of such terms does not imply a sequence of modified operations unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding, but may occur, for example, before, during, or in the time period overlapping with the second decoding.
[0078] Various methods and other aspects described in this application can be used to modify, such as Figure 2 and Figure 3 The modules of the video encoder 200 and decoder 300 shown are, for example, intra-frame prediction modules (260, 360). Furthermore, this aspect is not limited to VVC or HEVC and can be applied, for example, to other standards and recommendations, as well as any extensions of such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application may be used individually or in combination.
[0079] Various numerical values are used in this application. Specific values are used for illustrative purposes, and the aspects described are not limited to these specific values.
[0080] Various implementations involve decoding. As used herein, “decoding” can encompass all or part of a process performed, for example, on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, dequantization, inverse transform, and differential decoding. Whether the phrase “decoding process” is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description and is considered to be well understood by those skilled in the art.
[0081] Various implementations involve encoding. In a manner similar to the above discussion on "decoding," the term "encoding," as used in this application, can encompass all or part of the process performed, for example, on an input video sequence to produce an encoded bitstream.
[0082] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the features in question can be implemented in other forms (e.g., apparatus or program). Apparatus can be implemented, for example, in appropriate hardware, software, and firmware. Methods can be implemented, for example, in an apparatus, such as a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end users.
[0083] References to "an embodiment," "an embodiment," "an implementation," or "an implementation," and other variations thereof, mean that a particular feature, structure, characteristic, etc., described in connection with an embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment," "in one embodiment," "in one implementation," or "in one implementation," and any variations appearing throughout this application, do not necessarily all refer to the same embodiment.
[0084] Additionally, this application may refer to "determining" each piece of information. Determining information may include one or more of the following: estimated information, calculated information, predicted information, or information retrieved from memory.
[0085] Furthermore, this application may refer to "accessing" various pieces of information. Accessing information may include one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0086] Additionally, this application may refer to "receiving" various pieces of information. As with "access," "receiving" is intended to be a broad term. Receiving information may include one or more of, for example, accessing information or retrieving information (e.g., from memory). Further, "receiving" typically refers to something that occurs during operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0087] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of” (e.g., in the cases of “A / B”, “A and / or B”, and “at least one of A and B”) is intended to cover the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, this phrase is intended to cover the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be extended to as many items as are listed, as will be clear to those skilled in the art and related fields.
[0088] Moreover, as used herein, the term “signaling” refers, among other things, to instructing a corresponding decoder. For example, in some embodiments, the encoder signals the quantization matrix used for dequantization. In this way, in embodiments, the same parameters are used at both the encoder and decoder sides. Thus, for example, the encoder can transmit specific parameters (explicit signaling) to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has the specific parameters as well as other parameters, then signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameters. Bit saving is achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be done in a variety of ways. For example, in various embodiments, one or more syntax elements, tags, etc., are used to signal information to the corresponding decoder. Although the signature refers to the verb form of the term “signaling,” the term “signaling” may also be used as a noun herein.
[0089] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bit stream of the described embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is well known. The signal may be stored on a processor-readable medium.
Claims
1. A method for video decoding, comprising: Obtain the vertical and horizontal predictors for samples in an intra-block. A weighted average of the vertical predictor and the horizontal predictor is obtained to form a predictor for the sample, wherein, in forming the weighted average, a first weight for the vertical predictor is different from a second weight for the horizontal predictor. as well as The sample is decoded based on the predictor for the sample.
2. A video encoding method, comprising: Obtain the vertical and horizontal predictors for samples in an intra-block. A weighted average of the vertical predictor and the horizontal predictor is obtained to form a predictor for the sample, wherein when forming the weighted average, a first weight for the vertical predictor is different from a second weight for the horizontal predictor. as well as The sample is encoded based on the predictor for the sample.
3. An apparatus comprising: At least one memory; as well as One or more processors are coupled to the memory, wherein the one or more processors are configured to: Obtain the vertical and horizontal predictors for samples in an intra-block. A weighted average of the vertical predictor and the horizontal predictor is obtained to form a predictor for the sample, wherein, in forming the weighted average, a first weight for the vertical predictor is different from a second weight for the horizontal predictor. as well as The sample is decoded based on the predictor for the sample.
4. An apparatus comprising: At least one memory; as well as One or more processors are coupled to the memory, wherein the one or more processors are configured to: Obtain the vertical and horizontal predictors for samples in an intra-block. A weighted average of the vertical predictor and the horizontal predictor is obtained to form a predictor for the sample, wherein when forming the weighted average, a first weight for the vertical predictor is different from a second weight for the horizontal predictor. as well as The sample is encoded based on the predictor for the sample.
5. The method of claim 1 or 2 or the apparatus of claim 3 or 4, wherein the intra-frame block is encoded in DC mode or PLANA mode.
6. The method of any one of claims 1, 2 and 5 or the apparatus of any one of claims 3-5, wherein the height and width of the intra-frame block are not equal in length.
7. The method of any one of claims 1, 2, 5 and 6 or the apparatus of any one of claims 3-6, wherein the first and second weights are based on the width and height of the block.
8. The method of claim 7 or the apparatus of claim 7, wherein the first and second weights are based on the aspect ratio of the block.
9. The method of any one of claims 1, 2 and 5-8 or the apparatus of any one of claims 3-8, wherein the first and second weights are based on the values of the upper reference sample and the left reference sample.
10. The method of any one of claims 1, 2 and 5-9 or the apparatus of any one of claims 3-9, wherein the first and second weights are based on the difference between (1) the linear interpolation of the upper left and upper right reference samples and (2) the linear interpolation of the upper left and lower left reference samples.
11. The method of any one of claims 1, 2, and 5-10 or the apparatus of any one of claims 3-10, wherein the first and second weights are based on the availability of decoded neighboring blocks to the left and top of the current block.
12. The method of any one of claims 1, 2 and 5-11 or the apparatus of any one of claims 3-11, wherein the first and second weights are predetermined based on the aspect ratio of the block.
13. The method of any one of claims 1, 2, and 5-12 or the apparatus of any one of claims 3-12, wherein the first and second weights are selected from a list of weights based on the aspect ratio of the block, wherein a template is used to select the optimal weight pair from the list.
14. The method of any one of claims 1, 2, and 5-13 or the apparatus of any one of claims 3-13, wherein only the luminance samples are predicted using weights assigned to them by the corresponding vertical and horizontal predictors.
15. The method of any one of claims 1, 2, and 5-14 or the apparatus of any one of claims 3-14, wherein both luminance and chrominance samples are predicted using weights assigned to them by corresponding vertical and horizontal predictors, the weights being the same or different for the luminance and chrominance components.
16. The method of any one of claims 1, 2, and 5-15, or the apparatus of any one of claims 3-15, wherein the vertical predictor and the horizontal predictor for the sample are obtained based on vertical interpolation and horizontal interpolation, respectively.
17. The method of any one of claims 1, 2, and 5-16 or the apparatus of any one of claims 3-16, wherein the vertical predictor and the horizontal predictor for the sample correspond to a vertical plane mode and a horizontal plane mode, respectively.
18. The method of any one of claims 1, 2, and 5-16, or the apparatus of any one of claims 3-16, wherein the horizontal predictor and the vertical predictor for the sample correspond to the average value of the upper reference sample and the average value of the left reference sample, respectively.
19. The method of any one of claims 1, 2, and 5-15 or the apparatus of any one of claims 3-15, wherein one of the vertical predictor and the horizontal predictor for the sample is obtained based on interpolation, and the other is obtained based on a copy of a reference sample.
20. A signal comprising a bit stream, formed by performing the method as described in any one of claims 2 and 5-19.
21. A computer-readable storage medium having instructions stored thereon for encoding or decoding video according to any one of claims 1, 2, and 5-19.