Method and apparatus for picture encoding and decoding using position dependent intra prediction combination

By introducing the Position-Related Intra-Prediction Combination (PDPC) post-processing stage into video encoding and decoding, the calculation process of intra-prediction is simplified, the complexity and block artifacts in intra-prediction are solved, and the encoding and decoding efficiency and quality are improved.

CN114600450BActive Publication Date: 2025-11-21INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080058711.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-20
Filing Date
2020-06-15
Publication Date
2025-11-21
Estimated Expiration
2040-06-15

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies suffer from computational complexity and block artifacts in intra-frame prediction, especially at high compression ratios, leading to low efficiency.

Method used

The position-dependent intra-prediction combination (PDPC) post-processing stage simplifies computation and improves prediction quality by modifying the sampled values ​​of intra-prediction based on the left or top reference samples.

Benefits of technology

It improves the computational efficiency of video encoding and decoding while maintaining the same compression performance, reduces block artifacts, and enhances encoding and decoding quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114600450B_ABST
    Figure CN114600450B_ABST
Patent Text Reader

Abstract

A video coding system includes a post-processing stage using position dependent intra prediction combination, in which a predicted sample is modified based on a weighting between values of left or top reference samples and a predicted value of the obtained sample, wherein the left or top reference samples are determined based on an intra prediction angle. This provides better computational efficiency while maintaining the same compression performance. An encoding method, a decoding method, an encoding apparatus and a decoding apparatus based on the post-processing stage are proposed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one of these embodiments generally relates to a video encoding / decoding system, and more specifically, to a post-processing stage for location-related intra-frame prediction combination. Background Technology

[0002] To achieve high compression efficiency, image and video codec schemes typically employ prediction and transform to leverage spatial and temporal redundancy in the video content. Generally, intra-frame or inter-frame prediction is used to utilize intra-frame or inter-frame correlations, and then the difference between the original image block and the predicted image block (usually represented as prediction error or prediction residual) is transformed, quantized, and entropy encoded / decoded. During encoding, the original image block is typically segmented / divided into sub-blocks (e.g., possibly using quadtree segmentation). To reconstruct the video, the compressed data is decoded through the inverse processing corresponding to prediction, transform, quantization, and entropy encoding / decoding. Summary of the Invention

[0003] A video encoding / decoding system includes a post-processing stage for position-related intra-frame prediction combination, wherein the predicted samples are modified based on a weighted average between the values ​​of a left or top reference sample and the predicted values ​​of the obtained samples, wherein the left or top reference sample is determined based on the intra-frame prediction angle. This provides better computational efficiency while maintaining the same compression performance. Encoding methods, decoding methods, encoding devices, and decoding devices based on this post-processing stage are proposed.

[0004] According to a first aspect of at least one embodiment, a method for determining the value of a sampled block of an image, the value being intra-predicted based on a value representing an intra-prediction angle, the method comprising obtaining the predicted value of the sampled block, and determining the value of the sampled block based on a weighted average between the value of a left or top reference sample and the predicted value of the obtained sampled block, when the intra-prediction angle matches a criterion, wherein the left or top reference sample is determined based on the intra-prediction angle.

[0005] According to a second aspect of at least one embodiment, a video coding method includes, for each sample of a block of video, performing intra-frame prediction on the sample, modifying the value of the sample according to a first aspect, and encoding the block.

[0006] According to a third aspect of at least one embodiment, a video decoding method includes performing intra-frame prediction on each sample of a block of video, and modifying the value of the sample according to a first aspect.

[0007] According to a fourth aspect of at least one embodiment, a video encoding apparatus includes an encoder configured to perform intra-frame prediction on each sample of a block of video, modify the value of the sample according to a first aspect, and encode the block.

[0008] According to a fifth aspect of at least one embodiment, a video decoding apparatus includes a decoder configured to perform intra-frame prediction on each sample of a block of video, and modify the value of the sample according to a first aspect.

[0009] One or more embodiments of this embodiment also provide a non-transitory computer-readable storage medium storing instructions thereon for encoding or decoding video data according to at least a portion of any of the methods described above. One or more embodiments also provide a computer program product including instructions for performing at least a portion of any of the methods described above. Attached Figure Description

[0010] Figure 1A A block diagram of a video encoder according to one embodiment is shown.

[0011] Figure 1B A block diagram of a video decoder according to one embodiment is shown.

[0012] Figure 2 A block diagram illustrating an example of a system in which various aspects and embodiments are implemented is shown.

[0013] Figure 3A and 3B The symbols associated with intra-frame prediction of angles are shown.

[0014] Figure 4 The position-related intra-frame prediction combination of the lower left mode (mode 66) is shown.

[0015] Figure 5A and 5B The angular patterns adjacent to pattern 66 and pattern 2 are shown respectively.

[0016] Figure 6 An example of the angular mode of PDPC is shown, where there is a mismatch between the projection of the direction and the reference pixel.

[0017] Figure 7 Example flowcharts of PDPC processing for diagonal (modes 2 and 66) and adjacent diagonal modes are shown according to the example implementation of VTM 5.0.

[0018] Figure 8 An example of a PDPC in vertical diagonal mode 66 is shown, for which the left-side reference sample is not available for the target pixel.

[0019] Figure 9 An example flowchart of a first embodiment of the modified PDPC processing is shown.

[0020] Figure 10 An example flowchart of a second embodiment of the modified PDPC processing is shown.

[0021] Figure 11 The results of implementing the first embodiment are shown.

[0022] Figure 12 The results of implementing the second embodiment are shown.

[0023] Figure 13 Example flowcharts are shown according to various embodiments described in this disclosure. Detailed Implementation

[0024] Various embodiments relate to a post-processing method for predicting values ​​of samples of an image patch, the values ​​being predicted based on an intra-frame prediction angle, wherein the sampled values ​​are modified after the prediction such that they are determined based on a weighted sum of the difference between a left-side reference sample and the obtained predicted value of the sample, wherein the left-side reference sample is determined based on the intra-frame prediction angle. Encoding methods, decoding methods, encoding apparatuses, and decoding apparatuses based on this post-processing method are proposed.

[0025] Furthermore, although principles relating to specific drafts of the VVC (Video Common Coding) or HEVC (High-Efficiency Video Coding) specifications are described, this aspect is not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations, whether existing or future-developed, as well as any extensions of such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically excluded, the aspects described in this application may be used individually or in combination.

[0026] Figure 1A A video encoder 100 is shown. Variations of this encoder 100 are anticipated, but for clarity, encoder 100 is described below without describing all anticipated variations. Before encoding, the video sequence may undergo pre-coding (101), for example, applying color transformations to the input color picture (e.g., from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components to obtain a more compression-resistant signal distribution (e.g., using histogram equalization of one of the color components). Metadata can be associated with pre-processing and appended to the bitstream.

[0027] In encoder 100, the frame is encoded by encoder elements as described below. For example, the frame to be encoded is segmented (102) and processed in units of CUs. Each unit is encoded using, for example, an intra-frame or inter-frame mode. When a unit is encoded in intra-frame mode, intra-frame prediction (160) is performed. In inter-frame mode, motion estimation (175) and compensation (170) are performed. The encoder determines (105) which of the intra-frame or inter-frame modes to use for encoding the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (110) the prediction block from the original image block.

[0028] The predicted residual is then transformed (125) and quantized (130). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (145) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass the transform and quantization, i.e., encode and decode the residual directly without applying transform or quantization.

[0029] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inversely transformed (150) to decode the prediction residuals. The image blocks are reconstructed by combining (155) the decoded prediction residuals and the prediction blocks. An in-loop filter (165) is applied to the reconstructed image to perform filtering such as deblocking / SAO (Sampling Adaptive Offset) and Adaptive Loop Filter (ALF) to reduce coding artifacts. The filtered image is stored in a reference image buffer (180).

[0030] Figure 1B A block diagram of a video decoder 200 is shown. In decoder 200, the bitstream is decoded by decoder elements, as described below. The video decoder 200 typically performs... Figure 1AThe encoding traversal described herein is the reverse (reciprocal) decoding traversal. Encoder 100 typically also performs video decoding as part of the encoded video data. Specifically, the decoder input includes a video bitstream, which can be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other encoding / decoding information. The frame segmentation information indicates how the frame should be segmented. Therefore, the decoder can segment the frame based on the decoded frame segmentation information (235). The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The image blocks are reconstructed by combining (255) the decoded prediction residuals and prediction blocks. The prediction blocks (270) can be obtained from intra-frame prediction (260) or motion-compensated prediction (i.e., inter-frame prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference frame buffer (280).

[0031] The decoded image can undergo further post-decoding processing (285), such as inverse color transformation (e.g., from YCbCr4:2:0 to RGB 4:4:4), or perform an inverse remapping of the remapping process performed in the pre-encoding process (101). The post-decoding process can use metadata derived from the pre-encoding process and signaled in the bitstream.

[0032] Figure 2 A block diagram illustrating an example system in which various aspects and embodiments are implemented is shown. System 1000 can be implemented as a device including the various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1000 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described in this document.

[0033] System 1000 includes at least one processor 1010 configured to execute instructions loaded therein for implementing aspects such as those described in this document. Processor 1010 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 1000 includes at least one memory 1020 (e.g., volatile and / or non-volatile memory). System 1000 includes a storage device 1040 that may include non-volatile and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 1040 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0034] System 1000 includes an encoder / decoder module 1030 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents multiple modules that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 1030 may be implemented as a separate element of system 1000, or may be incorporated within processor 1010 as a combination of hardware and software known to those skilled in the art.

[0035] Program code to be loaded onto processor 1010 or encoder / decoder 1030 to execute the various aspects described in this document may be stored in storage device 1040 and subsequently loaded onto memory 1020 for execution by processor 1010. According to various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 may store one or more of various items during the execution of the processing described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0036] In some embodiments, the memory within processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 1010 or encoder / decoder module 1030) may be used for one or more of these functions. External memory may be memory 1020 and / or storage device 1040, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, while 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Codec; also known as H.265 and MPEG-H Part 2), or VVC (Universal Video Codec; a new standard being developed by the Joint Video Experts Group JVET).

[0037] As shown in block 1130, inputs to the components of system 1000 can be provided through various input devices. Such input devices include, but are not limited to: (i) an RF section that receives radio frequency (RF) signals, for example, transmitted over the air by a broadcaster; (ii) a component input (or a set of COMP inputs); (iii) a universal serial bus (USB) input; and / or (iv) a high-definition multimedia interface (HDMI) input. Figure 2 Other examples not shown include composite video.

[0038] In various embodiments, the input device of block 1130 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select, for example, a signal band (which may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a stream of desired data packets. The RF section in various embodiments includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners performing various of these functions, including, for example, down-converting received signals to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0039] Furthermore, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 1000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, can be implemented as needed, for example, within a separate input processing IC or processor 1010. Similarly, various aspects of USB or HDMI interface processing can be implemented as needed, within a separate interface IC or processor 1010. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010 and encoder / decoder 1030, which operate in combination with memory and storage elements to process the data stream as needed for presentation on the output device.

[0040] Various components of system 1000 can be provided within an integrated housing. Within the integrated housing, various components can be interconnected and transmit data therebetween using a suitable connection arrangement 1140 (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).

[0041] System 1000 includes a communication interface 1050 for communicating with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 1060 may be implemented within, for example, wired and / or wireless media.

[0042] In various embodiments, a wireless network such as Wi-Fi (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to system 1000. In these embodiments, the Wi-Fi signal is received via a communication channel 1060 and a communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, for applications that allow streaming and other over-the-top communication. Other embodiments use a set-top box to provide streaming data to system 1000, with the set-top box transmitting data via an HDMI connection of input block 1130. Still other embodiments use an RF connection of input block 1130 to provide streaming data to system 1000. As described above, various embodiments provide data in a non-streaming manner. Furthermore, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0043] System 1000 can provide output signals to various output devices, including a display 1100, a speaker 1110, and other peripheral devices 1120. The display 1100 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a flexible display, and / or a foldable display. The display 1100 can be used in a television, tablet computer, laptop computer, cellular phone (mobile phone), or other device. The display 1100 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 1120 include one or more of a stand-alone digital video disc (or digital multifunction disc) (DVR for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments utilize one or more peripheral devices 1120 that provide functionality based on the output of system 1000. For example, a disc player performs the function of playing the output of system 1000.

[0044] In various embodiments, signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention is used to communicate control signals between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120. Output devices can be communicatively coupled to system 1000 via dedicated connections through corresponding interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to system 1000 via communication interface 1050 using communication channel 1060. In electronic devices (such as, for example, televisions), display 1100 and speaker 1110 can be integrated into a single unit with other components of system 1000. In various embodiments, display interface 1070 includes a display driver, such as, for example, a timing controller (TCon) chip.

[0045] Display 1100 and speaker 1110 may alternatively be separated from one or more other components, for example, if the RF portion of input 1130 is part of a separate set-top box. In various embodiments where display 1100 and speaker 1110 are external components, output signals may be provided via dedicated output connections including, for example, HDMI ports, USB ports, or COMP outputs.

[0046] These embodiments can be implemented by computer software implemented by processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, embodiments can be implemented by one or more integrated circuits. As a non-limiting example, memory 1020 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical storage devices, magnetic storage devices, semiconductor-based storage devices, fixed memory, and removable memory. As a non-limiting example, processor 1010 can be of any type suitable for the technical environment and can include one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0047] In Universal Video Coding (VVC), intra-frame prediction is applied to intra-blocks in all intraframes and interframes, where the target block (called a codec unit (CU)) is spatially predicted from causal neighbor blocks within the same frame (i.e., the top and upper right blocks, the left and lower left blocks, and the upper left block). Based on the decoded pixel values ​​in these blocks, the encoder constructs different predictions for the target block and selects the prediction that yields the best rate distortion (RD) performance. Predictions are tested against 67 prediction modes, including one PLANA mode (indexed as mode 0), one DC mode (indexed as mode 1), and 65 angle modes. For rectangular blocks, angle modes can include only the regular angle prediction modes from mode 2 to mode 66, or they can include wide-angle modes defined outside the regular angle range of 45 degrees to -135 degrees.

[0048] Depending on the prediction mode, the predicted samples can also undergo post-processing stages, such as Position-Related Intra-Prediction Combination (PDPC). PDPC aims to smooth out discontinuities at block boundaries in certain prediction modes to improve the prediction of target blocks. DC and PLANA prediction modes, as well as several angular prediction modes (such as strictly vertical (mode 50), strictly horizontal (mode 18), diagonal mode 2 and VDIA_IDX (mode 66), and all other positive angular modes (including wide-angle modes)), can cause the predicted value on one side of the target block to differ significantly from the reference sample on the adjacent reference array. Without any type of post-filtering, the subsequent residual signal will cause block artifacts, especially at higher QP values. The purpose of PDPC is to prevent these block artifacts by smoothing the prediction at the target block boundaries in an elegant manner, making the intensity changes at the target block boundaries relatively gradual. However, PDPC achieves this with significant complexity. The predicted samples are combined with reference samples from the top and left reference arrays as a weighted average, which involves multiplication and clipping.

[0049] For a given target block to be intra-predicted, the encoder or decoder first constructs two reference arrays (one at the top and one at the left). The reference samples are taken from the decoded samples in the top, top-right, left, bottom-left, and top-left decoded blocks. If some of the top or left samples are unavailable, due to the corresponding CU not being in the same slice, or the current CU being located at a frame boundary, a method called reference sample replacement is performed, in which the lost samples are copied clockwise from the available samples. The reference samples are then filtered using a specified filter according to the current CU size and prediction mode.

[0050] For generality, we will assume a rectangular target block with a width of W pixels and a height of H pixels. We will denote the top and left reference arrays as refMain and refSide, respectively. The refMain array has 2*W+1 pixels and is indexed from refMain[0] to refMain[2W], where refMain[0] corresponds to the top-left reference pixel. Similarly, the refSide array has 2*H+1 pixels and is indexed from refSide[0] to refSide[2H], where refSide[0] again corresponds to the top-left reference pixel. For the special case of a square block with N pixels on each side, both reference arrays will have 2N+1 pixels. The horizontal prediction mode (i.e., the mode with an index less than the diagonal mode DIA_IDX, or mode 34) can be implemented by swapping the top and left reference arrays (again, the height and width of the target block). This is possible due to the symmetry of the vertical (horizontal) mode about the strictly vertical (horizontal) direction. Throughout this document, we will assume that, for horizontal prediction modes, refMain and refSide represent the top and left reference arrays after they have been swapped. Furthermore, since PDPC is only applied to some positive vertical and positive horizontal directions (besides the non-angular cases of PLANA and DC modes, and strictly vertical and strictly horizontal modes), we will limit our discussion in this disclosure to only positive prediction directions.

[0051] For any target pixel, the reference sample on refMain will be referred to as its predictor. For a given angular prediction pattern, the predictor sample on refMain is copied within the target PU along the corresponding direction. Some predictor samples may have integer positions, in which case they match the corresponding reference sample; other predictor positions will have a fractional part indicating that their position will fall between the two reference samples. In the latter case, the predictor samples are interpolated for the LUMA components using a 4-tap cubic or 4-tap Gaussian filter and linearly interpolated for the CHROMA components.

[0052] Mode 2 corresponds to a prediction direction from bottom left to top right at a 45-degree angle, while Mode 66 (also known as Mode VDIA_IDX) corresponds to the opposite direction. For both modes, all target pixels in the current block have their prediction factors at integer positions on refMain. That is, due to the 45-degree angle, each prediction sample is always consistent with a unique reference sample. For each target pixel, PDPC finds two reference pixels along the prediction direction, one on each reference array, and then combines their values ​​with the prediction value.

[0053] Figure 3Aand 3B The symbols associated with intra-frame prediction of angles are shown. Figure 3A The pattern number is shown, while Figure 3B The corresponding intraPredAngle value is shown.

[0054] Figure 4 The diagram illustrates the position-dependent intra-prediction combination (PDPC) for the lower left mode (mode 66). The top and left reference samples are located diagonally opposite each other. Target pixels where the left reference sample is unavailable are not subject to PDPC. It shows the two reference pixels for mode 66. Assuming the top and left references, as well as the height and width of the target block, have been swapped, the same diagram can be obtained for mode 2.

[0055] First, for mode 66, let's assume the target pixel P has coordinates (x, y), where 0 ≤ x < W and 0 ≤ y < H. PDPC finds reference pixels on refMain and refSide by extending the prediction direction. Let P top and P left This represents the values ​​of the top and left reference pixels corresponding to the target pixel P. Therefore:

[0056] P top =refMain[c+1];

[0057] P left =refSide[c+1];

[0058] Where c = 1 + x + y. If P pred (x, y) represents the predicted value, and the final predicted value at pixel P is obtained as follows:

[0059] P(x, y) = Clip(((64-wT-wL)*P pred (x, y) + wT*P top +wL*P left +32)>>6), (1)

[0060] The weights wT and wL are calculated as follows:

[0061] wT=16>>min(31, ((y<<1)>>scale));

[0062] wL=16>>min(31, ((x<<1)>>scale));

[0063] Where scale is pre-calculated as

[0064] scale=((log2(W)-2+log2(H)-2+2)>>2)

[0065] For a target pixel, if the second reference sample on the refSide is outside its length, the pixel is considered unavailable. In this case, the target pixel does not undergo the above modifications, and the first predicted value of the pixel remains unchanged. Figure 4 The gray pixels in the lower right corner of the target block that did not undergo PDPC are shown.

[0066] For mode 2, once the top and left reference arrays have been swapped along with the height and width of the target block, the PDPC processing is exactly the same as in mode 66.

[0067] Figure 5A The angular patterns adjacent to pattern 66 are shown, and Figure 5B The angled patterns adjacent to mode 2 are shown. In fact, the angled patterns adjacent to modes 2 and 66 can also benefit from PDPC. Specifically, the eight regular patterns adjacent to each of modes 2 and 66 are considered for PDPC. Furthermore, all wide-angle patterns are also considered for PDPC. Therefore, as... Figure 4 As shown, for the exemplary flat rectangular block with W = 2 * H, patterns 58 to 65 and wide-angle patterns 67 to 72 undergo PDPC. Similarly, for the exemplary tall rectangular block with H = 2 * W, as Figure 5B As shown, modes 3 to 10 and wide-angle modes -1 to -6 undergo PDPC. For a target block, the number of applicable wide-angle modes is a function of its aspect ratio (i.e., the ratio of width to height). In some implementations, all valid wide-angle modes can be considered for PDPC regardless of the target block's shape. Figure 4 and Figures 5A-5B The scenario shown is such that extreme angle modes 72 and -6 are effective because the required reference pixels are available. For rectangular blocks with W = 2H or H = 2W, the number of wide angle modes is 6.

[0068] Figure 6 An example of the PDPC's angular pattern is shown, where there is a mismatch between the projection of the direction and the reference pixel. In fact, in this example of pattern 68, the left reference pixel is not at the integer sampling position, necessitating the selection of the nearest reference pixel. More precisely, the nearest neighbor of the intersection point is selected as the left reference pixel.

[0069] In VVC, the index of the reference pixel is calculated as follows:

[0070] Let Δ x This indicates the horizontal displacement of the first reference sample from the target pixel position: Where A is the angle parameter intraPredAngle, which depends on the prediction mode index specified by VVC, as shown in Table 1 below. Positive prediction directions have positive A values, and negative prediction directions have negative A values. Δ x It can have a fractional part.

[0071] Table 1 shows an example mapping of the pattern index to the angle parameter A in VVC for the vertical direction. The mapping for the horizontal direction can be derived by swapping the vertical and horizontal indices.

[0072] Table 1

[0073]

[0074]

[0075] A similar process is applied to derive the left-side reference pixel. Let invAngle represent the inverse angle parameter corresponding to the prediction mode considered for PDPC. The invAngle value depends on the mode parameter and can be derived from intraPredAngle. Table 2 shows an example mapping of exemplary mode indices to inverse angle parameters in the VVC for the vertical direction. The mapping for the horizontal direction can be derived by swapping the vertical and horizontal indices.

[0076] Table 2

[0077]

[0078] Let P left This represents the value of the left reference pixel corresponding to the target pixel P at coordinates (x, y). left The following was obtained:

[0079] deltaPos=((1+x)invAngle+2)>>2;

[0080] deltaInt = deltaPos >> 6;

[0081] deltaFrac = deltaPos & 63;

[0082] deltay = 1 + y + deltaInt;

[0083] P left =refSide[deltay+(deltaFrac>>5)]

[0084] In the above text, deltaPos represents P leftThe distance to the reference pixel refSide[1+y] has a resolution of (1 / 64). deltaInt is the integer part of deltaPos with a resolution of 1, and deltaFrac represents the remaining fractional part (with a resolution of (1 / 64)) (equivalent to deltaPos = (deltaInt << 6) + deltaFrac). When deltaFrac is less than 32, P left It is the smaller integer neighbor refSide[deltay]; and when deltaFrac is greater than or equal to 32, P left It is the larger integer neighbor refSide[deltay+1].

[0085] If P pred (x, y) represents the initial predicted value, then the final predicted value at pixel P is obtained as:

[0086] P(x, y) = Clip(((64-wL)*P pred (x, y) + wL*P left +32)>>6) (2)

[0087] The weight wL is calculated as: wL = 32 >> min(31, ((x << 1) >> scale)), and scale is pre-calculated as: scale = ((log2(W) - 2 + log2(H) - 2 + 2) >> 2).

[0088] For a target pixel, if the left reference sample on the refSide is outside its length, the pixel is considered unusable. In this case, the target pixel does not undergo the above modifications, and the first predicted value of the pixel remains unchanged. In other words, PDPC post-processing is not applied in this case.

[0089] For the case of a horizontal prediction pattern undergoing PDPC around mode 2, the processing is similar once the top and left reference arrays have been swapped along with the height and width of the target block.

[0090] It's important to note here that wL is a non-negative decreasing function of x. If for x = x... n If wL = 0, then for x > x n We have wL = 0. According to equation (2), we observe that if wL = 0, the left-side reference pixels have no effect on the sum of weights. In this case, no PDPC operation is needed for the target pixels under consideration. Due to the above properties of wL, there is also no need to apply PDPC to the remaining target pixels in the same row of the block. Therefore, once wL becomes 0, the current PDPC tool terminates the PDPC operation.

[0091] Figure 7 Example flowcharts of PDPC processing for diagonal (modes 2 and 66) and adjacent diagonal modes are shown according to an example implementation of VTM 5.0. The flowcharts do not include PDPC for PLANAR, DC, strictly vertical, and strictly horizontal intra-prediction modes. Since VTM 5.0 swaps the reference arrays refMain and refSide when the mode is horizontal (i.e., mode index < 34 but not equal to 0 or 1), when the mode index is equal to 2 or any of its adjacent modes, refMain represents the reference array to the left of the current CU, and refSide represents the reference array to the top of the current CU. Therefore, in these cases, the parameters wL, wT, and P... left P top These actually correspond to the parameters wT, wL, and P, respectively. top P left The process includes multiple tests (steps 710, 730) and multiple multiplications required in each calculation (steps 743 and 753). In this process, firstly, when intraPredAngle is not greater than or equal to 12, in the "No" branch of step 710, intra-frame prediction terminates and no PDPC post-processing is applied. When intraPredAngle equals 32, in the "Yes" branch of step 730, wT and wL are calculated in step 741, and P is obtained in step 742. top and P left Then, PDPC post-processing is applied in step 743. In step 744, the processing iterates over the next pixel until all rows have been processed. When intraPredAngle is not equal to 32, in the "No" branch of step 730, wL is calculated in step 751, and P is obtained in step 752. left Then, in step 753, PDPC post-processing is applied. In step 754, the process iterates over the next pixel until all rows have been processed.

[0092] The embodiments described below have been designed with the foregoing in mind.

[0093] Figure 1A Encoder 100, Figure 1B Decoder 200 and Figure 2 The system 1000 is adapted to implement at least one of the following embodiments.

[0094] In at least one embodiment, the video encoding / decoding or decoding includes a post-processing stage for location-related intra-frame prediction combination, wherein computations are simplified to provide better computational efficiency while maintaining the same compression performance.

[0095] First simplification

[0096] refer to Figure 4 The predicted value for any target pixel is equal to the top reference pixel of that pixel. That is, P pred equals P top P top Substituting the value of into equation (1), we get:

[0097] P(x, y) = Clip(((64-wT-wL)*P pred (x, y) + wT*P pred (x, y) + wL*P left +32)>>6)

[0098] Cancel wT*P pred After adding the term (x, y), the above expression is simplified to

[0099] P(x, y) = Clip(((64-wL)*P pred (x, y) + wL*P left +32)>>6) (3)

[0100] Where wL = 16 >> min(31, ((x << 1) >> scale));

[0101] Except for the value of the weight wL, the expression is the same as equation (2). As we have seen, it is not necessary to calculate wT or P. top Therefore, it is unnecessary to calculate the term wT*P. top Due to this simplification, it is possible to merge the PDPC cases of Mode 2 and Mode 66 with the PDPC cases of other angle modes. In this case:

[0102] wL=wLmax>>min(31, ((x<<1)>>scale));

[0103] wLmax = 16 if predMode = 2 or predMode = 66;

[0104] =32, others

[0105] Since predMode equals 2 or 66 if and only if the angle parameter intraPredAngle equals 32, the above expression for wL can be restated as:

[0106] wL=wLmax>>min(31, ((x<<1)>>scale));

[0107] wLmax = 16 if intraPredAngle = 32;

[0108] =32, others

[0109] It should also be noted that even if the two cases are combined, the position of the left-hand reference sample in mode 2 or 66 will not change. The invAngle value for modes 2 and 66 can be equal to 256 because... By using this value, we obtain

[0110] deltaPos=((1+x)*256+2)>>2=(1+x)*64;

[0111] deltaInt=deltaPos>>6=(1+x);

[0112] deltaFrac=deltaPos&63=0;

[0113] deltay=1+y+deltaInt=1+y+1+x;

[0114] P left =refSide[1+y+1+x]=refSide[c+1] where c=1+x+y;

[0115] As we can see, this is the same value calculated previously in modes 2 and 66. Therefore, the above merging will result in the same outcome as the traditional implementation, but it is more efficient in terms of computational requirements because it requires fewer operations on the equation.

[0116] In fact, the deltaFrac term doesn't really serve any purpose. In the original PDPC proposal, deltaFrac was included because the position of the left-hand reference sample was linearly interpolated based on the value of deltaFrac. In another proposal, linear interpolation was replaced by nearest neighbor. However, for this point, calculating the deltaFrac term is quite redundant, as we will show in the next embodiment.

[0117] Second simplification

[0118] As calculated above, the deltay term causes the left reference sample to potentially occupy the smaller of the adjacent integer pairs. The value of deltaFrac determines the mapping from the left reference pixel to refSide[deltay] or refSide[deltay+1], depending on whether deltaFrac < 32 or whether deltaFrac >= 32, respectively. This process negatively impacts the checking of the availability of the left reference pixel. As previously mentioned, PDPC does not apply to target pixels where the left reference pixel is located outside the length of the left reference array. In this case, the left reference pixel is considered unavailable, and the initial predicted value is not modified.

[0119] Figure 8 An example of a PDPC in vertical diagonal mode 66 is shown, for which the left-side reference sample is not available for the target pixel.

[0120] Recall that the length of the left reference array refSide is 2H+1, with the first sample at index 0 (refSide[0]) and the last sample at index 2H (refSide[2H]). Using PDPC, if deltay > (2*H-1), deltay is checked to determine if the left sample is unavailable. This formula allows interpolation of the left pixels (if deltaFrac is not empty). For interpolation, a pair of refSide[deltay] and refSide[deltay+1] is needed. Here, since nearest neighbors are used instead of interpolation, the tool then determines one pixel in the pair based on the value of deltaFrac. However, in the PDPC cases of modes 2 and 66, since the left reference sample is located at an integer position (refSide[c+1] = refSide[2+x+y]), the tool checks c >= 2*H to determine if the left reference sample is unavailable. This introduces some ambiguity.

[0121] As we have seen before, the derivation of the left-side reference sample involves calculating the distance deltaPos: deltaPos = ((1+x)invAngle+2)>>2;

[0122] The terms on the right-hand side are simply recursive sums. Therefore, deltaPos can be calculated as follows, removing the multiplication with x:

[0123] invAngleSum[-1] = 2;

[0124] For 0 <= x < W

[0125] invAngleSum[x]=invAngleSum[x-1]+invAngle;

[0126] deltaPos=invAngleSum[x]>>2;

[0127] deltaInt = deltaPos >> 6;

[0128] deltaFrac = deltaPos & 63;

[0129] deltay = 1 + y + deltaInt;

[0130] P left=refSide[deltay+(deltaFrac>>5)]

[0131] In at least one embodiment, we propose modifying the process as follows:

[0132] invAngleSum[-1] = 128;

[0133] For 0 <= x < W

[0134] invAngleSum[x]=invAngleSum[x-1]+invAngle;

[0135] deltay=1+y+(invAngleSum[x]>>8),

[0136] P left =refSide[deltay]

[0137] Note that in the simplification above, we have combined the following three steps into one:

[0138] deltaPos=invAngleSum[x]>>2;

[0139] deltaInt = deltaPos >> 6;

[0140] deltay = 1 + y + deltaInt;

[0141] We initialize invAngleSum with 128 instead of 2, ensuring that the shift operation (invAngleSum >> 8) produces the closest integer due to rounding. This avoids calculating the deltaFrac parameter. Therefore, we only need to check if deltay > 2 * H to know if the left reference pixel is unavailable. This is the same check performed when PDPC is applied to mode 2 or mode 66.

[0142] The side effect of this simplification is better utilization of the last reference pixel in refSide, i.e., refSide[2H]. If the actual location of the reference sample is between 2H and 2H+1, and closer to refSide[2H], the above processing maps the location to 2H and uses pixel refSide[2H] as P. left In traditional PDPC implementations, this situation can be simply skipped based on deltay > 2H, thus the initial predicted value for the current pixel will remain unchanged. Aside from this special case, the proposed simplification will result in the same predicted value for the target pixel as in VTM 5.0.

[0143] Therefore, by observing the sum of weights without the need for clipping, equation (3) can be simplified:

[0144] P(x, y) = ((64-wL)*P pred (x, y) + wL*P left +32)>>6

[0145] For 0 ≤ x < W, 0 ≤ y < H;

[0146] This equation can be further simplified to

[0147] P(x, y) = (64 * P pred (x, y) - wL*P pred (x, y) + wL*P left +32)>>6

[0148] =P pred (x, y) + ((wL*(P) left -P pred (x, y))+32)>>6).

[0149] This simplification avoids the multiplication with (64-wL), which is not necessarily a power of 2. As we have seen, wL only has non-negative values. If wL > 0, then it is also a power of 2. Therefore, the only multiplication in the above equation can be performed using the shift shown in equation (4).

[0150] In at least one embodiment, the video encoding / decoding or decoding includes a post-processing stage using PDPC, wherein the prediction calculation uses the following equation (4):

[0151] P(x, y) = P pred (x, y) + ((((P) left -P pred (x, y))<<log2(wL))+32)>>6) (4)

[0152] Those skilled in the art will note that the above simplified proposal corresponds to VTM 5.0 code with an iterative loop and initialization of invAngleSum outside the loop to avoid multiplication. It can also be equivalently written as follows:

[0153] For 0 <= x < W

[0154] deltay=1+y+(((1+x)*invAngle+128)>>8)

[0155] P left =refSide[deltay]

[0156] Note that the refSide reference array has its 0th coordinate at the top-left pixel (i.e., at the reconstructed sample coordinates (-1, -1)). If it were initialized otherwise at coordinates (-1, 0), the above equation would be modified to:

[0157] For 0 <= x < W

[0158] deltay=y+(((1+x)*invAngle+128)>>8)

[0159] P left =refSide[deltay]

[0160] Therefore, through the above coordinate adjustments, the proposed PDPC for vertical mode 66 and its adjacent modes is simplified as follows:

[0161] For 0 <= x < W

[0162] deltay=y+(((1+x)*invAngle+128)>>8)

[0163] P left =refSide[deltay]

[0164] P(x, y) = P pred (x, y) + ((wL*(P) left -P pred (x, y) + 32) >> 6)

[0165] wL=wLmax>>min(31, ((x<<1)>>scale))

[0166] wLmax = 16 if mode = 66, otherwise wLmax = 32

[0167] Similarly, for horizontal prediction mode 2 and its neighboring modes, the proposed PDPC modifications are as follows:

[0168] For 0 <= y < H

[0169] deltax=x+(((1+y)*invAngle+128)>>8)

[0170] P top =refMain[deltax]

[0171] P(x, y) = P pred (x, y) + ((wT*(P) top -P pred (x, y) + 32) >> 6)

[0172] wT=wTmax>>min(31, ((y<<1)>>scale))

[0173] wTmax = 16 if mode = 2, otherwise wTmax = 32

[0174] Where refMain represents the reference array at the top of the target block, and P top This represents the reference pixel on the target block obtained by intersecting with the extension of the predicted direction.

[0175] Figure 9 An example flowchart of a first embodiment of modified PDPC processing is shown, following an example implementation based on VTM 5.0. In this modified processing, video encoding / decoding or decoding includes a post-processing stage using PDPC in angular mode, where two distinct prediction calculations are merged, and at least one pixel of a frame block is reconstructed based on its predicted value and a reference value in neighboring blocks weighted according to angular direction, as shown and described in the first simplification above. Figure 7 As shown in the right branch, when the prediction mode is horizontal, refSide represents the reference array at the top of the current CU. Therefore, when the mode index is equal to 2 or any of its neighboring modes, the parameters wL, wLMax, and P... left These actually correspond to the parameters wT, wTMax, and P, respectively. top wait.

[0176] In this embodiment, post-processing includes handling cases where the angle parameter intraPredAngle is greater than 12 (thus corresponding to a mode index greater than or equal to 58 or less than or equal to 10, as shown in Table 1). Figure 3A and 3B When (as shown), in the "Yes" branch of step 910, for a pixel, the weighted value wL is calculated in step 951, the left reference pixel is obtained in step 952 by extending the prediction direction to the left reference array, and the post-processing value of the predicted pixel is calculated in step 953 according to equation (3):

[0177] P(x, y) = Clip(((64-wL)*P pred (x, y) + wL*P left +32)>>6) (3)

[0178] In other words, the predicted pixel is modified based on a weighted average between its predicted value and a reference pixel selected from the column to the left of the block based on the angle prediction mode.

[0179] Note that the test performed in step 201 can be performed on the pattern index value instead of the intraPredAngle value. However, in this case, two or more comparisons will be required to check whether the predicted angle is within the range of values, where the range can correspond to either the vertical pattern set or the horizontal pattern set. Although in this embodiment, intraPredAngle is tested relative to a fixed value of 12 corresponding to pattern indices 10 and 58, other values ​​can be used for comparison, and these values ​​are compatible with the principles of this first embodiment. For example, in another embodiment, intraPredAngle is tested relative to a fixed value of 0 corresponding to pattern indices 18 and 50. In another embodiment, intraPredAngle is tested relative to a fixed value within the range of 0 to 12 corresponding to pattern indices between 10 and 18 and between 50 and 58.

[0180] Figure 10 An example flowchart of a second embodiment of modified PDPC processing is shown, following an example implementation based on VTM 5.0. In this second embodiment, video encoding / decoding or decoding includes a post-processing stage using PDPC in angle mode, where two different prediction calculations are merged, and at least one pixel of a frame block is reconstructed based on its predicted value and a reference value in neighboring blocks weighted according to the angle direction, such as... Figure 9 As shown and described above as a second simplification. Figure 9 As shown, when the prediction mode is horizontal, refSide represents the reference array at the top of the current CU. Therefore, when the mode index is equal to 2 or any of its neighboring modes, the parameters wL, wLMax, and P... lefi These actually correspond to the parameters wT, wTMax, and P, respectively. top wait.

[0181] In this embodiment, the post-processing includes, when the angle parameter intraPredAngle is greater than a predetermined value (e.g., 12), calculating a weighted value wL for a pixel in step 1003, obtaining a left reference pixel by extending the prediction direction to the left reference array in step 1004, and calculating the post-processed value of the predicted pixel according to equation (4) in step 1005:

[0182] P(x, y) = P pred (x, y) + ((((P) left -P pred (x, y))<<log2(wL))+32)>>6) (4)

[0183] In other words, the predicted pixel is modified based on a weighted average between its predicted value and a reference pixel selected from the column to the left of the block containing that pixel, based on the angle prediction pattern.

[0184] Traditional PDPC uses multiple checks and multiple multiplications, while the second embodiment uses only a single check and shift, addition, and subtraction operations, which are much more efficient than multiplication operations in terms of computational requirements.

[0185] Although in this embodiment, intraPredAngle is tested relative to a fixed value of 12, other values ​​for comparison can be used, and these values ​​are compatible with the principles of this first embodiment.

[0186] Third Embodiment

[0187] The third embodiment is based on the first or second embodiment, wherein for the target pixel at (x, y), the left reference sample is always refSide[c+1], where c = 1 + x + y, and is independent of the value of intraPredAngle.

[0188] Fourth embodiment

[0189] The fourth embodiment is based on one of the above embodiments, wherein the position of the left reference sample exceeds the length of refSide, and we use the last reference sample refSide[2H] instead of terminating PDPC at the current target pixel on the row.

[0190] Fifth Embodiment

[0191] The fifth embodiment is based on one of the above embodiments, wherein, in addition to the four fixed modes (PLANAR (mode 0), DC (mode 1), strictly vertical (mode 50), and strictly horizontal (mode 18)), all positive vertical and horizontal directions are considered as qualified modes of PDPC.

[0192] Sixth Embodiment

[0193] The sixth embodiment is based on one of the above embodiments, wherein any other non-negative and decreasing function is used to derive the weight wL. As an example, the weights can be derived based on the prediction direction as follows:

[0194] wL=wLmax>>((x<<1)>>scale);

[0195] wLmax=16*((intraPredAngle+32)>>5)

[0196] Seventh Embodiment

[0197] The seventh embodiment is based on one of the above embodiments, wherein the RD performance of the PDPC for the qualified mode is checked for both presence and absence. The use of the PDPC is signaled to the decoder at the CU level as a one-bit PDPC flag. This flag can be context-encoded, where the context can be a fixed value or can be derived from the neighborhood, prediction direction, etc.

[0198] Eighth embodiment

[0199] The eighth embodiment is based on one of the above embodiments, wherein PDPC is performed on all CUs in the stripe, and a one-bit flag in the stripe header is used to signal the decoder to the application of such PDPC.

[0200] Ninth Embodiment

[0201] The ninth embodiment is based on one of the above embodiments, wherein PDPC is performed on all CUs in the strip, and a one-bit flag in the Picture Parameter Set (PPS) header is used to signal the decoder to the application of such PDPC.

[0202] Tenth Embodiment

[0203] The tenth embodiment is based on one of the above embodiments, wherein PDPC is performed on any frame of the sequence, and a one-bit flag in the Sequence Parameter Set (SPS) header is used to signal to the decoder the application of such PDPC.

[0204] result

[0205] Experiments have been conducted with the VTM 5.0 codec in an intra-frame (AI) configuration with all the required test conditions.

[0206] Figure 11 The results of implementing the first embodiment are shown, and Figure 12 Results of the implementation of the second embodiment are shown. These figures illustrate the BD rate performance of the proposed simplified scheme compared to the VTM 5.0 anchor. In these tables, rows represent different categories of content to be encoded or decoded corresponding to samples of the original video content. The first column lists these categories. Columns 2, 3, and 4 (named Y, U, and V, respectively) show the size differences between the VTM 5.0 codec and the proposed embodiments for the Y, U, and V components of the original video, respectively. Therefore, these columns show the difference in compression efficiency compared to VTM 5.0. Columns 5 and 6 (named EncT and DecT, respectively) show the differences in encoding and decoding times relative to the encoding and decoding times of VTM 5.0, respectively. Comparing the encoding and decoding times compared to the VTM 5.0 implementation indicates the compression efficiency of the proposed embodiment. Figure 11 It shows an overall improvement in encoding and decoding time. Figure 12 The overall improvements regarding encoding and decoding time and encoding size are shown.

[0207] Figure 13 Example flowcharts according to various embodiments described in this disclosure are shown. This flowchart is activated when PDPC post-processing may occur and sampling of a sample block is iterated. First, in step 1301, information representing the prediction angle is tested and it is determined whether it matches a criterion. This angle corresponds to the angle used to perform intra-frame prediction and thus generate the current sample. This test can be performed based on an angle value, the value of the intra-predAngle parameter, or any value representing the prediction angle. In branch 1302, if the angle is incorrect, PDPC post-processing is not performed, and the predicted sample is not modified.

[0208] In at least one embodiment, the angle is considered correct when the intraPredAngle parameter is greater than or equal to 12. In a variant embodiment, the angle is considered correct when the intra-prediction mode is in the range of -6 to 10 or 58 to 72. Those skilled in the art will note that these two different matching criteria correspond to the same range of angles for intra-prediction.

[0209] Once the angle is deemed correct, a weighting is calculated in step 1303. This weighting will allow for balancing the correction amount between the sampled predicted value and the reference sample, thereby balancing the values ​​of the resulting samples.

[0210] In step 1305, a reference sample is obtained. In at least one embodiment, the reference sample is obtained from a left reference array containing the sampled values ​​of the left column of the block by extending the angular intra-frame prediction direction to the column to the left of the block and selecting the sample from the left reference array that is closest to the intersection between the angular intra-frame prediction direction and the column to the left of the block. The index of this array can be calculated in a simple manner, as previously mentioned under the name deltay.

[0211] In step 1307, the sampled values ​​are determined based on the obtained and calculated parameters. Equations 3 and 4 above provide different formulas for calculating the sampled values. These equations (and more specifically, equation 4) have the advantage of requiring less computation than conventional techniques. One reason is that reference samples from the top reference array are not considered, as they have already been taken into account to determine the predicted values ​​of the samples. At least another reason is that the fractional part is no longer considered, thus allowing for further computational simplification. Another reason is the elimination of an additional angle test. The determination of the sampled values ​​according to step 1307 can also be understood as a correction, modification, or post-processing of the previously predicted samples.

[0212] The operation in step 1301 can be performed once for each block. However, Figure 13Steps 1303 through 1307 need to be performed on each sample of the block, as the parameters involved depend on the location and / or value of the sample. Once all samples of the block have been processed, the block can be processed, for example, encoded.

[0213] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in detail and are generally described in a manner that may seem limiting, at least to illustrate their respective characteristics. However, this is for the purpose of clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in earlier documents.

[0214] The aspects described and anticipated in this application can be implemented in a variety of different forms. Figure 1A , 1B Figures 2 and 3 provide some embodiments, but other embodiments are contemplated, and the discussion of these figures does not limit the breadth of implementations. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having a bitstream generated according to any of the methods stored thereon.

[0215] This document describes various methods, each of which includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined.

[0216] The various methods and other aspects described in this application can be used to modify, for example... Figure 1A and Figure 1B The modules of the video encoder 100 and decoder 200 shown include, for example, motion compensation and motion estimation modules (170, 175, 275). Furthermore, this aspect is not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations, whether existing or future-developed, and any extensions of such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically excluded, the aspects described in this application can be used individually or in combination.

[0217] Various numerical values ​​are used in this application. These specific values ​​are for illustrative purposes, and the aspects described are not limited to these specific values.

[0218] Various implementations involve decoding. As used herein, "decoding" may include, for example, all or part of the processing performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, such processing includes one or more of the processing typically performed by a decoder. In various embodiments, such processing also includes, or alternatively includes, processing performed by the decoders of the various implementations described herein.

[0219] As a further example, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. It will be clear, and is considered fully understandable by those skilled in the art, whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to a broader decoding process, based on the context of the specific description.

[0220] Various implementations involve encoding. Similar to the discussion above regarding "decoding," the "encoding" used in this application may include, for example, all or part of the processing performed on the input video sequence to produce an encoded bitstream. In various embodiments, such processing includes one or more of the processing typically performed by an encoder. In various embodiments, such processing also includes, or alternatively includes, the processing performed by the encoders of the various implementations described in this application.

[0221] As a further example, in one embodiment, "encoding" refers only to entropy encoding; in another embodiment, "encoding" refers only to differential encoding; and in yet another embodiment, "encoding" refers to a combination of differential and entropy encoding. It will be clear, and is considered fully understandable by those skilled in the art, whether the phrase "encoding processing" is intended to specifically refer to a subset of operations or to refer to broader encoding processing, based on the context of the specific description.

[0222] Note that the grammatical elements used in this article are descriptive terms. Therefore, the use of other grammatical element names is not excluded.

[0223] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0224] Various embodiments involve rate distortion optimization. In particular, during encoding processing, a balance or trade-off between rate and distortion is typically considered, usually with constraints on computational complexity. Rate distortion optimization is generally formulated as minimizing a rate distortion function, which is a weighted sum of rate and distortion. Different approaches exist to address the rate distortion optimization problem. For example, these approaches can be based on extensive testing of all encoding options, including all considered modes or encoding / decoding parameter values, and a complete evaluation of their encoding / decoding costs and the associated distortion of the encoded / decoded and reconstructed signals. Faster methods can also be used to save encoding complexity, particularly by calculating approximate distortion based on predicting or predicting the residual signal rather than the reconstructed signal. A hybrid of these two approaches can also be used, such as by using approximate distortion only for some of the possible encoding options and full distortion for others. Other methods evaluate only a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding / decoding costs and associated distortion.

[0225] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in detail and are generally described in a manner that may seem limiting, at least to illustrate their respective characteristics. However, this is for the purpose of clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in earlier documents.

[0226] The implementations and aspects described herein can be implemented, for example, as methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the features under discussion can also be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, with appropriate hardware, software, and firmware. The methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, tablet computers, smartphones, mobile phones, portable / personal digital assistants, and other devices that facilitate information communication between end users.

[0227] References to “an embodiment” or “an embodiment” or “an implementation” or “an implementation”, and other variations thereof, mean that a particular feature, structure, characteristic, etc., described in connection with the embodiment is included in at least one embodiment. Therefore, the phrases “in an embodiment” or “in an embodiment” or “in an implementation” or “in an implementation” appearing throughout this application, and any other variations thereof, do not necessarily all refer to the same embodiment.

[0228] Furthermore, this application may involve "determining" various pieces of information. Determining information may include, for example, one or more of the following: estimated information, calculated information, predicted information, or information retrieved from memory.

[0229] Furthermore, this application may involve "accessing" various pieces of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0230] Furthermore, this application may relate to "receiving" various pieces of information. Like "access," "receiving" is intended as a broad term. Receiving information may include, for example, accessing information, or retrieving information (e.g., from memory) one or more of them. Moreover, "receiving" is typically referred to in one or more ways during operations such as, for example, storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0231] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "pixel" and "sample," and the terms "image," "picture," "frame," "strip," and "tile." Typically, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0232] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one of…” is intended to include selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” this wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many items as possible listed.

[0233] Furthermore, as used herein, the term "signal" refers, among other things, to indicating something to the corresponding decoder. For example, in some embodiments, the encoder signals a specific one of the illumination compensation parameters. In this way, in embodiments, the same parameter is used on both the encoder and decoder sides. Thus, for example, the encoder can send (explicit signaling) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be done in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the term "signal" has been referred to above, the term "signal" can also be used as a noun herein.

[0234] It will be apparent to those skilled in the art that various implementations can generate a variety of signals that are formatted to carry information, for example, that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the implementations. For example, the signal may be formatted to carry a bitstream of the embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.

[0235] In a first variation of the first aspect, an angle criterion is considered matched when the intra-frame predicted angle is less than or equal to a first value or greater than or equal to a second value, wherein the first value is less than the second value. In a variation of the first variation, the first value is 10 and the second value is 58.

[0236] In a second variation of the first aspect, the angle criterion is considered matched when the intraPredAngle value, representing the intra-predicted angle, is greater than or equal to a value of one. In a variation of the first variation, the angle criterion is considered matched when the intraPredAngle value, representing the intra-predicted angle, is greater than or equal to 12.

[0237] In another variation of the first aspect, for the vertical angular direction, the left reference sample is determined from a left reference array comprising sampled values ​​of the columns to the left of the block, which is determined by extending the vertical angular intra-prediction direction to the columns to the left of the block and selecting the sample from the left reference array that is closest to the intersection between the vertical angular intra-prediction direction and the columns to the left of the block; and for the horizontal angular direction, the top reference sample is determined from a top reference array comprising sampled values ​​of the rows to the top of the block, which is determined by extending the horizontal angular intra-prediction direction to the rows to the top of the block and selecting the sample from the top reference array that is closest to the intersection between the horizontal angular intra-prediction direction and the rows to the top of the block.

[0238] In another variation of the first aspect, for the vertical direction, according to P(x,y)=P pred (x, y) + ((((P) left -P pred (x, y)) << log2(wL)) + 32) >> 6) to determine the value of sampling P(x, y), and for the horizontal direction, according to P(x, y) = P pred (x, y) + ((((P) top -P pred (x, y)) << log2(wT)) + 32) >> 6) to determine the value of sampled P(x, y), where P pred (x, y) are the predicted values ​​of the obtained samples, P left or P top It is the value of the left reference sample or the top reference sample, and wL or wT is the weighted average.

[0239] In another variation of the first aspect, for the vertical angular direction, the left-side reference sample P left The value of P is determined by the following formula: left=refSide[deltay], where refSide is an array of sampled values ​​from the left column of the block, deltay = y + (((1+x)*invAngle+128)>>8), and for the horizontal angular direction, the top reference sample P top The value of P is determined by the following formula: top =refMain[deltax], where refMain is an array of sampled values ​​from the top row of the block, and deltax = x + (((1+y)*invAngle+128)>>8), where invAngle is the reciprocal of the angle parameter corresponding to the predicted angle.

[0240] In another variation of the first aspect, the weighting is determined according to the following:

[0241] wL = wLmax >> min(31, ((x << 1) >> scale)), where wLmax = 16 when intraPredAngle is 32, otherwise wLmax = 32, wT = wTmax >> min(31, ((y << 1) >> scale)), where wTmax = 16 when intraPredAngle is 32, otherwise wTmax = 32, and scale equals ((log2(W)-2+log2(H)-2+2) >> 2).

Claims

1. A method comprising: Obtain the predicted value P of the sampled patch of the image. pred (x, y), the predicted value is intra-predicted based on a value representing the intra-prediction angle, wherein the intra-prediction angle is a diagonal direction from the upper right to the lower left corresponding to the first prediction mode, the first prediction mode corresponding to prediction mode 66; and The sampled value P(x,y) is determined based on the following: P(x,y) = P pred (x,y) + ((((P left – P pred (x,y)) << log2(wL)) + 32) >> 6), where In the middle, P left It is determined based on the value of the left-side reference sample, and wL is weighted.

2. The method according to claim 1, wherein: For the vertical angle direction corresponding to the prediction mode 66, the left reference sample is determined from a left reference array, which includes the sampled values ​​of the left column of the block. The left reference sample is determined by extending the vertical angle intra-frame prediction direction to the left column of the block and selecting the sample from the left reference array that is closest to the intersection between the vertical angle intra-frame prediction direction and the left column of the block.

3. The method according to claim 1, wherein: The left reference sample P left The value is determined by the following formula: P left =refSide[deltay], where refSide is an array of sampled values ​​from the left column of the block, deltay = y + (((1+x)*invAngle+128)>>8), and invAngle is the reciprocal of the angle parameter corresponding to the predicted angle.

4. The method according to claim 1, wherein, The weighting is determined based on the following: wL = wLmax >> min(31, ((x << 1) >> scale)), where wLmax = 16. And where scale equals ((log2(W)–2+log2(H)-2+2)>>2), where W is the width of the block and H is the height of the block.

5. The method according to claim 1, wherein, The method is also applied to one of modes 58 to 72.

6. A method for encoding blocks of video, the method comprising: For each sample of the block of the video: Intra-frame prediction is performed on the sampled data. Modify the sampled value according to claim 1, and The block is encoded.

7. A method for decoding blocks of video, the method comprising: For each sample of the block of the video: Intra-frame prediction is performed on the sampled data. Modify the sampled value according to claim 1, and The block is decoded.

8. An apparatus including an encoder, the encoder being configured to, for blocks of video: For each sample of the block of the video: Intra-frame prediction is performed on the sampled data. Modify the sampled value according to claim 1, and The block is encoded.

9. An apparatus including a decoder, the decoder being configured to, for blocks of video: For each sample of the block of the video: Intra-frame prediction is performed on the sampled data. Modify the sampled value according to claim 1, and The block is decoded.

10. A non-transitory computer-readable medium comprising program code instructions that, when executed by a processor, are used to implement the steps of the method according to claim 1.