Interpolation filtering methods and apparatus for intra and inter prediction in video coding
Patent Information
- Application Number
- CN202210898128.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-07
- Filing Date
- 2019-09-06
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2039-09-06
AI Technical Summary
[0009]当前在BMS中描述的帧内模式译码方案被认为是复杂的,且非选择模式集的缺点在于索引列表总是恒定的,并且不能根据当前块属性(例如,对于其邻块帧内模式)进行调整
Smart Images

Figure CN115914624B_ABST
Abstract
Description
[0001] This application is a divisional application. The original application has the application number 201980046958.5 and the original application date is September 6, 2019. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This invention relates to the field of image and / or video encoding and decoding technology, and more specifically to interpolation filtering methods and apparatus for performing intra-frame prediction and inter-frame prediction. Background Technology
[0003] Since the advent of DVDs, digital video has been widely used. Video is encoded and then transmitted via a transmission medium. Viewers receive the video and use their viewing devices to decode and display it. Over the years, video quality has improved due to advancements in resolution, color depth, and frame rate. This has resulted in larger data streams, which are now typically transmitted via the internet and mobile communication networks.
[0004] However, higher resolution videos typically contain more information and therefore require more bandwidth. To reduce bandwidth requirements, video decoding standards involving video compression were introduced. When encoding video, the bandwidth requirement (or the corresponding memory requirement for storage) is reduced. This reduction often sacrifices quality. Therefore, video decoding standards attempt to find a balance between bandwidth requirements and quality.
[0005] High Efficiency Video Coding (HEVC) is an example of video decoding standards well-known to those skilled in the art. In HEVC, the coding unit (CU) is divided into a prediction unit (PU) or a transform unit (TU). Versatile Video Coding (VVC) is a recent joint video project between the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). These two standardization organizations collaborated in a partnership known as the Joint Video Exploration Team (JVET). VVC is also known as the ITU-T H.266 / Next Generation Video Coding (NGVC) standard. VVC eliminates the concept of multiple segmentation types, meaning it does not distinguish between CU, PU, and TU concepts (unless the CU size is too large for the maximum transform length) and supports more flexible CU segmentation shapes.
[0006] The processing of these coding units (CUs) (also called blocks) depends on their size, spatial location, and the coding mode specified by the encoder. Based on the type of prediction, coding modes can be divided into two categories: intra-frame prediction modes and inter-frame prediction modes. Intra-frame prediction modes use samples from the same picture / image (also called a frame) to generate reference samples to compute predicted values for the samples of the reconstructed block. Intra-frame prediction is also known as spatial prediction. Inter-frame prediction modes are designed for temporal prediction and use reference samples from previous, current (same), or subsequent images to predict samples for the blocks of the current image.
[0007] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) are studying the potential need for standardization of future video decoding technologies, which will significantly exceed the compression capabilities of the current HEVC standard (including current and recent extensions for screen content decoding and high dynamic range decoding). These groups are working with the Joint Video Exploration Team (JVET) on this exploration to evaluate compression technology designs proposed by their domain experts.
[0008] The Versatile Test Model (VTM) standard uses 35 intra-frame modes, while the Benchmark Set (BMS) uses 67 intra-frame modes.
[0009] The intra-mode decoding schemes currently described in BMS are considered complex, and the disadvantage of non-selective mode sets is that the index list is always constant and cannot be adjusted according to the current block attributes (e.g., for its neighboring block intra-modes). Summary of the Invention
[0010] This invention discloses an interpolation filtering method and apparatus for intra-frame prediction and inter-frame prediction. The apparatus and method employ the same sample interpolation process to unify the calculation flow for inter-frame prediction and intra-frame prediction, thereby improving decoding efficiency. The scope of protection is defined by the claims.
[0011] The above and other objectives are achieved through the subject matter claimed in the independent claims. Other implementations are apparent in the dependent claims, the specification, and the drawings.
[0012] Specific embodiments are set forth in the appended independent claims, and other embodiments are set forth in the dependent claims.
[0013] According to a first aspect, the present invention relates to a video decoding method, wherein the method comprises:
[0014] - Inter-frame prediction processing for the first block of an image or video, wherein the inter-frame prediction processing includes sub-pixel interpolation filtering of samples of a reference block (for fractional positions) (for the first block or for the first block);
[0015] - Intra-frame prediction processing for the second block of an image or video, wherein the intra-frame prediction processing includes sub-pixel interpolation filtering of a reference sample (for a fractional location) (for the second block or for the second block);
[0016] The method further includes:
[0017] - Based on the sub-pixel offset between the integer reference sample position and the fractional reference sample position, interpolation filter coefficients for the sub-pixel interpolation filter are selected, wherein for the same sub-pixel offset, the same interpolation filter coefficients are selected for intra-frame prediction processing and inter-frame prediction processing.
[0018] Subpixel interpolation filtering is performed on fractional (i.e., non-integer) reference sample positions because the corresponding values are typically not available from the decoded picture buffer (DPB), etc. Integer reference sample positions typically have values that can be obtained directly from the DPB, etc., thus eliminating the need for interpolation filtering. The method provided in the first aspect can also be referred to as an inter-frame prediction processing and intra-frame prediction processing method for video decoding, or a subpixel interpolation filtering method for inter-frame prediction processing and intra-frame prediction processing in video decoding.
[0019] In one implementation of the first aspect, the method may include, for example,: selecting a first set of interpolation filtering coefficients (e.g., c0 to c3) (e.g., for chroma samples) based on a first sub-pixel offset between an integer reference sample position and a fractional reference sample position, performing sub-pixel interpolation filtering for inter-frame prediction; and if there is a sub-pixel offset that is the same as the first sub-pixel offset, selecting the same first set of interpolation filtering coefficients (c0 to c3) (e.g., for luminance samples) for sub-pixel filtering for intra-frame prediction.
[0020] In one possible implementation of the method provided in the first aspect, the selected filtering coefficients are used to perform the sub-pixel interpolation filtering on the chroma samples for inter-frame prediction processing; the selected filtering coefficients are used to perform the sub-pixel interpolation filtering on the luminance samples for intra-frame prediction processing.
[0021] In one possible implementation of the method provided in the first aspect, the inter-frame prediction processing is intra-block copying processing.
[0022] In one possible implementation of the method provided in the first aspect, the interpolation filter coefficients used for inter-frame prediction processing and intra-frame prediction processing are obtained from a lookup table.
[0023] In one possible implementation of the method provided in the first aspect, a 4-tap filter is used for the sub-pixel interpolation filtering.
[0024] In one possible implementation of the method provided in the first aspect, selecting the interpolation filter coefficients includes: selecting the interpolation filter coefficients based on the following relationship between sub-pixel offsets and the interpolation filter coefficients:
[0025]
[0026] Wherein, the sub-pixel offset is defined with a 1 / 32 sub-pixel resolution, and c0 to c3 represent the interpolation filtering coefficients.
[0027] In one possible implementation of the method provided in the first aspect, selecting the interpolation filter coefficients includes: selecting the interpolation filter coefficients for fractional positions based on the following relationship between sub-pixel offsets and interpolation filter coefficients:
[0028]
[0029] Wherein, the sub-pixel offset is defined with a 1 / 32 sub-pixel resolution, and c0 to c3 represent the interpolation filtering coefficients.
[0030] According to a second aspect, the present invention relates to a video decoding method for obtaining predicted sample values of a current decoding block, wherein the method includes:
[0031] When the prediction sample of the current decoded block is obtained through the inter-frame prediction process, the following procedure (or steps) is executed to obtain the inter-frame prediction sample value.
[0032] The filtering coefficients are obtained from the lookup table based on the first sub-pixel offset value.
[0033] Based on the filtering coefficients, the inter-frame prediction sample values are obtained;
[0034] When the prediction sample of the current decoded block is obtained through the intra-frame prediction process, the following process (or steps) is executed to obtain the intra-frame prediction sample value.
[0035] The filtering coefficients are obtained from the lookup table based on the second sub-pixel offset value, wherein the lookup table used for inter-frame prediction is reused for intra-frame prediction.
[0036] Based on the filtering coefficients, the intra-frame prediction sample values are obtained.
[0037] As described in the first aspect, subpixel interpolation filtering is performed on fractional (i.e., non-integer) reference sample positions because the corresponding values are typically not available from the decoded picture buffer (DPB), etc. Integer reference sample positions typically have values that can be obtained directly from the DPB, etc., thus eliminating the need for interpolation filtering. The method provided in the second aspect can also be referred to as an inter-frame prediction processing and intra-frame prediction processing method for video decoding, or a subpixel interpolation filtering method for inter-frame prediction processing and intra-frame prediction processing in video decoding.
[0038] In one possible implementation of the method provided in the second aspect, filter coefficients from a lookup table are used in fractional sample position interpolation to perform an intra-frame prediction process or an inter-frame prediction process.
[0039] In one possible implementation of the method provided in the second aspect, the lookup table used in the intra-frame prediction process or used in the intra-frame prediction process is the same as the lookup table used in the inter-frame prediction process or used in the inter-frame prediction process.
[0040] In one possible implementation of the method provided in the second aspect, the lookup table is as follows:
[0041]
[0042] The "Subpixel Offset" column is defined with a 1 / 32 subpixel resolution, and c0, c1, c2, and c3 are filter coefficients.
[0043] In one possible implementation of the method provided in the second aspect, the lookup table is as follows:
[0044]
[0045] The "Subpixel Offset" column is defined with a 1 / 32 subpixel resolution, and c0, c1, c2, and c3 are filter coefficients.
[0046] In one possible implementation of the method provided in the second aspect, inter-frame prediction sample values are used for the chroma components of the current decoding block.
[0047] In one possible implementation of the method provided in the second aspect, intra-frame prediction sample values are used for the luminance component of the current decoding block.
[0048] In one possible implementation of the method provided in the second aspect, when the size of the primary reference edge used in intra-frame prediction is less than or equal to a threshold, a lookup table used in intra-frame prediction is selected.
[0049] In one possible implementation of the method provided in the second aspect, the threshold is 8 samples.
[0050] In one possible implementation of the method provided in the second aspect, the inter-frame prediction processing is intra-block copying processing.
[0051] According to a third aspect, the present invention relates to an encoder comprising processing circuitry for performing the methods provided by the first aspect, the second aspect, any possible embodiment of the first aspect, or any possible embodiment of the second aspect.
[0052] According to a fourth aspect, the present invention relates to a decoder comprising processing circuitry for performing the methods provided by the first aspect, the second aspect, any possible embodiment of the first aspect, or any possible embodiment of the second aspect.
[0053] According to a fifth aspect, the present invention relates to an apparatus for decoding a video stream, comprising a processor and a memory. The memory stores instructions that cause the processor to perform the methods provided in the first aspect, the second aspect, any possible embodiment of the first aspect, or any possible embodiment of the second aspect.
[0054] According to a sixth aspect, the present invention relates to an apparatus for decoding a video stream, comprising a processor and a memory. The memory stores instructions that cause the processor to perform the methods provided by the first aspect, the second aspect, any possible embodiment of the first aspect, or any possible embodiment of the second aspect.
[0055] According to a seventh aspect, a computer-readable storage medium is provided storing instructions that, when executed, cause one or more processors to decode video data. The instructions cause the one or more processors to perform the methods provided by the first aspect, the second aspect, any possible embodiment of the first aspect, or any possible embodiment of the second aspect.
[0056] According to an eighth aspect, the present invention relates to a computer program including program code for performing, when executed on a computer, the methods provided by the first aspect, the second aspect, any possible embodiment of the first aspect, or any possible embodiment of the second aspect.
[0057] One or more embodiments will be described in detail in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. Attached Figure Description
[0058] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings. In the drawings:
[0059] Figure 1 A block diagram illustrating an example of a video decoding system for implementing embodiments of the present invention;
[0060] Figure 2 A block diagram illustrating an example of a video encoder used to implement embodiments of the present invention;
[0061] Figure 3 This is a block diagram illustrating an example of a video decoder used to implement embodiments of the present invention;
[0062] Figure 4Examples of 67 intra-frame prediction modes are shown;
[0063] Figure 5 Examples of interpolation filters used for inter-frame prediction and intra-frame prediction are shown;
[0064] Figure 6 Another example of an interpolation filter used for inter-frame prediction and intra-frame prediction is shown;
[0065] Figure 7 Another example of an interpolation filter used for inter-frame prediction and intra-frame prediction is shown;
[0066] Figure 8 An embodiment of the present invention is shown that repeatedly uses a 4-tap interpolation filter for inter-frame prediction and intra-frame prediction;
[0067] Figure 9 Another embodiment of the invention is shown, in which a 4-tap interpolation filter is repeatedly used for inter-frame prediction and intra-frame prediction.
[0068] Figure 10 An embodiment of the present invention is shown that repeatedly uses 4-tap coefficients for inter-frame prediction and intra-frame prediction;
[0069] Figure 11 Examples of 35 intra-frame prediction modes are shown;
[0070] Figure 12 An example of interpolation filter selection is shown;
[0071] Figure 13 Examples of quadtree and binary tree partitioning are shown;
[0072] Figure 14 An example of rectangular block orientation is shown;
[0073] Figure 15 Another example of interpolation filter selection is shown;
[0074] Figure 16 Another example of interpolation filter selection is shown;
[0075] Figure 17 Another example of interpolation filter selection is shown;
[0076] Figure 18 This is a schematic diagram of a network device.
[0077] Figure 19 A block diagram of an apparatus is shown;
[0078] Figure 20 This is a flowchart of one embodiment of the present invention.
[0079] In the following text, unless otherwise expressly stated, the same reference numerals refer to the same or at least functionally equivalent features. Detailed Implementation
[0080] In the following description, reference is made to the accompanying drawings, which form part of the invention and illustrate specific aspects of embodiments of the invention or aspects that may be used with respect to embodiments of the invention. It should be understood that embodiments of the invention may be used in other aspects and may include structural or logical variations not depicted in the drawings. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of the invention is defined by the appended claims.
[0081] For example, it should be understood that the disclosure of the described method is equally applicable to corresponding devices or systems used to perform the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units (e.g., functional units) to perform the described one or more method steps (e.g., one unit performs one or more steps, or each of a plurality of units performs one or more of a plurality of steps), even if the one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a specific apparatus is described based on one or more units (e.g., functional units), the corresponding method may include a step to implement the function of one or more units (e.g., one step implements the function of one or more units, or each of a plurality of steps implements the function of one or more units of a plurality of units), even if the one or more steps are not explicitly described or illustrated in the drawings. Furthermore, it should be understood that, unless otherwise stated, features of the various exemplary embodiments and / or aspects described herein may be combined with each other.
[0082] Abbreviations and Terminology Definitions
[0083] JEMJoint Exploration Model (a software code library for future video decoding exploration)
[0084] JVET Joint Video Experts Team
[0085] LUTLook-Up Table
[0086] QTQuadTree quadtree
[0087] QTBTQuadTree plus Binary Tree (quadtree plus binary tree)
[0088] Rate-distortion optimization
[0089] ROM (Read-Only Memory)
[0090] VVC Test Model
[0091] VVCVersatile Video Coding is a standardized project developed by JVET.
[0092] CTU / CTB – Coding Tree Unit / Coding Tree Block
[0093] CU / CB – Coding Unit / Coding Block
[0094] PU / PB – Prediction Unit / Prediction Block
[0095] TU / TB –Transform Unit / Transform Block
[0096] HEVC – High Efficiency Video Coding
[0097] Video decoding schemes such as H.264 / AVC and HEVC are designed based on the successful principles of block-based hybrid video decoding. Using this principle, the image is first segmented into blocks, and then each block is predicted using intra-frame or inter-frame prediction.
[0098] Several video decoding standards since H.261 belong to the "lossy hybrid video codec" group (i.e., combining spatial and temporal prediction in the sample domain with 2D transform decoding in the transform domain for applying quantization). Each image in a video sequence is typically segmented into a set of non-overlapping blocks, and decoding is usually performed at the block level. In other words, at the encoder, the video is typically processed at the block (video block) level, i.e., encoded, for example, by generating prediction blocks through spatial (intra-frame image) prediction and temporal (inter-frame image) prediction; the prediction blocks are subtracted from the current block (the block currently being processed / to be processed) to obtain residual blocks; the residual blocks are transformed and quantized in the transform domain to reduce the amount of data to be transmitted (compressed). At the decoder, in contrast to the encoder, the inverse processing is partially applied to the encoded or compressed blocks to reconstruct the current block for representation. Furthermore, the encoder repeats the processing steps of the decoder, such that the encoder and decoder generate the same predictions (e.g., intra-frame and inter-frame predictions) and / or reconstructions for processing (i.e., decoding) of subsequent blocks.
[0099] As used herein, the term "block" can be a portion of an image or frame. For ease of description, embodiments of the invention are described herein with reference to the High-Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC) reference software developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that embodiments of the invention are not limited to HEVC or VVC. It can refer to CU, PU, and TU. In HEVC, the CTU is divided into CUs using a quadtree structure represented as a decoding tree. At the CU level, it is determined whether to use inter-frame (temporal) prediction or intra-frame (spatial) prediction to decode the image region. Each CU can be further divided into one, two, or four PUs depending on the PU partitioning type. The same prediction process is applied within a PU, and relevant information is transmitted to the decoder on a PU-by-PU basis. After obtaining residual blocks through the prediction process based on the PU partitioning type, the CU can be segmented into transform units (TUs) according to another quadtree structure similar to the decoding tree used for the CU. Recent advances in video compression technology utilize quadtree and binary tree (QTBT) segmentation to partition the decoded blocks. In the QTBT block structure, the CU can be square or rectangular. For example, the coding tree unit (CTU) is first segmented using a quadtree structure. The leaf nodes of the quadtree are then further segmented using a binary tree structure. The leaf nodes of the binary tree are called coding units (CUs), and this segmentation is used for prediction and transform processing without any further segmentation. That is, in the QTBT decoded block structure, the block sizes of CU, PU, and TU are the same. Furthermore, it has been proposed to combine multiple segmentations, such as ternary tree segmentation, with the QTBT block structure.
[0100] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) are studying the potential need for standardization of future video decoding technologies, which will significantly exceed the compression capabilities of the current HEVC standard (including current and recent extensions for screen content decoding and high dynamic range decoding). These groups are working with the Joint Video Exploration Team (JVET) on this exploration to evaluate compression technology designs proposed by their domain experts.
[0101] The Versatile Test Model (VTM) uses 35 intra-frame modes, while the Benchmark Set (BMS) uses 67 intra-frame modes. Intra-frame prediction is a mechanism used in many video decoding frameworks to improve compression efficiency when only a given frame is involved.
[0102] The video decoding used in this document refers to the processing of image sequences that constitute a video or video sequence. The terms "picture" or "frame" are used synonymously in the field of video decoding and in this application. Each image is typically segmented into a set of non-overlapping blocks. Image encoding / decoding is typically performed at the block level. For example, at the block level, prediction blocks are generated using inter-frame prediction or intra-frame prediction to subtract the prediction blocks from the current block (the currently processed block / the block to be processed) to obtain residual blocks. These residual blocks are further transformed and quantized to reduce the amount of data to be transmitted (compressed). At the decoding end, the encoded / compressed blocks are inversely processed to reconstruct the blocks for representation.
[0103] Figure 1 For illustrative purposes, an exemplary decoding system 10, such as a video decoding system 10, is shown that can utilize the techniques of this application (the present invention). The encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) of the video decoding system 10 represent examples of devices that can be used to perform various techniques according to the various examples described in this application. Figure 1 As shown, the decoding system 10 includes a source device 12, which provides encoded data 13 (e.g., encoded image 13) to a destination device 14, etc., for decoding the encoded data 13.
[0104] The source device 12 includes an encoder 20 and may additionally (optionally) include an image source 16, a preprocessing unit 18 (e.g., an image preprocessing unit 18), and a communication interface or communication unit 22.
[0105] Image source 16 may include or may be any type of image capture device, such as a device for capturing real-world images, and / or any type of image or commentary (for screen content decoding, some text on the screen is also considered part of the image to be encoded), such as a computer graphics processor for generating computer-animated images, or any type of device for acquiring and / or providing real-world images, computer-animated images (e.g., screen content, virtual reality (VR) images) and / or any combination thereof (e.g., augmented reality (AR) images).
[0106] A (digital) image is, or can be viewed as, a two-dimensional array or matrix of samples with intensity values. Samples in the array can also be called pixels (or pels) (short for image elements). The number of samples in the array or image along the horizontal and vertical directions (or axes) defines the image's size and / or resolution. Colors are typically represented using three color components; that is, the image can be represented as a three-sample array or include three sample arrays. In RGB format or color space, the image includes corresponding red, green, and blue sample arrays. However, in video decoding, each pixel is typically represented by a luma / chroma format or in a color space, such as YCbCr, including a luma component indicated by Y (sometimes also indicated by L) and two chroma components indicated by Cb and Cr. The luma (or simply luma) component Y represents the brightness or grayscale intensity (e.g., in a grayscale image), while the two chroma (or simply chroma) components Cb and Cr represent the chroma or color information components. Therefore, an image in YCbCr format includes a luma sample array of luma sample values (Y) and two chroma sample arrays of chroma values (Cb and Cr). An RGB format image can be converted or transformed to YCbCr format and vice versa; this process is also known as color transformation or conversion. If the image is monochrome, it may consist only of an array of luminance samples.
[0107] Image source 16 (e.g., video source 16) can be a camera for capturing images, a memory (e.g., an image memory) that includes or stores previously captured or generated images, and / or any type of (internal or external) interface for acquiring or receiving images. For example, the camera can be a local or integrated camera integrated into the source device, and the memory can be, for example, a local or integrated memory integrated into the source device. For example, the interface can be an external interface for receiving images from an external video source, such as an external image capture device like a camera, external memory, or an external image generation device such as an external computer graphics processor, computer, or server. The interface can be any type of interface according to any proprietary or standardized interface protocol, such as a wired or wireless interface, or an optical interface. The interface for acquiring image data 17 can be the same interface as communication interface 22, or as part of communication interface 22.
[0108] Unlike the preprocessing unit 18 and the processing performed by the preprocessing unit 18, the image or image data 17 (e.g., video data 16) can also be referred to as the raw image or raw image data 17.
[0109] The preprocessing unit 18 is used to receive (raw) image data 17 and preprocess the image data 17 to obtain a preprocessed image 19 or preprocessed image data 19. The preprocessing performed by the preprocessing unit 18 may include cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or noise reduction. It is understood that the preprocessing unit 18 may be an optional component.
[0110] Encoder 20 (e.g., video encoder 20) is used to receive preprocessed image data 19 and provide encoded image data 21 (hereinafter referred to as...) Figure 2 (Further details to be described).
[0111] The communication interface 22 of the source device 12 can be used to receive the encoded image data 21 and send it to other devices (e.g., the destination device 14 or any other device for storage or direct reconstruction); or to process the encoded image data 21 before storing the encoded data 13 and / or sending the encoded data 13 to other devices (e.g., the destination device 14 or any other device for decoding or storage).
[0112] The target device 14 includes a decoder 30 (e.g., a video decoder 30) and may additionally (i.e., optionally) include a communication interface or communication unit 28, a post-processing unit 32, and a display device 34.
[0113] The communication interface 28 of the destination device 14 is used to receive encoded image data 21 or encoded data 13, for example, directly from the source device 12 or any other source (e.g., a storage device such as an encoded image data storage device).
[0114] Communication interfaces 22 and 28 can be used to send or receive encoded image data 21 or encoded data 13 via a direct communication link between source device 12 and destination device 14 (e.g., a direct wired or wireless connection), or via any type of network (e.g., wired or wireless networks or any combination thereof, or any type of private and public network), or any combination thereof.
[0115] For example, the communication interface 22 can be used to package the encoded image data 21 into a suitable format (e.g., a data packet) for transmission over a communication link or communication network.
[0116] The corresponding part of the communication interface 22, the communication interface 28, can be used to unpack the encoded data 13 to obtain the encoded image data 21, etc.
[0117] Both communication interface 22 and communication interface 28 can be configured as unidirectional communication interfaces (e.g., Figure 1 The source device 12 points to the destination device 14 via an arrow indicating the encoded image data 13, or a bidirectional communication interface, and can be used to send and receive messages, such as establishing connections, acknowledging and exchanging any other information related to the communication link and / or data transmission (e.g., encoded image data transmission).
[0118] Decoder 30 is used to receive encoded image data 21 and provide decoded image data 31 or decoded image 31 (hereinafter referred to as...). Figure 3 (Further details to be described).
[0119] The post-processor 32 of the destination device 14 is used to post-process the decoded image data 31 (also known as reconstructed image data) (e.g., decoded image 31) to obtain post-processed image data 33 (e.g., post-processed image 33). The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming or resampling, or any other processing, such as preparing the decoded image data 31 for display by the display device 34, etc.
[0120] The display device 34 of the target device 14 is used to receive the post-processed image data 33 to display the image to a user or viewer. The display device 34 can be or includes any type of display for displaying the reconstructed image, such as an integrated or external display or monitor. For example, the display can include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0121] although Figure 1 Source device 12 and destination device 14 are described as separate devices; however, device embodiments may also include two devices or two functions, namely source device 12 or its corresponding function and destination device 14 or its corresponding function. In such embodiments, source device 12 or its corresponding function and destination device 14 or its corresponding function may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.
[0122] Based on the description, it is obvious to the technicians that... Figure 1 The presence and (precise) division of different units or functions in the source device 12 and / or destination device 14 shown may vary depending on the actual device and application.
[0123] Encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) can each be implemented as any of a variety of suitable circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology is implemented in part in software, the device can store the software instructions in a suitable non-transitory computer-readable medium and can use one or more processors to execute the instructions in the hardware to perform the technology of the present invention. Any of the foregoing circuits (including hardware, software, combinations of hardware and software, etc.) can be considered as one or more processors. Video encoder 20 and video decoder 30 can each be included in one or more encoders or decoders, wherein either can be integrated into the respective device as part of a combined encoder / decoder (encoder-decoder).
[0124] Figure 2 A schematic / conceptual block diagram of an exemplary video encoder 20 for implementing the technology of this application is shown. Figure 2 In the example, the video encoder 20 includes a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a buffer 216, a loop filter unit 220, a decoded picture buffer (DPB) 230, a prediction processing unit 260, and an entropy coding unit 270. The prediction processing unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254, and a mode selection unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The video encoder 20 shown can also be called a hybrid video encoder or a video encoder based on a hybrid video codec.
[0125] For example, the residual calculation unit 204, transform processing unit 206, quantization unit 208, prediction processing unit 260, and entropy encoding unit 270 form the forward signal path of encoder 20, while the inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, buffer 216, loop filter 220, decoded picture buffer (DPB) 230, and prediction processing unit 260 form the reverse signal path of encoder. The reverse signal path of encoder is connected to the decoder (see...). Figure 3The signal path corresponds to that of decoder 30 in the code.
[0126] Encoder 20 is used to receive images 201 or blocks 203 of images 201 (e.g., images forming a video or video sequence) via input terminal 202, etc. Image block 203 may also be referred to as current image block or image block to be decoded, and image 201 may also be referred to as current image or image to be decoded (especially in video decoding, in order to distinguish the current image from other images (e.g., previously encoded and / or decoded images of the same video sequence (i.e., video sequences that also include the current image)).
[0127] The prediction processing unit 260 (also known as the block prediction processing unit 260) is used to: receive or acquire block 203 (the current block 203 of the current image 201) and reconstructed image data, such as reference samples of the same (current) image from buffer 216 and / or reference image data 231 of one or more previously decoded images from decoded image buffer 230, and process such data to make predictions, i.e., provide prediction blocks 265, wherein the prediction blocks 265 may be inter-frame prediction blocks 245 or intra-frame prediction blocks 255.
[0128] The mode selection unit 262 can be used to select a prediction mode (e.g., intra-frame prediction or inter-frame prediction mode) and / or the corresponding prediction block 245 or 255 as prediction block 265 for calculating residual block 205 and reconstruction block 215.
[0129] Embodiments of the mode selection unit 262 can be used to select a segmentation and prediction mode (e.g., from prediction modes supported by the prediction processing unit 260), which provides the best match or minimum residual (minimum residual refers to better compression in transmission or storage), or provides minimum signaling overhead (minimum signaling overhead refers to better compression in transmission or storage), or considers or balances both. The mode selection unit 262 can be used to determine the prediction mode based on rate distortion optimization (RDO), i.e., selecting the prediction mode that provides minimum rate distortion optimization, or selecting the prediction mode with relevant rate distortion that at least meets the prediction mode selection criteria. In this document, terms such as "best," "minimum," and "optimal" do not necessarily refer to "best," "minimum," or "optimal" overall; they can also refer to situations where termination or selection criteria are met. For example, values exceeding or falling below a threshold or other constraints may lead to a "suboptimal selection," but this can reduce complexity and processing time.
[0130] Intra-prediction unit 254 is further configured to determine intra-prediction block 255 based on intra-prediction parameters (e.g., the selected intra-prediction mode). In any case, after selecting an intra-prediction mode for a block, intra-prediction unit 254 is also configured to provide intra-prediction parameters, i.e., information indicating the selected intra-prediction mode of the block, to entropy coding unit 270. In one example, intra-prediction unit 254 may be used to perform any combination of intra-prediction techniques described below.
[0131] Figure 3 An exemplary video decoder 30 is provided for implementing the technology of this application. The video decoder 30 is used to receive encoded image data (e.g., encoded bitstream) 21, for example, encoded by encoder 100, to obtain a decoded image 131. During the decoding process, the video decoder 30 receives video data from the video encoder 100, such as encoded video bitstreams representing image blocks of encoded video slices and associated syntax elements.
[0132] exist Figure 3 In one example, decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., a summer 314), a buffer 316, a loop filter 320, a decoded image buffer 330, and a prediction processing unit 360. Prediction processing unit 360 may include an inter-frame prediction unit 344, an intra-frame prediction unit 354, and a mode selection unit 362. In some examples, video decoder 30 may perform operations typically related to... Figure 2 The video encoder 100 describes the encoding process as the opposite of the decoding process.
[0133] Entropy decoding unit 304 is used to perform entropy decoding on the encoded image data 21 to obtain quantization coefficients 309 and / or decoded parameters. Figure 3 (Not shown in the image) , such as any or all of the (decoded) inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 is also used to forward inter-frame prediction parameters, intra-frame prediction parameters, and / or other syntax elements to the prediction processing unit 360. The video decoder 30 can receive syntax elements at the video stripe level and / or video block level.
[0134] The functions of the inverse quantization unit 310 and the inverse quantization unit 110 are the same; the functions of the inverse transform processing unit 312 and the inverse transform processing unit 112 are the same; the functions of the reconstruction unit 314 and the reconstruction unit 114 are the same; the functions of the buffer 316 and the buffer 116 are the same; the functions of the loop filter 320 and the loop filter 120 are the same; and the functions of the decoding image buffer 330 and the decoding image buffer 130 are the same.
[0135] The prediction processing unit 360 may include an inter-frame prediction unit 344 and an intra-frame prediction unit 354, wherein the function of the inter-frame prediction unit 344 may be similar to that of the inter-frame prediction unit 144, and the function of the intra-frame prediction unit 354 may be similar to that of the intra-frame prediction unit 154. The prediction processing unit 360 is typically used to perform block prediction based on the encoded data 21 and / or obtain the predicted block 365, and to receive or obtain (explicitly or implicitly) prediction-related parameters and / or information about the selected prediction mode from the entropy decoding unit 304, etc.
[0136] When decoding a video stripe into an intra-frame decoded (I) stripe, the intra-frame prediction unit 354 of the prediction processing unit 360 generates a prediction block 365 of the current video stripe's image blocks based on the indicated intra-frame prediction mode and data from previous decoded blocks of the current frame or image. When decoding a video frame into an inter-frame decoded (i.e., B or P) stripe, the inter-frame prediction unit 344 (e.g., a motion compensation unit) of the prediction processing unit 360 generates a prediction block 365 of the current video stripe's video blocks based on motion vectors and other syntax elements received from the entropy decoding unit 304. For inter-frame prediction, these prediction blocks can be generated from one of the reference images in one of the reference image lists. The video decoder 30 can construct the reference frame lists: list 0 and list 1, based on the reference images stored in the DPB 330 using a default construction technique.
[0137] The prediction processing unit 360 is used to determine prediction information for video blocks in the current video strip by parsing motion vectors and other syntax elements, and to generate prediction blocks for the decoded current video blocks using the prediction information. For example, the prediction processing unit 360 uses some received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) for decoding video blocks in the video strip, the inter-frame prediction stripe type (e.g., B-strip, P-strip, or GPB-strip), the construction information of one or more reference image lists for the strip, the motion vectors of each inter-frame coded video block in the strip, the inter-frame prediction state of each inter-frame decoded video block in the strip, and other information to decode video blocks in the current video strip.
[0138] The dequantization unit 310 is used to dequantize, or dequantize, the quantization transform coefficients provided in the bitstream and decoded by the entropy decoding unit 304. The dequantization process may include determining the degree of quantization using quantization parameters calculated by the video encoder 100 for each video block in the video strip, and similarly determining the degree of dequantization to be applied.
[0139] The inverse transform processing unit 312 is used to apply an inverse transform to the transform coefficients, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to generate a residual block in the pixel domain.
[0140] Reconstruction unit 314 (e.g., summer 314) is used to add inverse transform block 313 (i.e. reconstruction residual block 313) to prediction block 365 by adding the sample values of reconstruction residual block 313 and the sample values of prediction block 365, to obtain reconstruction block 315 in the sample domain.
[0141] Loop filter unit 320 (in or after the decoding loop) is used to filter the reconstructed block 315 to obtain filtered block 321, to smooth pixel transitions or otherwise improve video quality. In one example, loop filter unit 320 can be used to perform any combination of the filtering techniques described below. Loop filter unit 320 is used to represent one or more loop filters, such as deblocking filters, sample-adaptive offset (SAO) filters, or other filters, such as bilateral filters or adaptive loop filters (ALF), or sharpening or smoothing filters or cooperative filters. Although loop filter unit 320 is used in... Figure 3 The loop filter is shown in the diagram, but in other configurations, the loop filter unit 320 can be implemented as a post-loop filter.
[0142] Then, the decoded video block 321 in the given frame or image is stored in the decoded image buffer 330, which stores a reference image for subsequent motion compensation.
[0143] The decoder 30 is used to output the decoded image 331 through the output terminal 332, etc., to present to the user or allow the user to view it.
[0144] Other forms of video decoder 30 can be used to decode the compressed bitstream. For example, decoder 30 can generate an output video stream without loop filtering unit 320. For example, non-transform-based decoder 30 can directly dequantize the residual signals of certain blocks or frames without inverse transform processing unit 312. In another implementation, video decoder 30 may have a dequantization unit 310 and an inverse transform processing unit 312 combined into a single unit.
[0145] Figure 4 Examples of 67 intra-prediction modes proposed for VVC are shown. These 67 intra-prediction modes include: planar mode (index 0), DC mode (index 1), and angle modes (indexes 2 to 66). Figure 4 The bottom-left angle pattern in the text refers to index 2, and the index numbers are incremented until index 66 corresponds to... Figure 4 Up to the top right angle mode.
[0146] like Figure 4 As shown, the latest version of JEM has several modes corresponding to the tilted intra-prediction direction. For any of these modes, to predict samples within a block, if the corresponding position within the block edge is a fraction, interpolation of the adjacent reference sample set should be performed. In HEVC and VVC, linear interpolation between two adjacent reference samples is used. In JEM, a complex 4-tap interpolation filter is used. Gaussian or cubic filter coefficients are selected based on the block width or height value. The decision to use width or height is consistent with the decision to select the primary reference edge. When the value of the intra-prediction mode is greater than or equal to the value of the diagonal mode, the top edge of the reference sample is selected as the primary reference edge, and the width value is selected to determine the interpolation filter being used. When the value of the intra-prediction mode is less than the value of the diagonal mode, the primary reference edge is selected from the left side of the block, and the height value is used to control the filter selection process. Specifically, if the length of the selected edge is less than or equal to 8 samples, a 4-tap cubic filter is used. If the length of the selected edge is greater than 8 samples, a 4-tap Gaussian filter is used as the interpolation filter.
[0147] The specific filter coefficients used in JEM are shown in Table 1. Based on the sub-pixel offset and filter type, the predicted samples are calculated by performing convolution operations using the coefficients selected from Table 1, as shown below:
[0148]
[0149] In this equation, " indicates a bitwise right shift operation.
[0150] If a third filter is selected, the predicted sample (sample value) is further clipped to a range of allowed values, which is defined in the sequence parameter set (SPS) or derived from the bit depth of the selected component.
[0151] Table 1: Intra-frame prediction interpolation filters
[0152]
[0153] When the displacement of the reference block's pixels is fractional, the motion compensation process also uses filtering to predict sample values. In JEM, the luma component uses an 8-tap filter, and the chroma component uses a 4-tap length filter. First, a motion interpolation filter is applied horizontally, and then the output of the horizontal filter is further vertically filtered. Table 2 shows the coefficients of the 4-tap chroma filter.
[0154] Table 2: Chromaticity Motion Interpolation Filter Coefficients
[0155]
[0156] Many video decoding solutions also use different interpolation filters for intra-frame prediction and inter-frame prediction. Specifically, Figures 5-7 Different examples of interpolation filters are shown. Figure 5 An example of an interpolation filter used in JEM is shown. Figure 6 Another example of an interpolation filter proposed for Core-experiment CE 3–3.1.3, disclosed in ITU-JVET K1023, is shown. Figure 7 Another example of an interpolation filter proposed in ITU-JVET K0064 is shown.
[0157] The basic idea of this invention is to reuse the lookup table and / or hardware module of the chroma motion compensation subpixel filter. If the pixel value is in a fractional position between reference samples, the pixel value is interpolated into the intra-frame prediction value. Since the same hardware is expected to be used for inter-frame prediction and intra-frame prediction, the precision of the filter coefficients should be consistent; that is, the number of bits in the filter coefficients representing the intra-frame reference sample interpolation should be consistent with the precision of the motion subpixel motion compensation interpolation filter coefficients.
[0158] Figure 8 An embodiment of the present invention is shown. A 4-tap interpolation filter with 6-bit chroma coefficients (also known as a "unified intra / inter frame filter") can be used for two processes: interpolation of intra-frame prediction samples and interpolation of inter-frame prediction samples.
[0159] Figure 9 A specific embodiment utilizing this design is illustrated. In this implementation, the filtering module is implemented as a separate unit that participates in predicting chroma samples during motion compensation and predicting luma and chroma samples when performing intra-frame prediction. In this implementation, the hardware filtering portion is used for both intra-frame and inter-frame prediction processes.
[0160] Figure 10 Another embodiment is shown when only the LUT of the filter coefficients is reused (see Figure 10 ). Figure 10 This is an exemplary implementation of a LUT based on reused coefficients. In this implementation, a hardware filtering module loads coefficients from the LUT stored in ROM. The switch shown for the intra-frame prediction process determines the type of filter to be used based on the selected main side length for intra-frame prediction processing.
[0161] A practical embodiment of the provided application may use the following coefficients (see Table 3).
[0162] Table 3: Intra-frame and Inter-frame Interpolation Filters
[0163]
[0164] Based on the sub-pixel offset and filter type, intra-frame predicted samples are calculated by performing convolution operations using coefficients selected from Table 3, as shown below:
[0165]
[0166] In this equation, " indicates a bitwise right shift operation.
[0167] If you select "Unified Intra / Inter-Frame Filter", the predicted samples will be further corrected to be within the allowed range of values, which is defined in SPS or derived from the bit depth of the selected component.
[0168] The distinguishing features of the provided embodiments of the present invention are as follows:
[0169] For intra-frame reference sample interpolation and sub-pixel motion compensation interpolation, the same filters can be used to reuse hardware modules and reduce the overall memory size required.
[0170] In addition to reusing filters, the precision of the filter coefficients used for intra-frame reference sample interpolation should be consistent with the precision of the reused filter coefficients mentioned above.
[0171] Luminance processing in motion compensation doesn't necessarily require an 8-tap filter; a 4-tap filter can also be used. In this case, a 4-tap filter can be chosen for uniformity.
[0172] The embodiments of the present invention can be applied to different parts of the intra-frame prediction process, which may involve interpolation. In particular, when expanding the master reference sample, a uniform interpolation filter can also be used to filter the side reference samples (see Sections 2.1, 3.1, 4.1 and 5 of JVET-K0211 for details).
[0173] The intra-block copy operation also involves the use of interpolation steps that can be employed with the invention provided (for a description of intra-block copy, see [Xiaozhong Xu, Shan Liu, Tzu-Der Chuang, Yu-Wen Huang, Shawmin Lei, Krishnakanth Rapaka, Chao Pang, Vadim Seregin, Ye-Kui Wang, Marta Karczewicz: Intra Block Copy in HEVC Screen ContentCoding Extensions. IEEE J. Emerg. Sel. Topics Circuits Syst. 6(4): 409-419(2016)]).
[0174] Other embodiments may include a method for aspect ratio correlation filtering for intra-frame prediction, the method comprising:
[0175] Select an interpolation filter for the block to be predicted based on the block's aspect ratio.
[0176] In one example, the selection of the interpolation filter depends on the direction of the threshold of the intra-prediction mode used for the block to be predicted.
[0177] In one example, the direction corresponds to the angle of the main diagonal of the block to be predicted.
[0178] In one example, the angle of the direction is calculated as follows:
[0179] ,
[0180] in, These represent the width and height of the block to be predicted, respectively.
[0181] In one example, the aspect ratio R is determined. A For example, corresponding to the following equation:
[0182] ,in, These represent the width and height of the block to be predicted, respectively.
[0183] In one example, the angle of the main diagonal of the block to be predicted is determined based on the aspect ratio.
[0184] In one example, the threshold for the intra-prediction mode of the block is determined based on the angle of the main diagonal of the block to be predicted.
[0185] In one example, the choice of interpolation filter depends on which side the reference sample used belongs to.
[0186] In one example, the angle corresponds to a straight line in the intra-frame direction, dividing the block into two regions.
[0187] In one example, different interpolation filters are used to predict reference samples belonging to different regions.
[0188] In one example, the filter includes a cubic interpolation filter or a Gaussian interpolation filter.
[0189] In one implementation of this application, a frame is the same as an image.
[0190] In one implementation of this application, the value corresponding to VER_IDX is 50; the value corresponding to HOR_IDX is 18; the value corresponding to VDIA_IDX is 66, which can be the maximum value among the values corresponding to the angle mode; the value corresponding to intra-frame mode 2 can be the minimum value among the values corresponding to the angle mode; and the value corresponding to DIA_IDX is 34.
[0191] This invention provides an improvement to the intra-frame mode indication scheme. It proposes a video decoding method and a video decoder.
[0192] Figure 4 Examples of the 67 intra-prediction modes provided for VVC are shown. These 67 intra-prediction modes include: planar mode (index 0), DC mode (index 1), and angle modes (indexes 2 to 66), where... Figure 4 The bottom-left angle pattern in the text refers to index 2, and the index numbers are incremented until index 66 corresponds to... Figure 4 Up to the top right angle mode.
[0193] In another aspect of this application, a decoder including processing circuitry is disclosed for performing the above-described decoding method.
[0194] In another aspect of this application, a computer program product is disclosed, which includes program code for performing the above-described decoding method.
[0195] In another aspect of this application, a decoder for decoding video data is disclosed, the decoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing a program for execution by the processor, wherein, when the processor executes the program, the decoder is configured to perform the above-described decoding method.
[0196] The processing circuitry can be implemented in hardware or a combination of hardware and software, for example, through a software-programmable processor.
[0197] The processing circuitry can be implemented in hardware or a combination of hardware and software, for example, through a software-programmable processor.
[0198] Figure 11 A schematic diagram illustrates various intra-prediction modes used in the HEVC UIP scheme, which can be used by another embodiment. For a luma block, the intra-prediction modes can include up to 36 intra-prediction modes, including three non-directional modes and 33 directional modes. Non-directional modes can include planar prediction modes, mean (DC) prediction modes, and chroma prediction modes (LM) derived from luma prediction modes. Planar prediction modes perform prediction by assuming the block amplitude surface has horizontal and vertical slopes derived from the block boundaries. DC prediction modes perform prediction by assuming the flat block surface has values matching the mean of the block boundaries. LM prediction modes perform prediction by assuming the block's chroma values match the block's luma values. Directional modes can perform prediction based on neighboring blocks, such as... Figure 11 As shown.
[0199] H.264 / AVC and HEVC specify that a low-pass filter can be applied to a reference sample before being used in the intra-prediction process. Whether to use a reference sample filter is determined by the intra-prediction mode and block size. This mechanism can be called Mode Dependent Intra Smoothing (MDIS). Several methods related to MDIS also exist. For example, the Adaptive Reference Sample Smoothing (ARSS) method can explicitly (i.e., with the tag included in the bitstream) or implicitly (i.e., using data hiding to avoid including the tag in the bitstream to reduce indication overhead) indicate whether to filter the prediction samples. In this case, the encoder can decide to perform smoothing by testing the rate-distortion (RD) cost of all potential intra-prediction modes.
[0200] like Figure 4As shown, the latest version of JEM (JEM-7.2) has several modes corresponding to tilted intra-prediction directions. For any of these modes, if the corresponding position within a block edge is a fraction, interpolation of the adjacent set of reference samples should be performed to predict the sample within the block. HEVC and VVC use linear interpolation between two adjacent reference samples. JEM uses a more complex 4-tap interpolation filter. Gaussian or cubic filter coefficients are selected based on either the width or height value. The decision to use width or height is consistent with the decision to select the primary reference edge: when the intra-prediction mode is greater than or equal to the diagonal mode, the top edge of the reference sample is selected as the primary reference edge, and the width value is selected to determine the interpolation filter being used. Otherwise, the primary reference edge is selected from the left side of the block, and the height value is used to control the filter selection process. Specifically, if the length of the selected edge is less than or equal to 8 samples, a 4-tap cubic interpolation filter is used. Otherwise, a 4-tap Gaussian filter is used as the interpolation filter.
[0201] exist Figure 12 The figure shows an example of interpolation filter selection for the less than and greater than diagonal modes (denoted as 45°) in the case of a 32×4 block.
[0202] In VVC, a partitioning mechanism based on quadtrees and binary trees, called QTBT, is used. For example... Figure 13 As shown, QTBT segmentation can provide not only square blocks but also rectangular blocks. Of course, compared to the traditional quadtree-based segmentation used in the HEVC / H.265 standard, QTBT segmentation adds some indication overhead at the encoding end, increasing computational complexity. However, QTBT-based segmentation has better segmentation characteristics, and therefore, it is more efficient in decoding than traditional quadtree segmentation.
[0203] However, in its current state, VVC uses the same filter on both sides (left and top) of the reference sample. The reference sample filter is identical for both reference sample sides, regardless of whether the block is vertical or horizontal.
[0204] In this paper, the terms "vertical block" ("the vertical direction of the block") and "horizontal block" ("the horizontal direction of the block") are applied to rectangular blocks generated by the QTBT framework. The meanings of these terms are... Figure 14 Same as shown.
[0205] This invention proposes a mechanism for selecting different reference sample filters to take into account the orientation of the block. Specifically, the width and height of the block are examined separately so that different reference sample filters are applied to reference samples located on different edges of the block to be predicted.
[0206] In some examples, the choice of interpolation filter is described as consistent with the decision for the selection of the primary reference edge. Both decisions are currently based on a comparison between the intra-prediction mode and the diagonal (45-degree) direction.
[0207] However, it should be noted that this design has serious drawbacks for extended blocks. From Figure 15 As can be observed, even when the shorter side is selected as the primary reference according to the pattern comparison criteria, most predicted pixels will still be derived from the reference sample of the longer side (shown as the dashed area). Figure 15 An example of the selection of a reference filter related to side length is shown.
[0208] This invention provides a method for determining a threshold for an intra-frame prediction mode using an alternative direction during the interpolation filter selection process. Specifically, the direction corresponds to the angle of the main diagonal of the block to be predicted. For example, for blocks of sizes 32×4 and 4×32, this is used to determine the threshold mode of a reference sample filter. m T like Figure 16 The definition is shown.
[0209] The specific value of the intra-frame prediction angle for the threshold can be calculated using the following formula:
[0210] ,
[0211] in, W and H These are the block width and height, respectively.
[0212] Another embodiment of the invention uses different interpolation filters, depending on which side the reference sample used belongs to. Figure 17 An example of this determination is shown in the figure. Figure 17 Examples of using different interpolation filters are shown, depending on which side the reference sample used belongs to.
[0213] An angle corresponding to a straight line in the intra-frame direction m divides the prediction block into two regions. Different interpolation filters are used to predict samples belonging to different regions.
[0214] Example value m T (For the set of intra-prediction modes defined in BMS 1.0) and the corresponding angles are given in Table 4. Angles exist Figure 16 The information is provided in the text.
[0215] Table 4: Example values (for the set of intra-prediction modes defined in BMS 1.0)
[0216]
[0217] Compared with existing technologies and solutions, the present invention uses samples predicted within a block using different interpolation filters, wherein the interpolation filter used to predict the sample is selected according to the block shape, horizontal or vertical direction, and intra-frame prediction mode angle.
[0218] This invention can be applied in the reference sample filtering stage. Specifically, similar rules to those described above for the interpolation filter selection process can be used to determine the reference sample smoothing filter.
[0219] Figure 18 This is a schematic diagram of a network device 1300 (e.g., a decoding device) provided for an embodiment of the present invention. The network device 1300 is suitable for implementing the disclosed embodiments described herein. The network device 1300 includes: an input port 1310 and a receiver unit (Rx) 1320 for receiving data; a processor, logic unit, or central processing unit (CPU) 1330 for processing data; a transmitter unit (Tx) 1340 and an output port 1350 for transmitting data; and a memory 1360 for storing data. The network device 1300 may further include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the input port 1310, receiver unit 1320, transmitter unit 1340, and output port 1350, serving as an output or input source for optical or electrical signals.
[0220] Processor 1330 is implemented in both hardware and software. Processor 1330 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 1330 communicates with ingress port 1310, receiver unit 1320, transmitter unit 1340, egress port 1350, and memory 1360. Processor 1330 includes a decoding module 1370. Decoding module 1370 implements the embodiments disclosed above. For example, decoding module 1370 implements, processes, prepares, or provides various networking functions. Therefore, including decoding module 1370 significantly improves the functionality of network device 1300 and enables transitions between different states of network device 1300. Alternatively, decoding module 1370 can be implemented with instructions stored in memory 1360 and executed by processor 1330.
[0221] Memory 1360 includes one or more disks, tape drives, and solid-state drives, which can be used as overflow data storage devices to store programs when an executable program is selected, and to store instructions and data read during program execution. Memory 1360 can be volatile and / or non-volatile, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0222] While several embodiments have been provided in this invention, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the invention. The examples of this invention should be considered illustrative rather than restrictive, and the invention is not limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0223] Furthermore, without departing from the scope of the invention, the technologies, systems, subsystems, and methods described and illustrated as independent or separate in the various embodiments may be combined or integrated with other systems, modules, technologies, or methods. Other items shown or described as coupled to each other, or directly coupled, or communicating with each other may be indirectly coupled or communicated through some interface, device, or intermediate component, electrically, mechanically, or otherwise. Other examples of variations, substitutions, and modifications can be determined by those skilled in the art and may be exemplified without departing from the spirit and scope of the disclosure herein.
[0224] The following description is in conjunction with the accompanying drawings, which are part of the description and illustrate specific aspects in which the invention can be practiced.
[0225] According to the HEVC / H.265 standard, there are 35 available intra-frame prediction modes. For example... Figure 11 As shown, this set includes the following modes: planar mode (intra-prediction mode index is 0), DC mode (intra-prediction mode index is 1), and mode covering a 180° range with intra-prediction mode index values ranging from 2 to 34 (e.g., ...). Figure 11 The direction (angle) modes (shown by the black arrows in the image) are used in HEVC to capture arbitrary edge directions present in natural video. To capture these directions, the number of intra-frame direction modes used in HEVC has been expanded from 33 to 65. Other direction modes are shown in... Figure 4The dashed arrows represent the planar and DC modes, which remain unchanged. It's worth noting that the intra-frame prediction modes can cover a range greater than 180°. Specifically, the 62 directional modes with index values from 3 to 64 cover a range of approximately 230°, meaning that several pairs of modes have opposite directional properties. For example... Figure 4 As shown, in the case of the HEVC Reference Model (HM) and the JEM platform, only one pair of angular modes (i.e., modes 2 and 66) have opposite directions of directionality. To construct the predictor, conventional angular modes obtain sample predictions by acquiring reference samples and filtering them (if necessary). The number of reference samples required to construct the predictions depends on the length of the filter used for interpolation (e.g., the lengths of bilinear and cubic filters are 2 and 4, respectively).
[0226] For example, in an embodiment, based on the latest video decoding standard currently under development (called Versatile Video Coding (VVC)), a combination of quadtrees nested with multi-type trees (using binary and ternary trees) can be used to partition the structure, for example, for partitioning the coding tree unit (CTU). Within the coding tree structure of the CTU, the CU can be a square or a rectangle. For example, the CTU is first partitioned using a quadtree. Then, the leaf nodes of the quadtree can be further partitioned using a multi-type tree structure. There are four partitioning types for the multi-type tree structure: vertical binary tree partitioning (SPLIT_BT_VER), horizontal binary tree partitioning (SPLIT_BT_HOR), vertical ternary tree partitioning (SPLIT_TT_VER), and horizontal ternary tree partitioning (SPLIT_TT_HOR). The leaf nodes of the multi-type tree are called coding units (CUs), and unless the CU is greater than the maximum transform length, this partitioning is used for prediction and transform processing without any further partitioning. That is, in most cases, the CU, PU, and TU have the same block size in the quadtree-nested multi-type tree decoding block structure. This anomaly occurs when the maximum supported transform length is less than the width or height of the color component of the CU. A unique indication mechanism for the partitioning information in the decoding structure with quadtree-nested multi-type trees is established. In this indication mechanism, the coding tree unit (CTU) is processed as the root of the quadtree and is first partitioned through the quadtree structure. Then, each quadtree leaf node is further partitioned through the multi-type tree structure (when large enough to be partitioned). In the multi-type tree structure, a first flag (mtt_split_cu_flag) indicates whether to further partition the node; when further partitioning the node, a second flag (mtt_split_cu_vertical_flag) indicates the partitioning direction, and then a third flag (mtt_split_cu_binary_flag) indicates whether the partition is a binary tree partition or a ternary tree partition. Based on the values of `mtt_split_cu_vertical_flag` and `mtt_split_cu_binary_flag`, the decoder can derive the multi-type tree partitioning mode (MttSplitMode) of the CU according to predefined rules or tables. It should be noted that for certain designs, such as a 64×64 luma block and a 32×32 chroma pipeline design in a VVC hardware decoder, TT partitioning is not allowed when the width or height of the luma decoding block is greater than 64. Figure 6As shown. TT partitioning is also prohibited when the width or height of the chroma decoding block is greater than 32. Pipeline design divides the image into multiple virtual pipeline data units (VPDUs), defined as non-overlapping units in the image. In a hardware decoder, multiple pipeline stages process consecutive VPDUs simultaneously. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, thus requiring small VPDUs. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, ternary tree (TT) and binary tree (BT) partitioning can increase the VPDU size.
[0227] Additionally, it should be noted that when a portion of a tree node block extends beyond the bottom or right image boundary, the tree node block is forcibly partitioned until all samples of each decoded CU are within the image boundary.
[0228] For example, the Intra Sub-Partitions (ISP) tool can divide the luminance intra-prediction block vertically or horizontally into 2 or 4 sub-partitions based on the block size.
[0229] Intra-frame prediction
[0230] The intra-prediction mode set can include 35 different intra-prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined in HEVC; or it can include 67 different intra-prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined for VVC. In one example, several conventional angular intra-prediction modes are adaptively replaced with, for example, wide-angle intra-prediction modes for non-square blocks as defined in VVC. In another example, to avoid division operations in DC prediction, only the longer side is used to calculate the average value of non-square blocks. Furthermore, the intra-prediction results of planar mode can be modified using the position-dependent intra-prediction combination (PDPC) method.
[0231] The intra-prediction unit is used to generate intra-prediction blocks using reconstructed samples of neighboring blocks in the same current image, based on the intra-prediction modes in the intra-prediction mode set.
[0232] The intra-prediction unit (or typically the mode selection unit) is also used to output intra-prediction parameters (or typically information about the selected intra-prediction mode of the indicator block) to the entropy coding unit as syntax elements to be included in the encoded image data so that, for example, a video decoder can receive and use the prediction parameters for decoding.
[0233] Inter-frame prediction
[0234] In a possible implementation, the set of inter-frame prediction modes depends on the available reference image (i.e., a previously at least partially decoded image stored in the DPB) and other inter-frame prediction parameters, such as whether the entire reference image or only a portion of the reference image (e.g., a search window region near the current block) is used to search for the best matching reference block, and / or, for example, whether pixel interpolation (e.g., half-pixel, quarter-pixel, and / or 1 / 16-pixel interpolation) is applied.
[0235] In addition to the prediction modes mentioned above, skip mode, direct mode and / or other inter-frame prediction modes can also be applied.
[0236] For example, for extended merge prediction, the merge candidate list for this mode consists of five candidate types in sequence: spatial MVP of spatially adjacent CUs, temporal MVP of co-located CUs, history-based MVP of the FIFO table, pairwise average MVP, and zero MV. Decoder-side motion vector refinement (DMVR) based on bilateral matching can be applied to improve the accuracy of the MV in the merge mode. Merge mode with MVD (MMVD) originates from merge modes with motion vector differences. The MMVD flag is indicated immediately after sending the skip flag and merge flag to indicate whether the MMVD mode is used for the CU. An adaptive motion vector resolution (AMVR) scheme at the CU level can be applied. AMVR supports decoding the MVD of the CU with different accuracies. The MVD of the current CU can be adaptively selected based on its prediction mode. When decoding a CU in merge mode, a combined inter / intra prediction (CIIP) mode can be applied to the current CU. CIIP prediction is obtained by weighted averaging of inter-frame and intra-frame prediction signals. For affine motion compensation prediction, the affine motion field of the block is described by motion information from motion vectors of 2 control points (4 parameters) or 3 control points (6 parameters). Subblock-based temporal motion vector prediction (SbTMVP) is similar to temporal motion vector prediction (TMVP) in HEVC, but predicts the motion vector of sub-CUs within the current CU. Bidirectional optical flow (BDOF), formerly known as BIO, is a simplified version that reduces computation, particularly in terms of the number of multiplications and multiplier size. In the triangular partitioning mode, the CU is uniformly divided into two triangular parts using diagonal or anti-diagonal partitioning. Furthermore, the bidirectional prediction mode extends the simple averaging to support weighted averaging of the two prediction signals.
[0237] Inter-frame prediction units can include motion estimation (ME) units and motion compensation (MC) units (both in... Figure 2(Not shown in the image). The motion estimation unit can be used to receive or acquire image blocks (current image blocks of the current image) and decoded images, or at least one or more previously reconstructed blocks, such as reconstructed blocks of one or more other / different previously decoded images, for motion estimation. For example, a video sequence may include the current image and previously decoded images, or in other words, the current image and previously decoded images may be part of or form part of an image sequence that forms a video sequence.
[0238] For example, the encoder can be used to select a reference block from multiple reference blocks of the same or different images in multiple other images, and provide the offset (spatial offset) between the position (x-coordinate, y-coordinate) of the reference image (or reference image index) and / or the position of the reference block and the position of the current block as an inter-frame prediction parameter to the motion estimation unit. This offset is also called the motion vector (MV).
[0239] The motion compensation unit is used to acquire, for example, received inter-frame prediction parameters and perform inter-frame prediction based on or using these parameters to obtain inter-frame prediction blocks. Motion compensation performed by the motion compensation unit may involve extracting or generating prediction blocks based on motion / block vectors determined through motion estimation, and may also involve interpolating sub-pixel precision. Interpolation filtering can generate samples of additional pixels from samples of known pixels, potentially increasing the number of candidate prediction blocks available for decoding image blocks. Once the motion vector of the PU for the current image block is received, the motion compensation unit can locate the prediction block pointed to by the motion vector in one of the reference image lists.
[0240] The motion compensation unit can also generate syntax elements associated with blocks and video stripes for use by the video decoder 30 when decoding image blocks of the video stripes. In addition to stripes and corresponding syntax elements, or as a substitute for stripes and corresponding syntax elements, tile groups and / or tiles and corresponding syntax elements can also be received and / or used.
[0241] like Figure 20 As shown, in an embodiment of the present invention, a video decoding method may include:
[0242] S2001: Obtain video stream.
[0243] The decoding end receives the encoded video stream from the other end (encoding end or network sending end), or the decoding end reads the encoded video stream stored in the decoding end's memory.
[0244] The encoded video stream includes information for decoding the encoded image data, such as data representing image blocks of the encoded video and related syntax elements.
[0245] S2002: Determine whether the prediction sample for the current decoded block is obtained using intra-frame prediction or inter-frame prediction based on the video bitstream.
[0246] At the decoding end, the current decoded block is the block that the decoding end is currently reconstructing. The current decoded block is located in a frame or image of the video.
[0247] The prediction sample for the current decoded block can be determined by using intra-frame prediction or inter-frame prediction based on the syntax elements in the video bitstream.
[0248] A video stream can have a syntax element to indicate whether the current decoded block uses inter-frame prediction or intra-frame prediction. For example, the stream may have a flag indicating whether intra-frame prediction or inter-frame prediction is used for the current decoded block. When the flag value is 1 (or another value), intra-frame prediction is used to obtain the predicted sample for the current decoded block; when the flag value is 0 (or another value), inter-frame prediction is used to obtain the predicted sample for the current decoded block.
[0249] Two or more syntax elements can also be used to indicate whether the current decoded block uses inter-frame prediction or intra-frame prediction. For example, the bitstream may contain an indication (e.g., a flag) to indicate whether the current decoded block uses intra-frame prediction, and the bitstream may contain other indications (e.g., another flag) to indicate whether the current decoded block uses inter-frame prediction.
[0250] When it is determined that intra-frame prediction is used to obtain the prediction sample of the current decoding block, step S2003 is executed. When it is determined that inter-frame prediction is used to obtain the prediction sample of the current decoding block, step S2006 is executed.
[0251] S2003: Obtain the first sub-pixel offset value based on the intra-prediction mode of the current decoding block and the position of the prediction sample within the current decoding block.
[0252] In one example, the intra-frame prediction mode of the current decoded block can also be obtained from the video bitstream.
[0253] Figure 4 Examples of 67 intra-prediction modes proposed for VVC are shown. These 67 intra-prediction modes include: planar mode (index 0), DC mode (index 1), and angle modes (indexes 2 to 66). Figure 4 The bottom-left angle pattern in the text refers to index 2, and the index numbers are incremented until index 66 corresponds to... Figure 4 Up to the top right angle mode.
[0254] Figure 11This diagram illustrates various intra-prediction modes used in the HEVC UIP scheme. For a luma block, intra-prediction modes can include up to 36 modes, comprising three non-directional modes and 33 directional modes. Non-directional modes can include planar prediction modes, mean (DC) prediction modes, and chroma prediction modes (LM) derived from luma prediction modes. Planar prediction modes perform prediction by assuming the block amplitude surface has horizontal and vertical slopes derived from the block boundaries. DC prediction modes perform prediction by assuming the flat block surface has values matching the mean of the block boundaries. LM prediction modes perform prediction by assuming the block's chroma values match the block's luma values. Directional modes can perform prediction based on neighboring blocks, such as... Figure 11 As shown.
[0255] Based on the parsing of the video bitstream of the current decoded block, the intra-prediction mode of the current decoded block can be obtained. In one example, the value of the Most Probable Modes (MPM) flag for the current decoded block is obtained from the video bitstream. In another example, when the MPM flag is true (e.g., the MPM flag is 1), the index value is obtained, which represents the intra-prediction mode value of the current decoded block in the MPM.
[0256] In another example, when the MPM flag is true (e.g., the MPM flag is 1), the value of the second flag (e.g., the planar flag) is obtained. When the second flag is false (in one example, a false value of the second flag indicates that the intra-prediction mode of the current decoded block is not planar mode), the value of the index is obtained, which represents the intra-prediction mode value of the current decoded block in the MPM.
[0257] In one example, the syntax elements intra_luma_mpm_flag[x0][y0], intra_luma_mpm_idx[x0][y0], and intra_luma_mpm_remainder[x0][y0] represent the intra-prediction mode of the luminance sample. The array indices x0 and y0 represent the position (x0, y0) of the top-left luminance sample of the considered prediction block relative to the top-left luminance sample of the image. When intra_luma_mpm_flag[x0][y0] equals 1, the intra-prediction mode is inferred from the prediction units of adjacent intra-prediction blocks.
[0258] In one example, when the MPM flag is false (e.g., the MPM flag is 0), the value of the index is obtained, which represents the intra-prediction mode value of the current decoded block in non-MPM.
[0259] The position of the predicted sample within the current decoding block is obtained based on the slope of the intra-frame prediction mode. The position of the sample within the prediction block (e.g., the current decoding block) relative to the position of the top-left predicted sample is determined by a pair of integer values ( x p ,y p ) is defined, where, x p It is the horizontal offset of the predicted sample relative to the top-left predicted sample. y p This is the vertical offset of the predicted sample relative to the top-left predicted sample. The position of the top-left predicted sample is defined as follows: x p =0 ,y p =0.
[0260] To generate prediction samples from reference samples, the following steps are performed. Two intra-prediction mode ranges are defined. The first range of intra-prediction modes corresponds to the prediction in the vertical direction, and the second range corresponds to the mode in the horizontal direction. When the intra-prediction mode specified for a prediction block belongs to the first range, position (…) can also be used. x, y Addressing prediction sample blocks, where x set equals x p , y set equals y p When the intra-prediction mode specified for the prediction block belongs to the second range of intra-prediction modes, the position ( x, y Addressing prediction sample blocks, where x set equals y p , y set equals x p In some examples, the first range of the intra-prediction mode is defined as [34, 80]. The second range of the intra-prediction mode is defined as [–14, –1]∪[1, 33].
[0261] The inverse angle parameter invAngle is derived from intraPredAngle as follows:
[0262] invAngle = Round
[0263] Each intra-prediction mode has associated intra-prediction variables, also known as "intraPredAngle". This association is shown in Table 8-8.
[0264] The subpixel offset, denoted as "iFact" and also referred to as the "first subpixel offset value", is defined using the following equation:
[0265] iFact = (( y +1+refIdx) intraPredAngle)&31
[0266] In this equation, refIdx represents the offset of the reference sample set relative to the prediction block boundary. For the brightness component, this value can be obtained, for example, as shown below:
[0267]
[0268] The value of the syntax element “intra_luma_ref_idx” is indicated in the bitstream.
[0269] An embodiment of the process for obtaining prediction samples (as described in the VVC standard, JVET-O2001) is also provided, wherein, regardless of whether the intra-frame prediction mode is horizontal or vertical, the position (x, y) is always defined as x = x p and y=y p :
[0270] The values of the predicted samples predSamples[x][y] are derived as follows, where x = 0……nTbW–1, y = 0……nTbH–1:
[0271] – If predModeIntra is greater than or equal to 34, then use the following sequence of steps:
[0272] 1. The reference sample array ref[x] is specified as follows:
[0273] – The following applies:
[0274] ref[x] = p[–1–refIdx+x][–1–refIdx], where x = 0……nTbW+refIdx+1
[0275] – If intraPredAngle is less than 0, the master reference sample array is expanded as follows:
[0276] ref[x] = p[–1–refIdx][–1–refIdx+Min((x invAngle+256)>>9, nTbH)],
[0277] Where x = –nTbH……1
[0278] -otherwise,
[0279] ref[x] = p[–1–refIdx+x][–1–refIdx], where x = nTbW+2+refIdx……refW+refIdx
[0280] – The derivation of the additional sample ref[refW+refIdx+x] is as follows, where x = 1……(Max(1, nTbW / nTbH)) refIdx+2):
[0281] ref[refW+refIdx+x] = p[–1+refW][–1–refIdx]
[0282] 2. The values of the predicted samples predSamples[x][y] are derived as follows, where x = 0……nTbW–1, y = 0……nTbH–1:
[0283] – The derivation of the index variable iIdx and the multiplication factor iFact is as follows:
[0284] iIdx = (((y+1+refIdx) intraPredAngle)>>5)+refIdx
[0285] iFact = ((y+1+refIdx) intraPredAngle)&31
[0286] – If cIdx equals 0, then the following applies:
[0287] The interpolation filter coefficients fT[j] are derived as follows, where j = 0……3:
[0288] fT[j] = filterFlag?fG[iFact][j]:C[iFact][j]
[0289] The derivation of the values of the predicted samples predSamples[x][y] is as follows:
[0290] predSamples[x][y] = Clip1Y((( )+32)>>6)
[0291] Otherwise (cIdx is not equal to 0), the following applies based on the value of iFact:
[0292] – If iFact is not equal to 0, the value of the predicted sample predSamples[x][y] is derived as follows:
[0293] predSamples[x][y]=
[0294] ((32–iFact) ref[x+iIdx+1]+iFact ref[x+iIdx+2]+16)>>5
[0295] Otherwise, the value of the predicted sample predSamples[x][y] is derived as follows:
[0296] predSamples[x][y]= ref[x+iIdx+1]
[0297] - Otherwise (predModeIntra is less than 34), use the following sequence of steps:
[0298] 1. The reference sample array ref[x] is specified as follows:
[0299] – The following applies:
[0300] ref[x] = p[–1–refIdx][–1–refIdx+x], where x = 0……nTbH+refIdx+1
[0301] – If intraPredAngle is less than 0, the master reference sample array is expanded as follows:
[0302] ref[x] = p[–1–refIdx+Min((x invAngle+256)>>9, nTbW)][–1–refIdx],
[0303] Where x = –nTbW……–1
[0304] -otherwise,
[0305] ref[x] = p[–1–refIdx][–1–refIdx+x], where x = nTbH+2+refIdx……refH+refIdx
[0306] – The derivation of the additional sample ref[refH+refIdx+x] is as follows, where x = 1……(Max(1, nTbW / nTbH)) refIdx+2):
[0307] ref[refH+refIdx+x] = p[–1+refH][–1–refIdx]
[0308] 2. The values of the predicted samples predSamples[x][y] are derived as follows, where x = 0……nTbW–1, y = 0……nTbH–1:
[0309] – The derivation of the index variable iIdx and the multiplication factor iFact is as follows:
[0310] iIdx = (((x+1+refIdx) intraPredAngle)>>5)+refIdx
[0311] iFact = ((x+1+refIdx) intraPredAngle)&31
[0312] – If cIdx equals 0, then the following applies:
[0313] The interpolation filter coefficients fT[j] are derived as follows, where j = 0……3:
[0314] fT[j] = filterFlag?fG[iFact][j]:C[iFact][j]
[0315] The derivation of the values of the predicted samples predSamples[x][y] is as follows:
[0316] predSamples[x][y] = Clip1Y((( )+32)>>6)
[0317] Otherwise (cIdx is not equal to 0), the following applies based on the value of iFact:
[0318] – If iFact is not equal to 0, the value of the predicted sample predSamples[x][y] is derived as follows:
[0319] predSamples[x][y]=
[0320] ((32–iFact) ref[y+iIdx+1]+iFact ref[y+iIdx+2]+16)>>5
[0321] Otherwise, the value of the predicted sample predSamples[x][y] is derived as follows:
[0322] predSamples[x][y]= ref[y+iIdx+1].
[0323] S2004: Obtain the filtering coefficients based on the first sub-pixel offset value.
[0324] In one example, obtaining the filter coefficients based on the first sub-pixel offset value means obtaining the filter coefficients using a predefined lookup table and the first sub-pixel offset value. In one example, the first sub-pixel offset value is used as an index, and a predefined lookup table describes the mapping relationship between the filter coefficients and the sub-pixel offsets.
[0325] In one example, the predefined lookup table is described as follows:
[0326]
[0327] The "Subpixel Offset" column is defined at a subpixel resolution of 1 / 32. c 0 、c 1 、c 2 、c 3 represents the filter coefficient.
[0328] In another example, the predefined lookup table is described as follows:
[0329]
[0330] The "Subpixel Offset" column is defined at a subpixel resolution of 1 / 32. c 0 、c 1 、c 2 、c 3 represents the filter coefficient.
[0331] In another possible implementation, the derivation process of the interpolation filter coefficients used for intra-frame prediction and inter-frame prediction results in the coefficients of a 4-tap filter.
[0332] In one possible implementation, the interpolation filter coefficient derivation process is selected when the size of the primary reference edge used in intra-frame prediction is less than or equal to a threshold.
[0333] In one example, Gaussian or cubic filter coefficients are selected based on either the block's width or height. The decision to use width or height aligns with the decision made regarding the primary reference edge. When the intra-prediction mode value is greater than or equal to the diagonal mode value, the top edge of the reference sample is selected as the primary reference edge, and the width value is chosen to determine the interpolation filter being used. When the intra-prediction mode value is less than the diagonal mode value, the primary reference edge is selected from the left side of the block, and the height value controls the filter selection process. Specifically, if the length of the selected edge is less than or equal to 8 samples, a 4-tap cubic filter is used. If the length of the selected edge is greater than 8 samples, the interpolation filter is a 4-tap Gaussian filter.
[0334] For each intra-prediction mode, one value corresponds to one intra-prediction mode. Therefore, the primary reference edge can be selected using the value relationships (e.g., less than, equal to, or greater than) among the values of different intra-prediction modes.
[0335] Figure 12 This example shows the mode selection for both smaller and larger diagonal patterns (represented as 45°) in the case of a 32×4 block. Figure 12 As shown, if the value of the intra-prediction mode corresponding to the current decoded block is less than the value corresponding to the diagonal mode, the left (height) side of the current decoded block is selected as the primary reference side. In this case, the intra-prediction mode specified for the prediction block is horizontal, that is, the intra-prediction mode belongs to the second range of intra-prediction modes. Since the left side has 4 samples, which is less than the threshold (e.g., 8 samples), a cubic interpolation filter is selected.
[0336] If the value of the intra-prediction mode corresponding to the current decoded block is greater than or equal to the value corresponding to the diagonal mode, the top edge (width) of the current decoded block is selected as the master reference edge. In this case, the intra-prediction mode specified for the prediction block is vertical, i.e., the intra-prediction mode belongs to the first range of intra-prediction modes. Since the top edge has 32 samples, which is greater than the threshold (e.g., 8 samples), a Gaussian interpolation filter is selected.
[0337] In one example, if a third filter is selected, the predicted sample is further refined to be within an allowed range of values that are defined in the sequence parameter set (SPS) or derived from the bit depth of the selected component.
[0338] In one example, such as Figure 8 As shown, the "4-tap interpolation filter with 6-bit chroma coefficients" (also known as the "unified intra / inter frame filter") can be used for two processes: interpolation of intra-frame prediction samples and interpolation of inter-frame prediction samples.
[0339] Figure 9 An embodiment using this design is illustrated. In this implementation, the filtering module is implemented as a separate unit that participates in predicting chroma samples in motion compensation 906 and predicting luma and chroma samples when performing intra-frame prediction 907. In this implementation, a hardware filtering section (e.g., a 4-tap filter 904) is used for the intra-frame prediction and inter-frame prediction processes.
[0340] Another embodiment illustrates an implementation when reusing the LUT of the filter coefficients (see [link]). Figure 10 ). Figure 10 This is an exemplary implementation of a LUT based on reused coefficients. In this implementation, a hardware filtering module loads coefficients from the LUT stored in ROM. The switches shown during intra-frame prediction determine the type of filter to be used based on the selected main side length for intra-frame prediction processing.
[0341] In another example, Gaussian filter coefficients or cubic filter coefficients are selected based on the threshold.
[0342] In some examples, for blocks of sizes 32×4 and 4×32, the threshold mode used to determine the reference sample filter is... m T As it is Figure 16 The definition is shown.
[0343] The value of the intra-frame prediction angle for the threshold can be calculated using the following formula:
[0344] ,
[0345] in, W and H These are the block width and height, respectively.
[0346] In one example, the specifications for the INTRA_ANGULAR2……INTRA_ANGULAR66 intra-frame prediction modes are shown.
[0347] The input to this process is:
[0348] –Intra-frame prediction mode predModeIntra
[0349] – The variable refIdx represents the intra-frame prediction reference line index.
[0350] – The variable nTbW represents the transform block width.
[0351] – The variable nTbH represents the height of the transform block.
[0352] – The variable refW represents the reference sample width.
[0353] – The variable refH represents the height of the reference sample.
[0354] – The variable nCbW represents the decoding block width.
[0355] – The variable nCbH represents the decoded block height.
[0356] – The variable refFilterFlag represents the value of the reference filter flag.
[0357] – The variable cIdx represents the color component of the current block.
[0358] – Neighboring samples p[x][y], where x = –1–refIdx, y = –1–refIdx……refH–1, x = –refIdx……refW–1, y = –1–refIdx.
[0359] The output of this process is the predicted samples predSamples[x][y], where x = 0……nTbW–1, y = 0……nTbH–1.
[0360] Set the variable nTbS to (Log2(nTbW)+Log2(nTbH))>>1.
[0361] The derivation of the variable filterFlag is as follows:
[0362] - If one or more of the following conditions are true, filterFlag is set to 0.
[0363] –refFilterFlag equals 1
[0364] –refIdx is not equal to 0
[0365] –IntraSubPartitionsSplitType is not equal to ISP_NO_SPLIT
[0366] –Otherwise, the following applies:
[0367] – Set the variable minDistVerHor to Min(Abs(predModeIntra–50), Abs(predModeIntra–18)).
[0368] The variable intraHorVerDistThres[nTbS] is specified in Table 8-7.
[0369] The derivation of the variable filterFlag is as follows:
[0370] – If minDistVerHor is greater than intraHorVerDistThres[nTbS] and refFilterFlag is equal to 0, then filterFlag is set to 1.
[0371] Otherwise, filterFlag is set to 0.
[0372] Table 8-7 - Specifications of intraHorVerDistThres[nTbS] for various transform block sizes nTbS
[0373]
[0374] Table 8-8 is a mapping table between predModeIntra and the angle parameter intraPredAngle.
[0375] Table 8-8 – Specifications of intraPredAngle
[0376]
[0377] The inverse angle parameter invAngle is derived from intraPredAngle as follows:
[0378] invAngle = Round .
[0379] Specify the interpolation filter coefficients fC[phase][j] and fG[phase][j] in Table 8-9, where phase = 0……31 and j = 0……3.
[0380] Table 8-9 - Specifications of interpolation filter coefficients fC and fG
[0381]
[0382] S2005: Obtain intra-frame prediction sample values based on the filtering coefficients.
[0383] Intra-frame predicted sample values are used for the luminance component of the current decoded block.
[0384] In one embodiment, intra-frame predicted samples are computed by performing convolution operations using coefficients selected from Table 3, based on sub-pixel offset and filter type, as follows:
[0385]
[0386] In this equation, "Indicates a bitwise right shift operation, ci This represents one of the coefficients in a set of filter coefficients derived using the first sub-pixel offset value. s ( x ) indicates position ( x , y Intra-predicted samples at location ) ref i+x Let represent a set of reference samples, where ref 1+x Located in the location ( x r , y r At this location, the reference sample position is defined as follows:
[0387] xr = (((y+1+refIdx) intraPredAngle)>>5)+refIdx;
[0388] yr = –1–refIdx.
[0389] In one example, the values of the predicted samples predSamples[x][y] are derived as follows, where x = 0……nTbW–1, y = 0……nTbH–1:
[0390] – If predModeIntra is greater than or equal to 34, then use the following sequence of steps:
[0391] 3. The reference sample array ref[x] is specified as follows:
[0392] – The following applies:
[0393] ref[x] = p[–1–refIdx+x][–1–refIdx], where x = 0……nTbW+refIdx+1
[0394] – If intraPredAngle is less than 0, the master reference sample array is expanded as follows:
[0395] ref[x] = p[–1–refIdx][–1–refIdx+Min((x invAngle+256)>>9, nTbH)],
[0396] Where x = –nTbH……1
[0397] -otherwise,
[0398] ref[x] = p[–1–refIdx+x][–1–refIdx], where x = nTbW+2+refIdx……refW+refIdx
[0399] – The derivation of the additional sample ref[refW+refIdx+x] is as follows, where x = 1……(Max(1, nTbW / nTbH)) refIdx+2):
[0400] ref[refW+refIdx+x] = p[–1+refW][–1–refIdx]
[0401] 4. The values of the predicted samples predSamples[x][y] are derived as follows, where x = 0……nTbW–1, y = 0……nTbH–1:
[0402] – The derivation of the index variable iIdx and the multiplication factor iFact is as follows:
[0403] iIdx = (((y+1+refIdx) intraPredAngle)>>5)+refIdx
[0404] iFact = ((y+1+refIdx) intraPredAngle)&31
[0405] – If cIdx equals 0, then the following applies:
[0406] The interpolation filter coefficients fT[j] are derived as follows, where j = 0……3:
[0407] fT[j] = filterFlag?fG[iFact][j]:C[iFact][j]
[0408] The derivation of the predicted sample values predSamples[x][y] is as follows:
[0409] predSamples[x][y] = Clip1Y((( )+32)>>6)
[0410] Otherwise (cIdx is not equal to 0), the following applies based on the value of iFact:
[0411] – If iFact is not equal to 0, the value of the predicted sample predSamples[x][y] is derived as follows:
[0412] predSamples[x][y]=
[0413] ((32–iFact) ref[x+iIdx+1]+iFact ref[x+iIdx+2]+16)>>5
[0414] Otherwise, the value of the predicted sample predSamples[x][y] is derived as follows:
[0415] predSamples[x][y]= ref[x+iIdx+1]
[0416] - Otherwise (predModeIntra is less than 34), use the following sequence of steps:
[0417] 3. The reference sample array ref[x] is specified as follows:
[0418] – The following applies:
[0419] ref[x] = p[–1–refIdx][–1–refIdx+x], where x = 0……nTbH+refIdx+1
[0420] – If intraPredAngle is less than 0, the master reference sample array is expanded as follows:
[0421] ref[x] = p[–1–refIdx+Min((x invAngle+256)>>9, nTbW)][–1–refIdx],
[0422] Where x = –nTbW……–1
[0423] -otherwise,
[0424] ref[x] = p[–1–refIdx][–1–refIdx+x], where x = nTbH+2+refIdx……refH+refIdx
[0425] – The derivation of the additional sample ref[refH+refIdx+x] is as follows, where x = 1……(Max(1, nTbW / nTbH)) refIdx+2):
[0426] ref[refH+refIdx+x] = p[–1+refH][–1–refIdx]
[0427] 4. The values of the predicted samples predSamples[x][y] are derived as follows, where x = 0……nTbW–1, y = 0……nTbH–1:
[0428] – The derivation of the index variable iIdx and the multiplication factor iFact is as follows:
[0429] iIdx = (((x+1+refIdx) intraPredAngle)>>5)+refIdx
[0430] iFact = ((x+1+refIdx) intraPredAngle)&31
[0431] – If cIdx equals 0, then the following applies:
[0432] The interpolation filter coefficients fT[j] are derived as follows, where j = 0……3:
[0433] fT[j] = filterFlag?fG[iFact][j]:C[iFact][j]
[0434] The derivation of the predicted sample values predSamples[x][y] is as follows:
[0435] predSamples[x][y] = Clip1Y((( )+32)>>6)
[0436] Otherwise (cIdx is not equal to 0), the following applies based on the value of iFact:
[0437] – If iFact is not equal to 0, the value of the predicted sample predSamples[x][y] is derived as follows:
[0438] predSamples[x][y]=
[0439] ((32–iFact) ref[y+iIdx+1]+iFact ref[y+iIdx+2]+16)>>5
[0440] Otherwise, the value of the predicted sample predSamples[x][y] is derived as follows:
[0441] predSamples[x][y]= ref[y+iIdx+1].
[0442] S2006: Obtain the second sub-pixel offset value based on the motion information of the current decoding block.
[0443] Motion information for the current decoded block is indicated in the bitstream. Motion information may include motion vectors and other syntax elements used in inter-frame prediction.
[0444] In one example, the first subpixel offset value can be equal to the second subpixel offset value. In another example, the first subpixel offset value and the second subpixel offset value can be different.
[0445] S2007: Obtain the filtering coefficients based on the second sub-pixel offset value.
[0446] In a possible implementation, the derivation process for the interpolation filter coefficients used in inter-frame prediction is performed, as is the derivation of the predefined lookup table used in intra-frame prediction. In this example, obtaining the filter coefficients based on the first sub-pixel offset value means obtaining the filter coefficients based on the predefined lookup table and the second sub-pixel offset value. In one example, the second sub-pixel offset value is used as an index, and the predefined lookup table describes the mapping relationship between the filter coefficients and the sub-pixel offsets.
[0447] In one example, the predefined lookup table is described as follows:
[0448]
[0449] The "Subpixel Offset" column is defined at a subpixel resolution of 1 / 32. c 0 、c 1 、c 2 、c 3 represents the filter coefficient.
[0450] In another example, the predefined lookup table is described as follows:
[0451]
[0452] The "Subpixel Offset" column is defined at a subpixel resolution of 1 / 32. c 0 、c 1 、c 2 、c 3 represents the filter coefficient.
[0453] When the sub-pixel offset value is equal to 0, no filtering coefficients are needed to obtain inter-frame prediction samples. In a first alternative embodiment, the following steps can be performed:
[0454] predSampleLX C = >>shift1
[0455] In a second alternative embodiment, the following steps may be performed:
[0456] predSampleLX C = >>shift1
[0457] In a third alternative embodiment, the following steps may be performed:
[0458] The derivation of the sample array temp[n] is as follows, where n = 0……3:
[0459] temp[n] = >>shift1
[0460] –The derivation of the chromaticity sample prediction value predSampleLXC is as follows:
[0461] predSampleLX C = (f C [ [0] temp[0]+
[0462] f C [ [1] temp[1]+
[0463] f C [ [2] temp[2]+
[0464] f C [ [3] temp[3])>>shift2
[0465] In the three alternative embodiments described above, yFrac C and xFrac C Set to 0, fC[0][0] = 0, fC[0][1] = 64, fC[0][2] = 0, fC[0][3] = 0.
[0466] In another possible implementation, the derivation process of the interpolation filter coefficients used for intra-frame prediction and inter-frame prediction results in the coefficients of a 4-tap filter.
[0467] S2008: Obtain inter-frame prediction sample values based on the filtering coefficients.
[0468] In a possible implementation, the inter-frame predicted sample values are used for the chroma component of the current decoding block.
[0469] In one example, the chromaticity sample interpolation process is disclosed.
[0470] The input to this process is:
[0471] – Chromaticity position in the entire sample cell (xInt) C yInt C ),
[0472] Chromaticity position in –1 / 32 fractional sample unit (xFrac) C yFrac C ),
[0473] – The chromaticity position (xSbIntC, ySbIntC) in the full sample cell specifies the top-left sample of the boundary block used for reference sample filling relative to the top-left chromaticity sample of the reference image.
[0474] – The variable sbWidth specifies the width of the current child block.
[0475] – The variable sbHeight specifies the height of the current child block.
[0476] –Color reference sample array refPicLX C .
[0477] The output of this process is the predicted chromaticity sample value, predSampleLX. C .
[0478] The derivation of variables shift1, shift2, and shift3 is as follows:
[0479] – Set the variable shift1 to Min(4, BitDepth) C –8), set variable shift2 to 6, and variable shift3 to Max(2, 14 – BitDepth). C ).
[0480] – Change variable picW C Set it to pic_width_in_luma_samples / SubWidthC, and set the variable picH C Set it to pic_height_in_luma_samples / SubHeightC.
[0481] Table 8-13 specifies the chromaticity interpolation filter coefficients f for each 1 / 32 fraction sample position p. C [p] equals xFrac C or yFrac C .
[0482] The variable xOffset is set to (sps_ref_wraparound_offset_minus1+1). MinCbSizeY) / SubWidthC.
[0483] For i = 0……3, the chromaticity position (xInt) in the entire sample cell i yInt i The derivation is as follows:
[0484] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[0485] xInt i = Clip3(SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xInt L +i)
[0486] yInt i = Clip3(SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yInt L +i)
[0487] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following applies:
[0488] xInt i = Clip3(0, picW C –1, sps_ref_wraparound_enabled_flag?ClipH(xOffset, picW C , xInt C +i–1): xInt C +i–1)
[0489] yInt i = Clip3(0, picH C –1, yInt C +i–1)
[0490] For i = 0……3, the chromaticity positions in the full sample cell (xInti, yInti) are further modified as follows:
[0491] xInti = Clip3(xSbIntC–1, xSbIntC+sbWidth+2, xInt i )
[0492] yInt i = Clip3(ySbIntC–1, ySbIntC+sbHeight+2, yInt i )
[0493] Chromaticity sample prediction value predSampleLX C The derivation is as follows:
[0494] –If xFrac C and yFrac C If both are equal to 0, then predSampleLX C The value is derived as follows:
[0495] predSampleLX C = refPicLX C [xInt1][yInt1]< <shift3
[0496] Otherwise, if xFrac C Not equal to 0, yFrac C If it equals 0, then predSampleLX C The value is derived as follows:
[0497] predSampleLX C = >>shift1
[0498] Otherwise, if xFrac C and yFrac C If both are equal to 0, then predSampleLX C The value is derived as follows:
[0499] predSampleLX C = >>shift1
[0500] Otherwise, if xFrac C and yFrac C If none of them are equal to 0, then predSampleLX C The value is derived as follows:
[0501] The derivation of the sample array temp[n] is as follows, where n = 0……3:
[0502] temp[n] = >>shift1
[0503] –The derivation of the chromaticity sample prediction value predSampleLXC is as follows:
[0504] predSampleLX C = (f C [ [0] temp[0]+
[0505] f C [ [1] temp[1]+
[0506] f C [ [2] temp[2]+
[0507] f C [ [3] temp[3])>>shift2.
[0508] A decoder includes processing circuitry for performing the methods described above.
[0509] In this invention, a computer program product is disclosed, which includes program code for performing the above-described method.
[0510] In this invention, a decoder for decoding video data is disclosed, the decoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing a program for execution by the processor, wherein, when the processor executes the program, the decoder is configured to perform the above-described method.
[0511] Figure 18This is a schematic diagram of a network device 1300 provided in an embodiment of the present invention. The network device 1300 is suitable for implementing the disclosed embodiments described herein. The network device 1300 includes: an input port 1310 and a receiver unit (Rx) 1320 for receiving data; a processor, logic unit, or central processing unit (CPU) 1330 for processing data; a transmitter unit (Tx) 1340 and an output port 1350 for transmitting data; and a memory 1360 for storing data. The network device 1300 may further include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the input port 1310, receiver unit 1320, transmitter unit 1340, and output port 1350, serving as output or input points for optical or electrical signals.
[0512] Processor 1330 is implemented in both hardware and software. Processor 1330 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 1330 communicates with ingress port 1310, receiver unit 1320, transmitter unit 1340, egress port 1350, and memory 1360. Processor 1330 includes a decoding module 1370. Decoding module 1370 implements the embodiments disclosed above. For example, decoding module 1370 implements, processes, prepares, or provides various networking functions. Therefore, including decoding module 1370 significantly improves the functionality of network device 1300 and enables transitions between different states of network device 1300. Alternatively, decoding module 1370 can be implemented with instructions stored in memory 1360 and executed by processor 1330.
[0513] Memory 1360 includes one or more disks, tape drives, and solid-state drives, which can be used as overflow data storage devices to store programs when an executable program is selected, and to store instructions and data read during program execution. Memory 1360 can be volatile and / or non-volatile, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0514] Figure 19 This is a block diagram of an apparatus 1500 that can be used to implement various embodiments. Apparatus 1500 may be... Figure 1 The source device 102 shown, or Figure 2 The video encoder 200 shown, or Figure 1 The target device 104 shown, or Figure 3 The video decoder 300 is shown. Additionally, device 1100 may include one or more of the described elements. In some embodiments, device 1100 is equipped with one or more input / output devices, such as speakers, microphones, mice, touchscreens, keypads, keyboards, printers, displays, etc. Device 1500 may include one or more central processing units (CPUs) 1510, memory 1520, mass storage 1530, video adapters 1540, and I / O interfaces 1560 connected to a bus. The bus is one or more of several bus architectures of any type, including memory buses or memory controllers, peripheral buses, video buses, etc.
[0515] CPU 1510 may have any type of electronic data processor. Memory 1520 may have or may be any type of system memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), read-only memory (ROM), or combinations thereof. In one embodiment, memory 1520 may include ROM used at power-on and DRAM used to store programs and data during program execution. In one embodiment, memory 1520 is non-transitory memory. Mass storage 1530 includes any type of storage device that stores data, programs, and other information and enables data, programs, and other information to be accessed via a bus. For example, mass storage 1530 includes one or more of solid-state drives, hard disk drives, disk drives, optical disk drives, etc.
[0516] Video adapter 1540 and I / O interface 1560 provide interfaces for coupling external input and output devices to device 1100. For example, device 1100 may provide an SQL command interface to a client. As shown, examples of input and output devices include any combination of a monitor 1590 coupled to video adapter 1540 and a mouse / keyboard / printer 1570 coupled to I / O interface 1560. Other devices may be coupled to device 1100, and additional or fewer interface cards may be used. For example, a serial interface card (not shown) may be used to provide a serial interface for a printer.
[0517] The device 1100 also includes one or more network interfaces 1550, or one or more networks 1580, wherein the network interface 1550 includes a wired link such as an Ethernet cable, and / or a wireless link for access nodes. The network interface 1550 enables the device 1100 to communicate with remote units via the network 1580. For example, the network interface 1550 can provide communication with a database. In one embodiment, the device 1100 is coupled to a local area network (LAN) or a wide area network (WAN) for data processing and communication with remote devices such as other processing units, the Internet, or remote storage facilities.
[0518] A piecewise linear approximation is introduced to compute the weighting coefficients needed to predict pixels within a given block. This piecewise linear approximation significantly reduces the computational complexity of distance-weighted prediction mechanisms compared to direct weighting coefficient calculation, and also helps achieve higher accuracy in weighting coefficient values compared to existing simplification techniques.
[0519] The embodiments can be applied to other bidirectional and position-dependent intra-prediction techniques (e.g., different modifications of PDPC) and mechanisms that use weighted coefficients that depend on the distance from one pixel to another to blend different parts of an image (e.g., some blending methods in image processing).
[0520] The subject matter and operations described in this invention can be implemented in digital electronic circuits, computer software, firmware, or hardware, including implementation in the structures disclosed herein and their structural equivalents, or in combinations thereof. The subject matter described in this invention can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded in a computer storage medium for execution by a data processing device or to control the operation of a data processing device. Alternatively or additionally, the program instructions can be encoded on artificially generated propagating signals (e.g., machine-generated electrical, optical, or electromagnetic signals) to generate information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium, such as a computer-readable medium, can be or is included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access storage array or device, or a combination thereof. Furthermore, although the computer storage medium is not a propagating signal, it can be a source or destination of computer program instructions encoded into an artificially generated propagating signal. The computer storage medium can also be or be included in one or more separate physical and / or non-transient components or media (e.g., multiple CDs, disks, or other storage devices).
[0521] In some implementations, the operations described in this invention can be implemented as a hosted service provided on a server in a cloud computing network. For example, computer-readable storage media can be logically grouped and accessed within a cloud computing network. Servers in a cloud computing network may include cloud computing platforms for providing cloud services. Without departing from the scope of this invention, the terms "cloud," "cloud computing," and "cloud-based" may be used interchangeably as appropriate. Cloud services can be hosted services provided by servers and delivered over a network to client platforms to enhance, supplement, or replace applications running locally on client computers. Circuits can use cloud services to quickly receive software upgrades, applications, and other resources that would otherwise take a long time to arrive at the circuit.
[0522] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, declarative languages, or programming languages, and can be deployed in any form, such as as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment. A computer program may (but does not necessarily) correspond to a file in a file system. A program may be stored as part of a file that includes other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the related program, or multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on a single computer, or deployed to execute on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0523] The processes and logic flows described in this invention can be executed by one or more programmable processors that execute one or more computer programs to perform actions by manipulating input data and generating outputs. The processes and logic flows can also be executed by dedicated logic circuits such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), and the apparatus can also be implemented as said dedicated logic circuits.
[0524] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for performing actions according to instructions and one or more storage devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices (e.g., disks, magneto-optical disks, or optical disks) for storing data, or is operatively coupled to one or more mass storage devices for storing data to receive data from and / or transfer data to the mass storage devices. However, a computer does not necessarily have to have such devices. Furthermore, computers can be embedded in other devices, such as mobile phones, personal digital assistants (PDAs), mobile audio or video players, game consoles, Global Positioning System (GPS) receivers, or portable storage devices (e.g., universal serial bus (USB) flash drives), etc. Suitable devices for storing computer program instructions and data include various forms of non-volatile memory, media, and storage devices, such as semiconductor storage devices including EPROM, EEPROM, and flash memory devices; hard disks such as internal or removable hard disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory can be supplemented by or incorporated into dedicated logic circuitry.
[0525] While this invention includes details of many specific implementations, these should not be construed as limiting the scope of any implementation or the scope of the claims, but rather as descriptions of the features of a particular implementation. In the context of a single implementation, certain features described herein may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually in multiple implementations or in any suitable sub-combination. Furthermore, although features may be described above as being implemented in certain combinations, even initially claimed, in some cases one or more features may be removed from the claimed combination, and the claimed combination may be for sub-combinations or variations thereof.
[0526] Similarly, although the accompanying figures depict operations in a specific order, this should not be construed as requiring such operations to be performed in the specific order shown or sequentially, or requiring the execution of all the operations shown to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous. Furthermore, the separation of the various system components in the above implementations should not be construed as requiring such separation in all implementations. It should be understood that the described program components and systems can typically be integrated together into a single software product or encapsulated within multiple software products.
[0527] Therefore, specific implementations of this subject matter have been described. Other implementations are within the scope of the following claims. In some cases, the actions described in the claims can be performed in a different order and the desired result can still be achieved. Furthermore, the processes described in the drawings do not necessarily require execution in the specific order shown to achieve the desired result. In some implementations, multitasking and parallel processing can be advantageous.
[0528] While several embodiments have been provided in this invention, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the invention. The examples of this invention should be considered illustrative rather than restrictive, and the invention is not limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0529] Furthermore, without departing from the scope of the invention, the technologies, systems, subsystems, and methods described and illustrated as independent or separate in the various embodiments may be combined or integrated with other systems, modules, technologies, or methods. Other items shown or described as coupled to each other, or directly coupled, or communicating with each other may be indirectly coupled or communicated through some interface, device, or intermediate component, electrically, mechanically, or otherwise. Other examples of variations, substitutions, and modifications can be determined by those skilled in the art and may be exemplified without departing from the spirit and scope of the disclosure herein.
[0530] The following description of other embodiments of the invention is provided, wherein the numbering of the embodiments does not necessarily match the numbering used above.
[0531] Example 1. An intra-frame prediction method, wherein the method includes:
[0532] The interpolation filter used for the chroma component is used as the interpolation filter for intra-frame prediction of the block.
[0533] Example 2. According to the method described in Example 1, the lookup table for the interpolation filter used for the chroma component is the same as the lookup table for the interpolation filter used for intra-frame prediction.
[0534] Example 3. According to the method described in Example 1, the lookup table for the interpolation filter used for the chroma component is different from the lookup table for the interpolation filter used for intra-frame prediction.
[0535] Example 4. The method according to any one of Examples 1 to 3, wherein the interpolation filter is a 4-tap filter.
[0536] Example 5. The method according to any one of Examples 1 to 4, wherein the lookup table for the interpolation filter used for the chromaticity component is:
[0537]
[0538] Example 6. An intra-frame prediction method, wherein the method includes:
[0539] Select an interpolation filter from a set of interpolation filters used for intra-frame prediction of the block.
[0540] Example 7. According to the method described in Example 6, the set of interpolation filters includes a Gaussian filter and a cubic filter.
[0541] Example 8. The method according to Example 6 or 7, wherein the lookup table of the selected interpolation filter is the same as the lookup table of the interpolation filter used for the chromaticity component.
[0542] Example 9. The method according to any one of Examples 6 to 8, wherein the selected interpolation filter is a 4-tap filter.
[0543] Example 10. The method according to any one of Examples 6 to 9, wherein the selected interpolation filter is a cubic filter.
[0544] Example 11. The method according to any one of Examples 6 to 10, wherein the lookup table of the selected interpolation filter is:
[0545]
[0546] Example 12. An encoder including processing circuitry for performing the method according to any one of Examples 1 to 11.
[0547] Example 13. A decoder including processing circuitry for performing the method according to any one of Examples 1 to 11.
[0548] Example 14. A computer program product comprising program code for performing the method according to any one of Examples 1 to 11.
[0549] Example 15. A decoder, comprising:
[0550] One or more processors;
[0551] A non-transitory computer-readable storage medium is coupled to the processor and stores a program executed by the processor, wherein, when the processor executes the program, the decoder is configured to execute the method according to any one of embodiments 1 to 11.
[0552] Example 16. An encoder, comprising:
[0553] One or more processors;
[0554] A non-transitory computer-readable storage medium is coupled to the processor and stores a program executed by the processor, wherein, when the processor executes the program, the encoder is configured to perform the method according to any one of embodiments 1 to 11.
[0555] Example 17. A video decoding method, characterized in that the method includes:
[0556] - Inter-frame prediction processing for the first block, wherein the inter-frame prediction processing includes sub-pixel interpolation filtering of samples of the reference block;
[0557] - The second block of intra-frame prediction processing, wherein the intra-frame prediction processing includes sub-pixel interpolation filtering of reference samples;
[0558] The method further includes:
[0559] - Based on the sub-pixel offset between the integer reference sample position and the fractional reference sample position, interpolation filter coefficients for the sub-pixel interpolation filter are selected, wherein for the same sub-pixel offset, the same interpolation filter coefficients are selected for intra-frame prediction processing and inter-frame prediction processing.
[0560] Example 18. The method according to Example 17, characterized in that the (same) selected filtering coefficients are used to perform the sub-pixel interpolation filtering on the chroma samples for inter-frame prediction processing; the (same) selected filtering coefficients are used to perform the sub-pixel interpolation filtering on the luminance samples for intra-frame prediction processing.
[0561] Example 19. The method according to Example 17 or 18, wherein the inter-frame prediction processing is intra-block copying processing.
[0562] Example 20. The method according to any one of Examples 17-19, characterized in that the interpolation filter coefficients used for inter-frame prediction processing and intra-frame prediction processing are obtained from a look-up table (LUT).
[0563] Example 21. The method according to any one of Examples 17-20, characterized in that a 4-tap filter is used for the sub-pixel interpolation filtering.
[0564] Example 22. The method according to any one of Examples 17-21, characterized in that selecting the interpolation filter coefficients includes: selecting the interpolation filter coefficients based on the following relationship between sub-pixel offsets and interpolation filter coefficients:
[0565]
[0566] Wherein, the sub-pixel offset is defined with a 1 / 32 sub-pixel resolution, and c0 to c3 represent the interpolation filtering coefficients.
[0567] Example 23. The method according to any one of Examples 17-22, characterized in that selecting the interpolation filter coefficients includes: selecting the interpolation filter coefficients for fractional positions based on the following relationship between sub-pixel offsets and interpolation filter coefficients:
[0568]
[0569] Wherein, the sub-pixel offset is defined with a 1 / 32 sub-pixel resolution, and c0 to c3 represent the interpolation filtering coefficients.
[0570] Example 24. An encoder, characterized in that it includes processing circuitry for performing the method according to any one of Examples 17 to 23.
[0571] Example 25. A decoder, characterized in that it includes processing circuitry for performing the method according to any one of Examples 17 to 23.
[0572] Example 26. A computer program product, characterized in that it includes program code for performing the method according to any one of Examples 17 to 23.
[0573] Example 27. A decoder, characterized in that it comprises:
[0574] One or more processors;
[0575] A non-transitory computer-readable storage medium is coupled to the processor and stores a program executed by the processor, wherein, when the processor executes the program, the decoder is configured to perform the method of any one of embodiments 17 to 23.
[0576] Example 28. An encoder, characterized in that it comprises:
[0577] One or more processors;
[0578] A non-transitory computer-readable storage medium is coupled to the processor and stores a program executed by the processor, wherein, when the processor executes the program, the encoder is configured to perform the method of any one of embodiments 17 to 23.
[0579] In one embodiment, a video decoding method is disclosed, the method comprising:
[0580] The inter-frame prediction process for the block includes sub-pixel interpolation filters applied to the luminance and chrominance samples of the reference block (e.g., one or usually several filters can be defined for MC interpolation).
[0581] The intra-frame prediction process for a block includes sub-pixel interpolation filters applied to luma and chroma reference samples (e.g., one or usually several filters can be defined for intra-frame reference sample interpolation).
[0582] Specifically, the subpixel interpolation filter is selected based on the subpixel offset between the reference sample position and the interpolation sample position. For equal subpixel offsets in intra-frame prediction and inter-frame prediction processes, the filter in the intra-frame prediction process (e.g., one or more filters can be used for reference intra-frame sample interpolation) is selected to be the same as the filter used for inter-frame prediction processing.
[0583] In another embodiment, a filter for the intra-prediction process for a given sub-pixel offset is selected from a set of filters (e.g., one or more filters may be used for MC interpolation), wherein one of the filters in the set of filters is the same as the filter used for the inter-prediction process.
[0584] In another embodiment, the filter applied to the chroma samples during inter-frame prediction is the same as the filter applied to the luma and chroma reference samples during intra-frame prediction.
[0585] In another embodiment, the filters applied to the luminance and chrominance samples during inter-frame prediction are the same as those applied to the luminance and chrominance reference samples during intra-frame prediction.
[0586] In another embodiment, if the size of the main reference edge used in the intra-frame prediction process is less than a threshold, the filter for the intra-frame prediction process is selected to be the same as the filter used for the inter-frame prediction process.
[0587] In another embodiment, the edge size threshold is 16 samples.
[0588] In another embodiment, the inter-frame prediction process is an intra-block copying process.
[0589] In another embodiment, the filters used for inter-frame prediction and intra-frame prediction processes are finite impulse response filters, and their coefficients are obtained from a lookup table.
[0590] In another embodiment, the interpolation filter used in the intra-frame prediction process is a 4-tap filter.
[0591] In another embodiment, the filter coefficients depend on the sub-pixel offset, as follows:
[0592]
[0593] The "Subpixel Offset" column is defined with a subpixel resolution of 1 / 32.
[0594] In another embodiment, the set of filters includes a Gaussian filter and a cubic filter.
[0595] In another embodiment, the encoder includes processing circuitry for performing the methods described above.
[0596] In another embodiment, the decoder includes processing circuitry for performing the methods described above.
[0597] In another embodiment, a computer program product includes program code for performing the methods described above.
[0598] In another embodiment, the decoder includes: one or more processors; a non-transitory computer-readable storage medium coupled to the processors and storing a program for execution by the processors, wherein, when the processors execute the program, the decoder is configured to perform the methods described above.
[0599] In another embodiment, the encoder includes: one or more processors; a non-transitory computer-readable storage medium coupled to the processors and storing a program for execution by the processors, wherein, when the processors execute the program, the encoder is configured to perform the above-described encoding method.
Claims
1. A video encoding device, characterized in that, The device includes: Prediction processing unit and filtering unit; The prediction processing unit is used for inter-frame prediction processing of the first block, wherein the inter-frame prediction processing includes sub-pixel interpolation filtering of samples of the reference block. The prediction processing unit is further configured to perform intra-prediction processing on the second block according to the intra-prediction mode of the second block, wherein the intra-prediction processing includes sub-pixel interpolation filtering of reference samples, and the intra-prediction mode of the second block is one of the following intra-prediction modes: planar mode, DC mode, and angle mode, wherein the index of the planar mode is 0, the index of the DC mode is 1, and the index of the angle mode is 2 to 66. The filtering unit is used to select interpolation filtering coefficients for the sub-pixel interpolation filtering based on the sub-pixel offset between the integer reference sample position and the fractional reference sample position. For the same sub-pixel offset, the same interpolation filtering coefficients are selected for intra-frame prediction processing and inter-frame prediction processing. The selected interpolation filtering coefficients are used to perform the sub-pixel interpolation filtering on the chroma samples for inter-frame prediction processing, and the selected interpolation filtering coefficients are used to perform the sub-pixel interpolation filtering on the luminance samples for intra-frame prediction processing.
2. The apparatus according to claim 1, characterized in that, The interpolation filter coefficients used for inter-frame prediction processing and intra-frame prediction processing are obtained from a look-up table (LUT).
3. The apparatus according to claim 1 or 2, characterized in that, 4. Tap filters are used for the sub-pixel interpolation filtering.
4. The apparatus according to claim 3, characterized in that, The filtering unit is used to: select the interpolation filtering coefficients based on the following relationship between sub-pixel offsets and interpolation filtering coefficients: Wherein, the sub-pixel offset is defined with a 1 / 32 sub-pixel resolution, and c0 to c3 represent the interpolation filtering coefficients.
5. The apparatus according to claim 3, characterized in that, The filtering unit is used to: select the interpolation filtering coefficients for the fractional position based on the following relationship between sub-pixel offsets and interpolation filtering coefficients: Wherein, the sub-pixel offset is defined with a 1 / 32 sub-pixel resolution, and c0 to c3 represent the interpolation filtering coefficients.
6. A video encoding method, characterized in that, The method includes: The first block of inter-frame prediction processing includes sub-pixel interpolation filtering of samples from the reference block. Intra-prediction processing is performed on the second block according to the intra-prediction mode of the second block. The intra-prediction processing includes sub-pixel interpolation filtering of the reference sample. The intra-prediction mode of the second block is one of the following intra-prediction modes: planar mode, DC mode, and angle mode. The index of the planar mode is 0, the index of the DC mode is 1, and the index of the angle mode is 2 to 66. The method further includes: Based on the sub-pixel offset between the integer reference sample position and the fractional reference sample position, interpolation filtering coefficients for the sub-pixel interpolation filtering are selected. For the same sub-pixel offset, the same interpolation filtering coefficients are selected for intra-frame prediction processing and inter-frame prediction processing. The selected interpolation filtering coefficients are used to perform the sub-pixel interpolation filtering on the chroma samples for inter-frame prediction processing, and the selected interpolation filtering coefficients are used to perform the sub-pixel interpolation filtering on the luminance samples for intra-frame prediction processing.
7. The method according to claim 6, characterized in that, The interpolation filter coefficients used for inter-frame prediction processing and intra-frame prediction processing are obtained from a look-up table (LUT).
8. The method according to claim 6 or 7, characterized in that, 4. Tap filters are used for the sub-pixel interpolation filtering.
9. The method according to claim 8, characterized in that, The selection of the interpolation filter coefficients includes: selecting the interpolation filter coefficients based on the following relationship between sub-pixel offsets and interpolation filter coefficients: Wherein, the sub-pixel offset is defined with a 1 / 32 sub-pixel resolution, and c0 to c3 represent the interpolation filtering coefficients.
10. The method according to claim 8, characterized in that, The selection of the interpolation filter coefficients includes: selecting the interpolation filter coefficients for the fractional position based on the following relationship between sub-pixel offsets and interpolation filter coefficients: Wherein, the sub-pixel offset is defined with a 1 / 32 sub-pixel resolution, and c0 to c3 represent the interpolation filtering coefficients.
11. A video data encoder, characterized in that, The video data encoder includes: A memory for storing the video data in the form of a bitstream; An encoder for performing the method according to any one of claims 6-10.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program and a video stream, the computer program being executed by one or more processors to implement the method as described in any one of claims 6-10 to obtain the video stream.
13. A method for transmitting a video stream, characterized in that, Perform the method as described in any one of claims 6-10 to obtain the video stream and transmit the video stream.
14. A method for storing a video stream, characterized in that, Perform the method as described in any one of claims 6-10 to obtain the video stream and store the video stream.
15. A method for transmitting a video stream, characterized in that, Perform the method as described in any one of claims 6-10 to obtain the video stream, and send the video stream to the destination device.
16. A chip, characterized in that, The chip includes processing circuitry for performing the method according to any one of claims 6 to 10.
Citation Information
Patent Citations
Method of video coding using prediction based on intra picture block copy
CN106416243A
Amvp and merge candidate list derivation for intra bc and inter prediction unification
CN106797477A
Method and apparatus for interpolation filtering for intra- and inter-prediction in video coding
CN112425169A
Image encoding / decoding apparatus and method to which filter selection by precise units is applied
US20140192876A1