Method and apparatus for switching interpolation filters of intra-frame reference samples
By selecting different reference sample filters based on the block aspect ratio and intra-frame prediction mode during video coding, the problem of low coding efficiency for elongated blocks in existing technologies is solved, achieving more efficient video coding results.
Patent Information
- Application Number
- CN202411579743.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-23
- Filing Date
- 2019-09-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2039-09-23
AI Technical Summary
Existing video coding techniques have flaws in the selection of interpolation filters and the design of the primary reference side when processing long and thin blocks, causing predicted pixels to be obtained from the longer side reference samples, which affects coding efficiency.
By independently examining the width and height of the block, different reference sample filters are selected. Appropriate interpolation filters are selected based on the block's aspect ratio and intra-prediction mode. The intra-prediction mode is thresholded to improve the selection process of the reference sample filters.
It improves the coding efficiency of video coding, especially the coding performance of long and thin blocks, reduces computational complexity, and improves the accuracy of predicted pixels.
Smart Images

Figure CN119496891B_ABST
Abstract
Description
[0001] This application is a divisional application. The original application has the application number 201980062258.5 and the original application date is September 23, 2019. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the technical field of image and / or video encoding and decoding, and more particularly to methods and apparatus for aspect ratio correlation filtering for intra-frame prediction. Background Technology
[0003] Digital video has been widely used since the introduction of DVDs. Before transmission, video is encoded and sent using a transmission medium. Viewers receive the video and decode and display it using their viewing devices. Over the years, video quality has improved, for example, due to higher resolution, color depth, and frame rates. This has led to larger data streams, which are now typically transmitted via the Internet and mobile communication networks.
[0004] However, high-resolution video typically contains more information, thus requiring more bandwidth. To reduce bandwidth requirements, video coding standards involving video compression have been introduced. When video is encoded, bandwidth requirements (or corresponding memory requirements in the case of storage) are reduced. Typically, this reduction comes at the cost of quality. Therefore, video coding standards attempt to find a balance between bandwidth requirements and quality.
[0005] High efficiency video coding (HEVC) is an example of video coding standards generally known to those skilled in the art. In HEVC, the coding unit (CU) is divided into multiple prediction units (PU) or multiple transform units (TU). Versatile video coding (VVC) is the latest joint video project, a collaborative effort between the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) standardization organizations, known as the Joint Video Exploration Team (JVET). VVC is also known as the ITU-T H.266 / Next Generation Video Coding (NGVC) standard. In VVC, the concept of multiple partition types is removed; that is, the separation between the concepts of CU, PU, and TU is eliminated, except for CUs whose size exceeds the maximum transform length, and greater flexibility in CU partition shape is supported.
[0006] The processing of these coding units (CUs) (also called blocks) depends on their size, spatial location, and the coding mode specified by the encoder. Based on the prediction type, coding modes can be divided into two categories: intra-frame prediction modes and inter-frame prediction modes. Intra-frame prediction modes use samples from the same picture (also called a frame or image) to generate reference samples to compute predicted values for the samples of the reconstructed block. Intra-frame prediction is also called spatial prediction. Inter-frame prediction modes are designed for temporal prediction and use reference samples from the previous or next picture to predict samples for the blocks in the current picture.
[0007] The choice of interpolation filter is coordinated with the decision of the primary reference side. Currently, both decisions rely on a comparison between the intra-prediction mode and the diagonal (45-degree) direction. Summary of the Invention
[0008] Apparatus and methods for intra-frame prediction are disclosed. These apparatuses and methods threshold the intra-frame prediction mode during the interpolation filter or smoothing filter selection process using an alternative direction. Specifically, this direction corresponds to the angle of the main diagonal of the block to be predicted.
[0009] This embodiment proposes a mechanism for selecting different reference sample filters to take into account the orientation of a block. Specifically, the width and height of the block are examined independently, thereby applying different reference sample filters to reference samples located on different sides of the block to be predicted.
[0010] Implementations of this embodiment are described in the claims and the following description.
[0011] The scope of protection is defined by the claims. Attached Figure Description
[0012] In the following, exemplary embodiments are described in more detail with reference to the accompanying drawings, wherein:
[0013] Figure 1 A schematic diagram illustrating an example of a video encoding and decoding system 100 is shown.
[0014] Figure 2 A schematic diagram illustrating an example of a video encoder 200 is shown.
[0015] Figure 3 A schematic diagram illustrating an example of a video decoder 300 is shown.
[0016] Figure 4 A schematic diagram illustrating the proposed 67 intra-frame prediction modes is shown.
[0017] Figure 5 An example of QTBT is shown.
[0018] Figure 6 The orientation of the rectangular block is shown.
[0019] Figure 7 This is an example of thresholding intra-prediction modes during the interpolation filter selection process using an alternative direction.
[0020] Figure 8 This is an example of using different interpolation filters depending on which side the reference sample belongs to.
[0021] Figure 9 This is a schematic diagram illustrating an exemplary structure of the device. Detailed Implementation
[0022] In the following description, reference is made to the accompanying drawings, which form part of this disclosure, and which illustrate, by way of illustration, specific aspects of which embodiments may be arranged.
[0023] For example, it should be understood that the disclosures relating to the described method also apply to the corresponding device or system configured to perform the method, and vice versa. For example, if specific method steps are described, the corresponding device may include a unit that performs the described method steps, even if that unit is not explicitly described or illustrated in the figures. Furthermore, it should be understood that, unless otherwise specifically indicated, features of the various exemplary aspects described herein can be combined with each other.
[0024] Video coding generally refers to the processing of a sequence of images to form a video or video sequence. In the field of video coding and in this application, the terms picture, image, or frame can be used synonymously. Each picture is typically divided into a set of non-overlapping blocks. Image encoding / decoding is usually performed at the block level, where, for example, inter-frame prediction or intra-frame prediction is used to generate prediction blocks, which are then subtracted from the current block (the currently processed / pending block) to obtain a residual block. This residual block is further transformed and quantized to reduce the amount of data to be transmitted (compression). On the decoder side, the inverse processing is applied to the encoded / compressed block to reconstruct the block (video block) for representation.
[0025] Figure 1 This is a block diagram illustrating an exemplary video encoding and decoding system 100 that can utilize the techniques described in this disclosure, including techniques for encoding and decoding boundary partitions. System 100 is applied not only to video encoding and decoding but also to image encoding and decoding. Figure 1 As shown, system 100 includes source device 102, which generates encoded video data, which is then decoded by destination device 104 at a later time. Figure 2 The video encoder 200 shown is an example of the video encoder 108 of the source device 102. Figure 3 The video decoder 300 shown is an example of the video decoder 116 of the destination device 104. The source device 102 and destination device 104 can include any of a wide variety of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, so-called "smart" tablets, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source device 102 and destination device 104 can be equipped for wireless communication.
[0026] Destination device 104 can receive encoded video data to be decoded via link 112. Link 112 may include any type of medium or device capable of moving encoded video data from source device 102 to destination device 104. In one example, link 112 may include a communication medium enabling source device 102 to directly transmit encoded video data to destination device 104 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to destination device 104. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, wide area network, or global network such as the Internet. The communication medium may include a router, switch, base station, or any other device used to facilitate communication from source device 102 to destination device 104.
[0027] Alternatively, the encoded data can be output from output interface 110 to a storage device. Figure 1 (Not shown in the image). Similarly, encoded data can be accessed from the storage device via input interface 114. Destination device 104 can access the stored video data from the storage device via streaming or downloading. The technology disclosed herein is not limited to wireless applications or setups. This technology can be applied to video encoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., via the Internet), encoding of digital video stored on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, system 100 can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0028] exist Figure 1 In the example, source device 102 includes a video source 106, a video encoder 108, and an output interface 110. In some cases, the output interface 110 may include a modulator / demodulator (modem) and / or a transmitter. In source device 102, video source 106 may include sources such as: video capture devices, such as cameras, video archives containing previously captured video, video feed interfaces that receive video from video content providers, and / or computer graphics systems for generating computer graphics data as source video, or combinations of such sources. As an example, if video source 106 is a camera, source device 102 and destination device 104 may form a so-called camera phone or video phone. However, the techniques described in this disclosure are generally applicable to video encoding and can be applied to wireless and / or wired applications.
[0029] Captured, pre-captured, or computer-generated video can be encoded by video encoder 108. The encoded video data can be sent directly to destination device 104 via output interface 110 of source device 102. The encoded video data can also (or alternatively) be stored on a storage device for subsequent access by destination device 104 or other devices for decoding and / or playback.
[0030] Destination device 104 includes an input interface 114, a video decoder 116, and a display device 118. In some cases, the input interface 114 may include a receiver and / or a modem. The input interface 114 of destination device 104 receives encoded video data via link 112. The encoded video data transmitted via link 112 or provided on a storage device may include various syntax elements generated by video encoder 108 for use by video decoders, such as video decoder 116, to decode the video data. These syntax elements may be incorporated into the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0031] Display device 118 may be integrated with destination device 104 or located externally to destination device 104. In some examples, destination device 104 may include an integrated display device and is also configured to interface with an external display device. In other examples, destination device 104 may be a display device. Typically, display device 118 displays decoded video data to a user and may include any of a variety of display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0032] The video encoder 108 and the video decoder 116 can operate according to any kind of video compression standard, including but not limited to MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and ITU-T H.266 / Next Generation Video Coding (NGVC) standards.
[0033] It is generally expected that the video encoder 108 of the source device 102 can be configured to encode video data according to any of these current or future standards. Similarly, it is generally expected that the video decoder 116 of the destination device 104 can be configured to decode video data according to any of these current or future standards.
[0034] Both the video encoder 108 and the video decoder 116 can be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When these techniques are implemented in part in software, the device may store instructions for software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of the video encoder 108 and the video decoder 116 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.
[0035] In video coding standards, a video sequence typically comprises a series of images. However, it should be noted that this disclosure is also applicable to the field where interlaced scanning is used. Video encoder 108 can output a bitstream comprising a bitstream of representations forming encoded images and associated data. Video decoder 116 can receive the bitstream generated by video encoder 108. Furthermore, video decoder 116 can parse the bitstream to obtain syntax elements from it. Video decoder 116 can reconstruct images of video data, at least in part, based on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the inverse of the process performed by video encoder 108.
[0036] Figure 2 A schematic diagram illustrating an example of a video encoder 200 is shown. The video encoder 200 is applied not only to video encoding but also to image encoding. The video encoder 200 includes input blocks for receiving frames or images from a video stream and output for generating an encoded video bitstream. The video encoder 200 is adapted to apply prediction, transform, quantization, and entropy coding to the video stream. Transformation, quantization, and entropy coding are performed via transform unit 201, quantization unit 202, and encoding unit 203, respectively, to generate an encoded video bitstream as output.
[0037] The video stream corresponds to multiple frames, each of which is divided into blocks of a certain size that are either intra-frame or inter-frame encoded. Intra-frame prediction unit 209 performs intra-frame encoding on blocks, for example, the first frame of the video stream. Intra-frames are encoded using only information from the same frame, allowing them to be decoded independently and providing entry points for random access within the bitstream. Inter-frame prediction unit 210 performs inter-frame encoding on blocks of other frames of the video stream: information from the encoded frames (called reference frames) is used to reduce temporal redundancy, thereby predicting each block of the inter-frame encoded frame from blocks of the same size in the reference frames. Mode selection unit 208 is adapted to select whether a block of a frame is processed by intra-frame prediction unit 209 or inter-frame prediction unit 210.
[0038] To perform inter-frame prediction, the encoded reference frame is processed by inverse quantization unit 204, inverse transform unit 205, and filtering unit 206 (optionally) to obtain a reference frame, which is then stored in frame buffer 207. Specifically, these units can process reference blocks of the reference frame to obtain reconstructed reference blocks. The reconstructed reference blocks are then reassembled into the reference frame.
[0039] Inter-frame prediction unit 210 takes the current frame or image to be inter-coded and one or more reference frames or images from frame buffer 207 as input. Inter-frame prediction unit 210 applies motion estimation and motion compensation. Motion estimation is used to obtain motion vectors and reference frames based on a certain cost function. Then, motion compensation describes the current block of the current frame based on the transformation from the reference block of the reference frame to the current frame. Inter-frame prediction unit 210 outputs a prediction block for the current block, where the prediction block minimizes the difference between the current block to be encoded and its prediction block, i.e., minimizes the residual block. Minimizing the residual block is based, for example, on a rate distortion optimization process.
[0040] Then, the difference between the current block and its prediction block, i.e., the residual block, is transformed by the transform unit 201. The transform coefficients are quantized and entropy encoded by the quantization unit 202 and the encoding unit 203. The encoded video bitstream includes intra-frame coded blocks and inter-frame coded blocks.
[0041] Figure 3 A schematic diagram illustrating an example of a video decoder 300 is shown. The video decoder 300 is applied not only to video decoding but also to image decoding. The video decoder 300 specifically includes a frame buffer 307 and an inter-frame prediction unit 310. The frame buffer 307 is adapted to store at least one reference frame obtained from the encoded video bitstream. The inter-frame prediction unit 310 is adapted to generate a prediction block for the current block of the current frame from the reference block of the reference frame.
[0042] Decoder 300 is adapted to decode the encoded video bitstream generated by video encoder 200, and both decoder 300 and encoder 200 generate the same predictions. The frame buffer 307 and inter-frame prediction unit 310 have the same characteristics as... Figure 2 The frame buffer 207 and the inter-frame prediction unit 210 have similar characteristics.
[0043] Specifically, the video decoder 300 includes units also present in the video encoder 200, such as an inverse quantization unit 304, an inverse transform unit 305, a filtering unit 306 (optional), and an intra-frame prediction unit 309, which correspond to the inverse quantization unit 204, inverse transform unit 205, filtering unit 206, and intra-frame prediction unit 209 of the video encoder 200, respectively. The decoding unit 303 is adapted to decode the received encoded video bitstream and accordingly obtain the quantized residual transform coefficients. The quantized residual transform coefficients are fed to the inverse quantization unit 304 and the inverse transform unit 305 to generate residual blocks. The residual blocks are added to the prediction blocks, and finally fed to the filtering unit 306 to obtain the decoded video. The frames of the decoded video can be stored in the frame buffer 307 and used as reference frames for inter-frame prediction.
[0044] The video encoder 200 can divide the input video frames into blocks before encoding. The term "block" in this disclosure is used for any type of block or any depth block; for example, the term "block" includes, but is not limited to, root blocks, blocks, child blocks, leaf nodes, etc. Blocks to be encoded do not necessarily have the same size. A single image can include blocks of different sizes, and the block grid of different images in a video sequence can also be different.
[0045] According to the HEVC / H.265 standard, 35 intra-frame prediction modes can be used. For example... Figure 4 As shown, this set includes the following modes: planar mode (intra-prediction mode index 0), DC mode (intra-prediction mode index 1), and orientation (angle) mode, which covers a 180° range and has... Figure 4 The black arrows in the image indicate the range of intra-prediction mode index values from 2 to 34. To capture arbitrary edge directions present in natural video, the number of directional intra-prediction modes has been expanded from the 33 used in HEVC to 65. Other directional modes are... Figure 4 The dashed arrows indicate that the planar mode and DC mode remain unchanged. It is worth noting that the range covered by the intra-frame prediction mode can be greater than 180°. Specifically, the 62 directional modes with index values from 3 to 64 cover a range of approximately 230°, meaning that several pairs of modes have opposite directions. In the case of the HEVC reference model (HM) and the JEM platform, such as... Figure 4As shown, only one pair of angular patterns (i.e., patterns 2 and 66) have opposite directions of orientation. For constructing predictions, conventional angular patterns take reference samples and (if necessary) filter them to obtain sample predictions. The number of reference samples required to construct predictions depends on the length of the filters used for interpolation (e.g., bilinear and cubic filters have lengths of 2 and 4, respectively).
[0046] In VVC, a partitioning mechanism based on both quadtrees and binary trees, called QTBT, is used. Figure 5 As described, QTBT partitioning can provide not only square blocks but also rectangular blocks. Of course, compared to the traditional quadtree-based partitioning used in the HEVC / H.265 standard, QTBT partitioning comes at the cost of some signaling overhead and increased computational complexity at the encoder end. However, QTBT-based partitioning has better segmentation characteristics, thus significantly improving coding efficiency compared to traditional quadtrees.
[0047] In this paper, the terms "vertically oriented block" and "horizontally oriented block" apply to rectangular blocks generated by the QTBT framework. These terms have the same characteristics as... Figure 6 The same meaning is shown.
[0048] This invention proposes a mechanism for selecting different reference sample filters to take into account the orientation of a block. Specifically, the width and height of the block are examined independently, thereby applying different reference sample filters to reference samples located on different sides of the block to be predicted. Furthermore, this invention proposes a mechanism for selecting different reference sample filters based on the aspect ratio and intra-prediction mode of the block to be predicted.
[0049] The prior art review describes how the selection of the interpolation filter is coordinated with the decision of the primary reference side selection. Currently, both decisions rely on a comparison between the intra-prediction mode and the diagonal (45-degree) direction.
[0050] However, it can be noted that this design has serious drawbacks for elongated blocks. It was observed that even when the shorter side is selected as the primary reference using a pattern comparison standard, most predicted pixels are still derived from the reference sample of the longer side (shown as the dashed area).
[0051] This invention proposes using an alternative direction to threshold intra-prediction modes during the interpolation filter selection process. Specifically, this direction corresponds to the angle of the main diagonal of the block to be predicted. For example, for blocks of sizes 32×4 and 4×32, such as... Figure 7As shown, the threshold mode m for determining the reference sample filter is defined. T .
[0052] A specific value for the intra-frame prediction angle can be calculated using the following formula:
[0053]
[0054] Where W and H are the width and height of the block, respectively.
[0055] Another implementation of this embodiment uses different interpolation filters depending on which side the reference sample belongs to. This determined example is... Figure 8 As shown in the image.
[0056] A straight line with an angle corresponding to the intra-frame direction m divides the prediction block into two regions. Different interpolation filters are used to predict samples belonging to different regions.
[0057] Table 1 shows (for a set of intra-prediction modes defined in BMS 1.0) m T Example values and corresponding angles. For example... Figure 7 As shown, angle α is illustrated.
[0058] Table 1. (For a set of intra-frame prediction modes defined in BMS 1.0) m T Example values
[0059]
[0060] Different interpolation filters are used to predict samples within a block, with the interpolation filter selected based on the aspect ratio of the block to be predicted and the intra-prediction mode. The interpolation filter can be selected specifically based on the block shape, horizontal or vertical orientation, and the intra-prediction mode angle.
[0061] This embodiment can also be applied to the reference sample filtering stage. Specifically, similar rules described above for the interpolation filter selection process can be used to determine the reference sample smoothing filter (e.g., for intra-frame prediction). For example, the smoothing filter can be selected based on the aspect ratio of the block to be predicted and the intra-frame prediction mode. In particular, the smoothing filter can be selected based on the shape of the block, its horizontal or vertical orientation, and the angle of the intra-frame prediction mode.
[0062] A set of conditions for determining whether to filter the reference sample is controlled by mode-dependent intra-smoothing (MDIS) techniques. For example, a set of conditions may include MDIS conditions. Based on the width and height of the prediction block, it can be determined whether to apply the
[121] filter to the reference sample. MDIS decisions can also be used to switch between interpolation filters. For example, when the MDIS conditions are true, a Gaussian filter is selected for interpolation. Otherwise, when the MDIS conditions are not true, a cubic filter is used.
[0063] You can check the MDIS conditions by performing the following steps:
[0064] Step 1: Determine the table index by calculating (log2(W) + log2(H)) >> 1. Specifically, the index is determined based on the aspect ratio of the block to be predicted.
[0065] Step 2: Calculate the absolute value of the difference between the intra-prediction mode and the horizontal intra-prediction mode, and the absolute value of the difference between the intra-prediction mode and the vertical intra-prediction mode. Select the minimum value among these absolute values. Specifically, choose the minimum of the two absolute values.
[0066] Step 3: Compare the value determined in the previous step (i.e., the minimum value) with the value obtained from the table using the index calculated in Step 1. For example, this value can be obtained from the following table:
[0067]
[0068]
[0069] The proposed invention suggests modifying the above MDIS condition check to account for non-square blocks. Specifically, an offset value is added to the value obtained in step 3, i.e., added to the value obtained from the table. The value obtained from the table can be modified based on the offset value to obtain a threshold. This threshold can then be compared with the minimum value determined in step 2, instead of directly comparing the obtained value with the minimum value determined in step 2. It is worth noting that the MDIS condition is true if the minimum value determined in step 2 is greater than the threshold determined in the modified step 3. If all the following checks are successful, the offset value is applied to the obtained value, especially if the offset value is not zero:
[0070] a. Check whether the shorter side of the block, especially the block to be predicted, is at least twice as short as the longer side.
[0071] b. Check if the shorter side length of the block exceeds the threshold. This threshold length can be set to 2, 4, 8, or 16.
[0072] c. Check whether the intra-prediction mode direction is between the horizontal intra-prediction mode and the vertical intra-prediction mode.
[0073] d. If the block width is greater than the block height, check if the intra-prediction mode is less than DIA_IDX+dirOffset, i.e., less than the sum of the diagonal intra-prediction mode (see Table 1) and the direction offset; otherwise, check if the intra-prediction mode is greater than DIA_IDX+dirOffset. The value of dirOffset can be predetermined (e.g., set to equal to 5), or it can depend on the block's aspect ratio (e.g., for aspect ratios of 1:2, 1:4, 1:8, and 1:16, the value of dirOffset will be assigned to 5, 6, 7, and 7 respectively).
[0074] If all these checks are successful and the block width is greater than the block height, the value obtained in step 3 is decreased by a predetermined offset value (e.g., equal to dirOffset or equal to abs(dirOffset)). If all these checks are successful and the block width is less than the block height, the value obtained in step 3 is increased by a predetermined offset value (e.g., equal to dirOffset or equal to abs(dirOffset)).
[0075] Figure 9 This is a block diagram of an apparatus (or device) 1100 that can be used to implement various embodiments. For example, apparatus 1100 can be used to implement the mechanism described above in this invention. Apparatus 1100 can be configured to determine an intra-prediction mode and the aspect ratio of the block to be predicted, and to select an interpolation filter or a smoothing filter based on the determined intra-prediction mode and the determined aspect ratio. Specifically, apparatus 1100 can be configured to determine MDIS conditions based on the intra-prediction mode and the aspect ratio of the block to be predicted, and then select an interpolation filter or a smoothing filter based on the determined MDIS conditions. Apparatus 1100 can be as follows: Figure 1 The source device 102 shown, or as Figure 2 The video encoder 200 shown, or as... Figure 1 The destination device 104 shown, or as Figure 3The video decoder 300 is shown. Additionally, device 1100 may host one or more of the described elements. In some embodiments, device 1100 is equipped with one or more input / output devices, such as speakers, microphones, mice, touchscreens, keypads, keyboards, printers, displays, etc. Device 1100 may include one or more central processing units (CPUs) 1510, memory 1520, mass storage 1530, video adapters 1540, and I / O interfaces 1560 connected to a bus. The bus can be one or more of any type of bus architecture, including memory buses or memory controllers, peripheral buses, video buses, etc.
[0076] CPU 1510 can be any type of electronic data processor. Memory 1520 can be any type of system memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), read-only memory (ROM), or combinations thereof. In one embodiment, memory 1520 may include ROM used at startup and DRAM used for program and data storage during program execution. In this embodiment, memory 1520 is non-volatile. Mass storage 1530 includes any type of storage device that stores data, programs, and other information and makes such data, programs, and other information accessible via a bus. Mass storage 1530 includes one or more of, for example, solid-state drives, hard disk drives, disk drives, optical disk drives, etc.
[0077] Video adapter 1540 and I / O interface 1560 provide interfaces to couple external input and output devices to device 1100. For example, device 1100 may provide an SQL command interface to a client. As shown, examples of input and output devices include any combination of a monitor 1590 coupled to video adapter 1540 and a mouse / keyboard / printer 1570 coupled to I / O interface 1560. Other devices may be coupled to device 1100, and additional or fewer interface cards may be used. For example, a serial interface card (not shown) may be used to provide a serial interface for a printer.
[0078] The device 1100 also includes one or more network interfaces 1550, which include wired links such as Ethernet cables and / or wireless links for accessing nodes or one or more networks 1580. The network interface 1550 allows the device 1100 to communicate with remote units via the network 1580. For example, the network interface 1550 can provide communication to a database. In one embodiment, the device 1100 is coupled to a local area network or a wide area network for data processing and communication with remote devices such as other processing units, the Internet, remote storage facilities, etc.
[0079] A piecewise linear approximation is introduced to compute the weighting coefficients needed to predict pixels within a given block. On the one hand, compared to direct weighting coefficient computation, the piecewise linear approximation significantly reduces the computational complexity of the distance-weighted prediction mechanism; on the other hand, compared to the simplification of existing techniques, it helps to achieve higher weighting coefficient accuracy.
[0080] This embodiment can be applied to other bidirectional and position-dependent intra-prediction techniques (e.g., different modifications of PDPC) and mechanisms that use weighting coefficients that depend on the distance from one pixel to another to blend different parts of an image (e.g., some blending methods in image processing).
[0081] The subject matter and operations described in this disclosure can be implemented by digital electronic circuits, or by computer software, firmware, or hardware, including the structures disclosed in this disclosure and their equivalents, or combinations thereof. The subject matter described in this disclosure can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or additionally, the program instructions can be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium, such as a computer-readable medium, can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or combinations thereof, or is included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or combinations thereof. Furthermore, although the computer storage medium is not a propagating signal, it can be a source or destination of computer program instructions encoded in artificially generated propagating signals. Computer storage media may also be one or more separate physical and / or non-transitory components or media (e.g., multiple CDs, disks, or other storage devices), or may be included in one or more separate physical and / or non-transitory components or media.
[0082] In some implementations, the operations described in this disclosure can be implemented as hosted services provided on servers in a cloud computing network. For example, computer-readable storage media can be logically grouped and accessed within the cloud computing network. Servers within the cloud computing network may include cloud computing platforms for providing cloud-based services. The terms "cloud," "cloud computing," and "cloud-based" may be used interchangeably as appropriate without departing from the scope of this disclosure. Cloud-based services can be hosted services provided by servers and delivered over a network to client platforms to enhance, supplement, or replace applications that execute locally on client computers. Circuits can use cloud-based services to quickly receive software upgrades, applications, and other resources that would otherwise take a long time to be delivered to the circuit.
[0083] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages. Computer programs can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one or more computers located at a single site or distributed across multiple sites and interconnected via a communication network.
[0084] The processing and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs to perform actions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry (e.g., FPGA (Field-Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit)), and the apparatus can also be implemented as special-purpose logic circuitry (e.g., FPGA or ASIC).
[0085] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, as well as any type of digital computer or any other type of processor. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for performing actions according to instructions, and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, receiving data from or transferring data to one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily need to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.
[0086] While this disclosure contains numerous specific implementation details, these should not be construed as limiting any implementation or the scope of any possible claims, but rather as descriptions of features specific to a particular implementation. Certain features described in this disclosure within the context of individual implementations may also be implemented in combination within a single implementation. Conversely, various features described in the context of a single implementation may be implemented individually or in any suitable sub-combination in multiple implementations. Furthermore, although the foregoing may describe features as functioning in certain combinations or even those features originally claimed, in certain circumstances one or more features from a claimed combination may be removed from that combination, and the claimed combination may refer to a sub-combination or a variation of a sub-combination.
[0087] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential manner, or as requiring the execution of all shown operations to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous. Furthermore, the separation of the various system components in the above implementations should not be interpreted as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or encapsulated within multiple software products.
[0088] Therefore, specific implementations of the subject matter have been described. Other implementations are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequence shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
Claims
1. A method for switching intra-frame reference sample filters for intra-frame prediction, characterized in that, The method includes: Determine the intra-frame prediction mode and the aspect ratio of the block to be predicted, and Based on the intra-prediction mode and the aspect ratio, determine the mode-dependent intra-smoothing MDIS conditions; Based on the MDIS conditions, an interpolation filter or a reference sample smoothing filter is selected for intra-frame prediction; The determination of the MDIS conditions based on the intra-frame prediction mode and the aspect ratio includes: The index is determined based on (log2(W) + log2(H))>>1, where W and H are the width and height of the block to be predicted, respectively. Choose the minimum value from the absolute value of the difference between the intra-prediction mode and the horizontal intra-prediction mode and the absolute value of the difference between the intra-prediction mode and the vertical intra-prediction mode. Use the index to retrieve values from the table; A threshold is obtained based on the value; and The minimum value is compared with the threshold, wherein if the minimum value is greater than the threshold, the MDIS condition is true.
2. The method according to claim 1, wherein, The interpolation filter includes a cubic interpolation filter or a Gaussian interpolation filter.
3. The method according to claim 2, comprising: If the MDIS condition is true, then the cubic interpolation filter is selected, and If the MDIS condition is not true, then the Gaussian interpolation filter is selected.
4. An encoder (20), comprising: One or more processors; as well as A non-transitory computer-readable storage medium coupled to the one or more processors and storing a program executed by the one or more processors, wherein, when executed by the one or more processors, the program configures the encoder to perform the method according to any one of claims 1 to 3.
5. A decoder (30), comprising: One or more processors; as well as A non-transitory computer-readable storage medium coupled to the one or more processors and storing a program executed by the processors, wherein, when executed by the one or more processors, the program configures the decoder to perform the method according to any one of claims 1 to 3.
6. A computer-readable storage medium comprising a program that, when executed by one or more processors, causes the one or more processors to perform the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Intra smoothing filter for video coding
CN103141100A
Method and device for encoding / decoding multi-layer video signal
CN105659597A