Intra prediction for image and video compression
By using the focus and peripheral pixel lines to generate non-parallel prediction lines, a novel intra-frame prediction mode solves the problem of low efficiency in encoding parallel line images in the existing technology and improves the video compression effect.
Patent Information
- Application Number
- CN202080093804.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-13
- Filing Date
- 2020-05-14
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2040-05-14
AI Technical Summary
Existing intra-frame prediction modes have difficulty generating optimal prediction blocks when processing images containing parallel lines or checkerboard patterns, resulting in low coding efficiency.
A novel intra prediction mode is adopted to generate prediction blocks by using focal and peripheral pixel lines, and the value of each pixel is calculated according to different prediction angles to generate non-parallel prediction lines.
The coding efficiency of images containing parallel lines or checkerboard patterns is improved, the residual error is reduced, and the video compression effect is improved.
Smart Images

Figure CN115004703B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 976,001, filed February 13, 2020, the entire disclosure of which is incorporated herein by reference. Background Art
[0003] A digital video stream can use a series of frames or still images to represent a video. Digital video can be used for a variety of applications, including, for example, video conferencing, high-definition video entertainment, video advertising, or sharing of user-generated videos. A digital video stream can contain large amounts of data and consume a significant amount of computing or communication resources of a computing device for processing, transmitting, or storing the video data. Various methods have been proposed to reduce the amount of data in a video stream, including compression and other encoding techniques.
[0004] Spatial similarity-based coding can be performed by decomposing a frame or image into blocks that are predicted based on other blocks within the same frame or image. The difference between the block and the predicted block (i.e., the residual) is compressed and encoded in the bitstream. The decoder uses the difference and a reference frame to reconstruct the frame or image. Summary of the Invention
[0005] Disclosed herein are aspects, features, elements, and implementations for encoding and decoding blocks using intra prediction.
[0006] A first aspect is a method for coding a current block using an intra prediction mode. The method includes obtaining a focal point having coordinates (a, b) in a coordinate system; generating a prediction block for the current block using first and second peripheral pixels, wherein the first peripheral pixels form a first peripheral pixel line constituting an x-axis, wherein the second peripheral pixels form a second peripheral pixel line constituting a y-axis, and wherein the first and second peripheral pixel lines form a coordinate system having an origin; and coding a residual block corresponding to a difference between the current block and the prediction block. Generating the prediction block includes: determining at least one of an x-intercept or a y-intercept for each position of the prediction block at a position (i, j) of the prediction block, wherein the x-intercept is a first point at which a line formed by a point centered at each position of the prediction block and the focal point intersects the first peripheral pixel line, and wherein the y-intercept is a second point at which a line formed by a point centered at each position of the prediction block and the focal point intersects the second peripheral pixel line; and determining a predicted pixel value for each position of the prediction block using at least one of the x-intercept or the y-intercept.
[0007] Another aspect is an apparatus for decoding a current block. The apparatus includes a memory and a processor. The processor is configured to execute instructions stored in the memory to decode a focus point from a compressed bitstream; obtain a prediction block of prediction pixels of the current block, where each prediction pixel is located at a respective position within the prediction block; and reconstruct the current block using the prediction block. To obtain the prediction block includes, for each position of the prediction block, executing the instructions to obtain a line indicating a respective prediction angle that connects the focus point to the position; and determine a pixel value for the position using the line.
[0008] Another aspect is a method for encoding a current block. The method includes obtaining a prediction block of prediction pixels for the current block using peripheral pixels, where each prediction pixel is located at a respective position within the prediction block; and encoding a focus point in a compressed bitstream. Obtaining the prediction block includes obtaining the focus point having coordinates (a, b) in a coordinate system, the focus point being outside the current block, and the focus point not being any of the peripheral pixels; for each position of the prediction block, performing steps including obtaining a line indicating a respective prediction angle that connects the focus point to the position; and determining a pixel value for the position using the line.
[0009] It will be appreciated that aspects can be implemented in any convenient form. For example, aspects can be implemented by a suitable computer program which can be carried on a suitable carrier medium, which can be a tangible carrier medium (e.g. a disc) or an intangible carrier medium (e.g. a communication signal). Aspects can also be implemented using suitable apparatus, which can take the form of programmable computers running computer programs arranged to implement the described methods. Aspects can be combined such that features described in the context of one aspect can be implemented in another aspect. BRIEF DESCRIPTION OF DRAWINGS
[0010] The description herein makes reference to the accompanying drawings, wherein like reference numerals refer to like parts throughout the several views.
[0011] Figure 1 is a schematic diagram of a video encoding and decoding system.
[0012] Figure 2 is a block diagram of an example of a computing device that can implement a transmitting station or a receiving station.
[0013] Figure 3 is a diagram of a video stream to be encoded and subsequently decoded.
[0014] Figure 4 is a block diagram of an encoder according to an embodiment of the present disclosure.
[0015] Figure 5 is a block diagram of a decoder according to an embodiment of the present disclosure.
[0016] Figure 6 is a block diagram of a representation of a portion of a frame according to an embodiment of the disclosure.
[0017] Figure 7 is a diagram of an example of an intra prediction mode.
[0018] Figure 8 is an example of an image portion including a railroad track.
[0019] Figure 9 is an example of a flowchart of a technique for determining bit positions along a line of peripheral pixels for determining a predicted pixel value according to an embodiment of the disclosure.
[0020] Figure 10 is an example of bit positions calculated by the technique of Figure 9 according to an embodiment of the disclosure.
[0021] Figure 11 is an example of a predicted block calculated from Figure 10 according to an embodiment of the disclosure.
[0022] Figure 12 is a flowchart of a technique for intra prediction of a current block according to an embodiment of the disclosure.
[0023] Figure 13 is a flowchart of a technique for generating a predicted block for a current block using intra prediction according to an embodiment of the disclosure.
[0024] Figure 14 is a flowchart of a technique for decoding a current block using an intra prediction mode according to an embodiment of the disclosure.
[0025] Figure 15 is an example 1500 showing a focus point according to an embodiment of the disclosure.
[0026] Figure 16 is an example showing x and y intercepts according to an embodiment of the disclosure.
[0027] Figure 17 shows an example of a focus group according to an embodiment of the disclosure.
[0028] Figure 18 is an example of a flowchart of a technique for determining bit positions along a line of peripheral pixels for determining a predicted pixel value according to an embodiment of the disclosure.
[0029] Figure 19 is a flowchart of a technique for generating a predicted block for a current block using intra prediction according to an embodiment of the disclosure.
[0030] Figures 20A-20B is an example to illustrate the technique of Figure 19 according to an embodiment of the disclosure.
[0031] Figure 21 FIG. 1 shows when the above peripheral pixel is used as the primary peripheral pixel Figure 19 Examples of techniques. DETAILED DESCRIPTION
[0032] As described above, compression schemes related to coded video streams can include breaking up an image into blocks and using one or more techniques to generate a digital video output bitstream (i.e., an encoded bitstream) to limit the information included in the output bitstream. A received bitstream can be decoded to recreate the blocks and source image from the limited information. Encoding a video stream or a portion of a video stream such as a frame or a block can include using spatial similarities in the video stream to improve coding efficiency. For example, a current block of a video stream can be encoded based on identifying a previously coded pixel value or a combination of previously coded pixel values and a difference (a residual) between the previously coded pixel value or combination of previously coded pixel values and a pixel value in the current block.
[0033] Encoding using spatial similarities can be referred to as intra prediction. Intra prediction attempts to predict pixel values of a current block of a frame (i.e., an image, a picture) or a single image of a video stream using pixels that are peripheral to the current block; that is, pixels that are in the same frame as the current block but outside of the current block. Intra prediction can be performed along a prediction direction, referred to herein as a prediction angle, where each direction can correspond to an intra prediction mode. Intra prediction modes use pixels that are peripheral to the current block being predicted. Pixels that are peripheral to the current block are pixels that are outside of the current block. Intra prediction modes can be signaled by an encoder to a decoder.
[0034] Many different intra prediction modes can be supported. Some intra prediction modes use a single value for all pixels within a predicted block generated using at least one peripheral pixel. Others, referred to as directional intra prediction modes, each have a corresponding prediction angle. Intra prediction modes can include, for example, a horizontal intra prediction mode, a vertical intra prediction mode, and intra prediction modes for various other directions. For example, a codec can have available prediction modes corresponding to 50 to 60 prediction angles. With respect to Figure 7 Examples of intra prediction modes are described.
[0035] However, such as described above and with respect to Figure 7The intra prediction modes of current codecs that are described can not optimally code blocks of images or scenes that contain parallel lines. It is well known that a perspective representation of a scene (e.g., an image) or a scene viewed at an angle in which the image includes parallel lines can have one or more vanishing points. That is, the parallel lines can be perceived (i.e., seen) as converging to a vanishing point (or focal point) (or diverging from a vanishing point (or focal point)). Non-limiting examples of images that include parallel lines or checkerboard patterns include a striped shirt, bricks on a building, blinds, train tracks, panel wood flooring, sun rays, tree trunks, etc. While for ease of description, parallel lines are used herein, the present disclosure is not so limited. For example, the disclosure herein can be used with parallel edges such as the edges of a pencil in an image that is taken from a point of a point. Further, the present disclosure is not limited to straight lines. The parameters described below can be used to cover curves in the predicted lines.
[0036] Such patterns can be easily recognized by the eye. However, when viewed from a perspective angle that is not 90 degrees, such patterns can be significantly more difficult to programmatically discern and code. As already mentioned, parallel lines can appear to go to one point in the distance, such as with the converging train tracks shown in Figure 8 described.
[0037] While some intra prediction modes can have an associated direction, that same direction is used to generate each prediction pixel of the prediction block. However, with converging lines such as Figure 8 the train tracks, each line can have a different direction. Thus, a single direction intra prediction mode can not generate an optimal prediction block for coding the current block. By optimal prediction block is meant a prediction block that minimizes the residual error between the prediction block and the current block being coded.
[0038] Embodiments in accordance with the present disclosure use novel intra prediction modes that can be used to code blocks that include converging lines. As indicated above, intra prediction modes use pixels of the periphery of the current block. At a high level, prediction blocks generated using intra prediction modes in accordance with embodiments of the present disclosure can be such that one pixel of a row of the prediction block can be copied from one or more peripheral pixels in one direction, while another pixel of the same row of the prediction block can be copied from one or more other peripheral pixels in a different direction. In addition, as described further below, scaling (either enlargement or reduction) can optionally be applied in accordance with parameters of the intra prediction mode.
[0039] In some embodiments, to generate a prediction block according to an intra prediction mode of the present disclosure, the same set of surrounding pixels (i.e., the above and / or left surrounding pixels) that is typically used for intra prediction is repeatedly re-sampled for each pixel (i.e., pixel position) of the prediction block. The surrounding pixels are considered as a continuous line of pixel values, which is referred to herein as a surrounding pixel line for ease of reference. To generate a prediction pixel of the prediction block, different bits of the surrounding pixel line are considered. However, as further described below, each time a bit is considered, the bit is shifted from the immediately preceding bit according to parameters of the intra prediction mode. It is noted that only integer positions of the pixel positions of the surrounding pixel line have pixel values that are known: the surrounding pixels themselves. As such, sub-pixel (i.e., non-integer pixel) values of the surrounding pixel line are obtained from the surrounding pixels using, for example, interpolation or filtering operations.
[0040] In other embodiments, an initial prediction block can be generated using a directional prediction mode. A warping (e.g., a warping function, a set of warping parameters, etc.) can then be applied to the initial prediction block to generate the prediction block. In an example, the warping can be a perspective warping. In an example, the warping can be an affine warping.
[0041] In yet another embodiment, an intra prediction mode according to the present disclosure can use a focal point that is a distance away from the current block. The focal point can be considered as a point in space from which all pixels of the current block emanate or to which all pixels of the current block converge. For each prediction pixel position of the prediction block, a line connecting the pixel position and the focal point is drawn. An x-intercept and a y-intercept of the line with the x-axis and the y-axis of a coordinate system formed by the left and above surrounding pixels are determined. The x-intercept and the y-intercept are used to determine (e.g., identify, select, calculate, etc.) the surrounding pixels that are used to calculate the prediction pixel value.
[0042] In summary, for all pixels of the prediction block, such as with respect to Figure 7 The directional prediction modes described result in parallel prediction lines. However, the intra prediction modes according to embodiments of the present disclosure result in non-parallel prediction lines. As such, at least two prediction pixels can be calculated according to different prediction angles.
[0043] In contrast to the design and semantics of conventional directional prediction modes in which each prediction pixel is calculated according to the same prediction angle, while the intra prediction modes according to the present disclosure can result in two or more prediction pixels being derived (e.g., calculated) by using parallel prediction lines (i.e., the same prediction angle), this is only accidental. For example, while depending on the location of the focal point, more than one prediction pixel can have the same prediction angle, the intra prediction modes according to embodiments of the present disclosure are such that not all prediction pixels can have the same prediction angle.
[0044] After first describing an environment in which intra prediction for image and video compression disclosed herein can be implemented, details are described herein. Although the intra prediction mode according to the present disclosure is described with respect to a video encoder and a video decoder, the intra prediction mode can also be used in an image codec. The image codec can be or can share many aspects of the video codec described herein.
[0045] Figure 1 is a schematic diagram of a video encoding and decoding system 100. The sending station 102 can be, for example, a computer having a hardware internal configuration such as that described in Figure 2 However, other suitable implementations of the sending station 102 are possible. For example, the processing of the sending station 102 can be distributed among multiple devices.
[0046] The network 104 can connect the sending station 102 and the receiving station 106 to encode and decode a video stream. In particular, a video stream can be encoded in the sending station 102, and the encoded video stream can be decoded in the receiving station 106. The network 104 can be, for example, the Internet. The network 104 can also be a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a cellular telephone network, or any other means of transmitting a video stream from the sending station 102 to the receiving station 106 in the present example.
[0047] In one example, the receiving station 106 can be a computer having a hardware internal configuration such as that described in Figure 2 However, other suitable implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 can be distributed among multiple devices.
[0048] Other implementations of the video encoding and decoding system 100 are possible. For example, an implementation can omit the network 104. In another implementation, a video stream can be encoded and then stored for later transmission to the receiving station 106 or any other device having a memory. In one implementation, the receiving station 106 receives (e.g., via the network 104, a computer bus, and / or some communication pathway) an encoded video stream and stores the video stream for later decoding. In an example implementation, the Real-time Transport Protocol (RTP) is used to transmit the encoded video over the network 104. In another implementation, a transport protocol other than RTP can be used, such as a Hypertext Transfer Protocol (HTTP) based video streaming protocol.
[0049] When used in a video conferencing system, for example, the sending station 102 and / or the receiving station 106 can include the ability to both encode and decode video streams as described below. For example, the receiving station 106 can be a video conference participant that receives an encoded video bitstream from a video conference server (e.g., the sending station 102) to decode and view and further encode its own video bitstream and transmit its own video bitstream to the video conference server for decoding and viewing by other participants.
[0050] Figure 2 is a block diagram of an example of a computing device 200 that can implement a sending station or a receiving station. For example, the computing device 200 can implement one or both of the sending station 102 and the receiving station 106 of Figure 1 . The computing device 200 can be in the form of a computing system that includes multiple computing devices, or a single computing device, such as a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, etc.
[0051] The CPU 202 in the computing device 200 can be a central processing unit. Alternatively, the CPU 202 can be any other type of device or devices capable of manipulating or processing information now existing or developed in the future. Although the disclosed embodiments can be practiced with the illustrated single processor, such as the CPU 202, more than one processor can be used to achieve speed and efficiency advantages.
[0052] In one embodiment, the memory 204 in the computing device 200 can be a read only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of memory device can be used as the memory 204. The memory 204 can include code and data 206 that are accessed by the CPU 202 using a bus 212. The memory 204 can further include an operating system 208 and application programs 210, including at least one program that permits the CPU 202 to perform the methods described herein. For example, the application programs 210 can include applications 1 through N that further include a video coding application that performs the methods described herein. The computing device 200 can also include a secondary memory 214, which can be, for example, a memory card used with the removable computing device 200. Because video communication sessions can contain a significant amount of information, they can be stored in whole or in part in the secondary memory 214 and loaded into the memory 204 as needed for processing.
[0053] The computing device 200 can also include one or more input devices, such as a keyboard 216. The keyboard 216 can be a touch-sensitive keyboard that combines a keyboard with a touch-sensitive element operable to sense touch input. The keyboard 216 can be coupled to the CPU 202 via the bus 212. Other input devices that allow a user to program or otherwise use the computing device 200 can be provided in addition to or instead of the keyboard 216. When the input device is or includes a keyboard, the keyboard can be implemented in a variety of ways, including through a liquid crystal display (LCD) keyboard, a cathode ray tube (CRT) keyboard, or a light emitting diode (LED) keyboard, such as an organic LED (OLED) keyboard.
[0054] The computing device 200 can also include or be in communication with an image sensing device 220, such as a camera or any other image sensing device 220 now existing or hereafter developed that can sense an image, such as an image of a user operating the computing device 200. The image sensing device 220 can be positioned such that it is directed toward a user operating the computing device 200. In an example, the position and optical axis of the image sensing device 220 can be configured such that the field of view includes an area directly adjacent to and from which the display 218 is visible.
[0055] The computing device 200 can also include or be in communication with a sound sensing device 222, such as a microphone or any other sound sensing device now existing or hereafter developed that can sense sound in the vicinity of the computing device 200. The sound sensing device 222 can be positioned such that it is directed toward a user operating the computing device 200 and can be configured to receive sound, such as speech or other utterances, emitted by the user while operating the computing device 200.
[0056] Although Figure 2 While the CPU 202 and the memory 204 of the computing device 200 are depicted as integrated into a single unit, other configurations can be used. Operations of the CPU 202 can be distributed across multiple machines (each machine having one or more processors), which can be coupled directly or across a local area or other network. The memory 204 can be distributed across multiple machines, such as network-based memory or memory in multiple machines performing operations of the computing device 200. While depicted here as a single bus, the bus 212 of the computing device 200 can be composed of multiple buses. Also, the secondary storage 214 can be directly coupled to the other components of the computing device 200, or can be accessed via a network, and can comprise a single integrated unit, such as a memory card, or multiple units, such as multiple memory cards. The computing device 200 can therefore be implemented in a variety of configurations.
[0057] Figure 3 is a diagram of an example of a video stream 300 to be encoded and subsequently decoded. The video stream 300 includes a video sequence 302. At the next level, the video sequence 302 includes a number of neighboring frames 304. Although three frames are depicted as neighboring frames 304, the video sequence 302 can include any number of neighboring frames 304. The neighboring frames 304 can then be further subdivided into individual frames, such as frame 306. At the next level, the frame 306 can be divided into a series or planes of slices 308. For example, the slices 308 can be subsets of the frame that allow for parallel processing. The slices 308 can also be subsets of the frame that separate the video data into individual colors. For example, a frame 306 of color video data can include a luma plane and two chroma planes. The slices 308 can be sampled at different resolutions.
[0058] Whether or not the frame 306 is divided into slices 308, the frame 306 can be further subdivided into blocks 310, which can contain data corresponding to, for example, 16x16 pixels in the frame 306. The blocks 310 can also be arranged to include data from one or more slices 308 of pixel data. The blocks 310 can also have any other suitable size, such as 4x4 pixels, 8x8 pixels, 16x8 pixels, 8x16 pixels, 16x16 pixels, or larger.
[0059] Figure 4 is a block diagram of an encoder 400 according to an embodiment of the present disclosure. As described above, the encoder 400 can be implemented in the transmitting station 102, such as by providing a computer software program stored in a memory, such as the memory 204. The computer software program can include machine instructions that, when executed by a processor, such as the CPU 202, cause the transmitting station 102 to encode video data in the manner described herein. The encoder 400 can also be implemented as special-purpose hardware included in, for example, the transmitting station 102. The encoder 400 has the following stages to perform various functions in the forward path (shown by solid connection lines) to produce an encoded or compressed bitstream 420 using the video stream 300 as input: an intra / inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy encoding stage 408. The encoder 400 can also include a reconstruction path (shown by dotted connection lines) to reconstruct frames for use in encoding future blocks. In the Figure 4 , the encoder 400 has the following stages to perform various functions in the reconstruction path: a dequantization stage 410, an inverse transform stage 412, a reconstruction stage 414, and a loop filtering stage 416. Other structural variations of the encoder 400 can be used to encode the video stream 300.
[0060] When presenting the video stream 300 for encoding, the frames 306 can be processed in units of blocks. At an intra / inter-frame prediction stage 402, the blocks can be encoded using intra-frame prediction (also referred to as intra-prediction) or inter-frame prediction (also referred to as inter-prediction) or a combination of both. In any case, a predicted block can be formed. In the case of intra-prediction, all or a portion of the predicted block can be formed from samples in the current frame that have previously been encoded and reconstructed. In the case of inter-prediction, all or a portion of the predicted block can be formed from samples in one or more previously constructed reference frames determined using motion vectors.
[0061] Next, still referring to Figure 4 , the predicted block can be subtracted from the current block at the intra / inter-frame prediction stage 402 to produce a residual block (also referred to as a residual). A transform stage 404 transforms the residual using a block-based transform into transform coefficients, for example, in the frequency domain. Such block-based transforms include, for example, the discrete cosine transform (DCT) and the asymmetric discrete sine transform (ADST). Other block-based transforms are possible. Moreover, a combination of different transforms can be applied to a single residual. In one example of applying a transform, the DCT transforms the residual block into a frequency domain where the values of the transform coefficients are based on spatial frequencies. The lowest frequency (DC) coefficient is at the top left of the matrix, and the highest frequency coefficient is at the bottom right of the matrix. Notably, the size of the predicted block, and thus the resulting residual block, can be different than the size of the transform block. For example, the predicted block can be partitioned into smaller blocks to which separate transforms are applied.
[0062] A quantization stage 406 converts the transform coefficients into discrete quantum values, referred to as quantized transform coefficients, using a quantizer value or quantization level. For example, the transform coefficients can be divided by the quantizer value and truncated. The quantized transform coefficients are then entropy encoded by an entropy encoding stage 408. The entropy encoding can be performed using any number of techniques including tokens and binary trees. The entropy encoded coefficients are then output to a compressed bitstream 420 along with other information used to decode the blocks, which can include, for example, the type of prediction used, the type of transform, motion vectors, and quantizer values. The information to decode the blocks can be entropy encoded as block, frame, slice, and / or segment headers within the compressed bitstream 420. The compressed bitstream 420 can also be referred to as an encoded video stream or an encoded video bitstream, and these terms will be used interchangeably herein.
[0063] Figure 4The reconstruction path (shown by the dotted line) in the encoder 400 can be used to ensure that both the encoder 400 and the decoder 500 (described below) use the same reference frames and blocks to decode the compressed bitstream 420. The reconstruction path performs functions similar to those that occur during the decoding process discussed in more detail below, including dequantizing the quantized transform coefficients at the dequantization stage 410 and inverse transforming the dequantized transform coefficients at the inverse transform stage 412 to produce a derivative residual block (also referred to as a derivative residual). At the reconstruction stage 414, the predicted block predicted at the intra / inter prediction stage 402 can be added to the derivative residual to create a reconstructed block. The in-loop filtering stage 416 can be applied to the reconstructed block to reduce distortion, such as block artifacts.
[0064] Other variations of the encoder 400 can be used to encode the compressed bitstream 420. For example, a non-transform-based encoder 400 can directly quantize the residual signal without the transform stage 404 for certain blocks or frames. In another implementation, the encoder 400 can combine the quantization stage 406 and the dequantization stage 410 into a single stage.
[0065] Figure 5 is a block diagram of a decoder 500 according to an implementation of the present disclosure. The decoder 500 can be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the receiving station 106 to decode video data in the manner described herein. The decoder 500 can also be implemented with hardware included in, for example, the transmitting station 102 or the receiving station 106. Similar to the reconstruction path of the encoder 400 discussed above, the decoder 500 in one example includes the following stages to perform various functions to produce an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra / inter prediction stage 508, a reconstruction stage 510, an in-loop filtering stage 512, and a deblocking filtering stage 514. Other structural variations of the decoder 500 can be used to decode the compressed bitstream 420.
[0066] When the compressed bitstream 420 is presented for decoding, the data elements within the compressed bitstream 420 can be decoded by an entropy decoding stage 502 to produce a set of quantized transform coefficients. A dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by a quantizer value), and an inverse transform stage 506 inversely transforms the dequantized transform coefficients using a selected transform type to produce a derivative residual, which can be the same as that created by the inverse transform stage 412 in the encoder 400. Using the header information decoded from the compressed bitstream 420, the decoder 500 can use an intra / inter prediction stage 508 to create the same prediction block as that created in the encoder 400, for example, at the intra / inter prediction stage 402. At a reconstruction stage 510, the prediction block can be added to the derivative residual to create a reconstructed block. A loop filtering stage 512 can be applied to the reconstructed block to reduce blocking artifacts. Other filtering can be applied to the reconstructed block. In this example, a deblocking filter stage 514 is applied to the reconstructed blocks to reduce blocking artifacts, and the result is output as an output video stream 516. The output video stream 516 can also be referred to as a decoded video stream, and the terms will be used interchangeably herein.
[0067] Other variations of the decoder 500 can be used to decode the compressed bitstream 420. For example, the decoder 500 can produce the output video stream 516 without the deblocking filter stage 514. In some embodiments of the decoder 500, the deblocking filter stage 514 is applied before the loop filter stage 512. Additionally or alternatively, the encoder 400 includes a deblocking filter stage in addition to the loop filter stage 416.
[0068] Figure 6 It is a representation of an embodiment according to the present disclosure such as Figure 3 306. As shown, the portion of the frame 600 includes four 64x64 blocks 610 in two rows and two columns in a matrix or Cartesian plane, which may be referred to as a superblock. A superblock can have a larger or smaller size. Figure 6 While explained with respect to superblocks of size 64x64, the description is readily extended to larger (e.g., 128x128) or smaller superblock sizes.
[0069] In an example, and without loss of generality, a superblock can be a basic or maximum coding unit (CU). Each superblock can include four 32x32 blocks 620. Each 32x32 block 620 can include four 16x16 blocks 630. Each 16x16 block 630 can include four 8x8 blocks 640. Each 8x8 block 640 can include four 4x4 blocks 650. Each 4x4 block 650 can include 16 pixels, which can be represented in four rows and four columns in each respective block in a Cartesian plane or matrix. Pixels can include information representative of an image captured in a frame, such as luminance information, color information, and position information. In an example, a block such as the illustrated 16x16 pixel block can include a luminance block 660, which can include luminance pixels 662, and two chrominance blocks 670 / 680, such as a U or Cb chrominance block 670, and a V or Cr chrominance block 680. Chrominance blocks 670 / 680 can include chrominance pixels 690. For example, luminance block 660 can include 16x16 luminance pixels 662, and each chrominance block 670 / 680 can include 8x8 chrominance pixels 690, as illustrated. While one arrangement of blocks is illustrated, any arrangement can be used. While Figure 6 N x N blocks are illustrated, in some implementations, N x M blocks can be used, where N ≠ M. For example, 32x64 blocks, 64x32 blocks, 16x32 blocks, 32x16 blocks, or any other size of block can be used. In some implementations, N x 2N blocks, 2N x N blocks, or combinations thereof can be used.
[0070] In some implementations, video coding can include in-order block-level coding. In-order block-level coding can include coding blocks of a frame in order, such as in a raster scan order, where blocks can be identified and processed starting from a block in the upper left corner of the frame or portion of the frame, and proceeding to identify each block in order for processing from left to right along rows and from top rows to bottom rows. For example, superblocks in the top row and left column of a frame can be the first blocks coded, and superblocks immediately to the right of the first blocks can be the second blocks coded. The second row from the top can be the second row coded, such that superblocks in the left column of the second row can be coded after superblocks in the rightmost column of the first row.
[0071] In an example, coding a block can include using quadtree coding, which can include coding smaller block units with blocks in a raster scan order. For example, in Figure 6The 64x64 superblock shown in the lower left corner of the illustrated frame portion can be coded using quad-tree coding, where the upper left 32x32 block can be coded, then the upper right 32x32 block can be coded, then the lower left 32x32 block can be coded, and then the lower right 32x32 block can be coded. Each 32x32 block can be coded using quad-tree coding, where the upper left 16x16 block can be coded, then the upper right 16x16 block can be coded, then the lower left 16x16 block can be coded, and then the lower right 16x16 block can be coded. Each 16x16 block can be coded using quad-tree coding, where the upper left 8x8 block can be coded, then the upper right 8x8 block can be coded, then the lower left 8x8 block can be coded, and then the lower right 8x8 block can be coded. Each 8x8 block can be coded using quad-tree coding, where the upper left 4x4 block can be coded, then the upper right 4x4 block can be coded, then the lower left 4x4 block can be coded, and then the lower right 4x4 block can be coded. In some implementations, the 8x8 blocks can be omitted for 16x16 blocks, and the 16x16 blocks can be coded using quad-tree coding, where the upper left 4x4 block can be coded, and then the other 4x4 blocks in the 16x16 block can be coded in raster scan order.
[0072] In an example, video coding can include compressing information included in an original frame or input frame by omitting some of the information in the original frame from a corresponding encoded frame. For example, coding can include reducing spectral redundancy, reducing spatial redundancy, reducing temporal redundancy, or a combination thereof.
[0073] In an example, reducing spectral redundancy can include using a color model based on a luminance component (Y) and two chrominance components (U and V or Cb and Cr), which can be referred to as a YUV or YCbCr color model or color space. Using a YUV color model can include using a relatively large amount of information to represent a luminance component of a portion of a frame, and using a relatively small amount of information to represent each corresponding chrominance component of the portion of the frame. For example, a portion of a frame can be represented by a high resolution luminance component that can include a 16x16 block of pixels, and also by two lower resolution chrominance components that each represent the portion of the frame as an 8x8 block of pixels. Pixels can indicate values (e.g., values ranging from 0 to 255) and can be stored or transmitted using, for example, eight bits. While the disclosure is described with reference to a YUV color model, any color model can be used.
[0074] Reducing spatial redundancy can include transforming a block to a frequency domain as described above. For example, a unit of an encoder, such as a transform unit, can transform a block of pixels to a block of coefficients. The block of coefficients can be quantized, and then entropy coded. Figure 4The entropy encoding stage 408, which can use spatial frequency based transform coefficient values to perform a DCT.
[0075] Reducing temporal redundancy can include using similarities between frames to encode a frame using a relatively small amount of data based on one or more reference frames, which can be previously encoded, decoded, and reconstructed frames of the video stream. For example, a block or pixel of a current frame can be similar to a spatially corresponding block or pixel of a reference frame. A block or pixel of a current frame can be similar to a block or pixel of a reference frame at a different spatial location. As such, reducing temporal redundancy can include generating motion information indicating a spatial difference (e.g., a translation between a location of a block or pixel in a current frame and a corresponding location of a block or pixel in a reference frame).
[0076] Reducing temporal redundancy can include identifying a block or pixel in a reference frame or portion of a reference frame that corresponds to a current block or pixel of a current frame. For example, a reference frame or portion of a reference frame that can be stored in memory can be searched for a best block or pixel to use for encoding a current block or pixel of a current frame. For example, the search can identify a block in a reference frame for which a difference in pixel values between the reference block and the current block is minimized, and the search can be referred to as a motion search. The portion of the reference frame that is searched can be limited. For example, the portion of the reference frame that is searched, which can be referred to as a search area, can include a limited number of rows of the reference frame. In an example, identifying a reference block can include calculating a cost function, such as a sum of absolute differences (SAD), between pixels of a block in the search area and pixels of the current block.
[0077] As described above, a current block can be predicted using intra prediction. Intra prediction modes use pixels that are peripheral to the predicted current block. Pixels that are peripheral to the current block are pixels that are outside of the current block. Many different intra prediction modes can be available. Figure 7 A diagram that is an example of intra prediction modes.
[0078] Some intra prediction modes use a single value for all pixels within a prediction block that is generated using at least one peripheral pixel. For example, the VP9 codec includes an intra prediction mode, referred to as the true motion (TM_PRED) mode, in which all values of a prediction block have the value predicted pixel(x, y) = (top neighbor + left neighbor - top-left neighbor) for all x and y. As another example, the DC intra prediction mode (DC_PRED) is such that each pixel of a prediction block is set to the value predicted pixel(x, y) = average of the entire top row and left column.
[0079] Other intra prediction modes, which can be referred to as directional intra prediction modes, are such that each can have a corresponding prediction angle.
[0080] The intra prediction mode can be selected by the encoder as part of a rate-distortion loop. In brief, various intra prediction modes can be tested to determine which type of prediction has the lowest distortion for a given rate or number of bits to be transmitted in the coded video bitstream, including additional bits in the bitstream to indicate the type of prediction used.
[0081] In an example codec, the following 13 intra prediction modes can be available: DC_PRED, V_PRED, H_PRED, D45_PRED, D135_PRED, D117_PRED, D153_PRED, D207_PRED, D63_PRED, SMOOTH_PRED, SMOOTH_V_PRED, and SMOOTH_H_PRED, and PAETH_PRED. One of the 13 intra prediction modes can be used to predict a luma block.
[0082] Intra prediction mode 710 illustrates the V_PRED intra prediction mode, which is commonly referred to as the vertical intra prediction mode. In this mode, the first column of predicted block pixels is set to the value of the peripheral pixel A; the second column of predicted block pixels is set to the value of pixel B; the third column of predicted block pixels is set to the value of pixel C; and the fourth column of predicted block pixels is set to the value of pixel D.
[0083] Intra prediction mode 720 illustrates the H_PRED intra prediction mode, which is commonly referred to as the horizontal intra prediction mode. In this mode, the first row of predicted block pixels is set to the value of the peripheral pixel I; the second row of predicted block pixels is set to the value of pixel J; the third row of predicted block pixels is set to the value of pixel K; and the fourth row of predicted block pixels is set to the value of pixel L.
[0084] Intra prediction mode 730 illustrates the D117_PRED intra prediction mode, so called because the direction of the arrow along which the peripheral pixels forming a diagonal will propagate to generate the predicted block is at an angle of about 117° from the horizontal. That is, in D117_PRED, the prediction angle is 117°. Intra prediction mode 740 illustrates the D63_PRED intra prediction mode, which corresponds to a prediction angle of 63°. Intra prediction mode 750 illustrates the D153_PRED intra prediction mode, which corresponds to a prediction angle of 153°. Intra prediction mode 760 illustrates the D135_PRED intra prediction mode, which corresponds to a prediction angle of 135°.
[0085] Prediction modes D45_PRED and D207_PRED (not shown) correspond to prediction angles of 45° and 207°, respectively. DC_PRED corresponds to a prediction mode in which all of the prediction block pixels are set to a single value, which is a combination of the peripheral pixels A through M.
[0086] In the PAETH_PRED intra prediction mode, the prediction value of a pixel is determined as follows: 1) a base value is calculated as a combination of some of the peripheral pixels, and 2) the closest one of some of the peripheral pixels to the base value is used as the predicted pixel. The PAETH_PRED intra prediction mode is illustrated using, for example, pixel 712 (at position x=l, y=2). In an example of some of the peripheral pixel combinations, the base value can be calculated as base = B + K - M. That is, the base value is equal to: the value of the left peripheral pixel in the same row as the pixel to be predicted + the value of the above peripheral pixel in the same column as the pixel + the value of the top-left corner pixel.
[0087] In the SMOOTH_V intra prediction mode, the prediction pixels of the bottom-most row of the prediction block are estimated using the value of the last pixel in the left column (i.e., the value of the pixel at position L). The remaining pixels of the prediction block are calculated by quadratic interpolation in the vertical direction.
[0088] In the SMOOTH_H intra prediction mode, the prediction pixels of the right-most column of the prediction block are estimated using the value of the last pixel in the top row (i.e., the value of the pixel at position D). The remaining pixels of the prediction block are calculated by quadratic interpolation in the horizontal direction.
[0089] In the SMOOTH_PRED intra prediction mode, the prediction pixels of the bottom-most row of the prediction block are estimated using the value of the last pixel in the left column (i.e., the value of the pixel at position L), and the prediction pixels of the right-most column of the prediction block are estimated using the value of the last pixel in the top row (i.e., the value of the pixel at position D). The remaining pixels of the prediction block are calculated as a dilated weighted sum. For example, the value of the predicted pixel at position (i,j) of the prediction block can be calculated as a dilated weighted sum of the values of pixels L j , R, T i , and B. Pixel L j is the pixel in the left column and in the same row as the predicted pixel. Pixel R is the pixel provided by SMOOTH H. Pixel T i is the pixel in the above row and in the same column as the predicted pixel. Pixel B is the pixel provided by SMOOTH V. The weights can be equivalent to quadratic interpolation in the horizontal and vertical directions.
[0090] The intra prediction mode selected by the encoder can be sent to the decoder in the bitstream. The intra prediction mode can be entropy coded (encoded by the encoder and / or decoded by the decoder) using a context model.
[0091] Some codecs use the intra prediction mode of the left and above neighboring blocks as a context to code the intra prediction mode of the current block. Using Figure 7 For example, the left neighboring block can be a block containing pixels I through L, and the above neighboring block can be a block containing pixels A through D.
[0092] FIG. 770 illustrates the intra prediction modes available in the VP9 codec. The VP9 coding supports a set of 10 intra prediction modes for block sizes ranging from 4x4 to 32x32. These intra prediction modes are DC_PRED, TM_PRED, H_PRED, V_PRED, and 6 angular directional prediction modes: D45_PRED, D63_PRED, D117_PRED, D135_PRED, D153_PRED, D207_PRED, approximately corresponding to angles 45, 63, 117, 135, 153, and 207 degrees (measured counterclockwise from the horizontal axis).
[0093] Figure 8 is an example of an image portion 800 that includes a railroad track. The image portion 800 includes a first railroad track 802 and a second railroad track 804. In real life, the first railroad track 802 and the second railroad track 804 are parallel. However, in the image portion 800, the first railroad track 802 and the second railroad track 804 are shown converging at a focal point 803 outside of the image portion 800.
[0094] For purposes of illustration and clearer visualization, a current block 806 is superimposed on a portion of the image portion 800. Note that, generally, each location (i.e., unit) of a current block corresponds to one pixel or represents one pixel. However, for purposes of illustration and clarity, each unit of the current block 806 clearly includes significantly more than one pixel. Also note that, although not specifically labeled, each of the first track 802 and the second track 804 includes a pair of inner and outer lines. The lines in each pair of lines are also parallel and would converge at another focal point in the image portion 800.
[0095] The image portion 800 also includes a peripheral pixel 808. The peripheral pixel is shown above the current block 806. However, as noted above, the peripheral pixel can be an above pixel, a top-left pixel, a left pixel, or a combination thereof.
[0096] The unit 810 includes a portion of the first track 802. This portion propagates in a southwest direction into the current block 806. However, the portion of the second track 804 shown in the unit 812 propagates in a southeast direction into the current block 806.
[0097] As noted above, the unidirectional intra prediction mode cannot sufficiently predict the current block 806 from the peripheral pixels 808.
[0098] Figure 9 is an example of a flowchart of a technique 900 for determining (e.g., selecting, computing, deriving, etc.) positions along a line of peripheral pixels (i.e., a line of above peripheral pixels) for use in determining predicted pixel values in accordance with an embodiment of the present disclosure. The technique 900 for deriving positions is merely an example and other techniques are possible. For a prediction block (or equivalently, a current block) of size M x N, the technique 900 computes a block (e.g., a two-dimensional array) of size M x N. The two-dimensional array is referred to below as the array POSITIONS.
[0099] The technique 900 can be implemented, for example, as a software program that can be executed by a computing device such as the transmitting station 102 or the receiving station 106. The software program can include machine-readable instructions that can be stored in a memory such as the memory 204 or the secondary storage 214, and when executed by a processor such as the CPU 202, can cause the computing device to perform the technique 900. The technique 900 can be implemented in whole or in part in software and / or in hardware. Figure 4 The technique 900 can be implemented in the intra / inter prediction stage 402 of the encoder 400 of FIG. 1 and / or in the intra / inter prediction stage 508 of the decoder 500 of FIG. 2. The technique 900 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both can be used. Figure 5 The technique 900 can be implemented in the intra / inter prediction stage 402 of the encoder 400 of FIG. 1 and / or in the intra / inter prediction stage 508 of the decoder 500 of FIG. 2. The technique 900 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both can be used.
[0100] Given the current block and the set of above peripheral pixels (i.e., at integer peripheral pixel positions), the technique 900 determines, for each predicted pixel (i.e., or equivalently, each predicted pixel position) of the prediction block, a position along the line of above peripheral pixels from which to derive the value of the predicted pixel. As described further below, the position along the line of above peripheral pixels can be a sub-pixel position. Thus, the value at this position of the line of peripheral pixels can be derived (e.g., using interpolation) from the above peripheral pixels.
[0101] The technique 900 can be summarized as, for each row of the prediction block, resampling (e.g., repeatedly looking at, considering, etc.) the set of above peripheral pixels (e.g., i.e., the positions of the peripheral pixels), while at each resampling, shifting the positions according to one or more parameters of the intra prediction mode.
[0102] The positions can then be used to generate (e.g., compute, etc.) the prediction block for the current block. When the technique 900 is implemented by an encoder, the prediction block can be used to determine a residual block, which is then encoded in a compressed bitstream such as the compressed bitstream 420 of FIG. 1. Figure 4 When the technique 900 is implemented by a decoder, the prediction block can be generated by, for example, adding the prediction block to the current block from a compressed bitstream such as the compressed bitstream 420 of FIG. 1. Figure 5to reconstruct the current block from the compressed bitstream 420.
[0103] The peripheral pixel line can include the above peripheral pixels (e.g., Figure 7 pixels A to D, or pixels A to D and M. Figure 7 The peripheral pixel M of the current block can be considered as part of the above pixel, part of the left pixel, or can be referenced separately. In an example, the above peripheral pixels can include additional pixels (referred to herein as overhanging above pixels), such as Figure 7 and Figure 15 The peripheral pixels E to H of the current block. Note that, Figure 7 The overhanging left pixels are not shown in the current block.
[0104] The bit position calculated by the technique 900 can be related to a one-dimensional array that includes the peripheral pixels. For example, assume that the peripheral pixels available for predicting the current block are Figure 7 pixels A to H of the current block. The pixel values A to H can be stored in an array periph_pixels, such as periph_pixels = [0, 0, A, B, C, D, E, F, G, H]. An explanation of why the first two bit positions of the periph_pixels array are 0 is provided below.
[0105] In an example, the technique 900 can calculate negative bit positions that correspond to pixel positions outside of the block. The negative bit positions correspond to positions to the left of the block. In an example, pixel values that are too far away from the current pixel are not used as predictors for the pixel; whereas closer pixels can be used for prediction. In the case of a negative pixel position, as described further below, a pixel from the other border (e.g., a left peripheral pixel) can be determined to be closer (e.g., a better predictor).
[0106] As an implementation detail, periph_pixels can account for such cases by including empty positions in the array periph_pixels. The periph_pixels array is shown to include two (2) empty positions. Thus, if the technique 900 determines a bit position (e.g., calculated_position) of zero (0), then that bit position corresponds to the pixel value A (e.g., periph_pixels[empty_slots + calculated_position] = periph_pixels[2 + 0] = A). Similarly, if the technique 900 calculates a bit position of -2, then it corresponds to the pixel value periph_pixels[2 + (-2)] = [0] = 0. Similarly, the periph_pixels array can contain trained empty slots.
[0107] In an example, the technique 900 can receive as input, or have access to, one or more parameters of the intra prediction mode. The technique 900 can receive a size and a width of the current block (or equivalently, a size and a width of the prediction block). The technique 900 can also receive one or more of the following parameters: a horizontal offset (h off), a horizontal step (h st), a horizontal acceleration (h acc), a vertical offset (v off), a vertical step (v st), and a vertical acceleration (v acc). In an example, not receiving a parameter can be equivalent to receiving a zero value for the parameter. While the parameters are described below as additive quantities, in some examples, at least some of the parameters can equivalently be multiplicative values.
[0108] The horizontal offset (h off) parameter is an offset (i.e., a position offset) for each individual vertical pixel step. The horizontal offset (h off) can answer the question: for a new row (e.g., row = k) of the prediction block, along the line of the surrounding pixels, where does the bit of the first pixel of the previous row (e.g., row = k - 1) compared to (e.g., relative to) from which the value of the first pixel of the new row is derived? The horizontal offset can indicate an initial prediction angle. By initial prediction angle, it is meant the angle at which the first pixel of each row of the prediction block is predicted.
[0109] The horizontal step (h st) parameter is an offset that can be used for the next pixel in the horizontal direction. That is, for a given prediction pixel on a row of the prediction block, the horizontal step (h st) indicates a distance along the line of surrounding pixels to the next bit location relative to the bit location of the immediately preceding pixel on the same row. A horizontal step (h st) that is less than 1 (e.g., 0.95, 0.8, etc.) can achieve a zoom-out effect in the prediction block. For example, using a horizontal step (h st) that is less than 1, Figure 8 the rails will move away from each other in the prediction block. Similarly, using a horizontal step (h st) that is greater than 1 (e.g., 1.05, 1.2, etc.) can achieve a zoom-in effect in the prediction block.
[0110] The horizontal acceleration (h acc) parameter is a change that is added to each subsequent horizontal step (h st). The horizontal acceleration (h acc) parameter can be used as a shift to move the bit locations further and further apart rather than stepping from one bit location to the next along the line of surrounding pixels at a constant horizontal step. As such, the horizontal acceleration can achieve a transformation that is similar to a homographic transformation.
[0111] The vertical offset (v_off) parameter is the change in the horizontal offset (h_off) that will be applied to each subsequent row. That is, if h_off is used for the first row of the prediction block, then (h_off + v_off) is used for the second row, ((h_off + v_off) + v_off) is used for the third row, and so on. The vertical step size (v_st) parameter is the change in the horizontal step size (h_st) that will be applied to each subsequent row of the prediction block. The vertical acceleration (v_acc) parameter is the change in the horizontal acceleration (h_acc) that will be applied to each subsequent row of the prediction block.
[0112] Note that both accelerations in at least one direction (ie, the horizontal acceleration (h_acc) parameter and / or the vertical acceleration (v_acc) parameter) enable curve prediction. In other words, the acceleration parameters enable curve prediction.
[0113] At 902 , technique 900 initializes variables. Variable h_step_start can be initialized to a horizontal step size parameter: h_step_start=h_st. Variables h_offset and h_start can each be initialized to a horizontal offset: h_offset=h_off and h_start=h_off.
[0114] At 904, technique 900 initializes the outer loop variable i. Technique 900 performs 908 through 920 for each row of the prediction block. At 906, technique 900 determines whether there are more rows of the prediction block. If there are more rows, technique 900 proceeds to 908; otherwise, technique 900 ends at 922. When technique 900 ends at 922, each pixel position of the prediction position has a corresponding position along the peripheral pixel line from which the pixel value for each pixel position in the POSITIONS two-dimensional array is calculated (e.g., derived, etc.).
[0115] At 908 , technique 900 sets the bit variable p to the variable h_start (ie, p=h_start); and sets the variable h_step to the variable h_step_start (ie, h_step=h_step_start).
[0116] At 910, technique 900 initializes the inner loop variable j. Technique 900 performs 914 through 918 for each pixel in row i of the prediction block. At 912, technique 900 determines whether there are more pixel positions (i.e., more columns) for the row. If there are more columns, technique 900 proceeds to 914; otherwise, technique 900 proceeds to 920 to reset (i.e., update) the variables for the next row of the prediction block, if any.
[0117] At 914, the technique 900 sets the position (i,j) of the POSITIONS array to the bit position variable value p (i.e., POSITIONS(i,j) = p). At 916, the technique 900 advances the bit position variable p to the next horizontal bit position by adding h_step to the bit position variable p (i.e., p = h_step + p). At 918, in the presence of more unprocessed columns, the technique 900 prepares the variable h_step for the next column of row i of the prediction block. As such, the technique 900 adds the horizontal acceleration (h_acc) to the variable h_step (i.e., h_step = h_acc + h_step). The technique 900 returns from 918 to 912.
[0118] At 920, the technique 900 prepares (i.e., updates) the variables of the technique 900 in preparation for the next row of the prediction block, if any. Thus, for the next row (i.e., row = i + 1), the technique 900 updates h_start to h_start = h_offset + h_start; adds the vertical offset (v_off) to h_offset (i.e., h_offset = v_off + h_offset), adds the vertical step (v_st) to h_step_start (i.e., h_step_start = v_st + h_step_start), and adds the vertical acceleration (v_acc) to h_acc (i.e., h_acc = v_acc + h_acc).
[0119] Figure 10 is calculated by the technique of Figure 9 An example 1000 of bit positions (i.e., array POSITIONS) calculated by the technique of
[0120] Example 1000 shows, for each prediction block position (i, j), the position of the peripheral pixel lines from which the prediction value for position (i, j) should be derived, where i = 0, ..., columns - 1 and j = 0, ..., rows - 1. Predictor positions 1002 through 1008 show examples of position values for example 1000. Predictor position 1002, corresponding to prediction block position (3, 1), is to derive its prediction value from position 2.93 of the peripheral prediction line. Predictor position 1004, corresponding to prediction block position (6, 4), is to derive its prediction value from position 6.74 of the peripheral prediction line. Predictor position 1006, corresponding to prediction block position (0, 1), is to derive its prediction value from position -0.4 of the peripheral prediction line. Predictor position 1008, corresponding to prediction block position (0, 6), is to derive its prediction value from position -1.4 of the peripheral prediction line.
[0121] Figure 11 It is from Figure 10 Example 1100 of a prediction block calculated by example 1000. Example 1100 includes prediction block 1102, which is visualized as prediction block 1104. Prediction block 1102 (and, equivalently, prediction block 1104) is calculated using Figure 10 The example 1000 of FIG. 100 and the position of the peripheral predicted pixels 1106 which may be the top (ie, above) peripheral pixels are derived (eg, generated, calculated, etc.). The peripheral predicted pixels 1106 can be, for example, Figure 7 The peripheral predicted pixels 1108 are visualizations of the peripheral predicted pixels 1106 .
[0122] In the visualization, pixel value zero (0) corresponds to a black square, while pixel value (255) corresponds to a white square. Pixel values between 0 and 255 correspond to different shades of gray. As such, a luma block is shown as an example of an intra prediction mode according to the present disclosure. However, the present disclosure is not limited thereto. The disclosed techniques are also applicable to chroma blocks or any other color component blocks. In general, the techniques disclosed herein are applicable to any prediction block to be generated, which can be of any size M×N, where M and N are positive integers.
[0123] As mentioned above and as Figure 10 As shown in Example 1000, by Figure 9 The positions calculated by the technique 900 can be non-integer positions (i.e., sub-pixel positions) for peripheral pixel lines. Pixel values for the peripheral pixel lines at non-integer positions can be derived (e.g., calculated) from available integer pixel position values (i.e., peripheral pixels) such as the peripheral predicted pixels 1106.
[0124] Many techniques can be available for computing the sub-pixel location (i.e., the pixel value at the sub-pixel location). For example, a multi-tap (e.g., 4-tap, 6-tap, etc.) finite impulse response (FIR) filter can be used. For example, an average of surrounding pixels can be used. For example, bilinear interpolation can be used. For example, a 4-pixel bicubic interpolation of the surrounding pixels (e.g., the top row or left column) can be used. For example, a convolution operation can be used. The convolution operation can use pixels other than the surrounding pixels. In an example, a convolution kernel of size N x N can be used. Thus, N rows (columns) of the above (left) neighboring block can be used. To illustrate, a 4-tap filter or 4 x 1 convolution kernel with weights (-0.1, 0.6, 0.6, -0.1) can be used. Thus, the pixel value computed at the location between pixel1 and pixel2 of the set of four pixels (pixel0, pixel1, pixel2, pixel3) can be computed as clamp(-0.10*pixel0+0.6*pixel1+0.6*pixel2-0.10*pixel3, 0, 255), where the clamp() operation sets computed values less than zero to zero and computed values greater than 255 to 255.
[0125] The prediction block 1102 illustrates the use of bilinear interpolation. For a sub-pixel location of a surrounding pixel line, the two closest integer pixels are found. The prediction operator value is computed as a weighted sum of the two closest integer pixels. The weights are determined according to the distance of the sub-pixel location to the two integer pixel locations.
[0126] Given a location pos of a surrounding pixel line, the pixel value at pos, pix_val, can be computed as pix_val = left_weight * left_pixel + right_weight * right_pixel. left_pixel is the pixel value of the closest left-neighbor integer pixel to location pos. right_pixel is the pixel value of the closest right-neighbor integer pixel to location pos.
[0127] The location of left_pixel can be computed as left_pos = floor(pos), where floor() is a function that returns the largest integer less than or equal to pos. Thus, floor(6.74) = 6. The location 6.74 is for the prediction operator position 1004 of Figure 10 The location of right_pixel can be computed as right_pos = ceiling(pos), where ceiling() is a function that returns the smallest integer greater than or equal to pos. Thus, ceiling(6.74) = 7.
[0128] In the example, right_weight can be calculated as right_weight=pos−left_pos and left_weight can be calculated as left_weight=1−right_weight. Thus, for the position 6.74 of the predictor position 1004, right_weight=6.74−6=0.74 and left_weight=1−0.74=0.26. The pixel value 1112 of the prediction block 1102 is obtained from Figure 10 The value obtained by the predictor position 1004 of . As such, the pixel value 1112 is calculated as ((255×.26)+(0×0.74))=66.
[0129] Similarly, the predicted operator value 1110 is obtained from Figure 10 The predictor position 1006 (i.e., -0.4) is calculated. Therefore, left_pos and right_pos are -1 and 0, respectively. right_weight and left_weight are 0.6 and 0.4, respectively. The right neighboring pixel value is the peripheral pixel value of the peripheral predicted pixel 1106 at position 0. Therefore, the right neighboring pixel value is the pixel 1114 with a value of 0. The left neighboring pixel is not available. Therefore, the left neighboring pixel value is 0. Like this, the predictor value 1110 is calculated as ((0×0.4)+(0×0.6))=0.
[0130] The above about Figure 10 and 11 We discuss how to use the parameters (ie, parameter values) of the intra prediction modes. There are any number of ways to select parameter values.
[0131] In an example, the mode selection process of the encoder can test all possible parameter values to find the optimal combination of parameter values that produces the minimum residual. In an example, the minimum residual can be the residual that produces the best rate-distortion value. In an example, the minimum residual can be the residual that produces the minimum residual error. The minimum residual error can be the mean square error. The minimum residual error can be the sum of absolute difference errors. Any other suitable error measurement can be used. In addition to the indication of the intra-frame prediction mode itself, the encoder can also encode the parameter values of the optimal combination of parameter values in the encoded bitstream. The decoder can decode the parameter values of the optimal combination of parameter values. In an example, as described herein, the indication of the intra-frame prediction mode itself can be a number (e.g., an integer) that instructs the decoder to use the intra-frame prediction parameters to perform intra-frame prediction of the current block.
[0132] Testing all possible values for each parameter may be an impractical solution. As such, it is possible to select a limited number of values for each parameter and test a limited number of combinations of values.
[0133] For example, the horizontal offset (h_off) parameter can be selected from a limited range of values. In an example, the limited range of values can be [-4, +4]. A step size value can be used to select horizontal offset (h_off) parameter values to test within the limited range. In an example, the step size can be 0.25 (or some other value). Like this, the values -4, -3.75, -3.5, -3.25, ..., 3.75, 4 can be tested. In an example, the vertical offset (v_off) parameter can be selected similarly to the horizontal offset parameter.
[0134] With respect to the horizontal step size (h_st) and the vertical step size (v_st), a value relatively close to 1 can be selected. Any other value can result in too fast scaling of the prediction block. Therefore, values in the range of [0.9, 1.1] can be tested for the horizontal step size (h_st) and the vertical step size (v_st). However, typically, the horizontal step size (h_st) and the vertical step size (v_st) can each be selected from the range of [-4, 4] using a step size value that can be 0.25. The selected horizontal acceleration (h_acc) and vertical acceleration (v_acc) parameter values can be close to 0. In an example, the horizontal acceleration (h_acc) and vertical acceleration (v_acc) parameter values can each be 0 or 1. More generally, the horizontal parameter and the corresponding vertical parameter can have a value and / or a range of values.
[0135] In another example, the encoder can select parameter values based on a set of possible best parameter values. The set of possible best parameter values is also referred to herein as predicted parameter values. The set of possible best parameter values can be derived by predicting the peripheral pixels from their neighboring pixels. That is, in the case of upper peripheral pixels, the peripheral pixels constitute the bottommost row of the previously reconstructed block; and in the case of left peripheral pixels, the peripheral pixels constitute the rightmost column of the previously reconstructed block. Therefore, the neighboring rows, columns, or both (as the case may be) of the peripheral pixels can be used as prediction operators for the peripheral pixels. Since the prediction operators for the peripheral pixels and the peripheral pixels themselves are known, parameter values can be derived from them. In this case, the encoder does not need to encode the set of possible best parameter values in the compressed bitstream, because the decoder can perform an exactly similar process as the encoder to derive the set of possible best parameter values. Therefore, all the encoder needs to encode in the bitstream is an indication of the intra-frame prediction mode itself.
[0136] In another example, the differential parameter values can be encoded by the encoder. For example, as described above, the encoder can derive the optimal parameter values, and as also described above, can derive a set of possible optimal parameter values (i.e., predicted parameter values). Then, in addition to the indication of the intra prediction mode, the encoder encodes the respective difference between the optimal parameter values and the set of possible optimal parameter values. That is, for example, with respect to the horizontal offset (h off), the encoder can derive an optimal horizontal offset (opt h off) and a possible optimal horizontal offset (predicted h offset). The encoder then encodes the difference (opt h off - predicted h offset).
[0137] Figure 12 is a flowchart of a technique 1200 for intra prediction of a current block according to embodiments of the present disclosure. The intra prediction mode uses pixels of a current block periphery. The pixels of the current block periphery can be previously predicted pixels in the same video frame or picture as the current block. The current block can be a luma block, a chroma block, or any other color component block. The current block can be of size M x N, where M and N are positive integers. In an example, M is equal to N. In an example, M is not equal to N. For example, the current block can be of size 4 x 4, 4 x 8, 8 x 4, 8 x 8, 16 x 16, or any other current block size. The technique 1200 generates a prediction block for the current block that is of the same size as the current block. The technique 1200 can be implemented in an encoder, such as the encoder 400 of Figure 4 The technique 1200 can be implemented in a decoder, such as the decoder 500 of Figure 5 The technique 1200 can be implemented in a decoder, such as the decoder 500 of
[0138] The technique 1200 can be implemented as, for example, a software program executable by a computing device, such as the transmitting station 102 or the receiving station 106. The software program can include machine-readable instructions that can be stored in a memory, such as the memory 204 or the secondary storage 214, and executed by a processor, such as the CPU 202, to cause the computing device to perform the technique 1200. In at least some embodiments, the technique 1200 can be performed, in whole or in part, by the intra / inter prediction stage 402 of the encoder 400. Figure 4 The technique 1200 can be implemented in a decoder, such as the decoder 500 of Figure 5 The technique 1200 can be implemented in a decoder, such as the decoder 500 of
[0139] The techniques 1200 can be implemented using specially designed hardware or firmware. Some computing devices can have multiple memories, multiple processors, or both. The steps or operations of the techniques 1200 can be distributed using different processors, memories, or both. Use of the term "processor" or "memory" in singular form includes computing devices having one processor or one memory as well as devices having multiple processors or multiple memories that can be used to perform some or all of the described steps.
[0140] At 1202, the techniques 1200 select peripheral pixels of the current block. The peripheral pixels are used to generate a prediction block for the current block. In an example, the peripheral pixels can be pixels above the current block. In an example, the peripheral pixels can be pixels to the left of the current block. In an example, the peripheral pixels can be a combination of the pixels above and to the left of the current block. When implemented by a decoder, selecting the peripheral pixels can include reading (e.g., decoding) from a compressed bitstream an indication (e.g., syntax elements) that indicates which peripheral pixels are to be used.
[0141] For each position (i.e., pixel position) of the prediction block, the techniques 1200 perform 1206-1208. Thus, if the current block is of size M x N, the prediction block can include M x N pixel positions. As such, at 1204, the techniques 1200 determine whether there are more pixel positions of the prediction block that have not yet been performed 1206-1208. If there are more pixel positions, the techniques 1200 proceed to 1206; otherwise, the techniques 1200 proceed to 1210.
[0142] At 1206, the techniques 1200 select two respective pixels of the peripheral pixels for the pixel position of the prediction block. In an example, the techniques 1200 first select a position along a line of consecutive peripheral pixels, along which the peripheral pixels are integer pixel positions. At 1208, the techniques 1200 compute a predicted pixel (i.e., pixel value) for the pixel position of the prediction block by interpolating the two respective pixels.
[0143] In an example, selecting a position along a line of consecutive peripheral pixels can be as described with respect to Figure 9 Thus, selecting the two respective pixels of the peripheral pixels can include selecting a first two respective pixels of the peripheral pixels for computing a first predicted pixel of the prediction block, and selecting a second two respective pixels of the peripheral pixels for computing a second predicted pixel of the prediction block. The second predicted pixel can be a horizontally neighboring pixel of the first predicted pixel. The first two respective pixels and the second two respective pixels can be selected according to an intra-prediction mode parameter.
[0144] As noted above, the intra prediction mode parameters can include at least two of a horizontal offset, a horizontal step, or a horizontal acceleration. In an example, the intra prediction mode parameters can include the horizontal offset, the horizontal step, and the horizontal acceleration. As noted above, the horizontal offset can indicate an initial prediction angle; the horizontal step can be used as a subsequent offset for subsequent prediction pixels of the same row; and the horizontal acceleration can indicate a change to the horizontal offset that is added to each subsequent horizontal step.
[0145] In an example, the horizontal offset can be selected from a limited range based on a step value. In an example, the limited range can be -4 to 4. In an example, the step value can be 0.25. In an example, the horizontal (vertical) step can be selected from a range of -4 to 4 based on a step value. The step value can be 0.25 or some other value. In an example, the horizontal (vertical) acceleration can be 0. In another example, the horizontal (vertical) acceleration can be 1.
[0146] As further noted above, the intra prediction mode parameters can also include at least two of a vertical offset, a vertical step, or a vertical acceleration. In an example, the intra prediction mode parameters can include the vertical offset, the vertical step, and the vertical acceleration. The vertical offset can indicate a first change to the horizontal offset to be applied to each subsequent row of the prediction block. The vertical step can indicate a second change to the horizontal step to be applied to each subsequent row of the prediction block. The vertical acceleration can indicate a third change to the horizontal acceleration to be applied to each subsequent row of the prediction block.
[0147] In an example, calculating the prediction pixel by interpolating the two respective pixels can include using bilinear interpolation to calculate the prediction pixel.
[0148] At 1210, the technique 1200 encodes a residual block corresponding to a difference between the current block and the prediction block. When implemented by an encoder, the technique 1200 encodes the residual block in a compressed bitstream. When implemented by a decoder, the technique 1200 decodes the residual block from a compressed bitstream. The decoded residual block can be added to the prediction block to reconstruct the current block.
[0149] When implemented by a decoder, the technique 1200 can also include decoding the intra prediction mode parameters from the compressed bitstream. In another example, and as noted above, the intra prediction mode parameters can be derived by predicting the peripheral pixels from other pixels that include previously reconstructed peripheral pixels according to the intra prediction mode parameters.
[0150] Figure 13 is a flowchart of a technique 1300 for generating a prediction block for a current block using intra prediction according to an embodiment of the disclosure. The intra prediction mode uses pixels that are peripheral to the current block, which can be as described with respect to Figure 12The current block can be as described in relation to the technique 1200. Figure 12 The technology 1200 is described. The technology 1300 can be used in Figure 4 The technique 1300 can be implemented in an encoder such as Figure 5 The decoder 500 is implemented in the decoder.
[0151] Technique 1300 can be implemented, for example, as a software program that can be executed by a computing device such as sending station 102 or receiving station 106. The software program can include machine-readable instructions that can be stored in a memory such as memory 204 or secondary storage 214, and the machine-readable instructions can be executed by a processor such as CPU 202 to cause the computing device to perform technique 1300. In at least some embodiments, technique 1300 can be implemented in whole or in part by Figure 4 In other embodiments, the technique 1300 can be performed in whole or in part by the intra / inter prediction stage 402 of the encoder 400. Figure 5 This is performed by the intra / inter prediction stage 508 of the decoder 500 .
[0152] Technique 1300 can be implemented using specialized hardware or firmware. Some computing devices can have multiple memories, multiple processors, or both. The steps or operations of technique 1300 can be distributed using different processors, memories, or both. Use of the terms "processor" or "memory" in the singular includes computing devices having one processor or one memory, as well as devices having multiple processors or multiple memories used to perform some or all of the steps.
[0153] At 1302, the technique 1300 determines peripheral pixels for generating a prediction block for a current block. Determining peripheral pixels can mean selecting which peripheral pixels to use, such as with respect to Figure 12 As described in 1202. Peripheral pixels can be considered as integer pixel positions along a peripheral pixel line (ie, a continuous line of peripheral pixels).
[0154] At 1304, the technique 1300 determines, for each pixel of the prediction block, corresponding sub-pixel positions for the peripheral pixel lines. Determining corresponding sub-pixel positions can be as described with respect to Figure 9 As described. As used herein, sub-pixel positions of a peripheral pixel line also include integer pixel positions. That is, for example, the determined sub-pixel position can be the position of one of the peripheral pixels itself.
[0155] At 1306, for each prediction pixel of the prediction block, the technique 1300 calculates the prediction pixel as an interpolation of the integer pixels of the peripheral pixels corresponding to the respective sub-pixel position of each prediction pixel. In an example, the interpolation can be an interpolation of the closest integer pixels, as described above with respect to Figure 11 In an example, a bilinear interpolation can be used. In another example, a filtering of the integer pixels can be performed to obtain the prediction pixels of the prediction block.
[0156] In an example, determining the respective sub-pixel position of the peripheral pixels for each pixel of the prediction block can include, for each row pixel of a first row of the prediction block, determining the respective sub-pixel position using parameters including at least two of a horizontal offset, a horizontal step, or a horizontal acceleration. As described above, the horizontal offset can indicate an initial prediction angle. As described above, the horizontal step can be used as a subsequent offset for subsequent prediction pixels of the same row. As described above, the horizontal acceleration can indicate a change to the horizontal offset that is added to each subsequent horizontal step.
[0157] In an example, determining the respective sub-pixel position of the peripheral pixels for each pixel of the prediction block can include determining the respective sub-pixel position of each pixel of a second row of the prediction block, where the parameters further include at least two of a vertical offset, a vertical step, or a vertical acceleration. As described above, the vertical offset can indicate a first change to the horizontal offset to be applied to each subsequent row of the prediction block. As described above, the vertical step can indicate a second change to the horizontal step to be applied to each subsequent row of the prediction block. As described above, the vertical acceleration can indicate a third change to the horizontal acceleration to be applied to each subsequent row of the prediction block.
[0158] In an example, and when implemented by a decoder, the technique 1300 can include decoding the parameters from a compressed bitstream. In an example, decoding the parameters from the compressed bitstream can include, as described above, decoding the parameter differences; deriving the prediction parameter values; and for each parameter, adding the respective parameter difference to the respective prediction parameter value.
[0159] As described above, in embodiments, a directional prediction mode can be used to generate an initial prediction block. A warping (e.g., a warping function, a set of warping parameters, parameters, etc.) can then be applied to the initial prediction block to generate the prediction block.
[0160] In an example, the warp parameters can be derived using the current block and the initial prediction block. Any number of techniques can be used to derive the warp parameters. For example, a random sample consensus (RANSAC) method can be used to fit a model (i.e., warp model, parameters) to matching points between the current block and the initial prediction block. RANSAC is an iterative algorithm that can be used to estimate the warp parameters (i.e., parameters) between two blocks. In an example, the best matching pixels between the current block and the initial prediction block can be used to derive the warp parameters. The warp can be a homography warp, an affine warp, a similarity warp, or some other warp.
[0161] A homography warp can use eight parameters to project some pixels of the current block to some pixels of the initial prediction block. A homography warp is not constrained to be a linear transformation between two spatial coordinates. As such, the eight parameters that define a homography warp can be used to project pixels of the current block to a quadrilateral portion of the initial prediction block. Thus, a homography warp supports translation, rotation, scaling, aspect ratio change, shear, and other non-parallelogram warps.
[0162] An affine warp uses six parameters to project pixels of the current block to some pixels of the initial prediction block. An affine warp is a linear transformation between two spatial coordinates defined by six parameters. As such, the six parameters that define an affine warp can be used to project pixels of the current block to a parallelogram that is a portion of the initial prediction block. Thus, an affine warp supports translation, rotation, scaling, aspect ratio change, and shear.
[0163] A similarity warp uses four parameters to project pixels of the current block to pixels of the initial prediction block. A similarity warp is a linear transformation between two spatial coordinates defined by four parameters. For example, the four parameters can be a translation along an x-axis, a translation along a y-axis, a rotation value, and a scaling value. As such, the four parameters that define a similarity model can be used to project pixels of the current block to a square of the initial prediction block. Thus, a similarity warp supports a square-to-square transformation with rotation and scaling.
[0164] In an example, in addition to the directional intra prediction mode, the parameters of the warp can be sent from the encoder to the decoder in the compressed bitstream. The decoder can use the directional intra prediction mode to generate the initial prediction block. The decoder can then use the parameters of the sent parameters to decode the current block.
[0165] In another example, the parameters can be derived by the decoder. For example, as described above, the decoder can use previously decoded pixels to determine the parameters of the warp. For example, as described above, the pixels of the same block as the peripheral pixels can be used to predict the peripheral pixels, thereby determining the warp parameters. That is, since the peripheral pixels are known, the best warp parameters can be determined to predict the peripheral pixels from neighboring pixels of the peripheral pixels.
[0166] In another example, the differential warp parameters can be sent by the encoder, as described above. For example, the predicted warp parameters can be derived using neighboring pixels of the peripheral parameters, and the optimal warp parameters are derived as described above. The difference between the optimal warp parameters and the predicted warp parameters can be sent in the compressed bitstream.
[0167] As described above, at least with respect to Figures 9-11 is the case where only the above peripheral pixels are used. For example, in the case where other peripheral pixels (e.g., left peripheral pixels) are not available, only the above peripheral pixels can be used. For example, only the above peripheral pixels can be used even when the left peripheral pixels are available.
[0168] Using the left peripheral pixels can be similar to using the above peripheral pixels. In an example, only the left peripheral pixels can be used in the case where the above peripheral pixels are not available. In another example, only the left peripheral pixels can be used even when the above peripheral pixels are available.
[0169] Figure 18 is an example of a flowchart of a technique 1800 for determining predicted pixel values according to an implementation of the disclosure along a line of left peripheral pixel positions. For a prediction block (or equivalently, a current block) of size M x N, the technique 1800 computes a block (e.g., a two-dimensional array) of size M x N. The two-dimensional array is referred to below as the array POSITIONS.
[0170] Given the current block and a set of left peripheral pixel positions (i.e., at integer peripheral pixel positions), the technique 1800 determines, for each predicted pixel of the prediction block (i.e., or equivalently, each predicted pixel position), a position along the line of left peripheral pixel positions from which to derive the value of the predicted pixel. As described further below, the position along the line of left peripheral pixel positions can be a sub-pixel position. Thus, the value at that position of the line of left peripheral pixel positions can be derived (e.g., using interpolation) from the left peripheral pixel.
[0171] The technique 1800 can be summarized as, for each column of the prediction block, resampling (e.g., repeatedly looking at, considering, etc.) the set of left peripheral pixel positions (e.g., i.e., the positions of the left peripheral pixels), while at each resampling, shifting the positions according to one or more parameters of the intra-prediction mode. The positions can then be used to generate (e.g., compute, etc.) the prediction block for the current block.
[0172] A detailed description of technique 1800 is omitted because technique 1800 is very similar to technique 900. In technique 1800, the roles (i.e., uses) of the horizontal offset (h_off), the horizontal step (h_st), the horizontal acceleration (h_acc), the vertical offset (v_off), the vertical step (v_st), and the vertical acceleration (v_acc) are reversed from their roles in technique 900. That is, wherever a horizontally-related parameter is used in process 900, the corresponding vertically-related parameter is used instead in process 1800, and vice versa. Thus, 1802 through 1822 can be similar to 902 through 922, respectively. It should also be noted that while the outer iteration (at 906) of technique 900 iterates over the rows of the prediction block and the inner iteration (at 912) iterates over the columns of the prediction block, in technique 1800, the outer iteration (at 1906) iterates over the columns of the prediction block and the inner iteration (at 1812) iterates over the rows of the prediction block.
[0173] Figure 19 is a flowchart of a technique 1900 for generating a prediction block for a current block using intra prediction in accordance with an embodiment of the present disclosure. Technique 1900 uses both above and left neighboring pixels to generate the prediction block. The current block can be as described with respect to Figure 12 Technique 1200. Technique 1900 can be implemented in an encoder, such as Figure 4 encoder 400. Technique 1900 can be implemented in a decoder, such as Figure 5 decoder 500.
[0174] Technique 1900 can be implemented, for example, as a software program executable by a computing device, such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that can be stored in a memory, such as memory 204 or secondary storage 214, and executed by a processor, such as CPU 202, to cause the computing device to perform technique 1900. In at least some embodiments, technique 1900 can be performed, in whole or in part, by intra / inter prediction stage 402 of encoder 400. Figure 4 In other embodiments, technique 1900 can be performed, in whole or in part, by intra / inter prediction stage 508 of decoder 500. Figure 5
[0175] The techniques 1900 can be implemented using specially designed hardware or firmware. Some computing devices can have multiple memories, multiple processors, or both. The steps or operations of the techniques 1900 can be distributed using different processors, memories, or both. Use of the term "processor" or "memory" in singular form includes computing devices having one processor or one memory, as well as devices having multiple processors or multiple memories that can be used to perform some or all of the described steps.
[0176] The techniques 1900 are shown with reference to Figure 20A and Figure 20B . Figures 20A-20B With the following inputs: a current block size of 8x8, a horizontal offset h off = 0.25, a horizontal step h st = 1, a horizontal acceleration h acc = 1, a vertical offset v off = 4, a vertical step v st = 1, and a vertical acceleration v acc = 0.
[0177] At 1902, the process 1900 selects a first peripheral pixel of the current block. The first peripheral pixel is along a first edge of the current block. The first peripheral pixel is selected as the primary pixel to produce the prediction block. In an example, the first peripheral pixel can be an above peripheral pixel. In an example, the first peripheral pixel can be a left peripheral pixel. As described further below, the primary pixel is the pixel for which the bits along the first peripheral pixel line are computed, as described with respect to the techniques 900 (in the case that the first peripheral pixel is an above peripheral pixel) or the techniques 1800 (in the case that the first peripheral pixel is a left peripheral pixel).
[0178] When implemented by an encoder, the techniques 1900 can be performed once using the top peripheral pixel as the first peripheral pixel and performed a second time using the left peripheral pixel as the first peripheral pixel. The encoder can test both to determine which provides a better prediction of the current block. In an example, the techniques 1900 can send in a compressed bitstream, such as the compressed bitstream 420, a first indication of which of the first peripheral pixel or the second peripheral pixel is used as the primary pixel for generating the prediction block. That is, for example, the first indication can be a bit value of 0 (1) when the above (left) peripheral pixel is used as the primary pixel. Other values of the first indication are possible. Figure 4
[0179] Thus, when implemented by a decoder, the techniques 1900 can select the first peripheral pixel by decoding the first indication from the compressed bitstream. In another example, the decoder can derive whether the first peripheral pixel is the above peripheral pixel or the left peripheral pixel by predicting the above row (left column) from its neighboring above row (left column) according to the techniques described herein.
[0180] For each position (i.e., pixel position) of the prediction block, the technique 1900 performs 1906-1910. Thus, if the size of the current block is MxN, the prediction block can include M*N pixel positions. As such, at 1904, the technique 1900 determines whether there are more pixel positions of the prediction block that have not been performed 1906-1910. If there are more pixel positions, the technique 1900 proceeds to 1906; otherwise, the technique 1900 proceeds to 1910.
[0181] At 1906, the technique 1900 determines a first intercept along a first continuous line that includes a first peripheral pixel at a respective integer position. Thus, in the case where the first pixel position is a top peripheral pixel, then the first intercept is a y-axis intercept; and, in the case where the first pixel position is a left peripheral pixel, then the first intercept is an x-axis intercept.
[0182] Figure 20A A position 2010 of the values of the two-dimensional array POSITIONS described above is shown when a top peripheral pixel is used as the first peripheral pixel primary pixel. The position 2010 can be computed using the technique 900. Thus, the position 2010 provides a position within a top peripheral row.
[0183] Example Figure 20B A position 2050 of the values of the two-dimensional array POSITIONS is shown when a left peripheral pixel is used as the first peripheral pixel primary pixel. The position 2050 can be computed using the technique 1800. Thus, the position 2050 provides a position within a left peripheral column.
[0184] At 1908, the technique 1900 determines a second intercept along a second continuous line that includes a second peripheral pixel using the position of each prediction pixel and the first intercept. The second peripheral pixel is along a second edge of the current block that is perpendicular to the first edge.
[0185] In the case where the first prediction pixel is a top prediction pixel, then the second prediction pixel can be a left prediction pixel. In the case where the first prediction pixel is a left prediction pixel, then the second prediction pixel can be a top prediction pixel. Other combinations of first and second prediction pixels are possible. For example, combinations of a top-right peripheral pixel, a bottom-right peripheral pixel, or a bottom-left peripheral pixel are possible.
[0186] In an example, the second intercept can be computed by connecting a line between the prediction pixel position and the first intercept and extending the line towards the second continuous line.
[0187] Further described below Figure 16Now being used as an illustration. If the current prediction pixel is prediction pixel 1604 and the first intercept is y-intercept 1648, then the second intercept can be obtained by connecting prediction pixel 1604 and y-intercept 1648 and extending the line toward the x-axis. Thus, the second intercept is x-intercept 1646. Similarly, if the current prediction pixel is prediction pixel 1604 and the first intercept is x-intercept 1646, then the second intercept can be obtained by connecting prediction pixel 1604 and x-intercept 1646 and extending the line toward the y-axis. Thus, the second intercept is y-intercept 1648.
[0188] Figure 20A Block 2012 of FIG. 20 illustrates an x-intercept when the first peripheral pixel is the above peripheral pixel. Figure 20B Block 2052 of FIG. 20 illustrates a y-intercept when the first peripheral pixel is the left peripheral pixel. The intercept can be calculated using one of the formulas of equation (1).
[0189] For illustration, consider prediction pixel 2010A at position (1, 3) of the prediction. Thus, considering the top and left prediction pixels, prediction pixel 2010A is at position (2, 4) of a coordinate system having an origin at which the top and left peripheral pixels intersect. The first intercept (i.e., the y-intercept) is at 7.5. Thus, a line can be formed from the two points (2, 4) and (0, 7.5). Thus, as shown by x-intercept value 2012A, the x-intercept can be calculated as (-7.5 / ((7.5 - 4) / (0 - 2))) = 4.29.
[0190] As another example, prediction pixel 2050A at position (1, 1) of the prediction block is at position (2, 2) of a coordinate system including the top and left peripheral pixels. The first intercept (i.e., the x-intercept) is at 11. Thus, a line can be formed from the two points (2, 2) and (11, 0). Thus, as shown by y-intercept value 2052A, the y-intercept can be calculated as (-11 * (0 - 2) / (11 - 2)) = 2.44.
[0191] At 1910, the technique 1900 calculates a value for each prediction pixel using at least one of the first intercept and the second intercept.
[0192] In an example, the value of the prediction pixel can be calculated as a weighted sum of the first intercept and the second intercept. More particularly, the value of the prediction pixel can be calculated as a weighted sum of a first pixel value at the first intercept and a second pixel value at the second intercept. The weights can be inversely proportional to the distance from the position of the prediction pixel to each of the first intercept and the second intercept. As such, the value of the prediction pixel can be obtained as a bilinear interpolation of the first pixel value and the second pixel value. As is well known, the distance between two points (a, b) and (c, d) can be calculated as
[0193] In another example, the value of the prediction pixel can be computed based on the closest distance instead of the weighted sum. That is, the x-intercept and y-intercept that is closest (based on the computed distance) to the prediction pixel can be used to compute the value of the prediction pixel.
[0194] Figure 20A Distances 2014 and 2016 show the distance from each prediction pixel location to their respective y-intercept and x-intercept in the case where the first peripheral pixel is the left peripheral pixel. For example, for the prediction pixel at location (2, 4) (i.e., prediction pixel 2010A at location (1, 3) of the prediction block) has a y-intercept of 7.5 (i.e., point (0, 7.5)), the distance 2014A is and the distance 2016A to the x-intercept (4.29, 0) is given by Similarly, Figure 20B Distances 2054 and 2056 show the distance from each prediction pixel location to their respective y-intercept and x-intercept in the case where the first peripheral pixel is the left peripheral pixel.
[0195] Thus, for the prediction pixel at location (1, 4), the weights for the x-intercept and y-intercept are (4.61 / (4.61+4.03)) = 0.53, and (1 - 0.53) = 0.47, respectively, in the case where the weighted sum is used. The pixel value for each of the x-intercept and y-intercept can be computed (e.g., as an interpolation of the two closest integer pixel values) as described above.
[0196] In the case where the closest distance is used, the prediction pixel value at prediction block location (1, 4) can be computed using only the pixel at the y-intercept, since 4.03 (i.e., distance 2014A) is less than 4.61 (i.e., distance 2016A).
[0197] As described below with respect to Figure 14 In some cases, the x-intercept or y-intercept can not be valid (e.g., is negative). As also described with respect to Figure 14 In such cases, the prediction value can only be computed based on the available intercept value.
[0198] In the encoder, the technique 1900 can select one of the closest distance, the weighted sum, or some other function to generate the prediction block. In an example, the technique 1900 can generate a respective prediction block for each possible function and select one of the prediction blocks that produces the optimal prediction. As such, the technique 1900 can encode a second indication of the selected function in the compressed bitstream to combine the first peripheral pixel and the second peripheral pixel to obtain the prediction block. As described, the function can be the weighted sum, the closest distance, or some other function.
[0199] When implemented by a decoder, the technique 1900 receives, in a compressed bitstream, a second indication of a function for combining first peripheral pixels and second peripheral pixels to obtain a prediction block. The decoder is operable to generate the prediction block.
[0200] Figure 21 shows when the upper peripheral pixels are used as main peripheral pixels Figure 19 An example of the technique 2100. An example prediction block generated using the left peripheral block as the primary peripheral pixel is not shown.
[0201] Example 2100 illustrates different prediction blocks that can be generated using peripheral pixels. The prediction block of example 2100 is generated using upper peripheral pixels 2104 (visualized using upper pixels 2104), left peripheral pixels 2106 (visualized using left pixels 2108), or a combination thereof. As described above, upper left peripheral pixel 2110 can be the origin of the coordinate system.
[0202] Prediction block 2120 (visualized as prediction block 2122) shows that only the primary peripheral pixels (ie, the upper peripheral pixels) are used. Therefore, even if the left peripheral pixels are available, prediction block 2120 is as described with respect to Figure 10 and Figure 11 Generated as described.
[0203] The prediction block 2130 (visualized as prediction block 2132) shows that only non-primary peripheral pixels (i.e., left peripheral pixels) are used. Therefore, the prediction block 2130 can be generated by using primary peripheral pixels (i.e., upper peripheral pixels) to obtain the position in the upper row, as described with respect to Figure 20A 2010. Get the x-intercept, as described in Figure 20A as described in block 2012; and Figure 11 The predicted pixel value is calculated using the x-intercept in a similar manner as described for prediction block 1102 .
[0204] Prediction block 2140 (visualized as prediction block 2142) shows a prediction block generated using a weighted sum, as described above with respect to Figure 19 Prediction block 2150 (visualized as prediction block 2152) shows a prediction block generated using the nearest distance, as described above with respect to Figure 19 described.
[0205] As described above, intra-frame prediction modes according to the present disclosure can be defined with respect to a focal point. The focal point can be defined as the point from which all predicted points radiate. In other words, each point in the prediction block can be considered to be connected to the focal point. Similarly, the focal point can be considered to be a point within the distance where parallel lines in a perspective image intersect.
[0206] Figure 14is a flowchart of a technique 1400 for coding a current block using intra prediction modes according to embodiments of the present disclosure. The current block is coded using a focus point. The intra prediction modes use pixels that are peripheral to the current block. The current block can be coded as described with respect to Figure 12 Technique 1200. Technique 1400 can be implemented in an encoder, such as encoder 400. Figure 4 Technique 1400 can be implemented in a decoder, such as decoder 500. Figure 5
[0207] Technique 1400 can be implemented, for example, as a software program that can be executed by a computing device, such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that can be stored in a memory, such as memory 204 or secondary storage 214, that can be executed by a processor, such as CPU 202, to cause the computing device to perform technique 1400. In at least some embodiments, performing technique 1400 can be performed, in whole or in part, by intra / inter prediction stage 402 of encoder 400. Figure 4 In other embodiments, technique 1400 can be performed, in whole or in part, by intra / inter prediction stage 508 of decoder 500. Figure 5
[0208] Technique 1400 can be implemented using specialized hardware or firmware. Some computing devices can have multiple memories, multiple processors, or both. The steps or operations of technique 1400 can be distributed using different processors, memories, or both. Use of the term "processor" or "memory" in the singular does not exclude a computing device that has multiple processors or multiple memories, nor does it require that a single processor or memory be used to perform some or all of the steps.
[0209] Technique 1400 can be best understood with reference to Figure 15 and Figure 16
[0210] Figure 15 is an example 1500 that illustrates a focus point according to embodiments of the present disclosure. Example 1500 shows a current block 1502 that is to be predicted. That is, a prediction block is to be generated for current block 1502. Current block 1502 has a width 1504 and a height 1506. As such, the size of current block 1502 is W x H. In example 1500, the current block is shown as 8 x 4. However, the present disclosure is not so limited. The current block can have any size. For ease of illustration, pixel positions are shown as squares in example 1500. The actual values of the pixels can be more properly considered to be the value at the center (i.e., middle) of the square.
[0211] The current block 1502 will be predicted using peripheral pixels. The peripheral pixels can be or can include the above peripheral pixels 1508. The peripheral pixels can be or can include the left peripheral pixels 1512. The above peripheral pixels 1508 can include a number of pixels equal to the width 1504 (W). The left peripheral pixels 1512 can include a number of pixels equal to the height 1506 (H). For convenience or reference, the top-left peripheral pixel 1509 can be considered as part of the left peripheral pixels 1512, part of the above peripheral pixels 1508, or part of both the left peripheral pixels 1512 and the above peripheral pixels 1508.
[0212] The above peripheral pixels 1508 can include overhang above peripheral pixels 1510. The number of overhang above peripheral pixels 1508 is denoted as W0. In the example 1500, W0 is shown as equal to 8 pixels. However, the present disclosure is not limited thereto, and the overhang above peripheral pixels can include any number of pixels.
[0213] The left peripheral pixels 1512 can include overhang left peripheral pixels 1514. The number of overhang left peripheral pixels 1514 is denoted as H0. In the example 1500, H0 is shown as equal to 2 pixels. However, the present disclosure is not limited thereto, and the overhang left peripheral pixels 1514 can include any number of pixels.
[0214] While the left peripheral pixels 1512 are discrete pixels, the left peripheral pixels 1512 can be considered as pixel values at integer positions of a continuous line of peripheral pixels. Thus, the left peripheral pixels 1512 (e.g., first peripheral pixels) form a first line of peripheral pixels that constitutes an x-axis 1530. While the above peripheral pixels 1508 are discrete pixels, the above peripheral pixels 1508 can be considered as pixel values at integer positions of a continuous line of peripheral pixels. Thus, the above peripheral pixels 1508 (e.g., second peripheral pixels) form a second line of peripheral pixels that constitutes a y-axis 1532.
[0215] Three illustrative pixels of the current block 1502 are shown: a pixel 1518 at the top-right corner of the current block 1502, a pixel 1522 at the bottom-left corner of the current block 1502, and a pixel 1526. Each of the pixels 1518, 1522, 1526 can be considered as having coordinates (i, j), where the center can be at the top-left corner of the current block. Thus, the pixel 1518 is at coordinates (7, 0), the pixel 1522 is at coordinates (0, 3), and the pixel 1526 is at coordinates (5, 3).
[0216] The focus 1516 is shown as being outside and at a distance from the current block.The focus 1516 is at coordinates (a, b) in a coordinate system centered at the intersection between the x-axis 1530 and the y-axis 1532.
[0217] As described above, each pixel of current block 1502 emanates from focal point 1516. Thus, line 1520 connects pixel 1518 and focal point 1516, line 1524 connects pixel 1522 and focal point 1516, and line 1528 connects pixel 1526 and focal point 1516.
[0218] The x-intercept x0 of line 1520 (i.e., where line 1520 intersects x-axis 1530) is point 1534, and the y-intercept y0 of line 1520 (i.e., where line 1520 intersects y-axis 1532) is point 1535. The x-intercept x0 of line 1524 (i.e., where line 1524 intersects x-axis 1530) is point 1536, and the y-intercept y0 of line 1524 (i.e., where line 1524 intersects y-axis 1532) is point 1537. The x-intercept x0 of line 1528 (i.e., where line 1528 intersects x-axis 1530) is point 1538, and the y-intercept y0 of line 1528 (i.e., where line 1528 intersects y-axis 1532) is point 1539.
[0219] Point 1534 has a negative value (i.e., a negative x-intercept). Points 1536 and 1538 have positive values. Point 1535 has a positive value (i.e., a positive y-intercept). Points 1537 and 1539 are negative values.
[0220] It is well known that given two points on a line with coordinates (a,b) and (i,j), the x- and y-intercepts can be calculated using equation (1)
[0221]
[0222] Figure 16 is an example showing the x-intercept and the y-intercept according to an embodiment of the present disclosure. Figure 16 The examples of φ show the positive and negative x-intercepts and y-intercepts of the predicted pixel 1604 (or equivalently, the current pixel) at position (i, j) of the current block 1612 given different positions of focus.
[0223] Example 1600 shows a focus 1602 and a line 1606 that passes through (e.g., connects) the prediction pixel 1604 to the focus 1602. The x-intercept 1608 is negative. The y-intercept 1610 is positive. Example 1620 shows a focus 1622 and a line 1624 that passes through (e.g., connects) the prediction pixel 1604 to the focus 1622. The x-intercept 1626 is positive. The y-intercept 1628 is negative. Example 1640 shows a focus 1642 and a line 1644 that passes through (e.g., connects) the prediction pixel 1604 to the focus 1642. The x-intercept 1646 is positive. The y-intercept 1648 is positive.
[0224] Again, returning to Figure 14 At 1402, the technique 1400 obtains a focus. As described with respect to Figure 15 , the focus has coordinates (a, b) in a coordinate system.
[0225] When implemented by a decoder, obtaining the focus can include decoding an intra prediction mode from a compressed bitstream. The compressed bitstream can be the compressed bitstream 420. The intra prediction mode can indicate the focus. Figure 5
[0226] In an example, each of the available intra prediction modes can be associated with an index (e.g., a value). Decoding the index from the compressed bitstream instructs the decoder to perform intra prediction for the current block according to the intra prediction mode (i.e., the semantics of the intra prediction mode). In an example, the intra prediction mode can indicate the coordinates of the focus. For example, an intra prediction mode value of 45 can indicate that the focus is at coordinates (-1000, -1000), an intra prediction mode value of 46 can indicate that the focus is at coordinates (-1000, -850), and so on. Thus, for example, if 64 foci are possible, then 64 intra prediction modes are possible, each indicating the location of a focus. In an example, hundreds of foci (and equivalently, intra prediction modes) are available. While the location of the focus is given herein in Cartesian coordinates, the focus coordinates can be given in polar coordinates. The angle of the polar coordinates can be with respect to the x-axis, such as the x-axis 1530 of FIG. 15. Figure 15
[0227] In another example, obtaining the focus can include decoding the coordinates of the focus from the compressed bitstream. For example, the compressed bitstream can include an intra prediction mode that indicates an intra prediction using a focus followed by the coordinates of the focus.
[0228] When implemented by an encoder, obtaining the focus point can include selecting the focus point from a plurality of candidate focus points. That is, the encoder selects an optimal focus point for encoding the current block. The optimal focus point can be a focus point that results in an optimal encoding of the current block. In an example, the plurality of candidate focus points can be divided into candidate focus point groups. Each candidate focus point group can be arranged on a circumference of a respective circle. In an example, each candidate focus point group can include 16 candidate focus points.
[0229] Figure 17 An example 1700 of focus point groups is shown in accordance with an embodiment of the present disclosure. As described above, there can be hundreds of focus candidates, which can be located anywhere in the space outside the current block. The focus candidates are a subset of all possible focus points in the space outside the current block. The space outside the current block can be centered (i.e., have an origin) at the top-left peripheral pixel of the top-left peripheral pixel 1509. The center of the space can be any other point. In an example, the center can be the center point of the current block. Figure 15
[0230] To limit the search space, only a subset of all possible focus points can be considered as candidate focus points. There can be many ways to narrow down the search space to candidate focus points. In an example, the candidate focus points can be grouped into groups. Each candidate focus point group can be arranged on a circumference of a circle. In an example, three circles can be considered. The three circles can be the frame 1702, the frame 1704, and the frame 1706. The focus points are shown as black circles on each frame (such as focus points 1708-1712). However, any number of circles can be available. Each circle (or frame) generally corresponds to the slope (e.g., convergence rate) of the lines that connect the predicted pixels to the focus points on the circumference of the circle. Note that the example 1700 is merely illustrative and not drawn to scale.
[0231] The frame 1702 can correspond to far-away focus points. As such, given a focus point on the frame 1702, the slope of the lines from each predicted pixel location to the focus point can be generally the same. The frame 1702 can have a radius in the range of 1000 pixels. The farther away the focus points, the more the intra prediction using far-away focus points resembles (e.g., approximates) directional intra prediction.
[0232] The frame 1706 can correspond to nearby focus points. As such, given a focus point on the frame 1702, the lines from the focus point to each predicted pixel location can appear to diverge. Thus, the slopes of the lines can be very different. The frame 1706 can have a radius in the range of tens of pixels. For example, the radius can be 20, 30, or some other such number of pixels.
[0233] The frame 1704 can correspond to a circle with a medium-sized radius. The radius of the frame 1704 can be in the range of hundreds of pixels.
[0234] As mentioned above, although the circle (frame) can have an impractical number of focal points, only a sampling of the focal points are used as candidate focal points. The candidate focal points of each group (i.e., on each frame) can be equally spaced. For example, assuming N (e.g., 8, 16, etc.) candidate focal points are included in each group, the N focal points can be separated by 360 / N (e.g., 45, 22.5, etc.) degrees.
[0235] In an example, acquiring the focal points at 1402 in the encoder can include testing each candidate focal point to identify the best focal point. In another example, the outermost optimal focal point of the outermost frame (frame 1702) can be identified by performing intra prediction using each focal point of the outermost frame. The focal points corresponding (e.g., at the same angle) to the outermost optimal focal point can then be tried to determine whether any of them produces a more optimal prediction block. Other heuristic methods can also be used. For example, a binary search can be used.
[0236] Returning to Figure 14 At 1404, the technique 1400 can generate the prediction block using the first peripheral pixel and the second peripheral pixel. The first peripheral pixel can be a left-side peripheral pixel, such as the left-side peripheral pixel 1512 (including the top-left peripheral pixel 1509). The second peripheral pixel can be an above peripheral pixel 1508.
[0237] As mentioned above, the first peripheral pixels form a first peripheral pixel line constituting an x-axis, such as the x-axis 1530 of Figure 15 ; the second peripheral pixels form a second peripheral pixel line constituting a y-axis, such as the y-axis 1532 of Figure 15 ; and the first peripheral pixel line and the second peripheral pixel line form a coordinate system having an origin. Generating the prediction block can include performing 1404_4 to 1404_6 for each position (i.e., each pixel) of the prediction block. Each pixel of the prediction block is at position (i,j). If the block size is MxN, 1404_4 to 1404_6 are performed M*N times.
[0238] At 1404_2, the technique 1400 determines whether there are any more prediction block positions for which a pixel value has not been determined (e.g., calculated). If there are more pixel positions, the technique 1400 proceeds to 1404_4; otherwise, the technique 1400 proceeds to 1406.
[0239] At 1404_4, the technique 1400 can determine (e.g., calculate, identify, etc.) at least one of an x-intercept or a y-intercept of the predicted pixel at (i,j).
[0240] The x-intercept is a first point (e.g., x-intercept 1608, x-intercept 1626, x-intercept 1646) at which a line (e.g., line 1606, line 1624, line 1644) formed by a point centered at each location of the prediction block (e.g., prediction pixel 1604) and a focal point (e.g., focal point 1602, focal point 1622, focal point 1642) intersects a first peripheral pixel line (e.g., the x-axis).
[0241] The y-intercept is a second point (e.g., y-intercept 1609, y-intercept 1627, y-intercept 1647) at which the line (e.g., line 1606, line 1624, line 1644) formed by the point centered at each location of the prediction block (e.g., prediction pixel 1604) and the focal point (e.g., focal point 1602, focal point 1622, focal point 1642) intersects a second peripheral pixel line (e.g., the y-axis).
[0242] The x- and / or y-intercepts can be calculated using equation (1). However, in some cases, a line through a prediction pixel and a focal point can not be considered to intercept one of the axes. For example, a line nearly parallel to the x-axis can not be considered to intercept the x-axis; and a line nearly parallel to the y-axis can not be considered to intercept the y-axis. A line can not be considered to intercept the x-axis when b = j + e, and a line can not be considered to intercept the y-axis when a = i + e, where e is a small threshold value close to zero. As such, the x-intercept can be identified as i, and the y-intercept can be identified as j, without having to use equation (1).
[0243] At 1404_6, the technique 1400 can determine the prediction pixel value for each location (i.e., (i,j)) of the prediction block using at least one of the x-intercept or the y-intercept. From 1404_6, the technique 1400 returns to 1404_2.
[0244] In an example, determining the prediction pixel value can include, in a case where one of the at least one of the x-intercept or the y-intercept is negative, determining the prediction pixel value for each location using the other of the at least one of the x-intercept or the y-intercept. For example, with respect to example 1600, Figure 16 because the x-intercept 1608 is negative, the prediction pixel value for location (i,j) of the prediction block is calculated using only the y-intercept 1610, which is positive. For example, with respect to example 1620, Figure 16 because the y-intercept 1628 is negative, the prediction pixel value for location (i,j) of the prediction block is calculated using only the x-intercept 1626, which is positive.
[0245] In an example, determining the prediction pixel value can include, in a case where the x-intercept is positive and the y-intercept is positive, determining the prediction pixel value for each location as a weighted combination of a first pixel value at the x-intercept and a second pixel value at the y-intercept.
[0246] In an example, determining the predicted pixel value can include setting the pixel value at each position of the prediction block to the value of the first peripheral pixel at position i in the first peripheral pixel line if i is equal to a (i.e., i is very close to a). That is, if i≈a, setting p(i,j)=L[i]. That is, if the line is almost parallel to the y-axis, setting the predicted pixel value p(i,j) to the horizontally corresponding left peripheral pixel value L[i].
[0247] Similarly, in the example, determining the predicted pixel value can include setting the pixel value at each position of the prediction block to the value of the second peripheral pixel located at position j in the second peripheral pixel line if j is equal to b (i.e., j is very close to b). That is, if j ≈ b, then setting p(i, j) = T[i]. That is, if the line is almost parallel to the x-axis, then setting the predicted pixel value p(i, j) to the vertically corresponding upper (i.e., top) peripheral pixel value T[j].
[0248] In an example, determining the predicted pixel value can include setting the pixel value at each position of the prediction block to the pixel value located at the intersection of the first peripheral pixel line and the second peripheral pixel line when the x-intercept is zero and the y-intercept is zero. That is, if the x-intercept and the y-intercept are zero, the predicted pixel value can be set to the upper left peripheral pixel value.
[0249] The pseudo code of Table I shows an example of setting a predicted pixel value p(i,j) at position (i,j) of a prediction block using the focus at (a,b), where i=0,...,width-1 and j=0,...,height-1.
[0250] As described above, for a given pixel (i, j), a line is drawn connecting the focal point at (a, b) and (i, j). The x-intercept and y-intercept are calculated. Depending on these intercepts, the predicted pixel value p(i, j) is obtained by interpolating or extrapolating the intercept value from the top or left boundary pixels.
[0251] In particular, let L[k] denote the left boundary pixel at positive integer position k (e.g., Figure 15 , where k = 0, 1, ..., H + H0; and T[k] represents the top boundary pixel array of the positive integer position k (e.g., the upper peripheral pixels 1508 plus the upper left peripheral pixels 1509). Figure 15 The upper left peripheral pixels 1509), where k = 0, 1, ..., W + W0. Note that T [0] = L [0]. Also make f L (z) and f T(z) denotes an interpolation function at a high-precision (real- valued) point z obtained from the boundary pixels L[k], T[k] by a suitable interpolation (i.e. interpolation function). Note that at integer positions, for z = k = 0, 1,..., H+H0, f L (z) = L[k] for z = k = 0, 1,..., W+W0, f T (z) = T[k].
[0252]
[0253] In line 1 of Table I, if the focus and the predicted pixel at (i,j) are on the same horizontal line, then in line 2, the predicted pixel p(i,j) is set to the value L[i] of the left peripheral pixel located on the same horizontal line. More particularly, the focus and the predicted pixel can not be perfectly horizontally aligned. Thus, (i==a) can mean that the line connecting the focus and the predicted pixel passes through the square centered at L[i].
[0254] In line 3, if the focus and the predicted pixel are on the same vertical line, then in line 4, the predicted pixel p(i,j) is set to the value T[j] of the upper peripheral pixel located on the same vertical line. More particularly, the focus and the predicted pixel can not be perfectly vertically aligned. Thus, (j==b) can mean that the line connecting the focus and the predicted pixel passes through the square centered at T[j].
[0255] In lines 6-7, the x-intercept (x0) and the y-intercept (y0) are computed according to equation (1). In line 8, if the x-intercept (x0) and the y-intercept (y0) are at the origin, then in line 9, the predicted pixel p(i,j) is set to the top-left peripheral pixel L[0]. More particularly, the x-intercept (x0) and / or the y-intercept (y0) can not be exactly zero. Thus, x0==0 && y0==0 can mean that the x-intercept and the y-intercept are within the square (i.e. pixel) centered at the origin.
[0256] In line 10, if the x-intercept (x0) is positive but the y-intercept (y0) is negative, then in line 11, the predicted pixel p(i,j) uses the interpolation function f L (z) only from the x-intercept (x0). In line 12, if the x-intercept (x0) is negative and the y-intercept (y0) is positive, then in line 13, the predicted pixel p(i,j) uses the interpolation function f T (z) only from the y-intercept (y0).
[0257] In line 14, if the x-intercept (x0) and the y-intercept (y0) are both positive, then the predicted pixel p(i,j) is computed as the interpolation of the x-intercept (x0) (i.e. f L (x0)) and the interpolation of the y-intercept (y0) (i.e. f TThe weight of which depends on which of the x-intercept (x0) and y-intercept (y0) is farther from the prediction pixel p(i,j). If the x-intercept (x0) is farther (i.e., line 15), then in line 16, the y-intercept (y0) has a greater weight. On the other hand, if the y-intercept (y0) is farther (i.e., line 17), then in line 18, the x-intercept (x0) has a greater weight. Lines 20-21 are for completeness purposes and are intended to cover the case where both the x-intercept (x0) and y-intercept (y0) are negative, which is an impossible case.
[0258] In some cases, at least some of the top or left peripheral pixels can not be available. For example, the current block can be a block located at the top edge or left edge of the image. In such cases, the unavailable peripheral pixels can be considered to have a zero value.
[0259] The interpolation function f L and f T can be any interpolation function. The interpolation functions can be the same interpolation function or different interpolation functions. The interpolation functions can be as described above with respect to Figure 9 For example, the interpolation functions can be finite impulse response (FIR) filters. For example, as described above, the interpolation filters can be bilinear interpolation. That is, given an x-intercept (or y-intercept) value, the closest integer pixel can be determined and a weighted sum of the closest integer pixels can be used in the bilinear interpolation. As used herein, interpolation includes both interpolation and extrapolation.
[0260] Returning to Figure 14 At 1406, the technique 1400 encodes a residual block corresponding to a difference between the current block and the prediction block. When implemented by an encoder, the technique 1400 computes the residual block as the difference between the current block and the prediction block and encodes the residual block in a compressed bitstream. When implemented by a decoder, the technique 1400 encodes the residual block by decoding the residual block from the compressed bitstream. The decoder can then add the residual block to the prediction block to reconstruct the current block.
[0261] Another aspect of the disclosed implementations is a technique for encoding a current block. The technique includes obtaining a prediction block of prediction pixels of the current block using peripheral pixels. Each prediction pixel is located at a respective position (i,j) within the prediction block. Obtaining the prediction block can include obtaining a focal point having coordinates (a,b) in a coordinate system; and for each position of the prediction block, obtaining a line indicating a respective prediction angle and using the line to determine a pixel value for each position. As described above, the line connects the focal point to each position. As described with respect to Figures 15-17As described, the focus can be outside the current block, and the focus can be but can not be one of the peripheral pixels. As described above, each prediction pixel of the prediction block can have a prediction angle that is different from the prediction angle of each other prediction pixel of the prediction angle. While more than one prediction pixel can have the same prediction angle depending on the location of the focus, the intra prediction mode according to embodiments of the disclosure is such that not all prediction pixels can have the same prediction angle.
[0262] Encoding the current block can include encoding in the compressed bitstream the intra prediction mode that indicates the focus. As described above, the value associated with the intra prediction mode can indicate the location of the focus.
[0263] As described above, the peripheral pixels can include a left peripheral pixel and a top peripheral pixel. Determining the pixel value for each location using the line can include determining an x-intercept of the line; determining a y-intercept of the line; and determining the pixel value using the x-intercept and the y-intercept. The x-intercept is a first point at which the line intersects a left-side axis that includes the left peripheral pixel of the current block. The y-intercept is a second point at which the line intersects a top-side axis that includes the top peripheral pixel of the current block.
[0264] Obtaining the focus, as described above and with respect to Figure 17 As described, the obtaining the focus can include selecting the focus from a plurality of candidate focuses. The plurality of candidate focuses can be partitioned into candidate focus groups. Each candidate focus group can be arranged on a circumference of a respective circle.
[0265] Another aspect of the disclosed embodiments is a technique for decoding a current block. The technique includes decoding a focus from a compressed bitstream; obtaining a prediction block for prediction pixels of the current block; and reconstructing the current block using the prediction block. The obtaining the prediction block includes, for each location of the prediction block, obtaining a line representing a respective prediction angle, wherein the line connects the focus to the each location; and determining a pixel value for the each location using the line.
[0266] Determining the pixel value for each location using the line can include determining an x-intercept of the line; determining a y-intercept of the line; and determining the pixel value using the x-intercept and the y-intercept. The x-intercept is a first point at which the line intersects a left-side axis that includes a left peripheral pixel of the current block. The y-intercept is a second point at which the line intersects a top-side axis that includes a top peripheral pixel of the current block.
[0267] Determining the pixel value using the x-intercept and the y-intercept can include determining the pixel value as a weighted combination of a first pixel value at the x-intercept and a second pixel value at the y-intercept. Determining the pixel value using the x-intercept and the y-intercept can include, in a case where one of the x-intercept or the y-intercept is negative, using the other of the x-intercept or the y-intercept to determine the pixel value.
[0268] Decoding a focus point from a compressed bitstream can include decoding, from the compressed bitstream, an intra prediction mode that indicates the focus point.
[0269] For ease of explanation, each of techniques 900, 1200, 1300, 1400, 1800, and 1900 is depicted and described as a series of blocks, steps, or operations. However, the blocks, steps, or operations in accordance with this disclosure can occur in various orders and / or concurrently. Additionally, other steps or operations not presented and described herein can be used. Further, not all illustrated steps or operations can be required to implement a technique in accordance with the subject technology.
[0270] The above-described aspects of encoding and decoding illustrate some encoding and decoding techniques. However, it should be understood that encoding and decoding, as those terms are used in the claims, can represent compression, decompression, transformation, or any other data processing or change.
[0271] The word “example” or “implementation” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “example” or “implementation” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, use of the words “example” or “implementation” is intended to present concepts in a concrete fashion. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or as is clear from the context, the statement “X includes A or B” is intended to mean any of the natural inclusive permutations. That is, if X includes A; X includes B; or X includes both A and B, then the foregoing statement “X includes A or B” is satisfied by any of the foregoing instances. In addition, the articles “a” and “an” as used in this application and the appended claims should generally be construed to mean “one or more” unless specified otherwise or clear from context to be directed to a singular form. Moreover, use of the term “an implementation” or “one implementation” throughout is not intended to mean the same implementation or implementation unless so described. Rather, use of the term “an implementation” or “one implementation” is intended to mean “at least one implementation.” Further, the use of the term “implementation” or “one implementation” throughout is not intended to mean the same implementation or implementation unless so described.
[0272] The transmitting station 102 and / or the receiving station 106 (and algorithms, methods, instructions, etc. stored thereon and / or executed thereby, including the algorithms, methods, instructions, etc. executed by the encoder 400 and the decoder 500) can be implemented in hardware, software, or any combination thereof. The hardware can include for example a computer, an intellectual property (IP) core, an application-specific integrated circuit (ASIC), a programmable logic array, optical processors, programmable logic controllers, microcode, microcontrollers, servers, microprocessors, digital signal processors or any other suitable circuit. In the claims, the term "processor" should be understood as encompassing any of the foregoing hardware, whether alone or in combination. The terms "signal" and "data" can be used interchangeably. Moreover, portions of the transmitting station 102 and the receiving station 106 need not be implemented in the same manner.
[0273] Furthermore, in one aspect, for example, the transmitting station 102 or the receiving station 106 can be implemented using a computer or a processor with a computer program that, when executed, carries out any of the respective methods, algorithms and / or instructions described herein. In addition, or alternatively, for example, a special purpose computer / processor can be used that contains other hardware for carrying out any of the methods, algorithms or instructions described herein.
[0274] The transmitting station 102 and the receiving station 106 can be implemented, for example, on computers in a video conferencing system. Alternatively, the transmitting station 102 can be implemented on a server and the receiving station 106 can be implemented on a device separate from the server, such as a handheld communication device. In this example, the transmitting station 102 can encode content into an encoded video signal using the encoder 400 and transmit the encoded video signal to the communication device. In turn, the communication device can then decode the encoded video signal using the decoder 500. Alternatively, the communication device can decode content locally stored on the communication device, e.g., content that is not transmitted by the transmitting station 102. Other transmitting station 102 and receiving station 106 implementations are available. For example, the receiving station 106 can be a generally stationary personal computer, rather than a portable communication device and / or a device that includes the encoder 400 can also include the decoder 500.
[0275] Furthermore, all or portions of the embodiments of the present disclosure can take the form of a computer program product accessible from for example a tangible computer-usable or computer-readable medium. A computer-usable or computer-readable medium can be for example any device that can tangibly contain, store, communicate, or transport a program for use by or in connection with any processor. The medium can be for example an electronic, magnetic, optical, electromagnetic, or a semiconductor device. Other suitable mediums are available.
[0276] The above embodiments, implementations and aspects have been described to enable an easy understanding of the present disclosure and are not limiting thereof. Instead, the present disclosure is intended to encompass various modifications and equivalent arrangements included within the scope of the appended claims, which scope should be accorded a broadest interpretation so as to encompass all such modifications and equivalent structures legally permitted.
Claims
1. A method for coding a current block using an intra prediction mode, comprising: selecting a focus, the focus having coordinates (a, b) in a coordinate system and being selected from a plurality of candidate focuses, the plurality of candidate focuses being divided into candidate focus groups, each candidate focus group being arranged on a circumference of a corresponding circle centered at a point within the current block; Generate a prediction block of the current block using the first peripheral pixels and the second peripheral pixels, wherein the first peripheral pixels form a first peripheral pixel line constituting an x-axis, wherein the second peripheral pixels form a second peripheral pixel line constituting the y-axis, wherein the first peripheral pixel line and the second peripheral pixel line form the coordinate system having an origin, wherein the focus is outside the prediction block, is not any of the first peripheral pixel lines, and is not any of the second peripheral pixels, and The generating of the prediction block includes: for each position of the prediction block at position (i, j) in at least some of the positions of the prediction block, determining at least one of an x-intercept or a y-intercept, wherein the x-intercept is a first point at which a line formed by a point centered at each position of the prediction block and the focus intersects the first peripheral pixel line, and wherein the y-intercept is a second point at which a line formed by the point centered at each position of the prediction block and the focus intersects the second peripheral pixel line; and determining a predicted pixel value for each position of the prediction block using the at least one of the x-intercept or the y-intercept; and A residual block corresponding to a difference between the current block and the predicted block is coded.
2. The method according to claim 1, wherein Determining the predicted pixel value for each position of the prediction block using the at least one of the x-intercept or the y-intercept includes: The predicted pixel value at each position is determined using the other of the at least one of the x-intercept or the y-intercept under the condition that the other of the at least one of the x-intercept or the y-intercept is a negative value.
3. The method according to claim 2, wherein: Determining the predicted pixel value for each position of the prediction block using the at least one of the x-intercept or the y-intercept further comprises: Under the condition that the x-intercept is positive and the y-intercept is positive, the predicted pixel value at each position is determined as a weighted combination of a first pixel value at the x-intercept and a second pixel value at the y-intercept.
4. The method according to claim 1, wherein Determining the predicted pixel value for each position of the prediction block using the at least one of the x-intercept or the y-intercept includes: On the condition that i is equal to a, a pixel value at each position of the prediction block is set to a value of a first peripheral pixel located at position i of the first peripheral pixel line among the first peripheral pixels.
5. The method according to claim 1, wherein Determining the predicted pixel value for each position of the prediction block using the at least one of the x-intercept or the y-intercept includes: On the condition that j is equal to b, the pixel value at each position of the prediction block is set to the value of the second peripheral pixel located at position j of the second peripheral pixel line among the second peripheral pixels.
6. The method according to claim 1, wherein Determining the predicted pixel value for each position of the prediction block using the at least one of the x-intercept or the y-intercept includes: Under the condition that the x-intercept is zero and the y-intercept is zero, a pixel value at each position of the prediction block is set to a pixel value at an intersection of the first peripheral pixel line and the second peripheral pixel line.
7. The method according to claim 1, wherein Selecting the focus includes: The intra prediction mode is decoded from a compressed bitstream, wherein the intra prediction mode indicates the focus.
8. The method according to claim 1, wherein Each of the candidate focus groups includes 16 candidate focuses.
9. An apparatus for decoding a current block, comprising: Memory; as well as a processor configured to execute instructions stored in the memory for: decoding a focus from a compressed bitstream, wherein the focus is one of a plurality of candidate focuses, the plurality of candidate focuses being divided into candidate focus groups, each candidate focus group being arranged on a circumference of a corresponding circle centered at a point within the current block; Obtaining a prediction block of predicted pixels of the current block, wherein each predicted pixel is located at a corresponding position within the prediction block, wherein obtaining the prediction block comprises: For each of at least some positions of the prediction block, instructions are executed to: obtaining a line indicating a corresponding prediction angle, the line connecting the focus to each of the positions; and determining a pixel value at each of said locations using said line; and The current block is reconstructed using the prediction block.
10. The device according to claim 9, wherein Using the line to determine the pixel value at each location includes: determining an x-intercept of the line, wherein the x-intercept is a first point at which the line intersects a left axis including left peripheral pixels of the current block; determining a y-intercept of the line, wherein the y-intercept is a second point at which the line intersects a top axis including top peripheral pixels of the current block; and The pixel value is determined using the x-intercept and the y-intercept.
11. The device according to claim 10, wherein Determining the pixel value using the x-intercept and the y-intercept includes: The pixel value is determined as a weighted combination of a first pixel value at the x-intercept and a second pixel value at the y-intercept.
12. The device according to claim 10, wherein Determining the pixel value using the x-intercept and the y-intercept includes: Under the condition that one of the x-intercept or the y-intercept is a negative value, the pixel value is determined using the other of the x-intercept or the y-intercept.
13. The device according to claim 9, wherein Decoding the focus from the compressed bitstream comprises: An intra-prediction mode is decoded from the compressed bitstream, wherein the intra-prediction mode indicates the focus.
14. A method for encoding a current block, comprising: Obtaining a prediction block of predicted pixels of the current block using peripheral pixels, wherein each predicted pixel is located at a corresponding position within the prediction block, and wherein obtaining the prediction block comprises: selecting a focus, the focus having coordinates (a, b) in a coordinate system, the focus being outside the current block and not being any of the peripheral pixels, and the focus being selected from a plurality of candidate focuses, the plurality of candidate focuses being divided into candidate focus groups, each candidate focus group being arranged on a circumference of a corresponding circle centered at a point within the current block; For each of at least some positions of the prediction block, performing steps comprising: obtaining a line indicating a corresponding prediction angle, the line connecting the focus to each of the positions; and determining a pixel value at each of said locations using said line; and The focus is encoded in a compressed bitstream.
15. The method according to claim 14, wherein Encoding the focus in the compressed bitstream comprises: An intra-prediction mode indicating the focus is encoded in the compressed bitstream.
16. The method according to claim 14, wherein The peripheral pixels include left peripheral pixels and top peripheral pixels, and wherein using the line to determine the pixel value of each position includes: determining an x-intercept of the line, wherein the x-intercept is a first point at which the line intersects a left axis including the left peripheral pixels of the current block; determining a y-intercept of the line, wherein the y-intercept is a second point at which the line intersects a top axis including the top peripheral pixels of the current block; and The pixel value is determined using the x-intercept and the y-intercept.
Citation Information
Patent Citations
Method and device for encoding video to improve intra prediction processing speed, and method and device for decoding video
US20140328404A1
Intra-prediction for video coding using perspective information
WO2018231087A1