Method for deblocking HDR content and non-fugitive computer readable storage medium
Adaptive deblocking filtering based on pixel intensity addresses the uniform filtering issue in video coding, enhancing video quality by reducing artifacts in HDR content.
Patent Information
- Application Number
- JP2025149992
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-03-29
- Filing Date
- 2025-09-10
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2039-03-29
AI Technical Summary
Current video coding schemes apply deblocking and filtering uniformly across all content without considering image intensity, leading to suboptimal filtering results due to varying content intensity needs.
Implement deblocking filtering based on pixel intensity of coding units, adjusting filtering strength and criteria based on intensity information to improve filtering effectiveness.
Enhances deblocking performance by tailoring filtering to content intensity, reducing visual artifacts and improving video quality, particularly in high dynamic range (HDR) content.
Smart Images

Figure 2025175096000001_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to the field of video coding, and in particular to systems and methods for efficiently and effectively deblocking and filtering HDR content. [Background technology]
[0002] This application claims priority under U.S.C. §119(e) to earlier filed U.S. Provisional Application No. 62 / 650,252, filed March 29, 2018, the entirety of which is incorporated herein by reference.
[0003] Technical improvements in evolving video coding standards have shown a trend toward improved coding efficiency, enabling higher bit rates, higher resolutions, and better video quality. The Joint Video Exploration Team (JVET) has developed a new video coding scheme and is currently developing an even newer video coding scheme called Versatile Video Coding (VVC), where the complete content of the 7th edition of VVC in Draft 2 of the standard entitled Versatile Video Coding (Draft 2) by JVET, published on October 1, 2018, is incorporated herein by reference. Similar to other video coding schemes such as High Efficiency Video Coding (HEVC), both JVET and VVC are block-based hybrid spatial-temporal predictive coding schemes. However, compared to HEVC, JVET and VVC include many changes to the bitstream structure, syntax, constraints, and mapping for generating decoded pictures. JVET has been implemented in the Joint Exploration Model (JEM) encoder and decoder, but VVC is not expected to be implemented until early 2020.
[0004] Current video coding schemes implement deblocking and filtering without taking image intensity into account, resulting in filtering of content in a uniform manner across all content. However, data reveals that content intensity can affect the degree or level of filtering desired or necessary to reduce display issues. Therefore, what is needed are systems and methods for deblocking that are based at least in part on pixel intensity of coding units. [Brief explanation of the drawings]
[0005] [Figure 1] It shows the allocation of a frame into multiple coding tree units (CTUs). [Figure 2a] 1 illustrates an exemplary partitioning of a CTU into coding units (CUs). [Figure 2b] 1 illustrates an exemplary partitioning of a CTU into coding units (CUs). [Figure 2c] 1 illustrates an exemplary partitioning of a CTU into coding units (CUs). [Figure 3] A quadtree and binary tree (QTBT) representation is shown for the CU partitions in FIG. [Figure 4] 1 shows a simplified block diagram for CU coding in a JVET or VVC encoder. [Figure 5] 1 shows possible intra prediction modes for the luma component in JVET of VVC. [Figure 6] 1 shows a simplified block diagram for CU coding in JVET of a VVC decoder. [Figure 7] A block diagram of an HDR encoder / decoder system is shown. [Figure 8] 1 illustrates one embodiment of a normalized PQ vs. normalized intensity curve. [Figure 9] 1 illustrates one embodiment of a JND versus normalized intensity curve. [Figure 10]1 illustrates an embodiment of a block diagram of an encoding system that is based at least in part on intensity. [Figure 11] 1 illustrates one embodiment of a block diagram of a decoding system that is based at least in part on intensity. [Figure 12a] 10 and 11 show a series of exemplary β vs. QP and tc vs. QP curves that graphically represent the system described and shown in FIGS. [Figure 12b] 10 and 11 show a series of exemplary β vs. QP and tc vs. QP curves that graphically represent the system described and shown in FIGS. [Figure 12c] 10 and 11 show a series of exemplary β vs. QP and tc vs. QP curves that graphically represent the system described and shown in FIGS. [Figure 13] 1 illustrates one embodiment of a computer system adapted and configured to provide variable template sizes for template matching. [Figure 14] 1 illustrates an embodiment of a video encoder / decoder adapted and configured to provide variable template sizes for template matching. DETAILED DESCRIPTION OF THE INVENTION
[0006] One or more computer systems can be configured to perform specific operations or actions by having installed thereon software, firmware, hardware, or a combination thereof that, when operated, causes the system to perform the actions. One or more computer programs can be configured to perform specific operations or actions by including instructions that, when executed by a data processing device, cause the device to perform the actions. One such general aspect includes determining a coding unit, determining intensity information for pixels associated with a boundary of the coding unit, applying deblocking filtering to the coding unit prior to encoding based at least in part on the intensity information associated with the coding unit, and encoding the coding unit for transmission. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0007] Implementations may also include one or more of the following features: A method of encoding video, wherein stronger deblocking filtering is applied to coding units associated with intensity information having a value greater than a threshold value; A method of encoding video, wherein the threshold value is a predetermined value; A method of decoding video, further including identifying a neighboring coding unit proximate to the coding unit, determining intensity information of pixels associated with a boundary of the neighboring coding unit, and comparing the intensity information of pixels associated with the boundary of the coding unit with the intensity information of pixels associated with the neighboring coding unit, wherein the filtering is based at least in part on the comparison of the intensity information of pixels associated with the boundary of the coding unit with the intensity information of pixels associated with the neighboring coding unit; A method of encoding video, wherein stronger deblocking filtering is applied to coding units associated with intensity information having a value greater than a threshold value; A method of encoding video, wherein the threshold value is a predetermined value. Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.
[0008] One general aspect may include a method for decoding video, receiving a bitstream of encoded video, decoding the bitstream, determining a coding unit, determining intensity information for pixels associated with boundaries of the coding unit, applying deblocking filtering to the coding unit prior to encoding based at least in part on the intensity information associated with the coding unit, and encoding the coding unit for transmission. Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs recorded on one or more computer storage devices, each configured to perform the acts of the method.
[0009] Implementations may also include the same or similar features to the decoding process. Furthermore, implementations of the described techniques may include hardware, methods or processes, or computer software for computer-accessible media.
[0010] Further details of the invention will be explained with the aid of the accompanying drawings. FIG. 1 illustrates the allocation of a frame to multiple coding tree units (CTUs) 100. A frame can be an image in a video sequence. A frame can include a matrix or a set of matrices, with pixel values representing intensity measurements in the image. A set of these matrices can then generate a video sequence. Pixel values can be defined to represent color and luminance in full-color video coding, where pixels are allocated to three channels. For example, in the YCbCr color space, a pixel can have a luminance value Y, which represents the intensity of the gray levels in the image, and two chrominance values Cb and Cr, which represent how the color varies from gray to blue and red. In other embodiments, pixel values can be represented by values in different color spaces or color models. The resolution of the video can determine the number of pixels in a frame. Higher resolution can mean more pixels and better image definition, but it can also result in higher bandwidth, storage, and transmission requirements.
[0011] Frames of a video sequence can be encoded and decoded using JVET, a video coding scheme being developed by the Joint Video Exploration Team. A version of JVET has been implemented in the Joint Exploration Model (JEM) encoder and decoder. Similar to other video coding schemes such as High Efficiency Video Coding (HEVC), JVET is a block-based hybrid spatial-temporal predictive coding scheme. When coding using JVET, a frame is first divided into square blocks called CTUs 100, as shown in Figure 1. For example, a CTU 100 can be a block of 128 x 128 pixels.
[0012] 2a shows an exemplary partitioning of CTUs 100 into CUs 102. Each CTU 100 in a frame can be partitioned into one or more CUs (coding units) 102. The CUs 102 can be used for prediction and transform, as described below. Unlike HEVC, in JVET, CUs 102 can be rectangular or square and can be coded without further partitioning into prediction units or transform units. CUs 102 can be the same size as their root CTU 100, or can be smaller subdivisions of the root CTU 100, such as 4x4 blocks.
[0013] In JVET, a CTU 100 can be partitioned into CUs 102 according to a quadtree and binary tree (QTBT) scheme. The CTU 100 can be recursively partitioned into square blocks according to the quadtree, and the square blocks can then be recursively partitioned horizontally or vertically according to the binary tree. Parameters can be set to control the partitioning according to QTBT, such as CTU size, minimum size for quadtree leaf nodes and binary tree leaf nodes, maximum size for binary tree root node, and maximum depth for the binary tree. In VVC, a CTU 100 can also be partitioned into CUs by utilizing ternary partitioning. As a non-limiting example, FIG. 2a shows a CTU 100 partitioned into CUs 102, with solid lines indicating quadtree partitioning and dashed lines indicating binary tree partitioning. As shown, binary tree partitioning allows for horizontal and vertical partitioning, which can define the structure of the CTU and its subdivision into CUs. 2b and 2c show an alternative, non-limiting example of a third division of a CU, where the subdivision of the CU is not equal.
[0014] Figure 3 shows a QTBT representation for the section of Figure 2. The quadtree root node represents CTU 100, and each child node in the quadtree section represents one of four square blocks split from the parent square block. The square blocks represented by the quadtree leaf nodes can then be allocated zero or more times using a binary tree, with the quadtree leaf node being the root node of the binary tree. At each level of the binary tree section, blocks can be allocated either vertically or horizontally. A flag set to "0" indicates that the block is split horizontally, and a flag set to "1" indicates that the block is split vertically.
[0015] After quadtree and binary tree partitioning, the blocks represented by the leaf nodes of the QTBT represent the final CUs 102 to be coded, such as for coding using inter prediction or intra prediction. In the case of a slice or full frame coded by inter prediction, different partitioning structures can be used for the luma and chroma components. For example, in the case of an inter slice, CUs 102 may have coding blocks (CBs) for different color components, such as one luma CB and two chroma CBs. In the case of a slice or full frame coded by intra prediction, the partitioning structure can be the same for the luma and chroma components.
[0016] Figure 4 shows a simplified block diagram of CU coding in a JVET encoder. The main stages of video coding include partitioning to identify CUs 102 as described above, followed by encoding of the CUs 102 using prediction at 404 or 406, generating residual CUs 410 at 408, transforming at 412, quantizing at 416, and entropy coding at 420. The encoder and encoding process shown in Figure 4 also includes a decoding process, which is described in more detail below.
[0017] Given a current CU 102, the encoder can obtain a predicted CU 402 either spatially using intra prediction at 404 or temporally using inter prediction at 406. The basic idea of predictive coding is to transmit a difference or residual signal between an original signal and a prediction for the original signal. At the receiving end, the original signal can be reconstructed by adding the residual and the prediction, as described below. Because the difference signal is less correlated than the original signal, fewer bits are required for transmission.
[0018] A slice entirely coded by intra-predicted CUs, such as an entire picture or a portion of a picture, can be an I-slice that can be decoded without reference to other slices and can therefore be a possible point from which decoding can begin. A slice coded by at least some inter-predicted CUs can be a predictive (P) slice or a bi-predictive (B) slice that can be decoded based on one or more reference pictures. P slices can use intra-prediction and inter-prediction using previously coded slices. For example, P slices can be more compressed than I-slice by using inter-prediction, but require coding of previously coded slices to encode them. B slices can use intra-prediction or inter-prediction using interpolated prediction from two different frames to use data from previous and / or subsequent slices for their coding, which improves the accuracy of the motion estimation process. In some cases, P slices and B slices can be encoded together or alternately using intra-block copying, in which data from other parts of the same slice is used.
[0019] As described below, intra- or inter-prediction can be performed based on reconstructed CUs 434 from previously coded CUs 102, such as neighboring CUs 102 or CUs 102 in a reference picture.
[0020] When CU 102 is spatially coded using intra prediction at 404, an intra prediction mode can be found that best predicts pixel values of CU 102 based on samples from neighboring CUs 102 in the picture.
[0021] When coding the luma component of a CU, an encoder can generate a list of candidate intra prediction modes. While HEVC had 35 possible intra prediction modes for the luma component, JVET has 67 possible intra prediction modes for the luma component, and VVC has 85 prediction modes. These include a planar mode that uses a three-dimensional plane of values generated from neighboring pixels, a DC mode that uses values averaged from neighboring pixels, 65 directional modes that use values copied from neighboring pixels along the direction indicated by the solid line, as shown in Figure 5, and 18 wide-angle prediction modes that can be used in non-square blocks.
[0022] When generating a list of candidate intra-prediction modes for the luma component of a CU, the number of candidate modes on the list may depend on the size of the CU. The candidate list may include a subset of the 35 modes of HEVC with the lowest SATD (Sum of Absolute Transform Differences) cost, new directional modes added for JVET neighboring candidates found from the HEVC modes, and a set of six most probable modes (MPMs) for CU 102 identified based on the intra-prediction modes used for previously coded neighboring blocks and based on the list of default modes.
[0023] A list of candidate intra-prediction modes can also be generated when coding the chrominance component of a CU. The list of candidate modes can include modes generated using a cross-component linear model projection from luma samples, intra-prediction modes found for the luma CB at a particular arranged position of the chrominance block, and chrominance prediction modes previously found for neighboring blocks. The encoder can find the candidate modes on the list with the smallest rate-distortion cost and use these intra-prediction modes when coding the luma and chrominance components of the CU. Syntax can be coded in the bitstream indicating the intra-prediction mode used to code each CU 102.
[0024] After the best intra-prediction modes for CU 102 are selected, the encoder can use those modes to generate predicted CU 402. When the selected mode is a directional mode, the accuracy of the directionality can be improved by using a 4-tap filter. Columns or rows above or to the left of the prediction block can be adjusted using a boundary prediction filter, such as a 2-tap or 3-tap filter.
[0025] The predicted CU402 can be further smoothed by a position-dependent intra-prediction combining (PDPC) process, which adjusts the predicted CU402 generated based on filtered samples of neighboring blocks using unfiltered samples of neighboring blocks, or by adaptive reference sample smoothing using a 3-tap or 5-tap low-pass filter to process the reference samples.
[0026] When CU 102 is temporally coded using inter prediction at 406, a set of motion vectors (MVs) can be found that point to samples in reference pictures that best predict pixel values of CU 102. Inter prediction exploits temporal redundancy between slices by representing the displacement of pixel blocks within a slice. The displacement is determined according to pixel values in a previous or subsequent slice through a process called motion compensation. The motion vectors and associated reference indices that indicate pixel displacement relative to a particular reference picture can be provided to a decoder in a bitstream, along with residuals between the original and motion-compensated pixels. The decoder can reconstruct pixel blocks within a reconstructed slice using the residuals, the signaled motion vectors, and the reference indices.
[0027] In JVET, the precision of the motion vectors can be stored at 1 / 16 pixel, and the difference between the motion vector and the predicted motion vector of the CU can be coded at either quarter pixel or integer pixel resolution.
[0028] In JVET, techniques such as advanced temporal motion vector prediction (ATMVP), spatial-temporal motion vector prediction (STMVP), affine motion compensation prediction, pattern matching motion vector derivation (PMMVD), and / or bidirectional optical flow (BIO) can be used to find motion vectors for multiple sub-CUs within CU 102.
[0029] Using ATMVP, the encoder can find a temporal vector for CU 102 that points to a corresponding block in a reference picture. The temporal vector can be found based on the motion vectors and reference pictures found for previously coded neighboring CUs 102. Using the reference blocks pointed to by the temporal vector for the entire CU 102, a motion vector can be found for each sub-CU within CU 102.
[0030] STMVP can find motion vectors for sub-CUs by scaling and averaging motion vectors found for neighboring blocks previously coded using inter prediction with the temporal vector.
[0031] Using affine motion compensation prediction, a field of motion vectors for each sub-CU within a block can be predicted based on two control motion vectors found for the upper corners of the block. For example, motion vectors for sub-CUs can be derived based on the upper corner motion vectors found for each 4x4 block within CU 102.
[0032] PMMVD can use bilateral matching or template matching to find an initial motion vector for the current CU 102. Bilateral matching can look up the current CU 102 and reference blocks in two different reference pictures along the motion trajectory, while template matching can look up a corresponding block in the current CU 102 and a reference picture identified by the template. The initial motion vector found for CU 102 can then be refined for each sub-CU individually.
[0033] BIO can be used when performing inter prediction with bidirectional prediction based on a previous reference picture and a subsequent reference picture, and can find a motion vector for a sub-CU based on the gradient of the difference between the two reference pictures.
[0034] In some cases, local illumination compensation (LIC) can be used at the CU level, which finds values for scaling factor and offset parameters based on samples adjacent to the current CU 102 and based on corresponding samples adjacent to the reference block identified by the candidate motion vector. In JVET, the LIC parameters can be modified and signaled at the CU level.
[0035] For some of the above methods, the motion vectors found for each of the sub-CUs of a CU can be signaled to the decoder at the CU level. For other methods, such as PMMVD and BIO, the motion information is not signaled in the bitstream to save overhead, and the decoder can derive the motion vectors through the same process.
[0036] After the motion vectors for CU 102 are found, the encoder can use those motion vectors to generate the predicted CU 402. In some cases, when the motion vectors for individual sub-CUs are found, overlapping block motion compensation (OBMC) can be used in generating the predicted CU 402 by combining those motion vectors with previously found motion vectors for one or more neighboring sub-CUs.
[0037] When bidirectional prediction is used, JVET can find a motion vector by using decoder-side motion vector refinement (DMVR). DMVR can use a bidirectional template matching process to find a motion vector based on the two motion vectors found for bidirectional prediction. DMVR can find a weighted combination of the prediction CU 402 generated by each of the two motion vectors, and refine the two motion vectors by replacing them with a new motion vector that best points to the combined prediction CU 402. The two refined motion vectors can be used to generate the final prediction CU 402.
[0038] At 408, after the predicted CU 402 is found, as described above, by intra prediction at 404 or by inter prediction at 406, the encoder can subtract the predicted CU 402 from the current CU 102 to find the residual CU 410.
[0039] The encoder can use one or more transform operations at 412 to transform the residual CU 410 into transform coefficients 414 that represent the residual CU 410 in the transform domain. For example, a discrete cosine block transform (DCT) can be used to transform data into the transform domain. JVET allows more types of transform operations than HEVC, including DCT-II, DST-VII, DST-VII, DCT-VIII, DST-I, and DCT-V operations. The allowed transform operations can be grouped into subsets, and an indication of which subsets were used and which specific operations within those subsets were used can be signaled by the encoder. In some cases, using a large block size transform can zero out high-frequency transform coefficients in CUs 102 larger than a certain size, thereby retaining only low-frequency transform coefficients for those CUs 102.
[0040] In some cases, a mode-dependent non-separable quadratic transform (MDNSST) can be applied to the low-frequency transform coefficients 414 after the forward core transform. The MDNSST operation can use a Hypercube-Givens Transform (HyGT) based on rotated data. When in use, an index value that identifies the particular MDNSST operation can be signaled by the encoder.
[0041] At 416, the encoder can quantize the transform coefficients 414 into quantized transform coefficients 416. The quantization of each coefficient may be calculated by dividing the value of the coefficient by a quantization step derived from a quantization parameter (QP). In some embodiments, Qstep is 2 (QP-4) / 6 QP is defined as: Quantization can aid in data compression because it can convert high-precision transform coefficients 414 into quantized transform coefficients 416, which have a finite number of possible values. Quantizing the transform coefficients can therefore limit the amount of bits generated by the transform process and transmitted. However, quantization is a lossy operation, and although quantization losses cannot be recovered, the quantization process presents a trade-off between the quality of the reconstructed sequence and the amount of information needed to represent the sequence. For example, a lower QP value may require a greater amount of data for representation and transmission, but can result in better-quality decoded video. In contrast, a higher QP value may reduce the quality of the reconstructed video sequence, but requires less data and bandwidth.
[0042] JVET may utilize a variance-based adaptive quantization technique, in which every CU 102 may use a different quantization parameter for the coding process (instead of using the same frame QP in coding all CUs 102 of a frame). The variance-based adaptive quantization technique adaptively reduces the quantization parameter for certain blocks and increases the quantization parameter for other blocks. To select a particular QP for a CU 102, the variance of the CU is calculated. Simply put, if the variance of the CU is greater than the average variance of the frame, a larger QP than the frame QP may be set for that CU 102. If the CU 102 exhibits a lower variance than the average variance of the frame, a smaller QP may be assigned.
[0043] At 420, the encoder can find the final compressed bits 422 by entropy coding the quantized transform coefficients 418. Entropy coding aims to remove statistical redundancy in the information to be transmitted. In JVET, the quantized transform coefficients 418 can be coded using CABAC (Context Adaptive Binary Arithmetic Coding), which uses a probability measure to remove statistical redundancy. For CUs 102 that have non-zero quantized transform coefficients 418, the quantized transform coefficients 418 can be converted to binary. Each bit ("bin") of the binary representation can then be encoded using a context model. The CU 102 can be divided into three regions, each with its own set of context models to use for the pixels within that region.
[0044] Multiple scan passes can be performed to encode a bin. During the pass that encodes the first three bins (bin0, bin1, and bin2), the index value that indicates which context model to use for a bin can be found by finding the sum of that bin's position in up to five previously coded adjacent quantized transform coefficients 418 identified by the template.
[0045] The context model may be based on the probability that a bin's value is "0" or "1." As values are coded, the probabilities in the context model may be updated based on the actual number of "0" and "1" values encountered. While HEVC reinitialized the context model for each new picture by using a fixed table, in JVET the probabilities of the context model for a new inter-predicted picture may be initialized based on the context model developed for a previously coded inter-predicted picture.
[0046] The encoder may produce a bitstream that includes entropy-encoded bits 422 of the residual CU 410, prediction information such as a selected intra-prediction mode or motion vector, an indication of how CU 102 was partitioned from CTU 100 according to the QTBT structure, and / or other information about the encoded video. The bitstream can be decoded by a decoder, as described below.
[0047] In addition to using the quantized transform coefficients 418 to find the final compressed bits 422, the encoder can also use the quantized transform coefficients 418 to generate a reconstructed CU 434 by following the same decoding process that a decoder uses to generate a reconstructed CU 434. Thus, after the transform coefficients are calculated and quantized by the encoder, the quantized transform coefficients 418 can be sent to the encoder's decoding loop. After quantizing the transform coefficients of the CU, the decoding loop enables the encoder to generate the same reconstructed CU 434 that the decoder generates in the decoding process. Thus, when performing intra prediction or inter prediction on a new CU 102, the encoder can use the same reconstructed CU 434 that the decoder uses for a neighboring CU 102 or a reference picture. The reconstructed CU 102, a reconstructed slice, or a fully reconstructed frame can serve as a reference for further prediction stages.
[0048] In the decoding loop of the encoder to obtain pixel values for a reconstructed image (see below for the same operation in the decoder), an inverse quantization process can be performed. To inverse quantize a frame, for example, the quantization value for each pixel of the frame can be multiplied by a quantization step, such as Qstep described above, to obtain reconstructed inverse quantized transform coefficients 426. For example, in the decoding process shown in FIG. 4 in the encoder, the quantized transform coefficients 418 of the residual CU 410 can be inverse quantized at 424 to find the inverse quantized transform coefficients 426. If an MDNSST operation was performed during encoding, that operation can be reversed after inverse quantization.
[0049] At 428, the dequantized transform coefficients 426 may be inverse transformed to find a reconstructed residual CU 430, e.g., a DCT may be applied to the values to obtain a reconstructed image. At 432, the reconstructed residual CU 430 may be added to the corresponding predicted CU 402 found by intra prediction at 404 or inter prediction at 406, thereby finding a reconstructed CU 434.
[0050] At 436, one or more filters can be applied to the reconstructed data during the decoding process (either in the encoder or, as described below, in the decoder) at either the picture level or the CU level. For example, the encoder can apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). The encoder's decoding process can implement filters to estimate and transmit optimal filter parameters to the decoder to address potential artifacts in the reconstructed image. Such refinements improve the objective and subjective quality of the reconstructed video. Deblocking filtering can modify pixels near sub-CU boundaries, while SAO can modify pixels within the CTU 100 using either edge offset or band offset classification. JVET's ALF can use a circularly symmetric filter for each 2x2 block. An indication of the size and identity of the filter used for each 2x2 block can be signaled.
[0051] If the reconstructed pictures are reference pictures, they may be stored in a reference buffer 438 for inter-prediction of future CUs 102 at 406. During the above steps, JVET can use content-adaptive clipping operations to adjust color values to fit between lower and upper clipping boundaries. The clipping boundaries can change on a slice-by-slice basis, and parameters identifying the boundaries can be signaled in the bitstream.
[0052] 6 shows a simplified block diagram of CU coding in a JVET decoder. The JVET decoder can receive a bitstream containing information about encoded CUs 102. The bitstream can indicate how the CUs 102 of a picture were partitioned from the CTUs 100 according to the QTBT structure, prediction information about the CUs 102, such as intra-prediction modes or motion vectors, and bits 602 representing entropy-encoded residual CUs.
[0053] At 604, a decoder can decode the entropy encoded bits 602 using the CABAC context model signaled in the bitstream by the encoder. The decoder can update the probabilities of the context model using the parameters signaled by the encoder in the same way that they were updated during encoding.
[0054] After reversing the entropy encoding at 604 to find the quantized transform coefficients 606, the decoder can dequantize them at 608 to find the dequantized transform coefficients 610. If an MDNSST operation was performed during encoding, that operation can be reversed by the decoder after dequantization.
[0055] At 612, the dequantized transform coefficients 610 may be inverse transformed to find a reconstructed residual CU 614. At 616, the reconstructed residual CU 614 may be added to a corresponding predicted CU 626 found by intra prediction at 622 or inter prediction at 624 to find a reconstructed CU 618.
[0056] At 620, one or more filters can be applied to the reconstructed data at either the picture level or the CU level. For example, the decoder can apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). As described above, an in-loop filter located in the encoder's decode loop can be used to estimate optimal filter parameters to improve the objective and subjective quality of the frame. These parameters are transmitted to the decoder at 620 to filter the reconstructed frame to match the frame filtered and reconstructed in the encoder.
[0057] After the reconstructed picture is generated by finding the reconstructed CU 618 and applying the notified filter, the decoder can output the reconstructed picture as output video 628. If the reconstructed pictures are used as reference pictures, they can be stored in a reference buffer 630 for inter-prediction of future CUs 102 at 624.
[0058] FIG. 7 shows a block diagram 700 of HDR encoding 702 and decoding 704. One common HDR video format uses a linear light RGB domain, with each channel specified in a high bit-depth format, as a non-limiting example, the half-float format of the EXR file format. Because current video compression algorithms cannot directly handle HDR video formats, one approach to encoding HDR video is to first convert it to a format acceptable to the video encoder. The decoder video can then be converted back to the HDR format. An example of such a system is shown in FIG. 7, where the encoding 702 and decoding 704 modules correspond to the process described herein for JVET coding of SDR content.
[0059] The upper system of Figure 7 shows an example of converting an input HDR video format to a 10-bit 4:2:0 video format that can be encoded using a JVET encoder (or a Main10 HEVC encoder, etc.). To prepare the high-bit-depth input for conversion to a lower bit-depth, each RGB channel in the input HDR video first passes through a coding transfer function (TF) 706. The output R'G'B' is then converted to a color space more suitable for video coding, Y'CbCr 708. A perceptual mapping is then performed in step 710, and then each channel is quantized to 10 bits in step 712. After uniformly quantizing each channel to 10 bits in step 712, the chroma Cb and Cr channels are subsampled to a 4:2:0 format in step 714. The encoder then compresses the 10-bit 4:2:0 video using, for example, a Main10 HEVC encoder in step 716.
[0060] The lower system of Figure 7 reconstructs output HDR video from an input bitstream. In one embodiment, the bitstream is decoded in step 817, and a JVET decoder (or a Main10 HEVC decoder, or other known, convenient, and / or desired decoder) reconstructs 10-bit 4:2:0 video, which is then upsampled to a 4:4:4 format in step 720. After inverse quantization remapping of the 10-bit data in step 722, inverse perceptual mapping is applied in step 724 to generate Y'CbCr values. The Y'CbCr data can then be converted to R'G'B' color space in step 726, and the channels may undergo an inverse coding TF operation in step 728 before the HDR video data is output.
[0061] Blocking artifacts are primarily a result of the independent coding of adjacent units in block-based video coding. They tend to occur at low bit rates and be visible when adjacent blocks have different intra- or inter-coding types and in areas with low spatial activity. As a result, they introduce artificial discontinuities or boundaries, resulting in visual artifacts.
[0062] Deblocking filters, such as those in HEVC [1] and the current JVET, attempt to reduce visual artifacts by smoothing or low-pass filtering across PU / TU or CU boundaries. In some embodiments, vertical boundaries are filtered first, followed by horizontal boundaries. Up to four luminance pixel values reconstructed in a 4x4 region on either side of the boundary can be used to filter up to three pixels on either side of the boundary. Normal or weak filtering can filter up to two pixels on either side, while strong filtering filters three pixels on either side. The decision to filter a pixel may be based on the intra / inter mode decision, motion information, and residual information of neighboring blocks to generate a boundary strength value Bs of 0, 1, or 2. If Bs > 0, smoothness conditions can be checked in the first and last rows (or columns) of a 4x4 region on either side of the vertical (or horizontal) boundary. These conditions can determine the degree of deviation from the gradient across a given boundary. Generally, if the deviation is less than a threshold specified by the parameter β, deblocking filtering can be applied to the entire 4x4 region; large deviations may indicate the presence of true or intended boundaries, and therefore deblocking filtering may not be performed. The beta parameter is a non-decreasing function of the block QP value, such that larger QP values correspond to larger thresholds. In some embodiments, if Bs>0 and the smoothness condition is satisfied, a decision between strong and weak filtering can be made based on an additional smoothness condition and another parameter tc, which is also a non-decreasing function of QP. Generally, strong filtering is applied to smoother regions, since discontinuities are more visible in such regions.
[0063] In some embodiments, the deblocking filter operation is effectively a 4-tap or 5-tap filtering operation, where the difference between the input and the filtered output is first clipped and then added back to (or subtracted from) the input. The clipping attempts to limit over-smoothing, and the clipping level may be determined by tc and QP. For chroma deblocking, if at least one of the blocks is intra-coded, a 4-tap filter may be applied to one pixel on each side of the boundary.
[0064] Deblocking artifacts may result from inconsistencies in block boundaries (e.g., CU, prediction, transform boundaries, and / or other segmentation boundaries). These differences may be in DC levels, alignment, phase, and / or other data. Therefore, boundary differences may be considered noise added to the signal. As shown in FIG. 7, the original input HDR signal passes through both the coding TF and the inverse coding TF, but the deblocking noise passes through only the inverse coding TF. Conventional SDR deblocking artifacts are developed without considering this additional TF, and the decoder output in FIG. 7 is visible. In the case of HDR, the deblocking noise passes through the inverse coding TF and can change the visibility of the artifact. Therefore, the same discontinuous jump in both bright and dark areas may result in a larger or smaller discontinuous jump after the inverse coding TF operation.
[0065] Typical inverse coding TFs (often known as EOTFs), such as PQ, HLG, and Gamma, have the property that they are monotonically increasing functions of intensity. FIG. 8 shows a plotted curve 800 of a normalized PQ EOTF versus intensity (I) 802. For example, a normalized PQ EOTF curve 800 is shown in FIG. 8. Because the slope of the PQ EOTF curve 800 increases, discontinuous jumps are magnified by the EOTF in brighter versus darker areas, potentially making deblocking artifacts more visible. According to Weber's Law, it is understood that the larger the JND (Just Noticeable Difference), the greater the difference a viewer can tolerate in brighter areas. However, FIG. 9, which shows a normalized plot 900 of JND 902 plotted against intensity 802, shows that the JND decreases at higher intensities for the PQ EOTF, even when Weber's Law is taken into account. Figure 9, calculated based on α = 8% of the Weber's Law JND threshold, shows that the peak JND appears to be less sensitive to a wide range of PQ thresholds. Indeed, the peak JND for PQ appears to occur around the unity slope of the PQ EOTF in Figure 8, which occurs at approximately I = 78% of the (normalized) peak intensity. Alternative testing shows that for the HLG EOTF, the peak JND intensity appears to occur at approximately I = 50% of the (normalized) intensity, and the unity slope appears to occur at approximately 70% of the (normalized) intensity.
[0066] Based on this analysis and related visual observations, it becomes clear that an intensity-dependent deblocking filter operation results in improved performance. That is, by way of non-limiting example, the deblocking filter coefficients, the strength of the applied filtering (normal filtering vs. weak filtering), the number of input and output pixels used or affected, the decision to turn filtering on / off, and other filtering criteria may be influenced by, and thus based on, the intensity. The intensity may be relative to luma and / or chroma, and may be based on either nonlinear or linear intensities. In some embodiments, the intensity may be calculated based on local intensities, such as based on CU intensities or neighboring pixels around block boundaries. In some embodiments, the intensity may be a maximum, minimum, average, or some other statistical or metric value based on the luma / chroma of neighboring pixels. In alternative embodiments, the deblocking filtering may be based on intensities for a frame or group of frames per scene, sequence, or other inter- or intra-unit value.
[0067] In some embodiments, the deblocking operation may be determined based on a strength operation calculated at the encoder and / or decoder, or parameter(s) may be transmitted in the bitstream to the decoder for use in making the deblocking decision or filtering operation. The parameters may be transmitted at the CU, slice, picture, PPS, SPS level, and / or any other known, convenient, and / or desired level.
[0068] While intensity-based deblocking can also be applied to SDR content, it is expected that intensity-based deblocking will have a greater impact on HDR content due to the decoding TF applied to HDR. In some embodiments, deblocking may be based on the decoding TF (or coding TF). TF information may be signaled in the bitstream and used by the deblocking operation. As a non-limiting example, different deblocking strategies may be used based on whether the (local or collective) intensity is greater than or less than some threshold, which may be based on the TF. Additionally, in some embodiments, two or more thresholds may be identified and associated with multiple levels of filtering operations. In some embodiments, exemplary deblocking strategies may include filtering vs. no filtering, strong filtering vs. weak filtering, and / or different levels of filtering based on trigger values for different intensity levels. In some embodiments, it may be determined that deblocking filtering is not necessary after the decoding TF because artifacts may become less (or not) visible, thereby reducing computational requirements. The value of I* (a normalized intensity value) may be signaled, calculated, or specified based on the TF and used as a threshold in determining filtering. In some embodiments, more than one threshold may be used to modify the deblocking filter operation.
[0069] Modifications can be made to existing SDR deblocking in HEVC or JVET to incorporate intensity-based deblocking for HDR. As a non-limiting example, in HEVC, the deblocking parameters β (and tc) can be modified based on intensity to increase or decrease strong filtering / normal filtering or filtering on / off, and different β (and tc) parameter curves can be defined for HDR based on intensity values or ranges of intensity values. Alternatively, a shift or offset can be applied to the parameters and curves based on the intensity of a boundary, CU, region, or neighboring frame group. As a non-limiting example, a shift can be applied so that stronger filtering is applied to brighter areas.
[0070] FIG. 10 shows a block diagram of an encoding system 1000 in which strength is taken into consideration for purposes of determining filtering. In step 1002, information about a coding unit and adjacent / neighboring coding units may be obtained. A decision may then be made in step 1004 regarding whether to apply filtering. If it is determined in step 1004 that filtering is to be applied, in step 1006, strength values associated with the coding unit and / or adjacent / neighboring coding unit(s) may be evaluated. Based on the evaluation of the strength values in step 1006, a desired level of filtering may be applied to the coding unit in one of steps 1008a-1008c. In some embodiments, the selection of the level of filtering may be based on a comparison of the strength value of the coding unit and / or the strength values associated with the coding unit with strength values associated with one or more adjacent coding units. In some embodiments, this may be based on one or more established threshold strength values. After applying filtering in one of steps 1008a-1008c, the coding unit may be encoded for transmission in step 1010. However, if it is determined in step 1004 that filtering should not be applied, steps 1006-1008c may be bypassed, allowing the unfiltered coding unit to proceed directly to encoding in step 1010.
[0071] In alternative embodiments, step 1006 may precede step 1004, and the strength assessment may be used to determine filtering in step 1006, and step 1004 may be followed directly by either encoding in step 1010 if filtering is not desired, or one of steps 1008a-1008c if filtering is desired.
[0072] FIG. 11 shows a block diagram of a decoding system in which intensity is a factor taken into account in filtering for display. In the embodiment shown in FIG. 11, a bitstream may be received and decoded in step 1102. In some embodiments, an appropriate and / or desired level of deblocking may be determined in step 1104. However, in some alternative embodiments, step 1104 may determine whether filtering was applied during phase encoding. If filtering is determined to be desirable in step 1104 (or, in some embodiments, filtering was applied during phase encoding), a level of filtering is determined in step 1106. In some embodiments, this may be an offset value for use in establishing one or more factors associated with filtering and / or an indication of the level of filtering applied during phase encoding. Based at least in part on the determination of step 1106, levels 1108a-1108c of filtering are applied to render an image for display in step 1110. If filtering was not applied during phase encoding in step 1104, the image may be rendered for display in step 1110.
[0073] Figures 12a-12c show a series of exemplary β vs. QP and tc vs. QP curves 1200 that graphically represent the system described and shown in Figures 10 and 11. In the embodiment shown in Figure 12a, an exemplary pair of β vs. QP curves 1202 and tc vs. QP curves 1204 are presented, which may be used when the intensity is below a desired threshold x 1206. Thus, when the intensity value falls below the desired value x 1202, normal or typical values of β and tc can be used to determine the level of deblocking to apply. Figures 12b and 12c show β vs. QP curves 1212, 1222 and tc vs. QP curves 1214, 1224, which may be used when the intensity is determined to be equal to or greater than the desired value x 1208. FIG. 12b shows the same set of curves 1212, 1214 as shown in FIG. 12a, but shifted to the left, and FIG. 12c shows the same set of curves 1222, 1224 as shown in FIG. 12a, but shifted up. Thus, when the intensity value meets or exceeds (or exceeds) a desired value x, offset, non-standard, or modified values of β and tc can be used to determine the level of deblocking applied. Accordingly, as the intensity value increases, increased values of β and tc are selected, increasing the level of filtering applied. While FIGS. 12b and 12c illustrate variants where the intensity (I) is greater than or equal to a single value x, it should be appreciated that the present system can be extended to encompass systems where there are multiple sets of β vs. QP curves and tc vs. QP curves, each associated with various boundaries. That is, I<x、x≦I≦y、およびI> Conditions such as y and / or systems using multiple boundaries or regions are contemplated. Additionally, note that the use of <, >, <=, and >= is arbitrary, and that any logical boundary condition can be used. Finally, it should be appreciated that the curves represented in Figures 12a-12c are exemplary in nature, and that the same or similar techniques, methods, and logic can be applied to any known, convenient, and / or desired set of curves.
[0074] Execution of the instruction sequences necessary to implement the embodiments may be performed by a computer system 1300, as shown in Figure 13. In one embodiment, execution of the instruction sequences is performed by a single computer system 1300. According to other embodiments, two or more computer systems 1300 coupled by a communications link 1315 may execute the instruction sequences in coordination with one another. Although a description of only one computer system 1300 is presented below, it should be understood that any number of computer systems 1300 may be used to implement the embodiments.
[0075] A computer system 1300 according to one embodiment will now be described with reference to Figure 13, which is a block diagram of functional components of computer system 1300. As used herein, the term computer system 1300 is used broadly to describe any computing device that can store and independently execute one or more programs.
[0076] Each computer system 1300 may include a communication interface 1314 coupled to bus 1306. The communication interface 1314 provides two-way communication between the computer systems 1300. The communication interface 1314 of each computer system 1300 sends and receives electrical, electromagnetic, or optical signals containing data streams representing various types of signal information, such as commands, messages, and data. The communication link 1315 links one computer system 1300 to another computer system 1300. For example, the communication link 1315 may be a LAN, in which case the communication interface 1314 may be a LAN card, or the communication link 1315 may be a PSTN, in which case the communication interface 1314 may be an Integrated Services Digital Network (ISDN) card or modem, or the communication link 1315 may be the Internet, in which case the communication interface 1314 may be a dial-up, cable, or wireless modem.
[0077] Computer system 1300 can send and receive messages, data, and instructions, including programs, i.e., applications or code, via respective communications links 1315 and communications interface 1314. Received program code may be executed by respective processor 1307 as received, and / or stored in storage device 1310 or other associated non-volatile media for later execution.
[0078] In one embodiment, computer system 1300 operates in conjunction with data storage system 1331, for example, data storage system 1331 including database 1332 readily accessible by computer system 1300. Computer system 1300 communicates with data storage system 1331 via data interface 1333. Data interface 1333, coupled to bus 1306, sends and receives electrical, electromagnetic, or optical signals containing data streams representing various types of signal information, such as, for example, commands, messages, and data. In an embodiment, the functionality of data interface 1333 may be performed by communication interface 1314.
[0079] Computer system 1300 includes a bus 1306 or other communication mechanism for communicating information, collectively instructions, messages, and data, and one or more processors 1307 coupled to bus 1306 for processing information. Computer system 1300 also includes a main memory 1308, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 1306 for storing dynamic data and instructions that can be executed by processor(s) 1307. Main memory 1308 may also be used for storing temporary data or variables or other intermediate information during execution of instructions by processor(s) 1307.
[0080] Computer system 1300 may further include a read-only memory (ROM) 1309 or other static storage device, coupled to bus 1306, for storing static data and instructions for processor(s) 1307. A storage device 1310, such as a magnetic disk or optical disk, may also be provided and coupled to bus 1306 for storing data and instructions for processor(s) 1307.
[0081] Computer system 1300 may be coupled via bus 1306 to a display device 1311, such as, but not limited to, a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, for displaying information to a user. Input devices 1312, such as alphanumeric and other keys, are coupled to bus 1306 for communicating information and command selections to processor(s) 1307.
[0082] According to one embodiment, each computer system 1300 performs certain operations by its respective processor(s) 1307 executing one or more sequences of one or more instructions contained in main memory 1308. Such instructions may be read into main memory 1308 from another computer-usable medium, such as ROM 1309 or storage device 1310. Execution of the sequences of instructions contained in main memory 1308 causes processor(s) 1307 to perform the processes described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions. Thus, embodiments are not limited to any specific combination of hardware circuitry and / or software.
[0083] As used herein, the term "computer-usable medium" refers to any medium that provides information or is otherwise usable by the processor(s) 1307. Such media can take many forms, including but not limited to nonvolatile media, volatile media, and transmission media. Nonvolatile media, i.e., media that can retain information without electrical power, include ROM 1309, CD-ROM, magnetic tape, and magnetic disks. Volatile media, i.e., media that cannot retain information without electrical power, includes main memory 1308. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 1306. Transmission media can also take the form of carrier waves, i.e., electromagnetic waves that can be modulated in frequency, amplitude, or phase, etc., to transmit information signals. Additionally, transmission media can take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0084] In the foregoing specification, the embodiments have been described with reference to specific components thereof. However, it will be apparent that various modifications and changes can be made thereto without departing from the broader spirit and scope of the embodiments. For example, the reader will understand that the specific order and combination of process operations shown in the process flow diagrams described herein are merely exemplary, and that embodiments may be practiced using different or additional process operations, or using different combinations or orders of process operations. Accordingly, the specification and drawings should be regarded as illustrative rather than restrictive.
[0085] It should also be noted that the present invention may be implemented in various computer systems. The various techniques described herein may be implemented in hardware or software, or a combination of both. Preferably, the techniques are implemented in computer programs running on a programmable computer, each of which includes a processor, a processor-readable storage medium (including volatile memory, non-volatile memory, and / or storage elements), at least one input device, and at least one output device. The program code is applied to data entered using the input device to perform the functions described above and generate output information. The output information is applied to one or more output devices. Each program is preferably implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In either case, the language can be a compiled or interpreted language. Each such computer program is preferably stored on a storage medium or device (e.g., ROM or magnetic disk) that is readable by a general-purpose computer or a special-purpose programmable computer to configure and operate the computer when read by the computer to perform the procedures described above. The system may also be considered to be implemented as a computer-readable storage medium configured with a computer program, where the configured storage medium causes the computer to operate in a particular, predefined manner. Furthermore, the storage element of an exemplary computing application may be a relational or sequential (flat file) type computing database that may store data in various combinations and configurations.
[0086] 14 is a schematic diagram of a source device 1412 and a destination device 1410 that may incorporate features of the systems and devices described herein. As shown in FIG. 14, the exemplary video coding system 1410 includes a source device 1412 and a destination device 1414; in this example, the source device 1412 generates encoded video data. Accordingly, the source device 1412 may be referred to as a video encoding device. The destination device 1414 can decode the encoded video data generated by the source device 1412. Accordingly, the destination device 1414 may be referred to as a video decoding device. The source device 1412 and the destination device 1414 may be examples of video coding devices.
[0087] Destination device 1414 can receive encoded video data from source device 1412 via channel 1416. Channel 1416 can include any type of medium or device that can move encoded video data from source device 1412 to destination device 1414. In one example, channel 1416 can include a communication medium that allows source device 1412 to transmit encoded video data directly to destination device 1414 in real time.
[0088] In this example, source device 1412 may modulate the encoded video data according to a communication standard, such as a wireless communication protocol, and transmit the modulated video data to destination device 1414. The communication medium may include a wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or other equipment that facilitates communication from source device 1412 to destination device 1414. In another example, channel 1416 may correspond to a storage medium that stores the encoded video data generated by source device 1412.
[0089] 14 , source device 1412 includes a video source 1418, a video encoder 1420, and an output interface 1422. In some cases, output interface 1428 may include a modulator / demodulator (modem) and / or a transmitter. In source device 1412, video source 1418 may include a source such as a video capture device such as a video camera, a video archive containing previously captured video data, a video feed interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources.
[0090] The video encoder 1420 can encode captured, pre-captured, or computer-generated video data. Input images can be received by the video encoder 1420 and stored in an input frame memory 1421, from which a general-purpose processor 1423 can read information and perform the encoding. Programs for driving the general-purpose processor can be loaded from a storage device, such as the exemplary memory module shown in FIG. 14. The general-purpose processor can perform the encoding using the processing memory 1422, and the output of the information encoded by the general-purpose processor can be stored in a buffer, such as an output buffer 1426.
[0091] The video encoder 1420 may include a resampling module 1425 that may be configured to code (e.g., encode) the video data in a scalable video coding scheme that defines at least one base layer and at least one enhancement layer. The resampling module 1425 may resample at least some of the video data as part of the encoding process, and the resampling may be performed in an adaptive manner using a resampling filter.
[0092] The encoded video data, such as a coded bitstream, can be transmitted directly to the destination device 1414 via the output interface 1428 of the source device 1412. In the example of FIG. 14, the destination device 1414 includes an input interface 1438, a video decoder 1430, and a display device 1432. In some cases, the input interface 1428 can include a receiver and / or a modem. The input interface 1438 of the destination device 1414 receives the encoded video data via the channel 1416. The encoded video data can include various syntax elements that represent the video data and that are generated by the video encoder 1420. Such syntax elements can be included in the encoded video data transmitted over a communication medium or stored on a storage medium or file server.
[0093] The encoded video data may also be stored on a storage medium or file server for later access by the destination device 1414 for decoding and / or playback. For example, the coded bitstream may be temporarily stored in an input buffer 1431 and then loaded into a general-purpose processor 1433. A program for driving the general-purpose processor may be loaded from a storage device or memory. The general-purpose processor may perform the decoding by using processing memory 1432. The video decoder 1430 may also include a resampling module 1425 similar to the resampling module 1435 used in the video encoder 1420.
[0094] 14 shows the resampling module 1435 separate from the general-purpose processor 1433, those skilled in the art will understand that the resampling function may be performed by a program executed by the general-purpose processor, and that processing in a video encoder may be accomplished using one or more processors. The decoded image(s) may be stored in an output frame buffer 1436 and then sent to an input interface 1438.
[0095] The display device 1438 may be integrated with the destination device 1414 or may be located external to the destination device 1414. In some examples, the destination device 1414 may include an integrated display device and may be configured to interface to an external display device. In other examples, the destination device 1414 may be a display device. In general, the display device 1438 displays the decoded video data to a user.
[0096] The video encoder 1420 and the video decoder 1430 can operate according to a video compression standard. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) are studying the potential need for standardization of future video coding techniques with compression capabilities significantly exceeding those of the current High Efficiency Video Coding (HEVC) standard (including current and near-term extensions for screen content coding and high dynamic range coding). Both groups are working on this research effort in a collaborative effort known as the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by experts in the field. A recent capture of JVET developments is described in "Algorithm Description of Joint Exploration Test Model 5 (JEM 5)," JVET-E1001-V2, by J. Chen, E. Alshina, G. Sullivan, J. Ohm, and J. Boyce.
[0097] Additionally or alternatively, the video encoder 1420 and the video decoder 1430 may operate according to other proprietary or industry standards that function with the disclosed JVET functionality. Examples of such standards include the ITU-T H.264 standard, Part 10, Advanced Video Coding (AVC), alternatively referred to as MPEG-4, or extensions to those standards. Thus, while newly developed for JVET, the techniques of this disclosure are not limited to any particular coding standard or coding technique. Other examples of standards and techniques related to video compression include MPEG-2, ITU-T H.263, and proprietary or open-source compression and related formats.
[0098] The video encoder 1420 and the video decoder 1430 may be implemented in hardware, software, firmware, or any combination thereof. For example, the video encoder 1420 and the decoder 1430 may use one or more processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. If the video encoder 1420 and the decoder 1430 are implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and may execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of the video encoder 1420 and the video decoder 1430 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.
[0099] Aspects of the subject matter described herein may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer, such as the general-purpose processors 1423 and 1433 described above. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Aspects of the subject matter described herein may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including memory storage devices.
[0100] Examples of memory include random access memory (RAM), read-only memory (ROM), or both. The memory may store instructions, such as source code or binary code, for executing the techniques described above. The memory may also be used to store variables or other intermediate information during execution of instructions executed by a processor, such as processors 1423 and 1433.
[0101] The storage device may also store instructions, such as source code or binary code, for executing the techniques described above. The storage device may additionally store data used and manipulated by a computer processor. For example, the storage device in the video encoder 1420 or the video decoder 1430 may be a database accessed by the computer system 1423 or 1433. Other examples of storage devices include random access memory (RAM), read-only memory (ROM), a hard drive, a magnetic disk, an optical disk, a CD-ROM, a DVD, a flash memory, a USB memory card, or any other medium from which a computer can read.
[0102] The memory or storage device may be an example of a non-transitory computer-readable storage medium for use by or in connection with a video encoder and / or decoder. The non-transitory computer-readable storage medium includes instructions for controlling a computer system such that the computer system can be configured to perform the functions described in certain embodiments. The instructions, when executed by one or more computer processors, can be configured to perform the functions described in certain embodiments.
[0103] Also, note that some embodiments are described as processes that may be illustrated as flow diagrams or block diagrams. While each may be described as a sequential process of operations, many of the operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process may have additional steps not included in the diagrams.
[0104] Certain embodiments may be implemented in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to perform the methods described by certain embodiments. The computer system may include one or more computing devices. The instructions, when executed by one or more computer processors, may be configured to perform the operations described in certain embodiments.
[0105] As used in the description and throughout the claims that follow, "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used in the description and throughout the claims that follow, the meaning of "in" includes "in" and "on" unless the context clearly dictates otherwise.
[0106] While exemplary embodiments of the present invention have been described in detail in language specific to the structural features and / or methodological acts set forth above, it should be understood that those skilled in the art will readily appreciate that many additional modifications are possible in the exemplary embodiments without materially departing from the novel teachings and advantages of the present invention. It should further be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Accordingly, these and all such modifications are intended to be included within the scope of the present invention, which is to be broadly interpreted in accordance with the appended claims.
Claims
1. 1. A method of decoding video, comprising: (a) receiving an encoded video bitstream based on a coding tree unit associated with a coding unit; (b) decoding the bitstream of the encoded video; and (c) determining a quantization-coded rectangular coding unit in the encoded bitstream, in which a block of a luma component and a block of a chroma component are coded, the rectangular coding unit having a width and a height, the width and the height being different from each other, and the rectangular coding unit being neither a prediction unit nor a transform unit; (d) determining intensity information of pixels associated with vertical or horizontal boundaries of the rectangular coding unit; (e) applying deblocking filtering to the rectangular coding unit based at least in part on the intensity information associated with the rectangular coding unit, the applied deblocking filtering being based on filtering parameters β and tc that specify boundary filtering selectively modified based on an offset of a quantization parameter, the offset being based at least in part on the determined intensity information of pixels associated with the boundary; (f) a larger offset results in stronger filtering compared to not having the larger offset; (g) the offsets include the offsets of three or more quantization parameters that are different from one another; (h) the deblocking filtering is further based on clipping; (i) The method of decoding video, wherein the clipping is based on the quantization parameter.
2. 1. A method of encoding a bitstream by an encoder, comprising: (a) providing a bitstream that indicates how a rectangular coding unit in the encoded bitstream is coded by quantization based on a coding tree unit associated with a coding unit, wherein a block of a luminance component and a block of a chrominance component are coded in the rectangular coding unit, the rectangular coding unit has a width and a height, the width and the height being different from each other, and the rectangular coding unit is neither a prediction unit nor a transform unit; (b) pixel intensity information is associated with a vertical or horizontal boundary of the rectangular coding unit; (c) applying deblocking filtering to the rectangular coding unit based at least in part on the intensity information associated with the rectangular coding unit, the applied deblocking filtering being based on filtering parameters β and tc that specify boundary filtering selectively modified based on an offset of a quantization parameter, the offset being based at least in part on the determined intensity information of pixels associated with the boundary; (d) a larger offset results in stronger filtering compared to not having the larger offset; (e) the offsets include the offsets of three or more quantization parameters that are different from each other; (f) the deblocking filtering is further based on clipping; (g) the clipping is based on the quantization parameter.
3. 1. A non-transitory computer-readable storage medium having stored thereon instructions in the form of a bitstream for causing a video decoder to provide extracted video data, the instructions being provided to a processor of the video decoder to cause the processor to extract the video data, the instructions comprising: (a) determining a rectangular coding unit in the bitstream to be encoded based on a coding tree unit associated with a coding unit, the rectangular coding unit being coded by quantization, a block of a luminance component and a block of a chrominance component being coded in the rectangular coding unit, the rectangular coding unit having a width and a height, the width and the height being different from each other, and the rectangular coding unit being neither a prediction unit nor a transform unit; (b) determining that the pixel intensity information indicates that the pixel is associated with a vertical or horizontal boundary of the rectangular coding unit; (c) applying deblocking filtering to the rectangular coding unit based at least in part on the intensity information associated with the rectangular coding unit, the applied deblocking filtering being based on filtering parameters β and tc that specify boundary filtering selectively modified based on an offset of a quantization parameter, the offset being based at least in part on the determined intensity information of pixels associated with the boundary; (d) a larger offset results in stronger filtering compared to not having the larger offset; (e) the offsets include the offsets of three or more quantization parameters that are different from each other; (f) the deblocking filtering is further based on clipping; (g) the clipping is based on the quantization parameter.
Citation Information
Patent Citations
Intra PCM (IPCM) and Lossless Coding Mode Video Deblocking
JP2014531169A
Method and apparatus for encoding / decoding image information
JP2017163618A
Encoding device, decoding device, encoding method, and decoding method
WO2018097299A1