Method and apparatus for video encoding or decoding
Patent Information
- Application Number
- JP2024114077
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-07-26
- Filing Date
- 2024-07-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2039-12-11
AI Technical Summary
Current video encoding technologies, such as VTM3.0, face inefficiencies in decoding complexity and coding efficiency due to the use of a 6-tap interpolation filter for all luma samples in 4:2:0 YUV format, large block size calculations, and separate methods for deriving linear model parameters in cross-component linear models (CCLM) and local illumination correction (LIC), which increase decoder complexity without clear benefits.
Adapt the interpolation filter taps based on the YUV format of the video sequence, reduce the lookup table size by non-uniform quantization of difference values, and optimize parameter derivation methods for CCLM and LIC to improve coding efficiency and reduce decoder complexity.
The proposed methods enhance coding efficiency and reduce decoder complexity by tailoring interpolation filters and parameter derivation techniques to specific YUV formats, leading to improved video encoding and decoding performance.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 62 / 781,316, filed in the U.S. Patent and Trademark Office on December 18, 2018, U.S. Provisional Patent Application No. 62 / 785,678, filed on December 27, 2018, U.S. Provisional Patent Application No. 62 / 788,729, filed on January 4, 2019, U.S. Provisional Patent Application No. 62 / 789,992, filed on January 8, 2019, and U.S. Provisional Patent Application No. 16 / 523,258, filed on July 26, 2019, the disclosures of which are incorporated herein by reference in their entireties.
[0002] Methods and apparatus consistent with embodiments relate to video processing, and more specifically, to encoding or decoding video sequences with a focus on simplifying cross-component linear model prediction modes. [Background technology]
[0003] Recently, the Video Coding Experts Group (VCEG) of the ITU Telecommunication Standardization Sector (ITU-T), a sector of the International Telecommunication Union (ITU), and the ISO / IEC MPEG (JTC 1 / SC 29 / WG 11), a standardization subcommittee of the Joint Technical Committee ISO / IEC JTC 1 of the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC), published the H.265 / High Efficiency Video Coding (HEVC) standard (version 1) in 2013. The standard was updated to version 2 in 2014, version 3 in 2015, and version 4 in 2016.
[0004] In October 2017, they announced a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, 22 CfP responses had been submitted for Standard Dynamic Range (SDR), 12 for High Dynamic Range (HDR), and 12 for the 360 Video category. In April 2018, all received CfP responses were evaluated at the 122 MPEG / 10th JVET meeting. As a result of this meeting, JVET formally launched the standardization process for next-generation video coding beyond HEVC. The new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Expert Team.
[0005] Next, the intra prediction modes of the luma component in HEVC will be described. The intra prediction modes used in HEVC are shown in Figure 1. In HEVC, there may be a total of 35 intra prediction modes, of which mode 10 may be a horizontal mode, mode 26 may be a vertical mode, and modes 2, 18 and 34 may be diagonal modes. The intra prediction modes may be signaled by three most probable modes (MPMs) and 32 remaining modes.
[0006] Next, the intra prediction modes of the luma component in VVC will be described. In the current development of VVC, there may be a total of 95 intra prediction modes, as shown in FIG. 2, where mode 18 may be a horizontal mode, mode 50 may be a vertical mode, and modes 2, 34 and 66 may be diagonal modes. Modes 1 to 14 and modes 67 to 80 may be referred to as wide-angle intra prediction (WAIP) modes. As shown in FIG. 2, 35 intra prediction modes may be used in HEVC.
[0007] Next, the intra prediction mode of the luma component of VVC will be described. In the current development of VVC, there may be a total of 95 intra prediction modes, as shown in Figure 2, where mode 18 may be a horizontal mode, mode 50 may be a vertical mode, and modes 2, 34 and 66 may be diagonal modes. Modes 1 to 14 and modes 67 to 80 may be called wide-angle intra prediction (WAIP) modes.
[0008] Next, the Intra-mode mode of the chroma components of VVC will be described. In VTM, for the chroma components of the intra-PU, the encoder may select the optimal chroma prediction mode from eight modes, including Planar, DC, horizontal, vertical, direct duplication of the intra prediction mode (DM) from the luma components, left and upper cross-component linear mode (LT_CCLM), left cross-component linear mode (L_CCLM), and upper cross-component linear mode (T_CCLM). LT_CCLM, L_CCLM, and T_CCLM may be classified into a group of cross-component linear modes (CCLMs). The difference between these three modes is that different regions of neighboring samples may be used to derive the parameters α and β. For LT_CCLM, both the left and upper neighboring samples may be used to derive the parameters α and β. For L_CCLM, in general, only the left neighboring samples may be used to derive the parameters α and β. For T_CCLM, in general, only the upper neighboring samples may be used to derive the parameters α and β.
[0009] To reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode may be used, in which chroma samples are predicted based on reconstructed luma samples of the same CU using a linear model as follows: pred c (i,j)=α·rec L '(i,j)+β In the formula, pred c (i,j) represents the predicted chroma sample in a CU, and rec L(i,j) represent the downsampled reconstructed luma samples of the same CU. The parameters α and β may be derived by a linear equation, also called the max-min method. Because this computation process may be performed as part of the decoding process, not just as an encoder search operation, syntax is not necessarily required to communicate the α and β values.
[0010] There are various YUV formats, which are shown in Figures 3(A) to 3(D). In the 4:2:0 format, the LM prediction may apply a 6-tap interpolation filter to obtain downsampled luma samples corresponding to the chroma samples, as shown in Figures 3(A) to 3(D). In a formal way, the downsampled luma samples Rec'L[x,y] may be calculated from the reconstructed luma samples as follows: Rec' L [x,y]=(2×Rec L [2x,2y]+2×Rec L [2x,2y+1]+ Rec L [2x-1,2y]+Rec L [2x+1,2y]+ Rec L [2x-1,2y+1]+Rec L [2x+1,2y+1]+4)>>3
[0011] The downsampled luma samples may be used to find the maximum and minimum sample points. Two points (luma and chroma pairs) (A, B) may be the minimum and maximum values within a set of adjacent luma samples, as shown in FIG. The linear model parameters α and β may be obtained according to the following formula:
[0012]
number
[0013] Here, divisions may be avoided and replaced with multiplications and shifts. One look-up table (LUT) may be used to store pre-calculated values, and the absolute difference value between the maximum and minimum luma samples may be used to specify the entry index of the LUT, and the size of the LUT may be 512.
[0014] It was also proposed that the absolute difference between the maximum and minimum luma sample values in a specified region of adjacent samples, represented as diff_Y, may be non-uniformly quantized, and the quantized value of the absolute difference may be used to specify an entry index in a CCLM look-up table (LUT) such that the size of the LUT is reduced. The range of diff_Y may be divided into intervals, and different quantization step sizes may be used in the different intervals. In one example, the range of diff_Y may be divided into two intervals, and if diff_Y is less than or equal to a threshold value named Thres_1, one step size named Step_A may be used; otherwise, another step size named Step_B may be used. Thus, the parameters of the CCLM may be obtained as follows:
[0015]
number
[0016] Here, Thres_1, Step_A and Step_B can be any positive integers such as 1, 2, 3, 4, etc. Also, Step_A and Step_B are not equal.
[0017] To derive the chroma predictor, for the current VTM implementation, the multiplications are replaced with integer arithmetic as follows, where maxY, minY, maxC, and minC denote the maximum luma sample value, minimum luma sample value, maximum chroma sample value, and minimum chroma sample value, respectively. numSampL and numSampT denote the number of left and upper neighboring samples available, respectively. The following text is from VVC Draft 3, Section 8.2.4.2.8:
[0018] The variables a, b and k are derived as follows: If numSampL is equal to 0 and numSampT is equal to 0, the following applies: k=0 a=0 b=1<<(BitDepthC-1) Otherwise the following applies: shift=(BitDepthC>8)?BitDepthC-9:0 add=shift?1<<(shift-1):0 diff=(maxY-minY+add)>>shift k=16 If diff is greater than 0, the following applies: div=((maxC-minC)*(Floor(2 32 / diff)-Floor(2 16 / diff)*2 16 )+2 15 )>>16 a=((maxC-minC)*Floor(2 16 / diff)+div+add)>>shift Otherwise the following applies: a=0 b=minC-((a*minY)>>k)
[0019] The formula for diff greater than 0 can also be simplified to: a=((maxC-minC)*Floor(2 16 / diff)+add)>>shift
[0020] After deriving the parameters a and b, the chroma predictor is computed as follows: pred c (i,j)=(a·rec L '(i,j))>>S+b
[0021] The range of the variable "diff" is 1 to 512, so Floor(2 16The value of diff ( / diff) can be pre-calculated and stored in a look-up table (LUT) of size equal to 512. Furthermore, the value of diff is used to specify an entry index in the look-up table.
[0022] Because this computation process is performed as part of the decoding process and not just as an encoder search operation, no syntax is used to communicate the α and β values.
[0023] In T_CCLM mode, only the upper adjacent samples (including 2*W samples) are used to calculate the linear model coefficients. In L_CCLM mode, only the left adjacent samples (including 2*H samples) are used to calculate the linear model coefficients. This is shown in Figures 6A to 7B.
[0024] The CCLM prediction mode also includes prediction between two chroma components, i.e., the Cr component may be predicted from the Cb component. Instead of using the reconstructed sample signal, CCLM Cb-to-Cr prediction may be applied in the residual domain. This may be implemented by adding a weighted reconstructed Cb residual to the original Cr intra prediction to form the final Cr prediction:
[0025]
number
[0026] The CCLM luma-to-chroma prediction mode may be added as one additional chroma intra prediction mode. On the encoder side, one more RD cost check of the chroma components may be added to select the chroma intra prediction mode. If an intra prediction mode other than the CCLM luma-to-chroma prediction mode is used for the CU's chroma components, the CCLM Cb-to-Cr prediction may be used for the Cr component prediction.
[0027] Multi-model CCLM (MMLM) is another extension of CCLM. As the name suggests, there may be multiple models in MMLM, such as two models. In MMLM, the neighboring luma and chroma samples of the current block may be classified into two groups, and each group may be used as a training set to derive a linear model (i.e., a specific α and β are derived for a specific group). Furthermore, the samples of the current luma block may be classified based on the same rules as the classification of the neighboring luma samples.
[0028] 8 shows an example of classifying neighboring samples into two groups, where the threshold may be calculated as the average value of neighboring reconstructed luma samples. Neighboring samples with Rec'L[x,y]<=threshold may be classified into group 1, while neighboring samples with Rec'L[x,y]>threshold may be classified into group 2.
[0029]
number
[0030] Local Irradiance Correction (LIC) is based on a linear model of illumination changes with a scale factor a and an offset b. It is adaptively enabled or disabled for each coding unit (CU) coded in inter mode.
[0031] When applying LIC to a CU, a least square error method is used to derive parameters a and b using neighboring samples of the current CU and its corresponding reference samples. More specifically, as shown in FIG. 14, subsampled (2:1 subsampled) neighboring samples of the current CU 1410 (shown in part (a) of FIG. 14) and corresponding samples (identified by the motion information of the current CU or sub-CU) of a reference picture or block 1420 (shown in part (b) of FIG. 14) are used. IC parameters are derived and applied separately for each prediction direction.
[0032] If the CU is coded in merge mode, the LIC flag is copied from the neighboring blocks in a manner similar to the copying of motion information in merge mode, otherwise the LIC flag is signaled to the CU to indicate whether LIC is applied or not.
[0033] Despite the above progress, problems exist in the state of the art. Currently, in VTM3.0, in 4:2:0 YUV format, a 6-tap interpolation filter may be applied to all luma samples in a given contiguous region, but ultimately only the maximum and minimum two luma samples are used to derive the CCLM parameters, which increases the decoder complexity without any clear advantage in coding efficiency. Furthermore, in VTM3.0, the CCLM luma sample interpolation filter only supports 4:2:0 YUV format, while 4:4:4 and 4:2:2 YUV formats are still common and need to be supported.
[0034] Currently in VTM3.0, the size of the lookup table (LUT) of CCLM is 512, that is, the lookup table has 512 elements, and each element is represented by a 16-bit integer, which increases the decoder memory cost too much without any clear advantage in coding efficiency. Moreover, although the concepts of CCLM and LIC are similar, they use different methods to derive the parameters a and b of the linear model, which is undesirable.
[0035] Currently in VTM3.0, for large blocks such as 32x32 chroma blocks, the max and min are calculated using 64 samples of a given neighboring region, which increases the decoder complexity without any clear advantage in coding efficiency.
[0036] Currently in VTM3.0, for large blocks such as 64x64 blocks, the linear model parameters a and b are calculated using 64 samples of a given neighboring region, which increases the decoder complexity without any clear advantage in coding efficiency. Summary of the Invention [Means for solving the problem]
[0037] A method for encoding or decoding a video sequence may include applying a cross-component linear model (CCLM) to the video sequence and applying an interpolation filter in the cross-component linear model (CCLM), where the interpolation filter may depend on a YUV format of the video sequence.
[0038] According to one aspect of the present disclosure, in the above method, when applying the interpolation filter in the cross-component linear model (CCLM), the method further includes using taps of the interpolation filter that depend on the YUV format of the video sequence.
[0039] According to this aspect of the disclosure, the method further includes using taps of the interpolation filter used in the cross-component linear model (CCLM) that are in the same format as the YUV format of the video sequence.
[0040] According to one aspect of the present disclosure, in the above method, when applying an interpolation filter in a cross-component linear model (CCLM), the method may further include a step of setting the format of the interpolation filter to be the same as the format of the video sequence if the video sequence includes a 4:4:4: or 4:2:2 YUV format, and may include a step of setting the format of the interpolation filter to be different from the format of the video sequence if the video sequence includes a 4:2:0 YUV format.
[0041] According to one aspect of the present disclosure, in the above method, when applying the interpolation filter in the cross-component linear model (CCLM), the method may further include using different taps of the interpolation filter for various YUV formats of the video sequence.
[0042] According to one aspect of the present disclosure, in the above method, when applying the interpolation filter in the cross-component linear model (CCLM), the method may further include setting the interpolation filter differently for the reconstructed samples of the upper and left neighboring luma.
[0043] According to this aspect of the disclosure, the method may further include setting an interpolation filter, the interpolation filter being applied such that the upper and left neighboring luma reconstruction samples depend on the YUV format of the video sequence.
[0044] According to one aspect of the present disclosure, the method may further include setting a number of rows of upper neighboring luma samples and a number of columns of left neighboring luma samples used in the cross-component linear model (CCLM) to be dependent on the YUV format of the video sequence.
[0045] According to this aspect of the disclosure, the method may further include using one row of the upper adjacent region and / or one column of the left adjacent region in a cross-component linear model (CCLM) for a video sequence having one of 4:4:4 or 4:2:2 YUV formats.
[0046] According to this aspect of the disclosure, the method may further include using one row of the upper adjacent region and / or at least two columns of the left adjacent region in a cross-component linear model (CCLM) for a video sequence having a 4:2:2 YUV format.
[0047] According to one aspect of the present disclosure, a device for encoding or decoding a video sequence may include at least one memory configured to store program code and at least one processor configured to read the program code and operate as instructed by the program code, where the program code may include first encoding or decoding code configured to cause the at least one processor to apply a cross-component linear model (CCLM) to the video sequence and to apply an interpolation filter in the cross-component linear model (CCLM), where the interpolation filter may be dependent on a YUV format of the video sequence.
[0048] According to one aspect of the present disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to use taps of an interpolation filter applied in a cross-component linear model (CCLM), the taps being dependent on the YUV format of the video sequence.
[0049] According to this aspect of the disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to use taps of an interpolation filter applied in a cross-component linear model (CCLM) that is in the same format as the YUV format of the video sequence.
[0050] According to one aspect of the present disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to set a format of an interpolation filter applied in a cross-component linear model (CCLM) to be the same as a format of the video sequence if the video sequence includes a 4:4:4: or 4:2:2 YUV format, and to set the format of the interpolation filter to be different from the format of the video sequence if the video sequence includes a 4:2:0 YUV format.
[0051] According to one aspect of the present disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to use taps of an interpolation filter applied in a cross-component linear model (CCLM) that are different for various YUV formats of the video sequence.
[0052] According to one aspect of the present disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to set an interpolation filter applied in a cross-component linear model (CCLM) differently for reconstructed samples of upper and left neighboring luma.
[0053] According to this aspect of the disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to set interpolation filters applied to the reconstructed samples of the upper and left neighboring luma to be dependent on a YUV format of the video sequence.
[0054] According to one aspect of the disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to set a number of rows of upper neighboring luma samples and a number of columns of left neighboring luma samples used in the cross-component linear model (CCLM) to be dependent on a YUV format of the video sequence.
[0055] According to this aspect of the disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to use one row of the upper adjacent region and / or at least one column of the left adjacent region in a cross-component linear model (CCLM) for a video sequence having one of 4:4:4 or 4:2:2 YUV formats.
[0056] According to one aspect of the disclosure, a non-transitory computer-readable medium may be provided that stores program code, the program code including one or more instructions that, when executed by one or more processors of a device, may cause the one or more processors to apply a cross-component linear model (CCLM) to a video sequence and to apply an interpolation filter in the cross-component linear model (CCLM), where the interpolation filter is dependent on a YUV format of the video sequence.
[0057] Although the foregoing methods, devices, and non-transitory computer-readable medium have been described separately, this description is not intended to suggest any limitation as to the scope of their use or functionality, and in fact these methods, devices, and non-transitory computer-readable medium may be combined in other aspects of the present disclosure.
[0058] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]
[0059] [Figure 1] FIG. 2 is a diagram of a predictive model according to one embodiment. [Diagram 2] FIG. 2 is a diagram of a predictive model according to one embodiment. [Diagram 3] FIG. 2 is a diagram of a YUV format according to one embodiment. [Figure 4] FIG. 2 is a diagram of different luma values according to one embodiment. [Diagram 5] FIG. 1 is a sample diagram used in cross-component linear modeling according to one embodiment. [Figure 6] FIG. 1 is a sample diagram used in cross-component linear modeling according to one embodiment. [Figure 7] FIG. 1 is a sample diagram used in cross-component linear modeling according to one embodiment. [Figure 8]1 is an example of classification using multi-model CCLM, according to one embodiment. [Figure 9] FIG. 1 is a simplified block diagram of a communication system according to one embodiment. [Figure 10] FIG. 1 is a diagram of a streaming environment according to one embodiment. [Figure 11] FIG. 2 is a block diagram of a video decoder according to one embodiment. [Figure 12] FIG. 2 is a block diagram of a video encoder according to one embodiment. [Figure 13] 1 is a flowchart of an exemplary process for encoding or decoding a video sequence, according to one embodiment. [Figure 14] FIG. 1 is a diagram of adjacent samples used to derive illumination compensation (IC) parameters. [Figure 15] 1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Figure 16] 1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Figure 17] 1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Figure 18] 1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Figure 19] 1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Figure 20]1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Figure 21] 1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Figure 22] 1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Figure 23] 1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Figure 24] 1 is a diagram of a current coding unit (CU) and a subset of neighboring reconstructed samples of the current CU used to calculate maximum and minimum sample values, according to an embodiment. [Diagram 25] 4 is a diagram of a chroma block and selected neighboring samples of the chroma block used to calculate maximum and minimum sample values according to an embodiment. [Figure 26] 4 is a diagram of a chroma block and selected neighboring samples of the chroma block used to calculate maximum and minimum sample values according to an embodiment. [Figure 27] FIG. 13 is a diagram of filter positions for luma samples according to an embodiment. [Figure 28] FIG. 13 is a diagram of filter positions for luma samples according to an embodiment. [Figure 29] FIG. 13 is a diagram of a current CU and adjacent sample pairs of the current CU used to calculate parameters in a linear model prediction mode, according to an embodiment. [Diagram 30]FIG. 13 is a diagram of a current CU and adjacent sample pairs of the current CU used to calculate parameters in a linear model prediction mode, according to an embodiment. [Diagram 31] 4 is a diagram of a block and selected neighboring samples of the block used to calculate linear model parameters, according to an embodiment. [Diagram 32] 4 is a diagram of a block and selected neighboring samples of the block used to calculate linear model parameters, according to an embodiment. [Diagram 33] 4 is a diagram of a block and selected neighboring samples of the block used to calculate linear model parameters, according to an embodiment. [Diagram 34] 4 is a diagram of a block and selected neighboring samples of the block used to calculate linear model parameters, according to an embodiment. [Diagram 35] 4 is a diagram of a block and selected neighboring samples of the block used to calculate linear model parameters, according to an embodiment. [Diagram 36] 4 is a diagram of a block and selected neighboring samples of the block used to calculate linear model parameters, according to an embodiment. [Figure 37] FIG. 1 is a diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0060] FIG. 9 shows a simplified block diagram of a communication system (400) according to one embodiment of the present disclosure. The communication system (400) may include at least two terminals (410-420) interconnected via a network (450). In the case of one-way data transmission, a first terminal (410) may encode video data at a local location for transmission to another terminal (420) via the network (450). The second terminal (420) may receive the encoded video data of the other terminal from the network (450), decode the encoded data, and display the restored video data. One-way data transmission may be common in media serving applications, etc.
[0061] 9 illustrates a second pair of terminals (430, 440) provided to support two-way transmission of encoded video, such as might occur during a video conference. For two-way transmission of data, each terminal (430, 440) may encode video data captured at a local location for transmission over the network (450) to the other terminal. Each terminal (430, 440) may also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0062] In FIG. 9, the terminals (410-440) may be depicted as servers, personal computers, and smartphones, although the principles of the present disclosure are not so limited. Embodiments of the present disclosure find application with laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (450) represents any number of networks that convey encoded video data between the terminals (410-440), including, for example, wired and / or wireless communication networks. The communication network (450) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of the network (450) may not be important to the operation of the present disclosure unless otherwise described herein below.
[0063] 10 shows the arrangement of video encoders and decoders in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter can be equally applied to other video-enabled applications including, for example, video conferencing, digital television, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0064] The streaming system may include a video source (501) and a capture subsystem (513) that may include, for example, a digital camera that creates an uncompressed video sample stream (502). The sample stream (502), shown in bold to emphasize a larger amount of data when compared to an encoded video bitstream, may be processed by an encoder (503) coupled to the camera (501). The encoder (503) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream (504), shown as a thin line to emphasize a smaller amount of data compared to the sample stream, may be stored on a streaming server (505) for future use. One or more streaming clients (506, 508) may access the streaming server (505) to obtain copies (507, 509) of the encoded video bitstream (504). The client (506) may include a video decoder (510) that decodes an incoming copy of an encoded video bitstream (507) and creates an outgoing video sample stream (511) that can be rendered on a display (512) or other rendering device (not shown). In some streaming systems, the video bitstreams (504, 507, 509) may be encoded according to a particular video encoding / compression standard. Examples of these standards include H.265 HEVC. The developing video encoding standard is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0065] FIG. 11 may be a functional block diagram of a video decoder (510) according to an embodiment of the present invention.
[0066] The receiver (610) may receive one or more codec video sequences to be decoded by the decoder (610), in the same or another embodiment, one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences may be received from a channel (612), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (610) may receive the coded video data together with other data, e.g., coded audio data and / or auxiliary data streams, that may be forwarded to a respective using entity (not shown). The receiver (610) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (615) may be coupled between the receiver (610) and the entropy decoder / parser (620) (hereinafter, the "parser"). When the receiver (610) is receiving data from a store / forward device of sufficient bandwidth and controllability, or from an isosynchronous network, the buffer (615) may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer (615) may be needed and may be relatively large, and may advantageously be of adaptive size.
[0067] The video decoder (510) may include a parser (620) for reconstructing symbols (621) from the entropy-coded video sequence. These categories of symbols may include information used to manage the operation of the decoder (510) and information for controlling a rendering device, such as a display (512), that is not an integral part of the decoder but may be coupled to the decoder, as shown in FIG. 11. The control information for the rendering device(s) may be in the form of a Supplementary Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (620) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may follow any video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context-dependent coding, etc. The parser (620) may extract from the encoded video sequence at least one set of subgroup parameters for a subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser may also extract from the encoded video sequence information such as transform coefficients, quantization parameter (QP) values, motion vectors, etc.
[0068] The parser (620) may perform entropy decoding / parsing operations on the video sequence received from the buffer (615) to create symbols (621). The parser (620) may receive the encoded data and selectively decode particular symbols (621). Additionally, the parser (620) may determine whether a particular symbol (621) should be provided to a motion compensated prediction unit (653), a scaler / inverse transform unit (651), an intra prediction unit (652), or a loop filter unit (658).
[0069] The reconstruction of the symbols (621) can involve a number of different units, depending on the type of encoded video picture or part thereof (inter-picture and intra-picture, inter-block and intra-block, etc.), and other factors. Which units are involved and how can be controlled by subgroup control information parsed from the encoded video sequence by the parser (620). The flow of such subgroup control information between the parser (620) and the following units is not depicted for clarity.
[0070] Beyond the functional blocks already mentioned, the decoder (510) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:
[0071] The first unit is a scalar / inverse transform unit (651), which receives quantized transform coefficients as well as control information from the parser (620) including the transform used, block size, quantization factor, quantization scaling matrix, etc. as symbols (621). The scalar / inverse transform unit can output blocks containing sample values that can be input to an aggregator (655).
[0072] In some cases, the output samples of the scaler / inverse transform (651) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (652). In some cases, the intra-picture prediction unit (652) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current (partially reconstructed) picture (656). The aggregator (655) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (652) to the output sample information provided by the scaler / inverse transform unit (651).
[0073] In other cases, the output samples of the scalar / inverse transform unit (651) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensation prediction unit (653) may access the reference picture memory (657) to fetch samples used for prediction. After motion compensation of the fetched samples according to the symbols (621) associated with the block, these samples may be added to the output of the scalar / inverse transform unit by the aggregator (655) to generate output sample information (in this case referred to as residual samples or residual signals). The addresses in the reference picture memory form from which the motion compensation unit fetches the prediction samples may be controlled by a motion vector and are available to the motion compensation unit in the form of symbols (621) that may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.
[0074] The output samples of the aggregator (655) may be subject to various loop filtering techniques in a loop filter unit (658). Video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and made available to the loop filter unit (658) as symbols (621) from the parser (620), but may also be responsive to meta-information obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and may also be responsive to previously reconstructed and loop filtered sample values.
[0075] The output of the loop filter unit (658) may be a sample stream that can be output to the rendering device (512) as well as stored in a reference picture memory (656) for use in future inter-picture prediction.
[0076] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (620)), the current reference picture (656) can become part of the reference picture buffer (657), and a new current picture memory can be reallocated before beginning reconstruction of the next coded picture.
[0077] The video decoder (510) may perform decoding operations according to a given video compression technique, which may be documented in a standard, such as H.265 HEVC. The encoded video sequence may be compliant with the syntax specified by the video compression technique or standard being used, in the sense that it conforms to the syntax of the video compression technique or standard as specified in the video compression technique's document or standard, particularly the profile documents therein. Also required for compliance is that the complexity of the encoded video sequence is within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may in some cases be further limited by a hypothetical reference decoder (HRD) specification and HRD buffer management metadata conveyed in the encoded video sequence.
[0078] In one embodiment, the receiver (610) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0079] FIG. 12 may be a functional block diagram of a video encoder (503) according to one embodiment of the present disclosure.
[0080] The encoder (503) may receive video samples from a video source (501) (not part of the encoder) that may capture video images to be encoded by the encoder (503).
[0081] The video source (501) may provide a source video sequence to be encoded by the encoder (503) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (501) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (503) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of separate pictures that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art may easily understand the relationship between pixels and samples. The following description focuses on samples.
[0082] According to one embodiment, the encoder (503) may encode and compress pictures of a source video sequence into an encoded video sequence (743) in real time or under any other time constraint required by the application. Enforcing an appropriate encoding rate is one function of the controller (750). The controller (750) controls and is operatively coupled to other functional units as described below. For clarity, couplings are not depicted. Parameters set by the controller may include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art may readily identify other functions of the controller (750) as they may relate to a video encoder (503) optimized for a particular system design.
[0083] Some video encoders operate in what those skilled in the art will readily recognize as a "coding loop." As an oversimplified explanation, the coding loop can consist of a coding part of the encoder (730) (hereafter the "source coder") (responsible for creating symbols based on the input picture to be coded and the reference pictures), and a (local) decoder (733) embedded in the encoder (503) that reconstructs the symbols to create sample data that the (remote) decoder also creates (since the compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). That reconstructed sample stream is input to a reference picture memory (734). The decoding of the symbol stream results in bit-exact results regardless of the location of the decoder (local or remote), so the contents of the reference picture buffer are also bit-exact between the local and remote encoders. In other words, the predictive part of the encoder "sees" exactly the same sample values as the decoder "sees" when using prediction during decoding as reference picture samples. This basic principle of reference picture synchrony (and the drift that occurs when synchrony cannot be maintained, for example due to channel errors) is well known to those skilled in the art.
[0084] The operation of the "local" decoder (733) may be the same as the operation of the "remote" decoder (510), which is described in detail above in connection with Figure 10. However, with brief reference also to Figure 10, symbols may be available, and the encoding / decoding of the symbols into an encoded video sequence by the entropy coder (745) and parser (620) may be lossless, and the entropy decoding portion of the decoder (510), including the channel (612), receiver (610), buffer (615) and parser (620), may not be fully implemented in the local decoder (733).
[0085] An observation that can be made at this point is that any decoder techniques other than analysis / entropy decoding present in the decoder must necessarily be present in the corresponding encoder in substantially identical functional form. A description of the encoder techniques can be omitted, since they are the inverse of the decoder techniques described generically. Only in certain areas is a more detailed description necessary, which is provided below.
[0086] As part of its operation, the source coder (730) may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as “reference frames.” In this manner, the coding engine (732) codes differences between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as predictive references for the input frame.
[0087] The local video decoder (733) may decode the encoded video data of frames that may be designated as reference frames based on symbols created by the source coder (730). The operation of the encoding engine (732) may advantageously be a lossy process. If the encoded video data can be decoded in a video decoder (not shown in FIG. 11), the reconstructed video sequence may be a replica of the source video sequence, usually with some errors. The local video decoder (733) may replicate the decoding process that may be performed by the video decoder on the reference frames and cause the reconstructed reference frames to be stored in a reference picture cache (734). In this way, the encoder (503) may locally store replicas of reconstructed reference frames that have common content as reconstructed reference frames obtained by the far-end video decoder (without transmission errors).
[0088] The predictor (735) may perform the prediction search of the coding engine (732). That is, for a new frame to be coded, the predictor (735) may search the reference picture memory (734) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that may serve as suitable prediction references for the new picture. The predictor (735) may operate on one sample block per pixel block to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (735), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (734).
[0089] The controller (750) may manage the encoding operations of the video coder (730), including, for example, setting parameters and subgroup parameters used to encode the video data.
[0090] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (745), which converts the symbols produced by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0091] The transmitter (740) may buffer the encoded video sequence created by the entropy coder (745) and prepare it for transmission over a communication channel (760), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (740) may merge the encoded video data from the video coder (730) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0092] A controller (750) may manage the operation of the encoder (503). During encoding, the controller (750) may assign a particular encoded picture type to each encoded picture, which may affect the encoding technique that may be applied to each picture. For example, pictures may often be assigned as any of the following frame types:
[0093] An intra picture (I-picture) may be one that can be coded and decoded without using other frames in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, such as, for example, Independent Decoder Refresh Picture. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.
[0094] A predictive picture (P picture) may be one that can be encoded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict sample values for each block.
[0095] Bidirectionally predicted pictures (B-pictures) may be those that can be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0096] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and coded block by block. A block may be predictively coded with reference to other (already coded) blocks, as determined by a coding assignment applied to the respective picture of the block. For example, a block of an I picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). A pixel block of a P picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. A block of a B picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0097] The video coder (503) may perform encoding operations according to a given video encoding technique or standard, such as H.265 HEVC. In its operations, the video coder (503) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video encoding technique or standard being used.
[0098] In one embodiment, the transmitter (740) may transmit additional data along with the encoded video. The video coder (730) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplemental enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.
[0099] This disclosure is directed to various methods of cross-component linear model prediction modes.
[0100] As mentioned above, in this application, CCLM may refer to any variant of the cross-component linear model, such as LT_CCLM, T_CCLM, L_CCLM, MMLM, LT_MMLM, L_MMLM, T_MMLM, etc. Furthermore, in this application, a smoothing filter may be defined as a linear filter with an odd number of taps, whose filter coefficients may be symmetric. For example, a 3-tap filter [1, 2, 1] and a 5-tap filter [1, 3, 8, 3, 1] are both smoothing filters. An interpolation filter may be defined as a linear filter that uses integer-positioned samples to generate fractional-positioned samples.
[0101] It is proposed that the interpolation filter used for CCLM may be applied to a subset of neighboring luma reconstructed samples within a specified region.
[0102] In one embodiment, the specified adjacent sample region may be the left adjacent sample region used for L_CCLM and L_MMLM, or the upper adjacent sample region used for T_CCLM and T_MMLM, or the upper and left adjacent sample regions used for LT_CCLM and LT_MMLM.
[0103] In another aspect, the reconstructed luma samples in the specified region may be used directly to find the maximum and minimum values. After obtaining the maximum and minimum luma sample values, an interpolation (or smoothing) filter may be applied (e.g., simply) to these two samples.
[0104] In another aspect, after obtaining the maximum and minimum luma sample values, the co-located positions of the minimum and maximum luma samples within the chroma adjacent sample region may be recorded as (x_min, y_min) and (x_max, y_max), respectively, and a downsample filter may be applied to the co-located luma samples of the chroma samples having positions (x_min, y_min) and (x_max, y_max).
[0105] In another embodiment, for 4:2:0 YUV format, the downsample filter may be applied as follows: (x,y) in the following equation may be replaced with (x_min,y_min) or (x_max,y_max):
[0106]
number
[0107] denotes the downsampled luma sample value,
[0108]
number
[0109] denotes the luma sample constructed at a specified neighboring position (k·x, k·y), where k can be a positive integer such as 1, 2, 3, or 4.
[0110]
number
[0111] In another aspect, for a 4:2:0 YUV format, the same location in the luma sample domain for one chroma sample having location (x,y) can be any of the following six locations: (k*x,k*y), (k*x-1,k*y), (k*x+1,k*y), (k*x-1,k*y+1), (k*x,k*y+1) and (k*x+1,k*y+1), where k is a positive integer such as 1, 2, 3 or 4.
[0112] Alternatively, for 4:4:4 and 4:2:2 YUV formats, the above mentioned interpolation filters may be used here.
[0113] In another aspect, the reconstructed luma samples in the specified region may be used directly to find the maximum and minimum values.
[0114] In another aspect, for LT_CCLM mode, to find the maximum and minimum values, the neighboring luma samples of the first upper reference row (shown in Figures 6(A)-6(B)) and the second left reference column (shown in Figures 7A-7B) may be directly used.
[0115] In another aspect, for LT_CCLM mode, half of the adjacent luma samples in the first and second upper reference rows (shown in Figures 6(A)-6(B)) and the samples in the second left reference column (shown in Figures 7(A)-7(B)) may be used directly to find the maximum and minimum values.
[0116] In another aspect, for T_CCLM mode, the neighboring luma samples of the first upper reference row may be used directly to find the maximum and minimum values.
[0117] In another aspect, for T_CCLM mode, half of the neighboring luma samples in the first and second upper reference rows may be used directly to find the maximum and minimum values.
[0118] In another aspect, for the L_CCLM mode, the neighboring luma samples of the second left reference column may be used directly to find the maximum and minimum values.
[0119] In another aspect, if an interpolation filter is used for some (e.g., only some) of the neighboring luma samples, a different downsampling filter than that used in current CCLM may be applied, i.e., a 6-tap [1,2,1;1,2,1] / 8 filter, e.g., a 10-tap filter may be [1,4,6,4,1;1,4,6,4,1] / 16.
[0120] In another aspect, the downsampling filter may be an 8-tap filter, and the number of filter taps in the first (closest to the current block) reference row (as shown in FIGS. 6A-6B) may be different from the other (farther from the current block) reference columns (as shown in FIGS. 7A-7B). For example, the 8-tap filter may be [1,2,1;1,2,6,2,1] / 16, and the number of filter taps in the first reference row (or column) may be 5, while the number of taps in the second reference row (or column) may be 3.
[0121] In another aspect, it is proposed that the reconstructed chroma samples in a specified region may be used to find the maximum and minimum values.
[0122] In one aspect, after obtaining the maximum and minimum chroma sample values, an interpolation (or smoothing) filter may be applied to the co-located luma sample of these two chroma samples (e.g., only the luma sample). For example, in the case of 4:2:0 YUV format, the above-mentioned interpolation (or smoothing) filter may be used. As another example, in the case of 4:2:0 YUV format, the co-located position of one chroma sample having the above-mentioned position (x, y) may be used here. In yet another example, in the case of 4:4:4 and 4:2:2 YUV formats, the following described interpolation filter may be used here.
[0123] According to one aspect, it is proposed that the interpolation filter (or smoothing filter) used in the CCLM may depend on the YUV format, such as 4:4:4, 4:2:2, or 4:2:0 YUV format.
[0124] Figure 13 is a flow chart of an example process (800) for encoding or decoding a video sequence. In some implementations, one or more process blocks of Figure 13 may be performed by the decoder (510). In some implementations, one or more process blocks of Figure 13 may be performed by another device or group of devices separate from or including the decoder (510), such as the encoder (503).
[0125] As shown in FIG. 13, the process (800) may include encoding or decoding a video sequence (810).
[0126] As further shown in FIG. 13, the process (800) may further include applying a cross-component linear model (CCLM) to the video sequence (820).
[0127] As further shown in FIG. 13, the process (800) may further include applying an interpolation filter in a cross-component linear model (CCLM), where the interpolation filter depends on the YUV format of the video sequence (830).
[0128] 13 shows example blocks of process (800), in some implementations, process (800) may include additional, fewer, different, or differently arranged blocks than those shown in FIG 13. Additionally, or instead, two or more blocks of process (800) may be performed in parallel.
[0129] According to another aspect, the taps of the interpolation (or smoothing) filter used in the CCLM may depend on the YUV format.
[0130] In another aspect, the taps of the interpolation (or smoothing) filter used in the CCLM may be the same for the same YUV format, for example, the taps of the interpolation filter used in the CCLM may be the same as for the 4:2:0 YUV format.
[0131] In another aspect, the interpolation or smoothing filters for 4:4:4 and 4:2:2 YUV formats are the same but different from the filters for 4:2:0 YUV format. For example, a 3-tap (1,2,1) filter may be used for 4:4:4 and 4:2:2 YUV formats, but a 6-tap (1,2,1;1,2,1) filter may be used for 4:2:0 YUV format. As another example, a downsampling (or smoothing) filter may not be used for 4:4:4 and 4:2:2 YUV formats, but a 6-tap (1,2,1,1,2,1) or 3-tap (1,2,1) filter may be used for 4:2:0 YUV format.
[0132] According to another aspect, the taps of the interpolation (or smoothing) filter used in the CCLM may be different for various YUV formats. For example, a 5-tap (1,1,4,1,1) filter may be used for 4:4:4 YUV format, a (1,2,1) filter may be used for 4:2:2 YUV format, and a 6-tap (1,2,1,1,2,1) filter may be used for 4:2:0 YUV format.
[0133] In another aspect, the interpolation (or smoothing) filters may be different for the upper and left neighboring luma of the reconstructed sample.
[0134] According to another aspect, the interpolation (or smoothing) filters of the upper and left neighboring luma reconstruction samples may depend on the YUV format.
[0135] According to another aspect, for a 4:2:2 YUV format, the interpolation (or smoothing) filter may be different for the upper and left neighboring samples. For example, for a 4:2:2 YUV format, a (1,2,1) filter may be applied to the upper neighboring luma sample, and no interpolation (or smoothing) filter may be applied to the left neighboring luma sample.
[0136] According to another aspect, the number of rows of upper adjacent luma samples and the number of columns of left adjacent luma samples used in CCLM may depend on the YUV format. For example, in a 4:4:4 YUV or 4:2:2 YUV format, only one row of the upper adjacent region and / or one column of the left adjacent region may be used. As another example, one row of the upper adjacent region and / or three (or two) columns of the left adjacent region may be used for a 4:2:2 YUV format.
[0137] In another aspect, it is also proposed that the absolute difference between the maximum and minimum luma sample values in a specified adjacent sample region, represented by diff_Y, may be non-uniformly quantized and the quantized value of the absolute difference may be used to specify an entry index of a CCLM LUT such that the size of the LUT is reduced. The range of diff_Y may be divided into two intervals, and one step size named Step_A may be used if diff_Y is less than or equal to a threshold named Thres_1. Otherwise, another step size named Step_B may be used. Thus, the parameter a of the CCLM may be obtained as follows:
[0138]
number
[0139] Here, in this embodiment, Thres_1 may be set equal to 64, Step_A may be set equal to 1, and Step_B may be set equal to 8.
[0140] Next, we explain the modified CCLM parameter derivation process in addition to Section 8.2.4.2.8 of VVC Draft 3.
[0141] According to another aspect, in the immediately preceding formula, the variables a, b, and k may be derived as follows: -If numSampL is equal to 0 and numSampT is equal to 0, the following applies: k=0 a=0 b=1<<(BitDepth C -1) - Otherwise the following applies: Shift = (BitDepth C >8)?BitDepth C -9:0 add=shift?1<<(shift-1):0 diff=(maxY-minY+add)>>shift k=16 If -diff is greater than 0, the following applies: diff=(diff>64)?(56+(diff>>3)):diff a=((maxC-minC)*g_aiLMDivTableHigh[diff-1]+add)>>shift - Otherwise the following applies: a=0 b=minC-((a*minY)>>k)
[0142] Here, the predicted samples predSamples[x][y] with x=0..nTbW-1, y=0..nTbH-1 may be derived as follows: predSamples[x][y]=Clip1C(((pDsY[x][y]*a)>>k)+b)
[0143] Additionally, a simplified CCLM lookup table is shown below: int g_aiLMDivTableHigh[]={ 65536,32768,21845,16384,13107,10922,9362,8192,7281,6553,5957,5461,5041,4681,4369,4096, 3855,3640,3449,3276,3120,2978,2849,2730,2621,2520,2427,2340,2259,2184,2114,2048, 1985,1927,1872,1820,1771,1724,1680,1638,1598,1560,1524,1489,1456,1424,1394,1365, 1337,1310,1285,1260,1236,1213,1191,1170,1149,1129,1110,1092,1074,1057,1040,1024, 910,819,744,682,630,585,546,512,481,455,431,409,390,372,356,341, 327,315,303,292,282,273,264,256,248,240,234,227,221,215,210,204, 199,195,190,186,182,178,174,170,167,163,160,157,154,151,148,146, 143,141,138,136,134,132,130,128,};
[0144] Also in an embodiment, the absolute difference between the maximum and minimum luma sample values in a specified contiguous sample region, represented as diff_Y, is non-uniformly quantized, and the quantized absolute difference value is used to specify an entry index in a CCLM LUT such that the size of the LUT is reduced.
[0145] For example, the specified adjacent sample region may be the left adjacent sample region used for L_CCLM and L_MMLM, or the upper adjacent sample region used for T_CCLM and T_MMLM, or the upper and left adjacent sample regions used for TL_CCLM and TL_MMLM.
[0146] In another example, the range of diff_Y is divided into multiple intervals and different quantization step sizes are used in the different intervals. In a first embodiment, the range of diff_Y is divided into two intervals and one step size named Step_A is used if diff_Y is less than or equal to a threshold named Thres_1. Otherwise, another step size named Step_B may be used. Thus, the parameter a of the CCLM may be obtained as follows:
[0147]
number
[0148] Each of Thres_1, Step_A and Step_B can be any positive integer, such as 1, 2, 3, 4, etc. Step_A and Step_B are not equal. In an example, Thres_1 is set equal to 32, 48, or 64. In another example, Thres_1 is set equal to 64, Step_A is set equal to 1, and Step_B is set equal to 8. In another example, Thres_1 is set equal to 64, Step_A is set equal to 2 (or 1), and Step_B is set equal to 8 (or 4). In another example, the value of Step_A is less than the value of Step_B.
[0149] In a second embodiment, the range of diff_Y is divided into three intervals by two thresholds, namely Thres_1 and Thres_2. If diff_Y is less than or equal to Thres_1, one step size, namely Step_A, is used; otherwise, if diff_Y is less than or equal to Thres_2, another step size, namely Step_B, is used; otherwise, a third step size, namely Step_C, is used. Thres_1, Thres_2, Step_A, Step_B, and Step_C can be any positive integer, such as 1, 2, 3, 4, etc. Thres_1 is less than Thres_2. Step_A, Step_B, and Step_C are not equal. In an example, Thres_1 and Thres_2 are set equal to 64 and 256, respectively. In another example, Step_A, Step_B, and Step_C are set equal to 2 (or 1), 8 (or 4), and 32 (or 16), respectively.
[0150] In a third embodiment, the thresholds used to specify the intervals (such as Thres_1 and Thres_2) are powers of 2, for example, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024.
[0151] In a further embodiment, the LIC reuses the max-min method of CCLM to derive the parameters a and b.
[0152] For example, LIC reuses CCLM's lookup tables to avoid division operations.
[0153] In another example, LIC derives parameters a and b using the same adjacent reconstructed sample region as used for the luma component of TL_CCLM.
[0154] In another example, LIC reuses the same downsampling method in CCLM to filter neighboring reconstructed samples.
[0155] In a further embodiment, only a subset of neighboring reconstructed samples within a specified region is used to calculate maximum and minimum sample values for max-min method in CCLM prediction mode.
[0156] For example, only four samples of adjacent reconstructed sample regions are used to calculate the maximum and minimum sample values, regardless of the block size of the current block. In this example, the positions of the four selected samples are fixed and are shown in FIG. 15. Specifically, for the current CU 1510, only sample values of adjacent reconstructed samples at positions A, B, C, and D are used to calculate the maximum and minimum sample values.
[0157] In another example, for TL_CCLM mode, only the left (or upper) neighboring sample is used to calculate the minimum sample value, and only the upper (or left) neighboring sample is used to calculate the maximum sample value.
[0158] In another example, instead of finding the maximum and minimum sample values of the max-min method in CCLM prediction mode, the average values of the upper luma and chroma reference samples, i.e., top_mean_luma, top_mean_chroma, and the average values of the left luma and chroma reference samples, i.e., left_mean_luma, left_mean_chroma, may be used to derive the CCLM parameters.
[0159] In the following embodiment, with reference to FIG. 16, if the top left and bottom left samples of the current CU 1610 are used, these blocks may be replaced with their direct neighboring samples that are located outside the range of the upper and left samples of the current CU 1610. For example, in FIG. 16, samples A and B may be used.
[0160] In another example, in FIG. 17, samples A, B, and C may be used in the current CU 1710.
[0161] In the following embodiments, a chroma sample (in U or V component) together with its co-located luma sample is referred to as a sample pair, or in the case of LIC, adjacent samples of the current block and co-located adjacent samples of the reference block are referred to as a sample pair, such that the samples in the sample pair come from the same color component (luma, U or V).
[0162] In the following embodiments, for CCLM or any variant / extension of CCLM mode, the block size may be that of a chroma block. For LIC mode or any variant of LIC, the block size may be that of the current block of the current color component (luma, U, or V).
[0163] In an embodiment, up to N adjacent reconstructed sample pairs within a specified region are used to calculate the maximum and minimum sample values of the max-min method in CCLM prediction mode, where N is a positive integer such as 4, 8 or 16.
[0164] For example, for an Nx2 or 2xN chroma block, only one adjacent sample pair from the left and one adjacent sample pair from the top are used to calculate the maximum and minimum sample values. The positions of the adjacent sample pairs used are fixed. In one example, the positions of the chroma samples in the sample pairs are shown in Figures 18 and 19 for current CUs 1810 and 1910, respectively. In another example, the positions of the chroma samples in the sample pairs are shown in Figure 20 for current CU 2010.
[0165] In another example, the positions of the chroma samples in a sample pair are shown in Figure 21 for the current CU 2110, and sample B can be at any position on the long side of the current CU 2110. As shown in Figure 21, sample B is in the center of the long side of the current CU 2110. The x coordinate of sample B can be (width>>1)-1 or (width>>1) or (width>>1)+1.
[0166] In another example, the positions of the chroma samples in a sample pair are shown in Figure 22 for the current CU 2210, and sample B can be at any position on the long side of the current CU 2210. As shown in Figure 22, sample B is in the center of the long side of the current CU 2210. The x coordinate of sample B can be (width>>1)-1 or (width>>1) or (width>>1)+1.
[0167] In another example, for 2×N or N×2 chroma blocks, N can be any positive integer, such as 2, 4, 8, 16, 32, or 64. If N is equal to 2, only one adjacent sample pair from the left side and one adjacent sample pair from the top side are used to calculate the maximum and minimum sample values. Otherwise, if N is greater than 2, only one adjacent sample pair from the short side of the current CU and two or more adjacent sample pairs from the long side are used to calculate the maximum and minimum sample values. The positions of the adjacent sample pairs used are fixed. In one example, the positions of the chroma samples in the sample pairs are shown in FIG. 23 for the current CU 2310, where sample B is in the middle of the long side of the current CU 2310 and sample C is at the end of the long side.
[0168] In another example, for a chroma block, if both the width and height of the current block are greater than 2, two or more sample pairs from the left side and two or more sample pairs from the top side of the current block are used to calculate the maximum and minimum sample values. The positions of the adjacent sample pairs used are fixed. In one example, referring to FIG. 24, for each side of the current CU 2410, one sample is in the center of the side (such as samples B and C) and another sample is at the end of the side (such as samples A and D).
[0169] In another example, for CCLM mode, only a subset of the neighboring samples in the specified neighboring samples are used to calculate the maximum and minimum sample values. The number of samples selected varies depending on the block size. In one example, the number of neighboring samples selected is the same for the left side and the top side. For chroma blocks, the samples to be used are selected by a specified scan order. With reference to FIG. 25, for the left side of the chroma block 2510, the scan order is top to bottom, and for the top side, the scan order is left to right. The selected neighboring samples are highlighted with dark circles. With reference to FIG. 26, for chroma block 2610, the samples to be used are selected by a reverse scan order. That is, for the left side of the chroma block 2610, the scan order is bottom to top, and for the top side, the scan order is right to left. The number of neighboring samples selected is the same for the left side and the top side, and the selected neighboring samples are highlighted with dark circles.
[0170] In a further embodiment, for CCLM mode, different downsampling (or interpolation) filter types are used for the left neighboring luma samples and the upper neighboring luma samples.
[0171] For example, some filter types are the same for left and top neighboring samples, but the positions of the filters used for left and top neighboring samples are different. With reference to Figure 27, the positions of the filters for the luma samples on the left are shown, where W and X indicate the coefficients of the filter and the block indicates the position of the luma sample. With reference to Figure 28, the positions of the filters for the luma samples on the top are shown, where W and X indicate the coefficients of the filter and the block indicates the position of the luma sample.
[0172] In a further embodiment, up to N adjacent reconstructed sample pairs within the specified region are used to calculate the parameters a and b in the linear model prediction mode, where N is a positive integer such as 4, 8, or 16. The value of N may depend on the block size of the current block.
[0173] For example, LIC and CCLM (or any variant of CCLM such as L_CCLM, T_CCLM, LT_CCLM, etc.) share the same adjacent sample positions to calculate the linear model parameters a and b.
[0174] In another example, the sample pairs used to calculate the LIC parameters a and b may be located at different positions for different color components. The sample pairs for different color components may be determined independently, but the algorithm may be the same, for example, using minimum and maximum sample values to derive the model parameters.
[0175] In another example, a first sample pair used to calculate the LIC parameters a and b of one color component (e.g., luma) is located first, and then the sample pair used to calculate the LIC parameters a and b of another component (e.g., U or V) is determined as the samples co-located with the first sample pair.
[0176] In another example, LIC and CCLM (or any variant of CCLM, such as L_CCLM, T_CCLM, LT_CCLM, etc.) share the same method, such as the least square error method or the max-min method, to calculate the linear model parameters a and b. Both LIC and CCLM may use the max-min method to calculate the linear model parameters a and b. In the case of LIC, the sum of the absolute values of the two samples in each sample pair may be used to calculate the maximum and minimum values of the linear model.
[0177] In another example, there are three different LIC modes: LT_LIC, L_LIC, and T_LIC. For LT_LIC, both left and upper neighboring samples are used to derive the linear model parameters a and b. For L_LIC, only the left neighboring samples are used to derive the linear model parameters a and b. For T_LIC, only the upper neighboring samples are used to derive the linear model parameters a and b. If the current mode is a LIC mode, another flag may be signaled to indicate whether the LT_LIC mode is selected. Otherwise, one additional flag may be signaled to indicate whether the L_LIC mode or the T_LIC mode is signaled.
[0178] In another example, for an N×2 or 2×N block, only two adjacent sample pairs are used to calculate the parameters a and b (or calculate the maximum and minimum sample values) in the linear model prediction mode, and the positions of the adjacent sample pairs used are predefined and fixed.
[0179] In this example embodiment, for a linear model prediction mode that uses both sides to calculate parameters a and b, if both left and upper neighboring samples are available, only one neighboring sample pair from the left and one neighboring sample pair from the upper side are used to calculate parameters a and b in the linear model prediction mode.
[0180] In this example embodiment, for the linear model prediction mode using both sides to calculate parameters a and b, if the left or upper neighboring samples are not available, two neighboring sample pairs from the available side are used to calculate parameters a and b in the linear model prediction mode. Two examples of predefined positions of current CUs 2910 and 3010, respectively, are shown in Figures 29 and 30.
[0181] In this example embodiment, for a linear model prediction mode that uses both sides to calculate parameters a and b, if the left or upper neighboring samples are not available, parameters a and b are set to default values, such as a being set to 1 and b being set to 512.
[0182] In this example embodiment, the locations of the sample pairs are shown in FIGS.
[0183] In this example embodiment, the locations of the sample pairs are shown in FIG.
[0184] In this example embodiment, the positions of the sample pairs are shown in Figure 21, and sample B can be at any position on the long side of the current CU 2110. For example, sample B can be in the center of the long side. In another example, the x coordinate of sample B can be (width>>1)-1 or (width>>1) or (width>>1)+1.
[0185] In this example embodiment, the positions of the sample pairs are shown in FIG. 22, and sample B can be at any position on the long side of the current CU 2210. For example, sample B can be in the center of the long side. In another example, the x coordinate of sample B can be (width>>1)-1 or (width>>1) or (width>>1)+1.
[0186] In another example, for 2×N or N×2 blocks, N can be any positive integer, such as 2, 4, 8, 16, 32, or 64. If N is equal to 2, only one adjacent sample pair from the left side and one adjacent sample pair from the top side are used to calculate the linear model parameters a and b. Otherwise, if N is greater than 2, only one adjacent sample pair from the short side and two or more adjacent sample pairs from the long side are used to calculate the linear model parameters a and b. The positions of the adjacent sample pairs used are predefined and fixed. Referring to FIG. 23, the positions of the sample pairs are shown, and sample B can be at any position on the long side of the current CU 2310. Sample B is in the middle of the long side, and sample C is at the end of the long side.
[0187] In another example, K1 or more sample pairs from the left side and K2 or more sample pairs from the top side are used to calculate the linear model parameters a and b. The positions of the adjacent sample pairs used are predefined and fixed. Furthermore, the number of selected samples, such as K, depends on the block size. K1 and K2 are positive integers, such as 1, 2, 4, 8, etc., respectively. K1 and K2 may be equal.
[0188] In this example embodiment, the selected sample pairs on each side are evenly distributed.
[0189] In this example embodiment, the values of K1 and K2 should be less than or equal to the width and / or height of the current block.
[0190] In this example embodiment, the adjacent samples are selected by a specified scan order. When the number of selected samples reaches the specified number, the selection process ends. The number of selected samples on the left side and the top side may be the same. For the left side, the scan order is from top to bottom, and for the top side, the scan order is from left to right. With reference to Figures 25 and 31, the selected adjacent samples are highlighted with dark circles. In Figure 25, the adjacent samples are not downsampled, and in Figure 31, the adjacent samples are downsampled. Alternatively, the adjacent samples may be selected by a reverse scan order. That is, for the left side, the scan order is from bottom to top, and for the top side, the scan order is from right to left. With reference to Figures 26, 32 and 33, the positions of the selected samples are shown. In Figure 26, the adjacent samples are not downsampled, and in Figures 32 and 33, the adjacent samples are downsampled.
[0191] In this example embodiment, if both the width and height of the current block are greater than Th, K or more sample pairs from the left side and two or more sample pairs from the top side are used to calculate the linear model parameters a and b. The positions of the adjacent sample pairs used are predefined and fixed. Th is a positive integer such as 2, 4, 8, etc. K is a positive integer such as 2, 4, 8, etc.
[0192] In this example embodiment, if the left or upper neighboring samples are unavailable, then K samples from the available side are used to derive parameters a and b.
[0193] In another example, for linear models that use only left or upper neighboring samples, such as T_CCLM or L_CCLM, if the size of the side used, such as upper for T_CCLM or left for L_CCLM, is equal to 2, then only M samples are used, where M is a positive integer, such as 2 or 4, and the position is predefined. Otherwise, up to N samples are used, where N is a positive integer, such as 4 or 8. M may not be equal to N.
[0194] In this example embodiment, with reference to FIG. 34, if the size of the side used is greater than 2, a predetermined position is shown and the top side is used.
[0195] In this example embodiment, with reference to Figures 35 and 36, when the size of the side used is equal to 2, a predetermined position is shown, in Figure 35 the top side is used and in Figure 36 the left side is used.
[0196] Additionally, the proposed methods may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.
[0197] The above techniques may be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, Figure 37 illustrates a computer system (3700) suitable for implementing certain embodiments of the disclosed subject matter.
[0198] Computer software can be encoded using any suitable machine or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed directly or via translation, microcode execution, etc. by a computer central processing unit (CPU), graphics processing unit (GPU), etc.
[0199] The instructions may be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.
[0200] 37 for the computer system (3700) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of the computer system (3700).
[0201] The computer system (3700) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown). Human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as sound (speech, music, ambient sounds, etc.), images (scanned images obtained from a still image camera, photographic images, etc.), video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).
[0202] The input human interface devices may include one or more of a keyboard (3701), a mouse (3702), a trackpad (3703), a touch screen (3710), a data glove (3704), a joystick (3705), a microphone (3706), a scanner (3707), and a camera (3708).
[0203] The computer system (3700) may also include certain human interface output devices that may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (3710), data gloves (3704), or joystick (3705), although there may be haptic feedback devices that do not function as input devices), audio output devices (such as speakers 3709, headphones (not shown)), visual output devices (such as screens 3710 including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light emitting diode (OLED) screens, each with or without touchscreen input capability, each with or without haptic feedback capability, some of which may be capable of two-dimensional visual output or three or more dimensional output through means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0204] The computer system (3700) may also include human accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (3720) with CD / DVD or similar media (3721), thumb drives (3722), removable hard drives or solid state drives (3723), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD based devices (not shown) such as security dongles, etc.
[0205] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.
[0206] The computer system (3700) may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including Global System for Mobile Communications (GSM), Third Generation (3G), Fourth Generation (4G), Fifth Generation (5G), Long Term Evolution (LTE), etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial including CANBus, etc. A particular network typically requires an external network interface adapter (3754) that is attached to a particular general-purpose data port or peripheral bus (3749) (e.g., a universal serial bus (USB) port of the computer system (3700) for communicating with an external network (3755), others are generally integrated into the core of the computer system (3700) by connecting to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system, etc.). Using any of these networks, the computer system (3700) can communicate with other entities. Such communications may be one-way, receive only (e.g., television broadcast), one-way transmit only (e.g., CANbus to certain CANbus devices), or bidirectional, e.g., to other computer systems using local or wide area digital networks. As noted above, particular protocols and protocol stacks may be used with each of these networks and network interfaces.
[0207] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be connected to the core (3740) of the computer system (3700).
[0208] The cores (3740) may include one or more central processing units (CPUs) (3741), graphics processing units (GPUs) 3742, specialized programmable processing units in the form of field programmable gate areas (FPGAs) (3743), hardware accelerators for specific tasks (3744), and the like. These devices may connect through a system bus (3748) along with read only memory (ROM) (3745), random access memory (RAM) (3746), internal mass storage devices such as internal hard drives that are not accessible to the user, solid state drives (SSDs), and the like (3747). In some computer systems, the system bus (3748) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, and the like. Peripheral devices may be connected directly to the core's system bus (3748) or through a peripheral bus 3749. Peripheral bus architectures include peripheral component interconnect (PCI), USB, and the like.
[0209] The CPU (3741), GPU (3742), FPGA (3743), and accelerator (3744) can execute certain instructions that, in combination, can constitute the aforementioned computer code. The computer code can be stored in ROM (3745) or RAM (3746). Transient data can also be stored in RAM (3746), while permanent data can be stored, for example, in internal mass storage (3747). Rapid storage and retrieval of any memory device can be made possible through the use of cache memory, which can be closely associated with one or more of the CPU (3741), GPU (3742), mass storage (3747), ROM (3745), RAM (3746), etc.
[0210] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.
[0211] As an example, but not by way of limitation, a computer system having the architecture (3700), and in particular the core (3740), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be the user-accessible mass storage devices introduced above, as well as media associated with specific storage of the core (3740) of a non-transitory nature, such as the core internal mass storage (3747) or ROM (3745). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (3740). The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the core (3740), and in particular the processor therein (including a CPU, GPU, FPGA, etc.) to perform certain processes or certain portions of certain processes described herein, including defining data structures stored in RAM (3746) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of being logically hardwired or embodied in circuitry (e.g., accelerator 3744) that can operate in place of or together with software to perform certain processes or certain portions of certain processes described herein. References to software may include logic, and vice versa, as appropriate. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0212] While this disclosure describes several exemplary embodiments, there are modifications, permutations, and various substitute equivalents that are within the scope of the disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the disclosure and are therefore within its spirit and scope. [Explanation of symbols]
[0213] 400 Communication Systems 410 First Terminal 420 Second Terminal 430 Terminal 440 Terminal 450 Network 501 Video Sources 502 uncompressed video sample stream 503 Encoder 504 Encoded Video Bitstream 505 Streaming Server 506 Streaming Client 507 Copying of encoded video bitstreams 508 Streaming Client 509 Copying of encoded video bitstreams 510 Video Decoder 511 Outgoing Video Sample Stream 512 Display 513 Capture Subsystem 610 Receiver 612 Channels 615 Buffer Memory 620 Parser 621 Symbols 651 Scaler / Descaler Unit 652 intra prediction units 653 Motion Compensation Prediction Unit 655 Aggregation Device 656 Current (partially reconstructed) picture 657 Reference Picture Memory 658 Loop Filter Unit 730 Encoder 732 Encoding Engine 733 (local) decoder 734 Reference Picture Memory 735 Predictors 740 Transmitter 743 Video Sequence 745 Entropy Coder 750 Controller 760 Communication Channel 800 processes 810 Video Sequence 820 Video Sequence 830 Video Sequence 1410 Coding Unit (CU) 1420 Block 1510 Coding Unit (CU) 1610 Coding Unit (CU) 1710 Coding Unit (CU) 1810 Coding Unit (CU) 1910 Coding Unit (CU) 2010 Coding Unit (CU) 2110 Coding Unit (CU) 2210 Coding Unit (CU) 2310 Coding Unit (CU) 2410 Coding Unit (CU) 2510 Chroma Block 2610 Chroma Block 2910 Coding Unit (CU) 3010 Coding Unit (CU) 3700 Computer Systems 3701 Keyboard 3702 Mouse 3703 Trackpad 3704 Data Gloves 3705 Joystick 3706 Mike 3707 Scanner 3708 Camera 3709 Speaker 3710 Screen 3720 CD / DVD ROM / RW 3721 Similar media 3722 Thumb Drive 3723 Solid State Drive 3740 cores 3741 Central Processing Unit (CPU) 3742 Graphics Processing Unit (GPU) 3743 Field Programmable Gate Area (FPGA) 3744 Accelerator 3745 Read-Only Memory (ROM) 3746 Random Access Memory (RAM) 3747 Internal Mass Storage 3748 System Bus 3749 Surrounding bus 3754 External Network Interface Adapter 3755 External Network
Claims
1. 1. A method for decoding a video sequence, the method comprising: applying a cross-component linear model (CCLM) to the video sequence; applying an interpolation filter in the cross-component linear model (CCLM); obtaining an absolute difference between a maximum luma sample value and a minimum luma sample value in a specified contiguous sample region in the video sequence; performing a non-uniform quantization on the obtained absolute differences; dividing the non-uniform quantized absolute difference into a first interval and a second interval, the division being based on whether the non-uniform quantized absolute difference is greater than a first threshold or is less than or equal to the first threshold; deriving an entry index for a CCLM lookup table, for the first interval, deriving a first entry index of the CCLM lookup table using a first step size; deriving a second entry index for the CCLM lookup table using a first step size and a second step size different from the first step size for the second interval; retrieving CCLM parameters from the CCLM lookup table based on the first entry index and the second entry index; predicting samples of different chroma blocks among the chroma blocks associated with the video sequence based on the obtained CCLM parameters; Including, A method, wherein the interpolation filter depends on a YUV format of the video sequence.
2. When deriving an entry index for the CCLM lookup table, further comprising:
2. The method of claim 1, further comprising the step of obtaining a floor value using the non-uniform quantized absolute difference, and using the obtained floor value in deriving the entry index of the CCLM lookup table.
3. The method described in claim 1, wherein when applying the interpolation filter in the cross-component linear model (CCLM), the method further comprises a step of using taps of the interpolation filter that depend on the YUV format of the video sequence.
4. The method described in claim 3, further comprising a step of using taps of the interpolation filter used in the cross-component linear model (CCLM) that are the same for the same YUV format of the video sequence.
5. A method according to any one of claims 1 to 4, wherein when applying the interpolation filter in the cross-component linear model (CCLM), the method further comprises the step of setting the applied interpolation filter to be the same for the 4:4:4 and 4:2:2 YUV formats if the video sequence includes a 4:4:4 or 4:2:2 YUV format, and the step of setting the applied interpolation filter to be different from the interpolation filter applied for the 4:4:4 and 4:2:2 YUV formats if the video sequence includes a 4:2:0 YUV format.
6. A method according to any one of claims 1 to 5, wherein when applying the interpolation filter in the cross-component linear model (CCLM), the method further comprises a step of using different taps of the interpolation filter for various YUV formats of the video sequence.
7. A method according to any one of claims 1 to 6, wherein when applying the interpolation filter to the cross-component linear model (CCLM), the method further comprises a step of setting the interpolation filter differently for adjacent luma reconstruction samples above and to the left.
8. The method described in Claim 7, further comprising a step of setting the interpolation filter, wherein the interpolation filters of the upper and left adjacent luma reconstruction samples depend on the YUV format of the video sequence.
9. A method according to any one of claims 1 to 8, further comprising a step of using at least one row of the upper adjacent region and at least one column of the left adjacent region in the cross-component linear model (CCLM) for video sequences having a YUV format of 4:4:4 or 4:2:
2.
10. The method of claim 9, further comprising a step of using at least one of one row of the upper adjacent region and one column of the left adjacent region in the cross-component linear model (CCLM) for video sequences having one of 4:4:4 or 4:2:2 YUV formats.
11. The method described in claim 9, further comprising a step of using at least one of one row of the upper adjacent region and at least two columns of the left adjacent region in the cross-component linear model (CCLM) for a video sequence having a 4:2:2 YUV format.
12. The method of claim 11, further comprising: obtaining the maximum and minimum values using N adjacent sample pairs of the luma block and the chroma block, where N is a positive integer among 4, 8, and 16; 12. The method of claim 1, wherein each of the N pairs of adjacent samples includes a first adjacent sample at a first position adjacent to a first one of the luma block and the chroma block, and a second adjacent sample at a second position adjacent to a second one of the luma block and the chroma block and corresponding to the first position adjacent to the first one of the luma block and the chroma block.
13. The method of claim 12, further comprising the step of selecting the first adjacent sample by scanning the plurality of adjacent samples from bottom to top and / or right to left.
14. A method for obtaining a scale factor and an offset of a linear model of a local illumination correction (LIC) of a current block using N adjacent sample pairs of the current block and the reference block, where N is a positive integer among 4, 8, and 16; each of the N adjacent sample pairs includes a first adjacent sample at a first position adjacent to the current block and a second adjacent sample at a second position adjacent to the reference block and corresponding to the first position adjacent to the current block; performing the LIC of the current block using the obtained scale factor and the obtained offset; 14. The method of any one of claims 1 to 13, further comprising:
15. The method of claim 14, wherein N depends on the block size of the current block.
16. A device configured to carry out a method according to any one of claims 1 to 15.
17. A program for causing one or more processors to carry out a method according to any one of claims 1 to 15.