ON Adaptive Loop Filtering for Video Symbolization
Adaptive loop filtering techniques improve video encoding efficiency by refining encoding processes, reducing bit rate and enhancing quality, addressing bandwidth challenges in high-resolution videos.
Patent Information
- Application Number
- JP2021559943
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-16
- Filing Date
- 2020-04-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-04-16
AI Technical Summary
Existing video encoding technologies face challenges in efficiently managing bandwidth demand due to the increasing resolution of digital videos, necessitating improved encoding methods to reduce bit rate and enhance coding efficiency.
Adaptive loop filtering techniques are employed to refine video encoding processes, utilizing filter coefficients and intermediate results to enhance video quality and reduce bitstream representation complexity, including methods for encoding and decoding with adaptive loop filters and selective application of temporal filters.
Enhances video coding efficiency by reducing bit rate and improving quality through adaptive loop filtering, addressing the increasing bandwidth demands of high-resolution videos.
Smart Images

Figure 0007701271000067 
Figure 0007701271000068 
Figure 0007701271000069
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is based on International Patent Application No. PCT / CN2020 / 085075, filed on April 16, 2020, which claims the priority and benefit of International Patent Application No. PCT / CN2019 / 082855, filed on April 16, 2019. For all purposes under the US law, the entire disclosure of the above application is incorporated by reference herein as part of the disclosure of this specification.
[0002] This patent document relates to video encoding technology, devices, and systems.
Background Art
[0003] Despite the progress of video compression, digital video still occupies the largest bandwidth usage in the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for the use of digital video is expected to continue to grow.
Summary of the Invention
[0004] Regarding digital video encoding, specifically, devices, systems, and methods related to adaptive loop filtering for video encoding are described. The described methods can be applied to both existing video encoding standards (e.g., High Efficiency Video Coding (HEVC)) and future video encoding standards (e.g., Versatile Video Coding (VVC)), or both codecs.
[0005] The video coding standard has mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and both organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, in 2015, VCEG and MPEG jointly established the JVET (Joint Video Exploration Team). Since then, many new methods have been adopted by JVET and incorporated into the reference software called JEM (Joint Exploration Mode). In April 2018, the Joint Video Expert Team (JVET) was launched between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG), and is working on formulating the VVC standard with the goal of reducing the bit rate by 50% compared to HEVC.
[0006] In one representative embodiment, the disclosed technology may be used to provide a method of video processing. This method includes performing a filtering process on the current video block of the video, using filter coefficients and including two or more operations with at least one intermediate result, applying a clipping operation to the at least one intermediate result, and performing a conversion between the current video block and the bitstream representation of the video based on the at least one intermediate result, where the at least one intermediate result is based on the sum of the weighting of the filter coefficients and the difference between the current sample of the current video block and the samples in the vicinity of the current sample.
[0007] In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method includes encoding a current video block of a video into a bitstream representation of the video, where the current video block is encoded with an Adaptive Loop Filter (ALF), and selectively including an indication of a set of temporal adaptive filters within the one or more sets of temporal adaptive filters in the bitstream representation based on the availability or use of the one or more sets of temporal adaptive filters.
[0008] In yet another representative aspect, the disclosed technology may be used to provide a method for video processing. This method includes determining the availability or use of one or more sets of temporal adaptive filters comprising the set of temporal adaptive filters applicable to a current video block of a video encoded with an Adaptive Loop Filter (ALF) based on an indication of the set of temporal adaptive filters in the bitstream representation of the video, and generating the current video block decoded from the bitstream representation by selectively applying the set of temporal adaptive filters based on the determining.
[0009] In yet another representative aspect, the disclosed technology may be used to provide a method for video processing. This method includes determining the number of sets of temporal Adaptive Loop Filtering (ALF) coefficients for a current video block encoded with an Adaptive Loop Filter based on an available set of temporal ALF coefficients, where the available set of temporal ALF coefficients has been encoded or decoded prior to the determining, and the number of sets of ALF coefficients is used for a tile group, tile, slice, picture, Coding Tree Block (CTB), or video unit constituting the current video block, and performing a conversion between the current video block and the bitstream representation of the current video block based on the number of sets of temporal ALF coefficients.
[0010] In yet another representative aspect, the disclosed technology may be used to provide a method of video processing. This method includes determining that an indication of adaptive loop filtering (ALF) in a header of a video region of the video is equal to an indication of ALF in an adaptive parameter set (APS) network abstraction layer (NAL) unit associated with the bitstream representation for conversion between a current video block of the video and the bitstream representation of the video, and performing the conversion.
[0011] In yet another representative aspect, the disclosed technology may be used to provide a method of video processing. This method includes selectively enabling a non-linear adaptive loop filtering (ALF) operation based on a type of adaptive loop filter used in a video region of the video for conversion between a current video block of the video and the bitstream representation of the video, and performing the conversion after the selective enabling.
[0012] In yet another representative aspect, the above method is implemented in the form of code executable by a processing device and stored in a computer-readable program medium.
[0013] In yet another representative aspect, a device configured or operable to perform the above-described method is disclosed. The device may include a processing device programmed to implement this method.
[0014] In yet another representative aspect, a video decoder device may implement a method as described herein.
[0015] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the description, and the claims.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11A
Figure 11B
Figure 11C
Figure 11D
Figure 11E
Figure 11F
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0017] Due to the increasing demand for higher-resolution videos, video encoding methods and technologies are ubiquitous in modern technology. A video codec generally includes an electronic circuit or software for compressing or decompressing digital video and is constantly improved to provide higher encoding efficiency. A video codec converts uncompressed video into a compressed format or vice versa. There is a complex relationship among video quality, the amount of data used to represent the video (determined by the bitrate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, and end-to-end latency (waiting time). This compressed format usually conforms to standard video compression specifications, such as the High Efficiency Video Coding (HEVC) standard (also known as H.265 or MPEG-H Part 2), the Versatile Video Coding (VVC) standard to be completed, or other current and / or future video encoding standards.
[0018] In some embodiments, future video encoding technologies are explored using reference software known as the Joint Exploration Model (JEM). In JEM, sub-block-based prediction is applied with several encoding tools such as affine prediction, alternative temporal motion vector prediction (ATMVP), spatio-temporal motion vector prediction (STMVP), bidirectional optical flow (BIO), frame rate up-conversion (FRUC), local adaptive motion vector resolution (LAMVR), overlapped block motion compensation (OBMC), local illumination compensation (LIC), decoder-side motion vector refinement (DMVR), etc.
[0019] Embodiments of the disclosed technology may be applied to existing video encoding standards (e.g., HEVC, H.265) and future standards to improve runtime performance. In this specification, chapter headings are used to improve readability of the description, and the description or embodiments (and / or implementations) are not limited to each chapter only.
[0020] 1. Examples of Color Spaces and Chroma Subsampling A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes a range of colors as a numerical tuple, typically consisting of 3 or 4 values or color components (e.g., RGB). Basically, a color space is a refined combination of a coordinate system and a subspace.
[0021] In video compression, the most frequently used color spaces are YCbCr and RGB.
[0022] YCbCr, Y’CbCr, or Y Pb / Cb Pr / Cr, also called YCBCR or Y’CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photo systems. Y’ is the luminance component, and CB and CR are the blue-difference and red-difference chroma components. Y’ (with a prime) is distinguished from Y, where Y is luminance and means that the light intensity is non-linearly encoded based on gamma-corrected RGB primary colors.
[0023] Chroma subsampling is a method of encoding an image by taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance, so that the chroma information has a lower resolution than the luminance information.
[0024] 1.1 4:4:4 Color Format Each of the three Y’CbCr components has the same sample rate, so there is no chroma subsampling. This scheme may be used in high-end film scanners and movie post-production.
[0025] 1.2 4:2:2 Color Format The two chroma components are sampled at half the sample rate of luminance. For example, the horizontal chroma resolution is halved. This results in little or no visual difference and can reduce the bandwidth of an uncompressed video signal by 1 / 3.
[0026] 1.3 4:2:0 Color Format In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but in this scheme, the Cb and Cr channels are sampled only on every other line, so the vertical resolution is halved. Thus, the data rate is the same. Cb and Cr are each subsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme with different horizontal and vertical positions.
[0027] ○ In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are located between pixels in the vertical direction (located between the grids).
[0028] ○ In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located between the grids in the middle of alternating luminance samples.
[0029] ○ In 4:2:0 DV, Cb and Cr are co-located horizontally. In the vertical direction, they are co-located on lines that alternate.
[0030] 2 Examples of Encoding Flows of Typical Video Codecs FIG. 1 shows an example of an encoder block diagram of VVC including three in-loop filtering blocks, namely, a deblocking filter (DF), a sample adaptive offset (SAO), and an ALF. Different from the DF (which uses a pre-defined filter), SAO and ALF utilize the original samples of the current image, add offsets, and apply a finite impulse response (FIR) filter to reduce the mean squared error between the original samples and the reconstructed samples, respectively, along with the encoded side information that signals the offsets and filter coefficients. ALF is located at the last processing stage of each picture and can be regarded as a tool that tries to capture and correct the artifacts generated in the previous stage.
[0031] 3 Examples of Adaptive Loop Filters Based on Shape Conversion in JEM In JEM, an adaptive loop filter (GALF) based on shape conversion using block-based filter adaptation "3" is applied. For the luminance component, one of 25 filters is selected for each 2×2 block based on the direction and action of the local gradient.
[0032] 3.1 Examples of filter shapes In the present application, as the luminance component, up to three diamond filter shapes (5×5 diamond, 7×7 diamond, 9×9 diamond, as shown in FIGS. 2A, 2B, and 2C respectively) can be selected. To indicate the filter shape used for the luminance component, an index is signaled at the picture level. For the chroma component in one picture, a 5×5 diamond shape is always used.
[0033] 3.1.1 Block classification Each 2×2 block is classified into one of 25 classes. The classification index C is derived as follows based on the quantization values of its directionality D and activity A^.
[0034]
Equation
[0035] To calculate D and A^, first, a 1-D Laplacian is used to calculate the gradients in the horizontal, vertical, and two diagonal directions.
[0036]
Equation
[0037]
Equation
[0038]
Equation
[0039]
Number
[0040] i and j represent the coordinates of the upper left sample of the 2×2 block, and R(i, j) indicates the sample reconstructed at the coordinates (i, j). Then, set the maximum value Dmax and the minimum value Dmin of the gradients in the horizontal and vertical directions as follows.
[0041]
Number
[0042] And the maximum and minimum values of the gradients in the two diagonal directions are set as follows.
[0043]
Number
[0044] To derive the value of the directivity D, compare these values with each other and with two threshold values t1 and t2. Step 1.
[0045]
Number
[0046] If both are true, it is set to 0. Step 2.
[0047]
Number
[0048] In this case, continue from Step 3 or continue from Step 4. Step 3.
[0049]
Number
[0050] If it is the case, D is set to 2, or D is set to 1. Step 4.
[0051]
Number
[0052] If it is the case, D is set to 4, or D is set to 3. The activity value A is calculated as follows.
[0053]
Number
[0054] A is further quantized in the range of 0 to 4, and the quantized value is denoted as A^. For both chroma components in the picture, the classification method is not applied, that is, a single set of ALF coefficients is applied to each chroma component.
[0055] 3.1.2 Geometric Transformation of Filter Coefficients Before filtering each 2×2 luminance block, based on the gradient value calculated for that block, geometric transformations such as rotation or diagonal and vertical direction inversions are applied to the filter coefficients f(k, l). This is equivalent to applying these transformations to the samples within the filter support region. The idea is to make the different blocks to which ALF is applied more similar by aligning their directions.
[0056] Three geometric transformations including diagonal, vertical direction inversion and rotation are introduced.
[0057]
Number
[0058] Here, K is the size of the filter, 0 ≤ k, l ≤ K - 1 are the coefficient coordinates, the position (0, 0) is at the upper left corner, and the position (K - 1, K - 1) is at the lower right corner. This transformation is applied to the filter coefficient f(k, l) based on the gradient value calculated for that block. The relationship between the transformation and the four gradients in four directions is summarized in Table 1.
[0059]
Table 1
[0060] 3.1.3 Signal Notification of Filter Parameters In JEM, the GALF filter parameters are signaled for the first CTU, i.e., after the slice header and before the SAO parameters of the first CTU. Up to 25 sets of luminance filter coefficients can be signaled. To reduce the bit overhead, filter coefficients of different classifications can be merged. Also, the GALF coefficients of the reference picture can be stored and reused as the GALF coefficients of the current picture. The current picture may choose to use the GALF coefficients stored for the reference picture and avoid GALF coefficient signaling. In this case, only the index to one reference picture is signaled, and the stored GALF coefficients of the indicated reference picture are inherited by the current picture.
[0061] To support the prediction of GALF time, a candidate list of GALF filter sets is maintained. At the start of decoding a new sequence, the candidate list is empty. After decoding one picture, the corresponding set of filters may be added to the candidate list. When the size of the candidate list reaches the maximum allowed value (i.e., 6 in the current JEM), a new set of filters overwrites the oldest set in the order of decoding, i.e., the candidate list is updated applying the first-in-first-out (FIFO) rule. To avoid duplication, only one set can be added to the list if the corresponding picture does not use GALF time prediction. To support temporal scalability, there are candidate lists for multiple filter sets, and each candidate list is associated with one temporal layer. Specifically, each array assigned a temporal layer index (TempIdx) may consist of the filter sets of the previously decoded pictures with a smaller TempIdx. For example, the k-th array is assigned to be associated with a TempIdx equal to k, which includes only the filter sets from pictures with a TempIdx less than or equal to k. After encoding a particular picture, the filter sets associated with this picture are used to update the arrays associated with equal or higher TempIdx.
[0062] Temporal prediction of GALF coefficients is used for inter-coded frames to minimize signaling overhead. For intra-frames, temporal prediction is not available, and one set of 16 fixed filters is assigned to each class. To indicate the use of fixed filters, a flag for each class is signaled, and optionally, the index of the selected fixed filter is signaled. Even if a fixed filter is selected for a given class, the coefficients of the adaptive filter f(k, l) for this class can be transmitted, in which case the coefficients of the filter applied to the reconstructed image are the sum of both coefficient sets.
[0063] The filtering process of the luminance component can be controlled at the CU level. To indicate whether GALF is applied to the luminance component of the CU, one flag is signaled. In the case of the chroma component, whether GALF is applied is indicated only at the picture level.
[0064] 3.1.4 Filtering Process On the decoder side, when GALF is enabled for one block, each sample R(i,j) in this block is filtered, and as a result, the sample value R’(i,j) is obtained as shown below. Here, L represents the filter length, f m,n represents the filter coefficient, and f(k,l) represents the decoded filter coefficient.
[0065]
Equation
[0066] 3.1.5 Determination Process of Encoder-Side Filter Parameters Figure 3 shows the overall encoder determination process for GALF. For each luminance sample of each CU, the encoder determines whether GALF is applied and whether the appropriate signaling flag is included in the slice header. In the case of chroma samples, the decision to apply the filter is made based on the picture level rather than the CU level. Furthermore, the chroma GALF for a picture is checked only if the luminance GALF is enabled for this picture.
[0067] 4 Examples of Adaptive Loop Filters Based on Shape Conversion in VVC The current design of GALF in VVC has the following major changes compared to the design in JEM. 1) Remove the adaptive filter shape. Only a 7×7 filter shape is allowed for the luminance component, and only a 5×5 filter shape is allowed for the chroma component. 2) Both the temporal prediction of ALF parameters and the prediction from the fixed filter are removed. 3) For each CTU, a 1-bit flag is signaled regardless of whether the ALF is enabled or disabled. 4) The calculation of the class index is performed at the 4×4 level instead of 2×2. Also, as proposed in JVET-L0147, a subsampled Laplacian calculation method for ALF classification is used. Specifically, it is not necessary to calculate the horizontal / vertical / 45-degree diagonal / 135-degree gradients for each sample within one block. Instead, 1:2 subsampling is used.
[0068] 5 Examples of Region-Based Adaptive Loop Filter in AVS2 ALF is the final stage of in-loop filtering. This process has two stages. The first stage is the derivation of the filter coefficients. To train the filter coefficients, the encoder classifies the pixels of the reconstructed luminance component into 16 regions, uses the Wiener-Hopf equation to train one set of filter coefficients for each category, and minimizes the mean squared error between the original frame and the reconstructed frame. To reduce the redundancy among these 16 sets of filter coefficients, the encoder adaptively merges them based on the distortion rate performance. At its maximum value, 16 different filter sets can be assigned to the luminance component, and only one filter set can be assigned to the chrominance component. The second stage is the filter determination including both the frame level and the LCU level. First, the encoder determines whether to perform frame-level adaptive loop filtering. If the frame-level ALF is on, the encoder further determines whether to perform LCU-level ALF.
[0069] 5.1 Filter Shape The filter shape adapted to AVS-2 is a 7×7 cross shape, which is a superposition of 3×3 squares for both the luminance component and the chroma component as shown in FIG. 5. The squares in FIG. 5 correspond to samples respectively. Therefore, a total of 17 samples are used to derive the filtered value for the sample at position C8. Considering the overhead of transmitting coefficients, the point-symmetric filter uses {C0, C1, ···, C8} while leaving only 9 coefficients, thereby reducing the number of filter coefficients in filtering by half and reducing the number of multiplications. This point-symmetric filter can also reduce half of the calculation of one filtered sample. For example, only 9 multiplications and 14 addition operations are performed for one filtered sample.
[0070] 5.2 Region-based Adaptive Merge To adapt to different coding errors, AVS-2 employs multiple adaptive loop filters based on regions for the luminance component. The luminance component is divided into 16 approximately equally sized basic regions with each basic region aligned with the boundary of the largest coding unit (LCU) as shown in FIG. 6, and one Wiener filter is derived for each region. The more filters are used, the more distortion is reduced, but the bits used to code these coefficients increase with the number of filters. To achieve the best rate-distortion ratio, these regions can be merged into fewer and larger regions that share the same filter coefficients. To simplify the merge process, each region is assigned an index according to the Hilbert order modified based on the pre-correlation of the image. Based on the rate-distortion cost, two regions with consecutive indexes can be merged.
[0071] The mapping information between regions should be signaled to the decoder. In AVS-2, the number of basic regions is used to indicate the merge result, and the filter coefficients are sequentially compressed according to the order of those regions. For example, when merging {0,1}, {2,3,4}, {5,6,7,8,9} and the left basic region into one region each, only three integers (i.e., 2, 3, 5) are encoded to represent this merge map.
[0072] 5.3 Signaling of Supplementary Information Multiple switch flags are also used. The sequence switching flag adaptive_loop_filter_enable is a flag used to control whether to apply the adaptive loop filter throughout the sequence. The picture switching flag picture_alf_enble[i] controls whether to apply ALF to the corresponding i-th picture component. Only when picture_alf_enble[i] is enabled, the LCU-level flag and filter coefficients corresponding to that color component are transmitted. The LCU-level flag, lcu_alf_enable[k], controls whether ALF is enabled for the corresponding k-th LCU and is interleaved into the slice data. The determination of different levels of adjusted flags is all based on the distortion rate cost. Since it has high flexibility, ALF can further significantly improve the coding efficiency.
[0073] In some embodiments, for the luminance component, a set of up to 16 filter coefficients may exist.
[0074] In some embodiments, one set of filter coefficients may be transmitted for each chroma component (Cb and Cr).
[0075] 6 GALF in VTM-4 In VTM4.0, the filtering process of the adaptive loop filter is performed as follows.
[0076]
Number
[0077] Here, the sample I(x+i, y+j) is the input sample, O(x, y) is the filtered output sample (i.e., the filter result), and w(i, j) represents the filter coefficient. In practice, VTM4.0 is implemented using integer operations for fixed-point precision calculations.
[0078]
Number
[0079] Here, L represents the filter length, and w(i, j) is the filter coefficient in fixed-point precision.
[0080] 7 Nonlinear Adaptive Loop Filtering (ALF) 7.1 Re-formulation of Filtering Equation (11) can be reformulated with the following equation without affecting the coding efficiency.
[0081]
Number
[0082] Here, w(i, j) is the same as the filter coefficient in Equation (11) [except for w(0, 0), which is equal to 1 in Equation (13), but in Equation (11),
[0083]
Number
[0084] is equal to].
[0085] 7.2 Modified Filter By using the filter formula of (13) above, a simple clipping function is used to reduce the influence when the neighboring sample values (I(x+i,y+j)) are too different from the filtering of the current sample value (I(x,y)), thereby easily introducing non-linearity and making the ALF more efficient.
[0086] In this proposal, the ALF filter is modified as follows.
[0087]
Equation
[0088] Here, K(d,b)=min(b,max(-b,d)) is a clipping function, and k(i,j) is a clipping parameter, which depends on the (i,j) filter coefficient. The encoder performs optimization to find the best k(i,j).
[0089] In the implementation form of JVET-N0242, a clipping parameter k(i,j) is specified for each ALF filter, and one clipping value is signaled for each filter coefficient. This means that in the bitstream for each luminance filter, up to 12 clipping values can be signaled, and for the chroma filter, up to 6 clipping values can be signaled.
[0090] To limit the signaling cost and the complexity of the encoder, the evaluation of the clipping values is limited to a small set of possible values. In this proposal, only the same 4 fixed values are used for the INTER and INTRA tile groups.
[0091] Since the variance of the local difference is often larger in the case of luminance than in the case of chroma, two different sets of luminance filters and chroma filters are used. Each set includes the maximum sample value (here, 1024 for a 10-bit bit depth), and clipping can be disabled if not necessary.
[0092] Table 2 shows the set of clipping values used in the JVET-N0242 test. The four values were selected by approximately equally dividing the entire range of sample values for luminance (encoded in 10 bits) and the range of 4 to 1024 for chroma in the logarithmic domain.
[0093] More precisely, the luminance table of clipping values was obtained by the following formula.
[0094]
Equation
[0095] Similarly, the chroma table of clipping values is obtained according to the following formula.
[0096]
Equation
[0097]
Table 2
[0098] The selected clipping values are encoded into the "alf_data" syntax element using the Golomb coding scheme corresponding to the index of the clipping values in Table 2 above. This coding scheme is the same as the coding scheme for filter indexes.
[0099] 8 ALF Based on CTU in JVET-N0415 Slice-level temporal filter. Adaptive Parameter Sets (APSs) are adopted in VTM4. Each APS contains one set of signaled ALF filters, and up to 32 APSs are supported. In this proposal, the slice-level temporal filter is tested. One tile group can reduce overhead by reusing the ALF information from the APS. The APS is updated as a First-In-First-Out (FIFO) buffer.
[0100] CTB-based ALF. For the luma component, when the ALF is applied to the luma CTB, a selection from 16 fixed, 5 temporal, or 1 set of signaled filter sets is shown. Only the filter set index is signaled. For one slice, only one new set of 25 filters can be signaled. When a new set is signaled for one slice, all luma CTBs within the same slice share that set. A new slice-level filter set can be predicted using the fixed filter set and used as a candidate filter set for the luma CTB. The total number of filters is 64.
[0101] For the chroma component, when applying the ALF to the chroma CTB, if a new filter is signaled for one slice, the CTB uses this new filter; otherwise, it applies the latest temporal chroma filter that satisfies the temporal scalability constraint.
[0102] As a slice-level temporal filter, the APS is updated as a First-In-First-Out (FIFO) buffer.
[0103] Specification Based on JVET-K1001-v6, modify using {{fixed filter}}, [[temporal filters]], [[temporal filters]] and ((CTB-based filter index)), i.e., using double curly braces, double square brackets, and double parentheses.
[0104] 7.3.3.2 Adaptive Loop Filter Data Syntax
[0105] [Table 3]
[0106] 7.3.4.2 Encoding Tree Unit Syntax
[0107] [Table 4]
[0108] 7.4.4.2 Adaptive Loop Filter Data Semantics If ((alf_signal_new_filter_luma)) is 1, it indicates that a new luminance filter set is signaled. If alf_signal_new_filter_luma is 0, it indicates that a new luminance filter set is not signaled. If it does not exist, it is 0. If {{alf_luma_use_fixed_filter_flag}} is 1, it indicates that a fixed filter set is used to signal the adaptive loop filter. If alf_luma_use_fixed_filter_flag is 0, it indicates that no fixed filter is used to signal the adaptive loop filter. {{alf_luma_fixed_filter_set_index}} indicates the fixed filter set index. It can be 0...15. If {{alf_luma_fixed_filter_usage_pattern}} is 0, it indicates that all new filters use the fixed filter. If alf_luma_fixed_filter_usage_pattern is 1, it indicates that some of the new filters use the fixed filter and the rest do not. When {{alf_luma_fixed_filter_usage[i]}} is 1, it indicates that the i-th filter uses a fixed filter. When alf_luma_fixed_filter_usage[i] is 0, it indicates that the i-th filter does not use a fixed filter. If it does not exist, it is presumed to be 1. When ((alf_signal_new_filter_chroma)) is 1, it indicates that a new chroma filter is signaled. When alf_signal_new_filter_chroma is 0, it indicates that a new chroma filter is not signaled. ((alf_num_available_temporal_filter_sets_luma)) indicates the number of available temporal filter sets that can be used for the current slice. This number can be 0...5. If it does not exist, it is 0. The variable alf_num_available_filter_sets is derived as 16 + alf_signal_new_filter_luma + alf_num_available_temporal_filter_sets_luma. ((When alf_signal_new_filter_luma is 1, the following processing)) The variable filter Coefficients[sigFiltIdx][j] with sigFiltIdx = 0..alf_luma_num_filters_signalled_minus1 and j = 0..11 is initialized as follows. filterCoefficients[sigFiltIdx][j]=alf_luma_coeff_delta_abs[sigFiltIdx][j]*(1 - 2*alf_luma_coeff_delta_sign[sigFiltIdx][j]) (7 - 50) When alf_luma_coeff_delta_prediction_flag is 1, filterCoefficients[sigFiltIdx][j] with sigFiltIdx = 1..alf_luma_num_filters_signalled_minus1 and j = 0..11 is modified as follows. filterCoefficients[sigFiltIdx][j]+=filterCoefficients[sigFiltIdx-1][j] (7-51) The luminance filter coefficients AlfCoeff for elements with filtIdx = 0..NumAlfFilters-1, j = 0..11 L having AlfCoeff[filtIdx][j] L are derived as follows. AlfCoeff L [filtIdx][j]=filterCoefficients[alf_luma_coeff_delta_idx[filtIdx]][j] (7-52) { { When alf_luma_use_fixed_filter_flag is 1 and alf_luma_fixed_filter_usage[filtidx] is 1, the following applies. AlfCoeff L [filtIdx][j]=AlfCoeff L [filtIdx][j]+AlfFixedFilterCoeff[AlfClassToFilterMapping[alf_luma_fixed_filter_index][filtidx]][j]}} The last filter coefficients AlfCoeff for filtIdx = 0..NumAlfFilters-1 L [filtIdx]
[12] is derived as follows. AlfCoeff L [filtIdx]
[12] =128-Σ k (AlfCoeff L [filtIdx][k]<<1), with k = 0..11 (7-53) When filtIdx = 0..NumAlfFilters - 1 and j = 0..11, AlfCoeff L the value of [filtIdx][j] is in the range of -2 7 to 2 7 -1, and for AlfCoeff L the value of [filtIdx]
[12] is in the range of 0 to 2 8 -1, which is a requirement for bitstream compliance. ((Luminance filter coefficient)) For filtSetIdx = 0..15, filtSetIdx = 0..NumAlfFilters - 1, and j = 0..12, the element AlfCoeff LumaAll AlfCoeff with [filtSetIdx][filtIdx][j] LumaAll is derived as follows. AlfCoeff LumaAll [filtSetIdx][filtIdx][j] = {{AlfFixedFilterCoeff[AlfClassToFilterMapping[}}filtSetIdx{{][filtidx]][j]}} ((Luminance filter coefficient)) AlfCoeff with filtSetIdx = 16, filtSetIdx = 0..NumAlfFilters - 1 and j = 0..12 LumaAll with elements AlfCoeff LumaAll [filtSetIdx][filtIdx][j] is derived as follows. The variable closest_temporal_index is initialized to -1. Tid is the temporal layer index of the current slice. ((If alf_signal_new_filter_luma is 1)) AlfCoeff LumaAll
[16] [filtIdx][j] = AlfCoeff L [filtIdx][j] ((Otherwise, the following process is called.)) for(i = Tid; i >= 0; i--) { for (k = 0; k < temp_size_L; k++) { if (temp Tid_L [k] == i) { closest_temporal_index is set as k; break; } } } AlfCoeff LumaAll
[16] [filtIdx][j] = Temp L [closest_temporal_index][filtIdx][j] When ((luminance filter coefficient)) filtSetIdx = 17..alf_num_available_filter_sets - 1, filtSetIdx = 0..NumAlfFilters - 1 and j = 0..12, the element AlfCoeff LumaAll [filtSetIdx][filtIdx][j], the AlfCoeff with LumaAll is derived as follows. i = 17; for (k = 0; k < temp_size_L && i < alf_num_available_filter_sets; j++) { (If temp Tid_L [k] <= Tid and k is not equal to closest_temporal_index) { AlfCoeff LumaAll [i][filtIdx][j] = Temp L [k][filtIdx][j]; i++; } } {{AlfFixedFilterCoeff}}
[64]
[13] = { {0, 0, 2, -3, 1, -4, 1, 7, -1, 1, -1, 5, 112}, {0,0,0,0,0,-1,0,1,0,0,-1,2,126}, {0,0,0,0,0,0,0,1,0,0,0,0,126}, {0,0,0,0,0,0,0,0,0,0,-1,1,128}, {2,2,-7,-3,0,-5,13,22,12,-3,-3,17,34}, {-1,0,6,-8,1,-5,1,23,0,2,-5,10,80}, {0,0,-1,-1,0,-1,2,1,0,0,-1,4,122}, {0,0,3,-11,1,0,-1,35,5,2,-9,9,60}, {0,0,8,-8,-2,-7,4,4,2,1,-1,25,76}, {0,0,1,-1,0,-3,1,3,-1,1,-1,3,122}, {0,0,3,-3,0,-6,5,-1,2,1,-4,21,92}, {-7,1,5,4,-3,5,11,13,12,-8,11,12,16}, {-5,-3,6,-2,-3,8,14,15,2,-7,11,16,24}, {2,-1,-6,-5,-2,-2,20,14,-4,0,-3,25,52}, {3,1,-8,-4,0,-8,22,5,-3,2,-10,29,70}, {2,1,-7,-1,2,-11,23,-5,0,2,-10,29,78}, {-6,-3,8,9,-4,8,9,7,14,-2,8,9,14}, {2,1,-4,-7,0,-8,17,22,1,-1,-4,23,44}, {3,0,-5,-7,0,-7,15,18,-5,0,-5,27,60}, {2,0,0,-7,1,-10,13,13,-4,2,-7,24,74}, {3,3,-13,4,-2,-5,9,21,25,-2,-3,12,24}, {-5,-2,7,-3,-7,9,8,9,16,-2,15,12,14}, {0,-1,0,-7,-5,4,11,11,8,-6,12,21,32}, {3,-2,-3,-8,-4,-1,16,15,-2,-3,3,26,48}, {2,1,-5,-4,-1,-8,16,4,-2,1,-7,33,68}, {2,1,-4,-2,1,-10,17,-2,0,2,-11,33,74}, {1,-2,7,-15,-16,10,8,8,20,11,14,11,14}, {2,2,3,-13,-13,4,8,12,2,-3,16,24,40}, {1,4,0,-7,-8,-4,9,9,-2,-2,8,29,54}, {1,1,2,-4,-1,-6,6,3,-1,-1,-3,30,74}, {-7,3,2,10,-2,3,7,11,19,-7,8,10,14}, {0,-2,-5,-3,-2,4,20,15,-1,-3,-1,22,40}, {3,-1,-8,-4,-1,-4,22,8,-4,2,-8,28,62}, {0,3,-14,3,0,1,19,17,8,-3,-7,20,34}, {0,2,-1,-8,3,-6,5,21,1,1,-9,13,84}, {-4,-2,8,20,-2,2,3,5,21,4,6,1,4}, {2,-2,-3,-9,-4,2,14,16,3,-6,8,24,38}, {2,1,5,-16,-7,2,3,11,15,-3,11,22,36}, {1,2,3,-11,-2,-5,4,8,9,-3,-2,26,68}, {0,-1,10,-9,-1,-8,2,3,4,0,0,29,70}, {1,2,0,-5,1,-9,9,3,0,1,-7,20,96}, {-2,8,-6,-4,3,-9,-8,45,14,2,-13,7,54}, {1,-1,16,-19,-8,-4,-3,2,19,0,4,30,54}, {1,1,-3,0,2,-11,15,-5,1,2,-9,24,92}, {0,1,-2,0,1,-4,4,0,0,1,-4,7,120}, {0,1,2,-5,1,-6,4,10,-2,1,-4,10,104}, {3,0,-3,-6,-2,-6,14,8,-1,-1,-3,31,60}, {0,1,0,-2,1,-6,5,1,0,1,-5,13,110}, {3,1,9,-19,-21,9,7,6,13,5,15,21,30}, {2,4,3,-12,-13,1,7,8,3,0,12,26,46}, {3,1,-8,-2,0,-6,18,2,-2,3,-10,23,84}, {1,1,-4,-1,1,-5,8,1,-1,2,-5,10,112}, {0,1,-1,0,0,-2,2,0,0,1,-2,3,124}, {1,1,-2,-7,1,-7,14,18,0,0,-7,21,62}, {0,1,0,-2,0,-7,8,1,-2,0,-3,24,88}, {0,1,1,-2,2,-10,10,0,-2,1,-7,23,94}, {0,2,2,-11,2,-4,-3,39,7,1,-10,9,60}, {1,0,13,-16,-5,-6,-1,8,6,0,6,29,58}, {1,3,1,-6,-4,-7,9,6,-3,-2,3,33,60}, {4,0,-17,-1,-1,5,26,8,-2,3,-15,30,48}, {0,1,-2,0,2,-8,12,-6,1,1,-6,16,106}, {0,0,0,-1,1,-4,4,0,0,0,-3,11,112}, {0,1,2,-8,2,-6,5,15,0,2,-7,9,98}, {1,-1,12,-15,-7,-2,3,6,6,-1,7,30,50}, }; {{AlfClassToFilterMapping}}
[16]
[25] = { {8,2,2,2,3,4,53,9,9,52,4,4,5,9,2,8,10,9,1,3,39,39,10,9,52}, {11,12,13,14,15,30,11,17,18,19,16,20,20,4,53,21,22,23,14,25,26,26,27,28,10}, {16,12,31,32,14,16,30,33,53,34,35,16,20,4,7,16,21,36,18,19,21,26,37,38,39}, {35,11,13,14,43,35,16,4,34,62,35,35,30,56,7,35,21,38,24,40,16,21,48,57,39}, {11,31,32,43,44,16,4,17,34,45,30,20,20,7,5,21,22,46,40,47,26,48,63,58,10}, {12,13,50,51,52,11,17,53,45,9,30,4,53,19,0,22,23,25,43,44,37,27,28,10,55}, {30,33,62,51,44,20,41,56,34,45,20,41,41,56,5,30,56,38,40,47,11,37,42,57,8}, {35,11,23,32,14,35,20,4,17,18,21,20,20,20,4,16,21,36,46,25,41,26,48,49,58}, {12,31,59,59,3,33,33,59,59,52,4,33,17,59,55,22,36,59,59,60,22,36,59,25,55}, {31, 25, 15, 60, 60, 22, 17, 19, 55, 55, 20, 20, 53, 19, 55, 22, 46, 25, 43, 60, 37, 28, 10, 55, 52}, {12, 31, 32, 50, 51, 11, 33, 53, 19, 45, 16, 4, 4, 53, 5, 22, 36, 18, 25, 43, 26, 27, 27, 28, 10}, {5, 2, 44, 52, 3, 4, 53, 45, 9, 3, 4, 56, 5, 0, 2, 5, 10, 47, 52, 3, 63, 39, 10, 9, 52}, {12, 34, 44, 44, 3, 56, 56, 62, 45, 9, 56, 56, 7, 5, 0, 22, 38, 40, 47, 52, 48, 57, 39, 10, 9}, {35, 11, 23, 14, 51, 35, 20, 41, 56, 62, 16, 20, 41, 56, 7, 16, 21, 38, 24, 40, 26, 26, 42, 57, 39}, {33, 34, 51, 51, 52, 41, 41, 34, 62, 0, 41, 41, 56, 7, 5, 56, 38, 38, 40, 44, 37, 42, 57, 39, 10}, {16, 31, 32, 15, 60, 30, 4, 17, 19, 25, 22, 20, 4, 53, 19, 21, 22, 46, 25, 55, 26, 48, 63, 58, 55},}; ((When alf_signal_new_filter_chroma is 1, the following processing). When j = 0..5, the chroma filter coefficient AlfCoeff C [j] is derived as follows. AlfCoeff C [j]=alf_chroma_coeff_abs[j]*(1 - 2*alf_chroma_coeff_sign[j]) (7 - 57) The last filter coefficient when j = 6 is derived as follows. AlfCoeff C [6]=128 - Σ k (AlfCoeff C [k]<<1), with k = 0..5 (7 - 58) When filtIdx = 0..NumAlfFilters-1 and j = 0..5, AlfCoeff C [j] has a value in the range of -2 7 ~2 7 to -1, and AlfCoeff C [6] has a value in the range of 0 to 2 8 to -1, which is a requirement for bitstream compliance. Otherwise, the following is called if (((alf_signal_new_filter_chroma is 0))). for(i = Tid; i >= 0; i--) { for(k = 0; k < temp_size_C; k++) { if(temp Tid_C [k] == i) { closest_temporal_index is set as k; break; } } } When j = 0..6, the chroma filter coefficient AlfCoeff C [j] is derived as follows. AlfCoeff C [j] = Temp C [closest_temporal_index][j] 7.4.5.2 Syntax of Encoding Tree Unit ((alf_luma_ctb_filter_set_index[xCtb >> Log2CtbSize][yCtb >> Log2CtbSize])) specifies the filter set index of the luma CTB at position (xCtb, yCtb). When ((alf_use_new_filter)) is 1, it indicates that alf_luma_ctb_filter_set_index[xCtb>>Log2CtbSize][yCtb>>Log2CtbSize] is 16. When alf_use_new_filter is 0, it indicates that alf_luma_ctb_filter_set_index[xCtb>>Log2CtbSize][yCtb>>Log2CtbSize] is not equal to 16. When ((alf_use_fixed_filter)) is 1, it indicates using one of the fixed filter sets. When alf_use_fixed_filter is 0, it indicates that the current luma CTB does not use the fixed filter set. ((alf_fixed_filter_index)) indicates the index of the fixed filter set, and this index can range from 0 to 15. ((alf_temporal_index)) indicates the temporal filter set index, which can range from 0 to alf_num_available_temporal_filter_sets_luma - 1. [[8.5.1 General]] 1. When sps_alf_enabled_flag is 1, the following applies. - The temporal filter update process defined in clause 8.5.4.5 is called. - The adaptive loop filter process defined in clause 8.5.4.1 takes the reconstructed picture sample arrays S L , S Cb and S Cr as inputs, and the modified reconstructed picture sample arrays S' L , S' Cb and S' Cr are called as outputs. - Arrays S' L , S' Cb , S' Cr are arrays S L , SCb , S Cr is assigned to (representing the decoded picture). - The temporal filter update process defined in item 8.5.4.6 is called. ((8.5.4.2 Encoding Tree Block Filtering Process for Luminance Samples)) - For the array f[j] of luminance filter coefficients corresponding to the filter specified by filtIdx[x][y], with j = 0..12, it is derived as follows. f[j] = ((AlfCoeff LumaAll ))[alf_luma_ctb_filter_set_index[xCtb >> Log2CtbSize][yCtb >> Log2CtbSize]][filtIdx[x][y]][j] (8 - 732) [[8.5.4.5 Temporal Filter Update]] If any of the following conditions is true, - The current picture is an IDR picture. - The current picture is a BLA picture. - In the decoding order, the current picture is the first picture whose POC is greater than the POC of the previous decoded IRAP picture, i.e., a picture behind the first picture and behind the subsequent pictures. Next, set temp_size_L and temp_size_C to 0. [[8.5.4.6 Update of Temporal Filter]] When slice_alf_enabled_flag is 1 and alf_signal_new_filter_luma is 1, the following applies. If the luminance temporal filter buffer size, temp_size_L < 5, then temp_size_L = temp_size_L + 1. When i = temp_size_L - 1...1, j = 0...NumAlfFilters - 1, k = 0...12, Temp L[i][j][k] is updated as follows. Temp L [i][j][k] = Temp L [i - 1][j][k] When j = 0...NumAlfFilters - 1 and k = 0..12, Temp L [0][j][k] is updated as follows. Temp L [0][j][k] = AlfCoeff L [j][k] When i = temp_size_L - 1...1, Temp Tid_L [i] is updated as follows. Temp Tid_L [i] = Temp Tid_L [i - 1] Temp Tid_L Set [0] as the time layer index Tid of the current slice. When alf_chroma_idx is not 0 and alf_signal_new_filter_chroma is 1, the following applies. For i = temp_size_c - 1...1 and j = 0...6, Temp c [i][j] is updated as follows. Temp c [i][j] = Temp c [i - 1][j] When j = 0...6, Temp c [0][j] is updated as follows. Temp c [0][j] = AlfCoeff C [j] When i = temp_size_C - 1...1, Temp Tid_C [i] is updated as follows. Temp Tid_C [i] = Temp Tid_C [i - 1] Temp Tid_C Set [0] as the Tid of the current slice.
[0109]
Table 5
[0110]
Table 6
[0111] 9. In-loop reshaping (ILR) in JVET-M0427 The basic idea of in-loop reshaping (ILR) is to transform the original signal (prediction / reconstruction signal in the first domain) into a second domain (reshaped domain).
[0112] The in-loop luminance reshaper is implemented as a pair of lookup tables (LUTs), but since one of the two LUTs can be calculated from the other LUT that is signaled, only one of the two LUTs needs to be signaled. Each LUT is a one-dimensional 10-bit 1024-entry mapping table (1D-LUT). One LUT is the forward LUT, FwdLUT, which maps the input luminance code value Y i to the modified value Y r : Y r = FwdLUT[Y i . The other LUT is the inverse LUT, InvLUT, which maps the modified code value Y r to Y^ i : Y^ i = InvLUT[Y r . (Y^ i represents the reconstructed value of Y i .)
[0113] 9.1 PWL model Conceptually, piecewise linear (PWL) is implemented as follows.
[0114] Let x1 and x2 be two input pivots, and y1 and y2 be the output pivots corresponding to one piece. The output value y for any input value x between x1 and x2 can be interpolated by the following formula
[0115] y = ((y2 - y1) / (x2 - x1)) * (x - x1) + y1
[0116] In the fixed-point implementation, this equation can be rewritten as follows.
[0117] y = ((m * x + 2 * FP_PREC - 1) >> FP_PREC) + c
[0118] Here, m is a scalar, c is an offset, and FP_PREC is a constant for defining the precision.
[0119] Note that in the CE-12 software, the PWL model is used to pre-compute the 1024-entry FwdLUT mapping table and InvLUT mapping table. However, the PWL model also enables the on-the-fly calculation of the same mapping values in the implementation without pre-computing the LUT.
[0120] 9.2 Tests of CE12-2 in the 4th VVC Agreement 9.2.1 Luminance Reshaping The in-loop luminance reshaping test 2 (i.e., CE-12-2 in the proposal) provides a less complex pipeline and eliminates the decoding latency for block-level intra prediction in inter-slice reconstruction. Intra prediction is performed in the reshaped domain for both inter-slice and intra-slice.
[0121] Intra prediction is always performed in the reshaped domain regardless of the slice type. With such a configuration, intra prediction can be started immediately after the previous TU reconstruction. Such a configuration can also provide unified processing for the intra mode instead of depending on the slice. Figure 7 is a block diagram showing the decoding process based on the CE12-2 method.
[0122] CE12-2 may perform luminance and chroma residual scaling by using 16 piecewise linear (PWL) models instead of the 32 PWL models of CE12-1.
[0123] Inter-slice reconstruction using the in-loop luminance reshaper in CE-12-2 (The blocks with faint shadows indicate the signals in the reshaped domain. Luminance residual, intra-luminance prediction, and intra-luminance reconstruction)
[0124] 9.2.2 Luminance-dependent chroma residual scaling Luminance-dependent chroma residual scaling is a multiplication process implemented with fixed-point integer arithmetic. Chroma residual scaling compensates for the interaction between the chroma signal and the luminance signal. Chroma residual scaling is applied at the TU level. Specifically, the average value of the corresponding luminance prediction block is utilized.
[0125] This average value is used to specify the index in the PWL model. This index specifies the scaling coefficient cScaleInv. Multiply that number by the chroma residual.
[0126] Note that the chroma scaling coefficient is calculated from the forward-mapped predicted luminance value rather than the reconstructed luminance value.
[0127] 9.2.3 Signaling of ILR side information The parameters are (currently) signaled in the tile group header (similar to ALF). These are reported to require 40 - 100 bits. The following specifications are based on version 9 of JVET-L1001. The added syntax is highlighted in yellow. Sequence parameter set RBSP syntax in 7.3.2.1
[0128]
Table 7
[0129] Tile group header syntax in 7.3.3.1
[0130] [Table 8]
[0131] Add a new syntax table tile group reshaper.
[0132] [Table 9]
[0133] {Generally, in the semantics of the sequence parameter set RBSP, add the following semantics.} If sps_reshaper_enabled_flag is equal to 1, it specifies that the reshaper is used in the coded video sequence (CVS). If sps_reshaper_enabled_flag is equal to 0, it specifies that the reshaper is not used in the CVS. {In the tile group header syntax, add the following semantics.} If tile_group_reshaper_model_present_flag is equal to 1, it specifies that tile_group_reshaper_model() exists within the tile group. If tile_group_reshaper_model_present_flag is equal to 0, it specifies that tile_group_reshaper_model() does not exist in the tile group header. If tile_group_reshaper_model_present_flag does not exist, it is inferred to be equal to 0. When the tile_group_reshaper_enabled_flag is equal to 1, it stipulates that the reshaper is enabled for the current tile group. When the tile_group_reshaper_enabled_flag is equal to 0, it stipulates that the reshaper is not enabled for the current tile group. If the tile_group_reshaper_enable_flag does not exist, it is inferred to be 0. When the tile_group_reshaper_chroma_residual_scale_flag is equal to 1, it stipulates that chroma residual scaling is enabled for the current tile group. When the tile_group_reshaper_chroma_residual_scale_flag is equal to 0, it stipulates that chroma residual scaling is not enabled for the current tile group. If the tile_group_reshaper_chroma_residual_scale_flag does not exist, it is presumed to be 0. {{Add the tile_group_reshaper_model() syntax}} reshape_model_min_bin_idx stipulates that the minimum bin (or piece) index is used for the reshaper construction process. Assume that the value of reshape_model_min_bin_idx is within the range of 0 to MaxBinIdx. Assume that the value of MaxBinIdx is equal to 15. reshape_model_delta_max_bin_idx stipulates that the result of subtracting the maximum bin index from the maximum allowable bin (or piece) index MaxBinIdx is used in the reshaper construction process. The value of reshape_model_max_bin_idx is set to be equal to MaxBinIdx - reshape_model_delta_max_bin_idx. reshaper_model bin_delta_abs_cw_prec_minus1 + 1 defines the number of bits used for the representation of the syntax reshape_model_bin_delta_abs_CW[i]. reshape_model_bin_delta_abs_CW[i] defines the absolute delta code name value of the i-th bin. reshaper_model_bin_delta_sign_CW_flag[i] describes the sign of reshape_model_bin_delta_abs_CW[i] as follows. - When reshape_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is a positive value. - Otherwise (when reshape_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is a negative value. When reshape_model_bin_delta_sign_CW_flag[i] does not exist, it is assumed to be equal to 0. Variable RspDeltaCW[i] = (1 - 2 * reshape_model_bin_delta_sign_CW[i]) * reshape_model_bin_delta_abs_CW[i]; Variable RspCW[i] is derived as the following steps. Variable OrgCW is set equal to (1 << BitDepth Y ) / (MaxBinIdx + 1). - If reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx RspCW[i] = OrgCW + RspDeltaCW[i]. - Otherwise, RspCW[i] = 0. BitDepth Y When the value of BitDepth is equal to 10, the value of RspCW[i] falls within the range of 32 to 2 * OrgCW_1. The variable InputPivot[i] where i is in the range of 0 to MaxBinIdx + 1 is derived as follows. InputPivot[i] = i * OrgCW For the variable ReshapePivot[i] where i is in the range of 0 to MaxBinIdx + 1, variables ScaleCoef[i] and InvScaleCoeff[i] are each in the range of 0 to MaxBinIdx, and i is in the range of 0 to MaxBinIdx, it is derived as follows. shiftY = 14 ReshapePivot[0] = 0; for (i = 0; i <= MaxBinIdx; i++) { ReshapePivot[i + 1] = ReshapePivot[i] + RspCW[i] ScaleCoef[i] = (RspCW[i] * (1 << shiftY) + (1 << (Log2(OrgCW) - 1))) >> (Log2(OrgCW)) if (RspCW[i] == 0) InvScaleCoeff[i] = 0 else InvScaleCoeff[i] = OrgCW * (1 << shiftY) / RspCW[i] } The variable ChromaScaleCoef[i] where i is within the range of 0 to MaxBinIdx is derived as follows. ChromaResidualScaleLut
[64] = {16384, 16384, 16384, 16384, 16384, 16384, 16384, 8192, 8192, 8192, 8192, 5461, 5461, 5461, 5461, 4096, 4096, 4096, 4096, 3277, 3277, 3277, 3277, 2731, 2731, 2731, 2731, 2341, 2341, 2341, 2048, 2048, 2048, 1820, 1820, 1820, 1638, 1638, 1638, 1638, 1489, 1489, 1489, 1489, 1365, 1365, 1365, 1365, 1260, 1260, 1260, 1260, 1170, 1170, 1170, 1170, 1092, 1092, 1092, 1092, 1024, 1024, 1024, 1024}; shiftC = 11 - If (RspCW[i] == 0) ChromaScaleCoef[i] = (1 << shiftC) - Otherwise (RspCW[i] != 0), ChromaScaleCoef[i] = ChromaResidualScaleLut[RspCW[i] >> 1]
[0134] 9.2.4 Method of Using ILR On the encoder side, first, each picture (or tile group) is converted to a reshaped domain. Then, all encoding processes are performed in the reshaped domain. In the case of intra prediction, neighboring blocks are in the reshaped domain. In the case of inter prediction, first, the reference block (generated from the original domain in the decoded picture buffer) is converted to the reshaped domain. Then, the residual is generated and encoded into the bitstream.
[0135] After the encoding / decoding of the entire picture (or tile group) is completed, the samples in the reshaped domain are converted back to the original domain, and then the deblocking filter and other filters are applied.
[0136] In the following cases, forward reshaping to the prediction signal is disabled.
[0137] ○ The current block is intra-coded.
[0138] ○ The current block is coded as CPR (referencing the current picture, also known as intra-block copy, IBC).
[0139] ○ The current block is coded as combined inter-intra mode (CIIP), and forward reshaping is disabled for the intra-prediction block.
[0140] 10 Bidirectional Optical Flow (BDOF) 10.1 Overview and Analysis of BIO In BIO, first, motion compensation is performed to generate the first prediction (in each prediction direction) of the current block. The first prediction is used to derive the spatial gradient, temporal gradient, and optical flow of each sub-block or pixel within the block, and these are used to generate the second prediction, for example, the final prediction of the sub-block or pixel. The details are described below.
[0141] The bidirectional optical flow (BIO) method is an improvement of the motion at the sample unit executed on block-based motion compensation for bidirectional prediction. In some embodiments, the improvement of the motion at the sample level does not use signaling.
[0142] Let the luminance from the reference k (k = 0, 1) after block motion compensation be I (k) and
[0143]
Equation
[0144] be the horizontal component and vertical component of the I (k) gradient respectively. Assuming that the optical flow is valid, the motion vector field (vx , v y ) is obtained by the following equation.
[0145]
Number
[0146] By combining this optical flow equation for each sample's motion trajectory by Hermite interpolation, the two functional values I at both ends (k) and the derivative
[0147]
Number
[0148] A unique cubic polynomial that matches is obtained. The value of this polynomial at t = 0 is the BIO prediction, such as in the BIO equation.
[0149]
Number
[0150] Figure 8 shows an example of the trajectory of the optical flow in the bidirectional optical flow (BIO) method. Here, τ0 and τ1 indicate the distances to the reference frames. The distances τ0, τ1 are calculated based on the POC of Ref0 and Ref1: τ0 = POC(current) - POC(Ref0), τ1 = POC(current) - POC(Ref1). When both predictions come from the same time direction (either both from the past or both from the future), the signs are different (i.e., τ0·τ1 < 0). In this case, when the predictions are not from the same time (e.g., τ0 ≠ τ1), BIO is applied.
[0151] Motion vector field (v x , v y) is determined by minimizing the difference Δ between the values at points A and B. FIG. 8 shows an example at the intersection of the movement locus and the reference frame plane. The model uses only the first linear term of the local Taylor expansion for Δ as follows.
[0152]
Number
[0153] All values in the above formula depend on the sample position represented as (i’, j’). Assuming that the movement is consistent in the local surrounding area, Δ can be minimized inside a (2M + 1)×(2M + 1) square window Ω centered on the current prediction point (i, j). In the formula, M is equal to 2.
[0154]
Number
[0155] For this optimization problem, JEM uses a simple approach of first minimizing in the vertical direction and then in the horizontal direction. As a result, it becomes as follows.
[0156]
Number
[0157]
Number
[0158] Here
[0159]
Number
[0160] In order to avoid division by zero or a very small numerical value, regularization parameters r and m are introduced in equations (15) and (16). Wherein,
[0161]
Number
[0162]
Number
[0163] Here, d is the bit depth of the video sample.
[0164] To make the biometric memory access the same as normal bidirectional predictive motion compensation, for the positions within the current block, all prediction values and gradient values
[0165]
Number
[0166] are calculated. Figure 9A illustrates the access positions outside step 900. As shown in Figure 9A, in equation (17), a (2M + 1) × (2M + 1) square window Ω centered on the current prediction point on the boundary of the prediction block needs to access positions outside the block. In JEM, the
[0167]
Number
[0168] values are set to be equal to the nearest available numerical value inside the block. For example, this can be implemented as a padding area 901 as shown in Figure 9B.
[0169] By using BIO, the motion field can be improved for each sample. To reduce the computational complexity, block-based BIO design is used in JEM. The motion improvement can be calculated based on 4×4 blocks. In block-based BIO, the values of s in Equation (17) for all samples in a 4×4 block are integrated, and then the integrated value of s is used to derive the BIO motion vector offset for the 4×4 block. Specifically, the following equations can be used for block-based BIO derivation. n The value of s n is integrated, and then the integrated value of s n is used to derive the BIO motion vector offset for the 4×4 block. Specifically, the following equations can be used for block-based BIO derivation.
[0170]
Equation
[0171] where b k represents the set of samples belonging to the k-th 4×4 block of the predicted block, and s in Equations (15) and (16) n is replaced by ((s n,bk ) >> 4) to derive the relevant motion vector offset.
[0172] Depending on the scenario, the BIO MV regime may not be reliable due to noise or irregular motion. Therefore, in BIO, the magnitude of the MV regime is clipped to a threshold. The threshold is determined based on whether all reference pictures of the current picture are from one direction. For example, if all reference pictures of the current picture are from one direction, the threshold is set to 12×2 14-d , and if not, the threshold is set to 12×2 13-d .
[0173] The gradient of BIO may be calculated simultaneously with motion compensation interpolation using operations compliant with HEVC motion compensation processing (e.g., 2D separable finite impulse response (FIR)). In some embodiments, the input for the 2D separable FIR is the same reference frame sample as that for the motion compensation processing and the fractional position (fracX, fracY) according to the fractional part of the block motion vector. Horizontal gradient
[0174] [Number]
[0175] For the case of [Number], first, the signal is interpolated vertically using the BIOfilterS corresponding to the fractional position fracY with the descaling shift d - 8. Next, the horizontal gradient filter BIOfilterG corresponding to the fractional position fracXwith is applied with the descaling shift by 18 - d. Vertical gradient
[0176] [Number]
[0177] For the case of [Number], the gradient filter is applied vertically using the BIOfilterG corresponding to the fractional position fracY with the descaling shift d - 8. Then, the signal is shifted horizontally using the horizontal BIOfilterS corresponding to the fractional position fracX with the descaling shift by 18 - d. To maintain a reasonable level of complexity, the length of the interpolation filters for the gradient calculation BIOfilterG and the signal displacement BIOfilterF may be shorter (e.g., 6 - tap). Table 2 shows exemplary filters that can be used for gradient calculation at different fractional positions of the block motion vector in BIO. Table 3 shows exemplary interpolation filters that can be used for the generation of the predicted signal in BIO.
[0178] [Table 10]
[0179]
Table 11
[0180] In this JEM, when two predictions are from different reference pictures, BIO can be applied to all bi - directional prediction blocks. When enabling the local illumination compensation (LIC) of the CU, BIO can be disabled.
[0181] In some embodiments, OBMC is applied to one block after normal MC processing. To reduce the computational complexity, BIO may not be applied during OBMC processing. That is, BIO is applied in the MC processing of one block when using its own MV, and is not applied in the MC processing when using the MV of neighboring blocks in OBMC processing.
[0182] 11 Prediction fine - tuning by optical flow (PROF) in JVET - N0236 In this contribution, a method for fine - tuning sub - block - based affine motion compensation prediction using optical flow is proposed. After performing sub - block - based affine motion compensation, the prediction samples are fine - tuned by adding the difference derived from the optical flow equation, which is called optical flow prediction fine - tuning (PROF). The proposed method can achieve inter - prediction at the pixel - level granularity without increasing the memory access bandwidth.
[0183] To make the granularity of motion compensation finer, this contribution proposes a method for fine - tuning sub - block - based affine motion compensation prediction using optical flow. After performing sub - block - based affine motion compensation, the luminance prediction samples are fine - tuned by adding the difference derived from the optical flow equation. The proposed PROF is described in the following four steps.
[0184] Step 1) Perform sub-block based affine motion compensation to generate the sub-block prediction I(i,j).
[0185] Step 2) At each sample location, we use a 3-tap filter [-1,0,1] to find the spatial gradient of the subblock prediction g x (i,j) and g y Calculate (i,j).
[0186]
number
[0187]
number
[0188] The subblock prediction is expanded by one pixel on each side for gradient calculation. To reduce memory bandwidth and complexity, the pixels on the expanded boundary are copied from the nearest integer pixel location in the reference picture. Thus, additional interpolation for the padding area is avoided.
[0189] Step 3) Calculate the fine adjustment of the luminance prediction by the optical flow equation.
[0190]
number
[0191] Here, Δv(i,j) is the difference between the pixel MV calculated for sample position v(i,j), represented by (i,j), as shown in FIG. 10, and the sub-block MV of the sub-block MV to which pixel (i,j) belongs.
[0192] Since the affine model parameters and pixel positions with respect to the sub-block center do not change from sub-block to sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let the horizontal and vertical offsets from the pixel position to the center of the sub-block be x and y, then Δv(x,y) can be derived from the following formula.
[0193]
Number
[0194] For the 4-parameter affine model,
[0195]
Number
[0196] For the 6-parameter affine model,
[0197]
Number
[0198] Here, (v 0x , v 0y ), (v 1x , v 1y ), (v 2x , v 2y ) are the top-left, top-right, and bottom-left control point motion vectors, and w, h are the width and height of the CU.
[0199] Step 4) Finally, a fine adjustment of the luminance prediction is added to the sub-block prediction I(i,j). The final prediction I’ is generated as follows.
[0200]
Number
[0201] 12 Deficiencies of Existing Implementations The non-linear ALF (NLALF) in the design of JVET-N0242 has the following problems.
[0202] (1) In NLALF, many clipping operations are required.
[0203] (2) In the ALF based on CTU, when alf_num_available_temporal_filter_sets_luma is 0, there is no available temporal luma filter. However, alf_temporal_index may still be signaled.
[0204] (3) In the ALF based on CTU, when alf_signal_new_filter_chroma is 0, a new filter is not signaled for the chroma component, and it is assumed that the temporal chroma filter is used. However, it is not guaranteed that the temporal chroma filter is available.
[0205] (4) In the ALF based on CTU, alf_num_available_temporal_filter_sets_luma may be larger than the available temporal filter sets.
[0206] 13 Exemplary methods for adaptive loop filtering for video coding Embodiments of the technology of the present disclosure overcome the drawbacks of existing implementations, thereby providing video coding with higher coding efficiency. Techniques for adaptive loop filtering based on the disclosed technology can improve both existing and future video coding standards and are illustrated in the following examples for various embodiments. The examples of the disclosed technology provided below are illustrative of general concepts and should not be construed as limiting. In one example, various features described in these examples can be combined unless explicitly indicated otherwise or unless indicated to the contrary. 1. Instead of clipping the sample difference, it is proposed to apply the clipping operation to the intermediate result during the filtering process. Samples (adjacent or non - adjacent) in the vicinity of the current sample used in the filtering process may be classified into N (N≧1) groups. a. In one example, one or more intermediate results may be calculated for one group, and clipping may be performed on this one or more intermediate results. i. For example, for one group, first calculate the difference between each neighboring pixel and the current pixel, and then these differences may be weighted - averaged using the corresponding ALF coefficients (denoted as wAvgDiff). Clipping may be performed once on wAvgDiff for the group. b. Different clipping parameters may be used for different groups. c. In one example, clipping is applied to the final weighted sum of the filter coefficients multiplied by the sample difference. i. For example, when N = 1, and the clipping function is K(d,b)=min(b,max(-b,d)) and k is the clipping parameter, it may be performed as follows.
[0207]
Equation
[0208] 1) Further, alternatively, the weighted sum is as follows.
[0209]
Equation
[0210] It may be further rounded to an integer value, for example, rounded or shifted without rounding. 2. When filtering one sample, if N (N>1) neighboring samples share one filter coefficient, clipping (e.g., required by non-linear ALF) may be performed once for all N neighboring pixels. a. For example, if I(x+i1,y+j1) and I(x+i2,y+j2) share one filter coefficient w(i1,j1) (and / or one clipping parameter k(i1,j1)), clipping may be performed once as follows.
[0211]
Number
[0212] Using, in Equation (14),
[0213]
Number
[0214] may be replaced.
[0215] i. In one example, i1 and i2 may be in symmetric positions. Also, j1 and j2 may be in symmetric positions. 1. In one example, i1 is equal to (-i2) and j1 is equal to (-j2). ii. In one example, the distance between (x+i1,y+j1) and (x,y) and the distance between (x+i2,j+j2) and (x,y) may be the same. iii. The method disclosed in Brett 2 is activated when the filter shape is in symmetric mode. iv. Further alternatively, the clipping parameter associated with I(x+i1,y+j1) may be signaled / derived from a bit stream called ClipParam, and the above k(i1,j1) may be derived from the signaled clipping parameter such as 2*ClipParam. b. For example, if (i,j) ∈ C shares one filter coefficient w1 (and / or one clipping parameter k1), and C contains N elements, the clipping may be performed once as follows.
[0216]
Number
[0217] Here, k1 is the clipping parameter associated with C, and clipValue*w1 may be used to replace the following term in Equation (14).
[0218]
Number
[0219] i. Further alternatively, the clipping parameter associated with I(x+i,y+j) may be signaled / derived from a bitstream called ClipParam, and k1 may be derived from the signaled clipping parameter such as N*ClipParam. ii. Alternatively,
[0220]
Number
[0221] may be right-shifted and then clipped. c. In one example, the clipping may be performed once for M1 (M1 <= N) of the N neighboring samples. d. In one example, the N neighboring samples may be classified into M2 groups, and clipping may be performed once for each group. e. In one example, this method may be applied to specific or all color components. i. For example, it may be applied to the luminance component. ii. For example, it may be applied to the Cb and / or Cr components. 3. In the present disclosure, a clipping function K(min, max, input) that clips the input to the range [min, max] including min and max may be used. a. In one example, a clipping function K(min, max, input) that clips the input to the range (min, max) excluding min and max may be used for the above black dots. b. In one example, for the above black dots, a clipping function K(min, max, input) that clips the input to be in the range (min, max) including max but not including min may be used. c. In one example, for the above black dots, a clipping function K(min, max, input) that clips the input to be in the range [min, max] including min but not including max may be used. 4. When a temporal ALF coefficient set is not available (e.g., the ALF coefficients have not been encoded / decoded previously, or the encoded / decoded ALF coefficients are marked as "not available"), the signaling indicating which temporal ALF coefficient set is used may be omitted. a. In one example, when the temporal ALF coefficient set is not available, if neither new ALF coefficients nor fixed ALF coefficients are used for CTB / block / tile group / tile / slice / picture, it is presumed that ALF is not permitted for the CTB / block / tile group / tile / slice / picture. i. Further, alternatively, even if in this case it is shown that ALF is applied to the CTB / block / tile group / tile / slice / picture (e.g., alf_ctb_flag is true for the CTU / block), it may be presumed that ALF is ultimately not permitted for the CTB / block / tile group / tile / slice / picture. b. In one example, when the temporal ALF coefficient set is not available, it may be instructed to use only new ALF coefficients or fixed ALF coefficients, etc. for CTB / block / tile group / tile / slice / picture in the compliant bitstream. i. For example, assume that either alf_use_new_filter or alf_use_fixed_filter is true. c. In one example, if the condition is met that for a CTB / block / tile group / tile / slice / picture for which ALF adoption is indicated and when a temporal ALF coefficient set is not available, neither a new ALF coefficient nor a fixed ALF coefficient is indicated to be used, then that bitstream is considered a non - compliant bitstream. i. For example, a bitstream in which both alf_use_new_filter and alf_use_fixed_filter are false is considered a non - compliant bitstream. d. In one example, if alf_num_available_temporal_filter_sets_luma is 0, alf_temporal_index may not be signaled. e. The proposed method may be applied in different ways for different color components. 5. How many temporal ALF coefficient sets can be used for one tile group / tile / slice / picture / CTB / block / video unit depends on the available temporal ALF coefficient sets (denoted as ALF avai ), for example, on the ALF coefficient sets that were encoded / decoded before being marked as "available". a. In one example, for one tile group / tile / slice / picture / CTB / block, ALF avai the following temporal ALF coefficient sets can be used. b. For one tile group / tile / slice / picture / CTB / block, with N >= 0, less than or equal to min(N, ALF avai ) temporal ALF coefficient sets can be used. For example, N = 5. 6. The new set of ALF coefficients may be marked as "usable" after encoding / decoding. On the other hand, all "usable" sets of ALF coefficients may be marked as "unusable" when encountering an Intra Random Access Point (IRAP) access unit or / and an IRAP picture, or / and an Instantaneous Decoding Refresh (IDR) access unit or / and an IDR picture. a. A "usable" set of ALF coefficients may be used as a temporal ALF coefficient set for subsequent encoded pictures / tiles / tile groups / slices / CTBs / blocks. b. A "usable" set of ALF coefficients may be held in one list of ALF coefficient sets of maximum size equal to N (N>0). i. The list of ALF coefficient sets may be held in a first-in first-out order. c. If marked as "unusable", the related ALF APS information is removed from the bitstream or replaced with other ALF APS information. 7. One list of ALF coefficient sets may be held for each temporal layer. 8. One list of ALF coefficient sets may be held for K neighboring temporal layers. 9. Different lists of ALF coefficient sets may be held for different pictures, based on whether only prediction from the previous picture (in display order) is performed. a. For example, one list of ALF coefficient sets may be held for pictures predicted only from the previous picture. b. For example, one list of ALF coefficient sets may be held for pictures predicted from both the previous and the subsequent pictures. 10. The list of ALF coefficient sets may be emptied after encountering an IRAP access unit and / or an IRAP picture and / or an IDR access unit and / or an IDR picture. 11. Different lists of ALF coefficient sets may be held for different color components. a. In one example, one set of ALF coefficients is retained for the luminance component. b. In one example, one set of ALF coefficients is retained for the Cb and / or Cr components. 12. One set of ALF coefficients may be retained, but different indices (or priorities) may be assigned to the entries in the list for different pictures / tile groups / tiles / slices / CTUs. a. In one example, an up-index for the up absolute temporal layer difference between the current picture / tile group / tile / slice / CTU may be assigned to the ALF coefficient set. b. In one example, an ascending index for the ascending absolute POC (Picture Order Count) difference between the current picture / tile group / tile / slice / CTU may be assigned to the ALF coefficient set. c. In one example, there are K sets of ALF coefficients permitted by the current picture / tile group / tile / slice / CTU, and these K sets of ALF coefficients may be the K sets of ALF coefficients with the smallest indices. d. In one example, the indication of which temporal ALF coefficient set the current picture / tile group / tile / slice / CTU uses may depend on the assigned index instead of the original entry index in the list. 13. The neighboring samples used in ALF may be classified into K (K >= 1) groups, and one set of clipping parameters may be signaled for each group. 14. The clipping parameters may be predefined for specific or all fixed ALF filter sets. a. Alternatively, the clipping parameters may be signaled for specific or all fixed filter sets used by the current tile group / slice / picture / tile. i. In one example, the clipping parameters may be signaled only for a specific color component (e.g., the luminance component). b. Alternatively, when using a fixed ALF filter set, clipping may not be performed. i. In one example, clipping may be performed on a particular color component and not on other color components. 15. The clipping parameters may be stored with the ALF coefficients and may be inherited by following the encoded CTU / CU / tile / tile group / slice / picture. a. In one example, when a CTU / CU / tile / tile group / slice / picture uses a temporal ALF coefficient set, the corresponding ALF clipping parameters may be used. i. In one example, the clipping parameters may be inherited only for a particular color component (e.g., the luminance component). b. Alternatively, when a CTU / CU / tile / tile group / slice / picture uses a temporal ALF coefficient set, the clipping parameters may be signaled. i. In one example, the clipping parameters may be signaled only for a particular color component (e.g., the luminance component). c. In one example, the clipping parameters may be inherited for a particular color component and signaled for other color components. d. In one example, when a temporal ALF coefficient set is used, clipping is not performed. i. In one example, clipping may be performed on a particular color component and not on other color components. 16. Whether to use non - linear ALF may depend on the type of the ALF filter set (e.g., fixed ALF filter set, temporal ALF filter set, or signaled ALF coefficient set). a. In one example, when the current CTU uses a fixed ALF filter set or a temporal ALF filter set (i.e., the previously signaled filter set is used), non - linear ALF may not be used for the current CTU. b. In one example, when alf_luma_use_fixed_filter_flag is 1, non-linear ALF may be used for the current slice / tile group / tile / CTU. 17. Non-linear ALF clipping parameters may be signaled conditionally based on the type of ALF filter set (e.g., fixed ALF filter set, temporal ALF filter set, or signaled ALF coefficient set). a. In one example, non-linear ALF clipping parameters may be signaled for all ALF filter sets. b. In one example, non-linear ALF clipping parameters may be signaled only for the signaled ALF filter coefficient set. c. In one example, non-linear ALF clipping parameters may be signaled only for the fixed ALF filter coefficient set.
[0222] The examples described above may be included in the context of the methods described below, e.g., methods 1110, 1120, 1130, 1140, 1150, and 1160, which may be implemented in a video decoder and / or a video encoder.
[0223] FIG. 11A shows a flowchart of an exemplary video processing method. Method 1110 includes, in operation 1112, performing filtering processing on the current video block of a video, using filter coefficients and including two or more operations with at least one intermediate result.
[0224] Method 1110 includes, in operation 1114, applying a clipping operation to the at least one intermediate result.
[0225] Method 1110 includes, in operation 1116, performing a conversion between the current video block and the bitstream representation of the video based on at least one intermediate result. In some embodiments, the at least one intermediate result is based on the sum of the weightings of the filter coefficients and the difference between the current sample of the current video block and samples in the vicinity of the current sample.
[0226] FIG. 11B shows a flowchart of an exemplary video processing method. Method 1120 includes, in operation 1122, encoding a current video block of a video into a bitstream representation of the video, where the current video block is encoded with an adaptive loop filter (ALF).
[0227] Method 1120 includes, in operation 1124, selectively including an indication of a set of temporal adaptive filters within the one or more sets of temporal adaptive filters in the bitstream representation based on the availability or use of the one or more sets of temporal adaptive filters.
[0228] FIG. 11C shows a flowchart of an exemplary video processing method. This method 1130 includes, in operation 1132, determining the availability or use of one or more sets of the temporal adaptive filters that are applicable to a current video block of a video encoded with an adaptive loop filter (ALF) based on an indication of a set of temporal adaptive filters in the bitstream representation of the video.
[0229] Method 1130 includes, in operation 1134, generating a current video block decoded from the bitstream representation by selectively applying the set of temporal adaptive filters based on the determining.
[0230] FIG. 11D shows a flowchart of an exemplary video processing method. This method 1140 includes, in operation 1142, determining the number of temporal adaptive loop filtering (ALF) coefficient sets for a current video block encoded with an adaptive loop filter, based on an available set of temporal ALF coefficients, where the available set of temporal ALF coefficients has been encoded or decoded prior to the determining, and the number of ALF coefficient sets is used for a tile group, tile, slice, picture, coding tree block (CTB), or video unit that makes up the current video block.
[0231] Method 1140 includes, in operation 1144, performing a conversion between the current video block and a bitstream representation of the current video block, based on the number of the temporal ALF coefficient sets.
[0232] FIG. 11E shows a flowchart of an exemplary video processing method. This method 1150 includes, in operation 1152, determining that an indication of adaptive loop filtering (ALF) in a header of a video region of the video is equal to an indication of ALF in an adaptive parameter set (APS) network abstraction layer (NAL) unit associated with the bitstream representation, for performing a conversion between a current video block of the video and the bitstream representation of the video.
[0233] Method 1150 includes, in operation 1154, performing the conversion.
[0234] FIG. 11F shows a flowchart of an exemplary video processing method. This method 1160 includes, in operation 1162, selectively enabling a non-linear adaptive loop filtering (ALF) operation, based on a type of adaptive loop filter used in a video region of the video, for performing a conversion between a current video block and a bitstream representation of the video.
[0235] This method 1160 includes, in operation 1164, performing the conversion after the selectively enabling.
[0236] 10 Exemplary Implementations of the Disclosed Technology 10.1 Embodiment #1 Assume that one ALF coefficient set list is maintained for each of luminance and chroma, and the sizes of the two lists are lumaALFSetSize and chromaALFSetSize, respectively. The maximum sizes of the ALF coefficient set lists are lumaALFSetMax (e.g., lumaALFSetMax equals 5) and chromaALFSetMax (e.g., chromaALFSetMax equals 5), respectively.
[0237] Newly added parts are enclosed in double thick curly braces, i.e., {{a}} means that "a" has been added, and deleted parts are enclosed in double square brackets, i.e., [[a]] means that "a" has been deleted. 7.3.3.2 Adaptive Loop Filter Data Syntax
[0238] [Table 12]
[0239] 7.3.4.2 Encoding Tree Unit Syntax
[0240] [Table 13]
[0241] Specify that a new luminance filter set is signaled when alf_signal_new_filter_luma equals 1. Indicate that a new luminance filter set is not signaled when alf_signal_new_filter_luma equals 0. If it does not exist, it is 0. When alf_luma_use_fixed_filter_flag is equal to 1, it indicates that a fixed filter set is used to signal the adaptive loop filter. When alf_luma_use_fixed_filter_flag is equal to 0, it indicates that no fixed filter is used to signal the adaptive loop filter. alf_num_available_temporal_filter_sets_luma indicates the number of available temporal filter sets that can be used for the current slice. This number can be from 0 [[5]]{{lumaALFSetSize}}. If it does not exist, it is 0. {{When alf_num_available_temporal_filter_sets_luma is 0, there is a constraint that either alf_signal_new_filter_luma or alf_luma_use_fixed_filter_flag must be 1.}} When alf_signal_new_filter_chroma is equal to 1, it specifies that a new chroma filter is signaled. When alf_signal_new_filter_chroma is equal to 0, it indicates that no new chroma filter is signaled. {{There is a constraint that alf_signal_new_filter_chroma must be 1 when chromaALFSetSize is 0.}}
[0242] 10.2 Embodiment #2 Assume that one ALF coefficient set list is maintained for each of luminance and chroma, and the sizes of the two lists are lumaALFSetSize and chromaALFSetSize respectively. The maximum sizes of the ALF coefficient set lists are lumaALFSetMax (for example, lumaALFSetMax is equal to 5) and chromaALFSetMax (for example, chromaALFSetMax is equal to 5) respectively.
[0243] Newly added parts are enclosed by double thick braces, i.e., {{a}} means that "a" has been added, and deleted parts are enclosed by double square brackets, i.e., [[a]] means that "a" has been deleted. 7.3.3.2 Adaptive Loop Filter Data Syntax
[0244]
Table 14
[0245] 7.3.4.2 Encoding Tree Unit Syntax
[0246]
Table 15
[0247] When alf_signal_new_filter_luma is equal to 1, it specifies that a new luminance filter set is signaled. When alf_signal_new_filter_luma is equal to 0, it indicates that a new luminance filter set is not signaled. If it does not exist, it is 0. When alf_luma_use_fixed_filter_flag is equal to 1, it indicates that a fixed filter set is used to signal the adaptive loop filter. When alf_luma_use_fixed_filter_flag is equal to 0, it indicates that no fixed filter is used to signal the adaptive loop filter. alf_num_available_temporal_filter_sets_luma indicates the number of available temporal filter sets that can be used for the current slice, and this number can be from 0 [[5]]{{lumaALFSetSize}}. If it does not exist, it is 0. {There is a constraint that either alf_signal_new_filter_luma or alf_luma_use_fixed_filter_flag must be 1 when alf_num_available_temporal_filter_sets_luma is 0.} When alf_signal_new_filter_chroma is equal to 1, it specifies that a new chroma filter is signaled. When alf_signal_new_filter_chroma is equal to 0, it indicates that a new chroma filter is not signaled. {There is a constraint that alf_signal_new_filter_chroma must be 1 when chromaALFSetSize is 0.}
[0248] In some embodiments, the following technical solutions can be implemented.
[0249] A1. For a current video block of a video, perform filtering processing including two or more operations with filter coefficients and involving at least one intermediate result, apply a clipping operation to the at least one intermediate result, and perform a conversion between the current video block and a bitstream representation of the video based on the at least one intermediate result, where the at least one intermediate result is based on a sum of weights of the filter coefficients and a difference between a current sample of the current video block and samples in the vicinity of the current sample, a video processing method.
[0250] A2. The method according to solution A1, further including classifying samples in the vicinity of the current sample into a plurality of groups for the current sample, and applying the clipping operation to intermediate results in each of the plurality of groups using different parameters.
[0251] A3. The method according to solution A2, wherein the at least one intermediate result includes a weighted average of the differences between the current sample and the neighboring samples in each of the plurality of groups.
[0252] A4. The method according to solution A1, wherein a plurality of neighboring samples of the samples of the current video block share one filter coefficient, and the clipping operation is applied once to each of the plurality of neighboring samples.
[0253] A5. The method according to solution A4, wherein the positions of at least two samples among the plurality of neighboring samples are symmetric with respect to the sample of the current video block.
[0254] A6. The method according to solution A4 or A5, wherein the filter shape associated with the filtering process is in a symmetric mode.
[0255] A7. The method according to any one of solutions A4 to A6, wherein one or more parameters of the clipping operation are signaled in the bitstream representation.
[0256] A8. The method according to solution A1, wherein the samples of the current video block include N neighboring samples, the clipping operation is applied once to M1 neighboring samples of the N neighboring samples, and M1 and N are positive integers and M1 ≤ N.
[0257] A9. The method according to solution A1, further comprising classifying the N neighboring samples of the sample into M2 groups with respect to the sample of the current video block, and the clipping operation is applied once to each of the M2 groups, and M2 and N are positive integers.
[0258] A10. The method according to solution A1, wherein the clipping operation is applied to the luminance component associated with the current video block.
[0259] A11. The clipping operation is the method described in Solution A1 that is applied to the Cb component or the Cr component associated with the current video block.
[0260] A12. The clipping operation is defined as K(min, max, input), where input is the input to the clipping operation, min is the nominal minimum value of the output of the clipping operation, and max is the nominal maximum value of the output of the clipping operation, according to the method described in any of Solutions A1 to A11.
[0261] A13. The method described in Solution A12, where the actual maximum value of the output of the clipping operation is smaller than the nominal maximum value, and the actual minimum value of the output of the clipping operation is larger than the nominal minimum value.
[0262] A14. The method described in Solution A12, where the actual maximum value of the output of the clipping operation is equal to the nominal maximum value, and the actual minimum value of the output of the clipping operation is larger than the nominal minimum value.
[0263] A15. The method described in Solution A12, where the actual maximum value of the output of the clipping operation is smaller than the nominal maximum value, and the actual minimum value of the output of the clipping operation is equal to the nominal minimum value.
[0264] A16. The method described in Solution A12, where the actual maximum value of the output of the clipping operation is equal to the nominal maximum value, and the actual minimum value of the output of the clipping operation is equal to the nominal minimum value.
[0265] A17. The filtering process includes an adaptive loop filtering (ALF) process composed of a plurality of ALF filter coefficient sets, according to the method described in Solution A1.
[0266] A18. For one or more of the plurality of ALF filter coefficient sets, at least one parameter for the clipping operation is predefined, according to the method described in Solution A17.
[0267] A19. The method according to solution A17, wherein in the bitstream representation of a tile group, slice, picture or tile constituting the current video block, at least one parameter for the clipping operation is signaled.
[0268] A20. The method according to solution A19, wherein the at least one parameter is signaled only for one or more color components associated with the current video block.
[0269] A21. At least one of the plurality of ALF filter coefficient sets and one or more parameters for the clipping operation are stored in the same memory location, and the at least one or more of the plurality of ALF filter coefficient sets or the one or more parameters are inherited by a coding tree unit (CTU), coding unit (CU), tile, tile group, slice, or picture including the current video block. The method according to solution A17.
[0270] A22. If it is determined that the clipping operation uses a temporal ALF coefficient set for the filtering process of a CTU, CU, tile, tile group, slice, or picture constituting the current video block, it is configured to use one or more parameters corresponding to the temporal ALF coefficient set of the plurality of ALF filter coefficient sets. The method according to solution A21.
[0271] A23. The method according to solution A22, wherein the one or more parameters corresponding to the temporal ALF coefficient set are used only for one or more color components associated with the current video block.
[0272] A24. One or more parameters corresponding to the temporal ALF filter coefficient sets of the plurality of ALF filter coefficient sets are signaled in the bitstream representation when the temporal ALF coefficient set is determined to be used in the filtering process of the CTU, the CU, the tile, the tile group, the slice, or the picture that constitutes the current video block, according to the method described in Solution A21.
[0273] A25. The method according to Solution A24, wherein the one or more parameters corresponding to the temporal ALF coefficient set are signaled only for one or more color components associated with the current video block.
[0274] A26. Signal the parameters of the first set of the one or more parameters for the first color component associated with the current video block, and inherit the parameters of the second set of the one or more parameters for the second color component associated with the current video block, according to the method described in Solution A21.
[0275] A27. The method according to any one of Solutions A1 to A26, wherein the transformation generates the current block from the bitstream representation.
[0276] A28. The method according to any one of Solutions A1 to A26, wherein the transformation includes generating the bitstream representation from the current video block.
[0277] A29. An apparatus comprising a processing device and a non-transitory memory storing instructions therein, wherein when the instructions are implemented by the processing device, the processing device is caused to implement the method according to any one of Solutions A1 to A28 in a video system apparatus.
[0278] A30. A computer program product stored in a non-transitory computer-readable medium, comprising program code for executing the method according to any one of Solutions A1 to A28.
[0279] In some embodiments, the following technical solutions can be implemented.
[0280] B1. Encoding the current video block of the video into the bitstream representation of the video, where the current video block is encoded with an Adaptive Loop Filter (ALF), and selectively including an indication of the set of temporal adaptive filters within the one or more sets of temporal adaptive filters in the bitstream representation based on the availability or use of one or more sets of temporal adaptive filters. A video processing method comprising the above.
[0281] B2. The method according to solution B1, wherein when the set of temporal adaptive filters is not available, the indication of the set is excluded from the bitstream representation.
[0282] B3. The method according to solution B1 or B2, wherein when the set of temporal adaptive filters is not available, the indication of the set is included in the bitstream representation. The method according to any one of solutions B1 to B3, wherein when none of the one or more sets of temporal adaptive filters are available, this indication is excluded from the bitstream representation.
[0283] B4. The method according to any one of solutions B1 to B3, wherein each of the one or more sets of temporal adaptive filters is associated with one filter index.
[0284] B5. The method according to any one of solutions B1 to B3, wherein when none of the one or more sets of temporal adaptive filters are available, the indication to use a fixed filter is made equal to true.
[0285] B6. The method according to any one of solutions B1 to B3, wherein when none of the one or more sets of temporal adaptive filters are available, the indication to use a temporal adaptive filter is made equal to false.
[0286] B7. If none of the sets of the one or more temporal adaptive filters is available, the method according to any of Solutions B1 to B3, wherein an indication of the index of the fixed filter is included in the bitstream representation.
[0287] B8. Determining the availability or use of one or more sets of the temporal adaptive filters, each set comprising the set of temporal adaptive filters applicable to a current video block of video encoded with an adaptive loop filter (ALF), based on an indication of the set of temporal adaptive filters in the bitstream representation of the video; and generating the current video block decoded from the bitstream representation by selectively applying the set of temporal adaptive filters based on the determining.
[0288] B9. The method according to Solution B8, wherein when the set of the temporal adaptive filters is not available, the generating is performed without applying the set of the temporal adaptive filters.
[0289] B10. The method according to Solution B8 or B9, wherein when the set of the temporal adaptive filters is not available, performing the generating includes applying the set of the temporal adaptive filters.
[0290] B11. The method according to any of Solutions B1 to B10, wherein the one or more sets of the temporal adaptive filters are included in one adaptive parameter set (APS), and the indication is one APS index.
[0291] B12. The method according to any of Solutions B1 to B10, further comprising determining a filter index for at least one of the temporal adaptive filters of the one or more sets of the temporal adaptive filters based on gradient calculations in different directions.
[0292] B13. None of the one or more sets of the temporal adaptive filter are available, and it is determined that a new ALF coefficient set and a fixed ALF coefficient set are not used in the coding tree block (CTB), block, tile group, tile, slice, or picture that constitutes the current video block, and based on the determination, it is inferred that adaptive loop filtering is invalid, the method according to any of Solutions B1 to B11, further comprising.
[0293] B14. The bitstream representation includes a first indication of the use of a new ALF coefficient set and a second indication of the use of a fixed ALF coefficient set in response to at least one of the one or more sets of temporal adaptive filters being unavailable, and one of the first indication and the second indication is exactly true in the bitstream representation, the method according to any of Solutions B1 to B11.
[0294] B15. The bitstream representation complies with the format rules associated with the operation of the ALF, the method according to Solution B14.
[0295] B16. In response to none of the one or more sets of temporal adaptive filters being available, the bitstream representation includes an indication that the ALF is enabled and that a new ALF coefficient set and a fixed ALF coefficient set are not used in the coding tree block (CTB), block, tile group, tile, slice, or picture that constitutes the current video block, the method according to any of Solutions B1 to B11.
[0296] B17. The bitstream representation does not comply with the format rules associated with the operation of the ALF, the method according to Solution B16.
[0297] B18. The ALF is applied to one or more color components associated with the current video block, the method according to any of Solutions B1 to B17.
[0298] For a current video block encoded with an adaptive loop filter, determining the number of sets of temporal adaptive loop filtering (ALF) coefficients based on an available set of temporal ALF coefficients, wherein the available set of temporal ALF coefficients has been encoded or decoded prior to said determining, and the number of sets of ALF coefficients is determined for a tile group, tile, slice, picture, coding tree block (CTB), or video unit that constitutes the current video block; and performing a conversion between the current video block and a bitstream representation of the current video block based on said number of sets of temporal ALF coefficients. A video processing method.
[0299] Solution B20: The method according to solution B19, wherein the maximum number of sets of temporal ALF coefficients is set equal to the number of available sets of temporal ALF coefficients.
[0300] Solution B21: The method according to solution B20, wherein the number of sets of temporal ALF coefficients is set equal to the smaller of the number of available sets of temporal ALF coefficients and a predefined number N, where N is an integer and N ≥ 0.
[0301] Solution B22: The method according to solution B21, wherein N = 5.
[0302] For a current video block encoded with an adaptive loop filter, processing one or more new sets of adaptive loop filtering (ALF) coefficients as part of a conversion between the current video block of a video and a bitstream representation of the video; and after said processing, designating said one or more new sets of ALF coefficients as available sets of ALF coefficients. A video processing method.
[0303] Method according to solution B23, further comprising encountering an Intra Random Access Point (IRAP) access unit, an IRAP picture, an Instantaneous Decoding Refresh (IDR) access unit, or an IDR picture, and designating the available ALF coefficient set as an unavailable ALF coefficient set based on the encounter.
[0304] Method according to solution B23 or B24, wherein at least one of the available ALF coefficient sets is a temporal ALF coefficient set for a subsequent video block of the current video block.
[0305] Method according to any one of solutions B23 to B25, wherein the available ALF coefficient set is held in an ALF coefficient set list having a maximum size of N, where N is an integer.
[0306] Method according to solution B26, wherein the ALF coefficient set list is held in a First In First Out (FIFO) order.
[0307] Method according to any one of solutions B1 to B27, holding one ALF coefficient set list for each temporal layer associated with the current video block.
[0308] Method according to any one of solutions B1 to B27, holding one ALF coefficient set list for each of K neighboring temporal layers associated with the current video block.
[0309] Method according to any one of solutions B1 to B27, holding a first ALF coefficient set list for the current picture including the current video block and a second ALF coefficient set list for a subsequent picture of the current picture.
[0310] Solution B30: The method according to solution B30, wherein, based on the current picture, a prediction is made of the picture image that follows the current picture, and the first list of ALF coefficient sets is the same as the second list of ALF coefficient sets.
[0311] Solution B30: The method according to solution B30, wherein, based on the picture that follows the current picture and the picture that precedes the current picture, a prediction is made of the current picture, and the first list of ALF coefficient sets is the same as the second list of ALF coefficient sets.
[0312] Solution B23: The method according to solution B23, further comprising encountering an intra random access point (IRAP) access unit, an IRAP picture, an instantaneous decoding refresh (IDR) access unit, or an IDR picture, and after the encounter, emptying one or more lists of ALF coefficient sets.
[0313] Solution B23: The method according to solution B23, wherein different lists of ALF coefficient sets are maintained for different color components associated with the current video block.
[0314] Solution B34: The method according to solution B34, wherein the different color components include one or more of a luminance component, a Cr component, and a Cb component.
[0315] Solution B23: The method according to solution B23, wherein one list of ALF coefficient sets is maintained for a plurality of pictures, tile groups, tiles, slices, or coding tree units (CTUs), and the indexing of the one list of ALF coefficient sets is different for each of the plurality of pictures, tile groups, tiles, slices, or coding tree units (CTUs).
[0316] B37. The indexing is in ascending order and is based on a first temporal layer index associated with the current video block and a second temporal layer index associated with the current picture, tile group, tile, slice, or coding tree unit (CTU) that constitutes the current video block, the method according to solution B36.
[0317] B38. The indexing is in ascending order and is based on a picture order count (POC) associated with the current video block and a second POC associated with the current picture, tile group, tile, slice, or coding tree unit (CTU) that constitutes the current video block, the method according to solution B36.
[0318] B39. The indexing includes the smallest index assigned to the available ALF coefficient set, the method according to solution B36.
[0319] B40. The transformation includes a clipping operation, and the method further includes classifying samples in the vicinity of the samples of the current video block into a plurality of groups and performing a clipping operation for each of the plurality of groups using a single set of parameters signaled in the bitstream representation, the method according to solution B23.
[0320] B41. The transformation includes a clipping operation, and a set of parameters for the clipping operation is predefined for the one or more new ALF coefficient sets, the method according to solution B23.
[0321] B42. The transformation includes a clipping operation, and a set of parameters for the clipping operation is signaled in the bitstream representation for the one or more new ALF coefficient sets, the method according to solution B23.
[0322] For conversion between the current video block of a video and the bitstream representation of the video, determining that an indication of adaptive loop filtering (ALF) in a header of a video region of the video is equal to an indication of ALF in an adaptive parameter set (APS) network abstraction layer (NAL) unit associated with the bitstream representation, and performing the conversion. A video processing method including these steps.
[0323] The method according to solution B43, wherein the video region is one picture.
[0324] The method according to solution B43, wherein the video region is one slice.
[0325] For conversion between the current video block of a video and the bitstream representation of the video, selectively enabling a non-linear adaptive loop filtering (ALF) operation based on a type of adaptive loop filter used in a video region of the video, and performing the conversion after the selective enabling. A video processing method including these steps.
[0326] The method according to solution B46, wherein the video region is a coding tree unit (CTU), and when it is determined that the type of the adaptive loop filter includes a fixed ALF set or a temporal ALF set, the non-linear ALF operation is disabled.
[0327] The method according to solution B46, wherein the video region is a slice, a tile group, a tile or a coding tree unit (CTU), and the non-linear ALF operation is enabled when it is determined that the type of the adaptive loop filter includes a fixed ALF set.
[0328] The method according to solution B46, further comprising selectively signaling one or more clipping parameters in the bitstream representation for the non-linear ALF operation.
[0329] Solution B50. The method according to solution B49, wherein the one or more clipping parameters are signaled.
[0330] Solution B51. The method according to solution B49, wherein the one or more clipping parameters are signaled for an ALF filter coefficient set signaled in the bitstream representation.
[0331] Solution B52. The method according to solution B49, wherein the one or more clipping parameters are signaled when it is determined that the type of the adaptive loop filter includes a fixed ALF set.
[0332] Solution B53. The method according to any one of solutions B19 to B52, wherein the transformation generates the current block from the bitstream representation.
[0333] Solution B54. The method according to any one of solutions B19 to B52, wherein the transformation includes generating the bitstream representation from the current video block.
[0334] Solution B55. An apparatus comprising a processing device and a non-transitory memory storing instructions therein, wherein when the instructions are implemented by the processing device, the processing device is caused to perform the method according to any one of solutions B1 to A54 in a video system apparatus.
[0335] Solution B56. A computer program product stored on a non-transitory computer-readable medium, comprising program code for performing the method according to any one of solutions B1 to B54.
[0336] In some embodiments, the following technical solutions can be implemented.
[0337] C1. Performing a filtering process on the current video block that includes two or more operations with at least one intermediate result, applying a clipping operation to the at least one intermediate result, and performing a conversion between the current video block and a bitstream representation of the current video block based on the filtering operation, a video processing method.
[0338] C2. Further including classifying neighboring samples among the samples of the current video block into a plurality of groups, and applying the clipping operation to intermediate results in each of the plurality of groups using different parameters, the method according to solution C1.
[0339] C3. The at least one intermediate result includes a weighted average of differences between the current sample and the neighboring samples in each of the plurality of groups, the method according to solution C2.
[0340] C4. The filtering process uses filter coefficients, and the at least one intermediate result includes the sum of the weights of the filter coefficients and the differences between the current sample and the neighboring samples, the method according to solution C2.
[0341] C5. A plurality of neighboring samples among the samples of the current video block share filter coefficients, and applying a single clipping operation to each of the plurality of neighboring samples, the method according to solution C1.
[0342] C6. The filter shape associated with the filtering operation is in a symmetric mode, the method according to solution C5.
[0343] C7. One or more parameters of the clipping operation are signaled in the bitstream representation, the method according to solution C5 or C6.
[0344] C8. The clipping operation is defined as K(min, max, input), where input is the input to the clipping operation, min is the nominal minimum value of the output of the clipping operation, and max is the nominal maximum value of the output of the clipping operation, the method according to any one of solutions C1 to C7.
[0345] C9. The method according to solution C8, wherein the actual maximum value of the output of the clipping operation is smaller than the nominal maximum value, and the actual minimum value of the output of the clipping operation is larger than the nominal minimum value.
[0346] C10. The method according to solution C8, wherein the actual maximum value of the output of the clipping operation is equal to the nominal maximum value, and the actual minimum value of the output of the clipping operation is larger than the nominal minimum value.
[0347] C11. The method according to solution C8, wherein the actual maximum value of the output of the clipping operation is smaller than the nominal maximum value, and the actual minimum value of the output of the clipping operation is equal to the nominal minimum value.
[0348] C12. The method according to solution C8, wherein the actual maximum value of the output of the clipping operation is equal to the nominal maximum value, and the actual minimum value of the output of the clipping operation is equal to the nominal minimum value.
[0349] C13. A video processing method including performing conversion between the current video block and the bitstream representation of the current video block such that the bitstream representation omits an indication of the time adaptive loop filtering coefficient set based on the unavailability of the time adaptive loop filtering coefficient set.
[0350] In a coded tree block (CTB), block, tile group, tile, slice, or picture that constitutes the current video block, determining not to use new adaptive loop filtering (ALF) coefficients and fixed ALF coefficients, and further including inferring that adaptive loop filtering is disabled, the method according to solution C13.
[0351] C15. The compliant bitstream is the method according to solution C13, including an indication of new adaptive loop filtering (ALF) coefficients or an indication of fixed ALF coefficients.
[0352] C16. For the current video block, determining the number of one or more temporal adaptive loop filtering (ALF) coefficient sets based on the available set of temporal ALF coefficients, where the available set of temporal ALF coefficients has been encoded or decoded prior to the determination, and performing a conversion between the current video block and the bitstream representation of the current video block based on the one or more temporal ALF coefficient sets, the video processing method.
[0353] C17. The maximum number of the one or more temporal ALF coefficient sets is ALF available The method according to solution 16.
[0354] C18. The number of the one or more temporal ALF coefficient sets is min(N, ALF available ) where N is an integer and N≥0, the method according to solution C17.
[0355] C19. N = 5, the method according to solution C18.
[0356] C20. Processing one or more new adaptive loop filtering (ALF) coefficient sets for the current video block, and after this processing, designating the one or more new ALF coefficient sets as available ALF coefficient sets, and based on the available ALF coefficient sets, performing a conversion between the current video block and the bitstream representation of the current video block. A video processing method including the steps of.
[0357] C21. Further including encountering an intra-random access point (IRAP) access unit, an IRAP picture, an instantaneous decoding refresh (IDR) access unit, or an IDR picture, and designating the available ALF coefficient sets as unavailable ALF coefficient sets. The method according to solution C20.
[0358] C22. The method according to solution C20 or C21, wherein the available ALF coefficient sets are temporal ALF coefficient sets for subsequent video blocks of the current video block.
[0359] C23. The method according to any one of solutions C20 to C22, wherein the available ALF coefficient sets are held in an ALF coefficient set list having a maximum size of N, where N is an integer.
[0360] C24. The method according to solution C23, wherein the ALF coefficient set list is held in a first-in first-out (FIFO) order.
[0361] C25. Holding one ALF coefficient set list for each temporal layer associated with the current video block. The method according to any one of solutions C13 to C24.
[0362] C26. Holding one ALF coefficient set list for each of K neighboring temporal layers associated with the current video block. The method according to any one of solutions C13 to C24.
[0363] The method according to any one of Solutions C13 to C24, comprising: retaining a first set of ALF coefficients for the current picture including the current video block, and retaining a second set of ALF coefficients for a subsequent picture of the current picture.
[0364] The method according to Solution C27, comprising: predicting the picture image subsequent to the current picture based on the current picture, and the first set of ALF coefficients being the same as the second set of ALF coefficients.
[0365] The method according to Solution C20, further comprising: encountering an Intra Random Access Point (IRAP) access unit, an IRAP picture, an Instantaneous Decoding Refresh (IDR) access unit, or an IDR picture, and after the encounter, emptying one or more sets of ALF coefficient lists.
[0366] The method according to Solution C20, comprising: retaining different sets of ALF coefficients for different color components of the current video block.
[0367] An apparatus comprising a processing device and a non-transitory memory storing instructions therein, wherein when the instructions are implemented by the processing device, the processing device is caused to implement the method according to any one of Solutions C1 to C30 in a video system apparatus.
[0368] A computer program product stored in a non-transitory computer-readable medium, comprising program code for executing the method according to any one of Solutions C1 to C30.
[0369] FIG. 12 is a block diagram of a video processing apparatus 1200. The apparatus 1200 may be used to implement one or more of the methods described herein. The apparatus 1200 may be implemented by, for example, a smartphone, a tablet, a computer, an IoT (Internet of Things) receiver, or the like. The apparatus 1200 may include one or more processing devices 1202, one or more memories 1204, and video processing hardware 1206. The one or more processing devices 1202 may be configured to implement one or more of the methods described herein (including, but not limited to, methods 1100 and 1150). The memory(ies) 1204 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 1206 may be used to implement the techniques described herein in a hardware circuit.
[0370] In some embodiments, the video encoding method may be implemented using an apparatus implemented on a hardware platform, as described with reference to FIG. 12.
[0371] Some embodiments of the disclosed techniques include determining or deciding to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder uses or implements this tool or mode when processing one video block, but based on the use of this tool or mode, it is not necessarily required to modify the resulting bitstream. That is, the conversion from a video block to a bitstream representation of the video uses this video processing tool or mode when the video processing tool or mode is enabled based on a determination or decision. In another example, when a video processing tool or mode is enabled, the decoder knows that the bitstream has been modified based on the video processing tool or mode and processes the bitstream. That is, the conversion from the bitstream representation of the video to the video block is performed using the video processing tool or mode enabled based on a determination or decision.
[0372] Some embodiments of the disclosed technology include determining or deciding to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder does not use this tool or mode when converting a block of video into a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder knows that the bitstream has not been modified using the enabled video processing tool or mode based on the determination or decision, and processes the bitstream.
[0373] FIG. 13 is a block diagram showing an exemplary video processing system 1300 in which various technologies disclosed herein may be implemented. Various implementations may include some or all of the modules of system 1300. System 1300 may include an input unit 1302 for receiving video content. The video content may be received in an unprocessed or uncompressed format, such as 8 - or 10 - bit multi - module pixel values, or in a compressed or encoded format. Input unit 1302 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet®, Passive Optical Network (PON), etc., and wireless interfaces such as Wi - Fi® or cellular interfaces.
[0374] System 1300 may include an encoding module 1304 that can implement various encoding or encoding methods described herein. The encoding module 1304 may reduce the average bitrate of the video from the input unit 1302 to the output of the encoding module 1304 and generate an encoded representation of the video. Thus, this encoding technique may be referred to as video compression or video transcoding technique. The output of the encoding module 1304 may be stored or transmitted via connected communication as represented by module 1306. The bitstream (or encoded) representation of the video received, stored, or communicated in the input unit 1302 may be used by module 1308 to generate pixel values or a displayable video to be transmitted to the display interface unit 1310. The process of generating a video that a user can view from the bitstream representation may be referred to as video decompression (video expansion). Further, although specific video processing operations are referred to as "encoding" operations or tools, it will be understood that the encoding tools or operations and the corresponding decoding tools or operations that reverse the result of decoding are performed by a decoder.
[0375] Examples of the peripheral bus interface unit or the display interface unit may include a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI (registered trademark)) or a DisplayPort, etc. Examples of the storage interface include Serial Advanced Technology Attachment (SATA), PCI, IDE interface, etc. The techniques described herein may be implemented in various electronic devices such as mobile phones, notebook computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0376] Although specific embodiments of the technology of the present disclosure have been described for purposes of illustration above, it will be understood that various modifications are possible without departing from the scope of the present invention. Thus, the technology of the present disclosure is not limited except as by the appended claims.
[0377] The subject matter and the implementation forms of the functional operations described in this patent specification may be implemented in various systems, digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or may be implemented in a combination of one or more of them. The implementation forms of the subject matter described in this specification can be implemented as one or more modules of a computer program product, that is, computer program instructions encoded on a tangible non-transportable computer-readable medium for execution by a data processing device or for controlling the operation of a data processing device. This computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that provides a machine-readable propagated signal, or a combination of one or more of these. The term "data processing unit" or "data processing device" includes, for example, all devices, apparatuses, and machines for processing data, including programmable processing devices, computers, or multiple processing devices or computers. In addition to the hardware, this device can include code that creates the execution environment of the computer program, for example, code that constitutes processing device firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.
[0378] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including a compiled language or an interpreted language, and it can be distributed in any form, either as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program may be recorded as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), stored in a single file dedicated to the program, or stored in multiple coordinating files (e.g., files that hold one or more modules, subprograms, or portions of code). It is also possible to deploy one computer program to be executed on one computer located at one site or on multiple computers distributed across multiple sites and interconnected by a communication network.
[0379] The processes and logic flows described herein can be performed by one or more programmable processing devices that execute one or more computer programs to function by operating on input data and generating output. The processes and logic flows can also be performed by special-purpose logic circuits, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as special-purpose logic circuits.
[0380] A processing device suitable for the execution of a computer program includes, for example, both general-purpose and special-purpose microprocessors, as well as any one or more processing devices of any type of digital computer. Generally, the processing device receives instructions and data from a read-only memory or a random-access memory or both. The essential elements of a computer are a processor for executing instructions and one or more storage devices for storing instructions and data. Generally, a computer may include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or may be operatively coupled to receive data from or transfer data to these mass storage devices. However, a computer need not have such devices. A computer-readable medium suitable for storing computer program instructions and data includes any form of non-volatile memory, medium, and memory device, including, for example, semiconductor memory devices such as EPROM, EEPROM, flash memory devices, etc. The processing device and the memory may be supplemented by, or incorporated in, application-specific logic circuitry.
[0381] This patent specification includes many details, but these should not be construed as limiting the scope of any invention or the scope of the claims, but rather as descriptions and interpretations of features that may be specific to particular embodiments of a particular invention. Specific features described in the context of separate embodiments in this patent document may be implemented in combination in one example. Conversely, the various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0382] Similarly, although the operations are shown in a particular order in the drawings, it should not be understood that this requires that the operations be performed in the particular order shown or in a sequential order in order to achieve the desired result, or that all of the operations shown be performed. Also, the separation of the various system components in the examples described in this patent specification should not be understood to require such separation in all embodiments.
[0383] Only some implementations and examples are described, and other embodiments, extensions, and variations are possible based on the content described and illustrated in this patent document.
Claims
Claim 1 For conversion between a first video region of a video and the bitstream of the video, determining that a first syntax element corresponding to the first video region exists in the bitstream, the first syntax element indicating an index of an adaptive parameter set referred to by the first video region, the first video region being a video picture or a video slice, performing the conversion based on the determination, when the adaptive parameter set is invalid, the first syntax element does not exist in the bitstream, for samples of the first video region, a filter index is derived based on differences between samples in a plurality of different directions, and from the luminance filter set of the adaptive parameter set and another luminance filter not included in the luminance filter set, the filter index is allowed to be used to select a specific luminance filter based on the filter index, the video region includes a plurality of video coding tree blocks, the video coding tree blocks are divided into a plurality of M*M video blocks, and M is equal to 2 or 4, the same filter index is applied to each sample of the M*M video blocks, for each of the M*M video blocks, the differences between the samples in the plurality of different directions are derived based on a subsampling rate of 1:2, A video data processing method. Claim 2 The method according to claim 1, wherein the adaptive parameter set includes an adaptive loop filtering adaptive parameter set for adaptive loop filtering. Claim 3 The method according to claim 1, wherein the maximum value of the filter index is indicated by N, and the number of luminance filters in the luminance filter set is N+1 or less. Claim 4 The method according to claim 1, wherein the filter index is allowed to be used to derive another specific luminance filter included in a fixed filter set. Claim 5 The method according to any one of claims 1 to 4, wherein the conversion includes encoding the first video region into the bitstream. Claim 6 The method according to any one of claims 1 to 4, wherein the conversion includes decoding the first video region from the bitstream. Claim 7 An apparatus for processing video data, comprising a processing device and a non-transitory memory storing instructions, which when executed by the processing device, cause the processing device to determine that a first syntax element corresponding to the first video region of the video exists in the bitstream for conversion between the first video region of the video and the bitstream of the video, and the first syntax element indicates an index of an adaptive parameter set referred to by the first video region, and the first video region is a video picture or a video slice, perform the conversion based on the determination, when the adaptive parameter set is invalid, the first syntax element does not exist in the bitstream, a filter index is derived based on differences between samples in a plurality of different directions for samples of the first video region, and the filter index is allowed to be used to select a specific luminance filter from the luminance filter set of the adaptive parameter set and other luminance filters not included in the luminance filter set, the video region includes a plurality of video coding tree blocks, and the video coding tree blocks are divided into a plurality of M*M video blocks, where M is equal to 2 or 4, the same filter index is applied to each sample of the M*M video blocks, for each of the M*M video blocks, the differences between the samples in the plurality of different directions are derived based on a 1:2 subsampling rate, An apparatus for processing video data. Claim 8 A non-transitory computer-readable storage medium storing instructions, which when executed by a processing device, cause the processing device to determine that a first syntax element corresponding to the first video region of the video exists in the bitstream for conversion between the first video region of the video and the bitstream of the video, and the first syntax element indicates an index of an adaptive parameter set referred to by the first video region, and the first video region is a video picture or a video slice, perform the conversion based on the determination, when the adaptive parameter set is invalid, the first syntax element does not exist in the bitstream For samples in the first video region, a filter index is derived based on differences between samples in a plurality of different directions, and from the luminance filter set of the adaptive parameter set and other luminance filters not included in the luminance filter set, the filter index is used to select a specific luminance filter based on the filter index. The video region includes a plurality of video coding tree blocks, and the video coding tree blocks are divided into a plurality of M*M video blocks, where M is equal to 2 or 4. The same filter index is applied to each sample of the M*M video blocks. For each of the M*M video blocks, the differences between the samples in the plurality of different directions are derived based on a 1:2 subsampling rate. Non-transitory computer-readable storage medium. **Claim 9** A method for storing a bitstream of a video, the method comprising: determining that a first syntax element corresponding to the first video region of the video exists in the bitstream, the first syntax element indicating an index of an adaptive parameter set referred to by the first video region, and the first video region being a video picture or a video slice; generating the bitstream based on the determining; storing the bitstream in a non-transitory computer-readable recording medium, wherein when the adaptive parameter set is invalid, the first syntax element does not exist in the bitstream; For samples in the first video region, a filter index is derived based on differences between samples in a plurality of different directions, and from the luminance filter set of the adaptive parameter set and other luminance filters not included in the luminance filter set, the filter index is allowed to be used to select a specific luminance filter based on the filter index. The video region includes a plurality of video coding tree blocks, and the video coding tree blocks are divided into a plurality of M*M video blocks, where M is equal to 2 or 4. The same filter index is applied to each sample of the M*M video blocks. For each of the M*M video blocks, the differences between the samples in the plurality of different directions are derived based on a 1:2 subsampling rate. Method.