Separate constrained directionality enhancement filters.
The Separate Constrained Directional Enhancement Filter (SCDEF) addresses the challenge of filtering luma and chroma components independently, enhancing coding efficiency by allowing separate presets and block-level flexibility, particularly in scenarios with differing partitioning schemes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2026-03-10
AI Technical Summary
Existing video encoding technologies, such as AV1, face challenges in efficiently filtering luma and chroma components independently, particularly when they have different or semi-disconnected partitions, leading to limitations in coding efficiency and performance.
A Separate Constrained Directional Enhancement Filter (SCDEF) process is applied to filter luma and chroma components separately, allowing for different presets at the picture level and block level, and using the chroma block size to determine filter strength, independent of luma components.
This approach enhances coding efficiency by enabling independent filtering of luma and chroma components, improving performance in scenarios with varying partitioning schemes like semi-detached partitioning.
Smart Images

Figure 0007827803000021 
Figure 0007827803000022 
Figure 0007827803000023
Abstract
Description
[Technical Field]
[0001] Priority information This application claims the benefit of priority to U.S. Provisional Application No. 63 / 040,856, filed June 18, 2020, and U.S. Application No. 17 / 091,759, filed November 6, 2020, which are incorporated by reference herein in their entireties.
[0002] The present disclosure relates generally to the field of data processing, and more particularly to video encoding and / or decoding (eg, by a coder, decoder, or codec (decoder and encoder)). [Background technology]
[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to related technology codec extensions, for example. Summary of the Invention [Means for solving the problem]
[0004]
[0003] Embodiments relate to a method, a system, and a computer-readable medium for encoding and / or decoding video data. According to one aspect, a method for encoding and / or decoding video data is provided. The method may include receiving video data including chroma components and luma components, analyzing, deriving, or selecting a number of presets for the chroma components in a frame and a number of presets for the luma component in a frame, and decoding the video data, wherein the method includes performing a separate constrained directional emphasis filter (CDEF) process to filter the luma and chroma components independently of each other based on the number of presets for the chroma components in the frame and the number of presets for the luma component in the frame.
[0005] The method may include performing a separate Constrained Directional Emphasis Filter (CDEF) process to filter the luma and chroma components independent of each other when the luma and chroma components have different or semi-disconnected partitions, and obtaining an output of the separate CDEF process including filtered reconstructed samples of the luma / chroma components, wherein an input of the separate CDEF process is the reconstructed samples of the luma / chroma components, and an intermediate output of the separate CDEF process includes using the derived filter preset and per-block level preset index.
[0006] The number of presets derived for the luma component is different from the number of presets derived for the chroma components at the picture level.
[0007] The number of presets at the picture level may include one of 1, 2, 4, or 8.
[0008] The number of presets derived and selected for the luma component in one frame is two, and the number of presets derived and selected for the chroma component in one frame is one.
[0009] The number of presets derived and selected for the luma component is a positive integer N, and the number of presets for the chroma components is fixed to 1, which is derived as 1 in the decoder without signaling.
[0010] The selected preset index for the current luma block is different from the selected preset index for the current chroma block, and the inputs of the separate CDEF process are the luma / chroma reconstructed samples of the current block and the selected preset derived at the frame level. The output of this process is an index indicating which preset is selected for the current block.
[0011] The method may further include, at a frame level, selecting a preset index of luma block A as 7 and a preset index of chroma block B as 1 when the number of luma components corresponds to 8 presets and the number of chroma components corresponds to 4 presets, and luma block A and chroma block B are co-located or partially co-located.
[0012] The method may further include, when deriving the CDEF filtering strength for the chroma components, the input reconstructed samples being determined by a current chroma coding block size.
[0013] The method may further include, when the current chroma block is of a constant size, the input being chroma reconstructed sample values of the current block having a constant size.
[0014] The method may further include that when separate or semi-detached partitioning is applied to luma and chroma blocks, the luma and chroma blocks still share the same preset index, and only one of the luma or chroma block sizes is employed in the preset index derivation / signaling process.
[0015] The method may further include, when the luma and chroma components have the same coding block size, the CDEF filtering process for the luma and chroma components is performed separately.
[0016] Picture-level presets may be signaled separately for the luma and chroma components in a high-level parameter set, slice header, picture header, or Supplementary Enhancement Information (SEI) message.
[0017] The luma presets may be signaled first, then the chroma presets.
[0018] The block level preset index is signaled separately for the luma and chroma components.
[0019] The preset index for the luma component is signaled first, then the preset index for the chroma component.
[0020] A computer system for decoding video data may comprise one or more computer-readable non-transitory storage media configured to store computer program code; and one or more computer processors configured to access the computer program code and operate as instructed by the computer program code, the computer program code including: receiving code configured to cause the one or more computer processors to receive video data including chroma components and luma components; analyzing, deriving, or selecting code configured to cause the one or more computer processors to analyze, derive, or select a number of presets for the chroma components in a frame and a number of presets for the luma component in a frame; and decoding code configured to cause the one or more computer processors to decode the video data, wherein the method includes performing separate Constrained Directional Emphasis Filter (CDEF) processing to filter the luma and chroma components independent of each other based on the number of presets for the chroma components in a frame and the number of presets for the luma component in a frame.
[0021] A non-transitory computer-readable medium having stored thereon a computer program for decoding video data can be configured to cause one or more computer processors to receive video data including chroma components and luma components, including: analyzing, deriving, or selecting code configured to cause the one or more computer processors to analyze, derive, or select a number of presets for the chroma components in a frame and a number of presets for the luma component in a frame; and decoding code configured to cause the one or more computer processors to decode the video data, the method including performing a separate Constrained Directional Emphasis Filter (CDEF) process that filters the luma and chroma components independent of each other based on the number of presets for the chroma components in a frame and the number of presets for the luma component in a frame.
[0022] These and other objects, features and advantages will become apparent from the following detailed description of illustrative embodiments, which is to be read in connection with the accompanying drawings. The various features of the drawings are not to scale, as the illustrations are for clarity in facilitating understanding by those skilled in the art in conjunction with the detailed description. [Brief explanation of the drawings]
[0023] [Figure 1] 1 illustrates a networked computer environment in accordance with at least one embodiment. [Figure 2] The filter shape of the adaptive loop filter (ALF) is shown. [Figure 3A] Indicates the subsample position of the diagonal gradient. [Figure 3B] Indicates the subsample position of the diagonal gradient. [Figure 3C] Indicates the subsample position of the diagonal gradient. [Figure 3D] Indicates the subsample position of the diagonal gradient. [Figure 4] 10 shows modified block classification with virtual boundaries. [Figure 5] 10 illustrates modified ALF filtering for the luma component at the virtual boundary. [Figure 6] Indicates the position of the chroma samples relative to the luma samples. [Figure 7] An example of direction search for an 8x8 block is shown below. [Figure 8] An example of direction search for an 8x8 block is shown below. [Figure 9] An example of a coding tree structure (luma and chroma) is shown. [Figure 10] 1 shows a Separate Constrained Directional Enhancement Filter (SCDEF). [Figure 11] 1 is an operational flowchart illustrating steps performed by a program for encoding video data, according to at least one embodiment. [Figure 12] FIG. 2 is a block diagram of internal and external components of the computer and server illustrated in FIG. 1 according to at least one embodiment. [Figure 13] 2 is a block diagram of an exemplary cloud computing environment including the computer system illustrated in FIG. 1 according to at least one embodiment. [Figure 14] FIG. 14 is a block diagram of functional layers of the exemplary cloud computing environment of FIG. 13 in accordance with at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0024] Although detailed embodiments of the claimed structures and methods are disclosed herein, it should be understood that the disclosed embodiments are merely exemplary of the claimed structures and methods, which may be embodied in various forms. These structures and methods, however, may be embodied in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete, and will fully convey its scope to those skilled in the art. In this description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0025] FIELD OF THE INVENTION
[0002] Embodiments relate generally to the field of data processing, and more particularly to video encoding and / or decoding. Exemplary embodiments described below provide, among other things, systems, methods, and computer programs for encoding and / or decoding video data.
[0026] As mentioned above, AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as the successor to VP9 by the Alliance for Open Media (AOMedia), a consortium formed in 2015 that includes semiconductor companies, video-on-demand providers, video content creators, software developers, and web browser vendors.
[0027] Aspects are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0028]
[0023] Referring now to Figure 1, a functional block diagram of a networked computing environment is shown illustrating a video encoding system 100 (hereinafter "system") for encoding and / or decoding video data according to one embodiment. It should be understood that Figure 1 is provided only as an illustration of one implementation and is not meant to be limiting with respect to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made based on design and implementation requirements.
[0029] System 100 may include a computer 102 and a server computer 114. Computer 102 may communicate with server computer 114 via communications network 110 (hereinafter “network”). Computer 102 may include a processor 104 and a software program 108 stored on a data storage device 106, capable of interfacing with a user, and communicating with server computer 114. As described below with reference to FIG. 12 , computer 102 may include internal components 800A and external components 900A, respectively, and server computer 114 may include internal components 800B and external components 900B, respectively. Computer 102 may be, for example, a mobile device, a phone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of running programs, accessing a network, and accessing a database.
[0030] The server computer 114 may also operate in a cloud computing service model, such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), as described below with respect to Figures 13 and 14. The server computer 114 may also be located in a cloud computing deployment model, such as a private cloud, a community cloud, a public cloud, or a hybrid cloud.
[0031] The server computer 114, which may be used to encode video data, may execute a video encoding or decoding program 116 (hereinafter "program") that may interact with the database 112. The video encoding or decoding program method is described in more detail below with respect to FIG. 3. In one embodiment, the computer 102 may operate as an input device, including a user interface, and the program 116 may execute primarily on the server computer 114. In an alternative embodiment, the program 116 may execute primarily on one or more computers 102, with the server computer 114 being used to process and store data used by the program 116. Note that the program 116 may be a standalone program or may be integrated into a larger video encoding program. The video encoding or decoding program 116 may correspond to an encoder, a decoder, or both an encoder and a decoder.
[0032] It should be noted, however, that the processing of the program 116 may, in some cases, be shared between the computer 102 and the server computer 114 in any proportion. In another embodiment, the program 116 may run on multiple computers, server computers, or some combination of computers and server computers, e.g., multiple computers 102 communicating with a single server computer 114 over the network 110. In another embodiment, for example, the program 116 may run on multiple server computers 114 communicating with multiple client computers over the network 110. Alternatively, the program may run on a network server that communicates with the server and multiple client computers over the network.
[0033] Network 110 may include wired connections, wireless connections, fiber optic connections, or some combination thereof. In general, network 110 may be any combination of connections and protocols that support communication between computer 102 and server computer 114. Network 110 may include various types of networks, such as, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, a telecommunications network such as a public switched telephone network (PSTN), a wireless network, a public switched network, a satellite network, a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, an optical fiber-based network, etc., and / or a combination of these or other types of networks.
[0034] The number and arrangement of devices and networks shown in Figure 1 are provided as an example. In practice, there may be additional, fewer, different, or differently arranged devices and / or networks. Furthermore, two or more devices shown in Figure 1 may be implemented within a single device, or a single device shown in Figure 1 may be implemented as multiple distributed devices. Additionally, or alternatively, a set of devices (e.g., one or more devices) of system 100 may perform one or more functions that are described as being performed by another set of devices of system 100.
[0035] 1 Adaptive Loop Filter (ALF) Versatile Video Coding (VVC) (Draft 8) applies an adaptive loop filter (ALF) with block-based filter adaptation: for the luma component, one of 25 filters is selected for each 4x4 block based on local gradient direction and activity.
[0036] 1.1 Filter Shape VVC (Draft 8) allows the use of two diamond-shaped filter shapes (as shown in Figure 2): a 7x7 diamond shape is applied to the luma component, and a 5x5 diamond shape is applied to the chroma components.
[0037] 1.2 Block Classification For the luma component, each 4x4 block is classified into one of 25 classes. The classification index C is determined by its directionality D and the quantized value of the motion.
number
number
[0038] D and
number
number
[0039] where the indices i and j indicate the coordinates of the top-left sample in a 4x4 block, and R(i,j) indicates the reconstructed sample at coordinate (i,j).
[0040] To reduce the complexity of block classification, a subsampled one-dimensional Laplacian calculation is applied. As shown in Figures 3A-3D, the same subsampled positions are used for all gradient calculations (e.g., all-directional subsampled Laplacian calculations). For example, Figure 3A shows the subsampled positions of vertical gradients, Figure 3B shows the subsampled positions of horizontal gradients, and Figures 3C and 3D show the subsampled portions of diagonal gradients.
[0041] Next, set the maximum and minimum values of D for the horizontal and vertical gradients as follows:
number
[0042] The maximum and minimum values of the gradients in the two diagonal directions are set as follows:
number
[0043] These values are compared to each other and to two thresholds t1 and t2 to derive the value of the directionality D. Step 1.
number
number
number
number
number
[0044] The motion value A is calculated as follows:
number
[0045] Furthermore, A is quantized in the range of 0 to 4 (inclusive), and the quantized value is
number
[0046] No classification method is applied to the chroma components in a picture, i.e., a single set of ALF coefficients is applied to each chroma component.
[0047] 1.3 Geometric transformation of filter coefficients and clipping values Before filtering each 4x4 luma block, a geometric transformation such as a rotation or a diagonal or vertical flip is applied to the filter coefficients f(k,l) and the corresponding filter clipping values c(k,l) based on the gradient value calculated for that block. This is equivalent to applying these transformations to samples in the filter support region. The idea is to make different blocks that have ALF applied more similar by aligning their orientations.
[0048] Three geometric transformations are introduced: diagonal, vertical flip and rotation. Diagonal: f D (k, l)=f(l, k), c D (k, l)=c(l, k), (Equation 9) Vertical flip:f V (k, l)=f(k, Kl-1), c V (k, l)=c(k, Kl-1) (Equation 10) Rotation:f R (k, l)=f(Kl-1, k), c R (k, l)=c(Kl-1, k) (Equation 11) where K is the size of the filter, 0≦k, l≦K-1 are the coordinates of the coefficients, with position (0,0) being the top-left corner and position (K-1,K-1) being the bottom-right corner. The transformation is applied to the filter coefficients f(k,l) and clipping values c(k,l) based on the gradient values calculated for that block. The relationship between the transformation and the four gradients in the four directions is summarized in Table 1 below.
[0049] [Table 1]
[0050] 1.4 Filter Parameter Signaling In VVC (Draft 8), ALF filter parameters are signaled in an Adaptation Parameter Set (APS). One APS can signal up to 25 sets of luma filter coefficients and clipping value indices and up to 8 sets of chroma filter coefficients and clipping value indices. To reduce bit overhead, filter coefficients of different classifications of luma components can be merged. The slice header signals the index of the APS used for the current slice. In VVC (Draft 8), ALF signaling is CTU-based.
[0051] The clipping value index decoded from the APS can be used to determine the clipping values using the luma and chroma clipping value tables. These clipping values depend on the internal bit depth. More precisely, the clipping value tables are calculated by the following formula: AlfClip={round(2 B-α*n ), n∈[0..N-1]} (Equation 12) where B is equal to the internal bit depth, α is a predefined constant equal to 2.35, and N is the number of clipping values allowed in VVC (Draft 8) equal to 4.
[0052] Table 2 shows the output of Equation 12.
[0053] [Table 2]
[0054] The slice header can signal up to seven APS indices to specify the luma filter set to be used for the current slice. The filtering process can be further controlled at the coding tree block (CTB) level. A flag is always signaled to indicate whether the ALF is applied to the luma CTB. The luma CTB can select a filter set from 16 fixed filter sets and a filter set from an APS. The filter set index of the luma CTB is signaled to indicate which filter set to apply. The 16 fixed filter sets are predefined and hard-coded in both the encoder and decoder.
[0055] For chroma components, an APS index is signaled in the slice header to indicate the chroma filter set used in the current slice. At the CTB level, if an APS has multiple chroma filter sets, a filter index is signaled for each chroma CTB.
[0056] The filter coefficients can be quantized with a norm equal to 128. To limit the complexity of multiplications, bitstream conformance is applied so that coefficient values for non-center positions are in the range of -27 to 27-1 (inclusive). Center position coefficients are not signaled in the bitstream and are assumed to be equal to 128.
[0057] In VVC (Draft 8), the syntax and semantics of clipping indices and values are defined as follows: alf_luma_clip_idx[ sfIdx ][ j ] specifies the clipping index of the clipping value to use before multiplying the j-th coefficient of the signaled luma filter indicated by sfIdx. The bitstream conformance requirement is that the value of alf_luma_clip_idx[ sfIdx ][ j ] for sfIdx=0..alf_luma_num_filters_signalled_minus1 and j=0..11 is in the range 0 to 3 (inclusive).
[0058] The luma filter clipping values AlfClipL[ adaptation_parameter_set_id ], with elements AlfClipL[ adaptation_parameter_set_id ][ filtIdx ][ j ], where filtIdx=0..NumAlfFilters-1 and j=0..11, are derived as specified in Table 2 depending on bitDepth being set equal to BitDepthY and clipIdx being set equal to alf_luma_clip_idx[ alf_luma_coeff_delta_idx[ filtIdx ] ][ j ].
[0059] alf_chroma_clip_idx[ altIdx ][ j ] specifies the clipping index of the clipping value to use, which is then multiplied by the jth coefficient of the alternative chroma filter with index altIdx. It is a bitstream conformance requirement that the value of alf_chroma_clip_idx[ altIdx ][ j ] for altIdx=0..alf_chroma_num_alt_filters_minus1, j=0..5 be in the range 0..3 (inclusive).
[0060] The chroma filter clipping value AlfClipC[ adaptation_parameter_set_id ][ altIdx ] with element AlfClipC[ adaptation_parameter_set_id ][ altIdx ][ j ], where altIdx=0.. alf_chroma_num_alt_filters_minus1, j=0..5, is derived as specified in Table 2 depending on bitDepth being set equal to BitDepthC and clipIdx being set equal to alf_chroma_clip_idx[ altIdx ][ j ].
[0061] 1.5 Filtering Process On the decoder side, when ALF is enabled for the CTB, each sample R(i, j) in the CU is filtered to a sample value R'(i, j) as shown below.
number
number
number
[0062] 1.6 Virtual Boundary Filtering for Line Buffer Reduction To reduce the line buffer requirements of ALF, we employ modified block classification and filtering for samples near horizontal CTU boundaries. To this end, we can define a virtual boundary as a line by shifting the horizontal CTU boundary by "N" samples, as shown in Figure 4, where "N" is equal to 4 for the luma component and 2 for the chroma component.
[0063] As illustrated in Figure 4, modified block classification is applied to the luma component. The 1D Laplacian gradient calculation for 4x4 blocks above the virtual boundary uses only samples above the virtual boundary. Similarly, the 1D Laplacian gradient calculation for 4x4 blocks below the virtual boundary uses only samples below the virtual boundary. The quantization of the motion value A is scaled accordingly to account for the reduced number of samples used in the 1D Laplacian gradient calculation.
[0064] In the filtering process, both the luma and chroma components undergo a symmetric padding operation at the virtual boundary. As shown in Figure 5 ("Modified ALF Filtering for Luma Component at the Virtual Boundary"), when a sample to be filtered is located below the virtual boundary, the neighboring samples located above the virtual boundary are padded. Meanwhile, the other corresponding samples are also padded symmetrically.
[0065] 1.7 Maximum Coding Unit (LCU)-Aligned Picture Quadtree Partitioning To improve coding efficiency, JCTVC-C143 [3] proposes a coding unit-synchronized picture quadtree-based adaptive loop filter. The luma picture is divided into multiple multi-level quadtree partitions, and the boundaries of each partition are aligned with the boundaries of the largest coding unit (LCU). Each partition has its own filtering process and is therefore called a filter unit (FU).
[0066] The two-pass coding flow is described below. In the first pass, the quadtree division pattern and the optimal filter for each FU are determined. During the determination process, the filtering distortion is estimated by FFDE. The reconstructed picture is filtered according to the determined quadtree division pattern and the selected filters for all FUs. In the second pass, the CU-synchronized ALF on / off control is performed. According to the ALF on / off result, the initially filtered picture is partially restored by the reconstructed picture.
[0067] A top-down partitioning method is adopted to divide a picture into multi-level quadtree partitions using a rate-distortion criterion. Each partition is called a filter unit. The partitioning process aligns the quadtree partitions with the boundaries of LCUs. The coding order of FUs follows the z-scan order. For example, a picture can be divided into 10 FUs, and the coding order is FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, and FU9.
[0068] To indicate the picture quadtree partitioning pattern, the partition flags can be coded and transmitted in Z-order.
[0069] The filter for each FU can be selected from two filter sets based on a rate-distortion criterion. The first set can have newly derived ½-symmetric square and diamond filters for the current FU. The second set can be obtained from a time-delay filter buffer, which stores filters previously derived for FUs of the previous picture. The filter with the smallest rate-distortion cost of these two sets can be selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further divided into four child FUs, the rate-distortion costs of the four child FUs are calculated. By recursively comparing the rate-distortion costs for the division and no division cases, a picture quadtree division pattern can be determined.
[0070] In JCTVC-C143, the maximum quadtree division level is 2, which means the maximum number of FUs is 16. During the quadtree division decision, the correlation values for deriving the Wiener coefficients of the 16 FUs at the lowest quadtree level (smallest FU) can be reused. The remaining FUs can derive their Wiener filters from the correlations of the 16 FUs at the lowest quadtree level. Therefore, only one frame buffer access is required to derive the filter coefficients of all FUs.
[0071] After the quadtree partitioning pattern is determined, the CU-synchronized ALF on / off control is performed to further reduce the filtering distortion. By comparing the filtering distortion with the non-filtering distortion, leaf CUs can explicitly switch the ALF on / off in their local regions. Coding efficiency can be further improved by redesigning the filter coefficients according to the ALF on / off results. However, the redesign process requires additional frame buffer accesses. In the proposed CS-PQALF encoder design, there is no redesign process after the CU-synchronized ALF on / off decision to minimize the number of frame buffer accesses.
[0072] 2. Cross-component adaptive loop filter The cross-component adaptive loop filter (CC-ALF) utilizes the luma sample values to refine each chroma component.
[0073] CC-ALF works by applying a linear diamond-shaped filter to the luma channel of each chroma component. The filter coefficients are transmitted in APS and are expressed as 2 10The inputs are scaled by a factor of 16x16 and rounded for fixed-point representation. The application of the filters is controlled by variable block sizes and signaled by the context coding flags received for each block of samples. Block sizes are received at the slice level for each chroma component, along with the CC-ALF enable flag. The following block sizes (in chroma samples) are supported for contribution: 16x16, 32x32, 64x64.
[0074] The syntax changes for CC-ALF are explained below in Table 3.
[0075] [Table 3]
[0076] The semantics of CC-ALF related syntax are explained below.
[0077] alf_ctb_cross_component_cb_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] equal to 0 indicates that the cross-component Cb filter is not applied to the block of Cb color component samples at luma location (xCtb, yCtb).
[0078] If alf_cross_component_cb_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] is not equal to 0, it indicates that the alf_cross_component_cb_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]th cross-component Cb filter is applied to the block of Cb color component samples at luma location (xCtb, yCtb).
[0079] alf_ctb_cross_component_cr_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] equal to 0 indicates that no cross-component Cr filter is applied to the block of Cr color component samples at luma location (xCtb, yCtb). alf_cross_component_cr_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] not equal to 0 indicates that the alf_cross_component_cr_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]th cross-component Cr filter is applied to the block of Cr color component samples at luma location (xCtb, yCtb).
[0080] 3 Chroma Sampling Format Figure 6 ("Location of Chroma Samples Relative to Luma Samples") of this application Chroma shows the relative location of the top-left chroma sample when chroma_format_idc is 1 (4:2:0 chroma format) and chroma_sample_loc_type_top_field or chroma_sample_loc_type_bottom_field is equal to the value of the variable ChromaLocType. The area represented by the top-left 4:2:0 chroma sample (illustrated as a large red square with a large red dot in the center) is shown relative to the area represented by the top-left luma sample (illustrated as a small black square with a small black dot in the center). The area represented by the adjacent luma sample is shown as a small gray square with a small gray dot in its center.
[0081] 4. Constrained Directional Enhancement Filter The main goal of the in-loop Constrained Directional Enhancement Filter (CDEF) is to remove coding artifacts while preserving image details. In HEVC, the Pixel Adaptive Offset (SAO) algorithm achieves a similar goal by defining signaling offsets for different classes of pixels. Unlike SAO, CDEF is a nonlinear spatial filter. The filter design is constrained to be easily vectorizable (i.e., it can be implemented with SIMD operations), which is not the case for other nonlinear filters such as median filters or bilateral filters.
[0082] The design of CDEF is based on the following observations: The amount of ringing artifacts in a coded image tends to be roughly proportional to the quantization step size. Although the amount of detail is a property of the input image, the smallest detail retained in the quantized image also tends to be proportional to the quantization step size. For a given quantization step size, the amplitude of the ringing will generally be smaller than the amplitude of the detail.
[0083] CDEF identifies the orientation of each block and adaptively filters along the identified orientation and along a direction rotated 45 degrees from the identified orientation at a smaller angle. The filter strength is explicitly signaled, allowing for a high degree of control over blurring. An efficient encoder search is designed for the filter strength. CDEF is based on two previously proposed in-loop filters, and a combined filter has been adopted for the new AV1 codec.
[0084] 4.1 Direction search The direction search is performed on the reconstructed pixels immediately after the deblocking filter. Because these pixels are available to the decoder, the direction does not require signaling. The search operates on 8x8 blocks, which are small enough to properly handle nonlinear edges and large enough to reliably estimate direction when applied to a quantized image. Enforcing a certain directionality within the 8x8 region facilitates vectorization of the filter. For each block, the direction that best matches the pattern within the block is determined by minimizing the sum of squared differences (SSD) between the quantized block and the nearest fully directional block. A fully directional block is a block in which all pixels along a directional line have the same value. Figure 7 shows an example of a direction search for an 8x8 block. In this case, the 45-degree direction (indicated by the box surrounding column 12) that minimizes the error is selected.
[0085] 4.2 Nonlinear low-pass directional filters The main reason for identifying a direction is to align filter taps along that direction to reduce ringing while preserving directional edges or patterns. However, directional filtering alone may not be enough to reduce ringing. It is also desirable to use filter taps for pixels that are not aligned with the primary direction. These extra taps are treated more conservatively to reduce the risk of blurring. For this reason, CDEF defines primary and secondary taps. A complete 2D CDEF filter is expressed as follows:
number
[0086] 5. AV1 loop recovery Beyond traditional deblocking operations, a set of in-loop restoration schemes has been proposed for use in post-deblocking video coding, generally to remove noise and improve edge quality. These schemes are switchable within a frame for tiles of appropriate size. The specific schemes described are based on a separable symmetric Wiener filter and a dual self-guided filter with subspace projection. Because content statistics can change substantially within a frame, these tools are integrated into a switchable framework that allows different tools to be activated in different regions of the frame.
[0087] 5.1 Separable Symmetric Wiener Filter One restoration tool that has shown promise in the literature is the Wiener filter. Every pixel in a degraded frame can be reconstructed as a non-causal filtered version of the pixels in a w × w window around it, where w = 2r + 1 is odd with respect to the integer r. The 2D filter taps are expressed as a column vectorized form of w 2 If the vector is represented by a 1 × 1 element vector F, then direct LMMSE optimization gives F=H -1 The filter parameters are derived as given by M, where H=E[XX T ] is the autocovariance of x and w in a w × w window around the pixel 2 is a column-wise vectorized version of the samples of M=E[YX T] is the cross-correlation between x and the scalar source sample y to be estimated. The encoder can estimate H and M from the realization of the deblocked frame and the source, and send the resulting filter F to the decoder. However, doing so would require w 2 Not only does transmitting taps incur a significant bitrate cost, but non-separable filtering significantly complicates decoding. Therefore, several additional constraints are imposed on the properties of F. First, F is constrained to be separable, and the filtering can be implemented as separable horizontal and vertical w-tap convolutions. Second, each horizontal and vertical filter is constrained to be symmetric. Third, both horizontal and vertical filter coefficients are assumed to sum to one.
[0088] 5.2 Dual Self-Guided Filtering with Subspace Projection Guided filtering is one of the recent paradigms in image filtering, where a local linear model is y=Fx+G (Equation 15) The above formula is used to calculate the filtered output y from the unfiltered sample x. Here, F and G are determined based on the statistics of the guidance image in the neighborhood of the degraded image and the filtered pixel. If the guidance image is the same as the degraded image, the resulting so-called self-guided filtering has the effect of edge-preserving smoothing. The specific form of self-guided filtering proposed by the inventors depends on two parameters, the radius and the noise parameter e, and is listed as follows: 1. The mean μ and variance σ of pixels in a (2r+1) × (2r+1) window around each pixel 2 This can be efficiently implemented with box filtering based on integral imaging. 2. Calculate for all pixels: f = σ 2 / (σ 2 +e);g=(1-f)μ 3. Calculate F and G for every pixel as the average of the f and g values in a 3x3 window around the pixel being used.
[0089] Filtering is controlled by r and e, with larger r resulting in larger spatial variance and larger e resulting in larger range variance.
[0090] The principle of subspace projection is shown diagrammatically in Figure 8. Even if neither of the inexpensive reconstructions X1, X2 is close to the source Y, appropriate multipliers {α, β} can bring them quite close to the source, as long as they are somewhat shifted in the right direction. Figure 8 shows subspace projection using inexpensive reconstructions to produce a final reconstruction that is closer to the source.
[0091] 6. Semi-detached divisions Semi-decoupled partitioning (SDP), semi-separate tree (SST), or flexible block partitioning of the chroma component. In this method, the luma blocks and chroma blocks in one superblock (SB) can have the same or different block partitions depending on the luma coding block size or luma tree depth. Specifically, if the area size of the luma block is greater than a threshold T1 or the coding tree partition depth of the luma block is less than or equal to a threshold T2, the chroma block uses the same coding tree structure as the luma. Otherwise, if the block area size is less than or equal to T1 or the luma partition depth is greater than T2, the corresponding chroma block can have a different coding block partition from the luma component, which is called flexible block partitioning of the chroma component. T1 is a positive integer such as 128 or 256. T2 is a positive integer such as 1 or 2.
[0092] An improved semi-detached partitioning (SDP) scheme is proposed, in which the luma and chroma components can share a partial tree structure from the root node of a superblock, and the condition for starting separate tree partitions for luma and chroma depends on luma partition information. For example, Figure 9 shows an example of a coding tree structure for luma and chroma components.
[0093] In a constrained directional enhancement filter (CDEF), the luma and chroma components are constrained to share presets at the picture level. Furthermore, the luma and chroma components are constrained to have the same preset index at the block level. Finally, when deriving the filter strength of the chroma component, the luma block size is used to determine the input of the chroma component. These constraints can limit the coding efficiency of the CDEF.
[0094] In a conventional CDEF, one preset includes the primary / secondary intensities of luma and chroma. The number of allowed / available presets is signaled at the picture level. At the coding block level, an index indicating which preset is selected for the current block is signaled. The coding block sizes for CDEF include 128x128, 128x64, 64x64, and 64x128. The conventional CDEF has three limitations. First, the luma and chroma components must share a preset at the picture level. Second, the luma and chroma components must select the same preset index at the block level. Third, the luma block size must be used to determine the input of the chroma component when deriving the filter strength of the chroma component. Both of these limitations can limit the performance of the CDEF, especially in situations where the luma and chroma components have different partitioning schemes, such as the partitioning scheme in semi-detached partitioning (SDP).
[0095] This paper proposes a Separate Constrained Directional Enhancement Filter (SCDEF) that performs CDEF processing for the luma component and the chroma component separately. Compared with conventional CDEF, the SCDEF enables filtering of the luma component and the chroma component independently of each other. More specifically, the luma component and the chroma component may have different numbers of presets at the picture level, and further, the luma component and the chroma component may select different preset indexes at the block level, and the chroma block size is used to determine the input of the chroma component when deriving the filter strength of the chroma component.
[0096] As shown in Figure 12, when the luma and chroma components have different or semi-disconnected partitions, it is proposed to perform the CDEF filtering process for the luma and chroma components separately. The input of the CDEF filtering process is the reconstructed samples of the luma / chroma components. The intermediate outputs of this process include, but are not limited to, derived filter presets and per-block level preset indexes, as described in the proposed method above. The final output of this process is the filtered reconstructed samples of the luma / chroma components.
[0097] In one embodiment, the number of presets derived for the luma and chroma components may be different from each other at the picture level. The input of the CDEF filtering process is the reconstructed samples of the luma / chroma components. The output of this process is the presets derived at the picture level. Examples of the number of presets at the picture level include, but are not limited to, 1, 2, 4, and 8.
[0098] In one example, the number of presets derived and selected for the current luma component in a frame is two, and the number of presets derived and selected for the current chroma component of this frame is one.
[0099] In another example, the number of presets derived and selected for the luma component is N, where N is a positive integer such as 1, 2, 4, or 8, while the number of presets for the chroma components is fixed to 1. The number of presets for the chroma components does not need to be signaled in the bitstream and is derived as 1 in the decoder.
[0100] For example, FIG. 10 shows a Separate Constrained Directional Enhancement Filter (SCDEF).
[0101] In one embodiment, the selected preset indexes for the current luma block and chroma block may be different from each other. The inputs of this process are the luma / chroma reconstructed samples of the current block and the selected preset derived at the frame level. The output of this process is an index indicating which preset is selected for the current block.
[0102] In one example, when the luma component has 8 presets and the chroma component has 4 presets at the frame level, the preset index selected for luma block A is 7 and the preset index selected for chroma block B is 1. Luma block A and chroma block B are co-located or partially co-located.
[0103] In one embodiment, when deriving the CDEF filtering strength for the chroma components, the input reconstructed samples are determined by the current chroma coding block size.
[0104] In one example, if the current chroma block size is 32x64, the input is the chroma reconstructed sample values of the current 32x64 block.
[0105] In one embodiment, when disjoint or semi-detached partitioning is applied to luma and chroma blocks, the luma and chroma blocks still share the same preset index, and only the luma (or chroma) block size is adopted in the preset index derivation / signaling process.
[0106] In some embodiments, when the luma and chroma components have the same coding block size, the CDEF filtering processes for the luma and chroma components are performed separately.
[0107] In some embodiments, the signaling of the SCDEF is performed separately for the luma and chroma components.
[0108] In one embodiment, picture-level presets are signaled separately for the luma and chroma components, and these presets can be signaled in high-level parameter sets (DPS, VPS, SPS, PPS, APS), slice headers, picture headers, and SEI messages.
[0109] In one example, the luma presets are signaled first, then the chroma presets.
[0110] In one embodiment, the block level preset index is signaled separately for the luma and chroma components.
[0111] In one example, the preset index for the luma component is signaled first, followed by the preset index for the chroma component.
[0112] 11, an operational flowchart illustrating steps of a method 300 for decoding video data is shown. However, one skilled in the art will be able to understand how the encoding process operates based on FIG. 11. In some implementations, one or more processing blocks of FIG. 3 may be performed by computer 102 (FIG. 1) and server computer 114 (FIG. 1). In some implementations, one or more processing blocks of FIG. 3 may be performed by another device or group of devices separate from or including computer 102 and server computer 114.
[0113] At step 302, the method 300 includes receiving video data including chroma and luma components.
[0114] In step 304, the method 300 includes analyzing, deriving, or selecting a number of presets for the chroma components in a frame and a number of presets for the luma component in a frame.
[0115] At step 306, the method 300 includes encoding and / or decoding video data.
[0116] Operation 306 may be based on the number of presets for the chroma components in a frame and the number of presets for the luma component in a frame.
[0117] The method may further include performing a separate constrained directional emphasis filter (CDEF) process that filters the luma and chroma components independent of each other based on the number of presets for the chroma components in a frame and the number of presets for the luma component in a frame.
[0118] It should be understood that Figure 11 is provided as an illustration of only one implementation and is not intended to imply any limitations on how different embodiments may be implemented. Many modifications to the depicted environment may be made based on design and implementation requirements.
[0119] Figure 12 is a block diagram 400 of the internal and external components of the computer shown in Figure 1, according to an exemplary embodiment. It should be understood that Figure 4 is provided only as an example of one implementation and is not intended to be limiting with respect to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made based on design and implementation requirements.
[0120] The computer 102 (FIG. 1) and the server computer 114 (FIG. 1) may include respective sets of internal components 800A, 800B and external components 900A, 900B shown in FIG. 12. Each of the set of internal components 800 includes one or more processors 820, one or more computer-readable RAMs 822 and one or more computer-readable ROMs 824 on one or more buses 826, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.
[0121] The processor 820 may be implemented in hardware, firmware, or a combination of hardware and software. The processor 820 may be a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, the processor 820 includes one or more processors that can be programmed to perform functions. The bus 826 includes components that enable communication between the internal components 800A, 800B.
[0122] One or more operating systems 828, software programs 108 (FIG. 1), and video encoding programs 116 (FIG. 1) on server computer 114 (FIG. 1) are stored in one or more of respective computer-readable tangible storage devices 830 and executed by one or more of respective processors 820 via one or more of respective RAMs 822 (which typically include cache memory). In the embodiment shown in FIG. 12, each of computer-readable tangible storage devices 830 is an internal hard drive magnetic disk storage device. Alternatively, each of computer-readable tangible storage devices 830 is a semiconductor storage device such as ROM 824, EPROM, flash memory, optical disk, magneto-optical disk, solid-state disk, compact disk (CD), digital versatile disk (DVD), floppy disk, cartridge, magnetic tape, and / or another type of non-transitory computer-readable tangible storage device capable of storing computer programs and digital information.
[0123] Each set of internal components 800A, 800B also includes a R / W drive or interface 832 for reading from and writing to one or more portable computer-readable tangible storage devices 936, such as a CD-ROM, a DVD, a memory stick, a magnetic tape, a magnetic disk, an optical disk, or a semiconductor storage device. Software programs, such as software program 108 (FIG. 1) and video encoding program 116 (FIG. 1), may be stored on one or more of the respective portable computer-readable tangible storage devices 936, read via the respective R / W drive or interface 832, and loaded onto the respective hard drive 830.
[0124] Each set of internal components 800A, 800B also includes a network adapter or interface 836, such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G, 4G, or 5G wireless interface card or other wired or wireless communication link. The software program 108 (FIG. 1) and the video encoding program 116 (FIG. 1) on the server computer 114 (FIG. 1) can be downloaded to the computer 102 (FIG. 1) and the server computer 114 from an external computer via a network (e.g., the Internet, a local area network, or other wide area network) and the respective network adapter or interface 836. From the network adapter or interface 836, the software program 108 and the video encoding program 116 on the server computer 114 are loaded onto the respective hard drives 830. The network may include copper wire, optical fiber, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers.
[0125] Each of the sets of external components 900A, 900B may include a computer display monitor 920, a keyboard 930, and a computer mouse 934. The external components 900A, 900B may also include touch screens, virtual keyboards, touchpads, pointing devices, and other human interface devices. Each of the sets of internal components 800A, 800B also includes a device driver 840 for interfacing to the computer display monitor 920, the keyboard 930, and the computer mouse 934. The device driver 840, the R / W drive or interface 832, and the network adapter or interface 836 include hardware and software (stored in the storage device 830 and / or the ROM 824).
[0126] Although this disclosure includes detailed descriptions related to cloud computing, it should be understood in advance that implementations of the teachings recited herein are not limited to cloud computing environments. Rather, some embodiments may be practiced in conjunction with any other type of computing environment now known or later developed.
[0127] Cloud computing is a service delivery model that enables easy, on-demand network access to a shared pool of configurable computing resources (such as networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with the service provider. This cloud model includes at least five characteristics, at least three service models, and at least four deployment models.
[0128] The characteristics are as follows: On-demand self-service: Cloud consumers can independently provision computing capacity, such as server time or network storage, automatically as needed without the need for service provider interaction. Broad network access: Functionality is available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (eg, cell phones, laptops, and PDAs). Resource Pooling: Providers' computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated according to demand. Consumers generally have no control or knowledge over the exact location of the provided resources, but there is a sense of location independence in that they can specify location at a higher level of abstraction (e.g., country, state, or data center). Rapid Elasticity: Capacity can be provisioned quickly and elastically, sometimes automatically, and scaled out quickly or released quickly to scale in quickly. To the consumer, the capacity available for provisioning often appears unlimited, and they can purchase as much as they want, whenever they want. Measured Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported to provide transparency to both providers and consumers of utilized services.
[0129] The service model is as follows: Software as a Service (SaaS): The consumer is offered the ability to use a provider's applications that run on a cloud infrastructure. The applications are accessible from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, storage, or even individual application functions, except for limited user-specific application configuration settings. Platform as a Service (PaaS): The capability offered to consumers is the deployment of consumer-created or acquired applications, written using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does control the deployed applications and, in some cases, the application hosting environment configuration. Infrastructure as a Service (IaaS): The functionality offered to consumers is the provision of processing, storage, network, and other basic computing resources on which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does have control over the operating system, storage, deployed applications, and possibly limited control over select networking components (e.g., host firewalls).
[0130] The deployment models are as follows: Private Cloud: Cloud infrastructure is operated exclusively for an organization. It may be managed by the organization or a third party and may exist on-premise or off-premise. Community Cloud: Cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policy, and compliance considerations). It may be managed by the organization or a third party and may exist on-premises or off-premises. Public cloud: Cloud infrastructure is available to the general public or large industry groups and is owned by organizations that sell cloud services. Hybrid Cloud: A cloud infrastructure is a composition of two or more clouds (private, community, or public) that remain distinct entities but are tied together by standardized or proprietary technologies that allow for data and application portability (e.g., cloud bursting for load balancing between clouds).
[0131] Cloud computing environments are service-oriented with a focus on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0132] Referring to FIG. 13, an exemplary cloud computing environment 500 is illustrated. As shown, the cloud computing environment 500 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers can communicate, such as, for example, a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, and / or an automobile computer system 54N. The cloud computing nodes 10 can communicate with each other. They may be physically or virtually grouped (not shown) in one or more networks, such as the private cloud, community cloud, public cloud, hybrid cloud, or combinations thereof described above. This enables the cloud computing environment 600 to provide infrastructure, platform, and / or software as a service without the need for cloud consumers to maintain resources on their local computing devices. It should be understood that the types of computing devices 54A-N illustrated in FIG. 13 are for illustrative purposes only, and that the cloud computing nodes 10 and the cloud computing environment 500 can communicate with any type of computerized device over any type of network and / or network-addressable connection (e.g., using a web browser).
[0133] Referring to Figure 14, a set of functional abstraction layers 600 provided by the cloud computing environment 500 (Figure 13) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 6 are intended to be illustrative only, and embodiments are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0134] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (reduced instruction set computer) architecture-based servers 62, servers 63, blade servers 64, storage devices 65, and networks and network components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0135] The virtualization layer 70 provides an abstraction layer at which the following examples of virtual entities can be provided: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.
[0136] In one example, the management layer 80 may provide the functions described below. Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 provides allocation and management of cloud computing resources so that required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-allocation and procurement of cloud computing resources whose future requirements are anticipated according to SLAs.
[0137] The workload tier 90 provides examples of functions that can utilize a cloud computing environment. Examples of workloads and functions that can be provided from this tier include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and video encoding / decoding 96.
[0138] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of technical detail. The computer-readable media may include a computer-readable non-transitory storage medium having computer-readable program instructions for causing a processor to perform operations.
[0139] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge-in-groove structures with instructions recorded thereon, and any suitable combination of the above. As used herein, a computer-readable storage medium should not be construed as a transitory signal itself, such as an electric signal transmitted over a wire, an electromagnetic wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electric signal.
[0140] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.
[0141] The computer-readable program code / instructions for carrying out operations may be either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects or operations.
[0142] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to manufacture a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium having stored instructions includes an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0143] The computer-readable program instructions may also be loaded onto a computer, other programmable processing device, or other device to cause the computer, other programmable processing device, or other device to perform a series of operational steps to create a computer-implemented process, such that the instructions, executing on the computer, other programmable processing device, or other device, perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0144] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. The methods, computer systems, and computer-readable media may include additional, fewer, different, or differently arranged blocks than those shown in the figures. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may actually be executed concurrently or substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or executes a combination of dedicated hardware and computer instructions.
[0145] It will be apparent that the systems and / or methods described herein may be implemented in different forms, such as hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not intended to limit the scope of the invention. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it will be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0146] No element, act, or instruction used herein should be construed as critical or essential unless explicitly stated as such. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Furthermore, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, etc.) and may be used interchangeably with "one or more." Where only one item is intended, the term "one" or similar term is used. Also, as used herein, terms such as "has," "have," and "having" are intended to be open-ended terms. Furthermore, the phrase "based on" is intended to mean "based, at least in part on," unless otherwise specified.
[0147] The descriptions of various aspects and embodiments are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. While combinations of features are recited in the claims and / or disclosed herein, these combinations are not intended to limit the disclosure of possible implementations. Indeed, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. While each dependent claim listed below may depend directly on only one claim, the disclosure of possible implementations includes each dependent claim in combination with all other claims in the claim set. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been selected to best explain the principles of the embodiments, practical applications, or technical improvements to technology found in the marketplace, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0148] Acronyms used throughout this disclosure include the following: HEVC High Efficiency Video Coding HDR High Dynamic Range SDR Standard Dynamic Range VVC Versatile Video Coding JVET Joint Video Exploration Team MPM Most Probable Mode WAIP Wide-angle Intra Prediction CU Coding Unit CTB coding tree block PU Prediction Unit TU Conversion Unit CTU Coding Tree Unit PDPC Position-dependent Prediction Combination ISP Intra-subdivision SPS Sequence Parameter Set PPS Picture Parameter Set APS adaptive parameter set VPS Video Parameter Set DPS Decoding Parameter Set ALF Adaptive Loop Filter SAO pixel adaptive offset CC-ALF Cross-Component Adaptive Loop Filter CDEF Constrained Directional Enhancement Filter LR Loop Recovery Filter AV1 AOMedia Video 1 AV2 AOMedia Video 2 SDP Semi-Detached Partition SEI Supplemental Enhancement Information [Explanation of symbols]
[0149] 10 cloud computing nodes 54A Personal Digital Assistant (PDA) or Mobile Phone 54B Desktop Computer 54C Laptop Computer 54N Automotive Computer System 60 Hardware and Software Layers 61 Mainframe 62 RISC (Reduced Instruction Set Computer) architecture-based servers 63 servers 64 Blade Servers 65 Storage Devices 66 Networks and Network Components 67 Network Application Server Software 68 Database Software 70 Virtualization Layer 71 Virtual Servers 72 Virtual Storage 73 Virtual Networks 74 Virtual Applications and Operating Systems 75 Virtual Clients 80 Management layer 81 Resource Provisioning 82 Measurement and Pricing 83 User Portal 84 Service Level Management 85 Planning and Implementing Service Level Agreements (SLAs) 90 Workload Tier 91 Mapping and Navigation 92 Software Development and Lifecycle Management 93 Educational provision 94 Data Analysis Processing 95 Transaction Processing 96 Video Encoding / Decoding 100 Video Encoding System 102 Computer 104 processors 106 Data storage device 108 Software Programs 110 Communication Network 112 databases 114 Server Computer 116 Video Encoding Program 116 Programs 116 Video encoding or decoding program 500 Cloud Computing Environments 600 Functional Abstraction Layer 800A Internal Components 800B Internal Components 820 processor 822 computer readable RAM 824 computer readable ROM 826 Bus 828 Operating Systems 830 Hard Drives / Storage Devices 832 R / W drive or interface 836 Network Adapters or Interfaces 840 Device Drivers 900A External Components 900B External Components 920 Computer Display Monitor 930 keyboard 934 Computer Mouse 936 Portable computer-readable tangible storage device
Claims
1. 1. A processor-executable method of video decoding, the method comprising: receiving video data comprising a plurality of frames, including a current frame having chroma and luma components; determining that semi-detached partitioning (SDP) mode is enabled for the current frame; According to which SDP mode is enabled for the current frame, applying a first Constrained Directional Emphasis Filter (CDEF) process to the chroma components of the current frame; applying a second CDEF process to the luma component of the current frame, the first CDEF process being different from the second CDEF process; reconstructing the current frame according to the first CDEF process and the second CDEF process; A method comprising:
2. The method of claim 1 , wherein the first CDEF process and the second CDEF process are selected based on the chroma components having a different partition than the luma component.
3. 2. The method of claim 1 , wherein a first input of the first CDEF process comprises one or more reconstructed samples of the chroma components, and a second input of the second CDEF process comprises one or more reconstructed samples of the luma component.
4. The method further comprises: The method of claim 3 , comprising deriving a filtering strength of the first CDEF process based on a chroma-encoded block size of the chroma components.
5. After the step of applying a second CDEF process to the luma component of the current frame, deriving a first number of presets for the chroma components in a frame based on a first Constrained Directional Emphasis Filter (CDEF) process; deriving a second number of presets for the luma component in the one frame based on a second CDEF process, wherein the first number of presets is different from the second number of presets; The method of claim 1 , comprising:
6. The method of claim 5 , wherein the second number of presets is a positive integer and the first number of presets is fixed at one.
7. 6. The method of claim 5, wherein the step of deriving the first number of presets is separate and independent from the step of deriving the second number of presets.
8. The method further comprises: The method of claim 1 , comprising obtaining the same intermediate output of the first CDEF process and the second CDEF process.
9. The method of claim 1 , wherein the first CDEF process and the second CDEF process are selected based on a luma block size of the luma component.
10. The method of claim 1 , wherein the first CDEF process and the second CDEF process are selected based on a chroma block size of the chroma component.
11. The method of claim 1 , wherein the selected preset index for the current luma block is different from the selected preset index for the current chroma block.
12. the inputs of the separated CDEF processing are luma / chroma reconstructed samples of the current luma / chroma block and a selected preset derived at a frame level; The method of claim 11 , wherein the output of the separated CDEF processing is an index indicating which preset is selected for the current luma / chroma block.
13. Apparatus configured to carry out the method of any one of claims 1 to 12.
14. A computer program causing one or more computer processors to carry out the method of any one of claims 1 to 12.
Citation Information
Patent Citations
Video encoding or decoding device and video encoding or decoding method
JP2020107948A
Constrained directional enhancement filter selection for video coding
US20190045186A1
Adaptive in-loop filtering for video coding
US20190052877A1