Separated constrained directional enhancement filter
By independently filtering luma and chroma components with separate presets and block-level indices, the method addresses the inefficiencies of conventional CDEF, enhancing coding efficiency and performance in video encoding.
Patent Information
- Authority / Receiving Office
- JP Β· JP
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2026-02-26
- Publication Date
- 2026-05-19
AI Technical Summary
Existing video encoding formats like AV1 struggle with efficient filtering of chroma and luma components, particularly in situations where they have different or partially disconnected segments, leading to limitations in coding efficiency and performance.
The proposed method involves performing isolated constrained directional enhancement filtering (CDEF) to filter luma and chroma components independently, allowing for different numbers of presets at the picture level and different preset indices at the block level, with the chroma block size determining the input for filter intensity.
This approach enhances coding efficiency by enabling separate filtering of luma and chroma components, addressing the limitations of conventional CDEF and improving performance in scenarios with different partition schemes.
Smart Images

Figure 2026083172000001_ABST
Abstract
Description
[Technical Field]
[0001] Priority Information This application claims the benefit of priority from U.S. Provisional Application No. 63 / 040,856 filed on 18 June 2020 and U.S. Application No. 17 / 091,759 filed on 6 November 2020, which are incorporated herein by reference in their entirety.
[0002] This disclosure generally relates to the field of data processing, and more particularly to the encoding and / or decoding of video (e.g., by a coder, decoder, or codec (decoder and encoder)). [Background technology]
[0003] AOMedia Video 1 (AV1) is an open video encoding format designed for video transmission over the internet. It was developed, for example, as a successor to codec extensions of related technologies. [Overview of the project] [Means for solving the problem]
[0004] Embodiments relate to methods, systems, and computer-readable media for encoding and / or decoding video data. According to one embodiment, a method for encoding and / or decoding video data is provided. The method may include the steps of receiving video data including chroma and luma components, analyzing, deriving, or selecting the number of presets for chroma components and the number of presets for luma components in one frame, and decoding the video data, the method including performing isolated constrained directional enhancement filtering (CDEF) processing to filter the luma and chroma components independently of each other based on the number of presets for chroma components and the number of presets for luma components in one frame.
[0005] The method may include the steps of: performing a separated constrained directional emphasis filtering (CDEF) process to filter lumar and chroma components independently of each other when the lumar and chroma components have different or partially disconnected segments; and obtaining the output of the separated CDEF process, which includes filtered reconstructed samples of the lumar / chroma components, wherein the input to the separated CDEF process is a reconstructed sample of the lumar / chroma components, and the intermediate output of the separated CDEF process uses derived filter presets and block-by-block level preset indices.
[0006] The number of presets derived for the chroma component differs from the number of presets derived for the chroma component at the picture level.
[0007] The number of presets at the picture level can include one of 1, 2, 4, or 8.
[0008] The number of presets derived and selected for the lumens component within a single frame is 2, and the number of presets derived and selected for the chromens component within a single frame is 1.
[0009] The number of presets derived and selected for the lumen component is a positive integer N, while the number of presets for the chromen component is fixed at 1, which is derived as 1 in the decoder without signaling.
[0010] The preset index selected for the current lumen block differs from the preset index selected for the current chromen block. The input to the separated CDEF process is the lumen / chromen reconstruction sample of the current block, and the preset derived and selected at the frame level. The output of this process is an index indicating which preset is selected for the current block.
[0011] The method may further include selecting a preset index of 7 for lumen block A and a preset index of 1 for chroma block B when, at the frame level, the number of lumen components corresponds to 8 presets and the number of chroma components corresponds to 4 presets, such that lumen block A and chroma block B are in the same position or partially in the same position.
[0012] The method may further include the fact that, when deriving the CDEF filtering intensity of the chroma component, the input reconstructed sample is determined by the current chroma coding block size.
[0013] The method may further include the input being a chroma reconstruction sample value of the current block having a constant size, when the current chroma block is of a constant size.
[0014] The method may further include the fact that when separated or partially disconnected sections are applied to rumor and chroma blocks, the rumor and chroma blocks still share the same preset index, and in the preset index derivation / signaling process, only one of the rumor or chroma block sizes is adopted.
[0015] The method may further include the fact that when the lumern and chromar components have the same encoded block size, the CDEF filtering of the lumern and chromar components is performed separately.
[0016] Picture-level presets can be signaled separately for lumens and chromas within high-level parameter sets, slice headers, picture headers, or Supplementary Enhancement Information (SEI) messages.
[0017] The chroma preset may be signaled first, followed by the chroma preset.
[0018] The block-level preset index is signaled separately for the luma and chroma components.
[0019] The preset index of the luma component is signaled first, and then the preset index of the chroma component is signaled.
[0020] A computer system for decoding video data may include one or more computer-readable non-transitory storage media configured to store computer program code, and one or more computer processors configured to access the computer program code and operate as instructed by the computer program code. The computer program code may include reception code configured to cause one or more computer processors to receive video data including chroma and luma components, analysis, derivation or selection code configured to cause one or more computer processors to analyze, derive or select the number of presets of the chroma component and the number of presets of the luma component within one frame, and decoding code configured to cause one or more computer processors to decode the video data. The method includes performing separate constrained directional enhancement filter (CDEF) processing for filtering the luma and chroma components independently of each other based on the number of presets of the chroma component and the number of presets of the luma component within one frame.
[0021] A non - transient computer - readable medium storing a computer program for decrypting video data can be configured to cause one or more computer processors to receive video data including chroma components and luma components, and to cause one or more computer processors to perform analysis, derivation, or selection code configured to analyze, derive, or select the number of presets of chroma components in one frame and the number of presets of luma components in one frame, and to cause one or more computer processors to perform decryption code configured to decrypt the video data. The method includes performing separate constrained - directional enhancement filtering (CDEF) processing that filters luma and chroma components independently of each other based on the number of presets of chroma components in one frame and the number of presets of luma components in one frame.
[0022] These and other objects, features, and advantages will become apparent from the following detailed description of exemplary embodiments to be read in conjunction with the accompanying drawings. The illustrations are for the purpose of facilitating understanding by those skilled in the art in conjunction with the detailed description, and various features of the drawings are not to scale.
Brief Description of the Drawings
[0023] [Figure 1] Shows a networked computer environment according to at least one embodiment. [Figure 2] Shows the filter shape of an adaptive loop filter (ALF). [Figure 3A] Shows the subsample positions of a diagonal gradient. [Figure 3B] Shows the subsample positions of a diagonal gradient. [Figure 3C] Shows the subsample positions of a diagonal gradient. [Figure 3D] Shows the subsample positions of a diagonal gradient. [Figure 4] Shows the modified block classification at a virtual boundary. [Figure 5] Shows the modified ALF filtering for the luma component at a virtual boundary. [Figure 6] This shows the relative positions of the chromatic sample to the luminal sample. [Figure 7] Here is an example of a directional search using an 8x8 block. [Figure 8] Here is an example of a directional search using an 8x8 block. [Figure 9] An example of an encoded tree structure (lumer and chromer) is shown. [Figure 10] The separated constrained directional enhancement filter (SCDEF) is shown. [Figure 11] This is an operation flowchart illustrating the steps performed by a program that encodes video data, according to at least one embodiment. [Figure 12] This is a block diagram of the internal and external components of the computer and server shown in Figure 1, according to at least one embodiment. [Figure 13] This is a block diagram of an exemplary cloud computing environment, including the computer system shown in Figure 1, according to at least one embodiment. [Figure 14] This is a block diagram of the functional layer of the exemplary cloud computing environment shown in Figure 13, according to at least one embodiment. [Modes for carrying out the invention]
[0024] Detailed embodiments of the claimed structures and methods are disclosed herein, but it should be understood that the disclosed embodiments are merely illustrative of the claimed structures and methods, which may be embodied in various forms. These structures and methods may, however, be embodied in many different forms and should not be construed as being limited to the exemplary embodiments described herein. Rather, these exemplary embodiments are provided to ensure that this disclosure is detailed and complete and fully conveys its scope to those skilled in the art. In this description, well-known features and technical details may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0025] The embodiments generally relate to the field of data processing, and more particularly to the encoding and / or decoding of video. The exemplary embodiments described below provide, among other things, systems, methods, and computer programs for encoding and / or decoding video data.
[0026] As mentioned earlier, AOMedia Video 1 (AV1) is an open video encoding format designed for video transmission over the internet. It was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium established in 2015 with participation from semiconductor companies, video-on-demand providers, video content producers, software developers, and web browser vendors.
[0027] This specification describes various embodiments of methods, apparatuses (systems), and computer-readable media with reference to flowcharts and / or block diagrams. It will be understood that each block in the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0028] Referring here to Figure 1, a functional block diagram of a network computer environment is shown illustrating a video encoding system 100 (hereinafter, the "System") for encoding and / or decoding video data according to one embodiment. It should be understood that Figure 1 provides only an example of one embodiment and does not imply any limitation regarding the environment in which different embodiments may be implemented. Many modifications to the illustrated environment can be made based on design and implementation requirements.
[0029] System 100 may include a computer 102 and a server computer 114. Computer 102 can communicate with the server computer 114 via a communication network 110 (hereinafter referred to as the "network"). Computer 102 may include a processor 104 and a software program 108 that is stored in a data storage device 106, interfaces with a user, and can communicate with the server computer 114. As will be described later with reference to Figure 12, computer 102 may include internal components 800A and external components 900A, respectively, and the server computer 114 may include internal components 800B and external components 900B, respectively. Computer 102 may be, for example, a mobile device, telephone, personal digital assistant, netbook, laptop computer, tablet computer, desktop computer, or any type of computing device that can run programs, access a network, and access a database.
[0030] The server computer 114 can also operate in cloud computing service models such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), as will be discussed later with respect to Figures 13 and 14. The server computer 114 may also be deployed in cloud computing deployment models such as a private cloud, community cloud, public cloud, or hybrid cloud.
[0031] A server computer 114, which may be used to encode video data, can run a video encoding or decoding program 116 (hereinafter, the "program") that can interact with a database 112. The method of the video encoding or decoding program is described in more detail below with reference to Figure 3. In one embodiment, computer 102 can act as an input device including a user interface, and the program 116 can run primarily on the server computer 114. In an alternative embodiment, the program 116 can run primarily on one or more computers 102, and the server computer 114 can be used for processing and storing data used by the program 116. It should be noted that the program 116 may be a standalone program or may be integrated into a larger video encoding program. The video encoding or decoding program 116 may correspond to an encoder, decoder, or encoding (both encoder and decoder).
[0032] However, it should be noted that the processing of program 116 may, in some cases, be shared between computer 102 and server computer 114 in any ratio. In another embodiment, program 116 may run on multiple computers, server computers, or any combination of computers and server computers, for example, multiple computers 102 communicating with a single server computer 114 via network 110. In another embodiment, for example, program 116 may run on multiple server computers 114 communicating with multiple client computers via network 110. Alternatively, the program may run on a network server communicating with a server and multiple client computers via a network.
[0033] Network 110 may include wired connections, wireless connections, fiber optic connections, or any combination thereof. Generally, network 110 can be any combination of connections and protocols that support communication between computer 102 and server computer 114. Network 110 may include various types of networks, such as local area networks (LANs), wide area networks (WANs) such as the Internet, telecommunications networks such as public switched telephone networks (PSTNs), wireless networks, public switched networks, satellite networks, cellular networks (e.g., 5G networks, Long-Term Evolution (LTE) networks, 3G networks, Code Division Multiple Access (CDMA) networks, etc.), public land mobile networks (PLMNs), metropolitan area networks (MANs), private networks, ad hoc networks, intranets, fiber optic-based networks, and / or combinations of these or other types of networks.
[0034] The number and arrangement of devices and networks shown in Figure 1 are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks in a different arrangement than that shown in Figure 1. Furthermore, two or more devices shown in Figure 1 may be implemented within a single device, or a single device shown in Figure 1 may be implemented as multiple distributed devices. Furthermore, or alternatively, a set of devices in system 100 (e.g., one or more devices) may perform one or more functions that are described as being performed by another set of devices in system 100.
[0035] 1. Adaptive Loop Filter (ALF) In the Multipurpose Video Coding (VVC) (Draft 8), an Adaptive Loop Filter (ALF) with block-based filter adaptation is applied. For each 4x4 block, one of 25 filters is selected based on the direction and activity of the local gradient.
[0036] 1.1 Filter Shape VVC (Draft 8) allows the use of two diamond-shaped filter shapes (as shown in Figure 2). A 7x7 diamond shape is applied to the lumens component, and a 5x5 diamond shape is applied to the chromatic component.
[0037] 1.2 Block Classification For each lumen component, each 4x4 block is classified into one of 25 classes. The classification index C is determined by its direction D and the quantized value of its motion.
number
number
[0038] D and
number
number
[0039] Here, indices i and j represent the coordinates of the top-left sample within a 4x4 block, and R(i,j) represents the reconstructed sample at coordinates (i,j).
[0040] To reduce the complexity of block classification, a subsampled one-dimensional Laplacian calculation is applied. As shown in Figures 3A to 3D, the same subsample locations are used for gradient calculations in all directions (e.g., subsampled Laplacian calculations in all directions). For example, Figure 3A shows the subsample locations for vertical gradients, Figure 3B shows the subsample locations for horizontal gradients, and Figures 3C and 3D show the subsample locations for diagonal gradients.
[0041] Next, set the maximum and minimum values ββof the horizontal and vertical gradients D as follows:
number
[0042] The maximum and minimum values ββof the two diagonal gradients are set as follows:
number
[0043] To derive the value of directionality D, these values ββare compared to two thresholds t1 and t2. Step 1.
number
number
number
number
number
[0044] The motion value A is calculated as follows:
number
[0045] Furthermore, A is quantized in the range of 0 to 4 (including both ends), and its quantized value is
number
[0046] No classification method is applied to the chroma components within the picture; that is, a single set of ALF coefficients is applied to each chroma component.
[0047] 1.3 Geometric transformation of filter coefficients and clipping values Before filtering each 4x4 lumen block, geometric transformations such as rotation or diagonal or vertical flips are applied to the filter coefficients f(k, l) and the corresponding filter clipping values ββc(k, l) based on the gradient values ββcalculated for that block. This is equivalent to applying these transformations to samples in the filter support region. The idea is to make different blocks to which ALF has been applied more similar by aligning their orientations.
[0048] We introduce three geometric transformations: diagonal, vertical flip, and rotation. Diagonal: f D (k, l)=f(l, k), c D (k, l)=c(l, k), (Equation 9) Vertical flip: f V (k, l)=f(k, Kl-1), c V (k, l) = c(k, Kl-1) (Equation 10) Rotation: f R (k, l)=f(Kl-1, k), c R (k, l) = c(Kl-1, k) (Equation 11) Here, K is the filter size, 0β¦k, lβ¦K-1 are the coordinates of the coefficients, with (0, 0) being the top-left corner and (K-1, K-1) being the bottom-right corner. The transformation is applied to the filter coefficients f(k, l) and clipping value c(k, l) based on the gradient values ββcalculated in that block. The relationship between the transformation and the four gradients in the four directions is summarized in Table 1 below.
[0049] [Table 1]
[0050] 1.4 Filter Parameter Signaling In VVC (Draft 8), ALF filter parameters are signaled by an Adaptation Parameter Set (APS). A single APS can signal up to 25 sets of lumern filter coefficients and clipping value indices, and up to 8 sets of chromern filter coefficients and clipping value indices. To reduce bit overhead, filter coefficients for different classifications of lumern components can be merged. The slice header signals the index of the APS used for the current slice. In VVC (Draft 8), ALF signaling is CTU-based.
[0051] The clipping value index decoded from the APS can be used to determine the clipping values ββusing a table of lumen and chroma clipping values. These clipping values ββdepend on the internal bit depth. More precisely, the table of clipping values ββcan be obtained by the following formula: AlfClip={round(2 B-Ξ±*n ), nβ[0..N-1]} (Equation 12) Here, B is equal to the internal bit depth, Ξ± is a predefined constant equal to 2.35, and N is the number of clipping values ββallowed in VVC (Draft 8) equal to 4.
[0052] Table 2 shows the output of Equation 12.
[0053] [Table 2]
[0054] The slice header can signal up to seven APS indices to specify the rumor filter set used for the current slice. Filtering can be further controlled at the coding tree block (CTB) level. A flag is always signaled to indicate whether the ALF is applied to the rumor CTB. The rumor CTB can select a filter set from 16 fixed filter sets and filter sets from the APS. The rumor CTB's filter set index is signaled to indicate which filter set to apply. The 16 fixed filter sets are predefined and hardcoded in both the encoder and decoder.
[0055] The chroma component is signaled with an APS index in the slice header to indicate the chroma filter set used in the current slice. At the CTB level, if the APS has multiple chroma filter sets, each chroma CTB is signaled with a filter index.
[0056] The filter coefficients can be quantized to a norm equal to 128. To limit the complexity of the multiplication, bitstream fit is applied so that the non-centered coefficient values ββare in the range of -27 to 27-1 (including both ends). Centered coefficients are not signaled in the bitstream and are considered equal to 128.
[0057] In VVC (Draft 8), the syntax and semantics of the clipping index and value are defined as follows: alf_luma_clip_idx[ sfIdx ][ j ] specifies the clipping index of the clipping value to use before multiplying by the j-th coefficient of the signaled lumar filter indicated by sfIdx. The bitstream conformance requirement is that the value of alf_luma_clip_idx[ sfIdx ][ j ] is in the range of 0 to 3 (inclusive) when sfIdx = 0 ..alf_luma_num_filters_signalled_minus1 and j = 0 ..11.
[0058] The luma filter clipping value AlfClipL[adaptation_parameter_set_id], having the element AlfClipL[adaptation_parameter_set_id][filtIdx][j] with filtIdx=0..NumAlfFilters-1 and j=0..11, is derived as specified in Table 2, depending on bitDepth set to equal BitDepthY and clipIdx set to equal alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx]][j].
[0059] alf_chroma_clip_idx[altIdx][j] specifies the clipping index of the clipping value to use, and then multiplies it by the j-th coefficient of the alternative chroma filter at index altIdx. A bitstream compatibility requirement is that the value of alf_chroma_clip_idx[altIdx][j] for altIdx=0..alf_chroma_num_alt_filters_minus1 and j=0..5 be in the range of 0 to 3 (including both ends).
[0060] A chroma filter clipping value AlfClipC[adaptation_parameter_set_id][altIdx] with element AlfClipC[adaptation_parameter_set_id][altIdx] where altIdx=0..alf_chroma_num_alt_filters_minus1 and j=0..5 is derived as specified in Table 2, depending on bitDepth set to equal BitDepthC and clipIdx set to equal alf_chroma_clip_idx[altIdx][j].
[0061] 1.5 Filtering process On the decoder side, when ALF is enabled for CTB, each sample R(i, j) in CU is filtered, resulting in the sample value R'(i, j) shown below.
number
number
number
[0062] 1.6 Virtual boundary filtering process for line buffer reduction To reduce the ALF line buffer requirements, modified block classification and filtering are employed for samples near the horizontal CTU boundary. For this purpose, as shown in Figure 4, the virtual boundary can be defined as a line by shifting the horizontal CTU boundary by "N" samples, where "N" is equal to 4 for the lumens and 2 for the chromens.
[0063] As illustrated in Figure 4, a modified block classification is applied to the rumor component. For the one-dimensional Laplacian gradient calculation of 4x4 blocks above the virtual boundary, only samples above the virtual boundary are used. Similarly, for the one-dimensional Laplacian gradient calculation of 4x4 blocks below the virtual boundary, only samples below the virtual boundary are used. The quantization of the motion value A is scaled accordingly to account for the reduction in the number of samples used in the one-dimensional Laplacian gradient calculation.
[0064] In the filtering process, symmetrical padding operations are performed on both the lumern and chromar components at the virtual boundary. As shown in Figure 5 ("Modified ALF filtering for lumern components at the virtual boundary"), when a sample to be filtered is located below the virtual boundary, neighboring samples located above the virtual boundary are padded. Conversely, the corresponding samples on the other side are also padded symmetrically.
[0065] 1.7 Largest Encoding Unit (LCU) - Aligned Picture Quadtree Partitioning To improve coding efficiency, JCTVC-C143 [3] proposes an adaptive loop filter based on a coding unit-synchronous picture quadtree. The lumapicture is divided into multiple multilevel quadtree partitions, the boundaries of which are aligned to the boundaries of the largest coding unit (LCU). Each partition has its own filtering process and is therefore called a filter unit (FU).
[0066] The two-pass coding flow is described below. In the first pass, the quadtree partitioning pattern and the optimal filter for each FU are determined. During the determination process, filtering distortion is estimated by FFDE. The reconstructed picture is filtered according to the determined quadtree partitioning pattern and the selected filters for all FUs. In the second pass, the CU-synchronous ALF is turned on / off. According to the ALF on / off result, the initially filtered picture is partially restored by the reconstructed picture.
[0067] A top-down partitioning method is employed, and the picture is divided into multilevel quadtree partitions using a rate distortion criterion. Each partition is called a filter unit. In the partitioning process, the quadtree partitions are aligned to the LCU boundary. The encoding order of the FUs follows the z-scan order. For example, a picture can be divided into 10 FUs, and the encoding order is FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, FU9.
[0068] To indicate a picture quadtree partitioning pattern, the partitioning flags can be encoded and transmitted in Z-order.
[0069] The filter for each FU can be selected from two sets of filters based on rate distortion criteria. The first set may have newly derived 1 / 2 symmetric square and rhombus filters for the current FU. The second set can be obtained from a time-delay filter buffer, which stores filters previously derived for the FU of previous pictures. The filter with the minimum rate distortion cost from these two sets can be selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further divided into four child FUs, the rate distortion costs of the four child FUs are calculated. By recursively comparing the rate distortion costs with and without division, the picture quadtree division pattern can be determined.
[0070] In JCTVC-C143, the maximum quadtree partitioning level is 2, meaning the maximum number of FUs is 16. During quadtree partitioning decisions, the correlation values ββused to derive the Wiener coefficients of the 16 FUs at the lowest quadtree level (minimum FU) can be reused. The remaining FUs can have their Wiener filters derived from the correlations of the 16 FUs at the lowest quadtree level. Therefore, the framebuffer is accessed only once to derive the filter coefficients for all FUs.
[0071] After the quadtree partitioning pattern is determined, CU-synchronous ALF on / off control is performed to further reduce filtering distortion. By comparing the filtering distortion with the unfiltering distortion, leaf CUs can explicitly switch the ALF on or off in their local region. Encoding efficiency can be further improved by redesigning the filter coefficients according to the ALF on / off result. However, the redesign process requires access to an additional frame buffer. In the proposed CS-PQALF encoder design, to minimize the number of frame buffer accesses, there is no redesign process after the CU-synchronous ALF on / off determination.
[0072] 2. Cross-component adaptive loop filter Cross-component adaptive loop filtering (CC-ALF) refines each chromatic component using lumens sample values.
[0073] CC-ALF operates by applying a linear diamond-shaped filter to the lumen channel of each chroma component. The filter coefficients are transmitted via APS, and 2 10The values ββare scaled to double and rounded for fixed-point representation. Filter application is controlled by a variable block size and signaled by a context coding flag received for each block of samples. The block size, along with the CC-ALF enable flag, is received at the slice level of each chroma component. The following block sizes (in chroma samples) were supported in the contribution: 16x16, 32x32, and 64x64.
[0074] Table 3 below explains the syntax changes in CC-ALF.
[0075] [Table 3]
[0076] The semantics of CC-ALF related syntax are explained below.
[0077] If alf_ctb_cross_component_cb_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] is equal to 0, it indicates that the cross-component Cb filter is not applied to the block of Cb color component samples in the lumar location (xCtb, yCtb).
[0078] If alf_cross_component_cb_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] is not equal to 0, it indicates that the alf_cross_component_cb_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]-th cross-component Cb filter is applied to the block of Cb color component samples at the lumar location (xCtb, yCtb).
[0079] If alf_ctb_cross_component_cr_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] is equal to 0, it indicates that the cross-component Cr filter is not applied to the block of Cr color component samples in the lumaralocation (xCtb, yCtb). If alf_cross_component_cr_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] is not equal to 0, it indicates that the alf_cross_component_cr_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]-th cross-component Cr filter is applied to the block of Cr color component samples in the lumaralocation (xCtb, yCtb).
[0080] 3. Chroma sampling format Figure 6 of the Chroma application ("Position of Chroma Sample relative to Luma Sample") shows the relative position of the top-left chroma sample when chroma_format_idc is 1 (4:2:0 chroma format) and chroma_sample_loc_type_top_field or chroma_sample_loc_type_bottom_field is equal to the value of the variable ChromaLocType. The region represented by the top-left 4:2:0 chroma sample (illustrated as a large red square with a large red dot in the center) is shown relative to the region represented by the top-left luma sample (illustrated as a small black square with a small black dot in the center). Regions represented by adjacent luma samples are shown as small gray squares with a small gray dot in the center.
[0081] 4. Constrained Directional Enhancement Filter The primary purpose of a loop-constrained directional enhancement filter (CDEF) is to remove coding artifacts while preserving image detail. In HEVC, the pixel adaptive offset (SAO) algorithm achieves a similar objective by defining signaling offsets for pixels of different classes. Unlike SAO, CDEF is a nonlinear spatial filter. The filter design is constrained to be easily vectorizable (i.e., implementable in SIMD operation), which is not the case with other nonlinear filters such as median filters or bilateral filters.
[0082] The design of CDEF is based on the following observations: The amount of ringing artifacts in the encoded image tends to be approximately proportional to the quantization step size. While the amount of detail is a property of the input image, the minimum detail retained in the quantized image also tends to be proportional to the quantization step size. For a given quantization step size, the amplitude of ringing is generally smaller than the amplitude of detail.
[0083] CDEF identifies the orientation of each block and adaptively filters along the identified orientation and along the orientation rotated 45 degrees from the identified orientation, with the smallest possible angle. The filter intensity is explicitly signaled, allowing for high control over blurring. An efficient encoder search is designed for the filter intensity. CDEF is based on two previously proposed in-loop filters, and the combined filter has been adopted for the new AV1 codec.
[0084] 4.1 Direction search Direction exploration is performed immediately after the deblocking filter for the reconstructed pixels. Since these pixels are available at the decoder, the direction does not require signaling. The exploration operates on 8Γ8 blocks, which are small enough to appropriately handle non-linear edges and large enough to reliably estimate the direction when applied to the quantized image. By giving a certain directionality to the 8Γ8 region, vectorization of the filter becomes easier. For each block, the direction that best matches the pattern within the block is determined by minimizing the sum of squared differences (SSD) between the quantized block and the closest fully directional block. A fully directional block is a block in which all pixels along a line in a certain direction have the same value. Figure 7 is an example of direction exploration for an 8Γ8 block. In this case, the 45-degree direction (indicated by the box surrounding column 12) that minimizes the error is selected.
[0085] 4.2 Nonlinear Low-Pass Directional Filter The main reason for identifying the direction is to align the filter taps along that direction while maintaining the directional edges or patterns and reduce ringing. However, only directional filtering may not be sufficient to reduce ringing significantly. It is also desirable to use filter taps for pixels not along the main direction. To reduce the risk of blurring, these extra taps are handled more conservatively. For this reason, CDEF defines primary taps and secondary taps. The complete 2D CDEF filter is expressed as follows.
Equation
[0086] 5. AV1 Loop Restoration In general, to remove noise and improve edge quality, a set of in-loop restoration schemes have been proposed for use in post-deblocking video encoding, in addition to conventional deblocking operations. These schemes are switchable within a frame for each appropriately sized tile. The specific scheme described is based on a separable symmetric Wiener filter and a dual self-guided filter with subspace projection. Because content statistics can change substantially within a frame, these tools are integrated into a switchable framework that allows different tools to be invoked in different areas of the frame.
[0087] 5.1 Separable Symmetric Wiener Filters One restoration tool that has been shown to be promising in the literature is the Wiener filter. Every pixel of a degraded frame can be reconstructed as a non-causal filtered version of the pixels in a wΓw window around it, where w=2r+1 is odd for an integer r. The 2D filter tap is in the form of a column vectorized w 2 If it is represented by a single-element vector F, then by direct LMMSE optimization, F=H -1 The filter parameters given by M are derived here. Here, H = E[XX T ] is the autocovariance of x, and w in a wΓw window around the pixel. 2 This is a vectorized version of the sample in the column direction, where M=E[YX T] is the cross-correlation between x and the scalar source sample y that should be estimated. The encoder can estimate H and M from the realization of the deblocked frame and source and send the resulting filter F to the decoder. However, if we do that, w 2 Transmitting individual taps incurs a considerable bitrate cost, and the inseparable filtering makes decoding extremely complex. Therefore, several additional constraints are imposed on the properties of F. First, F is constrained to be separable, and filtering can be implemented as separable horizontal and vertical w-tap convolutions. Second, each horizontal and vertical filter is constrained to be symmetric. Third, it is assumed that the sum of the coefficients of both the horizontal and vertical filters is 1.
[0088] 5.2 Dual self-guided filtering using subspace projection Guided filtering is one of the recent paradigms in image filtering, where the local linear model is as follows: y = Fx + G (Equation 15) The above equation is used to calculate the filtered output y from an unfiltered sample x. Here, F and G are determined based on statistics of the degraded image and the guidance image of the neighboring pixels of the filtered image. If the guide image is the same as the degraded image, the resulting so-called self-guided filtering has the effect of edge-preserving smoothing. The specific forms of self-guided filtering proposed by the inventors depend on two parameters, radius and noise parameter e, and are listed below. 1. The mean ΞΌ and variance Ο of each pixel in a (2r+1)Γ(2r+1) window around each pixel. 2 We seek this. This can be efficiently implemented using box filtering based on integral imaging. 2. Calculate for all pixels: f = Ο 2 / (Ο 2 +e); g=(1-f)ΞΌ 3. Calculate F and G for all pixels as the average of the f and g values ββwithin a 3x3 window around the pixel being used.
[0089] Filtering is controlled by r and e, where a larger r results in greater spatial variance, and a larger e results in greater range variance.
[0090] The principle of subspace projection is schematically shown in Figure 8. Even if neither of the inexpensive reconstructions X1 and X2 are close to the source Y, a suitable multiplier {Ξ±,Ξ²} can bring them considerably closer to the source, as long as they are moving in a somewhat correct direction. Figure 8 shows a subspace projection that uses inexpensive reconstructions to produce a final reconstruction that is closer to the source.
[0091] 6. Sections that have been partially disconnected The semi-decoupled partitioning (SDP) scheme, or semi-separated tree (SST) or flexible block partitioning for chroma components. In this method, rumor blocks and chroma blocks within a single superblock (SB) can have the same or different block partitions, depending on the rumor coding block size or rumor tree depth. Specifically, if the area size of a rumor block is greater than one threshold T1, or if the coding tree partitioning depth of a rumor block is less than or equal to one threshold T2, the chroma block uses the same coding tree structure as the rumor. Otherwise, if the block area size is less than or equal to T1, or the rumor partitioning depth is greater than T2, the corresponding chroma block can have different coding block partitions from the rumor component, which is called flexible block partitioning for chroma components. T1 is a positive integer such as 128 or 256. T2 is a positive integer such as 1 or 2.
[0092] An improved semi-disconnected partition (SDP) scheme is proposed, in which the rumor and chroma components can share a subtree structure from the root node of the superblock, and the conditions for initiating separate tree partitions between the rumor and chroma depend on the rumor partition information. For example, Figure 9 shows an example of the encoded tree structure of the rumor and chroma components.
[0093] In a Constrained Directional Enhancement Filter (CDEF), the lumar and chroma components are restricted to sharing a preset at the picture level. Furthermore, the lumar and chroma components are restricted to having the same preset index at the block level. Finally, when deriving the filter intensity of the chroma component, the lumar block size is used to determine the input for the chroma component. These constraints may limit the coding efficiency of the CDEF.
[0094] In conventional CDEF, a single preset contains primary and secondary intensities for both lumern and chroman components. The number of allowed / available presets is signaled at the picture level. At the coding block level, an index indicating which preset is selected for the current block is signaled. CDEF coding block sizes include 128x128, 128x64, 64x64, and 64x128. Conventional CDEF has three limitations. The first limitation is that the lumern and chroman components must share a preset at the picture level. Another limitation is that the lumern and chroman components must select the same preset index at the block level. The third limitation is that when deriving the filter intensity of the chroman component, the lumern block size must be used to determine the input for the chroman component. The aforementioned limitations may together limit the performance of CDEF, especially in situations where the lumern and chroman components have different partition schemes, such as partition schemes in semi-disconnected partitions (SDPs).
[0095] This book proposes a Detached Constrained Directional Enhancement Filter (SCDEF) that performs CDEF processing on lumar and chroma components separately. Compared to conventional CDEF, SCDEF allows for independent filtering of lumar and chroma components. More specifically, the lumar and chroma components may have different numbers of presets at the picture level, and furthermore, they may select different preset indices at the block level, and the chroma block size is used to determine the input for the chroma component when deriving the filter intensity of the chroma component.
[0096] As shown in Figure 12, when the lumern and chroma components have different or partially disconnected segments, it has been proposed to perform CDEF filtering of the lumern and chroma components separately. The input to the CDEF filtering process is the reconstructed sample of the lumern / chroma components. The intermediate outputs of this process include, but are not limited to, the derived filter preset and the level preset index for each block, as described in the proposed method above. The final output of this process is the filtered reconstructed sample of the lumern / chroma components.
[0097] In one embodiment, the number of presets derived for the lumens and chroma components may differ from each other at the picture level. The input to the CDEF filtering process is a reconstructed sample of the lumens / chroma components. The output of this process is the presets derived at the picture level. Examples of the number of presets at the picture level include, but are not limited to, 1, 2, 4, and 8.
[0098] In one example, the number of presets derived and selected for the current lumen component within a frame is 2, and the number of presets derived and selected for the current chromen component of this frame is 1.
[0099] In another example, the number of presets derived and selected for the lumens component is N, where N is a positive integer such as 1, 2, 4, or 8, but the number of presets for the chromens component is fixed at 1. The number of presets for the chromens component does not need to be signaled in the bitstream and is derived as 1 in the decoder.
[0100] For example, Figure 10 shows a separated constrained direction-enhanced filter (SCDEF).
[0101] In one embodiment, the selected preset indices for the current rumor block and chroma block may be different from each other. The input to this process is the rumor / chroma reconstruction sample of the current block and the preset derived and selected at the frame level. The output of this process is an index indicating which preset is selected for the current block.
[0102] In one example, when the lumens component has 8 presets and the chromens component has 4 presets at the frame level, the preset index selected for lumens block A is 7, and the preset index selected for chromens block B is 1. Lumens block A and chromens block B are either in the same position or partially in the same position.
[0103] In one embodiment, when deriving the CDEF filtering intensity of the chroma component, the input reconstruction sample is determined by the current chroma coding block size.
[0104] For example, if the current chroma block size is 32x64, the input is the chroma reconstruction sample value for the current 32x64 block.
[0105] In one embodiment, when separated or partially disconnected sections are applied to a rumor block and a chroma block, the rumor block and chroma block still share the same preset index, and only the rumor (or chroma) block size is used in the preset index derivation / signaling process.
[0106] In some embodiments, when the lumern and chroma components have the same encoding block size, the CDEF filtering of the lumern and chroma components is performed separately.
[0107] In some embodiments, SCDEF signaling is performed separately for the lumern and chromatic components.
[0108] In one embodiment, picture-level presets are signaled separately for lumens and chroma components. These presets can be signaled with high-level parameter sets (DPS, VPS, SPS, PPS, APS), slice headers, picture headers, and SEI messages.
[0109] In one example, the chroma preset is signaled first, followed by the chroma preset.
[0110] In one embodiment, block-level preset indices are signaled separately for lumens and chroma components.
[0111] In one example, the preset index for the lumens component is signaled first, followed by the preset index for the chromens component.
[0112] Referring now to Figure 11, an operational flowchart illustrating the steps of method 300 for decoding video data is shown. However, those skilled in the art will be able to understand how the encoding process works based on Figure 11. In some implementations, one or more processing blocks in Figure 3 can be executed by computer 102 (Figure 1) and server computer 114 (Figure 1). In some implementations, one or more processing blocks in Figure 3 can be executed by a separate device or group of devices that is separate from or includes computer 102 and server computer 114.
[0113] In step 302, method 300 includes receiving video data that includes chroma and luma components.
[0114] In step 304, method 300 includes the steps of analyzing, deriving, or selecting the number of chroma component presets and the number of luma component presets within a single frame.
[0115] In step 306, method 300 includes the step of encoding and / or decoding video data.
[0116] Operation 306 may be based on the number of chroma component presets in a single frame and the number of lumen component presets in a single frame.
[0117] The method may further include the step of performing a separated constrained directional enhancement filter (CDEF) process that filters out the lumar and chroma components independently of each other, based on the number of chroma component presets within a frame and the number of lumar component presets within a frame.
[0118] Figure 11 provides only one example of an implementation and should be understood as not implying any limitations on how different embodiments may be implemented. Many modifications to the illustrated environment can be made based on design and implementation requirements.
[0119] Figure 12 is a block diagram 400 of the internal and external components of the computer shown in Figure 1, according to an exemplary embodiment. It should be understood that Figure 4 provides only an example of one embodiment and does not imply any limitation regarding the environment in which different embodiments may be implemented. Many modifications to the illustrated environment can be made based on design and implementation requirements.
[0120] Computer 102 (Figure 1) and server computer 114 (Figure 1) may include sets of internal components 800A, 800B and external components 900A, 900B, respectively, as shown in Figure 12. Each set of internal components 800 includes one or more processors 820, one or more computer-readable RAMs 822 and one or more computer-readable ROMs 824 on one or more buses 826, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.
[0121] The processor 820 is implemented as hardware, firmware, or a combination of hardware and software. The processor 820 is a central processing unit (CPU), graphics processing unit (GPU), acceleration unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or another type of processing component. In some embodiments, the processor 820 includes one or more processors that can be programmed to perform functions. The bus 826 includes components that enable communication between internal components 800A and 800B.
[0122] One or more operating systems 828, software programs 108 (Figure 1), and video encoding programs 116 (Figure 1) on the server computer 114 (Figure 1) are stored in one or more computer-readable tangible memory devices 830 and executed by one or more processors 820 via one or more RAMs 822 (typically including cache memory). In the embodiment shown in Figure 12, each of the computer-readable tangible memory devices 830 is a magnetic disk storage device of an internal hard drive. Alternatively, each of the computer-readable tangible memory devices 830 is a semiconductor storage device such as a ROM 824, EPROM, flash memory, optical disk, magneto-optical disk, solid-state disk, compact disk (CD), digital versatile disk (DVD), floppy disk, cartridge, magnetic tape, and / or another type of non-temporary computer-readable tangible memory device capable of storing computer programs and digital information.
[0123] Each set of internal components 800A and 800B also includes an R / W drive or interface 832 for reading from and writing to one or more portable computer-readable tangible storage devices 936, such as CD-ROMs, DVDs, memory sticks, magnetic tapes, magnetic disks, optical disks, or semiconductor storage devices. Software programs, such as software program 108 (Figure 1) and video encoding program 116 (Figure 1), can be stored in one or more of the respective portable computer-readable tangible storage devices 936, read via their respective R / W drives or interfaces 832, and loaded into their respective hard drives 830.
[0124] Each set of internal components 800A and 800B also includes a network adapter or interface 836, such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G, 4G, or 5G wireless interface card or other wired or wireless link. The software program 108 (Figure 1) and video encoding program 116 (Figure 1) on the server computer 114 (Figure 1) can be downloaded from an external computer to computer 102 (Figure 1) and the server computer 114 via the network (e.g., the Internet, a local area network, or another wide area network) and the respective network adapter or interface 836. From the network adapter or interface 836, the software program 108 and video encoding program 116 on the server computer 114 are loaded onto their respective hard drives 830. The network can include copper, fiber optic, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
[0125] Each of the sets of external components 900A and 900B may include a computer display monitor 920, a keyboard 930, and a computer mouse 934. External components 900A and 900B may also include a touchscreen, a virtual keyboard, a touchpad, a pointing device, and other human interface devices. Each of the sets of internal components 800A and 800B may also include a device driver 840 for interface with the computer display monitor 920, the keyboard 930, and the computer mouse 934. The device driver 840, the R / W drive or interface 832, and the network adapter or interface 836 include hardware and software (stored in the storage device 830 and / or ROM 824).
[0126] While this disclosure includes a detailed description of cloud computing, it should be understood in advance that the embodiments of the teachings enumerated herein are not limited to cloud computing environments. Rather, some embodiments can be implemented in conjunction with any other type of computing environment that is currently known or may be developed in the future.
[0127] Cloud computing is a service delivery model that enables easy, on-demand network access to a shared pool of configurable computing resources (such as networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be quickly prepared and deployed with minimal management effort and interaction with service providers. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.
[0128] The characteristics are as follows: On-demand self-service: Cloud consumers can independently prepare computing power, such as server time and network storage, automatically as needed, without requiring interaction with a service provider. Extensive network access: Functionality is available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs). Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated as needed. Consumers generally have a sense of location independence, in that they do not have control or knowledge of the exact location of the resources provided, but can specify the location at a higher level of abstraction (e.g., country, state, or data center). Rapid Flexibility: Capabilities can be prepared quickly and flexibly, and in some cases automatically, to scale out rapidly or release quickly to scale in rapidly. To consumers, the available capacity for preparation often appears unlimited and can be purchased as much as they like, whenever they like. Measured Services: Cloud systems automatically control and optimize resource usage by leveraging metric capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported to provide transparency to both the providers and consumers of the services being used.
[0129] The service model is as follows: Software as a Service (SaaS): The ability offered to consumers is the use of a provider's applications running on cloud infrastructure. These applications are accessible from various client devices via thin client interfaces, such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, storage, or even individual application functionalities, except for limited user-specific application configuration settings. Platform as a Service (PaaS): The functionality offered to consumers is the deployment of applications they have created or acquired, written using programming languages ββand tools supported by the provider, onto a cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they do control the deployed applications and, in some cases, the configuration of the application hosting environment. Infrastructure as a Service (IaaS): The functionality provided to consumers is the provision of processing, storage, networking, and other basic computing resources, which consumers can deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but have limited control over the operating system, storage, deployed applications, and possibly selected networking components (e.g., host firewalls).
[0130] The deployment model is as follows: Private Cloud: A cloud infrastructure is operated exclusively for a specific organization. It may be managed by the organization or a third party, and may reside on-premises or off-premises. Community Cloud: A cloud infrastructure is shared by several organizations and supports a specific community with shared interests (e.g., mission, security requirements, policies, and compliance considerations). It may be managed by an organization or a third party and may reside on-premises or off-premises. Public cloud: Cloud infrastructure is provided to the general public or large industry groups and is owned by organizations that sell cloud services. Hybrid Cloud: Cloud infrastructure is a configuration of two or more clouds (private, community, or public) that remain separate entities but are linked together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0131] Cloud computing environments are service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is infrastructure, including a network of interconnected nodes.
[0132] Referring to Figure 13, an exemplary cloud computing environment 500 is illustrated. As shown in the figure, the cloud computing environment 500 comprises one or more cloud computing nodes 10 that can communicate with local computing devices used by cloud consumers, such as a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, and / or an automotive computer system 54N. The cloud computing nodes 10 can communicate with each other. They may be physically or virtually grouped (not shown) in one or more networks, such as the private cloud, community cloud, public cloud, hybrid cloud, or a combination thereof. This allows the cloud computing environment 500 to provide infrastructure, platform, and / or software as a service, eliminating the need for cloud consumers to maintain resources on their local computing devices. The types of computing devices 54A-N shown in Figure 13 are for illustrative purposes only, and it should be understood that the cloud computing nodes 10 and the cloud computing environment 500 can communicate with any type of computerized device via any type of network and / or network addressable connection (e.g., using a web browser).
[0133] Referring to Figure 14, a set of functional abstraction layers 600 provided by the cloud computing environment 500 (Figure 13) is shown. It should be understood that the components, layers, and functionalities shown in Figure 6 are for illustrative purposes only, and embodiments are not limited thereto. As illustrated, the following layers and corresponding functionalities are provided:
[0134] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include a mainframe 61, a RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage devices 65, and a network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0135] The virtualization layer 70 provides an abstraction layer that may provide the following examples of virtual entities: virtual servers 71, virtual storage 72, virtual networks 73 including virtual private networks, virtual applications and operating systems 74, and virtual clients 75.
[0136] For example, the management layer 80 can provide the following functions: Resource provisioning 81 provides the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Measurement and pricing 82 provides cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. For example, these resources may include application software licenses. Security provides identification and verification for cloud consumers and tasks, as well as protection for data and other resources. The user portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 provides the allocation and management of cloud computing resources to ensure that the required service levels are met. Service level agreement (SLA) planning and execution 85 provides the pre-positioning and procurement of cloud computing resources whose future requirements are expected to conform to the SLA.
[0137] The workload layer 90 provides examples of capabilities that can leverage a cloud computing environment. Examples of workloads and capabilities that may be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, education delivery in virtual classrooms 93, data analysis processing 94, transaction processing 95, and video encoding / decoding 96.
[0138] Some embodiments may relate to systems, methods, and / or computer-readable media in integration at any possible level of technical detail. The computer-readable media may include computer-readable non-temporary storage media having computer-readable program instructions for causing a processor to perform an action.
[0139] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction-executing device. A computer-readable storage medium may, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random-access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital multipurpose disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or grooved raised structures on which instructions are recorded, and any suitable combination thereof. As used herein, computer-readable storage media should not be construed as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., optical pulses through optical fiber cables), or transient signals themselves, such as electrical signals transmitted through wires.
[0140] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.
[0141] The computer-readable program code / instructions for performing an operation may be either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or object-oriented programming languages ββsuch as Smalltalk and C++, and procedural programming languages ββsuch as the C programming language or similar programming languages. The computer-readable program instructions can run entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or it may be connected to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by personalizing the electronic circuit using state information of computer-readable program instructions in order to perform an action or operation.
[0142] These computer-readable program instructions may be provided to a general-purpose computer, a dedicated computer, or a processor of another programmable data processing device for manufacturing a machine, such that instructions executed via the processor of a computer or other programmable data processing device create means for performing functions / operations specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct computers, programmable data processing devices, and / or other devices to function in a particular way, and as a result, a computer-readable storage medium having stored instructions may include a product containing instructions that perform modes of functions / operations specified in one or more blocks of a flowchart and / or block diagram.
[0143] Computer-readable program instructions may also be loaded onto a computer, other programmable processing unit, or other device to generate a computer-executed process by causing the computer, other programmable processing unit, or other device to perform a series of action steps on the computer, other programmable processing unit, or other device so that the instructions executed on the computer, other programmable processing unit, or other device perform a function / operation specified in one or more blocks of a flowchart and / or block diagram.
[0144] The flowcharts and block diagrams in the figures illustrate the architecture, function, and operation of possible embodiments of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions containing one or more executable instructions for implementing a specified logical function. Methods, computer systems, and computer-readable media may include more blocks, fewer blocks, different blocks, or different arrangements of blocks than those shown in the figures. In some alternative embodiments, the functions described in the blocks may be performed in an order different from the order shown in the figures. For example, two blocks shown consecutively may actually be executed simultaneously or substantially simultaneously, or blocks may sometimes be executed in reverse order depending on the functions they relate to. It should also be noted that each block in a block diagram and / or flowchart, as well as combinations of blocks in a block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs a specified function or operation, or a combination of dedicated hardware and computer instructions.
[0145] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or combinations of hardware and software. Actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting to the embodiments. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware may be designed to implement the systems and / or methods based on the descriptions herein.
[0146] Any elements, actions, or instructions used herein should not be construed as important or essential unless expressly stated otherwise. Furthermore, where used herein, the articles βaβ and βanβ are intended to include one or more items and may be used interchangeably with βone or more.β Additionally, where used herein, the term βsetβ is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items) and may be used interchangeably with βone or more.β When only one item is intended, the term βoneβ or a similar term is used. Furthermore, where used herein, terms such as βhas,β βhave,β and βhavingβ are intended to be open-ended terms. Additionally, the phrase βbased onβ is intended to mean βat least partially based onβ unless otherwise specified.
[0147] The descriptions of various aspects and embodiments are presented for illustrative purposes only and are not intended to be exhaustive or limitful to the disclosed embodiments. While combinations of features are described in the claims and / or disclosed herein, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features may be combined in ways not specifically described in the claims and / or disclosed herein. Each dependent claim listed below may depend directly on only one claim, but the disclosure of possible embodiments includes each dependent claim in combination with all other claims in the claim set. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been selected to best describe the principles of the embodiments, their practical applications or technical improvements to the technology found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0148] The following are some of the acronyms used throughout this disclosure. HEVC High Efficiency Video Coding HDR (High Dynamic Range) SDR Standard Dynamic Range VVC (Variable Video Coding) JVET Joint Video Exploration Team MPM Highest Probability Mode WAIP Wide-Angle Intra Prediction CU Encoding Unit CTB Encoded Tree Block PU Prediction Unit TU Conversion Unit CTU Encoding Tree Unit PDPC location-dependent prediction combination ISP Intra-Subdivision SPS Sequence Parameter Set PPS Picture Parameter Set APS Adaptive Parameter Set VPS Video Parameter Set DPS Decryption Parameter Set ALF Adaptive Loop Filter SAO Pixel Adaptive Offset CC-ALF Cross-Component Adaptive Loop Filter CDEF Constrained Directional Enhancement Filter LR Loop Restoration Filter AV1 AOMedia Video 1 AV2 AOMedia Video 2 SDP (Sections that have been partially decoupled) SEI Supplemental Enhancement Information [Explanation of symbols]
[0149] 10 Cloud Computing Nodes 54A Personal Digital Assistant (PDA) or mobile phone 54B Desktop Computer 54C Laptop Computer 54N Automotive Computer System 60 hardware and software layers 61 Mainframe 62 RISC (Reduced Instruction Set Computer) architecture-based servers 63 servers 64 Blade Servers 65 Storage Devices 66 Networks and Network Components 67 Network Application Server Software 68 Database Software 70. Virtualization Layer 71 Virtual Servers 72 Virtual Storage 73 Virtual Network 74 Virtual Applications and Operating Systems 75 Virtual Clients 80 Management layer 81. Resource Provisioning 82 Measurement and Pricing 83 User Portal 84. Service Level Management 85. Planning and Implementation of Service Level Agreements (SLAs) 90 workload layers 91 Mapping and Navigation 92 Software Development and Lifecycle Management 93 Educational provision 94 Data Analysis Processing 95 Transaction Processing 96 Video Encoding / Decoding 100 Video Encoding Systems 102 Computer 104 Processors 106 Data Storage Devices 108 Software Programs 110 Communication Network 112 Databases 114 Server Computers 116 Video Encoding Programs 116 Programs 116 Video encoding or decoding programs 500 Cloud Computing Environments 600 Functional Abstraction Layer 800A Internal Components 800B Internal Components 820 processor 822 Computer-readable RAM 824 Computer-readable ROM 826 Bus 828 Operating Systems 830 Hard Drives / Storage Devices 832 R / W drive or interface 836 Network adapter or interface 840 Device Drivers 900A External Components 900B External Components 920 Computer Display Monitor 930 Keyboard 934 Computer Mouse 936 Portable Computer Readable Tangible Memory Device
Claims
[Claim 1] A method of video decoding that can be performed by a processor, The steps include receiving video data containing chroma and luma components, The steps include analyzing, deriving, or selecting the number of preset chroma components within a single frame and the number of preset luma components within a single frame. A method comprising the step of decoding the video data, wherein the method includes performing a separated constrained directional enhancement filter (CDEF) process that filters out mutually independent lumar and chroma components based on the number of presets for the chroma component in one frame and the number of presets for the lumar component in one frame.