Encoding and decoding methods for adaptive filtering using multi-criterion classification and corresponding devices
By using adaptive filtering technology based on multi-criteria classification to filter reconstructed samples, the problem of low efficiency in removing coding artifacts in existing technologies is solved, thereby improving the quality and compression efficiency of video encoding and decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2026-04-10
AI Technical Summary
Existing video coding techniques are inefficient at removing coding artifacts, especially block artifacts generated after using intra-frame or inter-frame prediction and transformation, which are difficult to reduce effectively.
Adaptive filtering techniques are employed to classify reconstructed samples. Filtering is performed based on a linear combination of gradient, frequency band index, residual-based classification index, and variance-based classification index. Adaptive filtering is achieved by using a multi-criteria classification selection filter.
It improves the efficiency of removing coding artifacts during video encoding and decoding, and enhances the quality and compression efficiency of reconstructed images.
Smart Images

Figure CN121844563A_ABST
Abstract
Description
Cross Reference to Related Applications
[0001] This application claims the benefit of European application 23306505.1, filed on September 11, 2023, which is incorporated by reference in its entirety. TECHNICAL FIELD
[0002] At least one of the embodiments of the present application relates generally to a method and apparatus for encoding and decoding a picture block using adaptive filtering, e.g., adaptive loop filtering. BACKGROUND
[0003] To achieve high compression efficiency, image and video coding schemes usually employ prediction and transform to exploit the spatial and temporal redundancy in the video content. Typically, inter or intra frame prediction is used to mine the inter or intra picture correlation, and then the difference between the original block and the predicted block, often denoted as prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by the inverse processes corresponding to the entropy coding, quantization, transform, and prediction. The reconstructed video can be filtered to remove some coding artifacts. SUMMARY
[0004] In one implementation, a reconstructed sample of a picture block is filtered. In an example, the filtering comprises: classifying the sample; and filtering the sample based on the classification. In an example, the classification is based on a linear combination of at least two values selected from a gradient-based value, a frequency band index, a residual-based classification index, and a variance-based classification index. BRIEF DESCRIPTION OF DRAWINGS
[0005] Figure 1 A block diagram of a system in which aspects of embodiments of the present application can be implemented is shown; Figure 2 A block diagram of an embodiment of a video encoder is shown; Figure 3 A block diagram of an embodiment of a video decoder is shown; Figure 4 A loop filter applied to reconstructed samples is shown; Figure 5 Diamond-shaped filters for luma (left) and chroma (right) are depicted; Figures 6 to 23 Classifications based on multiple criteria according to various examples are shown; and Figure 24 A flowchart of a decoding method according to an example is depicted. DETAILED DESCRIPTION
[0006] Various aspects are described, including tools, features, embodiments, models, methods, etc. Many of the aspects are described as having specific, and at least to show the various features, often described in a manner that can sound limiting. However, this is for clarity in description, and does not limit the application or scope of the aspects. In fact, all the different aspects can be combined and interchanged to provide additional aspects. Moreover, the aspects can also be combined and interchanged with aspects described in previous documents.
[0007] The aspects described and contemplated in this application can be implemented in many different forms. The following Figure 1 Figure 2 and Figure 3 Some embodiments are provided, but other embodiments are contemplated, and the discussion of Figure 1 Figure 2 and Figure 3 does not limit the breadth of implementation. At least one aspect relates generally to video encoding and decoding, and at least one other aspect relates generally to transmitting generated or encoded bitstreams. These and other aspects can be implemented as methods, devices, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon.
[0008] In this application, the terms “reconstruct” and “decode” are used interchangeably, the terms “encoded” or “coded” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image,” “picture,” and “frame” are used interchangeably. It is generally, but not necessarily, the case that the term “reconstruct” is used on the encoder side, while “decode” is used on the decoder side.
[0009] Various methods are described herein, and each of the methods includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, in various embodiments, terms such as “first,” “second,” etc. can be used to modify elements, components, steps, operations, etc., such as, for example, “first decode” and “second decode.” Unless specifically required, the use of such terms is not intended to imply ordering of the modified operations. Thus, in this example, the first decode need not be performed before the second decode, but can be performed, for example, in a time period that precedes, during, or overlaps with the second decode.
[0010] For clarity, in the embodiments described herein, meeting a condition, failing to meet a condition, and configuring a condition parameter are described as being relative to (e.g., greater than or less than) a threshold, a value (e.g., a threshold), a configured value (e.g., a threshold), and the like. For example, meeting a condition can be described as being above a value (e.g., a threshold), and failing to meet a condition (e.g., a performance criterion) can be described as being below a value (e.g., a threshold). The embodiments described herein are not limited to threshold-based conditions. Other conditions and parameters of any kind (e.g., belonging or not belonging to a range of values) can be applicable to the embodiments described herein.
[0011] The present aspects are not limited to VVC or HEVC, and can be applied to, for example, other standards and recommendations (whether existing earlier or developed in the future), as well as extensions of any such standards and recommendations, including VVC and HEVC. The aspects described in the present application can be used individually or in combination, unless otherwise noted or technically precluded.
[0012] Figure 1 A block diagram showing an example of a system in which various aspects and embodiments can be implemented is shown. System 100 can be implemented as a device including various components described below and configured to perform one or more of the aspects described in the present application. Examples of such a device include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100 can individually or collectively be implemented with a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and the encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in the present application.
[0013] System 100 includes at least one processor 110 configured to execute instructions loaded thereon to implement various aspects, such as those described in this application. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 140 may include internal storage, attached storage, and / or network-accessible storage.
[0014] System 100 includes an encoder / decoder module 130 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device can include one or both encoding and decoding modules. Alternatively, the encoder / decoder module 130 may be implemented as a separate element of system 100, or it may be incorporated into processor 110 as a combination of hardware and software known to those skilled in the art.
[0015] Program code to be loaded onto processor 110 or encoder / decoder 130 to execute the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.
[0016] In some embodiments, the memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 110 or encoder / decoder module 130) is used for one or more of these functions. External memory may be memory 120 and / or storage device 140, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, a fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Various Video Coding, i.e., a new standard developed by JVET (Joint Video Experts Group)).
[0017] As indicated in box 105, inputs can be provided to the components of system 100 via various input devices. Such input devices include, but are not limited to, i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1 Other examples not shown include composite video.
[0018] In various embodiments, the input device of block 105 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF section of various embodiments includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners performing various functions among these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering the signal again to the desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0019] Additionally, the USB and / or HDMI terminals may include their respective interface processors for connecting system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within processor 110 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within processor 110. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements (including, for example, processor 110 and encoder / decoder 130), which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0020] Various components of system 100 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected and transmit data between them using suitable connection means 115 (e.g., internal buses known in the art, including I2C buses, wiring, and printed circuit boards).
[0021] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 190. The communication interface (150) may include, but is not limited to, a modem or network interface card (NIC), and the communication channel (190) may be implemented, for example, in a wired and / or wireless medium.
[0022] In various embodiments, data is streamed to system 100 using a Wi-Fi network, such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received via a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 100, delivering data via an HDMI connection to input block 105. Still other embodiments use an RF connection to input block 105 to provide streaming data to system 100. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0023] System 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. The display 165 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 can be used in a television, tablet computer, laptop computer, mobile phone, or other device. The display 165 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 185 include one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), an optical disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 based on the output of system 100 to provide functionality. For example, an optical disc player performs the function of playing the output of system 100.
[0024] In various embodiments, signaling (such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention) is used to transmit control signals between system 100 and display 165, speaker 175, or other peripheral devices 185. Output devices may be communicatively coupled to system 100 via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices may be connected to system 100 via communication interface 150 using communication channel 190. Display 165 and speaker 175 may be integrated into a single unit along with other components of system 100 in an electronic device (e.g., a television). In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0025] For example, if the RF input section 105 is part of a separate set-top box, the display 165 and speaker 175 can alternatively be separate from one or more of the other components. In various embodiments where the display 165 and speaker 175 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0026] The embodiments can be executed by processor 110 or by computer software implemented by hardware or a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, memory 120 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 110 can be of any type suitable for the technical environment and can encompass one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0027] Figure 2 An exemplary video encoder 200, such as a VVC (Various Video Coding) encoder, is shown. Figure 2 It can also show encoders that improve upon the VVC standard, or encoders that employ VVC-like technology.
[0028] Before being encoded, the video sequence may undergo pre-coding (201) (e.g., applying color transformations to the input color picture (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components) to obtain a signal distribution more suited to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.
[0029] In encoder 200, the frame is encoded by encoder elements as described below. The frame to be encoded is partitioned (202) and processed in units such as CTU (Coding Tree Unit) and CU (Coding Unit). Each unit is encoded using, for example, an intra-frame or inter-frame mode. When a unit is encoded in intra-frame mode, intra-frame prediction is performed (260) using, for example, an intra-frame prediction tool such as decoder-side intra-frame mode derivation (DIMD) is used. In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which mode, intra-frame or inter-frame, to use to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the predicted block (also called the prediction block) from the original image block.
[0030] Then the prediction residual is transformed (225) into transformation coefficients c (also known as prediction residual transformation coefficients), and the transformation coefficients c are quantized (230) into quantization indexes. (Also referred to as transform coefficient level or quantized transform coefficient on the encoder side). For quantization level (also called quantization index). The encoder entropy-encoded (245) motion vectors and other syntax elements (such as screen partition information) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, that is, directly encode the residual without applying the transform or quantization process.
[0031] The encoder decodes the coded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) (also known as scaling) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (255) to reconstruct the image blocks. Block-based intra / inter-frame prediction and transform coding, as well as quantization, can produce various artifacts at medium and low bit rates. To reduce these artifacts, in-loop filters (265) such as deblocking filter (DBF), bilateral filter (BIF), sample adaptive offset (SAO), and adaptive loop filter (ALF) can be applied to the reconstructed image to perform filtering to reduce coding artifacts. The filtered image is stored in a reference frame buffer (280). Therefore, an in-loop filter (265) is used to enhance the reconstructed image before storing it in the reference frame buffer (280). In-loop filters form a complete series. Among them, the deblocking filter (DBF) aims to reduce block artifacts that occur along block boundaries, i.e., block discontinuities. Deblocking filters are typically designed to improve subjective quality, i.e., how well the human visual system perceives such coding errors. In common video coding standards such as HEVC and VVC, deblocking filters are pre-determined based on coded information (such as prediction patterns, motion vectors, and transform coefficients) and local variations at block boundaries. Sample Adaptive Offset (SAO) primarily aims to reduce artifacts caused by the quantization of transform coefficients. Adaptive Loop Filters (ALF) and Cross-Component Adaptive Loop Filters (CCALF) are adaptive filters that can enhance the reconstructed signal using coding methods such as Wiener filters. ALF and its variant CCALF can determine the optimal luma and chroma filters on the encoder side according to rate-distortion criteria. These filters can be sent with the bitstream and then decoded and used on the decoder side. Luma ALF can achieve further local adaptation based on luma classification, while chroma ALF and CCALF use region adaptation, where filter selection is performed at the CTU level. Typically, adaptive loop filters are applied at the CTU level, while deblocking filters are applied along block boundaries.
[0032] Figure 3 A block diagram of an exemplary video decoder 300 is shown. In decoder 300, bitstreams are decoded by decoder elements, as described below. Video decoder 300 typically performs operations similar to... Figure 2 The encoding process described herein is the inverse of the decoding process. Encoder 200 typically also performs video decoding as part of the encoded video data.
[0033] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy-decoded (330) to obtain the quantization level. (Also known as transform coefficient level or quantization level on the decoder side), prediction mode, motion vectors, and other encoded information. Picture partitioning information indicates how the picture is partitioned. Therefore, the decoder can divide the picture (335) based on the decoded picture partitioning information. The quantization level... Dequantization (340) is used to reconstruct the transform coefficients. Dequantization, also known as scaling, is applied to the reconstructed transform coefficients. An inverse transform (350) is performed to obtain the prediction residual. The prediction residual and the predicted block (also called the prediction block) are combined (355) to reconstruct the image block. The prediction block (370) can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference frame buffer (380). It should be noted that for a given frame, the contents of the reference frame buffer 380 on the decoder 300 side are the same as the contents of the reference image buffer 280 on the encoder 200 side (for the same frame).
[0034] The decoded image can also undergo post-decoding processing (385), such as inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping, which is the inverse of the remapping process performed in the pre-encoding process (201). Post-decoding processing can use metadata derived in the pre-encoding process and signaled in the bitstream.
[0035] Both the video encoder and decoder 200 and the video decoder 300 may include in-loop filters (265 and 365, respectively). Figure 4 The document describes an example of the workflow of an in-loop filter, such as in VVC. For instance, if / when the local deblocking condition is met, a deblocking filter (DBF) can first be used to filter the luma reconstruction samples and chroma reconstruction samples located along the block boundaries. Then, a sample adaptive offset (SAO) filter is used to locally add an offset based on the classification using a band-based or edge-based classifier. Finally, an adaptive loop filter (ALF) and / or a cross-component adaptive loop filter (CCALF) can be applied before storing the resulting sample values in a reference frame buffer (e.g., buffer 280 on the encoder side or buffer 380 on the decoder side). Only a subset of these filters can be applied.
[0036] In addition to DBF, SAO, ALF, and CCALF mentioned above, in-loop filters can also include bilateral filtering (BIF) and cross-component sample adaptive offset (CCSAO), for example, as in ECM (an abbreviation for augmented compression model).
[0037] Adaptive loop filter (ALF) ALF (Adaptive Filtering) is an adaptive filter that can be applied to reduce the MSE (Mean Squared Error) between the original and reconstructed samples using coding methods such as Wiener filters. The ALF filter (e.g., ALF filter coefficients and possible pruning indices) can be sent in the bitstream and decoded on the decoder side, and can be derived from parameters (e.g., previously) signaled or from (e.g., predefined) parameters before being applied to the reconstructed samples. The ALF filter can be point-symmetric, for example, having integer coefficients. Figure 5 As described, ALF can use a 7×7 diamond filter for brightness ( Figure 5 (on the left) and use a 5×5 diamond filter for chroma ( Figure 5 (The right side). An ALF filter can be DC neutral, that is, the center coefficient ( Figure 5 The shadow pixels in the image can be equal to one minus the sum of all other coefficients.
[0038] The filtering operation is further explained below. Let's assume... To reconstruct the samples and set This represents its ALF filter value. In the linear implementation of ALF, Calculated as:
[0039] in Let N represent the filter coefficients, and (N-1) be the number of filter coefficients, where:
[0040] in This indicates that it corresponds to the first coefficient The coordinate offset of the associated reconstructed sample.
[0041] In the nonlinear implementation of ALF, equation (1) becomes:
[0042] in:
[0043] It is with coefficient The associated clipping parameters are determined by the clipping index. To determine the trimming parameters. It can be exported as follows:
[0044] in This represents the sample bit depth, and It can be 0, 1, 2 or 3.
[0045] The ALF filter coefficients and pruning indices (if applicable) can be determined on the encoder side, for example, by minimizing the MSE (mean squared error) between the reconstructed sample and its original value by solving the Wiener-Hopf equations. For example, these coefficients and corresponding pruning indices (if applicable) can be encoded in an adaptive parameter set (APS) if / when the rate-distortion condition is satisfied. The ALF APS can contain a luma filter set and up to eight chroma filters, for example, as in VVC. Pre-trained (e.g., predefined) luma filters can also be used, which are stored (e.g., hard-coded) on both the encoder and decoder sides and therefore not transmitted in the bitstream.
[0046] The ALF for luminance can include further local adaptation utilizing classification based on local gradients, such as those computed over 4×4 blocks. Depending on the classification, a geometric filter transformation, such as a 90-degree rotation, diagonal, or vertical flip, can be applied to the filter prior to filtering. In the example, for the luminance component, each block (e.g., each 4×4 block) can be classified into one of K classes, where K is an integer (K=25). The classification index C can be derived based on the directionality value D computed for the block and the quantization value of activity  (e.g., Laplacian activity value), for example, An ALF (e.g., an ALF for luminance) can depend on a filter set. A filter set includes one or more filters and a list (e.g., a mapping list or lookup table) that associates a particular filter in the filter set with each class of the classification. On both the encoder and decoder sides, samples can be filtered by the filters associated with their respective classes.
[0047] Directional and activity values can be viewed as gradient-based features and can be calculated as disclosed below for 4×4 blocks. However, this principle can be applied to blocks of other sizes. First, the sample gradient values in the horizontal, vertical, and two diagonal directions can be calculated as follows:
[0048] Sub-block horizontal gradient based on sample gradient Vertical gradient and two diagonal gradients and Calculated as:
[0049] Where index and This represents the coordinates of the top-left sample in a 4×4 luminance sub-block.
[0050] Secondly, in order to determine directionality The ratio of the maximum to the minimum value of the block's horizontal and vertical gradients can be used. And the ratio of the maximum and minimum values of the diagonal gradients of the two blocks. Compare with each other and with a set of thresholds and Comparison, among which
[0051] and
[0052] For this purpose, in the first step (step 1), if / when and hour, Set as (The block is categorized as "texture").
[0053] In the second step (step 2), if / when In step 3, the directionality is calculated. Otherwise, calculate the directionality in step 4.
[0054] In the third step (step 3), if / when hour, Set as (The block is classified as "strong horizontal / vertical"); otherwise, Set as (The block is classified as "weak horizontal / vertical").
[0055] In the fourth step (step 4), if / when hour, Set as (The block is classified as "strong diagonal"); otherwise, Set as (The block is classified as "weak diagonal").
[0056] Activity value Calculated as:
[0057] Activity value It can be further mapped to the range of 0 to 4 (inclusive), and the quantized value is represented as .
[0058] Finally, each (e.g., 4×4) luminance block can be classified into one of 25 classes, each class having a specific filter assigned to the following:
[0059] Multiple pre-trained luma filter sets can also be available on the encoder and decoder sides. The encoder can choose whether to send a luma filter set optimized for the current slice / frame based on a rate-distortion criterion. At the CTU level, the encoder can decide whether to perform ALF and can choose the luma filter set to use between the pre-trained filter set and the already sent filter set. Unlike luma, ALF for chroma typically does not employ local classification. Instead, it is characterized by region adaptation. In the example, up to eight ALF chroma filters can be used simultaneously on the encoder and decoder sides, and each chroma block (e.g., CTB (Code Tree Block)) signals which filter it uses.
[0060] In other examples, such as in ECM, the chroma ALF rhombus size can be increased to 9×9. To filter the luminance samples, three different classifiers can be used ( , and And three different filter sets (F0, F1, and F2). Sets F0 and F1 can contain fixed filters, for example, with filters specific to the classifier. and The pre-trained coefficients. The coefficients of the filters in F2 can be signaled. Which filter from set Fi is used for a given sample can be determined by using a classifier. The class Ci assigned to the sample is determined by this. Two classifiers correspond to the fixed filter. and Both can be based on the Laplacian operator and applied to 2×2 sub-blocks. In both classifiers, activity and orientation values can be derived based on the vertical, horizontal, and diagonal gradients. Then, the class index can be determined based on the activity and orientation values.
[0061] Luminance classification can utilize additional alternative classifiers, such as band-based classifiers. For a signaled luminance filter set, a signal flag can be sent to indicate whether an alternative classifier is applied. In this example, no geometric transformation is applied to the alternative band-based classifier. When applying a band-based classifier, the sum of sample values for (e.g., 2×2) luminance blocks can be computed first. Then, the class index can be calculated as follows:
[0062] In addition to Laplacian classifiers and band-based classifiers, new classifiers suitable for residual samples can be used. The sum of the absolute values of residual samples in adjacent windows can be calculated, and the class index can be derived as follows:
[0063] The value of classIdx is in the range of 0 to 24, the same as the first two classifiers.
[0064] In another example, a classifier with a fixed filter is expanded. First, for at least one (e.g., each) 2×2 block, the mean of the surrounding window is computed. Then, for one (e.g., each) sample within that window, the difference between the sample value and the mean is computed. The scaling factor can be determined based on the activity values derived from the Laplacian classifier. The square root of the sum of squared differences (e.g., the variance) can be further quantized using the scaling factor. . The value is an integer between 0 and 7 (inclusive). In the case of, let This indicates the first one from ECM-8.0. A classifier with a fixed filter. Then, the proposed class index can be derived as follows. :
[0065] In-loop filtering can be performed using variance-based classification, namely, bilateral filtering (BIF) and adaptive loop filtering (ALF). For BIF, the variance of the transform unit (TU) can be used to identify texture intensity. More filtering intensity is introduced in BIF. The filtering intensity is determined by the variance.
[0066] For ALF, each classification unit can be jointly classified into two texture intensity levels based on variance and boundary location. These texture intensity levels are then combined with the existing classifier to output the final classification result.
[0067] Adaptive filters (such as SAO and ALF) are determined on the encoder side by optimizing rate-distortion criteria. These filters are sent with the bitstream and then decoded and used on the decoder side. Much of the gain in brightness provided by ALF is due to further local adaptation implemented for sample classification, i.e., filtering samples sharing the same local features using the same specific filter. This allows for the design of efficient, custom filters, such as non-separable 2D directional filters that remove artifacts in a given direction if / when classification is gradient-based. Other types of classification have been proposed: for example, based on color or component frequency bands, based on residual energy, or based on the variance of the reconstructed samples. Thus, sample classification depends on one or more criteria / functions. Filter parameters (e.g., coefficients and possible clipping indices) are then associated with each class ID.
[0068] In contrast, a method for decoding is disclosed below that enables the classification of samples in a general manner, making filter optimization as efficient as possible. For this purpose, multi-criteria classification, including novel combinations of classification criteria / functions, can be used for adaptive filtering (e.g., for signaling). Several criteria can be used, such as gradient-based or Sobel-based features, e.g., directionality and activity, single-component or multi-component frequency bands, local representations of residual samples, and locally computed variance of reconstructed samples.
[0069] In the following text, It is considered to be an index representing a local directional value, and It is a quantified activity value. Also set... Indicates frequency band index. Represents the index corresponding to the residual-based classification (e.g., subclassification), and sets... This represents an index corresponding to a variance-based classification (e.g., a subclassification). As previously mentioned, various criteria / functions can be used for classification, such as directionality, activity, frequency band, residuals, and variance of the reconstructed samples.
[0070] Directionality can be computed indiscriminately based on local gradients as previously disclosed or on any gradient approximation such as the Sobel operator. Directionality is determined based on the computation of local signal variations in one or more components. In one example, three local directionsality are defined, such as horizontal / vertical, diagonal, and a “texture” that is neither diagonal nor horizontal nor vertical.
[0071] Table 1 provides an example of the mapping of D values that depend on the directionality of local computation.
[0072] Table 1: Three-directional classification.
[0073]
[0074] In another example, five local orientations are defined, such as strong horizontal / vertical, weak horizontal / vertical, strong diagonal, weak diagonal, and texture.
[0075] Table 2 provides an example of the mapping of D values that depend on the directionality of local computation.
[0076] Table 2: Five-directional classification.
[0077]
[0078] Similar to directionality, activity can be computed indiscriminately based on local gradients or on any gradient approximation (such as the Sobel operator). Activity is determined based on the computation of local signal variations in one or more components.
[0079] In the following text, band-based classification (e.g., subclassification) depends on the definition of one or more component bands (i.e., luminance band, Cb band, Cr band, chromaticity Cb+Cr band, and luminance+Cb / Cr band, or YCbCr band).
[0080] Residual classification (e.g., subclassification) relies on a measure of local activity. As previously described, it can be derived from the L1 norm of the residuals near the current sample. In another variation, it can be derived from the local residual energy (i.e., the sum of squared residuals).
[0081] Classification based on reconstructed variance can be derived from a quantized version of the local variance of the reconstructed sample values.
[0082] Based on this principle, the classification of adaptive filtering is based on various combinations of the classification criteria disclosed above (such as directionality, activity, frequency band, residuals, and variance of the reconstructed samples). In the following text, "reconstructed sample" may indiscriminately refer to the sum of predictions and residuals, or the sum of predictions and residuals filtered by one or more in-loop or post-loop filters (such as DBF, SAO, CCSAO, and / or BIF). Furthermore, the "gradient" used to determine directionality and activity may refer to the exact calculation of the horizontal, vertical, and diagonal gradients, or indiscriminately refer to their approximations (such as the Sobel operator).
[0083] Figure 24 A flowchart illustrating the decoding method based on the example is provided.
[0084] At S100, samples of the image patch are reconstructed. Reconstruction may include dequantization (or inverse quantization), inverse transform, and the addition of a predictor. Optionally, a first filter (e.g., DBF, SAO, etc.) may be applied to the reconstructed samples.
[0085] At S102, at least one (e.g., each) reconstructed sample (possibly filtered by the first filter) is classified. Figures 6 to 23 In any of the disclosed examples, the classification is based on a linear combination of at least two values (e.g., at least two values associated with or computed for at least one reconstructed sample) selected from gradient-based values, band indexes, residual-based classification indexes, and variance-based classification indexes. The classification associates a class or index (e.g., a class index) denoted as C with each sample.
[0086] At S104, the at least one (e.g., each) reconstructed sample is filtered based on the classification (e.g., based on index C). Filtering the sample based on the classification includes applying a filter to the sample, the filter being associated with the class to which the sample to be filtered belongs.
[0087] In the example, the filtered sample value is equal to the sum of the offset and the unfiltered sample value, where the offset is equal to the weighted sum of the sample values of the component to be filtered.
[0088] Figure 24 A flowchart illustrating the decoding method according to the example is provided. However, since the encoding method includes a decoding loop, the same steps S100 to S104 apply to the encoding method.
[0089] In the example, the classification jointly depends on gradient-based features (such as directionality and activity) and frequency bands, as in Figure 6 As described in the text. Several examples of this classification based on gradient-based features (such as directionality and activity) and frequency bands are presented below.
[0090] In the example, three different directivities are defined, as described in Table 1. The activity is mapped to a range of 0 to 2 (inclusive), and three frequency bands are defined. In the example, the three frequency bands are defined by two distinct boundaries b0 and b1, where, for example, b1 > b0. The first frequency band consists of the range of values up to the first boundary b0, the second frequency band consists of the range of values between b0 and b1, and the third frequency band consists of the range of values greater than b1.
[0091] The 27-level classification can be achieved using a class index calculated in the following way:
[0092] In the variant, five different directivities are defined, as described in Table 2. The activity is mapped to a range of 0 to 2 (inclusive of end values), and two frequency bands are defined.
[0093] The 30-level classification can be achieved using a class index calculated in the following way:
[0094] In the variant, three frequency bands are defined, and only two activity levels are considered, i.e., mapped to a range of 0 to 1 (inclusive). The resulting 30-level classification can be achieved using a class index calculated as follows:
[0095] In the variant, five different directivities are defined, as described in Table 2, with activity mapped to a range of 0 to 2 (inclusive), and three frequency bands are defined. The resulting 45-level classification can be implemented using the same class index calculation as in Equation (16).
[0096] In another example, the classification jointly depends on directionality and frequency band, as in Figure 7 As depicted in the text.
[0097] In the example, three different directions of orientation are defined, as described in Table 1, and eight frequency bands are defined. The resulting 24-level classification can be achieved using a class index calculated as follows:
[0098] In the variant, five different directions are defined, as described in Table 2, and eight frequency bands are defined. The resulting 40-level classification can be achieved using a class index calculated as follows:
[0099] In the example, the classification jointly depends on gradient-based features (such as directionality and activity) and residual classification (e.g., sub-classification), as in Figure 8 As described in the text. Several examples of this classification based on gradient-based features (such as directionality and activity) and residuals are presented below.
[0100] In the example, three different orientations are defined, as described in Table 1. Activity is mapped to a range of 0 to 2 (inclusive), and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 2 (inclusive). The resulting 27-level classification can be achieved using class indices calculated as follows:
[0101] In the variant, five different orientations are defined, as described in Table 2. Activity is mapped to a range of 0 to 2 (inclusive), and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive). The resulting 30-level classification can be achieved using a class index calculated as follows:
[0102] In another variation, the residual classification (e.g., subclassing) index is mapped to a range of 0 to 2 (inclusive), and only the two activity levels are considered, i.e., mapped to a range of 0 to 1 (inclusive). The resulting 30-level classification can be achieved using a class index calculated as follows:
[0103] In the variant, five different orientations are defined, as described in Table 2. Activity is mapped to a range of 0 to 2 (inclusive of endpoints), and residual classification (e.g., subclassing) indices are mapped to a range of 0 to 2 (inclusive of endpoints). The resulting 45-level classification can be achieved using the same class indexes computed as in Equation (20).
[0104] In another example, the classification jointly depends on the orientation and residual classification, as in Figure 9 As described in the text. Several examples are presented below.
[0105] In the example, three different orientations are defined as described in Table 1, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 7 (inclusive). The resulting 24-level classification can be achieved using a class index calculated as follows:
[0106] In the variant, five different orientations are defined, as described in Table 2, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 7 (inclusive of endpoints). The resulting 40-level classification can be achieved using a class index calculated as follows:
[0107] In the example, classification depends on both gradient-based features (such as directionality and activity) and the local variance of the reconstructed samples, as in Figure 10 As described in the text. Several examples are presented below, where classification is based on gradient-based features (such as directionality and activity) and the local variance of the reconstructed samples.
[0108] In the example, three different orientations are defined, as described in Table 1. Activity is mapped to a range of 0 to 2 (inclusive), and quantized variance is mapped to a range of 0 to 2 (inclusive). The resulting 27-level classification can be achieved using a class index calculated as follows:
[0109] In the variant, five different orientations are defined, as described in Table 2. Activity is mapped to a range of 0 to 2 (inclusive), and quantified variance is mapped to a range of 0 to 1 (inclusive). The resulting 30-level classification can be achieved using a class index calculated as follows:
[0110] In another variation, the quantized variance is mapped to a range of 0 to 2 (inclusive), and only two activity levels are considered; that is, activity is mapped to a range of 0 to 1 (inclusive). The resulting 30-level classification can be achieved using a class index calculated as follows:
[0111] In the variant, five different orientations are defined, as described in Table 2. Activity is mapped to a range of 0 to 2 (inclusive), and quantified variance is mapped to a range of 0 to 2 (inclusive). The resulting 45-level classification can be achieved using the same class index calculation as in Equation (25).
[0112] In the example, classification depends on both directionality and the local variance of the reconstructed samples, as in... Figure 11 As described in the text. Several examples are presented below.
[0113] In the example, three different orientations are defined, as described in Table 1, and the quantized variance is mapped to a range of 0 to 1 (inclusive). The resulting 6-level classification can be achieved using a class index calculated as follows:
[0114] Alternatively:
[0115] In the variant, three distinct directions are defined as described in Table 1, and the quantified variance is mapped to a range of 0 to 2 (inclusive). The resulting 9-level classification can be achieved using the same class index calculation as in Equation (28), or alternatively:
[0116] In the variant, five different directions are defined, as described in Table 2, and the quantized variance is mapped to a range of 0 to 1 (inclusive). The resulting 10-level classification can be achieved using a class index calculated as follows:
[0117] Alternatively, the same class index as in equation (29) can be used for computation.
[0118] In the variant, three distinct orientations are defined as described in Table 1, and the quantified variance is mapped to a range of 0 to 3 (inclusive). The resulting 12-level classification can be achieved using the same class index calculation as in Equation (28), or alternatively:
[0119] In the variant, three distinct orientations are defined as described in Table 1, and the quantified variance is mapped to a range of 0 to 4 (inclusive). The resulting 15-level classification can be achieved using the same class index calculation as in Equation (28), or alternatively:
[0120] In another variant, five different orientations are defined as described in Table 2, and the quantified variance is mapped to a range of 0 to 2 (inclusive of the extreme values). The resulting 15-level classification can be achieved using the same class index calculation as in Equation (31), or alternatively, using the same class index calculation as in Equation (30).
[0121] In the variant, five different orientations are defined as described in Table 2, and the quantified variance is mapped to a range of 0 to 3 (inclusive). The resulting 20-level classification can be achieved using the same class index calculation as in Equation (31), or alternatively, using the same class index calculation as in Equation (32).
[0122] In the variant, three different directions are defined, as described in Table 1, and the quantized variance is mapped to a range of 0 to 7 (inclusive). The resulting 24-level classification can be achieved using a class index calculated as follows:
[0123] In the variant, five different orientations are defined as described in Table 2, and the quantified variance is mapped to a range of 0 to 4 (inclusive). The resulting 25-level classification can be achieved using the same class index calculation as in Equation (31), or alternatively, using the same class index calculation as in Equation (33).
[0124] In the variant, five different directions are defined, as described in Table 2, and the quantized variance is mapped to a range of 0 to 7 (inclusive). The resulting 40-level classification can be achieved using a class index calculated as follows:
[0125] In another example, the classification depends on both the frequency band and the residual classification, as in Figure 12 As described in the text. Several examples are presented below, where classification is based on frequency bands and residuals.
[0126] In the example, eight frequency bands are defined, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 2 (inclusive). The resulting 24-level classification can be achieved using class indices calculated as follows:
[0127] In the variant, five frequency bands are defined, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 4 (inclusive). The resulting 25-level classification can be achieved using a class index calculated as follows:
[0128] In the variant, eight frequency bands are defined, and the residual classification (e.g., subclass) index is mapped to a range of 0 to 3 (inclusive of end values). The resulting 32-level classification can be achieved using the same class indexes computed as in equation (36).
[0129] In the variant, eight frequency bands are defined, and the residual classification (e.g., subclass) index is mapped to a range of 0 to 4 (inclusive of end values). The resulting 40-level classification can be achieved using the same class indexes computed as in Equation (36).
[0130] In the example, the classification depends on both the definition based on the frequency band and the local variance of the reconstructed samples, as shown in... Figure 13 As described in the text. Several examples are presented below.
[0131] In the example, eight frequency bands are defined, and the quantized variance is mapped to a range of 0 to 2 (inclusive). The resulting 24-level classification can be achieved using a class index calculated as follows:
[0132] In the variant, five frequency bands are defined, and the quantized variance is mapped to a range of 0 to 4 (inclusive). The resulting 25-level classification can be achieved using a class index calculated as follows:
[0133] In the variant, eight frequency bands are defined, and the quantized variance is mapped to a range of 0 to 3 (inclusive). The resulting 32-level classification can be achieved using the same class index calculation as in equation (38).
[0134] In the variant, eight frequency bands are defined, and the quantized variance is mapped to a range of 0 to 4 (inclusive). The resulting 40-level classification can be achieved using the same class index calculation as in equation (38).
[0135] In the example, the classification depends on the residual classification and the local variance of the reconstructed samples, as shown in... Figure 14 As described in the text. Several examples are presented below, where classification is based on residual classification and the local variance of the reconstructed samples.
[0136] In the example, the residual classification (e.g., subclass) index is mapped to a range of 0 to 3 (inclusive), and the quantized variance is mapped to a range of 0 to 3 (inclusive). The resulting 16-level classification can be achieved using class indices calculated as follows:
[0137] In this variant, the residual classification (e.g., subclassing) index is mapped to a range of 0 to 4 (inclusive), and the quantized variance is also mapped to a range of 0 to 4 (inclusive). The resulting 25-level classification can be achieved using a class index calculated as follows:
[0138] In this variant, the residual classification (e.g., subclassing) index is mapped to a range of 0 to 5 (inclusive), and the quantized variance is also mapped to a range of 0 to 5 (inclusive). The resulting 36-level classification can be achieved using a class index calculated as follows:
[0139] In the example, the classification jointly depends on gradient-based features (such as directionality and activity), frequency bands, and residual classification, as in... Figure 15 As described in the text. Several examples are presented below.
[0140] In the example, three different orientations are defined, as described in Table 1. Activity is mapped to a range of 0 to 1 (inclusive of end values), two frequency bands are defined, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive of end values). The resulting 24-level classification can be achieved using class indices calculated as follows:
[0141] In the variant, three different directions are defined, as described in Table 1. Activity is mapped to a range of 0 to 2 (inclusive), two frequency bands are defined, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive). The resulting 36-level classification can be achieved using the class index calculated as follows:
[0142] In another variation, three distinct directions are defined, as described in Table 1. Activity is mapped to a range of 0 to 1 (inclusive of endpoints), three frequency bands are defined, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive of endpoints). The resulting 36-level classification can be achieved using a class index calculated as follows:
[0143] In another variant, three different orientations are defined, as described in Table 1: activity is mapped to a range of 0 to 1 (inclusive of end values), two frequency bands are defined, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 2 (inclusive of end values). The resulting 36-level classification can be achieved using the same class index computation as in Equation (43).
[0144] In the variant, three distinct directions are defined, as described in Table 1. Activity is mapped to a range of 0 to 2 (inclusive of endpoints), three frequency bands are defined, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive of endpoints). The resulting 54-level classification can be achieved using a class index calculated as follows:
[0145] In another variant, three different orientations are defined, as described in Table 1, with activity mapped to a range of 0 to 2 (inclusive of end values), two frequency bands defined, and residual classification (e.g., subclassification) indices mapped to a range of 0 to 2 (inclusive of end values). The resulting 54-level classification can be achieved using the same class indexes computed as in Equation (44).
[0146] In the example, the classification jointly depends on directionality, frequency band, and residual classification, as in... Figure 16 As described in the text. Several examples are presented below.
[0147] In the example, five different orientations are defined, as described in Table 2, four frequency bands are defined, and the residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive of end values). The resulting 40-level classification can be achieved using a class index calculated as follows:
[0148] In the example, the classification jointly depends on the band-based definition, residual classification, and the local variance of the reconstructed samples, as in... Figure 17 As described in the text. Several examples are presented below.
[0149] In the example, four frequency bands are defined. The residual classification (e.g., subclass) index is mapped to a range of 0 to 1 (inclusive of extreme values), and the quantized variance is mapped to a range of 0 to 1 (inclusive of extreme values). The resulting 16-level classification can be achieved using the class index calculated as follows:
[0150] In the variant, four frequency bands are defined, the residual classification (e.g., subclass) index is mapped to a range of 0 to 2 (inclusive of extreme values), and the quantized variance is mapped to a range of 0 to 1 (inclusive of extreme values). The resulting 24-level classification can be achieved using class indices calculated as follows:
[0151] In another variation, four frequency bands are defined, with the residual classification (e.g., subclass) index mapped to a range of 0 to 1 (inclusive of end values) and the quantized variance mapped to a range of 0 to 2 (inclusive of end values). The resulting 24-level classification can be achieved using the same class index computation as in Equation (48).
[0152] In the variant, four frequency bands are defined, the residual classification (e.g., subclass) index is mapped to a range of 0 to 2 (inclusive of end values), and the quantized variance is mapped to a range of 0 to 2 (inclusive of end values). The resulting 36-level classification can be implemented using the same class index calculation as in Equation (49).
[0153] In the example, the classification jointly depends on both gradient-based features (such as directionality and activity), residual classification, and the local variance of the reconstructed samples, as in... Figure 18 As described in the text. Several examples are presented below.
[0154] In the example, three different orientations are defined, as described in Table 1: activity is mapped to a range of 0 to 1 (inclusive of endpoints), residual classification (e.g., subclassing) indexes are mapped to a range of 0 to 1 (inclusive of endpoints), and quantized variance is mapped to a range of 0 to 1 (inclusive of endpoints). The resulting 24-level classification can be achieved using class indices calculated as follows:
[0155] In this variant, three different orientations are defined, as described in Table 1: activity is mapped to a range of 0 to 2 (inclusive of endpoints), residual classification (e.g., subclassing) indexes are mapped to a range of 0 to 1 (inclusive of endpoints), and quantized variance is mapped to a range of 0 to 1 (inclusive of endpoints). The resulting 36-level classification can be achieved using class indices calculated as follows:
[0156] In another variant, three different orientations are defined, as described in Table 1: activity is mapped to a range of 0 to 1 (inclusive of end values), residual classification (e.g., subclassing) indexes are mapped to a range of 0 to 1 (inclusive of end values), and quantized variance is mapped to a range of 0 to 2 (inclusive of end values). The resulting 36-level classification can be achieved using the same class indexes computed as in Equation (50).
[0157] In another variation, three different orientations are defined, as described in Table 1: activity is mapped to a range of 0 to 1 (inclusive of endpoints), residual classification (e.g., subclassing) indexes are mapped to a range of 0 to 2 (inclusive of endpoints), and quantized variance is mapped to a range of 0 to 1 (inclusive of endpoints). The resulting 36-level classification can be achieved using class indices calculated as follows:
[0158] In this variant, three different orientations are defined, as described in Table 1: activity is mapped to a range of 0 to 2 (inclusive), the residual classification (e.g., subclassing) index is mapped to a range of 0 to 2 (inclusive), and the quantized variance is mapped to a range of 0 to 1 (inclusive). The resulting 54-level classification can be achieved using class indices calculated as follows:
[0159] In another variant, three different orientations are defined, as described in Table 1: activity is mapped to a range of 0 to 2 (inclusive of endpoints), residual classification (e.g., subclassing) indexes are mapped to a range of 0 to 1 (inclusive of endpoints), and quantized variance is mapped to a range of 0 to 2 (inclusive of endpoints). The resulting 54-level classification can be achieved using the same class indexes computed as in Equation (51).
[0160] In the example, the classification jointly depends on the orientation, the residual classification, and the local variance of the reconstructed samples, as in... Figure 19 As described in the text. Several examples are presented below.
[0161] In this variant, five different orientations are defined, as described in Table 2. The residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive of extreme values), and the quantized variance is mapped to a range of 0 to 1 (inclusive of extreme values). The resulting 20-level classification can be achieved using the class index calculated as follows:
[0162] In this variant, five different orientations are defined, as described in Table 2. The residual classification (e.g., subclassing) index is mapped to a range of 0 to 2 (inclusive of extreme values), and the quantized variance is mapped to a range of 0 to 1 (inclusive of extreme values). The resulting 30-level classification can be achieved using the class index calculated as follows:
[0163] In another variant, five different orientations are defined, as described in Table 2, with residual classification (e.g., subclassing) indices mapped to a range of 0 to 1 (inclusive of end values) and quantized variance mapped to a range of 0 to 2 (inclusive of end values). The resulting 30-level classification can be achieved using the same class indexes computed as in Equation (54).
[0164] In the variant, five different orientations are defined as described in Table 2, with residual classification (e.g., subclassing) indices mapped to a range of 0 to 2 (inclusive of end values) and quantized variance mapped to a range of 0 to 2 (inclusive of end values). The resulting 45-level classification can be achieved using the same class indexes computed as in Equation (55).
[0165] In the example, the classification jointly depends on both gradient-based features (such as directionality and activity), frequency band, and the local variance of the reconstructed samples, as in Figure 20 As described in the text. Several examples are presented below.
[0166] In the example, three different directivities are defined, as described in Table 1. Activity is mapped to a range of 0 to 1 (inclusive), two frequency bands are defined, and the quantized variance is mapped to a range of 0 to 1 (inclusive). The resulting 24-level classification can be achieved using a class index calculated as follows:
[0167] In the variant, three different directivities are defined, as described in Table 1. Activity is mapped to a range of 0 to 2 (inclusive), two frequency bands are defined, and the quantized variance is mapped to a range of 0 to 1 (inclusive). The resulting 36-level classification can be achieved using a class index calculated as follows:
[0168] In another variant, three distinct directivities are defined, as described in Table 1. Activity is mapped to a range of 0 to 1 (inclusive), three frequency bands are defined, and the quantized variance is mapped to a range of 0 to 1 (inclusive). The resulting 36-level classification can be achieved using a class index calculated as follows:
[0169] In another variant, three different orientations are defined, as described in Table 1, with activity mapped to a range of 0 to 1 (inclusive of end values), two frequency bands defined, and residual classification (e.g., subclassification) indices mapped to a range of 0 to 2 (inclusive of end values). The resulting 36-level classification can be achieved using the same class indexes computed as in Equation (56).
[0170] In the variant, three distinct directivities are defined, as described in Table 1. Activity is mapped to a range of 0 to 2 (inclusive), three frequency bands are defined, and the quantized variance is mapped to a range of 0 to 1 (inclusive). The resulting 54-level classification can be achieved using a class index calculated as follows:
[0171] In another variant, three different directivities are defined, as described in Table 1, with activity mapped to a range of 0 to 2 (inclusive of end values), two frequency bands defined, and quantized variance mapped to a range of 0 to 2 (inclusive of end values). The resulting 54-level classification can be achieved using the same class index computation as in Equation (57).
[0172] In the embodiments, the classification jointly depends on directionality, frequency band, and the local variance of the reconstructed samples, as in Figure 21 As described in the text. Several examples are presented below.
[0173] In the variant, five different directivities are defined, as described in Table 2, four frequency bands are defined, and the quantized variance is mapped to a range of 0 to 1 (inclusive of the end values).
[0174] The 40-level classification can be achieved using a class index calculated in the following way:
[0175] In the example, the classification jointly depends on gradient-based features (such as directionality and activity), frequency bands, residual classification, and the local variance of the reconstructed samples, as in Figure 22 As described in the text. Several examples are presented below.
[0176] In the example, three different orientations are defined, as described in Table 1: activity is mapped to a range of 0 to 1 (inclusive of endpoints); two frequency bands are defined: the residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive of endpoints); and the quantized variance is mapped to a range of 0 to 1 (inclusive of endpoints). The resulting 48-level classification can be achieved using the class index calculated as follows:
[0177] In the example, the classification jointly depends on directionality, frequency band, residual classification, and the local variance of the reconstructed samples, as in... Figure 23 As described in the text. Several examples are presented below.
[0178] In the variant, three different directivities are defined, as described in Table 1, and two frequency bands are defined: the residual classification (e.g., subclassification) index is mapped to the range of 0 to 2 (inclusive of end values), and the quantized variance is mapped to the range of 0 to 1 (inclusive of end values).
[0179] The 48-level classification can be achieved using a class index calculated in the following way:
[0180] In another variation, three different orientations are defined, as described in Table 1, along with two frequency bands. The residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive of extreme values), and the quantized variance is mapped to a range of 0 to 2 (inclusive of extreme values). The resulting 48-level classification can be achieved using the class index calculated as follows:
[0181] In the variant, five different orientations are defined, as described in Table 2, and two frequency bands are defined. The residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive of extreme values), and the quantized variance is mapped to a range of 0 to 1 (inclusive of extreme values). The resulting 40-level classification can be achieved using the class index calculated as follows:
[0182] In the variant, three different directions are defined, as described in Table 1, and four frequency bands are defined. The residual classification (e.g., subclassing) index is mapped to a range of 0 to 1 (inclusive of extreme values), and the quantized variance is also mapped to a range of 0 to 1 (inclusive of extreme values). The resulting 48-level classification can be achieved using the class index calculated as follows:
[0183] Furthermore, this aspect is not limited to ECM, VVC, or HEVC, and can be applied to, for example, other standards and recommendations, as well as any extensions of such standards and recommendations. Unless otherwise stated or technically excluded, the aspects described in this application may be used alone or in combination.
[0184] Various numerical values are used in this application. Specific values are for illustrative purposes only, and the aspects described are not limited to these specific values.
[0185] It should be noted that the syntax elements used in this article (such as terms in equations and algorithms, signals (e.g., indices), coefficients, etc.) are descriptive terms. Therefore, they do not preclude the use of other syntax element names.
[0186] Various implementations involve decoding. As used herein, "decoding" can encompass all or part of a process performed, for example, on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes, or alternatively includes, processes performed by a decoder of the various implementations described herein, such as filtering reconstructed picture patches or classifying samples.
[0187] As another example, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding; and in yet another embodiment, "decoding" refers to the entire image reconstruction process, including entropy decoding. Based on the specific context of the description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to a broader decoding process, and it is believed that those skilled in the art will readily understand this.
[0188] Various implementations involve encoding. Similar to the discussion of "decoding" above, "encoding" as used herein can include, for example, all or part of a process performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various embodiments, such processes also include, or alternatively include, processes performed by an encoder of the various implementations described herein, such as determining filter coefficients, filtering reconstructed picture patches, and classifying samples.
[0189] As another example, in one embodiment, "encoding" refers only to entropy encoding; in another embodiment, "encoding" refers only to differential encoding; and in yet another embodiment, "encoding" refers to a combination of entropy encoding and differential encoding. It will be clear from the specific context of the description whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to refer to a broader encoding process, and it is believed that those skilled in the art will understand this well.
[0190] This disclosure describes various pieces of information (e.g., syntax) that can, for example, be transmitted or stored. This information can be encapsulated or arranged in a variety of ways, including those commonly used in video standards, such as placing the information in SPS, PPS, NAL units, headers (e.g., NAL unit headers or slice headers), or SEI messages. Other methods are also available, including those commonly used in system-level or application-level standards, such as placing the information in one or more of the following: a. SDP (Session Description Protocol) is used to describe the format of multimedia communication sessions for the purposes of session announcement and session invitation, for example, as described in the RFC and used in conjunction with RTP (Real-Time Transport Protocol) transmission.
[0191] b. DASH MPD (Media Presentation Description) descriptors, such as those used in DASH and transmitted via HTTP, are associated with representations or sets of representations to provide additional features to the content representation.
[0192] c. RTP header extensions, for example, used during RTP streaming.
[0193] d. ISO basic media file format, such as that used in OMAF, and uses boxes (also called “atoms” in some specifications) of object-oriented building blocks defined by unique type identifiers and lengths.
[0194] e. An HLS (HTTP Live Streaming) manifest transmitted over HTTP. The manifest can, for example, be associated with a version or set of versions of the content to provide characteristics of that version or set of versions.
[0195] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / device.
[0196] Some embodiments involve rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, usually under constraints of computational complexity. Rate-distortion optimization is generally expressed as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist to address the rate-distortion optimization problem. For example, these approaches can be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, where a comprehensive evaluation of the encoding cost and the associated distortion of the reconstructed signal is performed after encoding and decoding. Faster methods can also be used to save encoding complexity, particularly by calculating approximate distortion based on prediction or prediction of the residual signal rather than the reconstructed signal. A hybrid of these approaches can also be used, such as by using approximate distortion only for some possible encoding options and full distortion for others. Other methods evaluate only a subset of possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but the optimization is not necessarily a comprehensive evaluation of encoding cost and associated distortion.
[0197] The implementations and aspects described herein can be implemented in, for example, methods or processes, devices, software programs, data streams, or signals. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the features in question can be implemented in other forms (e.g., devices or programs). Devices can be implemented, for example, with appropriate hardware, software, and firmware. Methods can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0198] The references to "an embodiment" or "an embodiment" or "an implementation" or "an implementation," and their variations, mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in an embodiment" or "in an implementation" or "in an implementation," and any other variations appearing throughout this application, do not necessarily refer to the same embodiment.
[0199] Additionally, this application may involve "determining" various types of information. Determining information may include one or more of the following: for example, estimation information, calculation information, prediction information, and information retrieved from memory.
[0200] Furthermore, this application may involve "accessing" various types of information. Accessing information may include one or more of the following: for example, receiving information, retrieving information (e.g., retrieving information from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, and estimating information.
[0201] Additionally, this application may relate to "receiving" various types of information. Like "accessing," receiving is a broad term. Receiving information may include one or more of the following: for example, accessing information and retrieving information (e.g., retrieving information from memory). Furthermore, "receiving" is generally referred to in one or more ways during operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0202] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of” (e.g., in the cases of “A / B”, “A and / or B”, and “at least one of A and B”) is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As yet another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, this wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to a large number of listed items.
[0203] Furthermore, as used herein, the term "signaling" specifically refers to instructing the corresponding decoder to provide certain information. For example, in some embodiments, the encoder signals a particular one of multiple filter coefficients, a pruning index, etc. In this way, the same parameters are used on both the encoder and decoder sides in the embodiments. Thus, for example, the encoder can send (explicit signaling) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the word "signal" has been referred to above, the word "signal" can also be used as a noun herein.
[0204] It will be apparent to those skilled in the art that implementations can generate various signals that are formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
[0205] This document describes numerous examples. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more of the features, apparatuses, or aspects described herein, individually or in any combination across various claim classes and types. For example, the features described herein may be implemented in a bitstream or signal that includes information generated as described herein. This information may allow a decoder to decode the bitstream, and the encoder, bitstream, and / or decoder may be implemented according to any of the described embodiments. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the features described herein may be implemented by a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, the features described herein may be implemented by a TV, set-top box, cellular phone, tablet computer, or other electronic device performing decoding. The TV, set-top box, cellular phone, tablet computer, or other electronic device may (e.g., using a monitor, screen, or other type of display) display the resulting image (e.g., an image reconstructed from the residual of a video bitstream). The TV, set-top box, cellular phone, tablet computer, or other electronic device may receive a signal including the encoded image and perform decoding.
[0206] Many embodiments have been described above. Features of these embodiments may be provided individually or in any combination across various claim classes and types.
[0207] A decoding method is disclosed, the decoding method comprising: Reconstructing samples of image blocks; Classify at least one reconstructed sample; and The at least one reconstructed sample is filtered based on the classification. The classification of the at least one sample is based on a linear combination of at least two values selected from gradient-based values, frequency band indexes, residual-based classification indexes, and variance-based classification indexes.
[0208] An encoding method, comprising: Reconstructing samples of image blocks; Classify at least one reconstructed sample; and The at least one reconstructed sample is filtered based on the classification. The classification of the at least one sample is based on a linear combination of at least two values selected from gradient-based values, frequency band indexes, residual-based classification indexes, and variance-based classification indexes.
[0209] In the example, the classification of the sample is based on at least one linear combination of gradient-based values and frequency band indices.
[0210] In the example, the classification of the sample is based on a linear combination of at least one gradient-based value and a residual-based classification index.
[0211] In the example, the classification of the sample is based on a linear combination of at least one gradient-based value and a variance-based classification index.
[0212] In the example, the samples are classified based on a linear combination of directional values and variance-based classification indices.
[0213] In the example, the classification of the sample is based on equal to or The values of , where D is the directionality value and Cv is the variance-based classification index, and N and M are integer values.
[0214] In the example, the classification of the sample is based on a linear combination of at least one gradient-based value, a frequency band index, and a residual-based classification index.
[0215] In the example, the classification of the sample is based on a linear combination of at least one gradient-based value, a frequency band index, and a variance-based classification index.
[0216] In the example, the classification of the sample is based on a linear combination of at least one gradient-based value, a residual-based classification index, and a variance-based classification index.
[0217] In the example, the classification of the sample is based on a linear combination of at least one gradient-based value, a frequency band index, a residual-based classification index, and a variance-based classification index.
[0218] In the example, the classification of the samples is based on a linear combination of a frequency band index and a residual-based classification index.
[0219] In the example, the classification of the samples is based on a linear combination of a frequency band index and a variance-based classification index.
[0220] In the example, the classification of the samples is based on a linear combination of a frequency band index, a residual-based classification index, and a variance-based classification index.
[0221] In the example, the classification of the samples is based on a linear combination of a residual-based classification index and a variance-based classification index.
[0222] A decoding device is disclosed, the decoding device including one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform a method of any of the disclosed examples.
[0223] An encoding device is disclosed, the encoding device including one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform a method of any of the disclosed examples.
[0224] A computer program is disclosed, the computer program including program code instructions, which, when executed by a processor, are used to implement a decoding method of any of the disclosed examples.
[0225] A computer program is disclosed, the computer program including program code instructions, which, when executed by a processor, are used to implement the encoding method of any of the disclosed examples.
Claims
1. A decoding method, comprising: Reconstructing samples of image blocks; Classify at least one reconstructed sample; as well as The at least one reconstructed sample is filtered based on the classification. The classification of the at least one sample is based on a linear combination of at least two values selected from gradient-based values, frequency band indexes, residual-based classification indexes, and variance-based classification indexes.
2. The method according to claim 1, wherein, The classification of the samples is based on at least one linear combination of gradient-based values and frequency band indices.
3. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of at least one gradient-based value and a residual-based classification index.
4. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of at least one gradient-based value and a variance-based classification index.
5. The method according to claim 4, wherein, The classification of the samples is based on a linear combination of directional values and variance-based classification indices.
6. The method according to claim 5, wherein, The classification of the samples is based on equality or The values of , where D is the directionality value and Cv is the variance-based classification index, and N and M are integer values.
7. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of at least one gradient-based value, a frequency band index, and a residual-based classification index.
8. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of at least one gradient-based value, a frequency band index, and a variance-based classification index.
9. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of at least one gradient-based value, a residual-based classification index, and a variance-based classification index.
10. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of at least one gradient-based value, a frequency band index, a residual-based classification index, and a variance-based classification index.
11. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of frequency band index and residual-based classification index.
12. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of a frequency band index and a variance-based classification index.
13. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of a frequency band index, a residual-based classification index, and a variance-based classification index.
14. The method according to claim 1, wherein, The classification of the samples is based on a linear combination of a residual-based classification index and a variance-based classification index.
15. An encoding method, comprising: Reconstructing samples of image blocks; Classify at least one reconstructed sample; as well as The at least one reconstructed sample is filtered based on the classification. The classification of the at least one sample is based on a linear combination of at least two values selected from gradient-based values, frequency band indexes, residual-based classification indexes, and variance-based classification indexes.
16. A decoding device comprising one or more processors and at least one memory, said at least one memory being coupled to said one or more processors, wherein, The one or more processors are configured to perform the method according to any one of claims 1 to 14.
17. An encoding device comprising one or more processors and at least one memory, said at least one memory being coupled to said one or more processors, wherein, The one or more processors are configured to perform the method according to claim 15.
18. A computer program comprising program code instructions, which, when executed by a processor, are configured to implement the method according to any one of claims 1 to 14.
19. A computer-readable storage medium having instructions stored thereon for implementing the method according to claim 15.