Adaptive in-loop filtering in video coding

By employing a single-pass coding solution in video encoding, and utilizing modified rate-distortion optimization parameters and adaptive filtering decisions, the problems of in-loop filtering decision complexity and bit waste are solved, achieving more efficient coding and bit rate optimization.

CN120835160APending Publication Date: 2025-10-24INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302599.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-16
Filing Date
2025-03-14
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing video coding techniques suffer from increased complexity and bit waste when making in-loop filtering decisions, especially when using SAO and ALF filters, failing to effectively avoid overlapping decoding gains and complexity.

Method used

A single-pass coding solution is adopted, which makes block-level SAO decisions in intra-frame or scene-changing frames by modifying the bitrate-distortion optimization parameters, and adaptively disables picture-level or slice-level SAO decisions in non-intra-frame or non-scene-changing frames, combined with ALF decision to improve efficiency.

Benefits of technology

It reduces encoder complexity, optimizes the use of in-loop filter bits, improves coding efficiency, and reduces bit rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120835160A_ABST
    Figure CN120835160A_ABST
Patent Text Reader

Abstract

The invention relates to adaptive in-loop filtering in video coding. In some compression techniques, an in-loop filter may be included to effectively remove coding artifacts, and at the same time improve objective quality measurements. To avoid increasing complexity in the encoder, a single-pass encoding solution may be implemented to efficiently and effectively make a reasonable and optimal in-loop filtering decision. The solution may improve in-loop filtering bit usage (e.g., reduce bit rate), and at the same time reduce complexity in an encoder.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Video compression is a technology used to make video files smaller and more easily transmitted over the Internet. There are different methods and algorithms for video compression, which have different performance and trade-offs. Video compression involves encoding and decoding. Encoding is the process of transforming (uncompressed) video data into a compressed format. Decoding is the process of recovering video data from a compressed format. An encoder-decoder system is referred to as a codec. BRIEF DESCRIPTION OF DRAWINGS

[0002] The embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. For purposes of convenience and clarity, identical or similar reference numerals are used throughout the drawings to refer to like structures or methods. The embodiments are illustrated by way of example as being used in video compression, but the scope of the disclosure is not limited to video compression.

[0003] FIG. 1 An encoding system and multiple decoding systems according to some embodiments of the disclosure are shown.

[0004] FIG. 2 An exemplary encoder that encodes video frames and outputs an encoded bitstream according to some embodiments of the disclosure is shown.

[0005] FIG. 3 An exemplary decoder that decodes an encoded bitstream and outputs a decoded video according to some embodiments of the disclosure is shown.

[0006] FIG. 4 An in-loop filter and in-loop filter decision portion of an encoder according to some embodiments of the disclosure is shown.

[0007] FIG. 5 An in-loop filter decision portion according to some embodiments of the disclosure is shown.

[0008] FIG. 6 An in-loop filter decision process according to some embodiments of the disclosure is shown.

[0009] FIG. 7 A method for determining one or more in-loop filter decisions according to some embodiments of the disclosure is shown.

[0010] FIG. 8 A block diagram of an exemplary computing device according to some embodiments of the disclosure is depicted. DETAILED DESCRIPTION SUMMARY

[0011] Video coding or video compression is a process that compresses video data for storage, transmission, and playback. Video compression can involve taking a large amount of raw video data and applying one or more compression techniques to reduce the amount of data needed to represent the video while maintaining an acceptable level of visual quality. In some cases, video compression can provide for transmission and efficient storage of video content over limited bandwidth networks.

[0012] A frame can include an image or a single still image. A frame can have millions of pixels. For example, a frame of uncompressed 4K video can have a resolution of 3840 x 2160 pixels. Pixels can have luma / chroma component values. The terms "frame" and "picture" can be used interchangeably. There are several frame types or picture types. An I-frame or intra-frame can be the least compressible and does not depend on other frames for decoding. An I-frame can include a scene change frame. A P-frame can depend on data from a previous frame for decoding and can be more compressible than an I-frame. A B-frame can depend on data from both a previous frame and a subsequent frame for decoding and can be more compressible than an I-frame and a P-frame. Other frame types can include a reference B-frame and a non-reference B-frame. P-frames and B-frames can be referred to as inter-frame. The order of arrangement of I-frames, P-frames, and B-frames can be referred to as a group of pictures (GOP). A slice can be a spatially distinct region of a frame that is coded separately from any other region in the same frame.

[0013] In some cases, a frame can be partitioned into one or more blocks. Blocks can be used for block-based compression. Blocks can have much smaller sizes such as 512 x 512 pixels, 256 x 256 pixels, 128 x 128 pixels, 64 x 64 pixels, 32 x 32 pixels, 16 x 16 pixels, 8 x 8 pixels, 4 x 4 pixels, and so forth. A block can include a square or rectangular region of a frame. Different terminology can be used for blocks or different partition structures used to create blocks by various video compression techniques. In some video compression techniques, a frame can be partitioned into coding tree units (CTUs). A CTU can be split (separately for luma and chroma components) into coding tree blocks (CTBs). CTBs can have sizes of 64 x 64 pixels, 32 x 32 pixels, or 16 x 16 pixels. CTBs can be split into coding units (CUs). CUs can be split into prediction units (PUs) and / or discrete cosine transform (DCT) transform units (TUs).

[0014] One of the tasks of an encoder in a video codec is to make encoding decisions at different levels of the video (e.g., sequence level, GOP level, frame / picture level, slice level, CTU level, CTB level, block level, CU level, PU level, TU level, etc.) based on a desired bit rate and / or a desired (objective and / or subjective) quality. Making encoding decisions can include evaluating different options or parameter values for encoding data and determining the best option or parameter value that can achieve the desired bit rate and / or quality. The selected option and / or parameter value can be applied to encode the video to generate a bitstream. The selected option and / or parameter value will be encoded in the bitstream to inform a decoder how to decode the encoded bitstream according to the encoding decisions made by the encoder. Modern codecs provide a wide range of options and parameter values. While evaluating all possible combinations of options and parameter values can yield the optimal encoding decisions, the encoder does not have unlimited resources to bear the complexity that making globally optimal encoding decisions would require.

[0015] In some compression techniques, in-loop filters, luma mapping and chroma scaling (LMCS) filters, deblocking filters, sample adaptive offset (SAO) filters, and adaptive loop filters (ALF) can be included to effectively remove coding artifacts and at the same time improve objective quality measurements. For example, ALF can be applied to SAO filtered blocks to effectively remove coding artifacts and improve objective quality. While SAO plus ALF filtering can generally provide better coding gain, the coding gain of the two filters tends to overlap. In many cases, the ALF filter itself can provide similar quality improvement as the combined SAO and ALF filters. For these cases, bits used by SAO are wasted. Making the best joint filter decision can greatly increase complexity, so some encoders avoid complexity by making SAO decisions and ALF decisions in sequence. In other words, the SAO decisions are made without considering the subsequent ALF filter decisions. As a result, the decisions are not optimal. Bit and processing time waste cannot be avoided.

[0016] To avoid increasing complexity in the encoder, a single pass encoding solution can be implemented to efficiently and effectively make reasonably optimal in-loop filtering decisions. The solution can improve in-loop filtering bit usage (e.g., reduce bit rate) and at the same time reduce complexity in the encoder.

[0017] In some embodiments, if the frame is an intra frame or a scene change frame, block-level SAO decisions will be made with modified rate-distortion optimization (RDO) parameters. ALF decisions will be made after SAO decisions. In this context, an intra frame can be a frame that has been marked as an intra frame, or a frame that is to be all intra coded. In this context, a scene change frame can be a frame that represents a new scene captured in a sequence of video or video frames. A scene change can include content that is significantly different from, or has no (temporal) correlation with, the content in the previous frame.

[0018] In some embodiments, if the frame is not an intra frame or a scene change frame, picture-level or slice-level SAO decisions can be adaptively disabled if the corresponding picture-level or slice-level quantization parameter (QP) is greater than a content-dependent QP threshold (or a frame classification-dependent QP threshold). If the picture-level or slice-level SAO is disabled, all block-level SAO decisions and syntax coding are skipped. ALF decisions will be made while skipping SAO block-level SAO decisions and syntax coding.

[0019] The techniques described and illustrated herein for making in-loop filter decisions can be applied to a variety of codecs, such as AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), AV1 (AOMedia Video 1), and VVC (Versatile Video Coding). AVC (also known as “ITU-T H.264”) was approved in 2003 and last revised on 2021-08-22. HEVC (also known as “ITU-T H.265”) was approved in 2013 and last revised on 2023-09-13. AV1 is a video coding codec designed for video transport over the Internet. The “AV1 Bitstream & Decoding Process Specification” version 1.1.1 (with errata) was last modified in 2019. VVC (also known as “ITU-T H.266”) was finalized in 2020. While the techniques described herein relate to VVC, the present disclosure contemplates that the techniques can be applied to other codecs that have SAO filters and ALF as in-loop filters. Video compression

[0020] FIG. 1 An encoding system 130 and one or more decoding systems 150 according to some embodiments of the present disclosure are shown 1……D .

[0021] The techniques described and illustrated herein can be implemented in FIG. 8The encoding system 130 is implemented on a computing device 800. The encoding system 130 can be implemented in the cloud or in a data center. The encoding system 130 can be implemented on a device used to capture video. The encoding system 130 can be implemented on a standalone computing system. The encoding system 130 can perform encoding processes in video compression. The encoding system 130 can receive video (e.g., uncompressed video, initial video, raw video, etc.) that includes a sequence of video frames 104. The video frames 104 can include images or image frames that make up the video. The video can have a frame rate or frames per second (FPS) that defines the number of frames of video per second. The higher the FPS, the more real and smooth the video appears. Typically, the FPS is greater than 24 frames per second to give a human viewer a natural and realistic viewing experience. Examples of video can include television episodes, movies, short films, short videos (e.g., less than 15 seconds in length), videos that capture gaming experiences, computer screen content, video conferencing content, live event streaming content, sports content, surveillance videos, videos taken using a mobile computing device (e.g., a smartphone), etc. In some cases, the video can include a mix or combination of different types of video.

[0022] The encoding system 130 can include an encoder 102 that receives the video frames 104 and encodes the video frames 104 into an encoded bitstream 180. FIG. 2 An example implementation of the encoder 102 is shown in FIG. 1.

[0023] The encoded bitstream 180 can be compressed, meaning that the size of the encoded bitstream 180 can be smaller than the video frames 104. The encoded bitstream 180 can include a series of bits, e.g., having a number of 0s and a number of 1s. The encoded bitstream 180 can have header information, payload information, and footer information that can be encoded as bits in the bitstream. The header information can provide information about one or more of the following: the format of the encoded bitstream 180, the encoding process implemented in the encoder 102, parameters of the encoder 102, and metadata of the encoded bitstream 180. For example, the header information can include one or more of the following: resolution information, frame rate, aspect ratio, color space, etc. The payload information can include data representing the content of the video frames 104, such as sample frames, symbols, syntax elements, etc. For example, the payload information can include bits encoding one or more of motion predictors, transform coefficients, prediction modes, and quantization levels of the video frames 104. The footer information can indicate the end of the encoded bitstream 180. The footer information can include other information, including one or more of the following: checksums, error correction codes, and signatures. The format of the encoded bitstream 180 can vary depending on the specification of the encoding and decoding processes (i.e., codec).

[0024] The encoded bitstream 180 can include packets in which the encoded video data and signaling information can be packetized. One example format is the Open Bitstream Unit (OBU) used in AV1 encoded bitstreams. An OBU can include a header and a payload. The header can include information about the OBU, such as information that indicates the type of the OBU. Examples of OBU types can include a sequence header OBU, a frame header OBU, a metadata OBU, a time delimiter OBU, and a tile group OBU. The payload in an OBU can carry quantized transform coefficients and syntax elements that can be used in a decoder to correctly decode the encoded video data to regenerate video frames.

[0025] The encoded bitstream 180 can be transmitted to one or more decoding systems 150 via a network 140 1……D . The network 140 can be the Internet. The network 140 can include one or more of the following: a cellular data network, a wireless data network, a wired data network, a cable Internet network, a fiber optic network, a satellite Internet network, etc.

[0026] D decoding systems 150 are shown 1……D . At least one of the decoding systems 150 1……D may be implemented on the computing device 800 of FIG. 8 . Examples of the decoding systems 150 1……D may include a personal computer, a mobile computing device, a gaming device, an augmented reality device, a mixed reality device, a virtual reality device, a television, etc. Each of the decoding systems 150 1……D may perform a decoding process in video compression. Each of the decoding systems 150 1……D may include a decoder (e.g., decoders 1...D 162 1……D ) and one or more display devices (e.g., display devices 1...D 164 1……D ). FIG. 3 An example implementation of a decoder (e.g., decoder 1 1621) is shown in

[0027] For example, the decoding system 1 1501 can include the decoder 1 1621 and the display device 1 1641. The decoder 1 1621 can implement a decoding process of video compression. The decoder 1 1621 can receive the encoded bitstream 180 and produce a decoded video 1681. The decoded video 1681 can include a series of video frames, which can be versions or reconstructed versions of the video frames 104 encoded by the encoding system 130. The display device 1 1641 can output the decoded video 1681 for display to one or more human viewers or a user of the decoding system 1 1501.

[0028] For example, the decoding system 2 1502 can include a decoder 2 1622 and a display device 2 1642. The decoder 2 1622 can implement the decoding process of video compression. The decoder 2 1622 can receive the encoded bitstream 180 and produce a decoded video 1682. The decoded video 1682 can include a series of video frames, which can be versions or reconstructed versions of the video frames 104 encoded by the encoding system 130. The display device 2 1642 can output the decoded video 1682 for display to one or more human viewers or users of the decoding system 2 1502.

[0029] For example, the decoding system D 150 D may include a decoder D 162 D and a display device D 164 D . The decoder D 162 D may implement the decoding process of video compression. The decoder D 162 D may receive the encoded bitstream 180 and produce a decoded video 168 D . The decoded video 168 D may include a series of video frames, which can be versions or reconstructed versions of the video frames 104 encoded by the encoding system 130. The display device D 164 D may output the decoded video 168 D for display to one or more human viewers or users of the decoding system D 150 D . Video encoder

[0030] FIG. 2 An encoder 102 that encodes video frames and outputs an encoded bitstream is shown in accordance with some embodiments of the present disclosure. The encoder 102 can include one or more of the following: signal processing operations and data processing operations, including inter and intra prediction, transform, quantization, in-loop filtering, and entropy coding. The encoder 102 can include a reconstruction loop involving inverse quantization and inverse transform to ensure that the same reference blocks and frames will be seen by the decoder. The encoder 102 can receive the video frames 104 and encode the video frames 104 into the encoded bitstream 180. The encoder 102 can include one or more of the following: partitioning 206, transform and quantization 214, inverse transform and inverse quantization 218, in-loop filter 228, motion estimation 234, inter prediction 236, intra prediction 238, and entropy coding 216.

[0031] The partitioning 206 can split a frame in the video frame 104 into blocks of pixels. Different codecs can allow for different variable ranges of block sizes. In one codec, the frame can be partitioned by the partitioning 206 into blocks of 128x128 or 64x64 pixels. In some cases, the frame can be partitioned by the partitioning 206 into blocks of 32x32 pixels or 16x16 pixels. In some cases, the frame can be partitioned by the partitioning 206 into blocks of 256x256 or 512x512 pixels. Large blocks can be referred to as superblocks. The partitioning 206 can further split each superblock using a multiway partitioning tree structure. In some cases, the partitioning of a superblock can be further recursively split by the partitioning 206 using a multiway partitioning tree structure (e.g., all the way down to 4x4 size blocks). In another codec, the frame can be partitioned by the partitioning 206 into CTUs of 128x128 pixels. The partitioning 206 can split the CTUs into four CUs using a quadtree partitioning structure. The partitioning 206 can further recursively split the CUs using a quadtree partitioning structure. The partitioning 206 can subdivide the CUs using a multi-type tree structure (e.g., a quadtree, binary tree, or ternary tree structure). The smallest CUs can have a size of 4x4. The partitioning 206 can output the initial samples 208, e.g., as blocks of pixels.

[0032] Intra prediction 238 can predict samples of a block from reconstructed prediction samples of previously coded spatial neighboring / reference blocks of the same frame. Intra prediction 238 can receive reconstructed prediction samples 226 (of previously coded spatial neighboring blocks of the same frame). Reconstructed prediction samples 226 can be generated by summer 222 from prediction samples 212 and prediction residual 224. Intra prediction 238 can determine a suitable predictor for predicting samples from reconstructed prediction samples of previously coded spatial neighboring / reference blocks of the same frame. Intra prediction 238 can generate prediction samples 212 generated using the suitable predictor. Intra prediction 238 can output or identify the neighboring / reference blocks and the predictor used to generate prediction samples 212. The identified neighboring / reference blocks and predictor can be coded in coded bitstream 180 to enable a decoder to reconstruct the block using the same neighboring / reference blocks and predictor. In one codec, intra prediction 238 can support multiple diverse predictors, e.g., 56 different predictors. Some predictors, e.g., directional predictors, can capture different spatial redundancies in directional textures. Directional predictors can be used in intra prediction 238 to predict pixel values of a block by extrapolating pixel values of neighboring / reference blocks along a certain direction. Intra prediction 238 of different codecs can support different sets of predictors to exploit different spatial patterns within the same frame. Examples of predictors can include direct current (DC), planar, Paeth, smooth, vertical smooth, horizontal smooth, recursive-based filter pattern, generating chroma from luma, intra block copy, palette, multiple reference lines, intra sub-partition, matrix-based intra prediction (matrix coefficients can be defined by offline training using neural networks), wide-angle prediction, cross-component linear model, template matching, etc. In some cases, intra prediction 238 can perform block prediction, where a prediction block can be produced from reconstructed neighboring / reference blocks of the same frame using a vector. Optionally, some type of interpolation filter can be applied to the prediction block to blend the pixels of the prediction block. Vector compensation procedures can be used in intra prediction 238 to predict pixel values of a block by translating (within the same frame) a neighboring / reference block according to a vector (and optionally applying an interpolation filter to the neighboring / reference block) to produce prediction samples 212. Intra prediction 238 can output or identify the vector applied during generation of prediction samples 212. In some codecs, intra prediction 238 can code (1) a residual vector generated from the applied vector and a vector predictor candidate and (2) information identifying the vector predictor candidate, instead of coding the applied vector itself. Intra prediction 238 can output or identify the interpolation filter type applied during generation of prediction samples 212.

[0033] Motion estimation 234 and inter prediction 236 can predict samples of a block from samples of a previously encoded frame (e.g., a reference frame in the decoded picture buffer 232). Motion estimation 234 and inter prediction 236 can perform motion compensation, which can involve identifying a suitable reference block and a suitable motion predictor (or vector) for the block and, optionally, also an interpolation filter to be applied to the reference block. Motion estimation 234 can receive initial samples 208 from the partitioning 206. Motion estimation 234 can receive samples from the decoded picture buffer 232 (e.g., samples of a previously encoded frame or reference frame). Motion estimation 234 can use multiple reference frames to determine one or more suitable motion predictors. A motion predictor can include a reference block and a motion vector, which can be applied to generate a motion compensated block or predicted block. A motion predictor can include a motion vector that captures the movement of a block between frames in a video. Motion estimation 234 can output or identify one or more reference frames and one or more suitable motion predictors. Inter prediction 236 can apply the one or more suitable motion predictors determined in motion estimation 234 and the one or more reference frames to generate predicted samples 212. The identified reference frame(s) and motion predictor(s) can be encoded in the encoded bitstream 180 to enable a decoder to use the same reference frame(s) and motion predictor(s) to reconstruct the block. In one codec, motion estimation 234 can implement a single reference frame prediction mode, in which a single reference frame with a corresponding motion predictor is used for inter prediction 236. Motion estimation 234 can implement a multiple reference frame prediction mode, in which two reference frames with two corresponding motion predictors are used for inter prediction 236. In one codec, motion estimation 234 can implement techniques for searching and identifying good reference frame(s) that can yield the most efficient motion predictor. Techniques in motion estimation 234 can include searching good reference frame candidates in space (within the same frame) and in time (in previously encoded frames). Techniques in motion estimation 234 can include searching a spatial neighborhood for a pool of spatial candidates. Techniques in motion estimation 234 can include generating a pool of temporal candidates with a temporal motion field estimation mechanism. Techniques in motion estimation 234 can use a motion field estimation process. Thereafter, temporal and spatial candidates can be ranked, and a suitable motion predictor can be determined. In one codec, inter prediction 236 can support multiple diverse motion predictors.Examples of predictors can include geometric motion vectors (complex, non-linear motion), warped motion compensation (affine transform that captures non- translational object movement), overlapped block motion compensation, advanced compound prediction (compound wedge prediction, differential modulation masking prediction, frame distance based compound prediction, and inter-intra compound prediction), dynamic spatial and temporal motion vector referencing, affine motion compensation (capture higher order motion such as rotation, scaling, and shearing), adaptive motion vector resolution mode, geometric partitioning mode, optical flow, prediction refinement with optical flow, weighted bi-prediction, extended merge prediction, etc. Optionally, an interpolation filter of some type can be applied to the prediction block to blend the pixels of the prediction block. The pixel values of the block can be predicted using the motion predictors / vectors determined in the motion estimation 234 and inter prediction 236, and optionally applying an interpolation filter. In some cases, the inter prediction 236 can perform motion compensation, where the prediction block can be generated from a reconstructed reference block of a reference frame using the motion predictors / vectors. The inter prediction 236 can output or identify the motion predictors / vectors applied during the generation of the prediction samples 212. In some codecs, the inter prediction 236 can encode (1) residual vectors generated from the applied vectors and vector predictor candidates and (2) information identifying the vector predictor candidates, rather than the applied vectors themselves. The inter prediction 236 can output or identify the interpolation filter types applied during the generation of the prediction samples 212.

[0034] The mode selection 230 can be informed by components such as the motion estimation 234 to determine whether inter prediction 236 or intra prediction 238 can be more efficient for encoding the block (thus making an encoding decision). The inter prediction 236 can output prediction samples 212 for the prediction block. The inter prediction 236 can output the selected predictors and the selected interpolation filter (if applicable) that can be used to generate the prediction block. The intra prediction 238 can output prediction samples 212 for the prediction block. The intra prediction 238 can output the selected predictors and the selected interpolation filter (if applicable) that can be used to generate the prediction block. Regardless of the mode, the prediction residual 210 can be generated by the subtractor 220 by subtracting the initial samples 208 from the prediction samples 212. In some cases, the prediction residual 210 can include residual vectors from the inter prediction 236 and / or the intra prediction 238.

[0035] Transform and quantization 214 can receive prediction residual 210. Prediction residual 210 can be generated by subtractor 220, which takes initial sample 208 and subtracts prediction sample 212 to output prediction residual 210. Prediction residual 210 can be referred to as a prediction error (e.g., an error between initial sample and prediction sample 212) for intra prediction 238 and inter prediction 236. Prediction error has a smaller range of values than initial sample and can be coded into encoded bitstream 180 using fewer bits. Transform and quantization 214 can include one or more of transform and quantization. Transform can include converting prediction residual 210 from a spatial domain to a frequency domain. Transform can include applying one or more transform kernels. Examples of transform kernels can include DCT in horizontal and vertical forms, asymmetric discrete sine transform (ADST), flipped ADST, identity transform (IDTX), multiple transform selection, low-frequency non-separable transform, sub-block transform, non-square transform, DCT-VIII, discrete sine transform VII (DST-VII), discrete wavelet transform (DWT), etc. Transform can convert prediction residual 210 into transform coefficients. Quantization can quantize transform coefficients, e.g., by reducing precision of transform coefficients. Quantization can include using a quantization matrix (e.g., linear and non-linear quantization matrices). Elements in quantization matrix can be larger for higher frequency bands and smaller for lower frequency bands, which means that higher frequency coefficients are quantized more coarsely and lower frequency coefficients are quantized more finely. Quantization can include dividing each transform coefficient by a corresponding element in quantization matrix and rounding to the nearest integer. In practice, quantization matrix can achieve different QPs for different frequency bands and chroma planes, and can use spatial prediction. A suitable quantization matrix can be selected and signaled for each frame and encoded in encoded bitstream 180. Transform and quantization 214 can output quantized transform coefficients and syntax elements 278, which indicate coding modes and parameters used in encoding processes implemented in encoder 102.

[0036] Inverse transform and inverse quantization 218 can apply inverse operations performed in transform and quantization 214 to produce reconstructed prediction residuals 224 as part of a reconstruction path to generate a decoded picture buffer 232 for encoder 102. Inverse transform and inverse quantization 218 can receive quantized transform coefficients and syntax elements 278. Inverse transform and inverse quantization 218 can perform one or more inverse quantization operations, e.g., apply an inverse quantization matrix, to obtain unquantized / initial transform coefficients. Inverse transform and inverse quantization 218 can perform one or more inverse transform operations, e.g., inverse transforms (e.g., inverse DCT, inverse DWT, etc.), to obtain reconstructed prediction residuals 224. The reconstruction path is provided in encoder 102 to generate reference blocks and frames, which are stored in decoded picture buffer 232. The reference blocks and frames can match blocks and frames that will be generated in a decoder. The reference blocks and frames are used as reference blocks and frames by motion estimation 234, inter prediction 236, and intra prediction 238.

[0037] In-loop filter 228 may implement a filter that removes artifacts introduced by the encoding process in encoder 102 (e.g., processing performed by partitioning 206 and transform and quantization 214). In-loop filter 228 may receive reconstructed prediction samples 226 from summer 222 and output the frame to decoded picture buffer 232. Examples of in-loop filters may include a constrained low-pass filter, a directional de-ringing filter, an edge-directed conditional replacement filter, an in-loop restoration filter, a Wiener filter, a self-guided restoration filter, a constrained directional enhancement filter (CDEF), LMCS, a SAO filter, an ALF, a cross-component ALF, a low-pass filter, a deblocking filter, and the like. For example, applying a deblocking filter across the boundary between two blocks can address blocking artifacts caused by the Gibbs phenomenon. In some embodiments, in-loop filter 228 may draw data from a frame buffer containing reconstructed prediction samples 226 for various blocks of a video frame. In-loop filter 228 may determine whether to apply a particular in-loop filter. The in-loop filter 228 may determine one or more suitable filters to obtain good visual quality and / or one or more suitable filters to appropriately remove artifacts introduced by the encoding process in the encoder 102. The in-loop filter 228 may determine the type of in-loop filter to apply across the boundary between two blocks. The in-loop filter 228 may determine one or more strengths (e.g., filter coefficients) of the in-loop filter to apply across the boundary between the two blocks based on the reconstructed prediction samples 226 of the two blocks. In some cases, the in-loop filter 228 may take into account the desired bit rate when determining the one or more suitable filters. In some cases, the in-loop filter 228 may take into account the specified QP when determining the one or more suitable filters. The in-loop filter 228 may apply one or more (suitable) filters across the boundary separating the two blocks. After applying the one or more (suitable) filters, the in-loop filter 228 may write the (filtered) reconstructed samples to a frame buffer such as a decoded picture buffer 232. In FIG. 4 to FIG. 5 The decision making portion of the in-loop filter 228 for making one or more in-loop filtering decisions is shown and described in FIG. FIG. 6 to FIG. 7 The operations that may be performed by the decision making portion are shown and described in FIG.

[0038] Entropy coding 216 can receive quantized transform coefficients and syntax elements 278 (e.g., referred to as symbols herein) and perform entropy coding. Entropy coding 216 can generate and output an encoded bitstream 180. Entropy coding 216 can exploit statistical redundancies and apply lossless algorithms to encode the symbols and produce a compressed bitstream, e.g., encoded bitstream 180. Entropy coding 216 can implement some version of arithmetic coding. Different versions can have different advantages and disadvantages. In one codec, entropy coding 216 can implement adaptive multi-symbol arithmetic coding (symbols to symbols). In another codec, entropy coding 216 can implement a context-based adaptive binary arithmetic coder (CABAC). Binary arithmetic coding is different from multi-symbol arithmetic coding. Binary arithmetic coding encodes only one bit at a time, e.g., with a binary value of 0 or 1. Binary arithmetic coding can first convert each symbol to a binary representation (e.g., using a fixed number of bits per symbol). Dealing only with binary values of 0 or 1 can simplify the calculations and reduce complexity. Binary arithmetic coding can assign a probability to each binary value (e.g., a chance that the bit has a binary value of 0 and a chance that the bit has a binary value of 1). Multi-symbol arithmetic coding performs encoding on a symbol table with at least two or three symbol values and assigns a probability to each symbol value in the symbol table. Multi-symbol arithmetic coding can encode more bits at a time, which can result in fewer number of operations to encode the same amount of data. Multi-symbol arithmetic coding can require more calculations and storage (as probability estimates can be updated for each element in the symbol table). Maintaining and updating probabilities (e.g., cumulative probability estimates) for each possible symbol value in multi-symbol arithmetic coding can be more complex (e.g., complexity grows with the size of the symbol table). Multi-symbol arithmetic coding is not to be confused with binary arithmetic coding as the two different entropy coding processes are implemented differently and can result in different encoded bitstreams for the same set of quantized transform coefficients and syntax elements 278. Video decoder

[0039] FIG. 3A decoder 1 1621 that decodes the encoded bitstream and outputs a decoded video is shown in accordance with some embodiments of the present disclosure. The decoder 1 1621 can include one or more of the following: signal processing operations and data processing operations, including entropy decoding, inverse transform, inverse quantization, inter and intra prediction, in-loop filtering, etc. The decoder 1 1621 can have signal and data processing operations that are opposite to the operations performed in the encoder. The decoder 1 1621 can apply the signal and data processing operations signaled in the encoded bitstream 180 to reconstruct the video. The decoder 1 1621 can receive the encoded bitstream 180 and generate and output a decoded video 1681 having a plurality of video frames. The decoded video 1681 can be provided to one or more display devices for display to one or more human viewers. The decoder 1 1621 can include one or more of the following: entropy decoding 302, inverse transform and inverse quantization 218, in-loop filters 228, inter prediction 236, and intra prediction 238. Some of the functionality has been previously described and used in an encoder, such as the encoder 102 of FIG. 2

[0040] The entropy decoding 302 can decode the encoded bitstream 180 and output symbols coded in the encoded bitstream 180. The symbols can include quantized transform coefficients and syntax elements 278. The entropy decoding 302 can reconstruct the symbols from the encoded bitstream 180.

[0041] The inverse transform and inverse quantization 218 can receive the quantized transform coefficients and the syntax elements 278 and perform operations performed in the encoder. The inverse transform and inverse quantization 218 can output reconstructed prediction residuals 224. The summer 222 can receive the reconstructed prediction residuals 224 and the prediction samples 212 and generate reconstructed prediction samples 226. The inverse transform and inverse quantization 218 can output the syntax elements 278 with signaling information for signaling / instructing / controlling operations in the decoder 1 1621, such as the mode selection 230, the intra prediction 238, the inter prediction 236, and the in-loop filters 228.

[0042] Depending on the prediction mode signaled in the encoded bitstream 180 (e.g., as a syntax element in the quantized transform coefficients and the syntax elements 278), either the intra prediction 238 or the inter prediction 236 can be applied to generate the prediction samples 212.

[0043] The summer 222 can sum the prediction samples 212 of a decoded reference block and the reconstructed prediction residuals 224 to produce the reconstructed prediction samples 226 of a reconstructed block. For the intra prediction 238, the decoded reference block can be in the same frame as the block being decoded or reconstructed. For the inter prediction 236, the decoded reference block can be in a different (reference) frame in the decoded picture buffer 232.

[0044] ​Intra prediction 238 can determine a reconstruction vector based on the residual vector and the selected vector predictor candidate. Intra prediction 238 can apply the reconstruction predictor or vector to a reconstructed block, which can be generated using a decoded reference block of the same frame, e.g., according to signaled predictor information. Intra prediction 238 can apply an appropriate interpolation filter type to the reconstructed block, e.g., according to signaled interpolation filter information, to generate prediction samples 212.

[0045] Inter prediction 236 can determine a reconstruction vector based on the residual vector and the selected vector predictor candidate. Inter prediction 236 can apply the reconstruction predictor or vector to a reconstructed block, which can be generated using a decoded reference block of a different frame from decoded picture buffer 232, e.g., according to signaled predictor information. Inter prediction 236 can apply an appropriate interpolation filter type to the reconstructed block, e.g., according to signaled interpolation filter information, to generate prediction samples 212.

[0046] In-loop filters 228 can receive reconstructed prediction samples 226. In-loop filters 228 can apply one or more filters signaled in encoded bitstream 180 to reconstructed prediction samples 226. In-loop filters 228 can output decoded video 1681. In-loop filtering and making in-loop filtering decisions

[0047] FIG. 4 In-loop filters 228 and in-loop filter decision portion 440 of encoder 102 are shown in accordance with some embodiments of the disclosure. In-loop filters 228 can be applied to reconstructed prediction samples 226, and the filtered output can be output to decoded picture buffer 232. Decoded picture buffer 232 can be used as a reference for inter prediction (e.g., inter prediction 236) of subsequent frames. In some video compression techniques, in-loop filters 228 can include a pipeline or series of in-loop filters, which can be individually or selectively disabled or enabled by syntax elements in the encoded bitstream. FIG. 2

[0048] ​In VVC, the pipeline can include an LMCS filter 402, a deblocking filter 404, an SAO filter 406, and an ALF 408. The pipeline can apply filtering to the reconstructed prediction samples 226 in the order shown (e.g., from left to right). After applying the filter 404, the SAO filter 406 can modify the reconstructed prediction samples 226 by conditionally adding an offset to each sample. The offset can be applied based on edge direction / shape (Edge Offset (EO)). The offset can be applied based on pixel level (Band Offset (BO)). The SAO filter 406 can reduce ringing artifacts caused by large transforms and longer interpolation filters. The SAO filter 406 can reduce distortion between the original input signal and the coded signal. Syntax elements including a flag can be used to inform whether the SAO filter 406 is enabled or disabled. In VVC, sps_sao_enabled_flag is used to enable this feature in the sequence header, and the picture parameter set (PPS) flag sao_info_in_ph_flag is used to indicate whether the luma and chroma SAO control flags are in the picture header (PH) or in the slice header (SH). If the flag is on, ph_sao_luma_flag and ph_sao_chroma_flag are used to adaptively enable it per picture, and slice_sao_luma_flag and slice_sao_chroma_flag are used to adaptively enable it per slice. If the SAO filter 406 is enabled in the current slice, the SAO filter 406 is applied at the coded tree block (CTB) level. Each color component in the CTB has its own SAO parameters, including the SAO type (e.g., on / off, EO or BO classification) and corresponding EO-related parameters and BO-related parameters. While the SAO filter 406 can improve both subjective and objective quality, its bit usage (i.e., bits used to code the syntax elements, or SAO usage bits) as described herein is not trivial. The ALF 408 is applied to the SAO-filtered block at the output of the SAO filter 406 to further remove coding artifacts and improve objective quality. The ALF 408 can apply a diamond-shaped filter with specific coefficients selected based on the content of the data. While the SAO filter 406 plus ALF 408 solution can still provide better coding gain than a single filter, the coding gain of the SAO filter 406 plus ALF 408 often overlaps. In some, if not most, cases, the ALF 408 itself can provide a substantial portion, if not most, of the quality improvement as the combined solution of the SAO filter 406 plus ALF 408. For these cases, the bits used to inform SAO usage or SAO usage bits are wasted. It is also not practical to apply multi-pass encoding to determine the optimal joint filter decision for the SAO filter 406 and the ALF 408.

[0049] In-loop filter 228 can include an in-loop filter decision portion 440 to make one or more encoding decisions. In this case, in-loop filter decision portion 440 can make one or more in-loop filter decisions 460. In-loop filter decision portion 440 can output one or more in-loop filter decisions 460 to control LMCS filter 402, deblocking filter 404, SAO filter 406, and ALF 408. In-loop filter decisions 460 can be encoded as syntax elements in the encoded bitstream. For example, one or more in-loop filter decisions 460 can include decisions regarding whether to enable or disable a particular in-loop filter, and parameter value(s) to use with the particular in-loop filter. In-loop filter decision portion 440 can receive video frame 104. In-loop filter decision portion 440 can receive samples 490, such as reconstructed prediction samples 226, filtered samples 480, filtered samples 482, and filtered samples 484. In-loop filter decision portion 440 can use video frame 104 and / or samples 490 to determine a distortion amount for a certain option and / or a certain parameter value to use with an in-loop filter. When making in-loop filter decisions, in-loop filter decision portion 440 can perform rate-distortion optimization. RDO determines the tradeoff between bit rate (e.g., compression rate) and distortion (e.g., quality, objective quality, subjective quality, etc.) introduced by the compression process. The goal of RDO is to make the best encoding decisions (in this case, one or more in-loop filter decisions 460) that minimize a rate-distortion cost function that balances bit rate and distortion in the following equation: Cost = distortion + λ * bitrate (Equation 1)

[0050] Cost represents the rate-distortion cost. Distortion represents the distortion (e.g., mean squared error, sum of absolute differences, loss in objective quality, loss in subjective quality, etc.). Bitrate represents the bit rate or the number of bits used to encode the data. Lambda or λ is an RDO parameter (sometimes referred to as a Lagrange multiplier) that can control or adjust the relative importance of bit rate to distortion in the rate-distortion cost function. A higher value of λ means that it is more important to reduce the bit rate. A lower value of λ means that it is more important to reduce the distortion.

[0051] For the in-loop filter decision portion 440 to make in-loop filter decisions for the SAO filter 406 and the ALF 408 in single pass, a technical challenge is to identify scenarios where the in-loop filter decision for the SAO filter 406 and the use of the SAO filter 406 should be skipped so as not to waste the SAO usage bit. One insight is that if a certain frame is expected to have a relatively high bit rate, the SAO usage bit will be a small fraction of the total number of bits used to encode the data. Thus, it can be worthwhile to enable the SAO filter 406 to improve quality and reduce distortion. If the SAO filter 406 is to be considered, the RDO decision for the SAO filter 406 can be biased by increasing the RDO parameter λ. Increasing the RDO parameter λ can put more emphasis on reducing the bit rate, such that the SAO filter 406 can be enabled when the distortion is relatively low. The in-loop filter decision portion 440 can address this technical challenge by implementing the operations and portions as described in FIG. 4 to FIG. 7 "Techniques for In-Loop Filter Decisioning in Single Pass," which is incorporated by reference in its entirety.

[0052] FIG. 5 An in-loop filter decision portion 440 according to some embodiments of the present disclosure is shown. The in-loop filter decision portion 440 can use one or more heuristics to identify situations where the in-loop filter decision for the SAO filter 406 and the use of the SAO filter 406 should be skipped. The in-loop filter decision portion 440 can receive information such as the video frame 104 and the samples 490. The in-loop filter decision portion 440 can receive information such as the QP 556. The QP 556 can be a QP specified for a given frame or picture. The QP 556 can be determined or specified by the encoding application based on one or more target requirements for encoding the video frame 104. The QP 556 can be determined or specified by the quantization 214 of the encoder 102. The in-loop filter decision portion 440 can receive the current frame 550 to be encoded and encoding information associated with the current frame 550. The encoding information can be provided by one or more components or portions of the encoder 102. The encoding information can be determined by one or more components or portions of the encoder 102. The in-loop filter decision portion 440 can output one or more in-loop filter decisions 460. FIG. 2 FIG. 2 FIG. 2

[0053] ​​​In-loop filter decision portion 440 can include frame type determination 502. Frame type determination 502 can receive a current frame 550 to be encoded and encoding information associated with current frame 550. Frame type determination 502 can determine a frame type of current frame 550 from the encoding information associated with current frame 550. Current frame 550 can be associated with one or more frame types. Examples of frame types include: an intra frame or I-frame, a scene change frame (which can be an intra frame), a P-frame, a B-frame, etc. Frame type determination 502 can determine that current frame 550 is an intra frame. Frame type determination 502 can determine that current frame 550 is a scene change frame. Frame type determination 502 can determine that current frame 550 is either an intra frame or a scene change frame. Frame type determination 502 can determine that current frame 550 is neither an intra frame nor a scene change frame. For example, frame type determination 502 can determine that current frame 550 is a P-frame. In another example, frame type determination 502 can determine that current frame 550 is a B-frame.

[0054] In-loop filter decision portion 440 can include frame classification 504. Frame classification 504 can receive a current frame 550 to be encoded and encoding information associated with current frame 550. Frame classification 504 can receive QP 556. In some embodiments, frame classification 504 can be optional. In some embodiments, frame classification 504 can classify current frame 550 and determine a frame classification of current frame 550. Current frame 550 can be classified to determine which of several possible frame classifications current frame 550 belongs to.

[0055] Frame classification 504 can determine whether current frame 550 belongs to a certain frame classification based on one or more of spatial variation (amount of spatial variation) and temporal variation (amount of temporal variation). Frame classification 504 can analyze current frame 550 to determine one or more of spatial variation and temporal variation. In some cases, frame classification 504 can determine one or more of spatial variation and temporal variation from the encoding information associated with current frame 550. In some examples, possible frame classifications can include: (1) static or very small movement, and (2) high motion with strong edges. In some examples, possible frame classifications can include: (1) static or very small movement, (2) high motion with strong edges, (3) other frames that do not belong to classification (1) and classification (2).

[0056] Frame classification 504 can determine whether the current frame 550 belongs to a certain frame classification based on the resolution of the current frame 550. Frame classification 504 can compare the resolution of the current frame 550 to one or more thresholds, or determine whether the resolution of the current frame 550 falls within one or more ranges. In some examples, possible frame classifications can include: (1) relatively high resolution or resolution above a first resolution threshold, and (2) relatively low resolution or resolution below a first resolution threshold. In some examples, possible frame classifications can include: (1) resolution within a first range, and (2) resolution within a second range. In some examples, possible frame classifications can include: (1) relatively high resolution or resolution above a first resolution threshold, (2) medium resolution or resolution below the first resolution threshold and above a second resolution threshold, (3) relatively low resolution or resolution below the second resolution threshold. The first resolution threshold is greater than the second resolution threshold. In some examples, possible frame classifications can include: (1) resolution within a first range, (2) resolution within a second range, (3) resolution within a third range.

[0057] In-loop filter decision portion 440 can include frame classification dependent QP threshold setting 508. Frame classification dependent QP threshold setting 508 can set the QP threshold to a value corresponding to the frame classification determined in frame classification 504. The QP 556 of the current frame can be compared to the frame classification dependent QP threshold set in frame classification dependent QP threshold setting 508.

[0058] In embodiments where frame classification 504 can be omitted, frame classification dependent QP threshold setting 508 can set the QP threshold to a predetermined value.

[0059] In some embodiments, frame classification dependent QP threshold setting 508 can set the QP threshold higher when the frame classification indicates that more bits are expected to be used to encode content in the current frame 550. Frame classification dependent QP threshold setting 508 can set the QP threshold lower when the frame classification indicates that fewer bits are expected to be used to encode content in the current frame 550. Setting the frame classification dependent QP threshold based on the frame classification can ensure that the QP threshold can adapt to the content of the current frame 550.

[0060] One or more of frame type determination 502, frame classification 504, and frame classification dependent QP threshold setting 508 can be used to extract one or more heuristics about the current frame 550, which can influence how one or more in-loop filter decisions 460 will be made. The heuristics can indicate scenarios in which in-loop filter decisions for SAO filter 406 and use of SAO filter 406 should be skipped. The heuristics can indicate scenarios in which SAO usage bits can occupy too much of the bit rate budget to be used for encoding of the current frame 550.

[0061] The in-loop filter decision portion 440 can include an RDO lambda adjustment for SAO decision 506. In scenarios where the SAO RDO decision portion 520 is not skipped, the RDO lambda adjustment for SAO decision 506 can adjust the RDO parameter λ used in the rate-distortion cost calculation. As shown in Equation 1, the RDO parameter λ is a multiplier for the bit rate in the rate-distortion cost calculation. The RDO parameter λ can be preset according to the QP 556. The RDO lambda adjustment for SAO decision 506 can increase the preset value of the RDO parameter λ by some amount. The some amount can be increased using a multiplier or a scaling factor. The multiplier or scaling factor can be greater than 1. Increasing the preset value of the RDO parameter λ by some amount can include multiplying the preset value of the RDO parameter λ by the multiplier or scaling factor. In some cases, the RDO lambda adjustment for SAO decision 506 can decrease the preset value of the RDO parameter λ by some amount. The some amount can be predetermined. The RDO lambda adjustment for SAO decision 506 can determine the some amount based on the encoding information associated with the current frame 550. The RDO lambda adjustment for SAO decision 506 can determine the some amount based on the QP 556. The RDO lambda adjustment for SAO decision 506 can determine the some amount based on the frame classification determined in the frame classification 504. The RDO lambda adjustment for SAO decision 506 can determine the some amount based on the frame type determined in the frame type determination 502. The RDO lambda adjustment for SAO decision 506 can adjust the RDO parameter λ using a scaling factor S1 in response to determining that the frame type is an intra frame or a scene change frame. The RDO lambda adjustment for SAO decision 506 can adjust the RDO parameter λ using a scaling factor S2 in response to determining that the frame type is neither an intra frame nor a scene change frame (the frame type is an inter frame, such as a P frame or a B frame) and the QP parameter is less than or equal to a QP parameter threshold. The amount can have the following relationship: S2 > S1 > 1.

[0062] The in-loop filter decision portion 440 can include a SAO RDO decision portion 520. The SAO RDO decision portion 520 can evaluate options and / or parameter values and make a SAO RDO decision according to the rate-distortion cost function of Equation 1. The SAO RDO decision can specify whether to enable SAO filtering. The SAO RDO decision can specify one or more parameter values associated with a SAO filter to be applied. The SAO RDO decision portion 520 can evaluate various options and / or parameter values for a SAO filter according to the rate-distortion cost function with the RDO parameter λ. The SAO RDO decision portion 520 can select the option and / or parameter value(s) that result in the lowest rate-distortion cost and output the selected option and / or parameter value(s) as the SAO RDO decision in the one or more in-loop filter decisions 460.

[0063] The in-loop filter decision portion 440 can include an ALF RDO decision portion 530. The ALF RDO decision portion 530 can evaluate options and / or parameter values and make an ALF RDO decision according to the rate-distortion cost function of Equation 1. The ALF RDO decision can specify whether to enable ALF. The ALF RDO decision can specify one or more parameter values associated with an ALF to be applied. The ALF RDO decision portion 530 can evaluate various options and / or parameter values for an ALF filter according to the rate-distortion cost function with the RDO parameter λ. The ALF RDO decision portion 530 can select the option and / or parameter value(s) that result in the lowest rate-distortion cost and output the selected option and / or parameter value(s) as the ALF RDO decision in the one or more in-loop filter decisions 460. Exemplary process and method of making in-loop filtering decisions

[0064] FIG. 6 An in-loop filter decision process 600 according to some embodiments of the disclosure is shown. The in-loop filter decision process 600 can be implemented by one or more portions or components shown in the in-loop filter decision portion 440 of the video encoder 700 of FIG. 4 The in-loop filter decision process 600 can be encoded as instructions on the memory 804 that are executable by the processing device 802 of the computing device 800. FIG. 8 The in-loop filter decision process 600 can be encoded as instructions on the memory 804 that are executable by the processing device 802 of the computing device 800.

[0065] At 602, the frame type of the current frame 550 can be checked. It can be determined whether the current frame 550 is an intra frame or a scene change frame. In response to determining that the current frame 550 is an intra frame or a scene change frame, the process 600 proceeds from 602 to 612 via the "yes" path. Proceeding via the "yes" path can mean that SAO filtering is enabled at the picture level, and different options and / or parameter value(s) can be evaluated for the CTBs of the picture. In response to determining that the current frame 550 is neither an intra frame nor a scene change frame, the process 600 proceeds from 602 to 604 via the "no" path. Proceeding via the "no" path means that one or more other heuristics are evaluated for the current frame 550.

[0066] At 612, the RDO parameter λ can be adjusted (e.g., increased). Adjusting the RDO parameter λ can encourage the use of SAO filtering when the SAO filtering would result in little increase in syntax bits.

[0067] At 614, an SAO RDO decision can be made using the adjusted RDO parameter λ from 612.

[0068] At 604, the current frame 550 can be classified, or a frame classification can be determined for the current frame 550. In some cases, 604 can be omitted.

[0069] At 606, a QP threshold can be determined based on the frame classification determined at 604. In the case that 604 is omitted, the QP threshold can be set based on a predetermined value. In some embodiments, different frame classifications can have associated thresholds. The QP threshold can be adapted to the frame classification determined at 604.

[0070] In some embodiments, the possible frame classifications can include: (1) static or very small motion, (2) high motion with strong edges, and (3) other frames that do not belong to frame classification (1) or frame classification (2). The QP threshold for frame classification (1) can be set to a value Tl. Examples of frames belonging to frame classification (1) include video conference clips. Examples of frames belonging to frame classification (2) include live event videos (e.g., sports, concerts, news, etc.). The QP threshold for frame classification (2) can be set to a value T2. The QP threshold for frame classification (3) can be set to a value T3. These QP thresholds can have the following relationship: T2 > T3 > Tl.

[0071] In some embodiments, the possible frame classifications can include: (1) relatively high resolution or resolution above a first resolution threshold, and (2) relatively low resolution or resolution below the first resolution threshold. The QP threshold for frame classification (1) can be set to a value Tl. The QP threshold for frame classification (2) can be set to a value T2. These QP thresholds can have the following relationship: Tl > T2.

[0072] At 608, the QP 556 can be compared to the QP threshold determined at 606. It can be determined whether the QP 556 is greater than the QP threshold. It can be determined whether the QP 556 is less than or equal to the QP threshold. In response to determining that the QP 556 is greater than the QP threshold, the process 600 proceeds from 608 to 610 via the "NO" path. Proceeding via the "NO" path can mean that SAO filtering is skipped or disabled at the picture level, and different options and / or parameter value(s) for SAO filtering are not evaluated for the CTBs of the picture. SAO usage bits are not used. In response to determining that the QP 556 is less than or equal to the QP threshold, the process 600 proceeds from 608 to 612 via the "YES" path. Proceeding via the "YES" path means that SAO filtering is applied at the picture level, and different options and / or parameter value(s) for SAO filtering are evaluated for the CTBs of the picture. Proceeding via the "YES" path means that the RDO parameter λ can be adjusted for SAO RDO decision.

[0073] When the QP 556 is greater than the QP threshold (corresponding to the "NO" path out of 608), this means that the goal of the encoder is to encode the content of the current frame 550 using relatively fewer bits, and the SAO usage bits can occupy a substantial portion of the number of bits to be used for encoding the current frame 550. Therefore, it can be desirable to skip SAO filtering and not waste the SAO usage bits.

[0074] When the QP 556 is less than or equal to the QP threshold (corresponding to the "YES" path out of 608), this means that the goal of the encoder is to encode the content of the current frame 550 using relatively more bits, and the SAO usage bits can occupy a small or relatively small portion of the number of bits to be used for encoding the current frame 550. Therefore, it can be desirable to apply SAO filtering and bring the SAO usage bits.

[0075] The value of the QP threshold can affect whether the "YES" path out of 608 or the "NO" path out of 612 is taken. The value of the QP threshold can determine whether the SAO filter is skipped. Different types of content in the current frame 550 can result in more or fewer bits being used to encode the content of the current frame 550. Therefore, it can be desirable to adapt the QP threshold to the content of the current frame 550. When the content in the current frame 550 is complex spatially / temporally and / or has high resolution, more bits can be needed to encode the content of the current frame 550, and the QP threshold can be set higher. When the content in the current frame 550 is not complex spatially / temporally and / or has low resolution, fewer bits can be needed to encode the content of the current frame 550, and the QP threshold can be set lower.

[0076] When the QP threshold is higher (making it easier to find QPs 556 less than or equal to the QP threshold), SAO filtering is encouraged to be enabled. When the QP threshold is lower (making it easier to find QPs 556 greater than the QP threshold), SAO filtering is encouraged to be skipped.

[0077] At 610, an ALF RDO decision can be made using suitable RDO parameters. In some cases, if the ALF RDO decision is to use or enable ALF, the ALF RDO decision can cause the SAO RDO decision to enable picture-level SAO filtering.

[0078] FIG. 7 A method 700 for determining one or more in-loop filter decisions is shown in accordance with some embodiments of the present disclosure. The method 700 can be implemented by one or more portions or components of the in-loop filter decision portion 440 of the video encoder 400. FIG. 4 The method 700 can be encoded as instructions on the memory 804 that can be executed by the processing device 802 of the computing device 800. FIG. 8 The method 700 can be encoded as instructions on the memory 804 that can be executed by the processing device 802 of the computing device 800.

[0079] At 702, the in-loop filter decision portion 440 can determine whether the frame is an intra-frame or a scene change frame.

[0080] At 704, in response to determining that the frame is an intra-frame or a scene change frame, the in-loop filter decision portion 440 can adjust the RDO parameters by a first amount. The in-loop filter decision portion 440 can use the adjusted RDO parameters to make an SAO RDO decision for the frame.

[0081] At 706, in response to determining that the frame is neither an intra-frame nor a scene change frame, the in-loop filter decision portion 440 can determine a QP threshold. The in-loop filter decision portion 440 can determine whether a QP specified for the frame is greater than the QP threshold.

[0082] At 708, in response to determining that the QP is greater than the QP threshold, the SAO RDO decision is skipped for the frame.

[0083] At 710, in response to determining that the frame is an intra-frame or a scene change frame, the in-loop filter decision portion 440 can make an ALF RDO decision for the frame.

[0084] At 710, in response to determining that the QP is greater than the QP threshold, the in-loop filter decision portion 440 can make an ALF RDO decision for the frame.

[0085] In 704, in response to determining that the QP is less than or equal to the QP threshold, the in-loop filter decision portion 440 can adjust the RDO parameter by a second amount. The second amount can be greater than the first amount. The in-loop filter decision portion 440 can make an SAO RDO decision for the frame using the adjusted RDO parameter. In 710, the in-loop filter decision portion 440 can make an ALF RDO decision for the frame. Exemplary computing device

[0086] FIG. 8 is a block diagram of a device or system (e.g., exemplary computing device 800) in accordance with some embodiments of the present disclosure. One or more computing devices 800 can be used to implement the functionality described herein with respect to the accompanying FIG. 1 drawings. The various components illustrated in the figures can be included in the computing device 800, but any one or more of these components can be omitted or replicated in order to fit application. In some embodiments, some or all of the components included in the computing device 800 can be attached to one or more motherboards. In some embodiments, some or all of the components are fabricated onto a single system on a chip (SoC) die. In addition, in various embodiments, the computing device 800 can not include one or more of the components illustrated in the figures, and the computing device 800 can include interface circuitry for coupling to one or more components. For example, the computing device 800 can not include a display device 806, and can instead include display device interface circuitry (e.g., a connector and driver circuitry) to which a display device 806 can be coupled. In another set of examples, the computing device 800 can not include an audio input device 818 or an audio output device 808, and can instead include audio input or output device interface circuitry (e.g., connectors and support circuitry) to which an audio input device 818 or an audio output device 808 can be coupled. FIG. 8

[0087] The computing device 800 can include a processing device 802 (e.g., one or more processing devices, one or more same type of processing devices, one or more different type of processing devices). The processing device 802 can include processing circuitry or electronic circuitry that processes electronic data from data storage elements (e.g., registers, memory, resistors, capacitors, qubit units) to transform that electronic data into other electronic data that can be stored in registers and / or memory. Examples of the processing device 802 can include a CPU, a GPU, a quantum processor, a machine learning processor, an artificial intelligence processor, a neural network processor, an artificial intelligence accelerator, an application specific integrated circuit (ASIC), an analog signal processor, an analog computer, a microprocessor, a digital signal processor, a field programmable gate array (FPGA), a tensor processing unit (TPU), a data processing unit (DPU), etc.​

[0088] The computing device 800 may include a memory 804, which may itself include one or more memory devices, such as volatile memory (e.g., DRAM), non-volatile memory (e.g., read-only memory (ROM)), high-bandwidth memory (HBM), flash memory, solid-state memory, and / or a hard drive. The memory 804 includes one or more non-transitory computer-readable storage media. In some embodiments, the memory 804 may include a memory that shares a die with the processing device 802.

[0089] In some embodiments, memory 804 includes one or more non-transitory computer-readable media storing instructions executable to perform the operations described herein, such as FIG. 1 to FIG. 7 , the in-loop filter decision process 600, and the method 700 shown in FIG. 1 . In some embodiments, the memory 804 includes one or more non-transitory computer-readable media storing instructions that are executable to perform one or more operations of the encoder 102. In some embodiments, the memory 804 includes one or more non-transitory computer-readable media storing instructions that are executable to perform one or more operations of the in-loop filter decision portion 440. In some embodiments, the memory 804 includes one or more non-transitory computer-readable media storing instructions that are executable to perform one or more operations of the in-loop filter 228. The instructions stored in the memory 804 may be executed by the processing device 802.

[0090] In some embodiments, the memory 804 can store data, such as data structures, binary data, bits, metadata, files, binary large objects (blobs), etc., as described herein with reference to the accompanying FIG. 1 The memory 804 may include one or more non-transitory computer-readable media that store one or more of the following: input frames to the encoder (e.g., video frames 104), intermediate data structures calculated by the encoder, bitstreams generated by the encoder (encoded bitstream 180), bitstreams received by the decoder (encoded bitstream 180), intermediate data structures calculated by the decoder, and reconstructed frames generated by the decoder. The memory 804 may include one or more non-transitory computer-readable media that store one or more of the following: FIG. 6 The memory 804 may include one or more non-transitory computer-readable media that store one or more of the following: FIG. 7 The memory 804 may include the decoded picture buffer 232.

[0091] In some embodiments, computing device 800 can include a communication device 812 (e.g., one or more communication devices). For example, communication device 812 can be configured for managing wired and / or wireless communications in order to transmit data to and from computing device 800. The term "wireless" and its derivatives can be used to describe circuits, devices, systems, methods, techniques, communications channels, and / or the like that can communicate data through the use of modulated electromagnetic radiation through a non-solid medium. This can not involve the use of any wires, although wireless can, in some embodiments, include the use of wires for the reception and / or transmission of data. Communication device 812 can implement any of a number of wireless standards or protocols, including but not limited to: Bluetooth®, a wireless personal area network (PAN) based on the IEEE 802.15 standard; a wireless display (WiDi) standard; an Institute of Electrical and Electronics Engineers (IEEE) standard, including the IEEE 802.10 family of standards such as the IEEE 802.10e standard and / or the IEEE 802.1 la standard; an IEEE 802.16 standard (e.g., the IEEE 802.16-2005 amendment); a Long Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., the LTE-Advanced project, the Ultra Mobile Broadband (UMB) project (also referred to as "3GPP2"), and / or the like). IEEE 802.16-compatible Broadband Wireless Access (BWA) networks are generally referred to as WiMAX networks (an acronym for Worldwide Microwave Access), which is an authentication mark for products that pass conformance and interoperability testing for the IEEE 802.16 standards. Communication device 812 can operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunication System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. Communication device 812 can operate in accordance with Enhanced Data for GSM Evolution (EDGE), a GSM EDGE Radio Access Network (GERAN), a Universal Terrestrial Radio Access Network (UTRAN), or an Evolved UTRAN (E-UTRAN). Communication device 812 can operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO) and its derivatives, and any other wireless protocol named 3G, 4G, 5G, and beyond. In other embodiments, communication device 812 can operate in accordance with other wireless protocols. Computing device 800 can include an antenna 822 to facilitate wireless communication and / or to receive other wireless communications (such as radio frequency transmissions). Computing device 800 can include receiver circuitry and / or transmitter circuitry. In some embodiments, communication device 812 can manage wired communications, such as electrical, optical, or any other suitable communication protocol (e.g., Ethernet). As described above, communication device 812 can include multiple communication chips.For example, a first communication device 812 can be dedicated to short-range wireless communications, such as Wi-Fi or Bluetooth, while a second communication device 812 can be dedicated to long-range wireless communications, such as Global Positioning System (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, and so on. In some embodiments, a first communication device 812 can be dedicated to wireless communications, while a second communication device 812 can be dedicated to wired communications.

[0092] The computing device 800 can include a power supply / power circuitry 814. The power supply / power circuitry 814 can include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry to couple components of the computing device 800 to an energy source separate from the computing device 800 (e.g., DC power, AC power, etc.).

[0093] The computing device 800 can include a display device 806 (or corresponding interface circuitry, as discussed above). For example, the display device 806 can include any visual indicators, such as a heads-up display, computer monitor, projector, touchscreen display, liquid crystal display (LCD), light-emitting diode display, or flat panel display.

[0094] The computing device 800 can include an audio output device 808 (or corresponding interface circuitry, as discussed above). For example, the audio output device 808 can include any device that generates an audible indicator, such as speakers, headphones, or earbuds.

[0095] The computing device 800 can include an audio input device 818 (or corresponding interface circuitry, as discussed above). The audio input device 818 can include any device that generates a signal representative of a sound, such as a microphone, microphone array, or digital instrument (e.g., an instrument with a musical instrument digital interface (MIDI) output).

[0096] The computing device 800 can include a GPS device 816 (or corresponding interface circuitry, as discussed above). The GPS device 816 can communicate with satellite-based systems, and can receive a location of the computing device 800, as is known in the art.

[0097] The computing device 800 can include a sensor 830 (or one or more sensors). The computing device 800 can include a corresponding interface circuitry, as described above. The sensor 830 can sense a physical phenomenon and convert the physical phenomenon into an electrical signal that can be processed, for example, by the processing device 802. Examples of the sensor 830 can include a capacitive sensor, an inductive sensor, a resistive sensor, an electromagnetic field sensor, a light sensor, a camera, an imager, a microphone, a pressure sensor, a temperature sensor, a vibration sensor, an accelerometer, a gyroscope, a strain sensor, a moisture sensor, a humidity sensor, a distance sensor, a range sensor, a time-of-flight sensor, a pH sensor, a particulate matter sensor, an air quality sensor, a chemical sensor, a gas sensor, a biological sensor, an ultrasonic sensor, a scanner, or the like.

[0098] The computing device 800 can include other output devices 810 (or corresponding interface circuitry, as described above). Examples of the other output devices 810 can include an audio codec, a video codec, a printer, a wired or wireless transmitter to provide information to other devices, a haptic output device, a gas output device, a vibration output device, an illumination output device, a home automation controller, or an additional storage device.

[0099] The computing device 800 can include other input devices 820 (or corresponding interface circuitry, as described above). Examples of the other input devices 820 can include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a barcode reader, a quick response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.

[0100] The computing device 800 can have any desired form factor, such as a handheld or mobile computer system (e.g., a cellular telephone, a smartphone, a mobile Internet device, a music player, a tablet computer, a laptop computer, a netbook computer, a personal digital assistant (PDA), an ultra-mobile personal computer, a remote control, a wearable device, a headset, eyewear, footwear, an electronic garment, or the like), a desktop computer system, a server or other networked computing component, a printer, a scanner, a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital camera, a digital video recorder, an Internet of Things device, or a wearable computer system. In some embodiments, the computing device 800 can be any other electronic device that processes data. Selection example

[0101] Example 1 provides a method comprising determining whether a frame is an intra frame or a scene change frame; in response to determining that the frame is an intra frame or a scene change frame, adjusting a rate-distortion optimization (RDO) parameter by a first amount and making a sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter; and in response to determining that the frame is neither an intra frame nor a scene change frame: determining a quantization parameter (QP) threshold; determining whether a QP specified for the frame is greater than the QP threshold; and in response to determining that the QP is greater than the QP threshold, skipping a SAO RDO decision for the frame.

[0102] Example 2 provides the method of example 1, further comprising: in response to determining that the frame is an intra frame or a scene change frame, making an adaptive loop filter (ALF) RDO decision for the frame.

[0103] Example 3 provides the method of example 1 or 2, wherein: the RDO parameter is a multiplier for a bit rate; and adjusting the RDO parameter comprises increasing a value of the RDO parameter by the first amount.

[0104] Example 4 provides the method of any of examples 1-3, further comprising: in response to determining that the QP is greater than the QP threshold, making an ALF RDO decision for the frame.

[0105] Example 5 provides the method of any of examples 1-4, further comprising: in response to determining that the QP is less than or equal to the QP threshold, adjusting the RDO parameter by a second amount, making a sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter, and making an ALF RDO decision for the frame.

[0106] Example 6 provides the method of example 5, wherein the second amount is greater than the first amount.

[0107] Example 7 provides the method of any of examples 1-6, wherein determining the QP threshold comprises setting the QP threshold to a predetermined value.

[0108] Example 8 provides the method of any of examples 1-7, wherein determining the QP threshold comprises: determining a frame classification for the frame based on one or more of a spatial variation and a temporal variation; and setting the QP threshold to a value corresponding to the frame classification.

[0109] Example 9 provides the method of any of examples 1-8, wherein determining the QP threshold comprises: determining a frame classification for the frame based on a resolution of the frame; and setting the QP threshold to a value corresponding to the frame classification.

[0110] Example 10 provides one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to: determine whether a frame is an intra frame or a scene change frame; in response to determining that the frame is an intra frame or a scene change frame, adjust a rate-distortion optimization (RDO) parameter by a first amount and make a sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter; and in response to determining that the frame is neither an intra frame nor a scene change frame: determine a quantization parameter (QP) threshold; determine whether a QP specified for the frame is greater than the QP threshold; and in response to determining that the QP is greater than the QP threshold, skip a SAO RDO decision for the frame.

[0111] A variation of Example 10 provides one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to: determine whether a frame is an intra frame or a scene change frame; in response to determining that the frame is an intra frame or a scene change frame, adjust a rate-distortion optimization (RDO) parameter by a first amount and make a sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter, wherein the RDO parameter is a multiplier for a bit rate; and in response to determining that the frame is neither an intra frame nor a scene change frame: determine a quantization parameter (QP) threshold; determine whether a QP specified for the frame is greater than the QP threshold; and in response to determining that the QP is greater than the QP threshold, skip a SAO RDO decision for the frame.

[0112] Example 11 provides the one or more non-transitory computer-readable media of Example 10, wherein the instructions further cause the one or more processors to: in response to determining that the frame is an intra frame or a scene change frame, make an adaptive loop filter (ALF) RDO decision for the frame.

[0113] Example 12 provides the one or more non-transitory computer-readable media of Example 10 or 11, wherein: the RDO parameter is a multiplier for a bit rate; and adjusting the RDO parameter comprises increasing a value of the RDO parameter by the first amount.

[0114] Example 13 provides the one or more non-transitory computer-readable media of any of Examples 10-12, wherein the instructions further cause the one or more processors to: in response to determining that the QP is greater than the QP threshold, make an adaptive loop filter (ALF) RDO decision for the frame.

[0115] Example 14 provides one or more non-transitory computer-readable media of any of Examples 10-13, wherein the instructions further cause the one or more processors to: in response to determining that the QP is less than or equal to the QP threshold, adjust the RDO parameter by a second amount, make a sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter, and make an ALF RDO decision for the frame.

[0116] Example 15 provides one or more non-transitory computer-readable media of Example 14, wherein the second amount is greater than the first amount.

[0117] Example 16 provides one or more non-transitory computer-readable media of any of Examples 10-15, wherein determining the QP threshold comprises setting the QP threshold to a predetermined value.

[0118] Example 17 provides one or more non-transitory computer-readable media of any of Examples 10-16, wherein determining the QP threshold comprises: determining a frame classification of the frame based on one or more of a spatial variation and a temporal variation; and setting the QP threshold to a value corresponding to the frame classification.

[0119] Example 18 provides one or more non-transitory computer-readable media of any of Examples 10-17, wherein determining the QP threshold comprises: determining a frame classification of the frame based on a resolution of the frame; and setting the QP threshold to a value corresponding to the frame classification.

[0120] Example 19 provides a system comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to: determine whether a frame is an intra-frame or a scene change frame; in response to determining that the frame is an intra-frame or a scene change frame, adjust a rate-distortion optimization (RDO) parameter by a first amount and make a sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter; and in response to determining that the frame is neither an intra-frame nor a scene change frame: determine a quantization parameter (QP) threshold; determine whether a QP specified for the frame is greater than the QP threshold; and in response to determining that the QP is greater than the QP threshold, skip the SAO RDO decision for the frame.

[0121] A variation of Example 19 provides a system comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to: determine whether a frame is an intra-frame or a scene change frame; in response to determining that the frame is an intra-frame or a scene change frame, increase a rate-distortion optimization (RDO) parameter by a first amount and make a sample adaptive offset (SAO) RDO decision for the frame using the increased RDO parameter; and in response to determining that the frame is neither an intra-frame nor a scene change frame: determine a quantization parameter (QP) threshold; determine whether a QP specified for the frame is greater than the QP threshold; and in response to determining that the QP is greater than the QP threshold, skip a SAO RDO decision for the frame.

[0122] Example 20 provides the system of Example 19, wherein the instructions further cause the one or more processors to: in response to determining that the frame is an intra-frame or a scene change frame, make an adaptive loop filter (ALF) RDO decision for the frame.

[0123] Example 21 provides the system of Example 19 or 20, wherein: the RDO parameter is a multiplier for a bit rate; and adjusting the RDO parameter comprises increasing a value of the RDO parameter by the first amount.

[0124] Example 22 provides the system of any of Examples 19-21, wherein the instructions further cause the one or more processors to: in response to determining that the QP is greater than the QP threshold, make an ALF RDO decision for the frame.

[0125] Example 23 provides the system of any of Examples 19-22, wherein the instructions further cause the one or more processors to: in response to determining that the QP is less than or equal to the QP threshold, adjust the RDO parameter by a second amount, make a sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter, and make an ALF RDO decision for the frame.

[0126] Example 24 provides the system of Example 23, wherein the second amount is greater than the first amount.

[0127] Example 25 provides the system of any of Examples 19-24, wherein determining the QP threshold comprises: setting the QP threshold to a predetermined value.

[0128] Example 26 provides the system of any of Examples 19-25, wherein determining the QP threshold comprises: determining a frame classification for the frame based on one or more of a spatial variation and a temporal variation; and setting the QP threshold to a value corresponding to the frame classification.

[0129] Example 27 provides the system of any of Examples 19-26, wherein determining the QP threshold comprises: determining a frame classification of the frame based on a resolution of the frame; and setting the QP threshold to a value corresponding to the frame classification.

[0130] Example A provides a device comprising means to perform or for performing any of the methods provided in Examples 1-9 and the methods / processes described herein.

[0131] Example B provides one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods provided in Examples 1-9 and the methods / processes described herein.

[0132] Example C provides a device comprising: one or more processors to execute instructions; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods provided in Examples 1-9 and the methods / processes described herein.

[0133] Example D provides an encoder to generate an encoded bitstream using the operations described herein.

[0134] Example E provides an encoder to perform any of the methods provided in Examples 1-9 and the methods / processes described herein.

[0135] Example F provides an in-loop filtering decision portion 440 as described herein. Variations and other notes

[0136] Although the operations of the example methods shown and described are each illustrated as occurring one time in a specific order, it will be recognized that some operations can be executed concurrently, in a different order, or with various degrees of repetition. FIG. 6 to FIG. 7 The operations shown and described in the example methods shown and described can be combined, or can include more or fewer details than described. FIG. 6 to FIG. 7

[0137] The above description of the illustrated implementations of the present disclosure, including what is described in the abstract, is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. While specific implementations of, and examples for, the present disclosure are described herein, the scope of the present disclosure is not limited to these described specific implementations or examples. Many modifications and variations will be apparent to those of ordinary skill in the art, having the benefit of the teachings described in this detailed description. The functions or operations of the present disclosure can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media include computer-readable storage media. A computer-readable storage medium can be any available medium or ​

[0138] For purposes of explanation, specific numbers, materials, and configurations are set forth in order to provide a thorough understanding of the illustrative implementations. However, a person having ordinary skill in the art will appreciate that the disclosure can be practiced without the specific details set forth herein, and that multiple implementations can be made with only some of the described aspects. In other instances, well-known features are omitted or simplified in order not to obscure the illustrative implementations.

[0139] Furthermore, reference is made to the accompanying drawings, which form a part of this disclosure, and in which are shown, by way of illustration, implementations that can be practiced. It is to be understood that other implementations can be utilized, and that structural or logical changes can be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense.

[0140] Various operations can be described as multiple discrete actions or operations in turn, in a manner that can most closely correspond to the drawings and description of the disclosure. However, the order of description is not to be construed as limiting, and any number of the described operations can be reordered, added, or omitted, according to alternative implementations. Further, descriptions and depictions of various

[0141] For purposes of the present disclosure, the phrase "A or B" or the phrase "A and / or B" means (A), (B), or (A and B). For purposes of the present disclosure, the phrase "A, B, or C" or the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B and C). The term "between" when used in reference to a range of measurements includes the end points of the range of measurements.

[0142] For purposes of the present disclosure, "A is less than or equal to a first threshold" is equivalent to "A is less than a second threshold," provided that the first threshold and the second threshold are set in such a way that, for any value of A, both statements result in the same logical outcome. For purposes of the present disclosure, "B is greater than a first threshold" is equivalent to "B is greater than or equal to a second threshold," provided that the first threshold and the second threshold are set in such a way that, for any value of B, both statements result in the same logical outcome.

[0143] The description uses the phrases "in an embodiment," or "in embodiments," which can each refer to one or more embodiments among the same or different embodiments. The terms "comprise," "comprising," "have," "having," "include," "including," and "contain," "containing," or variants thereof, as used in the disclosure, are synonymous with each other. The disclosure can use perspective-based descriptions such as "above," "below," "top," "bottom," and "side" that can be used to explain the various features illustrated in the figures. However, these terms are used only to facilitate discussion, and not to imply or require any specific orientation. The drawings are not necessarily to scale. Unless otherwise specified, the use of ordinal adjectives such as "first," "second," and "third," etc., to describe a common object merely indicate different instances of similar objects, and do not imply that the objects are in any given order, either temporally, spatially, in ranking, or in any other manner.

[0144] In the following detailed description, various aspects of the illustrative implementations will be described using terminology generally understood by those skilled in the art of the field to which the implementations pertains.

[0145] As described herein or as known in the art, the terms "substantially," "near," "approximately," "close," and "about" generally mean within + / - 20% of a target value. Similarly, as described herein or as known in the art, terms indicating the orientation of various elements, such as "co-planar," "perpendicular," "orthogonal," "parallel," or any other angle between elements, generally mean within + / - 5% to 20% of a target value.

[0146] In addition, the terms "comprise," "comprising," "include," "including," "have," "having," or any other variation thereof, are intended to cover non-exclusive inclusion. For example, a process, process, or apparatus that comprises a list of elements is not necessarily limited to only those elements but can include other elements not expressly listed or inherent to such process, process, or apparatus. Also, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or."

[0147] The systems, methods, and apparatuses of the present disclosure each have several innovative aspects, none of which is, by itself, solely responsible for the overall desirable attributes of the implementations disclosed herein. The details of one or more implementations of the subject matter described in this specification are set forth in the description and the drawings.

Claims

1. A method comprising: determining whether a frame is an intra frame or a scene change frame; in response to determining that the frame is an intra frame or a scene change frame, adjusting a rate-distortion optimization (RDO) parameter by a first amount and making a sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter; and in response to determining that the frame is neither an intra frame nor a scene change frame: determining a quantization parameter (QP) threshold; determining whether a QP specified for the frame is greater than the QP threshold; and in response to determining that the QP is greater than the QP threshold, skipping the SAO RDO decision for the frame.

2. The method of claim 1, further comprising: in response to determining that the frame is an intra frame or a scene change frame, making an adaptive loop filter (ALF) RDO decision for the frame.

3. The method of claim 1 or 2, wherein: the RDO parameter is a multiplier for a bit rate; and adjusting the RDO parameter comprises increasing a value of the RDO parameter by the first amount.

4. The method of claim 1 or 2, further comprising: in response to determining that the QP is greater than the QP threshold, making an ALF RDO decision for the frame.

5. The method of any of claims 1-4, further comprising: in response to determining that the QP is less than or equal to the QP threshold, adjusting the RDO parameter by a second amount, making the sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter, and making an ALF RDO decision for the frame. the second amount is greater than the first amount.

6. The method of claim 5, wherein, determining the QP threshold comprises:

7. The method of claim 1 or 2, wherein, setting the QP threshold to a predetermined value. determining the QP threshold comprises:

8. The method of claim 1 or 2, wherein, determining a frame classification of the frame based on one or more of a spatial variation and a temporal variation; and setting the QP threshold to a value corresponding to the frame classification. determining the QP threshold comprises:

9. The method of claim 1 or 2, wherein, determining a frame classification of the frame based on a resolution of the frame; and setting the QP threshold to a value corresponding to the frame classification.

10. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to: determine whether a frame is an intra frame or a scene change frame; and in response to determining that the frame is neither an intra frame nor a scene change frame: in response to determining that the frame is an intra frame or a scene change frame, adjusting a rate-distortion optimization (RDO) parameter by a first amount, and using the adjusted RDO parameter to make a sample adaptive offset (SAO) RDO decision for the frame; determine a quantization parameter (QP) threshold; determine whether a QP specified for the frame is greater than the QP threshold; and in response to determining that the QP is greater than the QP threshold, skip the SAO RDO decision for the frame. The instructions further cause the one or more processors to: in response to determining that the frame is an intra frame or a scene change frame, make an adaptive loop filter (ALF) RDO decision for the frame.

11. The one or more non-transitory computer-readable media of claim 10, wherein, 12. The one or more non-transitory computer-readable media of claim 10 or 11, wherein: the RDO parameter is a multiplier for a bit rate; and ​ ​ adjusting the RDO parameter includes increasing a value of the RDO parameter by the first amount.

13. The one or more non-transitory computer-readable media of claim 10 or 11, wherein, the instructions further cause the one or more processors to: in response to determining that the QP is greater than the QP threshold, make an ALF RDO decision for the frame.

14. The one or more non-transitory computer-readable media of claim 10 or 11, wherein, the instructions further cause the one or more processors to: in response to determining that the QP is less than or equal to the QP threshold, adjust the RDO parameter by a second amount, make the sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter, and make an ALF RDO decision for the frame.

15. The one or more non-transitory computer-readable media of claim 14, wherein, the second amount is greater than the first amount.

16. The one or more non-transitory computer-readable media of claim 10 or 11, wherein, determining the QP threshold includes: setting the QP threshold to a predetermined value.

17. The one or more non-transitory computer-readable media of claim 10 or 11, wherein, determining the QP threshold includes: determining a frame classification of the frame based on one or more of spatial variation and temporal variation; and setting the QP threshold to a value corresponding to the frame classification.

18. The one or more non-transitory computer-readable media of claim 10 or 11, wherein, determining the QP threshold includes: determining a frame classification of the frame based on a resolution of the frame; and setting the QP threshold to a value corresponding to the frame classification.

19. A system comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to: determine whether a frame is an intra-frame or a scene change frame; in response to determining that the frame is an intra-frame or a scene change frame, adjust a rate-distortion optimization (RDO) parameter by a first amount and make a sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter; and in response to determining that the frame is neither an intra-frame nor a scene change frame: determine a quantization parameter (QP) threshold; determine whether a QP specified for the frame is greater than the QP threshold; and in response to determining that the QP is greater than the QP threshold, skip the SAO RDO decision for the frame.

20. The system of claim 19, wherein, the instructions further cause the one or more processors to: in response to determining that the frame is an intra-frame or a scene change frame, make an adaptive loop filter (ALF) RDO decision for the frame.

21. The system of claim 19 or 20, wherein: the RDO parameter is a multiplier for a bit rate; and adjusting the RDO parameter includes increasing a value of the RDO parameter by the first amount.

22. The system of claim 19 or 20, wherein, the instructions further cause the one or more processors to: in response to determining that the QP is greater than the QP threshold, make an ALF RDO decision for the frame.

23. The system of claim 19 or 20, wherein, the instructions further cause the one or more processors to: in response to determining that the QP is less than or equal to the QP threshold, adjust the RDO parameter by a second amount, make the sample adaptive offset (SAO) RDO decision for the frame using the adjusted RDO parameter, and make an ALF RDO decision for the frame.

24. The system of claim 23, wherein, the second amount is greater than the first amount.

25. A computer program product comprising instructions which, when executed by a processor, cause the processor to carry out the method according to any one of claims 1-9.