Adaptive coding tool selection with content classification
By selecting a lightweight adaptive encoding tool based on video content classification, the system solves the problem of high encoder computational complexity, improves encoding efficiency and video quality, and adapts to different types of video content.
Patent Information
- Application Number
- CN202510303084.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2025-03-14
- Publication Date
- 2025-10-24
AI Technical Summary
Existing encoders suffer from high computational complexity and difficulty in effectively detecting and classifying video content types during the encoding process, resulting in low encoding efficiency. This is especially true when processing screen content and natural content, making it difficult to select the most suitable encoding tool.
A lightweight adaptive encoding tool selection system is adopted. By classifying video frames into content, screen content, weak screen content, and natural content, and generating encoding tool control signals based on this, specific encoding tools and configuration parameters are selectively turned on or off, reducing the computational complexity of the encoder.
It achieves improved encoding efficiency under low latency and real-time operation conditions, adapts to different types of video content, reduces the computational complexity of the encoder, and maintains video quality.
Smart Images

Figure CN120835146A_ABST
Abstract
Description
BACKGROUND
[0001] Video compression is a technology used to make video files smaller and easier to send over the internet. There are different methods and algorithms for video compression with different performance and trade-offs. Video compression involves encoding and decoding. Encoding is the process of transforming (uncompressed) video data into a compressed format. Decoding is the process of recovering video data from a compressed format. An encoder-decoder system is called a codec. BRIEF DESCRIPTION OF DRAWINGS
[0002] Embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. To facilitate this description, like reference numerals designate like structural elements in the drawings. Embodiments are illustrated by way of example in the figures of the drawing in which:
[0003] Figure 1 An encoding system and multiple decoding systems are shown in accordance with some embodiments of the present disclosure.
[0004] Figure 2 An exemplary encoder for encoding video frames and outputting an encoded bitstream is shown in accordance with some embodiments of the present disclosure.
[0005] Figure 3 An exemplary decoder for decoding an encoded bitstream and outputting a decoded video is shown in accordance with some embodiments of the present disclosure.
[0006] Figure 4 An exemplary encoder and adaptive coding tool selector are shown in accordance with some embodiments of the present disclosure.
[0007] Figure 5 An exemplary implementation of an adaptive coding tool selector is shown in accordance with some embodiments of the present disclosure.
[0008] Figure 6 An exemplary implementation of a content classifier is shown in accordance with some embodiments of the present disclosure.
[0009] Figure 7 An exemplary process for classifying content is depicted in accordance with some embodiments of the present disclosure.
[0010] Figure 8 An exemplary process for setting one or more coding tool control signals is depicted in accordance with some embodiments of the present disclosure.
[0011] Figure 9 An exemplary process for setting one or more coding tool control signals is depicted in accordance with some embodiments of the present disclosure.
[0012] Figure 10An exemplary process for setting one or more coding tool control signals is depicted in accordance with some embodiments of the present disclosure.
[0013] Figure 11 A method for adaptively controlling coding tools of an encoder based on content classification is shown in accordance with some embodiments of the present disclosure.
[0014] Figure 12 A block diagram of an exemplary computing device in accordance with some embodiments of the present disclosure is depicted. DETAILED DESCRIPTION
[0015] SUMMARY
[0016] Video encoding or video compression is the process of compressing video data for storage, transmission, and playback. Video compression can involve taking a large amount of raw video data and applying one or more compression techniques to reduce the amount of data needed to represent the video while maintaining an acceptable level of visual quality. In some cases, video compression can provide efficient storage and transmission of video content over limited bandwidth networks.
[0017] A video includes video frames or one or more (temporal) sequences of frames. A frame can include an image or a single still image. A frame can have millions of pixels. For example, a frame for uncompressed 4K video can have a resolution of 3840 x 2160 pixels. A pixel can have a luma / luminance value and a chroma / chrominance value. The terms “frame” and “picture” can be used interchangeably. There are several frame types of picture types. An I-frame or intra-frame can be the most non-compressible and does not depend on other frames for decoding. An I-frame can include a scene change frame. An I-frame can be a reference frame for one or more other frames. A P-frame can depend on data from a previous frame for decoding and can be more compressible than an I-frame. A P-frame can be a reference frame for one or more other frames. A B-frame can depend on data from a previous frame and a forward frame for decoding and can be more compressible than an I-frame and a P-frame. A B-frame can refer to two or more frames, such as one frame in the future and one frame in the past. Other frame types can include a reference B-frame and a non-reference B-frame. A reference B-frame can be used as a reference for another frame. A non-reference B-frame is not used as a reference for any frame. A reference B-frame is stored in a decoded picture buffer, while a non-reference B-frame does not need to be stored in a decoded picture buffer. A P-frame and a B-frame can be referred to as an inter-frame. An order or coding hierarchy in which I-frames, P-frames, and B-frames are arranged can be referred to as a group of pictures (GOP). In some cases, a frame can be an instantaneous decoder refresh (IDR) frame within a GOP. An IDR-frame can indicate that any frame after the IDR-frame cannot reference any frame before the IDR-frame. Thus, an IDR-frame can signal to a decoder that the decoder can clear a decoded picture buffer. Each IDR-frame can be an I-frame, but an I-frame can or can not be an IDR-frame. A closed GOP can start from an IDR-frame. A slice can be a spatially distinct area in a frame that is coded separately from any other area in the same frame.
[0018] In some cases, a frame can be partitioned into one or more blocks. The blocks can be used for block-based compression. The blocks of pixels resulting from the partitioning can be referred to as tiles. The blocks can have a much smaller size such as, for example, 512x512 pixels, 256x256 pixels, 128x128 pixels, 64x64 pixels, 32x32 pixels, 16x16 pixels, 8x8 pixels, 4x4 pixels, etc. The blocks can include square or rectangular areas of a frame. Various video compression technologies can use different terminology for blocks or different partition structures for creating blocks. In some video compression technologies, a frame can be partitioned into Coding Tree Units (CTUs). The CTUs can be split into Coding Tree Blocks (CTBs) (for luma and chroma components, respectively). The CTBs can have sizes of 64x64 pixels, 32x32 pixels, or 16x16 pixels. The CTBs can be split into Coding Units (CUs). The CUs can be split into Prediction Units (PUs) and / or discrete cosine transform (DCT) transform units (TUs). The CTUs, CTBs, CUs, PUs, and TUs can be considered as blocks or tiles herein.
[0019] Modern codecs can support multiple coding tools to improve the quality of the encoded bitstream. The coding tools can have one or more parameter values that can be used to adjust the coding tools for the data being encoded. One of the tasks of an encoder in a video codec is to make coding decisions related to the coding tools at different levels of the video (e.g., sequence-level, GOP-level, frame / picture-level, slice-level, CTU-level, CTB-level, block-level, CU-level, PU-level, TU-level, etc.) based on the desired bitrate and / or the desired (objective and / or subjective) quality. Making coding decisions can include evaluating different coding tool options or parameter values for encoding the data and determining the best coding tool options or parameter values that can achieve the desired bitrate and / or quality. The selected coding tool options and / or parameter values can be applied to encode the video to generate the bitstream. The selected coding tool options and / or parameter values will be encoded in the bitstream to signal the decoder how to decode the encoded bitstream according to the coding decisions made by the encoder. While evaluating all possible combinations of options and parameter values can yield the best coding decisions, the encoder does not have unlimited resources to undertake the complexity required to evaluate every available coding tool option and parameter value. While some codecs can achieve significant subjective quality improvement at similar bitrates compared to earlier codecs, the price of the improvement is an increase in complexity in the encoder and the decoder. It is a technical challenge to reduce the complexity in the encoder while having little impact on the quality of the video.
[0020] Among the coding tools supported by modern codecs, many are broadly applicable to different types of video content, but some may be particularly well-suited for specific types of video content. For example, motion compensated temporal filters (MCTFs) can help significantly improve encoder efficiency for natural (non-screen) content, but may result in quality loss for screen content. Intrablock copy (IBC) and palette coding in intra-frame prediction can perform well for screen content, but may result in quality loss if applied to natural (non-screen) content. Luma mapping with chroma scaling (LMCS) in loop filtering can perform well for natural content, but is not very efficient for screen content. Furthermore, the complexity associated with these coding tools is relatively high. To reduce complexity in the encoder while maintaining the quality improvements achievable by modern codecs, these coding tools can be selectively enabled or disabled, adapted, or configured in a specific manner that best suits the type of content being encoded. One of the technical challenges is analyzing the content before the encoding process and, based on this analysis, effectively and efficiently detecting or classifying the content being encoded. Another technical challenge is to implement a scheme that can control or configure an encoder based on and appropriate to the content classification.
[0021] Some neural network-based solutions can analyze video content and perform content classification and screen content detection. However, as pre-analysis operations before the encoder, neural network-based solutions can be too heavy, incur long latency, and be computationally intensive to be practical. For hardware-based encoder solutions, implementing lightweight pre-analysis operations would be effective. In one solution involving statistics based on 4x4 pixel blocks, many calculations are required to calculate the statistics, and the statistics can lead to misclassification of natural content with many homogeneous black or dark areas.
[0022] To address some of these issues, a lightweight but effective adaptive coding tool selection system with content classification can be implemented. The system can have low complexity and can be hardware friendly. In particular, content classification can detect screen content or classify one or more frames among at least three classifications: a screen content classification (indicating strong screen content), a weak screen content classification (indicating weak screen content), and a natural content classification (indicating non-screen content). The content classification can utilize two statistics, e.g., the number of colors and the variance, for blocks of size 8x8 pixels or larger (e.g., 8x8 blocks, 16x16 blocks, and 32x32 blocks). In some embodiments, the two statistics are computed for each 8x8 block of the current frame. In some embodiments, the two statistics are computed for each 16x16 block of the current frame. The statistics can be used to compute three frame-level statistics, e.g., the proportion / percentage of blocks with few colors, the proportion / percentage of blocks with zero variance, and the proportion / percentage of blocks with large / great variance. The frame-level statistics are used to determine which classification one or more frames belong to. The frame-level statistics can be used to determine whether the current frame has strong, weak, or non-screen content. Based on the classification or detection results, coding tool control flags or control signals can be generated accordingly to cause the encoding system to configure, e.g., turn on or off certain coding tools, and / or use certain parameter values for the coding tools.
[0023] By adaptively selecting and configuring the coding tools, coding efficiency improvement and complexity reduction can be achieved with a wide range of video content. The solutions described herein can be implemented as a standalone component in front of the encoder that supports the coding tools. The lightweight and hardware friendly solution can benefit encoders that operate with low latency and in real time (e.g., in data center, server, or video streaming use cases).
[0024] The techniques described and illustrated herein for adaptive coding tool selection with content classification can be applied to various codecs, such as AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), AV1 (AOMedia Video 1), and VVC (Versatile Video Coding). AVC, also known as “ITU-T H.264,” was approved in 2003 and last revised on 2021-08-22. HEVC, also known as “ITU-T H.265,” was approved in 2013 and last revised on 2023-09-13. AV1 is a video coding codec designed for video transmission over the Internet. The “AV1 Bitstream & Decoding Process Specification” version 1.1.1 with Errata was last modified in 2019. VVC, also known as “ITU-T H.266,” was finalized in 2020.
[0025] Video compression
[0026] Figure 1 The encoding system 130 and one or more decoding systems 150 according to some embodiments of the present disclosure are shown. 1…D .
[0027] The encoding system 130 can be Figure 12 The encoding system 130 may be implemented on a computing device 1200. The encoding system 130 may be implemented in the cloud or in a data center. The encoding system 130 may be implemented on a device used to capture video. The encoding system 130 may be implemented on a standalone computing system. The encoding system 130 may perform the encoding process in video compression. The encoding system 130 may receive a video comprising a sequence of video frames 104 (e.g., uncompressed video, original video, raw video, etc.). The video frames 104 may include image frames or images that make up the video. The video may have a frame rate or frames per second (FPS) that defines the number of frames per second of the video. The higher the FPS, the more realistic and smooth the video appears. Typically, for human viewers, the FPS is greater than 24 frames per second to provide a natural, realistic viewing experience. Examples of videos may include television series, movies, short films, short videos (e.g., less than 15 seconds long), video captured gaming experiences, computer screen content, video conferencing content, live event broadcast content, sports content, surveillance video, videos shot using a mobile computing device (e.g., a smartphone), etc. In some cases, a video may include a mix or combination of different types of videos.
[0028] The encoding system 130 may include an encoder 102 that receives a video frame 104 and encodes the video frame 104 into an encoded bitstream 180. Figure 2 An exemplary implementation of the encoder 102 is shown in FIG.
[0029] The encoded bitstream 180 can be compressed, which indicates that the size of the encoded bitstream 180 can be smaller than the video frame 104. The encoded bitstream 180 can include a series of bits, e.g., with 0s and Is. The encoded bitstream 180 can have header information, payload information, and trailer information, which can be encoded as bits in the bitstream. The header information can provide information about one or more of the following: the format of the encoded bitstream 180, the encoding process implemented in the encoder 102, the parameters of the encoder 102, and the metadata of the encoded bitstream 180. For example, the header information can include one or more of the following: resolution information, frame rate, aspect ratio, color space, etc. The payload information can include data representing the content of the video frame 104, such as sample frames, symbols, syntax elements, etc. For example, the payload information can include bits encoding one or more of the following in the video frame 104: motion predictors, transform coefficients, prediction modes, and quantization levels. The trailer information can indicate the end of the encoded bitstream 180. The trailer information can include other information including one or more of the following: checksum, error correction code, and signature. The format of the encoded bitstream 180 can vary depending on the specification of the encoding and decoding process (i.e., codec).
[0030] The encoded bitstream 180 can include packets, where the encoded video data and signaling information can be packetized. One example format is the Open Bitstream Unit (OBU), which is used in AV1 encoded bitstreams. The OBU can include a header and a payload. The header can include information about the OBU, such as information indicating the type of the OBU. Examples of OBU types can include a sequence header OBU, a frame header OBU, a metadata OBU, a time delimiter OBU, and a tile group OBU. The payload in the OBU can carry quantized transform coefficients and syntax elements, which can be used in the decoder to properly decode the encoded video data to regenerate the video frames.
[0031] The encoded bitstream 180 can be transmitted to one or more decoding systems 150 1…D via a network 140. The network 140 can be the Internet. The network 140 can include one or more of the following: a cellular data network, a wireless data network, a wired data network, a cable Internet network, a fiber optic network, a satellite Internet network, etc.
[0032] D decoding systems 150 1…D are shown. At least one of the decoding systems 150 1…D may be located in a different country than the encoder 102. Figure 12implemented on computing devices 1200. Decoding system 150 1…D Examples can include personal computers, mobile computing devices, gaming devices, augmented reality devices, mixed reality devices, virtual reality devices, televisions, and the like. Decoding system 150 1…D Each of decoding systems 150 1…D may perform a decoding process in video compression. Decoding system 150 1…D may include a decoder (e.g., decoders 1...D 162 1…D ) and one or more display devices (e.g., display devices 1...D 164 D ). Figure 3 An example implementation of a decoder (e.g., decoder 1 1621) is shown.
[0033] For example, decoding system 1 1501 can include decoder 1 1621 and display device 1 1641. Decoder 1 1621 can implement a decoding process of video compression. Decoder 1 1621 can receive encoded bitstream 180 and produce decoded video 1681. Decoded video 1681 can include a series of video frames that can be versions or reconstructed versions of video frames 104 encoded by encoding system 130. Display device 1 1641 can output decoded video 1681 for display to one or more human viewers or users of decoding system 1 1501.
[0034] For example, decoding system 2 1502 can include decoder 2 1622 and display device 2 1642. Decoder 2 1622 can implement a decoding process of video compression. Decoder 2 1622 can receive encoded bitstream 180 and produce decoded video 1682. Decoded video 1682 can include a series of video frames that can be versions or reconstructed versions of video frames 104 encoded by encoding system 130. Display device 2 1642 can output decoded video 1682 for display to one or more human viewers or users of decoding system 2 1502.
[0035] For example, decoding system D 150 D may include decoder D 162 D and display device D 164 D . Decoder D 162 D may implement a decoding process of video compression. Decoder D 162 D may receive encoded bitstream 180 and produce decoded video 168 D . Decoded video 168 DA series of video frames can be included, which can be a version or a reconstructed version of the video frames 104 encoded by the encoding system 130. The display device D 164 D The decoded video 168 can be output D for display to the decoding system D 150 D one or more human viewers or users.
[0036] Video encoder
[0037] Figure 2 An encoder 102 for encoding video frames 104 and outputting an encoded bitstream is shown in accordance with some embodiments of the present disclosure. The encoder 102 can include one or more of signal processing operations and data processing operations, including inter-prediction and intra-prediction, transforms, quantization, loop filtering, and entropy encoding. The encoder 102 can include a reconstruction loop involving inverse quantization and inverse transforms to guarantee that the decoder will see the same reference blocks and frames. The encoder 102 can receive the video frames 104 and encode the video frames 104 into an encoded bitstream 180. The encoder 102 can include one or more of partitioning 206, transforms and quantization 214, inverse transforms and inverse quantization 218, loop filter 228, motion estimation 234, inter-prediction 236, intra-prediction 238, and entropy encoding 216.
[0038] In some embodiments, the video frames 104 can be processed by pre-analysis 290 prior to the encoding process being applied by the encoder 102. The pre-analysis 290 and the encoder 102 can form a system as shown in Figure 1The illustrated encoding system 130. Pre-analysis 290 can analyze video frames 104 to determine picture statistics, which can be used to inform one or more encoding processes to be performed by one or more components in the encoder 102. Pre-analysis 290 can determine information that can be used for quantization parameter (QP) adaptation, scene cut detection, and frame type adaptation. Pre-analysis 290 can determine a recommended frame type for each frame. Pre-analysis 290 can apply MCTF to denoise video frames 104. Filtered versions of video frames 104 that have applied MCTF can be provided to the encoder 102 as input video frames (instead of video frames 104), e.g., to partition 206. MCTF can include a motion estimation analysis operation and a bilateral filtering operation. MCTF can attenuate random picture components in a motion-aware manner to improve coding efficiency. MCTF can operate on blocks of 8x8 pixels, or 16x16 pixels. MCTF can operate on luma values and chroma values separately. MCTF can be applied in three dimensions (e.g., spatial directions and temporal direction). MCTF can produce noise estimates for various blocks.
[0039] Partition 206 can partition frames in video frames 104 (or filtered versions of video frames 104 from pre-analysis 290) into blocks of pixels. Different codecs can allow for different variable ranges of block sizes. In one codec, frames can be divided by partition 206 into blocks of 128x128 or 64x64 pixels. In some cases, frames can be divided by partition 206 into blocks of 256x256 or 512x512 pixels. In some cases, frames can be partitioned by partition 206 into blocks of 32x32 or 16x16 pixels. Large blocks can be referred to as superblocks, macroblocks, or CTBs. Partition 206 can further divide each large block using a multi-typed tree structure. In some cases, partitions of superblocks can be further recursively divided (e.g., subdivided into blocks / partitions of 4x4 size) by partition 206 using a multi-typed tree structure. In another codec, frames can be partitioned by partition 206 into CTUs of 128x128 pixels. Partition 206 can divide CTUs into four CUs using a quad-tree partitioning structure. Partition 206 can further recursively divide CUs using a quad-tree partitioning structure. Partition 206 can further subdivide CUs using a multi-typed tree structure (e.g., a quad-tree, binary tree, or ternary tree structure). Minimum CUs can have a size of 4x4 pixels. CUs can be referred to herein as blocks or partitions. Partition 206 can output original samples 208, e.g., as blocks or partitions of pixels.
[0040] In VVC, a frame in a video frame 104 can be partitioned into multiple non-overlapping CTUs. A CTU has a specified size, for example, 128x128 pixels or 64x64 pixels. Different types of partitioning shapes can be used to recursively split a CTU into smaller blocks or partitions. A quadtree partitioning structure can be used to partition a CTU into four CUs. One or more of the CUs obtained by the quadtree partitioning structure can be recursively divided (for example, up to three times) into smaller CUs using one of multiple types of structures (including, for example, quadtree, binary tree, or ternary tree structures) to support non-square partitioning. The quadtree partitioning structure can partition a CU into four CUs. The binary tree partitioning structure can partition a CU into two CUs (for example, horizontally or vertically). The ternary tree structure can partition a CU into three CUs (for example, horizontally or vertically). The smallest CU (for example, called a block or partition) can have a size of 4x4 pixels. A CU can be larger than 4x4 pixels. It will be appreciated that a CTU can be partitioned into CUs using many different possible partitioning combinations. A CTU can be partitioned in many different ways, resulting in many different partitioned results.
[0041] In some cases, one or more operations in segmentation 206 may be implemented in intra prediction 238 and / or inter prediction 236 .
[0042] Intra prediction 238 can predict samples of a block or partition from reconstructed predicted samples of previously encoded spatial neighboring / reference blocks of the same frame. Intra prediction 238 can receive reconstructed predicted samples 226 (of previously encoded spatial neighboring blocks of the same frame). Reconstructed predicted samples 226 can be generated by summer 222 from reconstructed predicted residuals 224 and predicted samples 212. Intra prediction 238 can determine a suitable predictor for predicting samples (and thus make an intra prediction decision) from reconstructed predicted samples of previously encoded spatial neighboring / reference blocks of the same frame. Intra prediction 238 can generate predicted samples 212 generated using the suitable predictor. Intra prediction 238 can output or identify the neighboring / reference blocks and predictor used to generate predicted samples 212. The identified neighboring / reference blocks and predictor can be encoded in encoded bitstream 180 to enable a decoder to use the same neighboring / reference blocks and predictor to reconstruct the block. In one codec, intra prediction 238 can support a number of distinct predictors, e.g., 56 different predictors. In one codec, intra prediction 238 can support a number of distinct predictors, e.g., 95 different predictors. Some predictors (e.g., directional predictors) can capture different spatial redundancies in directional textures. A directional predictor can be used in intra prediction 238 to predict pixel values of a block by extrapolating pixel values of a neighboring / reference block along a particular direction. Different codecs’ intra prediction 238 can support different sets of predictors to exploit different spatial patterns within the same frame. Examples of predictors can include direct current (DC), planar, Paeth, smooth, smooth vertical, smooth horizontal, recursive-based filtering pattern, chroma from luma, IBC, color palette or palette coding, multiple reference lines, intra sub-partition, matrix-based intra prediction (matrix coefficients can be defined by offline training using neural networks), angular prediction, wide-angle prediction, cross-component linear model, template matching, etc. IBC works by copying a reference block within the same frame to predict a current block. Palette coding or palette mode works by using a color palette with a few colors (e.g., 2-8 colors) and encoding a current block using indices for the color palette. In some cases, intra prediction 238 can perform block prediction, where a predicted block can be produced from reconstructed neighboring / reference blocks of the same frame using a vector. Optionally, a particular type of interpolation filter can be applied to the predicted block to blend pixels of the predicted block. A vector compensation process can be used in intra prediction 238 to predict pixel values of a block from a vector (within the same frame) by translating a neighboring / reference block (and optionally applying an interpolation filter to the neighboring / reference block) to produce predicted samples 212.Intra prediction 238 can output or identify the vector applied in generating the predicted samples 212. In some codecs, intra prediction 238 can encode: (1) a residual vector resulting from the applied vector and a vector predictor candidate; and (2) information identifying the vector predictor candidate, rather than encoding the applied vector itself. Intra prediction 238 can output or identify the interpolation filter type applied in generating the predicted samples 212.
[0043] Motion estimation 234 and inter prediction 236 can predict samples of a block from samples of a previously encoded frame (e.g., a reference frame in decoded picture buffer 232). Motion estimation 234 and inter prediction 236 can perform operations to make an intra prediction decision or an inter prediction decision. Motion estimation 234 can perform motion analysis and determine motion information for a current frame. Motion estimation 234 can determine a motion field for a current frame. The motion field can include motion vectors for blocks of the current frame. Motion estimation 234 can determine an average amount of motion vectors for the current frame. Motion estimation 234 can determine motion information that can indicate how much motion is present in the current frame (e.g., large motion, very dynamic motion, small / little motion, very static).
[0044] Motion estimation 234 and inter prediction 236 can perform motion compensation, which can involve identifying a suitable reference block and a suitable motion predictor (or motion vector predictor) for the block and optionally an interpolation filter to be applied to the reference block. Motion estimation 234 can receive original samples 208 from partitioning 206. Motion estimation 234 can receive samples from a decoded picture buffer 232 (e.g., samples of a previously encoded frame or reference frame). Motion estimation 234 can use multiple reference frames for determining one or more suitable motion predictors. A motion predictor can include a motion vector and a reference block that can be applied to generate a motion compensated block or predicted block. A motion predictor can include a motion vector that captures the movement of a block between frames in a video. Motion estimation 234 can output or identify one or more reference frames and one or more suitable motion predictors. Inter prediction 236 can apply the one or more suitable motion predictors and the one or more reference frames determined in motion estimation 234 to generate predicted samples 212. The identified reference frame(s) and motion predictor(s) can be encoded in encoded bitstream 180 to enable a decoder to use the same reference frame(s) and motion predictor(s) to reconstruct the block. In one codec, motion estimation 234 can implement a single reference frame prediction mode, where a single reference frame with a corresponding motion predictor is used for inter prediction 236. Motion estimation 234 can implement a compound reference frame prediction mode, where two reference frames with two corresponding motion predictors are used for inter prediction 236. In one codec, motion estimation 234 can implement techniques for searching and identifying good reference frame(s) that can yield the most efficient motion predictor. Techniques in motion estimation 234 can include searching good reference frame candidate(s) spatially (within the same frame) and temporally (in previously encoded frames). Techniques in motion estimation 234 can include searching a depth spatial neighborhood to find a pool of spatial candidates. Techniques in motion estimation 234 can include generating a pool of temporal candidates with a temporal motion field estimation mechanism. Techniques in motion estimation 234 can use a motion field estimation process. After temporal and spatial candidates can be ranked, a suitable motion predictor can be determined. In one codec, inter prediction 236 can support multiple distinct motion predictors.Examples of predictors can include geometric motion vectors (complex, non-linear motion), warped motion compensation (affine transforms that capture non- translational object movement), overlapped block motion compensation, advanced compound prediction (compound wedge prediction, differential modulation masking prediction, compound prediction based on frame distance, and compound inter-intra prediction), dynamic spatial and temporal motion vector referencing, affine motion compensation (capture high order motion such as rotation, scaling, and skew), adaptive motion vector resolution mode, geometric partition mode, optical flow, prediction refinement with optical flow, bi-prediction with weights, extended merge prediction, etc. Optionally, a particular type of interpolation filter can be applied to the predicted block to blend the pixels of the predicted block. The pixel values of the block can be predicted using the motion predictors / vectors determined in the motion estimation 234 and inter-prediction 236 and optionally applying the interpolation filter. In some cases, the inter-prediction 236 can perform motion compensation, where the predicted block can be generated from a reconstructed reference block of a reference frame using the motion predictors / vectors. The inter-prediction 236 can output or identify the motion predictors / vectors applied in generating the predicted samples 212. In some codecs, the inter-prediction 236 can encode: (1) a residual vector generated from the applied vector and vector predictor candidates; and (2) information identifying the vector predictor candidates, rather than the applied vector itself. The inter-prediction 236 can output or identify the interpolation filter type applied in generating the predicted samples 212.
[0045] The mode selection 230 can be informed by components such as the motion estimation 234 to determine whether inter-prediction 236 or intra-prediction 238 can be more efficient to use to encode a block (thus making an encoding decision). The inter-prediction 236 can output predicted samples 212 of a predicted block. The inter-prediction 236 can output a selected predictor and a selected interpolation filter (if applicable) that can be used to generate the predicted block. The intra-prediction 238 can output predicted samples 212 of a predicted block. The intra-prediction 238 can output a selected predictor and a selected interpolation filter (if applicable) that can be used to generate the predicted block. Regardless of the mode, the predicted residual 210 can be generated by the subtractor 220 by subtracting the original samples 208 from the predicted samples 212. In some cases, the predicted residual 210 can include a residual vector from the inter-prediction 236 and / or the intra-prediction 238.
[0046] Transform and quantization 214 can receive the predicted residual 210. The predicted residual 210 can be generated by a subtractor 220 that takes the original sample 208 and subtracts the predicted sample 212 to output the predicted residual 210. The predicted residual 210 can be referred to as a prediction error (e.g., an error between the original sample and the predicted sample 212) of intra prediction 238 and inter prediction 236. The prediction error has a range of smaller values than the original sample and can be encoded with fewer bits in the encoded bitstream 180. Transform and quantization 214 can include one or more of a transform and quantization. The transform can include converting the predicted residual 210 from a spatial domain to a frequency domain. The transform can include applying one or more transform kernels. Examples of transform kernels can include DCT in horizontal and vertical forms, asymmetrical discrete sine transform (ADST), flipped ADST, and identity transform (IDTX), multiple transform selection, low-frequency non-separable transform, sub-block transform, non-square transform, DCT-VIII, discrete sine transform VII (DST-VII), discrete wavelet transform (DWT), etc. The transform can convert the predicted residual 210 to transform coefficients. The quantization can quantize the coefficients of the transform, e.g., by reducing the precision of the transform coefficients. The quantization can include using a quantization matrix (e.g., linear and non-linear quantization matrices). The elements in the quantization matrix can be larger for higher frequency bands and smaller for lower frequency bands, which indicates that the higher frequency coefficients are coarser quantized and the lower frequency coefficients are finer quantized. The quantization can include dividing each transform coefficient by the corresponding element in the quantization matrix and rounding to the nearest integer. Effectively, the quantization matrix can implement different QPs for different frequency bands and chroma planes, and can use spatial prediction. A suitable quantization matrix can be selected and signaled for each frame and encoded in the encoded bitstream 180. The transform and quantization 214 can output quantized transform coefficients and syntax elements 278 that indicate the encoding modes and parameters used in the encoding process implemented in the encoder 102.
[0047] Inverse transform and inverse quantization 218 can apply inverse operations performed in transform and quantization 214 to produce reconstructed predicted residuals 224 as part of a reconstruction path to produce decoded picture buffer 232 for encoder 102. Inverse transform and inverse quantization 218 can receive quantized transform coefficients and syntax elements 278. Inverse transform and inverse quantization 218 can perform one or more inverse quantization operations (e.g., apply an inverse quantization matrix) to obtain unquantized / original transform coefficients. Inverse transform and inverse quantization 218 can perform one or more inverse transform operations, e.g., inverse transforms (e.g., inverse DCT, inverse DWT, etc.), to obtain reconstructed predicted residuals 224. The reconstruction path is provided in encoder 102 to generate reference blocks and frames stored in decoded picture buffer 232. The reference blocks and frames can match blocks and frames to be generated in a decoder. The reference blocks and frames are used as reference blocks and frames by motion estimation 234, inter prediction 236, and intra prediction 238.
[0048] The loop filter 228 can implement filters to remove artifacts introduced by the encoding process in the encoder 102 (e.g., processing performed by the partitioning 206 and the transform and quantization 214). The loop filter 228 can receive the reconstructed predicted samples 226 from the summer 222 and output frames to a decoded picture buffer 232. Examples of loop filters can include a constrained low pass filter, a directional de-noising filter, an edge-directed conditional substitution filter, a loop restoration filter, a Wiener filter, a self- guided restoration filter, a constrained directional enhancement filter (CDEF), an LMCS filter, a Sample Adaptive Offset (SAO) filter, an Adaptive Loop Filter (ALF), a cross-component ALF, a low pass filter, a deblocking filter, etc. For example, applying a deblocking filter across a boundary between two blocks can address blocking artifacts caused by the Gibbs phenomenon. In some embodiments, the loop filter 228 can obtain data from a frame buffer having reconstructed predicted samples 226 of individual blocks of a video frame. The loop filter 228 can determine whether to apply a loop filter. The loop filter 228 can determine one or more suitable filters that achieve good visual quality and / or one or more suitable filters that appropriately remove artifacts introduced by the encoding process in the encoder 102. The loop filter 228 can determine a type of loop filter to apply across a boundary between two blocks. The loop filter 228 can determine one or more strengths (e.g., filter coefficients) of the loop filter to apply across a boundary between two blocks based on the reconstructed predicted samples 226 of the two blocks. In some cases, the loop filter 228 can consider a desired bitrate when determining one or more suitable filters. In some cases, the loop filter 228 can consider a specified QP when determining one or more suitable filters. The loop filter 228 can apply the one or more (suitable) filters across the boundary separating the two blocks. After applying the one or more (suitable) filters, the loop filter 228 can write the (filtered) reconstructed samples to a frame buffer, such as the decoded picture buffer 232.
[0049] Entropy coding 216 can receive quantized transform coefficients and syntax elements 278 (e.g., referred to as symbols herein) and perform entropy coding. Entropy coding 216 can generate and output encoded bitstream 180. Entropy coding 216 can exploit statistical redundancies and apply lossless algorithms to encode the symbols and produce a compressed bitstream, e.g., encoded bitstream 180. Entropy coding 216 can implement some version of arithmetic coding. Different versions can have different advantages and disadvantages. In one codec, entropy coding 216 can implement (symbol-to-symbol) adaptive multi-symbol arithmetic coding. In another codec, entropy coding 216 can implement a context-based adaptive binary arithmetic coder (CABAC). Binary arithmetic coding differs from multi-symbol arithmetic coding. Binary arithmetic coding encodes only one bit at a time, e.g., with a binary value of 0 or 1. Binary arithmetic coding can first convert each symbol to a binary representation (e.g., using a fixed number of bits per symbol). Processing only binary values of 0 or 1 can simplify calculations and reduce complexity. Binary arithmetic coding can assign probabilities to each binary value (e.g., a chance that a bit has a binary value of 0 and a chance that a bit has a binary value of 1). Multi-symbol arithmetic coding performs encoding on an alphabet with at least two or three symbol values and assigns probabilities to each symbol value in the alphabet. Multi-symbol arithmetic coding can encode more bits at a time, which can result in a smaller number of operations to encode the same amount of data. Multi-symbol arithmetic coding can require more calculations and storage (as probability estimates can be updated for each element in the alphabet). Maintaining and updating probabilities (e.g., cumulative probability estimates) for each possible symbol value can be more complex in multi-symbol arithmetic coding (e.g., complexity grows with alphabet size). Multi-symbol arithmetic coding does not confuse binary arithmetic coding as the two different entropy coding processes are implemented differently and can result in different encoded bitstreams for the same set of quantized transform coefficients and syntax elements 278.
[0050] Video decoder
[0051] Figure 3A decoder 11621 for decoding an encoded bitstream and outputting a decoded video is shown in accordance with some embodiments of the present disclosure. The decoder 11621 can include one or more of signal processing operations and data processing operations, including entropy decoding, inverse transform, inverse quantization, inter prediction and intra prediction, loop filtering, etc. The decoder 11621 can have signal and data processing operations that mirror the operations performed in the encoder. The decoder 11621 can apply the signal and data processing operations signaled in the encoded bitstream 180 to reconstruct the video. The decoder 11621 can receive the encoded bitstream 180 and generate and output a decoded video 1681 having a plurality of video frames. The decoded video 1681 can be provided to one or more display devices for display to one or more human viewers. The decoder 11621 can include one or more of entropy decoding 302, inverse transform and inverse quantization 218, loop filter 228, inter prediction 236, and intra prediction 238. Some of the functions were previously described and used in the encoder 102, such as Figure 2 the encoder 102 of FIG. 1.
[0052] The entropy decoding 302 can decode the encoded bitstream 180 and output symbols encoded in the encoded bitstream 180. The symbols can include quantized transform coefficients and syntax elements 278. The entropy decoding 302 can reconstruct the symbols from the encoded bitstream 180.
[0053] The inverse transform and inverse quantization 218 can receive the quantized transform coefficients and the syntax elements 278 and perform operations performed in the encoder. The inverse transform and inverse quantization 218 can output reconstructed predicted residuals 224. The summer 222 can receive the reconstructed predicted residuals 224 and the predicted samples 212 and generate reconstructed predicted samples 226. The inverse transform and inverse quantization 218 can output the syntax elements 278 having signaling information to inform / indicate / control operations in the decoder 11621, such as mode selection 230, intra prediction 238, inter prediction 236, and loop filter 228.
[0054] Depending on the prediction mode signaled in the encoded bitstream 180 (e.g., as a syntax element in the quantized transform coefficients and the syntax elements 278), either the intra prediction 238 or the inter prediction 236 can be applied to generate the predicted samples 212.
[0055] The summer 222 can sum the predicted samples 212 of the decoded reference block with the reconstructed predicted residuals 224 to produce reconstructed predicted samples 226 of the reconstructed block. For intra prediction 238, the decoded reference block can be in the same frame as the block being decoded or reconstructed. For inter prediction 236, the decoded reference block can be in a different (reference) frame in the decoded picture buffer 232.
[0056] The intra prediction 238 can determine a reconstructed vector based on the residual vector and the selected vector predictor candidate. The intra prediction 238 can apply the reconstructed predictor or vector (e.g., according to signaled predictor information) to the reconstructed block, which can be produced using a decoded reference block of the same frame. The intra prediction 238 can apply an appropriate interpolation filter type (e.g., according to signaled interpolation filter information) to the reconstructed block to generate the predicted samples 212.
[0057] The inter prediction 236 can determine a reconstructed vector based on the residual vector and the selected vector predictor candidate. The inter prediction 236 can apply the reconstructed predictor or vector (e.g., according to signaled predictor information) to the reconstructed block, which can be produced using a decoded reference block of a different frame than the decoded picture buffer 232. The inter prediction 236 can apply an appropriate interpolation filter type (e.g., according to signaled interpolation filter information) to the reconstructed block to generate the predicted samples 212.
[0058] The in-loop filter 228 can receive the reconstructed predicted samples 226. The in-loop filter 228 can apply one or more filters signaled in the encoded bitstream 180 to the reconstructed predicted samples 226. The in-loop filter 228 can output the decoded video 1681.
[0059] Encoding tools in an encoding system that are not suitable for all types of content
[0060] Referring back to Figure 2 , the MCTF in pre-analysis 290 can help greatly improve encoder efficiency for natural (non-screen) content, but can cause quality loss for screen content. Also, MCTF is computationally very intensive. In some cases, MCTF can be implemented in the encoder 102 of FIG. 1, which can also share similar characteristics as the MCTF in pre-analysis 290. Figure 1
[0061] Continuing to refer to Figure 2 IBC in intra prediction 238 and palette coding can perform well with respect to screen content, but can result in quality loss if applied to natural (non-screen) content. Evaluating IBC and palette coding at both the frame level and the block level increases complexity for the encoder. Each coding tool can have a significant search space to consider when determining the best reference block and / or parameter values to apply. Additionally, the coding tools add additional options for the encoder to consider for rate-distortion optimization.
[0062] With continued reference to Figure 2 LMCS in loop filter 228 can perform well with respect to natural content, but is not very efficient with respect to screen content. Whether to apply LMCS is an additional coding decision that the encoder considers for rate-distortion optimization.
[0063] An encoding system (e.g., Figure 1 It would be beneficial for an encoding system 130, or an encoding system formed by pre-analysis 290 and encoder 102, to be able to disable or enable certain encoding tools, or configure encoding tools, based on the type of content in the video frames 104 being encoded. The complexity of the encoder would be significantly reduced while maintaining the quality of the encoder. However, it is not simple to implement an efficient (e.g., accurate) and efficient scheme for content classification and control flag / signal generation. Figures 4-11 The various embodiments described and illustrated in
[0064] Encoding system with pre-analysis and encoder
[0065] Figure 4 An encoder 102 and an adaptive coding tool selector 402 according to some embodiments of the disclosure are shown. In some embodiments, the adaptive coding tool selector 402 can be included before the encoder 102 as part of the pre-analysis 290. The encoder 102 and the adaptive coding tool selector 402 can form the encoding system 130. The video frames 104 can be provided to the adaptive coding tool selector 402. The video frames 104 can also be provided to the encoder 102.
[0066] The video frame 104 can be processed and / or analyzed by the adaptive coding tool selector 402. The adaptive coding tool selector 402 can set, generate, and / or output one or more coding tool control flags 404. The one or more coding tool control flags 404 can be used as control signals or can be used by the coding system 130 (e.g., the encoder 102 and the pre-analysis 290) with one or more coding tools to configure the one or more coding tools. The one or more coding tool control flags 404 can indicate one or more on / off decisions for the one or more coding tools. The one or more coding tool control flags 404 can include one or more parameter values (or configuration settings) for the one or more coding tools. The encoder 102 can comply with the one or more coding tool control flags 404 to correspondingly configure the one or more coding tools. The pre-analysis 290 can comply with the one or more coding tool control flags 404 to correspondingly configure the one or more coding tools. The configuration of the one or more coding tools can include turning on or enabling the coding tool. The configuration of the one or more coding tools can include turning off or disabling the coding tool. The configuration of the one or more coding tools can include configuring the coding tool to use one or more parameter values (e.g., one or more configuration settings).
[0067] In some cases, the one or more coding tool control flags 404 can signal information about the type of content in the video frame 104. One or more components in the pre-analysis 290 and the encoder 102 can derive one or more configuration commands based on the one or more coding tool control flags 404. The configuration commands can be executed by one or more components in the pre-analysis 290 and the encoder 102 to configure one or more components in the pre-analysis 290 and the encoder 102 in a particular way based on the information about the type of content in the video frame 104.
[0068] The one or more coding tool control flags 404 can be provided to the MCTF 406 in the pre-analysis 290. The one or more coding tool control flags 404 can turn on or off (e.g., enable or disable) the MCTF 406 in the pre-analysis 290. In some cases, the one or more coding tool control flags 404 can turn on or off (e.g., enable or disable) the MCTF implemented in the encoder 102. The one or more coding tool control flags 404 can specify a parameter value, e.g., a strength parameter value for the MCTF 406 in the pre-analysis 290 (e.g., which can affect the strength or coefficients of the filtering performed in the MCTF 406). The one or more coding tool control flags 404 can specify a parameter value, such as a strength parameter for the MCTF implemented in the encoder 102.
[0069] One or more coding tool control flags 404 can be provided to one or more components in the encoder 102. The one or more coding tool control flags 404 can turn on or off (e.g., enable or disable) one or more coding tools in the encoder 102. The one or more coding tool control flags 404 can specify parameter values, e.g., strength parameters (e.g., which can affect the strength or coefficients of filtering performed in the encoder 102), for one or more coding tools in the encoder 102.
[0070] If the one or more coding tool control flags 404 enable the MCTF 406, the MCTF 406 can generate a filtered version of the video frame 104. The filtered version of the video frame 104 generated by the MCTF 406 can be provided to the encoder 102. The encoder 102 can receive and use the filtered version of the video frame 104 instead of the original video frame 104. If the one or more coding tool control flags 404 disable the MCTF 406, the encoder 102 can receive and use the original video frame 104.
[0071] Adaptive encoding tool selector
[0072] Figure 5 An exemplary implementation of the adaptive coding tool selector 402 is shown in accordance with some embodiments of the present disclosure. The adaptive coding tool selector 402 can include a content classifier 502. The adaptive coding tool selector 402 can include a coding tool control flag determination 504.
[0073] The content classifier 502 can receive the video frame 104. The content classifier 502 can analyze the video frame 104 and detect whether the current frame has strong screen content, weak screen content, or non-screen content (or natural content). In some cases, the content classifier 502 can analyze the video frame 104 to classify the current frame as one of three classifications: a strong screen content classification, a weak screen content classification, and a natural content classification. The content classifier 502 can analyze the video frame 104 to determine whether the current frame belongs to the strong screen content classification, the weak screen content classification, or the natural content classification. The content classifier 502 can compute or calculate statistics about the frame and utilize the statistics to perform the classification. The content classifier 502 can generate and / or output a content classification 510. Implementation details of the content classifier 502 are further shown in Figures 6-7
[0074] Encoding tool control flag determination 504 can receive content classification 510 and generate one or more encoding tool control flags 404 based on content classification 510. For example, encoding tool control flag determination 504 can set one or more encoding tool control flags 404 based on the classification performed in content classifier 502 or content classification 510. Based on content classification 510, encoding tool control flag determination 504 can make one or more decisions regarding one or more encoding tools based on content classification 510. Encoding tool control flag determination 504 can set one or more values for one or more encoding tool control flags 404 according to the one or more decisions.
[0075] In some embodiments, encoding tool control flag determination 504 can decide whether to enable or disable MCTF based on content classification 510. Encoding tool control flag determination 504 can accordingly set the value for MCTF_FLAG. Encoding tool control flag determination 504 can make the decision for MCTF on a frame-by-frame basis.
[0076] In some embodiments, if MCTF is enabled, encoding tool control flag determination 504 can decide the strength of MCTF based on content classification 510. Encoding tool control flag determination 504 can accordingly set the value for MCTF_STRENGTH. Encoding tool control flag determination 504 can make the decision for MCTF strength on a frame-by-frame basis.
[0077] In some embodiments, encoding tool control flag determination 504 can decide whether to enable or disable LMCS based on content classification 510. Encoding tool control flag determination 504 can accordingly set the value for LMCS_FLAG. Encoding tool control flag determination 504 can make the decision for LMCS on a frame-by-frame basis.
[0078] In some embodiments, encoding tool control flag determination 504 can decide whether to enable or disable IBC and palette coding based on content classification 510. Encoding tool control flag determination 504 can accordingly set the values for IBC_FLAG and PALETTECODING_FLAG. In some cases, encoding tool control flag determination 504 can make the decisions regarding IBC and palette coding based on the first frame and / or an IDR-frame in a GOP or set of video frames, and the decisions can be applied as default to all other frames in the GOP or set of video frames. Encoding tool control flag determination 504 can modify the default decisions in certain scenarios.
[0079] Content classification
[0080] Figure 6An exemplary implementation of the content classifier 502 is shown, in accordance with some embodiments of the present disclosure.
[0081] The content classifier 502 can include a block-level statistics operation 602. The block-level statistics operation 602 can divide a current video frame in the video frames 104 into blocks (or blocks of pixels), where the size of the blocks can be 8x8 pixels. Using 8x8 blocks can be particularly effective because some encoding tools (e.g., MCTF) can operate on 8x8 pixel blocks. Using 8x8 blocks or larger blocks can be particularly effective because 8x8 blocks or larger blocks can better represent the current frame than 4x4 blocks, which can be prone to capturing false positives or noise information in the frame. In some cases, the size of the blocks can be 8x8 pixels or larger, such as 16x16 pixels. The block-level statistics operation 602 can compute or operate on two or more statistics for one or more blocks of pixels of the current frame. The block-level statistics operation 602 can compute the two or more statistics for each block. In some cases, the one or more blocks of pixels include luminance / brightness pixel values. In some cases, the one or more blocks of pixels include chrominance / chromaticity pixel values. In some cases, the block-level statistics operation 602 can only consider luminance / brightness pixel values. In some cases, the two or more statistics can already be available or computed in the encoder system for other purposes. In some cases, the two or more statistics can be easily and quickly computed without requiring heavy computations. The block-level statistics operation 602 can generate and / or output block-level statistics 690.
[0082] The two or more statistics can include a color number. The block-level statistics operation 602 can include a color number operation 604 to compute or operate on the color number for each block. The color number is a count or number of unique values in the block. Computing or operating on the color number for a block can include determining the number or count of unique pixel values (e.g., luminance / brightness values) in the block. The block can include 8x8 pixels = 64 pixels, and the color number can have a value of 1 to 64. A color number of 1 can indicate that there is only one unique value in the block. A color number of 64 can indicate that each value in the block is different / unique from each other.
[0083] The two or more statistics can include a variance. The variance is a statistical measure of how the pixel values of a block are distributed from the mean or average pixel value. The block-level statistics operation 602 can include a variance computation 606. Computing or operating on the variance for a block can include determining the average squared deviation from the mean of the pixel values of the block. In some cases, the standard deviation (e.g., square root of the variance) of the block can be computed or operated on in the variance computation 606.
[0084] The content classifier 502 can include frame-level statistics operations 610. The frame-level statistics operations 610 can receive block-level statistics 690 from the block-level statistics operations 602. Based on the block-level statistics 690, the frame-level statistics operations 610 can operate frame-level statistics. In some embodiments, the frame-level statistics operations 610 can determine three or more frame-level statistics. The three or more frame-level statistics can help accurately identify the type of content present in the current video frame. The three or more frame-level statistics can help classify the type of content present in the current frame with good precision and recall. In some cases, the three or more statistics can already be available or computed in the encoder system for other purposes. In some cases, the three or more statistics can be easily and quickly operated without requiring heavy computation.
[0085] The insight behind the implementation of the frame-level statistics operations 610 is that strong screen content can have one or more of the following frame-level characteristics when compared to natural content: (1) more blocks with only a few colors; (2) more blocks with zero variance (artificially flat regions); and (3) more blocks with large variance (very sharp edges). The frame-level statistics operations 610 can determine frame-level statistics 692 based on the block-level statistics 690, which can be used to help accurately distinguish different types of content. For example, the frame-level statistics operations 610 can determine a first proportion of pixel blocks in the current frame with a number of colors less than a first number (A) (%CL). The frame-level statistics operations 610 can determine a second proportion of pixel blocks in the current frame with a variance of (or very close to) zero (%ZV). The frame-level statistics operations 610 can determine a third proportion of pixel blocks in the current frame with a variance greater than a second number (B) (%BV). The frame-level statistics operations 610 can output the frame-level statistics 692.
[0086] The three or more statistics can include a first proportion of pixel blocks in the current frame with a number of colors less than a first number (A) (%CL). The first proportion can measure how many blocks (N1) of the total number of blocks (M) in the current frame have only a few colors or unique pixel values. The proportion can be measured as a fraction, a decimal number between 0 and 1, a percentage, etc. The frame-level statistics operations 610 can include a %CL operation 612 to operate the first proportion %CL. Computing or operating %CL can include determining the number or count of blocks with a number of colors less than A based on the block-level statistics 690, and dividing the number / count by M. In some cases, the result can be multiplied by 100 to obtain a percentage. The first proportion of pixel blocks %CL can be operated as:
[0087] %CL = 100 * N1 / M
[0088] The three or more statistics can include a second proportion of pixel blocks in the current frame that have a variance of zero (%ZV). The second proportion can measure how many blocks (N2) of the total number of blocks (M) of the current frame are artificially flat. The proportion can be measured as a fraction, a decimal number between 0 and 1, a percentage, etc. The frame-level statistics operation 610 can include a %ZV operation 614 to operate the second proportion %ZV. Computing or operating %ZV can include determining a number or count of blocks that have a variance of zero based on the block-level statistics 690 and dividing the number / count by M.
[0089] In some cases, the result can be multiplied by 100 to obtain a percentage. The second proportion of pixel blocks %ZV can be operated as:
[0090] %ZV = 100 * N2 / M
[0091] The three or more statistics can include a third proportion of pixel blocks in the current frame that have a variance greater than a second number (B) (%BV). The second proportion can measure how many blocks (N3) of the total number of blocks (M) of the current frame have very sharp edges. The proportion can be measured as a fraction, a decimal number between 0 and 1, a percentage, etc. The frame-level statistics operation 610 can include a %BV operation 616 to operate the third proportion %BV. Computing or operating %BV can include determining a number or count of blocks that have a variance greater than B based on the block-level statistics 690 and dividing the number / count by M. In some cases, the result can be multiplied by 100 to obtain a percentage. The third proportion of pixel blocks %BV can be operated as:
[0092] %BV = 100 * N3 / M
[0093] In some embodiments, the first number A can be set to a value of 4. In some embodiments, the first number A can be set to a value of 2. In some embodiments, the second number B can be set to 60 (e.g., for 8-bit pixel values). In some embodiments, the second number B can be set to 120 (e.g., for 8-bit pixel values).
[0094] The content classifier 502 can include a condition checker 680. The condition checker 680 can receive the frame-level statistics 692, which can include the first proportion %CL, the second proportion %ZV, and the third proportion %BV. Specifically, the condition checker 680 can check the frame-level statistics 692 against one or more conditions. The condition checker 680 can then use the results from checking the frame-level statistics 692 against the one or more conditions to classify the current frame as a strong screen content classification, a weak screen content classification, or a natural content classification. The condition checker 680 can classify the current frame as one of the three classifications based on the first proportion %CL, the second proportion %ZV, and the third proportion %BV.
[0095] Condition checker 680 can check for one or more conditions that can indicate that the current frame has strong screen content. Condition checker 680 can include strong screen content condition checker 620 to check whether frame-level statistics 692 satisfy one or more conditions that indicate that the current frame has strong screen content.
[0096] Condition checker 680 can check for one or more conditions that can indicate that the current frame has weak screen content. Condition checker 680 can include weak screen content condition checker 630 to check whether frame-level statistics 692 satisfy one or more conditions that indicate that the current frame has weak screen content.
[0097] Condition checker 680 can check for one or more conditions that can indicate that the current frame has non-screen or natural content. Condition checker 680 can include natural content condition checker 640 to check whether frame-level statistics 692 satisfy one or more conditions that indicate that the current frame has non-screen content.
[0098] Figure 7 Exemplary conditions used in condition checker 680 are shown in Table 1. In some embodiments, conditions are specified to identify frames that can have strong screen content. Conditions that are strong or more likely to be an indicator of strong screen content can be checked before other conditions that are weaker or less likely to be an indicator of strong screen content (e.g., conditions that explicitly indicate that the current frame has strong screen content).
[0099] Figure 7 An exemplary process 700 for classifying content is depicted in accordance with some embodiments of the present disclosure. Process 700 can be performed by content classifier 502 of FIG. 1. Figures 5-6
[0100] In 702, frame-level statistics can be computed. For example, frame-level statistics 692 can be computed by block-level statistics operation 602 and frame-level statistics operation 610 of FIG. 1. Figure 6
[0101] In 720, the current frame can be classified as strong screen content, or belong to the strong screen content classification.
[0102] In 724, the current frame can be classified as weak screen content, or belong to the weak screen content classification.
[0103] In 726, the current frame can be classified as natural content, or belong to the natural content classification.
[0104] Block 788 of process 700 includes one or more conditions, if any of the one or more conditions are satisfied, then a "yes" path will be followed to 720 and cause the current frame to be classified as strong screen content. The conditions can be arranged in order of stronger indicator to weaker indicator. The conditions checked in 704 can indicate a likely condition of strong screen content where many of the blocks have strong screen content of several colors. The conditions checked in 706 can indicate a likely condition of strong screen content where many of the blocks have sharp edges.
[0105] In 704, Figure 6 Strong screen content condition checker 620 of FIG. 1 can determine whether the first proportion of %CL is greater than a first color number threshold THCL1. In response to determining that the first proportion of %CL is greater than the first color number threshold THCL1 (e.g., following a "yes" path from 704 to 720), strong screen content condition checker 620 can determine that the current frame belongs to a strong screen content classification. In response to determining that the first proportion of %CL is not greater than (e.g., less than or equal to) the first color number threshold THCL1, process 700 can follow a "no" path from 704 to 706.
[0106] In 706, Figure 6 Strong screen content condition checker 620 of FIG. 1 can determine whether the third proportion of %BV is greater than a first large variance threshold THBV1. In response to determining that the third proportion of %BV is greater than the first large variance threshold THBV1 (e.g., following a "yes" path from 706 to 720), strong screen content condition checker 620 can determine that the current frame belongs to a strong screen content classification. In response to determining that the third proportion of %BV is not greater than (e.g., less than or equal to) the first large variance threshold THBV1, process 700 can follow a "no" path from 706 to 708.
[0107] In 708, Figure 6The strong screen content condition checker 620 can determine whether the first proportion %CL is greater than a second color number threshold THCL2 and whether the second proportion %ZV is greater than a first zero variance threshold THZV1. In response to determining that the first proportion %CL is greater than the second color number threshold THCL2 and that the second proportion %ZV is greater than the first zero variance threshold THZV1 (e.g., following the “yes” path from 708 to 720), the strong screen content condition checker 620 can determine that the current frame belongs to the strong screen content classification. In response to determining that the first proportion %CL is not greater than (e.g., less than or equal to) the second color number threshold THCL2 and / or that the second proportion %ZV is not greater than (e.g., less than or equal to) the first zero variance threshold THZV1, the process 700 can follow the “no” path from 708 to 710. If (%CL > THCL2 && %ZV > THZV1) = TRUE, then the process 700 can follow the “yes” path from 708 to 720. If (%CL > THCL2 && %ZV > THZV1) = FALSE, then the process 700 can follow the “no” path from 708 to 710.
[0108] In 708, Figure 6 The strong screen content condition checker 620 can determine whether the first proportion %CL is greater than a third color number threshold THCL3, whether the third proportion %BV is greater than a second large variance threshold TBV2, and whether the second proportion %ZV is greater than a second zero variance threshold TZV2. In response to determining that the first proportion %CL is greater than the third color number threshold THCL3, that the third proportion %BV is greater than the second large variance threshold TBV2, and that the second proportion %ZV is greater than the second zero variance threshold TZV2 (e.g., following the “yes” path from 710 to 720), the strong screen content condition checker 620 can determine that the current frame belongs to the strong screen content classification. In response to determining that the first proportion %CL is not greater than (e.g., less than or equal to) the third color number threshold THCL3, that the third proportion %BV is not greater than (e.g., less than or equal to) the second large variance threshold TBV2, and / or that the second proportion %ZV is not greater than (e.g., less than or equal to) the second zero variance threshold TZV2, the process 700 can follow the “no” path from 710 to 712. If (%CL > THCL3 && %BV > THBV2 && %ZV > THZV2) = TRUE, then the process 700 can follow the “yes” path from 710 to 720. If (%CL > THCL3 && %BV > THBV2 && %ZV > THZV2) = FALSE, then the process 700 can follow the “no” path from 710 to 712.
[0109] In 712, Encoding tool control flag determinationThe weak screen content condition checker 630 can determine whether the first proportion %CL is greater than the fourth color count threshold THCL4, the third proportion %BV is greater than the third large variance threshold THBV3, and the second proportion %ZV is greater than the third zero variance threshold THZV3. In response to determining that the first proportion %CL is greater than the fourth color count threshold THCL4, the third proportion %BV is greater than the third large variance threshold THBV3, and the second proportion %ZV is greater than the third zero variance threshold THZV3 (e.g., following the “yes” path from 712 to 724), the weak screen content condition checker 630 can determine that the current frame belongs to the weak screen content classification. In response to determining that the first proportion %CL is not greater than (e.g., less than or equal to) the fourth color count threshold THCL4, the third proportion %BV is not greater than (e.g., less than or equal to) the third large variance threshold THBV3, and the second proportion %ZV is not greater than (e.g., less than or equal to) the third zero variance threshold THZV3, the process 700 can follow the “no” path from 712 to 726. If (%CL > THCL4 && %BV > THBV3 && %ZV > THZV3) = TRUE, the process 700 can follow the “yes” path from 710 to 720. If (%CL > THCL4 && %BV > THBV3 && %ZV > THZV3) = FALSE, the process 700 can follow the “no” path from 712 to 726.
[0110] As a result of checking whether the first proportion %CL, the second proportion %ZV, and the third proportion %BV satisfy one or more conditions indicative of strong screen content for the current frame (e.g., the condition(s) checked in 704, 706, 708, and 710) and one or more conditions indicative of weak screen content for the current frame (e.g., the condition(s) checked in 712), the natural content condition checker 640 can determine whether the current frame has natural content. For example, in response to determining that the first proportion %CL, the second proportion %ZV, and the third proportion %BV do not satisfy one or more conditions indicative of strong screen content for the current frame (e.g., the condition(s) checked in 704, 706, 708, and 710) and do not satisfy one or more conditions indicative of weak screen content for the current frame (e.g., the condition(s) checked in 712), in 726, the natural content condition checker 640 can determine that the current frame belongs to the natural content classification.
[0111] The natural content condition checker 640 implementing part of the process 700 can not necessarily check for a particular condition(s) that explicitly indicate natural content. The process 700 can determine that the current frame has natural content because the current frame is not otherwise classified as having strong screen content in 720 and does not have weak screen content in 724.
[0112] The threshold values used can have one or more of the following relationships:
[0113] THCL1 > THCL2 > THCL3 > THCL4
[0114] THBV1 > THBV2 > THBV3
[0115] THZV1 > THZV2 > THZV3
[0116] Figure 8
[0117] Figure 5 An exemplary process 800 for setting one or more coding tool control signals is depicted in accordance with some embodiments of the present disclosure. Process 800 can be performed by a video encoder 102 of a video encoder 100 Figure 9 determined 504.
[0118] In 802, one or more coding tool control flags can be set to default values:
[0119] MCTF_FLAG = 1
[0120] LMCS_FLAG = 1
[0121] IBC_FLAG = 0
[0122] PALETTECODING_FLAG = 0
[0123] By default, MCTF and LMCS can be enabled, and control flags MCTF_FLAG and LMCS_FLAG can be set to have a value of 1. By default, IBC and palette coding can be disabled, and control flags IBC_FLAG and PALETTECODING_FLAG can be set to 0.
[0124] In 804, coding tool control flag determination 504 can determine whether the current frame has been classified as belonging to the strong screen content classification. If the current frame has been classified as belonging to the strong screen content classification, process 800 can follow the "yes" path from 804 to 806. If the current frame has not been classified as belonging to the strong screen content classification, process 800 can follow the "no" path from 804 to 806.
[0125] In 806, the encoding tool control flag determination 504 can set one or more encoding tool control flags to one or more values that disable one or more of MCTF and LMCS for the current frame. The encoding tool control flag determination 504 can set one or more encoding tool control flags to one or more values that disable both MCTF and LMCS for the current frame. The encoding tool control flag determination 504 can set MCTF_FLAG = 0 and LMCS_FLAG = 0.
[0126] In 808, the encoding tool control flag determination 504 can determine whether the current frame has been classified as belonging to the weak screen content classification. If the current frame has been classified as belonging to the weak screen content classification, the process 800 can follow the "yes" path from 808 to 810. If the current frame has not been classified as belonging to the weak screen content classification (or if the current frame has been classified as belonging to the natural content classification), the process 800 can follow the "no" path from 806 to 812.
[0127] In 810, the encoding tool control flag determination 504 can set one or more encoding tool control flags to one or more values that enable one or more of MCTF and LMCS at a weak strength for the current frame. The encoding tool control flag determination 504 can set one or more encoding tool control flags to one or more values that enable both MCTF and LMCS at a weak strength for the current frame. The encoding tool control flag determination 504 can set MCTF_FLAG = 1, MCTF_STRENGTH = WEAK, and LMCS_FLAG = 1.
[0128] In 812, the encoding tool control flag determination 504 can set one or more encoding tool control flags to one or more values that enable one or more of MCTF and LMCS for the current frame in response to classifying the current frame as belonging to the natural content classification. The encoding tool control flag determination 504 can set one or more encoding tool control flags to one or more values that enable both MCTF and LMCS for the current frame in response to classifying the current frame as belonging to the natural content classification. The encoding tool control flag determination 504 can set MCTF_FLAG = 1 and LMCS_FLAG = 1.
[0129] After 806 and 810, the process 800 can proceed to 808. Proceeding to 808 indicates that the current frame has been classified as belonging to the strong screen content classification or the weak screen content classification.
[0130] After 812, the process 800 can proceed to 814. Proceeding to 814 indicates that the current frame has been classified as belonging to the natural content classification.
[0131] Figure 5 An example process 900 for setting one or more coding tool control signals is depicted in accordance with some embodiments of the present disclosure. Process 900 can be performed by Figure 8 coding tool control flag determination 504. Process 900 can begin at 802. Figure 10
[0132] At 902, the coding tool control flag determination 504 can determine whether the current frame is the first frame in a GOP that includes the current frame. If the current frame is the first frame, process 900 can follow the “yes” path from 902 to 906. If the current frame is not the first frame, process 900 can follow the “no” path from 902 to 904.
[0133] At 904, the coding tool control flag determination 504 can determine whether the current frame is an IDR-frame. If the current frame is an IDR-frame, process 900 can follow the “yes” path from 904 to 906. If the current frame is not an IDR-frame (and is not the first frame of the GOP), process 900 can follow the “no” path from 904 to 908.
[0134] At 906, in response to determining that the current frame is the first frame in the GOP or an IDR-frame, the coding tool control flag determination 504 can set one or more coding tool control flags to one or more values that enable one or more of IBC and palette coding for all frames in the GOP. In response to determining that the current frame is the first frame in the GOP or an IDR-frame, the coding tool control flag determination 504 can set one or more coding tool control flags to one or more values that enable both IBC and palette coding for all frames in the GOP. The coding tool control flag determination 504 can set IBC_FLAG = 1, and PALETTECODING_FLAG = 1. The coding tool control flag determination 504 can enable or turn on IBC and palette coding for all frames in the GOP. The coding tool control flag determination 504 can set one or more syntax elements in a sequence parameter set for the GOP to indicate that IBC and palette coding are turned on for all frames in the GOP.
[0135] At 908, in response to determining that the current frame is not the first frame in the GOP and is not an IDR-frame, the coding tool control flag determination 504 can leave the values of one or more coding tool control flags IBC_FLAG and PALETTECODING_FLAG unchanged (or leave the one or more coding tool control flags at default values unchanged).
[0136] Figure 5 An exemplary process 1000 for setting one or more coding tool control signals is depicted in accordance with some embodiments of the present disclosure. Process 1000 can be performed by Figure 8 coding tool control flag determination 504. Process 1000 can begin at 1002. Figures 8-10
[0137] At 1002, the coding tool control flag determination 504 can determine whether the current frame is an IDR-frame. If the current frame is an IDR-frame, process 1000 can follow the “yes” path from 1002 to 1004. If the current frame is not an IDR-frame (nor is it the first frame of the GOP), process 900 can follow the “no” path from 1002 to 1006.
[0138] At 1004, in response to determining that the current frame is an IDR-frame, the coding tool control flag determination 504 can set one or more coding tool control flags to one or more values that disable one or more of IBC and palette coding for all frames in the GOP. In response to determining that the current frame is an IDR-frame, the coding tool control flag determination 504 can set one or more coding tool control flags to one or more values that disable both IBC and palette coding for all frames in the GOP. The coding tool control flag determination 504 can set IBC_FLAG = 0, and PALETTECODING_FLAG = 0. The coding tool control flag determination 504 can disable or turn off IBC and palette coding for all frames in the GOP. The coding tool control flag determination 504 can set one or more syntax elements in a sequence parameter set for the GOP to indicate that IBC and palette coding are turned off for all frames in the GOP.
[0139] At 1006, the coding tool control flag determination 504 can determine whether the sequence parameter set for the current GOP indicates or signals that IBC and palette coding have been enabled or set to on. In response to determining that the sequence parameter set indicates that IBC and palette coding are enabled (and the current frame is not an IDR-frame), process 1000 can follow the “yes” path from 1006 to 1008. In response to determining that the sequence parameter set indicates that IBC and palette coding are not enabled (and the current frame is not an IDR-frame), process 1000 can follow the “no” path from 1006 to 1010.
[0140] In 1008, in response to determining that the sequence parameter set indicates that IBC and palette coding are not enabled (and the current frame is not an IDR-frame), the coding tool control flag determination 504 can maintain the values of one or more coding tool control flags, IBC_FLAG and PALETTECODING_FLAG, unchanged (or leave the one or more coding tool control flags at default values unchanged).
[0141] In 1010, in response to determining that the sequence parameter set indicates that IBC and palette coding are enabled (and the current frame is not an IDR-frame), the coding tool control flag determination 504 can set the one or more coding tool control flags to one or more values that (cause the encoder) to skip IBC and palette coding for blocks of the current frame (decision making). The coding tool control flag determination 504 can skip all block-level IBC and palette coding decisions at block level for the current frame, even though IBC and palette coding are enabled at GOP level. Skipping all block-level IBC and palette coding decisions for the current frame can indicate that IBC and palette coding can not be selected as predictors in intra prediction for encoding the current frame. Skipping block-level IBC and palette coding decisions for the current frame can help reduce the complexity of intra prediction without considering IBC and palette coding decisions for at least the current frame, while maintaining encoder quality because the current frame has already been classified as having natural content.
[0142] Although Exemplary method for adaptively controlling encoding tools based on content classification The process in 1000 shows joint decision of MCTF, LMCS, IBC, and palette coding based on the classification of the current frame, in some embodiments, one or more decisions related to MCTF, LMCS, IBC, and palette coding can be decided independently and / or separately. In some embodiments, IBC can be enabled if the current frame is encoded using all intra coding for the current frame.
[0143] Figure 11
[0144] Exemplary computing device A method 1100 for adaptively controlling coding tools of an encoder based on content classification according to some embodiments of the disclosure is shown. The method 1100 can be performed by the adaptive coding tool selector 402 in the figures.
[0145] In 1102, for one or more blocks of pixels of a current frame, a color number and a variance can be computed. The one or more blocks of pixels can be 8x8 pixels or larger in size. The color number and the variance can be computed for each block of pixels of the current frame. The current frame can have a plurality of blocks of pixels. In some cases, the color number and the variance can be computed for each block of pixels in a subsampled set of blocks of pixels of the current frame.
[0146] At 1104, a first proportion of pixel blocks in the current frame having a number of colors less than a first number, a second proportion of pixel blocks in the current frame having a variance of zero, and a third proportion of pixel blocks in the current frame having a variance greater than a second number can be determined.
[0147] At 1106, the current frame can be classified as a strong screen content classification, a weak screen content classification, or a natural content classification based on the first proportion, the second proportion, and the third proportion.
[0148] At 1108, one or more coding tool control flags can be set based on the classification, wherein the one or more coding tool control flags configure one or more coding tools used by an encoding system.
[0149] Figure 12
[0150] Figure 12 is a block diagram of a device or system (e.g., example computing device 1200) in accordance with some embodiments of the present disclosure. One or more computing devices 1200 can be used to implement the functionality described in the figures and herein. A number of the components shown in the figure can be included in the computing device 1200, although any one or more of these components can be omitted or repeated for applications as appropriate. In some embodiments, some or all of the components included in the computing device 1200 can be attached to one or more motherboards. In some embodiments, some or all of the components are manufactured on a single system on a chip (SoC) die. Additionally, in various embodiments, the computing device 1200 can not include one or more of the components illustrated in FIG. 12, but the computing device 1200 can include an interface circuitry module to couple to the one or more components. For example, the computing device 1200 can not include a display device 1206, but can include a display device interface circuitry module (e.g., a connector and driver circuitry module) to which a display device 1206 can be coupled. In another set of examples, the computing device 1200 can not include an audio input device 1218 or an audio output device 1208, but can include an audio input or output device interface circuitry module (e.g., a connector and support circuitry module) to which an audio input device 1218 or an audio output device 1208 can be coupled. Figures 1-11
[0151] The computing device 1200 can include a processing device 1202 (e.g., one or more processing devices, one or more of the same type of processing device, one or more different types of processing devices). The processing device 1202 can include a processing circuitry or an electronic circuitry that processes electronic data from a data storage element (e.g., a register, a memory, a resistor, a capacitor, a qubit unit) for transforming that electronic data into other electronic data that can be stored in a register and / or a memory. Examples of the processing device 1202 can include a CPU, a GPU, a quantum processor, a machine learning processor, an artificial intelligence processor, a neural network processor, an artificial intelligence accelerator, an application specific integrated circuit (ASIC), an analog signal processor, an analog computer, a microprocessor, a digital signal processor, a field programmable gate array (FPGA), a tensor processing unit (TPU), a data processing unit (DPU), and the like.
[0152] The computing device 1200 can include a memory 1204, which can itself include one or more memory devices, such as volatile memory (e.g., DRAM), non-volatile memory (e.g., read-only memory (ROM)), high bandwidth memory (HBM), flash memory, solid-state memory, and / or a hard disk drive. The memory 1204 includes one or more non-transitory computer-readable storage media. In some embodiments, the memory 1204 can include a memory that shares a die with the processing device 1202.
[0153] In some embodiments, the memory 1204 includes one or more non-transitory computer-readable media storing instructions that are executable to implement operations described herein, such as Figure 7, process 700, process 800, process 900, process 1000, and the operations shown in method 1100. In some embodiments, memory 1204 includes one or more non-transitory computer-readable media storing instructions that are executable to implement one or more operations of encoder 102. In some embodiments, memory 1204 includes one or more non-transitory computer-readable media storing instructions that are executable to implement one or more operations of pre-analysis 290. In some embodiments, memory 1204 includes one or more non-transitory computer-readable media storing instructions that are executable to implement one or more operations of adaptive encoding tool selector 402. The instructions stored in memory 1204 can be executed by processing device 1202.
[0154] In some embodiments, memory 1204 may store data, such as data structures, binary data, bits, metadata, files, binary large objects (blobs), etc., as described in the figures and herein. Memory 1204 may include one or more non-transitory computer-readable media that store one or more of the following: input frames to the encoder (e.g., video frames 104), intermediate data structures computed by the encoder, a bitstream generated by the encoder (encoded bitstream 180), a bitstream received by the decoder (encoded bitstream 180), intermediate data structures computed by the decoder, and reconstructed frames generated by the decoder. Memory 1204 may include one or more non-transitory computer-readable media that store one or more of the following: data generated by pre-analysis 290 and / or data received. Memory 1204 may include one or more non-transitory computer-readable media that store one or more of the following: data generated by and / or data received by adaptive coding tool selector 402. Memory 1204 may include one or more non-transitory computer-readable media that store one or more of the following: data generated by Figure 8 The data generated and / or received by the process 700. The memory 1204 may include one or more non-transitory computer-readable media storing one or more of the following: Figure 9 The data generated and / or received by the process 800. The memory 1204 may include one or more non-transitory computer-readable media storing one or more of the following: Figure 10 The data generated and / or received by the process 900. The memory 1204 may include one or more non-transitory computer-readable media storing one or more of the following: Figure 11 The data generated and / or received by the process 1000. The memory 1204 may include one or more non-transitory computer-readable media storing one or more of the following:Selected examples data generated by the method 1100 and / or received data.
[0155] In some embodiments, computing device 1200 can include a communication device 1212 (e.g., one or more communication devices). For example, communication device 1212 can be configured for managing limited and / or wireless communication, for communicating data to and from computing device 1200. The term “wireless” and its derivatives can be used to describe circuits, devices, systems, methods, techniques, communication channels, and / or the like that can communicate data through the use of modulated electromagnetic radiation through a non-solid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Communication device 1212 can implement any of a number of wireless standards or protocols, including but not limited to the Institute for Electrical and Electronic Engineers (IEEE) standards, including Wi-Fi (the IEEE 802.10 family of standards), the IEEE 802.16 standards (e.g., the IEEE 802.16-2005 amendment), the Long-Term Evolution (LTE) project and any amendments, updates and / or revisions thereof (e.g., the LTE-Advanced project, the Ultra Mobile Broadband (UMB) project (also referred to as “3GPP2”), and / or the like). IEEE 802.16 compatible Broadband Wireless Access (BWA) networks are commonly referred to as WiMAX networks, the abbreviation being short for Worldwide Microwave Access, which is an authentication logo for products that pass conformance and interoperability testing for the IEEE 802.16 standards. Communication device 1212 can operate in accordance with a Global System for Mobile Communication (GSM), a General Packet Radio Service (GPRS), a Universal Mobile Telecommunications System (UMTS), a High Speed Packet Access (HSPA), an Evolved HSPA (E-HSPA), or an LTE network.The communication devices 1212 can operate according to Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). The communication devices 1212 can operate according to Code-division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and its derivatives, as well as any other wireless protocol that is designated as 3G, 4G, 5G, and beyond. In other embodiments, the communication devices 1212 can operate according to other wireless protocols. The computing device 1200 can include an antenna 1222 to facilitate
[0156] The computing device 1200 can include a power supply / power circuitry 1214. The power supply / power circuitry 1214 can include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry to couple the components of the computing device 1200 to an energy source separate from the computing device 1200 (e.g., a DC power supply, an AC power supply, etc.).
[0157] The computing device 1200 can include a display device 1206 (or corresponding interface circuitry, as discussed above). The display device 1206 can include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.
[0158] The computing device 1200 can include an audio output device 1208 (or corresponding interface circuitry, as discussed above). The audio output device 1208 can include any device that generates an audible indicator, such as speakers, headsets, or earbuds.
[0159] The computing device 1200 can include an audio input device 1218 (or corresponding interface circuitry, as discussed above). The audio input device 1218 can include any device that generates a signal representative of an acoustic sound, such as microphones, microphone arrays, or digital musical instruments (e.g., musical instruments with musical instrument digital interface (MIDI) output).
[0160] The computing device 1200 can include a GPS device 1216 (or corresponding interface circuitry, as discussed above). The GPS device 1216 can communicate with satellite-based systems and can receive a location of the computing device 1200, as is known in the art.
[0161] The computing device 1200 can include a sensor 1230 (or one or more sensors). As discussed above, the computing device 1200 can include corresponding interface circuitry. The sensor 1230 can sense a physical phenomenon and convert the physical phenomenon into an electrical signal that can be processed by, for example, the processing device 1202. Examples of the sensor 1230 can include a capacitive sensor, an inductive sensor, a resistive sensor, an electromagnetic field sensor, a light sensor, a camera, an imager, a microphone, a pressure sensor, a temperature sensor, a vibration sensor, an accelerometer, a gyroscope, a strain sensor, a moisture sensor, a humidity sensor, a distance sensor, a range sensor, a time-of-flight sensor, a pH sensor, a particulate sensor, an air quality sensor, a chemical sensor, a gas sensor, a biological sensor, an ultrasonic sensor, a scanner, and the like.
[0162] The computing device 1200 can include another output device 1210 (or a corresponding interface circuitry, as noted above). Examples of the other output device 1210 can include an audio codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, a haptic output device, a gas output device, a vibration output device, an illumination output device, a home automation controller, or an additional storage device.
[0163] The computing device 1200 can include another input device 1220 (or a corresponding interface circuitry, as noted above). Examples of the other input device 1220 can include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a bar code reader, a Quick Response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.
[0164] The computing device 1200 can have any desired form factor, such as a handheld or mobile computer system (e.g., a cell phone, a smart phone, a mobile Internet device, a music player, a tablet computer, a laptop computer, a netbook computer, a personal digital assistant (PDA), an ultra-mobile personal computer, a remote control, a wearable computer, a headgear, eyewear, footwear, electronic apparel, etc.), a desktop computer system, a server or other networked computing component, a printer, a scanner, a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital camera, a digital video recorder, an Internet of Things device, or a wearable computer system. In some embodiments, the computing device 1200 can be any other electronic device that processes data.
[0165] Variations and other notes
[0166] Example 1 provides a method comprising: computing a color number and a variance for one or more blocks of pixels of a current frame, wherein the one or more blocks of pixels are of a size of 8x8 pixels or greater; determining a first proportion of the blocks of pixels in the current frame in which the color number is less than a first number, a second proportion of the blocks of pixels in the current frame in which the variance is zero, and a third proportion of the blocks of pixels in the current frame in which the variance is greater than a second number; classifying the current frame as a strong screen content classification, a weak screen content classification, or a natural content classification based on the first proportion, the second proportion, and the third proportion; and setting one or more coding tool control flags based on the classification, wherein the one or more coding tool control flags configure one or more coding tools used by an encoding system.
[0167] Example 2 provides the method of example 1, wherein the pixel blocks of the one or more pixel blocks comprise luminance values.
[0168] Example 3 provides the method of example 1 or 2, wherein: calculating the number of colors comprises determining a count of unique pixel values in the pixel blocks.
[0169] Example 4 provides the method of any of examples 1-3, wherein classifying the current frame comprises: checking the first proportion, the second proportion, and the third proportion against one or more conditions indicative of strong screen content in the current frame, one or more conditions indicative of weak screen content in the current frame, and one or more conditions indicative of non-screen content in the current frame.
[0170] Example 5 provides the method of any of examples 1-4, wherein classifying the current frame comprises: determining whether the first proportion is greater than a first number of colors threshold; and in response to determining that the first proportion is greater than the first number of colors threshold, determining that the current frame belongs to a strong screen content classification.
[0171] Example 6 provides the method of any of examples 1-5, wherein classifying the current frame comprises: determining whether the third proportion is greater than a first large variance threshold; and in response to determining that the third proportion is greater than the first large variance threshold, determining that the current frame belongs to a strong screen content classification.
[0172] Example 7 provides the method of any of examples 1-6, wherein classifying the current frame comprises: determining whether the first proportion is greater than a second number of colors threshold and whether the second proportion is greater than a first zero variance threshold; and in response to determining that the first proportion is greater than the second number of colors threshold and that the second proportion is greater than the first zero variance threshold, determining that the current frame belongs to a strong screen content classification.
[0173] Example 8 provides the method of any of examples 1-7, wherein classifying the current frame comprises: determining whether the first proportion is greater than a third number of colors threshold, whether the third proportion is greater than a second large variance threshold, and whether the second proportion is greater than a second zero variance threshold; and in response to determining that the first proportion is greater than the third number of colors threshold, that the third proportion is greater than the second large variance threshold, and that the second proportion is greater than the second zero variance threshold, determining that the current frame belongs to a strong screen content classification.
[0174] Example 9 provides the method of any of examples 1-8, wherein classifying the current frame comprises: determining whether the first proportion is greater than a fourth number of colors threshold, whether the third proportion is greater than a third large variance threshold, and whether the second proportion is greater than a third zero variance threshold; and in response to determining that the first proportion is greater than the fourth number of colors threshold, that the third proportion is greater than the third large variance threshold, and that the second proportion is greater than the third zero variance threshold, determining that the current frame belongs to a weak screen content classification.
[0175] Example 10 provides the method of any of examples 1-9, wherein classifying the current frame comprises: determining whether the first proportion, the second proportion, and the third proportion satisfy one or more conditions indicative of strong screen content for the current frame and one or more conditions indicative of weak screen content for the current frame; and in response to determining that the first proportion, the second proportion, and the third proportion do not satisfy the one or more conditions indicative of strong screen content for the current frame and do not satisfy the one or more conditions indicative of weak screen content for the current frame, determining that the current frame belongs to the natural content classification.
[0176] Example 11 provides the method of any of examples 1-10, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the strong screen content classification, setting the one or more coding tool control flags to one or more values that disable a temporal filter with motion compensation and luma mapping with chroma scaling for the current frame.
[0177] Example 12 provides the method of any of examples 1-11, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the weak screen content classification, setting the one or more coding tool control flags to one or more values that enable a temporal filter with weak strength motion compensation and luma mapping with chroma scaling for the current frame.
[0178] Example 13 provides the method of any of examples 1-12, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the natural content classification, setting the one or more coding tool control flags to one or more values that enable a temporal filter with motion compensation and luma mapping with chroma scaling for the current frame.
[0179] Example 14 provides the method of any of examples 1-13, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the strong screen content classification or the weak screen content classification: determining whether the current frame is a first frame in a group of pictures or an instantaneous decoder refresh frame, the group of pictures including the current frame; and in response to determining that the current frame is the first frame in the group of pictures or the instantaneous decoder refresh frame, setting the one or more coding tool control flags to one or more values that enable intra block copy and palette coding for all frames in the group of pictures.
[0180] Example 15 provides the method of any of examples 1-14, wherein setting the one or more coding tool control flags comprises, in response to classifying the current frame as belonging to the natural content classification: determining whether the current frame is an instantaneous decoder refresh frame; and in response to determining that the current frame is an instantaneous decoder refresh frame, setting the one or more coding tool control flags to one or more values that disable intra block copy and palette coding for all frames in a group of pictures, the group of pictures including the current frame.
[0181] Example 16 provides the method of any of examples 1-15, wherein setting the one or more coding tool control flags comprises, in response to classifying the current frame as belonging to the natural content classification: determining whether the current frame is not an instantaneous decoder refresh frame and whether a sequence parameter set indicates that intra block copy and palette coding are enabled; and in response to determining that the current frame is not an instantaneous decoder refresh frame and that the sequence parameter set indicates that intra block copy and palette coding are enabled, setting the one or more coding tool control flags to one or more values that skip intra block copy and palette coding for blocks of the current frame.
[0182] Example 17 provides one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following operations: compute a number of colors and a variance for one or more blocks of pixels of a current frame, wherein the one or more blocks of pixels are 8x8 pixels or larger in size; determine a first proportion of blocks of pixels in the current frame that have a number of colors less than a first number, a second proportion of blocks of pixels in the current frame that have a variance of zero, and a third proportion of blocks of pixels in the current frame that have a variance greater than a second number; classify the current frame as a strong screen content classification, a weak screen content classification, or a natural content classification based on the first proportion, the second proportion, and the third proportion; and set one or more coding tool control flags based on the classification, wherein the one or more coding tool control flags configure one or more coding tools used by an encoding system.
[0183] Example 18 provides the one or more non-transitory computer-readable media of example 17, wherein a block of pixels of the one or more blocks of pixels includes a luminance value.
[0184] Example 19 provides the one or more non-transitory computer-readable media of example 17 or 18, wherein: computing the number of colors comprises determining a count of unique pixel values in the block of pixels.
[0185] Example 20 provides the one or more non-transitory computer- readable media of any of Examples 17-19, wherein classifying the current frame comprises checking the first proportion, the second proportion, and the third proportion against one or more conditions indicative of strong screen content in the current frame, one or more conditions indicative of weak screen content in the current frame, and one or more conditions indicative of non-screen content in the current frame.
[0186] Example 21 provides the one or more non-transitory computer- readable media of any of Examples 17-20, wherein classifying the current frame comprises determining whether the first proportion is greater than a first color number threshold, and in response to determining that the first proportion is greater than the first color number threshold, determining that the current frame belongs to a strong screen content classification.
[0187] Example 22 provides the one or more non-transitory computer- readable media of any of Examples 17-21, wherein classifying the current frame comprises determining whether the third proportion is greater than a first large variance threshold, and in response to determining that the third proportion is greater than the first large variance threshold, determining that the current frame belongs to a strong screen content classification.
[0188] Example 23 provides the one or more non-transitory computer- readable media of any of Examples 17-22, wherein classifying the current frame comprises determining whether the first proportion is greater than a second color number threshold and the second proportion is greater than a first zero variance threshold, and in response to determining that the first proportion is greater than the second color number threshold and the second proportion is greater than the first zero variance threshold, determining that the current frame belongs to a strong screen content classification.
[0189] Example 24 provides the one or more non-transitory computer- readable media of any of Examples 17-23, wherein classifying the current frame comprises determining whether the first proportion is greater than a third color number threshold, the third proportion is greater than a second large variance threshold, and the second proportion is greater than a second zero variance threshold, and in response to determining that the first proportion is greater than the third color number threshold, the third proportion is greater than the second large variance threshold, and the second proportion is greater than the second zero variance threshold, determining that the current frame belongs to a strong screen content classification.
[0190] Example 25 provides the one or more non-transitory computer- readable media of any of Examples 17-24, wherein classifying the current frame comprises determining whether the first proportion is greater than a fourth color number threshold, the third proportion is greater than a third large variance threshold, and the second proportion is greater than a third zero variance threshold, and in response to determining that the first proportion is greater than the fourth color number threshold, the third proportion is greater than the third large variance threshold, and the second proportion is greater than the third zero variance threshold, determining that the current frame belongs to a weak screen content classification.
[0191] Example 26 provides the one or more non-transitory computer-readable media of any of Examples 17-25, wherein classifying the current frame comprises: determining whether the first proportion, the second proportion, and the third proportion satisfy one or more conditions indicative of strong screen content for the current frame and one or more conditions indicative of weak screen content for the current frame; and in response to determining that the first proportion, the second proportion, and the third proportion do not satisfy the one or more conditions indicative of strong screen content for the current frame and do not satisfy the one or more conditions indicative of weak screen content for the current frame, determining that the current frame belongs to the natural content classification.
[0192] Example 27 provides the one or more non-transitory computer-readable media of any of Examples 17-26, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the strong screen content classification, setting the one or more coding tool control flags to one or more values that disable a temporal filter with motion compensation and luma mapping with chroma scaling for the current frame.
[0193] Example 28 provides the one or more non-transitory computer-readable media of any of Examples 17-27, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the weak screen content classification, setting the one or more coding tool control flags to one or more values that enable a temporal filter with motion compensation at a weak strength and luma mapping with chroma scaling for the current frame.
[0194] Example 29 provides the one or more non-transitory computer-readable media of any of Examples 17-28, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the natural content classification, setting the one or more coding tool control flags to one or more values that enable a temporal filter with motion compensation and luma mapping with chroma scaling for the current frame.
[0195] Example 30 provides the one or more non-transitory computer-readable media of any of Examples 17-29, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the strong screen content classification or the weak screen content classification: determining whether the current frame is a first frame in a group of pictures or an instantaneous decoder refresh frame, the group of pictures including the current frame; and in response to determining that the current frame is the first frame in the group of pictures or the instantaneous decoder refresh frame, setting the one or more coding tool control flags to one or more values that enable intra block copy and palette coding for all frames in the group of pictures.
[0196] Example 31 provides the one or more non-transitory computer-readable media of any of Examples 17-30, wherein setting the one or more coding tool control flags comprises, in response to classifying the current frame as belonging to the natural content classification: determining whether the current frame is an instantaneous decoder refresh frame; and in response to determining that the current frame is an instantaneous decoder refresh frame, setting the one or more coding tool control flags to one or more values that disable intra block copy and palette coding for all frames in a group of pictures, the group of pictures including the current frame.
[0197] Example 32 provides the one or more non-transitory computer-readable media of any of Examples 17-31, wherein setting the one or more coding tool control flags comprises, in response to classifying the current frame as belonging to the natural content classification: determining whether the current frame is not an instantaneous decoder refresh frame and whether a sequence parameter set indicates that intra block copy and palette coding are enabled; and in response to determining that the current frame is not an instantaneous decoder refresh frame and that the sequence parameter set indicates that intra block copy and palette coding are enabled, setting the one or more coding tool control flags to one or more values that skip intra block copy and palette coding for blocks of the current frame.
[0198] Example 33 provides a system comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to: compute a color number and a variance for one or more blocks of pixels of a current frame, wherein the one or more blocks of pixels are 8x8 pixels or larger in size; determine a first proportion of blocks of pixels in the current frame that have a color number less than a first number, a second proportion of blocks of pixels in the current frame that have a variance of zero, and a third proportion of blocks of pixels in the current frame that have a variance greater than a second number; classify the current frame as a strong screen content classification, a weak screen content classification, or a natural content classification based on the first proportion, the second proportion, and the third proportion; and set one or more coding tool control flags based on the classification, wherein the one or more coding tool control flags configure one or more coding tools used by an encoding system.
[0199] Example 34 provides the system of Example 33, wherein a block of pixels of the one or more blocks of pixels includes a luminance value.
[0200] Example 35 provides the system of Example 33 or 34, wherein: computing the color number comprises determining a count of unique pixel values in the block of pixels.
[0201] Example 36 provides the system of any of Examples 33-35, wherein classifying the current frame comprises checking the first proportion, the second proportion, and the third proportion against one or more conditions indicative of strong screen content in the current frame, one or more conditions indicative of weak screen content in the current frame, and one or more conditions indicative of non-screen content in the current frame.
[0202] Example 37 provides the system of any of Examples 33-36, wherein classifying the current frame comprises determining whether the first proportion is greater than a first color number threshold, and in response to determining that the first proportion is greater than the first color number threshold, determining that the current frame belongs to the strong screen content classification.
[0203] Example 38 provides the system of any of Examples 33-37, wherein classifying the current frame comprises determining whether the third proportion is greater than a first large variance threshold, and in response to determining that the third proportion is greater than the first large variance threshold, determining that the current frame belongs to the strong screen content classification.
[0204] Example 39 provides the system of any of Examples 33-38, wherein classifying the current frame comprises determining whether the first proportion is greater than a second color number threshold and the second proportion is greater than a first zero variance threshold, and in response to determining that the first proportion is greater than the second color number threshold and the second proportion is greater than the first zero variance threshold, determining that the current frame belongs to the strong screen content classification.
[0205] Example 40 provides the system of any of Examples 33-39, wherein classifying the current frame comprises determining whether the first proportion is greater than a third color number threshold, the third proportion is greater than a second large variance threshold, and the second proportion is greater than a second zero variance threshold, and in response to determining that the first proportion is greater than the third color number threshold, the third proportion is greater than the second large variance threshold, and the second proportion is greater than the second zero variance threshold, determining that the current frame belongs to the strong screen content classification.
[0206] Example 41 provides the system of any of Examples 33-40, wherein classifying the current frame comprises determining whether the first proportion is greater than a fourth color number threshold, the third proportion is greater than a third large variance threshold, and the second proportion is greater than a third zero variance threshold, and in response to determining that the first proportion is greater than the fourth color number threshold, the third proportion is greater than the third large variance threshold, and the second proportion is greater than the third zero variance threshold, determining that the current frame belongs to the weak screen content classification.
[0207] Example 42 provides the system of any of Examples 33-41, wherein classifying the current frame comprises: determining whether the first proportion, the second proportion, and the third proportion satisfy one or more conditions indicative of strong screen content for the current frame and one or more conditions indicative of weak screen content for the current frame; and in response to determining that the first proportion, the second proportion, and the third proportion do not satisfy the one or more conditions indicative of strong screen content for the current frame and do not satisfy the one or more conditions indicative of weak screen content for the current frame, determining that the current frame belongs to the natural content classification.
[0208] Example 43 provides the system of any of Examples 33-42, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the strong screen content classification, setting the one or more coding tool control flags to one or more values that disable a temporal filter with motion compensation and luma mapping with chroma scaling for the current frame.
[0209] Example 44 provides the system of any of Examples 33-43, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the weak screen content classification, setting the one or more coding tool control flags to one or more values that enable a temporal filter with motion compensation at a weak strength and luma mapping with chroma scaling for the current frame.
[0210] Example 45 provides the system of any of Examples 33-44, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the natural content classification, setting the one or more coding tool control flags to one or more values that enable a temporal filter with motion compensation and luma mapping with chroma scaling for the current frame.
[0211] Example 46 provides the system of any of Examples 33-45, wherein setting the one or more coding tool control flags comprises: in response to classifying the current frame as belonging to the strong screen content classification or the weak screen content classification: determining whether the current frame is a first frame in a group of pictures or an instantaneous decoder refresh frame, the group of pictures including the current frame; and in response to determining that the current frame is the first frame in the group of pictures or the instantaneous decoder refresh frame, setting the one or more coding tool control flags to one or more values that enable intra block copy and palette coding for all frames in the group of pictures.
[0212] Example 47 provides the system of any of Examples 33-46, wherein setting the one or more coding tool control flags comprises: responsive to classifying the current frame as belonging to the natural content classification: determining whether the current frame is an instantaneous decoder refresh frame; and responsive to determining that the current frame is an instantaneous decoder refresh frame, setting the one or more coding tool control flags to one or more values that disable intra block copy and palette coding for all frames in a group of pictures, the group of pictures including the current frame.
[0213] Example 48 provides the system of any of Examples 33-47, wherein setting the one or more coding tool control flags comprises: responsive to classifying the current frame as belonging to the natural content classification: determining whether the current frame is not an instantaneous decoder refresh frame and whether a sequence parameter set indicates that intra block copy and palette coding are enabled; and responsive to determining that the current frame is not an instantaneous decoder refresh frame and that the sequence parameter set indicates that intra block copy and palette coding are enabled, setting the one or more coding tool control flags to one or more values that skip intra block copy and palette coding for a block of the current frame.
[0214] Example A provides an apparatus comprising means or circuitry for performing any of the methods recited in Examples 1-16 and the methods / processes described herein.
[0215] Example B provides one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods recited in Examples 1-16 and the methods / processes described herein.
[0216] Example C provides an apparatus comprising: one or more processors for executing instructions, and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods recited in Examples 1-16 and the methods / processes described herein.
[0217] Example D provides an encoder system that generates an encoded bitstream using the operations described herein.
[0218] Example E provides an encoder system for performing any of the methods recited in Examples 1-16 and the methods / processes described herein.
[0219] Embodiment F provides pre-analysis 290 as described herein.
[0220] Example G provides adaptive coding tool selector 402 as described herein.
[0221] Example H provides pre-analysis 290 and an encoder 102 as described herein.
[0222] Figures 7-11
[0223] Although shown and described in Figures 7-11 Figures 7-11 The operations of the example methods described herein are illustrated in or other figures can be combined or can include more or fewer details than shown. Embodiments of the disclosure are operational with numerous other general purpose or special purpose computing system environments.
[0224] The above description of illustrative implementations of the disclosure, including what is described in the Abstract, is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. While specific implementations of, and examples for, the disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure, as those skilled in the relevant art will recognize. These modifications can be made to the disclosure in light of the above detailed description of the disclosure.
[0225] For purposes of explanation, specific numbers, materials, and configurations are set forth in order to provide a thorough understanding of the illustrative implementations. However, as would be readily apparent to one skilled in the art, the
[0226] Furthermore, with reference to the drawings, there is shown in schematic form embodiments that can be practiced. It should be understood that other embodiments can be utilized and structural or logical changes can be made without departing from the scope of the present disclosure. The following detailed description, therefore, is not to be taken in a limiting sense.
[0227] Various operations can be described as multiple discrete actions or operations in a manner that is most helpful in understanding the claimed subject matter. However, the order of description should not be construed as to imply that these operations are necessarily order dependent. In particular, these operations can not be performed in the order of presentation. Operations from different embodiments can be performed in an alternating manner. Operations described can be performed in a different order than described.
[0228] For purposes of this disclosure, the phrase "A or B" or the phrase "A and / or B" means (A), (B), or (A and B). For purposes of this disclosure, the phrase "A, B, or C" or the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B and C). When the term "between" is used to refer to a range of measurements, it includes the two ends of the range of measurements.
[0229] For purposes of this disclosure, "A is less than or equal to a first threshold" is equivalent to "A is less than a second threshold," provided the first threshold and the second threshold are set in a manner such that both statements result in the same logical outcome for any value of A. For purposes of this disclosure, "B is greater than a first threshold" is equivalent to "B is greater than or equal to a second threshold," provided the first threshold and the second threshold are set in a manner such that both statements result in the same logical outcome for any value of B.
[0230] The phrases "in an embodiment," "in embodiments," "in various embodiments" or "in some embodiments" as used throughout this description means one or more embodiments. The terms "comprising," "including," "containing," etc. (e.g., as used in the embodiments of the disclosure) are synonymous with each other and are used synonymously herein. The disclosure can use perspective-based descriptions such as "above," "below," "top," "bottom," and "side" that can be used to explain various features of the figures, but these terms are used only to facilitate discussing the figures, and do not imply or require any particular orientation of the devices described herein. The figures are not drawn to scale. Unless otherwise stated, the use of ordinal adjectives such as "first," "second," and "third," etc. to describe common objects does not indicate a required or required sequence, as distinct from any other sequence. The use of the articles "a" and "an" does not exclude a plurality, and "each" and "the" do not exclude a plurality and mean "at least one."
[0231] In the following detailed description, various aspects of the illustrative implementations will be described using terms commonly employed by those skilled in the art to convey the substance of their work to others skilled in the art.
[0232] The terms “substantially,” “close,” “approximately,” “near,” and “about” generally mean within + / - 20% of the target value as described herein or as known in the art. Similarly, the terms indicating the orientation of various elements, such as “coplanar,” “perpendicular,” “orthogonal,” “parallel,” or any other angle between elements, generally mean within + / - 5-20% of the target value as described herein or as known in the art.
[0233] In addition, the terms “comprise,” “comprising,” “include,” “including,” “have,” “having,” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Also, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or” unless expressly indicated otherwise.
[0234] The systems, methods, and devices of the disclosure each have several innovative aspects, none of which is solely responsible for the desirable attributes disclosed herein. The details of one or more implementations of the subject matter described in this specification are set forth in the description and the drawings.
Claims
1. A method comprising: computing, for one or more blocks of pixels of a current frame, a color number and a variance, wherein the one or more blocks of pixels are of a size of 8x8 pixels or greater; determining a first proportion of blocks of pixels in the current frame in which the color number is less than a first number, a second proportion of blocks of pixels in the current frame in which the variance is zero, and a third proportion of blocks of pixels in the current frame in which the variance is greater than a second number; classifying the current frame as a strong screen content classification, a weak screen content classification, or a natural content classification based on the first proportion, the second proportion, and the third proportion; and setting one or more coding tool control flags based on the classification, wherein the one or more coding tool control flags configure one or more coding tools used by an encoding system.
2. The method of claim 1, wherein, A block of pixels of the one or more blocks of pixels includes a luminance value.
3. The method of claim 1 or 2, wherein: computing the color number includes determining a count of unique pixel values in a block of pixels.
4. The method of claim 1 or 2, wherein, classifying the current frame includes: checking the first proportion, the second proportion, and the third proportion against one or more conditions indicative of strong screen content in the current frame, one or more conditions indicative of weak screen content in the current frame, and one or more conditions indicative of non-screen content in the current frame.
5. The method of claim 1 or 2, wherein, classifying the current frame includes: determining whether the first proportion is greater than a first color number threshold; and in response to determining that the first proportion is greater than the first color number threshold, determining that the current frame belongs to the strong screen content classification.
6. The method of claim 1 or 2, wherein, classifying the current frame includes: determining whether the third proportion is greater than a first large variance threshold; and in response to determining that the third proportion is greater than the first large variance threshold, determining that the current frame belongs to the strong screen content classification.
7. The method of claim 1 or 2, wherein, classifying the current frame includes: determining whether the first proportion is greater than a second color number threshold and the second proportion is greater than a first zero variance threshold; and in response to determining that the first proportion is greater than the second color number threshold and the second proportion is greater than the first zero variance threshold, determining that the current frame belongs to the strong screen content classification.
8. The method of claim 1 or 2, wherein, classifying the current frame includes: determining whether the first proportion is greater than a third color number threshold, the third proportion is greater than a second large variance threshold, and the second proportion is greater than a second zero variance threshold; and in response to determining that the first proportion is greater than the third color number threshold, the third proportion is greater than the second large variance threshold, and the second proportion is greater than the second zero variance threshold, determining that the current frame belongs to the strong screen content classification.
9. The method of claim 1 or 2, wherein, classifying the current frame includes: determining whether the first proportion is greater than a fourth color number threshold, the third proportion is greater than a third large variance threshold, and the second proportion is greater than a third zero variance threshold; and in response to determining that the first proportion is greater than the fourth color number threshold, the third proportion is greater than the third large variance threshold, and the second proportion is greater than the third zero variance threshold, determining that the current frame belongs to the strong screen content classification. determining that the current frame belongs to the strong screen content classification in response to determining that the first proportion is greater than the fourth color number threshold, that the third proportion is greater than the third large variance threshold, and that the second proportion is greater than the third zero variance threshold.
10. The method of claim 1 or 2, wherein, classifying the current frame includes: determining whether the first proportion, the second proportion, and the third proportion satisfy one or more conditions indicative of strong screen content for the current frame and one or more conditions indicative of weak screen content for the current frame; and determining that the current frame belongs to the natural content classification in response to determining that the first proportion, the second proportion, and the third proportion do not satisfy the one or more conditions indicative of strong screen content for the current frame and do not satisfy the one or more conditions indicative of weak screen content for the current frame.
11. The method of claim 1 or 2, wherein, setting the one or more coding tool control flags includes: setting the one or more coding tool control flags to one or more values that disable a temporal filter for motion compensation and luma mapping with chroma scaling for the current frame in response to classifying the current frame as belonging to the strong screen content classification.
12. The method of claim 1 or 2, wherein, setting the one or more coding tool control flags includes: setting the one or more coding tool control flags to one or more values that enable a temporal filter for motion compensation at a weak strength and luma mapping with chroma scaling for the current frame in response to classifying the current frame as belonging to the weak screen content classification.
13. The method of claim 1 or 2, wherein, setting the one or more coding tool control flags includes: setting the one or more coding tool control flags to one or more values that enable a temporal filter for motion compensation and luma mapping with chroma scaling for the current frame in response to classifying the current frame as belonging to the natural content classification.
14. The method of claim 1 or 2, wherein, setting the one or more coding tool control flags includes: setting the one or more coding tool control flags to one or more values in response to classifying the current frame as belonging to the strong screen content classification or the weak screen content classification: determining whether the current frame is a first frame in a group of pictures or an instantaneous decoder refresh frame, the group of pictures including the current frame; and setting the one or more coding tool control flags to one or more values that enable intra block copy and palette coding for all frames in the group of pictures in response to determining that the current frame is the first frame in the group of pictures or the instantaneous decoder refresh frame.
15. The method of claim 1 or 2, wherein, setting the one or more coding tool control flags includes: setting the one or more coding tool control flags to one or more values in response to classifying the current frame as belonging to the natural content classification: determining whether the current frame is an instantaneous decoder refresh frame; and setting the one or more coding tool control flags to one or more values that disable intra block copy and palette coding for all frames in a group of pictures, the group of pictures including the current frame, in response to determining that the current frame is the instantaneous decoder refresh frame.
16. The method of claim 1 or 2, wherein, setting the one or more coding tool control flags includes: in response to classifying the current frame as belonging to the natural content classification: determining whether the current frame is not an instant decoder refresh frame and a sequence parameter set indicates that intra block copy and palette coding are enabled; and in response to determining that the current frame is not the instant decoder refresh frame and the sequence parameter set indicates that intra block copy and palette coding are enabled, setting the one or more coding tool control flags to one or more values that skip intra block copy and palette coding for blocks of the current frame.
17. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to implement a method according to any of the preceding claims.
18. A system comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to implement a method according to any of claims 1-16.
19. A computer program comprising instructions that, when executed by a processor, cause the processor to perform a method according to any of claims 1-16.