Domain-adaptive, data-efficient generation of partitioning decisions and mode decisions for video coding

The video coding system addresses the challenge of achieving high compression efficiency and video quality by using detected image features to optimize encoding operations, resulting in improved data usage efficiency and compression rate.

DE102018129344B4Active Publication Date: 2025-06-26INTEL CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102018129344
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-12-28
Filing Date
2018-11-21
Publication Date
2025-06-26
Estimated Expiration
2038-11-21

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in achieving high compression efficiency while maintaining video quality, particularly in optimizing data usage efficiency and coding speed.

Method used

The proposed solution involves a video coding system that modifies encoding operations based on detected features of an image, such as luma and chroma averages, edge detection, and motion analysis, to selectively use chroma information and generate partial transform coefficient blocks, thereby improving data usage efficiency.

Benefits of technology

This approach enhances data usage efficiency and compression rate while maintaining or improving video quality by tailoring encoding decisions to specific image features, reducing computational resources, and minimizing artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented method (1600) for video coding, comprising: Receiving (1601) an input video (111) for encoding, wherein the input video (111) comprises a plurality of images (200) and a first image (301) of the plurality of images (200) comprises a region comprising an individual block, the individual block comprising a plurality of partitions; applying (1602) one or more detectors to the region, the individual block, and / or one or more of the plurality of partitions to generate one or more detection indicators; Generating (1603) a partitioning decision for the individual block and coding mode decisions for partitions of the individual block corresponding to the partitioning decision using the detection indicators based on generating an evaluation decision for luma and chroma or only for luma for a first partition of the partitions, generating a merge mode decision or drop mode decision for a second partition of the partitions having an initial merge mode decision, generating only a portion of a transform coefficient block for a third partition of the partitions, and / or evaluating 4×4 modes only for a fourth partition of the partitions that is an initial 8×8 coding partition; and encoding (1604) the individual block based at least on the partitioning decision to generate a portion of an output bitstream (113);wherein the detection indicators comprise indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, and an average of a second chroma channel of the first partition exceeds a third threshold, and generating the partitioning decision and the encoding mode decisions comprises generating the evaluation decision for luma and chroma or only for luma for the first partition by applying an evaluation decision for only luma for the first partition if the luma average does not exceed the first threshold, the average of the first chroma channel does not exceed the second threshold, and the average of the second chroma channel does not exceed the third threshold;
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUNDIn compression / decompression (codec) systems, compression efficiency, data usage efficiency, and video quality are important performance criteria. Optical quality is an important aspect of user experience in many video applications, and compression efficiency, affected by data usage efficiency, affects the amount of data storage required to store video files and / or the amount of bandwidth required to send and / or stream video content. For example, a video encoder compresses video information such that more information can be transmitted over a given bandwidth, stored in a given memory location, or the like. The compressed signal or data may then be decoded by a decoder that decodes or decompresses the signal or data for display to a user. In most implementations, higher optical quality with more compression is desirable. Moreover, the coding speed and the coding efficiency are important aspects of video coding.CHEN, Jianle [et al.]: "Algorithm Description of Joint Exploration Test Model 7 (JEM 7)" describes the algorithm of the Joint Exploration Model 7 (JEM 7). It describes the coding functions studied as a potentially improved video coding technology in a coordinated test model study of the Joint Video Exploration Team (JVET). The description of the coding strategies used in experiments for studying the new technology in JEM is also provided.US 2014 / 0 003 495 A1 discloses a method and an apparatus for scalable video coding, wherein the video data is configured into a base layer (BL) and an extension layer (EL), and wherein the EL has a higher spatial resolution or better video quality than the BL. According to embodiments of the present invention, information from the base layer is used to encode the enhancement layer. The information coding for the enhancement layer includes a CU structure, information on a motion vector predictor (MVP), MVP / merge candidates, an intra prediction mode, residual-quadrature information, texture information, residual information, context-adaptive entropy coding, an adaptive Lop filter (ALF), a sample adaptive offset (SAO), and a deblocking filter.It may be advantageous to improve data usage efficiency and compression rate by data usage efficiency while maintaining or even improving video quality. With these and other considerations in mind, the present improvements have been required. Such improvements may be essential as the desire to compress and transmit video data is becoming more prevalent.BRIEF DESCRIPTION OF THE DRAWINGSThe fabric described herein is illustrated by way of example and not limitation in the accompanying figures. For simplification and clarification of the illustration, elements represented in the figures are not necessarily drawn to scale. For example, for clarity, the dimensions of some elements may be exaggerated relative to other elements. Further, reference numerals have been repeated between the figures where deemed appropriate to designate corresponding or corresponding elements. The figures show: FIG. 1 is an illustrative diagram of an example system for providing video coding; FIG. 2 shows an example group of images; FIG. 3 is an example video image; FIG. 4 is an illustrative diagram of an example partitioning decision module and example mode decision module for providing LCU partitions and internal / cross mode data; FIG. 5 is an illustrative diagram of an example encoder for generating a bitstream; FIG. 6 is a block diagram of an example integrated coding system; FIG. 7 is a flow chart depicting an example process for selectively using chroma information in partitioning decisions and coding mode decisions; FIG. 8 is a flow chart depicting an example process for generating a merge mode decision or exit mode decision for a partition having an initial merge mode decision; FIG. 9 is a flow chart illustrating an example process for determining a partitioning decision and encoding mode decisions for a block by generating only a portion of a transform coefficient block for a portion of the block; FIG. 10 shows an example data structure corresponding to an example subtransformation; FIG. 11 shows an example data structure corresponding to a further example subtransformation; FIG. 12 is a flow chart illustrating an example process for determining a partitioning decision and encoding mode decisions for a block by generating only a portion of a transform coefficient block for a partition of the block based on whether the partition is in an optically important area; FIG. 13 is a flow chart illustrating an example process for determining a partitioning decision and encoding mode decisions for a block by generating only a portion of a transform coefficient block for a partition of the block based on edge detection in the block; FIG. 14 is a flow chart illustrating an example process for selectively evaluating 4×4 partitions in video coding; FIG. 15 is an illustrative diagram of an example detector for flat and noisy regions; FIG. 16 is a flow diagram illustrating an example process for video coding; FIG. 17 is an illustrative diagram of an example system for video coding; FIG. 18 is an illustrative diagram of an example system; and FIG. 19 illustrates an example device corresponding to at least some implementations of the present disclosure.DETAILED DESCRIPTIONOne or more embodiments or implementations will now be described with reference to the accompanying figures. While certain configurations and arrangements are discussed, it is to be understood that this is done for purposes of illustration only. Those skilled in the art will recognize that other configurations and arrangements may be employed without departing from the spirit and scope of the description. It will be apparent to those skilled in the art that techniques and / or arrangements described herein may also be employed in a variety of other systems and applications other than those described herein.While the following description sets forth various implementations that may occur in architectures such as system-on-a-chip (SoC) architectures, implementations of the techniques and / or arrangements described herein are not limited to particular architectures and / or computing systems and may be implemented for similar purposes by any architecture and / or computing system. For example, various architectures employing, e.g., multiple integrated circuit (IC) chips and / or integrated circuit (IC) packages, various computing devices and / or consumer electronics (CE) devices such as set-top boxes, smart phones, etc., may implement the techniques and / or arrangements described herein. Further, while the following description may set forth numerous specific details such as logic implementations, types and relationships of system components, logic partitioning decisions / logic integration decisions, etc., the claimed subject matter may be practiced without such specific details. In other instances, certain matter such as control structures and full software instruction sequences may not be shown in detail in order not to obscure the matter disclosed herein.The substance disclosed herein may be implemented in hardware, firmware, software, or any combinations thereof. The substance disclosed herein may also be implemented as instructions stored on a machine readable medium and read and executed by one or more processors. A machine-readable medium may include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include read-only memory (ROM); random access memory (RAM); magnetic disk storage media; optical storage media; flash memory devices; electrical, optical, acoustical or other forms of propagating signals (e.g., carrier waves, infrared signals, digital signals, etc.), and so forth.References in the specification to "any implementation", "an implementation", "an example implementation", etc., indicate that the described implementation may include a particular feature, structure, or characteristic, but not every embodiment necessarily needs to include the particular feature, structure, or characteristic. Moreover, such word connections relate to the same implementation. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is applicable that causing such feature, structure, or characteristic in connection with other implementations be within the skill of a person of ordinary skill in the art, regardless of whether explicitly described herein.Methods, apparatus, computing platforms, and articles are described herein in relation to video coding, and more particularly, to implementing detectors of video properties to alter video coding to achieve improved efficiency.Techniques discussed herein provide improved data usage efficiency by modifying encoding operations based on detected features of a portion of an image. As used herein, the term region may include any block of an image, a coding unit of an image, a largest coding unit of an image, a region including multiple contiguous blocks of an image, a partition of a block or coding unit, a portion of an image, or the image itself. Moreover, the term partition may indicate a partition for encoding or a partition for transformation. The detected features that may be displayed by detection indicators may include any features discussed herein, such as a luma average of a range (i.e., the average of luma values for one range), an average of a chroma channel of a range (i.e., the average of chroma values for a particular chroma channel), and / or an average of a second chroma channel of a range (i.e., the average of chroma values for another particular chroma channel), indicators of the result of comparing such values to thresholds (e.g., whether the average exceeds a threshold), the time level of the range (e.g., whether the range is an I-portion, a B portion of a base layer, a B portion of no base layer, etc.), an amount of difference between encoding cost of an initial exhaust mode and encoding cost of an initial merging mode for an area, an indicator of whether an area includes an edge, an indicator of the strength of such an edge, an indicator of whether an area is in a revealed area or is a revealed area, an initial best internal mode of an area, or others, as discussed herein.Such detected features or detection indicators are then used to alter the encoding as further discussed herein. Such coding changes may include the evaluation of luma alone or the evaluation of luma and chroma for partitioning decisions and / or coding modes for a block, the use of luma and chroma only for merge mode decisions or skip mode decisions, the use of initial merge mode decision or skip mode decision without further evaluation in a coding pass, the generation of only portions of transform coefficient blocks in a local decoding loop (i.e., no generation of full transform coefficient blocks in some cases to improve efficiency), an evaluation of 4×4 internal modes in addition to an evaluation of 8×8 coding modes, and as discussed herein, further.The detected features or detection indicators discussed may be generated using original video content (e.g., without use of pixels reconstructed in the local decoding loop) and may be implemented in the context of a decoupled video encoder that decouples the generation of a final partitioning decision and associated initial coding mode decisions based on the use of only source samples from fully standard compliant coding with a compliant local decoding loop, or in the context of an integrated encoder that generates partitioning decisions and coding mode decisions using reconstructed samples from a local decoding loop. As used herein, the term sample or pixel sample may be any suitable pixel value. The term original pixel sample is used to indicate samples or values of input video and to render them to reconstructed pixel samples that are not original pixel samples but are reconstructed after encoding and decoding operations in a standard compliant encoder.FIG. 1 is an illustrative diagram of an example system 100 for providing video coding, configured in accordance with at least some implementations of the present disclosure. As shown in FIG. 1, system 100 includes a partitioning and mode decision module 101 and an encoder 102. As shown, partitioning and mode decision module 101, which may be characterized as a partitioning, motion estimation, and mode decision module or the like, captures input video 111 and optionally reconstructed images 114, and partitioning and mode decision module 101 generates partitions of the largest encoding unit (LCU partitions) and corresponding encoding mode data 112 (internal / cross-across modes). For example, the partitioning and mode decision module 101 provides, for each LCU of each picture of the input video 111, a partitioning decision (i.e., data indicating how to partition the LCU into encoding units / prediction units / transformation units (CU / PU / TU)), an encoding mode for each CU (i.e., an overlapping mode, an internal mode, or the like), and, if necessary, information for the encoding mode (i.e., a motion vector for the overlapping encoding). As used herein, the term partition is used to indicate any sub-block or sub-region of a block, such as a partition for encoding, a partition for transformation, or the like. For example, in the context of a block that is a largest coding unit, a partition may be a coding unit (e.g., CU) or a transformation unit (e.g., TU). A transformation unit may be equal to or smaller than an encoding unit.As shown, the encoder 102 captures LCU partitions and internal / cross-over mode data 112 and the encoder 102 generates the bitstream 113 such as a standard compliant bitstream and reconstructed images 114. For example, the encoder 102 implements LCU partitions and internal / cross-over mode data 112. In decoupled encoder embodiments, the encoder 102 implements final decisions made by the partitioning and mode decision module 101, selectively adjusts any initial mode decisions made by the partitioning and mode decision module 101 and implements such partitioning and mode decisions to generate the standard compliant bitstream 113. In such embodiments, reconstructed images 114 may be generated to serve as reference images in encoder 102, however, such reconstructed images 114 are not used in partitioning and mode decision module 101. In integrated encoder implementations, the encoder 102 implements decisions made by the partitioning and mode decision module 101 and implements such partitioning decisions and mode decisions to generate the standard compliant bitstream 113 and reconstructed images 114. Such reconstructed images 114 are used in the generation of partitioning decisions and mode decisions for subsequent LCUs of the input video 111.As shown, the system 100 receives input video 111 for encoding and the system provides video compression to generate the bitstream 113, such that the system 100 may be a video encoder implemented by a computer, computing device, or the like. Bitstream 113 may be any suitable bitstream such as a standard compliant bitstream. For example, bitstream 113 may be standard compliant to H.264 / MPEG-4 advanced video coding (AVC), standard compliant to H.265 high efficient video coding (HEVC), standard compliant to VP9, etc. The system 100 may be implemented by any device such as a personal computer, a laptop computer, a tablet, a phablet, a smartphone, a digital camera, a game console, a wearable device, a multifunction device, a dual function device, or the like, or a platform such as a mobile platform, or the like. For example, as used herein, a system, device, computer, or computing device may include any such device or platform.The input video 111 may include any suitable video frames, video images, a sequence of video frames, a group of images, groups of images, video data, or the like, at any suitable resolution. For example, the video may be video graphics assembly (VGA), high-resolution (HD), full-HD (e.g., 1080p), 4K resolution video, 8K resolution video, or the like, and the video may include any number of video frames, sequences of video frames, images, groups of images, or the like. Techniques discussed herein are discussed with respect to images and blocks and / or encoding units for clarity of presentation. However, such pictures may be characterized as frames, video frames, sequences of frames, video sequences, or the like, and such blocks and / or encoding units may be characterized as encoding blocks, macroblocks, subunits, sub-blocks, regions, sub-regions, etc. Typically, the terms block and coding unit are used interchangeably herein. For example, an image or frame of color video data may include a luma plane or luma component (i.e., luma pixel values) and two chroma planes or chroma components (i.e., chroma pixel values) at the same or different resolutions with respect to the luma plane. The input video 111 may include images or frames that may be divided into blocks and / or encoding units of any size including data corresponding to, for example, M×N blocks and / or encoding units of pixels. Such blocks and / or encoding units may contain data from one or more levels or color channels of pixel data. As used herein, the term block may include macroblocks, encoding units, or the like, of any suitable size. As will be appreciated, such blocks for prediction, transformation, etc. may also be divided into sub-blocks.FIG. 2 illustrates an example set of images 200 configured in accordance with at least some implementations of the present disclosure. As shown in FIG. 2, the group of images 200 may include any number of images 201, such as 64 images (where 0-16 are shown) or the like. Moreover, the images 201 may be provided in a temporal order 202 such that the images 201 are presented in a temporal order while the images 201 are encoded in an encoding order (not shown) such that the encoding order is different with respect to the temporal order 202. Moreover, the images 201 can be provided in an image hierarchy 203 such that a base layer (L0) of images 201 contains the images 0, 8, 16, etc.; a non-base layer (L1) of images 201 contains the images 4, 12, etc.; a non-base layer (L2) of images 201 contains the images 2, 6, 10, 14, etc.; and a non-base layer (L3) of images 201 contains the images 1, 3, 5, 7, 9, 11, 13, 15, etc. For example, when moving through the cross mode hierarchy, the images of L0 may only refer to other images of L0, the images of L1 may only refer to images of L0, the images of L2 may only refer to images of L0 or L1, and the images of L3 may refer to images of arbitrary layers of L0-L2. For example, the images 201 include base layer images and non-base layer images such that base layer images are reference images for non-base layer images, but non-base layer images are not reference images for base layer images, as shown. In one embodiment, the input video 111 includes the group of images 200 and / or the system 100 implements the group of images 200 with respect to the input video 111. Although illustrated with respect to the example group of images 200, the input video 111 may have any suitable structure implementing the group of images 200, another format of a group of images, etc. In one embodiment, a prediction structure for video coding includes groups of images, such as the group of images 200. For example, in the context of broadcast and streaming implementations, the prediction structure may be periodic and may include periodic groups of pictures (GOPs). In one embodiment, a GOP includes about 1 second of pictures organized in the structure described in Figure 2, followed by another GOP beginning with an I-picture, and so on.FIG. 3 illustrates an example video image 301 configured in accordance with at least some implementations of the present disclosure. The video image 301 may be any image of a video sequence or clip, such as a VGA, HD, full HD, 4K, 8K video image, etc. For example, the video image 301 may be any of the images 201. As shown, the video image 301 may be segmented or partitioned into one or more portions, as illustrated with respect to portion 302 of the video image 301. Moreover, the video image 301 may be segmented or partitioned into one or more LCUs as illustrated with respect to LCU 303, which in turn may be segmented into one or more encoding units as illustrated with respect to CUs 305, 306 and / or prediction units (PUs) and transform units (TUs) not shown. As used herein, the term partition may refer to a CU, a PU, or a TU. Although illustrated with respect to portion 302, LCU 303, and CUs 305, 306, which corresponds to HEVC encoding, the techniques discussed herein may be implemented in any encoding context. As used herein, an area may include a portion, an LCU, a CU, an image, or an area of an image. Moreover, a partition as used herein includes a portion of a block, region, or the like. For example, in the context of HEVC, a CU is a partition of an LCU. However, a partition may be any sub-region of a region, a sub-block of a block, etc. The terminology according to HEVC is used herein for clarity of illustration, but is not intended to be limiting.FIG. 4 is an illustrative diagram of an example partitioning and mode decision module 101 for providing LCU partitions and internal / cross mode data 112 configured in accordance with at least some implementations of the present disclosure. For example, FIGS. 4 and 5 illustrate an example embodiment of a decoupled encoder, while FIG. 6 illustrates an example embodiment of an integrated encoder. Both embodiments may be used in implementing the techniques discussed herein.As shown in FIG. 4, the partitioning and mode decision module 101 may include or implement an LCU loop 421 including a source sample motion (SS) estimation module 401, an SS-internal search module 402, a fast CU loop processing module 403, a full CU loop processing module 404, an overlapping depth decision module 405, an internal / overlapping 4×4 refinement module 406, and an exhaust / merge decision module 407. As shown, the LCU loop 421 receives the input video 111 and the LCU loop 421 generates the final LCU partitioning data and initial mode decision data 418. Final LCU partitioning data and initial mode decision data 418 may be any suitable data that indicates or describes the partitioning for the LCU into CUs and a coding mode decision for each CU of the LCU. In one embodiment, final LCU partitioning data and initial mode decision data 418 include final partitioning data that is implemented without changes by the encoder 102 and initial mode decisions that can be changed. For example, the encoding mode decisions may include an internal mode (i.e., one of the available internal modes based on the standard that is implemented) or an inter-mode (i.e., skip, merge, or motion estimate, ME). In addition, the LCU partitioning data and mode decision data 418 may include any additional data required for the particular mode (e.g., a motion vector for a cross mode). For example, in the context of HEVC, a coding tree unit may include 64×64 pixels that may define an LCU. An LCU may be partitioned into CUs for encoding by quad-tree partitioning such that the CUs may be 32×32, 16×16 pixels, or 8×8 pixels. Such partitioning may be indicated by LCU partitioning data and mode decision data 418. Moreover, such partitioning is used to evaluate candidate partitions (candidate CUs) of an LCU.As shown, the SS motion estimation module 401 captures the input video 111, and the SS motion estimation module 401 performs a motion search for CUs or candidate partitions of a current image of the input video 111 using one or more reference images of the input video 111 such that the reference images only include original pixel samples of the input video 111. As shown, the SS motion estimation module 401 generates motion estimation candidates 411 (i.e., MVs) corresponding to CUs of a particular partitioning of a currently evaluated LCU. For example, one or more MVs may be provided for each CU. In addition, the SS-internal search module 402 captures the input video 111, and the SS-internal search module 402 generates internal modes for CUs of a current image of the input video 111 using the current image of the input video 111 by comparing the CU with an internal prediction block generated (based on the current internal mode being evaluated) using original pixel samples of the current input video 111. As shown, the SS-internal search module 402 generates internal candidates 412 (i.e., selected internal modes) that correspond to CUs of a particular partitioning of a currently evaluated LCU. For example, one or more internal candidates may be provided for each CU. In one embodiment, a best partitioning decision and corresponding internal and / or overlapping candidates (e.g., having a lowest distortion, lowest rate distortion cost, or the like) are provided from motion estimation candidates 411 and internal candidates 412 for use by the encoder 102 as discussed herein. For example, the following processing may be omitted.In addition, a detector module 408 receives the input video 111 and / or data from the SS motion estimation module 401 and / or the SS internal search module 402. Detector module 408 applies one or more detectors to input video 111 and / or data so received, and detector module 408 generates and provides detection indicators 419 for use by other modules of partitioning and mode decision module 101 and / or encoder 102, as discussed with respect to FIG. 5.The fast CU loop processing module 403 receives motion estimation candidates 411, internal candidates 412, detection indicators 419, and neighbor data 416 and, as shown, generates MV merge candidates, generates advanced motion vector prediction (AMVP) candidates, and makes a CU mode decision. Neighbor data 416 includes any suitable data for spatially adjacent CUs of the current CUs being evaluated, such as internal and / or cross-over modes of the spatial neighbors. The fast CU loop processing module 403 generates MV merge candidates using any one or more suitable techniques. For example, the merging mode may provide motion impairment candidates using MVs from spatially adjacent CUs of a current CU. For example, one or more MVs from spatially adjacent CUs may be provided (e.g., inherited) as MV candidates for the current CU. Moreover, the fast CU loop processing module 403 generates AMVP candidates using any one or more suitable techniques. In one embodiment, the fast CU loop processing module 403 may use data from a reference image and data from neighboring CUs to generate candidate AMVP MVs. Moreover, non-standard-conformity techniques may be used in generating MV merging candidates and / or AMVP candidates. Predictions for the MV merge candidates and AMVP candidates are generated using only source samples.As shown, the fast CU loop processing module 403 makes a coding mode decision for each CU for the current partitioning based on motion estimation candidates 411, internal candidates 412, MV merging candidates, and AMVP candidates. The coding mode decision may be made using one or more of any suitable techniques. In one embodiment, a sum of a distortion measurement and a weighted rate estimate is used to evaluate the internal and cross modes for the CUs. For example, distortion between the current CU and prediction CUs (generated using the corresponding mode) may be determined and combined with an estimated coding rate to determine the best candidates. As shown, a subset 413 of ME, internal, and merging / AMVP candidates may be generated as a subset of all available candidates.The subset 413 of ME, internal and merging / AMVP candidates is provided to the full CU loop processing module 404. As shown, for each coding mode of the subset 413 of ME, internal, and merging / AMVP candidates (i.e., the residue is a difference between the CU and the predicted CU generated using the current mode), the full CU loop processing module 404 performs a forward transform, a forward quantization, an inverse quantization, and an inverse transform to form a reconstructed residue. Subsequently, the full CU loop processing module 404 generates a reconstruction of the CU (i.e., by adding the reconstructed residue to the predicted CU) and measures the distortion for each mode of the subset 413 of ME, internal and merging / AMVP candidates. The mode with optimal rate distortion is selected as the CU modes 414.CU modes 414 are provided to cross depth decision module 405 which may evaluate the available partitions of the current LCU to generate LCU partitioning data 415. As shown, the LCU partitioning data 415 is provided to the 4×4 internal / cross refinement module 406, which may evaluate 4×4 partitions using internal and / or cross modes. For example, prior processing evaluates partitioning down to an encoding unit size of 8×8, and the 4×4 internal / over-across refinement module 406 evaluates 4×4 partitioning and internal and / or over-across modes for such 4×4 partitions in different contexts. As shown, the 4×4 intra / overarching refinement module 406 provides the skip / merge decision module 407, which determines whether the CU is a skip CU or a merge CU, final LCU partitioning data 417 for any CUs having a coding mode corresponding to a merge MV. For example, for a merging CU, the MV is inherited from a spatially adjacent CU and a remainder is sent for the CU. For an egress CU, the MV is inherited from a spatially adjacent CU (as in the merge mode), but no remainder is sent for the CU. As shown, after such merge / skip decisions, the LCU loop 421 provides final LCU partitioning data and initial mode decision data 418.FIG. 5 is an illustrative diagram of an example encoder 102 for generating the bitstream 113 configured in accordance with at least some implementations of the present disclosure. As shown in FIG. 5, the encoder 102 may include or implement an LCU loop 521 (e.g., an LCU loop for an encoding pass) including a CU loop processing module 501 and an entropy encoding module 502. As also shown, the encoder 102 may include a packetizing module 503. As shown, the LCU loop 521 receives the input video 111 and final LCU partitioning data and initial mode decision data 418, and the LCU loop 521 generates quantized transform coefficients, control data, and parameters 513 that may be entropy encoded by the entropy encoding module 502 and packaged by the packetizing module 503 to generate the bitstream 113.For example, the CU loop processing module 501 receives the input video 111, final LCU partitioning and initial mode decision data 418, and detection indicators 419. Based on the final LCU partitioning data and initial mode decision data 418, the CU loop processing module 501 generates internal reference pixel samples for internal CUs (as appropriate), as shown. For example, internal reference pixel samples may be generated using adjacent reconstructed pixel samples (generated via a local decoding loop). As shown, the CU loop processing module 501 generates a prediction CU for each CU as needed using neighbor data 511 (e.g., data from neighbors of the current CU). For example, the cross mode prediction CU may be generated by retrieving previously reconstructed pixel samples for a CU indicated by an MV or MVs of one or more reconstructed reference images and, if necessary, combining the retrieved reconstructed pixel samples to generate the prediction CU. For internal modes, the prediction CU may be generated using the adjacent reconstructed pixel samples from the image of the CU based on the internal mode of the current CU. As shown, a remainder is generated for the current CU. For example, the remainder may be generated by differentiating the current CU and the prediction CU.The remainder is then forward transformed and forward quantized to produce quantized transform coefficients that are included in the quantized transform coefficients, control data, and parameters 513. Moreover, in a local decoding loop, for example, the transform coefficients are inversely quantized and inversely transformed to generate a reconstructed residual for the current CU. As shown, the CU loop processing module 501 performs a reconstruction for the current CU, e.g., by adding the reconstructed residue and the prediction CU (as discussed above) to generate a reconstructed CU. The reconstructed CU may be combined with other CUs to reconstruct the current image or portions thereof using additional techniques such as sample adaptive offset (SAO) filtering, which may include generating SAO parameters (included in the quantized transform coefficients, control data, and parameters 513) and implementing the SAO filter on reconstructed CUs, and / or deblock loop filtering (DLF) which may include generating DLF parameters (included in the quantized transform coefficients, control data, and parameters 513) and implementing the DLF filter on reconstructed CUs. Such reconstructed CUs may be provided as reference images (e.g., stored in a reconstructed image buffer), for example. Such reference images or portions thereof are provided as reconstructed samples 512 which are used for generating prediction CUs (in cross and internal modes) as discussed above.As shown, the quantized transform coefficients, the control data and parameters 513, the transform coefficients for residual encoding units, control data such as final LCU partitioning and mode decision data (i.e., final LCU partitioning and initial mode decision data 418), and parameters such as SAO / DLF filter parameters, may be entropy encoded and packaged to form bitstream 113. Bitstream 113 may be any suitable bitstream such as a standard compliant bitstream. For example, bitstream 113 may be standard compliant to H.264 / MPEG-4 advanced video coding (AVC), standard compliant to H.265 high efficient video coding (HEVC), standard compliant to VP9, etc.FIG. 6 illustrates a block diagram of an example integrated coding system 600 configured in accordance with at least some implementations of the present disclosure. As discussed, encoding system 600 provides an integrated encoder implementation such that LCU partitions and internal / cross-over mode data 112 may be determined using detection indicators 419, and an evaluation of partitions and encoding modes using reconstructed pixel data. For example, the encoding system 600 may implement the system 100 discussed herein. As shown, the encoding system 600 may include the detector module 408, a controller 601, a motion estimation and compensation module 602, an internal prediction module 603, a deblocking (deblocking) and sample adaptive offset (SAO) module 605, a selector 607, a differentiator 606, an adder 608, a transform (T) module 609, a quantization (Q) module 610, an inverse quantization (IQ) module 611, an inverse transform (IT) module 612, an entropy coder (EE) module 613, and a frame buffer 604 for storing reconstructed frames 114. The encoding system 600 may include additional modules and / or connections, which are not shown for clarity of illustration.As shown, encoding system 600 captures input video 111 and encoding system 600 generates bitstream 113, which may have any characteristics, as discussed herein. For example, encoding system 600 divides images of input video 111 into LCUs, which in turn are partitioned into candidate partitions. After evaluating such candidate partitions, a partitioning decision for the LCU and coding mode decisions for partitions of the individual block corresponding to the partitioning decision are generated by the controller 601 as LCU partitions and internal / cross-over mode data 112 provided to other components of the coding system 600 for encoding the LCU and inclusion in the bitstream 113. As shown, the detection indicators 419 are used in the generation of LCU partitions and internal / cross mode data 112 and bitstream 113 for improved efficiency, as discussed below.Still referring to FIG. 6, the encoding system 600 may perform an LCU loop analogous to the LCU loop 421 via the motion estimation and compensation module 602, which receives the input video 111 and reconstructed images 114 (not shown in FIG. 6 ) and performs motion estimation for candidate CUs or partitions of a current LCU of an image of the input video 111 using one or more reference images of reconstructed images 114 such that the reference images include reconstructed pixel samples (e.g., after a local decoding loop 614 is applied), as generated by the motion estimation and compensation module 602 and the controller 601 analogous to motion estimation candidates 411. Moreover, the internal prediction module 603 receives the input video 111, and the reconstructed picture sample post-adder 608, the internal prediction module 603, and the controller 601 generate internal modes for CUs of a current picture of the input video 111 using the current picture of the input video 111 by comparing the CU to an internal prediction block generated using the reconstructed pixel sample (e.g., after the local decoding loop 614 has been applied). The controller 601 then generates an LCU partition decision (e.g., defining partitioning of an LCU) and a corresponding mode decision for each partition (e.g., an internal or a cross mode for each partition) to encode the LCU. For example, for each partition (e.g., CU) of a block (e.g., LCU), controller 601 may control selector 607 using LCU partitions and internal / cross mode data 112 to generate a predicted partition (e.g., CU) for each block (e.g., LCU) based on the best mode (e.g., the lowest cost mode) for the partition (e.g., CU).After the decision as to whether a partition is encoded internally or across (and the corresponding mode from the internal and across candidates) has been made, a difference to the source pixels is formed by means of the differentiator 606. For example, a difference is formed between a partition (e.g., CU) or a block (e.g., LCU) of original pixel samples from the input video 111 and a reconstructed partition or block at a differentiator 606 such that the reconstructed partition or block is generated using the corresponding best coding mode as implemented by a local decoding loop 614. The difference (e.g., the residual partition or residual block) is converted to the frequency domain (e.g., using a discrete cosine transform or another transform) by a transform module 609 to generate transform coefficients, and the transform coefficients are quantized by the quantization module 610 to generate quantized transform coefficients. Transform coefficients thus quantized are entropy encoded together with various control signals (including LCU partitions and internal / overlapping mode data 112) by means of the entropy encoder module 613 to generate the bitstream 113, which may be sent to a decoder or stored in a data store. In addition, the quantized transform coefficients from the quantization module 610 are inversely quantized by the inverse quantization module 612 and inversely transformed by the inverse transform module 612 to generate reconstructed differences or residual partitions or residual blocks. The reconstructed partitions or blocks are combined with reference blocks (e.g., reconstructed reference blocks as selected via selector 607) by adder 608 to generate reconstructed partitions or blocks which are provided to internal prediction module 603 for use in internal prediction as shown. Moreover, the reconstructed differences or residues may be filtered using the deblocking and sample adaptive offset module 605 with a deblocking filter and / or a sample adaptive offset filter, reconstructed into a picture, and stored in the picture buffer 604 for use in cross prediction.As discussed, a decoupled coding system or an integrated coding system may implement the detection indicators 419 for improved data usage efficiency. The discussion now addresses detected features, indicators, and implementations thereof.FIG. 7 is a flow diagram illustrating an example process 700 for selectively using chroma information in partitioning decisions and coding mode decisions, and configured in accordance with at least some implementations of the present disclosure. The process 700 may include one or more operations 701- 710, as shown in FIG. 7. Process 700 may be performed by a system (e.g., system 100, encoding system 600, etc.) to improve data usage efficiency by selectively using chroma in partitioning decisions and encoding mode decisions in video encoding. For example, using only luma information offers the advantage of faster processing and lower complexity at the expense of reduced accuracy (e.g., by eliminating chroma from cost computations). Alternatively, using luma information and chroma information offers the advantage of accuracy at the expense of a decreased computational speed (e.g., by adding chroma to cost computations). The process 700 provides a trade-off between computational cost and accuracy by efficiently generating an evaluation decision for luma and chroma or an evaluation decision for luma only for blocks of an image.The process 700 begins at a decision process 701, where detectors are applied to a region, block, or partition of a current image of an input video to generate detected features or detection indicators. For example, the operation 701 may be performed by the detector module 408. As shown, the detected features for an area, block (e.g., LCU), or partition (e.g., CU) include a luma average of the area, block, or partition, an average of a first chroma channel (e.g., Cb channel) of the area, block, or partition, an average of a second chroma channel (Cr average) of the area, block, or partition, an indicator of whether the area, block, or partition includes an edge, an indicator of whether the area, block, or partition is in a revealed area or not, and a temporal layer of the area, block, or partition.The detection indicators determined at operation 701 may be generated using one or more of any suitable techniques. In one embodiment, the luma average is an average of luma values at pixel locations of the region, block, or partition. In one embodiment, the average of the first chroma channel is an average of chroma values at pixel locations of the region, block, or partition for a first chroma channel and the average of the second chroma channel is an average of chroma values at pixel locations of the region, block, or partition for a second chroma channel. For example, the pixels may include one luma component and two chroma components such as Cb and Cr components, although any suitable color space may be implemented. Although discussed with respect to averages for all pixel locations, in some embodiments, some pixel values (e.g., high and low, outliers, etc.) may be discarded prior to generating the averages. An edge feature for the region, block, or partition may be detected using any suitable edge detection techniques such as Canny edge detection.Moreover, whether or not the area, block, or partition is in a revealed area is intended to detect the areas, blocks, or partitions that are in areas that have been revealed because there is some motion in the input video 111. For example, a moving person would uncover a revealed area that was previously behind it. Such a determination as to whether the region, block, or partition is in a revealed region may be made using one or more suitable techniques. In one embodiment, a difference between a best motion estimation sum of difference amounts (SAD) and a best internal prediction SAD is taken for the region, block, or partition, and if the best internal prediction SAD plus a threshold is less than the best motion estimation SAD, the region, block, or partition is displayed as being in a revealed area. For example, adding a threshold, bias, or the like to the best internal prediction SAD and the sum being less than the best motion estimation SAD may indicate that the internal prediction SAD is much less than the best motion estimation SAD, which in turn indicates that the region, block, or partition is in a revealed region because accurate motion estimation compensation for the block cannot be found. For example, the best motion estimation SAD may be the SAD corresponding to the best motion estimation mode as determined by the SS motion estimation module 401 or the motion estimation and compensation module 602, and the best internal prediction SAD may be the SAD corresponding to the best internal mode as determined by the SS internal search module 402 or the internal prediction module 603. That is, either an open-loop prediction SAD (using only original pixel samples) or a closed-loop prediction SAD (using reconstructed pixel samples) may be used.Processing continues with decision operation 702 where the luma average of the region, block, or partition, the average of the first chroma channel of the region, block, or partition, and the average of the second chroma channel of the region, block, or partition are compared to corresponding thresholds. If each of the luma average of the range, block or partition, the average of the first chroma channel of the range, block or partition, and the average of the second chroma channel of the range, block or partition does not exceed its corresponding threshold (e.g., if they are inconvenient compared to the thresholds), processing continues to operation 703. Although discussed herein with respect to detection indicators of averages of blocks being compared to thresholds, in further embodiments, the detection indicators include indicators of whether or not each of the averages exceeds, reaches, or exceeds (e.g., is beneficial compared to the threshold) the corresponding threshold (e.g., 1 or 0, true, or false). In such embodiments, the decision process 702 may simply determine whether any such indicators are false.In operation 703, only luma information is used for partitioning decisions and encoding mode decisions for the current region, block (e.g., LCU), or partition (e.g., CU). For example, comparing an original partition or block to a predicted partition or block (predicted using only original pixel samples or predicted using reconstructed pixels), only luma pixel values are used while chroma pixel values are discarded. That is, when distortion measurements, comparisons, etc. are made between the block or partition and a predicted block or partition, only luma information is used. Such techniques may be implemented using any suitable modules or components discussed herein that participate in partitioning decisions and coding mode decisions for the current block, such as modules 401- 407 of LCU loop module 402 and / or modules 601- 612 of coding system 600. Such modules are not expressly listed here for clarity of illustration. That is, any operation used in partitioning decisions and coding mode decisions may only operate on luma information (e.g., samples) while chroma information is discarded. It should be noted that modules and operations associated with encoding operations to generate bitstream 113 to generate bitstream 113 still operate on both luma information and chroma information (e.g., both luma residues and chroma residues are formed, etc.). For example, the CU loop processing module 501 operates on both luma and chroma to generate quantized transform coefficients from quantized transform coefficients, control data, and parameters 513. Moreover, in the context of the encoding system 600, the modules 602, 603, 606, 609, 610 operate to generate the bitstream 113 on luma and chroma information. Such modules may therefore only use luma in the context of partitioning decisions and coding mode decisions, while both use luma in the context of generating bitstream 113 as needed. For example, such modules discard chroma in the context of partitioning decisions and coding mode decisions to save substantial computational resources, and then use chroma information as needed to use such partitioning decisions and coding mode decisions to generate bitstream 113. For example, it may be advantageous to disregard chroma for relatively dark blocks to save computational resources.Returning to decision operation 702, if the luma average of the region, block, or partition, the average of the first chroma channel of the region, block, or partition, or the average of the second chroma channel of the region, block, or partition exceeds or reaches its corresponding threshold (e.g., is favorably compared to the threshold), processing continues at decision operation 704, where a determination is made whether the region, block, or partition contains an edge (as determined based on the edge detection indicator of operation 701) and / or whether the region, block, or partition is in a revealed region (as determined based on the revealed region indicator of operation 701).If either is true, processing continues at operation 705, using luma and chroma information for partitioning decisions and coding mode decisions for the current region, block (e.g., LCU), or partition (e.g., CU). For example, in comparing an original partition or block to a predicted partition or block (either predicted using only original pixel samples or predicted using reconstructed pixels), both luma and chroma pixel values are used. That is, when distortion measurements, comparisons, etc. are made between the partition or block and a predicted partition or block, both luma and chroma information are used. Such techniques may be implemented using any suitable modules or components discussed herein that participate in partitioning decisions and coding mode decisions for the current block, such as modules 401- 407 of LCU loop module 402 and / or modules 601- 612 of coding system 600. For example, each operation used in partitioning decisions and coding mode decisions may operate using luma as well as chroma information (e.g., samples). For example, for increased accuracy or noise reduction, it may be advantageous to use both luma and chroma information for blocks that have edges or that are in revealed areas.Returning to decision process 704, if the region, block, or partition does not contain an edge and is not in a revealed region, processing continues at decision process 706, and a determination is made as to whether the region, block, or partition is part of an I-slice or an I-picture. For example, an I-portion or an I-picture may be any portion or picture encoded without reference to another picture. Referring to FIG. 2, an I-portion or image may be image 0 or a portion of image 0 or a portion of any other image encoded without reference to another image. As shown, if the region, block, or partition is part of an I-slice or image, then process 700 continues at operation 703, where only luma information is used for partitioning decisions and coding mode decisions for the current region, block (e.g., LCU), or partition (e.g., CU), as discussed.Then, if the region, block, or partition is not part of an I-slice or an I-picture, process 700 continues at decision process 707 where a determination is made as to whether the region, block, or partition is part of a base layer B-slice or a base layer B-picture. For example, a base layer B portion or a base layer B image may be any portion or image that is part of a base layer (e.g., that only references other images in the same base layer but does not reference non-base layer images). Referring to FIG. 2, a base layer B portion or a base layer B image may be image 8, 16,..., such that the base layer B portion or the base layer B image may be reference I image 0 or other base layer B images but not non-base layer B images. As shown, if the region, block, or partition is part of a base layer B-portion or base layer B-image, then process 700 continues at operation 705, using luma and chroma information for partitioning decisions and coding mode decisions for the current region, block (e.g., LCU), or partition (e.g., CU), as discussed above.Then, if the region, block, or partition is not part of a base layer B portion, the process 700 continues at decision process 708, where a determination is made as to whether the region, block, or partition is part of a non-base layer B portion or non-base layer B image. For example, a non-base layer B portion or a non-base layer B image may be any portion or image that is part of a non-base layer (e.g., refers to other images in the base layer, the same layer, and deeper layers, but is not related to base layer images or deeper layers). Referring to FIG. 2, a non-base layer B portion or non-base layer B image may be an L1 non-base layer image (4, 12,...), a layer L2 non-base layer image (2, 6, 7, 14,...) or a layer L3 non-base layer image (1, 3, 5, 7, 8, 11, 13, 15,...) such that the non-base layer B portion or non-base layer B image may refer to images in the same or lower layers. Then, if the region, block, or partition is not part of a non-base layer B portion or a non-base layer B image, processing is completed at operation 710.As shown, if the region, block, or partition is part of a non-base layer B portion or non-base layer B image, then process 700 continues at operation 709, where both luma information and chroma information are used only for merge / skip decisions to merge / skip decisions, while further partitioning decisions and coding mode decisions are made using only luma information and without using chroma plane or chroma component. For example, in evaluating the internal, cross, AMVP, and candidate merge modes for the block, only luma pixel samples or luma values are used (and chroma pixel samples or values are discarded). This means that when distortion measurements, comparisons, etc. are made between a partition or a block and a predicted partition or a predicted block for the above modes, only luma information is used. Then, when the selected encoding mode is a candidate merging mode, the decision is made between the merging mode (using the merging MV and the transmission of residual data) and the exhaust mode (using the merging MV and the transmission of no residual data) using the luma information and chroma information. For example, the merge / exit decision may be performed as a last step in the partitioning and mode decision process to decide whether to encode a partition or block as a merge partition or block or as an exit partition or block. Such merging / dropping decision can be made by comparing the cost of the two modes such that the cost is obtained using distortion values and coding rate estimates for each of the two modes. For the skip mode, the distortion is assumed to be zero, such that no transform coefficients need to be encoded (nevertheless, the distortion used to determine the cost of the skip mode is not zero). For the merge mode, the distortion is a measure of the difference between the partition or block being encoded and the predicted partition or block. In the context of operation 709, such a distortion measurement is generated using both luma pixel samples or luma pixel values and chroma pixel samples or chroma values.For example, in the SS motion estimation module 401, the SS internal search module 402, the fast CU loop processing module 403, the full CU loop processing module 404, and the cross depth decision module 405, only luma pixel values are used for a block that is part of a non-base layer B portion or a non-base layer B image. However, in the skip / merge decision module 407, both luma and chroma pixel values are used when a block is part of a non-base layer B portion or non-base layer B image.Similarly, in the control unit 601, the motion estimation and compensation module 602, in the internal prediction module 603, and in corresponding modules used for reconstruction, only luma pixel values are used for partitioning decisions and coding mode decisions other than the skip / merge decision, and in the control unit 601, both luma pixel samples and chroma pixel samples are used for skip / merge decision for a block that has been set as a merge encoded block.The chroma inclusion techniques discussed herein with respect to process 700 and elsewhere may offer reduced processing requirements (since chroma is not used for all mode decisions) while minimizing spurious signals caused by complete removal of the chroma's use. For example, removing the use of chroma information can result in visible spurious signals such as color tracking or color bleeding and blocking artifacts. The techniques discussed herein may reduce or remove interference signals. The described techniques may provide switching between different chroma processing modes corresponding to varying levels of chroma information usage. For example, such switching is based on luma levels and chroma levels, edge detection, revealed area detection, and time-slice information (e.g., base-slice information or non-base-slice information) such that images having different detected features use different amounts of chroma information. In one embodiment, the various modes are defined as full chroma, chroma only for merging / dropping decision, or no chroma. For full chroma, all cost computations used in mode decisions for partitioning decisions and coding mode decisions use the full chroma data (e.g., 4:2:0 input video). In chroma for merging / dropping only, chroma information is used for only the purpose of deciding between the merging mode and the dropping mode, so that the decision between the merging mode and the dropping mode for a merging candidate is made based on the cost associated with each candidate using full luma information and chroma information. If chroma is not turned off as the name indicates, chroma information is not used for mode decisions. In one embodiment, chroma is not used for I-sections or I-pictures; full chroma is used for base layer B-sections; and chroma is used for merge / skip decision only for non-base layer B-sections.FIG. 8 is a flow diagram illustrating an example process 800 for generating a merge mode decision or exit mode decision for a partition having an initial merge mode decision, configured in accordance with at least some implementations of the present disclosure. Process 800 may include one or more of operations 801- 805, as shown in FIG. 8. Process 800 may be performed by a system (e.g., system 100, encoding system 600, etc.) to improve data usage efficiency by generating a merge mode decision or exhaust mode decision for a partition based on initial merge mode encoding cost and exhaust mode encoding cost. For example, using initial merge mode encoding cost and exit mode encoding cost offers the advantage of faster processing and less complexity during encoding.Process 800 begins at decision process 801, where detectors are applied for a block or region of a current image of the input video to generate detected features or detection indicators. For example, the operation 801 may be performed by the detector module 408. As shown, the detected feature for a block (e.g., an LCU) or a partition of a block (e.g., a CU) includes an amount of difference between initial exhaust mode encoding cost and initial merge mode encoding cost for a partition having an initial merge mode decision. For example, in the context of a decoupled encoder, fast CU loop processing module 403 and / or full CU loop processing module 404 may determine that an initial best coding mode decision for a partition (e.g., CU) is a merge mode. Such a merge mode indicates that motion interference candidates are to be used to encode the partition using MVs from spatially adjacent partitions (e.g., CUs) of a partition (e.g., CU). Both the exhaust and merge coding modes use the derived MV, but the exhaust and merge coding modes differ in that no remainder is sent for the partition in the exhaust mode, while a remainder is sent for the partition in the merge mode. In the context of an integrated coding system, the controller 601, motion estimation and compensation module 602, and internal prediction module 603 may determine an initial merging mode for the partition while again delaying the determination of whether to use the skip or merging mode for the initial merging mode partition.The detection indicator determined at operation 801 may be generated using one or more suitable techniques. In one embodiment, initial merge encoding cost and initial exhaust mode encoding cost for the partition are determined using original pixel samples, approximated reconstructed pixel samples, etc. In one embodiment, the encoding cost is rate distortion cost. As shown, processing continues at decision process 802, where the amount of difference between the initial merge mode encoding cost and the exhaust mode encoding cost is compared to a threshold. As shown, if the amount of difference exceeds the threshold (or meets or exceeds the threshold or is beneficial compared to the threshold), processing continues at operation 803 where the mode having the lower encoding cost is selected for encoding. Moreover, the comparison of the encoding cost at full encoding (e.g., as performed by the CU loop processing module 501) is skipped in response to the amount of difference exceeding the threshold, for example. Such techniques provide the efficiency advantages as full coding loop operations are reduced in such contexts.Returning to decision operation 802, if the magnitude of the difference does not exceed the threshold (e.g., is adverse compared to the threshold), processing continues at operation 804 in which the skip mode decision or merge mode decision is moved to a full encoding pass. That is, the skip mode decision or merge mode decision is not made based on the initial encoding cost. Instead, processing continues at operation 805, where only the exhaust or merge modes for the partition (e.g., CU) are evaluated at a full encode pass and the lower cost mode is selected for encoding. For example, the complete encoding pass may be performed by the CU loop processing module 501 in the context of a decoupled encoder. In the context of an integrated coding system, the complete coding pass may be performed by the control unit 601 and the motion estimation and compensation module 602. In any case, during the complete coding pass, only the exit and merge modes for the partition are evaluated, such that the evaluation of further overlapping and / or internal modes is skipped. The skip or merge mode evaluation at the full encoding pass may be performed using one or more of any suitable techniques such as differentiating the partition (e.g., CU) with a reconstructed partition (e.g., CU) reconstructed using a local decoding loop based on the merge mode candidate motion vector, transforming the resulting residual, quantizing the transformed residual coefficients, and generating skip mode costs associated with not including the resulting transformed residual coefficients in the encoding, and merge mode costs associated with including the resulting transformed residual coefficients in the encoding. As discussed, such costs may be rate distortion costs that include distortion costs and rate costs of the modes. The resulting costs may then be compared and the mode corresponding to the lower cost selected than the final mode for the partition (e.g., CU). The partition is then encoded into bitstream 113 using the resulting final mode.The merge mode selection techniques or skip mode selection techniques discussed herein with respect to process 800 and elsewhere may offer decreased processing requirements (in view of cases where initial mode cost indicates use of the merge mode or skip mode, full coding pass evaluation is skipped), while interfering signals caused by eliminating use of such full coding pass evaluation in cases where selection of the merge mode or skip mode is not resolved using the initial cost are minimized.FIG. 9 is a flow diagram illustrating an example process 900 for determining a partitioning decision and encoding mode decisions for a block by generating only a portion of a transform coefficient block for a portion of the block, configured in accordance with at least some implementations of the present disclosure. The process 900 may include one or more of operations 901- 905, as shown in FIG. 9. Process 900 may be performed by a system (e.g., system 100, encoding system 600, etc.) to improve data usage efficiency by reducing the computations of transform coefficients. For example, by reducing the number of transform coefficients available, if transformations are performed during partitioning decisions and coding mode decisions, the computations are reduced for more efficient processing.Process 900 begins at decision operation 901, where a difference is formed between a partition (e.g., CU or PU) and a predicted partition (e.g., CU or PU). The predicted partition may be generated using one or more of any suitable techniques. For example, the predicted partition may be a predicted candidate partition corresponding to a candidate coding mode for a candidate partition (e.g., CU) of a block (e.g., LCU). The predicted partition is generated using internal or cross techniques based on the appropriate test coding mode. In one embodiment, in the context of a decoupled decoder, the predicted partition may include a partition that was generated using only original pixel samples as discussed herein. In further embodiments, the predicted partition is generated using reconstructed pixel samples. In either case, a difference is formed between the partition of the input video and the predicted partition to generate a residual partition. A difference may be formed between the predicted partition (e.g., CU or PU) and an original partition to generate a residual partition. The remaining partitions may then be further partitioned into partitions (e.g., TUs) for purposes of transformation processing. For example, the discussed difference formation may be performed at the CU level or PU level, with the subsequent transformation processing performed at the TU level.Processing continues at operation 902 in which a partial transformation is performed on the remaining partitions (e.g., TU) to generate transformation coefficients such that the number of available transformation coefficients is less than the number of remaining values in the remaining partitions (e.g., TU). For example, the residual partition has 64 values (although some may be zero) if the residual partition (e.g., TU) is an 8×8 partition. In such an example, the number of available transform coefficients after a subtransform is less than 64, such as 36 (e.g., for a 6×6 transform coefficient block), 16 (e.g., for a 4×4 transform coefficient block), etc. As with the available residual values, some of the transform coefficient values determined using the subtransform may be zero; however, such values are still available when applying the subtransform. Those transform coefficients that are not determined as part of the application of the subtransform may be set to zero. For example, the application of the subtransform may compute some available transform coefficient values to zero and those that are not available are set to zero such that a resulting transform coefficient block has the same number of values as the residual partition (e.g., TU). In one embodiment, at operation 902, a transform coefficient block is generated based on a residual partition (e.g., TU) of operation 901 by performing a partial transform on the residual partition to generate transform coefficients from a portion of the transform coefficient block such that the number of transform coefficients in the portion is less than a number of values of the residual partition, and setting the remaining transform coefficients of the transform coefficient block to zero.The subtransformation performed in operation 902 may be performed using one or more arbitrary techniques. In one embodiment, performing the subtransform includes applying only those transformation calculations necessary to generate transformation coefficients for those coefficients that are to be available after the subtransformation, while skipping those transformation calculations necessary to generate transformation coefficients that are not to be available after the subtransformation. The subtransform discussed herein may be characterized as a subfrequency transform, a constrained transform, a constrained frequency transform, a reduced frequency transform, or the like.FIG. 10 illustrates an example data structure 1000 corresponding to an example subtransform 1010 configured in accordance with at least some implementations of the present disclosure. As shown in FIG. 10, an example 4×4 residue block 1001 includes 16 available residue values (labeled R11, R12, R13,..., R44). For example, the residual block 1001 may be a partition such as TU. Although illustrated herein with respect to a 4×4 residual block 1001, the residual block 1001 may be of any suitable size, such as 8×8, 16×106, 32×32, etc. As also shown in FIG. 10, the subtransform 1010 transforms the residual values of the residual block 1001 into the frequency domain such that the resulting transform coefficient block 1002 has fewer available transform coefficients 1003 (in the illustrated example 4, denoted as tc11, tc12, tc21, tc22) than the number of available residual values of the residual block 1001. In the illustrated example, the residual block 1001 has 16 available residual values and the transform coefficient block 1002 has 4 available transform coefficient values. However, the number of available residual values and the number of available transform coefficient values after the sub-transformation may be any suitable values as long as the number of available transform coefficient values is less than the number of available residual values. In one embodiment, the residual block 1001 is an 8×8 block and the transform coefficient block 1002 is a 4×4 block. In one embodiment, the residual block 1001 is a 16×16 block and the transform coefficient block 1002 is an 8×8 block. In one embodiment, the residual block 1001 is a 32×32 block and the transform coefficient block 1002 is a 16×16 block.Moreover, FIG. 10 illustrates unavailable transform coefficient values 1004 that are unavailable due to the application of a subtransform rather than a full transform. As shown, coefficients 1003 available after a subtransform 1010 may be those in an upper left corner of the full transform coefficients from a full transform. Such available transform coefficients 1003 hold lower frequency information in transform coefficient block 1002, while higher frequency information is effectively discarded. Such techniques can provide more accurate representations of lower frequency residual blocks. However, the available transform coefficients 1003 may be any part of the full transform coefficients from a full transform and may correspond to any frequency transform coefficients.FIG. 11 illustrates an example data structure 1100 corresponding to another example subtransform 1110 configured in accordance with at least some implementations of the present disclosure. As shown in FIG. 11, the subtransform 1110 transforms the residual values of the residual block 1001 into the frequency domain such that the resulting transform coefficient block 1102 has fewer available transform coefficients 1103 (in the illustrated example 9, denoted tc11, tc12, tc13,..., tc33) than the number of available residual values of the residual block 1001. In the illustrated example, the residual block 1001 has 16 available residual values and the transform coefficient block 1102 has 9 available transform coefficient values. However, the number of available residual values and the number of available transform coefficient values after the sub-transformation may be any suitable values as long as the number of available transform coefficient values is less than the number of available residual values. In one embodiment, the residual block 1001 is an 8×8 block and the transform coefficient block 1102 is a 6×6 block. In one embodiment, the residual block 1001 is a 16×16 block and the transform coefficient block 1102 is a 12×12 block. In one embodiment, the residual block 1001 is a 32×32 block and the transform coefficient block 1002 is a 16×16 block. Moreover, FIG. 11 illustrates unavailable transform coefficient values 1104 that are unavailable due to the application of a subtransform rather than a full transform, as discussed with respect to FIG. 10. In addition, available transform coefficients 1103 after subtransform 1100, as discussed with respect to FIG. 10, may be those in an upper left corner of the full transform coefficients from a full transform.As has been shown with reference to FIGS. 10 and 11, the application of the subtransforms 1001, 1100 may have different levels of reduction of the available transform coefficient values 1003, 1103. For example, subtransform 1001 reduces the number of available transform coefficient values 1003 to 4, while subtransform 1100 reduces the number of available transform coefficient values 1103 to 9. Due to such variations in the number of transform coefficient values available, the transform coefficient block 1102 represents the residual block 1001 better compared to the transform coefficient block 1002 because the transform coefficient block 1102 has more lost information. Therefore, subtransform 1010 may be described as more aggressive or lossy as compared to subtransform 1110, which may be described as more moderate, less aggressive, or less lossy. As discussed further herein with respect to FIGS. 12 and 13, more or less aggressive subtransformations may be performed for residual partitions or residual blocks depending on detected features or characteristics of the blocks corresponding to the residual partitions or residual blocks.Returning to FIG. 9, processing continues at operation 903, where the transform coefficients generated at operation 902 are quantized to generate quantized transform coefficients. The transform coefficients may be quantized using any one or more suitable techniques. The number of quantized transform coefficients is equal to the number of transform coefficients. Therefore, the number of available quantized transform coefficients is also less than the number of residue values in the residue partition generated in operation 902.Processing continues at operation 904 where the quantized transform coefficients are inversely quantized and inversely transformed to generate a reconstructed residual partition (e.g., TU). The inverse quantization and inverse transformation may be performed using any one or more suitable techniques that reverse operations performed in operations 902, 903. For example, an inverse transform may be performed to generate a reconstructed residue block (e.g., TU) having the same number of available values as the number of residues generated in operation 901. For example, the inverse transform may take into account the fact that some of the inverse quantized coefficients are zero to reduce the number of computations, but the reconstructed residues may be the complete array of residues. As such, the reconstructed residual block has the same number of values and the same block shape (e.g., size) as the residual block generated in operation 902. In one embodiment, multiple TUs may be combined to form a CU or a PU.Processing continues at operation 905, where the reconstructed residual block generated at operation 904 is added to the predicted partition (as discussed at operation 901) to generate a reconstructed partition (e.g., CU or PU) corresponding to the original partition (as discussed again at operation 901). The reconstructed partition may then be used in partitioning decisions and coding mode decisions for the block of which the partition is a part. For example, process 900 may be repeated for any number of candidate (cross and internal) coding modes and any number of candidate partitions (e.g., candidate PUs or candidate CUs) of a block (e.g., LCU) to select partitioning for the block (e.g., LCU) and coding modes for partitions (e.g., CUs) corresponding to the partitioning, and the best partitioning decision as well as the best or some best coding modes (e.g., to be further evaluated) may be selected for the partitions (e.g., CUs).In another embodiment, operation 904 includes merely inverse quantizing the quantized transform coefficients to generate inverse quantized transform coefficients, which may also be characterized as reconstructed transform coefficients. The reconstructed transform coefficients may then be compared to the transform coefficients generated in operation 902 for purposes of partitioning decisions and encoding mode decisions. For example, distortions may be generated based on a sum of the squares of the differences between the transform coefficients from operation 902 and the output of the inverse quantization (i.e., the reconstructed transform coefficients). For example, the reconstructed transform coefficients may be used in partitioning decisions and coding mode decisions for the block of which the partition is a part, as discussed above. In one embodiment, a distortion measure corresponding to the predicted partition (as discussed with respect to operation 901) is generated based on the inverse quantized transform coefficients based on a sum of the squares of differences between the transform coefficients from operation 902 and the output of the inverse quantization (i.e., the reconstructed transform coefficients).The subtransformation techniques discussed with respect to process 900 may be performed for all residue blocks or only for residue blocks in certain contexts. Moreover, the strength of the subtransformation may be varied in certain contexts as discussed herein. The subtransformation techniques save computation time and resources in determining partitioning decisions and selecting coding modes by reducing transformation computations as well as computations of quantization, inverse quantization, and inverse transformation. As discussed, the subtransformation techniques can be used in determining a partitioning decision and encoding mode decisions for a block by generating only a portion of a transform coefficient block for a partition of the block. In some embodiments, full transformations are used for the full coding pass to generate standard-compliant quantized transformation coefficients for inclusion in bitstream 113.FIG. 12 is a flow diagram illustrating an example process 1200 for determining a partitioning decision and encoding mode decisions for a block by generating only a portion of a transform coefficient block for a partition of the block based on whether the partition is in an optically important region, configured in accordance with at least some implementations of the present disclosure. The process 1200 may include one or more of the operations 1201- 1204, as shown in FIG. 12. Process 1200 may be performed by a system (e.g., system 100, encoding system 600, etc.) to improve data usage efficiency by reducing computations of transform coefficients. For example, computations are reduced by reducing the number of available transform coefficients in performing transforms during partitioning decisions and encoding mode decisions for more efficient processing.Process 1200 begins at operation 1301, where detectors are applied to a region, block, or partition of a current image of an input video to generate detected features or detection indicators. For example, operation 1201 may be performed by detector module 408. As shown, the detected features for a region, a block or a partition and / or the detection indicator show whether the region, the block or the partition is an optically important region or is located in one.The determination of whether the region, block, or partition is an optically important region or is in one may be made using one or more of any suitable techniques. In one embodiment, the determination is made based on whether the region, block, or partition contains an edge. For example, edge detection for the region, block, or partition may be performed using one or more of any suitable techniques such as Canny edge detection techniques, and when an edge is detected, the region, block, or partition is displayed as or as being in an optically detected region.In one embodiment, the determination is made based on whether the area, block, or partition is a still background area of a video. Such a determination may be made by determining whether an in-place region, block, or partition, or area including the region, block, or partition, has low distortions (e.g., as measured by a sum of difference amounts, SAD) over frames (e.g., over time over two or more consecutive frames). For example, if the SAD is less than a threshold for one or more temporal previous images and the current image based on the difference between the current region, the current block, the current partition, or the current surface and a co-location region, a co-location block, a co-location partition, or a co-location surface (e.g., predicted using original pixel samples), a determination is made that the region, block, partition, or surface is in a non-moving background and thus an optically important region.In one embodiment, the determination is made based on whether the area, block, or partition is in an Aura area. Such a determination may be made by determining that a motion estimation distortion (e.g., SAD based on a difference between the current region, the current block or partition and a best candidate of a predicted ME region, a predicted ME block or partition) is greater than a first threshold, the best candidate of a motion vector corresponding to the best candidate of the predicted ME has an amount greater than a second threshold, and at least one spatially adjacent region, block or partition of the current region, block or partition has a motion estimation distortion greater than a third threshold, and then the current region, block or partition has an Aura region, an Aura block or an Aura partition and thus identified as an optically important area. For example, if the current region, block, or partition has a motion estimation distortion that is large (e.g., greater than a first threshold), a long motion vector (e.g., having an amount greater than a second threshold), and an adjacent region, block, or partition that also has a large motion estimation distortion (e.g., greater than the first threshold or third threshold), the region, block, or partition is an Aura region, an Aura block, or an Aura partition and is indicated as optically important.As shown, processing continues at decision operation 1202, where a determination is made as to whether the region, block, or partition is optically important or whether the region, block, or partition is in an optically important region. As discussed, if the region, block, or partition includes or is in an area that includes an edge, stationary background, or aura, then the region, block, or partition is displayed as being optically important. If the region, block, or partition is not optically important, processing continues at operation 1203, where a most aggressive or more aggressive subtransformation is applied to partitions (e.g., TUs) of the block (e.g., LCU). For example, the partitions of the block may be subjected to subtransforms and further processing discussed with respect to the process 900 for partitioning decisions and encoding mode decisions, such that the partitions of the block are subjected to more aggressive subtransforms as compared to operation 1204. For example, as discussed with respect to FIGS. 10 and 11, more aggressive transformations may reduce the number of available transformation coefficients more than less aggressive transformations. In one embodiment, the most aggressive subtransforms applied in operation 1203 provide a number of available transform coefficients that is a quarter of the number of residue values of residue blocks. For example, for 4×4 residual blocks (partitions), the most aggressive subtransformation yields 4 transform coefficients, for 8×8 residual blocks (partitions), the most aggressive subtransformation yields 16 transform coefficients, etc.If the region, block, or partition is optically important, processing continues at operation 1204, where a moderate or less aggressive subtransform (as compared to that used in operation 1203) is applied to partitions (e.g., TUs) of the block (e.g., LCU) or no subtransform is applied (e.g., a full transform is applied). For example, partitions of the block may be subjected to subtransforms and further processing discussed with respect to the process 900 for partitioning decisions and encoding mode decisions such that the partitions of the block are subjected to less aggressive subtransforms as compared to operation 1203. For example, as discussed with respect to FIGS. 10 and 11, more aggressive transformations may reduce the number of available transformation coefficients more than less aggressive transformations. As discussed, the most aggressive subtransforms applied in operation 1203 may provide a number of available transform coefficients that is one-fourth the number of residue values of residue blocks. In contrast, the less aggressive subtransforms applied in operation 1204 may provide a number of available transform coefficients that is more than half the number of residue values of residue blocks. For example, the less aggressive subtransforms for 4×4 residue blocks may result in 9 transform coefficients, the less aggressive subtransforms for 8×8 residue blocks may result in 36 transform coefficients, etc.As discussed with respect to operations 1201, 1202, if a region, block, or partition is optically important, then a less aggressive subtransform (or full transform) is applied to partitions (e.g., TUs), and if the region, block, or partition is not optically important, then a more aggressive subtransform (or full transform) is applied to partitions (e.g., TUs) of a current block.FIG. 13 is a flow diagram illustrating an example process 1300 for determining a partitioning decision and encoding mode decisions for a block by generating only a portion of a transform coefficient block for a partition of the block based on edge detection in the block, configured in accordance with at least some implementations of the present disclosure. The process 1300 may include one or more of the operations 1301- 1307, as illustrated in FIG. 13. Process 1300 may be performed by a system (e.g., system 100, encoding system 600, etc.) to improve data usage efficiency by reducing transform coefficient computations. For example, by reducing the number of transform coefficients available in performing transformations during partitioning decisions and coding mode decisions, computations are reduced for more efficient processing.Process 1300 begins at operation 1301, where detectors are applied to a region, block, or partition of a current image of an input video to generate detected features or detection indicators. For example, the operation 1301 may be performed by the detector module 408. As shown, the detected features for an area, a block, or a partition indicate whether the block includes an edge and, if so, an edge strength corresponding to the edge. The determination of whether the region, block, or partition includes an edge may be made using any one or more suitable edge detection techniques such as canny edge detection. If the region, block, or partition includes an edge, the edge strength may be generated using one or more of any suitable techniques. In one embodiment, the edge strength is a variance of the region, the block, or the partition. In one embodiment, the edge strength is a measure of contrast across the edge. In some embodiments, the variance or measurement may be categorized using thresholding to mark the edge as weak (e.g., when the variance or contrast measurement is less than a corresponding threshold), strong (e.g., when the variance or contrast measurement is greater than a corresponding threshold), etc. For example, the edge may be categorized as strong or weak, strong, moderate, or weak, or the like.Processing continues at operation 1302, where a determination is made as to whether the region, block (e.g., LCU), or partition (CU) contains an edge. If not, processing continues at operation 1303, where a most aggressive subtransformation is applied to partitions (e.g., TUs) of the block (e.g., LCU). For example, partitions of the region, block, or partition may be subjected to subtransforms and further processing discussed with respect to the process 900 for partitioning decisions and encoding mode decisions such that the partitions of the block are subjected to more aggressive subtransforms as compared to operation 1306. For example, as discussed with respect to FIGS. 10 and 11, more aggressive transformations may reduce the number of available transformation coefficients more than less aggressive transformations. In one embodiment, the most aggressive subtransforms applied in operation 1303 provide a number of available transform coefficients that is a quarter of the number of residue values of residue blocks. For example, the most aggressive subtransforms for 4×4 residue blocks (partitions) may result in 4 transform coefficients, the most aggressive subtransforms for 8×8 residue blocks may result in 16 transform coefficients, etc.Returning to decision operation 1302, if the region, block, or partition contains an edge, processing continues at operation 1304, where a determination is made whether the edge is a weak edge. If so, processing continues at operation 1303, as discussed above, where the most aggressive subtransformations are applied to remaining partitions (e.g., TUs) of the region, block, or partition. If not, processing continues at decision process 1305, where a determination is made as to whether the edge is a strong edge. If so, processing continues at operation 1307, where no partial transformation is applied to remaining partitions (e.g., TUs) of the block (e.g., LCU). That is, for blocks with a strong edge, the partitions are evaluated for partitioning decision and coding mode decisions using full transformations such that the number of available transformation coefficients for the full transformation corresponds to the number of residue values of the residue partitions. For example, partitions of the block may be subjected to full transformations and further processing (e.g., quantization, inverse quantization, inverse transformation) for partitioning decisions and coding mode decisions.If the region, block, or partition does not have a strong edge (e.g., if the block has a medium or moderate edge), processing continues at operation 1306, where a moderate or less aggressive subtransformation (as compared to that applied in operation 1303) is applied to partitions (e.g., TUs) of the block (e.g., LCU). For example, partitions of the block may be subjected to subtransforms and further processing discussed with respect to process 900 for partitioning decisions and encoding mode decisions such that the partitions of the block are subjected to less aggressive subtransforms as compared to operation 1303. For example, as discussed with respect to FIGS. 10 and 11, more aggressive transformations may reduce the number of available transformation coefficients more than less aggressive transformations. As discussed, the most aggressive subtransforms applied in operation 1303 may provide a number of available transform coefficients that is one-fourth the number of residue values of residue blocks. In contrast, the less aggressive subtransforms applied in operation 1306 may provide a number of available transform coefficients that is more than half the number of residue values of residue blocks. For example, the less aggressive subtransforms for 4×4 residue blocks may result in 9 transform coefficients, the less aggressive subtransforms for 8×8 residue blocks may result in 36 transform coefficients, etc.As discussed, operations 1303, 1306, 1307 may include applying different levels of subtransformations in evaluating partitioning decisions and coding mode decisions. Such partitioning mode decision evaluation and encoding mode decision evaluation may provide any other characteristics discussed herein, such as quantization operations, inverse quantization operations, inverse subtransform operations, comparisons of costs for various candidate partitions, candidate encoding modes, etc. The subtransforms discussed reduce computational resources and time required for such partitioning decisions and encoding mode decisions. As will be appreciated, full transformations are used for the full coding pass to generate standard-compliant quantized transformation coefficients for inclusion in the bitstream 113.FIG. 14 is a flow diagram illustrating an example process 1400 for selectively evaluating 4×4 partitions in video coding, configured in accordance with at least some implementations of the present disclosure. Process 1400 may include one or more of operations 1401-1409, as shown in FIG. 14. Process 1400 may be performed by a system (e.g., system 100, encoding system 600, etc.) to improve data usage efficiency by selectively reducing partition evaluation. For example, by reducing the number of partition scores in partitioning and encoding mode evaluation, computations are reduced for more efficient processing. In the context of decoupled coding systems, process 1400 may be implemented by components of LCU loop 421. In the context of integrated coding systems, process 1400 may be implemented by controller 601, detector module 408, and internal prediction module 603.Process 1400 begins at operation 1401, where an initial partitioning decision is made for a block by evaluating smallest candidate partitions (e.g., CUs) down to a size of 8×8 pixels (and not less than 8×8). For example, a block (e.g., LCU) may be partitioned into candidate partitions and the candidate partitions may be evaluated using cross and internal coding modes as discussed herein and such that the smallest available candidate partitions are 8×8 partitions. In particular, 4×4 partitions are not evaluated to save computational resources in generating the initial partitioning decision. In the context of a decoupled coding system, operation 1401 may be performed by components of the LCU loop 421 (e.g., the SS motion estimation module 401, the SS internal search module 402, the fast CU loop processing module 403, the full CU loop processing module 404, and / or the comprehensive depth decision module 405). For example, operation 1401 may generate LCU partitioning data 415 and CU modes 414. In the context of integrated coding systems, operation 1401 may be performed by controller 601, motion estimation and compensation module 602, internal prediction module 603, and components of local decoding loop 614 to generate an initial partitioning decision. Moreover, operation 1401 may generate initial encoding mode decisions for the initial partitions of the block corresponding to the initial partitioning decision. In either case, operation 1401 generates an initial partitioning decision for a block (e.g., LCU) such that the smallest candidate partitions down to a size of 8×8 partitions are evaluated and the evaluation of smaller partitions is skipped.Processing continues at decision operation 1402, where a determination is made as to whether one of the candidate partitions (e.g., CUs) of the initial partitioning decision of the block (e.g., LCU) has 8×8 partitions. If not, processing is ended and the initial partitioning decision is used as the final partitioning decision for the block (e.g., LCU). Additionally, the initial encoding mode decisions for the partitions (e.g., CUs) are used as final encoding mode decisions. For example, the initial partitioning decision and the initial encoding mode decisions may be made a final partitioning decision and final encoding mode decisions to generate final LCU partitioning data and CU encoding mode data 1421. For example, if the current block (e.g., LCU) does not have 8×8 partitions (e.g., CUs) as part of the initial partitioning decision, then the encoding modes for 4×4 partitions (e.g., CUs) are not evaluated. That is, process 1400 may provide a 4×4 coding mode evaluation (e.g., CU4×4) as a refinement level only. An internal and / or cross coding mode check for 4×4 partitions is performed only after partitioning and coding mode evaluation of a block (e.g., LCU, 64×64) to partition sizes (e.g., CU sizes) of 8×8. After such processing (as discussed above), if no partitions (e.g., CUs) of the block (e.g., LCU) are 8×8 in size, testing of coding modes for 4×4 size coding units is bypassed. In the discussion of Figure 14, the terms block and partitions are used for clarity. As discussed herein, processing may be performed on one or any suitable LCU, CU, macroblock, etc., and partitions thereof may be referred to as sub-blocks, CUs, blocks, etc.Returning to decision operation 1402, if one of the partitions (e.g., CUs) of the block (e.g., LCU) is an 8×8 partition (e.g., CU), processing continues such that checking internal and / or cross modes is evaluated for 4×4 size coding units. In one embodiment, such continued processing is provided for 8×8 partitions (e.g., CUs) that correspond to a cross mode or an internal mode. In another embodiment, such continued processing is provided only for 8×8 partitions (e.g., CUs) to which an internal mode corresponds. In another embodiment, such continued processing is provided only for 8×8 partitions (e.g., CUs) to which a cross mode corresponds. For example, the decision process 1402 may include determining whether the current block unit (e.g., LCU unit) has 8×8 partitions (e.g., CUs) of an internal encoding mode. If not, processing ends as discussed above (even if the block has an 8×8 cross coding mode coding unit). If so, processing continues with checking internal and / or cross modes for 4×4 size partitions (e.g., CUs).As shown, processing continues at operation 1403, where a first 8×8 partition (e.g., CU) is selected using one or more of any suitable techniques. Processing continues at optional decision operation 1404, where a determination is made as to whether to partition the selected 8×8 partition (e.g., CU) into 4×4 partitions (e.g., CUs) and evaluate according to the results of operation 1405. As shown, at operation 1405, one or more detectors may be applied to the current block (e.g., LCU). For example, operation 1405 may be performed by detector module 408.In one embodiment, a flat and noisy block (e.g., LCU) or region detector (e.g., a region containing the block or portions of the image) may be applied in operation 1405 (and via detector module 408). The flat and noisy block or region detector may be applied using one or more of any suitable techniques, such as those discussed with respect to FIG. 15.FIG. 15 is an illustrative diagram of an example flat and noisy region detector 1500 configured in accordance with at least some implementations of the present disclosure. As shown in FIG. 15, the example flat and noisy region detector 1500 may include a noise canceller 1501, a differentiator 1502, a flatness check module 1503, and a noise check module 1504. As shown, the noise canceller 1501 receives an input region 1511 and suppresses the noise in the input region 1511 using any one or more suitable techniques, such as filtering techniques, to generate a noise-suppressed region 1512. The input area 1511 may be any suitable area such as a block (e.g., LCU), an area including the block and other blocks (e.g., LCUs), such as an area of 9×9 blocks with the target block in the middle of the area, or a portion including a block. The noise-suppressed region 1512 is provided to the flatness checking module 1503, which checks the noise-suppressed region 1512 for flatness using one or more of any suitable techniques. In one embodiment, flatness check module 1503 determines a variance of noise-suppressed region 1512 and compares the variance to a predetermined threshold. If the variance does not exceed the threshold, a flatness indicator 1513 indicating that the noise suppressed region 1512 is flat is provided. Moreover, the input region 1511 and the noise-suppressed region 1512 are provided to the differentiator 1502, which may form a difference between the input region 1511 and the noise-suppressed region 1512 using one or more of any suitable techniques to generate the difference 1514. As shown, the difference 1514 is provided to the noise testing module 1504, which tests the difference 1514 to determine whether the input region 1511 is a noisy region using one or more suitable techniques. In one embodiment, the noise check module 1504 determines a variance of the difference 1514 and compares the variance to a predetermined threshold. When the variance reaches or exceeds the threshold, a noise indicator 1515 is provided indicating that the input region 1511 is noisy. When both the flatness indicator 1513 and the noise indicator 1515 are confirmed for the input area 1511, the input area 1511 is determined to be a flat and noisy area.Returning to decision operation 1404 of FIG. 14, if the block is a flat, noisy block (or if the block is in a flat, noisy region), evaluating internal and / or cross coding modes for 4×4 partitions (e.g., CUs) is bypassed such that processing may continue at decision operation 1409, as discussed below. For example, disabling 4×4 partition refinement (e.g., evaluating coding modes for 4×4 partitions) for flat, noisy LCUs may provide the advantage of bypassing such an evaluation when the 4×4 internal modes are unlikely to improve optical quality with respect to the compressed video.Returning to operation 1405, additionally or alternatively, an edge detector may be applied to the block (e.g., LCU) or an area including the block at operation 1405. For example, edge detection may be applied by the detector module 408. The edge detector may be applied using any one or more suitable techniques, such as Canny edge detection techniques. In one embodiment, if an edge is detected in the current block (e.g., LCU) (or an area containing the current block), then an evaluation of internal and / or cross coding modes is provided for 4×4 partitions (e.g., CUs) and, if not, the evaluation of internal and / or cross coding modes is bypassed for 4×4 partitions (e.g., CUs). Providing 4×4 partition refinement (e.g., an evaluation of coding modes for 4×4 partitions) for blocks (e.g., LCUs) that include an edge provides improved optical quality and reduced spurious signals.The discussed detection techniques and decisions as to whether to provide coding modes for 4×4 partitions (e.g., CUs) may be combined using one or more of any suitable techniques. In one embodiment, all 8×8 partitions (e.g., CUs) are evaluated. However, in another embodiment, all of the 8×8 internal partitions (e.g., CUs) are evaluated (overlapping 8×8 partitions (e.g., CUs) not). In one embodiment, all 8×8 partitions (e.g., CUs) are evaluated in a block (e.g., LCU) that includes an edge. In another embodiment, only 8×8 internal partitions (e.g., CUs) are evaluated in a block (e.g., LCU) that includes an edge. In one embodiment, all 8×8 partitions (e.g., CUs) except those that are flat and noisy are evaluated. In another embodiment, only 8×8 internal partitions (e.g., CUs) that are not flat and noisy are evaluated.For cases where the current 8×8 partition (e.g., CU) is to be evaluated, processing continues at operation 1406, where the internal and / or cross coding modes are evaluated for each of the 4×4 partitions (e.g., CUs) partitioned by the current 8×8 partition (e.g., CU) selected in operation 1403. In one embodiment, if the 8×8 partition (e.g., CU) has an initial encoding mode that is an internal mode, only internal modes for the 4×4 partitions (e.g., CUs) are evaluated in operation 1406. Accordingly, in one embodiment, if the 8×8 partition (e.g., CU) has an initial encoding mode that is an overlapping mode, only overlapping modes are evaluated for the 4×4 partitions (e.g., CUs) in operation 1406.In embodiments where internal modes are evaluated for the 4×4 partitions (e.g., CUs), all available internal coding modes may be evaluated or a constrained set of the available internal coding modes may be evaluated. In one embodiment, the evaluated internal coding modes are limited to those provided by optional operation 1407. As shown in operation 1407, operation 1406 may implement a constrained subset of available internal coding modes such that the subset includes only the best internal coding mode for the current 8×8 coding unit (if applicable), the DC internal mode, the planar internal mode, and one or more adjacent modes of the best internal coding mode for the current 8×8 coding unit. For example, for a particular internal directional mode, the immediately adjacent modes are those that are directionally adjacent to the particular internal directional mode, and adjacent modes include immediately adjacent modes and a limited number of immediately adjacent modes from the immediately adjacent modes. For example, with respect to HEVC internal mode 5, immediately adjacent modes are internal modes 4 and 6, and additional adjacent modes are modes 3 and 7 (and 2 and 8, etc.). In one embodiment, the one or more adjacent modes include only the two immediately adjacent modes. In one embodiment, the one or more adjacent modes include the two immediately adjacent modes and two additional immediate neighbors of the two immediately adjacent modes (i.e., one neighbor each of the immediately adjacent modes). In one embodiment, the one or more adjacent modes include the two immediately adjacent modes and four additional immediate neighbors of the two immediately adjacent modes (i.e., two neighbors of the immediately adjacent modes each). However, any number of adjacent modes may be used.In embodiments where cross-modes are evaluated for the 4×4 partitions (e.g., CUs), a full motion estimation search may be performed or the motion estimation search may be limited to an area centered around a location indicated by the best best cross-mode motion vector candidate for the 8×8 partition (or two areas centered around two locations when Bi prediction is the best cross-mode). In one embodiment, the cross coding modes and motion estimation search are limited to those provided by optional operation 1407. As shown in operation 1407, operation 1406 may implement a constrained subset of available cross-coding modes and a motion estimation search such that the subset or constraint only searches a constrained area centered around a location indicated by the best cross-mode motion vector candidate for the current 8×8 partition. As discussed, if the best cross-across mode for the current 8×8 partition is bi-directional prediction, then two surfaces centered around two motion vector candidates are used. For example, for a particular cross mode motion vector, a search range for the 4×4 partitions is defined as a region of a reference image (e.g., original pixel samples or reconstructed pixel samples) centered around the location in the reference image indicated by the motion vector of the current 8×8 partition. The search area or area centered around the location of the reference image indicated by the motion vector of the current 8×8 partition may be limited to any search area such as a search 36×36 pixel search area centered at the location or a 100×100 pixel search area. However, a search area of any size (e.g., a square search area) that is less than a full search may be used.As shown, processing continues from operation 1406 at operation 1408, where a better candidate is selected between the encoding mode received for the 8×8 partition (e.g., CU) and the encoding modes for the 4×4 partitions (e.g., CUs). The better coding mode candidate may be selected using one or more of any suitable techniques such as rate distortion optimization techniques or the like. In the context of a decoupled encoding system, the candidate generation and selection may be done using only original pixel samples (e.g., without full decode loop reconstruction) such that either only luma samples or both luma samples and chroma samples are used, as discussed elsewhere herein. In the context of an integrated coding system, candidate generation and selection may be performed using reconstructed pixel samples (e.g., using local decode loop 614) such that either only luma samples or both luma samples and chroma samples are used, as discussed elsewhere herein. When the four 4×4 partitions (e.g., CUs) are selected (each with a corresponding internal or cross coding mode), the partitioning decision and the CU coding mode decision data are updated to generate a final partitioning and final coding modes 1421. For example, the final partitioning and coding modes 1421 indicate 4×4 partitioning and the internal or coding mode for each of the 4×4 coding units.Processing continues from operation 1408 at decision operation 1409, where a determination is made as to whether the current 8×8 partition (e.g., selected at operation 1403) is the last 8×8 partition (e.g., CU) in the current block (e.g., LCU). If so, processing ends and a final LCU partitioning and CU encoding modes 1421 are generated for the block (e.g., LCU). If no, processing continues at operation 1410, in which a next 8×8 partition (e.g., CU) is selected, and process 1400 continues as discussed above (beginning at decision operation 1404) until a last 8×8 partition (e.g., CU) has been processed.FIG. 16 is a flow diagram illustrating an example process 1600 for video coding configured in accordance with at least some implementations of the present disclosure. The process 1600 may include one or more of operations 1601-1604, as shown in FIG. 16. The process 1600 may form at least a portion of a video coding process. By way of non-limiting example, process 1600 may form at least a portion of a video coding process performed by any device or system such as system 160 discussed herein. Moreover, process 1600 is described herein with respect to system 1700 of FIG. 17.FIG. 17 is an illustrative diagram of an example system 1700 for video coding, configured in accordance with at least some implementations of the present disclosure. As shown in FIG. 17, the system 1700 may include a central processor 1701, a video preprocessor 1702, a video processor 1703, and a data storage 1704. As also shown, the video preprocessor 1702 may include or implement the partitioning and mode decision module 101 and the video processor 1703 may include or implement the encoder 102. Additionally or alternatively, the video processor 1703 may include or implement the encoder 600. In the example of system 1700, data store 1704 may store video data or similar content such as video data, image data, partitioning data, mode data, and / or any other data as discussed herein.As shown, in some embodiments, partitioning and mode decision module 101 is implemented using video preprocessor 1702. In further embodiments, the partitioning and mode decision module 101 or portions thereof are implemented by the central processor 1701 or another processing unit such as an image processor, a graphics processor, or the like. As also shown, in some embodiments, encoder 102 is implemented using video processor 1703. In further embodiments, the encoder 102 or portions thereof is implemented by means of the central processor 1701 or a further processing unit such as an image processor, a graphics processor or the like. Moreover, as shown, in some embodiments, the encoding system 600 (referred to as encoder 600 in FIG. 17 ) is implemented via the video processor 1703. In further embodiments, the encoder 600 or parts thereof is implemented by means of the central processor 1701 or a further processing unit such as an image processor, a graphics processor or the like.The video preprocessor 1702 may include any number and type of video, image, or graphics processing units that can provide the operations discussed herein. Such operations may be implemented by software or hardware, or a combination thereof. For example, the video preprocessor 1702 may include circuitry dedicated to manipulate images, image data, or the like obtained from the memory 1704. Similarly, video processor 1703 may include any number and type of video, image, or graphics processing units that may provide the operations as discussed herein. Such operations may be implemented by software or hardware, or a combination thereof. For example, video processor 1703 may include circuitry dedicated to manipulate images, image data, or the like obtained from memory 1704. Central processor 1701 may include any number and types of processing units or modules that may provide high level control and other functions for system 1700 and / or provide any operations as discussed herein. The memory 1704 may be any type of memory such as volatile memory (e.g., static read / write memory (SRAM), dynamic read / write memory (DRAM), etc.), or nonvolatile memory (e.g., flash memory, etc.), etc. In a non-limiting example, the data store 1704 may be implemented by a cache.In one embodiment, one or more or portions of partitioning and mode decision module 101, encoder 102, and encoder 600 are implemented using an execution unit (EU). The EU may include, for example, programmable logic or circuitry such as a logic core or cores that may provide a wide array of programmable logic functions. In one embodiment, one or more or portions of partitioning and mode decision module 101, encoder 102, and encoder 600 are implemented using dedicated hardware such as fixed functionality circuitry, or the like. Fixed functionality circuitry may include dedicated logic or circuitry and may provide a set of fixed function entry points that may be mapped to the dedicated fixed purpose logic or fixed function. In one embodiment, the partitioning and mode decision module 101 is implemented using a field programmable gate array (FPGA).Returning to FIG. 16, process 1600 may begin at operation 1601, where an input video is received for encoding. For example, the input video may include multiple images such that a first image of the multiple images includes an area including an individual block such that the individual block includes multiple partitions. As discussed herein, the partitions may be any one or combinations of encoding units, prediction units, transform units, or the like.Processing continues at operation 1602, where one or more detectors are applied to the region, the individual block, and / or one or more partitions to generate one or more detection indicators. The detection indicators may include any indicators discussed herein, such as those discussed with respect to operation 1603.Processing continues at operation 1603, where a partitioning decision for the individual block is generated and coding mode decisions are generated for partitions of the individual block corresponding to the partitioning decision using the detection indicators. As shown, the partitioning decision and the encoding mode decisions are made based at least on generating a luma and chroma evaluation decision or luma only evaluation decision for a first partition of the partitions, generating a merge mode decision or exhaust mode decision for a second partition of the partitions that has an initial merge mode decision, generating only a portion of a transform coefficient block for a third partition of the partitions, and / or evaluating 4×4 modes only for a fourth partition of the partitions that is an initial 8×8 encoding partition.In one embodiment, the detection indicators include indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, and an average of a second chroma channel of the first partition exceeds a third threshold, and generating the partitioning decision and the coding mode decisions includes generating the luma and chroma evaluation decision or luma only for the first partition by applying a luma only evaluation decision for the first partition if the luma average does not exceed the first threshold, the average of the first chroma channel does not exceed the second threshold, and the average of the second chroma channel does not exceed the third threshold. For example, the evaluation decision for luma only limits the partitioning decisions and encoding mode decisions to use luma information only. In one embodiment, the detection indicators further include indicators of whether the first partition includes an edge and whether the first partition is in a revealed area, and generating the partitioning decision and the coding mode decisions includes generating the luma and chroma evaluation decision or only luma for the first partition by applying a luma and chroma evaluation decision for the first partition in response to the luma average, the average of the first chroma channel, or the average of the second chroma channel exceeding its corresponding threshold and the first partition includes an edge or is in a revealed area. For example, the luma and chroma evaluation decision provides that the partitioning decisions and encoding mode decisions use luma and chroma information.In one embodiment, the image includes an I-portion including the first partition, and generating the partitioning decision and the encoding mode decisions includes generating the evaluation decision for luma and chroma or only luma for each partition of the image by indicating use of only luma for the first partition in response to the first partition being in the I-portion. In one embodiment, the plurality of images include base layer images and non-base layer images such that base layer images are reference images for non-base layer images but non-base layer images are not reference images for base layer images, the image being a base layer image including a B portion including the first partition, and generating the partitioning decision and the coding mode decisions includes generating the luma and chroma evaluation decision or only luma evaluation decision for each partition of the image by displaying the use of luma and chroma for the first partition in response to the first partition being in the base layer B portion. In another embodiment, the image is a non-base layer image that includes a B portion that includes the first partition, and generating the partitioning decision and the encoding mode decisions includes generating the evaluation decision for luma and chroma or only luma for each partition of the image by indicating use of luma and chroma only for the first partition to select between a merge mode and an exhaust mode in response to the first partition being in the non-base layer B portion and the partitions having initial merge mode decisions.In one embodiment, the detection indicators include a determination whether an amount of a difference between initial exhaust mode encoding cost and initial merging mode encoding cost for the second partition exceeds a threshold, and generating the partitioning decision and the encoding mode decisions includes generating the merging mode decision or exhaust mode decision by selecting the exhaust mode encoding or the merging mode encoding for the second partition if the amount of the difference exceeds the threshold to generate a final exhaust mode decision or merging mode decision or shifting the selection of the exhaust mode encoding or merging mode encoding to a merging mode decision or an exhaust mode decision of a full encoding pass if the amount of the difference does not exceed the threshold.In one embodiment, generating the partitioning decision and the encoding mode decisions includes generating the encoding mode decisions by evaluating an encoding mode for the third partition of the individual block by forming a difference between the third partition and a predicted partition corresponding to the encoding mode to generate a residual partition, generating a transform coefficient block based on the residual partition by performing a subtransform on the residual partition to generate transform coefficients of a part of the transform coefficient block such that a number of transform coefficients in the part is less than a number of values of the residual partition, and setting the remaining transform coefficients of the transform coefficient block to zero, quantizing the transform coefficient block to generate quantized transform coefficients, inverse quantizing the quantized transform coefficients and generating a distortion measure corresponding to the predicted partitions based on the inverse quantized transform coefficients. For example, the third partition may be a TU. In one embodiment, the detection indicators include an indicator of whether the region, the individual block, or the third partition is optically important, and generating the partitioning decision and the encoding mode decisions includes generating only the portion of the transform coefficient block by generating a first transform coefficient block having a first number of available transform coefficients if the region, the individual block, or the third partition is optically important, or generating a second transform coefficient block having a second number of available transform coefficients if the region, the individual block, or the third partition is not optically important, such that the second number is less than the first number.In one embodiment, generating the partitioning decision includes determining an initial partitioning decision for the individual block that evaluates smallest candidate partitions of 8×8 candidate partitions of the individual block, the initial partitioning decision partitions the individual block into the fourth partition and one or more partitions, and generating the partitioning decision further includes evaluating, responsive to the fourth partition being an 8×8 partition, 4×4 sub-partitions of the fourth partition. In one embodiment, the detection indicators further include a best mode for the 8×8 fourth partition and evaluating the 4×4 sub-partitions includes evaluating only cross modes for the 4×4 sub-partitions when the best mode is a cross mode and evaluating only internal modes for the 4×4 sub-partitions when the best mode is an internal mode. In one embodiment, the detection indicators further include a selected best cross mode motion vector for the fourth 8×8 partition, and evaluating the 4×4 sub-partitions includes performing a motion estimation for each of the 4×4 sub-partitions using the selected motion vector to define a search center for the motion estimation searches. In one embodiment, the detection indicators further include a best internal mode corresponding to the fourth 8×8 partition, and evaluating the 4×4 sub-partitions uses only the best internal mode corresponding to the fourth 8×8 partition, a DC mode, a planar mode, and one or more internal modes adjacent to the best internal mode.Processing continues at operation 1604, where the individual block is encoded based at least in part on the partitioning decision to generate a portion of an output bitstream. The individual block may be encoded using any one or more suitable techniques, and the bitstream may be any suitable bitstream such as a standard compliant bitstream.The process 1600 may be repeated in series or in parallel for any number of input video sequences, pictures, coding units, blocks, etc., any number of times. As discussed, process 1600 may provide improved video data usage efficiency by restricting the information used in partitioning decision and encoding mode decision.Various components of the systems described herein may be implemented in software, firmware, and / or hardware, and / or any combinations thereof. For example, various components of the systems or devices discussed herein may be provided, at least in part, by hardware of a computing system sound a-chip (SoC), such as found in a computing system such as a smartphone. Those skilled in the art will appreciate that systems described herein may include additional components not shown in the corresponding figures. For example, the systems illustrated herein may include additional components such as bi-stream multiplexer modules or bitstream demultiplexer modules, and the like, which have not been illustrated for clarity.While implementations of the example operations discussed herein may include performing all illustrated operations in the illustrated order, the present disclosure is not limited in this respect, and in various examples, the implementation of the example processes herein may include only a subset of the illustrated operations, operations performed in other than the illustrated order, or additional operations.Additionally, one or more of the operations discussed herein may be performed in response to instructions provided by one or more computer program products. Such program products may include signal bearing media that provide instructions that, when executed by, for example, a processor, may provide the functionality described herein. The computer program products may be provided in any form of one or more machine readable media. As such, for example, a processor including one or more graphics processing units or one or more processors may perform one or more of the blocks of any of the example operations herein in response to program code and / or instructions or instruction sets conveyed to the processor through one or more machine readable media. In general, a machine-readable medium may convey software in the form of program code and / or instructions or instruction sets that may cause one of the devices and / or systems described herein to implement at least portions of the operations discussed herein and / or portions of the devices, systems, or any modules or components as discussed herein.As used in any implementation described herein, the term "module" refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry to provide the functionality described herein. The software may be embodied as a software package, code, and / or instruction set or instructions, and "hardware", as used in any implementation described herein, may include, for example, single or in any combination, hard-wired circuitry, programmable circuitry, state machine circuitry, fixed function circuitry, unit-for-execution circuitry, and / or firmware storing instructions executed by programmable circuitry. The modules may be embodied jointly or individually as circuitry forming part of a larger system, e.g., an integrated circuit (IC), a system-on-a-chip (SoC), etc.FIG. 18 is an illustrative diagram of an example system 1800 configured in accordance with at least some implementations of the present disclosure. In various implementations, system 1800 may be a mobile system, although system 1800 is not limited in this context. For example, the system 1800 may be incorporated into a personal computer (PC), a laptop computer, an ultra-laptop computer, a tablet, a touch-sensitive panel, a portable computer, a handheld computer, a palmtop computer, a personal digital assistant (PDA), a mobile phone, a combination of a mobile phone and a PDA, a television, a smart device (e.g., a smart phone, a smart tablet or smart TV), a mobile internet appliance (MID), a communication appliance, a data communication appliance, cameras (e.g., compact cameras, super-zoom cameras, digital reflex cameras (DSLR)), etc.In various implementations, the system 1800 includes a platform 1802 coupled to a display 1820. Platform 1802 may receive content from a content device, such as one or more content services devices 1830, one or more content delivery devices 1840, or other similar content sources. A navigation controller 1850 that includes one or more navigation components may be used to interact with platform 1802 and / or display 1820, for example. Each of these components will be described in more detail below.In various implementations, platform 1802 may include any combination of a chipset 1805, processor 1810, memory 1812, antenna 1813, memory 1814, graphics subsystem 1815, applications 1816 and / or radio 1818. Chipset 1805 may provide communication between processor 1810, memory 1812, memory 1814, graphics subsystem 1815, applications 1816 and / or radio 1818. For example, chipset 1805 may include a storage adapter (not shown) that may provide communication with storage 1814.Processor 1810 may be implemented as complex instruction set computer (CISC) processors or reduced instruction set computer (RISC) processors, x86 instruction set compatible processors, multi-cores, or any other microprocessor or central processing unit (CPU). In various implementations, processor 1810 may be one or more dual core processors, one or more dual core mobile processors, etc.The data memory 1812 may be implemented as a volatile data storage device such as, but not limited to, random access memory (RAM), dynamic random access memory (DRAM), or static RAM (SRAM).The memory 1814 may be implemented as a non-volatile data storage device such as, but not limited to, a magnetic media drive, an optical media drive, a tape drive, an internal storage device, a mounted storage device, a flash memory, a battery back-up SDRAM (synchronous DRAM), and / or a network-accessible storage device. In various implementations, the memory 1814 may include technology to increase improved storage performance protection for valuable digital media, e.g., when multiple hard drives are included.Graphics subsystem 1815 may perform processing of images such as still images or video for display. The graphics subsystem 1815 may be, for example, a graphics processing unit (GPU) or an optics processing unit (VPU). An analog or digital interface may be used to couple graphics subsystem 1815 and display 1820. For example, the interface may be a high resolution multimedia interface, a DisplayPort, wireless HDMI, and / or wireless HD compliant techniques. Graphics subsystem 1815 may be integrated with processor 1810 or chipset 1805. In some implementations, graphics subsystem 1815 may be a stand-alone device communicatively coupled to chipset 1805.The graphics processing techniques and / or video processing techniques described herein may be implemented in various hardware architectures. For example, graphics functionality and / or video functionality may be integrated into a chipset. Alternatively, a discrete graphics processor and / or video processor may be used. As yet another implementation, the graphics function and / or video functions may be provided by a general purpose processor that includes a multi-core processor. In further embodiments, the functions may be implemented in an consumer electronics device.The radio 1818 may include one or more radios capable of transmitting and receiving signals using various suitable wireless communication techniques. Such techniques may include communications over one or more wireless networks. Examples of wireless networks include (but are not limited to) local area wireless networks (WLANs), wireless personal area networks (WPANs), wireless metropolitan area networks (WMANs), cellular networks, and satellite networks. When communicating over such networks, the radio 1818 may operate in accordance with one or more applicable standards in any version.In various implementations, the display 1820 may include any television type monitor or display. The display 1820 may include, for example, a computer display screen, a touch screen display, a video surveillance device, a television-type device, and or a television. The display 1820 may be digital and / or analog. In various implementations, the display device 1820 may be a holographic display device. Additionally, the display 1820 may be a transparent surface that may receive an optical projection. Such projections may convey various forms of information, images, and / or objects. For example, such projections may be an optical overlay for a mobile augmented reality (MAR) application. Under the control of one or more software applications 1816, platform 1802 may display user interface 1822 on display 1820.In various implementations, one or more content services devices 1830 may be hosted by a national, international, and / or independent service and therefore accessible to platform 1802, e.g., via the Internet. One or more content services 1830 may be coupled to platform 1802 and / or display 1820. Platform 1802 and / or one or more content services devices 1830 may be coupled to a network 1860 to communicate (e.g., transmit and / or receive) media information to and from network 1860. One or more content providers 1840 may also be coupled to platform 1802 and / or display 1820.In various implementations, the one or more content services devices 1830 may include a cable television box, a personal computer, a network, a telephone, Internet enabled devices, or a device capable of communicating digital information and / or content, and any other device capable of unidirectionally or bidirectionally communicating content between content providers and the platform 1802 and / or the display 1820 via the network 1860, or directly. It will be appreciated that the content may be communicated unidirectionally and / or bidirectionally to and from any component in system 1800 and a content provider via network 1860. Examples of content may include any media information including, for example, video, music, medical information, gaming information, etc.One or more content services devices 1830 may receive content such as cable television programming including media information, digital information, and / or other content. Examples of content providers may include any cable or satellite television providers, radio providers, or Internet content providers. The examples provided are not intended to limit implementations according to the present disclosure in any way.In various implementations, platform 1802 may receive control signals from navigation controller 1850 that have one or more navigation components. The navigation components may be used to interact with, for example, user interface 1822. In various embodiments, navigation may be a pointing device, which may be a computer hardware component (particularly a human interface device) that allows a user to input spatial (e.g., continuous and multi-dimensional) data to a computer. Many systems, such as graphical user interfaces (GUI), televisions, and monitoring devices, allow the user to control the computer or television using physical gestures and provide data therefor.Movements of the navigation components may be replicated on a display device (e.g., display device 1820) by movements of a pointer, cursor, focus ring, or other visual indicators displayed on the display device. For example, the navigation components that are located on the navigation may be mapped to virtual navigation components displayed at user interface 1822, under the control of software applications 1816, for example. In various embodiments, it may not be a single component, but may be integrated into platform 1802 and / or display 1820. However, the present disclosure is not limited to the elements or context shown or described herein.In various embodiments, drivers (not shown) may include technology to allow users, for example, to turn on and off platform 1802 with the touch of a button immediately like a television after the initial launch. Program logic may enable platform 1802 to stream content to media adapters or one or more other content services devices 1830 or one or more content delivery devices 1840, even when the platform is powered off. Additionally, chipset 1805 may include, for example, hardware and / or software support for 5.1 surround sound audio and / or high resolution 7.1 surround sound audio. The drivers may include a graphics driver for integrated graphics platforms. In various embodiments, the graphics driver may include an Express Peripheral Component Interconnect (PCI) graphics card.In various implementations, one or more of the components shown in system 1800 may be integrated. For example, platform 1802 and one or more content services devices 1830 may be integrated, platform 1802 and one or more content providers 1840 may be integrated, or platform 1802, one or more content services devices 1830 and one or more content providers 1840 may be integrated. In various embodiments, platform 1802 and display 1820 may be an integrated unit. For example, the display device 1820 and one or more content services devices 1830 may be integrated, or the display device 1820 and one or more content providers 1840 may be integrated. These examples are not intended to limit the present disclosure.In various embodiments, the system 1800 may be implemented as a wireless system, a wired system, or a combination of both. When implemented as a wireless system, system 1800 may include components and interfaces suitable for communicating over shared wireless media such as one or more antennas, transmitters, receivers, transceivers, amplifiers, filters, control logic, etc. An example of shared wireless media may include portions of a wireless spectrum such as the RF spectrum, etc. When implemented as a wired system, system 1800 may include components and interfaces suitable for communicating over wired communication media such as input / output (I / O) adapters, physical connectors to connect the I / O adapters to a corresponding wired communication medium, a network interface card (NIC), a disk controller, a video controller, an audio controller, and the like. Examples of wired communication media may include a wire, a cable, metal conductors, a printed circuit board (PCB), a bus board, a circuit structure, a semiconductor material, a twisted pair cable, a coaxial cable, a fiber optic, etc.Platform 1802 may establish one or more logical or physical channels to communicate information. The information may include media information and control information. Media information may refer to any data representing content intended for a user. Examples of content may include, for example, data from a voice conversation, a video conference, a video stream, an electronic mail ("email") message, voice message, alphanumeric symbols, graphics, image data, video data, text, etc. Data from a voice conversation may be, for example, voice information, pause periods, background noise, comfort noise, tones, etc. Control information may refer to any data representing commands, instructions, or control words intended for an automated system. For example, control information may be used to route media information through a system or to instruct a node to process the media information in a predetermined manner. However, the embodiments are not limited to the elements or context shown or described in FIG. 18.As described above, the system 1800 may be embodied in varying physical styles or form factors. FIG. 19 illustrates a small form factor example device 1900 configured in accordance with at least some implementations of the present disclosure. In some examples, system 1800 may be implemented via device 1900. In further examples, the system 100 or portions thereof may be implemented by means of the apparatus 1900. In various embodiments, device 1900 may be implemented, for example, as a mobile computing device having wireless capabilities. A mobile computing device may refer to any device having, for example, a processing system and a mobile power source or power supply such as one or more batteries.Examples of a mobile computing device may include a personal computer (PC), a laptop computer, an ultra-laptop computer, a tablet, a touch panel, a portable computer, a handheld computer, a palmtop computer, a personal digital assistant (PDA), a mobile phone, a combination of a mobile phone and a PDA, a smart device (e.g., a smart phone, a smart tablet or a smart mobileTV), a mobile internet device (MID), a communication device, a data communication device, cameras, etc.Examples of a mobile computing device may also include computers configured to be worn by a person, such as wrist computers, finger computers, ring computers, glasses computers, belt buckle computers, bracelet computers, shoe computers, apparel computers, and other wearable computers. In various embodiments, a mobile computing device may be implemented, for example, as a smartphone that can execute computer applications as well as voice communications and / or data communications. Although some embodiments may be described as exemplary with a mobile computing device implemented as a smartphone, it may be appreciated that other embodiments may also be implemented using other wireless mobile computing devices. The embodiments are not limited in this context.As shown in FIG. 19, the device 1900 may include a housing having a front side 1901 and a back side 1902. The apparatus 1900 includes a display device 1904, an input / output (I / O) device 1906, and an integrated antenna 1908. The device 1900 may also include navigation components 1912. The I / O device 1906 may include any suitable I / O device for inputting information to a mobile computing device. Examples of the I / O device 1906 may include an alphanumeric keyboard, a numeric keypad, a touch panel, input keys, buttons, switches, microphones, speakers, a speech recognizer, and speech recognition software, etc. Information may also be input to device 1900 in the form of a microphone (not shown) or may be digitized by a speech recognizer. As shown, device 1900 may include a camera 1905 (e.g., including a lens, aperture, and image sensor) and flash 1910 integrated into the back 1902 (or elsewhere) of device 1900. In further examples, camera 1905 and flash 1910 may be integrated into front 1901 of device 1900 or both front and rear cameras may be provided. Camera 1905 and flash 1910 may be components of a camera module to generate image data that is processed into a video stream that is output to display device 1904 and / or communicated remotely from device 1900, e.g., via antenna 1908.Various embodiments may be implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements may include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASIC), programmable logic devices (PLD), digital signal processors (DSP), field programmable gate array (FPGA), logic gates, registers, semiconductor device, chips, microchips, chip sets, etc. Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application programming interfaces (API), instruction sets, computing code, machine code, code segments, machine code segments, words, values, symbols, or any combinations thereof. Determining whether an embodiment is implemented using hardware elements and / or software elements may vary according to any number of factors such as desired computational rate, power levels, thermal tolerances, processing cycle equipment, input data rates, output data rates, data storage resources, data bus speeds, and other design or performance dependencies.One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium that represents various logic in the processor and that, when read by a machine, causes the machine to fabricate logic to perform the techniques described herein. Such representations, known as IP cores, may be stored on a tangible, machine-readable medium and provided to various customers or manufacturing facilities for loading into the manufacturing machines that actually make the logic or processor.While certain features set forth herein have been described with respect to various implementations, this description is not intended to be construed in a limiting sense. Therefore, various modifications to the implementations described herein, as well as other implementations that will be apparent to those skilled in the art regarding the present disclosure, are deemed to be within the spirit and scope of the present disclosure.The following embodiments relate to further embodiments.In one or more first embodiments, a computer-implemented method for video coding includes receiving an input video for coding, wherein the input video is comprised of a plurality of images and a first image of the plurality of images includes an area including an individual block such that the individual block includes a plurality of partitions, applying one or more detectors to the area, the individual block, and / or one or more of the plurality of partitions to generate one or more detection indicators, generating a partitioning decision for the individual block and coding mode decisions for partitions of the individual block corresponding to the partitioning decision using the detection indicators based on generating an evaluation decision for luma and chroma or only for luma for a first partition of the partitions, generating a merge mode decision or exit mode decision for a second partition of the partitions that has an initial merge mode decision, generating only a portion of a transform coefficient block for a third partition of the partitions, or evaluating 4×4 modes only for a fourth partition of the partitions that is an initial 8×8 encoding partition, and encoding the individual block based at least on the partitioning decision to generate a portion of an output bitstream.In one or more second embodiments, the detection indicators for each of the first embodiments include indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, and an average of a second chroma channel of the first partition exceeds a third threshold, and generating the partitioning decision and the coding mode decisions includes generating the luma and chroma evaluation decision or only luma evaluation decision for the first partition by applying a luma only evaluation decision for the first partition if the luma average does not exceed the first threshold, the average of the first chroma channel does not exceed the second threshold, and the average of the second chroma channel does not exceed the third threshold.In one or more third embodiments, the detection indicators for each of the first or second embodiments include indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, an average of a second chroma channel of the first partition exceeds a third threshold, the first partition includes an edge, and the first partition is in a revealed area, and generating the partitioning decision and the coding mode decisions includes generating the luma and chroma evaluation decision or only luma for the first partition by applying luma and chroma evaluation decision for the first partition in response to the luma average, the average of the first chroma channel or the average of the second chroma channel exceed their respective thresholds and the first partition contains an edge or is in a revealed area.In one or more fourth embodiments, the image includes, for each of the first through third embodiments, an I-portion including the first partition, and generating the partitioning decision and the encoding mode decisions includes generating the evaluation decision for luma and chroma or only luma for each partition of the image by indicating use of only luma for the first partition in response to the first partition being in the I-portion.In one or more fifth embodiments, the plurality of images for each of the first to fourth embodiments include base layer images and non-base layer images such that base layer images are reference images for non-base layer images, but non-base layer images are not reference images for base layer images, the image being a base layer image including a B portion including the first partition, and generating the partitioning decision and the coding mode decisions includes generating the evaluation decision for luma and chroma or only luma for each partition of the image by displaying the use of luma and chroma for the first partition in response to the first partition being in the base layer B portion.In one or more sixth embodiments, the plurality of images for each of the first to fifth embodiments include base layer images and non-base layer images such that base layer images are reference images for non-base layer images, but non-base layer images are not reference images for base layer images, the image being a non-base layer image including a B portion including the first partition, and generating the partitioning decision and the coding mode decisions includes generating the evaluation decision for luma and chroma or only luma for each partition of the image by displaying the use of luma and chroma for the first partition to respond only to it, the first partition is in the non-base layer B portion and the partitions have initial merge mode decisions to select between a merge mode and an exit mode.In one or more seventh embodiments, the detection indicators for each of the first through sixth embodiments include a determination as to whether an amount of a difference between initial exhaust mode encoding cost and initial merging mode encoding cost for the second partition exceeds a threshold, and generating the partitioning decision and the encoding mode decisions includes generating the merging mode decision or exhaust mode decision by selecting the exhaust mode encoding or merging mode encoding for the second partition when the amount of the difference exceeds the threshold to generate a final exhaust mode decision or merging mode decision, or shifting the selection of the exhaust mode encoding or merging mode encoding to a merging mode decision or an exhaust mode decision of a full encoding pass when the amount of the difference does not exceed the threshold.In one or more eighth embodiments, generating the partitioning decision and the encoding mode decisions includes, for each of the first to seventh embodiments, generating the encoding mode decisions by evaluating an encoding mode for the third partition of the individual block by forming a difference between the third partition and a predicted partition corresponding to the encoding mode to generate a residual partition, generating a transform coefficient block based on the residual partition by performing a partial transform on the residual partition to generate transform coefficients from a part of the transform coefficient block such that a number of transform coefficients in the part is less than a number of values of the residual partition, and setting the remaining transform coefficients of the transform coefficient block to zero, quantizing the transform coefficient block, to generate quantized transform coefficients, inverse quantize the quantized transform coefficients, and generate a distortion measure corresponding to the predicted partition based on the inverse quantized transform coefficients.In one or more ninth embodiments, the detection indicators for each of the first through eighth embodiments include an indicator of whether the region, the individual block, or the third partition is optically important, and generating the partitioning decision and the encoding mode decisions includes generating only the portion of the transform coefficient block by generating a first transform coefficient block having a first number of available transform coefficients when the region, the individual block, or the third partition is optically important, or generating a second transform coefficient block having a second number of available transform coefficients when the region, the individual block, or the third partition is not optically important, such that the second number is less than the first number.In one or more tenth embodiments, generating the partitioning decision for each of the first through ninth embodiments includes determining an initial partitioning decision for the individual block that evaluates smallest candidate partitions of 8×8 candidate partitions of the individual block, wherein the initial partitioning decision partitions the individual block into the fourth partition and one or more further partitions, and generating the partitioning decision further includes evaluating 4×4 sub-partitions of the fourth partition in response to the fourth partition being an 8×8 partition.In one or more eleventh embodiments, the detection indicators for each of the first through tenth embodiments include a best mode for the fourth 8×8 partition and evaluating the 4×4 sub-partitions includes evaluating only cross modes for the 4×4 sub-partitions when the best mode is a cross mode and evaluating only internal modes for the 4×4 sub-partitions when the best mode is an internal mode.In one or more twelfth embodiments, the detection indicators for each of the first through eleventh embodiments include a selected best cross mode motion vector for the fourth 8×8 partition, and evaluating the 4×4 sub-partitions includes performing a motion estimation for each of the 4×4 sub-partitions using the selected motion vector to define a search center for the motion estimation searches.In one or more thirteenth embodiments, the detection indicators for each of the first to twelfth embodiments include a best internal mode corresponding to the fourth 8×8 partition, and evaluating the 4×4 sub-partitions uses only the best internal mode corresponding to the fourth 8×8 partition, a DC mode, a planar mode, and one or more internal modes adjacent to the best internal mode.In one or more fourteenth embodiments, a system for video coding includes a data store to store an input video for coding, the input video including a plurality of images and a first image of the plurality of images including an area including an individual block such that the individual block includes a plurality of partitions, and one or more processors coupled to the data store, the one or more processors configured to apply one or more detectors to the area, the individual block, and / or one or more of the plurality of partitions to generate one or more detection indicators, a partitioning decision for the individual block, and coding mode decisions for partitions of the individual block corresponding to the partitioning decision, using the detection indicators based on at least one of the one or more processors to generate a luma and chroma evaluation decision or only luma evaluation decision for a first partition of the partitions, generate a merge mode decision or exhaust mode decision for a second partition of the partitions having an initial merge mode decision, generate only a portion of a transform coefficient block for a third partition of the partitions, or evaluate 4×4 modes only for a fourth partition of the partitions being an initial 8×8 encoding partition, and encode the individual block based at least on the partitioning decision to generate a portion of an output bitstream.In one or more fifteenth embodiments, the detection indicators for each of the fourteenth embodiments include indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, an average of a second chroma channel of the first partition exceeds a third threshold, the first partition includes an edge, and the first partition is in a revealed area, and generating, by the one or more processors, the partitioning decision and the coding mode decisions includes generating, by the one or more processors, the luma and chroma evaluation decision or only luma evaluation decision for the first partition, by applying an luma only evaluation decision for the first partition, when the luma average does not exceed the first threshold, the average of the first chroma channel does not exceed the second threshold, and the average of the second chroma channel does not exceed the third threshold, and applying an evaluation decision for luma and chroma to the first partition in response to the luma average, the average of the first chroma channel, or the average of the second chroma channel exceeding its corresponding threshold, and the first partition including an edge or being in a revealed area.In one or more sixteenth embodiments, the detection indicators for each of the fourteenth or fifteenth embodiments include a determination as to whether an amount of a difference between initial outlet mode encoding cost and initial merging mode encoding cost for the second partition exceeds a threshold, and the one or more processors for generating the partitioning decision and the encoding mode decisions include the one or more processors for generating the merging mode decision or outlet mode decision by selecting the outlet mode encoding or the merging mode encoding for the second partition when the amount of the difference exceeds the threshold to generate a final outlet mode decision or shifting the selection of the outlet mode encoding or the merging mode encoding to a merging mode decision or an outlet mode decision of a full encoding pass, if the amount of the difference does not exceed the threshold.In one or more seventeenth embodiments, the detection indicators for each of the fourteenth through sixteenth embodiments include an indicator of whether the region, the individual block or the third partition is optically important and the one or more processors for generating the partitioning decision and the coding mode decisions include the one or more processors for generating only the portion of the transform coefficient block by the one or more processors for generating a first transform coefficient block having a first number of available transform coefficients if the region, the individual block or the third partition is optically important or the one or more processors for generating a second transform coefficient block having a second number of available transform coefficients if the region, the individual block or the third partition is not optically important, such that, the second number being less than the first number.In one or more eighteenth embodiments, the one or more processors to generate the partitioning decision for each of the fourteenth through seventeenth embodiments include the one or more processors to determine an initial partitioning decision for the individual block that evaluates smallest candidate partitions of 8×8 candidate partitions of the individual block, wherein the initial partitioning decision partitions the individual block into the fourth partition and one or more further partitions and to generate the partitioning decision further includes, in response to the fourth partition being an 8×8 partition, the evaluation of 4×4 sub-partitions of the fourth partition.In one or more nineteenth embodiments, the detection indicators for each of the fourteenth through eighteenth embodiments include a best mode for the fourth 8×8 partition and the one or more processors for evaluating the 4×4 sub-partitions include evaluating only cross modes for the 4×4 sub-partitions when the best mode is a cross mode and evaluating only internal modes for the 4×4 sub-partitions when the best mode is an internal mode such that the evaluating includes only cross modes such that the one or more processors evaluate the 4×4 sub-partitions through a motion estimation search for each of the 4×4 sub-partitions using a selected best cross mode motion vector for the fourth 8×8 partition, to define a search center for motion estimation searches, and such that the evaluation includes only internal modes, such that the one or more processors evaluate the 4×4 sub-partitions using only the best internal mode corresponding to the fourth 8×8 partition, a DC mode, a planar mode, and one or more internal modes adjacent to the best internal mode.In one or more twentieth embodiments, at least one machine-readable medium may include a plurality of instructions that, in response to being executed on a computing device, cause the computing device to perform a method according to any of the above embodiments.In one or more twenty-first embodiments, an apparatus may include means for performing a method according to any of the above embodiments.It will be appreciated that the embodiments are not limited to the described embodiments, but may be practiced with modification and modification without departing from the scope of the appended claims. For example, the above embodiments may include a particular combination of features. However, the above embodiments are not limited in this regard, and in various implementations, the above embodiments may include performing only a subset of such features, performing a different order of such features, performing different combinations of such features, and / or performing additional features to those expressly listed. The scope of the embodiments should therefore be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

Claims

A computer-implemented method (1600) for video coding, comprising: receiving (1601) an input video (111) for coding, wherein the input video (111) comprises multiple images (200) and a first image (301) of the multiple images (200) comprises an area comprising an individual block, wherein the individual block comprises multiple partitions; applying (1602) one or more detectors to the area, the individual block and / or one or more of the multiple partitions to generate one or more detection indicators; generating (1603) a partitioning decision for the individual block and encoding mode decisions for partitions of the individual block corresponding to the partitioning decision using the detection indicators based on generating an evaluation decision for luma and chroma or only luma for a first partition of the partitions, generating a merging mode decision or exhaust mode decision for a second partition of the partitions that has an initial merging mode decision, generating only a portion of a transform coefficient block for a third partition of the partitions, and / or evaluating 4×4 modes only for a fourth partition of the partitions that is an initial 8×8 encoding partition; and encoding (1604) the individual block at least based on the partitioning decision to generate a portion of an output bitstream (113); wherein the detection indicators comprise indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, and an average of a second chroma channel of the first partition exceeds a third threshold, and generating the partitioning decision and the coding mode decisions comprises generating the evaluation decision for luma and chroma or only luma for the first partition by applying an evaluation decision only for luma for the first partition if the luma average does not exceed the first threshold, the average of the first chroma channel does not exceed the second threshold, and the average of the second chroma channel does not exceed the third threshold.The method (1600) of claim 1, wherein the detection indicators comprise indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, an average of a second chroma channel of the first partition exceeds a third threshold, the first partition includes an edge, and the first partition is in a revealed area, and generating the partitioning decision and the coding mode decisions comprises generating the luma and chroma evaluation decision or only luma evaluation decision for the first partition by applying luma and chroma evaluation decision for the first partition in response to the luma average, the average of the first chroma channel or the average of the second chroma channel exceed their respective thresholds and the first partition includes an edge or is in a revealed area.The method (1600) of any one of claims 1 or 2, wherein the image comprises an I-portion comprising the first partition, and generating the partitioning decision and the encoding mode decisions comprises generating the evaluation decision for luma and chroma or only luma for each partition of the image by indicating use of only luma for the first partition in response to the first partition being in the I-portion.The method (1600) of any one of claims 1-3, wherein the plurality of images comprise base layer images and non-base layer images such that base layer images are reference images for non-base layer images, but non-base layer images are not reference images for base layer images, wherein the image is a base layer image comprising a B portion comprising the first partition, and generating the partitioning decision and the coding mode decisions comprises generating the evaluation decision for luma and chroma or only luma for each partition of the image by displaying the use of luma and chroma for the first partition in response to the first partition being in the base layer B portion.The method (1600) of any of claims 1-4, wherein the plurality of images comprise base layer images and non-base layer images such that base layer images are reference images for non-base layer images, but non-base layer images are not reference images for base layer images, wherein the image is a non-base layer image comprising a B portion comprising the first partition, and generating the partitioning decision and the coding mode decisions comprises generating the luma and chroma evaluation decision or only luma evaluation decision for each partition of the image by displaying the use of luma and chroma for the first partition to only in response to the first partition being in the non-base layer B portion and the partitions having initial merging mode decisions, selecting between a merging mode and an outlet mode.The method (1600) of any of claims 1-5, wherein the detection indicators comprise a determination of whether an amount of a difference between initial outlet mode coding cost and initial merging mode coding cost for the second partition exceeds a threshold, and generating the partitioning decision and the coding mode decisions comprises generating the merging mode decision or outlet mode decision by selecting the outlet mode coding or merging mode coding for the second partition if the amount of the difference exceeds the threshold to generate a final outlet mode decision or merging mode decision, or shifting the selection of the outlet mode coding or merging mode coding to a merging mode decision or an outlet mode decision of a full coding pass if the amount of the difference does not exceed the threshold.The method (1600) of any of claims 1-6, wherein generating the partitioning decision and the encoding mode decisions comprises generating the encoding mode decisions by evaluating an encoding mode for the third partition of the individual block by: forming a difference between the third partition and a predicted partition corresponding to the encoding mode to generate a residual partition; generating a transform coefficient block based on the residual partition by: performing a partial transform on the residual partition to generate transform coefficients from a part of the transform coefficient block, wherein a number of transform coefficients in the part is less than a number of values of the residual partition; and setting the remaining transform coefficients of the transform coefficient block to zero; quantizing the transform coefficient block to generate quantized transform coefficients; inverse quantizing the quantized transform coefficients; and generating a distortion measure corresponding to the predicted partition based on the inverse quantized transform coefficients.The method (1600) of any of claims 1-7, wherein the detection indicators comprise an indicator of whether the region, the individual block, or the third partition is optically important, and generating the partitioning decision and the coding mode decisions comprises generating only the portion of the transformation coefficient block by generating a first transformation coefficient block having a first number of available transformation coefficients if the region, the individual block, or the third partition is optically important, or generating a second transformation coefficient block having a second number of available transformation coefficients if the region, the individual block, or the third partition is not optically important, wherein the second number is less than the first number.The method (1600) of any of claims 1-8, wherein generating the partitioning decision comprises determining an initial partitioning decision for the individual block that evaluates smallest candidate partitions of 8×8 candidate partitions of the individual block, wherein the initial partitioning decision partitions the individual block into the fourth partition and one or more further partitions, and generating the partitioning decision further comprises evaluating 4×4 sub-partitions of the fourth partition in response to the fourth partition being an 8×8 partition.The method (1600) of any of claims 1-9, wherein the detection indicators comprise a best mode for the fourth 8×8 partition, and evaluating the 4×4 sub-partitions comprises evaluating only cross modes for the 4×4 sub-partitions if the best mode is a cross mode, and evaluating only internal modes for the 4×4 sub-partitions if the best mode is an internal mode.The method (1600) of any of claims 1-10, wherein the detection indicators comprise a selected best cross mode motion vector for the fourth 8×8 partition, and evaluating the 4×4 sub-partitions comprises performing a motion estimation search for each of the 4×4 sub-partitions using the selected motion vector to define a search center for the motion estimation searches.The method (1600) of any of claims 1-11, wherein the detection indicators comprise a best internal mode corresponding to the fourth 8×8 partition, and evaluating the 4×4 sub-partitions uses only the best internal mode corresponding to the fourth 8×8 partition, a DC mode, a planar mode, and one or more internal modes adjacent to the best internal mode.A system (1700) for video coding, comprising: a data store (1704) to store an input video for coding, wherein the input video (111) comprises a plurality of images (200) and a first image (301) of the plurality of images (200) comprises an area comprising an individual block, wherein the individual block comprises a plurality of partitions; and one or more processors (1701) coupled to the data store, wherein the one or more processors are configured to apply one or more detectors to the area, the individual block, and / or one or more of the plurality of partitions to generate one or more detection indicators; a partitioning decision for the individual block and encoding mode decisions for partitions of the individual block corresponding to the partitioning decision using the detection indicators based on at least one of the one or more processors to generate an evaluation decision for luma and chroma or only luma for a first partition of the partitions, generate a merging mode decision or exhaust mode decision for a second partition of the partitions that has an initial merging mode decision, generate only a portion of a transform coefficient block for a third partition of the partitions, and / or evaluate 4×4 modes only for a fourth partition of the partitions that is an initial 8×8 encoding partition; Encoding the individual block at least based on the partitioning decision to generate a portion of an output bitstream (113); wherein the detection indicators comprise indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, an average of a second chroma channel of the first partition exceeds a third threshold, the first partition includes an edge, and the first partition is in a revealed area, Generating, by the one or more processors, the partitioning decision and the encoding mode decisions includes generating, by the one or more processors, the evaluation decision for luma and chroma or only luma for the first partition by applying an evaluation decision for luma only for the first partition if the luma average does not exceed the first threshold, the average of the first chroma channel does not exceed the second threshold, and the average of the second chroma channel does not exceed the third threshold.The system (1700) of claim 13, wherein the applying an evaluation decision for luma and chroma for the first partition in response to the luma average, the average of the first chroma channel, or the average of the second chroma channel exceeding its corresponding threshold and the first partition including an edge or being in a revealed area comprises.The system (1700) of claim 13 or 14, wherein the detection indicators comprise a determination of whether an amount of a difference between initial outlet mode coding cost and initial merging mode coding cost for the second partition exceeds a threshold, and the one or more processors to generate the partitioning decision and the coding mode decisions comprise the one or more processors (1701) to generate the merging mode decision or outlet mode decision by selecting the outlet mode coding or the merging mode coding for the second partition if the amount of the difference exceeds the threshold to generate a final outlet mode decision or merging mode decision, or shifting the selection of the outlet mode coding or the merging mode coding to a merging mode decision or an outlet mode decision of a full coding pass, if the amount of the difference does not exceed the threshold value.The system (1700) of any of claims 13-15, wherein the detection indicators comprise an indicator of whether the region, the individual block, or the third partition is optically important, and the one or more processors (1701) for generating the partitioning decision and the coding mode decisions comprise the one or more processors (1701) for generating only the portion of the transform coefficient block by the one or more processors for generating a first transform coefficient block having a first number of available transform coefficients if the region, the individual block, or the third partition is optically important, or the one or more processors (1701) for generating a second transform coefficient block having a second number of available transform coefficients if the region, the individual block, or the third partition is not optically important, wherein the second number is less than the first number.The system (1700) of any of claims 13-16, wherein the one or more processors (1701) for generating the partitioning decision comprise the one or more processors (1701) for determining an initial partitioning decision for the individual block that evaluates smallest candidate partitions of 8×8 candidate partitions of the individual block, wherein the initial partitioning decision partitions the individual block into the fourth partition and one or more further partitions, and for generating the partitioning decision further comprises evaluating 4×4 sub-partitions of the fourth partition in response to the fourth partition being an 8×8 partition.The system (1700) of any of claims 13-17, wherein the detection indicators comprise a best mode for the fourth 8×8 partition and the one or more processors (1701) for evaluating the 4×4 sub-partitions comprise evaluating only cross modes for the 4×4 sub-partitions if the best mode is a cross mode and evaluating only internal modes for the 4×4 sub-partitions if the best mode is an internal mode, wherein evaluating only cross modes comprises the one or more processors evaluating the 4×4 sub-partitions by a motion estimation search for each of the 4×4 sub-partitions using a selected best cross mode motion vector for the fourth 8×8 partition, To define a search center for motion estimation searches, and wherein evaluating only internal modes comprises the one or more processors evaluating the 4×4 sub-partitions using only the best internal mode corresponding to the fourth 8×8 partition, a DC mode, a planar mode, and one or more internal modes adjacent to the best internal mode.At least one machine readable medium comprising a plurality of instructions that, in response to being executed on a computing device, cause the computing device to perform video coding by: receiving an input video (111) for coding, wherein the input video (111) comprises a plurality of images (200) and a first image (301) of the plurality of images (200) comprises an area comprising an individual block, wherein the individual block comprises a plurality of partitions; applying one or more detectors to the area, the individual block, and / or one or more of the plurality of partitions to generate one or more detection indicators; generating a partitioning decision for the individual block and encoding mode decisions for partitions of the individual block corresponding to the partitioning decision using the detection indicators based on generating an evaluation decision for luma and chroma or only luma for a first partition of the partitions, generating a merging mode decision or exhaust mode decision for a second partition of the partitions that has an initial merging mode decision, generating only a portion of a transform coefficient block for a third partition of the partitions, and / or evaluating 4×4 modes only for a fourth partition of the partitions that is an initial 8×8 encoding partition; encoding the individual block at least based on the partitioning decision to generate a portion of an output bitstream (113); wherein the detection indicators comprise indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, an average of a second chroma channel of the first partition exceeds a third threshold, the first partition includes an edge, and the first partition is in a revealed area, and generating the partitioning decision and the encoding mode decisions comprises generating the luma and chroma evaluation decision or only luma evaluation decision for the first partition by applying a luma only evaluation decision for the first partition, when the luma average does not exceed the first threshold, the average of the first chroma channel does not exceed the second threshold, and the average of the second chroma channel does not exceed the third threshold.The machine readable medium of claim 19, further comprising instructions to apply a luma and chroma evaluation decision for the first partition in response to the luma average, the average of the first chroma channel, or the average of the second chroma channel exceeding its corresponding threshold and the first partition including an edge or being in a revealed area.The machine readable medium of claim 19 or 20, wherein the detection indicators comprise a determination of whether an amount of a difference between initial outlet mode encoding cost and initial merge mode encoding cost for the second partition exceeds a threshold, and generating the partitioning decision and the encoding mode decisions comprises generating the merge mode decision or outlet mode decision by selecting the outlet mode encoding or the merge mode encoding for the second partition if the amount of the difference exceeds the threshold to generate a final outlet mode decision or merge mode decision, or shifting the selection of the outlet mode encoding or the merge mode encoding to a merge mode decision or outlet mode decision of a full encoding pass if the amount of the difference does not exceed the threshold.The machine readable medium of any of claims 19-21, wherein the detection indicators comprise an indicator of whether the region, the individual block, or the third partition is optically important, and generating the partitioning decision and the encoding mode decisions comprises generating only the portion of the transform coefficient block by generating a first transform coefficient block having a first number of available transform coefficients if the region, the individual block, or the third partition is optically important, or generating a second transform coefficient block having a second number of available transform coefficients if the region, the individual block, or the third partition is not optically important, the second number being less than the first number.The machine readable medium of any of claims 19-22, wherein generating the partitioning decision comprises determining an initial partitioning decision for the individual block that evaluates smallest candidate partitions of 8×8 candidate partitions of the individual block, wherein the initial partitioning decision partitions the individual block into the fourth partition and one or more further partitions, and generating the partitioning decision further comprises, in response to the fourth partition being an 8×8 partition, evaluating 4×4 sub-partitions of the fourth partition.The machine readable medium of any of claims 19-23, wherein the detection indicators comprise a best mode for the fourth 8×8 partition, and evaluating the 4×4 sub-partitions comprises evaluating only cross modes for the 4×4 sub-partitions if the best mode is a cross mode, and evaluating only internal modes for the 4×4 sub-partitions if the best mode is an internal mode, wherein evaluating only cross modes comprises evaluating the 4×4 sub-partitions by performing a motion estimation search for each of the 4×4 sub-partitions using a selected best cross mode motion vector for the fourth 8×8 partition to define a search center for the motion estimation searches, wherein evaluating only internal modes comprises evaluating the 4×4 sub-partitions using only the best internal mode corresponding to the fourth 8×8 partition, a DC mode, a planar mode, and one or more internal modes adjacent the best internal mode.A system (1800), comprising: means for receiving an input video (111) for encoding, wherein the input video (111) comprises a plurality of images (200) and a first image (301) of the plurality of images (200) comprises an area comprising an individual block, wherein the individual block comprises a plurality of partitions; means for applying one or more detectors to the area, the individual block, and / or one or more of the plurality of partitions to generate one or more detection indicators; means for generating a partitioning decision for the individual block and encoding mode decisions for partitions of the individual block corresponding to the partitioning decision using the detection indicators based on generating an evaluation decision for luma and chroma or only luma for a first partition of the partitions, generating a merging mode decision or exhaust mode decision for a second partition of the partitions that has an initial merging mode decision, generating only a portion of a transform coefficient block for a third partition of the partitions, and / or evaluating 4×4 modes only for a fourth partition of the partitions that is an initial 8×8 encoding partition; and means for encoding the individual block based at least on the partitioning decision to generate a portion of an output bitstream (113); wherein the detection indicators comprise indicators of whether a luma average of the first partition exceeds a first threshold, an average of a first chroma channel of the first partition exceeds a second threshold, an average of a second chroma channel of the first partition exceeds a third threshold, the first partition includes an edge, and the first partition is in a revealed area, and generating the partitioning decision and the encoding mode decisions comprises generating the evaluation decision for luma and chroma or only luma for the first partition by applying an evaluation decision only for luma for the first partition; and, when the luma average does not exceed the first threshold, the average of the first chroma channel does not exceed the second threshold, and the average of the second chroma channel does not exceed the third threshold.The system (1800) of claim 25, comprising means for applying a luma and chroma evaluation decision for the first partition in response to the luma average, the average of the first chroma channel, or the average of the second chroma channel exceeding its corresponding threshold and the first partition including an edge or being in a revealed area.The system (1800) of claim 25 or 26, wherein the detection indicators comprise a determination of whether an amount of a difference between initial exhaust mode encoding cost and initial merging mode encoding cost for the second partition exceeds a threshold, and generating the partitioning decision and the encoding mode decisions comprises generating the merging mode decision or exhaust mode decision by selecting the exhaust mode encoding or the merging mode encoding for the second partition if the amount of the difference exceeds the threshold to generate a final exhaust mode decision or merging mode decision, or shifting the selection of the exhaust mode encoding or the merging mode encoding to a merging mode decision or an exhaust mode decision of a full encoding pass if the amount of the difference does not exceed the threshold.The system (1800) of claims 25-27, wherein the detection indicators comprise an indicator of whether the region, the individual block, or the third partition is optically important, and generating the partitioning decision and the encoding mode decisions comprises generating only the portion of the transform coefficient block by generating a first transform coefficient block having a first number of available transform coefficients if the region, the individual block, or the third partition is optically important, or generating a second transform coefficient block having a second number of available transform coefficients if the region, the individual block, or the third partition is not optically important, wherein the second number is less than the first number.The system (1800) of claims 25-28, wherein generating the partitioning decision comprises determining an initial partitioning decision for the individual block that evaluates smallest candidate partitions of 8×8 candidate partitions of the individual block, wherein the initial partitioning decision partitions the individual block into the fourth partition and one or more further partitions, and generating the partitioning decision further comprises, responsive to the fourth partition being an 8×8 partition, evaluating 4×4 sub-partitions of the fourth partition.

Citation Information

Patent Citations

  • Method and Apparatus of Scalable Video Coding

    US20140003495A1