Decoupled Prediction and Coding Structures for Video Coding
By decoupling prediction and video encoding, using standards to adapt to coding to generate partition and pattern decisions, the problems of encoding speed and visual quality artifacts in the prior art are solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN201811533419.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-12-20
- Filing Date
- 2018-12-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2038-12-14
AI Technical Summary
While maintaining video quality, existing video encoding technologies are difficult to improve encoding speed and compression speed, and there are problems with visual quality artifacts.
By decoupling prediction and video encoding, only source samples from all-standard compliant encodings with local decoding loops are used to generate final partition decisions and initial coding mode decisions and transmit them to standard compliant encoders, improving parallelism in the coding process.
It improves encoding speed, reduces the latency and memory requirements of the encoder, and reduces visual quality artifacts, and improves the efficiency and quality of video encoding.
Smart Images

Figure CN110035290B_ABST
Abstract
Description
Background Art
[0001] In a compression / decompression (codec) system, compression efficiency and video quality are important performance criteria. Visual quality is an important aspect of the user experience in many video applications, and compression efficiency affects the amount of memory storage required to store video files and / or the amount of bandwidth required to transmit and / or stream video content. For example, a video encoder compresses video information so that more information can be sent over a given bandwidth, or more information can be stored in a given storage space, etc. The compressed signal or data can then be decoded by a decoder that decodes or decompresses the signal or data for display to the user. In most implementations, higher visual quality with a higher degree of compression is desired. Additionally, encoding speed and efficiency are important aspects of video encoding.
[0002] It may be advantageous to increase video encoding speed and compression rate while maintaining or even improving video quality. Current improvements are needed with respect to these and other considerations. Such improvements may become critical as the need to compress and transmit video data becomes more prevalent. Brief Description of the Drawings
[0003] The content described herein is shown by way of example and not as a limitation to the drawings. For simplicity and clarity of illustration, the elements shown in the drawings are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Additionally, where considered appropriate, reference numerals are repeated in the drawings to indicate corresponding or similar elements.
[0004] In the drawings:
[0005] Figure 1 is a schematic diagram of an exemplary system for providing video encoding;
[0006] Figure 2 is a schematic diagram of another exemplary system for providing video encoding;
[0007] Figure 3 shows an exemplary group of pictures;
[0008] Figure 4 shows an exemplary video image;
[0009] Figure 5 is a schematic diagram of an exemplary partitioning and mode decision module for providing LCU partitioning and intra / inter mode data;
[0010] Figure 6 is a schematic diagram of an exemplary encoder for generating a bitstream;
[0011] Figure 7is a flowchart showing an exemplary process for reducing flicker artifacts;
[0012] Figure 8 is a schematic diagram of an exemplary flat and noisy region detector;
[0013] Figure 9 is a flowchart showing an exemplary process for reducing line artifacts;
[0014] Figure 10 is a flowchart showing an exemplary process for video coding;
[0015] Figure 11 is a schematic diagram of an exemplary system for video coding;
[0016] Figure 12 is a schematic diagram of an exemplary system; and
[0017] Figure 13 shows an exemplary device arranged according to at least some implementations of the present disclosure.
[0018] Specific Implementations
[0019] One or more embodiments or implementations are now described with reference to the accompanying drawings. Although specific configurations and arrangements are discussed, it should be understood that this is for illustrative purposes only. Those skilled in the relevant art will recognize that other configurations and arrangements can be employed without departing from the spirit and scope of this specification. It will be apparent to those skilled in the relevant art that the techniques and / or arrangements described herein can also be used in a variety of other systems and applications beyond those described herein.
[0020] Although the following description sets forth various implementations that can be exhibited in an architecture such as a system-on-chip (SoC) architecture, the implementations of the techniques and / or arrangements described herein are not limited to a particular architecture and / or computing system, and can be implemented by any architecture and / or computing system for similar purposes. For example, various architectures and / or various computing devices and / or consumer electronics (CE) devices such as set-top boxes, smart phones, etc., employing, for example, multiple integrated circuit (IC) chips and / or packages can implement the techniques and / or arrangements described herein. Additionally, although the following description may set forth many specific details such as logical implementations, types of system components and their interrelationships, logical partitioning / integration choices, etc., the claimed subject matter can be practiced without these specific details. For example, in other instances, some content may not be shown in detail, such as control structures and complete software instruction sequences, so as not to obscure the content disclosed herein.
[0021] The subject matter disclosed herein can be implemented in hardware, firmware, software, or any combination thereof. The subject matter disclosed herein can also be implemented as instructions stored on a machine-readable medium, which can be read and executed by one or more processors. A machine-readable medium can include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium can include read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.), and the like.
[0022] References in the specification to "one implementation," "an implementation," "an exemplary implementation," etc., mean that the described implementation may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include that particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same implementation. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is considered within the knowledge of one of ordinary skill in the art to implement such feature, structure, or characteristic in connection with other implementations whether or not explicitly described.
[0023] Methods, devices, apparatuses, computing platforms, and articles related to video coding are described herein, and specifically, relate to decoupled prediction and video coding.
[0024] The techniques discussed herein provide for the generation of decoupling of the final partitioning decision and the associated initial coding mode decision based on using only source samples from a fully compliant coding with compliant local decoding loops. For example, the mode decision module generates the final partitioning decision and the associated initial coding mode decision without using any data generated by a compliant encoder with a compliant decoding loop. After generating such final partitioning and initial coding mode decisions, an encoder such as a compliant encoder with a compliant local decoding loop employs such decisions to generate a compliant bitstream. That is, the final partitioning decision and the initial coding mode decision are generated independently and then transmitted to the compliant encoder, where the prediction error information (i.e., the residual data and the corresponding transform coefficients) is generated in a compliant manner. For example, by generating the final prediction data including the partitioning / mode decision and associated data such as motion vectors using only the input source image (e.g., source samples), the decoupling of the mode decision from the compliant coding can allow each of the two processes to operate on separate hardware devices, which improves the parallelism in the coding process, thereby increasing the coding speed and reducing the latency and memory requirements of the encoder. As used herein, the term "sample" or "pixel sample" can be any suitable pixel value. The term "original pixel sample" is used to indicate a sample or value from the input video and is contrasted with a reconstructed pixel sample, which is not an original pixel sample but is reconstructed after the encoding and decoding operations in a compliant encoder.
[0025] As further discussed herein, in various embodiments, the independent motion estimation and mode decision module may employ various approximation tools and techniques to generate the prediction data and the partitioning / mode decision, which provides a substantial gain in terms of speed. Using only the source input image in such mode decision processing and using various approximations may introduce video quality artifacts as compared to using a fully compliant encoder for such prediction and partitioning / mode decision. The techniques discussed herein employ only source samples and introduce various approximations while minimizing or eliminating such artifacts.
[0026] By decoupling the partitioning / mode decision from the compliant coding that uses the partitioning / mode decision, two key aspects arise. First, the functionality for the partitioning / mode decision does not need to use compliant techniques (although the actual coding will be compliant). Second, the information flow is unidirectional such that information only flows from the partitioning / mode decision module or process to the coding module or process. For example, the reconstructed pixel samples using compliant coding are not used for the partitioning / mode decision transmitted to the compliant encoder.
[0027] Figure 1 is a schematic diagram of an exemplary system 100 arranged according to at least some implementations of the present disclosure for providing video coding. As Figure 1As shown, system 100 includes a partitioning and mode decision module 101 and an encoder 102. As shown, the partitioning and mode decision module 101 (which may be characterized as a partitioning, motion estimation, and mode decision module, etc.) receives an input video 111 and generates maximum coding unit (LCU) partitioning and corresponding coding mode (intra / inter mode) data 112, which may be characterized as final partitioning / mode decision data, final partitioning / initial mode decision data, etc. For example, for each LCU of each image of the input video 111, the partitioning and mode decision module 101 may provide a final partitioning decision (i.e., data indicating how to divide the LCU into coding units (CUs)), the coding mode of each CU (i.e., inter mode, intra mode, etc.), and (if needed) information about the coding mode (i.e., motion vectors for inter coding).
[0028] As shown, the encoder 102 receives the LCU partitioning and intra / inter mode data 112 and generates a bitstream 113, such as a standard-compliant bitstream. For example, the encoder 102 implements the LCU partitioning and intra / inter mode data 112 such that the encoder 102 does not make any such decisions, the encoder 102 does not make any partitioning decisions, etc. Instead, the encoder implements the final decisions made by the partitioning and mode decision module 101, optionally adjusts any initial mode decisions made by the partitioning and mode decision module 101, and implements such partitioning and mode decisions to generate a standard-compliant bitstream 113.
[0029] As shown, system 100 receives an input video 111 to be encoded, and the system provides video compression to generate a bitstream 113, where system 100 may be a video encoder implemented via a computer or computing device, etc. The bitstream 113 may be any suitable bitstream, such as a standard-compliant bitstream. For example, the bitstream 113 may be compliant with the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, the H.265 High Efficiency Video Coding (HEVC) standard, the VP9 standard, etc. For example, system 100 may be implemented via any suitable device, such as a personal computer, laptop computer, tablet, phablet, smartphone, digital camera, gaming console, wearable device, all-in-one device, two-in-one device, etc. or a platform such as a mobile platform. For example, as used herein, a system, device, computer, or computing device may include any such device or platform.
[0030] The input video 111 may include any suitable video frames, video images, sequences of video frames, sets of images, multiple sets of images, video data, etc. having any suitable resolution. For example, the video may be video graphics array (VGA), high definition (HD), full high definition (e.g., 1080p), 4K resolution video, 8K resolution video, etc., and the video may include any number of video frames, sequences of video frames, images, sets of images, etc. For clarity of presentation, the techniques discussed herein are discussed with respect to images and blocks and / or coding units. However, such images may be characterized as frames, video frames, sequences of frames, video sequences, etc., and such blocks and / or coding units may be characterized as coding blocks, macroblocks, sub-units, sub-blocks, regions, sub-regions, etc. Generally, the terms "block" and "coding unit" may be used interchangeably herein. For example, an image or frame of color video data may include a luminance plane or component (i.e., luminance pixel values) and two chrominance planes or components (i.e., chrominance pixel values) of the same or different resolution relative to the luminance plane. The input video 111 may include an image or frame that may be partitioned into blocks and / or coding units of any size, which contain data corresponding to, for example, M×N blocks and / or coding units of pixels. Such blocks and / or coding units may include data from one or more planes or color channels of the pixel data. As used herein, the term "block" may include macroblocks, coding units, etc. of any suitable size. It will be understood that these blocks may also be partitioned into sub-blocks for prediction, transformation, etc.
[0031] Figure 2 is a schematic diagram of another exemplary system 200 arranged according to at least some implementations of the present disclosure for providing video coding. As Figure 2As shown, system 200 includes the components of system 100, with the addition of bit depth limiter / color sampler reduction module 201. For example, bit depth limiter / color sampler reduction module 201 can receive input video 111 and perform one or both of bit depth limitation and color sample reduction. Bit depth limitation is used to reduce the bit depth of input video 111 (i.e., by keeping the most significant bits and discarding the least significant bits), and color sample reduction is used to reduce the color sampling of input video 111 to provide a reduced bit depth and / or subsampled video 212, which is provided as an 8-bit 4:2:0 video in the illustrated embodiment. For example, partitioning and mode decision module 101 can operate on the reduced bit depth and / or subsampled video for generating LCU partitioning and intra / inter mode data 112, while encoder 102 operates on the full bit depth and / or full color sampled input video. In an embodiment, input video 111 is a 10-bit 4:2:2 video, and as shown, the reduced bit depth and / or subsampled video 212 is an 8-bit 4:2:0 video. However, input video 111 and the reduced bit depth and / or subsampled video 212 can be any video data, where input video 111 is at a higher bit depth and / or higher color sampling than the reduced bit depth and / or subsampled video 212. This reduction in bit depth and / or color sampling can reduce the computational resources and / or memory transfer requirements of partitioning and mode decision module 101.
[0032] For example, input video 111 can be received at a bit depth of at least 8 bits (i.e., the luminance and chrominance values associated with a given source pixel / sample are represented using at least 8 bits per value (e.g., 10 bits per value)). Data with more than 8 bits per pixel / sample requires more memory transfer (for moving data between memory and the processor) and more complex arithmetic operations compared to 8-bit data. To reduce the impact of high bit depth on the required memory and computational resources, input video 111 can be converted to 8-bit data by keeping the eight most significant bits. For example, for 10-bit input video data, the two least significant bits are discarded. Additionally, input video 111 with a higher color representation (e.g., 4:2:2 or 4:4:4) includes increased chrominance information. However, processing the chrominance information in 4:2:0 video data (where the number of chrominance samples is half the number of luminance samples) can provide a balance between the video quality value using the chrominance information and its computational and memory transfer costs.
[0033] Figure 3 An exemplary group of pictures 300 arranged in accordance with at least some implementations of the present disclosure is shown. As Figure 3As shown, the picture group 300 may include any number of pictures 301, such as 64 pictures (shown as 0 - 16), etc. In addition, the pictures 301 may be provided in chronological order 302 such that the pictures 301 are presented in chronological order, while the pictures 301 are encoded in an encoded order (not shown) such that the encoded order is different from the chronological order 302. In addition, the pictures 301 may be provided in a picture hierarchy 303 such that the base layer (L0) of the pictures 301 includes pictures 0, 8, 16, etc., the non - base layer (L1) of the pictures 301 includes pictures 4, 12, etc., the non - base layer (L2) of the pictures 301 includes pictures 2, 6, 10, 14, etc., and the non - base layer (L3) of the pictures 301 includes pictures 1, 3, 5, 7, 9, 11, 13, 15, etc. For example, moving through the hierarchy, for the inter - frame mode, the pictures in L0 may only refer to other pictures in L0, the pictures in L1 may only refer to the pictures in L0, the pictures in L2 may only refer to the pictures in L0 or L1, and the pictures in L3 may refer to the pictures in any one of L0 - L2. For example, as shown, the pictures 301 include base - layer pictures and non - base - layer pictures such that the base - layer pictures are reference pictures for the non - base - layer pictures, but the non - base - layer pictures are not reference pictures for the base - layer pictures. In an embodiment, the input video 111 includes the picture group 300, and / or the systems 100, 200 implement the picture group 300 with respect to the input video 111. Although shown with respect to the exemplary picture group 300, the input video 111 may have any suitable structure implementing the picture group 300, other picture format groups, etc.
[0034] In an embodiment, the prediction structure for encoding video includes a picture group such as the picture group 300. For example, in the context of broadcast and streaming implementations, the prediction structure may be periodic and may include a periodic picture group (GOP). In an embodiment, the GOP includes approximately 1 second of pictures organized in the structure described in Figure 3 and then another GOP starting with an I - picture, and so on.
[0035] Figure 4An exemplary video image 401 arranged according to at least some implementations of the present disclosure is shown. The video image 401 may include any image of a video sequence or clip, such as a video image in VGA, HD, full HD, 4K, 8K, etc. For example, the video image 401 may be any one of the images 301 of the group of pictures 300. As shown, the video image 401 may be segmented or partitioned into one or more segments, as shown by segment 402 of the video image 401. Additionally, as shown with respect to the LCU 403, the video image 401 may be segmented or partitioned into one or more LCUs, which in turn may be segmented into one or more coding units, such as shown by CUs 405, 406, and / or prediction units (PUs) and transform units (TUs) (not shown).
[0036] Figure 5 is a schematic diagram of an exemplary partitioning and mode decision module 101 arranged according to at least some implementations of the present disclosure for providing LCU partitioning and intra / inter mode data 112. As Figure 5As shown, the partitioning and mode decision module 101 may include or implement an LCU loop 521, which includes a source sample (SS) motion estimation module 501, an SS intra search module 502, a CU fast loop processing module 503, a CU full loop processing module 504, an inter depth decision module 505, and a skip-merge decision module 507. As shown, the LCU loop 521 receives an input video 111 or a reduced bit depth and / or color subsampled video 212, and the LCU loop 521 generates final LCU partitioning and initial mode decision data 518. The final LCU partitioning and initial mode decision data 518 may be any suitable data indicating or describing the LCU partitioning into CUs and the coding mode decision for each CU of the LCU. In an embodiment, the final LCU partitioning and initial mode decision data 518 includes final partitioning data that will be implemented without modification by the encoder 102 and initial mode decisions that may be modified by the encoder final LCU partitioning and mode decision data 518 as discussed herein. In an embodiment, although described herein as initial mode decisions, such mode decisions may be final and implemented without modification by the encoder 102. For example, the mode decisions may be initial or final. For example, the coding mode decision may include an intra mode (i.e., one of the available intra modes based on the standard being implemented) or an inter mode (i.e., skip, merge, or motion estimation, ME). Additionally, the LCU partitioning and mode decision data 518 may include any additional data required for a particular mode (e.g., motion vectors for an inter mode). For example, in the context of HEVC, a coding tree unit may be 64×64 pixels, which may define an LCU. The LCU may be partitioned into CUs for encoding via quadtree partitioning such that a CU may be 32×32, 16×16 pixels, or 8×8 pixels. Such partitioning may be indicated by the LCU partitioning and mode decision data 518. Additionally, such partitioning is used to evaluate candidate partitions (candidate CUs) of the LCU.
[0037] As shown in the figure, the SS motion estimation module 501 receives the input video 111 or the video 212 with reduced bit depth and / or subsampled color. In the following discussion, for the sake of clear presentation, the input video 111 or the video 212 with reduced bit depth and / or subsampled color is characterized as the input videos 111, 212. The SS motion estimation module 501 performs a motion search on the CU or candidate partition of the current image of the input videos 111, 212 using one or more reference images of the input videos 111, 212. That is, the SS motion estimation module 501 performs a motion search on the CU of the current image by searching for matching CUs or one or more reference images of the input videos 111, 212, such that the reference images only include the original pixel samples of the input videos 111, 212. For example, the SS motion estimation module 501 performs a motion search without reconstructing the pixels into a reconstructed reference image using a local decoding loop. For example, the SS motion estimation module 501 evaluates multiple inter-frame modes (such as different motion vectors, reference images, etc.) of the candidate partition of a block or CU by comparing the candidate partition with a search partition that only includes the original pixel samples from the reference image input video 111. For example, for a specific candidate partition, such search partitions are those searched during motion estimation. As shown in the figure, the SS motion estimation module 501 generates motion estimation candidates 511 (i.e., MVs) corresponding to the CUs of the specific partitions of the current LCU being evaluated. For example, for each CU, one or more MVs can be provided. In an embodiment, the SS motion estimation module 501 uses a non-standard compliant interpolation filter to generate an interpolation search region for sub-pixel MV search.
[0038] In addition, the SS intra search module 502 receives the input videos 111, 212, and the SS intra search module 502 generates an intra mode for a CU of the current image of the input videos 111, 212 using the current image of the input videos 111, 212. That is, the SS intra search module 502 performs an intra mode evaluation on the CU or candidate partition of the current image by comparing the CU with an intra prediction block generated using the original pixel samples of the current image of the input videos 111, 212 (based on the current intra mode being evaluated). For example, the SS intra search module 502 performs an intra mode evaluation without reconstructing the pixels into reconstructed pixel samples (e.g., of a previously encoded CU) using a local decoding loop. As shown, the SS intra search module 502 generates an intra candidate 512 (i.e., the selected intra mode) corresponding to the CU of a particular partition of the current LCU being evaluated. For example, for each CU, one or more intra candidates may be provided. In an embodiment, as discussed herein, the best partition decision and corresponding best intra and / or inter candidates (e.g., having the lowest distortion or lowest rate distortion cost, etc.) from the motion estimation candidates 511 and intra candidates 512 are provided for use by the encoder 102. For example, subsequent processing may be skipped. The rate distortion cost may include any suitable rate distortion cost, such as the sum of the distortion and the product of the Lagrange multiplier and the rate. For example, the distortion may be a measure of squared error distortion, HVS weighted squared error distortion, etc.
[0039] The CU fast loop processing module 503 receives the motion estimation candidates 511, the intra candidates 512, and the neighboring data 516, and as shown, generates MV merge candidates, generates advanced motion vector prediction (AMVP) candidates, and makes a CU mode decision. The neighboring data 516 includes any suitable data of the spatially neighboring CUs of the current CU being evaluated, such as the intra and / or inter modes of the spatially neighboring CUs. The CU fast loop processing module 503 uses any suitable technique to generate the MV merge candidates. For example, the merge mode may use the MVs of the spatially neighboring CUs of the current CU to provide motion inference candidates. For example, one or more MVs from the spatially neighboring CUs may be provided (e.g., inherited) as MV candidates for the current CU. In addition, the CU fast loop processing module 503 uses any suitable technique to generate the AMVP candidates. In an embodiment, the CU fast loop processing module 503 may use data from the reference image and data from the neighboring CUs to generate the AMVP candidate MVs. In addition, non-standard compliant techniques may be used when generating the MV merge and / or AMVP candidates. Only the source samples are used to generate the predictions for the MV merge and AMVP candidates.
[0040] As shown in the figure, the CU fast loop processing module 503 makes an encoding mode decision for each CU of the current partition based on motion estimation candidates 511, intra candidates 512, MV merge candidates, and AMVP candidates. Any suitable technique can be used to make the encoding mode decision. In an embodiment, the sum of the distortion measurement value and the weighted rate estimation value is used to evaluate the intra and inter modes of the CU. For example, the distortion between the current CU and the predicted CU (generated using the corresponding mode) can be determined and combined with the estimated encoding rate to determine the best candidate. As shown in the figure, a subset 513 of the ME, intra, and merge / AMVP candidates can be generated as a subset of all available candidates.
[0041] The subset 513 of the ME, intra, and merge / AMVP candidates is provided to the CU full loop processing module 504. As shown in the figure, the CU full loop processing module 504 performs forward transform, forward quantization, inverse quantization, and inverse transform on the residual block of each encoding mode of the subset 513 of the ME, intra, and merge / AMVP candidates (i.e., the residual is the difference between the CU and the predicted CU generated using the current mode) to form a reconstructed residual. Then, the CU full loop processing module 504 generates the reconstruction of the CU (i.e., by adding the reconstructed residual to the predicted CU) and measures the distortion of each mode of the subset 513 of the ME, intra, and merge / AMVP candidates. The mode with the best rate distortion is selected as the CU mode 514.
[0042] The CU mode 514 is provided to the inter-depth decision module 505, which can evaluate the available partitions of the current LCU to generate LCU partition data 515. As shown in the figure, the LCU partition data 515 is provided to the skip-merge decision module 507, which determines whether a CU is a skip CU or a merge CU for any CU having an encoding mode corresponding to a merged MV. For example, for a merge CU, the MV is inherited from spatially adjacent CUs, and the residual is sent for that CU. For a skip CU, the MV is inherited from spatially adjacent CUs (as in the merge mode), but the residual is not sent for that CU.
[0043] As shown in the figure, after such a merge-skip decision, the LCU loop 521 provides the final LCU partition and the initial mode decision data 518, as discussed, which indicates or describes partitioning the LCU into CUs and the encoding mode decision for each CU (and any information required for that mode decision).
[0044] Figure 6 is a schematic diagram of an exemplary encoder 102 arranged according to at least some implementations of the present disclosure for generating a bitstream 113. As Figure 6As shown, the encoder 102 may include or implement an LCU loop 621 (e.g., an LCU loop for one - pass encoding), which includes a CU loop processing module 601 and an entropy encoding module 602. Additionally, as shown, the encoder 102 may include a packetization module 603. As shown, the LCU loop 621 receives the input video 111 and the final LCU partition and initial mode decision data 518, and the LCU loop 621 generates quantized transform coefficients, control data, and parameters 613, which may be entropy - encoded by the entropy encoding module 602 and packetized by the packetization module 603 to generate a bitstream 113.
[0045] For example, the CU loop processing module 601 receives the input video 111 and the final LCU partition and initial mode decision data 518. As shown, based on the final LCU partition and initial mode decision data 518, the CU loop processing module 601 generates intra - reference pixel samples for intra CUs (as needed). For example, adjacent reconstructed pixel samples (generated via a local decoding loop) may be used to generate intra - reference pixel samples. As shown, for each CU, the CU loop processing module 601 generates a predicted CU using neighbor data 611 (e.g., data from adjacent CUs of the current CU) as needed. For example, for an inter - frame mode, a predicted CU may be generated by obtaining previously reconstructed pixel samples of the CU indicated by one or more MVs from one or more reconstructed reference images, and if needed, by combining the obtained reconstructed pixel samples to generate the predicted CU. For an intra - frame mode, a predicted CU may be generated using adjacent reconstructed pixel samples of the image of the CU based on the intra - frame mode of the current CU. As shown, a residual is generated for the current CU. For example, a residual may be generated by taking the difference between the current CU and the predicted CU.
[0046] Then, forward transformation and forward quantization are performed on the residual to generate quantized transform coefficients, which are included in the quantized transform coefficients, control data, and parameter 613. Additionally, in a local decoding loop, for example, the transform coefficients are inverse quantized and inverse transformed to generate a reconstructed residual for the current CU. As shown, the CU loop processing module 601 performs reconstruction for the current CU by, for example, adding the reconstructed residual to the predicted CU (as described above) to generate a reconstructed CU. The reconstructed CU can be combined with other CUs to use additional techniques such as sample adaptive offset (SAO) filtering to reconstruct the current image or a portion thereof, which can include generating SAO parameters (which are included in the quantized transform coefficients, control data, and parameter 613) and implementing the SAO filter on the reconstructed CU and / or deblocking loop filter (DLF), which can include generating DLF parameters (which are included in the quantized transform coefficients, control data, and parameter 613) and implementing the DLF filter on the reconstructed CU. For example, such a reconstructed CU can be provided as a reference image (e.g., stored in a reconstructed image buffer). Such a reference image or a portion thereof is provided as a reconstructed sample 612, which is used as described above to generate a predicted CU (in both inter-frame and intra-frame modes).
[0047] As shown, the quantized transform coefficients, control data, and parameter 613 (which includes the transform coefficients of the residual coding unit, control data such as the final LCU partition and mode decision data (i.e., from the final LCU partition and initial mode decision data 518), and parameters such as SAO / DLF filter parameters) can be entropy encoded and grouped to form a bitstream 113. The bitstream 113 can be any suitable bitstream, such as a standard-compliant bitstream. For example, the bitstream 113 can be compliant with the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, the H.265 High Efficiency Video Coding (HEVC) standard, the VP9 standard, etc.
[0048] As discussed, the partitioning and mode decision module 101 generates partitioning and mode decision data for the LCU of the input video 111 using the original pixel samples, and the encoder 102 implements the partitioning and mode decision data on the input video 111 (including using a local decoding loop) to generate a bitstream 113. Generating the partitioning and mode decision data using only the original pixel samples can provide decoupling between the partitioning and mode decision module 101 (which can be implemented as hardware such as an application-specific integrated circuit) and the encoder 102 (which can be implemented as separate hardware, such as a separate application-specific integrated circuit). This decoupling provides efficiencies such as hardware decoupling, parallelization, processing speed, etc. However, generating the partitioning and mode decision data using only the original pixel samples may degrade the visual quality. Now, the discussion turns to techniques for improving the visual quality in a decoupled system.
[0049] Figure 7is a flowchart showing an exemplary process 700 arranged according to at least some implementations of the present disclosure for reducing flicker artifacts. As Figure 7 shown, process 700 may include one or more operations 701 - 706. Process 700 may be executed by a system (e.g., system 100 as discussed herein) to reduce flicker artifacts in video coding. In an embodiment, process 700 is implemented by the CU fast loop processing module 503 and the CU full loop processing module 504 of the partitioning and mode decision module 101.
[0050] For example, flicker in a noisy flat region may occur in a region that is spatially flat and noisy. Examples of such regions include very noisy sky areas in an image. Since the coding modes (e.g., intra, inter) are non-uniformly distributed both spatially (i.e., the region contains a mixture of intra- and inter-coded blocks) and temporally (the coding mode changes over time), flicker appears in such regions. For example, the appearance of flicker may be due to a switch from one frame to the next between the intra mode and the inter mode. Blocks encoded using the intra mode may be smooth and not contain much noise, while blocks encoded using the inter mode may be sharper due to better noise reproduction. The change in the coding mode between adjacent blocks may cause a discontinuity in sharpness and thus result in flicker in the image. Process 700 can address such flicker artifacts.
[0051] Processing begins at operation 701, where a region is selected for evaluation. The region can be any suitable region of the input image, such as the whole image, a segment of the image, a quadrant of the image, a predefined grid portion of the image, or an LCU of the image. The processing continues at operation 702, where it can be determined whether the selected region is a flat and noisy region by performing flat and noisy region detection. The flat and noisy region detection can be performed using any suitable technique, such as those Figure 8 discussed.
[0052] Figure 8 is a schematic diagram of an exemplary flat and noisy region detector 800 arranged according to at least some implementations of the present disclosure. As Figure 8As shown, the flat and noisy region detector 800 may include a denoiser 801, a differencer 802, a flatness checking module 803, and a noise checking module 804. As shown, the denoiser 801 receives the input region 811 and denoises the input region 811, using any suitable technique (such as a filtering technique) to produce a denoised region 812. The denoised region 812 is provided to the flatness checking module 803, which checks the flatness of the denoised region 812 using any suitable technique. In an embodiment, the flatness checking module 803 determines the variance of the denoised region 812 and compares the variance with a predetermined threshold. If the variance does not exceed the threshold, a flatness indicator 813 is provided, indicating that the denoised region 812 is flat. Additionally, the input region 811 and the denoised region 812 are provided to the differencer 802, which may use any suitable technique to differ the input region 811 and the denoised region 812 to produce a difference 814. As shown, the difference 814 is provided to the noise checking module 804, which checks the difference 814 using any suitable technique to determine whether the input region 811 is a noisy region. In an embodiment, the noise checking module 804 determines the variance of the difference 814 and compares the variance with a predetermined threshold. If the variance meets or exceeds the threshold, a noise indicator 815 is provided, indicating that the input region 811 is noisy. If both the flatness indicator 813 and the noise indicator 815 are confirmed for the input region 811, the input region 811 is determined to be a flat and noisy region.
[0053] Referring again to Figure 7 , the process continues at operation 703, where edge detection is performed on the region to determine whether any edges exist within the region or block. Any suitable technique may be used to perform the edge detection. The process continues at operation 704, where the change in the region over time is evaluated. For example, for the current image, the change in the region is determined. Additionally, a plurality of images with respect to the current image are selected. Such images may be adjacent to the current image in time (before or after or both) or adjacent in the coding order. For the juxtaposed regions of such images, the change in each such region is also determined. The variance of the plurality of changes (i.e., including the change in the region of the current image and the changes in the juxtaposed regions of the plurality of images) is determined. Then the variance of the changes is compared with a threshold. If the variance does not exceed the threshold, the region is considered to be uniform over time. If the variance of the changes exceeds the threshold, the region is discarded from the processing discussed in connection with operation 705 below.
[0054] Processing continues at operation 705, where, during the coding mode decision for coding units within a region that is flat and noisy (as detected at operation 702), includes one or more edges (as detected at operation 703), and is temporally uniform (as indicated by the variance of spatial variations across multiple images being less than a threshold, as detected in operation 704), a bias can be applied to the inter-frame coding mode (as opposed to the intra-frame coding mode) of the coding units. For example, if the region is not flat and noisy, no edges are detected, or the region is not temporally uniform, no bias is applied to the coding units within the region. Any suitable technique can be used to apply the bias. In an embodiment, a predetermined weighting factor is added to or multiplied by the cost of the intra-frame mode, and not to the cost of the inter-frame mode. In an embodiment, additionally or alternatively, a predetermined additional term can be subtracted from the cost of the inter-frame mode, or the cost of the inter-frame mode can be divided by a factor. For example, the inter-frame and intra-frame mode costs can be rate-distortion costs. Processing continues at operation 706, where a coding mode is selected for coding units within the flat and noisy region. For example, any suitable technique can be used to select the coding mode while applying the bias discussed with respect to operation 703. For example, rate-distortion optimization techniques can be used to select the coding mode while applying a bias to the use of the inter-frame mode. Such inter-frame modes can include any inter-frame mode, such as inter-frame, merge, or AMVP modes.
[0055] As discussed with respect to process 700, flat region detection and noisy region detection can be performed for regions of an input image to determine whether the region is a flat and noisy region, and generating a coding mode decision for a block of the region can include: in response to the region being a flat and noisy region, applying a bias to one or both of an intra-frame mode result and an inter-frame mode result to favor selecting the inter-frame mode result for coding the block.
[0056] For any number of regions of an input image, process 700 or portions thereof can be repeated any number of times, serially or in parallel.
[0057] Figure 9 is a flowchart showing an exemplary process 900 arranged according to at least some implementations of the present disclosure for reducing line artifacts. As Figure 9 shown, process 900 can include one or more operations 901 - 904. Process 900 can be performed by a system (e.g., system 100 as discussed herein) to reduce line artifacts in video coding. In an embodiment, process 900 is implemented by the CU loop processing module 601 of the encoder 102. For any number of intra-frame coded blocks or coding units of an input image, process 900 or portions thereof can be repeated any number of times, serially or in parallel.
[0058] For example, a line artifact may appear as short line segments in an intra-coded block, such as in a flat area next to an edge. Such a line artifact may be due to an error in a reference sample used in intra prediction during an encoding, such that the reference sample from the reconstructed reference image may contain an error caused by quantization. The pixel error can be magnified to a linear error in some intra-coded modes (e.g., vertical or horizontal modes) by, for example, pixel replication.
[0059] Processing begins at operation 901, where a transform, such as a Hadamard transform, is applied to the residual corresponding to an intra-coded block encoded using an intra mode selected by the partitioning and mode decision module 101. For example, during an encoding implemented by the encoder 102, a Hadamard transform or other suitable transform is applied to the residual generated based on implementing an intra-coded mode (e.g., an intra-coded mode selected using only original pixel samples) selected by the partitioning and mode decision module 101. The processing continues at operation 902, where a transform is applied to the residual corresponding to an intra-coded block encoded using one or more intra modes other than the one selected by the partitioning and mode decision module 101. The other intra modes can be any suitable modes, such as a DC mode, a planar mode, etc.
[0060] The processing continues at operation 903, where the high-frequency energy is determined for each of the transformed residual blocks generated at operations 901, 902. Any suitable technique can be used to determine the high-frequency energy. For example, the high-frequency energy can be determined as a measure (e.g., an average value, a weighted average value, etc.) of the high-frequency components of the transformed residual block (e.g., the residuals not in the upper left quadrant of the transformed residual block, the residuals in the lower right quadrant of the transformed residual block, etc.). For example, the transforms performed at operations 901, 902 generate transformed residual coefficients (e.g., a block of transformed residual coefficients having the same size as the residual block). Then a high-frequency coefficient energy value is generated as a high-frequency feature of the transformed residual coefficients. The high-frequency coefficient energy value can be any suitable value, such as an average value of the coefficients corresponding to the high-energy components, a median value of these coefficients, etc. As discussed, the coefficients selected as the high-frequency coefficient energy value can be any such coefficients that do not include the DC coefficient of the transformed residual coefficients (e.g., the upper left value of the transformed residual coefficient block). In an embodiment, the coefficients in the lower right corner of the transformed residual coefficient block are used. For example, for a 16×16 transformed residual coefficient block, the lower right 4×4 transformed residual coefficients can be used, such that those transformed residual coefficients correspond to the highest frequency components. In an embodiment, all of the transformed residual coefficients of the transformed residual coefficient block are used except for the DC coefficient (e.g., which corresponds to the lowest frequency component). In an embodiment, determining the high-frequency coefficient energy value includes: averaging a plurality of transformed residual coefficients (from the block) except for the DC transformed coefficient of the transformed residual coefficients.
[0061] Processing continues at operation 904, where the intra mode corresponding to the lowest high-frequency energy of those determined at operation 903 can be used to encode the block. In an embodiment, before performing operation 902, the high-frequency energy of the transform residual block generated at operation 901 can be compared with a threshold, and if the high-frequency energy exceeds the threshold, the processing can be performed only at operation 902. Otherwise, the intra mode selected by the partitioning and mode determination module 101 is used to encode the block.
[0062] As discussed with respect to process 900, a transform can be applied at the video encoder to a residual block corresponding to an individual block and a predicted block generated using the intra mode selected by the partitioning and mode decision module 101 (e.g., the residual block is the difference between the individual block and the predicted block) to generate a corresponding transform residual. For example, the transform can be a Hadamard transform, and the intra mode can be intra horizontal or intra vertical. Additionally, a transform can be applied to other residual blocks corresponding to the individual block and predicted blocks generated using other intra modes to generate additional transform residuals. A high-frequency energy value can be determined for each transform residual. In response to one of the other transform residuals having a lower high-frequency energy value than the energy value of the transform residual using the intra mode selected by the partitioning and mode decision module 101, the intra mode selected by the partitioning and mode decision module 101 is discarded and the intra mode with the lowest high-frequency energy value is used to encode the individual block.
[0063] Figure 10 is a flowchart showing an exemplary process 1000 for video coding arranged according to at least some implementations of the present disclosure. As Figure 10 shown, process 1000 can include one or more operations 1001 - 1005. Process 1000 can form at least a part of a video coding process. As a non-limiting example, process 1000 can form at least a part of a video coding process performed by any device or system (e.g., system 100) as discussed herein. Additionally, process 1000 will be described herein with reference to Figure 11 system 1100.
[0064] Figure 11 is a schematic diagram of an exemplary system 1100 for video coding arranged according to at least some implementations of the present disclosure. As Figure 11As shown, system 1100 may include a central processor 1101, a video pre-processor 1102, a video processor 1103, and a memory 1104. Also as shown, the video pre-processor 1102 may include or implement a partitioning and mode decision module 101, and the video pre-processor 1102 may include or implement an encoder 102. In an example of system 1100, the memory 1104 may store video data or related content, such as input video data, image data, partitioning data, mode data, and / or any other data discussed herein.
[0065] As shown, in some embodiments, the partitioning and mode decision module 101 is implemented via the video pre-processor 1102. In other embodiments, the partitioning and mode decision module 101 or portions thereof are implemented via the central processor 1101 or another processing unit (e.g., an image processor, a graphics processor, etc.). Also as shown, in some embodiments, the encoder 102 is implemented via the video processor 1103. In other embodiments, the encoder 102 or portions thereof are implemented via the central processor 1101 or another processing unit (e.g., an image processor, a graphics processor, etc.).
[0066] The video pre-processor 1102 may include any number and type of video, image, or graphics processing units that may provide the operations discussed herein. These operations may be implemented via software or hardware or a combination thereof. For example, the video pre-processor 1102 may include circuitry dedicated to manipulating images, image data, etc. obtained from the memory 1104. Similarly, the video processor 1103 may include any number and type of graphics or image processing units that may provide the operations discussed herein. These operations may be implemented via software or hardware or a combination thereof. For example, the video processor 1103 may include circuitry dedicated to manipulating images, image data, etc. obtained from the memory 1104. The central processor 1101 may include any number and type of processing units or modules that may provide control and other high-level functions for system 1100 and / or provide any of the operations discussed herein. The memory 1104 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, the memory 1104 may be implemented by a cache memory.
[0067] In an embodiment, one or more or portions of the partitioning and mode decision module 101 and the encoder 102 are implemented via an execution unit (EU). The EU may include, for example, programmable logic or circuitry (e.g., logic cores), which may provide a variety of programmable logic functions. In an embodiment, one or more or portions of the partitioning and mode decision module 101 and the encoder 102 are implemented via dedicated hardware such as fixed function circuitry. The fixed function circuitry may include dedicated logic or circuitry and may provide a set of fixed function entry points, which may be mapped to dedicated logic for fixed purposes or functions. In an embodiment, the partitioning and mode decision module 101 is implemented via a field programmable grid array (FPGA).
[0068] Returning to Figure 10 the discussion, the process 1000 may begin at operation 1001, where an input video is received for encoding, where the input video includes a plurality of images. Any suitable technique may be used to receive the input video, and the input video may be video data in any suitable format. In an embodiment, the process 1000 further includes: reducing one of the bit depth and chroma subsampling of the input video to generate a second input video, where the encoding mode decision generated at operation 1002 is made using the lower bit depth or chroma subsampling video, and the difference taking discussed with respect to operation 1004 uses the higher bit depth or chroma subsampling video.
[0069] The processing continues at operation 1002, where a final partitioning decision is generated for a single block of the first image among the plurality of images by evaluating a plurality of intra modes for candidate partitions for a single block using a second image among the plurality of images and evaluating a plurality of inter modes for candidate partitions for a single block. In an embodiment, evaluating a plurality of intra modes for candidate partitions for a single block includes: comparing the candidate partition with an intra prediction partition generated using only the original pixel samples from the first image, and evaluating a plurality of inter modes for the candidate partition includes: comparing the candidate partition with a plurality of search partitions including only the original pixel samples from the second image among the plurality of images. In an embodiment, comparing the candidate partition with the intra prediction partition generated using only the original pixel samples from the first image includes: generating a first intra prediction partition using only the original pixel samples according to the current intra mode evaluated for the first candidate partition, and taking the difference between the first candidate partition and the first intra prediction partition.
[0070] In an embodiment, generating the final partitioning decision further includes: generating an initial encoding mode decision for the partitioning of the single block corresponding to the final partitioning decision based on the evaluation of the plurality of intra modes and the plurality of inter modes.
[0071] In an embodiment, a single block is part of a region of a first image, and processing 1000 further includes: performing flat region detection and noisy region detection on the region, performing edge detection on the region to determine whether the region includes detected edges, and evaluating a temporal variance of changes of the region over a plurality of images including the first image to determine whether the region has a low variance of change, and generating an initial coding mode decision for the single block includes: applying a bias to one or both of an intra mode result and an inter mode result in response to the region being a flat and noisy region, the region including edges, and the region having a low temporal variance of change, to favor selection of the inter mode result for one or more partitions of the single block. For example, performing flat region detection may include: denoising the region and determining that a change of the denoised region is less than a first threshold, and performing noisy region detection includes: taking a difference between the region and the denoised region and determining that a variance of the difference region is greater than a second threshold, and evaluating the temporal variance of changes of the region over a plurality of images may include: determining a change of each juxtaposed region over the plurality of images relative to the region, determining a variance of the changes, and determining that the variance of the changes is less than a third threshold. The intra and inter mode results may include any suitable results or costs, such as a rate distortion cost, etc. Any suitable technique may be used to apply the bias, such as adding a first factor to the intra mode result or multiplying a second factor by the intra mode result.
[0072] In an embodiment, a first initial coding mode decision for a first partition of a single block is a first intra mode among a plurality of intra modes, and process 1000 further includes: at a video encoder, applying a transform to a first residual partition corresponding to the first partition and the first intra mode to generate a first transformed residual; at the video encoder, applying a transform to a second residual partition corresponding to the first partition and a second intra mode to generate a second transformed residual; determining a first high-frequency energy value of the first transformed residual and a second high-frequency energy value of the second transformed residual, discarding the first intra mode in response to the second high-frequency value being less than the first high-frequency value; and at the video encoder, encoding the first partition using the second intra mode. The transform applied to the first and second residual partitions may be any suitable transform, such as a Hadamard transform. In an embodiment, determining the first high-frequency coefficient energy value includes averaging a plurality of first transformed residual coefficients other than a DC transform coefficient of the first transformed residual coefficients.
[0073] Processing continues at operation 1003, where a final partition decision generated at operation 1002 is provided to the video encoder. In some embodiments, an initial coding mode decision may also be provided to the encoder. The video encoder may be any suitable video encoder. In an embodiment, the video encoder is a standard compliant video encoder.
[0074] Processing continues at operation 1004, where, at the video encoder, a single block is subtracted from a reconstructed block generated based on a final partitioning decision to generate a residual block corresponding to the single block. For example, a standard-compliant local decoding loop is applied based on the final partitioning decision to generate the reconstructed block. The coding mode used to generate the reconstructed block can be the initial coding mode decision for the partitioning of the single block, or the initial coding mode decision can be modified by the encoder, as discussed herein.
[0075] Processing continues at operation 1005, where the transform coefficients of the residual block are encoded into an output bitstream. Any suitable technique can be used to encode the transform coefficients to generate any suitable output bitstream. In an embodiment, the output bitstream is a standard (e.g., AVC, HEVC, VP9, etc.)-compliant bitstream. In an embodiment, the transform coefficients are generated by performing a forward transform and forward quantization on the residual block.
[0076] For any number of input video sequences, images, coding units, blocks, etc., process 1000 can be repeated any number of times serially or in parallel. As discussed, process 1000 can provide video coding that generates coding mode decisions using only raw pixel samples or values, where the coding mode decisions are implemented by an encoder such as a standard-compliant encoder.
[0077] The various components of the systems described herein can be implemented in software, firmware, and / or hardware and / or any combination thereof. For example, the various components of the systems or devices discussed herein can be provided at least in part by hardware such as a computing system on a chip (SoC) that can be present in a computing system such as a smart phone. Those skilled in the art will recognize that the systems described herein can include additional components not depicted in the corresponding figures. For example, the systems discussed herein can include additional components such as a bitstream multiplexer or demultiplexer module, etc., which are not described for clarity.
[0078] Although the implementation of the exemplary processes discussed herein can include the performance of all the operations shown in the order shown, the present disclosure is not limited in this regard, and in various examples, the implementation of the exemplary processes herein can include only a subset of the shown operations, operations performed in an order different from that shown, or additional operations.
[0079] Additionally, any one or more operations discussed herein can be performed in response to instructions provided by one or more computer program products. Such program products can include a signal-bearing medium providing the instructions that, when executed by, for example, a processor, can provide the functionality described herein. The computer program products can be provided in any form of one or more machine-readable media. Thus, for example, in response to program code and / or instructions or instruction sets transmitted to a processor via one or more machine-readable media, a processor including one or more graphics processing units or processor cores can perform one or more blocks of the exemplary processes described herein. Generally, the machine-readable media can transmit software in the form of program code and / or instructions or instruction sets that can cause any device and / or system described herein to at least implement portions of the operations discussed herein or any portion of the devices, systems, or any modules or components discussed herein.
[0080] As used in any implementation described herein, the term "module" refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. The software can be embodied as software groupings, code, and / or instruction sets or instructions, and "hardware" as used in any implementation described herein can include, alone or in any combination, for example, hardwired circuitry, programmable circuitry, state machine circuitry, fixed function circuitry, execution unit circuitry, and / or firmware storing instructions executed by the programmable circuitry. The modules can be embodied, jointly or individually, as circuitry forming part of a larger system, for example, an integrated circuit (IC), a system on a chip (SoC), etc.
[0081] Figure 12 is a schematic diagram of an exemplary system 1200 arranged in accordance with at least some implementations of the present disclosure. In various implementations, system 1200 can be a mobile system, but system 1200 is not limited thereto. For example, system 1200 can be incorporated into a personal computer (PC), laptop computer, ultra-portable computer, tablet, touchpad, portable computer, handheld computer, palmtop computer, personal digital assistant (PDA), cellular phone, combination cellular phone / PDA, television, smart device (e.g., smart phone, smart tablet, or smart TV), mobile Internet device (MID), messaging device, data communication device, camera (e.g., point-and-shoot camera, super-zoom camera, digital single-lens reflex (DSLR) camera), etc.
[0082] In various implementations, system 1200 includes platform 1202 coupled to display 1220. Platform 1202 may receive content from content devices such as content service device 1230 or content distribution device 1240 or other similar content sources. Navigation controller 1250, which includes one or more navigation features, may be used to interact with, for example, platform 1202 and / or display 1220. Each of these components is described in more detail below.
[0083] In various implementations, platform 1202 may include any combination of chipset 1205, processor 1210, memory 1212, antenna 1213, storage 1214, graphics subsystem 1215, applications 1216, and / or radio 1218. Chipset 1205 may provide intercommunication between processor 1210, memory 1212, storage 1214, graphics subsystem 1215, applications 1216, and / or radio 1218. For example, chipset 1205 may include a storage adapter (not shown) capable of providing intercommunication with storage 1214.
[0084] Processor 1210 may be implemented as a complex instruction set computer (CISC) or reduced instruction set computer (RISC) processor, an x86 instruction set compliant processor, a multi-core, or any other microprocessor or central processing unit (CPU). In various implementations, processor 1210 may be a dual-core processor, a dual-core mobile processor, etc.
[0085] Memory 1212 may be implemented as a volatile storage device, such as but not limited to random access memory (RAM), dynamic random access memory (DRAM), or static RAM (SRAM).
[0086] Storage 1214 may be implemented as a non-volatile storage device, such as but not limited to a disk drive, an optical disk drive, a tape drive, an internal storage device, an attached storage device, a flash memory, SDRAM (synchronous DRAM) with a backup battery, and / or a network accessible storage device. In various implementations, for example, storage 1214 may include techniques that enhance protection of storage performance for valuable digital media when including multiple hard disk drives.
[0087] The graphics subsystem 1215 can perform image processing such as still images or videos for display. For example, the graphics subsystem 1215 can be a graphics processing unit (GPU) or a vision processing unit (VPU). An analog or digital interface can be used to communicatively couple the graphics subsystem 1215 and the display 1220. For example, the interface can be any one of high definition multimedia interface, DisplayPort, wireless HDMI, and / or wireless HD compliant technology. The graphics subsystem 1215 can be integrated into the processor 1210 or the chipset 1205. In some implementations, the graphics subsystem 1215 can be a stand-alone device communicatively coupled to the chipset 1205.
[0088] The graphics and / or video processing techniques described herein can be implemented in a variety of hardware architectures. For example, the graphics and / or video functions can be integrated within a chipset. Alternatively, discrete graphics and / or video processors can be used. As yet another implementation, the graphics and / or video functions can be provided by a general-purpose processor, including a multi-core processor. In a further embodiment, the functions can be implemented in a consumer electronic device.
[0089] The radio device 1218 can include one or more radio devices capable of transmitting and receiving signals using a variety of suitable wireless communication technologies. These technologies can involve communication across one or more wireless networks. Exemplary wireless networks include (but are not limited to) wireless local area network (WLAN), wireless personal area network (WPAN), wireless metropolitan area network (WMAN), cellular network, and satellite network. When communicating through these networks, the radio device 1218 can operate in accordance with one or more applicable standards of any version.
[0090] In various implementations, the display 1220 can include any television-type monitor or display. The display 1220 can include, for example, a computer display screen, a touch screen display, a video monitor, a television-like device, and / or a television. The display 1220 can be digital and / or analog. In various implementations, the display 1220 can be a holographic display. Moreover, the display 1220 can be a transparent surface that can receive a visual projection. Such a projection can convey various forms of information, images, and / or objects. For example, such a projection can be a visual overlay of a mobile augmented reality (MAR) application. Under the control of one or more software applications 1216, the platform 1202 can display a user interface 1222 on the display 1220.
[0091] In various implementations, the content service device 1230 can be hosted by any national, international, and / or independent service and can thus be accessible via the Internet access platform 1202, for example. The content service device 1230 can be coupled to the platform 1202 and / or the display 1220. The platform 1202 and / or the content service device 1230 can be coupled to the network 1260 to communicate media information (e.g., send and / or receive) with the network 1260. The content distribution device 1240 can also be coupled to the platform 1202 and / or the display 1220.
[0092] In various implementations, the content service device 1230 can include a cable TV box, a personal computer, a network, a telephone, an Internet device or apparatus capable of distributing digital information and / or content, and any other similar device capable of one-way or two-way content communication via the network 1260 or directly between the content provider and the platform 1202 and / or the display 1220. It should be understood that one-way and / or two-way content communication can be performed via the network 1260 with any one of the components in the system 1200 and the content provider. Examples of content can include any media information, including, for example, video, music, medical, and game information, etc.
[0093] The content service device 1230 can receive content such as cable TV programs, including media information, digital information, and / or other content. Examples of content providers can include any cable or satellite TV or radio device or Internet content provider. The provided examples are not meant to limit the implementations according to the present disclosure in any way.
[0094] In various implementations, the platform 1202 can receive control signals from a navigation controller 1250 having one or more navigation features. For example, the navigation features can be used to interact with the user interface 1222. In various embodiments, the navigation can be a pointing device, which can be a computer hardware component (specifically, a human-machine interface device) that allows a user to input spatial (e.g., continuous and multi-dimensional) data into a computer. Many systems such as graphical user interfaces (GUIs), TVs, and monitors allow users to use physical gestures to control and provide data to a computer or TV.
[0095] The movement of the navigation features can be replicated on a display (e.g., the display 1220) by the movement of a pointer, cursor, focus ring, or other visual indicator displayed on the display. For example, under the control of the software application 1216, the navigation features located on the navigation can be mapped to virtual navigation features displayed, for example, on the user interface 1222. In various embodiments, it can not be a separate component but can be integrated into the platform 1202 and / or the display 1220. However, the present disclosure is not limited to the elements or backgrounds shown or described herein.
[0096] In various implementations, for example, when enabled, a driver (not shown) can include techniques that enable a user to immediately turn on and off a TV-like platform 1202 via a touch button after an initial startup. Even when the platform is "off," program logic can allow the platform 1202 to stream content to a media adapter or other content service device 1230 or content distribution device 1240. Additionally, for example, the chipset 1205 can include hardware and / or software for supporting 5.1 surround sound audio and / or high-definition 7.1 surround sound audio. The driver can include a graphics driver for an integrated graphics platform. In various embodiments, the graphics driver can include a Peripheral Component Interconnect (PCI) Express graphics card.
[0097] In various implementations, any one or more of the components shown in system 1200 can be integrated. For example, platform 1202 and content service device 1230 can be integrated, or platform 1202 and content distribution device 1240 can be integrated, or for example, platform 1202, content service device 1230, and content distribution device 1240 can be integrated. In various embodiments, platform 1202 and display 1220 can be an integrated unit. For example, display 1220 and content service device 1230 can be integrated, or display 1220 and content distribution device 1240 can be integrated. These embodiments are not meant to limit the present disclosure.
[0098] In various embodiments, system 1200 can be implemented as a wireless system, a wired system, or a combination of both. When implemented as a wireless system, system 1200 can include components and interfaces suitable for communicating via a wireless shared medium, such as one or more antennas, transmitters, receivers, transceivers, amplifiers, filters, control logic, etc. Examples of wireless shared media can include portions of the wireless spectrum, such as the RF spectrum, etc. When implemented as a wired system, system 1200 can include components and interfaces suitable for communicating via a wired communication medium, such as an input / output (I / O) adapter, a physical connector for connecting the I / O adapter to the corresponding wired communication medium, a network interface card (NIC), an optical disc controller, a video controller, an audio controller, etc. Examples of wired communication media can include leads, cable wires, printed circuit boards (PCBs), backplanes, switched light rays, semiconductor materials, twisted pairs, coaxial cables, optical fibers, etc.
[0099] Platform 1202 may establish one or more logical or physical channels for information communication. The information may include media information and control information. Media information may refer to any data representing content for a user. Examples of content may include, for example, data from a voice conversation, a video conference, a streaming video, an email ("mail") message, a voicemail message, alphanumeric symbols, graphics, images, video, text, etc. Data from a voice conversation may be, for example, voice information, silence periods, background noise, comfort noise, tones, etc. Control information may refer to any data representing commands, instructions, or control words for an automated system. For example, control information may be used to route media information through the system or to instruct a node to process media information in a predetermined manner. However, the embodiments are not limited to Figure 12 the elements or scenarios shown or described herein.
[0100] As described above, system 1200 may be embodied in different physical forms or form factors. Figure 13 Exemplary small form factor device 1300 arranged according to at least some implementations of the present disclosure is shown. In some examples, system 1200 may be implemented via device 1300. In other examples, system 100 or a portion thereof may be implemented via device 1300. In various embodiments, for example, device 1000 may be implemented as a mobile computing device having wireless capabilities. For example, a mobile computing device may refer to any device having a processing system and a mobile power source or power supply (such as one or more batteries).
[0101] Examples of mobile computing devices may include a personal computer (PC), a laptop computer, an ultra-portable computer, a tablet computer, a touchpad, a portable computer, a handheld computer, a palmtop computer, a personal digital assistant (PDA), a cellular phone, a combination cellular phone / PDA, a smart device (e.g., a smart phone, a smart tablet, or a smart mobile TV), a mobile Internet device (MID), a messaging device, a data communication device, a camera, etc.
[0102] Examples of mobile computing devices may also include computers arranged to be worn by a person, such as a wrist computer, a finger computer, a ring computer, a glasses computer, a clip-on computer, an armband computer, a shoe computer, a clothing computer, and other wearable computers. In various embodiments, for example, a mobile computing device may be implemented as a smart phone capable of executing computer applications as well as voice communication and / or data communication. Although some embodiments may be described by way of example using a mobile computing device implemented as a smart phone, it can be understood that other embodiments may also be implemented using other wireless mobile computing devices. The embodiments are not limited in this regard.
[0103] As Figure 13As shown, device 1300 may include a housing having a front portion 1301 and a rear portion 1302. Device 1300 includes a display 1304, an input / output (I / O) device 1306, and an integrated antenna 1308. Device 1300 may also include a navigation feature 1312. The I / O device 1306 may include any suitable I / O device for inputting information into the mobile computing device. Examples of the I / O device 1306 may include an alphanumeric keyboard, a numeric keypad, a touchpad, input keys, buttons, switches, a microphone, a speaker, a voice recognition device, and software, etc. Information may also be input into device 1300 through a microphone (not shown), or may be digitized by a voice recognition device. As shown, device 1300 may include a camera 1305 (e.g., including a lens, an aperture, and an imaging sensor) and a flash 1310 integrated into the rear portion 1302 (or elsewhere) of device 1300. In other examples, the camera 1305 and the flash 1310 may be integrated into the front portion 1301 of device 1300, or front and rear cameras may be provided. The camera 1305 and the flash 1310 may be components of a camera module to initiate processing of image data into streaming video, which is output to the display 1304 and / or remotely transmitted from device 1300 via, for example, the antenna 1308.
[0104] Various embodiments may be implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements may include a processor, a microprocessor, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chip sets, etc. Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. Determining whether to implement an embodiment using hardware elements and / or software elements may vary according to any number of factors, such as desired computing speed, power level, heat tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance constraints.
[0105] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium that represents various logic within a processor. The instructions, when read by a machine, cause the machine to fabricate the logic to implement the techniques described herein. Such representations, referred to as “IP cores,” may be stored on a tangible machine-readable medium and provided to various customers or manufacturing facilities to load into the manufacturing machines that actually fabricate the logic or processor.
[0106] Although certain features described herein have been referenced with respect to various implementations, the description is not intended to be construed in a limiting sense. Accordingly, various modifications of the implementations described herein as well as other implementations will be apparent to those of ordinary skill in the art and are considered to be within the spirit and scope of the present disclosure.
[0107] The following embodiments pertain to further embodiments.
[0108] In one or more first embodiments, a computer-implemented method for video encoding includes: receiving an input video to be encoded, the input video including a plurality of images; evaluating a plurality of intra modes for a candidate partition of a single block of a first image among the plurality of images by comparing the candidate partition of the single block of the first image with an intra prediction partition generated using only original pixel samples from the first image, and evaluating a plurality of inter modes for the candidate partition by comparing the candidate partition with a plurality of search partitions including only original pixel samples from a second image among the plurality of images, thereby generating a final partition decision for the single block; providing the final partition decision to a video encoder; at the video encoder, taking a difference between the single block and a reconstructed block generated using the final partition decision to generate a residual block corresponding to the single block; and at the video encoder, encoding at least transform coefficients of the residual block into an output bitstream.
[0109] In one or more second embodiments, for any of the first embodiments, the single block is part of a region of the first image, and the method further includes: performing flat region detection and noisy region detection on the region to determine whether the region is a flat and noisy region; performing edge detection to determine whether the region includes detected edges; and evaluating a temporal variance of changes of the region over the plurality of images including the first image to determine whether the region has a low temporal variance of changes, wherein evaluating the intra mode and the inter mode includes applying a bias to one or both of the intra mode result and the inter mode result in response to the region being a flat and noisy region, the region including edges, and the region having a low temporal variance of changes to favor selection of the inter mode result for one or more partitions of the single block.
[0110] In one or more third embodiments, for any one of the first and second embodiments, applying a bias includes one of adding a first factor to an intra-mode result or an inter-mode result and multiplying a second factor by the intra-mode result or the inter-mode result.
[0111] In one or more fourth embodiments, for any one of the first to third embodiments, the intra-mode and inter-mode results include a rate-distortion cost.
[0112] In one or more fifth embodiments, for any one of the first to fourth embodiments, performing flat region detection includes denoising a region and comparing a change in the denoised region with a first threshold, and performing noisy region detection includes taking a difference between the region and the denoised region and comparing a variance of the difference region with a second threshold.
[0113] In one or more sixth embodiments, for any one of the first to fifth embodiments, generating a final partition decision further includes generating an initial coding mode decision for partitioning a single block corresponding to the final partition decision based on an evaluation of a plurality of intra-modes and a plurality of inter-modes.
[0114] In one or more seventh embodiments, for any one of the first to sixth embodiments, a first initial coding mode decision for a first partition of a single block includes a first intra-mode among a plurality of intra-modes, and the method further includes: at a video encoder, applying a transform to a first residual partition corresponding to the first partition and the first intra-mode to generate first transformed residual coefficients; at the video encoder, applying a transform to a second residual partition corresponding to the first partition and a second intra-mode to generate second transformed residual coefficients; determining a first high-frequency coefficient energy value of the first transformed residual coefficients and a second high-frequency coefficient energy value of the second transformed residual coefficients; discarding the first intra-mode in response to the second high-frequency value being less than the first high-frequency value; and at the video encoder, encoding the first partition using the second intra-mode.
[0115] In one or more eighth embodiments, for any one of the first to seventh embodiments, determining the first high-frequency coefficient energy value includes averaging a plurality of first transformed residual coefficients other than a DC transform coefficient of the first transformed residual coefficients.
[0116] In one or more ninth embodiments, for any one of the first to eighth embodiments, a video encoder uses the final partition decision and the initial coding mode decision to generate a reconstructed block.
[0117] In one or more tenth embodiments, for any of the first to ninth embodiments, comparing a candidate partition with an intra prediction partition generated only using original pixel samples from a first image includes: generating a first intra prediction partition only using the original pixel samples according to a current intra mode evaluated for a first candidate partition, and taking a difference between the first candidate partition and the first intra prediction partition.
[0118] In one or more eleventh embodiments, a system for video coding includes: a memory for storing an input video to be coded, the input video including a plurality of images; one or more first processors; and a second processor coupled to the one or more first processors and implementing a video encoder, the one or more first processors being configured to: generate a final partition decision for a single block based on an evaluation of a plurality of intra modes for a candidate partition of a single block of a first image among the plurality of images and an evaluation of a plurality of inter modes for the candidate partition using a second image among the plurality of images, wherein the one or more first processors use a comparison between the candidate partition and an intra prediction partition generated only using original pixel samples from the first image to evaluate the plurality of intra modes, and the one or more first processors use a comparison between the candidate partition and a plurality of search partitions including only original pixel samples from the second image among the plurality of images to evaluate the plurality of inter modes; and provide the final partition decision to the second processor, the second processor being configured to: take a difference between the single block and a reconstructed block generated using the final partition decision at the video encoder to generate a residual block corresponding to the single block; and at the video encoder, encode at least transform coefficients of the residual block as an output bitstream.
[0119] In one or more twelfth embodiments, for any of the eleventh embodiments, the single block is part of a region of the first image, and the one or more first processors are further configured to: perform flat region detection and noisy region detection on the region to determine whether the region is a flat and noisy region; perform edge detection to determine whether the region includes detected edges; and evaluate a temporal variance of changes of the region over a plurality of images including the first image to determine whether the region has a low temporal variance of changes, wherein the one or more first processors generating an initial coding mode decision for the single block includes: the one or more first processors applying a bias to one or both of an intra mode result and an inter mode result in response to the region being a flat and noisy region, the region including edges, and the region having a low temporal variance of changes, to favor selecting the inter mode result for one or more partitions of the single block.
[0120] In one or more thirteenth embodiments, for any one of the eleventh and twelfth embodiments, the one or more first processors applying a bias includes: the one or more first processors adding a first factor to an intra-mode result or an inter-mode result or multiplying a second factor by an intra-mode result or an inter-mode result.
[0121] In one or more fourteenth embodiments, for any one of the eleventh to thirteenth embodiments, the intra and inter mode results include rate-distortion cost.
[0122] In one or more fifteenth embodiments, for any one of the eleventh to fourteenth embodiments, the one or more first processors performing flat region detection includes: the one or more first processors denoising a region and comparing a change of the denoised region with a first threshold, and the one or more first processors performing noisy region detection includes: the one or more first processors taking a difference between a region and the denoised region and comparing a variance of the difference region with a second threshold.
[0123] In one or more sixteenth embodiments, for any one of the eleventh to fifteenth embodiments, the one or more first processors generating a final partition decision includes: the one or more first processors generating an initial coding mode decision for a partition of a single block corresponding to the final partition decision based on an evaluation of a plurality of intra-modes and a plurality of inter-modes.
[0124] In one or more seventeenth embodiments, for any one of the eleventh to sixteenth embodiments, the first initial coding mode decision for a first partition of a single block includes a first intra-mode among a plurality of intra-modes, and the second processor is further configured to: at a video encoder, apply a transform to a first residual partition corresponding to the first partition and the first intra-mode to generate a first transformed residual; at the video encoder, apply a transform to a second residual partition corresponding to the first partition and a second intra-mode to generate a second transformed residual; determine a first high-frequency energy value of the first transformed residual and a second high-frequency energy value of the second transformed residual; discard the first intra-mode in response to the second high-frequency value being less than the first high-frequency value; and at the video encoder, encode the first partition using the second intra-mode.
[0125] In one or more eighteenth embodiments, for any one of the eleventh to seventeenth embodiments, the second processor determining the first high-frequency coefficient energy value includes: the second processor averaging a plurality of first transformed residual coefficients other than a DC transform coefficient of the first transformed residual coefficient.
[0126] In one or more nineteenth embodiments, for any one of the eleventh to eighteenth embodiments, the video encoder uses the final partition decision and the initial coding mode decision to generate a reconstructed block.
[0127] In one or more twentieth embodiments, for any one of the eleventh to nineteenth embodiments, comparing a candidate partition with an intra prediction partition generated only using raw pixel samples from a first image includes: the one or more first processors generating a first intra prediction partition only using the raw pixel samples and taking the difference between the first candidate partition and the first intra prediction partition.
[0128] In one or more twenty - first embodiments, at least one machine - readable medium may include a plurality of instructions that, in response to execution on a computing device, cause the computing device to perform a method according to any one of the above - described embodiments.
[0129] In one or more twenty - second embodiments, an apparatus may include a module for performing a method according to any one of the above - described embodiments.
[0130] It will be recognized that the embodiments are not limited to the embodiments so described, but may be practiced with modifications and variations without departing from the scope of the appended claims. For example, the above embodiments may include a particular combination of features. However, the above embodiments are not limited in this regard, and in various implementations, the above embodiments may include only a subset of these features, a different order of implementing these features, a different combination of implementing these features, and / or implementing other features in addition to the explicitly listed features. Therefore, the scope of the embodiments should be determined with reference to the appended claims and the full scope of equivalents of these claims.
Claims
1. A computer-implemented method for video encoding, comprising: receiving an input video to be encoded, the input video including a plurality of images; generating a final partition and an encoding mode decision for a single block of a first image among the plurality of images through the following steps: evaluating a plurality of intra modes for the candidate partition by comparing the candidate partition of the single block with an intra prediction partition generated only using original pixel samples from the first image; evaluating a plurality of inter modes for the candidate partition by comparing the candidate partition with a plurality of search partitions including only original pixel samples from a second image among the plurality of images; and determining a final partition decision for the single block and an encoding mode for each partition of the final partition based on the evaluation of the intra mode and the inter mode, wherein a first encoding mode of a first partition of the single block is a first intra encoding mode among the plurality of intra modes; providing the final partition decision to a video encoder; applying a transform to a first residual partition including a difference between the first partition and a first reconstructed partition generated using the first intra mode to generate first transform residual coefficients, and applying a transform to a second residual partition including a difference between the first partition and a second reconstructed partition generated using the second intra mode to generate second transform residual coefficients; determining a first high-frequency coefficient energy value and a second high-frequency coefficient energy value representing the first transform residual coefficients and the second transform residual coefficients respectively using only high-frequency transform residual coefficients among the first transform residual coefficients and the second transform residual coefficients; selecting a final intra mode for the first partition as one of the first intra mode and the second intra mode having a lower high-frequency coefficient energy value; at the video encoder, taking a difference between the single block and a reconstructed block generated using the final partition decision and the selected final intra mode for the first partition to generate a residual block corresponding to the single block; and at the video encoder, encoding at least transform coefficients of the residual block into an output bitstream.
2. The method according to claim 1, wherein the single block is a part of an area of the first image, and the method further includes: performing flat area detection and noisy area detection on the area to determine whether the area is a flat and noisy area; performing edge detection to determine whether the area includes detected edges; and evaluating a temporal variance of changes of the area on a plurality of images including the first image to determine whether the area has a low temporal variance of changes, wherein evaluating the intra mode and the inter mode includes: in response to the area being a flat and noisy area, the area including edges, and the area having a low temporal variance of changes, applying a bias to one or both of an intra mode result and an inter mode result to favor selecting an inter mode result for one or more partitions of the single block.
3. The method according to claim 2, wherein applying the bias includes: One of adding a first factor to the intra-mode result or the inter-mode result and multiplying a second factor by the intra-mode result or the inter-mode result.
4. The method according to claim 2, wherein, the intra-mode result and the inter-mode result include rate-distortion cost.
5. The method according to claim 2, wherein, performing flat region detection includes: denoising the region, and comparing the change of the denoised region with a first threshold, and performing noisy region detection includes: taking the difference between the region and the denoised region and comparing the variance of the difference region with a second threshold.
6. The method according to claim 1, wherein, the transform includes Hadamard transform.
7. The method according to claim 1, wherein, the first transform residual coefficient includes 16x16 first transform residual coefficients, and the first high-frequency coefficient energy value includes the average value of the 4x4 transform residual coefficients in the lower right group from the 16x16 first transform residual coefficients.
8. The method according to claim 7, wherein, the first high-frequency coefficient energy value includes the average value of a plurality of first transform residual coefficients except the DC transform coefficient in the first transform residual coefficients.
9. The method according to claim 1, wherein, the video encoder uses the final partition decision and the initial coding mode decision to generate the reconstructed block.
10. The method according to any one of claims 1-9, wherein, comparing the candidate partition with the intra-prediction partition generated only using the original pixel samples from the first image includes: generating a first intra-prediction partition only using the original pixel samples according to the current intra-mode evaluated for the first candidate partition; and taking the difference between the first candidate partition and the first intra-prediction partition.
11. A system for video coding, comprising: a memory for storing an input video to be encoded, the input video including a plurality of images; one or more first processors; and a second processor, coupled to the one or more first processors and implementing a video encoder, the one or more first processors being configured to: Generate a final partition and an encoding mode decision for the single block based on an evaluation of multiple intra modes for a candidate partition of a single block of the first image among the multiple images and an evaluation of multiple inter modes for the candidate partition using a second image among the multiple images, wherein the one or more first processors: evaluate the multiple intra modes using a comparison of the candidate partition with an intra prediction partition generated using only raw pixel samples from the first image; evaluate the multiple inter modes using a comparison of the candidate partition with multiple search partitions each including only raw pixel samples from the second image among the multiple images; determine, based on the evaluation of the intra modes and the inter modes, a final partition decision for the single block and an encoding mode for each partition of the final partition, wherein a first encoding mode of a first partition of the single block is a first intra encoding mode among the multiple intra modes; and Provide the final partition decision to the second processor, The second processor is configured to: Apply a transform to a first residual partition including a difference between the first partition and a first reconstructed partition generated using the first intra mode to generate first transform residual coefficients, and apply a transform to a second residual partition including a difference between the first partition and a second reconstructed partition generated using the second intra mode to generate second transform residual coefficients; Determine a first high-frequency coefficient energy value representing the first transform residual coefficients and a second high-frequency coefficient energy value representing the second transform residual coefficients, respectively using only high-frequency transform residual coefficients among the first transform residual coefficients and the second transform residual coefficients; Select a final intra mode for the first partition as the one of the first intra mode and the second intra mode having a lower high-frequency coefficient energy value; At the video encoder, take a difference between the single block and a reconstructed block generated using the final partition decision and the selected final intra mode for the first partition to generate a residual block corresponding to the single block; and At the video encoder, encode at least transform coefficients of the residual block as an output bitstream.
12. The system according to claim 11, wherein, The single block is a part of a region of the first image, and the one or more first processors are further configured to: Perform flat region detection and noisy region detection on the region to determine whether the region is a flat and noisy region; Perform edge detection to determine whether the region includes detected edges; and Evaluate a temporal variance of changes of the region over multiple images including the first image to determine whether the region has a low temporal variance of changes, wherein the one or more first processors generating the initial encoding mode decision for the single block includes: The one or more first processors apply a bias to one or both of the intra-mode result and the inter-mode result in response to the region being a flat and noisy region, the region including edges, and the region having a low temporal variance of change, to favor selection of the inter-mode result for one or more partitions of the single block.
13. The system according to claim 12, wherein, the one or more first processors applying the bias includes: the one or more first processors adding a first factor to the intra-mode result or the inter-mode result, or multiplying a second factor by the intra-mode result or the inter-mode result.
14. The system according to claim 13, wherein, the intra-mode result and the inter-mode result include rate-distortion costs.
15. The system according to claim 13, wherein, the one or more first processors performing flat region detection includes: the one or more first processors denoising the region and comparing the change of the denoised region with a first threshold, and the one or more first processors performing noisy region detection includes: the one or more first processors taking the difference between the region and the denoised region and comparing the variance of the difference region with a second threshold.
16. The system according to claim 11, wherein, the transform includes a Hadamard transform.
17. The system according to claim 11, wherein, the first transform residual coefficients include 16x16 first transform residual coefficients, and the first high-frequency coefficient energy value includes the average of the lower right 4x4 transform residual coefficients from the 16x16 first transform residual coefficients.
18. The system according to claim 11, wherein, the first high-frequency coefficient energy value includes the average of a plurality of first transform residual coefficients other than the DC transform coefficient in the first transform residual coefficients.
19. The system according to claim 11, wherein, the reconstructed block is generated by the video encoder using a final partition decision and an initial coding mode decision.
20. The system according to any one of claims 11-19, wherein, comparing the candidate partition with an intra-predicted partition generated using only the original pixel samples from the first image includes: the one or more first processors: generating a first intra-predicted partition using only the original pixel samples according to the current intra-mode evaluated for the first candidate partition; and taking the difference between the first candidate partition and the first intra-predicted partition.
21. At least one machine-readable medium, including a plurality of instructions that, when executed on a computing device, cause the computing device to perform the method according to any one of claims 1-10.
22. An apparatus, including: a module for performing the method according to any one of claims 1-10.
Citation Information
Patent Citations
Image encoding device and encoding method, and image decoding device and decoding method
US20090110070A1
Moving image encoding device
WO2016152446A1