Apparatus and method for encoding video data
By generating and selecting multiple intra-frame candidate modes, and optimizing video coding using template prediction and gradient filters, the problem of low coding efficiency in TIMD is solved, achieving more efficient block cell prediction and reconstruction.
Patent Information
- Application Number
- CN202210759282.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-29
- Filing Date
- 2022-06-29
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2042-06-29
AI Technical Summary
Existing template-based intra-frame pattern derivation (TIMD) is not efficient enough in video coding and is difficult to accurately predict or reconstruct target blocks.
By receiving the bit stream, block units are determined and multiple intra-frame candidate modes are generated. Template angle and magnitude are generated using template prediction and gradient filters. The optimal prediction mode is selected and the block unit is reconstructed by combining weighted parameters.
It improves the coding efficiency of video coding, enables more accurate block cell prediction and reconstruction, and enhances video quality.
Smart Images

Figure CN115550643B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This disclosure claims the benefit and priority of Provisional U.S. Patent Application No. 63 / 216191, filed June 29, 2021, entitled “PROPOSED BLENDING INDEXDERIVATION OF DECODER SIDE INTRA PREDICTION MODE DERIVATION” (hereinafter referred to as “'191 Provisional Case”). The disclosure of '191 Provisional Case is hereby incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates generally to video coding, and more particularly to techniques for using template prediction in template-based intra-mode derivation (TIMD). Background Technology
[0004] Template-based intra-frame mode derivation (TIMD) is an encoding tool used for video coding. In conventional video coding methods, the encoder and decoder can use previously reconstructed samples adjacent to the target block to generate one of several intra-frame default modes for predicting the target block.
[0005] However, TIMD prediction for the target block is based solely on a prediction pattern selected using template prediction, which may result in insufficient coding efficiency when TIMD is used to predict the target block. Therefore, the encoder and decoder may require new TIMD to predict or reconstruct the target block more accurately. Summary of the Invention
[0006] This disclosure relates to an apparatus and method for predicting block cells in an image frame using template prediction in TIMD.
[0007] In a first aspect of this disclosure, a method for decoding a bitstream and an electronic device for performing the method are provided. The method includes: receiving the bitstream; determining block units from image frames based on the bitstream; selecting a plurality of intra-frame default modes from a plurality of intra-frame default modes for the block units; generating a template prediction for each of the plurality of intra-frame candidate modes; selecting a plurality of prediction modes from the plurality of intra-frame candidate modes based on the template prediction; and reconstructing the block units based on the plurality of prediction modes.
[0008] The first aspect of the implementation further includes: determining a plurality of template blocks adjacent to the block unit; determining a plurality of template references adjacent to the plurality of template blocks; and predicting the plurality of template blocks based on the plurality of template references by using the plurality of intra-frame candidate modes to generate the template prediction.
[0009] The first aspect of the implementation further includes: determining a plurality of generation values by comparing the plurality of template blocks with each of the template predictions; and selecting the plurality of prediction patterns based on the plurality of generation values.
[0010] In another embodiment of the first aspect, each of the plurality of cost values determined by the cost function corresponds to one of the template predictions.
[0011] In another embodiment of the first aspect, the plurality of template blocks include a top adjacent block and a left adjacent block, and each is adjacent to the block unit.
[0012] The first aspect of the implementation further includes: predicting the block unit based on the plurality of prediction modes to generate a plurality of prediction blocks, each prediction block corresponding to one of the plurality of prediction modes; weightedly combining the plurality of prediction blocks to generate a prediction block having a plurality of weighting parameters; and reconstructing the block unit based on the prediction blocks.
[0013] In another embodiment of the first aspect, the plurality of prediction modes are selected based on a plurality of template blocks; and the plurality of weighting parameters are determined based on the plurality of template blocks.
[0014] In another embodiment of the first aspect, the plurality of template blocks are predicted to generate template predictions for selecting the plurality of prediction modes based on the template predictions of the plurality of template blocks; the plurality of weighting parameters are determined based on a plurality of cost values; and the plurality of cost values are determined by comparing the plurality of template blocks with each template prediction in the template predictions, respectively.
[0015] In another embodiment of the first aspect, the plurality of intra-frame candidate modes are a plurality of most probable modes (MPMs) selected from the plurality of intra-frame default modes.
[0016] The first aspect of the implementation further includes: determining a plurality of template regions adjacent to the block unit; filtering the plurality of template regions using a gradient filter to generate a plurality of template angles and a plurality of template amplitudes, wherein each of the plurality of template angles corresponds to one of the plurality of template amplitudes; and generating a gradient histogram (HoG) based on the plurality of template angles and the plurality of template amplitudes to select the plurality of intra-frame candidate modes.
[0017] The first aspect of the implementation further includes: mapping each of the plurality of template angles to one of the plurality of intra-default modes based on a predefined relationship to generate at least one mapping mode; and generating the HoG by accumulating the plurality of template amplitudes based on the at least one mapping mode, wherein the plurality of intra-candidate modes are selected from the plurality of intra-default modes based on the accumulated amplitudes in the HoG.
[0018] In a second aspect of this disclosure, a method for decoding a bitstream and an electronic device for performing the method are provided. The method includes: receiving the bitstream; determining block units and a plurality of adjacent regions adjacent to the block units from an image frame based on the bitstream; selecting a plurality of intra-frame candidate modes from a plurality of intra-frame default modes based on the adjacent regions; generating a template prediction for each of the plurality of intra-frame candidate modes; selecting a plurality of prediction modes from the plurality of intra-frame candidate modes based on the template prediction; and reconstructing the block units based on the plurality of prediction modes.
[0019] The second aspect of the implementation further includes: determining a plurality of template blocks adjacent to the block unit; determining a plurality of template references adjacent to the plurality of template blocks; and predicting the plurality of template blocks based on the plurality of template references by using the plurality of intra-frame candidate modes to generate the template prediction.
[0020] The second aspect of the implementation further includes: determining a plurality of generation values by comparing the plurality of template blocks with each of the template predictions; and selecting the plurality of prediction patterns based on the plurality of generation values.
[0021] In another embodiment of the second aspect, each of the plurality of cost values determined by the cost function corresponds to one of the template predictions.
[0022] In another embodiment of the second aspect, the plurality of template blocks include a top adjacent block and a left adjacent block, and each of them is adjacent to the block unit.
[0023] The second aspect of the implementation further includes: predicting the block unit based on the plurality of prediction modes to generate a plurality of prediction blocks, each prediction block corresponding to one of the plurality of prediction modes; weightedly combining the plurality of prediction blocks to generate a prediction block with a plurality of weighting parameters; and reconstructing the block unit based on the prediction blocks.
[0024] In another embodiment of the second aspect, the plurality of prediction modes are selected based on a plurality of template blocks; and the plurality of weighting parameters are determined based on the plurality of template blocks.
[0025] In another embodiment of the second aspect, the plurality of template blocks are predicted to generate template predictions for selecting the plurality of prediction modes based on the template predictions of the plurality of template blocks; the plurality of weighting parameters are determined based on a plurality of cost values; and the plurality of cost values are determined by comparing the plurality of template blocks with each template prediction in the template predictions, respectively.
[0026] In another embodiment of the second aspect, the plurality of adjacent regions are a plurality of reconstructed blocks adjacent to the block unit; the plurality of reconstructed blocks are reconstructed based on at least one reconstruction mode before the block unit is reconstructed; and the plurality of intra-frame candidate modes are a plurality of most probable modes (MPMs) selected from the plurality of intra-frame default modes based on the at least one reconstruction mode.
[0027] The second aspect of the implementation further includes: filtering the plurality of adjacent regions using a gradient filter to generate a plurality of template angles and a plurality of template amplitudes, wherein each of the plurality of template angles corresponds to one of the plurality of template amplitudes; and generating a gradient histogram (HoG) based on the plurality of template angles and the plurality of template amplitudes for selecting the plurality of intra-frame candidate modes.
[0028] The second aspect of the implementation further includes: mapping each of the plurality of template angles to one of the plurality of intra-default modes based on a predefined relationship to generate at least one mapping mode; and generating the HoG by accumulating the plurality of template amplitudes based on the at least one mapping mode, wherein the plurality of intra-candidate modes are selected from the plurality of intra-default modes based on the accumulated amplitudes in the HoG. Attached Figure Description
[0029] The various aspects of this disclosure can be best understood from the following detailed disclosure and corresponding drawings. The different features are not drawn to scale, and for clarity of discussion, the sizes of the various features may be arbitrarily increased or decreased.
[0030] Figure 1 A block diagram of a system configured to encode and decode video data according to an embodiment of the present disclosure is shown.
[0031] Figure 2 The embodiments according to this disclosure are shown in Figure 1 The block diagram of the decoder module of the second electronic device is shown in the figure.
[0032] Figure 3 A flowchart is shown of a method for decoding video data via an electronic device according to an embodiment of the present disclosure.
[0033] Figure 4A and Figure 4B This is a schematic diagram of an exemplary implementation of an adjacent region of a block unit.
[0034] Figure 5A and Figure 5B This is a schematic diagram of an exemplary implementation of a block unit, multiple template blocks, and a reference area.
[0035] Figure 6 The embodiments according to this disclosure are shown in Figure 1 The diagram shows a block diagram of the encoder module of the first electronic device. Detailed Implementation
[0036] The following disclosure includes specific information relating to embodiments described herein. The accompanying drawings and corresponding detailed disclosure are directed to exemplary embodiments. However, this disclosure is not limited to these exemplary embodiments. Other variations and embodiments of this disclosure will occur to those skilled in the art.
[0037] Unless otherwise specified, the same or corresponding elements in the accompanying drawings may be indicated by the same or corresponding reference numerals. The drawings and descriptions are generally not drawn to scale and are not intended to correspond to actual relative dimensions.
[0038] For the purposes of consistency and ease of understanding, similar features are identified by reference numerals in the exemplary drawings (but are not shown in some examples). However, features in different embodiments may differ in other respects and should not be narrowly limited to what is shown in the drawings.
[0039] The phrases “in one embodiment” or “in some embodiments” as used in this disclosure may each refer to one or more of the same or different embodiments. The term “coupled” is defined as a connection, whether direct or indirect through intermediate components, and is not necessarily limited to a physical connection. The term “comprising” means “including but not limited to”; it specifically indicates open inclusion or membership in combinations, groups, series, and equivalents so described.
[0040] For purposes of explanation and non-restriction, specific details such as functional entities, technologies, protocols, and standards are described to provide an understanding of the disclosed technologies. Detailed disclosures of well-known methods, technologies, systems, and architectures have been omitted to avoid making the disclosure unclear due to unnecessary details.
[0041] Those skilled in the art will recognize that any coded functions or algorithms described in this disclosure can be implemented by hardware, software, or a combination of both. The described functions may correspond to modules, which are software, hardware, firmware, or any combination thereof.
[0042] Software implementations may include programs having computer-executable instructions stored on a computer-readable medium such as memory or other types of storage devices. For example, one or more microprocessors or general-purpose computers with communication processing capabilities may be programmed using the executable instructions to perform the described functions or algorithms.
[0043] These microprocessors or general-purpose computers may be formed from application-specific integrated circuits (ASICs), programmable logic arrays, and / or using one or more digital signal processors (DSPs). While some of the disclosed embodiments are directed to software installed and executed on computer hardware, alternative embodiments, implemented as firmware or hardware or a combination of hardware and software, are also fully within the scope of this disclosure. Computer-readable media include, but are not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, compact disc read-only memory (CD-ROM), magnetic tape, magnetic tape, disk storage, or any other equivalent medium capable of storing computer-readable instructions.
[0044] Figure 1 A block diagram of a system 100 configured to encode and decode video data according to an embodiment of the present disclosure is shown. System 100 includes a first electronic device 110, a second electronic device 120, and a communication medium 130.
[0045] The first electronic device 110 may be a source device, including any device configured to encode video data and transmit the encoded video data to the communication medium 130. The second electronic device 120 may be a destination device, including any device configured to receive and decode the encoded video data via the communication medium 130.
[0046] The first electronic device 110 can communicate with the second electronic device 120 via a communication medium 130, either wired or wirelessly. The first electronic device 110 may include a source module 112, an encoder module 114, and a first interface 116. The second electronic device 120 may include a display module 122, a decoder module 124, and a second interface 126. The first electronic device 110 may be a video encoder, and the second electronic device 120 may be a video decoder.
[0047] The first electronic device 110 and / or the second electronic device 120 may be a mobile phone, tablet computer, desktop computer, laptop or other electronic device. Figure 1 An example of a first electronic device 110 and a second electronic device 120 is shown. The first electronic device 110 and the second electronic device 120 may include more or fewer components than shown, or different configurations of the components shown in the various illustrations.
[0048] Source module 112 may include a video capture device for capturing new video, a video archive for storing previously captured video, and / or a video feed interface for receiving video from a video content provider. Source module 112 may generate computer graphics-based data as source video, or generate a combination of real-time video, archived video, and computer-generated video as source video. The video capture device may be a charge-coupled device (CCD) image sensor, a complementary metal-oxide-semiconductor (CMOS) image sensor, or a camera.
[0049] Encoder module 114 and decoder module 124 can each be implemented as any of a variety of suitable encoder / decoder circuits, such as one or more microprocessors, central processing units (CPUs), graphics processing units (GPUs), systems-on-chips (SoCs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented in part in software, the device may store a program having instructions for software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in the hardware to perform the disclosed methods. Each of encoder module 114 and decoder module 124 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) in the device.
[0050] The first interface 116 and the second interface 126 may utilize custom protocols or comply with existing or de facto standards, including but not limited to Ethernet, IEEE 802.11 or IEEE 802.15 series, wireless USB, or telecommunications standards, including but not limited to Global System for Mobile Communication (GSM), Code Division Multiple Access 2000 (CDMA), Time Division Synchronous Code Division Multiple Access (TD-SCDMA), Worldwide Interoperability for Microwave Access (WiMAX), Third Generation Partnership Project Long-Term Evolution (3GPP-LTE), or Time-Division LTE (TD-LTE). The first interface 116 and the second interface 126 may each include any device configured to transmit and / or store compatible video bitstreams via communication medium 130 and receive compatible video bitstreams via communication medium 130.
[0051] The first interface 116 and the second interface 126 may include a computer system interface that enables compatible video bitstreams to be stored on or received from a storage device. For example, the first interface 116 and the second interface 126 may include a chipset supporting Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, proprietary bus protocols, Universal Serial Bus (USB) protocols, Inter-Integrated Circuit (I2C) protocols, or any other logical and physical architecture that can be used to interconnect peer devices.
[0052] Display module 122 may include a display using liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED), or light-emitting polymer display (LPD) technology, as well as other display technologies used in other embodiments. Display module 122 may include a high-definition display or an ultra-high-definition display.
[0053] Figure 2 The embodiments according to this disclosure are shown in Figure 1 The block diagram shown is of the decoder module 124 of the second electronic device 120. The decoder module 124 includes an entropy decoder (e.g., entropy decoding unit 2241), a prediction processor (e.g., prediction processing unit 2242), an inverse quantization / inverse transform processor (e.g., inverse quantization / inverse transform unit 2243), a summer (e.g., summer 2244), a filter (e.g., filter unit 2245), and a decoded image buffer (e.g., decoded image buffer 2246). The prediction processing unit 2242 further includes an intra-frame prediction processor (e.g., intra-frame prediction unit 22421) and an inter-frame prediction processor (e.g., inter-frame prediction unit 22422). The decoder module 124 receives a bitstream and decodes the bitstream to output a decoded video.
[0054] Entropy decoding unit 2241 can be derived from... Figure 1 The second interface 126 receives a bitstream including multiple syntax elements and performs a parsing operation on the bitstream to extract the syntax elements. As part of the parsing operation, the entropy decoding unit 2241 can entropy decode the bitstream to generate quantized transform coefficients, quantization parameters, transform data, motion vectors, intra-frame modes, segmentation information, and other syntax information.
[0055] The entropy decoding unit 2241 can perform Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Vinary Arithmetic Coding (CABAC), Syntax-based Context-adaptive Binary Arithmetic Coding (SBAC), Probability Interval Partitioning Entropy (PIPE) coding, or another entropy coding technique to generate quantized transform coefficients. The entropy decoding unit 2241 can provide the quantized transform coefficients, quantization parameters, and transform data to the inverse quantization / inverse transform unit 2243, and provide motion vectors, intra-frame modes, segmentation information, and other syntactic information to the prediction processing unit 2242.
[0056] The prediction processing unit 2242 may receive syntax elements, such as motion vectors, intra-frame modes, segmentation information, and other syntax information, from the entropy decoding unit 2241. The prediction processing unit 2242 may receive syntax elements including segmentation information and segment image frames according to the segmentation information.
[0057] Based on the segmentation information, each image frame can be divided into at least one image block. This at least one image block may include a luminance block for reconstructing multiple luminance samples and at least one chrominance block for reconstructing multiple chrominance samples. The luminance block and at least one chrominance block may be further subdivided to generate macroblocks, coding tree units (CTUs), coding blocks (CBs), their sub-segments, and / or another equivalent coding unit.
[0058] During the decoding process, the prediction processing unit 2242 may receive predicted data, which includes the intra-frame mode or motion vector of the current image block in a specific image frame. The current image block may be one of the luma blocks or chroma blocks in the specific image frame.
[0059] Intra-prediction unit 22421 can perform intra-predictive coding of the current block unit for one or more adjacent blocks in the same frame as the current block unit, based on syntax elements associated with the intra-frame mode, to generate a prediction block. The intra-frame mode can specify the position of reference samples selected from adjacent blocks within the current frame. When chroma components are reconstructed by prediction processing unit 2242, intra-prediction unit 22421 can reconstruct multiple chroma components of the current block unit based on multiple luma components of the current block unit.
[0060] When the luminance component of the current block is reconstructed by the prediction processing unit 2242, the intra-prediction unit 22421 can reconstruct multiple chrominance components of the current block unit based on multiple luminance components of the current block unit.
[0061] Inter-frame prediction unit 22422 can perform inter-frame prediction coding of the current block unit on one or more blocks in one or more reference image blocks based on syntax elements associated with motion vectors, in order to generate a prediction block. The motion vectors indicate the displacement of the current block unit within the current image block relative to a reference block unit within the reference image block. A reference block unit is a block determined to closely match the current block unit. Inter-frame prediction unit 22422 can receive reference image blocks stored in the decoded image buffer 2246 and reconstruct the current block unit based on the received reference image blocks.
[0062] The inverse quantization / inverse transform unit 2243 can apply inverse quantization and inverse transform to reconstruct the residual block in the pixel domain. The inverse quantization / inverse transform unit 2243 can apply inverse quantization to the residual quantized transform coefficients to generate residual transform coefficients, and then apply inverse transform to the residual transform coefficients to generate the residual block in the pixel domain.
[0063] The inverse transform can be applied in reverse through transform processes such as the discrete cosine transform (DCT), discrete sine transform (DST), adaptive multiple transform (AMT), mode-dependent non-separable secondary transform (MDNSST), hypercube-given transform (HyGT), signal-dependent transform, Karhunen-Loéve transform (KLT), wavelet transform, integer transform, subband transform, or conceptually similar transforms. The inverse transform can convert residual information from the transform domain (e.g., the frequency domain) back to the pixel domain. The degree of inverse quantization can be modified by adjusting the quantization parameters.
[0064] The summer 2244 adds the reconstructed residual block to the prediction block provided from the prediction processing unit 2242 to generate the reconstructed block.
[0065] Filtering unit 2245 may include a deblocking filter, a sample adaptive offset (SAO) filter, a bilateral filter, and / or an adaptive loop filter (ALF) to remove block artifacts from the reconstructed blocks. In addition to the deblocking filter, SAO filter, bilateral filter, and ALF, additional filters (in-loop or post-loop) may be used. For simplicity, these filters are not explicitly described, but they may filter the output of summer 2244. After filtering the reconstructed blocks of a specific image frame, filtering unit 2245 may output the decoded video to display module 122 or other video receiving units.
[0066] The decoded image buffer 2246 may be a reference image memory that stores reference blocks for the prediction processing unit 2242 to decode the bitstream (in inter-frame coding mode). The decoded image buffer 2246 may be formed from any of a variety of memory devices, such as dynamic random-access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded image buffer 2246 may be on-chip along with other components of the decoder module 124, or off-chip relative to those components.
[0067] Figure 3 A flowchart of a method 300 for decoding video data via an electronic device according to an embodiment of the present disclosure is shown. Method 300 is merely an example, as various ways exist to perform video data decoding.
[0068] Method 300 can be used in Figure 1 and Figure 2 The configuration shown in the figures is used to execute the method, and various elements in these figures are referenced with respect to method 300. Figure 3 Each box shown can represent one or more processes, methods, or subroutines being executed.
[0069] Figure 3 The order of the boxes in this document is illustrative only and may be changed. Additional boxes may be added or fewer boxes may be used without departing from this disclosure.
[0070] In box 310, decoder module 124 receives video data. The video data received by decoder module 124 may be a bitstream.
[0071] Reference Figure 1 and Figure 2The second electronic device 120 can receive bitstreams from an encoder or other video provider, such as the first electronic device 110, via a second interface 126. The second interface 126 can provide bitstreams to the decoder module 124.
[0072] Entropy decoding unit 2241 can decode the bitstream to determine multiple prediction indicators and multiple segmentation indicators for multiple image frames. Decoder module 124 can then further reconstruct the multiple image frames based on the prediction indicators and segmentation indicators. The prediction indicators and segmentation indicators may include multiple flags and multiple indices.
[0073] In box 320, decoder module 124 determines block units from image frames based on video data.
[0074] Reference Figure 1 and Figure 2 The decoder module 124 can determine image frames based on the bitstream and divide the image frames according to the segmentation instructions in the bitstream to determine block units. For example, the decoder module 124 can segment the image frames to generate multiple CTUs (coding tree units), and can further divide one of the CTUs to determine a block unit based on any video coding standard according to the segmentation instructions.
[0075] In box 330, decoder module 124 selects multiple intra-frame candidate modes from multiple intra-frame default modes used for block units.
[0076] Reference Figure 1 and Figure 2Decoder module 124 can determine an intra-default mode for predicting block units via intra-frame prediction. The intra-default mode may include multiple non-angular modes and multiple angular modes. Non-angular modes may include planar modes and DC modes. Furthermore, when decoder module 222 decodes block units in High Efficiency Video Coding (HEVC), the number of angular modes may be equal to 32 for method 300. When decoder module 124 decodes block units in Versatile Video Coding (VVC) or VVC Test Model (VTM), the number of angular modes may be equal to 65 for method 300. Furthermore, when decoder module 124 decodes block units in Enhanced Compression Model (ECM), the number of angular modes may be equal to 129 for method 300. Therefore, for method 300 in HEVC, the number of intra-frame default modes can be equal to 34; for method 300 in VVC or VTM, the number of intra-frame default modes can be equal to 67; and for method 300 in ECM, the number of intra-frame default modes can be equal to 130.
[0077] Figure 4A and Figure 4B This is a schematic diagram of an exemplary implementation of multiple adjacent regions of a block unit. Figure 4AThis is a schematic diagram of an exemplary implementation of block unit 4100 and multiple adjacent regions 4110 and 4120. Adjacent regions 4110-4120 can be multiple reconstructed blocks of adjacent block unit 4100. Before reconstructing block unit 4100, the reconstructed blocks can be reconstructed based on at least one reconstruction mode. Adjacent regions 4110 and 4120 adjacent to block unit 4100 are two different reconstructed blocks reconstructed before reconstructing block unit 4100. Decoder module 124 selects adjacent regions 4110 and 4120 based on multiple adjacent positions 4111 and 4121. Adjacent position 4111 can be located to the left of the lower left corner of block unit 4100, and adjacent position 4121 can be located above the upper right corner of block unit 4100. When a reconstructed block adjacent to block unit 4100 covers adjacent position 4111, the reconstructed block can be considered as adjacent region 4110. When a reconstructed block adjacent to block unit 4100 covers adjacent location 4121, the reconstructed block can be considered as adjacent region 4120. At least one reconstruction mode of adjacent regions 4110 and 4120 can be used to determine intra-frame candidate modes. When the reconstruction mode of adjacent region 4110 is the same as the reconstruction mode of adjacent region 4120, the number of at least one reconstruction mode of adjacent regions 4110 and 4120 is equal to 1. When the reconstruction mode of adjacent region 4110 is different from the reconstruction mode of adjacent region 4120, the number of at least one reconstruction mode of adjacent regions 4110 and 4120 is equal to 2. Intra-frame candidate modes can be multiple most probable modes (MPM) selected from the intra-frame default mode based on at least one reconstruction mode of adjacent regions 4110 and 4120. The MPM can be selected from the intra-frame default mode by using at least one reconstruction mode, based on any video coding standard (such as VVC, HEVC, and Advanced Video Coding (AVC) selection scheme) or any reference software of the video coding standard (such as VCM and ECM).
[0078] Figure 4BThis is a schematic diagram of an exemplary embodiment of block unit 4200 and a plurality of adjacent regions 4210 adjacent to block unit 4200. Decoder module 124 determines the adjacent regions 4210 adjacent to block unit 4200. The adjacent regions 4210 can be a plurality of adjacent regions adjacent to block unit 4200. The top adjacent region included in the adjacent regions may be located above block unit 4200, and the left adjacent region included in the adjacent regions may be located to the left of block unit 4200. In addition, there may be a top-left adjacent region located to the top left of the top left corner of block unit 4200. The adjacent regions 4210 may contain a plurality of reconstructed samples. The height of the top adjacent region may be equal to the number of reconstructed samples Nrt in the vertical direction, and the width of the top adjacent region may be equal to the width of block unit 4200. The height of the left adjacent region may be equal to the height of block unit 4200, and the width of the left adjacent region may be equal to the number of reconstructed samples Nrl in the horizontal direction. Furthermore, the height of the upper left adjacent region can be equal to the number of reconstructed samples Nrt along the vertical direction, and the width of the upper left adjacent region can be equal to the number of reconstructed samples Nrl along the horizontal direction. In one embodiment, the numbers Nrt and Nrl can be positive integers. Furthermore, the numbers Nrt and Nrl can be equal to each other. Further, the numbers Nrt and Nrl can be greater than or equal to 3.
[0079] All reconstructed samples in the neighboring region 4210 can be configured to be included in multiple template regions. Multiple template gradients are generated by filtering the template regions using a gradient filter. In other words, the neighboring regions can be filtered. In one implementation, the gradient filter can be a Soble filter. The template gradients are generated by filtering the reconstructed samples in the neighboring region 4210 based on the following filtering equation:
[0080] or
[0081] or Here, the operator * represents a two-dimensional signal processing convolution operation, and matrix A represents one of multiple filtered blocks 4211 in a neighboring region. In other words, each template gradient is generated based on one of the filtered blocks. Each filtered block includes Nf reconstructed samples. The number Nf can be a positive integer. For example, when the size of the filtered block is 3×3, the number Nf equals 9.
[0082] The template gradients of the filtered blocks can be further calculated to generate multiple template amplitudes and multiple template angles. Thus, the template region can be filtered using gradient filters to generate template angles and amplitudes. Each template amplitude can be generated by deriving the absolute value of the sum of the corresponding one in the template gradients. Furthermore, each template angle can be derived based on the partitioning results of the two fractional gradients Gx and Gy. The template amplitudes and template angles can be derived using the following equations:
[0083] Amp = abs(G x )+abs(G y )
[0084]
[0085] A predefined relationship between template angles and intra-frame default modes can be predefined in the first electronic device 110 and the second electronic device 120. For example, this relationship can be stored in the form of a look-up table (LUT), an equation, or a combination thereof. Thus, when a template angle is determined, the decoder module 124 can generate at least one mapping mode by mapping each of the multiple template angles to one of the multiple intra-frame default modes based on the predefined relationship. In other words, the at least one mapping mode can be generated by mapping each of the multiple template angles to multiple intra-frame default modes. For example, the number of at least one mapping mode can be equal to 1 when each template angle of block unit 4200 corresponds to the same intra-frame default mode. Conversely, the number of at least one mapping mode can be greater than 1 when some template angles of block unit 4200 correspond to different intra-frame default modes. In one embodiment, 360 degrees can be divided into multiple segments, and each segment represents an intra-frame prediction index. Therefore, if a template angle falls into a segment, the intra-frame prediction index corresponding to that segment can be derived according to the mapping rules.
[0086] The template gradient of a specific block within the filtered block can be calculated to generate a specific template amplitude in the template amplitude and a specific template angle in the template angle. Thus, a specific template amplitude can correspond to a specific template angle. In other words, each template angle in the filtered block can correspond to a corresponding template amplitude. Therefore, when at least one mapping mode is determined, the decoder module 124 can generate a histogram of gradients (HoG) by accumulating template amplitudes based on at least one mapping mode. For example, when two different template angles correspond to the same intra-frame default mode, the two template amplitudes of the two template angles can be accumulated for one mapping mode corresponding to the two template angles. Thus, the HoG can be generated by accumulating template amplitudes based on at least one mapping mode. The horizontal axis of the HoG can represent the intra-frame prediction mode index, and the vertical axis of the HoG can represent the accumulated intensity (e.g., amplitude). In this embodiment, the HoG is generated based on the template angle and template amplitude for selecting multiple candidate intra-frame modes.
[0087] Some intra-frame default modes can be selected as intra-frame candidate modes based on the cumulative amplitude in HoG. When the number of intra-frame candidate modes is equal to 6, 6 intra-frame prediction indices can be selected based on the first 6 amplitudes. When the number of intra-frame candidate modes is equal to 3, 3 intra-frame prediction indices can be selected based on the first 3 amplitudes. Thus, when the number of intra-frame candidate modes is equal to X, X intra-frame prediction indices can be selected based on the first X amplitudes. The number X can be a positive integer. In one implementation, non-angular modes from the intra-frame default modes can be directly added to the intra-frame candidate modes. For example, a non-angular mode can be a planar mode. In another implementation, a non-angular mode can be a DC mode.
[0088] Continue to refer to Figure 3 In box 340, decoder module 124 generates template predictions for each of the multiple intra-frame candidate modes.
[0089] Reference Figure 1 and Figure 2 The decoder module 124 can identify multiple template blocks adjacent to the block unit. Figure 5A and Figure 5B This is a schematic diagram of an exemplary implementation of a block unit, multiple template blocks, and a reference area. Figure 5A This is a schematic diagram of an exemplary embodiment of a block unit 5100, a plurality of template blocks 5101-5103 adjacent to the block unit 5100, and a reference region 5130 adjacent to the template blocks 5101-5103. In this embodiment, reference is made to... Figure 4B and Figure 5AThe adjacent region 4210 may be the same as multiple template blocks 5101-5103. The first template block in template block 5101 may be the left adjacent block located to the left of block unit 5100, the second template block in template block 5102 may be the top adjacent block located above block unit 5100, and the third template block in template block 5103 may be the upper left adjacent block located to the upper left of block unit 5100. The height of the top adjacent block may be equal to the number of reconstructed samples Nbt of the top adjacent block along the vertical direction, and the width of the top adjacent block may be equal to the width of block unit 4200. The height of the left adjacent block may be equal to the height of block unit 4200, and the width of the left adjacent block may be equal to the number of reconstructed samples Nbl of the left adjacent block along the horizontal direction. Additionally, the height of the upper left adjacent block may be equal to the number of reconstructed samples Nbt of the top adjacent block along the vertical direction, and the width of the upper left adjacent block may be equal to the number of reconstructed samples Nbl of the left adjacent block along the horizontal direction. In one embodiment, the numbers Nbt and Nbl may be positive integers. Furthermore, the quantities Nbt and Nbl can be the same or different from each other. Additionally, the quantities Nbt and Nbl can be greater than or equal to 2. For example, the quantity Nbt can be equal to 2, 3, or 4, and the quantity Nbt can be equal to 2, 3, or 4.
[0090] In some embodiments, decoder module 124 may identify template blocks 5101-5102 as template units 5110 to generate template predictions. In another embodiment, decoder module 124 may identify template blocks 5101-5103 as template units 5120 to generate template predictions. Decoder module 124 identifies a plurality of template references in reference region 5130 that are adjacent to the plurality of template blocks. Template references may be a plurality of reference samples reconstructed prior to reconstruction block unit 5100. Furthermore, template units may include a plurality of template samples reconstructed prior to reconstruction block unit 5100.
[0091] Block unit 5100 may have a block width W0 and a block height H0. First template block 5101 may have a first template width W1 and a first template height H0, second template block 5102 may have a second template width W0 and a second template height H2, and third template block 5103 may have a third template width W1 and a third template height H2. Reference area 5130 may have a reference width M and a reference height N. Furthermore, the reference width M may be equal to 2×(W0+W1)+1, and the reference height N may be equal to 2×(H0+H2)+1. In embodiments, the values W0, H0, W1, H2, M, and N may be positive integers. In one embodiment, the value W1 may be equal to the value H2. In another embodiment, the value W1 may be different from the value H2.
[0092] Figure 5BThis is a schematic diagram of an exemplary embodiment of block unit 5200, a plurality of template blocks 5201-5203 adjacent to block unit 5200, and a reference region 5230 adjacent to template blocks 5201-5203. The first template block in template block 5201 may be a left-adjacent block located to the left of block unit 5200, the second template block in template block 5202 may be a top-adjacent block located above block unit 5200, and the third template block in template block 5203 may be a top-left adjacent block located to the upper left of block unit 5200. In some embodiments, decoder module 124 may determine template blocks 5201-5202 as template units 5210 to generate template predictions. In another embodiment, decoder module 124 may determine template blocks 5201-5203 as template units 5220 to generate template predictions. Decoder module 124 determines a plurality of template references adjacent to the plurality of template blocks in reference region 5230. The template reference can be multiple reference samples reconstructed prior to the reconstruction block unit 5200. Furthermore, the template unit can include multiple template samples reconstructed prior to the reconstruction block unit 5200.
[0093] Block unit 5200 may have a block width W0 and a block height H0. First template block 5201 may have a first template width W1 and a first template height H1 greater than the block height H0; second template block 5202 may have a second template width W2 greater than the block width W0 and a second template height H2; and third template block 5203 may have a third template width W1 and a third template height H2. Reference region 5230 may have a reference width M and a reference height N. Furthermore, the reference width M may be equal to 2×(W1+W2)+1, and the reference height N may be equal to 2×(H1+H2)+1. In implementation, the values W0, H0, W1, H1, W2, H2, M, and N may be positive integers. In one embodiment, the value W1 may be equal to the value H2. In another embodiment, the value W1 may be different from the value H2.
[0094] Decoder module 124 can generate template predictions by predicting template blocks in a template unit based on a reference region with a template reference using intra-candidate modes. Decoder module 124 can also generate one template prediction by predicting a template block in a template unit using one of the intra-candidate modes based on a template reference. Thus, the number of intra-candidate modes can be equal to the number of template predictions. For example, when the number of intra-candidate modes is 6, the number of template predictions can also be 6.
[0095] Continue to refer to Figure 3 In box 350, decoder module 124 selects multiple prediction modes from multiple intra-frame candidate modes based on template prediction.
[0096] The prediction pattern is selected based on template blocks. (Refer to...) Figure 1 and Figure 2 The decoder module 124 can compare the template prediction with the template sample in the template unit. Since the template sample in the template unit is reconstructed before the reconstruction block unit, the template sample is also reconstructed before the template prediction is generated. Thus, when the template prediction is generated, the decoder module 124 is allowed to compare the template prediction with the reconstructed template sample in the template unit.
[0097] Decoder module 124 compares the template prediction of the template block with the template unit by selecting a prediction mode from intra-candidate modes using a cost function. Thus, the template block is predicted to generate a template prediction for selecting the prediction mode based on the template prediction of the template block. Decoder module 124 can determine multiple cost values by comparing the reconstructed template block with the template prediction. For example, decoder module 124 can compare the reconstructed template block with one of the template predictions generated using one of the intra-candidate modes to generate one cost value. Thus, each cost value determined by the cost function corresponds to one of the template predictions generated using one of the intra-candidate modes.
[0098] Cost functions may include, but are not limited to, Sum of Absolute Difference (SAD), Sum of Absolute Transformed Difference (SATD), Mean Absolute Difference (MAD), Mean Squared Difference (MSD), and Structural Similarity (SSIM). It should be noted that any cost function may be used without departing from the scope of this invention.
[0099] Decoder module 124 can select prediction modes from intra-candidate modes based on the cost value of template prediction, which is generated based on template blocks. When the number of prediction modes is equal to 2, 2 intra-prediction indices can be selected based on the 2 lowest cost values. When the number of intra-candidate modes is equal to 3, 3 intra-prediction indices can be selected based on the 3 lowest cost values. Therefore, when the number of prediction modes is equal to Y, Y intra-prediction indices can be selected based on Y lowest cost values. The number Y can be a positive integer.
[0100] When a prediction mode is selected, decoder module 124 can determine multiple weighting parameters based on the cost of template predictions generated based on template blocks. Thus, the weighting parameters are determined based on template blocks. Template blocks can be predicted to generate template predictions for selecting the prediction mode using template predictions based on template blocks. The weighting parameters can be determined based on cost values, and these cost values are determined by comparing the template block with each of the template predictions, respectively.
[0101] Decoder module 124 can compare the cost values of the predicted modes to determine the weighting parameters. For example, when the number of predicted modes is equal to 2, the weighting parameters of the two predicted modes can be determined based on the following function:
[0102]
[0103] Among them, w1, w2, C1, and C2 are the weighted parameters and cost values of the prediction model.
[0104] Decoder module 124 can predict block units based on a reference line of the block unit using prediction modes and weighting parameters. Decoder module 124 can predict block units based on prediction modes to generate multiple prediction blocks. Each prediction block corresponds to one of the prediction modes, and therefore each prediction block also corresponds to one of the weighting parameters. Decoder module 124 can generate prediction blocks of block units by weightedly combining prediction blocks and weighting parameters.
[0105] Reference Figure 1 and Figure 2 The decoder module 124 can directly weight and combine template predictions using multiple weighting parameters. These weighting parameters can be determined based on HoG (Homogeneous Geometry). For example, weighting parameters can be determined based on the cumulative magnitude of template blocks in HoG. Thus, multiple weighting parameters are determined based on multiple template blocks. For example, when the number of combined template predictions is equal to 3, the weighting parameters for three intra-frame candidate modes can be determined based on the following function:
[0106]
[0107] Among them, the values w1, w2, w3, A1, A2, and A3 are weighted parameters and cumulative magnitudes of intra-candidate modes selected based on template blocks. In the implementation, the weighted parameter w i It can be equal to Furthermore, the value p can be equal to the number of template predictions in the combination.
[0108] When the number of intra-candidate modes is equal to 3, the decoder module 124 can generate multiple intermediate predictions for block units based on the 3 intra-candidate modes. There can be 3 intermediate predictions, each generated using 2 of the intra-candidate modes. Alternatively, 1 intermediate prediction can be generated using 3 intra-candidate modes. When the number of intra-candidate modes is equal to Y, the number of intermediate predictions NI can be less than or equal to... For example, when the quantity Y equals 2, the quantity NI equals 1. When the quantity Y equals 3, the quantity NI can be less than or equal to 4. Furthermore, when the quantity Y equals 4, the quantity NI can be less than or equal to 11. For example, when the quantity Y equals 3 and the decoder module 124 selects two intra-frame candidate modes to generate each intermediate prediction, the quantity NI can be equal to 3. In one implementation, the function... This represents the number of m combinations of Y elements, where the number NI is a positive integer greater than or equal to 1, and the number m is a positive integer greater than or equal to 2.
[0109] When intermediate predictions are generated using m of Y intra-candidate patterns, the weighting parameters can be determined based on the cumulative magnitude of the m intra-candidate patterns. For example, the quantity Y equals 4, and the quantity m equals 2. Then, the two weighting parameters can be determined based solely on the two cumulative magnitudes of the two intra-candidate patterns used to generate one of the intermediate predictions. For example, the two weighting parameters could be equal to A1 / (A1+A2) and A2 / (A1+A2).
[0110] Decoder module 124 can select a prediction block from intermediate predictions by comparing intermediate predictions with template units using a cost function. Decoder module 124 can determine multiple cost values by comparing the reconstructed template block with intermediate predictions. For example, decoder module 124 can compare the reconstructed template block with one of the intermediate predictions to generate one cost value. Thus, each of the cost values determined by the cost function corresponds to one of the intermediate predictions generated using at least two intra-candidate modes.
[0111] Cost functions may include, but are not limited to, Sum of Absolute Difference (SAD), Sum of Absolute Transformed Difference (SATD), Mean Absolute Difference (MAD), Mean Squared Difference (MSD), and Structural Similarity (SSIM). It should be noted that any cost function may be used without departing from the scope of this invention.
[0112] Decoder module 124 can select a prediction block from intermediate predictions based on the cost value of the intermediate predictions generated based on the template block. Decoder module 124 can select a specific one from the intermediate predictions with the lowest cost value as the prediction block. Thus, the intra-frame candidate mode used to generate a specific intermediate prediction can be regarded as the prediction mode.
[0113] Return to Figure 3 In box 360, decoder module 124 reconstructs block units based on multiple prediction modes.
[0114] Further reference Figure 1 and Figure 2 Decoder module 124 can determine multiple residual components from the bitstream of a block unit and add the residual components to the prediction block to reconstruct the block unit. Decoder module 124 can reconstruct all other block units in the image frame in order to reconstruct the image frame and the video.
[0115] Figure 6 Exemplary embodiments according to this disclosure are shown in Figure 1 The diagram shows a block diagram of an encoder module 114 of a first electronic device 110. The encoder module 114 may include a prediction processor (e.g., prediction processing unit 6141), at least a first summer (e.g., first summer 6142) and a second summer (e.g., second summer 6145), a transform / quantization processor (e.g., transform / quantization unit 6143), an inverse quantization / inverse transform processor (e.g., inverse quantization / inverse transform unit 6144), a filter (e.g., filter unit 6146), a decoded image buffer (e.g., decoded image buffer 6147), and an entropy encoder (e.g., entropy coding unit 6148). The prediction processing unit 6141 of the encoder module 114 may further include a segmentation processor (e.g., segmentation unit 61411), an intra-frame prediction processor (e.g., intra-frame prediction unit 61412), and an inter-frame prediction processor (e.g., inter-frame prediction unit 61413).
[0116] Encoder module 114 can receive source video and encode the source video to output a bitstream. Encoder module 114 can receive source video comprising multiple image frames, and then divide the image frames according to the encoding structure. Each image frame can be divided into at least one image block.
[0117] At least one image block may include a luminance block having multiple luminance samples and at least one chrominance block having multiple chrominance samples. The luminance block and at least one chrominance block may be further subdivided to generate macroblocks, coding tree units (CTUs), coding blocks (CBs), their sub-segments, and / or another equivalent coding unit.
[0118] Encoder module 114 can perform additional sub-segmentation of the source video. It should be noted that the disclosed implementation is generally applicable to video encoding, regardless of how the source video is segmented before and / or during encoding.
[0119] During the encoding process, the prediction processing unit 6141 may receive the current image block of a specific image frame. The current image block may be one of the luma blocks or chroma blocks in the specific image frame.
[0120] Segmentation unit 61411 can divide the current image block into multiple block units. Intra-frame prediction unit 61412 can perform intra-frame prediction coding of the current block unit relative to one or more adjacent blocks in the same frame as the current block unit to provide spatial prediction. Inter-frame prediction unit 61413 can perform inter-frame prediction coding of the current block unit relative to one or more blocks in one or more reference image blocks to provide temporal prediction.
[0121] The prediction processing unit 6141 may select one of the coding results generated by the intra-prediction unit 61412 and the inter-prediction unit 61413 based on a mode selection method (e.g., a cost function). The mode selection method may be a rate-distortion optimization (RDO) process.
[0122] The prediction processing unit 6141 can determine the selected encoding result and provide the prediction block corresponding to the selected encoding result to the first summer 6142 for generating the residual block, and to the second summer 6145 for reconstructing the encoded block unit. The prediction processing unit 6141 can further provide syntax elements such as motion vectors, intra-frame mode indicators, segmentation information, and other syntax information to the entropy coding unit 6148.
[0123] Intra-prediction unit 61412 can perform intra-prediction on the current block unit. Intra-prediction unit 61412 can determine the intra-prediction mode for the reconstructed samples adjacent to the current block unit so as to encode the current block unit.
[0124] Intra-prediction unit 61412 can encode the current block cell using various intra-prediction modes. Intra-prediction unit 61412 of prediction processing unit 6141 can select an appropriate intra-prediction mode from the selected modes. Intra-prediction unit 61412 can encode the current block cell using a cross-component prediction mode to predict one of the two chrominance components of the current block cell based on the luma component of the current block cell. Intra-prediction unit 61412 can predict the first of the two chrominance components of the current block cell based on the second of the two chrominance components of the current block cell.
[0125] As an alternative to intra-prediction performed by intra-prediction unit 61412, inter-prediction unit 61413 may perform inter-prediction on the current block unit. Inter-prediction unit 61413 may perform motion estimation to estimate the motion of the current block unit used to generate motion vectors.
[0126] The motion vector indicates the displacement of the current block cell within the current image block relative to the reference block cell within the reference image block. The inter-frame prediction unit 61413 may receive at least one reference image block stored in the decoded image buffer 6147 and estimate motion based on the received reference image block to generate a motion vector.
[0127] The first summer 6142 can generate a residual block by subtracting the predicted block determined by the prediction processing unit 6141 from the original current block unit. The first summer 6142 may represent one or more components performing the subtraction.
[0128] The transform / quantization unit 6143 can apply a transform to the residual block to generate residual transform coefficients, and then quantize the residual transform coefficients to further reduce the bit rate. The transform can be one of DCT, DST, AMT, MDNSST, HyGT, signal correlation transform, KLT, wavelet transform, integer transform, subband transform, or a conceptually similar transform.
[0129] This transformation converts residual information from the pixel value domain to the transform domain, such as the frequency domain. The degree of quantization can be modified by adjusting the quantization parameters.
[0130] The transform / quantization unit 6143 can perform a scan of a matrix including quantized transform coefficients. Alternatively, the entropy encoding unit 6148 can perform the scan.
[0131] The entropy coding unit 6148 can receive multiple syntax elements, including quantization parameters, transform data, motion vectors, intra-frame modes, segmentation information, and other syntax information, from the prediction processing unit 6141 and the transform / quantization unit 6143. The entropy coding unit 6148 can encode the syntax elements into a bit stream.
[0132] Entropy coding unit 6148 can entropy code the quantized transform coefficients to generate an encoded bitstream by performing CAVLC, CABAC, SBAC, PIPE coding, or another entropy coding technique. The encoded bitstream can be transmitted to another device (i.e., Figure 1 The second electronic device 120 in the system may be archived for later transmission or retrieval.
[0133] The inverse quantization / inverse transform unit 6144 can apply inverse quantization and inverse transform to reconstruct the residual block in the pixel domain for later use as a reference block. The second summer 6145 can add the reconstructed residual block to the prediction block provided from the prediction processing unit 6141 to generate the reconstructed block stored in the decoded image buffer 6147.
[0134] Filtering unit 6146 may include a deblocking filter, a SAO filter, a bilateral filter, and / or an ALF to remove block artifacts from the reconstructed block. In addition to the deblocking filter, SAO filter, bilateral filter, and ALF, additional filters (in-loop or post-loop) may be used. For simplicity, these filters are not described, and the output of the second summer 6145 may be filtered.
[0135] The decoded image buffer 6147 may be a reference image memory that stores reference blocks for the encoder module 614 to encode video in modes such as intra-frame or inter-frame coding. The decoded image buffer 6147 may include various memory devices, such as DRAM (e.g., including SDRAM, MRAM, RRAM) or other types of memory devices. The decoded image buffer 6147 may be on-chip along with other components of the encoder module 114, or off-chip relative to those components.
[0136] Encoder module 114 can receive video data and predict multiple image frames in the video data using multiple intra-frame default modes through method 300. The video data can be the video to be encoded. Encoder module 114 can determine a block unit from one of the image frames based on the video data.
[0137] Encoder module 114 can select multiple intra-candidate modes from the intra-default mode used for the block unit. The intra-candidate modes can be multiple most probable modes (MPMs) determined based on at least one reconstruction mode for multiple neighboring regions of the block unit. The intra-candidate modes can be selected based on gradient histograms (HoGs) generated by multiple template angles and multiple template amplitudes derived from multiple neighboring regions.
[0138] Encoder module 114 can generate template predictions for each intra-candidate mode. Template cells adjacent to block cells can be predicted based on a reference region using each intra-candidate mode. The number of template predictions can be equal to the number of intra-candidate modes.
[0139] Encoder module 114 can select multiple prediction modes from intra-candidate modes based on template prediction. The template prediction can be compared with multiple reconstructed samples in the template unit using a cost function for selecting the prediction mode. In another embodiment, the template predictions can be directly weighted and combined to generate multiple combined template predictions. The combined template predictions can be compared with reconstructed samples in the template unit using a cost function for selecting the prediction mode.
[0140] Encoder module 114 can determine prediction blocks based on a prediction mode and compare multiple pixel elements in a block unit with the prediction block to determine multiple residual values. Encoder module 114 can encode the residual values into a bitstream for transmission to second electronic device 120. Furthermore, to further encode other blocks in the image frame and other image frames, encoder module 114 can further reconstruct block units based on the prediction block and residual values. Thus, encoder module 114 can also use an intra-frame default mode to predict image frames in the video data using method 300.
[0141] The disclosed embodiments should be considered illustrative rather than restrictive in all respects. It should also be understood that while this disclosure is not limited to the specific disclosed embodiments, many rearrangements, modifications, and substitutions are possible without departing from the scope of this disclosure.
Claims
1. A method for decoding a bitstream by an electronic device, the method comprising: receiving the bitstream; determining a block unit from image frames according to the bitstream; selecting a plurality of intra-candidate modes from a plurality of intra-default modes for the block unit, wherein: the block unit is adjacent to a plurality of template regions, a plurality of template angles and a plurality of template magnitudes are generated by filtering the plurality of template regions using a gradient filter, each of the plurality of template angles corresponds to one of the plurality of template magnitudes, and the plurality of intra-candidate modes are selected based on a histogram of gradients (HoG) generated by the plurality of template angles and the plurality of template magnitudes; determining a plurality of template blocks adjacent to the block unit; determining a plurality of template references adjacent to the plurality of template blocks; generating template predictions for each of the plurality of intra-candidate modes, wherein the plurality of template blocks are predicted based on the plurality of template references using the plurality of intra-candidate modes to generate the template predictions; determining a plurality of cost values by comparing the plurality of template blocks to each of the template predictions, respectively; selecting a plurality of prediction modes from the plurality of intra-candidate modes based on the plurality of cost values; predicting the block unit based on the plurality of prediction modes to generate a plurality of prediction blocks, wherein one of the plurality of prediction blocks corresponds to one of the plurality of prediction modes; weighting combining the plurality of prediction blocks to generate a prediction block having a plurality of weighting parameters; and reconstructing the block unit based on the generated prediction block having the plurality of weighting parameters. 2.The method of claim 1, wherein: the plurality of cost values are determined using a cost function, and each of the plurality of cost values corresponds to one of the template predictions. 3.The method of claim 1, wherein: the plurality of template blocks include top and left neighboring blocks, and each of the plurality of template blocks is adjacent to the block unit. 4.The method of claim 1, wherein: the plurality of prediction modes are selected based on the plurality of template blocks; and the plurality of weighting parameters are determined based on the plurality of template blocks.
5. The method of claim 1, wherein, the method further comprises: mapping each of the plurality of template angles to one of the plurality of intra-default modes based on a predefined relationship to generate at least one mapping mode; and generating the HoG by accumulating the plurality of template magnitudes based on the at least one mapping mode, and selecting the plurality of intra-candidate modes from the plurality of intra-default modes based on accumulated magnitudes in the HoG. 6.An electronic device for decoding a bitstream, the electronic device comprising: at least one processor; and a memory coupled to the at least one processor and storing a plurality of instructions that, when executed by the at least one processor, cause the electronic device to: receive the bitstream; determine a block unit from image frames according to the bitstream; Multiple intra-candidate modes are selected from multiple intra-default modes used for the block unit, wherein: The block unit is adjacent to multiple template regions. Multiple template angles and multiple template amplitudes are generated by filtering the multiple template regions using a gradient filter. Each of the plurality of template angles corresponds to one of the plurality of template amplitudes, and The multiple intra-frame candidate modes are selected based on gradient histograms (HoG) generated by the multiple template angles and the multiple template amplitudes; Identify multiple template blocks adjacent to the block unit; Determine multiple template references adjacent to the multiple template blocks; For each of the plurality of intra-candidate modes, a template prediction is generated, wherein the template prediction is generated by using the plurality of intra-candidate modes and predicting the plurality of template blocks based on the plurality of template references; Multiple cost values are determined by comparing each of the multiple template blocks with each of the template predictions; Multiple prediction modes are selected from the multiple intra-frame candidate modes based on the multiple cost values; The block unit is predicted based on the plurality of prediction modes to generate a plurality of prediction blocks, wherein one of the plurality of prediction blocks corresponds to one of the plurality of prediction modes; The multiple prediction blocks are weighted and combined to generate a prediction block with multiple weighting parameters; and The block cell is reconstructed based on the generated prediction block with the multiple weighting parameters.
7. The electronic device according to claim 6, characterized in that, The multiple prediction patterns are selected based on the multiple template blocks; and The weighted parameters are determined based on the template blocks.
8. The electronic device of claim 6, wherein, The plurality of instructions, when executed by the at least one processor, also cause the electronic device to: Based on a predefined relationship, each of the multiple template angles is mapped to one of the multiple intra-frame default modes to generate at least one mapping mode. as well as The HoG is generated by accumulating the magnitudes of the plurality of templates based on the at least one mapping mode, and the plurality of intra-frame candidate modes are selected from the plurality of intra-frame default modes based on the magnitudes accumulated in the HoG.
Citation Information
Patent Citations
Template matching for JVET intra prediction
US20170339404A1