Method for decoding, encoding and electronic device for blocks in a video bitstream - Patent Application 20070122997
Intra-warped prediction using weighted basis functions addresses redundancy in video coding, enhancing compression efficiency and bandwidth optimization in video transmission.
Patent Information
- Application Number
- JP2025540042
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2023-10-31
- Publication Date
- 2026-01-16
AI Technical Summary
Existing video coding techniques struggle to effectively reduce redundancy in uncompressed digital video signals, particularly in intra-prediction methods, which can lead to inefficient compression and transmission bandwidth requirements.
Intra-warped prediction is employed using a weighted sum of basis functions at pixel coordinate locations within a video block, utilizing polynomial and trigonometric functions, along with predefined eigenvectors and clipping mechanisms to enhance prediction accuracy and compression efficiency.
This approach improves video coding efficiency by reducing redundancy and optimizing bandwidth usage through advanced intra-prediction techniques, enabling more effective compression and decoding processes.
Smart Images

Figure 2026501770000001_ABST
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 443,901, entitled "Intra Warp Mode," filed February 7, 2023, which is incorporated herein by reference in its entirety, and to U.S. Non-Provisional Patent Application No. 18 / 497,726, entitled "A warp Mode for Intra Prediction," filed October 30, 2023.
[0002] [Technical field] This disclosure describes a set of advanced video coding techniques, and in particular relates to intra-warped prediction of samples, in which samples are intra-predicted using a weighted sum of basis functions at the pixel coordinate locations of the samples within a video block. [Background technology]
[0003] Uncompressed digital video may include a series of pictures and may have specific bit rate requirements for transmission bandwidth in storage, data processing, and streaming applications. One goal of video coding and decoding may be to reduce redundancy in the uncompressed input video signal through various compression techniques. Summary of the Invention
[0004] This disclosure describes a set of advanced video coding techniques, and in particular relates to intra-warped prediction of samples, in which samples are intra-predicted using a weighted sum of basis functions at the pixel coordinate locations of the samples within a video block.
[0005] In some implementations, a method for decoding a block in a video bitstream is disclosed, the method comprising: determining, based on a syntax element in the video bitstream, that intra-warp prediction is applied to the block; determining a set of weighting factors to be applied to a set of basis functions used in the intra-warp prediction; predicting values of the samples of the block based on a weighted sum of the set of basis functions on pixel coordinate locations of the samples, the set of basis functions being each weighted according to the set of weighting coefficients; reconstructing a block containing said samples based on said predicted values; may include:
[0006] In the above-described exemplary implementation, the set of basis functions includes power functions of the pixel coordinate locations of the samples up to order N, where N is an integer; The weighted sum of the basis functions includes polynomials of the pixel coordinate locations of the samples up to order N.
[0007] In any one of the example implementations described above, N=2 and the set of weighting factors is: Six coefficients of a two-dimensional polynomial term of the horizontal and vertical pixel coordinates of said sample; five coefficients of a two-dimensional polynomial term in the horizontal and vertical pixel coordinates of said sample, excluding one cross term; or four coefficients of two-dimensional polynomial terms that depend on the horizontal and vertical pixel coordinates of the sample, excluding one cross term, and at least one further coordinate-independent term; may include:
[0008] In any one of the above example implementations, the at least one coordinate-independent term includes an offset term that includes a DC value obtained from a reconstruction neighborhood of the sample.
[0009] In any one of the above-described exemplary implementations, the set of weighting coefficients further includes an offset coefficient, and the at least one coordinate-independent term includes an offset term that is a product of the offset coefficient and a reconstructed value of a sample in an upper-left corner of the block; or the set of weighting coefficients further includes a first offset coefficient and a second offset coefficient, and the at least one coordinate independent term includes a first offset term and a second offset term, the first offset term being a product of the first offset coefficient and an above reconstruction neighbor of the sample, and the second offset term being a product of the second offset coefficient and a left reconstruction neighbor of the sample; or The set of weighting coefficients further includes first, second, third, and fourth offset coefficients, and the at least one coordinate independent term includes first, second, third, and fourth offset terms, where the first offset term is a product of the first offset coefficient and an above reconstruction neighborhood of the sample, the second offset term is a product of the second offset coefficient and a left reconstruction neighborhood of the sample, the third offset term is a product of the third offset coefficient and an above-right reconstruction neighborhood of the sample, and the fourth offset term is a product of the fourth offset coefficient and a below-left reconstruction neighborhood of the sample.
[0010] In any one of the example implementations described above, N=2 and the weighted sum of the basis functions may include:
number
[0011] In any one of the above example implementations, the set of basis functions includes up to N trigonometric basis functions for the pixel coordinate locations of the samples, where N is an integer, and the N trigonometric basis functions are: the first N type-2 discrete cosine transform bases, the first N type 7 discrete sine transform bases, or Any N discrete sine transform bases of types 1 to 8 and discrete cosine transform bases of types 1 to 8, may include:
[0012] In any one of the above example implementations, the set of basis functions includes up to N Karhunen-Loeve Transform (KLT) bases as a function of the pixel coordinate location of the sample, where N is an integer.
[0013] In any one of the above example implementations, the N KLT bases include eigenvectors of a covariance matrix derived from a reconstruction neighborhood of the sample, or are predefined eigenvectors known to both the encoder and decoder.
[0014] In any one of the above example implementations, the set of basis functions includes a subset of power functions and a subset of trigonometric functions of pixel coordinate positions of the samples.
[0015] In any one of the above example implementations, the method further includes parsing the video bitstream to determine the set of weighting factors signaled in the video bitstream.
[0016] In any one of the above example implementations, the set of weighting factors is derived by the encoder using a multi-line template above and to the left of the block.
[0017] In any one of the above-described exemplary implementations, parsing the video bitstream to determine the weighting factors comprises: parsing the video bitstream to obtain an index of the set of weighting factors from the video bitstream; identifying a set of weight confidences from among a plurality of sets of weight coefficients according to the index; may include:
[0018] In any one of the above example implementations, the set of weighting factors is clipped to a predetermined range of values.
[0019] In any one of the above-described exemplary implementations, the predicted value of the sample is clipped to a predetermined range of values, and the clipped minimum and clipped maximum values of the predicted value are determined by an internal bit depth range.
[0020] In any one of the above example implementations, the set of weighting factors is quantized according to one or more predetermined precisions.
[0021] In any one of the above example implementations, at least two of the sets of weighting factors are quantized to different precisions.
[0022] In any one of the above example implementations, terms in a weighted sum of the set of basis functions associated with the same precision are calculated and rounded together.
[0023] In any one of the above example implementations, the precision of the set of weighting factors is signaled in a high-level syntax in the video bitstream.
[0024] In some other implementations, an apparatus for processing video information is disclosed, which may include circuitry configured to perform any one of the implementations of the methods described above.
[0025] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding and / or encoding. [Brief explanation of the drawings]
[0026] Further features, characteristics, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings.
[0027] [Figure 1] 1 shows a simplified block diagram schematic of a communication system (100) according to an exemplary embodiment.
[0028] [Figure 2] 2 shows a simplified block diagram schematic of a communication system (200) according to an exemplary embodiment.
[0029] [Figure 3] 1 shows a simplified block diagram schematic of a video decoder according to an exemplary embodiment;
[0030] [Figure 4] 1 shows a schematic diagram of a simplified block diagram of a video encoder, according to an example embodiment;
[0031] [Figure 5] 1 shows a block diagram of a video encoder according to another example embodiment.
[0032] [Figure 6] 1 illustrates a block diagram of a video decoder according to another exemplary embodiment.
[0033] [Figure 7] 1 shows a schematic diagram of an exemplary subset of intra-prediction directional modes.
[0034] [Figure 8] 10 shows the nominal angle in directional intra prediction.
[0035] [Figure 9] 1 shows a diagram of exemplary intra-prediction directions.
[0036] [Figure 10] Indicates the top, left and top left positions of the PAETH mode of the coding block.
[0037] [Figure 11] 1 illustrates an exemplary recursive intra-filtering mode.
[0038] [Figure 12]10 shows a flowchart for performing an exemplary intra prediction in intra warp mode.
[0039] [Figure 13] 1 shows a schematic diagram of a computer system according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0040] Throughout this specification and claims, terms may have nuanced meanings that are suggested or implied in context beyond their explicitly stated meaning. The phrases "in one embodiment / implementation" or "in some embodiments / implementations" used herein do not necessarily refer to the same embodiment / implementation, and the phrases "in another embodiment / implementation" or "in other embodiments" used herein do not necessarily refer to different embodiments. For example, claimed subject matter is intended to include combinations of example embodiments / implementations, in whole or in part.
[0041] Generally, terms can be understood, at least in part, from their use in context. For example, as used herein, terms such as "and," "or," or "and / or" can include a variety of meanings, depending on the context. Typically, when used to relate a list such as A, B, or C, "or" is intended to mean A, B, and C, which is used herein in an inclusive sense, as well as A, B, or C, which is used herein in an exclusive sense. Furthermore, as used herein, the terms "one or more," "at least one," "a," "an," or "the" can be used in a singular or plural sense, depending at least in part on the context. Furthermore, the terms "based on" or "determined by" can be understood as not necessarily intended to convey an exclusive set of elements, but instead can allow for the presence of additional elements, again not necessarily explicitly recited, depending at least in part on the context.
[0042] FIG. 1 illustrates a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The communication system (100) includes multiple terminal devices, e.g., 110, 120, 130, and 140, that can communicate with each other, e.g., via a network (150). In the example of FIG. 1, a first pair of terminal devices (110) and (120) may perform unidirectional transmission of data. For example, the terminal device (110) may code video data in the form of one or more coded bitstreams (e.g., streams of video pictures captured by the terminal device (110)) for transmission over the network (150). The terminal device (120) may receive the coded video data from the network (150), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. The unidirectional data transmission may be implemented in a media serving application, for example.
[0043] In another example, a second pair of terminal devices 130 and 140 may perform a bidirectional transmission of coded video data, such as during a video conference, where each of the terminal devices 130 and 140 can code video data for transmission (e.g., of a stream of video pictures captured by the terminal device) and can receive coded video data from the other of the terminal devices 130 and 140 to recover and display the video pictures.
[0044] In the example of FIG. 1 , the terminal devices may be implemented as servers, personal computers, and smartphones, although the underlying principles of the present disclosure are not limited thereto. Embodiments of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, and the like. Network 150 represents any number or type of network that conveys coded video data between terminal devices, including, for example, wired and / or wireless communication networks. Communication network 150 may exchange data over circuit-switched, packet-switched, and / or other types of channels. Exemplary networks include electronic communication networks, local area networks, wide area networks, and / or the Internet.
[0045] 2 illustrates the arrangement of a video encoder and a video decoder in a video streaming environment as an example of an application of the disclosed subject matter, which is equally applicable to other video applications including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0046] As shown in Figure 2, a video streaming system may include a video source (201), e.g., a video capture subsystem (213), which may include a digital camera, for creating a stream of uncompressed video pictures or images (202). In one example, the video picture stream (202) includes samples recorded by the digital camera of the video source (201). The video picture stream (202), shown in bold to emphasize its high data capacity when compared to the encoded video data (204) (or coded video bitstream), may be processed by an electronic device (220) including a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination thereof, and may enable or implement aspects of the disclosed subject matter, as described in more detail below. The coded video data 204 (or coded video bitstream 204), shown with thin lines to emphasize its lower data volume compared to the uncompressed video picture stream 202, can be stored on the streaming server 205 for future use or directly on a downstream video device (not shown). One or more streaming client subsystems, such as the client subsystems 206 and 208 of FIG. 2, can access the streaming server 205 to retrieve copies 207 and 209 of the coded video data 204. The client subsystem 206 may include a video decoder 210, for example, within the electronic device 230. The video decoder 210 decodes the input copy of the coded video data 207 and generates an uncompressed output video picture stream 211 that can be rendered on a display 212 (e.g., a display screen) or other rendering device (not shown).
[0047] 3 shows a block diagram of a video decoder (310) of an electronic device (330) according to any of the following embodiments of the present disclosure. The electronic device (330) may include a receiver (331) (e.g., a receiving circuit). The video decoder (310) can be used in place of the video decoder (210) in the example of FIG. 2.
[0048] As shown in FIG. 3, the receiver (331) may receive one or more coded video sequences from the channel (301). To address network jitter and / or handle playback timing, a buffer memory (315) may be disposed between the receiver (331) and an entropy decoder / parser (320) (hereinafter, "parser (320)"). The parser (320) may reconstruct symbols (321) from the coded video sequence. Categories of these symbols include information used to manage the operation of the video decoder (310) and information for controlling a rendering device such as a display (312) (e.g., a display screen). The parser (320) may parse / entropy decode the coded video sequence. The parser (320) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder. Subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (320) may also extract information from the coded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantization parameter values, motion vectors, etc. The reconstruction of the symbols (321) may involve multiple different processing or functional units. The units included and how they are included can be controlled by subgroup control information parsed from the coded video sequence by the parser (320).
[0049] The first unit may include a scalar / inverse transform unit (351), which may receive quantized transform coefficients and control information from the parser (320) as symbols (321), including information indicating which type of inverse transform should be used, block size, quantization coefficients / parameters, quantization scaling matrices, etc. The scalar / inverse transform unit (351) may output blocks containing sample values that may be input to an aggregator (355).
[0050] In some examples, the output samples of the scaler / inverse transform (351) are derived relative to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) may generate a block of the same size and shape as the block being reconstructed using already reconstructed surrounding block information stored in the current picture buffer (358). The current picture buffer (358), for example, buffers the reconstructed current picture partially and / or completely. In some implementations, the aggregator (355) may add the prediction information generated by the intra-prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a sample-by-sample basis.
[0051] In other cases, the output samples of the scaler / inverse transform unit (351) may relate to an inter-coded, possibly motion-compensated, block. In such cases, the motion-compensated prediction unit (353) can access a reference picture memory (357) based on a motion vector to fetch samples used for inter-picture prediction. After motion-compensating the fetched reference samples according to the symbols (321) associated with the block, these samples can be added by an aggregator (355) to the output of the scaler / inverse transform unit (351) to generate output sample information (the output of unit 351 may be referred to as residual samples or a residual signal).
[0052] The output samples of the aggregator (355) may undergo various loop filtering techniques in a loop filter unit (356), which may include several types of loop filters. The output of the loop filter unit (356) may be a sample stream that can be output to a rendering device (312) and stored in a reference picture memory (357) for use in future inter-picture prediction.
[0053] 4 shows a block diagram of a video encoder (403) according to an exemplary embodiment of the present disclosure. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may further include a transmitter (440) (e.g., a transmitting circuit). The video encoder (403) may be used in place of the video encoder (403) in the example of FIG. 4.
[0054] The video encoder (403) may receive video samples from a video source (401). According to some example embodiments, the video encoder (403) may code and compress pictures of a source video sequence into a coded video sequence (443) in real time or under any other time constraint required by the application. Enforcing an appropriate coding rate constitutes one function of the controller (450). In some embodiments, the controller (450) may be functionally coupled to and control other functional units, as described below. Parameters set by the controller (450) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.
[0055] In some example embodiments, the video encoder (403) may be configured to operate within a coding loop. The coding loop may include a source encoder (430) and a (local) decoder (433) embedded in the video encoder (403). The decoder (433) reconstructs symbols to create sample data in a manner similar to that of a (remote) decoder, even if the embedded decoder 433 processes a video stream coded by the source encoder 430 without entropy coding (because the compression between the symbols in entropy coding and the coded video bitstream may be lossless in the video compression techniques considered in the disclosed subject matter). A consideration in this regard is that any decoder technology, with the exception of parsing / entropy decoding, which may exist only in the decoder, must also exist in substantially the same functional form as in the corresponding encoder. For this reason, the disclosed subject matter may focus on decoder operations related to the decoding portion of the encoder. Descriptions of encoder technology may therefore be omitted, as they are the reverse of the decoder technology, which is described generically. Only in certain areas or aspects will a more detailed description of the encoder be provided below.
[0056] In operation, in some example implementations, the source coder (430) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from the video sequence designated as "reference pictures."
[0057] The local video decoder (433) may decode coded video data for pictures that may be designated as reference pictures. The local video decoder (433) may replicate the decoding process that may be performed by a video decoder on the reference pictures, resulting in reconstructed reference pictures to be stored in a reference picture cache (434). In this way, the video encoder (403) may store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures obtained by a far-end (remote) video decoder (in the absence of transmission errors).
[0058] The predictor (435) may perform a predictive search for the coding engine (432). That is, for a new picture to be coded, the predictor (435) may search the reference picture memory (434) for sample data (such as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that can serve as suitable prediction references for the new picture.
[0059] The control unit (450) may manage the coding operations of the source coder (430), including, for example, setting parameters and subgroup parameters used for encoding the video data.
[0060] The output of all of the above functional units may undergo entropy coding in an entropy coder (445). The transmitter (440) may buffer the coded video sequence produced by the entropy coder (445) for transmission over a communication channel (460), which may be a hardware / software link to a storage device that may store the coded video data. The transmitter (440) may merge the coded video data from the video coder (403) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0061] The controller (450) may manage the operation of the video encoder (403). During coding, the controller (450) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures are often assigned as one of the following picture types: intra picture (I picture), predicted picture (P picture), bidirectionally predicted picture (B picture), or multi-predicted picture. A source picture may generally be spatially subdivided into multiple sample coding blocks, as described in more detail below.
[0062] 5 shows a diagram of a video encoder (503) according to another exemplary embodiment of this disclosure. The video encoder (503) is configured to receive a processed block (e.g., a predictive block) of sample values in a current video picture in a video picture sequence and encode the processed block into a coded picture that is part of the coded video sequence. The exemplary video encoder (503) can be used in place of the video encoder (403) in the example of FIG. 4.
[0063] For example, the video encoder (503) receives a matrix of sample values for a processing block. The video encoder (503) then determines whether the processing block is best coded using intra-mode, inter-mode, or bi-predictive mode, for example, using rate-distortion optimization (RDO).
[0064] In the example of Figure 7, the video encoder (503) includes an inter-encoder (530), an intra-encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general control unit (521), and an entropy encoder (525) coupled together as shown in the exemplary configuration of Figure 7.
[0065] The inter-encoder (530) is configured to receive samples of a current block (e.g., a block being processed), compare the block with one or more reference blocks in a reference picture (e.g., blocks in previous and subsequent pictures in display order), generate inter-prediction information (e.g., a description of redundant information due to inter-coding techniques, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique.
[0066] The intra encoder (522) is configured to receive samples of a current block (e.g., a block being processed), compare the block to previously coded blocks in a sample picture, generate quantized coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques).
[0067] The general control unit (521) may be configured to determine general control data and control other components of the video encoder (503) based on the general control data, for example, to determine a prediction mode for the block and provide a control signal to the switch (526) based on the prediction mode.
[0068] The residual calculator (523) may be configured to calculate the difference (residual data) between the received block and a prediction result of the selected block from the intra-encoder (522) or inter-encoder (530). The residual encoder (524) may be configured to encode the residual data to generate transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various exemplary embodiments, the video encoder (503) also includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform and generate decoded residual data. The entropy encoder (525) may be configured to format the bitstream to include the coded blocks and perform entropy coding.
[0069] 6 shows a diagram of an exemplary video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. By way of example, the video decoder (610) can be used in place of the video decoder (410) in the example of FIG. 4.
[0070] In the example of Figure 6, the video decoder (610) includes an entropy decoder (671), an inter-decoder (680), a residual decoder (673), a reconstruction module (674), and an intra-decoder (672), coupled together as shown in the exemplary configuration of Figure 6.
[0071] The entropy decoder (671) may be configured to reconstruct, from the coded picture, certain symbols representing generated syntax elements of the coded picture. The inter decoder (680) may be configured to receive inter prediction information and generate inter prediction results based on the inter prediction information. The intra decoder (672) may be configured to receive intra prediction information and generate prediction results based on the intra prediction information. The residual decoder (673) may be configured to perform inverse quantization, extract inverse quantized transform coefficients, and process the inverse quantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The reconstruction module (674) may be configured to combine, in the spatial domain, the residual as output by the residual decoder (673) and the prediction results (possibly as output by the inter or intra prediction module) to form reconstructed blocks that form part of the reconstructed picture as part of the reconstructed video.
[0072] It is noted that the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610) may be implemented using any suitable technology. In some exemplary embodiments, the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610) may be implemented using one or more processors executing software instructions.
[0073] Turning to block partitioning for coding and decoding, general partitioning can start from a base block and follow a predefined set of rules, a specific pattern, a partition tree, or any partition structure or scheme. Partitioning can be hierarchical and recursive. After dividing or partitioning the base block according to any of the exemplary partitioning procedures described below or other procedures, or a combination thereof, a final set of partitions or coding blocks can be obtained. Each of these partitions can be one of various partitioning levels within the partitioning hierarchy and can have various shapes. Each partition can be referred to as a coding block (CB). In various exemplary partitioning implementations described further below, each resulting CB can be of any allowable size and partitioning level. Such partitions are referred to as coding blocks because they form the units at which some basic coding / decoding decisions are made, coding / decoding parameters are optimized and determined, and signaled in the coded video bitstream. The highest or deepest level of the final partitions represents the depth of the coding block partitioning structure in the tree. A coding block may be a luma coding block or a chroma coding block. The CB tree structure for each color may be referred to as a coding block tree (CBT). The coding blocks of all color channels may be collectively referred to as a coding unit (CU). The hierarchical structures of all color channels may be collectively referred to as a coding tree unit (CTU). The partitioning pattern or structure of various color channels within a CTU may or may not be the same.
[0074] In some implementations, the partition tree scheme or structure used for the luma and chroma channels may not need to be the same. In other words, the luma and chroma channels may have separate coding tree structures or patterns. Furthermore, whether the luma and chroma channels use the same or different coding partition tree structures, and the actual coding partition tree structure used, may depend on whether the slice being coded is a P, B, or I slice. For example, for an I slice, the chroma and luma channels may have separate coding partition tree structures or coding partition tree structure modes, while for a P or B slice, the luma and chroma channels may share the same coding partition tree scheme. When separate coding partition tree structures or modes are applied, the luma channel may be partitioned into CBs by one coding partition tree structure, and the chroma channels may be partitioned into chroma CBs by another coding partition tree structure.
[0075] A video block (PB or CB, also referred to as PB when not further partitioned into multiple predictive blocks) is predicted in various ways rather than being directly coded, thereby exploiting various correlations and redundancies in the video data to improve compression efficiency. Correspondingly, such prediction may be performed in various modes. For example, a video block may be predicted via intra-prediction or inter-prediction. In particular, in an inter-prediction mode, a video block may be predicted by one or more other reference blocks or inter-predicted blocks from one or more other frames via either single-reference or mixed-reference inter-prediction. For inter-prediction implementations, a reference block may be specified by its frame identifier (the temporal location of the reference block) and a motion vector (the spatial location of the reference block) that indicates the spatial offset between the reference block and a current block being coded or decoded. The reference frame identification and the motion vector may be signaled in the bitstream. The motion vector, as a spatial block offset, may be signaled directly or may itself be predicted by another reference motion vector or a predicted motion vector. For example, the current motion vector may be predicted directly by a reference motion vector (e.g., of a candidate neighboring block), or by a combination of the reference motion vector and the motion vector difference (MVD) between the current motion vector and the reference motion vector. The latter may be called merge mode with motion vector difference (MMVD). The reference motion vector may be identified in the bitstream, for example, as a pointer to a spatial neighboring block of the current block or a temporal neighboring but spatially co-located block.
[0076] Returning to the intra-prediction process, samples within a block (e.g., a luma or chroma prediction block, or a coding block if not further divided into prediction blocks) are predicted by samples from a neighbor, next neighbor, or one or more other lines, or a combination thereof, to generate a prediction block. The residual between the actual block being coded and the prediction block may then be processed by a transform followed by quantization. Various intra-prediction modes may be made available, and parameters related to intra-mode selection and other parameters may be signaled in the bitstream. The various intra-prediction modes may be associated, for example, with one or more line positions for predicting samples, the direction in which prediction samples are selected from one or more prediction lines, and other special intra-prediction modes.
[0077] For example, the set of intra-prediction modes (interchangeably referred to as "intra modes") may include a predefined number of directional intra-prediction modes. As illustrated in the example implementation of Figure 7, these intra-prediction modes may correspond to a predefined number of directions in which an out-of-block sample is selected as a prediction for a sample predicted within a particular block. In another particular example implementation, eight main directional modes may be supported and predefined, corresponding to angles from 45 degrees to 207 degrees relative to the horizontal axis, as also shown in Figure 7.
[0078] In some other implementations of intra prediction, the directional intra modes can be further expanded to a finer set of angles to further exploit more types of spatial redundancy in directional textures. For example, the eight-angle implementation described above can be configured to provide eight nominal angles designated V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, as shown in FIG. 8, with a predetermined number (e.g., seven) of finer angles added for each nominal angle. Such expansion can result in a larger total number of directional angles (e.g., 56 in this example) available for intra prediction, corresponding to the same number of predefined directional intra modes. The prediction angle can be represented by the sum of the nominal intra angle and an angle delta. In the specific example described above with seven finer angle directions for each nominal angle, the angle delta can be -3 to 3 times the 3-degree step size.
[0079] As a further example, several other angle schemes for directional prediction may be used, as shown in FIG. 9, with 65 different prediction angles.
[0080] In some exemplary implementations, a predefined number of non-directional intra-prediction modes may be predefined and made available as an alternative or addition to the above-described directional intra-modes. For example, five non-directional intra-modes called smooth intra-prediction modes may be specified. These non-directional intra-mode prediction modes may be specifically called DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H intra-modes. Prediction of samples of a particular block under these exemplary non-directional modes is shown in FIG. 10. As an example, FIG. 10 shows a 4×4 block 1002 predicted by samples from an upper neighboring line and / or a left neighboring line. A particular sample 1010 in the block 1002 may correspond to a sample 1004 directly above the sample 1010 in the upper neighboring line of the block 1002, a sample 1006 above and to the left of the sample 2011 as the intersection of the upper and left neighboring lines, and a sample 1008 directly to the left of the sample 1010 in the left neighboring line of the block 1002. In an example of a DC intra prediction mode, the average of the left and above neighboring samples 1008 and 1004 can be used as the predictor of sample 1010. In an example of a PAETH intra prediction mode, the above, left, and above-left reference samples 1004, 1008, and 1006 can be fetched, and then the closest value (above + left - above-left) among these three reference samples can be set as the predictor of sample 1010. In an example of a SMOOTH_V intra prediction mode, sample 1010 can be predicted by quadratic interpolation in the vertical direction of the above-left neighboring sample 1006 and the left neighboring sample 1008. In an example of a SMOOTH_H intra prediction mode, sample 1010 can be predicted by quadratic interpolation in the horizontal direction of the above-left neighboring sample 2006 and the above neighboring sample 1004. In an example of a SMOOTH intra prediction mode, sample 1010 can be predicted by the average of quadratic interpolation in the vertical and horizontal directions. The above implementations of non-directional intra modes are merely illustrative and non-limiting examples. Other neighboring lines, other non-directional selections of samples, and other ways of combining prediction samples to predict a particular sample within a prediction block are also possible.
[0081] The encoder's selection of a particular intra-prediction mode from the above directional or non-directional modes at various coding levels (picture, slice, block, unit, etc.) can be signaled in the bitstream. In some exemplary implementations, eight exemplary nominal directional modes and five non-angle smooth modes (a total of 13 options) may be signaled first. Then, if the signaled mode is one of the eight nominal angle intra-modes, an index is further signaled indicating the selected angle delta relative to the corresponding signaled nominal angle. In some other exemplary implementations, all intra-prediction modes may be indexed together for signaling (e.g., 56 directional modes plus 5 non-directional modes to generate 61 intra-prediction modes).
[0082] In some exemplary implementations, the exemplary 56 or other number of directional intra-prediction modes may be implemented with a joint directional predictor that projects each sample of a block to a reference sub-sample position and interpolates the reference samples, for example, by a 2-tap bilinear filter.
[0083] In some embodiments, additional filter modes, called FILTER INTRA modes, may be designed to capture the decaying spatial correlation with edge references. In these modes, prediction samples within a block, in addition to out-of-block samples, may be used as intra-prediction reference samples for some patches within the block. These modes may, for example, be predefined and made available for intra-prediction of at least the luma block (or only the luma block). A predefined number (e.g., 5) of filter intra modes may be predesigned, each represented by a set of n-tap filters (e.g., 7-tap filters) that reflect the correlation between samples in a 4x2 patch and its n=7 neighbors. In other words, the weight coefficients of the n-tap filters may depend on the position. Taking an 8x8 block, a 4x2 patch, and 7-tap filtering as an example, as shown in FIG. 11, the 8x8 block 2002 may be divided into eight 4x2 patches. These patches are denoted by B0, B1, B1, B3, B4, B5, B6, and B7 in FIG. 11. For each patch, its seven neighboring samples, denoted R0-R6 in FIG. 11, can be used to predict the samples in the current patch. For patch B0, all neighboring samples may already be reconstructed and used to predict patch B0. However, for other patches, some of the neighboring samples may be within the current block and therefore not yet reconstructed, in which case their predicted values can be used as reference. For example, as shown in FIG. 11, all seven neighboring samples of patch B7 in the current block have not yet been reconstructed, so the predicted samples of these neighboring samples are used instead.
[0084] The above exemplary intra-prediction implementations, including the directional intra-prediction mode, various non-directional modes, and recursive filtering intra-prediction mode, generally rely on linear prediction, relying on estimating or predicting a particular pixel based on linear operations of other intra-reference samples, e.g., neighboring samples. In some situations, texture patterns within an image may be highly dynamic and may not necessarily exhibit linear relationships between neighboring pixels. Thus, intra-prediction based on non-linear relationships between neighboring samples may be designed to better capture such relationships and provide improved intra-prediction coding gain, as described in more detail below.
[0085] The following exemplary methods / implementations / embodiments may be used separately or combined in any order. Furthermore, each of the methods / implementations / embodiments, part or all of the decoder of the encoder, may be implemented by processing circuitry (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors may execute a program stored on a non-transitory computer-readable medium. Hereinafter, the term "block" may refer to a prediction block, a coding block, or a coding unit, i.e., a CU.
[0086] In some embodiments, intra-prediction samples may be derived using a predefined function that uses horizontal and vertical coordinate values and / or surrounding sample values as input, and the function parameters (or model parameters) may be derived using neighboring reconstructed samples or explicitly signaled. In other words, linear and / or nonlinear relationships between neighboring samples that reflect texture information may be parameterized with a predefined model. Such functions may alternatively be referred to as prediction functions or prediction models. The term "predefined function" may be used to refer to the form of the function. In other words, the form of the prediction function may be predefined as shown in the following example:
[0087] In such an intra prediction method, the predictor of a sample depends on its pixel coordinate within its coding block, so it is similar in some sense to warped inter prediction, in which an affine model can be used to transform a block into a warped reference block (by spatial warping) so that a predictor sample in a reference frame for a sample being predicted in a current block depends on the pixel coordinate of the sample being predicted. Due to this pixel coordinate dependency of the predictor sample, intra prediction as described above and in further detail below may be referred to as intra warping. The intra prediction mode for such a prediction method may correspondingly be referred to as an intra warp mode. The intra warp mode may be considered as an additional intra prediction mode to the other intra prediction modes mentioned above, such as a directional intra prediction mode, a non-directional intra prediction mode, a filtered intra prediction mode, etc. In some implementations, such a warped intra prediction mode may be parallel to the non-warped intra prediction mode in which the other intra prediction modes mentioned above exist. Correspondingly, signaling may be provided in the bitstream by the encoder to indicate to the decoder which of multiple intra-prediction modes are used at a particular coding level. Alternatively, signaling may be provided in the bitstream by the encoder to indicate to the decoder whether a warped intra-prediction mode is used at a particular coding level, followed by additional signaling as to which of the other non-warped intra-prediction modes are used at a particular coding level if not coded under the warped intra-prediction mode.
[0088] In one example, the predefined prediction function may be a polynomial function with degree up to N.
[0089] In one example, the prediction function may have model parameters a, b, c, d, e, and f and may be of the following form:
number
[0090] In one example, the prediction function may have model parameters a, b, d, e, f and may be of the form:
number
[0091] In one example, the prediction function may have model parameters a, b, d, and e and may be of the form:
number
[0092] In one example, the prediction function is a function of the model parameters a, b, d, e, and w TL and may be in the following form:
number
[0093] In one example, the prediction function is the model parameters a, b, d, e, w T , and w L and may be in the following form:
number
[0094] In one example, the prediction function is the model parameters a, b, d, e, w T , w L , w TR , and w BL and may be in the following form:
number
[0095] In the above example, the degree N of the above exemplary polynomial can be 1 (linear), 2 (degree 2), 3, 4, ... Additional model parameters can be introduced.
[0096] In some other exemplary implementations, rather than a polynomial function, the predefined prediction function may be in the form of a weighted sum of multiple triangular basis functions with up to N bases (each triangular function is referred to as a basis function or basis). The triangular basis functions or bases may, for example, be orthogonal to one another. Each triangular basis function may again be, for example, a function of (x, y), which are sample coordinates relative to the upper-left corner of the block. The triangular basis function may, for example, be a 2D triangular basis function or a combination of 2D triangular basis functions.
[0097] In one example, the basis can include the first N Discrete Cosine Transform Type 2 (DCT-2) bases. In one example, the basis can include the first N Discrete Sine Transform Type 7 (DST-7) bases. In one example, the basis can be derived from any of the bases of DCT Types 1-8 and DST Types 1-8.
[0098] In some other embodiments, the predefined prediction function may be in the form of a weighted sum of multiple Karhunen Loeve Transform (KLT) bases, with up to N KLT bases. The KLT bases may be, for example, eigenvectors of a covariance matrix derived from neighboring reconstructed samples. The KLT bases may be, for example, predefined eigenvectors known to both the encoder and the decoder. Each KLT basis may again be a function of (x, y), for example, the sample coordinates relative to the top-left corner of the block.
[0099] In some other embodiments, the predefined prediction function may be in the form of a mixture of the polynomial and trigonometric functions described above, and the number of terms with corresponding model parameters may be predefined.
[0100] The various model parameters described above may be predefined or signaled. In some embodiments, the model parameters may be predefined as an indexed set. The selection of a particular block or set of indexes may be made by the encoder at various levels (picture, frame, slice, macroblock, block, etc.) in real time and signaled in the bitstream, for example. In some embodiments, the model parameters may be derived by the encoder in real time and signaled in the bitstream at various levels.
[0101] In some embodiments, model parameters used in predefined prediction functions may be clipped to a predefined range (to limit the number of bits representing the values of these parameters). Model parameters with clipped ranges help reduce computational load. Furthermore, in situations where model parameters are signaled, clipping the parameter value range helps further reduce the number of bits for signaling the parameters. In some embodiments, the parameter value range may be signaled rather than predefined.
[0102] In some embodiments, the predicted values produced by a predefined prediction function may be clipped to a predefined range to reduce the computational load at the encoder and / or decoder.
[0103] In the above embodiments, the clipped minimum and maximum values of the predicted value of the predefined function may depend on the range of the internal bit depth. For example, if the internal bit depth is 8 (or 10) for calculation purposes, the minimum and maximum range may be from 0 to 255 (1023). In some embodiments, the clipped minimum and maximum values or range may be signaled in a high-level syntax.
[0104] In some embodiments, the model parameters used in the predefined prediction functions may be quantized to a particular precision, such as 1 (an integer), 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, 1 / 128, 1 / 256, etc.
[0105] In some embodiments, different model parameters of a predefined prediction function may be quantized in different ways. For example, in the example prediction function of equation (1) above, parameters a, b, and c may be quantized to a precision of 1 / 16, and parameters d, e, and f may be quantized to a precision of 1 / 4. Thus, this prediction function may be implemented as:
number
[0106] In some embodiments, the following weighted terms of the above prediction function associated with the same precision may be calculated together and rounded together:
number
[0107] In some embodiments, the precision of model parameters may be signaled in a high-level syntax and at various levels: a set of available precisions may be predefined, and the selection of precision for a parameter or group of parameters may be signaled by an index within the predefined set of available precisions.
[0108] In some embodiments, the coefficients of the above functions may be derived at both the encoder and decoder sides according to a predefined or signaled algorithm or optimization procedure. The derivation may be based on using single-line or multi-line templates above and to the left of the current block containing the reconstructed samples. The templates representing the selected samples for the derivation of the optimization may be predefined or signaled. For example, a set of templates may be predefined, and the templates to be used may be signaled by their index within the set of templates at various coding levels in the bitstream.
[0109] In some embodiments, the current block may be divided into multiple sub-blocks (with pixel width / height of a predetermined value, e.g., 2 or more), and the above model parameters or coefficients may be calculated / derived for each sub-block using available surrounding information of the current sub-block.
[0110] In some embodiments, the coefficients of the above model prediction function may be explicitly signaled in the bitstream and parsed at the decoder side to perform the reconstruction.
[0111] FIG. 12 shows a flowchart of an example method 1200 for decoding a block in a video stream. Method 1200 begins at S1201. In step S1210, it is determined based on a syntax element in the video bitstream that intra-warp prediction is applied to the block. In step S1220, a set of weighting factors to be applied to the set of basis functions used in the intra-warp prediction is determined. In step S1230, values of samples of the block are predicted based on a weighted sum of the set of basis functions at the pixel coordinate locations of the samples, where the set of basis functions are each weighted according to the set of weighting factors. In step S1240, a block containing the samples is reconstructed based on the predicted values. Method flow 1200 ends at S1299.
[0112] The above-described operations may be combined or arranged in any quantity or order as desired. Two or more steps and / or operations may be performed in parallel. The embodiments and implementations of the present disclosure may be used separately or combined in any order. Furthermore, each of the method (or embodiment), encoder, and decoder may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium. The embodiments of the present disclosure may apply to luma blocks or chroma blocks. The term "block" may refer to a prediction block, a coding block, or a coding unit, i.e., a CU. The term "block" herein may also be used to refer to a transform block. In the following sections, references to block size may refer to either the width or height of the block, the maximum value of the width and height, the minimum value of the width and height, the area size of the block (width x height), or the aspect ratio (width:height or height:width).
[0113] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 13 illustrates a computer system (1300) suitable for implementing certain embodiments of the subject matter of this disclosure.
[0114] Computer software can be coded using any suitable machine code or computer language that can be processed by mechanisms such as assembly, compilation, linking, etc. to generate code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly or through interpretation, microcode execution, etc.
[0115] The instructions may be executed by a variety of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0116] 13 of the computer system (1300) are exemplary in nature and do not suggest any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Furthermore, the arrangement of components should not be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of the computer system (1300).
[0117] The computer system (1300) may include certain human interface input devices, which may include one or more of a keyboard (1301), a mouse (1302), a trackpad (1303), a touchscreen (1310), a data glove (not shown), a joystick (1305), a microphone (1306), a scanner (1307), and a camera (1308) (only one of which is shown).
[0118] The computer system (1300) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses through, for example, sensory output, sound, light, and smell / taste. Such human interface output devices may include sensory output devices (e.g., sensory feedback via a touchscreen (1310), a data grab (not shown), or a joystick (1305; however, sensory feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers (1309), headphones (not shown)), and visual output devices (e.g., a screen (1310), including a CRT screen, an LCD screen, a plasma screen, and an OLED screen, each with or without touchscreen input capability and each with or without sensory feedback capability, some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher-dimensional output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown)).
[0119] The computer system (1300) may also include human-accessible storage devices and associated media such as optical media including CD / DVD ROM / RW (1320) with media such as CD / DVD (1321), thumb drives (1322), removable hard drives or solid state drives (1323), legacy magnetic media such as tape and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), etc.
[0120] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0121] The computer system 1300 may also include interfaces 1354 to one or more communications networks 1355. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CAN Bus, etc.
[0122] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to the core (1340) of the computer system (1300).
[0123] The core (1340) may include one or more central processing units (CPUs) (1341), graphics processing units (GPUs) (1342), dedicated programmable processing units (1343) in the form of FPGAs, task-specific hardware accelerators (1344), graphics adapters (1350), etc. These devices may be connected through a system bus (1348), along with read-only memory (ROM) (1345), random access memory (1346), and internal mass storage devices (1347) such as internal non-user-accessible hard drives, SSDs, etc. In some computer systems, the system bus (1348) is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core's system bus (1348) or through a peripheral bus (1349). In an example, a screen (1310) can be connected to the graphics adapter (1350). Peripheral bus architectures include PCI, USB, and the like.
[0124] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0125] While this disclosure has described several exemplary embodiments, alterations, permutations, and various substitute equivalents exist, and are encompassed within the scope of this disclosure. Those skilled in the art will appreciate that numerous systems and methods can be devised that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore are within the spirit and scope of the present disclosure.
Claims
1. 1. A method for decoding a block in a video bitstream in a decoder, the method comprising: determining, based on a syntax element in the video bitstream, that intra-warp prediction is applied to the block; determining a set of weighting factors to be applied to a set of basis functions used in the intra-warp prediction; predicting values of the samples of the block based on a weighted sum of the set of basis functions on pixel coordinate locations of the samples, the set of basis functions being each weighted according to the set of weighting coefficients; reconstructing a block containing said samples based on said predicted values; A method comprising:
2. the set of basis functions includes power functions of the pixel coordinate locations of the samples up to order N, where N is an integer; The method of claim 1 , wherein the weighted sum of basis functions comprises a polynomial of the pixel coordinate locations of the samples up to order N.
3. N=2, and the set of weighting factors is Six coefficients of a two-dimensional polynomial term of the horizontal and vertical pixel coordinates of the sample; five coefficients of a two-dimensional polynomial term in the horizontal and vertical pixel coordinates of the sample, excluding one cross term; or four coefficients of a two-dimensional polynomial term that depend on the horizontal and vertical pixel coordinates of the sample, excluding one cross term, and at least one further coordinate-independent term; The method of claim 2 , comprising:
4. The method of claim 3 , wherein the at least one coordinate-independent term includes an offset term that includes a DC value obtained from a reconstruction neighborhood of the sample.
5. the set of weighting coefficients further comprises an offset coefficient, and the at least one coordinate independent term comprises an offset term that is a product of the offset coefficient and a reconstructed value of a sample in an upper left corner of the block; or the set of weighting coefficients further includes a first offset coefficient and a second offset coefficient, and the at least one coordinate independent term includes a first offset term and a second offset term, the first offset term being a product of the first offset coefficient and an above reconstructed neighborhood of the sample, and the second offset term being a product of the second offset coefficient and a left reconstructed neighborhood of the sample; or the set of weighting coefficients further includes first, second, third, and fourth offset coefficients, and the at least one coordinate independent term includes first, second, third, and fourth offset terms, the first offset term being a product of the first offset coefficient and an above reconstructed neighborhood of the sample, the second offset term being a product of the second offset coefficient and a left reconstructed neighborhood of the sample, the third offset term being a product of the third offset coefficient and an above-right reconstructed neighborhood of the sample, and the fourth offset term being a product of the fourth offset coefficient and a below-left reconstructed neighborhood of the sample. The method of claim 3.
6. N=2, and the weighted sum of the basis functions comprises: [Equation 1] 3. The method of claim 2, wherein a, b, c, d, e, and f represent polynomial coefficients, (x, y) represent sample coordinates, and p(x, y) represents the weighted sum.
7. The set of basis functions includes up to N trigonometric basis functions for the pixel coordinate locations of the samples, where N is an integer, and the N trigonometric basis functions are the first N type-2 discrete cosine transform bases, the first N type-7 discrete sine transform bases, or Any N discrete sine transform bases of types 1 to 8 and discrete cosine transform bases of types 1 to 8, The method according to any one of claims 1 to 6, comprising:
8. 7. The method of claim 1, wherein the set of basis functions comprises up to N Karhunen-Loeve Transform (KLT) bases as a function of the pixel coordinate location of the sample, where N is an integer.
9. 9. The method of claim 8, wherein the N KLT bases comprise eigenvectors of a covariance matrix derived from a reconstruction neighborhood of the sample, or are predefined eigenvectors known to both the encoder and the decoder.
10. The method of any one of claims 1 to 6, wherein the set of basis functions comprises a subset of power functions and a subset of trigonometric functions of pixel coordinate positions of the samples.
11. The method of claim 1 , further comprising parsing the video bitstream to determine the set of weighting factors signaled in the video bitstream.
12. The method of any one of claims 1 to 6, wherein the set of weighting factors is derived by an encoder using a multi-line template above and to the left of the block.
13. Parsing the video bitstream to determine the weighting factors comprises: parsing the video bitstream to obtain an index of the set of weighting factors from the video bitstream; identifying a set of weight confidences from among a plurality of sets of weight coefficients according to the index; The method according to any one of claims 1 to 6, comprising:
14. The method of any one of claims 1 to 6, wherein the set of weighted beliefs is clipped to a predetermined range of values.
15. 7. The method according to claim 1, wherein the predicted values of the samples are clipped to a predetermined range of values, the minimum and maximum clipped values of the predicted values being determined by an internal bit depth range.
16. The method according to any one of claims 1 to 6, wherein the set of weighting factors is quantized according to one or more predetermined precisions.
17. The method of claim 16 , wherein at least two of the sets of weighting factors are quantized to different precisions.
18. 17. The method of claim 16, wherein terms in a weighted sum of the set of basis functions associated with the same precision are calculated and rounded together.
19. The method of claim 16 , wherein the precision of the set of weighting factors is signaled in a high-level syntax in the video bitstream.
20. 1. An electronic device comprising: a memory for storing instructions; and a processor, the processor executing the instructions to: determining, based on syntax elements in the video stream, that intra-warp prediction is to be applied to a block in the video stream; determining a set of weighting factors to be applied to a set of basis functions used in the intra-warp prediction; predicting values of the samples based on a weighted sum of the set of basis functions on pixel coordinate locations of the samples of the block, the set of basis functions each being weighted according to the set of weighting coefficients; reconstructing a block containing said samples based on said predicted values; An electronic device that performs a
21. 1. A method for encoding a block in a video bitstream in an encoder, the method comprising: determining, based on a syntax element in the video bitstream, that intra-warp prediction is applied to the block; determining a set of weighting factors to be applied to a set of basis functions used in the intra-warp prediction; predicting values of the samples of the block based on a weighted sum of the set of basis functions on pixel coordinate locations of the samples, the set of basis functions being each weighted according to the set of weighting coefficients; coding a block containing said samples into a bitstream based on said predicted values; A method comprising:
22. 1. A method for encoding a block in a video bitstream in an encoder, the method comprising: determining, based on a syntax element in the video bitstream, that intra-warp prediction is applied to the block; determining a set of weighting factors to be applied to a set of basis functions used in the intra-warp prediction; predicting values of the samples of the block based on a weighted sum of the set of basis functions on pixel coordinate locations of the samples, the set of basis functions being each weighted according to the set of weighting coefficients; coding a block containing said samples into a bitstream based on said predicted values; transmitting the bitstream; A method comprising:
Citation Information
Patent Citations
Intra-prediction apparatus, encoder, decoder, and program
JP2011172185A
luminance-based coding tool for video compression
JP2017511045A
Intra prediction using polynomial model
US20210297670A1
Smooth surface prediction
WO2022211715A1