System and method for intra prediction mode based on adaptive extrapolation filter
By selecting different EIP parameter sets according to boundary conditions in video encoding, the problem of large signaling overhead in the prior art is solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202480004299.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2024-04-22
- Publication Date
- 2025-05-30
AI Technical Summary
When the existing video encoding technology uses intra prediction mode based on extrapolation filters, it is difficult to effectively manage boundary conditions, resulting in an increase in signaling overhead.
By determining the EIP mode parameters from the first EIP parameter set when the current block satisfies the boundary condition, and when the boundary condition is not met, the parameters are determined from the second EIP parameter set, the second EIP parameter set includes parameters not in the first set to reduce signaling overhead.
It effectively reduces signaling overhead and improves the efficiency of video encoding, especially when dealing with scenes with complex boundary conditions.
Smart Images

Figure CN120077654A_ABST
Abstract
Description
Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 461,547, entitled "Adaptive Extrapolation Filter Based Intra Prediction Mode," filed on Apr. 24, 2023, and this application is a continuation of and claims priority to U.S. Patent Application No. 18 / 641,215, entitled "Systems and Methods for Adaptive Extrapolation Filter Based Intra Prediction Mode," filed on Apr. 19, 2024. Technical Field
[0002] The disclosed embodiments generally relate to video coding, including but not limited to systems and methods for implementing an Extrapolation Filter-Based Intra Prediction (EIP) mode. Background Art
[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital imaging devices, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise convey digital video data over a communication network and / or store the digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited memory resources of the storage device, video coding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data. Video coding can be performed by a server providing cloud services or by hardware and / or software on an electronic / client device.
[0004] Video coding typically uses prediction methods that exploit the redundancy inherent in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. A variety of video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended as a successor to HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verified version 1.0.0 of the specification with errata 1 was released. Summary of the Invention
[0005] Among other things, the present disclosure describes a set of methods for video (image) compression, more specifically related to an intra-frame prediction mode based on an extrapolation filter (or "extrapolation filter intra-frame prediction mode"). Intra-frame prediction based on an extrapolation filter can be processed in two steps. First, extrapolation filter coefficients are obtained from adjacent reconstructed pixels of the current block using a predetermined template. Second, predictions are generated position-by-position from the top left to the bottom right within the current block. Embodiments described herein include restricting EIP parameters when the current block meets boundary conditions. The advantage of determining whether the current block meets boundary conditions when using the intra-frame prediction mode based on an extrapolation filter is that it can reduce signaling overhead (e.g., by signaling fewer indicators and / or signaling a smaller set).
[0006] According to some embodiments, a method for video decoding includes (i) receiving a video bitstream including a plurality of blocks; when the EIP mode is valid for a current block among the plurality of blocks and a boundary condition is satisfied for the current block, (ii) determining one or more EIP mode parameters from a first EIP parameter set; when the boundary condition is not satisfied for the current block, (iii) determining one or more EIP mode parameters from a second EIP parameter set, where the second EIP parameter set includes one or more parameters not included in the first EIP parameter set; and (iv) reconstructing the current block using the one or more EIP mode parameters.
[0007] According to some embodiments, a method for video encoding includes (i) receiving video data including a plurality of video blocks; when the EIP mode is valid for a current block among the plurality of blocks and a boundary condition is satisfied for the current block, (ii) determining one or more EIP mode parameters from a first EIP parameter set; when the boundary condition is not satisfied for the current block, (iii) determining one or more EIP mode parameters from a second EIP parameter set, where the second EIP parameter set includes one or more parameters not included in the first EIP parameter set; and (iv) encoding the current block using the one or more EIP mode parameters.
[0008] According to some embodiments, a method for processing visual media data includes: (i) obtaining a source video sequence including a plurality of frames; and (ii) performing a conversion between the source video sequence and a video bitstream of the visual media data, where the bitstream includes: (a) a plurality of encoded blocks corresponding to a plurality of video blocks, the plurality of encoded blocks including a first block encoded using the EIP mode; when a boundary condition is satisfied for the first block, (b) an indication of a first EIP mode parameter set for the first block; and when the boundary condition is not satisfied for the first block, (iii) an indication of a second EIP mode parameter set for the first block, where the second EIP parameter set includes one or more parameters not included in the first EIP parameter set.
[0009] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic devices. The computing system includes control circuitry and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder). According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets for execution by the computing system. The one or more instruction sets include instructions for performing any of the methods described herein.
[0010] Accordingly, apparatuses, systems, and methods for encoding and decoding video are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding / decoding. The features and advantages described in the specification are not necessarily exhaustive, and in particular, given the figures, specification, and claims provided in the present disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in this specification has been selected primarily for readability and guidance purposes, and not for the purpose of depicting or limiting the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] For a more detailed understanding of the present disclosure, a more specific description may be made by referring to the features of various embodiments, some of which are shown in the drawings. However, the drawings only show relevant features of the present disclosure and are therefore not necessarily considered restrictive, as those skilled in the art will understand, after reading the present disclosure, that the description may permit other effective features.
[0012] Figure 1 is a block diagram showing an example communication system according to some embodiments.
[0013] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.
[0014] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.
[0015] Figure 3 is a block diagram showing an example server system according to some embodiments.
[0016] Figure 4A , Figure 4B and Figure 4C show aspects of an intra prediction mode based on an extrapolation filter.
[0017] Figure 5A , Figure 5B and Figure 5C show applications of an intra prediction mode based on an extrapolation filter according to some embodiments.
[0018] Figure 6A shows an example video decoding process according to some embodiments.
[0019] Figure 6B shows an example video encoding process according to some embodiments.
[0020] By convention, the various features shown in the drawings are not necessarily drawn to scale, and throughout the specification and drawings, the same reference numerals may be used to denote the same features. Detailed Description
[0021] The present disclosure describes video / image compression techniques related to an extrapolation filter-based intra prediction mode. Some embodiments include determining whether a current block satisfies one or more boundary conditions, and restricting EIP mode parameters / settings when the current block satisfies one or more boundary conditions. For example, when the current block is at an upper boundary and upper reconstructed samples are not available for EIP mode calculation, the EIP reconstruction region option is restricted to reduce / minimize options that rely on upper reconstructed samples. The advantage of determining whether the current block satisfies boundary conditions when using the EIP mode is that signaling overhead can be reduced (e.g., by signaling fewer indicators and / or signaling a smaller set). Example systems and devices
[0022] Figure 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0023] The source device 102 includes a video source 104 (e.g., a camera device component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera device (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams based on the video stream. The video stream from the video source 104 may be of high data volume compared to the encoded video bitstream generated by the encoder component 106. Since the encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video to the network 110).
[0024] One or more networks 110 represent any number of networks that convey information between source device 102, server system 112, and / or electronic devices 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0025] One or more networks 110 include server system 112 (e.g., a distributed / cloud computing system). In some embodiments, server system 112 is a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from source device 102) or includes such a streaming server. Server system 112 includes codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, codec component 114 includes an encoder component and / or a decoder component. In various embodiments, codec component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, codec component 114 is configured to decode encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, server system 112 is configured to generate multiple video formats and / or encodings based on encoded video bitstream 108. In some embodiments, server system 112 functions as a Media-Aware Network Element (MANE). For example, server system 112 may be configured to trim encoded video bitstream 108 to customize potentially different bitstreams for one or more of electronic devices 120. In some embodiments, the MANE is provided separately from server system 112.
[0026] Electronic device 120-1 includes decoder component 122 and display 124. In some embodiments, decoder component 122 is configured to decode encoded video data 116 to generate an outgoing video stream that can be presented on a display or other type of rendering device. In some embodiments, one or more of electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media memory). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.
[0027] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, one or more of the electronic devices 120 and / or the source device 102 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0028] In an example operation of the communication system 100, the source device 102 transmits an encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using the codec component 114. For example, the server system 112 may apply an encoding that is more optimal for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.
[0029] Figure 2A is a block diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the device of the encoder component 106). The video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 YCrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that produce motion when viewed in sequence. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can readily understand the relationship between pixels and samples.
[0030] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder component 106 is configured to perform a conversion between a source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skipping, quantizer, and / or λ value of rate distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may relate to optimizing the encoder component 106 for a particular system design.
[0031] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, e.g., a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder 210. (In the case where the compression between the symbols and the encoded video bitstream is lossless) the decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data. The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the same sample values as the sample values that the decoder will interpret during decoding when using prediction as reference picture samples.
[0032] The operation of the decoder 210 can be the same as that of a remote decoder, e.g., the decoder component 122 described in detail below in connection with Figure 2B However, briefly referring to Figure 2B , since the symbols are available and the symbols are encoded into an encoded video sequence by the entropy encoder 214 and the decoding of the symbols by the parser 254 can be lossless, the entropy decoding part including the buffer memory 252 and the parser 254 of the decoder component 122 may not be fully implemented in the local decoder 210.
[0033] Except for parsing / entropy decoding, the decoder techniques described herein may exist in corresponding encoders in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. Additionally, the description of encoder techniques may be simplified as they may be the opposite of decoder techniques.
[0034] As part of its operation, the source encoder 202 may perform motion-compensated predictive coding that predicts an input frame with reference to one or more previously encoded frames designated as reference frames from a video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the reference frame, which may be selected as a prediction reference for the input frame. The controller 204 may manage the encoding operation of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.
[0035] The decoder 210 decodes the encoded video data of frames that may be designated as reference frames based on the symbols created by the source encoder 202. The operation of the encoding engine 212 may advantageously be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 2A not shown), the reconstructed video sequence may be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that may be performed by a remote video decoder on a reference frame and may cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 stores a copy of the reconstructed reference frame locally, which has common content (in the absence of transmission errors) with the reconstructed reference frame that would be obtained by a remote video decoder.
[0036] The predictor 206 may perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that may be used as an appropriate prediction reference for the new picture. The predictor 206 may operate on a per-sample block, pixel block basis to find an appropriate prediction reference. As determined by the search results obtained by the predictor 206, the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory 208.
[0037] The outputs of all of the above-mentioned functional units may undergo entropy coding in the entropy encoder 214. The entropy encoder 214 converts these symbols into an encoded video sequence by losslessly compressing the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).
[0038] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence created by the entropy encoder 214 in preparation for transmission via a communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may transmit additional data along with the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set segments, and the like.
[0039] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a specific encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). Intra pictures may be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are familiar with those variations of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. Predictive pictures may be encoded and decoded using inter prediction or intra prediction that uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures may be encoded and decoded using inter prediction or intra prediction that uses at most two motion vectors and a reference index to predict the sample values of each block. Similarly, multi-predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0040] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and encoded on a block-by-block basis. Predictive encoding of a block can be made with reference to other (already encoded) blocks, which are determined by the encoding assignment for the corresponding picture to which the block belongs. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be non-predictively encoded, or can be encoded via spatial prediction or via temporal prediction (with reference to a previously decoded reference picture). Blocks of a B picture can be non-predictively encoded, or can be encoded via spatial prediction or via temporal prediction (with reference to one or two previously decoded reference pictures).
[0041] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (commonly abbreviated to intra prediction) exploits spatial correlations within a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an example, a particular picture in encoding / decoding, which is referred to as the current picture, is segmented into blocks. In a case where a block in the current picture is similar to a reference block in a reference picture that has been previously decoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0042] The encoder component 106 can perform encoding operations according to a predetermined video coding technique or standard (such as any of the video coding techniques or standards described herein). In the operation of the encoder component 106, the encoder component 106 can perform various compression operations, including predictive encoding operations that utilize temporal redundancy and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.
[0043] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to the loop filter 256 and is configured to transmit data to the display 124 (e.g., via a wired or wireless connection).
[0044] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. Decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0045] According to some embodiments, decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. Decoder component 122 may be implemented at least in part in software.
[0046] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to counter network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to counter network jitter). When receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be needed, or buffer memory 252 may be small. For use on a best-effort packet network such as the Internet, buffer memory 252 may be needed, which may be relatively large and / or have an adaptive size, and may be implemented at least in part in the operating system or a similar element external to decoder component 122.
[0047] The parser 254 is configured to reconstruct symbols 270 from an encoded video sequence. The symbols may include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The decoding of the encoded video sequence may be performed according to a video decoding technique or standard and may follow principles well known to those skilled in the art, including: variable length decoding, Huffman coding, arithmetic decoding with or without context sensitivity, etc. The parser 254 may extract a subgroup parameter set for at least one of the subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroups may include a group of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser 254 may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0048] Depending on the type of the encoded video picture or a portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbol 270 may involve multiple different units. Which units are involved and the manner of involvement may be controlled by the subgroup control information parsed by the parser 254 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 254 and the multiple units below are not depicted.
[0049] The decoder component 122 may be conceptually subdivided into multiple functional units, and in some embodiments, these units interact closely with each other and may be at least partially integrated with each other. However, for the sake of clarity, the conceptual subdivision of the functional units is maintained herein.
[0050] The scaler / inverse transform unit 258 receives, from the parser 254, the quantized transform coefficients as the symbol 270 and control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix). The scaler / inverse transform unit 258 may output a block including sample values, which may be input into the aggregator 268.
[0051] In some cases, the output samples of the scaler / inverse transform unit 258 belong to an intra-coded block; that is: a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed part of the current picture. Such predictive information may be provided by the intra picture prediction unit 262. The intra picture prediction unit 262 may use the surrounding already reconstructed information obtained from the current (partially reconstructed) picture in the current picture memory 264 to generate a block of the same size and shape as the block being reconstructed. The aggregator 268 may add the prediction information already generated by the intra picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 on a per sample basis.
[0052] In other cases, the output samples of the scaler / inverse transform unit 258 belong to an inter-coded and possibly motion-compensated block. In such a case, the motion compensation prediction unit 260 may access the reference picture memory 266 to obtain samples for prediction. After motion compensating the obtained samples according to the sign 270 belonging to the block, these samples may be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signal in this case) to generate output sample information. The address in the reference picture memory 266 from which the motion compensation prediction unit 260 obtains the prediction samples may be controlled by a motion vector. The motion vector may be available to the motion compensation prediction unit 260 in the form of the sign 270, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include, for example, interpolation of sample values obtained from the reference picture memory 266 when using sub-sampled accurate motion vectors, i.e., a motion vector prediction mechanism.
[0053] The output samples of the aggregator 268 may be subject to various loop filtering techniques in the loop filter unit 256. The video compression technique may include in-loop filter techniques, which are controlled by parameters included in the coded video bitstream and available to the loop filter unit 256 as the sign 270 from the parser 254, but the video compression technique may also respond to meta-information obtained during the decoding of previous (in decoding order) parts of the coded picture or coded video sequence, and to the sample values of the previously reconstructed and loop-filtered samples. The output of the loop filter unit 256 may be a sample stream, which may be output to a rendering device such as the display 124 and stored in the reference picture memory 266 for future inter-picture prediction.
[0054] Once reconstructed, certain coded pictures can be used as reference pictures for future prediction. Once a coded picture has been reconstructed and the coded picture (by, for example, parser 254) has been identified as a reference picture, the current reference picture can become part of the reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent coded pictures.
[0055] The decoder component 122 can perform decoding operations according to a predetermined video compression technique that can be recorded in any standard such as the standards described herein. As specified in a video compression technique document or standard and in particular in a profile therein, an encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard. Additionally, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within a range defined by the levels of the video compression technique or standard. In some cases, the levels limit the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the levels can be further restricted by the Hypothetical Reference Decoder (HRD) specification and metadata signaled in the encoded video sequence for HRD buffer management.
[0056] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes one or more processors (e.g., a CPU, GPU, and / or DPU). In some embodiments, the control circuitry includes a field programmable gate array, a hardware accelerator, and / or an integrated circuit (e.g., an application specific integrated circuit).
[0057] The network interface 304 can be configured to interface with one or more communication networks (e.g., wireless, wired, and / or optical networks). The communication networks can be local, wide area, metropolitan area, vehicular, and industrial, real-time, delay tolerant, etc. Examples of communication networks include: local area networks such as Ethernet; wireless LAN (Local Area Network); cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus, etc. Such communication can be one-way reception only (e.g., broadcast TV), one-way transmission only (e.g., CAN bus to certain CAN bus devices), or two-way (e.g., using local or wide area digital networks to other computer systems). Such communication can include communication to one or more cloud computing networks.
[0058] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 can include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera device, etc. The output device 308 can include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., monitors or display screens), etc.
[0059] The memory 314 can include high-speed random access memory (e.g., DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices remote from the control circuitry 302. The memory 314, or alternatively, the non-volatile solid-state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · An operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks; · A network communication module 318 for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); · A codec module 320 for performing various functions related to encoding and / or decoding data such as video data. In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: ο A decoding module 322 for performing various functions related to decoding encoded data, such as those previously described for the decoder component 122; and ο An encoding module 340 for performing various functions related to encoding data, such as those previously described for the encoder component 106; and ● A picture memory 352 for storing pictures and picture data, e.g., for use with the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.
[0060] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described for the parser 254), a transform module 326 (e.g., configured to perform various functions previously described for the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions previously described for the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described for the loop filter 256).
[0061] In some embodiments, the encoding module 340 includes an encoding module 342 (e.g., configured to perform various functions previously described for the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described for the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the modules shown. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.
[0062] Each of the foregoing modules stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The foregoing modules (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the transcoding module 320 optionally does not include separate decoding and encoding modules, but instead uses the same set of modules to perform two sets of functions. In some embodiments, the memory 314 stores a subset of the foregoing modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above.
[0063] Although Figure 3 FIG. 112 shows a server system 112 according to some embodiments, Figure 3 it is intended more as a functional description of the various features that may exist in one or more server systems than as a structural schematic of the embodiments described herein. In fact, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in FIG. 112 may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among them will vary according to the implementation, and optionally, it depends in part on the amount of data traffic processed by the server system during peak usage periods as well as during average usage periods. Example encoding and decoding techniques
[0064] The transcoding processes and techniques described below may be performed at the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). Hereinafter, a transform may refer to a primary transform (e.g., Multiple Transform Selection (MTS) or Non-Separable Primary Transform (NSPT)) or a secondary transform (e.g., Non-Separable Secondary Transform (NSST) or Low Frequency Non-Separable Transform (LFNST)).
[0065] Transform coding can be applied to the prediction residuals to remove potential spatial correlations. Some examples of transform kernels include type 2DCT (Discrete Cosine Transform-2, DCT-2), type 7DST (DST-7), and type 8DCT (DCT-8). In cases where the residuals have a non-uniform distribution, DST-7 and DCT-8 can be more efficient than DCT-2 because the basis functions of DST-7 and DCT-8 can better match such statistics. Therefore, due to the diverse nature of image or video content, the coding efficiency can be improved by not using a single transform kernel for all prediction residuals.
[0066] As described below, the EIP mode can be used to decode the current block. The following description is based on decoding blocks in a frame, but similarly applies to encoding one or more blocks using the EIP mode. When decoding the current block using the EIP mode (e.g., optionally signaling an indicator in the bitstream to indicate that the bitstream contains information encoded using the EIP mode (e.g., an encoded block)), the intra prediction mode used to decode the current block (e.g., Figure 4A one or more of the intra prediction modes shown in Figure 4B and Figure 4C and Figure 5A and Figure 5B and / or other directional prediction mode information) is not signaled in the video bitstream. Decoding the current block using the EIP mode can involve using a two-step process to select the transform kernel for the current block. For example, the transform kernel can be selected from a set of transforms related to the intra prediction mode, where each set has multiple candidates for the transform kernel. The extrapolation filter coefficients can be obtained using a predetermined template from the reconstructed region adjacent to the current block (e.g., containing the reconstructed pixels that are adjacent pixels of the current block), as described below with reference to Figure 5A and
[0067] Next, as shown in Figure 4B and Figure 4C the predicted value can be generated by performing a position-by-position (e.g., sample-by-sample) extrapolation from the top left to the bottom right within the current block. In some embodiments, the average value is removed when feeding the input to the EIP filter. For example, the value of the DC mode for the current block can be used as the average value for EIP prediction. The minimum and maximum values can be searched for among the reconstructed pixels in the reconstructed region (e.g., consisting of thirteen columns and thirteen rows). Figure 4BShows three different reconstruction regions 404, 408, and 410 according to some embodiments. The reconstruction region 404 is a first type of reconstruction region in an L shape and is adjacent to the top edge and the left edge of the current block 406 (e.g., prediction unit). The reconstruction region 408 is a second type of reconstruction region in a rectangular shape, having a width greater than its height (e.g., 3 columns by 8 rows), and is adjacent to the top edge of the current block 406. The reconstruction region 410 is a third type of reconstruction region in a rectangular shape, having a height greater than its width (e.g., 8 columns by 3 rows), and is adjacent to the left edge of the current block 406. In some embodiments, the EIP mode includes one or more additional reconstruction regions (e.g., having different numbers of columns and / or rows).
[0068] Figure 4C Shows different example filter shapes according to some embodiments. The filter shape 412 is a first filter shape in a square shape. The filter shape 412 includes 16 samples (e.g., positions), fifteen shaded samples 418 in the filter shape 412 are provided as the input to the EIP mode, and the EIP mode provides a prediction output 420 at the sixteenth position in the filter shape 412. The filter shape 414 is a second type of filter shape that also includes 16 samples or positions. The filter shape 414 is in a rectangular shape, having a width greater than its height. The filter shape 416 is a third filter shape in a rectangular shape, having a height greater than its width. Each of the filter shape 412, the filter shape 414, and the filter shape 416 includes 16 samples (or positions), where fifteen shaded samples 418 are provided as the input to the EIP to generate a prediction output 420 at the sixteenth position. In some embodiments, the EIP mode includes one or more additional filter shapes (e.g., having different numbers of columns and / or rows). In some embodiments, when the EIP mode is used for prediction in the current block 406, the decoder decodes one or more relevant syntax elements to determine the type of the reconstruction region and the filter shape selected for the current block. In some embodiments, the selected filter slides in the selected reconstruction region with a one-pixel step to collect the input samples and output samples of the EIP mode. In some embodiments, while removing the average value from the input samples and output samples, an autocorrelation matrix and a cross-correlation vector are constructed. In some embodiments, the EIP coefficients are obtained in a method similar to the method in the Convolutional Cross-Component Model (CCCM) for predicting chrominance samples based on reconstructed luminance samples.
[0069] Instead of being limited to a restricted set of intra prediction modes for EIP (e.g., only the planar mode (e.g., mode 0) or the DC mode (e.g., mode 1)), using information from adjacent reconstructed blocks enables the EIP mode to be more adaptable and can provide a more efficient decoding process. For example, in cases where the adjacent reconstructed samples do not have much directionality (e.g., decoded using the DC mode or the planar mode), one or more features associated with the EIP mode of the current block can reflect the lack of directionality. In contrast, in cases where the adjacent reconstructed samples have directionality (e.g., decoded using a 45-degree angle mode such as mode 2 or mode 34, or having strong directionality), one or more features associated with the EIP mode of the current block can reflect that directionality. Thus, using one or more features associated with the EIP mode enables the current block to be encoded using an intra prediction mode that adapts to the directionality of the adjacent reconstructed samples. In some embodiments, one or more features associated with the EIP mode include a directionality indicator that is derived on-the-fly (e.g., derived during the decoding process and not signaled in the bitstream). The directionality indicator can be used to indicate whether directionality exists and / or to specify the angle of a directional intra prediction mode (e.g., the modes 14 to 80 depicted in Figure 4A for decoding the current block). The angle of the derived directional intra prediction mode can be an angle that matches the texture pattern of the current block.
[0070] In some embodiments, deriving one or more features may be computationally more complex. In some embodiments, one or more features (e.g., the directionality indicator) associated with the EIP mode are signaled in the bitstream. Although signaling one or more features may have a higher signaling cost, signaling can provide the encoder with greater flexibility in indicating to the decoder how to select the transform kernel for the current block. In some embodiments, one or more features associated with the EIP mode of the current block decoded using the EIP mode are indices signaled in the bitstream for the current block. The signaled indices can be used to select the transform kernel at both the encoder and the decoder.
[0071] The derived or signaled directionality indicator can be used to map the EIP mode of the current block to one of a directional intra prediction mode (e.g., the modes 14 to 80 depicted in Figure 4A or a non-directional intra prediction mode (e.g., the planar mode or the DC mode). Different directional intra prediction modes can have different transform kernel preferences. By providing the ability to select the transform kernel for a specific directional intra prediction, the characteristics of the current block can be represented more accurately, which can improve the coding efficiency.
[0072] In some embodiments, one or more features associated with the EIP mode include one or more intra prediction modes of adjacent blocks, which can be used to map the EIP mode of the current block to one of a directional intra prediction mode or a non-directional intra prediction mode (e.g., planar mode or DC mode), and to select a transform kernel for the current block. For example, if the upper neighbor and the left neighbor of the current block are encoded using a 45-degree intra prediction mode (e.g., Figure 4A mode 34 in
[0073] ), the same intra prediction mode can be used for the EIP mode of the current block.
[0074] Figure 5A In addition to, or instead of, using one or more features associated with the EIP mode of the current block to derive a directionality indicator for selecting a transform kernel for the current block, one or more features can also be used to generate a prediction of the intra prediction mode of a subsequent block (e.g., the next block encoded using a conventional intra prediction mode rather than the EIP mode). In some embodiments, one or more features associated with the EIP mode of the current block include one or more intra prediction modes of previously decoded adjacent blocks and the EIP mode of the current block. Then, for the next block encoded using an intra prediction mode, one or more features are mapped to one of a directional intra prediction mode or a non-directional intra prediction mode (e.g., planar or DC).
[0075] Template-based Intra Mode Derivation (TIMD) is a method for (e.g., by decoder component 122) deriving an intra prediction mode of a sample based on information from neighboring samples (e.g., neighboring or non-neighboring samples) in a template region. The neighboring samples can be reconstructed samples or previous predicted samples and are collectively referred to hereinafter as "reconstructed neighboring samples". For each intra prediction mode in the Most Probable Mode (MPM) list, the sum of absolute transformed differences (SATD) between the predicted samples of the template and the reconstructed neighboring samples can be calculated. The intra prediction mode with the minimum SATD can be selected as the TIMD mode and used for prediction of the current block. For example, to determine the intra prediction mode of the current block 550 in Figure 5C , two or more of the templates 554, 556, and 558 can be used. Each of the L-shaped templates 554, 556, and 558 includes a plurality of samples (e.g., 13 samples, 13 samples in each horizontal and each vertical portion of the L-shaped template). For each template, a difference is derived for the current block 550, which represents the difference between a candidate predicted value based on a candidate prediction mode in the MPM list and a value from the reconstructed neighboring samples. The candidate prediction mode with the minimum prediction error can be selected as the intra prediction mode for the current block. For example, if the direction of the candidate prediction mode is aligned with the reconstructed neighboring samples, the error (e.g., SATD) is small, and thus information about the directionality of the current block can be inferred. Then, the candidate prediction mode is selected as the TIMD mode and used for prediction of the current block 550. In some embodiments, in addition to information about the current block 550, the encoded information from template 554 is also considered, and this encoded information is used to provide a prediction of template 556, which can then be used to provide a prediction of template 558.
[0076] Another method involves using Decoder-Side Intra Mode Derivation (DIMD) to derive an intra prediction mode for the current block based on information from the reconstructed neighboring samples. For example, the gradient of the encoded information in each reconstructed neighboring sample is calculated and used to populate a histogram. The prediction mode with the highest frequency can be selected as the intra prediction mode for the current block.
[0077] In some embodiments, one or more features associated with the EIP mode include a directional indicator specified by an intra prediction mode derived using the TIMD or DIMD method. Thus, one or more features include a directional indicator derived from neighboring reconstructed samples. In some embodiments, it may be signaled whether to select TIMD or DIMD to derive the intra prediction mode. In some embodiments, TIMD is always selected and a pair of templates (e.g., template 554 and template 556) are used to predict each other. In this way, the intra prediction mode with the minimum prediction error can be identified as the intra prediction mode to be used for the current block. Furthermore, the derived intra prediction can also be used as the directional indicator of the EIP mode for the current block. In some embodiments, one or more features include a directional indicator derived from the intra prediction mode of neighboring samples (e.g., the signaled syntax indicating the intra prediction mode of neighboring samples).
[0078] In some embodiments, one or more features associated with the EIP mode include a directional indicator derived using the coefficient values of the filter used in the EIP mode. For example, each of the grayscale samples (e.g., 418-1, 418-2, 418-3 in filter shapes 412, 414, and 416 as shown in Figure 4C has a corresponding coefficient, and the coefficient and the grayscale sample can be jointly used as the input of the EIP mode. The prediction outputs (e.g., 420-1, 420-2, and 420-3) can be a weighted sum of the grayscale samples, and the coefficients of the grayscale samples can present directional information. In some embodiments, the coefficient values of the filter used in the EIP mode are further quantized into a finite set of combinations, and each combination can optionally be mapped to a value of the directional indicator. In some embodiments, the magnitude of the coefficient of the filter used in the EIP mode (optionally, the quantized magnitude) is optionally also provided to a look-up table (e.g., as its input) to determine the value of the directional indicator. In some embodiments, the signed value of the coefficient of the filter used in the EIP mode (optionally, the quantized signed value, which can be negative) is optionally also provided to a look-up table (e.g., as its input) to determine the value of the directional indicator.
[0079] In some embodiments, one or more features associated with the EIP mode are derived based on the filter shape. For example, different filter shapes can have different supported transform types. In some embodiments, the supported transform type is derived from the aspect ratio of the filter shape (e.g., mapped to the aspect ratio of the filter shape).
[0080] In some embodiments, one or more EIP mode parameters are selected from a first EIP parameter set when the EIP mode is valid for the current block and boundary conditions are met for the current block. When the boundary conditions are not met for the current block, one or more EIP mode parameters are selected from a second EIP parameter set. The second EIP parameter set may include one or more parameters not included in the first EIP parameter set. For example, the second EIP parameter set includes the three types of reconstruction regions shown in Figure 4B and the first EIP parameter set is a subset of the second EIP parameter set (e.g., the first set includes only reconstruction regions 404 and 408). In some embodiments, the second EIP parameter set includes more than the three types of reconstruction regions shown in Figure 4B . In some embodiments, the use of the first EIP parameter set is not signaled (e.g., determined at decoder component 122).
[0081] Figure 5B Unit 514 is shown which may correspond to a picture, sub - picture, slice, or tile. In some embodiments, the boundary conditions are met when the current block is at a first relative position with respect to a picture boundary, sub - picture boundary, slice boundary, and / or tile boundary. For example, the current block 510 has a first relative position with respect to the top boundary portion 508 (e.g., within a threshold number of rows from the top edge and / or within the top boundary portion 508). Depending on the nature of unit 514, the top boundary portion may correspond to a picture boundary / sub - picture boundary / slice boundary / tile boundary. When the current block 510 is within the top boundary portion 508 and / or within a threshold number of rows from the top edge of the top boundary portion 508, the EIP mode parameters corresponding to the type of reconstruction region may be limited to reconstruction regions that include only left - hand reconstruction samples (e.g., the reconstruction region 410 shown in Figure 4B ). For example, when the number of available top rows of reconstruction samples above the current block is less than a threshold, only the reconstruction region 410 with left - hand reconstruction samples is used to generate the prediction output 420. When the number of available top rows is less than a threshold (e.g., two or three lines), the results generated from that top region (e.g., similar to reconstruction region 408) may be unreliable. In some embodiments, when the current block 510 meets the boundary conditions, the type of reconstruction region is not signaled but is determined to correspond to Figure 4BThe reconstructed region 410 shown in []. Optionally, the EIP mode is not valid for the current block within the top boundary portion 508. Additionally or alternatively, when the current block 510 is within the top boundary portion 508 and / or within a threshold number of rows from the top edge of the top boundary portion 508, the EIP mode parameters corresponding to the filter shape can be restricted to a subset of the filter shapes. For example, the restricted subset of filter shapes can include only Figure 4C the filter shape 416 shown in [].
[0082] For example, the current block 512 has a second relative position with respect to the left boundary portion 506 (e.g., within a threshold number of columns from the left edge and / or within the left boundary portion 506). Depending on the nature of the unit 514, the left boundary portion 506 can correspond to a picture boundary / sub-picture boundary / slice boundary / tile boundary. When the current block 512 is within the left boundary portion 506 and / or within a threshold number of lines (e.g., columns) from the left edge of the left boundary portion 506, the type of the reconstructed region includes only top reconstructed samples (e.g., Figure 4B the reconstructed region 408 shown in []. For example, when the number of available left columns of the reconstructed samples to the left of the current block is less than the threshold, only the reconstructed region 408 with top reconstructed samples is used to generate the prediction output 420. In some embodiments, when the current block 512 satisfies the boundary condition, instead of signaling the type of the reconstructed region, it is derived as corresponding to Figure 4B the reconstructed region 408 shown in []. Optionally, the EIP mode is not valid for the current block within the left boundary portion 506. Additionally or alternatively, when the current block 512 is within the left boundary portion 506 and / or within a threshold number of columns from the left edge of the left boundary portion 508, the EIP mode parameters corresponding to the filter shape can be restricted to a subset of the filter shapes. For example, the restricted subset of filter shapes can include only Figure 4C the filter shape 414 shown in [].
[0083] For a set with three reconstruction regions and a set with three filter shapes, there are nine combinations of reconstruction region - filter shape pairs (e.g., three combinations from each pairing of reconstruction region 404 with filter shape 412, filter shape 414, and filter shape 416; three combinations from each pairing of reconstruction region 408 with filter shape 412, filter shape 414, and filter shape 416; and three combinations from each pairing of reconstruction region 410 with filter shape 412, filter shape 414, and filter shape 416). In some embodiments, the combinations are limited to a subset of the nine combinations when boundary conditions are met for the current block. For example, each type of reconstruction region may be paired with only two of the three filter shapes, or each type of reconstruction region may be paired with only one of the filter shapes. In some embodiments, reconstruction region 404 is paired only with filter shape 412, reconstruction region 408 is paired only with filter shape 414, and reconstruction region 410 is paired only with filter shape 416.
[0084] In some embodiments, for an encoded block located in region 516 of unit 514 (e.g., at a partition boundary or at the upper left corner of unit 514), an indicator for the EIP mode is not signaled, indicating that the EIP mode is not valid for the encoded block.
[0085] In some embodiments, when encoding the current block in EIP mode and only partial samples of the requested reconstruction region type (e.g., and associated filter shape) are available, the missing samples are filled (e.g., using a predefined value, using a copy of another sample, by extrapolating from available samples, or by interpolating using available samples) to construct a reconstruction region with a complete set of samples. For example, in Figure 5B where the current block 512 meets the boundary conditions (e.g., the current block 512 is within the left partial boundary 506), and the EIP mode parameters for the filter shape corresponding to the current block may indicate the use of filter shape 414 for the current block 512, and optionally include the use of reconstruction region 408. As described with respect to Figure 4C fifteen input samples (shaded squares) are used to generate the prediction output 420 - 4 from filter shape 414. In Figure 5CIn [the figure], eight samples in region 518 are missing from the filter shape 414 (e.g., outside the reconstruction region and / or outside cell 514). In some embodiments, predefined (e.g., constant) values are used to fill in the missing samples in the filter shape (e.g., the eight missing samples in region 518). The advantage of filling in the missing samples (e.g., within the reconstruction region and / or within the filter samples) is that a unified processing scheme can be used for both the reconstruction region or filter shape with missing samples and the reconstruction region or filter shape with a complete set of samples. Having a unified processing scheme can reduce hardware requirements (e.g., using a shared or common pipeline for different processing scenarios).
[0086] In some embodiments, the predefined value is set to 1<<(bitdepth - 1), where bitdepth is the bit depth of the luminance or chrominance samples. For example, when the bit depth (bitdepth) of the luminance samples is 4, the predefined value is 1 * 2^3 = 8. In some embodiments, the available rows and columns are extended towards the missing rows and columns in the template of the reconstruction region. For example, in Figure 5B [the figure], seven samples to the right of region 518 are shifted into region 518 to help construct the eight missing samples (e.g., filling the missing eighth sample, and / or using a copy of one of the seven samples for the eighth sample).
[0087] In some embodiments, the filled samples are derived line by line (e.g., row by row, column by column). For example, the missing samples in the line closest to the available reconstruction region are derived from N (N >= 1) available adjacent samples. In some embodiments, the average value of the N adjacent samples is used to fill in the missing samples in the line closest to the N adjacent samples. In a manner similar to TIMD (but the derived values are used for filling instead of prediction), this operation is repeated on the newly filled line to obtain the values for filling the missing samples in the second closest available reconstruction region line (e.g., followed by the third closest line and the fourth closest line, etc.) until all missing samples are filled. For example, in Figure 5B the enlarged version of the filter shape 414 shown at the bottom of [the figure], region 518 includes the first line 526 closest to the available reconstruction region. Figure 5B An example is shown where N = 4 and the values of the samples in the first line 526 are derived (e.g., average, maximum, or minimum) using the four available samples within region 530. Subsequently, the values of the samples in the second line 528 are derived using the next four available samples within region 532 that include the newly filled samples from the first line 526. This process is repeated until all four missing lines are filled.
[0088] One or more features of the EIP mode can be used for one or more of the following: (1) selecting a primary transform set or a primary transform type, (2) selecting a secondary transform set or a secondary transform type, and / or (3) deriving the most likely intra prediction mode of other coding blocks (e.g., the coding block after the current block). The secondary transform is an additional transform process after the primary transform. For example, in NSST, an in-separable secondary transform is applied to the lower frequency coefficients, so that the computational complexity of the in-separable transform can be reduced.
[0089] The LFNST set indicates a set of transform kernel options that can be selected in the LFNST. In some embodiments (e.g., in VVC), four LFNST sets represented as lfnstSetIdx are defined, and the selection of the set can depend on the intra prediction mode. Three different options of the LFNST kernel are provided in each of the four LFNST sets, and an index (e.g., between 0 and 2) is used to indicate which one of the three kernels is to be used. For example, when the index is 0, the LFNST may not be applied. Otherwise, the LFNST is applied using one of the two kernels in the LFNST set, and this selection is indicated by the LFNST index.
[0090] As another example, if the intra prediction mode of the current block is the planar vertical mode, the horizontal intra prediction mode can be used to derive the transform kernels in the MTS set and the LFNST set. Additionally, if the intra prediction mode of the current block is the planar horizontal mode, the vertical intra prediction mode can be used to derive the transform kernels in the MTS set and the LFNST set.
[0091] In some embodiments, a separate transform set is used for the EIP mode. For example, in the case where the EIP mode is selected for the current block, a separate primary transform set and / or a secondary transform set can be applied. Alternatively, there may be some overlap between the primary transform set and / or the secondary transform set used for the EIP mode and the primary transform set and / or the secondary transform set used for the non-EIP mode. For example, some existing transform sets used for the non-EIP mode may not be optimal for the EIP mode (and thus are excluded from consideration for the EIP mode), while other existing transform sets can also be used for the EIP mode.
[0092] In some embodiments, multiple cores are available for each transform set. For example, an indicator (e.g., an optionally signaled indicator) determines the transform set to be used (e.g., one of 12 transform sets), and a second indicator (e.g., optionally signaled) is used to determine which core within the set will be used for the EIP mode. In some embodiments, depending on whether the IP mode is valid, different cores are preferred for the intra prediction mode for the same angle. As an example, the 45° intra prediction mode may be used to encode one sample, and that sample includes very sharp features, while the 45° intra prediction mode may also be used to encode another sample, but that sample includes very smooth features. The two samples may preferably use different transform sets and / or different cores within a particular transform set. In some embodiments, there are 12 transform sets, and each set has 3 cores (e.g., all three cores are for the same angle). Given an intra prediction mode for the EIP mode, one of the 12 sets is selected, and the signaled indicator may indicate which one of the three cores is to be used. In some embodiments, additional transform sets are added for the EIP mode (e.g., resulting in a total of 15 or more transform sets).
[0093] Figure 6A FIG. is a flow diagram illustrating a method 600 for decoding video according to some embodiments. Method 600 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 600 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.
[0094] The system receives (602) a video bitstream that includes a plurality of blocks. When an extrapolation filter-based intra prediction (EIP) mode is valid for a current block among the plurality of blocks, the system determines (604) one or more EIP mode parameters from a first EIP parameter set when boundary conditions are satisfied for the current block. When the boundary conditions are not satisfied for the current block, the system determines (606) one or more EIP mode parameters from a second EIP parameter set, where the second EIP parameter set includes one or more parameters not included in the first EIP parameter set. The system uses the one or more EIP mode parameters to reconstruct (610) the current block. For example, when encoding the current block by the EIP mode, only a subset of all predefined types of reconstruction regions is applicable and signaled under certain conditions. For example, these conditions are related to the relative position of the current block within a picture / sub-picture / slice / tile. In some embodiments, one or more EIP mode parameters are determined from the first EIP parameter set according to the determination that the boundary conditions are satisfied. In some embodiments, one or more EIP mode parameters are determined from the second EIP parameter set according to the determination that the boundary conditions are not satisfied.
[0095] Figure 6B is a flowchart illustrating a method 650 for encoding video according to some embodiments. Method 650 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) that has control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 650 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.
[0096] The system receives (652) video data that includes a plurality of video blocks. When an extrapolation filter-based intra prediction (EIP) mode is valid for a current block among the plurality of blocks, the system determines (654) one or more EIP mode parameters from a first EIP parameter set when boundary conditions are satisfied for the current block. When the boundary conditions are not satisfied for the current block, the system determines (656) one or more EIP mode parameters from a second EIP parameter set, where the second EIP parameter set includes one or more parameters not included in the first EIP parameter set. The system encodes (658) the current block using the one or more EIP mode parameters. As previously described, the encoding process may mirror the decoding process described herein (e.g., the EIP embodiments described above). For the sake of brevity, these details are not repeated here.
[0097] Although Figure 6A and Figure 6BA number of logical stages are shown in a particular order, but stages that are not order-dependent can be reordered, and other stages can be combined or split. For those of ordinary skill in the art, some reorderings or other groupings that are not specifically mentioned will be obvious, so the orderings and groupings presented herein are not exhaustive. In addition, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.
[0098] Turning now to some example embodiments.
[0099] (A1) In one aspect, some embodiments include a method of video decoding (e.g., method 600). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving a video bitstream (e.g., an encoded video sequence) including a plurality of blocks; when an EIP mode is valid for a current block among the plurality of blocks: when boundary conditions are satisfied for the current block, (ii) determining one or more EIP mode parameters from a first set of EIP parameters; when boundary conditions are not satisfied for the current block, (iii) determining one or more EIP mode parameters from a second set of EIP parameters, wherein the second set of EIP parameters includes one or more parameters not included in the first set of EIP parameters; and (iv) reconstructing the current block using the one or more EIP mode parameters. For example, in the case of encoding the current block by an EIP mode, only a subset of all predefined types of reconstruction regions is applicable and signaled under certain conditions. For example, these conditions are related to the relative position of the current block within a picture / sub-picture / slice / tile. In some embodiments, one or more EIP mode parameters are determined from the first set of EIP parameters according to the determination that boundary conditions are satisfied. In some embodiments, one or more EIP mode parameters are determined from the second set of EIP parameters according to the determination that boundary conditions are not satisfied.
[0100] (A2) In some embodiments of A1, one or more EIP mode parameters include parameters for a reconstruction region of the EIP mode. For example, the parameters of the reconstruction region are the index, size, or shape of the reconstruction region.
[0101] (A3) In some embodiments of A1 or A2, one or more EIP mode parameters include parameters for a filter shape of the EIP mode. For example, the parameters of the filter shape are the index, size, or shape of the filter shape.
[0102] (A4)In some embodiments of any one of A1 to A3, the method further includes parsing, from the video bitstream, an indicator indicating whether to use an extrapolation filter intra prediction (EIP) mode to decode a current block among the plurality of blocks. In some embodiments, one or more EIP mode parameters are determined based on a determination that the EIP mode will be used to decode the current block as indicated by the indicator. In some embodiments, one or more EIP mode parameters are not determined (e.g., the computing system abandons determining one or more EIP mode parameters) based on a determination that the EIP mode will not be used to decode the current block as indicated by the indicator.
[0103] (A5)In some embodiments of any one of A1 to A4, the boundary conditions relate to picture boundaries, sub-picture boundaries, slice boundaries, and / or tile boundaries. For example, different subsets of a predefined type of reconstruction region are applicable depending on the relative position of the current block within the picture / sub-picture / slice / tile.
[0104] (A6)In some embodiments of any one of A1 to A5, the boundary conditions include that the current block is located at a top split boundary. For example, if an encoded block is located at the top picture boundary, sub-picture boundary, slice boundary, or tile boundary, only a subset of a predefined type of reconstruction region (and / or filter type) is applicable, e.g., only the left reconstruction samples are used (or weighted).
[0105] (A7)In some embodiments of any one of A1 to A6, the boundary conditions include that the current block is located at a left split boundary. For example, if an encoded block is located at the left picture boundary, sub-picture boundary, slice boundary, or tile boundary, only a subset of a predefined type of reconstruction region (and / or filter type) is applicable, e.g., only the top reconstruction samples are used (or weighted).
[0106] (A8)In some embodiments of A1 to A7, the method further includes: determining that the EIP mode is invalid for the current block in the case where the current block is located at a split boundary. In some embodiments, determining that the current block is located at a split boundary is determining that the current block is within a predetermined threshold number of lines from the split boundary. In some embodiments, based on the determination that the current block is located at a split boundary, it is determined that the EIP mode is invalid for the current block.
[0107] (A9)In some embodiments of A8, the split boundary is an upper left split boundary. For example, if an encoded block is located at the upper left picture corner, sub-picture corner, slice corner, or tile corner, the EIP mode is not applicable or not signaled.
[0108] (A10)In some embodiments of A1 to A9, the boundary condition corresponds to the number of available reconstruction samples for the current block. For example, the condition is related to the number of available top rows of reconstruction samples and / or the number of available left columns of reconstruction samples. Depending on the number of available top rows of reconstruction samples and / or the number of available left columns of reconstruction samples, different subsets of a predefined type of reconstruction region (and / or filter shape) are applicable.
[0109] (A11)In some embodiments of A10, the boundary condition is whether the number of available top rows is less than a predetermined threshold. For example, if the number of available top rows of reconstruction samples is less than a given threshold, only a subset of a predefined type of reconstruction region (and / or filter shape) is applicable. For example, the reconstruction region only includes left reconstruction samples. In some embodiments, the predetermined threshold is a non - negative integer value (e.g., of samples, lines, rows, or columns).
[0110] (A12)In some embodiments of A10 or A11, the boundary condition is whether the number of available left columns is less than a predetermined threshold. For example, if the number of available left columns of reconstruction samples is less than a given threshold, only a subset of a predefined type of reconstruction region (and / or filter shape) is applicable. For example, the reconstruction region only includes top reconstruction samples. In some embodiments, the predetermined threshold is a non - negative integer value (e.g., of samples, lines, rows, or columns).
[0111] (A13)In some embodiments of any one of A1 to A12, in the case where the EIP mode is valid for the current block and one or more samples in the reconstruction region for the EIP mode are unavailable, one or more samples are filled for the EIP mode. For example, when encoding the current block in EIP mode and only partial samples in the requested reconstruction region type are available, the missing samples are filled to construct a full - size reconstruction region. In some embodiments, the one or more samples are filled based on the determination that one or more samples for the EIP mode are unavailable.
[0112] (A14)In some embodiments of A13, filling one or more samples includes filling the one or more samples with a predefined constant value. For example, the missing samples are filled with a predefined constant value.
[0113] (A15)In some embodiments of A14, the predefined constant value is based on the sample bit depth. For example, the value can be set to 1<<(bitdepth - 1), where bitdepth is the bit depth of the luminance or chrominance sample, and << is the left - shift operator.
[0114] (A16)In some embodiments of any one of A13 to A15, filling one or more samples includes extending available sample values to fill the one or more samples. For example, extending available rows and columns towards missing rows and columns in a template of a reconstruction region. In some embodiments, filled samples are derived line by line (e.g., row by row or column by column), where samples in the line closest to the available reconstruction region are derived from N (N >= 1) adjacent samples (e.g., the average of N adjacent samples). The same operation can be applied to newly filled lines in order to derive the second closest line, the third closest line, etc., until the full-sized reconstruction region is filled.
[0115] (A17)In some embodiments of any one of A1 to A16, a first set of EIP parameters includes a first set of reconstruction region types and filter shapes for an EIP mode, a second set of EIP parameters includes a second set of reconstruction region types and filter shapes for the EIP mode, and the first set of reconstruction region types and filter shapes for the EIP mode includes a subset of the second set of reconstruction region types and filter shapes for the EIP mode. For example, for boundary cases, only predefined combinations of filter shapes and reconstruction regions are allowed. For example, the number of combinations of filter shapes and reconstruction regions can be less than 9 (e.g., for the second set), e.g., 3 or 6.
[0116] (B1)In another aspect, some embodiments include a method of video encoding (e.g., method 650). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving video data including a plurality of video blocks; when the EIP mode is valid for a current block among the plurality of blocks: when boundary conditions are met for the current block, (ii) determining one or more EIP mode parameters from a first set of EIP parameters; when boundary conditions are not met for the current block, (iii) determining one or more EIP mode parameters from a second set of EIP parameters, where the second set of EIP parameters includes one or more parameters not included in the first set of EIP parameters; and (iv) encoding the current block using the one or more EIP mode parameters.
[0117] (B2)In some embodiments of B1, the system further includes instructions for signaling via a video bitstream an indicator indicating whether to decode a current block among the plurality of blocks using the EIP mode.
[0118] (C1)In another aspect, some embodiments include a method for visual media data processing. In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) obtaining a source video sequence including a plurality of frames; and (ii) performing a conversion between the source video sequence and a video bitstream of visual media data, where the video bitstream includes: (a) a plurality of coded blocks corresponding to a plurality of video blocks, the plurality of coded blocks including a first block encoded using an extrapolation filter-based intra prediction (EIP) mode; in the case where boundary conditions are satisfied for the first block, (b) an indication of a first set of EIP mode parameters for the first block; and in the case where the boundary conditions are not satisfied for the first block, (c) an indication of a second set of EIP mode parameters for the first block, where the second set of EIP parameters includes one or more parameters not included in the first set of EIP parameters.
[0119] (D1)In one aspect, some embodiments include a method for video decoding. In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving a video bitstream including a plurality of blocks; in the case where the EIP mode is valid for a current block among the plurality of blocks: (ii) identifying a reconstruction region for the EIP mode; in the case where the reconstruction region includes one or more unavailable samples, (iii) deriving one or more filled samples by filling the one or more unavailable samples, and (iv) reconstructing the current block according to the EIP mode using the one or more filled samples. For example, each missing sample is filled using a predefined constant value (e.g., based on the sample bit depth). As another example, available rows and columns are extended towards missing rows and columns in a template of the reconstruction region. As another example, the filled samples are derived line by line, where the samples in the line closest to the available reconstruction region are derived from N adjacent samples.
[0120] (D2)In some embodiments of D1, the method further includes: in the case where boundary conditions are satisfied for the current block, (i) determining one or more EIP mode parameters from a first set of EIP parameters; and in the case where the boundary conditions are not satisfied for the current block, (ii) determining one or more EIP mode parameters from a second set of EIP parameters, where the second set of EIP parameters includes one or more parameters not included in the first set of EIP parameters; and where the one or more EIP mode parameters are used to reconstruct the current block.
[0121] (D3) In some embodiments of D1 or D2, the method further includes any one of the various techniques described above with respect to A1 to A17.
[0122] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more instruction sets configured to be executed by the control circuitry, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 to A17, B1 to B2, C1, and D1 to D3 above).
[0123] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more instruction sets for execution by control circuitry of a computing system, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 to A17, B1 to B2, C1, and D1 to D3 above).
[0124] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0125] As used herein, the term "when" can be interpreted, depending on the context, as "if the precondition is true" or "after the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "in response to detecting that the precondition is true". Similarly, the phrase "if it is determined that [the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" can be interpreted, depending on the context, as "after determining that the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "after detecting that the precondition is true" or "in response to detecting that the precondition is true". As used herein, N refers to a variable number. Unless explicitly stated, different instances of N can refer to the same number (e.g., the same integer value, such as the number 2) or different numbers.
[0126] For purposes of illustration, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best illustrate the operating principles and practical applications, thereby enabling others skilled in the art to implement.
Claims
1. A method of video decoding, the method being performed at a computing system having a memory and one or more processors, the method comprising: receiving a video bitstream comprising a plurality of blocks; In case an extrapolation filter based intra prediction (EIP) mode is valid for a current block among the plurality of blocks: When a boundary condition is satisfied for the current block, determining one or more EIP mode parameters from a first EIP parameter set; When the boundary condition is not satisfied for the current block, determining the one or more EIP mode parameters from a second EIP parameter set, wherein the second EIP parameter set includes one or more parameters not included in the first EIP parameter set; as well as The current block is reconstructed using the one or more EIP mode parameters.
2. The method according to claim 1, wherein: The one or more EIP mode parameters include parameters for a reconstruction region of the EIP mode.
3. The method according to claim 1, wherein: The one or more EIP mode parameters include parameters of a filter shape for the EIP mode.
4. The method according to claim 1, further comprising: An indicator is parsed from the video bitstream to indicate whether a current block of the plurality of blocks is to be decoded using an extrapolation filter intra prediction (EIP) mode.
5. The method according to claim 1, wherein: The boundary conditions relate to picture boundaries, sub-picture boundaries, slice boundaries and / or tile boundaries.
6. The method according to claim 1, wherein: The boundary condition includes that the current block is located at a top partition boundary.
7. The method according to claim 1, wherein: The boundary condition includes that the current block is located at a left partition boundary.
8. The method according to claim 1, further comprising: In case the current block is located at a partition boundary, it is determined that the EIP mode is invalid for the current block.
9. The method according to claim 8, wherein: The segmentation boundary is an upper left segmentation boundary.
10. The method according to claim 1, wherein: The boundary condition corresponds to the number of available reconstructed samples of the current block.
11. The method according to claim 10, wherein: The boundary condition is whether the number of available top rows is less than a predetermined threshold.
12. The method according to claim 10, wherein: The boundary condition is whether the number of available left columns is less than a predetermined threshold.
13. The method according to claim 1, further comprising: In case the EIP mode is valid for the current block and one or more samples in a reconstruction area for the EIP mode are not available, the one or more samples are filled for the EIP mode.
14. The method according to claim 13, wherein: Padding the one or more samples includes padding the one or more samples using a predefined constant value.
15. The method according to claim 14, wherein: The predefined constant value is based on the sample bit depth.
16. The method according to claim 13, wherein: Filling the one or more samples includes extending available sample values to fill the one or more samples.
17. The method according to claim 1, wherein: The first EIP parameter set comprises a first set of reconstruction area types and filter shapes for the EIP mode, wherein the second EIP parameter set comprises a second set of reconstruction area types and filter shapes for the EIP mode, and wherein the first set of reconstruction area types and filter shapes for the EIP mode comprises a subset of the second set of reconstruction area types and filter shapes for the EIP mode.
18. A computing system comprising: Control circuit system; Memory; as well as one or more instruction sets stored in the memory and configured for execution by the control circuit system, the one or more instruction sets comprising instructions for: receiving video data comprising a plurality of video blocks; In case an extrapolation filter based intra prediction (EIP) mode is valid for a current block among the plurality of blocks: When a boundary condition is satisfied for the current block, determining one or more EIP mode parameters from a first EIP parameter set; When the boundary condition is not satisfied for the current block, determining the one or more EIP mode parameters from a second EIP parameter set, wherein the second EIP parameter set includes one or more parameters not included in the first EIP parameter set; as well as The current block is encoded using the one or more EIP mode parameters.
19. The computing system of claim 18, further comprising: An indicator indicating whether a current block among the plurality of blocks is to be decoded using the EIP mode is signaled via a video bitstream.
20. A non-transitory computer-readable storage medium storing one or more instruction sets configured for execution by a computing device having a control circuit system and a memory, the one or more instruction sets comprising instructions for: Obtaining a source video sequence comprising a plurality of video blocks; as well as performing a conversion between the source video sequence and a bitstream of visual media data, wherein: The bitstream includes: a plurality of coding blocks corresponding to the plurality of video blocks, the plurality of coding blocks comprising a first block encoded using an extrapolation filter based intra prediction (EIP) mode; an indication of a first EIP mode parameter set for the first block if a boundary condition is satisfied for the first block; and An indication of a second EIP mode parameter set for the first block if the boundary condition is not satisfied for the first block, wherein the second EIP parameter set includes one or more parameters not included in the first EIP parameter set.