Intra block copy for filtering based on gradient and location
By applying the coefficients of linear filters, gradient filters and nonlinear values in the intra-block copy mode of video encoding and decoding, and combining position information, the problem of difficulty in effectively utilizing gradient and position information in the prior art is solved, and prediction efficiency and video quality are improved.
Patent Information
- Application Number
- CN202480004074.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-26
- Filing Date
- 2024-04-23
- Publication Date
- 2025-06-03
AI Technical Summary
The existing video encoding and decoding technology is difficult to effectively utilize gradient and position information in the intra-block replication mode, resulting in low prediction efficiency.
A method of processing visual media data is adopted to determine the predicted value of the current sample in the current block by applying coefficients of linear filters, gradient filters and nonlinear values in the intra-block copy mode, combined with position information.
Improve the prediction efficiency of intra-block replication mode, reduce encoding errors and improve video quality by utilizing gradient and position information more accurately.
Smart Images

Figure CN120092445A_ABST
Abstract
Description
Related Applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 462,235, filed on Apr. 26, 2023, entitled “Intra Block Copy Based on Gradient and Location Filtering,” which is incorporated herein by reference in its entirety. Technical Field
[0002] Embodiments of the present disclosure generally relate to video coding / decoding. Background Art
[0003] The background description provided herein is for the purpose of generally presenting the background of the present disclosure. To the extent that the work of the presently named inventors, which is described in the background art section and in various aspects of this specification, was done, it does not, at the time of the filing of this disclosure, constitute prior art, and is never expressly or implicitly admitted as prior art to this disclosure.
[0004] Image / video compression can help to transmit image / video data across different devices, storage, and networks while minimizing quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra prediction, which can compress an image based on spatial redundancy. For example, intra prediction can use reference data from the currently reconstructed picture for sample prediction. In another example, a video codec can use a technique called inter prediction, which can compress an image based on temporal redundancy. For example, inter prediction can predict samples in the current picture from a previously reconstructed picture through motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention
[0005] Aspects of the present disclosure include methods and apparatuses for video encoding / decoding.
[0006] In one aspect, a method of processing visual media data includes: processing a bitstream of the visual media data according to formatting rules. The bitstream includes syntax elements that indicate use of a filtered intra-block copy (FIBC) mode to predict a current block in a current picture. The formatting rules specify that a linear prediction value of a current sample in the current block is determined by applying a linear filter to samples predicted using one of an intra-block copy (IBC) mode and an intra-template matching (IntraTMP) mode. The formatting rules specify that at least one gradient filter is used to determine a gradient value associated with the current sample in the current block. The formatting rules specify that a non-linear value associated with the current sample is determined using a non-linear relationship between the non-linear value and a value of at least one of the current sample and adjacent samples of the current sample, and based on the current sample and the at least one of the adjacent samples. The formatting rules specify that a position value is determined based on a position of a central sample located at a center of the linear filter. The formatting rules specify that a predicted value of the current sample is determined based on a sum of the linear prediction value and at least one modification value, the at least one modification value including the gradient value, the non-linear value, and the position value, and the FIBC filter in the FIBC mode includes the linear filter, the at least one gradient filter, a coefficient of the non-linear value, and a coefficient of the position. The formatting rules specify that the current sample is processed based on the predicted value of the current sample.
[0007] In an example, the linear filter includes a bias term.
[0008] In an example, the linear filter is configured to add an average value of the current block and remove the average value of the current block from each of the samples predicted using one of the IBC mode and the IntraTMP mode.
[0009] In one aspect, a method for video coding includes: determining a linear prediction value of a current sample in a current block by applying a linear filter to samples predicted using one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode, the current block being predicted using a filtered intra block copy (FIBC) mode; determining a gradient value associated with the current sample in the current block using at least one gradient filter; determining a non - linear value associated with the current sample using a non - linear relationship between the non - linear value and a value of at least one of the current sample and adjacent samples of the current sample, and based on the current sample and the at least one of the adjacent samples; determining a predicted value of the current sample based on a sum of the linear prediction value and at least one modification value, the at least one modification value including the gradient value and the non - linear value, the FIBC filter in the FIBC mode including the linear filter, the at least one gradient filter, and coefficients of the non - linear value; and encoding the current sample according to the predicted value of the current sample.
[0010] In an example, the method for video coding further includes: determining a position value using a position of a central sample located at a center of the linear filter; and determining a predicted value of the current sample based on a sum of the linear prediction value and the at least one modification value, the at least one modification value including the gradient value, the non - linear value, and the position value, the FIBC filter in the FIBC mode including the linear filter, the at least one gradient filter, coefficients of the non - linear value, and coefficients of the position.
[0011] In an example, the linear filter includes a bias term; or the linear filter is configured to add an average value of the current block and remove the average value of the current block from each of the samples predicted using one of the IBC mode and the IntraTMP mode.
[0012] According to one aspect of the present disclosure, an apparatus for video decoding includes processing circuitry. The processing circuitry is configured to: receive encoded information that indicates prediction of a current block in a current picture using a filtered intra block copy (FIBC) mode; determine a linear prediction value of a current sample in the current block by applying a linear filter to prediction samples predicted using one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode; determine a gradient value associated with the current sample in the current block using at least one gradient filter; determine a prediction value of the current sample based on a sum of the linear prediction value and at least one modification value including the gradient value, wherein the FIBC filter in the FIBC mode includes the linear filter and the at least one gradient filter; and reconstruct the current sample according to the prediction value of the current sample.
[0013] In an example, the processing circuitry is configured to: determine a position value using a position of a central sample located at a center of the linear filter; and determine the prediction value of the current sample based on a sum of the linear prediction value and the at least one modification value, wherein the at least one modification value includes the gradient value and the position value, and the FIBC filter in the FIBC mode includes the linear filter, the at least one gradient filter, and a coefficient of the position.
[0014] In an example, the processing circuitry is configured to: determine a non - linear value associated with the current sample using a non - linear relationship between a non - linear value and a value of at least one of the current sample and an adjacent sample of the current sample, and according to the current sample and the at least one of the adjacent samples; and determine the prediction value of the current sample based on a sum of the linear prediction value and the at least one modification value, wherein the at least one modification value includes the gradient value and the non - linear value, and the FIBC filter in the FIBC mode includes the linear filter, the at least one gradient filter, and a coefficient of the non - linear value.
[0015] In an example, the linear filter includes a bias term.
[0016] In an example, the linear filter adds an average value of the current block and removes the average value of the current block from each of the samples predicted using one of the IBC mode and the IntraTMP mode.
[0017] In an example, the processing circuitry is configured to clip the prediction value of the current sample.
[0018] In an example, the processing circuit is configured to determine coefficients of an FIBC filter in the FIBC mode based on a current template of the current block and a reference template of a reference block indicated by a block vector of the current block.
[0019] In an example, the processing circuit is configured to use LDL decomposition to determine coefficients of the FIBC filter in the FIBC mode.
[0020] In an example, the linear filter has a cross shape, and the cross shape includes: (i) five samples, the five samples including a center sample with an offset of (0, 0) of the linear filter, a north sample N with an offset of (0, -1), a south sample S with an offset of (0, 1), an east sample E with an offset of (1, 0), and a west sample W with an offset of (-1, 0), and the offsets of the five samples in the linear filter are relative to the center sample; or (ii) nine samples, the nine samples including a center sample with an offset of (0, 0) of the linear filter, two north samples with offsets of (0, -1) and (0, -2) respectively, two south samples with offsets of (0, 1) and (0, 2) respectively, two east samples with offsets of (1, 0) and (2, 0) respectively, and two west samples with offsets of (-1, 0) and (-2, 0) respectively, and the offsets of the nine samples in the linear filter are relative to the center sample.
[0021] In an example, when the linear filter has the five samples, the current sample is located at one of the five positions of the five samples; and when the linear filter has the nine samples, the current sample is located at one of the nine positions of the nine samples.
[0022] In an example, the shape of the linear filter is predefined, and the shape of at least one gradient filter is predefined.
[0023] In an example, the samples in the linear filter are spatially separated from all the remaining samples in the linear filter.
[0024] In an example, when the at least one gradient filter consists of a horizontal gradient filter, the gradient value is a horizontal gradient value, and the number of first input samples and the positions of the first input samples in the horizontal gradient filter are independent of the linear filter setting; when the at least one gradient filter consists of a vertical gradient filter, the gradient value is a vertical gradient value, and the number of second input samples and the positions of the second input samples in the vertical gradient filter are independent of the linear filter setting; when the at least one gradient filter includes a horizontal gradient filter and a vertical gradient filter, the gradient value is the sum of the horizontal gradient value and the vertical gradient value, and the number of the first input samples and the positions of the first input samples in the horizontal gradient filter, and the number of the second input samples and the positions of the second input samples in the vertical gradient filter are set independently of each other and independently of the linear filter setting.
[0025] In an example, the at least one gradient filter includes a horizontal gradient filter; the gradient value includes a horizontal gradient value, which is the sum of the horizontal gradients of the respective first input samples in the horizontal gradient filter; the processing circuit is configured to determine each horizontal gradient of the respective first input samples based on one of the following: (i) the difference between the respective first input samples and their left neighbors; (ii) the difference between the left neighbors and the right neighbors of the first input samples; and (iii) the difference between a first value and a second value, the first value being the sum determined based on the upper left neighbor, the left neighbor, and the lower left neighbor of the first input sample, and the second value being the sum determined based on the upper right neighbor, the right neighbor, and the lower right neighbor of the second input sample.
[0026] In an example, the processing circuit is configured to determine which difference to use to calculate the horizontal gradient of the respective first input samples based on the positions of the respective first input samples.
[0027] In an example, the at least one gradient filter includes a vertical gradient filter; the gradient value includes a vertical gradient value, which is the sum of the vertical gradients of the respective second input samples in the vertical gradient filter. The processing circuit is configured to determine the vertical gradient of each of the respective second input samples based on one of the following: (i) the difference between the respective second output samples and the top neighbors of the respective second output samples; (ii) the difference between the top neighbors of the respective second input samples and the bottom neighbors of the respective second output samples; and (iii) the difference between a first value and a second value, where the first value is a sum determined based on the upper left neighbor, the top neighbor, and the upper right neighbor of the second input sample, and the second value is a sum determined based on the lower left neighbor, the bottom neighbor, and the lower right neighbor of the respective second input samples.
[0028] In an example, the processing circuit is configured to determine which difference to use to calculate the vertical gradient of the respective second output samples based on the position of the respective second input samples.
[0029] Aspects of the present disclosure also provide an apparatus for video coding. The apparatus for video coding includes a processing circuit configured to implement any one of the video coding methods described above.
[0030] Aspects of the present disclosure also provide a method for video decoding. The method includes any method implemented by the apparatus for video decoding.
[0031] Aspects of the present disclosure also provide a non - volatile computer - readable medium storing instructions that, when executed by a computer, cause the computer to execute any one of the video decoding / encoding methods described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0033] Figure 1 is a schematic diagram of an example of a block diagram of a communication system (100).
[0034] Figure 2 is a schematic diagram of an example of a block diagram of a decoder.
[0035] Figure 3 is a schematic diagram of an example of a block diagram of an encoder.
[0036] Figure 4 illustrates an example of a convolutional filter according to one aspect of the present disclosure.
[0037] Figure 5Shows an example of a reference region for deriving filter coefficients according to one aspect of the present disclosure.
[0038] Figure 6 Shows an example of spatial samples for a gradient and location based convolutional cross-component model (GL-CCCM) according to one aspect of the present disclosure.
[0039] Figure 7A Shows an example of an intra-template matching prediction (IntraTMP) mode according to one aspect of the present disclosure.
[0040] Figure 7B Shows an example of a modification of a filtered intra-block copy (FIBC) model according to one aspect of the present disclosure.
[0041] Figures 8 - 10 Shows an example of available filters in the FIBC mode according to aspects of the present disclosure.
[0042] Figure 11 Shows an example of neighboring samples of input sample C for calculating a gradient according to one aspect of the present disclosure.
[0043] Figure 12 Shows an example of a method for selecting a gradient calculation according to one aspect of the present disclosure.
[0044] Figure 13 Shows an example of the positions of samples C, A, L, and AL according to one aspect of the present disclosure.
[0045] Figure 14 Shows a flowchart outlining a decoding process according to some aspects of the present disclosure.
[0046] Figure 15 Shows a flowchart outlining an encoding process according to some aspects of the present disclosure.
[0047] Figure 16 Is a schematic diagram of a computer system according to one aspect. Detailed Description
[0048] Figure 1 Shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is an example of an application for the disclosed subject matter, video encoders, and video decoders in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0049] The video processing system (100) includes an acquisition subsystem (113), and the acquisition subsystem may include a video source (101) such as a digital camera, and the video source creates an uncompressed video picture stream (102). In an embodiment, the video picture stream (102) includes samples taken by the digital camera. Compared with the encoded video data (104) (or encoded video bitstream), the video picture stream (102) is depicted as a thick line to emphasize the high-data-volume video picture stream. The video picture stream (102) can be processed by an electronic device (120), and the electronic device (120) includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the video picture stream (102), the encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize the lower-data-volume encoded video data (104) (or encoded video bitstream), which can be stored on the streaming server (105) for future use. At least one streaming client subsystem, such as Figure 1 the client subsystem (106) and the client subsystem (108) in
[0050] can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in the electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an output video picture stream (111) that can be presented on a display (112) (such as a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), video data (107), and video data (109) (such as video bitstreams) may be encoded according to certain video coding / compression standards. Embodiments of such standards include ITU-T H.265. In an embodiment, a video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application can be used in the context of the VVC standard.
[0051] Figure 2Shows an example of a block diagram of a video decoder (210). The video decoder (210) can be provided in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used to replace Figure 1 the video decoder (110) in the embodiment.
[0052] The receiver (231) can receive at least one encoded video sequence, for example included in a bitstream, to be decoded by the video decoder (210). In one aspect, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence can be received from a channel (201), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (231) can receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that can be forwarded to their respective using entities (not labeled). The receiver (231) can separate the encoded video sequence from other data. To prevent network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) can be provided outside the video decoder (210) (not labeled). And in other cases, a buffer memory (not labeled) is provided outside the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) can be configured inside the video decoder (210) to, for example, handle the playback timing. And when the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may also not be necessary to configure the buffer memory (215), or the buffer memory can be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (215) may also be required, which can be relatively large and can have an adaptive size, and can be at least partially implemented in an operating system or a similar element (not labeled) outside the video decoder (210).
[0053] The video decoder (210) can include a parser (220) to reconstruct symbols (221) according to the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210), and potential information for controlling a display device (212) (e.g., a display screen), etc., which is not a component of the electronic device (230), but can be coupled to the electronic device (230), such as Figure 2As shown in [reference]. The control information for the display device may be a parameter set segment (not labeled) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be based on a video coding technology or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (220) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0054] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0055] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve at least two different units. Which units are involved and the way they are involved may be controlled by the subgroup control information parsed by the parser (220) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (220) and at least two units below are not described.
[0056] In addition to the functional blocks already mentioned, the video decoder (210) may be conceptually divided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and may be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the functional units below.
[0057] The first unit is a scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block including sample values, and the sample values can be input into the aggregator (255).
[0058] In some cases, the output samples of the scaler / inverse transform unit (251) can belong to intra-coded blocks. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but can use predictive information from a previously reconstructed part of the current picture. Such predictive information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a surrounding block with the same size and shape as the block being reconstructed using the reconstructed information extracted from the current picture buffer (258). For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.
[0059] In other cases, the output samples of the scaler / inverse transform unit (251) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (253) can access the reference picture memory (257) to extract samples for prediction. After motion-compensating the extracted samples according to the symbol (221), these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (referred to as residual samples or residual signals in this case), thereby generating output sample information. The motion compensation prediction unit (253) obtaining the prediction samples from an address within the reference picture memory (257) can be controlled by a motion vector, and the motion vector is in the form of the symbol (221) for use by the motion compensation prediction unit (253), and the symbol (221) includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and so on.
[0060] The output samples of the aggregator (255) can be adopted by various loop filtering techniques in the loop filter unit (256). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream), and the parameters can be used by the loop filter unit (256) as symbols (221) from the parser (220). However, in other embodiments, video compression techniques can also respond to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0061] The output of the loop filter unit (256) can be a sample stream that can be output to the display device (212) and stored in the reference picture memory (257) for subsequent inter-picture prediction.
[0062] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.
[0063] The video decoder (210) can perform decoding operations according to a predetermined video compression technique or standard (such as ITU-T H.265). In the sense that the encoded video sequence conforms to the syntax of the video compression technique or standard and the meaning of the profile recorded in the video compression technique or standard, the encoded video sequence can comply with the syntax specified by the video compression technique or standard used. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0064] In an embodiment, a receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of an encoded video sequence. The additional data may be used by a video decoder (210) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, a temporal, spatial, or signal noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, and the like.
[0065] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is provided in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) may be used to replace Figure 1 the video encoder (103) in the embodiment.
[0066] The video encoder (303) may receive video samples from a video source (301) (which is not Figure 3 part of the electronic device (320) in the embodiment), and the video source may capture video images to be encoded by the video encoder (303). In another embodiment, the video source (301) is part of the electronic device (320).
[0067] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as at least two separate pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial pixel arrays, and each pixel may include at least one sample depending on the sampling structure, color space, etc. used. The following focuses on the description of the samples.
[0068] According to an embodiment, a video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other desired time constraint. Enforcing an appropriate encoding speed is a function of a controller (350). In some embodiments, the controller (350) controls and is functionally coupled to other functional units as described below. For simplicity, couplings are not labeled in the figures. Parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be used for other suitable functions that relate to optimizing the video encoder (303) for a certain system design.
[0069] In some embodiments, the video encoder (303) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data. The reconstructed sample stream (sample data) is input into a reference picture memory (334). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related technologies.
[0070] The operation of the "local" decoder (333) may be the same as that of the "remote" decoder that has been described in detail above in connection with Figure 2 the video decoder (310). However, briefly referring additionally to Figure 2 , when symbols are available and the entropy encoder (345) and parser (220) can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).
[0071] In an embodiment, in addition to parsing / entropy decoding that exists in the decoder, decoder techniques also exist in the corresponding encoder in the same or substantially the same functional form. Therefore, this application focuses on decoder operations. The description of encoder techniques can be simplified because encoder techniques are inverse to the decoder techniques described comprehensively. In some areas, a more detailed description is provided below.
[0072] During operation, in some embodiments, the source encoder (330) may perform motion-compensated predictive coding. Referring to at least one previously encoded picture designated as a "reference picture" in a video sequence, the motion-compensated predictive coding performs predictive coding on an input picture. In this way, the coding engine (332) encodes the difference between a pixel block of the input picture and a pixel block of the reference picture, and the reference picture can be selected as the prediction reference for the input picture.
[0073] The local video decoder (333) may decode the encoded video data of a picture that can be designated as a reference picture based on the symbols created by the source encoder (330). The operation of the coding engine (332) may be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 3 not shown), the reconstructed video sequence is generally a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that can be performed by the video decoder on the reference picture and may store the reconstructed reference picture in the reference picture memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference picture, which has the same content (without transmission errors) as the reconstructed reference picture that will be obtained by a remote video decoder.
[0074] The predictor (335) may perform a prediction search for the coding engine (332). That is, for a new picture to be encoded, the predictor (335) may search in the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor (335) may operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (335), it may be determined that the input picture may have prediction references obtained from at least two reference pictures stored in the reference picture memory (334).
[0075] The controller (350) may manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0076] The outputs of all the above functional units can be entropy encoded in an entropy encoder (345). The entropy encoder (345) losslessly compresses the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0077] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) to prepare for transmission over a communication channel (360), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0078] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applicable to the corresponding picture. For example, a picture can generally be assigned to any of the following picture types:
[0079] An intra picture (I picture) can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0080] A predictive picture (P picture) can be encoded and decoded using intra prediction or inter prediction, which uses motion vectors and reference indices to predict the sample values of each block.
[0081] A bi - predictive picture (B picture) can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values of each block. Similarly, at least two predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0082] Source pictures can typically be spatially subdivided into at least two sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the coding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0083] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (303) can perform various compression operations, including prediction coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0084] In an embodiment, the transmitter (340) can transmit additional data when transmitting the encoded video. The source encoder (330) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0085] The captured video can be at least two source pictures (video pictures) in a time series. Intra-picture prediction (often simplified to intra-frame prediction) utilizes the spatial correlation in a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In an embodiment, the specific picture being encoded / decoded is segmented into blocks, and the specific picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using at least two reference pictures, the motion vector can have a third dimension that identifies the reference picture.
[0086] In some embodiments, bidirectional prediction techniques can be used for inter - picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be predicted by a combination of the first reference block and the second reference block.
[0087] In addition, the merge mode technique can be used for inter - picture prediction to improve coding efficiency.
[0088] According to some embodiments disclosed in the present application, predictions such as inter - picture prediction and intra - picture prediction are performed on a block - by - block basis. For example, according to the HEVC standard, a picture in a video picture sequence is segmented into coding tree units (CTUs) for compression. CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be split into at least one coding unit (CU) by a quadtree. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In an embodiment, each CU is analyzed to determine the prediction type for the CU, such as an inter - frame prediction type or an intra - frame prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into at least one prediction unit (PU). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed on a prediction - block basis. Taking the luminance prediction block as the prediction block as an example, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0089] Note that any suitable technology can be used to implement video encoders (103) and (303) and video decoders (110) and (210). In an embodiment, at least one integrated circuit can be used to implement video encoders (103) and (303) and video decoders (110) and video decoder (210). In another embodiment, at least one processor executing software instructions can be used to implement video encoders (103) and video encoder (303) and video decoders (110) and video decoder (210).
[0090] The convolutional cross-component intra prediction model can be used, for example, in an enhanced compression model (ECM). The convolutional cross-component model (CCCM) can be used to predict chrominance samples from reconstructed luma samples in a manner similar to the current CCLM mode. Similar to CCLM, when chrominance subsampling is used, the reconstructed luma samples can be downsampled to match the lower-resolution chrominance grid. Similar to CCLM, the top reference sample, the left reference sample, or both the top and left reference samples can be used as templates for model derivation.
[0091] Similar to CCLM, there can be an option to use single-model or multi-model variants of CCCM. The multi-model variant can use two models, one model derived for samples above an average luma reference value and another model derived for the remaining samples (following the spirit of the CCLM design). For example, for a PU with at least 128 available reference samples, the multi-model CCCM mode can be selected.
[0092] The convolutional filter (e.g., a convolutional 7-tap filter) can include (e.g., consist of) a 5-tap plus-shaped spatial component, a non-linear term, and a bias term. The input to the spatial 5-tap component of the filter can include (e.g., consist of) the central (C) luma sample co-located with the chrominance sample to be predicted and its above or north (N) neighbor, below or south (S) neighbor, left or west (W) neighbor, and right or east (E) neighbor, as Figure 4 shown.
[0093] The non-linear term NP can be represented as the second power of the central luma sample C and can be scaled to the sample value range of the content, as described in, for example, Equation 1. NP = (C 2 + midVal) >> bitDepth Equation 1
[0094] That is, for 10-bit content, Equation 2 can be used to calculate the non-linear term NP. NP = (C 2 + 512) >> 10 Equation 2
[0095] The mid value (midVal) is 210 / 2, i.e., 512.
[0096] The bias term B can represent a scalar offset between the input and the output (e.g., similar to the offset term in CCLM), and can be set to an intermediate chroma value (e.g., 512 for 10-bit content).
[0097] The output of the filter can be calculated as the convolution between the filter coefficients c i and the input values, and can be clipped to the range of valid chroma samples using Equation 3. predChromaVal = c 0 C + c 1 N + c 2 S + c 3 E + c 4 W + c 5 P + c 6 B Equation 3
[0098] For example, the filter coefficients c can be calculated by minimizing the mean squared error (MSE) between the predicted chroma samples and the reconstructed chroma samples in the reference region. i . Figure 5 An example of a reference region (and its padding) for deriving filter coefficients according to one aspect of the present disclosure is shown. The reference region can include (e.g., consist of) 6 rows of chroma samples above and to the left of the PU. The reference region can extend one PU width to the right and one PU height below the PU boundary. The region can be adjusted to include only the available samples. The extension of this region can be used to support the "side samples" of the plus-shaped spatial filter and is filled when in the unavailable region.
[0099] The MSE minimization can be achieved by calculating the autocorrelation matrix of the luminance input and the cross-correlation vector between the luminance input and the chroma output. The autocorrelation matrix can be subjected to LDL decomposition, and back-substitution can be used to calculate the final filter coefficients. This process generally follows the calculation of the ALF filter coefficients used in, for example, ECM. However, LDL decomposition is chosen instead of Cholesky decomposition, for example, to avoid using square root operations.
[0100] The autocorrelation matrix can be calculated using the reconstructed values of the luma samples and the chroma samples. The luma samples and the chroma samples can be in the full range (e.g., between 0 and 1023 for 10-bit content), resulting in relatively large values in the autocorrelation matrix. This may use high-bit-depth operations during the model parameter calculation. For each model, a fixed offset can be removed from the luma samples and the chroma samples in each PU. This can reduce the magnitude of the values used in model creation and allow for a reduction in the precision of fixed-point arithmetic. Thus, in some examples, 16-bit decimal precision can be used instead of the 22-bit precision of the original CCCM implementation.
[0101] In some examples, for simplicity, the reference sample values outside the upper left corner of the PU can exactly be used as the offsets (offsetLuma, offsetCb, and offsetCr). The sample values used in model creation and final prediction (e.g., the luma and chroma in the reference region, and the luma in the current PU) can be subtracted by a fixed value as follows: C' = C - offsetLuma; N' = N - offsetLuma; S' = S - offsetLuma; E' = E - offsetLuma; W' = w - offsetLuma; P' = nonLinear(C'); B = midValue = 1 << (bitDepth - 1); and the chroma values are predicted using Equation 4, where offsetChroma is equal to offsetCr for the Cr component and offsetCb for the Cb component respectively. predChromaVal = c 0 C' + c 1 N' + c 2 S' + c 3 E' + c 4 W' + c 5 P' + c 6 B + offsetChroma Equation 4
[0102] In the example, to avoid any additional sample-level operations, the luma offset is removed during the luma reference sample interpolation. This can be achieved, for example, by replacing the rounding term used in the luma reference sample interpolation with an updated offset that includes the rounding term and offsetLuma. The chroma offset can be removed by directly subtracting the chroma offset from the reference chroma samples. As an alternative, the influence of the chroma offset can be removed from the cross-component vectors, resulting in the same result. To add the chroma offset back to the output of the convolution prediction operation, the chroma offset can be added to the bias term of the convolution model.
[0103] The process of calculating CCCM model parameters can use division operations. In some examples, division operations may not be conducive to implementation. Division operations can be replaced by multiplication (with a scaling factor) and shift operations, where the scaling factor and the number of shifts can be calculated based on the denominator. For example, similar to the method used when calculating CCLM parameters.
[0104] The gradient and location based convolutional cross-component model (GL-CCCM) can map luminance values to chrominance values using a filter. The input to the filter includes (for example, consists of) a spatial luminance sample, two gradient values, two location information, a non-linear term, and a bias term. The GL-CCCM method can use gradient and location information instead of the 4 spatially adjacent samples used in the CCCM filter. The GL-CCCM filter for prediction can be described using Equation 5. predChromaVal = c 0 C + c 1 G y + c 2 G x + c 3 Y + c 4 X + c 5 P + c 6 B Equation 5
[0105] Figure 6 Shows an example of spatial samples for GL-CCCM according to one aspect of the present disclosure. Gy and Gx are the vertical gradient and the horizontal gradient respectively, and are calculated using Equation 6. G y = (2N + NW + NE) – (2S + SW + SE) G x = (2W + NW + SW) – (2E + NE + SE) Equation 6
[0106] Y and X are the spatial coordinates of the central luminance sample.
[0107] The remaining parameters can be the same as those of the CCCM tool. The reference region for parameter calculation can be the same as that of the CCCM method.
[0108] The use of the GL-CCCM mode can be signaled using a flag such as a PU-level flag of CABAC coding. In terms of signaling, the GL-CCCM mode can be regarded as a sub-mode of CCCM. For example, the GL-CCCM flag is signaled only when the original CCCM flag is true.
[0109] Similar to CCCM, in some examples, the GL-CCCM tool has six modes for calculating parameters: single-model GL-CCCM from the upper and left templates, single-model GL-CCCM from the upper template, single-model GL-CCCM from the left template, multi-model GL-CCCM from the upper and left templates, multi-model GL-CCCM from the upper template, and multi-model GL-CCCM from the left template. The encoder can perform a search (e.g., sum of absolute transform differences (SATD) search) on the six GL-CCCM modes and the existing CCCM modes to find the best candidate for the full-rate distortion (RD) test.
[0110] The intra-block copy (IBC) mode and the intra-template matching prediction (intraTMP) mode can be used to predict, for example, a block in the current picture based on a reference block (e.g., a reconstructed block) in the current picture.
[0111] In one aspect, the IBC mode is a tool adopted in the HEVC screen content coding (SCC) extension. In some examples, the IBC mode can significantly improve the coding efficiency of screen content materials. Since the IBC mode can be implemented as a block-level coding mode, block matching (BM) can be performed at the encoder to find the best block vector (BV) for each CU. In the IBC mode, the BV can be used to indicate the displacement from the current block to a reference block that has been reconstructed within the current picture. In an example, the BV can be regarded as a motion vector, where the reference picture is the current picture. The luminance BV of an IBC-encoded CU can be integer precision. The chrominance BV can be rounded to integer precision. When combined with the adaptive motion vector resolution (AMVR) mode, the IBC mode can switch between 1-pixel precision and 4-pixel precision (e.g., also referred to as 1-pixel motion vector precision and 4-pixel motion vector precision). In addition to the intra prediction mode and the inter prediction mode, the IBC mode for encoding an IBC-encoded CU can be regarded as a third prediction mode. The IBC mode can be applicable to CUs with a width and height less than or equal to 64 luminance samples.
[0112] On the encoder side, hash-based motion estimation can be performed on the IBC mode. The encoder can perform a rate distortion (RD) check on blocks with a width or height not greater than 16 luminance samples. For non-merge modes, block vector search can first be performed using hash-based search. If the hash search does not return a valid candidate, a block-matching-based local search can be performed.
[0113] In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference block can be extended to all allowed block sizes. The hash key calculation for each position in the current picture can be based on 4×4 sub-blocks. For a current block of a larger size, when all the hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions, it can be determined that the hash key matches the hash key of the reference block. If it is found that the hash keys of at least two reference blocks match the hash key of the current block, the block vector cost of each matching reference block can be calculated, and the one with the minimum cost can be selected.
[0114] In block matching search, the search range can be set to include the previous CTU and the current CTU (e.g., the previously reconstructed CTU and the current CTU).
[0115] At the CU level, the IBC mode can be signaled with a flag and can be signaled as an IBC adaptive motion vector prediction (AMVP) mode or an IBC skip / merge mode as follows.
[0116] For the IBC skip / merge mode: The merge candidate index can be used to indicate which of the block vectors from adjacent candidate IBC-encoded blocks in the merge list is used to predict the current block. The merge list can include at least one spatial candidate, at least one HMVP candidate, and at least one pairwise candidate. In an example, the merge list consists of at least one spatial candidate, at least one HMVP candidate, and at least one pairwise candidate.
[0117] For the IBC AMVP mode: The block vector difference (BVD) can be encoded in the same way as the motion vector difference. The block vector prediction method can use two candidates as prediction values, one from the left neighbor (if it is IBC-encoded) and one from the upper neighbor (if it is IBC-encoded). When either neighbor is not available, the default BV can be used as the prediction value. A flag can be signaled to indicate the block vector prediction value index.
[0118] Figure 7AAn example of the Intra-frame Template Matching Prediction (IntraTMP) mode according to one aspect of the present disclosure is shown. In one aspect, for example, in an Enhanced Compression Model (ECM) software, IntraTMP is a special intra-frame prediction mode that can copy the best prediction block (e.g., the matching block (721)) from the reconstructed part of the current frame (or current picture), where the template (e.g., the L-shaped template) (720) of the best prediction block can match the current template (730) of the current block (731) (e.g., the current PU or current CU). For a predefined search range, the encoder can search in the reconstructed part of the current frame for the template that is most similar to the current template and can use the corresponding block as the prediction block. The encoder can signal the use of the IntraTMP mode and can perform the same prediction operation on the decoder side.
[0119] A prediction signal can be generated by matching the current template (730) (e.g., the L-shaped causal neighbor of the current block (731)) with the template of another block in a predefined search area. Figure 7A The example search area shown in Figure 7A can include at least two CTUs (or super blocks). Referring to
[0120] In each area, the decoder can search for the template with the minimum cost (e.g., minimum SAD) relative to the current template and can use the block associated with the template having the minimum cost as the prediction block.
[0121] The dimensions of the area indicated by (SearchRange_w, SearchRange_h) can be set to be proportional to the block dimensions (BlkW, BlkH) so that each pixel has a fixed number of SAD comparisons. Thus, SearchRange_w = a × BlkW, SearchRange_h = a × BlkH.
[0122] The parameter "a" can be a constant used to control the trade-off between gain and complexity. In one example, "a" is 5.
[0123] In an example, to accelerate the template matching process, the search range (e.g., the search ranges of all search regions) is subsampled by a factor of 2, which results in a 4-fold reduction in the template matching search. After finding the best match (or initial best match), a refinement process can be performed. The refinement is done by performing a second template matching search in a reduced range around the best match (or initial best match). The reduced range is defined as min(BlkW, BlkH) / 2.
[0124] The intra-frame template matching tool can be enabled for CUs with width and height less than or equal to 64. The maximum CU size (e.g., 64) for intra-frame template matching is configurable.
[0125] The filtered intra-block copy (FIBC) model can be used together with, for example, the IBC mode or the IntraTMP mode. In one aspect, the predicted samples of the IBC mode or the IntraTMP mode can be enhanced by applying a linear filter. Return reference Figure 4 , in the example, the linear filter consists of five spatial terms and one bias term, as shown in Equation 7. These five spatial terms consist of the center (C) position, the upper / north neighbor (N), the lower / south neighbor (S), the left / west neighbor (W), and the right / east neighbor (E). predVal = α 0 ·C + α 1 ·N + α 2 ·S + α 3 ·W + α 4 ·E + α 5 ·β Equation 7
[0126] α i is a coefficient (e.g., i ranges from 0 to 5), and β is an offset related to the bias term. Up to 4 rows / columns of samples above and to the left of the current CU can be applied to derive the filter coefficients (e.g., including α 0 to α 5 ). The filter coefficients can be derived based on the minimization of the difference between the template samples and the corresponding reference samples via a regression-based minimization technique (e.g., the same regression-based minimization technique used in other tools such as CCCM described in the present disclosure and used in ECM).
[0127] For signaling, an additional indication flag (also referred to as the FIBC flag) can be introduced for the FIBC mode, and this additional indication flag can be signaled after the IBC-local illumination compensation (LIC) flag. In the example, when the IBC-LIC flag is true, the FIBC flag can be signaled and used to indicate whether the FIBC mode is applied to the current block.
[0128] In the FIBC mode described above, a filtered IBC model can be used, where the prediction samples of the IBC mode are enhanced by applying a linear filter. Up to 4 rows / columns of samples above and to the left of the current CU can be applied to derive the filter coefficients. In various examples, the filter design in FIBC as described in Equation 7 may not be accurate enough because the filter consists only of linear terms of sample values and does not use additional information such as gradient information and position information. In addition, the filter in the above FIBC mode does not include non-linear terms.
[0129] Aspects of the present disclosure provide techniques for improving the FIBC mode by including at least one of gradient information, position information, non-linear terms, etc. when deriving filter coefficients for the FIBC mode.
[0130] The methods, aspects, and examples described in the present disclosure can be used alone or in any order of combination. The term "IBC mode" may refer to the IBC mode or variants described in the present disclosure. The term "IntraTMP mode" may refer to the IntraTMP mode or variants described in the present disclosure. In addition, the methods, aspects, and examples can be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit). In one example, the at least one processor executes a program stored in a non-volatile computer-readable medium.
[0131] Figure 7B An example of a modification of the FIBC mode according to an aspect of the present disclosure is shown. One of the IBC mode and the IntraTMP mode can be used to predict the current block (701) in the current picture (700). In the example, one of the IBC mode and the IntraTMP mode is used to determine the BV (702). The BV (702) may indicate the reference block (703) of the current block (701). The reference block (703) is located in the current picture (700) and may have been reconstructed. The reference samples in the reference block (703) have been reconstructed and can be used to predict the current block (701). Therefore, the reference samples in the reference block (703) can be referred to as IBC prediction samples. According to an aspect of the present disclosure, in the FIBC mode (e.g., an updated FIBC mode modified according to Equation 7), the IBC prediction samples can be filtered by an FIBC filter. Thus, a filtered reference block (703) can be generated and can include the IBC prediction samples filtered by the FIBC filter. The IBC prediction samples filtered by the FIBC filter can be used to predict the current samples in the current block (701).
[0132] Refer to Figure 7B, in the example, the FIBC mode updated based on additional information (such as gradient information, position information, and / or non - linear terms) will be used to predict the current sample (710) in the current block (701). The reference sample (711) in the reference block (703) corresponds to the current sample (710). For example, BV (702) indicates the displacement between the reference sample (711) and the current sample (710). The IBC predicted sample (711) filtered by the FIBC filter can be used to predict the current sample (710). After being filtered by the FIBC filter, the reference sample (711) can be referred to as the FIBC - filtered predicted sample (711). For example, the predicted value of the current sample (710) is the value of the FIBC - filtered predicted sample (711). The value of the FIBC - filtered predicted sample can also be referred to as the predicted value because, for example, this value can be directly used as the predicted value of the current sample (710).
[0133] In one aspect, in addition to the linear filter similar or identical to the filter in Equation 7, the FIBC filter can further include at least one of the following: (i) at least one gradient filter (e.g., indicating the gradient information related to each IBC predicted sample), (ii) a position filter (e.g., indicating the position information related to the linear filter), (iii) non - linear terms (also called non - linear values) related to each IBC predicted sample, etc.
[0134] The coefficients of the FIBC filter (also called filter coefficients) can include the linear coefficients of the linear filter and at least one of the following: (i) the gradient coefficients of at least one gradient filter, (ii) the position coefficients of the position filter, (iii) at least one non - linear coefficient of the non - linear terms, etc. If other information is included in the FIBC filter, additional coefficients can be included in the filter coefficients of the FIBC filter.
[0135] In one aspect, the coefficients of the FIBC filter can be determined according to the current template (704) of the current block (701) and the reference template (705) of the reference block (703).
[0136] The details of the method, aspects, and examples of the FIBC mode using the FIBC filter are described below.
[0137] In one aspect, gradient information including the gradients (G x , G y ) of adjacent reconstruction samples can be used to filter the IBC predicted samples. The predicted value pred0(x,y) at (x, y) can be determined (e.g., defined) using Equations 8 to 12. pred0(x,y) = P0 + GX+GY + B Equation 8
[0138] pred0(x,y) can be the predicted value at (x, y). The four parameter sets associated with P0, GX, GY, and B can respectively correspond to the input sample value, horizontal gradient information, vertical gradient information, and bias. P0 can be a linear prediction value obtained using a linear filter as described in Equation 9. GX can be a horizontal gradient value obtained using a horizontal gradient filter as described in Equation 10. GY can be a vertical gradient value obtained using a vertical gradient filter as described in Equation 11. B can be a bias term as described in Equation 12. In an example, the linear filter can include a bias term. In the example shown by Equation 8, the at least one gradient filter can include a horizontal gradient filter and a vertical gradient filter.
[0139] In an example, the current sample (710) is located at the position (x, y) in the current block (701). The reference sample or IBC predicted sample (711) can be located at the same position (x, y) in the reference block (703). By applying the FIBC filter as described in Equations 8 - 12 to the IBC predicted sample (711) at (x, y), the predicted value pred0(x,y) at the position (x, y) can be determined.
[0140] The reconstructed sample located at the position (x - xoffset 0,i , y - yoffset 0,i ) in the reference block (703) can be referred to as the first input sample or N0 input sample, where i ranges from 1 to N0. The reconstructed sample located at the position (x - xoffset 1,j , y - yoffset 1,j ) in the reference block (703) can be referred to as the second input sample or N1 input sample, where j ranges from 1 to N1. The reconstructed sample located at the position (x - xoffset 2,k , y - yoffset 2,k ) can be referred to as the third input sample or N2 input sample, where k ranges from 1 to N2. The N0 input sample, N1 input sample, and N2 input sample can be used to obtain the predicted value pred0(x,y). The N0 input sample, N1 input sample, and N2 input sample can include: (i) adjacent reconstructed samples of the reference sample (711) (e.g., Figure 7BThe samples N, S, W, and E in the reference block (703) shown, (ii) reconstructed samples not adjacent to the reference sample (711), etc. After determining the predicted value pred0(x, y), for example, for the reference sample (711), the current sample (710) can be predicted based on the predicted value pred0(x, y). For example, the predicted value of the current sample (710) is equal to pred0(x, y). The current sample (710) can be reconstructed based on the predicted value of the current sample (710).
[0141] The first set or the first set of parameters can be related to the sample value (e.g., the N1 input sample value), which can include the parameters of the N0 input samples (e.g., including the coefficient c 0,i ). (x - xoffset 0,i , y - yoffset 0,i ) is the position of the i-th input sample, c 0,i is the coefficient of the i-th input sample, and t(x - xoffset 0,i , y - yoffset 0,i ) is the value of the i-th input sample.
[0142] The linear filter used in the FIBC filter can be based on the coefficient c 0,i of the N0 input samples and the position. The linear coefficients can include the coefficients {c 0,i}. The shape of the linear filter can be based on the offsets {(xoffset 0,i , yoffset 0,i )} relative to the reference position (such as (x, y)).
[0143] Reference Figure 7B , if the N0 input samples include the N, S, W, and E samples in the reference block (703), then N0 is 4. If the N0 input samples include the samples N, S, W, E, and (711) in the reference block (703), then N0 is 5. For the N sample, xoffset 0,i is 0, and yoffset 0,i is –1.
[0144] The second set or the second set of parameters can be related to the horizontal gradient information, which can include the parameters of at least one N1 input sample. c 1,j is the coefficient of the j-th input sample, and G x (x - xoffset 1,j , y - yoffset 1,j ) is the value of the horizontal gradient G x of the j-th input sample. In an example, the at least one N1 input sample can be the same as the N0 input sample. In an example, the at least one N1 input sample can be different from the N0 input sample.
[0145] The horizontal gradient filter used in the FIBC filter can be based on the coefficients c of N1 input samples 1,j and positions. The horizontal gradient coefficients can include the coefficients {c 1,j}. The shape of the horizontal gradient filter can be determined based on the offset {(x - xoffset 1,j , y - yoffset 1,j )} relative to a reference position (such as (x, y)).
[0146] A third set or third parameter set can be related to the vertical gradient information, which can include parameters of at least one N2 input sample. c 2,k is the coefficient of the kth input sample, and G y (x - xoffset 2,k , y - yoffset 2,k ) is the value of the vertical gradient G y of the kth input sample. In an example, the at least one N2 input sample can be the same as the N0 input sample. In an example, the at least one N2 input sample can be different from the N0 input sample. In an example, the at least one N2 input sample can be the same as the N1 input sample. In an example, the at least one N2 input sample can be different from the N1 input sample.
[0147] The vertical gradient filter can be based on the coefficients c of N2 input samples 2,k and positions. The vertical gradient coefficients can include the coefficients {c 2,k}. The shape of the vertical gradient filter can be determined based on the offset {(x - xoffset 2,k , y - yoffset 2,k )} relative to a reference position (such as (x, y)).
[0148] A fourth set or third parameter set can be related to the bias B and thus can include parameters of respective biases b l . b l is the lth bias. c 3,l is the coefficient of the lth bias b l . In an example, N3 can be zero and the predicted value pred0(x, y) does not include a bias. In an example, the linear filter can include the bias B.
[0149] According to one aspect of the present disclosure, the linear prediction value of the current sample in the current block can be determined by applying a linear filter to the predicted sample predicted using one of the IBC mode and the IntraTMP (Intra Template Matching) mode, as described in Equation 9. The gradient value associated with the current sample in the current block can be determined using at least one gradient filter (such as a horizontal gradient filter and / or a vertical gradient filter), as described in Equations 10-11. The predicted value of the current sample can be determined based on the sum of the linear prediction value and at least one modified value including the gradient value. The gradient value is determined based on at least one of GX and GY. The FIBC filter in the FIBC mode can include a linear filter and at least one gradient filter. The current sample can be reconstructed according to the predicted value of the current sample.
[0150] In one aspect, the available filters in the FIBC mode can have any suitable shape and / or size. In an example, the linear filter has a cross shape and includes 5 samples, and these 5 samples include a central sample with an offset of (0, 0) of the linear filter, a north sample N with an offset of (0, -1), a south sample S with an offset of (0, 1), an east sample E with an offset of (1, 0), and a west sample W with an offset of (-1, 0), as Figure 8 shown. The offsets of the 5 samples in the linear filter are relative to the central sample. When the linear filter has 5 samples, the current sample can be located at one of the 5 positions of the corresponding 5 samples.
[0151] In an example, the linear filter has a cross shape and includes 9 samples, and these 9 samples include a central sample with an offset of (0, 0) of the linear filter, two north samples with offsets of (0, -1) and (0, -2) respectively, two south samples with offsets of (0, 1) and (0, 2) respectively, two east samples with offsets of (1, 0) and (2, 0) respectively, and two west samples with offsets of (-1, 0) and (-2, 0) respectively, as Figure 9 shown. The offsets of the 9 samples in the linear filter are relative to the central sample. When the linear filter has 9 samples, the current sample is located at one of the 9 positions of the corresponding 9 samples.
[0152] Figures 8 - 10 Examples of available filters in the FIBC mode according to aspects of the present disclosure are shown. In an example, the available filters include Figures 8 - 9 the cross-shaped filters (801)-(802) shown. In Figure 8 the example shown, when the filter (801) is applied to the N0 input samples, N0 is 5, and the five samples (for example, five input samples) include: a central sample (C), where c 0,i is the coefficient of C, xoffset0,i = 0, yoffset 0,y = 0 (e.g., i = 1); North sample (N), where c 0,i is the coefficient of N, xoffset 0,i = 0, yoffset 0,i = –1 (e.g., i = 2); South sample (S), where c 0,i is the coefficient of S, xoffset 0,i = 0, yoffset 0,i = 1 (e.g., i = 3); East sample (E), where c 0,i is the coefficient of E, xoffset 0,i = 1, yoffset 0,i = 0 (e.g., i = 4); and West sample (W), where c 0,i is the coefficient of W, xoffset 0,i = –1, yoffset 0,i = 0 (e.g., i = 5).
[0153] Figure 9 The shown filter shape (802) can be interpreted in a similar way to the shape of filter (801). For example, if the filter shape (802) is applied to N0 input samples, then N0 is 9, and the N0 input samples include C, N, S, E, W which are the same as the samples (e.g., C, N, S, E, and W) described in Figure 8 . The N0 input samples in Figure 9 further include samples NN, SS, EE, and WW. In the example, the offset of the NN sample is (0, –2).
[0154] In the above example as in Figures 8 - 9 , the position of the current sample and the position of the corresponding reference sample can be at the center sample of the filter (e.g., (801) or (802)), so the offset of the center sample of the filter (e.g., (801) or (802)) with respect to (x, y) is (0, 0). The position of the current sample and the position of the corresponding reference sample can be not limited to the positions shown in the example (e.g., the center position), and other input samples can be used as reference samples or current samples. Referring to the filter (801) in Figure 8 , the current sample or the corresponding reference sample can be at one of the 5 positions of the corresponding 5 input samples (e.g., C, N, S, E, or W). In the example, the reference sample is at N, so the offset of N with respect to (x, y) is (0, 0), the offset of C is (0, 1), and the offsets of the other samples are shifted accordingly. Referring toFigure 9 For the filter (802) in [description], the current sample or the corresponding reference sample can be located at one of the nine positions of the nine input samples. In the example, the reference sample corresponding to the current sample is located at W, so the offset of W is (0, 0), the offset of C is (1, 0), and the offsets of the other samples are shifted accordingly.
[0155] In one aspect, for the first parameter set, the second parameter set, and the third parameter set associated with P0, GX, and GY respectively, for each sample i (e.g., each of the N0 input samples), each G x element j and each G y element k, the allowed values of the predefined offset (e.g., indicated by xoffset and yoffset) can be predefined. In the example, the positions of the N0 input samples are predefined. Therefore, the shape of the linear filter is predefined. The shape of at least one of the at least one gradient filter is predefined. In the example, the position of the N1 input samples and thus the shape of the horizontal gradient filter are predefined. In the example, the position of the N2 input samples and thus the shape of the vertical gradient filter are predefined.
[0156] In one aspect, the distribution of the current sample and the adjacent samples of the current sample defined by (x - xoffset, y - yoffset) does not have to be continuous. As described above, the position of the current sample in the current block and the position of the corresponding reference sample in the reference block are (x, y), the distribution of the reference sample and the adjacent samples of the reference sample is defined by (x - xoffset, y - yoffset), and does not have to be continuous. In the example, the input samples in one of the linear filter, the horizontal gradient filter, and the vertical gradient filter are spatially separated from all the remaining input samples in this filter. For example, the samples in the linear filter are spatially separated from all the remaining samples in the linear filter.
[0157] In the example, none of the input samples in one of the linear filter, the horizontal gradient filter, and the vertical gradient filter are adjacent to another input sample in this filter, as Figure 10 shown.
[0158] Figure 10 An example of a filter shape (803) without continuous samples according to one aspect of the present disclosure is shown. The filter shape or filter (803) can include five input samples located at C, +2 x , –2 x , +2 y and –2 y The four sample positions (e.g., +2 x , –2 x , +2 yand –2 y Each of them in y is shifted 2 samples from C, so the offset is 2, and x / y is the direction indicator. For example, the position +2 x has an offset from C of (+2, 0). The filter shape (803) can be used to derive P0, GX, and / or GY.
[0159] In one aspect, several variants of the linear prediction value P0 can be used. For example, when feeding an input (e.g., N0 input samples) to the FIBC input reference samples, a mean removal operation can be applied to P0. P1 represents a variant of the linear prediction value, which can be P0 after the mean removal operation is applied, as described in Equation 13. The mean removal operation can be similar to the method used in CCCM. Then, Equations 13 - 14 can be used to generate the predicted value pred1(x,y) of (x, y). pred1(x,y) = P1 + GX + GY Equation 14
[0160] In an example, the linear filter (described in Equation 13) adds the mean of the current block (e.g., mean in Equation 13), and removes the mean of the current block from each sample predicted using one of the IBC mode and IntraTMP mode (e.g., the IBC predicted sample t(x - xoffset 0,i ,y - yoffset 0,i ))
[0161] In an example, a clipping operation can be applied to generate the final prediction pred2(x,y). pred2(x,y) = clip(min, max, predi(x,y)) Equation 15
[0162] The min value and max value can be the minimum and maximum values of the samples in the template respectively. predi(x,y) can be pred0(x,y), pred1(x,y), etc. In an example, the predicted value predi(x,y) of the current sample can be clipped, as shown in Equation 15.
[0163] Details related to the gradient information are further described below.
[0164] The number and positions of different parameter sets can be set independently. In an example, the positions of N0 and N0 input samples are set independently. In an example, the positions of N1 and N1 input samples are set independently. In an example, the positions of N2 and N2 input samples are set independently. In an example, the linear filter, the horizontal gradient filter, and the vertical gradient filter can be set independently.
[0165] In one aspect, both GX and GY can be used as inputs to the filter in the FIBC mode, N1 is equal to N2, and the same sample positions are used as the inputs for both sets. In an example, the horizontal gradient filter and the vertical gradient filter have the same size and the same shape.
[0166] In an example, when at least one gradient filter includes a horizontal gradient filter and a vertical gradient filter, the gradient value is the sum of the horizontal gradient value (e.g., GX) and the vertical gradient value (e.g., GY), and the number N1 of the first input samples (e.g., N1 input samples) and the position of the first input samples in the horizontal gradient filter, and the number N2 of the second input samples (e.g., N2 input samples) and the position of the second input samples in the vertical gradient filter are set independently of each other and independently of the linear filter settings (e.g., based on N0 and N0 input samples).
[0167] In one aspect, both GX and GY can be used as inputs to the filter in the FIBC mode, N1 is equal to N2, but different sample positions are used as the inputs for GX and GY respectively. In an example, the position of the first input samples (e.g., N1 input samples) in the horizontal gradient filter and the position of the second input samples (e.g., N2 input samples) in the vertical gradient filter are different. For example, the horizontal gradient filter and the vertical gradient filter have the same size and different shapes.
[0168] When at least one gradient filter consists of a horizontal gradient filter, the gradient value is the horizontal gradient value, and the number N1 of the first input samples and the position of the first input samples in the horizontal gradient filter are independent of the linear filter settings.
[0169] In one aspect, in the FIBC mode, only GX can be used as the input to the filter. In an example, when at least one gradient filter consists of a horizontal gradient filter, the gradient value is the horizontal gradient value (e.g., GX), and the number N1 of the first input samples (e.g., N1 input samples) and the position of the first input samples in the horizontal gradient filter are independent of the linear filter settings (e.g., based on N0 and N0 input samples). For example, pred0(x,y) = P0 + GX or pred1(x,y) = P1 + GX.
[0170] When at least one gradient filter consists of a vertical gradient filter, the gradient value is the vertical gradient value, and the number of the second input samples and the position of the second input samples in the vertical gradient filter are independent of the linear filter settings.
[0171] In one aspect, in the FIBC mode, only GY can be used as the input to the filter. In an example, when at least one of the gradient filters consists of a vertical gradient filter, the gradient value is a vertical gradient value (e.g., GY), and the number N2 of the second input samples (e.g., N2 input samples) and the positions of the second input samples in the vertical gradient filter are independent of the linear filter settings (e.g., based on N0 and N0 input samples). For example, pred0(x,y) = P0 + GY or pred1(x,y) = P1 + GY.
[0172] The gradient value can be determined based on at least one of a horizontal gradient value (GX) or a vertical gradient value (GY), where the horizontal gradient value (GX) is based on the horizontal gradient G x ) of each of the first input samples (e.g., N1 input samples), and the vertical gradient value (GY) is based on the vertical gradient G y of each of the second input samples (e.g., N2 input samples), as shown in Equation 8, Equation 10, and Equation 11.
[0173] The gradient associated with the input sample, such as the horizontal gradient G x or the vertical gradient G y , can be calculated using any suitable method. Some examples are described below.
[0174] Figure 11 Shows an example of adjacent samples of the input sample C for calculating the gradient according to one aspect of the present disclosure. Given the input sample C and the adjacent samples of the input sample C (e.g., using those represented by "N", "S", "E", "W", "NW", "NE", "SW", and "SE" as described above in Figures 8 - 9 ), a single gradient (e.g., G x and G y ) can be calculated as follows.
[0175] In one aspect, G x and G y are calculated based on samples C, W, and N: G y = (C - N) G x = (C - W) Equation 16
[0176] In one aspect, G x and G y can be calculated based on samples N, S, W, and E. G y = (N - S) G x = (W - E) Equation 17
[0177] In one aspect, G can be calculated based on N, S, W, E, NW, SW, NE, and SE x and G y . G y =(2N + NW + NE) – (2S + SW + SE) G x =(2W + NW + SW) – (2E + NE + SE) Equation 18
[0178] In one aspect, at least one gradient filter includes a horizontal gradient filter. The gradient value includes a horizontal gradient value (e.g., GX), which is the sum of the horizontal gradients G of each first input sample in the horizontal gradient filter (e.g., Equation 10). Each horizontal gradient of each first input sample can be determined based on one of the following: (i) the difference between each first input sample and its left neighbor (e.g., Equation 16); (ii) the difference between the left neighbor of the first input sample and its right neighbor (e.g., Equation 17); and (iii) the difference between a first value (2W + NW + SW) and a second value (2E + NE + SE), as shown in Equation 18. The first value is a weighted sum determined based on the upper left neighbor, left neighbor, and lower left neighbor of the first input sample. The second value is a weighted sum determined based on the upper right neighbor, right neighbor, and lower right neighbor of the first input sample. x Which difference to use to calculate the horizontal gradient of each first input sample can be determined based on the position of each first input sample.
[0179] In an example, at least one gradient filter includes a vertical gradient filter. The gradient value includes a vertical gradient value, which is the sum of the vertical gradients of each second input sample in the vertical gradient filter (e.g., Equation 11). Each vertical gradient of each second input sample can be determined based on one of the following: (i) the difference between each second input sample and its top neighbor; (ii) the difference between the top neighbor of each second input sample and its bottom neighbor; and (iii) the difference between a first value and a second value, as shown in Equation 18. The first value is a weighted sum determined based on the upper left neighbor, top neighbor, and upper right neighbor of the second input sample. The second value is a weighted sum determined based on the lower left neighbor, bottom neighbor, and lower right neighbor of each second input sample.
[0180] Which difference to use to calculate the vertical gradient of each second input sample is determined based on the position of each second input sample.
[0181]
[0182] In one aspect, G can be calculated according to the position of the input sample Cx and G y 。 Figure 12 illustrates an example of selecting a gradient calculation method (e.g., one of the methods described in Equations 16 - 18) based on the position of an input sample according to an aspect of the present disclosure. Different gradient calculation methods can be used to calculate the gradient {G x , G y} of the input sample (e.g., input samples (1201)-(1209)) according to the position of the input sample.
[0183] Input samples (1201)-(1205) are located at Type 1 positions. For each of the input samples (1201)-(1205) at Type 1 positions, the sample C, W, and N of each input sample can be used to calculate G x and G y of the input sample, for example, using Equation 16. For example, the sample C, W, and N of input sample (1202) can be samples (1202), (1203), and (1206) respectively. For example, the sample C, W, and N of input sample (1204) can be samples (1204), (1206), and (1205) respectively.
[0184] Input sample (1206) is located at a Type 2 position. Samples N, S, W, and E can be used to calculate G x and G y of the input sample (1206) at the Type 2 position, for example, using Equation 17.
[0185] Input samples (1207)-(1209) are located at Type 3 positions. For each of the input samples (1207)-(1209) at Type 3 positions, the corresponding samples N, S, W, E, NW, SW, NE, and SE of each input sample can be used to calculate G x and G y of the input sample, for example, using Equation 18.
[0186] The position of the central luminance sample can be used to filter the IBC prediction samples. In an example, the central luminance sample is within a filter (e.g., filter (801), filter (802), filter (803), etc.). In an example, the central luminance sample is within a linear filter (e.g., filter (801), filter (802), filter (803), etc.). When applying the filter to each input sample, the position of the central luminance sample may move. Referring to Figure 8 , the central luminance sample can be at position C.
[0187] In one aspect, the predicted value includes, but is not limited to, pred0(x,y) or pred1(x,y). The predicted value can be further refined by using Equation 19 to add information on the coordinates of the central luminance sample.
[0188] predi(x,y) can be an original prediction (e.g., a predicted value without position information), including but not limited to pred0(x,y), pred1(x,y), etc. LX and LY can be the position information of the central luminance sample in the horizontal and vertical directions respectively. LX and LY can be obtained (e.g., defined) using Equation 20. LX = c 4 X LY = c 5 Y Equation 20
[0189] c 4 and c 5 are coefficients of the position information, and the X parameter and Y parameter are the spatial coordinates of the central luminance sample.
[0190] As described in Equation 20, the position value (e.g., LX + LY) can be determined using the position of the central sample located at the center of the filter (e.g., a linear filter).
[0191] Based on Equation 19, when predi(x,y) = pred0(x,y), when, if B is 0, then
[0192] Based on Equation 19, when predi(x,y) = pred1(x,y), when,
[0193] Therefore, the predicted value of the current sample can be determined based on the sum of a linear predicted value (e.g., P0 or P1) and at least one modified value including a gradient value (e.g., GX + GY) and a position value (e.g., LX + LY). The FIBC filter in the FIBC mode includes a linear filter (e.g., for P0 or P1), at least one gradient filter (e.g., for GX and / or GY), and coefficients for the position (e.g., c 4 and c 5 ).
[0194] The IBC predicted samples can be filtered using a non - linear term derived from adjacent reconstructed samples.
[0195] In one aspect, the predicted value predi(x,y) can be further refined by adding a non-linear term (also referred to as a non-linear value) NP using Equation 21.
[0196] predi(x,y) is the original prediction (e.g., the predicted value without position information), including but not limited to pred0(x,y), pred1(x,y), etc. In an example, Equation 21 can be used to obtain a refined prediction based on the predicted value predi(x,y) (e.g., pred0(x,y), pred1(x,y), etc.) and the non-linear term.
[0197] In one example, Equation 22 is used to generate the non-linear term based on the original predicted value of the current prediction sample. NP = (C × C + midVal) >> bitDepth Equation 22
[0198] C in Equation 22 can be the original predicted value of the current sample, which can be the IBC predicted value of the current sample obtained from one of the IBC mode and the IntraTMP mode. Thus, C in Equation 22 can be the value of the reference sample (IBC prediction sample) without the FIBC filter.
[0199] In an example, the middle value (midVal) of 10-bit content is 2 10 / 2, i.e., 512. Thus, for 10-bit content, the non-linear term NP can be calculated using Equation 23. NP = (C × C + 512) >> 10 Equation 23
[0200] Then, a clipping operation can be applied to generate the final prediction pred2(x,y) using Equation 24.
[0201] In another example, the non-linear term can be generated based on the reconstructed value or the previous predicted value in the immediate neighborhood of the current sample and the current prediction sample. Figure 13 An example of the respective positions of samples A, L, and AL according to one aspect of the present disclosure is shown. Samples A, L, and AL are adjacent samples of the current sample C. M1 can be the average value of the values of samples A, L, AL, and C, e.g., M1 = mean(A, L, AL, C). M1 can be the median value of the values of samples A, L, AL, and C, e.g., M1 = median(A, L, AL, C). The non-linear term NP is defined using Equation 25. NP = (M1 × M1 + midVal) >> bitDepth Equation 25
[0202] In another example, the non-linear term can be generated based on the reconstructed value or the previous predicted value in the immediate neighborhood of the current predicted sample. The non-linear term NP is defined using Equation 26. NP = (M2 × M2 + midVal) >> bitDepth Equation 26
[0203] M2 can be the average value of the values of samples A, L, and AL, e.g., M2 = mean(A, L, AL). M2 can be the median value of the values of samples A, L, and AL, e.g., M2 = median(A, L, AL).
[0204] The non-linear value or non-linear term NP associated with the current sample can use the non-linear relationship (e.g., Equation 23, Equation 25, and Equation 26) between the non-linear value NP and the value of at least one of the current sample and the adjacent samples of the current sample (e.g., C, A, L, and AL), and is determined according to the current sample (e.g., Figure 13 C in Figure 13 and at least one of the adjacent samples (e.g., A, L, AL in 6 ). The predicted value of the current sample
[0205] can be determined based on the sum of the linear predicted value (e.g., P0 or P1) and at least one modified value (including the gradient value (e.g., GX + GY) and the non-linear value NP). The FIBC filter in the FIBC mode includes a linear filter, at least one gradient filter, and the coefficient of the non-linear value (e.g., c 6 ).
[0205] In one aspect, the coefficient (e.g., the filter coefficient of the FIBC filter) is derived from the template (e.g., the current template) of the coded block (e.g., the current block) and the corresponding reference template. Return reference Figure 7B , in one aspect, the filter coefficient of the FIBC filter can be determined based on the current template (704) of the current block (701) and the reference template (705) of the reference block (703). In the example, up to 4 rows and 4 columns of samples above and to the left of the current block (701) can be applied to derive the filter coefficient of the FIBC filter.
[0206] In one aspect, the filter coefficient of the FIBC filter includes the linear filter coefficient {c 0,i} of the linear filter and at least one of the following: (i) the gradient filter coefficient {c 1,j} of the horizontal gradient filter and / or the gradient filter coefficient {c 2,k} of the vertical gradient filter, (ii) the position coefficients c 4 and c 5 of the position filter, and (iii) the non-linear coefficient c 6 of the non-linear term。In an example, if a bias B is included, the filter coefficients include {c 3,l}. In an example, the linear filter includes {c 3,l}. If other information is included in the FIBC filter, additional coefficients may be included in the filter coefficients of the FIBC filter.
[0207] The filter coefficients of the FIBC filter may be determined (e.g., derived) based on the current template (704) of the current block (701) and the reference template (705) of the reference block (703). In Figure 7B the example shown, the current template (704) includes 4 rows above the current block (701) and 4 columns to the left of the current block (701). The samples in the current template (704) have been reconstructed. The predicted samples of each sample in the current template (704) can be generated by applying the FIBC filter (e.g., described by at least one of Equation 8-15, Equation 19, and Equation 21) to the corresponding samples in the reference template (705). For example, the filter coefficients are determined by minimizing an error function (e.g., MSE) between the predicted samples and the reconstructed samples in the current template, as used in CCCM.
[0208] In one aspect, the LDL method (also known as LDL decomposition) or variance used in CCCM is used to determine (e.g., derive) these coefficients. In LDL decomposition, the matrix A can be decomposed into A = LDL T . L is a lower unit triangular (unit triangular) matrix, D is a diagonal matrix, and L T is the transpose of L. The LDL decomposition can be a closely related variant of the Cholesky decomposition. In some examples, the LDL decomposition may be advantageous over the Cholesky decomposition because the LDL decomposition can avoid taking square roots.
[0209] Figure 14 FIG. shows a flowchart of an overview process (1400) according to an embodiment of the present disclosure. The process (1400) can be used in a device such as a video decoder. In various embodiments, the process (1400) is executed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), etc. In some embodiments, the process (1400) is implemented as software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1400). The process starts at (S1401) and proceeds to (S1410).
[0210] At (S1410), encoded information is received, and the encoded information indicates that the current block in the current picture is predicted using the filtered intra block copy (FIBC) mode.
[0211] In (S1420), a linear prediction value of a current sample in the current block is determined by applying a linear filter to a prediction sample predicted using one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode.
[0212] In an example, the linear filter includes a bias term.
[0213] In an example, the linear filter is configured to add an average value of the current block and remove the average value of the current block from each of the samples predicted using one of the IBC mode and the IntraTMP mode.
[0214] In an example, the linear filter has a cross shape, and the cross shape includes: (i) five samples, the five samples including a center sample with an offset of (0, 0) of the linear filter, a north sample N with an offset of (0, -1), a south sample S with an offset of (0, 1), an east sample E with an offset of (1, 0), and a west sample W with an offset of (-1, 0), and the offsets of the five samples in the linear filter are relative to the center sample; or (ii) nine samples, the nine samples including a center sample with an offset of (0, 0) of the linear filter, two north samples with offsets of (0, -1) and (0, -2) respectively, two south samples with offsets of (0, 1) and (0, 2) respectively, two east samples with offsets of (1, 0) and (2, 0) respectively, and two west samples with offsets of (-1, 0) and (-2, 0) respectively, and the offsets of the nine samples in the linear filter are relative to the center sample.
[0215] When the linear filter has the five samples, the current sample is located at one of the five positions of the five samples. When the linear filter has the nine samples, the current sample is located at one of the nine positions of the nine samples.
[0216] In (S1430), at least one gradient filter is used to determine a gradient value associated with a current sample in the current block.
[0217] In an example, when at least one gradient filter consists of a horizontal gradient filter, the gradient value is a horizontal gradient value, and the number of first input samples and the positions of the first input samples in the horizontal gradient filter are independent of the linear filter setting. When at least one gradient filter consists of a vertical gradient filter, the gradient value is a vertical gradient value, and the number of second input samples and the positions of the second input samples in the vertical gradient filter are independent of the linear filter setting. When at least one gradient filter includes both a horizontal gradient filter and a vertical gradient filter, the gradient value is the sum of the horizontal gradient value and the vertical gradient value, and the number of first input samples and the positions of the first input samples in the horizontal gradient filter, and the number of second input samples and the positions of the second input samples in the vertical gradient filter are set independently of each other and independently of the linear filter setting.
[0218] In an example, at least one gradient filter includes a horizontal gradient filter. The gradient value includes a horizontal gradient value, which is the sum of the horizontal gradients of the respective first input samples in the horizontal gradient filter. Each horizontal gradient of each respective first input sample is determined based on one of the following: (i) the difference between each respective first input sample and the left neighbor of the respective first input sample; (ii) the difference between the left neighbor of the first input sample and the right neighbor of the first input sample; and (iii) the difference between a first value and a second value, where the first value is the sum determined based on the upper left neighbor, the left neighbor, and the lower left neighbor of the first input sample, and the second value is the sum determined based on the upper right neighbor, the right neighbor, and the lower right neighbor of the first input sample.
[0219] In an example, which difference is used to calculate the horizontal gradient of each respective first input sample is determined based on the position of each respective first input sample.
[0220] In an example, at least one gradient filter includes a vertical gradient filter. The gradient value includes a vertical gradient value, which is the sum of the vertical gradients of the respective second input samples in the vertical gradient filter. Each vertical gradient of each respective second input sample is determined based on one of the following: (i) the difference between each respective second input sample and the top neighbor of the respective second input sample; (ii) the difference between the top neighbor of each respective second input sample and the bottom neighbor of the respective second input sample; and (iii) the difference between a first value and a second value, where the first value is the sum determined based on the upper left neighbor, the top neighbor, and the upper right neighbor of the second input sample, and the second value is the sum determined based on the lower left neighbor, the bottom neighbor, and the lower right neighbor of the respective second input sample.
[0221] In an example, which difference is used to calculate the vertical gradient of each respective second input sample is determined based on the position of each respective second input sample.
[0222] At (S1440), a predicted value of a current sample is determined based on a sum of a linear prediction value and at least one modified value including a gradient value. The FIBC filter in the FIBC mode includes a linear filter and at least one gradient filter.
[0223] In an example, the predicted value of the current sample is clipped.
[0224] At (S1450), the current sample is reconstructed according to the predicted value of the current sample.
[0225] Then, the process proceeds to (S1499) and ends.
[0226] The process (1400) can be adjusted appropriately. At least one step in the process (1400) can be modified and / or omitted. At least one additional step can be added. Any suitable order of implementation can be used.
[0227] In an example, the position of a central sample located at the center of the linear filter is used to determine a position value, and the predicted value of the current sample is determined based on a sum of the linear prediction value and at least one modified value including the gradient value and the position value. The FIBC filter in the FIBC mode includes a linear filter, at least one gradient filter, and a position filter (e.g., coefficients of positions).
[0228] In an example, a non - linear relationship between a non - linear value and a value of at least one of the current sample and adjacent samples of the current sample is used to determine a non - linear value associated with the current sample according to the current sample and at least one of the adjacent samples of the current sample, and the predicted value of the current sample is determined based on a sum of the linear prediction value and at least one modified value including the gradient value and the non - linear value. The FIBC filter in the FIBC mode includes a linear filter, at least one gradient filter, and coefficients of the non - linear value.
[0229] In an example, coefficients of the FIBC filter in the FIBC mode are determined according to a current template of a current block and a reference template of a reference block indicated by a block vector of the current block.
[0230] In an example, LDL decomposition is used to determine coefficients of the FIBC filter in the FIBC mode.
[0231] In an example, the shape of the linear filter is predefined, and at least one shape of at least one gradient filter is predefined.
[0232] In an example, samples in the linear filter are spatially separated from all remaining samples in the linear filter.
[0233] Figure 15A flowchart showing an overview process (1500) according to an embodiment of the present disclosure is presented. The process (1500) can be used in a video encoder. In various embodiments, the process (1500) is executed by a processing circuit, such as a processing circuit that executes the functions of the video encoder (103), a processing circuit that executes the functions of the video encoder (303), etc. In some embodiments, the process (1500) is implemented by software instructions. Thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1500). The process starts at (S1501) and proceeds to (S1510).
[0234] At (S1510), a linear prediction value of the current sample in the current block is determined by applying a linear filter to samples predicted using one of the intra block copy (IBC) mode and the intra template matching (IntraTMP) mode. The current block uses the filtered intra block copy (FIBC) mode and is predicted using the FIBC filter. The FIBC filter includes a linear filter.
[0235] In an example, the linear filter includes a bias term, or the linear filter is configured to add the average value of the current block and subtract the average value of the current block from each of the samples predicted using one of the IBC mode and the IntraTMP mode.
[0236] At (S1520), at least one gradient filter is used to determine a gradient value associated with the current sample in the current block. The FIBC filter includes at least one gradient filter
[0237] At (S1530), a non - linear value associated with the current sample is determined using a non - linear relationship between a non - linear value and the value of the current sample and at least one of the adjacent samples of the current sample, and based on the current sample and the at least one of the adjacent samples.
[0238] At (S1540), a prediction value of the current sample is determined based on the sum of the linear prediction value and at least one modification value, where the at least one modification value includes the gradient value and the non - linear value. The FIBC filter in the FIBC mode includes the linear filter, the at least one gradient filter, and the coefficients of the non - linear value.
[0239] At (S1550), the current sample is encoded according to the prediction value of the current sample.
[0240] Then, the process proceeds to (S1599) and ends.
[0241] The process (1500) can be adjusted appropriately. At least one step in the process (1500) can be modified and / or omitted. At least one additional step can be added. Any suitable order of implementation can be used.
[0242] In an example, the position value is determined using the position of the central sample located at the center of the linear filter; and based on the sum of the linear prediction value and the at least one modification value, the prediction value of the current sample is determined, the at least one modification value including the gradient value, the non - linear value, and the position value. The FIBC filter in the FIBC mode includes the linear filter, the at least one gradient filter, the coefficient of the non - linear value, and the position filter (e.g., the coefficient of the position).
[0243] In one aspect, a method for processing visual media data is disclosed. The method includes: processing a bitstream of the visual media data according to formatting rules. The bitstream includes syntax elements that indicate using a filtered intra - block copy (FIBC) mode to predict a current block in a current picture. The formatting rules specify that a linear prediction value of a current sample in the current block is determined by applying a linear filter to samples predicted using one of an intra - block copy (IBC) mode and an intra - template matching (IntraTMP) mode. At least one gradient filter is used to determine a gradient value associated with the current sample in the current block; a non - linear value associated with the current sample is determined using a non - linear relationship between the non - linear value and at least one of the value of the current sample and the value of an adjacent sample of the current sample, and based on the current sample and the at least one adjacent sample; a position value is determined based on the position of the central sample located at the center of the linear filter; the prediction value of the current sample is determined based on the sum of the linear prediction value and at least one modification value, the at least one modification value including the gradient value, the non - linear value, and the position value, and the FIBC filter in the FIBC mode includes the linear filter, the at least one gradient filter, the coefficient of the non - linear value, and the coefficient of the position; and processing the current sample according to the prediction value of the current sample.
[0244] In an example, the linear filter includes a bias term.
[0245] In an example, the linear filter is configured to add an average value of the current block and remove the average value of the current block from each of the samples predicted using one of the IBC mode and the IntraTMP mode.
[0246] Aspects and / or examples in the present disclosure can be used alone or in any combination. For example, some aspects and / or examples performed by a decoder can be performed by an encoder, and vice versa. Each method (or aspect), encoder, and decoder can be implemented by processing circuitry (e.g., at least one processor or at least one integrated circuit). In one example, at least one processor executes a program stored in a non-volatile computer-readable medium.
[0247] The above technologies can be implemented as computer software by computer-readable instructions and physically stored in at least one computer-readable medium. For example, Figure 16 A computer system (1600) is shown, which is suitable for implementing certain embodiments of the disclosed subject matter.
[0248] The computer software can be encoded by any suitable machine code or computer language, and code including instructions is created through mechanisms such as assembly, compilation, and linking. The instructions can be directly executed by at least one computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through decoding, microcode, etc.
[0249] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0250] Figure 16 The components shown for the computer system (1600) are exemplary and do not impose any limitations on the scope of use or functions of the computer software implementing the embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependence on or requirement for any one component or combination thereof shown in the exemplary embodiments of the computer system (1600).
[0251] The computer system (1600) can include certain human-machine interface input devices. Such human-machine interface input devices can respond to inputs from at least one human user through tactile inputs (such as keyboard inputs, swipes, data glove movements), audio inputs (such as sounds, applause), visual inputs (such as gestures), olfactory inputs (not shown). The human-machine interface device can also be used to capture certain media, which does not have to be directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), videos (e.g., two-dimensional videos, three-dimensional videos including stereoscopic videos).
[0252] The human-machine interface input device may include at least one of the following (only one is shown): keyboard (1601), mouse (1602), touchpad (1603), touch screen (1610), data glove (not shown), joystick (1605), microphone (1606), scanner (1607), camera (1608).
[0253] The computer system (1600) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through the touch screen (1610), data glove (not shown), or joystick (1605), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (1609), headphones (not shown)), visual output devices (e.g., screens (1610) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light emitting diode screens, each of which has or does not have touch screen input function, each of which has or does not have tactile feedback function - some of which can output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).
[0254] The computer system (1600) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical discs with CD / DVD (CD / DVD ROM / RW) (1620) or similar media (1621), thumb drives (1622), removable hard disk drives or solid state drives (1623), traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.
[0255] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0256] The computer system (1600) may also include an interface (1654) to at least one communication network (1655). For example, the network may be wireless, wired, optical. The network may also be a local area network, wide area network, metropolitan area network, vehicular network, and industrial network, real-time network, delay tolerant network, etc. The network also includes local area networks such as Ethernet, wireless local area network, cellular network (GSM, 3G, 4G, 5G, LTE, etc.), television cable or wireless wide area digital network (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), etc. Some networks typically require an external network interface adapter for connection to certain common data ports or peripheral buses (1649) (e.g., the USB port of the computer system (1600)); other systems are typically integrated into the core of the computer system (1600) by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks, the computer system (1600) can communicate with other entities. The communication can be one-way, only for receiving (e.g., wireless television), one-way only for sending (e.g., CAN bus to certain CAN bus devices), or two-way, e.g., via a local or wide area digital network to other computer systems. Each of the above networks and network interfaces can use certain protocols and protocol stacks.
[0257] The above-mentioned human-machine interface device, human-accessible storage device, and network interface can be connected to the core (1640) of the computer system (1600).
[0258] The core (1640) may include at least one central processing unit (CPU) (1641), a graphics processing unit (GPU) (1642), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1643), a hardware accelerator for specific tasks (1644), a graphics adapter (1650), etc. These devices, as well as read-only memory (ROM) (1645), random access memory (1646), internal mass storage (e.g., internal non-user-accessible hard disk drive, solid state drive, etc.) (1647), etc. can be connected via a system bus (1648). In some computer systems, the system bus (1648) can be accessed in the form of at least one physical plug for expansion with additional central processing units, graphics processing units, etc. Peripherals can be directly attached to the system bus (1648) of the core or connected via a peripheral bus (1649). In the example, the screen (1610) can be connected to the graphics adapter (1650). The architecture of the peripheral bus includes external peripheral component interconnect PCI, universal serial bus USB, etc.
[0259] The CPU (1641), GPU (1642), FPGA (1643), and accelerator (1644) can execute certain instructions that, when combined, can form the aforementioned computer code. This computer code can be stored in the ROM (1645) or RAM (1646). Transitional data can also be stored in the RAM (1646), while permanent data can be stored in, for example, the internal mass storage (1647). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with at least one of the CPU (1641), GPU (1642), mass storage (1647), ROM (1645), RAM (1646), etc.
[0260] The computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well-known and available to those skilled in the computer software arts.
[0261] By way of example and not limitation, a computer system having an architecture (1600), particularly a core (1640), can provide the functionality of a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in at least one tangible computer-readable medium. Such a computer-readable medium can be a medium associated with the aforementioned user-accessible mass storage, as well as a specific memory of the non-volatile core (1640), such as the core internal mass storage (1647) or ROM (1645). The software implementing various embodiments of this disclosure can be stored in such a device and executed by the core (1640). Depending on specific requirements, the computer-readable medium can include one or more storage devices or chips. The software can cause the core (1640), particularly the processors therein (including the CPU, GPU, FPGA, etc.), to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in the RAM (1646) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality that is logically hardwired or otherwise included in circuitry (e.g., accelerator (1644)) that can operate instead of or in conjunction with the software to execute the specific processes or specific parts of the specific processes described herein. In appropriate instances, references to software can include logic, and vice versa. In appropriate instances, references to the computer-readable medium can include circuitry (such as an integrated circuit (IC)) that stores the software for execution, circuitry that contains the execution logic, or both. This disclosure encompasses any suitable combination of hardware and software.
[0262] As used in this disclosure, "at least one" or "one of" is intended to include any one or combination of the recited elements. For example, reference to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A through C is intended to include only A, only B, only C, or any combination thereof. Reference to one of A or B and one of A and B is intended to include A or B or (A and B). Where applicable, use of "one of" does not exclude any combination of the recited elements, such as when the elements are not mutually exclusive.
[0263] Although this disclosure has described at least two exemplary embodiments, various changes, permutations, and various equivalent replacements of the embodiments are within the scope of this disclosure. Therefore, it should be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of this disclosure and are thus within the spirit and scope of this disclosure.
Claims
1. A method for processing visual media data, characterized in that: The method comprises: The code stream of the visual media data is processed according to the format rules, wherein The code stream includes a syntax element indicating that a filtered intra block copy (FIBC) mode is used to predict a current block in a current picture; and The formatting rules specify that Determining a linear prediction value of a current sample in the current block by applying a linear filter to samples predicted using one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode; determining, using at least one gradient filter, a gradient value associated with the current sample in the current block; determining a nonlinear value associated with the current sample using a nonlinear relationship between the nonlinear value and a value of at least one of the current sample and a neighboring sample of the current sample and based on the at least one of the current sample and the neighboring sample; The position value is determined based on the position of a center sample located at the center of the linear filter; the prediction value of the current sample is determined based on a sum of the linear prediction value and at least one modification value, the at least one modification value comprising the gradient value, the non-linearity value and the position value, the FIBC filter in the FIBC mode comprising the linear filter, the at least one gradient filter, a coefficient of the non-linearity value and a coefficient of the position; and The current sample is processed according to the predicted value of the current sample.
2. The method according to claim 1, characterized in that The linear filter includes a bias term.
3. The method according to claim 1 or 2, characterized in that: The linear filter is configured to add a mean value of the current block and remove the mean value of the current block from each of the samples predicted using one of the IBC mode and the IntraTMP mode.
4. A video encoding method, characterized in that: include: determining a linear prediction value of a current sample in a current block predicted using a filtered intra block copy (FIBC) mode by applying a linear filter to samples predicted using one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode; determining, using at least one gradient filter, a gradient value associated with the current sample in the current block; determining a nonlinear value associated with the current sample using a nonlinear relationship between the nonlinear value and a value of at least one of the current sample and a neighboring sample of the current sample and based on the at least one of the current sample and the neighboring sample; determining a prediction value of the current sample based on a sum of the linear prediction value and at least one modification value, the at least one modification value comprising the gradient value and the non-linear value, the FIBC filter in the FIBC mode comprising coefficients of the linear filter, the at least one gradient filter and the non-linear value; as well as The current sample is encoded according to the predicted value of the current sample.
5. The method according to claim 4, characterized in that Further including: determining a position value using the position of a center sample located at the center of the linear filter; as well as Based on the sum of the linear prediction value and the at least one modification value, the prediction value of the current sample is determined, the at least one modification value includes the gradient value, the nonlinear value and the position value, and the FIBC filter in the FIBC mode includes the linear filter, the at least one gradient filter, the coefficient of the nonlinear value and the coefficient of the position.
6. The method according to claim 4 or 5, characterized in that: The linear filter includes a bias term; or The linear filter is configured to add a mean value of the current block and remove the mean value of the current block from each of the samples predicted using one of the IBC mode and the IntraTMP mode.
7. A video decoding device, characterized in that: include: A processing circuit, the processing circuit being configured to: Receiving encoded information indicating that a current block in a current picture is predicted using a filtered intra block copy (FIBC) mode; Determining a linear prediction value of a current sample in the current block by applying a linear filter to a prediction sample predicted using one of an intra block copy (IBC) mode and an intra template matching (IntraTMP) mode; determining, using at least one gradient filter, a gradient value associated with the current sample in the current block; determining a prediction value of the current sample based on a sum of the linear prediction value and at least one modification value including the gradient value, the FIBC filter in the FIBC mode including the linear filter and the at least one gradient filter; as well as The current sample is reconstructed according to the predicted value of the current sample.
8. The device according to claim 7, characterized in that The processing circuit is configured to: determining a position value using the position of a center sample located at the center of the linear filter; and Based on the sum of the linear prediction value and the at least one modification value, the prediction value of the current sample is determined, the at least one modification value includes the gradient value and the position value, and the FIBC filter in the FIBC mode includes the coefficients of the linear filter, the at least one gradient filter and the position.
9. The device according to claim 7, characterized in that The processing circuit is configured to: determining a nonlinear value associated with the current sample using a nonlinear relationship between the nonlinear value and a value of at least one of the current sample and a neighboring sample of the current sample and based on the at least one of the current sample and the neighboring sample; as well as Based on the sum of the linear prediction value and the at least one modification value, a prediction value of the current sample is determined, the at least one modification value includes the gradient value and the nonlinear value, and the FIBC filter in the FIBC mode includes coefficients of the linear filter, the at least one gradient filter and the nonlinear value.
10. The device according to any one of claims 7 to 9, characterized in that The linear filter includes a bias term.
11. The device according to any one of claims 7 to 10, characterized in that The linear filter adds a mean value of the current block and removes the mean value of the current block from each of the samples predicted using one of the IBC mode and the IntraTMP mode.
12. The device according to any one of claims 7 to 11, characterized in that The processing circuit is configured to clip the predicted value of the current sample.
13. The device according to any one of claims 7 to 12, characterized in that The processing circuit is configured to determine coefficients of the FIBC filter in the FIBC mode based on a current template of the current block and a reference template of a reference block indicated by a block vector of the current block.
14. The device according to any one of claims 7 to 13, characterized in that The linear filter has a cross shape, and the cross shape includes: (i) 5 samples, the 5 samples including a center sample of the linear filter with an offset of (0, 0), a north sample N with an offset of (0, -1), a south sample S with an offset of (0, 1), an east sample E with an offset of (1, 0), and a west sample W with an offset of (-1, 0), the offsets of the 5 samples in the linear filter being relative to the center sample; or (ii) 9 samples, the 9 samples include a center sample of the linear filter offset (0, 0), two north samples offset (0, -1) and (0, -2) respectively, two south samples offset (0, 1) and (0, 2) respectively, two east samples offset (1, 0) and (2, 0) respectively, and two west samples offset (-1, 0) and (-2, 0) respectively, the offsets of the 9 samples in the linear filter are relative to the center sample.
15. The device according to claim 14, characterized in that When the linear filter has the 5 samples, the current sample is located at one of the 5 positions of the 5 samples; and When the linear filter has the 9 samples, the current sample is located at one of the 9 positions of the 9 samples.