Intra prediction based on extrapolation filter

By introducing gradient information and nonlinear values ​​into the intra prediction of video encoding, the problem of inaccurate prediction in the prior art is solved, and the encoding efficiency and accuracy are improved.

CN119968845APending Publication Date: 2025-05-09TENCENT AMERICA LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202480004211.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-20
Filing Date
2024-04-19
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing video encoding technology is difficult to effectively utilize gradient information and nonlinear relationships in intra prediction, resulting in inaccurate prediction values ​​and affecting encoding efficiency.

Method used

The predicted value of the current sample is determined by introducing gradient information and nonlinear values ​​in the intra prediction. The specific method includes: determining the gradient information associated with the current sample according to the format rules; determining the nonlinear value using the nonlinear relationship between the nonlinear value and the values ​​of the adjacent sample; and determining the predicted value of the current sample based on the initial predicted value, gradient information and nonlinear value.

Benefits of technology

The accuracy and encoding efficiency of intra prediction are improved, and the amount of encoded data is reduced by more precisely utilizing the local characteristics of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968845A_ABST
    Figure CN119968845A_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure include methods and apparatus for video decoding and video encoding, and methods of processing visual media data. An apparatus for video decoding includes processing circuitry configured to: receive prediction information indicating that a current block in a current picture is predicted using an extrapolation filter-based intra prediction (EIP) mode; determining gradient information associated with a current sample in the current block; determining a prediction value of the current sample based on an initial prediction value predicted using the EIP mode and additional information including gradient information; and reconstructing the current sample according to the predicted value of the current sample.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 460,886, filed on April 20, 2023, entitled “Improvement of extrapolation filter based intra prediction,” which is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure describes aspects generally related to video encoding. Background Art

[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work of the presently named inventors described in this background section and in various aspects of this specification was performed, it does not indicate that it qualifies as prior art at the time of filing, and it is never explicitly or implicitly admitted that it is prior art to the present disclosure.

[0004] Image / video compression can help transmit image / video data between different devices, storage, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial redundancy and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress an image based on spatial redundancy. For example, intra-frame prediction can use reference data from a current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress an image based on temporal redundancy. For example, inter-frame prediction can predict samples in a current picture based on a previously reconstructed picture using motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the invention

[0005] Aspects of the present disclosure include methods and apparatus for video encoding / decoding.

[0006] In one aspect, a method for processing visual media data includes converting between a visual media file and a code stream of the visual media data according to a format rule, wherein the code stream includes prediction information indicating that a current block in a current picture is predicted using an extrapolation filter-based intra prediction (EIP) mode. The format rule specifies: determining gradient information associated with a current sample in a current block predicted using an EIP mode; determining the nonlinear value associated with the current sample based on neighboring samples of the current sample using a nonlinear relationship between the nonlinear value and values ​​of neighboring samples of the current sample; determining a prediction value of the current sample based on an initial prediction value predicted based on the EIP mode, the gradient information, and the nonlinear value; when the gradient information includes horizontal gradient information, the number of first input samples used to determine the horizontal gradient information and the position of the first input samples are set independently from the number and position of input samples used to determine the initial prediction value; when the gradient information includes vertical gradient information, the number of second input samples used to determine the vertical gradient information and the position of the second input samples are set independently from the number and position of input samples used to determine the initial prediction value; and when the gradient information includes horizontal gradient information and vertical gradient information, (i) the number and position of the first input samples and (ii) the number and position of the second input samples are set independently from each other and are set independently from the number and position of input samples used to determine the initial prediction value.

[0007] In one example, the gradient information includes a sum of horizontal gradient information and vertical gradient information; and the number of the first input samples is equal to the number of the second input samples.

[0008] In one example, (i) the number of first input samples or the positions of the first input samples for determining horizontal gradient information in the gradient information, (ii) the number of second input samples or the positions of the second input samples for determining vertical gradient information in the gradient information, and one of (i) and (ii) depends on the block shape of the current block or the filter shape of the EIP mode.

[0009] In one example, determining the gradient information includes determining the gradient information based on at least one of: (i) horizontal gradient information determined based on a horizontal gradient of a corresponding first input sample and (ii) vertical gradient information determined based on a vertical gradient of a corresponding second input sample.

[0010] In one example, the horizontal gradient of one of the first input samples is determined based on the following difference values: (i) a difference between the one first input sample and a left adjacent sample of the one first input sample; (ii) a difference between the left adjacent sample of the one first input sample and a right adjacent sample of the one first input sample; and (iii) a difference between a first value and a second value, the first value being a sum of upper left adjacent samples, left adjacent samples, and lower left adjacent samples of the one first input sample, and the second value being a sum of upper right adjacent samples, right adjacent samples, and lower right adjacent samples of the one first input sample.

[0011] In one example, based on the position of the one of the first input samples, it is determined which difference to use to calculate the horizontal gradient of the one first input sample.

[0012] In one example, a vertical gradient of one of the second input samples is determined based on the following difference values: (i) a difference between the one second input sample and an upper neighboring sample of the one second input sample; (ii) a difference between an upper neighboring sample of the one second input sample and a lower neighboring sample of the one second input sample; and (iii) a difference between a first value and a second value, the first value being based on a sum of upper left neighboring samples, upper neighboring samples, and upper right neighboring samples of the one second input sample, and the second value being based on a sum of lower left neighboring samples, lower neighboring samples, and lower right neighboring samples of the one second input sample.

[0013] In one example, based on the position of the one of the second input samples, it is determined which difference to use to calculate the vertical gradient of the one of the second input samples.

[0014] In one aspect, a method for video encoding includes determining gradient information associated with a current sample in a current block predicted using an extrapolation filter-based intra prediction (EIP) mode; determining a nonlinear value from neighboring samples of the current sample based on a nonlinear relationship between a nonlinear value associated with the current sample and values ​​of neighboring samples of the current sample; and determining a prediction value of the current sample based on an initial prediction value predicted using the EIP mode, the gradient information, and the nonlinear value.

[0015] In one example, when the gradient information includes horizontal gradient information, the number of first input samples for determining the horizontal gradient information and the position of the first input samples are set independently from the number and position of input samples for determining the initial prediction value. When the gradient information includes vertical gradient information, the number of second input samples for determining the vertical gradient information and the position of the second input samples are set independently from the number and position of input samples for determining the initial prediction value. When the gradient information includes horizontal gradient information and vertical gradient information, (i) the number and position of the first input samples and (ii) the number and position of the second input samples are set independently from each other and independently from the number and position of input samples for determining the initial prediction value.

[0016] In one example, the gradient information includes a sum of horizontal gradient information and vertical gradient information; and the number of the first input samples is equal to the number of the second input samples.

[0017] In one example, (i) the number of first input samples or the positions of the first input samples for determining horizontal gradient information in the gradient information, (ii) the number of second input samples or the positions of the second input samples for determining vertical gradient information in the gradient information, and one of (i) and (ii) depends on the block shape of the current block or the filter shape of the EIP mode.

[0018] In one example, determining the gradient information includes determining the gradient information according to at least one of horizontal gradient information determined based on a horizontal gradient of a corresponding first input sample and vertical gradient information determined based on a vertical gradient of a corresponding second input sample.

[0019] In one example, the horizontal gradient of one of the first input samples is determined based on the following difference values: (i) the difference between the one first input sample and a left adjacent sample of the one first input sample; (ii) the difference between the left adjacent sample of the one first input sample and a right adjacent sample of the one first input sample; and (iii) the difference between a first value and a second value, the first value being based on the sum of the upper left adjacent samples, the left adjacent samples, and the lower left adjacent samples of the one first input sample, and the second value being based on the sum of the upper right adjacent samples, the right adjacent samples, and the lower right adjacent samples of the one first input sample.

[0020] In one example, based on the position of the one of the first input samples, it is determined which difference to use to calculate the horizontal gradient of the one first input sample.

[0021] In one example, a vertical gradient of one of the second input samples is determined based on the following difference values: (i) a difference between the one second input sample and an upper neighboring sample of the one second input sample; (ii) a difference between an upper neighboring sample of the one second input sample and a lower neighboring sample of the one second input sample; and (iii) a difference between a first value and a second value, the first value being a sum of upper left neighboring samples, upper neighboring samples, and upper right neighboring samples of the one second input sample, and the second value being a sum of lower left neighboring samples, lower neighboring samples, and lower right neighboring samples of the one second input sample.

[0022] In one example, based on the position of the one of the second input samples, it is determined which difference to use to calculate the vertical gradient of the one of the second input samples.

[0023] In one aspect, an apparatus for video decoding includes a processing circuit. The processing circuit is configured to: receive prediction information indicating that a current block in a current picture is predicted using an extrapolation filter-based intra prediction (EIP) mode; determine gradient information associated with a current sample in the current block; determine a prediction value of the current sample based on an initial prediction value predicted using the EIP mode and additional information including the gradient information; and reconstruct the current sample according to the prediction value of the current sample.

[0024] In one example, the processing circuit is configured to: determine the nonlinear value according to the neighboring samples of the current sample based on the nonlinear relationship between the nonlinear value associated with the current sample and the values ​​of the neighboring samples of the current sample; and determine the prediction value of the current sample according to the initial prediction value predicted using the EIP mode and additional information including gradient information and the nonlinear value.

[0025] In one example, when the gradient information includes horizontal gradient information, the number of first input samples for determining the horizontal gradient information and the position of the first input samples are set independently from the number and position of input samples for determining the initial prediction value. When the gradient information includes vertical gradient information, the number of second input samples for determining the vertical gradient information and the position of the second input samples are set independently from the number and position of input samples for determining the initial prediction value. When the gradient information includes horizontal gradient information and vertical gradient information, (i) the number and position of the first input samples and (ii) the number and position of the second input samples are set independently from each other and independently from the number and position of input samples for determining the initial prediction value.

[0026] In one example, the gradient information includes a sum of horizontal gradient information and vertical gradient information; and the number of the first input samples is equal to the number of the second input samples.

[0027] In one example, (i) the number of first input samples or the positions of the first input samples for determining horizontal gradient information in the gradient information, (ii) the number of second input samples or the positions of the second input samples for determining vertical gradient information in the gradient information, and one of (i) and (ii) depends on the block shape of the current block or the filter shape of the EIP mode.

[0028] In one example, the processing circuit is configured to determine the gradient information based on at least one of horizontal gradient information determined based on a horizontal gradient of a corresponding first input sample and vertical gradient information determined based on a vertical gradient of a corresponding second input sample.

[0029] In one example, the processing circuit is configured to determine the horizontal gradient of one of the first input samples based on the following difference values: (i) a difference between the one first input sample and a left adjacent sample of the one first input sample; (ii) a difference between the left adjacent sample of the one first input sample and a right adjacent sample of the one first input sample; and (iii) a difference between a first value and a second value, the first value being a sum of upper left adjacent samples, left adjacent samples, and lower left adjacent samples of the one first input sample, and the second value being a sum of upper right adjacent samples, right adjacent samples, and lower right adjacent samples of the one first input sample.

[0030] In one example, the processing circuit is configured to determine which difference to use to calculate the horizontal gradient of the one of the first input samples based on the position of the one of the first input samples.

[0031] In one example, the processing circuit is configured to determine a vertical gradient of one of the second input samples based on the following difference values: (i) a difference between the one second input sample and an upper adjacent sample of the one second input sample; (ii) a difference between an upper adjacent sample of the one second input sample and a lower adjacent sample of the one second input sample; and (iii) a difference between a first value and a second value, the first value being a sum of upper left adjacent samples, upper adjacent samples, and upper right adjacent samples of the one second input sample, and the second value being a sum of lower left adjacent samples, lower adjacent samples, and lower right adjacent samples of the one second input sample.

[0032] In one example, the processing circuit is configured to determine, based on the position of the one of the second input samples, which difference to use to calculate the vertical gradient of the one of the second input samples.

[0033] In one example, the processing circuit is configured to: calculate a first value, which is an average or median of upper left neighboring samples, upper neighboring samples, and left neighboring samples of the current sample; and determine the nonlinearity value based on the square of the first value.

[0034] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions, which, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0036] Figure 1 is a schematic illustration of an exemplary block diagram of a communication system (100).

[0037] Figure 2 is a schematic illustration of an exemplary block diagram of a decoder.

[0038] Figure 3 is a schematic illustration of an exemplary block diagram of an encoder.

[0039] Figure 4 Examples of intra prediction modes (eg, 67 intra prediction modes) according to an aspect of the present disclosure are shown.

[0040] Figure 5 An example of a matrix-weighted intra prediction (MIP) process according to an aspect of the present disclosure is shown.

[0041] Figure 6 An example of a spatial portion of a convolution filter according to an aspect of the present disclosure is shown.

[0042] Figure 7 An example of a reference region for deriving filter coefficients according to an aspect of the present disclosure is shown.

[0043] Figure 8 An example of spatial samples for a gradient and location-based convolutional cross-component model (GL-CCCM) according to an aspect of the present disclosure is shown.

[0044] Fig. 9 Examples of three types of reconstructed regions used in an extrapolation filter based intra prediction (EIP) mode according to an aspect of the present disclosure are shown.

[0045] Fig.10 Examples of three types of filter shapes used in EIP mode according to an aspect of the present disclosure are shown.

[0046] Figures 11 to 13 An example of predicting samples in a current block based on an EIP mode according to an aspect of the present disclosure is shown.

[0047] Fig.14 An example of input samples used in an EIP filter according to an aspect of the present disclosure is shown.

[0048] Fig.15 An example of spatial samples of input samples C for calculating gradient information according to an aspect of the present disclosure is shown.

[0049] Fig.16 An example of selecting a gradient calculation method based on the position of an input sample according to an aspect of the present disclosure is shown.

[0050] Fig.17 An example of the positions of neighboring samples of a current sample according to an aspect of the present disclosure is shown.

[0051] Fig.18 A flow chart outlining a decoding process according to some aspects of the present disclosure is shown.

[0052] Fig.19 A flow chart outlining an encoding process according to some aspects of the present disclosure is shown.

[0053] Fig. 20 is a schematic illustration of a computer system according to an aspect. DETAILED DESCRIPTION

[0054] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application of the disclosed subject matter, namely a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.

[0055] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101), such as a digital camera, which creates, for example, an uncompressed video picture stream (102). In one example, the video picture stream (102) includes samples captured by a digital camera. The video picture stream (102), which is depicted as a thick line to emphasize the high amount of data compared to the encoded video data (104) (or the encoded video bitstream), can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or the encoded video bitstream), which is depicted as a thin line to emphasize the lower amount of data compared to the video picture stream (102), can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 A client subsystem (106) and a client subsystem (108) in a streaming server (105) may access a copy (107) and a copy (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an output video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0056] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0057] Figure 2 An exemplary block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 A video decoder (110) in an example of FIG.

[0058] The receiver (231) may receive one or more encoded video sequences, for example, included in a bitstream to be decoded by the video decoder (210). In one aspect, the encoded video sequences are received one at a time, wherein the decoding of each encoded video sequence is independent of the decoding of the other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not depicted). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be located outside the video decoder (210) (not depicted). In still other applications, a buffer memory (not depicted) may be provided outside the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be provided inside the video decoder (210) to, for example, handle playback timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be required, or the buffer memory (215) may be made smaller. For use on a traffic packet network such as the Internet, the buffer memory (215) may be required, and the buffer memory (215) may be relatively large, advantageously may have an adaptive size, and may be implemented at least partially in an operating system or similar element (not depicted) external to the video decoder (210).

[0059] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210) and potential information for controlling a presentation device such as a presentation device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (220) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0060] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0061] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by parser (220) through subgroup control information parsed from the coded video sequence. For clarity, such subgroup control information flow between parser (220) and the multiple units below is not depicted.

[0062] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual subdivision into the following multiple functional units is appropriate.

[0063] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) may output a block including sample values, which may be input into an aggregator (255).

[0064] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from a current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0065] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to a block that is inter-coded and potentially motion compensated. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) by the aggregator (255), thereby generating output sample information. The extraction of prediction samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, which may be provided to the motion compensated prediction unit (253) in the form of symbols (221), which may have, for example, an X component, a Y component and a reference picture component. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0066] The output samples of the aggregator (255) may be used by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of the coded video sequence, and to previously reconstructed and loop filtered sample values.

[0067] The output of the loop filter unit (256) may be a sample stream that may be output to a rendering device (212) and stored in a reference picture memory (257) for future inter-picture prediction.

[0068] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0069] The video decoder (210) may perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence may also be required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.

[0070] In one aspect, the receiver (231) can receive additional (redundant) data when receiving the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can take the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0071] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used to replace Figure 1 A video encoder (103) in an example of FIG.

[0072] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In an example, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source (301) can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is a part of the electronic device (320).

[0073] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), the digital video sample stream may have any suitable bit depth (e.g. 8-bit, 10-bit, 12-bit, ...), any color space (e.g. BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g. Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The following description focuses on the samples.

[0074] According to one aspect, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required. Implementing an appropriate encoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to the other functional units described. For clarity, the coupling is not depicted in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions that are related to the video encoder (303) optimized for a certain system design.

[0075] In some aspects, the video encoder (303) is configured to operate in an encoding loop. As an oversimplified description, in one example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.

[0076] The operation of the "local" decoder (333) may be similar to that described above in conjunction with Figure 2 The "remote" decoder of the video decoder (210) described in detail is identical. However, additional brief reference is made to Figure 2 , since the symbols are available and the entropy encoder (345) and the parser (220) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).

[0077] On the one hand, except for the parsing / entropy decoding present in the decoder, the decoder technology is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.

[0078] During operation, in some examples, the source encoder (330) may perform motion compensated predictive coding that predictively encodes an input picture by referencing one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.

[0079] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data may be decoded at the video decoder ( Figure 3 When the video sequence is decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.

[0080] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (335) may operate on a pixel block by pixel block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (334).

[0081] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0082] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) converts the symbols generated by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0083] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device that may store the encoded video data. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0084] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0085] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0086] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses a motion vector and a reference index to predict the sample values ​​of each block.

[0087] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0088] The source picture may typically be spatially subdivided into blocks of samples (e.g. blocks of 4×4, 8×8, 4×8 or 16×16 samples) and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively coded, or blocks of an I picture may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0089] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0090] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0091] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often shortened to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is partitioned into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0092] In some aspects, a bidirectional prediction technique may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that precede the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0093] In addition, merge mode technology can be used for inter-picture prediction to improve coding efficiency.

[0094] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the High-Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. The CU is divided into one or more prediction units (PUs) based on temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of values ​​for pixels (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0095] It should be noted that the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using any suitable technology. In one aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more integrated circuits. In another aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more processors executing software instructions.

[0096] Intra-frame prediction can be used, for example, in VVC. For example, advanced intra-frame prediction techniques used in VVC may include DC mode and planar mode similar to HEVC, additional finer-grained angle prediction with more angles than HEVC (for example, the number of angle prediction modes that can be used is increased from 33 in HEVC to 93), additional matrix-based prediction modes for luma components, and cross-component prediction modes for chroma components. For example, new (e.g., additional) intra-frame coding tools used in VVC may include: 67 intra-frame modes with wide-angle mode extension; 4-tap interpolation filters that depend on block size and mode; position dependent intra prediction combination (PDPC); cross component linear model (CCLM) intra-frame prediction; multi-reference line (MRL) intra-frame prediction; intra sub-partitions (ISP); and weighted intra-frame prediction using matrix multiplication.

[0097] In one example, intra-mode coding with 67 intra-prediction modes is described as follows. To capture arbitrary edge directions present in natural video, the number of directional intra modes used in VVC is extended from 33 as used in HEVC to 65. New directional modes not in HEVC are Figure 4 , while the planar mode and DC mode remain unchanged. The denser directional intra prediction mode may be applicable to various block sizes (e.g., all block sizes) and to luma intra prediction and chroma intra prediction. In one example, for example in VVC, for non-square blocks, a plurality of traditional angular intra prediction modes are adaptively replaced with a wide-angle intra prediction mode. In one example, the traditional angular intra prediction mode may include an angular intra prediction mode used in HEVC.

[0098] In one example, such as in HEVC, an intra-coded block (e.g., each intra-coded block) has a square shape, and the length of each side may be a power of 2. Therefore, no division operation is required when generating intra-prediction values ​​using the DC mode. In one example, such as in VVC, the block may have a rectangular shape. In some examples, a division operation for each block may be used (e.g., a division operation for each block may be required). To avoid division operations for DC prediction, in some examples, for non-square blocks, only the longer sides are used to calculate the average value.

[0099] An example of intra-mode encoding is described as follows. In order to keep the complexity of the most probable mode (MPM) list generation low, an intra-mode encoding method with 6 MPMs can be used by considering two available adjacent intra-modes. The following three aspects can be considered to construct the MPM list: default intra-mode; adjacent intra-mode; and derived intra-mode. In one example, a unified 6-MPM list is used for intra-blocks regardless of whether the MRL encoding tool and the ISP encoding tool are applied. The MPM list can be constructed based on the intra-modes of the left adjacent block and the upper adjacent block of the current block.

[0100] In one example, a 4-tap interpolation filter and reference sample smoothing may be applied. A four-tap intra-frame interpolation filter (IF) may be used to improve directional intra prediction accuracy. In HEVC, a two-tap linear interpolation filter is used to generate intra prediction blocks in directional prediction modes (e.g., excluding planar prediction and DC prediction). In VVC, in some examples, two sets of 4-tap IFs may replace the low-precision linear interpolation in HEVC, wherein one set of 4-tap IFs is a DCT-based interpolation filter (DCTIF) and the other set of 4-tap IFs is a 4-tap smoothing interpolation filter (SIF). DCTIF can be constructed in the same way as DCTIF for chroma component motion compensation in HEVC and VVC. SIF can be obtained by convolving a 2-tap linear interpolation filter with a [1 2 1] / 4 filter.

[0101] Depending on the intra prediction mode, in some examples the following reference sample processing may be performed: - Directional intra prediction modes are classified into one of the following groups: - Group A: vertical mode or horizontal mode (HOR_IDX, VER_IDX), - Group B: Directivity modes and planar modes for non-fractional angles (-14, -12, -10, -6, 2, 34, 66, 72, 76, 78, 80), - Group C: Other directional modes; - if the directional intra prediction mode is classified as belonging to group A, no filter is applied to the reference samples to generate the prediction samples; Otherwise, if the mode falls into group B, the mode is a directional mode, and all of the following conditions are true, a [1, 2, 1] reference sample filter (subject to the mode-dependent intra smoothing (MDIS) condition) may be applied to the reference samples to further copy the filtered values ​​into intra prediction values ​​according to the selected direction, but no interpolation filter is applied: -refIdx equals 0 (no MRL) -TU size is greater than 32 -brightness - No ISP block - Otherwise, if the mode is classified as belonging to group C, the MRL index is equal to 0, and the current block is not an ISP block, then only the intra reference sample interpolation filter is applied to the reference samples to generate prediction samples that fall into fractional or integer positions between the reference samples according to the selected direction (no reference sample filtering is performed). The interpolation filter type is determined as follows: - Set minDistVerHor equal to Min(Abs(predModeIntra-50), Abs(predModelntra-18)) - Set nTbS equal to (Log2(W)+Log2(H))>>1 - Set intraHorVerDistThres[nTbS] as specified below: -If minDistVerHor is greater than intraHorVerDistThres[nTbS], use SIF for interpolation - Otherwise, use DCTIF to do the interpolation.

[0102] Matrix weighted intra prediction (MIP) can be used. The MIP method is a new intra prediction technique in VVC. For predicting samples of a rectangular block with width and height, the MIP mode can take a row of H reconstructed adjacent boundary samples on the left side of the block and a row of reconstructed adjacent boundary samples above the block as input. If the reconstructed samples are not available, the samples can be generated in the same way as the traditional intra prediction generates samples. The generation of the prediction signal can be based on the following three steps, including averaging, matrix-vector multiplication, and linear interpolation, such as Figure 5 shown.

[0103] A convolutional cross-component intra prediction model may be used, for example, in ECM. A convolutional cross-component model (CCCM) may predict chroma samples from reconstructed luma samples in a manner similar to how the current CCLM mode is applied. As with CCLM, when chroma subsampling is used, the reconstructed luma samples may be downsampled to match a lower resolution chroma grid. Similar to CCLM, an upper reference sample, a left reference sample, or both an upper reference sample and a left reference sample may be used as templates for model derivation.

[0104] Similar to CCLM, you can choose to use a single model or a multi-model variant of CCCM. The multi-model variant can use two models: one model is derived for samples above the average brightness reference value, and another model is derived for the remaining samples (following the spirit of CCLM design). For PUs with at least 128 available reference samples, for example, the multi-model CCCM mode can be selected.

[0105] The convolution filter (e.g., a convolution 7-tap filter) may include a 5-tap plussign shape spatial component, a nonlinear term, and a bias term (e.g., composed of a 5-tap plussign shape spatial component, a nonlinear term, and a bias term). The input of the spatial 5-tap component of the filter may include a central luminance sample (C) and its upper or northern neighboring samples (N), lower or southern neighboring samples (S), left or western neighboring samples (W), and right or eastern neighboring samples (E) (e.g., composed of a central luminance sample (C) and its upper or northern neighboring samples (N), lower or southern neighboring samples (S), left or western neighboring samples (W), and right or eastern neighboring samples (E)), wherein the central luminance sample (C) is co-located with the chrominance sample to be predicted, such as Figure 6 shown.

[0106] The non-linear term NP may be expressed as a power of 2 of the center luminance sample C and may be scaled to the sample value range of the content, such as described in Equation 1. NP=(C 2 + midVal)>>bitDepth Equation 1

[0107] That is, for 10-bit content, the nonlinear term NP can be calculated using Equation 2. NP=(C 2 +512)>>10 Equation 2

[0108] The middle value (midVal) is 2 10 / 2, which is 512.

[0109] The bias term B may represent a scalar offset between the input and output (eg, similar to an offset term in CCLM) and may be set to an intermediate chroma value (eg, 512 for 10-bit content).

[0110] The output of the filter can be calculated as the filter coefficient c i The convolution between and the input value, and the output of the filter can be clipped to the range of valid chroma samples using Equation 3. predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B Equation 3

[0111] For example, the filter coefficient c i It can be calculated by minimizing the mean square error (MSE) between the predicted chroma samples and the reconstructed chroma samples in the reference region. Figure 7 An example of a reference region (and its filling) for deriving filter coefficients according to an aspect of the present disclosure is shown. The reference region may include (e.g., consist of) 6 rows of chroma samples above and to the left of the PU. The reference region may extend one PU width to the right and one PU height below the PU boundary. The region may be adjusted to include only available samples. An extension (701) to this region may be used to support "side samples" of a cross-shaped spatial filter and to be filled when located in an unavailable area.

[0112] MSE minimization can be performed by calculating the autocorrelation matrix of the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix can be LDL decomposed, and back substitution can be used to calculate the final filter coefficients. This process roughly follows the calculation of the ALF filter coefficients used in, for example, ECM, however, LDL decomposition is chosen instead of Cholesky decomposition, for example to avoid the use of square root operations.

[0113] The autocorrelation matrix can be calculated using the reconstructed values ​​of the luma samples and chroma samples. The luma samples and chroma samples can be in the full range (e.g., between 0 and 1023 for 10-bit content), resulting in relatively large values ​​in the autocorrelation matrix. This allows the use of high bit depth operations when calculating model parameters. Fixed offsets can be removed from the luma samples and chroma samples in each PU of each model. This reduces the size of the values ​​used in model creation and allows the precision of fixed point arithmetic to be reduced. Therefore, in some examples, 16 bits of fractional precision can be used instead of the 22 bits of precision of the original CCCM implementation.

[0114] In some examples, to simplify the operation, the reference sample values ​​just outside the upper left corner of the PU can be used as offsets (offsetLuma, offsetCb, and offsetCr). The sample values ​​used in model creation and final prediction (e.g., luma and chroma in the reference region, and luma in the current PU) can be reduced by fixed values ​​as follows: C'=C-offsetLuma; N'=N-offsetLuma; S'=S-offsetLuma; E'=E-offsetLuma; W'=W-offsetLuma; P'=nonLinear(C'); B=midValue=1<<(bitDepth-1); and the chroma values ​​are predicted using Equation 4, where offsetChroma is equal to offsetCr and offsetCb for the Cr component and Cb component, respectively. predChromaVal = c0C' + c1N' + c2S' + c3E' + c4W' + c5P' + c6B + offsetChroma Equation 4

[0115] In one example, to avoid any additional sample-level operations, the luma offset is removed during luma reference sample interpolation. For example, this can be achieved by replacing the rounding term used in the luma reference sample interpolation with an updated offset that includes the rounding term and offsetLuma. The chroma offset can be removed by subtracting the chroma offset directly from the reference chroma samples. As an alternative, the effect of the chroma offset can be removed from the cross-component vector to obtain the same result. In order to add the chroma offset back to the output of the convolution prediction operation, the chroma offset can be added to the bias term of the convolution model.

[0116] The CCCM model parameter calculation process may use a division operation. In some examples, the division operation may not be user-friendly to implement. The division operation may be replaced by a multiplication (using a scaling factor) and a shift operation, where the scaling factor and the number of shifts may be calculated based on the denominator, for example, similar to the method used when calculating CCLM parameters.

[0117] A gradient and position based convolutional cross-component model (GL-CCCM) can map luma values ​​to chroma values ​​using a filter with inputs including one spatial luma sample, two gradient values, two position information, a nonlinear term, and a bias term (e.g., consisting of one spatial luma sample, two gradient values, two position information, a nonlinear term, and a bias term). The GL-CCCM method can use gradient and position information instead of the four spatial neighboring samples used in the CCCM filter. The GL-CCCM filter used for prediction can be described using Equation 5. predChromaVal = c0C + c1Gy + c2Gx + c3Y + c4X + c5P + c6B Equation 5

[0118] Figure 8 An example of a spatial sample for GL-CCCM according to an aspect of the present disclosure is shown. Gy and Gx are the vertical gradient and the horizontal gradient, respectively, and are calculated using Eq. 6. G y =(2N+NW+NE)–(2S+SW+SE) G x =(2W+NW+SW)–(2E+NE+SE) Equation 6

[0119] Y and X are the spatial coordinates of the center brightness sample.

[0120] The remaining parameters can be the same as the CCCM tool. The reference area used for parameter calculation can be the same as the CCCM method.

[0121] A flag (e.g., a PU-level flag for CABAC encoding) may be used to signal the use of GL-CCCM mode. From a signaling perspective, GL-CCCM mode may be considered a sub-mode of CCCM, e.g., the GL-CCCM flag is signaled only if the original CCCM flag is true.

[0122] Similar to CCCM, in some examples, the GL-CCCM tool has 6 modes for calculating parameters: single-model GL-CCCM based on the upper template and the left template; single-model GL-CCCM based on the upper template; single-model GL-CCCM based on the left template; multi-model GL-CCCM based on the upper template and the left template; multi-model GL-CCCM based on the upper template; and multi-model GL-CCCM based on the left template. The encoder can perform a search (e.g., a sum of absolute transform differences (SATD) search) on these 6 GL-CCCM modes and the existing CCCM modes to find the best candidate for the entire rate-distortion (RD) test.

[0123] An intra prediction (EIP) mode based on an extrapolation filter may be used. In one example, the EIP mode includes two steps. In the first step, the extrapolation filter coefficients may be obtained from the neighboring reconstructed pixels (or samples) of the current block using a predetermined template. In the second step, the extrapolation may generate prediction values ​​position by position, for example, from the upper left sample to the lower right sample within the current block, generating prediction values ​​position by position.

[0124] In one example, the average, minimum and maximum values ​​may be searched in the following manner. Similar to the CCCM mode, in the EIP mode, the average value may be removed when the input is fed to the EIP filter. The value of the DC mode for the current block may be used as the average value for the EIP prediction. The minimum and maximum values ​​may be searched from the reconstructed pixels in a reconstructed area having, for example, thirteen columns and thirteen rows.

[0125] The filter coefficients may be calculated as follows. Fig. 9 An example of three types of reconstruction regions defined according to an aspect of the present disclosure is shown. The three types of reconstruction regions or reference regions (901)-(903) may include thirteen columns or rows of reconstruction pixels. Fig.10 An example of three types of filter shapes defined according to one aspect of the present disclosure is shown. The three types of filter shapes (1001)-(1003) may include fifteen inputs (also referred to as input samples) and generate one output. When the current block is predicted using the EIP mode, the decoder may decode the relevant syntax elements to determine the reconstruction region type and filter shape selected for the current block.

[0126] The selected filter can be slid in the selected reconstruction area with a step size of one pixel to collect the input samples and output samples of the EIP mode. When the mean value is removed from the input samples and the output samples, the autocorrelation matrix and the cross-correlation vector can be constructed. Then, the EIP coefficients can be obtained by the same method as in CCCM.

[0127] Figures 11 to 13 An example of predicting samples in a current block (1100) based on an EIP mode according to an aspect of the present disclosure is shown. The EIP mode can predict samples in the current block position by position.

[0128] refer to Fig.11 , all inputs of the EIP are reconstructed samples. For a position located above the left of the current block (e.g., the upper left position) (1101), the input of the EIP filter is a reconstructed sample, such as a reconstructed reference sample in the reference area (1110). Fig.12 , for locations along the border of the current block (1100), part of the input to the EIP filter is reference samples that have been reconstructed in the reference region (1110), and part of the input to the EIP filter is previously predicted samples in the current block (1100). Fig.13 , all inputs of the EIP filter are predicted samples in the current block (1100), for example, for other positions in the current block (1100), the input of the EIP filter may include previously predicted samples in the current block (1100).

[0129] To reduce the prediction error, the searched minimum and maximum values ​​may be applied to limit the output range of each prediction value, as described in Equation 7.

[0130] pred (x,y) is the predicted value at (x, y) in the current block (1100), min, max are the minimum and maximum values ​​searched from, for example, the reference area (1110) (e.g., thirteen reconstructed columns and rows), c i represents the i-th coefficient of the derived EIP filter, t (x-xoffset_i,y-yoffset_i) It is used to predict the reconstructed or predicted value of the current position or current sample, and mean is the average value calculated by the DC prediction mode.

[0131] In some examples, such as in VVC, multiple intra prediction modes are defined to generate prediction values, such as planar mode, DC mode, and angular intra prediction mode. The general intra prediction process can be described as a process of extrapolating a current sample from a reference sample through Equation 8.

[0132] n is the number of reference samples, ref i is the i-th reference sample, c i is the i-th reference sample ref i Given a particular intra prediction mode, the filter coefficients of the sample and the reference sample can be determined. In some examples, only predetermined filter coefficients are allowed to be used, and the filter coefficients cannot be adaptively adjusted according to the video content. In addition, a significant bit rate overhead is used to signal the selected intra prediction mode.

[0133] As described above, the extrapolation filter in EIP mode can be derived from the neighboring reconstructed samples of the current block and used to perform intra prediction. Some examples of extrapolation filters may not be accurate enough because the extrapolation filter consists only of linear terms of sample values ​​without using gradient information or including nonlinear terms.

[0134] Aspects of the present disclosure provide techniques including improving the EIP mode, for example, by including at least one of gradient information, nonlinear terms, and / or the like as additional inputs to the EIP filter. In one aspect, gradient information associated with a current sample in a current block may be determined. The nonlinear value (also interchangeably referred to as a nonlinear term) may be determined based on neighboring samples of the current sample based on a nonlinear relationship between the nonlinear value associated with the current sample and the values ​​of neighboring samples of the current sample. A predicted value of the current sample may be determined based on an initial predicted value predicted using the EIP mode and additional information. The additional information may include at least one of the gradient information and the nonlinear value. The current sample may be encoded based on the predicted value of the current sample.

[0135] In one aspect, the gradient information of adjacent reconstructed samples (e.g., the gradient G of adjacent reconstructed samples) may be utilized. x and G y ) is used to derive the intra prediction of the current block. The predicted value pred0(x,y) at (x, y) can be defined as: pred0(x,y)=P0+GX+GY Equation 9

[0136] pred0(x, y) is the predicted value at (x, y). (x, y) may be the position of the current sample. The three parameter sets associated with P0, GX, and GY may correspond to input sample values, horizontal gradient information, and vertical gradient information, respectively. The gradient information may include horizontal gradient information GX and vertical gradient information GY.

[0137] The first set may be related to the input sample values, which may include parameters of N0 input samples, (x-xoffset 0,i ,y-yoffset 0,i ) is the position of the i-th input sample, c 0,i is the coefficient of the i-th input sample, t(x-xoffset 0,i ,y-yoffset 0,i ) is the value of the i-th input sample.

[0138] In one example, N0 and the position of the input sample are determined based on the filter (or filter shape) used in EIP mode to determine P0, e.g. Fig.10 As shown, the N0 of the filter shapes (1001)-(1003) is 15. If different filters are used, the N0 may be different.

[0139] The second set may be related to horizontal gradient information, which may include parameters of N1 input samples. 1,j is the coefficient of the jth input sample, G x (x-xoffset 1,j ,y-yoffset 1,j ) is the horizontal gradient G of the jth input sample x In one example, the N1 input samples may be the same as the N0 input samples. In one example, the N1 input samples may be different from the N0 input samples.

[0140] The third set may be related to vertical gradient information, which may include parameters for N2 input samples. 2,k is the coefficient of the kth input sample, G y (x-xoffset 2,k ,y-yoffset 2,k ) is the vertical gradient G of the kth input sample y In one example, the N2 input samples may be the same as the N0 input samples. In one example, the N2 input samples may be different from the N0 input samples. In one example, the N2 input samples may be the same as the N1 input samples. In one example, the N2 input samples may be different from the N1 input samples.

[0141] The N0 input samples, the N1 input samples, and the N2 input samples may include samples that are reconstructed samples and previously predicted samples. Fig.11 In the example shown, N0 input samples include reconstructed samples in the reference region ( 1110 ), and N1 input samples and N2 input samples may also include reconstructed samples in the reference region ( 1110 ).

[0142] exist Fig.12 In the example shown, N0 input samples include reconstructed samples in the reference area (1110) and previously predicted samples in the current block (1100) (indicated by dark gray), and N1 input samples and N2 input samples may include reconstructed samples in the reference area (1110) and / or previously predicted samples in the current block (1100) (indicated by dark gray).

[0143] exist Fig.13 In the example shown, N0 input samples, N1 input samples, and N2 input samples may include previous prediction samples (indicated by dark grey) in the current block (1100).

[0144] In one example, gradient information (e.g., horizontal gradient information GX and / or vertical gradient information GY) associated with a current sample in a current block may be determined. A prediction value (e.g., pred0(x, y)) of the current sample may be determined based on an initial prediction value (e.g., P0) predicted using an EIP mode and additional information including gradient information (e.g., GX and / or GY). The current sample may be encoded according to the prediction value of the current sample.

[0145] As described above, several variations of P0 may be used. For example, when the input (e.g., N0 input samples) is fed to the EIP filter, a mean removal operation may be applied to P0. P1 indicates an initial prediction value, which may be P0 after the mean removal operation is applied, such as described in Equation 13. Then, the prediction value pred1(x, y) at (x, y) may be generated using Equations 13 and 14. pred1(x,y)=P1+GX+GY Equation 14

[0146] In one example, a clipping operation may be applied to generate a final prediction pred2(x,y). pred2(x,y)=clip(min,max,Pi+GX+GY) Equation 15

[0147] The clipping operation can limit the value of the final prediction pred2(x,y) to the range from the minimum to the maximum value.

[0148] The initial prediction value Pi in Equation 15 can be P0, P1, or other variations of P0.

[0149] Gradient information may be obtained using any suitable method, such as described below.

[0150] In one example, when the gradient information includes horizontal gradient information (eg, GX), the number of first input samples (eg, N1) and the position of the first input sample (eg, (x-xoffset in Equation 11) used to determine the horizontal gradient information is 0,j ,y-yoffset 0,j ) indicates) for example, the number (e.g., N0) and position (e.g., indicated by (x-xoffset in Equation 10) of input samples used to determine the initial prediction value (e.g., P0 or P1) 0,i ,y-yoffset 0,i ) indicates) is independently set. When the gradient information includes vertical gradient information (e.g., GY), the number of second input samples (e.g., N2) and the position of the second input sample (e.g., by (x-xoffset in Equation 12) used to determine the vertical gradient information 0,k,y-yoffset 0,k ) indicates that) is set independently of the number and position of the input samples used to determine the initial prediction value, for example. When the gradient information includes horizontal gradient information and vertical gradient information, (i) the number and position of the first input samples and (ii) the number and position of the second input samples are set independently of each other and independently of the number and position of the input samples used to determine the initial prediction value, for example, as described below.

[0151] The number and location of different parameter sets can be set independently. Fig.14 2 shows an example of input samples used in an EIP filter according to an aspect of the present disclosure. Fig.14 In one example, the values ​​of 15 grayscale samples (including (1402) and (1403)) may be used as inputs to an EIP filter (e.g., equations 9 and 10 or equations 13 and 14) to generate P0 or P1 for the current sample (1401), so N0 is 15. To generate gradient information (e.g., GX and GY), samples (1402) and (1403) may be used as inputs to an EIP filter (e.g., equations 9, 11, and 12), so N1 and N2 are 2. The EIP filter may include a spatial filter described in equation 10 to generate P0 (or a variant, such as P1) and a gradient filter described in equations 10 and 11 to generate GX and GY, respectively. In one example, the values ​​of 15 grayscale samples may be used by the EIP filter, but only a subset (1402) and (1403) may be used in actual processing.

[0152] In one aspect, the gradient information includes the sum of horizontal gradient information and vertical gradient information. The number (eg, N1) of first input samples (also referred to as N1 input samples) is equal to the number (eg, N2) of second input samples (also referred to as N2 input samples), eg, N1=N2.

[0153] In one aspect, N1 and N2 input samples may be used to generate GX and GY, respectively, for the EIP filter. N1 may be equal to N2, and the same sample positions may be used as inputs to GX and GY (eg, to generate GX and GY). Fig.14 An example is shown where N1 and N2 are 2, the same sample positions (1402) and (1403) are used as N1 input samples to generate GX, and the same sample positions (1402) and (1403) are used as N2 input samples to generate GY.

[0154] In one aspect, N1 and N2 input samples may be used to generate GX and GY, respectively, for the EIP filter. N1 is equal to N2. Different sample positions are used as inputs to GX and GY. Fig.14, N1 and N2 are 1, sample (1402) is used to generate GX, and sample (1403) is used to generate GY.

[0155] In one aspect, GX is generated using N1 input samples, and only GX is included in the EIP filter, such as shown in Equation 16. For example, only G x Can be used as the input of the EIP filter. In this example, the gradient information consists of horizontal gradient information, and N2 is 0. pred0(x,y)=P0+GX Equation 16

[0156] In one aspect, GY is generated using N2 input samples, and only GY is included in the EIP filter, such as shown in Equation 17. For example, only G y Can be used as the input of the EIP filter. In this example, the gradient information consists of vertical gradient information, and N1 is 0. pred0(x,y)=P0+GY Equation 17

[0157] In one aspect, (i) the number (N1) of first input samples (N1 input samples) or the position of the first input samples for determining horizontal gradient information in the gradient information, (ii) the number (N2) of second input samples (N2 input samples) or the position of the second input samples for determining vertical gradient information in the gradient information, one of (i) and (ii) depending on the block shape of the current block or the filter shape of the EIP mode (e.g., Fig.10 one of (1001)-(1003) shown).

[0158] In one aspect, the number (e.g., N1 and N2) or positions of input samples for gradient information may depend on block shape. For example, the number N1 or the N1 positions of N1 input samples for horizontal gradient information may depend on block shape. For example, the number N2 or the N2 positions of N2 input samples for vertical gradient information may depend on block shape.

[0159] In one example, if the block shape is a rectangle with width ≤ height or vice versa (e.g., height ≤ width), the number (e.g., N1 and / or N2) or position of input samples (e.g., N1 input samples and / or N2 input samples) used to generate gradient information may be adjusted to align with the block size. In one example, if the block width i ≤ the block height, then N1 ≤ N2.

[0160] In one aspect, the number or position of input samples for gradient information may depend on the filter shape. For example, the number N1 or the N1 position of N1 input samples for horizontal gradient information may depend on the filter shape. For example, the number N2 or the N2 position of N2 input samples for vertical gradient information may depend on the filter shape. Fig.10 Some examples of filter shapes are shown.

[0161] In one example, if the filter shape is rectangular (e.g., Fig.10 If the filter shape (1002) or (1003) shown is used, the number or position of input samples used to generate gradient information can be adjusted to align with the filter shape.

[0162] The horizontal gradient (G x ) determines the horizontal gradient information (GX) and the vertical gradient G based on the corresponding second input sample y At least one of the determined vertical gradient information (GY) determines the gradient information, such as shown in Equation 9, Equation 11 and Equation 12.

[0163] Any suitable method may be used to calculate the gradient associated with the input sample, such as the horizontal gradient G x Or vertical gradient G y Some examples are described below.

[0164] Fig.15 An example of spatial samples of an input sample C for calculating gradient information according to an aspect of the present disclosure is shown. In one example, an EIP filter is applied to a current sample (1501). Given an input sample C and neighboring samples of the input sample C (e.g., indicated using "N", "S", "E", "W", "NW", "NE", "SW", and "SE"), a separate gradient, such as G, can be calculated as follows: x and G y .

[0165] In one aspect, G can be calculated based on sample C, sample W, and sample N. x and G y . G y =(CN) G x =(CW) Equation 18

[0166] In one aspect, G can be calculated based on sample N, sample S, sample W, and sample E. x and G y . G y =(NS) Gx =(WE) Equation 19

[0167] In one aspect, G may be calculated based on sample N, sample S, sample W, sample E, sample NW, sample SW, sample NE, and sample SE. x and G y . G y =(2N+NW+NE)–(2S+SW+SE) G x =(2W+NW+SW)–(2E+NE+SE)Equation 20

[0168] In one example, the input sample C is a first input sample among the first input samples (eg, N1 input samples). The horizontal gradient G of the first input sample x The difference may be determined based on the following differences: (i) the difference between the first input sample and the left neighboring sample (e.g., W) of the first input sample (e.g., as shown in Equation 18); (ii) the difference between the left neighboring sample of the first input sample and the right neighboring sample (e.g., E) of the first input sample (e.g., as shown in Equation 19); and (iii) the difference between the first value and the second value (e.g., as shown in Equation 20), the first value (e.g., (2W+NW+SW)) may be the sum (e.g., weighted sum) of the upper left neighboring sample (e.g., NW), the left neighboring sample (e.g., W), and the lower left neighboring sample (e.g., SW) of the first input sample, and the second value (e.g., (2E+NE+SE)) may be the sum (e.g., weighted sum) of the upper right neighboring sample (e.g., NE), the right neighboring sample (e.g., E), and the lower right neighboring sample (e.g., SE) of the first input sample. In one example, it may be determined which difference is used to calculate the horizontal gradient of the first input sample based on the position of the first input sample.

[0169] In one example, the input sample C is a second input sample among the second input samples (eg, N2 input samples). The vertical gradient G of the second input sample y It can be determined based on the following difference values: (i) the difference between the second input sample and the upper adjacent sample of the second input sample (e.g., N) (e.g., as shown in Equation 18); (ii) the difference between the upper adjacent sample of the second input sample and the lower adjacent sample of the second input sample (e.g., S) (e.g., as shown in Equation 19); and (iii) the difference between the first value and the second value (e.g., as shown in Equation 20). yThe first value (e.g., (2N+NW+NE)) may be a sum (e.g., a weighted sum) based on the upper left neighboring sample (e.g., NW), the upper neighboring sample (e.g., N), and the upper right neighboring sample (e.g., NE) of the one second input sample. The second value (e.g., (2S+SW+SE)) may be a sum (e.g., a weighted sum) based on the lower left neighboring sample (e.g., SW), the lower neighboring sample (e.g., S), and the lower right neighboring sample (e.g., SE) of the one second input sample. In one example, which difference to use to calculate the vertical gradient G of the one second input sample may be determined based on the position of the one second input sample. y .

[0170] In one aspect, G can be calculated based on the position of the input sample C x and G y . Fig.16 An example of selecting a gradient calculation method (e.g., one of the methods described using equations 18-20) based on the position of the input samples according to an aspect of the present disclosure is shown. Different gradient calculation methods may be used to calculate the gradient {G of the eight input samples (1602)-(1609) according to the positions of the eight input samples (1602)-(1609). x , G y}. Input samples (1602)-(1605) are located at type 1 positions. For each input sample in input samples (1602)-(1605) at type 1 positions, G of the input sample can be calculated using sample C, sample W, and sample N for the corresponding input sample. x and G y , for example, using Equation 18 to calculate G x and G y . For example, sample C, sample W, and sample N for input sample (1602) are sample (1602), sample (1603), and sample (1606), respectively. For example, sample C, sample W, and sample N for input sample (1604) are sample (1604), sample (1606), and sample (1605), respectively. Input sample (1606) is located at a type 2 position. Sample N, sample S, sample W, and sample E can be used to calculate G of input sample (1606) at a type 2 position. x and G y , for example, using Equation 19 to calculate G x and G y Input samples (1607)-(1609) are located at type 3 positions. For each input sample in input samples (1607)-(1609) at type 3 positions, G of the input sample can be calculated using sample N, sample S, sample W, sample E, sample NW, sample SW, sample NE, and sample SE for the corresponding input sample. x and Gy , for example, using Equation 20 to calculate G x and G y .

[0171] In one aspect, the intra prediction of the current block may be derived by utilizing a non-linear term derived from neighboring reconstructed samples.

[0172] In one aspect, a nonlinear term (also referred to as a nonlinear value NP) can be used to further refine the predicted value, such as pred1(x,y) or pred(x,y) (e.g., pred0(x,y)). The nonlinear term can be generated based on, for example, the reconstructed or previously predicted values ​​of the immediate neighborhood of the current sample (e.g., the current prediction sample).

[0173] In one example, the refined prediction can be obtained based on the predicted values ​​predi(x,y) (e.g., pred0(x,y), pred1(x,y), etc.) and the nonlinear term using Equation 21:

[0174] c nonlinear is the coefficient of the nonlinear term. The nonlinear term NP can be defined using Equation 22. NP = (M × M + midVal) >> bitDepth Equation 22

[0175] M may be determined based on an average value (M=mean(A, L, AL)) or a median value (M=median(A, L, AL)) of neighboring samples A, L, and AL of the current sample. Fig.17 An example of the positions of sample A, sample L, and sample AL is shown. Sample A, sample L, and sample AL are neighboring samples of the current sample (1701).

[0176] In one example, for 10-bit content, the midVal is 2 10 / 2, which is 512. Therefore, for 10-bit content, the nonlinear term NP can be calculated using Equation 23. NP=(M×M+512)>>10 Equation 23

[0177] Can A clipping operation is applied to generate the final prediction pred2(x,y) using Equation 24.

[0178] In one example, based on the nonlinear relationship (eg, quadratic relationship) between the nonlinear value NP and the values ​​of the adjacent samples, the adjacent samples (eg, Fig.17The nonlinear value associated with the current sample is determined based on the sample A, sample L, and sample AL in the EIP mode, such as described in Equation 22. The predicted value of the current sample can be determined based on the initial predicted value predicted using the EIP mode and the additional information including the gradient information and the nonlinear value, such as shown in Equation 21. For example, predi(x,y) (e.g., pred0(x,y) or pred1(x,y)) may include the initial predicted value (e.g., P0 or P1) predicted using the EIP mode and the gradient information (e.g., GX and / or GY). In one example, reference Fig.17 , calculate a first value (e.g., M in Equation 22), the first value is the average or median of the upper left neighboring sample (AL), the upper neighboring sample (A), and the left neighboring sample (L) of the current sample (1701), and a square based on the first value (e.g., M 2 ) determines the nonlinear value.

[0179] In one aspect, the predicted value of the current sample predicted using the EIP mode is determined based on at least one of GX (e.g., calculated using Equation 11), GY (e.g., calculated using Equation 12), and the nonlinear term NP (e.g., calculated using Equation 22) and the initial predicted value Pi (e.g., calculated using Equation 10, Equation 13, etc.). In some examples, the predicted value of the current sample may be clipped.

[0180] The EIP filter may include different sets of coefficients, such as a first set (eg, c in Equation 10 or Equation 13). 0,i , which is associated with N0 input samples with or without the mean removal operation), a second set associated with GX (e.g., c in Equation 11) 1,j ), a third set related to GY (e.g., c in Eq. 12 2,k ), c related to the nonlinear term nonlinear When additional coefficient sets are included in the EIP filter, the same Figure 7 The filter coefficients in the EIP filter are obtained according to the neighboring reconstructed pixels (or samples) of the current block using a predetermined template similar to or the same as described in . For example, if the predicted value pred0(x, y) is calculated using equation 9, the EIP filter includes a first coefficient set {c 0,i}, the second coefficient set {c 1,j} and the third coefficient set {v 2,k}, so a predetermined template can be used to obtain {c 0,i}、{v 1,j} and {c 2,k}.

[0181] Fig.18A flow chart outlining a process (1800) according to an aspect of the present disclosure is shown. The process (1800) may be used in a video decoder. In various aspects, the process (1800) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some aspects, the process (1800) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (1800). The process starts at (S1801) and proceeds to (S1810).

[0182] At (S1810), prediction information is received, the prediction information indicating that a current block in a current picture is predicted using an extrapolation filter based intra prediction (EIP) mode.

[0183] At (S1820), gradient information associated with a current sample in a current block is determined.

[0184] In one example, when the gradient information includes horizontal gradient information, the number of first input samples for determining the horizontal gradient information and the position of the first input samples are set independently from the number and position of input samples for determining the initial prediction value. When the gradient information includes vertical gradient information, the number of second input samples for determining the vertical gradient information and the position of the second input samples are set independently from the number and position of input samples for determining the initial prediction value. When the gradient information includes horizontal gradient information and vertical gradient information, (i) the number and position of the first input samples and (ii) the number and position of the second input samples are set independently from each other and independently from the number and position of input samples for determining the initial prediction value.

[0185] In one example, the gradient information includes a sum of horizontal gradient information and vertical gradient information, and the number of first input samples is equal to the number of second input samples.

[0186] In one example, (i) the number of first input samples or the positions of the first input samples for determining horizontal gradient information in the gradient information, or (ii) the number of second input samples or the positions of the second input samples for determining vertical gradient information in the gradient information, depends on the block shape of the current block or the filter shape of the EIP mode.

[0187] In one example, the gradient information is determined based on at least one of horizontal gradient information determined based on a horizontal gradient of a corresponding first input sample and vertical gradient information determined based on a vertical gradient of a corresponding second input sample.

[0188] In one example, the horizontal gradient of one of the first input samples is determined based on the following differences: (i) the difference between the one first input sample and a left neighboring sample of the one first input sample; (ii) the difference between the left neighboring sample of the one first input sample and a right neighboring sample of the one first input sample; and (iii) the difference between a first value and a second value, the first value being based on the sum of the upper left neighboring samples, the left neighboring samples, and the lower left neighboring samples of the one first input sample, and the second value being based on the sum of the upper right neighboring samples, the right neighboring samples, and the lower right neighboring samples of the one first input sample. Based on the position of the one first input sample, it is determined which difference is used to calculate the horizontal gradient of the one first input sample.

[0189] In one example, the vertical gradient of one of the second input samples is determined based on the following differences: (i) the difference between the one second input sample and an upper adjacent sample of the one second input sample; (ii) the difference between the upper adjacent sample of the one second input sample and a lower adjacent sample of the one second input sample; and (iii) the difference between a first value and a second value, the first value being based on the sum of upper left adjacent samples, upper adjacent samples, and upper right adjacent samples of the one second input sample, and the second value being based on the sum of lower left adjacent samples, lower adjacent samples, and lower right adjacent samples of the one second input sample. Based on the position of the one second input sample, it is determined which difference is used to calculate the vertical gradient of the one second input sample.

[0190] At (S1830), a prediction value of the current sample is determined based on the initial prediction value predicted using the EIP mode and the additional information including the gradient information.

[0191] At (S1840), the current sample is reconstructed according to the predicted value of the current sample.

[0192] The process then proceeds to (S1899) and terminates.

[0193] The process (1800) may be adapted as appropriate. Steps in the process (1800) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0194] In one example, based on a nonlinear relationship between a nonlinear value associated with a current sample and values ​​of neighboring samples of the current sample, the nonlinear value is determined according to the neighboring samples; and a prediction value of the current sample is determined according to an initial prediction value predicted using an EIP mode and additional information including gradient information and the nonlinear value.

[0195] In one example, a first value is calculated, which is an average or median value of upper left neighboring samples, upper neighboring samples, and left neighboring samples of the current sample, and the nonlinearity value is determined based on the square of the first value.

[0196] Fig.19 A flow chart outlining a process (1900) according to an aspect of the present disclosure is shown. The process (1900) may be used for a video encoder. In various aspects, the process (1900) is performed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some aspects, the process (1900) is implemented in software instructions, so that when the processing circuit executes the software instructions, the processing circuit performs the process (1900). The process starts at (S1901) and proceeds to (S1910).

[0197] At (S1910), gradient information associated with a current sample in a current block is determined. The current block is predicted using an extrapolation filter based intra prediction (EIP) mode.

[0198] In one example, when the gradient information includes horizontal gradient information, the number of first input samples for determining the horizontal gradient information and the position of the first input samples are set independently from the number and position of input samples for determining the initial prediction value. When the gradient information includes vertical gradient information, the number of second input samples for determining the vertical gradient information and the position of the second input samples are set independently from the number and position of input samples for determining the initial prediction value. When the gradient information includes horizontal gradient information and vertical gradient information, (i) the number and position of the first input samples and (ii) the number and position of the second input samples are set independently from each other and independently from the number and position of input samples for determining the initial prediction value.

[0199] In one example, the gradient information includes a sum of horizontal gradient information and vertical gradient information, and the number of first input samples is equal to the number of second input samples.

[0200] In one example, (i) the number of first input samples or the positions of the first input samples for determining horizontal gradient information in the gradient information, (ii) the number of second input samples or the positions of the second input samples for determining vertical gradient information in the gradient information, and one of (i) and (ii) depends on the block shape of the current block or the filter shape of the EIP mode.

[0201] In one example, the gradient information is determined based on at least one of horizontal gradient information determined based on a horizontal gradient of a corresponding first input sample and vertical gradient information determined based on a vertical gradient of a corresponding second input sample.

[0202] In one example, the horizontal gradient of one of the first input samples is determined based on the following differences: (i) the difference between the one first input sample and a left neighboring sample of the one first input sample; (ii) the difference between the left neighboring sample of the one first input sample and a right neighboring sample of the one first input sample; and (iii) the difference between a first value and a second value, the first value being based on the sum of the upper left neighboring samples, the left neighboring samples, and the lower left neighboring samples of the one first input sample, and the second value being based on the sum of the upper right neighboring samples, the right neighboring samples, and the lower right neighboring samples of the one first input sample. Based on the position of the one first input sample, it is determined which difference is used to calculate the horizontal gradient of the one first input sample.

[0203] In one example, the vertical gradient of one of the second input samples is determined based on the following differences: (i) the difference between the one second input sample and an upper adjacent sample of the one second input sample; (ii) the difference between the upper adjacent sample of the one second input sample and a lower adjacent sample of the one second input sample; and (iii) the difference between a first value and a second value, the first value being based on the sum of upper left adjacent samples, upper adjacent samples, and upper right adjacent samples of the one second input sample, and the second value being based on the sum of lower left adjacent samples, lower adjacent samples, and lower right adjacent samples of the one second input sample. Based on the position of the one second input sample, it is determined which difference is used to calculate the vertical gradient of the one second input sample.

[0204] At (S1920), the nonlinear value is determined from neighboring samples of the current sample based on a nonlinear relationship between the nonlinear value associated with the current sample and values ​​of the neighboring samples of the current sample.

[0205] At (S1930), a prediction value of the current sample is determined based on the initial prediction value predicted using the EIP mode, the gradient information, and the nonlinear value.

[0206] Then, the process proceeds to (S1999) and terminates.

[0207] The process (1900) may be adapted as appropriate. Steps in the process (1900) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0208] In one aspect, a method for processing visual media data is disclosed. The method includes: performing conversion between a visual media file and a code stream of the visual media data according to a format rule. The code stream includes prediction information indicating that a current block in a current picture is predicted using an intra prediction (EIP) mode based on an extrapolation filter. The format rule specifies: determining gradient information associated with a current sample in a current block predicted using an EIP mode; determining the nonlinear value associated with the current sample based on a nonlinear relationship between a value of a neighboring sample of the current sample and the neighboring sample; determining a prediction value of the current sample based on an initial prediction value predicted based on the EIP mode, the gradient information, and the nonlinear value; when the gradient information includes horizontal gradient information, the number of first input samples used to determine the horizontal gradient information and the position of the first input samples are set independently from the number and position of input samples used to determine the initial prediction value; when the gradient information includes vertical gradient information, the number of second input samples used to determine the vertical gradient information and the position of the second input samples are set independently from the number and position of input samples used to determine the initial prediction value; and when the gradient information includes horizontal gradient information and vertical gradient information, (i) the number and position of the first input samples and (ii) the number and position of the second input samples are set independently from each other and are set independently from the number and position of input samples used to determine the initial prediction value.

[0209] In one example, the gradient information includes a sum of horizontal gradient information and vertical gradient information, and the number of first input samples is equal to the number of second input samples.

[0210] In one example, (i) the number of first input samples or the positions of the first input samples for determining horizontal gradient information in the gradient information, or (ii) the number of second input samples or the positions of the second input samples for determining vertical gradient information in the gradient information, depends on the block shape of the current block or the filter shape of the EIP mode.

[0211] In one example, determining the gradient information includes determining the gradient information based on at least one of: (i) horizontal gradient information determined based on a horizontal gradient of a corresponding first input sample and (ii) vertical gradient information determined based on a vertical gradient of a corresponding second input sample.

[0212] In one example, the horizontal gradient of one of the first input samples is determined based on the following difference values: (i) a difference between the one first input sample and a left adjacent sample of the one first input sample; (ii) a difference between the left adjacent sample of the one first input sample and a right adjacent sample of the one first input sample; and (iii) a difference between a first value and a second value, the first value being a sum of upper left adjacent samples, left adjacent samples, and lower left adjacent samples of the one first input sample, and the second value being a sum of upper right adjacent samples, right adjacent samples, and lower right adjacent samples of the one first input sample.

[0213] In one example, based on the position of the one of the first input samples, it is determined which difference to use to calculate the horizontal gradient of the one first input sample.

[0214] In one example, a vertical gradient of one of the second input samples is determined based on the following difference values: (i) a difference between the one second input sample and an upper neighboring sample of the one second input sample; (ii) a difference between an upper neighboring sample of the one second input sample and a lower neighboring sample of the one second input sample; and (iii) a difference between a first value and a second value, the first value being a sum of upper left neighboring samples, upper neighboring samples, and upper right neighboring samples of the one second input sample, and the second value being a sum of lower left neighboring samples, lower neighboring samples, and lower right neighboring samples of the one second input sample.

[0215] In one example, based on the position of the one of the second input samples, it is determined which difference to use to calculate the vertical gradient of the one of the second input samples.

[0216] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig. 20 A computer system (2000) suitable for implementing certain aspects of the disclosed subject matter is shown.

[0217] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.

[0218] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0219] Fig. 20 The components of the computer system (2000) shown are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary aspects of the computer system (2000).

[0220] The computer system (2000) may include certain human interface input devices. Such human interface input devices may be responsive to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, captured images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0221] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (2001), mouse (2002), touch pad (2003), touch screen (2010), data gloves (not shown), joystick (2005), microphone (2006), scanner (2007), camera (2008).

[0222] The computer system (2000) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (2010), a data glove (not shown), or a joystick (2005), but may also be a tactile feedback device that is not an input device), audio output devices (e.g., speakers (2009), headphones (not depicted)), visual output devices (e.g., screens (2010) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual outputs or outputs exceeding three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), and printers (not depicted).

[0223] The computer system (2000) may also include human-machine accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2020) having CD / DVD and other media (2021), thumb drives (2022), removable hard drives or solid-state drives (2023), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security software dogs (not depicted), etc.

[0224] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0225] The computer system (2000) may also include an interface (2054) to one or more communication networks (2055). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some general data port or peripheral bus (2049) (e.g., a USB port of the computer system (2000)); other network interfaces are typically integrated into the kernel of the computer system (2000) by attaching to a system bus as described below (e.g., connected to an Ethernet interface in a PC computer system or connected to a cellular network interface in a smartphone computer system). The computer system (2000) can use any of these networks to communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANBus connected to certain CANBus devices), or bidirectional, for example, using a LAN or WAN digital network to connect to other computer systems. Certain protocols and protocol stacks may be used on each of those networks and network interfaces as described above.

[0226] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel (2040) of the computer system (2000).

[0227] The kernel (2040) may include one or more central processing units (CPUs) (2041), graphics processing units (GPUs) (2042), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (2043), hardware accelerators (2044) for certain tasks, graphics adapters (2050), etc. These devices, as well as read-only memory (ROM) (2045), random access memory (2046), internal mass storage (2047) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (2048). In some computer systems, the system bus (2048) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the kernel's system bus (2048) or to the kernel's system bus (2048) via a peripheral bus (2049). In one example, a screen (2010) may be connected to a graphics adapter (2050). The architecture of the peripheral bus includes PCI, USB, etc.

[0228] The CPU (2041), GPU (2042), FPGA (2043) and accelerator (2044) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (2045) or RAM (2046). Transition data can also be stored in RAM (2046), while permanent data can be stored, for example, in internal mass storage (2047). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more CPUs (2041), GPUs (2042), mass storage (2047), ROM (2045), RAM (2046), etc.

[0229] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The medium and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer codes may be of a type well known and available to those skilled in the art of computer software.

[0230] As an example, and not by way of limitation, a computer system (2000) having an architecture, particularly a kernel (2040), can provide functionality because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage as described above, as well as memories of certain non-temporary kernels (2040), such as kernel internal mass storage (2047) or ROM (2045). Software implementing various aspects of the present disclosure can be stored in such devices and executed by the kernel (2040). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the kernel (2040), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes described herein or specific parts of specific processes described herein, including defining data structures stored in RAM (2046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality due to logic hardwired or otherwise embodied in a circuit (e.g., accelerator (2044)) that may replace software or operate in conjunction with software to perform specific processes described herein or specific portions of specific processes described herein. Where appropriate, references to portions of software may include logic and vice versa. Where appropriate, references to portions of computer-readable media may include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0231] As used in this disclosure, "at least one of" or "one of" is intended to include any one or combination of the listed elements. For example, references to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A to C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B, and one of A and B are intended to include A or B or (A and B). Where applicable, the use of "one of" does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.

[0232] Although the present disclosure has described several exemplary aspects, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.

Claims

1. A method for processing visual media data, the method comprising: Conversion between the visual media file and the code stream of the visual media data is performed according to the format rules, wherein: The code stream includes prediction information indicating that a current block in a current picture is predicted using an extrapolation filter-based intra prediction (EIP) mode; and The format rules specify: determining gradient information associated with a current sample in the current block predicted using the EIP mode; Determine the nonlinear value based on the neighboring samples of the current sample using a nonlinear relationship between the nonlinear value associated with the current sample and values ​​of neighboring samples of the current sample; Determining a prediction value of the current sample according to an initial prediction value predicted based on the EIP mode, the gradient information, and the nonlinear value; When the gradient information includes horizontal gradient information, the number of first input samples used to determine the horizontal gradient information and the positions of the first input samples are set independently from the number and positions of input samples used to determine the initial prediction value; When the gradient information includes vertical gradient information, the number of second input samples used to determine the vertical gradient information and the positions of the second input samples are set independently from the number and positions of the input samples used to determine the initial prediction value; and When the gradient information includes the horizontal gradient information and the vertical gradient information, (i) the number and positions of the first input samples and (ii) the number and positions of the second input samples are set independently of each other and independently of the number and positions of the input samples used to determine the initial prediction value.

2. The method according to claim 1, wherein: The gradient information includes the sum of the horizontal gradient information and the vertical gradient information; and The number of the first input samples is equal to the number of the second input samples.

3. The method according to claim 1, wherein: (i) for determining the number of the first input samples of the horizontal gradient information in the gradient information or the position of the first input samples, (ii) for determining the number of the second input samples of the vertical gradient information in the gradient information or the position of the second input samples, one of (i) and (ii) depends on the block shape of the current block or the filter shape of the EIP mode.

4. The method according to claim 1, wherein: Determining the gradient information includes determining the gradient information based on at least one of: (i) the horizontal gradient information determined based on the horizontal gradient of the corresponding first input sample; and (ii) the vertical gradient information determined based on the vertical gradient of the corresponding second input sample.

5. A method for video encoding, comprising: determining gradient information associated with a current sample in a current block, the current block being predicted using an extrapolation filter based intra prediction (EIP) mode; Determining the nonlinear value according to the neighboring samples of the current sample based on a nonlinear relationship between the nonlinear value associated with the current sample and values ​​of neighboring samples of the current sample; as well as A prediction value of the current sample is determined based on an initial prediction value predicted using the EIP mode, the gradient information, and the nonlinear value.

6. The method according to claim 5, wherein: When the gradient information includes horizontal gradient information, the number of first input samples used to determine the horizontal gradient information and the positions of the first input samples are set independently from the number and positions of input samples used to determine the initial prediction value; When the gradient information includes vertical gradient information, the number of second input samples used to determine the vertical gradient information and the positions of the second input samples are set independently from the number and positions of the input samples used to determine the initial prediction value; as well as When the gradient information includes the horizontal gradient information and the vertical gradient information, (i) the number and positions of the first input samples and (ii) the number and positions of the second input samples are set independently of each other and independently of the number and positions of the input samples used to determine the initial prediction value.

7. The method according to claim 5, wherein: (i) for determining the number of first input samples of horizontal gradient information in the gradient information or the position of the first input samples, (ii) for determining the number of second input samples of vertical gradient information in the gradient information or the position of the second input samples, one of (i) and (ii) depends on the block shape of the current block or the filter shape of the EIP mode.

8. The method according to claim 5, wherein: Determining the gradient information includes determining the gradient information according to at least one of horizontal gradient information determined based on a horizontal gradient of a corresponding first input sample and vertical gradient information determined based on a vertical gradient of a corresponding second input sample.

9. A device for video decoding, comprising: The processing circuit is configured to: Receiving prediction information indicating that a current block in a current picture is predicted using an extrapolation filter based intra prediction (EIP) mode; determining gradient information associated with a current sample in the current block; Determining a predicted value of the current sample based on an initial predicted value predicted using the EIP mode and additional information including the gradient information; as well as The current sample is reconstructed according to the predicted value of the current sample.

10. The device according to claim 9, wherein: The processing circuit is configured to: Determining the nonlinear value according to the neighboring samples of the current sample based on a nonlinear relationship between the nonlinear value associated with the current sample and values ​​of neighboring samples of the current sample; as well as The predicted value of the current sample is determined according to the initial predicted value predicted using the EIP mode and the additional information including the gradient information and the nonlinear value.

11. The device according to claim 9 or 10, wherein: When the gradient information includes horizontal gradient information, the number of first input samples used to determine the horizontal gradient information and the positions of the first input samples are set independently from the number and positions of input samples used to determine the initial prediction value; When the gradient information includes vertical gradient information, the number of second input samples used to determine the vertical gradient information and the positions of the second input samples are set independently from the number and positions of the input samples used to determine the initial prediction value; as well as When the gradient information includes the horizontal gradient information and the vertical gradient information, (i) the number and positions of the first input samples and (ii) the number and positions of the second input samples are set independently of each other and independently of the number and positions of the input samples used to determine the initial prediction value.

12. The device according to claim 11, wherein The gradient information includes the sum of the horizontal gradient information and the vertical gradient information; and The number of the first input samples is equal to the number of the second input samples.

13. The device according to claim 9 or 10, wherein: (i) for determining the number of first input samples of horizontal gradient information in the gradient information or the position of the first input samples, (ii) for determining the number of second input samples of vertical gradient information in the gradient information or the position of the second input samples, one of (i) and (ii) depends on the block shape of the current block or the filter shape of the EIP mode.

14. The device according to claim 9 or 10, wherein: The processing circuit is configured to determine the gradient information based on at least one of horizontal gradient information determined based on a horizontal gradient of a corresponding first input sample and vertical gradient information determined based on a vertical gradient of a corresponding second input sample.

15. The device according to claim 10, wherein: The processing circuit is configured to: Calculating a first value, where the first value is an average or median value of an upper left adjacent sample, an upper adjacent sample, and a left adjacent sample of the current sample; and The nonlinearity value is determined based on the square of the first value.

Citation Information

Cited By

  • Encoding method, decoding method, bitstream, encoders, decoders and storage medium

    WO2026193696A1