Intra prediction with bias based extrapolation

The recursive filter with a bias term improves video codec efficiency by stabilizing predictions in large blocks, addressing inefficiencies in existing extrapolation-based methods.

WO2025153216A1PCT designated stage expired Publication Date: 2025-07-24NOKIA TECHNOLOGIES OY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/083056
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-17
Filing Date
2024-11-21
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing video codecs face challenges in accurately predicting pixel values using extrapolation-based filters, especially when the size of the prediction block is large, leading to inefficiencies due to mismatched filter coefficients and disconnect between trained and actual content.

Method used

Implementing a recursive filter that combines a spatial kernel with a bias term for intra sample prediction, using inputs that include both reconstructed and previously predicted samples, along with a constant value, to stabilize the prediction process.

Benefits of technology

Enhances coding efficiency by stabilizing filter outputs and improving prediction accuracy, particularly in large prediction blocks, thereby optimizing video compression and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000022_0001
    Figure IMGF000022_0001
  • Figure IMGF000023_0001
    Figure IMGF000023_0001
  • Figure IMGF000023_0002
    Figure IMGF000023_0002
Patent Text Reader

Abstract

An apparatus configured to: determine a block of one or more samples to be predicted; determine a filter, wherein inputs to the filter comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determine at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and apply the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.
Need to check novelty before this filing date? Find Prior Art

Description

INTRA PREDICTION WITH BIAS BASED EXTRAPOLATIONTECHNICAL FIELD

[0001] The example and non-limiting embodiments relate generally to intra sample prediction and, more particularly, to selection of inputs to a filter for performing recursive intra sample prediction.BACKGROUND

[0002] It is known, for performing intra prediction, to use an extrapolation based filter.SUMMARY

[0003] The following summary is merely intended to be illustrative. The summary is not intended to limit the scope of the claims.

[0004] In accordance with one aspect, an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine a block of one or more samples to be predicted; determine a filter for performing intra sample prediction, wherein inputs to the filter comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determine at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and apply the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0005] In accordance with one aspect, a method comprising: determining, with a user equipment, a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0006] In accordance with one aspect, an apparatus comprising means for: determining a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0007] In accordance with one aspect, a non-transitory computer-readable medium comprising program instructions stored thereon for performing at least the following: determining a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one ormore samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0008] According to some aspects, there is provided the subject matter of the independent claims. Some further aspects are defined in the dependent claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The foregoing aspects and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:

[0010] FIG. 1 is a block diagram of one possible and non-limiting example system in which the example embodiments may be practiced;

[0011] FIG. 2 is a block diagram of one possible and non-limiting exemplary system in which the example embodiments may be practiced;

[0012] FIG. 3 is a diagram illustrating features as described herein;

[0013] FIG. 4 is a diagram illustrating features as described herein;

[0014] FIG. 5 is a diagram illustrating features as described herein;

[0015] FIG. 6 is a diagram illustrating features as described herein;

[0016] FIG. 7 is a diagram illustrating features as described herein; and

[0017] FIG. 8 is a flowchart illustrating steps as described herein.DETAILED DESCRIPTION OF EMBODIMENTS

[0018] The following abbreviations that may be found in the specification and / or the drawing figures are defined as follows:3 GPP third generation partnership project4G fourth generation5G fifth generation5GC 5G core networkAR augmented realityAVC Advanced Video Coding (ITU-T H.264 video coding standard)BCW Bi-prediction with CU-level weightCABAC Context Adaptive Binary Arithmetic CoderCCCM Convolutional Cross-Component ModelCCRM Cross-Component Residual Model / Cross-ComponentReconstruction ModelCDMA code division multiple accessCIIP Combined Inter and Intra PredictionCPU central processing unit cRAN cloud radio access networkCTU coding tree unitCU coding unitDCT discrete cosine transformDPB decoded picture bufferDST Discrete Sine TransformECM Enhanced Compression Model (JVET’s exploratory video codec) eNB (or eNodeB) evolved Node B (e.g., an LTE base station)EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN- DCE-UTRA evolved universal terrestrial radio access, i.e., the LTE radio access technologyFDMA frequency division multiple accessgNB (or gNodeB) base station for 5G / NR, i.e., a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCGPU graphical processing unit GSM global systems for mobile communicationsHEVC High Efficiency Video Coding (ITU-T H.265 video coding standard)HMD head-mounted displayIBC intra block copy IEEE Institute of Electrical and Electronics EngineersIMD integrated messaging deviceIMS instant messaging service loT Internet of ThingsLCU largest coding unit LIC Local Illumination CompensationLTE long term evolutionMMS multimedia messaging serviceMPEG-I Moving Picture Experts Group immersive codec familyMR mixed reality MSE mean squared error ng or NG new generation ng-eNB or NG-eNB new generation eNBNR new radioN / W or NW network O-RAN open radio access networkPC personal computerPDA personal digital assistantPU prediction unitRGB Red, Green, Blue SMS short messaging serviceSNR signal-to-noise ratioTCP -IP transmission control protocol-internet protocolTDMA time division multiple accessTU transform unitUE user equipment (e.g., a wireless, typically mobile device)UMTS universal mobile telecommunications systemUSB universal serial busVNR virtualized network functionVR virtual realityWC Versatile Video Coding (ITU-T H.266 video coding standard)WEAN wireless local area networkWP Weighted PredictionYUV / YCbCr a color model based on one luminance and two chrominance / color difference channels (typically used in many video coding applications)

[0019] The following describes suitable apparatus and possible mechanisms for practicing example embodiments of the present disclosure. Accordingly, reference is first made to FIG. 1, which shows an example block diagram of an apparatus 50. The apparatus may be configured to perform various functions such as, for example, gathering information by one or more sensors, encoding and / or decoding information, receiving and / or transmitting information, analyzing information gathered or received by the apparatus, or the like. A device configured to encode a video scene may (optionally) comprise one or more microphones for capturing the scene and / or one or more sensors, such as cameras, for capturing information about the physical environment in which the scene is captured. Alternatively, a device configured to encode a video scene may be configured to receive information about an environment in which a scene is captured and / or a simulated environment. A device configured to decode and / or render the video scene may be configured to receive a Moving Picture Experts Group immersive codec family (MPEG-I) bitstream comprising the encoded video scene. A device configured to decode and / or render thevideo scene may comprise one or more speakers / audio transducers and / or displays, and / or may be configured to transmit a decoded scene or signals to a device comprising one or more speakers / audio transducers and / or displays. A device configured to decode and / or render the video scene may comprise a user equipment, a head / mounted display, or another device capable of rendering to a user an AR, VR and / or MR experience.

[0020] The electronic device 50 may for example be a mobile terminal or user equipment of a wireless communication system. Alternatively, the electronic device may be a computer or part of a computer that is not mobile. It should be appreciated that example embodiments of the present disclosure may be implemented within any electronic device or apparatus which may process data. The electronic device 50 may comprise a device that can access a network and / or cloud through a wired or wireless connection. The electronic device 50 may comprise one or more processors 56, one or more memories 58, and one or more transceivers 52 interconnected through one or more buses. The one or more processors 56 may comprise a central processing unit (CPU) and / or a graphical processing unit (GPU). Each of the one or more transceivers 52 includes a receiver and a transmitter. The one or more buses may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. A “circuit” may include dedicated hardware or hardware in association with software executable thereon. The one or more transceivers may be connected to one or more antennas 44. The one or more memories 58 may include computer program code. The one or more memories 58 and the computer program code may be configured to, with the one or more processors 56, cause the electronic device 50 to perform one or more of the operations as described herein.

[0021] The electronic device 50 may connect to a node of a network. The network node may comprise one or more processors, one or more memories, and one or more transceivers interconnected through one or more buses. Each of the one or more transceivers includes a receiver and a transmitter. The one or more buses may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. The one or more transceivers may be connected to one or more antennas. The one or more memories may include computerprogram code. The one or more memories and the computer program code may be configured to, with the one or more processors, cause the network node to perform one or more of the operations as described herein.

[0022] The electronic device 50 may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The electronic device 50 may further comprise an audio output device 38 which in example embodiments of the present disclosure may be any one of: an earpiece, speaker, or an analogue audio or digital audio output connection. The electronic device 50 may also comprise a battery (or in other example embodiments of the present disclosure the device may be powered by any suitable mobile energy device such as solar cell, fuel cell, or clockwork generator). The electronic device 50 may further comprise a camera 42 or other sensor capable of recording or capturing images and / or video. Additionally or alternatively, the electronic device 50 may further comprise a depth sensor. The electronic device 50 may further comprise a display 32. The electronic device 50 may further comprise an infrared port for short range line of sight communication to other devices. In other example embodiments of the present disclosure the apparatus 50 may further comprise any suitable short-range communication solution such as for example a BLUETOOTH™ wireless connection or a USB / firewire wired connection.

[0023] It should be understood that an electronic device 50 configured to perform example embodiments of the present disclosure may have fewer and / or additional components, which may correspond to what processes the electronic device 50 is configured to perform. For example, an apparatus configured to encode a video might not comprise a speaker or audio transducer and may comprise a microphone, while an apparatus configured to render the decoded video might not comprise a microphone and may comprise a speaker or audio transducer.

[0024] Referring now to FIG. 1, the electronic device 50 may comprise a controller 56, processor or processor circuitry for controlling the apparatus 50. The controller 56 may be connected to memory 58 which in example embodiments of the present disclosure may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable forcarrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.

[0025] The electronic device 50 may further comprise a card reader 48 and a smart card 46, for example a UICC and UICC reader, for providing user information and being suitable for providing authentication information for authentication and authorization of the user / electronic device 50 at a network. The electronic device 50 may further comprise an input device 34, such as a keypad, one or more input buttons, or a touch screen input device, for providing information to the controller 56.

[0026] The electronic device 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).

[0027] The electronic device 50 may comprise a microphone 38, camera 42, and / or other sensors capable of recording or detecting audio signals, image / video signals, and / or other information about the local / virtual environment, which are then passed to the codec 54 or the controller 56 for processing. The electronic device 50 may receive the audio / image / video signals and / or information about the local / virtual environment for processing from another device prior to transmission and / or storage. The electronic device 50 may also receive either wirelessly or by a wired connection the audio / image / video signals and / or information about the local / virtual environment for encoding / decoding. The structural elements of electronic device 50 described above represent examples of means for performing a corresponding function.

[0028] The memory 58 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devicesand systems, fixed memory and removable memory. The memory 58 may be a non-transitory memory. The memory 58 may be means for performing storage functions. The controller 56 may be or comprise one or more processors, which may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multicore processor architecture, as non-limiting examples. The controller 56 may be means for performing functions.

[0029] The electronic device 50 may be configured to perform capture of a volumetric scene according to example embodiments of the present disclosure. For example, the electronic device 50 may comprise a camera 42 or other sensor capable of recording or capturing images and / or video. The electronic device 50 may also comprise one or more transceivers 52 to enable transmission of captured content for processing at another device. Such an electronic device 50 may or may not include all the modules illustrated in FIG. 1.

[0030] The electronic device 50 may be configured to perform processing of volumetric video content according to example embodiments of the present disclosure. For example, the electronic device 50 may comprise a controller 56 for processing images to produce volumetric video content, a controller 56 for processing volumetric video content to project 3D information into 2D information, patches, and auxiliary information, and / or a codec 54 for encoding 2D information, patches, and auxiliary information into a bitstream for transmission to another device with radio interface 52. Such an electronic device 50 may or may not include all the modules illustrated in FIG. 1.

[0031] The electronic device 50 may be configured to perform encoding or decoding of 2D information representative of volumetric video content according to example embodiments of the present disclosure. For example, the electronic device 50 may comprise a codec 54 for encoding or decoding 2D information representative of volumetric video content. Such an electronic device 50 may or may not include all the modules illustrated in FIG. 1.

[0032] The electronic device 50 may be configured to perform rendering of decoded 3D volumetric video according to example embodiments of the present disclosure. For example, the electronic device 50 may comprise a controller for projecting 2D information to reconstruct 3D volumetric video, and / or a display 32 for rendering decoded 3D volumetric video. Such an electronic device 50 may or may not include all the modules illustrated in FIG. 1.

[0033] With respect to FIG. 2, an example of a system within which example embodiments of the present disclosure can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, E-UTRA, LIE, CDMA, 4G, 5G network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a BLUETOOTH™ personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and / or the Internet. A wireless network may implement network virtualization, which is the process of combining hardware and software network resources and network functionality into a single, software-based administrative entity, a virtual network. Network virtualization involves platform virtualization, often combined with resource virtualization. Network virtualization is categorized as either external, combining many networks, or parts of networks, into a virtual unit, or internal, providing network-like functionality to software containers on a single system. For example, a network may be deployed in a tele cloud, with virtualized network functions (VNF) running on, for example, data center servers. For example, network core functions and / or radio access network(s) (e.g. CloudRAN, O-RAN, edge cloud) may be virtualized. Note that the virtualized entities that result from the network virtualization are still implemented, at some level, using hardware such as processors and memories, and also such virtualized entities create technical effects.

[0034] It may also be noted that operations of example embodiments of the present disclosure may be carried out by a plurality of cooperating devices (e.g. cRAN).

[0035] The system 10 may include both wired and wireless communication devices and / or electronic devices suitable for implementing example embodiments of the present disclosure.

[0036] For example, the system shown in FIG. 2 shows a mobile telephone network 11 and a representation of the internet 28. Connectivity to the internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.

[0037] The example communication devices shown in the system 10 may include, but are not limited to, an apparatus 15, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22, and a head-mounted display (HMD) 17. The electronic device 50 may comprise any of those example communication devices. In an example embodiment of the present disclosure, more than one of these devices, or a plurality of one or more of these devices, may perform the disclosed process(es). These devices may connect to the internet 28 through a wireless connection 2.

[0038] The example embodiments of the present disclosure may also be implemented in a set-top box; i.e. a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding. The example embodiments of the present disclosure may also be implemented in cellular telephones such as smart phones, tablets, personal digital assistants (PDAs) having wireless communication capabilities, portable computers having wireless communication capabilities, image capture devices such as digital cameras having wireless communication capabilities, gaming devices having wireless communication capabilities, music storage and playback appliances having wireless communication capabilities, Internet appliances permitting wireless Internet access and browsing, tablets with wireless communication capabilities, as well as portable units or terminals that incorporate combinations of such functions.

[0039] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24, which may be, for example, an eNB, gNB, access point, access node, other node, etc. The base station 24 may beconnected to a network server 26 that allows communication between the mobile telephone network 11 and the internet 28. The system may include additional communication devices and communication devices of various types.

[0040] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), BLUETOOTH™, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various example embodiments of the present disclosure may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.

[0041] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, which may be a MPEG-I bitstream, from one or several senders (or transmitters) to one or several receivers.

[0042] Having thus introduced one suitable but non-limiting technical context for the practice of the example embodiments of the present disclosure, example embodiments will now be described with greater specificity.

[0043] Features as described herein may generally relate to coding and decoding of digital video material.

[0044] A video codec consists of an encoder that transforms an input video into a compressed representation suited for storage / transmission, and a decoder that can decompress the compressed video representation back into a viewable form. Typically, the encoder discards some informationin the original video sequence to be able to represent the video in a more compact form (that is, at lower bitrate).

[0045] Typical hybrid video codecs, such as H.264 / AVC, H.265 / HEVC and H.266 / VVC, encode the video information in two phases. Firstly, pixel values in a certain picture area (or “block”) (310) may be predicted (320) for example by motion compensation means (finding and indicating an area in one of the previously coded pictures that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, may be coded (330). This may typically be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of it) (340), quantizing the resulting transform coefficients (350), and entropy coding the quantized coefficients (360). By varying the fidelity of the quantization process, the encoder may control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate). An example of the encoding process is illustrated in FIG. 3.

[0046] In some video codecs, such as H.265 / HEVC and H.266 / VVC, the video pictures may be divided into coding units (CU) covering the area of the picture. A CU may consist of one or more prediction units (PU) defining the prediction process for the samples within the CU and one or more transform units (TU) defining the prediction error coding process for the samples in the said CU. Typically, a CU may consist of a rectangular block of samples with a size selectable from a predefined set of possible CU sizes. A CU with the maximum allowed size may typically be named as LCU (largest coding unit) or CTU (coding tree unit), and the video picture may be divided into non-overlapping CTUs. A CTU may be further split into a combination of smaller CUs, e.g. by recursively splitting the CTU and resultant CUs. Each resulting CU typically may have at least one PU and at least one TU associated with it. Each PU and TU may be further split into smaller PUs and TUs to increase granularity of the prediction and prediction error coding processes, respectively. Each PU may have prediction information associated with it defining what kind of a prediction is to be applied for the pixels within that PU (e.g. motion vector information for inter predicted PUs and intra prediction directionality information for intrapredicted PUs). Similarly, each TU may be associated with information describing the prediction error decoding process for the samples within the TU (including e.g. DCT coefficient information). It is typically signaled at CU level whether prediction error coding may be applied or not for each CU. In the case there is no prediction error residual associated with the CU, it may be considered that there are no TUs for the said CU. The division of the image into CUs, and division of CUs into PUs and Tus, may typically be signaled in the bitstream, and may allow the decoder to reproduce the intended structure of these units.

[0047] The decoder may reconstruct the output video by applying prediction means (410) similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain) (420). After applying prediction and prediction error decoding means, the decoder may sum up the prediction and prediction error signals (pixel values) to form the output video frame (440). The decoder (and encoder) may also apply additional filtering means (430) to improve the quality of the output video before passing it for display and / or storing it as a prediction reference for the forthcoming frames in the video sequence. An example of the decoding process is illustrated in FIG. 4.

[0048] Instead, or in addition to approaches utilizing sample value prediction and transform coding for indicating the coded sample values, a color palette based coding may be used. Palette based coding refers to a family of approaches for which a palette, i.e. a set of colors and associated indexes, is defined and the value for each sample within a coding unit is expressed by indicating its index in the palette. Palette based coding may typically achieve good coding efficiency in coding units with a relatively small number of colors (such as image areas which are representing computer screen content, like text or simple graphics). In order to improve the coding efficiency of palette coding, different kinds of palette index prediction approaches may be utilized, or the palette indexes may be run-length coded to be able to represent larger homogenous image areas efficiently. Also, in the case the CU contains sample values that are not recurring within the CU, escape coding may be utilized. Escape coded samples may be transmitted without referring to anyof the palette indexes. Instead, their values may be indicated individually for each escape coded sample.

[0049] In typical video codecs the motion information may be indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors may represent the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side), and the prediction source block in one of the previously coded or decoded pictures. To represent motion vectors efficiently, those may typically be coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors may be created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions may be to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture may be predicted. The reference index may typically be predicted from adjacent blocks and / or or co-located blocks in a temporal reference picture. Moreover, typical high efficiency video codecs may employ an additional motion information coding / decoding mechanism, often called merging or merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, may be predicted and used without any modification / correction. Similarly, predicting the motion field information may be carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures, and the used motion field information may be signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.

[0050] Typically, video codecs may support motion compensated prediction from at least one source image (uni-prediction) and two sources (bi-prediction). In the case of uni-prediction, a single motion vector may be applied, whereas in the case of bi-prediction, two motion vectors may be determined and the motion compensated predictions from two sources may be combined to create the final sample prediction. In the case of weighted prediction, the relative weights of the two predictions may be adjusted, or a signaled offset may be added to the prediction signal.

[0051] In addition to applying motion compensation for inter picture prediction, a similar approach may be applied to intra picture prediction. In this case, the displacement vector may indicate from where the same picture a block of samples may be copied to form a prediction of the block to be coded or decoded. This kind of intra block copying (IBC) method(s) may improve the coding efficiency substantially in presence of repeating structures within the frame, such as text or other graphics.

[0052] In typical video codecs, the prediction residual after motion compensation or intra prediction may first be transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual, and transform may in many cases help reduce this correlation and provide more efficient coding.

[0053] Typical video encoders may utilize Lagrangian cost functions to find optimal coding modes, e.g. the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor A. to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + R Eq. 1

[0054] Where C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).

[0055] Scalable video coding refers to coding structure(s) where one bitstream may contain multiple representations of the content at different bitrates, resolutions, or frame rates. In these cases, the receiver may extract the desired representation depending on its characteristics (e.g. resolution that matches best the display device). Alternatively, a server or a network element may extract the portions of the bitstream to be transmitted to the receiver depending on e.g. the network characteristics or processing capabilities of the receiver. A scalable bitstream may typically consist of a “base layer” providing the lowest quality video available, and one or moreenhancement layers that enhance the video quality when received and decoded together with the lower layers. To improve coding efficiency for the enhancement layers, the coded representation of that layer may typically depend on the lower layers. For example, the motion and mode information of the enhancement layer may be predicted from lower layers. Similarly, the pixel data of the lower layers may be used to create prediction for the enhancement layer.

[0056] A scalable video codec for quality scalability (also known as Signal-to-Noise or SNR) and / or spatial scalability may be implemented as follows. For a base layer, a conventional non- scalable video encoder and decoder may be used. The reconstructed / decoded pictures of the base layer may be included in the reference picture buffer for an enhancement layer. In H.264 / AVC, H.265 / HEVC, and similar codecs using reference picture list(s) for inter prediction, the base layer decoded pictures may be inserted into a reference picture list(s) for coding / decoding of an enhancement layer picture, similarly to the decoded reference pictures of the enhancement layer. Consequently, the encoder may choose a base-layer reference picture as inter prediction reference, and indicate its use typically with a reference picture index in the coded bitstream. The decoder decodes from the bitstream, for example from a reference picture index, that a base-layer picture may be used as an inter prediction reference for the enhancement layer. When a decoded baselayer picture is used as prediction reference for an enhancement layer, it may be referred to as an inter-layer reference picture.

[0057] In addition to quality scalability, examples of other scalability modes include:

[0058] - Spatial scalability: Enhancement layer pictures are coded at a higher resolution than the base layer pictures

[0059] - Bit-depth scalability: Enhancement layer pictures are coded at higher bit-depth (e.g.10 or 12 bits) than base layer pictures (e.g. 8 bits)

[0060] - Chroma format scalability: Enhancement layer pictures provide higher fidelity in chroma (e.g. coded in 4:4:4 chroma format) than base layer pictures (e.g. 4:2:0 format)

[0061] In all of the above scalability cases, base layer information may be used to code enhancement layer to minimize the additional bitrate overhead.

[0062] Scalability may be enabled in two basic ways: either by introducing new coding modes for performing prediction of pixel values or syntax from lower layers of the scalable representation, or by placing the lower layer pictures to the reference picture buffer (decoded picture buffer, DPB) of the higher layer. The first approach is more flexible and thus may provide better coding efficiency in most cases. However, the second, reference frame based scalability approach may be implemented very efficiently, with minimal changes to single layer codecs while still achieving majority of the coding efficiency gains available. Essentially, a reference frame based scalability codec may be implemented by utilizing the same hardware or software implementation for all the layers, just taking care of the DPB management by external means.

[0063] To be able to utilize parallel processing, images may be split into independently codable and decodable image segments (e.g. slices or tiles). Slices typically refer to image segments constructed of a certain number of basic coding units that are processed in default coding or decoding order, while tiles typically refer to image segments that have been defined as rectangular image regions that are processed, at least to some extent, as individual frames.

[0064] Typically, video may be encoded in the YUV or YCbCr color space, as that is found to reflect some characteristics of the human visual system and allows using lower quality representation for Cb and Cr channels, as human perception is less sensitive to the chrominance fidelity those channels represent.

[0065] Typical video codecs, such as H.265 / HEVC (ITU-T recommendation H.265: “High efficiency video coding”, https: / / www.itu.int / rec / T-REC-H.265) and H.266 / WC (ITU-T recommendation H.266: “Versatile video coding”, http: / / www.itu.int / rec / T-REC-H.266) standards split images into blocks of samples which are predicted in different ways from reconstructed samples of adjacent blocks. Such prediction processes typically extrapolate the samples of adjacent blocks to fill the sample block to be predicted with values generated with a determined filtering process. The process is typically iterative, relying on the sample values ofthose adjacent blocks, allowing any of the sample(s) in the prediction block to be calculated independently from others.

[0066] Alternatively, recursive filters may be used. In the case of recursive filtering, the output samples of earlier steps of the filtering process may be used as inputs to predict one or more new sample values. This kind of process is proposed, for example, in JVET document JVET- AF0080: “EE2-2.7: An extrapolation filter-based intra prediction mode”, October 2023, where NxM neighboring sample values are used to generate a new predicted sample value. In this case, the neighboring sample values used as inputs may include both samples from the adjacent blocks of samples, as well as predicted sample values for the block of samples that is being predicted. The filter may typically be constructed or trained using the reconstructed samples of the neighboring blocks and then applied to predict samples of the current block.

[0067] Extrapolation based filters are sensitive to difference of the range of values used in training the filter, and the range of values that the filter is expected to predict. This becomes an issue especially if the size of the prediction block is relatively large, as the filter coefficients trained using the reconstructed samples of the neighboring blocks may no longer be able to represent properly the content in the prediction block.

[0068] In addition, quite often in modern codecs there are so-called intra prediction “merge modes” enabled. Those modes copy parameters of earlier prediction blocks and apply those parameters to the current prediction block. When using a traditional extrapolation based filter that has been constructed for another prediction block and trained using neighbors of that other prediction block, the disconnect between the parameters and the content of the actual prediction block becomes even larger.

[0069] In an example embodiment, a recursive filter may be used that combines a spatial kernel to a specifically trained bias term to perform intra sample prediction for a block of samples. Alternative configurations may use symmetric filter kernels, where each may omit input samples from two opposite corners (i.e. not include those input samples). Additionally or alternatively, alternative configurations may use filter kernels with more input samples directly above or left ofthe output sample position, compared to the number of input samples on the opposite border of the kernel.

[0070] In an example embodiment, a block of samples may be predicted using a filter of the following form:

[0071] Where p(x, y) is the predicted sample at location (x, y), q are the filter coefficients determined using reconstructed samples from a neighborhood of the block to be predicted, and Si(x,y) are the input samples to the filter when predicting p(x, y).

[0072] The input samples to the filter may be determined in different ways. For example, the input samples may be obtained from fixed locations with respect to p(x, y). There may also be different mode(s) of operation with different selection of the location(s) of the input samples. In such case, an encoder may decide the most suitable mode, for example by evaluating a cost function and selecting the mode that minimizes the cost function for the block of samples to be predicted. The encoder may then include a syntax element or syntax elements to the bitstream, and a decoder may identify the selected mode by decoding and interpreting those syntax elements.

[0073] One of the input samples, for example the last one SN-i(x,y), may be determined to have a constant value instead of a sample value from the prediction block or a sample value from outside of the prediction block. Such constant may be referred to as a bias term, as it, when multiplied by its corresponding filter coefficient, may represent a constant bias to the filter output. As the results of the multiplication of a constant input with a determined filter coefficient that is also constant within one block of samples is a constant, it may be pre-calculated and included as a separate term in the filtering equation:P (%, y ) = b + £ =-o2 si ( y)ciEq-3

[0074] It may be noted that Eq. 3 includes a bias term instead of a last input sample.

[0075] In the present disclosure, the terms “bias term” and “constant value” may be used interchangeably.

[0076] Input samples Si(x,y) typically represent samples in a still picture or a picture in a video sequence. Such samples typically have a certain bit depth that determines how many bits are needed to describe values of the samples. Typical selections for the bit depths used in video coding include 8, 10, 12 or 16 bits per samples. If the samples have a bit depth of B bits, a convenient selection for the bias input SN-i(x,y) may be half of the maximum of the sample value range 2B-1or (1 « B - 1) using “«” to represent left bitshift operation. This selection may allow processing of the bias input as a typical sample value with relatively similar value to the other input samples. Naturally also other selections may be made. For example, constant values, such as 128, 512, 1024 or 65536 may be selected. In this case, where 2B-1is used, the filtering equation may be further represented as:

[0077] It may be noted that, in Eq. 4, 2Bis multiplied by the filter coefficient for N-l.

[0078] As mentioned above, the spatial sample inputs Si(x, y) with i ranging from 0 to N-2 may be determined with preselected offsets with respect to the position of the predicted sample p(x, y). In the case of recursive filtering progressing diagonally through the prediction block, samples directly above, samples directly to the left, and samples in the above-left direction may already be predicted or available from already coded / decoded neighboring blocks. Thus, those may be used as input samples to the filter. Naturally, the sample at position x, y cannot be used, as it is the output sample of the filtering process. Thus, for example filter shapes such as the ones represented in FIG. 5 may be used. In the case of examples of FIG. 5 the input samples may be given as follows:

[0079] Where s(x, y) represents an input sample in location x, y and may either be an already predicted sample, if the position of s(x, y) is within the prediction block, or a reconstructed samplefrom a neighboring block. If the neighboring block is unavailable, such sample values may be substituted with other sample values, for example using the closest available sample values. The delta coordinates may depend on the prediction mode, and different selections may be made for different coding modes. For example, the following arrays may be used to determine the delta parameters (d(x, d^~) in the case of the three example modes of FIG. 5 (510, 520, 530) defining a rectangular filter kernel, where the first element in the array corresponds to (dj, d^), second to (d , d^), and so on:

[0080] For mode 0 (510):

[0081] [ (-1, 0), (-2, 0), (-3, 0),(0,-1), (-1,-1), (-2,-1), (-3,-1),(0,-2), (-1,-2), (-2,-2), (-3,-2),(0,-3), (-1,-3), (-2,-3), (-3,-3) ]

[0082] For mode 1 (520):

[0083] [ (-1, 0), ( 0,-1), (-2, 0),(-1,-1), (-3, 0), (-2,-1), (-4, 0),(-3,-1), (-5, 0), (-4,-1), (-6, 0),(-5,-1), (-7, 0), (-6,-1), (-7,-1) ]

[0084] For mode 2 (530):

[0085] [ (-1, 0), ( 0,-1), (-1,-1),(0,-2), (-1,-2), ( 0,-3), (-1,-3),(0,-4), (-1,-4), ( 0,-5), (-1,-5),(0,-6), (-1,-6), ( 0,-7), (-1,-7) ]

[0086] The above input sample locations may be coordinates relative to p at (0, 0).

[0087] In another example embodiment, a symmetric selection for the spatial input samples may be made, where the last spatial location may be omitted from the arrays (i.e. not included).In general, such selection may be made by including all samples within an WxH rectangle, except the bottom-right sample at coordinates (0, 0) with respect to the output sample position and topleft sample at coordinates (-[W-l], -[H-l]) with respect to the output sample position. This selection benefits from a lower complexity of both the filtering operation itself, as well as lower complexity of solving the filter coefficients, as there is one less coefficient to determine. At the same time, the support window may still contain samples within the same W*H rectangle. Such an example is depicted in FIG. 6, with mode 0 of 4x4 (610), mode 1 of 8x2 (620), and mode 2 of 2x8 (630) rectangular input kernels, and may be implemented, for example, using the following input arrays for the delta parameters (d(x, d^~)

[0088] For mode 0 (610):

[0089] [ (-1, 0), (-2, 0), (-3, 0),(0,-1), (-1,-1), (-2,-1), (-3,-1),(0,-2), (-1,-2), (-2,-2), (-3,-2),(0,-3), (-1,-3), (-2,-3) ]

[0090] Compared to 510 of FIG. 5, (-3, -3) is omitted.

[0091] For mode 1 (620):

[0092] [ (-1, 0), ( 0,-1), (-2, 0),(-1,-1), (-3, 0), (-2,-1), (-4, 0),(-3,-1), (-5, 0), (-4,-1), (-6, 0),(-5,-1), (-7, 0), (-6,-1) ]

[0093] Compared to 520 of FIG. 5, (-7, -1) is omitted.

[0094] For mode 2 (630):

[0095] [ (-1, 0), ( 0,-1), (-1,-1),(0,-2), (-1,-2), ( 0,-3), (-1,-3),(0,-4), (-1,-4), ( 0,-5), (-1,-5),(0,-6), (-1,-6), ( 0,-7) ]

[0096] Compared to 530 of FIG. 5, (-1, -7) is omitted.

[0097] In another example embodiment, some other spatial sample may be omitted from the example configurations of FIG. 5 (e.g. instead of the last sample location). Similarly to the example above, with such selection the bias term may be included in the filter without affecting negatively the computational aspect relating to solving and applying the filter, while achieving coding efficiency gains due to the stabilizing properties of the bias term.

[0098] In another example embodiment, input samples may be selected in a way that, for a vertically aligned input where there are more input samples directly above the output sample compared to input samples on the left of the output sample, a larger amount of input samples may be included above the output sample compared to the input samples included above the bottomleft input sample. This may be helpful in replicating patterns along one prediction direction, while still providing support for the other direction, but with a minimal number of filter coefficients and associated complexities. Such selection is depicted for mode 2 (730) in FIG. 7.

[0099] Similarly, transposed selection may be made for horizontally aligned input, as illustrated for mode 1 (720) in FIG. 7. For a symmetric input, such as the mode 0 (710) of FIG. 7, the same rule may be applied for both horizontal and vertical directions, resulting in a filter kernel omitting samples on the top and left borders of the filter, but keeping more input samples both directly above and directly left of the output samples, compared to at least some other rows and columns of the kernel.

[0100] Kernels illustrated in FIG. 7 may be implemented, for example, by using the following tables to determine the delta parameters (d(x, d^~)

[0101] For mode 0 (710):

[0102] [ (-1, 0), (-2, 0), (-3, 0),(0,-1), (-1,-1), (-2,-1),(0,-2), (-1,-2), (-2,-2),(0,-3) ]

[0103] Five sample locations are omitted, in comparison to 510 of FIG. 5.

[0104] For mode 1 (720):

[0105] [ (-1, 0), ( 0,-1), (-2, 0),(-1,-1), (-3, 0), (-2,-1), (-4, 0),(-5, 0), (-6, 0),(-7, 0) ]

[0106] Five sample locations are omitted, in comparison to 520 of FIG. 5.

[0107] For mode 2 (730):

[0108] [ (-1, 0), ( 0,-1), (-1,-1),(0,-2), (-1,-2), ( 0,-3),(0,-4), ( 0,-5),(0,-6), ( 0,-7) ]

[0109] Five sample locations are omitted, in comparison to 530 of FIG. 5.

[0110] To determine the filter parameters, a certain amount of training samples may need to be determined. For example, a fixed or block size dependent neighborhood of the block to be predicted may be chosen as the reference sample area used in the process. The samples in the reference sample area may then be used to build an autocorrelation matrix and a cross-correlation vector between the filter input and output. Once the autocorrelation matrix and cross-correlation vector are determined, solving of the filter parameters may be done in different ways. For example, approaches based on linear regression, LDL decomposition, other matrix decompositions, or Gaussian elimination may be used. Naturally also other approaches, either using or not using autocorrelation matrices and / or cross-correlation vectors, may be used.

[0111] A reference sample area for training a filter may be selected in different ways. For example, the number of lines of samples included above and left of the prediction block may be determined to be equal to the minimum of three values: width of the prediction block, height of the prediction block, and a pre-determined value specifying the maximum number of lines that may be included in the reference area. The value specifying the maximum number of lines that may be included in the reference area may be, for example, 4 or 8 or 12 or 16, but naturally any feasible number can be selected.

[0112] Spatial sample values, including predicted sample values for the current block to be predicted and reconstructed sample values of neighboring areas of the block, may be used as input to the filter. Alternatively, modified sample values may be used. Modifications may include, for example, deducting a constant from the input values or different filtering operations, including but not limited to, low-pass filtering, high-pass filtering, band-pass filtering and other linear or nonlinear operation applied to either single or multiple sample values.

[0113] Similarly, when determining the filter coefficients, reconstructed sample values in a determined picture area may be used as inputs to the filter training, or those may be modified prior to or during the determination of the filter parameters in similar or different ways, than what is used during the filtering process.

[0114] Some of the inputs to the filter may be exponential or polynomial functions of the input samples, or values derived from the input samples. For example, one of the inputs to the filter may be derived based on a square of an input sample value.

[0115] In addition to creating a filter using sample values in neighboring areas of a block, alternative approaches may be used to determine the filter parameters. For example, filter parameters determined for other prediction blocks of the same or different pictures may be selected for use.

[0116] In an example embodiment, a filter including at least one input value that is determined based on a predicted sample value or a reconstructed sample value, and at least one constant input value, may be used for recursive intra sample prediction.

[0117] In an example embodiment, a filter including at least one spatial input sample value and at least one constant input value, may be used for recursive intra sample prediction.

[0118] In an example embodiment, a filter including at least one constant input value and multiple spatial input sample values, where the spatial sample values may be determined based on a rectangular filter kernel which omits the bottom-right sample and the top-left sample of the kernel, may be used for intra sample prediction.

[0119] In an example embodiment, a filter including at least one constant input value and multiple spatial input sample values, where the spatial sample values may be determined based on a rectangular filter kernel which omits the bottom-right sample and the top-left sample of the kernel, may be trained to be used for intra sample prediction.

[0120] In an example embodiment, a filter including at least one constant input value and multiple spatial input sample values, where the spatial sample values may be determined based on a rectangular filter kernel which contains more input samples directly above or left of the output sample position, compared to the number of input samples on the opposite border of the kernel.

[0121] In an example embodiment, a codec may determine a block of samples to be predicted. In an example embodiment, a codec may determine a filter, for which input may consist of at least one constant value and at least one value that is determined based on a predicted sample value or a reconstructed sample value. In an example embodiment, a codec may determine filter coefficients for the filter using reconstructed sample values outside of the block to be predicted. In an example embodiment, a codec may apply the filter with the determined filter coefficients to predict samples of the block recursively, and may use samples outside the block, predicted samples inside the block, and the constant value as inputs.

[0122] In an example embodiment, the spatial input to the filter may form a rectangular kernel with bottom-right and top-left samples omitted.

[0123] In an example embodiment, the spatial input to the filter may form a rectangular kernel, which may contain more input samples directly above or left of the output sample position compared to a number of input samples on the opposite border of the kernel.

[0124] In an example embodiment, a codec may have a set of recursive intra prediction modes, each with a different filter kernel, where each of those kernels may include a bias term.

[0125] In an example embodiment, a codec may have a set of recursive intra prediction modes, each with a different filter kernel, where at least one of those filter kernels may include a bias term, and at least one of those filter kernels may not include a bias term.

[0126] FIG. 8 illustrates the potential steps of an example method 800. The example method 800 may include: determining a block of one or more samples to be predicted, 810; determining a filter for performing intra sample prediction, wherein inputs to the filter comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block, 820; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block, 830; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block, 840. The example method 800 may be performed, for example, with a codec, an encoder, a decoder, a module or device configured to perform encoding, a module or device configured to perform decoding, a UE, a network node, etc.

[0127] In accordance with one example embodiment, an apparatus may comprise: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine a block of one or more samples to be predicted; determine a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samplesinside the block; determine at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and apply the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0128] The at least one of the one or more reconstructed samples, or the one or more previously predicted samples, may form a rectangular kernel with bottom-right and top-left samples omitted.

[0129] The at least one of the one or more reconstructed samples, or the one or more previously predicted samples, may form a rectangular kernel comprising a greater number of input samples directly above, or left, of an output sample position compared to a number of input samples on an opposite border of the rectangular kernel.

[0130] The example apparatus may be further configured to: determine a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples; and provide an indication of the determined mode for inclusion in one or more syntax elements of a bitstream.

[0131] The mode may be selected based, at least partially, on evaluation of a cost function for predicting the one or more samples of the block.

[0132] The example apparatus may be further configured to: parse at least one syntax element of a bitstream; and determine a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on the at least one syntax element.

[0133] The example apparatus may be further configured to: determine the at least one constant value based, at least partially, on at least one of: the at least one filter coefficient, a maximum of a sample value range, or a bit depth of the one or more samples outside the block, or the one or more previously predicted samples inside the block.

[0134] The example apparatus may be further configured to: select the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on one or more preselected offsets with respect to an output sample position.

[0135] The example apparatus may be further configured to: determine a reference sample area relative to the determined block; determine an autocorrelation matrix based, at least partially, on the reference sample area; determine a cross-correlation vector based, at least partially, on the reference sample area; and determine the at least one filter coefficient based, at least partially, on the autocorrelation matrix and the cross-correlation vector.

[0136] The at least one filter coefficient may be determined using at least one of: linear regression, LDL decomposition, matrix decomposition, or Gaussian elimination.

[0137] The example apparatus may be further configured to: modify the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, using at least one of: deduction of a constant from respective values of the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, filtering, low-pass filtering, high-pass filtering, or band-pass filtering.

[0138] In accordance with one aspect, an example method may be provided comprising: determining, with a user equipment, a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0139] The at least one of the one or more reconstructed samples, or the one or more previously predicted samples, may form a rectangular kernel with bottom-right and top-left samples omitted.

[0140] The at least one of the one or more reconstructed samples, or the one or more previously predicted samples, may form a rectangular kernel comprising a greater number of input samples directly above, or left, of an output sample position compared to a number of input samples on an opposite border of the rectangular kernel.

[0141] The example method may further comprise: determining a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples; and providing an indication of the determined mode for inclusion in one or more syntax elements of a bitstream.

[0142] The mode may be selected based, at least partially, on evaluation of a cost function for predicting the one or more samples of the block.

[0143] The example method may further comprise: parsing at least one syntax element of a bitstream; and determining a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on the at least one syntax element.

[0144] The example method may further comprise: determining the at least one constant value based, at least partially, on at least one of: the at least one filter coefficient, a maximum of a sample value range, or a bit depth of the one or more samples outside the block, or the one or more previously predicted samples inside the block.

[0145] The example method may further comprise: selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on one or more preselected offsets with respect to an output sample position.

[0146] The example method may further comprise: determining a reference sample area relative to the determined block; determining an autocorrelation matrix based, at least partially, on the reference sample area; determining a cross-correlation vector based, at least partially, on thereference sample area; and determining the at least one filter coefficient based, at least partially, on the autocorrelation matrix and the cross-correlation vector.

[0147] The at least one filter coefficient may be determined using at least one of: linear regression, LDL decomposition, matrix decomposition, or Gaussian elimination.

[0148] The example method may further comprise: modifying the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, using at least one of: deduction of a constant from respective values of the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, filtering, low-pass filtering, high-pass filtering, or band-pass filtering.

[0149] In accordance with one example embodiment, an apparatus may comprise: circuitry configured to perform: determining, with a user equipment, a block of one or more samples to be predicted; circuitry configured to perform: determining a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; circuitry configured to perform: determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and circuitry configured to perform: applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0150] In accordance with one example embodiment, an apparatus may comprise: processing circuitry; memory circuitry including computer program code, the memory circuitry and the computer program code configured to, with the processing circuitry, enable the apparatus to: determine a block of one or more samples to be predicted; determine a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samplesoutside the block, or one or more previously predicted samples inside the block; determine at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and apply the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0151] As used in this application, the term “circuitry” or “means” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0152] In accordance with one example embodiment, an apparatus may comprise means for: determining a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least onereconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0153] The at least one of the one or more reconstructed samples, or the one or more previously predicted samples, may form a rectangular kernel with bottom-right and top-left samples omitted.

[0154] The at least one of the one or more reconstructed samples, or the one or more previously predicted samples, may form a rectangular kernel comprising a greater number of input samples directly above, or left, of an output sample position compared to a number of input samples on an opposite border of the rectangular kernel.

[0155] The means may be further configured for: determining a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples; and providing an indication of the determined mode for inclusion in one or more syntax elements of a bitstream.

[0156] The mode may be selected based, at least partially, on evaluation of a cost function for predicting the one or more samples of the block.

[0157] The means may be further configured for: parsing at least one syntax element of a bitstream; and determining a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on the at least one syntax element.

[0158] The means may be further configured for: determining the at least one constant value based, at least partially, on at least one of: the at least one filter coefficient, a maximum of a sample value range, or a bit depth of the one or more samples outside the block, or the one or more previously predicted samples inside the block.

[0159] The means may be further configured for: selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on one or more preselected offsets with respect to an output sample position.

[0160] The means may be further configured for: determining a reference sample area relative to the determined block; determining an autocorrelation matrix based, at least partially, on the reference sample area; determining a cross-correlation vector based, at least partially, on the reference sample area; and determining the at least one filter coefficient based, at least partially, on the autocorrelation matrix and the cross-correlation vector.

[0161] The at least one filter coefficient may be determined using at least one of: linear regression, LDL decomposition, matrix decomposition, or Gaussian elimination.

[0162] The means may be further configured for: modifying the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, using at least one of: deduction of a constant from respective values of the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, filtering, low-pass filtering, high-pass filtering, or band-pass filtering.

[0163] A processor, memory, and / or example algorithms (which may be encoded as instructions, program, or code) may be provided as example means for providing or causing performance of operation.

[0164] In accordance with one example embodiment, a non-transitory computer-readable medium comprising instructions stored thereon which, when executed with at least one processor, cause the at least one processor to: determine a block of one or more samples to be predicted; determine a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determine at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and apply the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may beapplied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0165] In accordance with one example embodiment, a non-transitory computer-readable medium comprising program instructions stored thereon for performing at least the following: determining a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0166] In accordance with another example embodiment, a non-transitory program storage device readable by a machine may be provided, tangibly embodying instructions executable by the machine for performing operations, the operations comprising: determining a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0167] In accordance with another example embodiment, a non-transitory computer-readable medium comprising instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: determining a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0168] A computer implemented system comprising: at least one processor and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: determining a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0169] A computer implemented system comprising: means for determining a block of one or more samples to be predicted; means for determining a filter for performing intra sample prediction, wherein inputs to the filter may comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; means for determining atleast one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and means for applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter may be applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

[0170] The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e. tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).

[0171] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications can be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modification and variances which fall within the scope of the appended claims.

Claims

CLAIMSWhat is claimed is:

1. An apparatus comprising: at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine a block of one or more samples to be predicted; determine a filter for performing intra sample prediction, wherein inputs to the filter comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determine at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and apply the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, orrespective values of the one or more previously predicted samples inside the block.

2. The apparatus of claim 1, wherein the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, form a rectangular kernel with bottom-right and top-left samples omitted.

3. The apparatus of claim 1, wherein the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, form a rectangular kernel comprising a greater number of input samples directly above, or left, of an output sample position compared to a number of input samples on an opposite border of the rectangular kernel.

4. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: determine a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples; and provide an indication of the determined mode for inclusion in one or more syntax elements of a bitstream.

5. The apparatus of claim 4, wherein the mode is selected based, at least partially, on evaluation of a cost function for predicting the one or more samples of the block.

6. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: parse at least one syntax element of a bitstream; and determine a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on the at least one syntax element.

7. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: determine the at least one constant value based, at least partially, on at least one of: the at least one filter coefficient, a maximum of a sample value range, or a bit depth of the one or more samples outside the block, or the one or more previously predicted samples inside the block.

8. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: select the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on one or more preselected offsets with respect to an output sample position.

9. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: determine a reference sample area relative to the determined block; determine an autocorrelation matrix based, at least partially, on the reference sample area; determine a cross-correlation vector based, at least partially, on the reference sample area; and determine the at least one filter coefficient based, at least partially, on the autocorrelation matrix and the cross-correlation vector.

10. The apparatus of claim 9, wherein the at least one filter coefficient is determined using at least one of:linear regression,LDL decomposition, matrix decomposition, orGaussian elimination.

11. The apparatus of claim 1 , wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: modify the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, using at least one of: deduction of a constant from respective values of the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, filtering, low-pass filtering, high-pass filtering, or band-pass filtering.

12. A method comprising: determining, with a user equipment, a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter comprise at least: at least one constant value, andat least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

13. The method of claim 12, wherein the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, form a rectangular kernel with bottom-right and top-left samples omitted.

14. The method of claim 12, wherein the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, form a rectangular kernel comprising a greater number of input samples directly above, or left, of an output sample position compared to a number of input samples on an opposite border of the rectangular kernel.

15. The method of claim 12, further comprising: determining a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples; andproviding an indication of the determined mode for inclusion in one or more syntax elements of a bitstream.

16. The method of claim 15, wherein the mode is selected based, at least partially, on evaluation of a cost function for predicting the one or more samples of the block.

17. The method of claim 12, further comprising: parsing at least one syntax element of a bitstream; and determining a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on the at least one syntax element.

18. The method of claim 12, further comprising: determining the at least one constant value based, at least partially, on at least one of: the at least one filter coefficient, a maximum of a sample value range, or a bit depth of the one or more samples outside the block, or the one or more previously predicted samples inside the block.

19. The method of claim 12, further comprising: selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on one or more preselected offsets with respect to an output sample position.

20. The method of claim 12, further comprising: determining a reference sample area relative to the determined block;determining an autocorrelation matrix based, at least partially, on the reference sample area; determining a cross-correlation vector based, at least partially, on the reference sample area; and determining the at least one filter coefficient based, at least partially, on the autocorrelation matrix and the cross-correlation vector.

21. The method of claim 20, wherein the at least one filter coefficient is determined using at least one of: linear regression,LDL decomposition, matrix decomposition, orGaussian elimination.

22. The method of claim 12, further comprising: modifying the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, using at least one of: deduction of a constant from respective values of the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, filtering, low-pass filtering, high-pass filtering, or band-pass filtering.

23. An apparatus comprising means for: determining a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value, respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

24. The apparatus of claim 23, wherein the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, form a rectangular kernel with bottom-right and top-left samples omitted.

25. The apparatus of claim 23, wherein the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, form a rectangular kernel comprising a greater number of input samples directly above, or left, of an output sampleposition compared to a number of input samples on an opposite border of the rectangular kernel.

26. The apparatus of claim 23, wherein the means are further configured for: determining a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples; and providing an indication of the determined mode for inclusion in one or more syntax elements of a bitstream.

27. The apparatus of claim 26, wherein the mode is selected based, at least partially, on evaluation of a cost function for predicting the one or more samples of the block.

28. The apparatus of claim 23, wherein the means are further configured for: parsing at least one syntax element of a bitstream; and determining a mode for selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on the at least one syntax element.

29. The apparatus of claim 23, wherein the means are further configured for: determining the at least one constant value based, at least partially, on at least one of: the at least one filter coefficient, a maximum of a sample value range, or a bit depth of the one or more samples outside the block, or the one or more previously predicted samples inside the block.

30. The apparatus of claim 23, wherein the means are further configured for: selecting the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, based on one or more preselected offsets with respect to an output sample position.

31. The apparatus of claim 23, wherein the means are further configured for: determining a reference sample area relative to the determined block; determining an autocorrelation matrix based, at least partially, on the reference sample area; determining a cross-correlation vector based, at least partially, on the reference sample area; and determining the at least one filter coefficient based, at least partially, on the autocorrelation matrix and the cross-correlation vector.

32. The apparatus of claim 31, wherein the at least one filter coefficient is determined using at least one of: linear regression,LDL decomposition, matrix decomposition, orGaussian elimination.

33. The apparatus of claim 23, wherein the means are further configured for: modifying the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, using at least one of:deduction of a constant from respective values of the at least one of the one or more reconstructed samples, or the one or more previously predicted samples, filtering, low-pass filtering, high-pass filtering, or band-pass filtering.

34. A non-transitory computer-readable medium comprising program instructions stored thereon for performing at least the following: determining a block of one or more samples to be predicted; determining a filter for performing intra sample prediction, wherein inputs to the filter comprise at least: at least one constant value, and at least one value determined based on at least one of: one or more reconstructed samples outside the block, or one or more previously predicted samples inside the block; determining at least one filter coefficient for the filter based, at least partially, on at least one reconstructed sample value outside of the block; and applying the filter, using the at least one filter coefficient, to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of: the at least one constant value,respective values of the one or more reconstructed samples outside the block, or respective values of the one or more previously predicted samples inside the block.

Citation Information

Patent Citations

  • Video intra prediction using hybrid recursive filters

    US20160373785A1

  • Prediction systems and methods for video coding based on filtering nearest neighboring pixels

    US20190182482A1

  • Offset-based refinement of intra prediction (ORIP) of video coding

    US20220141459A1

  • A method, an apparatus and a computer program product for encoding and decoding of digital media content

    WO2023194647A1