Intra prediction using bias-based extrapolation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2026-08-11
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The exemplary and non-limiting embodiments relate generally to intra-sample prediction, and more specifically, to selecting the input to the filter used to perform intra-sample prediction. Background Technology
[0002] As is well known, extrapolation-based filters are used for performing intra-frame prediction. Summary of the Invention
[0003] The following description of the invention is for illustrative purposes only. This invention is not intended to limit the scope of the claims.
[0004] According to one aspect, an apparatus includes: at least one processor; and at least one memory storing instructions, which, when executed by the at least one processor, cause the apparatus to at least: determine a block of one or more samples to be predicted; determine a filter for performing intra-sample prediction, wherein input to the filter includes at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determine at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and apply the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter is recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0005] According to one aspect, a method includes: determining, using a user equipment, a block of one or more samples to be predicted; determining a filter for performing intra-sample prediction, wherein the input to the filter includes at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determining at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter is recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0006] According to one aspect, an apparatus includes components for: determining a block of one or more samples to be predicted; determining a filter for performing intra-sample prediction, wherein the input to the filter includes at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determining at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter is recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0007] According to one aspect, a non-transitory computer-readable medium includes program instructions stored thereon for performing at least the following: determining a block of one or more samples to be predicted; determining a filter for performing intra-frame sample prediction, wherein the input to the filter includes at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determining at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter is recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0008] The independent claims are provided with respect to several aspects. Several other aspects are defined in the dependent claims. Attached Figure Description
[0009] The above aspects and other features are explained in the following description in conjunction with the accompanying drawings, in which:
[0010] Figure 1 This is a block diagram of a possible non-limiting example system in which example embodiments can be practiced;
[0011] Figure 2 This is a block diagram of a possible non-limiting example system in which example embodiments can be practiced;
[0012] Figure 3 It is a diagram illustrating the features described in this article;
[0013] Figure 4 It is a diagram illustrating the features described in this article;
[0014] Figure 5It is a diagram illustrating the features described in this article;
[0015] Figure 6 It is a diagram illustrating the features described in this article;
[0016] Figure 7 It is a diagram illustrating the features described in this article; and
[0017] Figure 8 This is a flowchart illustrating the steps described in this article. Detailed Implementation
[0018] The following abbreviations, which can be found in the instruction manual and / or accompanying drawings, are defined as follows: 3GPP: Third Generation Partnership Project 4G: Fourth Generation 5G: Fifth Generation 5GC: 5G Core Network AR: Augmented Reality AVC: Advanced Video Coding (ITU-T H.264 video coding standard) BCW: Dual Prediction Using CU-Level Weights CABAC: Context-Adaptive Binary Arithmetic Encoder CCCM: Convolutional Cross-Component Model CCRM: Cross-component residual model / Cross-component reconstruction model CDMA: Code Division Multiple Access CIIP: Inter- and Intra-Frame Joint Prediction CPU: Central Processing Unit cRAN: Cloud Radio Access Network CTU: Coding Tree Unit CU: Encoding Unit DCT: Discrete Cosine Transform DPB: Decode Image Cache DST: Discrete Sine Transform ECM: Enhanced Compression Model (JVET's Exploratory Video Codec) eNB (or eNodeB): Evolved Node B (e.g., LTE base station) EN-DC: E-UTRA-NR Dual Connection en-gNB or EN-gNB: A node that provides NR user plane and control plane protocol termination to the UE and acts as a secondary node in the EN-DC. E-UTRA: Evolved Universal Terrestrial Radio Access, also known as LTE radio access technology. FDMA: Frequency Division Multiple Access gNB (or gNodeB): A base station used for 5G / NR, specifically a node that provides NR user plane and control plane protocol termination to the UE and connects to the 5GC via the NG interface. GPU: Graphics Processing Unit GSM: Global System for Mobile Communications HEVC: High Efficiency Video Coding (ITU-T H.265 video coding standard) HMD: Head-mounted display IBC: Intra-Block Copy IEEE: Institute of Electrical and Electronics Engineers IMD: Integrated Messaging Device IMS: Instant Messaging Service IoT: Internet of Things LCU: Maximum Coding Unit LIC: Local Illumination Compensation LTE: Long Term Evolution MMS: Multimedia Messaging Service MPEG-I: Moving Picture Experts Group Immersive Codec Series MR: Mixed Reality MSE: Mean Squared Error ng or NG: the new generation ng-eNB or NG-eNB: Next-generation eNB NR: New Radio N / W or NW: Network O-RAN: Open Radio Access Network PC: Personal Computer PDA: Personal Digital Assistant PU: Prediction Unit RGB: Red, Green, Blue SMS: Short Message Service SNR: Signal-to-noise ratio TCP / IP: Transmission Control Protocol, Internet Protocol TDMA: Time Division Multiple Access TU: Transformation Unit UE: User Equipment (e.g., wireless equipment, typically mobile equipment) UMTS: Universal Mobile Telecommunications System USB: Universal Serial Bus VNR: Virtualized Network Function VR: Virtual Reality VVC: Universal Video Coding (ITU-T H.266 video coding standard) WLAN: Wireless Local Area Network WP: Weighted Prediction YUV / YCbCr: A color model based on one luminance channel and two chrominance / color difference channels (commonly used in many video encoding applications).
[0019] The following describes suitable apparatus and possible mechanisms for practicing exemplary embodiments of this disclosure. Therefore, reference is made first to... Figure 1 , Figure 1 An example block diagram of device 50 is shown. This device can be configured to perform various functions, such as, for example, collecting information via one or more sensors, encoding and / or decoding information, receiving and / or transmitting information, analyzing information collected or received by the device, and so on. A device configured to encode a video scene may (optionally) include: one or more microphones for capturing the scene, and / or one or more sensors, such as a camera, for capturing information about the physical environment in which the scene is captured. Alternatively, a device configured to encode a video scene may be configured to receive information about the environment and / or simulated environment in which the scene is captured. A device configured to decode and / or render a video scene may be configured to receive a Moving Picture Experts Group Immersive Codec Series (MPEG-I) bitstream including the encoded video scene. A device configured to decode and / or render a video scene may include one or more speaker / audio transducers and / or displays, and / or may be configured to send the decoded scene or signal to a device including one or more speaker / audio transducers and / or displays. Devices configured to decode and / or render video scenes may include: user devices, head-mounted displays, or other devices capable of rendering AR, VR, and / or MR experiences to users.
[0020] Electronic device 50 may be, for example, a mobile terminal or user equipment of a wireless communication system. Alternatively, electronic device may be a non-mobile computer or part of a computer. It should be understood that the exemplary embodiments of this disclosure can be implemented within any electronic device or apparatus capable of processing data. Electronic device 50 may include devices that can access a network and / or cloud via a wired or wireless connection. Electronic device 50 may include one or more processors 56, one or more memories 58, and one or more transceivers 52 interconnected via one or more buses. The one or more processors 56 may include a central processing unit (CPU) and / or a graphics processing unit (GPU). Each of the one or more transceivers 52 includes a receiver and a transmitter. The one or more buses may be address, data, or control buses and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optic cable, or other optical communication device. “Circuit” may include dedicated hardware or hardware associated with software that can be executed thereon. The one or more transceivers may be connected to one or more antennas 44. The one or more memories 58 may include computer program code. The one or more memories 58 and the computer program code may be configured, together with the one or more processors 56, to cause electronic device 50 to perform one or more of the operations described herein.
[0021] Electronic device 50 can be connected to a node in a network. A network node may include one or more processors, one or more memories, and one or more transceivers interconnected via one or more buses. Each of the one or more transceivers includes a receiver and a transmitter. The one or more buses may be address, data, or control buses and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optic cables, or other optical communication devices. The one or more transceivers may be connected to one or more antennas. The one or more memories may include computer program code. The one or more memories and computer program code may be configured, together with the one or more processors, to cause the network node to perform one or more of the operations described herein.
[0022] Electronic device 50 may include microphone 36 or any suitable audio input, which may be a digital or analog signal input. Electronic device 50 may also include audio output device 38, which in the example embodiments of this disclosure may be headphones, a speaker, or any analog or digital audio output connection. Electronic device 50 may also include a battery (or, in other example embodiments of this disclosure, the device may be powered by any suitable mobile energy device, such as a solar cell, fuel cell, or clockwork generator). Electronic device 50 may also include camera 42 or other sensors capable of recording or capturing images and / or video. Additionally or alternatively, electronic device 50 may also include a depth sensor. Electronic device 50 may also include display 32. Electronic device 50 may also include an infrared port for short-range line-of-sight communication with other devices. In other example embodiments of this disclosure, device 50 may also include any suitable short-range communication solution, such as, for example, BLUETOOTH™ wireless connectivity or USB / FireWire wired connectivity.
[0023] It should be understood that the electronic device 50 configured to perform the example embodiments of this disclosure may have fewer and / or more components, which may correspond to the process that the electronic device 50 is configured to perform. For example, means configured to encode video may not include a speaker or audio transducer, but may include a microphone, while means configured to render decoded video may not include a microphone, but may include a speaker or audio transducer.
[0024] Now for reference Figure 1 The electronic device 50 may include a controller 56, a processor, or a processor circuitry for controlling the device 50. The controller 56 may be connected to a memory 58, which, in exemplary embodiments of this disclosure, may store data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may also be connected to a codec circuitry 54 adapted to perform encoding and / or decoding of audio and / or video data, or to assist in encoding and / or decoding performed by the controller.
[0025] Electronic device 50 may also include a card reader 48 and a smart card 46, such as a UICC and a UICC reader, which are used to provide user information and are suitable for providing authentication information for authenticating and authorizing the user / electronic device 50 at the network. Electronic device 50 may also include an input device 34 for providing information to controller 56, such as a keypad, one or more input buttons, or a touchscreen input device.
[0026] Electronic device 50 may include a radio interface circuitry system 52 connected to a controller and adapted to generate wireless communication signals, for example, for communication with a cellular communication network, a wireless communication system, or a wireless local area network. Device 50 may also include an antenna 44 connected to the radio interface circuitry system 52 for transmitting radio frequency signals generated at the radio interface circuitry system 52 to / from other devices and / or for receiving radio frequency signals from / from other devices.
[0027] Electronic device 50 may include microphone 38, camera 42, and / or other sensors capable of recording or detecting audio signals, image / video signals, and / or other information about the local / virtual environment, which is then passed to codec 54 or controller 56 for processing. Electronic device 50 may receive audio / image / video signals and / or information about the local / virtual environment from another device for processing before transmission and / or storage. Electronic device 50 may also receive audio / image / video signals and / or information about the local / virtual environment wirelessly or via a wired connection for encoding / decoding. The structural elements of the above-described electronic device 50 represent examples of components used to perform corresponding functions.
[0028] Memory 58 may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. Memory 58 may be non-transitory memory. Memory 58 may be a component for performing storage functions. Controller 56 may be or include one or more processors, which may be of any type suitable for the local technical environment and, as a non-limiting example, may include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and a processor based on a multi-core processor architecture. Controller 56 may be a component for performing functions.
[0029] According to exemplary embodiments of this disclosure, electronic device 50 can be configured to perform volumetric scene capture. For example, electronic device 50 may include camera 42 or other sensors capable of recording or capturing images and / or video. Electronic device 50 may also include one or more transceivers 52 to enable the transmission of captured content for processing at another device. Such electronic device 50 may or may not include these features. Figure 1 All modules shown.
[0030] According to exemplary embodiments of this disclosure, electronic device 50 may be configured to perform volumetric video content processing. For example, electronic device 50 may include a controller 56 for processing images to generate volumetric video content, a controller 56 for processing volumetric video content to project 3D information into 2D information, patches, and auxiliary information, and / or a codec 54 for encoding 2D information, patches, and auxiliary information into a bitstream for transmission to another device having a radio interface 52. Such electronic device 50 may or may not include these features. Figure 1 All modules shown.
[0031] According to exemplary embodiments of this disclosure, electronic device 50 may be configured to perform encoding or decoding of 2D information representing volumetric video content. For example, electronic device 50 may include a codec 54 for encoding or decoding 2D information representing volumetric video content. Such electronic device 50 may or may not include such codecs. Figure 1 All modules shown.
[0032] According to an example embodiment of this disclosure, electronic device 50 may be configured to perform the rendering of decoded 3D volumetric video. For example, electronic device 50 may include a controller for projecting 2D information to reconstruct the 3D volumetric video, and / or a display 32 for rendering the decoded 3D volumetric video. Such electronic device 50 may or may not include these features. Figure 1 All modules shown.
[0033] about Figure 2This illustrates an example of a system in which exemplary embodiments of the present disclosure can be utilized. System 10 includes multiple communication devices that can communicate over one or more networks. System 10 may include any combination of wired or wireless networks, including but not limited to wireless cellular telephone networks (such as GSM, UMTS, E-UTRA, LTE, CDMA, 4G, 5G networks, etc.), wireless local area networks (WLANs) such as those defined by any standard in the IEEE 802.x standard, BLUETOOTH™ personal area networks, Ethernet LANs, token ring LANs, wide area networks, and / or the Internet. Wireless networks can implement network virtualization, which is the process of combining hardware and software network resources and network functions into a single software-based management entity (virtual network). Network virtualization involves platform virtualization, which is often combined with resource virtualization. Network virtualization is divided into two categories: one is external, which combines many networks or network parts into virtual units; the other is internal, which provides network-like functionality to software containers on a single system. For example, a network may be deployed in a remote cloud, where virtualized network functions (VNFs) run on, for example, data center servers. For example, core network functions and / or (multiple) radio access networks (e.g., CloudRAN, O-RAN, edge cloud) can be virtualized. Note that the virtualized entities resulting from network virtualization are still implemented to some extent using hardware such as processors and memory, and such virtualized entities also produce technical effects.
[0034] It should also be noted that the operation of the example embodiments of this disclosure can be performed by multiple cooperating devices (e.g., cRAN).
[0035] System 10 may include both wired and wireless communication devices and / or electronic devices suitable for implementing example embodiments of this disclosure.
[0036] For example, Figure 2 The system shown illustrates a representation of mobile phone network 11 and Internet 28. Connection to Internet 28 can include, but is not limited to, long-range wireless connections, short-range wireless connections, and various wired connections, including but not limited to telephone lines, power lines, and similar communication paths.
[0037] The example communication devices shown in System 10 may include, but are not limited to, device 15, a combination of a personal digital assistant (PDA) and mobile phone 14, PDA 16, integrated messaging device (IMD) 18, desktop computer 20, laptop computer 22, and head-mounted display (HMD) 17. Electronic device 50 may include any of these example communication devices. In the example embodiments of this disclosure, more than one of these devices, or multiple devices of one or more of these devices, may perform the disclosed processes(s). These devices may be connected to the Internet 28 via wireless connection 2.
[0038] The exemplary embodiments of this disclosure can also be implemented in set-top boxes; i.e., digital television receivers, which may / may not have a display or wireless capabilities; in tablet computers or (laptop) personal computers (PCs) having hardware and / or software for processing neural network data; in various operating systems; and in chipsets, processors, DSPs, and / or embedded systems that provide hardware / software-based coding. The exemplary embodiments of this disclosure can also be implemented in: cellular phones (such as smartphones), tablet computers, personal digital assistants (PDAs) with wireless communication capabilities, portable computers with wireless communication capabilities, image capture devices (such as digital cameras) with wireless communication capabilities, gaming devices with wireless communication capabilities, music storage and playback devices with wireless communication capabilities, internet devices allowing wireless internet access and browsing, tablet computers with wireless communication capabilities, and portable units or terminals combining such functions.
[0039] Some or additional devices can send and receive calls and messages, and communicate with service providers via a wireless connection 25 to base station 24, which may be, for example, an eNB, gNB, access point, access node, or other node. Base station 24 may connect to a network server 26 that allows communication between mobile phone network 11 and the Internet 28. The system may include additional communication equipment and various types of communication devices.
[0040] Communication devices can communicate using various transmission technologies, including but not limited to Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Transmission Control Protocol / Internet Protocol (TCP-IP), Short Message Service (SMS), Multimedia Messaging Service (MMS), Email, Instant Messaging Service (IMS), BlueTouch™, IEEE 802.11, 3GPP Narrowband IoT, and any similar wireless communication technologies. Communication devices implementing various example embodiments of this disclosure can communicate using various media, including but not limited to radio, infrared, laser, cable connections, and any suitable connection.
[0041] In telecommunications and data networks, a channel can refer to either a physical channel or a logical channel. A physical channel can refer to a physical transmission medium such as a wired connection, while a logical channel can refer to a logical connection via a multiplexed medium capable of transmitting several logical channels. A channel can be used to transmit information signals (e.g., bitstreams, which may be MPEG-I bitstreams) from one or more transmitters to one or more receivers.
[0042] Therefore, having introduced a suitable but non-limiting technical context for practicing exemplary embodiments of this disclosure, the exemplary embodiments will now be described in more detail.
[0043] The features described in this article are generally related to the encoding and decoding of digital video materials.
[0044] A video codec consists of an encoder and a decoder. The encoder transforms the input video into a compressed representation suitable for storage / transmission, and the decoder decompresses the compressed video representation back into a visual form. Typically, the encoder discards some information from the original video sequence in order to represent the video in a more compact form (i.e., at a lower bit rate).
[0045] Typical hybrid video codecs (such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC) encode video information in two stages. First, the pixel values in a certain picture region (or “block”) (310) can be predicted (320) by means of, for example, motion compensation (finding and indicating a region in a previously encoded picture that closely corresponds to the block being encoded) or spatial means (using the pixel values around the block to be encoded in a specified manner). Second, the prediction error (i.e., the difference between the predicted pixel block and the original pixel block) can be encoded (330). This is typically achieved by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant thereof) (340), quantizing the resulting transform coefficients (350), and entropy encoding the quantization coefficients (360). By varying the fidelity of the quantization process, the encoder can control the balance between the precision of the pixel representation (image quality) and the size of the final encoded video representation (file size or transmission bit rate). An example of the encoding process is shown in Figure 3 The diagram is shown in the image.
[0046] In some video codecs (such as H.265 / HEVC and H.266 / VVC), video frames can be divided into coding units (CUs) covering the frame area. A CU may include one or more prediction units (PUs) defining the prediction process for samples within the CU, and one or more transform units (TUs) defining the prediction error coding process for samples within the CU. Typically, a CU may include rectangular sample blocks, the size of which can be selected from a predefined set of possible CU sizes. CUs with the largest allowed size are typically named LCUs (Maximum Coding Units) or CTUs (Coding Tree Units), and video frames can be divided into non-overlapping CTUs. CTUs can be further subdivided into combinations of smaller CUs, for example, by recursively splitting CTUs and resulting CUs. Each resulting CU may typically have at least one PU and at least one TU associated with it. Each PU and TU may be further subdivided into smaller PUs and TUs to correspondingly increase the granularity of the prediction and prediction error coding processes. Each PU may have associated prediction information defining which predictions will be applied to pixels within the PU (e.g., motion vector information for inter-frame prediction PUs, and intra-frame prediction directionality information for intra-frame prediction PUs). Similarly, each TU can be associated with information describing the prediction error decoding process of samples within the TU (including, for example, DCT coefficient information). Typically, signaling at the CU level indicates whether prediction error coding can be applied to each CU. Without prediction error residuals associated with a CU, the aforementioned CU can be considered to have no TU. Signaling in the bitstream to divide the image into CUs and CUs into PUs and TUs can typically be done, allowing the decoder to reproduce the intended structure of these units.
[0047] The decoder can reconstruct the output video by applying a prediction component (410) similar to that of the encoder to form a predictive representation of pixel blocks (using motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (the inverse operation of prediction error encoding, which recovers the quantized prediction error signal in the spatial pixel domain) (420). After applying the prediction and prediction error decoding components, the decoder can add the prediction and prediction error signals (pixel values) to form an output video frame (440). The decoder (and encoder) can also apply an additional filtering component (430) to improve the quality of the output video before passing it for display and / or storing it as a prediction reference for upcoming frames in a video sequence. An example of the decoding process is as follows: Figure 4 As shown.
[0048] Alternatives to or in addition to methods that use sample value prediction and transform coding to indicate coded sample values, palette-based coding can be used. Palette-based coding refers to a set of methods for defining a palette (i.e., a set of colors and associated indices), and the value of each sample within a coding unit is represented by indicating its index in the palette. Palette-based coding can generally achieve good coding efficiency in coding units with a relatively small number of colors (such as image regions representing computer screen content, like text or simple graphics). To improve the coding efficiency of palette coding, different types of palette index prediction methods can be used, or run-length coding of the palette indices can be performed to efficiently represent larger uniform image regions. Furthermore, escape coding can be used when a CU contains non-repeating sample values within the CU. Escape-coded samples can be sent without referencing any palette index. Instead, their values can be indicated individually for each escape-coded sample.
[0049] In a typical video codec, motion information can be indicated using motion vectors associated with each motion-compensated image block. Each of these motion vectors can represent the displacement of an image block in the picture to be encoded (on the encoder side) or decoded (on the decoder side), and the displacement of a predicted source block in one of the previously encoded or decoded images. To efficiently represent motion vectors, these vectors are typically differentially encoded relative to block-specific predicted motion vectors. In a typical video codec, predicted motion vectors can be created in a predefined manner, such as by calculating the median of the encoded or decoded motion vectors of neighboring blocks. Another approach to creating motion vector predictions is to generate a list of candidate predictions based on neighboring and / or juxtaposed blocks in a time-referenced image, and to signal the selected candidates as motion vector predictors. In addition to predicting motion vector values, reference indices of previously encoded / decoded images can also be predicted. Reference indices can typically be predicted based on neighboring and / or juxtaposed blocks in a time-referenced image. Furthermore, a typical high-efficiency video codec can employ an additional motion information encoding / decoding mechanism, often referred to as a merging or fusion mode, where all motion field information (including motion vectors and corresponding reference image indices for each available list of reference images) can be predicted and used without any modification / correction. Similarly, the prediction of motion field information can be performed using motion field information from adjacent and / or juxtaposed blocks in a temporal reference image, and the motion field information used can be signaled in a motion field candidate list populated with motion field information from available adjacent / juxtaposed blocks.
[0050] Typically, video codecs can support motion-compensated prediction from at least one source image (unidirectional prediction) and from two sources (bidirectional prediction). In the case of unidirectional prediction, a single motion vector can be applied, while in the case of bidirectional prediction, two motion vectors can be determined, and the motion-compensated predictions from the two sources can be combined to create the final sample prediction. In the case of weighted prediction, the relative weights of the two predictions can be adjusted, or a signaling offset can be added to the prediction signal.
[0051] Besides applying motion compensation for inter-image prediction, similar methods can be applied to intra-image prediction. In this case, the displacement vector can indicate the predicted location where sample blocks can be copied from the same image to form the block to be encoded or decoded. This intra-block copying (IBC) method can significantly improve coding efficiency when there are repetitive structures (such as text or other graphics) within the frame.
[0052] In typical video codecs, the prediction residuals after motion compensation or intra-frame prediction can first be transformed using a transform kernel (such as DCT) before encoding. This is done because there is often some correlation between the residuals, and in many cases, the transform can help reduce this correlation and provide more efficient coding.
[0053] A typical video encoder might use a Lagrange cost function to find the optimal coding pattern, such as the desired macroblock pattern and associated motion vectors. This cost function uses a weighting factor λ to correlate the (exact or estimated) image distortion caused by lossy coding methods with the (exact or estimated) amount of information needed to represent pixel values in an image region. (Equation 1)
[0054] Where C is the Lagrange cost to be minimized, D is the image distortion (e.g., mean squared error) considering pattern and motion vectors, and R is the number of bits required to represent the data needed to reconstruct the image patch in the decoder (including the amount of data used to represent candidate motion vectors).
[0055] Scalable video coding refers to a coding structure in which a bitstream can contain multiple representations of content at different bit rates, resolutions, or frame rates. In these cases, a receiver can extract the desired representation based on its characteristics (e.g., the resolution best matched to the display device). Alternatively, a server or network element can extract portions of the bitstream to be sent to the receiver based on, for example, the receiver's network characteristics or processing capabilities. A scalable bitstream typically includes a "base layer" providing the lowest quality video and one or more "enhancement layers" that can enhance video quality when received and decoded together with the lower layer. To improve the coding efficiency of an enhancement layer, its coded representation can often depend on the lower layer. For example, motion and pattern information of the enhancement layer can be predicted based on the lower layer. Similarly, pixel data from the lower layer can be used to create predictions for the enhancement layer.
[0056] Scalable video codecs for quality scalability (also known as signal-to-noise ratio or SNR) and / or spatial scalability can be implemented as follows. For the base layer, a conventional non-scalable video encoder and decoder can be used. The reconstructed / decoded image of the base layer can be included in the reference image buffer for the enhancement layer. In H.264 / AVC, H.265 / HEVC, and similar codecs that use multiple reference image lists for inter-frame prediction, the image decoded from the base layer can be inserted into the multiple reference image lists for encoding / decoding enhancement layer images, similar to the reference images used for decoding enhancement layers. Therefore, the encoder can select a base layer reference image as an inter-frame prediction reference, typically indicated by a reference image index in the encoded bitstream. Based on the bitstream, for example, based on the reference image index, the decoder decodes the base layer image, which can be used as an inter-frame prediction reference for the enhancement layer. When the decoded base layer image is used as a prediction reference for the enhancement layer, it can be referred to as an inter-layer reference image.
[0057] Besides quality scalability, examples of other scalability patterns include:
[0058] - Spatial scalability: Enhancement layer images are encoded at a higher resolution than base layer images.
[0059] - Bit depth scalability: Enhancement layer images are encoded with a higher bit depth (e.g., 10 or 12 bits) than the base layer images (e.g., 8 bits).
[0060] - Chroma format scalability: Enhancement layer images offer higher fidelity in chroma than base layer images (e.g., encoded in 4:4:4 chroma format).
[0061] In all the aforementioned scalability scenarios, base layer information can be used to encode enhancement layers to minimize additional bit rate overhead.
[0062] Scalability can be enabled in two basic ways: by introducing a new coding scheme for performing pixel value prediction, or by introducing a syntax from a lower layer of the scalable representation; or by placing lower-layer images into a reference image buffer (Decoding Image Buffer (DPB)) at a higher layer. The first approach is more flexible and therefore offers better coding efficiency in most cases. However, the second approach, the reference frame-based scalability approach, can be implemented very efficiently with minimal changes to the single-layer codec while still achieving most of the available coding efficiency gains. Essentially, a reference frame-based scalable codec can be implemented using the same hardware or software implementation for all layers and handles DPB management only externally.
[0063] To enable parallel processing, images can be divided into independent coded and decodeable image segments (e.g., slices or tiles). Slices typically refer to image segments consisting of a certain number of basic coding units that are processed in a default coding or decoding order, while tiles typically refer to image segments defined as rectangular image regions that are processed at least to some extent as individual frames.
[0064] Typically, video can be encoded in the YUV or YCbCr color space, which reflects some characteristics of the human visual system and allows for lower-quality representations for the Cb and Cr channels because human perception is not very sensitive to the chromaticity fidelity of these channel representations.
[0065] Typical video codecs (such as H.265 / HEVC (ITU-T Recommendation H.265: "Highefficiency video coding"), https: / / www.itu.int / rec / T-REC-H.265 and H.266 / VVC (ITU-T Recommendation H.266: "Versatile video coding"), http: / / www.itu.int / rec / T-REC-H.266) standards) break down an image into sample blocks, which are predicted differently based on reconstructed samples from neighboring blocks. This prediction process typically extrapolates samples from neighboring blocks to fill the block to be predicted with values generated by a defined filtering process. This process is usually iterative, depending on the sample values of these neighboring blocks, to allow any sample(s) in the predicted block to be computed independently of the other samples.
[0066] Alternatively, recursive filters can be used. In the case of recursive filtering, the output samples of earlier steps in the filtering process can be used as input for predicting one or more new sample values. For example, such a process is proposed in JVET document JVET-AF0080: “EE2-2.7: An extrapolation filter-based intra prediction mode” (October 2023), where new predicted sample values are generated using N×M neighboring sample values. In this case, the neighboring sample values used as input can include both samples from neighboring sample blocks and predicted sample values from the sample block being predicted. The filter can typically be constructed or trained using reconstructed samples from neighboring blocks and then applied to predict the sample of the current block.
[0067] Extrapolation-based filters are sensitive to the difference between the range of values used when training the filter and the range of values the filter is expected to predict. This becomes a problem, specifically, when the prediction block size is relatively large, because filter coefficients trained using reconstructed samples from neighboring blocks may no longer be able to accurately represent the contents of the prediction block.
[0068] Furthermore, modern codecs often employ so-called intra-prediction "merging modes." These modes copy the parameters of earlier predicted blocks and apply them to the current predicted block. When using conventional extrapolation-based filters constructed for another predicted block and trained using the neighbors of that other predicted block, the discrepancy between the parameters and the actual predicted block becomes even greater.
[0069] In an example embodiment, a recursive filter can be used, which combines a spatial kernel with a specially trained bias term to perform intra-sample prediction on a block of samples. An alternative configuration can use symmetric filter kernels, where each kernel can omit input samples from two opposite corners (i.e., exclude these input samples). Additionally or alternatively, an alternative configuration can use a filter kernel that has more input samples directly above or to the left of the output sample location compared to the number of input samples on the opposite boundaries of the kernel.
[0070] In the example embodiment, the sample block can be predicted using a filter of the following form: (Equation 2)
[0071] Where p(x, y) is the predicted sample at position (x, y), and c i The filter coefficients, s, are determined using reconstructed samples from the neighborhood of the block to be predicted. i (x,y) is the input sample of the filter when predicting p(x,y).
[0072] The input samples of the filter can be determined in different ways. For example, the input samples can be obtained from a fixed position relative to p(x,y). For different selections of the (multiple) positions of the input samples, there can also be (multiple) different operation modes. In this case, the encoder can decide the most suitable mode, for example, by evaluating a cost function and selecting the mode that minimizes the cost function for the sample block to be predicted. The encoder can then include one or more syntax elements in the bitstream, and the decoder can identify the selected mode by decoding and interpreting these syntax elements.
[0073] One of the input samples in the input samples (e.g., the last s N-1 (x,y)) can be determined to have a constant value, rather than the sample value from the prediction block or the sample value outside the prediction block. Such a constant can be referred to as a bias term because it can represent a constant bias to the filter output when multiplied by its corresponding filter coefficient. Since the result of multiplying the constant input by the determined filter coefficients that are also constant within a sample block is constant, it can be pre-computed and included as a separate term in the filtering equation: (Equation 3)
[0074] It can be noted that Equation 3 includes the bias term rather than the last input sample.
[0075] In the present disclosure, the terms "bias term" and "constant value" can be used interchangeably.
[0076] The input sample s i (x,y) generally represents a sample in a still picture or a picture in a video sequence. Such a sample typically has a certain bit depth, which determines how many bits are needed to describe the sample value. Typical selections of the bit depth used in video coding include 8, 10, 12, or 16 bits per sample. If the sample has a bit depth of B bits, a convenient selection for the bias input s N-1 (x,y) can be half of the maximum value of the sample value range 2 B-1 , or (1 << B - 1), where "<<" represents the left shift operation. This selection can allow the bias input to be processed as a typical sample value with relatively similar values to other input samples. Of course, other selections can also be made. For example, constants such as 128, 512, 1024, or 65536 can be selected. In this case, when using 2 B-1 , the filtering equation can be further expressed as: (Equation 4)
[0077] It can be noted that in equation 4, 2 B-1 Multiply by N-1 filter coefficients.
[0078] As mentioned above, i is a spatial sample input s in the range of 0 to N-2. i (x, y) can be determined by a pre-selected offset relative to the position of the predicted sample p(x, y). In the case of recursive filtering along the diagonal through the prediction block, the sample directly above, directly to the left, and in the upper left direction can have already been predicted, or can be obtained from adjacent blocks that have already been encoded / decoded. Therefore, these samples can be used as input samples to the filter. Of course, the sample at position x, y cannot be used, as it is the output sample of the filtering process. Therefore, for example, samples such as... Figure 5 The filter shape is shown. In Figure 5 In the case of the example, the input sample can be given as follows: (Equation 5)
[0079] Here, s(x, y) represents the input sample at positions x and y. If the position of s(x, y) is within the prediction block, then s(x, y) can be a already predicted sample or a reconstructed sample from a neighboring block. If a neighboring block is unavailable, this sample value can be replaced with another sample value, such as using the closest available sample value. The incremental coordinates can depend on the prediction mode and can be chosen differently for different coding modes. For example, in defining the rectangular filter kernel... Figure 5 In the three example modes (510, 520, 530), the following array can be used to determine the incremental parameters. The first element in the array corresponds to The second element corresponds to ,etc:
[0080] For mode 0 (510):
[0081] [ (-1, 0), (-2, 0), (-3, 0), (0,-1), (-1,-1), (-2,-1), (-3,-1), (0,-2), (-1,-2), (-2,-2), (-3,-2), (0,-3), (-1,-3), (-2,-3), (-3,-3) ]
[0082] For pattern 1 (520):
[0083] [ (-1, 0), ( 0,-1), (-2, 0), (-1,-1), (-3, 0), (-2,-1), (-4, 0), (-3,-1), (-5, 0), (-4,-1), (-6, 0), (-5,-1), (-7, 0), (-6,-1), (-7,-1) ]
[0084] For pattern 2 (530):
[0085] [ (-1, 0), ( 0,-1), (-1,-1), (0,-2), (-1,-2), ( 0,-3), (-1,-3), (0,-4), (-1,-4), ( 0,-5), (-1,-5), (0,-6), (-1,-6), ( 0,-7), (-1,-7) ]
[0086] The above input sample position can be the coordinates relative to p at (0, 0).
[0087] In another example embodiment, spatial input samples can be symmetrically selected, where the last spatial position can be omitted from the array (i.e., not included). Typically, this selection can be made by including all samples within a W×H rectangle, except for the bottom-right sample at coordinates (0,0) relative to the output sample position and the top-left sample at coordinates (-[W-1], -[H-1]) relative to the output sample position. This selection benefits from both the lower complexity of the filtering operation itself and the lower complexity of solving for the filter coefficients, since one less coefficient needs to be determined. Meanwhile, the support window can still contain the same W... Samples within rectangle H. Such examples are... Figure 6 The diagram describes a rectangular input kernel where mode 0 is 4×4 (610), mode 1 is 8×2 (620), and mode 2 is 2×8 (630), and the following parameters can be used for incremental parameters. Implemented using an input array:
[0088] For mode 0 (610):
[0089] [ (-1, 0), (-2, 0), (-3, 0), (0,-1), (-1,-1), (-2,-1), (-3,-1), (0,-2), (-1,-2), (-2,-2), (-3,-2), (0,-3), (-1,-3), (-2,-3) ]
[0090] and Figure 5 Compared to 510, (-3, -3) is omitted.
[0091] For pattern 1 (620):
[0092] [ (-1, 0), ( 0,-1), (-2, 0), (-1,-1), (-3, 0), (-2,-1), (-4, 0), (-3,-1), (-5, 0), (-4,-1), (-6, 0), (-5,-1), (-7, 0), (-6,-1) ]
[0093] and Figure 5 Compared to 520, (-7, -1) is omitted.
[0094] For mode 2 (630):
[0095] [ (-1, 0), ( 0,-1), (-1,-1), (0,-2), (-1,-2), ( 0,-3), (-1,-3), (0,-4), (-1,-4), ( 0,-5), (-1,-5), (0,-6), (-1,-6), ( 0,-7) ]
[0096] and Figure 5 Compared to 530, (-1, -7) has been omitted.
[0097] In another example embodiment, it is possible to obtain from Figure 5 In the example configuration, some other spatial samples are omitted (e.g., instead of the last sample position). Similar to the example above, this choice allows the bias term to be included in the filter without negatively impacting the computational aspects related to solving and applying the filter, while simultaneously improving coding efficiency due to the stable nature of the bias term.
[0098] In another example embodiment, the input samples can be selected such that, for vertically aligned inputs, there are more input samples directly above the output sample compared to the input samples to the left of the output sample, and more input samples above the output sample compared to the input samples above the lower left input sample. This can help replicate the pattern along one prediction direction while still providing support for the other direction, but with a minimal number of filter coefficients and associated complexity. Figure 7 This choice of mode 2 (730) is described in the text.
[0099] Similarly, transpose selection can be performed on horizontally aligned inputs, such as... Figure 7 As shown in the diagram for mode 1 (720). For symmetric inputs, as... Figure 7In mode 0 (710), the same rule can be applied to the horizontal and vertical directions, causing the filter kernel to omit samples on the top and left boundaries of the filter, but retaining more input samples directly above and to the left of the output samples compared to at least some other rows and columns of the kernel.
[0100] For example, Figure 7 The kernel shown can be determined using the incremental parameters in the table below. To achieve:
[0101] For mode 0 (710):
[0102] [ (-1, 0), (-2, 0), (-3, 0), (0,-1), (-1,-1), (-2,-1), (0,-2), (-1,-2), (-2,-2), (0,-3) ]
[0103] and Figure 5 Compared to 510, five sample locations were omitted.
[0104] For pattern 1 (720):
[0105] [ (-1, 0), ( 0,-1), (-2, 0), (-1,-1), (-3, 0), (-2,-1), (-4, 0), (-5,0), (-6, 0), (-7, 0) ]
[0106] and Figure 5 Compared to 520, five sample locations were omitted.
[0107] For mode 2 (730):
[0108] [ (-1, 0), ( 0,-1), (-1,-1), (0,-2), (-1,-2), ( 0,-3), (0,-4), ( 0,-5), (0,-6), ( 0,-7) ]
[0109] and Figure 5 Compared to 530, five sample locations were omitted.
[0110] To determine the filter parameters, a certain number of training samples may need to be determined. For example, a fixed-size or block-size-related neighborhood of the block to be predicted can be selected as a reference sample region used in the process. Samples in the reference sample region can then be used to construct the autocorrelation matrix and cross-correlation vector between the filter input and output. After determining the autocorrelation matrix and cross-correlation vector, the filter parameters can be solved in different ways. For example, methods based on linear regression, LDL decomposition, other matrix factorizations, or Gaussian elimination can be used. Of course, other methods, with or without using the autocorrelation matrix and / or cross-correlation vector, can also be used.
[0111] The reference sample region used to train the filter can be selected in different ways. For example, the number of sample rows included above and to the left of the prediction block can be determined to be equal to the minimum of the following three values: the width of the prediction block, the height of the prediction block, and a predetermined value for the maximum number of rows that can be included in the specified reference region. The value for the maximum number of rows that can be included in the specified reference region can be, for example, 4, 8, 12, or 16, but naturally any feasible number can be chosen.
[0112] Spatial sample values (including predicted sample values for the current block to be predicted and reconstructed sample values for the block's neighboring regions) can be used as input to the filter. Alternatively, modified sample values can be used. For example, modifications may include subtracting a constant from the input values, or different filtering operations, including but not limited to low-pass filtering, high-pass filtering, band-pass filtering, and other linear or nonlinear operations applied to single or multiple sample values.
[0113] Similarly, when determining filter coefficients, the reconstructed sample values in the determined image region can be used as input for filter training, or these sample values can be modified in a manner similar to or different from that used in the filtering process before or during the determination of filter parameters.
[0114] Some inputs to a filter can be exponential or polynomial functions of the input samples, or values derived from the input samples. For example, one input to a filter can be derived based on the square of the input sample values.
[0115] In addition to creating filters using sample values from neighboring regions of a block, alternative methods can be used to determine filter parameters. For example, filter parameters determined for other predicted blocks of the same or different images can be selected.
[0116] In an example embodiment, a filter including at least one input value and at least one constant input value can be used for recursive intra-frame sample prediction, the at least one input value being determined based on predicted or reconstructed sample values.
[0117] In an example embodiment, a filter including at least one spatial input sample value and at least one constant input value can be used for recursive intra-frame sample prediction.
[0118] In an example embodiment, a filter comprising at least one constant input value and multiple spatial input sample values can be used for intra-frame sample prediction, wherein the spatial sample values can be determined based on a rectangular filter kernel that omits the lower right and upper left samples of the kernel.
[0119] In an example embodiment, a filter comprising at least one constant input value and multiple spatial input sample values can be trained for intra-frame sample prediction, wherein the spatial sample values can be determined based on a rectangular filter kernel that omits the lower right and upper left samples of the kernel.
[0120] In an example embodiment, a filter comprising at least one constant input value and a plurality of spatial input sample values can be used for intra-frame sample prediction, wherein the spatial sample values can be determined based on a rectangular filter kernel, wherein the rectangular filter kernel contains more input samples directly above or to the left of the output sample location compared to the number of input samples on the relative boundary of the kernel.
[0121] In an example embodiment, the codec may determine a block of samples to be predicted. In an example embodiment, the codec may determine a filter, for which the input may include at least one constant value and at least one value determined based on predicted or reconstructed sample values. In an example embodiment, the codec may use reconstructed sample values outside the block to be predicted to determine the filter coefficients. In an example embodiment, the codec may apply a filter with the determined filter coefficients to recursively predict samples of the block, and may use samples outside the block, predicted samples within the block, and constant values as input.
[0122] In an example embodiment, the spatial input of the filter can form a rectangular kernel with the lower right sample and upper left sample omitted.
[0123] In an example embodiment, the spatial input of the filter can form a rectangular kernel, wherein the rectangular kernel can contain more input samples directly above or to the left of the output sample location compared to the number of input samples on the relative boundaries of the kernel.
[0124] In an example embodiment, the codec may have a set of recursive intra-prediction modes, each of which has a different filter kernel, wherein each kernel may include a bias term.
[0125] In an example embodiment, the codec may have a set of recursive intra-prediction modes, each recursive intra-prediction mode having a different filter kernel, wherein at least one of these filter kernels may include a bias term, and at least one of these filter kernels may not include a bias term.
[0126] Figure 8 The illustration shows the potential steps of example method 800. Example method 800 may include: determining a block of one or more samples to be predicted (810); determining a filter for performing intra-frame sample prediction, wherein the input to the filter includes at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block (820); determining at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block (830); and applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter is recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block (840). Example method 800 may be performed, for example, by a codec, encoder, decoder, module or device configured to perform encoding, module or device configured to perform decoding, UE, network node, etc.
[0127] According to one example embodiment, an apparatus may include: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: determine a block of one or more samples to be predicted; determine a filter for performing intra-frame sample prediction, wherein the input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determine at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and apply the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0128] One or more reconstructed samples, or at least one of one or more previously predicted samples, can form a rectangular kernel with the bottom right and top left samples omitted.
[0129] One or more reconstructed samples, or at least one or more previously predicted samples, can form a rectangular kernel, wherein the rectangular kernel includes a larger number of input samples directly above or to the left of the output sample location compared to the number of input samples on the relative boundaries of the rectangular kernel.
[0130] The example apparatus can also be configured to: determine a pattern for selecting one or more reconstructed samples, or at least one of one or more previously predicted samples; and provide an indication of the determined pattern for inclusion in one or more syntax elements of the bitstream.
[0131] The pattern can be selected, at least in part, based on the evaluation of a cost function for one or more samples used to predict the block.
[0132] The example apparatus can also be configured to: parse at least one syntax element of a bitstream; and, based on at least one syntax element, determine a pattern for selecting at least one of one or more reconstructed samples, or at least one of one or more previously predicted samples.
[0133] The example apparatus may also be configured to determine at least one constant value based at least in part on at least one of the following: at least one filter coefficient, the maximum value of the range of sample values, or the bit depth of one or more samples outside the block or the bit depth of one or more previously predicted samples within the block.
[0134] The example device can also be configured to select one or more reconstructed samples, or at least one of one or more previously predicted samples, based on one or more pre-selected offsets relative to the output sample location.
[0135] The example apparatus can also be configured to: determine a reference sample region relative to the determined block; determine an autocorrelation matrix based at least in part on the reference sample region; determine a cross-correlation vector based at least in part on the reference sample region; and determine at least one filter coefficient based at least in part on the autocorrelation matrix and the cross-correlation vector.
[0136] At least one filter coefficient can be determined using at least one of the following: linear regression, LDL decomposition, matrix decomposition, or Gaussian elimination.
[0137] The example device can also be configured to modify at least one of the reconstructed samples or at least one of the previously predicted samples using at least one of the following: subtracting a constant, filtering, low-pass filtering, high-pass filtering, or band-pass filtering from the corresponding value of the reconstructed samples or at least one of the previously predicted samples.
[0138] According to one aspect, an example method may be provided, comprising: determining, using a user equipment, a block of one or more samples to be predicted; determining a filter for performing intra-sample prediction, wherein the input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determining at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0139] One or more reconstructed samples, or at least one of one or more previously predicted samples, can form a rectangular kernel with the bottom right and top left samples omitted.
[0140] One or more reconstructed samples or at least one of one or more previously predicted samples can form a rectangular kernel, wherein the rectangular kernel includes a larger number of input samples directly above or to the left of the output sample location than the number of input samples on the relative boundary of the rectangular kernel.
[0141] Example methods may also include: determining a pattern for selecting one or more reconstructed samples, or at least one of one or more previously predicted samples; and providing an indication of the determined pattern for inclusion in one or more syntax elements of the bitstream.
[0142] The pattern can be selected, at least in part, based on the evaluation of a cost function for one or more samples used to predict the block.
[0143] Example methods may also include: parsing at least one syntax element of a bitstream; and determining a pattern for selecting at least one of one or more reconstructed samples, or at least one of one or more previously predicted samples, based on at least one syntax element.
[0144] Example methods may also include determining at least one constant value based at least in part on at least one of the following: at least one filter coefficient, the maximum value of the range of sample values, or the bit depth of one or more samples outside the block, or the bit depth of one or more previously predicted samples within the block.
[0145] Example methods may also include: selecting one or more reconstructed samples, or at least one of one or more previously predicted samples, based on one or more pre-selected offsets relative to the output sample location.
[0146] The example method may also include: determining a reference sample region relative to the determined block; determining an autocorrelation matrix based at least in part on the reference sample region; determining a cross-correlation vector based at least in part on the reference sample region; and determining at least one filter coefficient based at least in part on the autocorrelation matrix and the cross-correlation vector.
[0147] At least one filter coefficient can be determined using at least one of the following: linear regression, LDL decomposition, matrix decomposition, or Gaussian elimination.
[0148] Example methods may also include modifying one or more reconstructed samples, or one or more previously predicted samples, using at least one of the following: subtracting a constant, filtering, low-pass filtering, high-pass filtering, or band-pass filtering from the corresponding value of one or more reconstructed samples, or one or more previously predicted samples.
[0149] According to one example embodiment, an apparatus may include: circuitry configured to determine a block of one or more samples to be predicted; circuitry configured to determine a filter for performing intra-sample prediction, wherein input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; circuitry configured to determine at least one filter coefficient for the filter based at least in part on at least one reconstructed sample value outside the block; and circuitry configured to apply the filter to predict one or more samples of the block using at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0150] According to one example embodiment, an apparatus may include: a processing circuitry system; and a memory circuitry system including computer program code configured, together with the processing circuitry system, to enable the apparatus to: determine a block of one or more samples to be predicted; determine a filter for performing intra-frame sample prediction, wherein input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determine at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and apply the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0151] As used in this application, the terms "circuit system" or "component" may refer to one or more or all of the following: (a) a purely hardware circuit implementation (such as an implementation solely in analog and / or digital circuit systems), and (b) a combination of hardware circuitry and software, such as (if applicable): (i) a combination of (multiple) analog and / or digital hardware circuitry with software / firmware; and (ii) any portion of (multiple) hardware processors having software, including (multiple) digital signal processors, software, and (multiple) memories, which work together to enable a device such as a mobile phone or server to perform various functions; and (c) (multiple) hardware circuitry and (multiple) / or processors, such as (multiple) microprocessors or portions thereof, which require software (e.g., firmware) for operation, but may be absent when the software is not required for operation. This definition of circuit system applies to all uses of the term in this application, including in any claim. As another example, as used in this application, the term circuit system also covers implementations of hardware circuitry or processors (or multiple processors) or portions thereof and their accompanying software and / or firmware. For example, if applicable to a particular claim element, the term "circuit system" also includes a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device.
[0152] According to one example embodiment, an apparatus may include components for: determining a block of one or more samples to be predicted; determining a filter for performing intra-sample prediction, wherein the input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determining at least one filter coefficient for the filter based at least in part on at least one reconstructed sample value outside the block; and applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0153] One or more reconstructed samples, or at least one of one or more previously predicted samples, can form a rectangular kernel with the bottom right and top left samples omitted.
[0154] One or more reconstructed samples, or at least one or more previously predicted samples, can form a rectangular kernel, wherein the rectangular kernel includes a larger number of input samples directly above or to the left of the output sample location compared to the number of input samples on the relative boundaries of the rectangular kernel.
[0155] The component can also be configured to: determine a pattern for selecting one or more reconstructed samples, or at least one of one or more previously predicted samples; and provide an indication of the determined pattern for inclusion in one or more syntax elements of the bitstream.
[0156] The pattern can be selected, at least in part, based on the evaluation of a cost function for one or more samples used to predict the block.
[0157] The component can also be configured to: parse at least one syntax element of the bitstream; and, based on at least one syntax element, determine a pattern for selecting at least one of one or more reconstructed samples, or at least one of one or more previously predicted samples.
[0158] The component can also be configured to determine at least one constant value based at least in part on at least one of the following: at least one filter coefficient, the maximum value of the range of sample values, or the bit depth of one or more samples outside the block or the bit depth of one or more previously predicted samples within the block.
[0159] The component can also be configured to: select one or more reconstructed samples, or at least one of one or more previously predicted samples, based on one or more pre-selected offsets relative to the output sample location.
[0160] The component can also be configured to: determine a reference sample region relative to the determined block; determine an autocorrelation matrix based at least in part on the reference sample region; determine a cross-correlation vector based at least in part on the reference sample region; and determine at least one filter coefficient based at least in part on the autocorrelation matrix and the cross-correlation vector.
[0161] At least one filter coefficient can be determined using at least one of the following: linear regression, LDL decomposition, matrix decomposition, or Gaussian elimination.
[0162] The component can also be configured to modify one or more reconstructed samples, or one or more previously predicted samples, using at least one of the following: subtracting a constant, filtering, low-pass filtering, high-pass filtering, or band-pass filtering from the corresponding value of one or more reconstructed samples, or one or more previously predicted samples.
[0163] Processors, memory, and / or example algorithms (which may be encoded as instructions, programs, or code) may be provided as example components for providing or causing the execution of operations.
[0164] According to one example embodiment, a non-transitory computer-readable medium includes instructions stored thereon that, when executed by at least one processor, cause at least one processor to: determine a block of one or more samples to be predicted; determine a filter for performing intra-frame sample prediction, wherein the input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determine at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and apply the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0165] According to one example embodiment, a non-transitory computer-readable medium includes program instructions stored thereon for performing at least the following: determining a block of one or more samples to be predicted; determining a filter for performing intra-frame sample prediction, wherein the input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determining at least one filter coefficient for the filter based at least in part on at least one reconstructed sample value outside the block; and applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0166] According to another example embodiment, a machine-readable non-transitory program storage device may be provided, the device tangibly embodying machine-executable instructions for performing operations including: determining a block of one or more samples to be predicted; determining a filter for performing intra-frame sample prediction, wherein the input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determining at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0167] According to another example embodiment, a non-transitory computer-readable medium includes instructions that, when executed by a means, cause the means to perform at least the following: determining a block of one or more samples to be predicted; determining a filter for performing intra-frame sample prediction, wherein the input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determining at least one filter coefficient for the filter based at least in part on at least one reconstructed sample value outside the block; and applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0168] A computer-implemented system includes: at least one processor and at least one non-transitory memory storing instructions, which, when executed by the at least one processor, cause the system to at least: determine a block of one or more samples to be predicted; determine a filter for performing intra-frame sample prediction, wherein the input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; determine at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and apply the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0169] A computer-implemented system includes: components for determining a block of one or more samples to be predicted; components for determining a filter for performing intra-sample prediction, wherein input to the filter may include at least: at least one constant value, and at least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; components for determining at least one filter coefficient for the filter based at least in part on the at least one reconstructed sample value outside the block; and components for applying the filter to predict one or more samples of the block using the at least one filter coefficient, wherein the filter may be recursively applied using at least one of: at least one constant value, a corresponding value of one or more reconstructed samples outside the block, or a corresponding value of one or more previously predicted samples within the block.
[0170] The term “non-transient” as used in this article refers to a limitation on the medium itself (i.e., tangible, not signaling), rather than a limitation on the persistence of data storage (e.g., RAM vs. ROM).
[0171] It should be understood that the above description is illustrative only. Those skilled in the art can devise various alternatives and modifications. For example, the features described in the various dependent claims can be combined with each other in any suitable combination. Furthermore, features from the different embodiments described above can be selectively combined to form new embodiments. Therefore, this specification is intended to include all such alternatives, modifications, and variations that fall within the scope of the appended claims.
Claims
1. An apparatus comprising: At least one processor; as well as At least one non-transitory memory, the at least one non-transitory memory storing instructions, the instructions, when executed by the at least one processor, cause the device to at least: Identify blocks of one or more samples to be predicted; Determine a filter for performing intra-frame sample prediction, wherein the inputs to the filter include at least: At least one constant value, and At least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; At least one filter coefficient for the filter is determined, based at least in part on at least one reconstructed sample value outside the block; and The filter is applied using the at least one filter coefficient to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of the following: The at least one constant value The corresponding values of the one or more reconstructed samples outside the block, or The corresponding values of the one or more previously predicted samples within the block.
2. The apparatus of claim 1, wherein the one or more reconstructed samples, or at least one of the one or more previously predicted samples, form a rectangular kernel in which the lower right sample and the upper left sample are omitted.
3. The apparatus of claim 1, wherein at least one of the one or more reconstructed samples or the one or more previously predicted samples forms a rectangular kernel, wherein the rectangular kernel includes a larger number of input samples directly above or to the left of the output sample location compared to the number of input samples on the relative boundaries of the rectangular kernel.
4. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: Determine a pattern for selecting at least one of the one or more reconstructed samples, or the one or more previously predicted samples; and Provides an indication of the determined pattern for inclusion in one or more syntax elements of the bitstream.
5. The apparatus of claim 4, wherein the mode is selected at least in part based on an evaluation of a cost function for predicting the block of one or more samples.
6. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: Parse at least one syntax element of the bitstream; and Based on the at least one syntax element, a pattern is determined for selecting at least one of the one or more reconstructed samples or the one or more previously predicted samples.
7. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: The at least one constant value is determined based at least in part on at least one of the following: The at least one filter coefficient, The maximum value in the range of sample values, or The bit depth of the one or more samples outside the block, or the bit depth of the one or more previously predicted samples within the block.
8. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: Based on one or more pre-selected offsets relative to the output sample location, select at least one of the one or more reconstructed samples or the one or more previously predicted samples.
9. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to: Determine a reference sample region relative to the determined block; The autocorrelation matrix is determined at least in part based on the reference sample region; The cross-correlation vector is determined at least in part based on the reference sample region; as well as The at least one filter coefficient is determined based at least in part on the autocorrelation matrix and the cross-correlation vector.
10. The apparatus of claim 9, wherein the at least one filter coefficient is determined using at least one of the following: Linear regression, LDL decomposition Matrix factorization, or Gaussian elimination.
11. The apparatus of claim 1, wherein the at least one memory stores instructions, which, when executed by the at least one processor, cause the apparatus to: Use at least one of the following to modify at least one of the one or more reconstructed samples, or at least one of the one or more previously predicted samples: Subtract the constant from the corresponding value of at least one of the one or more reconstructed samples or the one or more previously predicted samples. Filtering Low-pass filter, High-pass filter, or Bandpass filtering.
12. A method comprising: Use user equipment to identify blocks of one or more samples to be predicted; Determine a filter for performing intra-frame sample prediction, wherein the inputs to the filter include at least: At least one constant value, and At least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; At least one filter coefficient for the filter is determined, based at least in part on at least one reconstructed sample value outside the block; and The filter is applied using the at least one filter coefficient to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of the following: The at least one constant value The corresponding values of the one or more reconstructed samples outside the block, or The corresponding values of the one or more previously predicted samples within the block.
13. The method of claim 12, wherein the one or more reconstructed samples, or at least one of the one or more previously predicted samples, form a rectangular kernel in which the lower right sample and the upper left sample are omitted.
14. The method of claim 12, wherein at least one of the one or more reconstructed samples or the one or more previously predicted samples forms a rectangular kernel, wherein the rectangular kernel includes a larger number of input samples directly above or to the left of the output sample location compared to the number of input samples on the relative boundaries of the rectangular kernel.
15. The method of claim 12, further comprising: Determine a pattern for selecting at least one of the one or more reconstructed samples or the one or more previously predicted samples; as well as Provides an indication of the determined pattern for inclusion in one or more syntax elements of the bitstream.
16. The method of claim 15, wherein the mode is selected at least in part based on an evaluation of a cost function for predicting the block of one or more samples.
17. The method of claim 12, further comprising: Parse at least one syntax element of the bitstream; as well as Based on the at least one syntax element, a pattern is determined for selecting at least one of the one or more reconstructed samples or the one or more previously predicted samples.
18. The method of claim 12, further comprising: The at least one constant value is determined based at least in part on at least one of the following: The at least one filter coefficient, The maximum value in the range of sample values, or The bit depth of the one or more samples outside the block, or the bit depth of the one or more previously predicted samples within the block.
19. The method of claim 12, further comprising: Based on one or more pre-selected offsets relative to the output sample location, select at least one of the one or more reconstructed samples or the one or more previously predicted samples.
20. The method of claim 12, further comprising: Determine a reference sample region relative to the determined block; The autocorrelation matrix is determined at least in part based on the reference sample region; The cross-correlation vector is determined at least in part based on the reference sample region; as well as The at least one filter coefficient is determined based at least in part on the autocorrelation matrix and the cross-correlation vector.
21. The method of claim 20, wherein the at least one filter coefficient is determined using at least one of the following: Linear regression, LDL decomposition Matrix factorization, or Gaussian elimination.
22. The method of claim 12, further comprising: Use at least one of the following to modify at least one of the one or more reconstructed samples, or at least one of the one or more previously predicted samples: Subtract the constant from the corresponding value of at least one of the one or more reconstructed samples or the one or more previously predicted samples. Filtering Low-pass filter, High-pass filter, or Bandpass filtering.
23. An apparatus comprising components for: Identify blocks of one or more samples to be predicted; Determine a filter for performing intra-frame sample prediction, wherein the inputs to the filter include at least: At least one constant value, and At least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; At least one filter coefficient for the filter is determined, based at least in part on at least one reconstructed sample value outside the block; and The filter is applied using the at least one filter coefficient to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of the following: The at least one constant value The corresponding values of the one or more reconstructed samples outside the block, or The corresponding values of the one or more previously predicted samples within the block.
24. The apparatus of claim 23, wherein at least one of the one or more reconstructed samples, or the one or more previously predicted samples, forms a rectangular kernel in which the lower right sample and the upper left sample are omitted.
25. The apparatus of claim 23, wherein at least one of the one or more reconstructed samples or the one or more previously predicted samples forms a rectangular kernel, wherein the rectangular kernel includes a larger number of input samples directly above or to the left of the output sample location compared to the number of input samples on the relative boundaries of the rectangular kernel.
26. The apparatus of claim 23, wherein the component is further configured to: Determine a pattern for selecting at least one of the one or more reconstructed samples, or the one or more previously predicted samples; and Provides an indication of the determined pattern for inclusion in one or more syntax elements of the bitstream.
27. The apparatus of claim 26, wherein the mode is selected at least in part based on an evaluation of a cost function for predicting the block of one or more samples.
28. The apparatus of claim 23, wherein the component is further configured to: Parse at least one syntax element of the bitstream; and Based on the at least one syntax element, a pattern is determined for selecting at least one of the one or more reconstructed samples or the one or more previously predicted samples.
29. The apparatus of claim 23, wherein the component is further configured to: The at least one constant value is determined based at least in part on at least one of the following: The at least one filter coefficient, The maximum value in the range of sample values, or The bit depth of the one or more samples outside the block, or the bit depth of the one or more previously predicted samples within the block.
30. The apparatus of claim 23, wherein the component is further configured to: Based on one or more pre-selected offsets relative to the output sample location, select at least one of the one or more reconstructed samples or the one or more previously predicted samples.
31. The apparatus of claim 23, wherein the component is further configured to: Determine a reference sample region relative to the determined block; The autocorrelation matrix is determined at least in part based on the reference sample region; The cross-correlation vector is determined at least in part based on the reference sample region; as well as The at least one filter coefficient is determined based at least in part on the autocorrelation matrix and the cross-correlation vector.
32. The apparatus of claim 31, wherein the at least one filter coefficient is determined using at least one of the following: Linear regression, LDL decomposition Matrix factorization, or Gaussian elimination.
33. The apparatus of claim 23, wherein the component is further configured to: Use at least one of the following to modify at least one of the one or more reconstructed samples or the one or more previously predicted samples: Subtract the constant from the corresponding value of at least one of the one or more reconstructed samples or the one or more previously predicted samples. Filtering Low-pass filter, High-pass filter, or Bandpass filtering.
34. A non-transitory computer-readable medium comprising program instructions stored thereon, the program instructions being configured to perform at least the following: Identify blocks of one or more samples to be predicted; Determine a filter for performing intra-frame sample prediction, wherein the inputs to the filter include at least: At least one constant value, and At least one value determined based on at least one of one or more reconstructed samples outside the block, or one or more previously predicted samples within the block; At least one filter coefficient for the filter is determined, based at least in part on at least one reconstructed sample value outside the block; and The filter is applied using the at least one filter coefficient to predict the one or more samples of the block, wherein the filter is applied recursively using at least one of the following: The at least one constant value The corresponding values of the one or more reconstructed samples outside the block, or The corresponding values of the one or more previously predicted samples within the block.