Apparatus, method and computer program for video encoding and decoding
By extending decoder-side intra mode derivation to multiple regions of support with specific parameter-based predictions, the method addresses suboptimal predictions in existing DIMD methods, enhancing prediction accuracy and efficiency in video codecs.
Patent Information
- Application Number
- JP2025518954
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-04
- Filing Date
- 2023-08-23
- Publication Date
- 2025-10-15
AI Technical Summary
Existing decoder-side intra mode derivation (DIMD) methods in video codecs often result in suboptimal prediction for current blocks due to reliance on a single region of support, leading to inefficiencies in signaling intra prediction modes.
The proposed method involves determining multiple regions of support for video sample data, performing decoder-side intra mode derivation processes for each region, and computing predictions based on specific intra mode derivation parameters, including analysis of directionality and gradient histograms, to enhance prediction accuracy.
This approach improves the accuracy of intra prediction by accounting for position-dependent intra-prediction modes and sample-based weighting, resulting in more optimal block predictions.
Smart Images

Figure 2025534412000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus, a method and a computer program for video encoding and decoding. [Background technology]
[0002] In video coding, decoder-side intra mode derivation (DIMD) techniques have been shown to have a beneficial impact on prior art video codecs. These methods generally rely on estimation processes that operate on a region of support formed from previously reconstructed samples around the current block. Gradient estimation techniques can be used to predict the direction and strength of edges in the region of support. These parameters are used to estimate the intra prediction direction and ultimately derive a directional intra prediction (DIMD) mode, which is used to predict the current block based on either a single DIMD mode or a Fusion 2 or 3 DIMD mode. The use of DIMD is generally signaled for a given block.
[0003] When DIMD is signaled to be used, no other information is required to perform intra prediction for the current block, and the intra prediction mode is inferred instead. Thus, DIMD can successfully reduce the overhead required to signal a given intra prediction mode.
[0004] However, when using known DIMD methods, the resulting prediction for the current block is often not optimal. Summary of the Invention [Problem to be solved by the invention]
[0005] Now, to at least alleviate the above problems, an extended method is introduced herein.
[0006] The scope of protection sought for various embodiments of the invention is set forth in the independent claims. Embodiments and features described herein that are not within the scope of the independent claims, if any, should be interpreted as examples useful for understanding various embodiments of the invention. [Means for solving the problem]
[0007] A method according to a first aspect includes means for decoding coded samples of a block of video sample data, means for determining two or more regions of support for the block of video sample data, each region of support comprising a set of reconstructed samples at a predetermined position relative to the block of video sample data, means for performing a decoder-side intra mode derivation process for each region of support to obtain decoder-side intra mode derivation parameters specific to each region of support, and means for computing predictions for samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each region of support.
[0008] According to one embodiment, the apparatus comprises means for performing an analysis of the directionality of samples belonging to at least one region of support as part of the decoder-side intra mode derivation process.
[0009] According to one embodiment, the apparatus includes means for obtaining one or more intra-prediction modes from a decoder-side intra-mode derivation process.
[0010] According to one embodiment, the apparatus includes means for inferring whether one or more intra-prediction modes are position-dependent for a particular region of support based on decoder-side intra-mode derivation parameters specific to each region of support.
[0011] According to one embodiment, the apparatus includes means for estimating a strength of position dependency of one or more intra prediction modes for a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.
[0012] According to one embodiment, the apparatus includes means for deriving a histogram of gradients of samples of at least one region of support as part of a decoder-side intra-mode derivation process.
[0013] According to one embodiment, the apparatus includes means for deriving a decoder-side intra-mode derivation mode specific to each region of support based on a decoder-side intra-mode derivation parameter specific to each region of support.
[0014] According to one embodiment, the apparatus includes means for computing a prediction for a current block based on a decoder-side intra mode derivation mode specific to each region of support.
[0015] An apparatus according to a second aspect includes means for decoding coded samples of a block of video sample data, means for performing a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters, means for obtaining one or more intra prediction modes from the decoder-side intra mode derivation parameters, means for computing at least two predictors for the block of video sample data based on decoder-side intra prediction mode derivation parameters specific to each region of support, and means for combining the at least two predictors using blending to form a prediction for the block, wherein the blending is performed using sample-based weights.
[0016] According to one embodiment, an apparatus includes means for determining two or more regions of support for a block of video sample data, each region of support comprising a set of reconstructed samples at a predetermined position relative to the block of video sample data, and means for performing a decoder-side intra mode derivation process to derive decoder-side intra mode derivation parameters specific to each region of support.
[0017] According to one embodiment, the apparatus includes means for inferring whether one or more intra prediction modes are position-dependent for a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.
[0018] According to one embodiment, the sample-based weights depend on the position of the sample in the block.
[0019] According to one embodiment, the sample-based weights depend on the position dependency of a given intra-prediction mode relative to a particular region of support.
[0020] According to one embodiment, the apparatus includes means for inferring the use of sample-based weighting based on characteristics of a block of video data.
[0021] According to one embodiment, the apparatus includes means for inferring whether one or more intra-prediction modes are position-dependent for a particular region of support based on decoder-side intra-mode derivation parameters specific to each region of support.
[0022] According to one embodiment, the sample-based weighting is inferred depending on the strength of the position dependency of the decoder-side intra-mode derivation mode for a particular region of support.
[0023] According to one embodiment, the apparatus includes means for computing a prediction for a current block using one or more predefined intra prediction modes in combination with one or more intra prediction modes derived using a decoder-side intra mode derivation process.
[0024] According to one embodiment, the apparatus includes means for computing a prediction for a current block using sample-based weights, wherein the weights used for one or more predefined intra-prediction modes depend on decoder-side intra-mode derivation parameters specific to each region of support.
[0025] A method according to a third aspect includes decoding coded samples of a block of video sample data, determining two or more regions of support for the block of video sample data, each region of support comprising a set of reconstructed samples at a predetermined position relative to the block of video sample data, performing a decoder-side intra mode derivation process for each region of support to obtain decoder-side intra mode derivation parameters specific to each region of support, and computing predictions for samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each region of support.
[0026] A method according to a fourth aspect includes decoding coded samples of a block of video sample data, performing a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters, obtaining one or more intra prediction modes from the decoder-side intra mode derivation parameters, computing at least two predictors for the block of video sample data based on the intra prediction modes, and combining the at least two predictors using blending to form a prediction for the block, wherein the blending is performed using sample-based weights.
[0027] Thus, as described above, an apparatus and a computer-readable storage medium having code stored thereon is configured to perform one or more of the above-described methods and related embodiments.
[0028] For a better understanding of the present invention, reference will now be made, by way of example, to the accompanying drawings in which: [Brief explanation of the drawings]
[0029] [Figure 1] 1 is a schematic diagram illustrating an electronic device utilizing an embodiment of the present invention. [Figure 2]FIG. 1 is a schematic diagram illustrating user equipment suitable for utilizing embodiments of the present invention. [Figure 3] FIG. 1 is a schematic diagram illustrating electronic devices utilizing embodiments of the present invention connected using wireless and wired network connections. [Figure 4a] 1 is a schematic diagram of an encoder suitable for implementing embodiments of the present invention; [Figure 4b] FIG. 1 is a schematic diagram of a decoder suitable for implementing embodiments of the present invention; [Figure 5] 1 is a flow diagram of a decoding method according to one embodiment. [Figure 6] FIG. 1 illustrates an example of multiple regions of support for predicting a current image block. [Figure 7] FIG. 10 shows an example of using a gradient filter to analyze the directionality of pixels belonging to a region of support. [Figure 8] 4 is a flow diagram of a decoding method according to another embodiment; [Figure 9] FIG. 1 illustrates an example of sample-based blending of at least two predictors according to one embodiment. [Figure 10] 1 is a schematic diagram illustrating an exemplary multimedia communication system in which various embodiments may be implemented; DETAILED DESCRIPTION OF THE INVENTION
[0030] The following describes in detail suitable apparatus and possible mechanisms for predicting chroma samples. In this regard, reference is first made to Figures 1 and 2, which shows a block diagram of a video encoding system according to an example embodiment as a schematic block diagram of an example apparatus or electronic device 50 that can incorporate a codec according to one embodiment of the present invention. Figure 2 shows an apparatus layout according to an example embodiment. The elements of Figures 1 and 2 will now be described.
[0031] The electronic device 50 may be, for example, a mobile terminal or user equipment of a wireless communication system, but it will be appreciated that embodiments of the present invention may be implemented in any electronic device or apparatus that may require the encoding and / or decoding of video images.
[0032] The device 50 may include a housing 30 that houses and protects the device. The device 50 may further include a display 32 in the form of a liquid crystal display. In other embodiments of the invention, the display may be any suitable display technology suitable for displaying images or video. The device 50 may further include a keypad 34. In other embodiments of the invention, any suitable data or user interface mechanism may be utilized. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0033] The device 50 may include a microphone 36 or any suitable audio input, which may be a digital or analog signal input. The device 50 may further include an audio output device, which in embodiments of the present invention may be an earpiece 38, a speaker, or any one of an analog audio or digital audio output connection. The device 50 may also include a battery (or in other embodiments of the present invention, the device may be powered by any suitable mobile energy device, such as a solar cell, a fuel cell, or a clockwork generator). The device may further include a camera capable of recording or capturing images and / or video. The device 50 may further include an infrared port for short-range line-of-sight communication with other devices. In other embodiments, the device 50 may further include any suitable short-range communication solution, such as a Bluetooth wireless connection or a USB / Firewire wired connection.
[0034] The device 50 may include a controller 56, processor, or processor circuitry that controls the device 50. The controller 56 may be connected to a memory 58, which in embodiments of the present invention may store data, both in the form of image and audio data, and / or may also store instructions for implementation by the controller 56. The controller 56 may further be connected to a codec circuit 54 suitable for performing encoding and decoding of audio and / or video data or for assisting in the encoding and decoding performed by the controller.
[0035] The device 50 may further include a card reader 48 and a smart card 46, for example a UICC and UICC reader suitable for providing user information and credentials for authentication and authorization of the user on the network.
[0036] The device 50 may include a radio interface circuit 52 coupled to the controller and suitable for generating wireless communication signals for communication with, for example, a cellular communication network, a wireless communication system, or a wireless local area network. The device 50 may further include an antenna 44 coupled to the radio interface circuit 52 for transmitting radio frequency signals generated by the radio interface circuit 52 to other devices and for receiving radio frequency signals from other devices.
[0037] Device 50 may include a camera capable of recording or detecting individual frames that are then passed to a codec 54 or controller for processing. The device may receive video image data for processing from another device before transmission and / or storage. Device 50 may also receive images for encoding / decoding either wirelessly or via a wired connection. The structural elements of device 50 described above represent examples of means for performing the corresponding functions.
[0038] With reference to Figure 3, an example of a system in which embodiments of the present invention may be utilized is shown. System 10 includes a plurality of communication devices capable of communicating over one or more networks. System 10 may include any combination of wired or wireless networks, including, but not limited to, wireless cellular telephone networks (such as GSM, UMTS, and CDMA networks), wireless local area networks (WLANs) as defined by any of the IEEE 802.x standards, Bluetooth personal area networks, Ethernet local area networks, token ring local area networks, wide area networks, and the Internet.
[0039] System 10 may include both wired and wireless communication devices and / or apparatus 50 suitable for implementing embodiments of the present invention.
[0040] For example, the system shown in Figure 3 shows a representation of a mobile telephone network 11 and the Internet 28. Connectivity to the Internet 28 can include, but is not limited to, long-range wireless connections, short-range wireless connections, and various wired connections, including, but not limited to, telephone lines, cable lines, electrical wires, and similar communication paths.
[0041] Exemplary communication devices shown in system 10 may include, but are not limited to, electronic devices or devices 50, combination personal digital assistants (PDAs) and mobile phones 14, PDAs 16, integrated messaging devices (IMDs) 18, desktop computers 20, and notebook computers 22. Devices 50 may be stationary or mobile when carried by individuals on the move. Devices 50 may also be located in modes of transportation, including, but not limited to, cars, trucks, taxis, buses, trains, boats, airplanes, bicycles, motorcycles, or any similar suitable mode of transportation.
[0042] Embodiments may also be implemented in set-top boxes, i.e., digital TV receivers that may or may not have display or wireless capabilities, in tablets or (laptop) personal computers (PCs) with hardware or software or a combination of encoder / decoder implementations in various operating systems, and also in chipsets, processors, DSPs and / or embedded systems that provide hardware / software based encoding.
[0043] Some or other devices may send and receive calls and messages and may communicate with a service provider via a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that enables communication between the mobile telephone network 11 and the Internet 28. The system may include additional communication devices and different types of communication devices.
[0044] The communication devices may communicate using a variety of transmission technologies, including, but not limited to, Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Transmission Control Protocol-Internet Protocol (TCP-IP), Short Messaging Service (SMS), Multimedia Messaging Service (MMS), email, Instant Messaging Service (IMS), Bluetooth, IEEE 802.11, and any similar wireless communication technologies. The communication devices involved in implementing various embodiments of the present invention may communicate using a variety of mediums, including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0045] In telecommunications and data networks, a channel can refer to either a physical channel or a logical channel. A physical channel can refer to a physical transmission medium such as a wire, while a logical channel can refer to a logical connection through a multiplexer medium that can carry several logical channels. A channel can be used to transmit an information signal, e.g., a bit stream, from one or more senders (or transmitters) to one or more receivers.
[0046] The MPEG-2 Transport Stream (TS), specified in ISO / IEC 13818-1 or equivalently in ITU-T Recommendation H.222.0, is a format for carrying audio, video, and other media, as well as program or other metadata, in a multiplexed stream. Packet Identifiers (PIDs) are used to identify elementary streams (also known as packetized elementary streams) within a TS. Thus, logical channels within an MPEG-2 TS can be considered to correspond to specific PID values.
[0047] Available media file format standards include the ISO Base Media File Format (ISO / IEC 14496-12, which can be abbreviated as ISOBMFF) and the File Format for NAL Unit Structured Video (ISO / IEC 14496-15), which is derived from ISOBMFF.
[0048] A video codec consists of an encoder that converts input video into a compressed representation suitable for storage / transmission and a decoder that can restore the compressed video representation to a viewable form. The video encoder and / or video decoder can also be separate from each other, i.e., they do not need to form a codec. Typically, the encoder discards some information from the original video sequence in order to represent the video in a more compact form (i.e., at a lower bitrate).
[0049] A typical hybrid video encoder, such as many encoder implementations of ITU-T H.263 and H.264, encodes video information in two phases. First, pixel values of a certain picture area (or "block") are predicted, for example, by motion compensation (locating and pointing to an area in one of the previously coded video frames that closely corresponds to the block being coded) or by spatial means (using pixel values around the block being coded in a specified manner). Next, the prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the pixel value differences using a specified transform (e.g., the discrete cosine transform (DCT) or a variant thereof), quantizing the coefficients, and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the precision of the pixel representation (picture quality) and the size of the resulting coded video representation (file size or transmission bit rate).
[0050] In temporal prediction, the source of prediction is a previously decoded picture (also known as a reference picture). In intra block copy (IBC, also known as intra block copy prediction), prediction is similarly applied to temporal prediction, but the reference picture is the current picture, and only previously decoded samples are referenced in the prediction process. Inter-layer or inter-view prediction can be similarly applied to temporal prediction, but the reference picture is a decoded picture from another scalable layer or another view, respectively. In some cases, inter-prediction can refer only to temporal prediction; in other cases, inter-prediction can collectively refer to temporal prediction and any of intra block copy, inter-layer, and inter-view prediction, which are performed by the same or similar process as temporal prediction. Inter-prediction or temporal prediction is sometimes also referred to as motion compensation or motion-compensated prediction.
[0051] Motion compensation can be performed with either full-sample or sub-sample accuracy. In full-sample accuracy motion compensation, motion can be represented as a motion vector with integer values for horizontal and vertical displacement, and the motion compensation process uses these displacements to efficiently copy samples from a reference picture. In sub-sample accuracy motion compensation, the motion vector is represented by fractional or decimal values for the horizontal and vertical components of the motion vector. When a motion vector indicates a non-integer position in the reference picture, a sub-sample interpolation process is typically invoked to calculate predicted sample values based on the reference sample and a selected sub-sample position. The sub-sample interpolation process typically consists of horizontal filtering to compensate for the horizontal offset relative to the full sample position, followed by vertical filtering to compensate for the vertical offset relative to the full sample position. However, vertical processing can also occur before horizontal processing in some environments.
[0052] Inter prediction, motion compensation, or motion-compensated prediction, which may also be called temporal prediction, reduces temporal redundancy. In inter prediction, the source of prediction is a previously decoded picture. Intra prediction exploits the fact that adjacent pixels in the same picture may be correlated. Intra prediction can be performed in the spatial or transform domain, i.e., it can predict either sample values or transform coefficients. Intra prediction is generally utilized in intra coding, where inter prediction is not applicable.
[0053] One result of the encoding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be efficiently entropy coded if they are first predicted from spatially or temporally neighboring parameters. For example, motion vectors can be predicted from spatially adjacent motion vectors, and only the difference relative to the motion vector predictor is coded. Prediction of coding parameters and intra-prediction can be collectively referred to as in-picture prediction.
[0054] Figures 4a and 4b show an encoder and decoder suitable for utilizing embodiments of the present invention. A video codec consists of an encoder that converts input video into a compressed representation suitable for storage / transmission, and a decoder that can decode the compressed video representation into a viewable form. Typically, the encoder discards and / or removes some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bit rate). An example of an encoding process is shown in Figure 4a. Figure 4a shows the image (I) to be encoded. n ), predicted representations for image blocks (P ’n ), prediction error signal (D n ), the reconstructed prediction error signal (D ’n ), preliminary reconstructed image (I ’n ), the final reconstructed image (R ’n ), transformation (T) and inverse transformation (T -1 ), quantization (Q) and dequantization (Q -1 ), entropy coding (E), reference frame memory (RFM), inter prediction (P inter ), intra prediction (P intra ), mode selection (MS) and filtering (F).
[0055] An example of the decoding process is shown in Figure 4b. Figure 4b shows the predicted representation (P ’n ), the reconstruction prediction error signal (D ’n ), preliminary reconstructed image (I ’n ), the final reconstructed image (R ’n ), inverse transformation (T ―1 ), inverse quantization (Q -1 ), entropy decoding (E -1 ), Reference Frame Memory (RFM), Prediction (either Inter or Intra) (P), and Filtering (F).
[0056] Many hybrid video encoders encode video information in two phases. First, pixel values of a certain picture area (or "block") are predicted, for example, by motion compensation (locating and pointing to an area in one of the previously coded video frames that closely corresponds to the block being coded) or by spatial means (using pixel values around the block being coded in a specified manner). Next, the prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the pixel value differences using a specified transform (e.g., the discrete cosine transform (DCT) or its variants), quantizing the coefficients, and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the precision of the pixel representation (picture quality) and the size of the resulting coded video representation (file size or transmission bitrate). Video codecs may also provide a transform skip mode that the encoder can choose to use. In transform skip mode, prediction errors are coded in the sample domain, for example, by deriving sample-wise difference values for certain neighboring samples and encoding the sample-wise difference values with an entropy coder.
[0057] Entropy coding / decoding can be performed in many ways. For example, context-based coding / decoding can be applied, in which both the encoder and decoder modify the context state of the coding parameters based on previously coded / decoded coding parameters. The context-based coding can be, for example, context-adaptive binary arithmetic coding (CABAC) or context-based variable-length coding (CAVLC) or any similar entropy coding. Entropy coding / decoding can alternatively or additionally be performed using a variable-length coding scheme such as Huffman coding / decoding or Exponential-Golomb coding / decoding. The decoding of coding parameters from an entropy-coded bitstream or codeword can be called parsing.
[0058] A phrase following a bitstream (e.g., indicating according to a bitstream) can be defined to indicate out-of-band transmission, signaling, or storage in a manner in which the out-of-band data is associated with the bitstream. A phrase decoding according to a bitstream, etc., can indicate decoding of referenced out-of-band data associated with the bitstream (which can be obtained from out-of-band transmission, signaling, or storage). For example, an indication following a bitstream can refer to metadata of a container file that encapsulates the bitstream.
[0059] The H.264 / AVC standard was developed by the Joint Working Group (JVT) of the Video Coding Experts Group (VCEG) of the International Telecommunication Union's Telecommunication Normalization Sector (ITU-T) and the International Organization for Normalization (ISO) / International Electrotechnical Commission (IEC) Moving Picture Experts Group (MPEG). The H.264 / AVC standard was published by both parent normalization organizations and is known as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). Several versions of the H.264 / AVC standard have existed, incorporating new extensions and features into the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0060] Version 1 of the High-Definition Video Coding (H.265 / HEVC, also known as HEVC) standard was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard, published by both parent normalization organizations, is called ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, and is also known as MPEG-H Part 2 High-Definition Video Coding (HEVC). Later versions of H.265 / HEVC include scalable, multiview, fidelity range, 3D, and screen content coding extensions, which can be abbreviated as SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.
[0061] Versatile Video Coding (VVC) (MPEG-I Part 3), also known as ITU-T H.266, is a video compression standard developed by the Joint Working Group of Moving Picture Experts (MPEG) (formally ISO / IEC JTC1 SC29 WG11) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union (ITU), and is the successor to HEVC / H.265.
[0062] Some important definitions, bitstream and coding structures, and concepts of H.264 / AVC and HEVC are presented in this paragraph as examples of video encoders, decoders, encoding methods, decoding methods, and bitstream structures, so that embodiments can be implemented. Some important definitions, bitstream and coding structures, and concepts of H.264 / AVC are the same as those of HEVC, and therefore they are described together below. Although aspects of the present invention are not limited to H.264 / AVC or HEVC, the description is presented for one possible basis, in addition to which the present invention can be partially or completely realized.
[0063] Like many earlier video coding standards, H.264 / AVC and HEVC specify bitstream syntax and semantics, as well as the decoding process for error-free bitstreams. Although the encoding process is not specified, the encoder must produce a conforming bitstream. Conformance between the bitstream and the decoder can be verified by a hypothetical reference decoder (HRD). The standard includes coding tools to help deal with transmission errors and losses, but the use of the tools in encoding is optional, and the decoding process is not prescribed for erroneous bitstreams.
[0064] The elementary units of input to an H.264 / AVC or HEVC encoder and output of an H.264 / AVC or HEVC decoder are pictures. A picture provided as input to an encoder can also be called a source picture, and a picture decoded by a decoder can also be called a decoded picture.
[0065] The source and decoded pictures each consist of one or more sample arrays, such as one of the following sets of sample arrays: -Luma(Y) only (monochrome). -Luma and two chroma (YCbCr or YCgCo). -Green, Blue and Red (GBR, also known as RGB). - Arrays representing other unspecified monochrome or tristimulus color sampling (e.g., also known as YZX, XYZ).
[0066] In H.264 / AVC and HEVC, a picture can be either a frame or a field. A frame contains a matrix of luma samples and possibly corresponding chroma samples. A field is a set of alternating sample rows of a frame and can be used as encoder input when the source signal is interlaced. The chroma sample array can be absent (thus allowing monochrome sampling to be used) or can be sub-sampled when compared to the luma sample array. The chroma format can be summarized as follows: -In monochrome sampling, there is only one sample array, which can nominally be thought of as the luma array. In -4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array. In -4:2:2 sampling, each of the two chroma arrays has the same height as the luma array and half the width of the luma array. In 4:4:4 sampling, when separate color planes are not used, each of the two chroma arrays has the same height and width as the luma array.
[0067] In H.264 / AVC and HEVC, the sample arrays can be coded into the bitstream as separate color planes, and each coded color plane can be decoded separately from the bitstream. When separate color planes are used, each of these color planes is processed separately (by the encoder and / or decoder) as a picture by monochrome sampling.
[0068] Partitioning can be defined as the division of a set into subsets such that each element of the set is in the correct subset of the subsets.
[0069] The following representations may be used when describing HEVC encoding and / or decoding operations: A coding block may be defined as an NxN sample block for some value of N, such that the division of a coding tree block into coding blocks is a partition. A coding tree block (CTB) may be defined as an NxN sample block for some value of N, such that the division of components into coding tree blocks is a partition. A coding tree unit (CTU) may be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples for a picture having three sample arrays, or a coding tree block of samples for a picture coded using three separate color planes and syntax structures used to code a monochrome picture or sample. A coding unit (CU) may be defined as a coding block of luma samples, two corresponding coding tree blocks of chroma samples for a picture having three sample arrays, or a coding block of samples for a picture coded using three separate color planes and syntax structures used to code a monochrome picture or sample. A CU with the largest allowed size can be named an LCU (Largest Coding Unit) or a Coding Tree Unit (CTU), and a video picture is divided into non-overlapping LCUs.
[0070] A CU consists of one or more prediction units (PUs), which define a prediction process for samples within the CU, and one or more transform units (TUs), which define a prediction error coding process for the samples of the CU. Generally, a CU consists of a square block of samples, with a size selectable from a predefined set of possible CU sizes. Each PU and TU is further divided into smaller PUs and TUs to increase the granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated with it that defines the type of prediction to apply to pixels within this PU (e.g., motion vector information for an inter-predicted PU and intra prediction direction information for an intra-predicted PU).
[0071] Each TU may have associated information (e.g., including DCT coefficient information) that describes the prediction error decoding process for the samples within the TU. Typically, whether prediction error coding is applied to each CU is signaled at the CU level. If there are no prediction error residuals associated with a CU, the CU is considered to have no TU. The division of an image into CUs, and the division of CUs into PUs and TUs, is typically signaled in the bitstream so that a decoder can reproduce the intended structure of these units.
[0072] In HEVC, a picture can be partitioned into tiles, which are rectangular and contain an integer number of LCUs. In HEVC, the partitioning into tiles forms a regular grid, and the heights and widths of the tiles differ from each other by at most one LCU. In HEVC, a slice is defined as an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any) in the same access unit. In HEVC, a slice segment is defined as an integer number of coding tree units that are ordered consecutively in tile traversal and contained in a single NAL unit. The division of each picture into slice segments is called partitioning. In HEVC, an independent slice segment is defined as a slice segment in which the values of syntax elements in its slice segment header are not inferred from the values of preceding slice segments, and a dependent slice segment is defined as a slice segment in which the values of some syntax elements in its slice segment header are inferred from the values of preceding independent slice segments in decoding order. In HEVC, a slice header is defined to be the slice segment header of an independent slice segment that is either the current slice segment or an independent slice segment preceding the current dependent slice segment, and a slice segment header is defined to be the part of a coded slice segment that contains data elements for the first or all coding tree units represented in the slice segment. When tiles are not used, CUs are scanned in the raster scan order of LCUs within a tile or picture. Within an LCU, CUs have a specific scan order.
[0073] The decoder reconstructs the output video by applying prediction means similar to the encoder (using motion or spatial information generated by the encoder and stored in the compressed representation) and prediction error decoding (the inverse operation of prediction error encoding, which recovers the quantized prediction error signal in the spatial pixel domain) to form a predicted representation for a block of pixels. After applying the prediction and prediction error decoding means, the decoder sums the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) may also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as a predictive reference for subsequent frames in the video sequence.
[0074] The filtering may include, for example, one or more of the following: deblocking, sample adaptive offset (SAO), and / or adaptive loop filtering (ALF). H.264 / AVC includes deblocking, while HEVC includes both deblocking and SAO.
[0075] In a typical video codec, motion information is represented by a motion vector associated with each motion-compensated image block, such as a prediction unit. Each of these motion vectors represents the displacement of an image block in a picture to be coded (at the encoder side) or decoded (at the decoder side) and a prediction source block in one of the previously coded or decoded pictures. To efficiently represent a motion vector, the motion vector is typically differentially coded with respect to a block-specific predicted motion vector. In a typical video codec, the predicted motion vector is generated in a predefined manner, for example, by calculating the median of the coded or decoded motion vectors of neighboring blocks. Another method for generating a motion vector prediction is to generate a list of candidate predictions from neighboring blocks and / or co-located blocks of a temporal reference picture and signal a selected candidate as the motion vector predictor. In addition to predicting a motion vector value, the reference picture used for the motion-compensated prediction can be predicted, and this prediction information can be represented, for example, by a reference index of a previously coded / decoded picture. The reference index is typically predicted from neighboring blocks and / or co-located blocks of a temporal reference picture. Furthermore, common high-efficiency video codecs utilize an additional motion information encoding / decoding mechanism, often called merging / merge mode, in which all of the motion field information, including the motion vectors and corresponding reference picture indexes of each available reference picture list, is predicted and used without any modification / correction. Similarly, the prediction of the motion field information is performed using the motion field information of neighboring blocks and / or co-located blocks of the temporal reference picture, and the motion field information to be used is signaled between lists of motion field lists filled with the motion field information of the available neighboring / co-located blocks.
[0076] In common video codecs, the prediction residual after motion compensation is first transformed by a transform kernel (such as a DCT) and then coded, because there is often some correlation between the residuals, and the transformation can help reduce this correlation and provide more efficient coding.
[0077] Video coding standards and specifications allow encoders to divide coded pictures into coded slices or the like. In-picture prediction is generally disabled across slice boundaries. Thus, slices can be considered a way of dividing coded pictures into independently decodable pieces. In H.264 / AVC and HEVC, in-picture prediction can be disabled across slice boundaries. Thus, slices can be considered a way of dividing coded pictures into independently decodable pieces, and thus slices are often considered elementary units for transmission. Often, encoders can indicate in the bitstream the type of in-picture prediction that is turned off across slice boundaries, and decoder operations take this information into account when concluding, for example, which prediction sources are available. For example, if neighboring CUs reside in different slices, samples from the neighboring CUs are considered unavailable for intra prediction.
[0078] The elementary units at the output of an H.264 / AVC or HEVC encoder and at the input of an H.264 / AVC or HEVC decoder are Network Abstraction Layer (NAL) units, respectively. For transport over packet-oriented networks or storage into structured files, NAL units can be encapsulated in packets or similar structures. A byte stream format is specified in H.264 / AVC and HEVC for transmission or storage environments that do not provide a framing structure. The byte stream format separates NAL units from each other by prepending a start code to each NAL unit. To avoid false detection of NAL unit boundaries, the encoder implements a byte-oriented start code emulation prevention algorithm and adds an emulation prevention byte to the NAL unit payload if a start code occurs otherwise. To enable direct gateway operation between packet-oriented and stream-oriented systems, start code emulation prevention can always be performed regardless of whether the byte stream format is used. A NAL unit can be defined as bytes containing data in the form of RBSPs, interspersed where necessary with a syntax structure containing an indication of the type of data that follows and emulation prevention bytes. A Raw Byte Sequence Payload (RBSP) can be defined as a syntax structure containing an integer number of bytes encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits containing syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0079] A NAL unit consists of a header and a payload. In H.264 / AVC and HEVC, the NAL unit header indicates the type of the NAL unit.
[0080] In HEVC, a two-byte NAL unit header is used for all indicated NAL unit types. The NAL unit header contains one reserved bit, a six-bit NAL unit type indication, a three-bit nuh_temporal_id_plus1 indication for the temporal level (which can be required to be greater than or equal to one), and a six-bit nuh_layer_id syntax element. The temporal_id_plus1 syntax element can be considered as a temporal identifier for the NAL unit, and the zero-based TemporalID variable can be derived as follows: TemporalId = temporal_id_plus1 - 1. The abbreviation TID can be used synonymously with the TemporalId variable. A TemporalId equal to 0 corresponds to the lowest temporal level. The value of temporal_id_plus1 must be non-zero to prevent start code emulation involving two NAL unit header bytes. A bitstream generated by excluding all VCL NAL units with a TemporalId greater than or equal to the selected value and including all other VCL NAL units remains conformant. As a result, a picture with TemporalId equal to tid_value does not use any picture with TemporalId greater than tid_value as an inter-prediction reference. A sub-layer or temporal sub-layer can be defined to be a temporal scalable layer (or temporal layer, TL) of a temporal scalable bitstream consisting of non-VCL NAL units associated with VCL NAL units having a particular value of the TemporalId variable. nuh_layer_id can be understood as a scalability layer identifier.
[0081] NAL units can be categorized into video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units. In HEVC, a VCL NAL unit contains a syntax element that represents one or more CUs.
[0082] A non-VCL NAL unit can be, for example, one of the following types: sequence parameter set, picture parameter set, supplemental enhancement information (SEI) NAL unit, access unit delimiter, end of sequence NAL unit, end of bitstream NAL unit, or filler data NAL unit. While a parameter set may be necessary for the reconstruction of a decoded picture, many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.
[0083] Parameters that remain unchanged throughout a coded video sequence can be included in a sequence parameter set. In addition to parameters that may be required by the decoding process, a sequence parameter set can optionally contain video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation. In HEVC, a sequence parameter set RBSP contains parameters that can be referenced by one or more SEI NAL units that contain one or more picture parameter sets RBSP or buffering period SEI messages. A picture parameter set contains actionable parameters that do not change across several coded pictures. A picture parameter set RBSP can contain parameters that can be referenced by coded slice NAL units of one or more coded pictures.
[0084] In HEVC, a video parameter set (VPS) can be defined as a syntax structure that contains zero or more syntax elements that apply to all coded video sequences, as determined by the content of syntax elements found in the SPS referenced by syntax elements found in the PPS referenced by syntax elements found in each slice segment header.
[0085] A video parameter set RBSP may contain parameters that may be referenced by one or more sequence parameter sets RBSP.
[0086] The relationship and hierarchy among video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS) can be described as follows: A VPS exists at one level above an SPS in the parameter set hierarchy and in the context of scalability and / or 3D video. A VPS may contain parameters that are common to all slices across all (scalability or view) layers in all coded video sequences. An SPS contains parameters that are common to all slices in a particular (scalability or view) layer in all coded video sequences and may be shared by multiple (scalability or view) layers. A PPS contains parameters that are common to all slices in a particular layer representation (one scalability or view layer representation of one access unit) and may be shared by all slices in multiple layer representations.
[0087] The VPS can provide information about layer dependencies in the bitstream, as well as many other pieces of information that are applicable to all slices in all (scalability or view) layers in the entire coded video sequence. The VPS can be thought of as including two parts: a base VPS and a VPS extension part, where the VPS extension part can optionally be present.
[0088] The out-of-band transmission, signaling, or storage may additionally or alternatively be used for purposes other than robustness to transmission errors, such as ease of access or session negotiation. For example, a sample entry in a track of a file conforming to the ISO Base Media File Format may include a parameter set, and the coded data of the bitstream may be stored elsewhere in the file or in a separate file. The phrases according to the bitstream (e.g., indicating according to the bitstream) or according to a coding unit of the bitstream (e.g., indicating according to a coding tile) may be used in the claims and described embodiments to indicate out-of-band transmission, signaling, or storage in a manner in which the out-of-band data is associated with a bitstream or a coding unit, respectively. The phrases according to the bitstream or decoding according to a coding unit of a bitstream or the like may refer to decoding of the indicated out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) associated with the bitstream or coding unit, respectively.
[0089] An SEI NAL unit can contain one or more SEI messages that are not necessary for decoding the output picture but can assist in related processes such as picture output timing, rendering, error detection, error concealment, and resource reservation.
[0090] A coded picture is a coded representation of a picture. In HEVC, a coded picture can be defined as a coded representation of a picture that includes all coding tree units of the picture. In HEVC, an access unit (AU) can be defined as a set of NAL units that are consecutive in decoding order and associated with each other according to specified classification rules, and that include at most one picture with any particular value of nuh_layer_id. In addition to including VCL NAL units of a coded picture, an access unit can also include non-VCL NAL units. The specified classification rules can, for example, associate pictures with the same output time or picture output count value to the same access unit.
[0091] A bitstream can be defined as a sequence of bits, in the form of a NAL unit stream or a byte stream, forming a representation of a coded picture and associated data forming one or more coded video sequences. A first bitstream can be followed by a second bitstream of the same logical channel, such as in the same file or the same connection of a communication protocol. An elementary stream (in the context of video coding) can be defined as a sequence of one or more bitstreams. The end of a first bitstream can be indicated by a specific NAL unit, which can indicate the end of bitstream (EOB) NAL unit and is the last NAL unit of the bitstream. In HEVC and this current draft extension, the EOB NAL unit must have nuh_layer_id equal to 0.
[0092] In H.264 / AVC, a coded video sequence is defined to be a sequence of consecutive access units in decoding order from an IDR access unit, inclusively, to the next IDR access unit, exclusively, or to the end of the bitstream, even if it appears earlier.
[0093] In HEVC, a coded video sequence (CVS) can be defined as a sequence of access units, for example, consisting of, in decoding order, an IRAP access unit with NoRaslOutputFlag equal to 1, followed by zero or more access units that are not IRAP access units with NoRaslOutputFlag equal to 1, up to and including all subsequent access units but not including any subsequent access units that are IRAP access units with NoRaslOutputFlag equal to 1. An IRAP access unit can be defined as an access unit whose base layer picture is an IRAP picture. The value of NoRaslOutputFlag is equal to 1 for each IDR picture, each BLA picture, and each IRAP picture that is the first picture in this particular layer of the bitstream in decoding order and that is the first IRAP picture following, in decoding order, the last NAL unit of the sequence with the same value of nuh_layer_id. There can be a means to provide the value of HandleCraAsBlaFlag to the decoder from an external entity, such as a player or receiver, that can control the decoder. HandleCraAsBlaFlag can be set to 1, for example, by a player that seeks a new position in the bitstream or tunes to a broadcast and starts decoding, then begins decoding from a CRA picture. When HandleCraAsBlaFlag is equal to 1 at a CRA picture, the CRA picture is treated and decoded as if a BLA picture were present.
[0094] In HEVC, a coded video sequence is indicated as ending when a specific NAL unit, which may additionally or alternatively be referred to (as specified above) as the end-of-sequence (EOS) NAL unit, appears in the bitstream and has a nuh_layer_id equal to 0.
[0095] A group of pictures (GOP) and its properties can be defined as follows: A GOP can be decoded regardless of whether any previous pictures have been decoded. An open GOP is a group of pictures in which, when decoding starts from the initial intra picture of the open GOP, pictures preceding the initial intra picture in output order cannot be correctly decoded. In other words, pictures in an open GOP can refer to pictures belonging to previous GOPs (in inter prediction). A specific NAL unit type, the CRA NAL unit type, can be used for this coded slice so that an HEVC decoder can recognize the intra picture that starts the open GOP. A closed GOP is a group of pictures in which all pictures can be correctly decoded when decoding starts from the initial intra picture of the closed GOP. In other words, pictures in a closed GOP do not refer to any pictures of previous GOPs. In H.264 / AVC and HEVC, a closed GOP can start with an IDR picture. In HEVC, a closed GOP can also start from a BLA_W_RADL or BLA_N_LP picture. Due to the high flexibility in the selection of reference pictures, an open GOP coding structure is potentially more efficient in compression compared to a closed GOP coding structure.
[0096] A decoded picture buffer (DPB) can be used in an encoder and / or decoder. There are two reasons for buffering decoded pictures: reference for inter-prediction and reordering decoded pictures into output order. Because H.264 / AVC and HEVC provide a lot of flexibility for both reference picture marking and output reordering, separate buffers for reference picture buffering and output picture buffering may waste memory resources. Therefore, the DPB can include a unified decoded picture buffering process for reference pictures and output reordering. Decoded pictures can be removed from the DPB when they are not used as references and are not needed for output.
[0097] In many coding modes of H.264 / AVC and HEVC, reference pictures for inter prediction are indicated by the index of a reference picture list. The index can usually be coded by variable length coding, where a small index has a short value of the corresponding syntax element. In H.264 / AVC and HEVC, two reference picture lists (reference picture list 0 and reference picture list 1) are generated for each bi-predicted (B) slice, and one reference picture list (reference picture list 0) is formed for each inter-coded (P) slice.
[0098] Many coding standards, including H.264 / AVC and HEVC, may have a decoding process that derives a reference picture index of a reference picture list that can be used to indicate one of multiple reference pictures used for inter prediction of a particular block. The reference picture index may be coded into the bitstream by the encoder in some inter coding modes, or may be derived (by the encoder and decoder) using, for example, neighboring blocks in some other inter coding modes.
[0099] The motion parameter types or motion information may include, but are not limited to, one or more of the following types: - an indication of the prediction type (e.g., intra-prediction, uni-prediction, bi-prediction) and / or the number of reference pictures; an indication of the prediction direction, such as inter (aka temporal) prediction, inter-layer prediction, inter-view prediction, view synthesis prediction (VSP), and inter-component prediction (which may be indicated per reference picture and / or per prediction type, and in some embodiments, inter-view and view synthesis prediction may be considered together as one prediction direction); and / or - an indication of the reference picture type, such as short-term reference picture and / or long-term reference picture and / or inter-layer reference picture (which may for example be indicated per reference picture); a reference index of a reference picture list and / or any other identifier of the reference picture (for example an identifier that can be indicated per reference picture and whose type can depend on the prediction direction and / or the reference picture type, and can be accompanied by other relevant pieces of information such as the reference picture list to which the reference index applies), Horizontal motion vector components (e.g., components that can be indicated per prediction block or per reference index, etc.), - vertical motion vector components (e.g., components that can be indicated per prediction block or per reference index, etc.); one or more parameters, such as a picture order count difference and / or a relative camera separation between the picture containing or associated with the motion parameter and its reference picture, which can be used for scaling horizontal and / or vertical motion vector components in one or more motion vector prediction processes (the one or more parameters can be indicated, for example, for each reference picture or each reference index, etc.), - the coordinates of the block given by the motion parameters and / or motion information, e.g. the coordinate of the top left sample of the block in luma sample units, - The motion parameters and / or the extent of the block that the motion information gives (e.g. width and height).
[0100] Compared to previous video coding standards, the versatile video codec (H.266 / VVC) introduces several new coding tools, such as: Intra prediction - 67 Intra-mode with wide-angle mode extension -Block size and mode dependent 4-tap interpolation filter - Position-dependent intra-prediction combining (PDPC) -Cross-Component Linear Model Intra-Prediction (CCLM) -Multiple reference line intra-prediction -Intra subdivision - Weighted intra prediction with matrix multiplication Inter-picture prediction -Block motion copy with spatial, temporal, history-based, and pairwise average merging candidates -Affine motion inter-prediction - Sub-block based temporal motion vector prediction - Adaptive motion vector resolution - 8x8 block-based motion compression for temporal motion prediction High-definition (1 / 16 pel) motion vector storage and motion compensation with an 8-tap interpolation filter for the luma component and a 4-tap interpolation filter for the chroma component -triangular division - Intra and inter combined prediction -Merge with MVD (MMVD) -Symmetric MVD coding - Bidirectional Optical Flow -Decoder side motion vector adjustment - Bi-prediction by CU level weights Transformation, quantization and coefficient coding -Multiple primary transform selection with DCT2, DCT7 and DCT8 -Secondary transformation of low frequency zone - Sub-block transform of inter prediction residual Dependent quantization with max QP increased from -51 to 63 - Transform coefficient coding with sign data hiding - Transform skip residual coding Entropy coding -Arithmetic coding engine with adaptive double window probability update In-loop filter -In-loop reshaping - Deblocking filter with powerful length filter -Sample adaptive offset -Adaptive Loop Filter Screen content encoding - Current picture reference by reference area restriction 360-degree video encoding -Horizontal wraparound motion compensation High-level syntax and parallel processing Reference picture management via direct reference picture list signaling - Tile groups with rectangular tile groups
[0101] The new coding tools listed above, lacking decoder-side intra mode derivation (DIMD), have been considered for inclusion in the VVC / H.266 video codec. DIMD techniques have been shown to have beneficial impacts on prior art video codecs. These methods generally rely on estimation operations on a region of support formed from already reconstructed samples around the current block. Gradient estimation techniques can be used to predict the direction and strength of edges in the region of support. These are then used to estimate the intra prediction direction and ultimately derive the directional intra prediction mode. These modes are used to predict the current block. The use of DIMD is generally signaled for a given block. When the use of DIMD is signaled, no information is needed to perform intra prediction on the current block; instead, the intra prediction mode is inferred.
[0102] However, while DIMD can successfully reduce the overhead required to signal a given intra-prediction mode, the resulting prediction is often suboptimal.
[0103] Here, an improved method for performing decoder-side intra mode derivation is introduced.
[0104] One embodiment of a method is shown in FIG. 5 and includes decoding coded samples of a block of video sample data (500); determining two or more regions of support for the block of video sample data (502), each region of support including a set of reconstructed samples at a predetermined position relative to the block of video sample data; performing a decoder-side intra mode derivation process for each region of support to obtain decoder-side intra mode derivation parameters specific to each region of support (504); and computing predictions for samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each region of support (506).
[0105] Thus, the method generates an intra prediction for a given block and performs at least two or more decoder-side intra mode derivation (DIMD) processes, each operating on reconstructed samples extracted from a particular region of support formed from a set of samples at a particular position relative to the current block, and each process outputs DIMD parameters specific to the region of support.
[0106] Thus, at least two regions of support around the current block are identified. The regions of support contain previously reconstructed decoded samples (or pixels). Each region of support can be formed from samples at a specific location with respect to the current block. The regions of support can overlap each other.
[0107] An example of support regions is shown in Figure 6, where three support regions for a current coding unit (CU) are identified: Support Region 1, which corresponds to the samples immediately above the current block; Support Region 3, which corresponds to the samples immediately to the left of the current block; and Support Region 2, which corresponds to the samples in the top-left area of the current block. In this example, the first support region (Support Region 1) may form a three-pixel-high area of reconstructed samples immediately above the current block; the second support region (Support Region 2) may form a three-by-three pixel area of reconstructed samples located at the top-left of the current block; and the third support region (Support Region 3) may form a three-pixel-wide area of reconstructed samples immediately to the left of the current block.
[0108] For each support region, a decoder-side intra-mode derivation (DIMD) process is performed to derive DIMD parameters specific to each region.
[0109] According to one embodiment, the method comprises, as part of the decoder-side intra mode derivation process, performing an analysis of the directionality of samples belonging to at least one region of support.
[0110] Therefore, an analysis of the orientation of pixels belonging to a particular region of support can be performed as part of the DIMD processing, after which DIMD parameters specific to each region are derived.
[0111] According to one embodiment, the method includes deriving a histogram of gradients of samples of at least one region of support as part of the decoder-side intra mode derivation process.
[0112] As an example of directional analysis, 3x3 Sobel gradient filters (horizontal and vertical) can be used, which are convolved with samples in the support region to obtain a given histogram of gradients.
[0113] According to one embodiment, the method includes obtaining one or more intra-prediction modes from a decoder-side intra-mode derivation process.
[0114] The histogram has a number of bins corresponding to the number M of possible intra-prediction modes. The amplitude of a given bar in the histogram at a particular index m=0...M-1 represents the cumulative amplitude of gradients estimated to have the same direction as intra-prediction mode m. In the example shown in Figure 7, the shaded (gray) area indicates the pixel used as the center of the sliding window used for region of support 1, and the dotted line indicates a 3x3 sliding window of samples convolved with a Sobel kernel.
[0115] According to one embodiment, the method includes inferring whether one or more intra-prediction modes are position-dependent for a particular region of support based on decoder-side intra-mode derivation parameters specific to each region of support.
[0116] The DIMD parameters obtained for each region of support are then used to determine whether a given derived DIMD mode is position dependent, and if so, on which position. As one example, a given DIMD mode that can be obtained using any conventional technique can be classified as position dependent, and its position can be identified by considering, and possibly comparing, the DIMD parameters output from different regions of support.
[0117] As an example, two regions of support can be considered, one forming samples located immediately above the current block and the other forming samples located immediately to the left of the current block, corresponding to Support Region 1 and Support Region 3 in Figure 6, respectively. Two histograms of gradients, H specific to the upper support region, above and H specific to the left support region leftThen, for a given DIMD mode m, the amplitudes of the two histograms obtained in the two regions of support can be used to determine whether mode m is position dependent. left (m)==0 and H above If (m)≠0, this can be taken as an indication that mode m is position-dependent in the upper support region. Conversely, H left (m) ≠ 0 and H above If (m)==0, then m can be classified as position-dependent in the left support region.
[0118] This can be generalized as follows, where N support regions are considered and N histograms H0, H1, ... H N-1 results, a given mode m can be determined to be position-dependent in region i if Hi(m)≠0 and Hj(m)==0 for j=0, 1, ..., N-1 and j≠1.
[0119] The height of the histogram bars specific to the top and left support regions can also be used to determine the position dependence of a given mode m. As one example, mode m can be determined to be position dependent on region i, considering the factor K∈
[01] : JPEG2025534412000002.jpg14150.
[0120] According to one embodiment, the method includes performing a normalization of the DIMD parameter output for each region of support according to the number of samples that belong to the region of support. As an example, for N regions of support, the resulting N histograms H0, H1, ...H N-1 Considering s i Let m be the number of pixels that belong to support region i. Let M be the number of bins in each histogram, and let m be the cumulative number of pixels in all support regions. Consider JPEG2025534412000003.jpg14150.
[0121] Then the normalized histogram JPEG2025534412000004.jpg24150 JPEG2025534412000005.jpg6150 is It can be obtained as JPEG2025534412000006.jpg10150.
[0122] According to one embodiment, the method includes estimating the strength of positional dependency of one or more intra-prediction modes for a particular region of support based on decoder-side intra-mode derivation parameters specific to each region of support.
[0123] As an alternative or in addition to determining whether a given DIMD mode is classified as position-dependent for a given region of support, the strength of the position dependence of a given mode for a given region of support can be determined as a function of the DIMD parameters output for each region of support. As one example, given N regions of support, the resulting N histograms H0, H1, ... H, which may or may not be normalized, can be calculated. N-1 Consider a given mode m to be classified as position-dependent with respect to region i. The strength of the position-dependence of mode m with respect to region of support i can be obtained by calculating the ratio: JPEG2025534412000007.jpg16150
[0124] According to one embodiment, the method includes deriving a decoder-side intra-mode derivation mode specific to each region of support based on decoder-side intra-mode derivation parameters specific to each region of support.
[0125] According to this approach, which can be used exclusively or in combination with other embodiments, the DIMD parameters obtained for each region of support are used to determine a DIMD mode specific to each region of support. As one example, a given DIMD mode can be obtained using any conventional method, and the reconstructed pixels used to compute the DIMD parameters are restricted to pixels belonging to a particular region of support. As one example, given N regions of support, the resulting N histograms H0, H1, ...H N-1 Considering that, and M is the number of bins in each histogram, for a given support region i, a particular DIMD mode mi can be obtained as the mode with the highest peak in Hi, i.e., JPEG2025534412000008.jpg7150
[0126] According to one embodiment, the method includes computing a prediction for the current block based on a decoder-side intra-mode derivation mode specific to each region of support.
[0127] A prediction for the current block can be computed using a DIMD mode specific to each support region. As one example, an index identifying the specific support region to use to determine the DIMD mode for the current block can be signaled in the bitstream. As another example, a flag can be signaled to identify whether a prediction for the current block is computed as a result of blending the specific DIMD modes determined for each support region.
[0128] This approach can also be used in combination with other embodiments, for example, to find a unique DIMD mode m for a given region of support. i and the strength of the position dependence on the region of support i can be obtained as JPEG2025534412000009.jpg16150
[0129] According to one embodiment, which can be implemented independently or in combination with other embodiments, the method includes computing at least two predictors for a block of video sample data based on decoder-side intra-mode derived parameters specific to each region of support, and combining the at least two predictors with each other using blending to form a prediction for the block, wherein the blending is performed using sample-based weights.
[0130] When implemented independently, it can perform the method illustrated in the flowchart of Figure 8, which includes decoding coded samples of a block of video sample data (800), performing a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters (802), obtaining one or more intra prediction modes from the decoder-side intra mode derivation parameters (804), computing at least two predictors for the block of video sample data based on the intra prediction modes (806), and combining the at least two predictors with each other using blending to form a predictor for the block (808), wherein the blending is performed using sample-based weights.
[0131] Therefore, according to one approach, sample-based blending of several DIMD modes can be performed to obtain a prediction for the current block. Sample-based blending can operate based on determining weights specific to each predictor and each sample in the current block. The weights can be determined according to the position of each sample in the block, and different samples in the block can be blended using different weights. However, blending does not necessarily require determining DIMD parameters specific to each region of support.
[0132] According to one embodiment, the method includes determining two or more regions of support for a block of video sample data, each region of support comprising a set of reconstructed samples at a predetermined position relative to the block of video sample data, and performing a decoder-side intra mode derivation process to derive decoder-side intra mode derivation parameters specific to each region of support.
[0133] Thus, DIMD parameters specific to each region of support can be derived without limiting the use of these parameters solely to determining intra-prediction modes, but instead can be used to determine other things, such as determining the weights used to combine predictors or inferring whether to use sample-based weighting.
[0134] According to one embodiment, the method includes obtaining the at least two predictors based on two or more regions of support, performing a decoder-side intra mode derivation process to derive a decoder-side intra mode derivation mode specific to each region of support, and computing the at least two predictors based on the decoder-side intra mode derivation modes specific to the two or more regions of support.
[0135] As an example, suppose three regions of support are considered as shown in Figure 6. Suppose two DIMD modes, m0 and m1, are considered. These modes can be determined according to any of the methods described herein. Suppose two predictors P0 and P1 are obtained by performing intra prediction processes according to modes m0 and m1, respectively. Then, the final predictor P of the current block can be obtained as follows: P(x,y) = w1(x,y)P1(x,y) +w0(x,y)P0(x,y) where P i (x,y) is the predictor P at location (x,y) i The pixel is shown.
[0136] According to one embodiment, the method includes inferring the use of sample-based weighting based on characteristics of a block of video data.
[0137] According to one embodiment, the sample-based weights depend on the position of the sample in the block.
[0138] According to one embodiment, the sample-based weights depend on the position dependency of a given intra-prediction mode relative to a particular region of support.
[0139] Thus, the use of sample-based blending can be inferred depending, for example, on whether the DIMD modes are position-dependent on a particular support region and / or the strength of the position-dependence of each DIMD mode. w i (x,y) is the position (x,y) The weight can depend on w i (x,y) is also m i is classified as position-dependent on a particular region of support. In the above example, assume that mode m0 is classified as dependent on the region of support above, and m1 is classified as dependent on the region of support to the left, as shown in Figure 9. Assume a block of size HxW, where W is the width and H is the height. The predictor weights can then be determined for each sample as follows: JPEG2025534412000010.jpg8150 JPEG2025534412000011.jpg6150
[0140] Integer precision representation of the weights can also be considered. Assuming a 6-bit representation of the weights, they can be defined as follows: JPEG2025534412000012.jpg9150 JPEG2025534412000013.jpg6150
[0141] Also, division-free operations can be used to determine the weights, for example by scaling and shifting.
[0142] The weights can also depend on the strength of the position dependence of a given mode in a given region of support. In the above example, the strength of the position dependence of mode m from the upper region of support is above (m0), and similarly, the strength of the position dependence of m1 from the left support region is F left (m1). Next, F above (m0) and F left Using (m1), the weight w i (x, y) can be determined. As an example, two intermediate parameters can be derived, Δ x and Δ y The parameters can be derived as follows: JPEG2025534412000014.jpg9150 JPEG2025534412000015.jpg9150
[0143] Other methods of deriving the intermediate parameters can be used, which can then be used to compute the weights as follows: JPEG2025534412000016.jpg9150 JPEG2025534412000017.jpg6150
[0144] The weights can be clipped within a predefined range, for example, the weights can be clipped between 0 and 1.
[0145] The weights can also be pre-computed and stored, for example, in a look-up table that is reused during the decoding process.
[0146] According to one embodiment, the method includes computing a prediction for a current block using one or more predefined intra-prediction modes in combination with one or more intra-prediction modes derived using a decoder-side intra-mode derivation process.
[0147] Therefore, additional pre-determined predictors can also be considered and used. As an example, in addition to the two DIMD-derived predictors P0 and P1, an additional predictor P2 obtained by performing a pre-defined planar mode on the current block can be used. pin When using these pre-derived predictors, the two DIMD-derived predictors and the additional predictor can be blended using appropriate weights. These weights can be derived according to the position dependency of each mode in a given region of support and / or the strength of such position dependency, and / or the DIMD parameters output during DIMD processing.
[0148] According to one embodiment, the method includes computing a prediction for the current block using sample-based weights, where the weights used for one or more predefined intra-prediction modes depend on decoder-side intra-mode derivation parameters specific to each region of support.
[0149] Thus, if a predefined (eg, planar) mode is used in combination with DIMD, sample-based weights can be used, where the weights used for the planar mode also depend on the DIMD parameters.
[0150] According to one aspect, an apparatus includes means for decoding coded samples of a block of video sample data, means for determining two or more regions of support for the block of video sample data, each region of support including a set of reconstructed samples at a predetermined location relative to the block of video sample data, means for performing a decoder-side intra mode derivation process for each region of support to obtain decoder-side intra mode derivation parameters specific to each region of support, and means for computing predictions for samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each region of support.
[0151] According to one embodiment, the apparatus comprises means for performing an analysis of the directionality of samples belonging to at least one region of support as part of the decoder-side intra mode derivation process.
[0152] According to one embodiment, the apparatus includes means for obtaining one or more intra-prediction modes from a decoder-side intra-mode derivation process.
[0153] According to one embodiment, the apparatus includes means for inferring whether one or more intra-prediction modes are position-dependent for a particular region of support based on decoder-side intra-mode derivation parameters specific to each region of support.
[0154] According to one embodiment, the apparatus includes means for estimating the strength of positional dependency of one or more intra prediction modes in a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.
[0155] According to one embodiment, the apparatus includes means for deriving a histogram of gradients of samples of at least one region of support as part of a decoder-side intra-mode derivation process.
[0156] According to one embodiment, the apparatus includes means for deriving a decoder-side intra-mode derivation mode specific to each region of support based on decoder-side intra-mode derivation parameters specific to each region of support.
[0157] According to one embodiment, the apparatus includes means for computing a prediction for a current block based on a decoder-side intra mode derivation mode specific to each region of support.
[0158] An apparatus according to a second aspect includes means for decoding coded samples of a block of video sample data, means for performing a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters, means for obtaining one or more intra prediction modes from the decoder-side intra mode derivation parameters, means for computing at least two predictors for the block of video sample data based on decoder-side intra mode derivation parameters specific to each region of support, and means for combining the at least two predictors using blending to form a prediction for the block, wherein the blending is performed using sample-based weights.
[0159] According to one embodiment, the apparatus includes means for determining two or more regions of support for a block of video sample data, each region of support comprising a set of reconstructed samples at a predetermined position relative to the block of video sample data, and means for performing a decoder-side intra mode derivation process to derive decoder-side intra mode derivation parameters specific to each region of support.
[0160] According to one embodiment, the apparatus includes means for inferring whether one or more intra-prediction modes are position-dependent for a particular region of support based on decoder-side intra-mode derivation parameters specific to each region of support.
[0161] According to one embodiment, the sample-based weights depend on the position of the sample in the block.
[0162] According to one embodiment, the sample-based weights depend on the position dependency of a given intra-prediction mode relative to a particular region of support.
[0163] According to one embodiment, the apparatus includes means for inferring the use of sample-based weighting based on characteristics of a block of video data.
[0164] According to one embodiment, the apparatus includes means for inferring whether one or more intra prediction modes are position-dependent for a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.
[0165] According to one embodiment, the sample-based weighting is inferred depending on the strength of the position dependency of the decoder-side intra-mode derivation mode for a particular region of support.
[0166] According to one embodiment, the apparatus includes means for computing a prediction for a current block using one or more predefined intra prediction modes in combination with one or more intra prediction modes derived using a decoder-side intra mode derivation process.
[0167] According to one embodiment, the apparatus includes means for computing a prediction for a current block using sample-based weights, the weights used for one or more predefined intra-prediction modes depending on decoder-side intra-mode derivation parameters specific to each region of support.
[0168] In another aspect, there is provided an apparatus including at least one processor and at least one memory, the at least one memory storing code that, when executed by the at least one processor, causes the apparatus to at least: decode coded samples of a block of video sample data; determine two or more regions of support for the block of video sample data, each region of support comprising a set of reconstructed samples at a predetermined location with respect to the block of video sample data; perform, for each region of support, a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters specific to each region of support; and compute predictions for samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each region of support.
[0169] According to one embodiment, the apparatus includes code configured to cause the apparatus to perform an analysis of directionality of samples belonging to at least one region of support as part of a decoder-side intra-mode derivation process.
[0170] According to one embodiment, an apparatus includes code configured to cause the apparatus to obtain one or more intra-prediction modes from a decoder-side intra-mode derivation process.
[0171] According to one embodiment, the device includes code configured to cause the device to infer whether one or more intra prediction modes are position-dependent for a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.
[0172] According to one embodiment, the device includes code configured to cause the device to infer the strength of positional dependency of one or more intra prediction modes for a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.
[0173] According to one embodiment, the apparatus includes code configured to cause the apparatus to derive a histogram of gradients of samples of at least one region of support as part of a decoder-side intra-mode derivation process.
[0174] According to one embodiment, the device includes code configured to cause the device to derive a decoder-side intra-mode derivation mode specific to each support region based on decoder-side intra-mode derivation parameters specific to each support region.
[0175] According to one embodiment, the apparatus includes code configured to cause the apparatus to compute a prediction for a current block based on a decoder-side intra-mode derivation mode specific to each region of support.
[0176] An apparatus according to a fourth aspect includes at least one processor and at least one memory, the at least one memory storing code that, when executed by the at least one processor, causes the apparatus to at least: decode coded samples of a block of video sample data; perform a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters; obtain one or more intra-prediction modes from the decoder-side intra mode derivation parameters; compute at least two predictors for the block of video sample data based on decoder-side intra mode derivation parameters specific to each region of support; and combine the at least two predictors using blending to form a prediction for the block, wherein the blending is performed using sample-based weights.
[0177] According to one embodiment, an apparatus includes code configured to cause the apparatus to determine two or more regions of support for a block of video sample data, each region of support including a set of reconstructed samples at a predetermined position relative to the block of video sample data, and code configured to perform a decoder-side intra mode derivation process to cause the apparatus to derive decoder-side intra mode derivation parameters specific to each region of support.
[0178] According to one embodiment, the device includes code configured to cause the device to infer whether one or more intra prediction modes are position-dependent for a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.
[0179] According to one embodiment, the sample-based weights depend on the position of the sample in the block.
[0180] According to one embodiment, the sample-based weights depend on the position dependency of a given intra-prediction mode relative to a particular region of support.
[0181] According to one embodiment, the apparatus includes code configured to cause the apparatus to infer the use of sample-based weighting based on characteristics of a block of video data.
[0182] According to one embodiment, the device includes code configured to cause the device to infer whether one or more intra prediction modes are position-dependent for a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.
[0183] According to one embodiment, the sample-based weighting is inferred depending on the strength of the position dependency of the decoder-side intra-mode derivation mode for a particular region of support.
[0184] According to one embodiment, the device includes code configured to cause the device to compute a prediction for a current block using one or more predefined intra-prediction modes in combination with one or more intra-prediction modes derived using a decoder-side intra-mode derivation process.
[0185] According to one embodiment, the device includes code configured to cause the device to compute a prediction for a current block using sample-based weights, where the weights used for one or more predefined intra-prediction modes depend on decoder-side intra-mode derivation parameters specific to each region of support.
[0186] Such an apparatus may, for example, include the functional units disclosed in any of Figures 1, 2, 4a and 4b for implementing the embodiments.
[0187] Such an apparatus further includes code stored in the at least one memory that, when executed by the at least one processor, causes the apparatus to perform one or more of the embodiments disclosed herein.
[0188] FIG. 10 is a graphical representation of an exemplary multimedia communication system in which various embodiments can be implemented. A data source 1510 provides a source signal in analog, uncompressed digital, or compressed digital format, or any combination of these formats. An encoder 1520 may include or be connected to pre-processing, such as data format conversion and / or filtering of the source signal. The encoder 1520 encodes the source signal into a coded media bitstream. The decoded bitstream can be received directly or indirectly from a remote device located virtually on any type of network. In addition, the bitstream can be received from local hardware or software. The encoder 1520 may encode more than one media type, such as audio and video, or more than one encoder 1520 may be requested to encode different media types of the source signal. The encoder 1520 may also receive synthesized input, such as graphics and text, or may generate a coded bitstream of mixed media. In the following, the processing of only one coded media bitstream of one media type will be considered for simplicity. It should be noted, however, that a real-time broadcast service typically includes several streams (typically at least one audio, video, and text subtitle stream). It should also be noted that a system may include many encoders, but only one encoder 1520 is shown in the figure for simplicity of explanation without loss of generality. Furthermore, while the text and examples contained herein may specifically depict encoding processes, those skilled in the art should understand that the same concepts and principles will also apply to the corresponding decoding processes, and vice versa.
[0189] The encoded media bitstreams can be transferred to storage 1530. Storage 1530 can include any type of mass memory for storing encoded media bitstreams. The format of the encoded media bitstreams in storage 1530 can be an elementary self-contained bitstream format, or one or more encoded media bitstreams can be encapsulated in a container file, or the encoded media bitstreams can be encapsulated in a segment format suitable for DASH (or a similar streaming system) and stored as a sequence of segments. When one or more media bitstreams are encapsulated in a container file, a file generator (not shown) can be used to store the one or more media bitstreams in a file and generate file format metadata, which can also be stored in the file. Encoder 1520 or storage 1530 can include a file generator, or the file generator can be operably attached to either encoder 1520 or storage 1530. Some systems operate "live," i.e., omitting storage and transferring the coded media bitstreams directly from the encoder 1520 to the sender 1540. The coded media bitstreams can then be transferred to the sender 1540, also called a server, as needed. The format used for transfer can be an elementary self-contained bitstream format, a packet stream format, a segment format suitable for DASH (or similar streaming systems), or one or more coded media bitstreams can be encapsulated in a container file. The encoder 1520, storage 1530, and server 1540 can reside on the same physical device, or they can be included in different devices.The encoder 1520 and server 1540 operate on live real-time content, in which case the encoded media bitstream is generally not stored persistently, but rather is buffered for short periods in the content encoder 1520 and / or server 1540 to eliminate processing delays, transmission delays, and fluctuations in the encoded media bitrate.
[0190] The server 1540 transmits the coded media bitstream using a communication protocol stack. The stack may include, but is not limited to, one or more of Real-Time Transport Protocol (RTP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Transmission Control Protocol (TCP), and Internet Protocol (IP). When the communication protocol stack is packet-oriented, the server 1540 encapsulates the coded media bitstream into packets. For example, when RTP is used, the server 1540 encapsulates the coded media bitstream into RTP packets according to the RTP payload format. Generally, each media type has its own RTP payload format. Also, note that while the system can include more than one server 1540, for simplicity, the following description considers only one server 1540.
[0191] If the media content is encapsulated in a container file for storage 1530 or for inputting the data to the transmitter 1540, the transmitter 1540 may include or be operatively attached to a "send file parser" (not shown). In particular, if a container file is not transmitted but the enclosed coded media bitstream is encapsulated for transport over a communication protocol, the sender file parser locates the appropriate portions of the coded media bitstream to be conveyed over the communication protocol. The sender file parser may also assist in generating the correct format for the communication protocol, such as packet headers and payloads. The multimedia container file may include encapsulation instructions, such as an ISOBMFF hint track, for encapsulation of at least one of the enclosed media bitstreams in the communication protocol.
[0192] The server 1540 may or may not be connected to the gateway 1550 through a communication network, which may be, for example, a CDN, the Internet, and / or a combination of one or more access networks. The gateway may also be, or alternatively may be, a middlebox. In DASH, the gateway may be an edge server (of a CDN) or a web proxy. Note that while a system may generally include any number of gateways, for simplicity, the following description considers only one gateway 1550. The gateway 1550 may perform various types of functions, such as transitioning a packet stream according to one communication protocol stack to another, merging and forking data streams, and editing data streams according to downlink and / or receiver functions, such as controlling the bit rate of the forward stream according to pre-determined downlink network conditions. The gateway 1550 may be a server entity in various embodiments.
[0193] The system typically includes one or more receivers 1560 capable of receiving, demodulating, and decapsulating the transmitted signal into a coded media bitstream. The coded media bitstream can be transferred to archival storage 1570. The archival storage 1570 can include any type of mass memory for storing the coded media bitstream. Alternatively, or additionally, the archival storage 1570 can include computational memory, such as random access memory. The format of the coded media bitstreams in the archival storage 1570 can be an elementary self-contained bitstream format, or one or more coded media bitstreams can be encapsulated in a container file. When there are multiple coded media bitstreams, such as audio and video streams, that are related to each other, a container file is typically used, and the receiver 1560 includes or is attached to a container file generator that generates the container file from the input streams. Some systems operate "live," i.e., omitting the archival storage 1570 and transferring the coded media bitstream directly from the receiver 1560 to the decoder 1580. In some systems, only the most recent portion of the recorded stream, for example the most recent 10 minute excerpt of the recorded stream, is kept in recording storage 1570, and any earlier recorded data is discarded from recording storage 1570.
[0194] The coded media bitstreams can be transferred from the recording storage 1570 to the decoder 1580. If there are many coded media bitstreams, such as audio and video streams, that are associated with each other and encapsulated in a container file, or if a single media bitstream is encapsulated in a container file, e.g., for ease of access, a file parser (not shown) is used to decapsulate each coded media bitstream from the container file. The recording storage 1570 or the decoder 1580 can include the file parser, or the file parser is attached to either the recording storage 1570 or the decoder 1580. It should also be noted that while a system can include many decoders, only one decoder 1570 is discussed herein for simplicity of explanation without loss of generality.
[0195] The encoded media bitstream may be further processed by a decoder 1570, the output of which is one or more uncompressed media streams. Finally, a renderer 1590 may play the uncompressed media streams, for example, via loudspeakers or a display. The receiver 1560, recording storage 1570, decoder 1580, and renderer 1590 may reside on the same physical device, or they may be included in different devices.
[0196] The transmitter 1540 and / or the gateway 1550 can be configured to switch between different representations, for example, to switch between different viewpoints of the 360-degree video content, view switching, bitrate adaptation, and / or fast startup, and / or to select the representation to be transmitted. Switching between different representations can be done for multiple reasons, such as in response to a request from the receiver 1560 or preconditions, such as the throughput of the network over which the bitstream is transmitted. In other words, the receiver 1560 can initiate the switch between representations. A request from the receiver can be, for example, a request for a segment or subsegment from a different representation than before, a request for a change in the scalability layer and / or sublayer to be transmitted, or a change in a rendering device with different capabilities compared to before. A request for a segment can be an HTTP GET request. A request for a subsegment can be an HTTP GET request with a byte range. Additionally or alternatively, bitrate adjustment or bitrate adaptation can be used, for example, to provide so-called fast start-up in streaming services, where the bitrate of the transmitted stream is lower than the channel bitrate after the start of streaming or random access in order to achieve a buffer occupancy level that starts playback immediately and tolerates occasional packet delays and / or retransmissions. Bitrate adaptation can include multiple representation or layer up-switching and representation or layer down-switching operations performed in various orders.
[0197] The decoder 1580 may be configured to switch between different representations, e.g., switching between different viewpoints of 360-degree video content, view switching, bitrate adaptation, and / or fast startup, and / or the decoder 1580 may be configured to select a representation to be transmitted. Switching between different representations may be performed for several reasons, such as to achieve fast decoding operations or to adapt the transmitted bitstream, e.g., in terms of bitrate, due to preconditions such as the throughput of the network over which the bitstream is transmitted. Fast decoding operations may be necessary, for example, when a device including the decoder 1580 multitasks and uses computing resources for purposes other than decoding the video bitstream. In another embodiment, fast decoding operations may be necessary when content is played at a pace faster than normal playback speed, e.g., two or three times faster than conventional real-time playback speed.
[0198] While some embodiments have been described above with respect to and / or using HEVC and / or VVC, it should be understood that the embodiments may be implemented by any video encoder and / or decoder as well.
[0199]
[0033] Where exemplary embodiments are described above with respect to an encoder, it should be understood that the resulting bitstream and decoder may have corresponding elements. Similarly, where exemplary embodiments are described with respect to a decoder, it should be understood that the encoder may have a structure and / or computer program for generating a bitstream that is decoded by the decoder. For example, some embodiments are described with respect to generating predictive blocks as part of encoding. Embodiments may similarly be implemented by generating predictive blocks as part of decoding, with the difference being that coding parameters such as horizontal and vertical offsets are decoded from the bitstream rather than being determined by the encoder.
[0200] The above-described embodiments of the present invention describe the codec in terms of separate encoder and decoder devices to aid in understanding the processes involved. However, it will be understood that the devices, structures, and operations may be implemented as a single encoder-decoder device / structure / operation. Furthermore, the encoder and decoder may share some or all common elements.
[0201] Although the above examples show embodiments of the invention operating within a codec within an electronic device, it will be appreciated that the invention as defined in the claims may be implemented as part of any video codec. Thus, for example, embodiments of the invention may be implemented in a video codec that is capable of performing video encoding over a fixed or wired communication path.
[0202] Thus, the user equipment may include a video codec as set forth in the above embodiments of the present invention. It should be understood that the term user equipment encompasses any suitable type of wireless user equipment, such as a mobile telephone, a portable data processing device, or a portable web browser.
[0203] Furthermore, elements of a public land mobile network (PLMN) may also include video codecs such as those mentioned above.
[0204] In general, various embodiments of the present invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, although the present invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flow charts, or using some other graphical representations, it should be understood that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of example and not limitation, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing device, or some combination thereof.
[0205] Embodiments of the present invention can be implemented by computer software that can be executed by a data processor of a mobile device, such as a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of the logic flow shown in the figures can represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. Software can be stored on physical media, such as memory chips, or memory blocks embodied within a processor, magnetic media, such as hard disks or floppy disks, and optical media, such as DVDs and their data variants, CDs.
[0206] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memory, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting example, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), and a processor based on a multi-core processor architecture.
[0207] Embodiments of the present invention can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available for converting logic-level designs into semiconductor circuit designs to be etched and formed on semiconductor substrates.
[0208] Programs such as those offered by Synopsys, Inc. of Mountain View, Calif., and Cadence Design of San Jose, Calif., automatically route conductors and place components on a semiconductor chip using properly established design rules and pre-stored libraries of design modules. Once the design of a semiconductor circuit is complete, the resulting design can be sent in a normalized electronic format (e.g., Opus, GDSII, etc.) to a semiconductor manufacturing facility or "fab" for manufacturing.
[0209] The foregoing description provides a complete and informative description of exemplary embodiments of the present invention, by way of illustrative and not limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art from the foregoing description read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention are intended to fall within the scope of this invention. [Explanation of symbols]
[0210] Decodes the encoded samples of a block of 500 video sample data 502 Determine two or more regions of support for a block of video sample data Each region of support includes a set of reconstructed samples at a predetermined location relative to the block of video sample data. 504. For each region of support, perform a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters specific to each region of support. 506. Compute predictions for samples of the block of video sample data based on decoder-side intra mode derived parameters specific to each support region.
Claims
1. means for decoding coded samples of blocks of video sample data; means for determining two or more regions of support for the block of video sample data, each region of support comprising a set of reconstructed samples at a predetermined position relative to the block of video sample data; means for performing a decoder-side intra mode derivation process for each of the regions of support to obtain decoder-side intra mode derivation parameters specific to each of the regions of support; means for computing predictions for samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each region of support; An apparatus comprising:
2. The apparatus of claim 1 , comprising: means for performing an analysis of directionality of the samples belonging to at least one region of support as part of the decoder-side intra mode derivation process.
3. 3. The apparatus of claim 1, comprising means for obtaining one or more intra-prediction modes from the decoder-side intra-mode derivation process.
4. The apparatus of claim 3 , comprising: means for inferring whether the one or more intra-prediction modes are position-dependent for a particular region of support based on the decoder-side intra-mode derivation parameters specific to each region of support.
5. 5. The apparatus of claim 3, further comprising means for estimating a strength of positional dependency of the one or more intra-prediction modes for a particular region of support based on the decoder-side intra-mode derivation parameters specific to each region of support.
6. Apparatus according to any preceding claim, comprising means for deriving a histogram of gradients of samples of said at least one region of support as part of said decoder-side intra mode derivation process.
7. The apparatus according to any one of claims 1 to 6, comprising means for deriving a decoder-side intra mode derivation mode specific to each of the support regions based on the decoder-side intra mode derivation parameters specific to each of the support regions.
8. The apparatus of claim 7 , comprising: means for computing a prediction for the current block based on the decoder-side intra-mode derivation mode specific to each region of support.
9. means for decoding coded samples of blocks of video sample data; means for performing a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters; means for obtaining one or more intra prediction modes from the decoder-side intra mode derivation parameters; means for computing at least two predictors for the block of video sample data based on the decoder-side intra mode derived parameters specific to each region of support; means for combining the at least two predictors using blending to form a prediction for the block, the blending being performed using sample-based weights; and An apparatus comprising:
10. means for determining two or more regions of support for the block of video sample data, each region of support comprising a set of reconstructed samples at a predetermined position relative to the block of video sample data; means for performing the decoder-side intra mode derivation process to derive decoder-side intra mode derivation parameters specific to each of the regions of support; The apparatus of claim 9, comprising:
11. 10. The apparatus of claim 9, comprising: means for inferring whether the one or more intra-prediction modes are position-dependent for a particular region of support based on the decoder-side intra-mode derivation parameters specific to each region of support.
12. The apparatus of claim 9 , wherein the sample-based weights depend on the position of the sample in the block.
13. The apparatus of claim 9 , wherein the sample-based weights depend on a position dependency of a given intra-prediction mode relative to a particular region of support.
14. Apparatus according to any one of claims 9 to 13, comprising means for inferring the use of sample-based weighting based on characteristics of the block of video data.
15. 15. The apparatus of claim 14, comprising: means for inferring whether the one or more intra-prediction modes are position-dependent for a particular region of support based on the decoder-side intra-mode derivation parameters specific to each region of support.
16. The apparatus of claim 14 , wherein the sample-based weighting is inferred according to the strength of position dependency of a decoder-side intra-mode derivation mode for a particular region of support.
17. 7. The apparatus of claim 1, comprising: means for computing a prediction for the current block using one or more predefined intra-prediction modes in combination with one or more intra-prediction modes derived using the decoder-side intra-mode derivation process.
18. 20. The apparatus of claim 17, comprising: means for computing the prediction for the current block using sample-based weights, wherein the weights used for the one or more predefined intra-prediction modes depend on the decoder-side intra-mode derivation parameters specific to each region of support.
19. decoding coded samples of the block of video sample data; determining two or more regions of support for the block of video sample data, each region of support including a set of reconstructed samples at a predetermined position relative to the block of video sample data; performing a decoder-side intra mode derivation process for each region of support to obtain decoder-side intra mode derivation parameters specific to each region of support; computing predictions for samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each region of support; A method comprising:
20. decoding coded samples of the block of video sample data; performing a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters; obtaining one or more intra-prediction modes from the decoder-side intra-mode derivation parameters; computing at least two predictors for the block of video sample data based on the intra-prediction mode; combining the at least two predictors using blending to form a prediction for the block, the blending being performed using sample-based weights; and A method comprising:
Citation Information
Patent Citations
Image encoding / decoding method and device
JP2019535211A
Method and apparatus for interaction between decoder-side intra mode derivation and adaptive intra prediction modes
JP2022529645A
Method and apparatus for video coding using decoder side intra prediction derivation
US20190215521A1
Method and apparatus for encoding / decoding an image
US20190379891A1
Method and apparatus for interactions between decoder-side intra mode derivation and adaptive intra prediction modes
US20210243452A1