Apparatus, method and computer program for video encoding and decoding

By performing decoder-side intra-mode derivation on multiple support areas of video sample data blocks, obtaining decoder-side intra-mode derivation parameters of each support area, and calculating sample prediction of video sample data blocks, the suboptimal prediction problem caused by the DIMD method in the prior art is solved, and the prediction quality is improved.

CN119999213APending Publication Date: 2025-05-13NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070826.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-04
Filing Date
2023-08-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When DIMD methods are used in the prior art, the final prediction of the current block is usually suboptimal.

Method used

By decoding the encoded samples of the video sample data block, multiple support areas are determined, and the decoder-side intra-mode derivation process is performed for each support area, decoder-side intra-mode derivation parameters specific to each support area are obtained, and sample prediction of the video sample data block is calculated based on these parameters.

Benefits of technology

Through the decoder-side intra-mode derivation parameter calculation of multiple supporting areas, the prediction quality of the current block is improved and the suboptimality of the prediction is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119999213A_ABST
    Figure CN119999213A_ABST
Patent Text Reader

Abstract

A method comprising: decoding encoded samples of a block of video sample data; determining two or more support regions for a block of video sample data, wherein each support region comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; performing decoder-side intra mode derivation processing on each support region to obtain decoder-side intra mode derivation parameters specific to each support region; and calculating a prediction of samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each support region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to apparatus, methods and computer programs for video encoding and decoding. Background Art

[0002] In video coding, decoder-side intra mode derivation (DIMD) techniques have been shown to have a beneficial impact on state-of-the-art video codecs. These methods typically rely on an inference process that operates on a support region formed by previously reconstructed samples around the current block. Gradient estimation techniques can be used to predict the directionality and strength of edges in the support region. These parameters are used to infer the intra prediction direction and ultimately derive a directional intra prediction (DIMD) mode that is used to predict the current block based on a single DIMD mode or a fusion of two or three DIMD modes. The use of DIMD is typically signaled for a given block.

[0003] When DIMD is signaled to be used, no more information is needed to perform intra prediction for the current block, but the intra prediction mode is inferred. Therefore, DIMD can successfully reduce the overhead required to signal a given intra prediction mode.

[0004] However, when using the known DIMD method, the final prediction for the current block is often suboptimal. Summary of the invention

[0005] Now in order to at least alleviate the above problems, this paper introduces an enhanced method.

[0006] The scope of protection sought by the various embodiments of the invention is given by the independent claims. Embodiments and features described in this specification that do not fall within the scope of the independent claims, if any, are to be interpreted as examples useful for understanding the various embodiments of the invention.

[0007] The method according to the first aspect comprises means for decoding encoded samples of a block of video sample data; means for determining two or more regions of support for the block of video sample data, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; means for performing a decoder side intra mode derivation process for each region of support to obtain decoder side intra mode derivation parameters specific to each region of support; and means for computing a prediction of samples of the block of video sample data based on the decoder side intra mode derivation parameters specific to each region of support.

[0008] According to an embodiment, the apparatus comprises means for performing a directionality analysis of samples belonging to at least one region of support as part of an intra mode derivation process at a decoder side.

[0009] According to an embodiment, the apparatus comprises means for deriving one or more intra prediction modes from a decoder-side intra mode derivation process.

[0010] According to an embodiment, the apparatus comprises means for inferring whether one or more intra prediction modes are position-dependent on a specific region of support based on a decoder-side intra mode derivation parameter specific to each region of support.

[0011] According to an embodiment, the apparatus comprises means for inferring the strength of positional dependency of one or more intra prediction modes over a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.

[0012] According to an embodiment, the apparatus comprises means for deriving a gradient histogram of samples of at least one region of support as part of a decoder-side intra-mode derivation process.

[0013] According to an embodiment, the apparatus comprises means for deriving a decoder-side intra-mode derivation mode specific to each support region based on decoder-side intra-mode derivation parameters specific to each support region.

[0014] According to an embodiment, the apparatus comprises means for computing a prediction for a current block based on a decoder-side intra mode derivation mode specific to each region of support.

[0015] The apparatus according to the second aspect comprises means for decoding coded samples of a block of video sample data; means for performing a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters; means for deriving one or more intra prediction modes from the decoder-side intra mode derivation parameters; means for computing at least two predictors for the block of video sample data based on the decoder-side intra mode derivation parameters specific to each region of support; and means for combining the at least two predictors together using blending to form a prediction for the block, wherein the blending is performed using sample based weights.

[0016] According to an embodiment, the apparatus comprises means for determining two or more regions of support for a block of video sample data, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; and means for performing a decoder-side Intra mode derivation procedure to derive decoder-side Intra mode derivation parameters specific to each region of support.

[0017] According to an embodiment, the apparatus comprises means for inferring whether one or more intra prediction modes are position-dependent on a specific region of support based on a decoder-side intra mode derivation parameter specific to each region of support.

[0018] According to an embodiment, the sample-based weight depends on the position of the sample in the block.

[0019] According to an embodiment, the sample-based weights depend on the positional dependency of a given intra prediction mode over a specific support region.

[0020] According to an embodiment, the apparatus comprises means for inferring the use of sample-based weightings based on features of a block of video data.

[0021] According to an embodiment, the apparatus comprises means for inferring whether one or more intra prediction modes are position-dependent on a specific region of support based on a decoder-side intra mode derivation parameter specific to each region of support.

[0022] According to an embodiment, the sample based weighting is inferred depending on the strength of the position dependency of the decoder-side intra-mode derivation mode over a specific support region.

[0023] According to an embodiment, the apparatus comprises means for combining one or more intra prediction modes derived using a decoder-side intra mode derivation process with computing a prediction for a current block using one or more predetermined intra prediction modes.

[0024] According to an embodiment, the apparatus comprises means for computing a prediction of a current block using sample based weights, wherein the weights for one or more predefined intra prediction modes depend on decoder-side intra mode derivation parameters specific to each region of support.

[0025] The method according to the third aspect comprises decoding encoded samples of a block of video sample data; determining two or more support regions for the block of video sample data, wherein each support region comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; performing a decoder side intra mode derivation process for each support region to obtain decoder side intra mode derivation parameters specific to each support region; and computing a prediction of the samples of the block of video sample data based on the decoder side intra mode derivation parameters specific to each support region.

[0026] The method according to the fourth aspect comprises decoding coded samples of a block of video sample data; performing a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters; obtaining one or more intra prediction modes from the decoder-side intra mode derivation parameters; computing at least two predictors for the video block sample data based on the intra prediction modes; and combining at least the two predictors together using blending to form a prediction for the block, wherein the blending is performed using sample-based weights.

[0027] As described above, the apparatus and computer readable storage medium having code stored thereon are therefore arranged to perform the above-described method and one or more embodiments related thereto. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] For a better understanding of the present invention, reference will now be made, by way of example, to the accompanying drawings, in which:

[0029] Figure 1 An electronic device using an embodiment of the present invention is schematically shown;

[0030] Figure 2 Schematically illustrates a user equipment suitable for using an embodiment of the present invention;

[0031] Figure 3 Further schematically illustrating the use of embodiments of the present invention to connect electronic devices using wireless and wired network connections;

[0032] Figure 4a and Figure 4b Schematically shows an encoder and a decoder suitable for implementing an embodiment of the present invention;

[0033] Figure 5 A flowchart of a decoding method according to an embodiment is shown;

[0034] Figure 6 An example of multiple support regions for predicting a current picture block is shown;

[0035] Figure 7 An example of using a gradient filter to analyze the directionality of pixels belonging to a region of support is shown;

[0036] Figure 8 A flowchart of a decoding method according to another embodiment is shown;

[0037] Fig. 9 An example of a sample-based blend of at least two predictors according to an embodiment is shown; and

[0038] Fig.10 A schematic diagram showing an example multimedia communication system in which various embodiments may be implemented is shown. DETAILED DESCRIPTION

[0039] Suitable devices and possible mechanisms for chroma sampling prediction are described in detail below. Figure 1 and Figure 2 ,in Figure 1 A block diagram of a video encoding system according to an exemplary embodiment is shown as a schematic block diagram of an exemplary device or electronic device 50, which may combine a codec according to an embodiment of the present invention. Figure 2 The layout of the device according to the exemplary embodiment is shown. Figure 1 and Figure 2 elements.

[0040] The electronic device 50 may be, for example, a mobile terminal or user equipment of a wireless communication system. However, it should be understood that the embodiments of the present invention may be implemented in any electronic device or apparatus that may need to encode and decode video pictures or encode or decode video pictures.

[0041] The device 50 may include a housing 30 for combining and protecting the device. The device 50 may also include a display 32 in the form of a liquid crystal display. In other embodiments of the present invention, the display may be any suitable display technology suitable for displaying pictures or videos. The device 50 may also include a keyboard 34. In other embodiments of the present invention, any suitable data or user interface mechanism may be adopted. For example, the user interface may be implemented as a virtual keyboard or a data input system as part of a touch-sensitive display.

[0042] The device may include a microphone 36 or any suitable audio input that may be a digital or analog signaling input. The device 50 may also include an audio output device, which in embodiments of the present invention may be any of the following: headphones 38, speakers, or analog audio or digital audio output connections. The device 50 may also include a battery (or in other embodiments of the present invention, the device may be powered by any suitable mobile energy device, such as a solar cell, a fuel cell, or a clock generator). The device may also include a camera capable of recording or capturing pictures and / or videos. The device 50 may also include an infrared port for short-range line-of-sight communication with other devices. In other embodiments, the device 50 may also include any suitable short-range communication solution, such as a Bluetooth wireless connection or a USB / FireWire wired connection.

[0043] The apparatus 50 may include a controller 56, a processor or processor circuit for controlling the apparatus 50. The controller 56 may be connected to a memory 58, which in embodiments of the invention may store data in the form of both pictures and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may also be connected to a codec circuit 54, which may be adapted to perform encoding and decoding of audio and / or video data, or to assist in encoding and decoding performed by the controller.

[0044] The device 50 may also include a card reader 48 and a smart card 46, such as a UICC and a UICC reader, for providing user information and adapted to provide authentication information for authenticating and authorizing a user on a network.

[0045] The apparatus 50 may include a radio interface circuit 52 connected to the controller and adapted to generate wireless communication signaling, for example for communicating with a cellular communication network, a wireless communication system or a wireless local area network. The device 50 may also include an antenna 44 connected to the radio interface circuit 52 for transmitting radio frequency signaling generated at the radio interface circuit 52 to (multiple) other devices, and for receiving radio frequency signaling from (multiple) other devices.

[0046] The device 50 may include a camera capable of recording or detecting individual frames, which are then transmitted to the codec 54 or controller for processing. The device may receive video picture data for processing from another device before transmission and / or storage. The device 50 may also receive pictures for encoding / decoding wirelessly or via a wired connection. The structural elements of the above-mentioned device 50 represent examples of components for performing corresponding functions.

[0047] about Figure 3 , showing an example of a system in which embodiments of the present invention may be utilized. System 10 includes a plurality of communication devices capable of communicating over one or more networks. System 10 may include any combination of wired or wireless networks, including but not limited to wireless cellular telephone networks (such as GSM, UMTS, CDMA networks, etc.), such as wireless local area networks (WLANs) defined by any IEEE 802.x standards, Bluetooth personal area networks, Ethernet local area networks, token ring local area networks, wide area networks, and the Internet.

[0048] System 10 may include devices and / or apparatus 50 suitable for implementing both wired and wireless communications of embodiments of the present invention.

[0049] For example, Figure 3 The illustrated system shows a representation of the mobile telephone network 11 and the Internet 28. Connections to the Internet 28 may include, but are not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communications paths.

[0050] The example communication devices shown in system 10 may include, but are not limited to, electronic devices or apparatuses 50, a combination of a personal digital assistant (PDA) and a mobile phone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22. The apparatus 50 may be stationary or mobile when carried by a moving individual. The apparatus 50 may also be located in a mode of transportation, including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle, or any similar suitable mode of transportation.

[0051] Embodiments may also be implemented in set-top boxes; i.e., digital TV receivers that may or may not have a display or wireless capabilities, in tablet or (laptop) personal computers (PCs) with hardware or software or a combination of encoder / decoder implementations, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems providing hardware / software based encoding.

[0052] Some or more of the devices may send and receive calls and messages and communicate with service providers via a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that allows communication between the mobile phone network 11 and the Internet 28. The system may include additional communication devices and various types of communication devices.

[0053] The communication devices may communicate using various transmission technologies, including but not limited to code division multiple access (CDMA), global system for mobile communications (GSM), universal mobile telecommunications system (UMTS), time division multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol Internet protocol (TCP-IP), short message service (SMS), multimedia message service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11 and any similar wireless communication technology. The communication devices involved in implementing various embodiments of the present invention may communicate using various media, including but not limited to radio, infrared, laser, cable connection and any suitable connection.

[0054] In telecommunications and data networks, a channel can refer to a physical channel or a logical channel. A physical channel can refer to a physical transmission medium such as a wire, while a logical channel can refer to a logical connection on a multiplexed medium capable of transmitting several logical channels. A channel can be used to transmit information signaling (e.g., a bit stream) from one or several senders (or transmitters) to one or several receivers.

[0055] The MPEG-2 Transport Stream (TS), specified in ISO / IEC 13818-1 or equivalently in ITU-T Recommendation H.222.0, is a format for carrying audio, video and other media, as well as program metadata or other metadata in a multiplexed stream. A packet identifier (PID) is used to identify elementary streams (also called packetized elementary streams) within a TS. Therefore, a logical channel within an MPEG-2 TS can be considered to correspond to a specific PID value.

[0056] Available media file format standards include ISO Base Media File Format (ISO / IEC 14496-12, which may be abbreviated as ISOBMFF) and a file format for NAL unit structured video derived from ISOBMFF (ISO / IEC 14496-15).

[0057] A video codec consists of an encoder that converts the input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. The video encoder and / or the video decoder can also be separate from each other, i.e. they do not need to form a codec. Typically, an encoder discards some information in the original video sequence in order to represent the video in a more compact form (i.e. at a lower bit rate).

[0058] A typical hybrid video encoder (e.g., many encoder implementations of ITU-TH.263 and H.264) encodes video information in two stages. First, the pixel values ​​in a specific picture area (or "block") are predicted, for example, by a motion compensation component (finding and indicating an area that closely corresponds to the block being encoded in one of the previously encoded video frames) or by a spatial component (using pixel values ​​around the block encoded in a specified manner). Second, the prediction error is encoded, that is, the difference between the predicted pixel block and the original pixel block. This is usually done by transforming the difference in pixel values ​​using a specified transform (e.g., discrete cosine transform (DCT) or its variant), quantizing coefficients, and entropy encoding the quantized coefficients. By changing the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and the size of the resulting encoded video representation (file size or transmission bit rate).

[0059] In temporal prediction, the source of the prediction is a previously decoded picture (also called a reference picture). In intra block copy (IBC; also called intra-block copy prediction), prediction is similarly applied to temporal prediction, but the reference picture is the current picture, and only previously decoded samples can be referenced during the prediction process. Inter-layer or inter-view prediction can be similarly applied to temporal prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter-frame prediction may refer only to temporal prediction, and in other cases, inter-frame prediction may collectively refer to temporal prediction and any intra-frame block copy, inter-layer prediction, and inter-view prediction, as long as they are performed in the same or similar process as temporal prediction. Inter-frame prediction or temporal prediction may sometimes be referred to as motion compensation or motion compensated prediction.

[0060] Motion compensation can be performed with full sample or sub-sample accuracy. In the case of full sample accurate motion compensation, the motion can be represented as a motion vector with integer values ​​for horizontal and vertical displacements, which the motion compensation process uses to effectively copy samples from a reference picture. In the case of sub-sample accurate motion compensation, the motion vector is represented by fractional or decimal values ​​of the horizontal and vertical components of the motion vector. In the case where the motion vector refers to a non-integer position in the reference picture, a sub-sample interpolation process is usually called to calculate the predicted sample value based on the reference sample and the selected sub-sample position. The sub-sample interpolation process typically includes horizontal filtering compensation for the horizontal offset relative to the full sample position, followed by vertical filtering compensation for the vertical offset relative to the full sample position. However, in some environments, vertical processing can also be completed before horizontal processing.

[0061] Inter prediction (also known as temporal prediction, motion compensation or motion compensated prediction) reduces temporal redundancy. In inter prediction, the prediction source is a previously decoded picture. Intra prediction exploits the fact that adjacent pixels within the same picture may be related. Intra prediction can be performed in the spatial domain or in the transform domain, i.e., sample values ​​or transform coefficients can be predicted. Intra prediction is typically used for intra-frame coding, where inter prediction is not applied.

[0062] One result of the encoding process is a set of coded parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy coded more efficiently if they are first predicted from spatially or temporally adjacent parameters. For example, a motion vector can be predicted from spatially adjacent motion vectors, and only the difference relative to the motion vector predictor can be encoded. Prediction of coding parameters and intra-frame prediction can be collectively referred to as intra-picture prediction.

[0063] Figure 4a and Figure 4b An encoder and decoder suitable for use with embodiments of the present invention are shown. A video codec consists of an encoder that converts an input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically, an encoder discards and / or loses some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bit rate). An example of the encoding process is shown in FIG. Figure 4a shown. Figure 4a Displays the picture to be encoded (I n ); predicted representation of the image block (P' n ); prediction error signaling (D n ); reconstructed prediction error signaling (D' n ); Preliminary reconstruction of the image (I' n ); Final reconstructed image (R' n ); Transform (T) and Inverse Transform (T -1), quantization (Q) and inverse quantization (Q -1 ); Entropy coding (E); Reference frame memory (RFM); Inter-frame prediction (P inter ); intra prediction (P intra ); mode selection (MS) and filtering (F).

[0064] exist Figure 4b An example of the decoding process is shown in . Figure 4b The picture block (P' n ) prediction representation; reconstruct prediction error signaling (D' n ); Preliminary reconstruction of the image (I' n ); Final reconstructed image (R' n ); inverse transform (T -1 ); inverse quantization (Q -1 ); Entropy decoding (E -1 ); reference frame memory (RFM); prediction (inter or intra) (P); and filtering (F).

[0065] Many hybrid video encoders encode video information in two stages. First, the pixel values ​​in a specific picture area (or "block") are predicted, for example, by a motion compensation component (finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being encoded) or by a spatial component (using the pixel values ​​around the block encoded in a specified manner). Second, the prediction error, that is, the difference between the predicted pixel block and the original pixel block, is encoded. This is usually done by transforming the difference in pixel values ​​using a specified transform (e.g., discrete cosine transform (DCT) or its variant), quantizing the coefficients and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and the size of the resulting encoded video representation (file size or transmission bit rate). The video codec can also provide a transform skip mode that the encoder can choose to use. In the transform skip mode, the prediction error is encoded in the sample domain, for example by deriving a sample-by-sample difference relative to certain adjacent samples and encoding the sample-by-sample difference with an entropy encoder.

[0066] Entropy coding / decoding can be performed in a variety of ways. For example, context-based coding / decoding can be applied, in which both the encoder and the decoder modify the context state of the coding parameters based on the coding parameters of the previous coding / decoding. Context-based coding can be, for example, context-adaptive binary arithmetic coding (CABAC) or context-based variable length coding (CAVLC) or any similar entropy coding. Entropy coding / decoding can alternatively or additionally be performed using a variable length coding scheme, such as Huffman coding / decoding or Exp-Golomb coding / decoding. The decoding of the coding parameters of the bit stream or codeword from the entropy coding can be referred to as parsing.

[0067] The phrase along the bitstream (e.g., indication along the bitstream) may be defined to refer to out-of-band transmission, signaling, or storage in a manner that out-of-band data is associated with the bitstream. The phrase decoding along the bitstream, etc., etc. may refer to decoding the mentioned out-of-band data associated with the bitstream (which may be obtained from the out-of-band transmission, signaling, or storage). For example, indication along the bitstream may refer to metadata in a container file that encapsulates the bitstream.

[0068] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunication Standardization Sector of the International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard was published by two parent standardization organizations and is known as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard that integrate new extensions or features into the specification. These extensions include scalable video coding (SVC) and multi-view video coding (MVC).

[0069] Version 1 of the High Efficiency Video Coding (H.265 / HEVC, also known as HEVC) standard was developed by the Joint Collaboration Team on Video Coding (JCT-VC) of VCEG and MPEG. The standard was published by two parent standardization organizations and is known as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Later versions of H.265 / HEVC include scalable, multi-view, fidelity range, three-dimensional, and screen content coding extensions, which can be abbreviated as SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.

[0070] Versatile Video Coding (VVC) (MPEG-1 Part 3), also known as: ITU-T H.266, is a video compression standard developed by the Joint Video Experts Group (JVET) of the Motion Picture Experts Group (MPEG) (formally ISO / IEC JTC1 SC29WG11) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union (ITU), and is the successor to HEVC / H.265.

[0071] In this section, some key definitions, bitstreams and coding structures and concepts of H.264 / AVC and HEVC are described as examples of video encoders, decoders, encoding methods, decoding methods and bitstream structures, in which these embodiments can be implemented. Some key definitions, bitstreams and coding structures and concepts of H.264 / AVC are the same as those in HEVC, so they are described together below. Aspects of the present invention are not limited to H.264 / AVC or HEVC, but are described for one possible basis on which the present invention can be partially or fully implemented.

[0072] Similar to many earlier video coding standards, the bitstream syntax and semantics and the decoding process for error-free bitstreams are specified in H.264 / AVC and HEVC. The encoding process is not specified, but the encoder must generate a consistent bitstream. Bitstream and decoder consistency can be verified with a hypothetical reference decoder (HRD). The standard includes coding tools that help cope with transmission errors and losses, but their use in encoding is optional and the decoding process is not specified for erroneous bitstreams.

[0073] The basic unit for the input of an H.264 / AVC or HEVC encoder and the output of an H.264 / AVC or HEVC decoder, respectively, is a picture. A picture given as input to an encoder may also be referred to as a source picture, and a picture decoded by a decoder may be referred to as a decoded picture.

[0074] The source picture and the decoded picture each include one or more sample arrays, such as one of the following sample array sets:

[0075] - Luminance (Y) only (monochrome).

[0076] - Luma and dual chroma (YCbCr or YCgCo).

[0077] - Green, Blue and Red (GBR, also known as RGB).

[0078] -An array representing other unspecified monochrome or tristimulus color samples (e.g., YZX, also called XYZ).

[0079] In H.264 / AVC and HEVC, pictures can be either frames or fields. A frame consists of a matrix of luma samples and possibly corresponding chroma samples. A field is a set of alternating sample rows of a frame, which can be used as encoder input when the source signaling is interlaced. The chroma sample array may not be present (so monochrome sampling may be used), or the chroma sample array may be subsampled compared to the luma sample array. The chroma formats can be summarized as follows:

[0080] - In monochrome sampling, there is only one sample array, which nominally can be considered as the brightness array.

[0081] - In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array.

[0082] - In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array.

[0083] - In 4:4:4 sampling, when separate color planes are not used, each of the two chroma arrays has the same height and width as the luma array.

[0084] In H.264 / AVC and HEVC, sample arrays can be encoded into the bitstream as separate color planes, and the separate encoded color planes from the bitstream are decoded separately. When separate color planes are used, each of them is processed separately (by the encoder and / or decoder) as a picture with monochrome sampling.

[0085] Partitioning can be defined as partitioning a set into subsets such that each element of the set is in exactly one of the subsets.

[0086] When describing the operation of HEVC encoding and / or decoding, the following terms may be used. A coding block may be defined as a block of N x N samples for a certain value of N, such that partitioning a coding tree block into coding blocks is a partition. A coding tree block (CTB) may be defined as a block of N x N samples for a certain value of N, such that partitioning a component into coding tree blocks is a partition. A coding tree unit (CTU) may be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples of a picture with three sample arrays, or a coding tree block of samples of a monochrome picture or a picture encoded using three separate color planes and syntax structures for encoding samples. A coding unit (CU) may be defined as a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture with three sample arrays, or a coding block of samples of a monochrome picture or a picture encoded using three separate color planes and syntax structures for encoding samples. A CU with the maximum allowed size may be named LCU (maximum coding unit) or coding tree unit (CTU), and a video picture is divided into non-overlapping LCUs.

[0087] A CU consists of one or more prediction units (PUs) that define the prediction process for samples within the CU and one or more transform units (TUs) that define the prediction error coding process for samples in the CU. Typically, a CU consists of square blocks of samples, the size of which can be selected from a predetermined set of possible CU sizes. Each PU and TU can also be divided into smaller PUs and TUs to increase the granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated with it that defines which type of prediction will be applied to the pixels within the PU (e.g., motion vector information for inter-prediction PUs and intra-prediction directionality information for intra-prediction PUs).

[0088] Each TU may be associated with information describing the prediction error decoding process for the samples within the TU, including, for example, DCT coefficient information. Signaling is typically issued at the CU level to indicate whether prediction error coding is applied to each CU. In the case where there is no prediction error residual associated with a CU, the CU may be considered to have no TUs. The partitioning of a picture into CUs and the partitioning of a CU into PUs and TUs is typically signaled in the bitstream, allowing the decoder to reproduce the expected structure of these units.

[0089] In HEVC, a picture can be partitioned into rectangular tiles ("tiles") and include an integer number of LCUs. In HEVC, the partitioning of tiles forms a regular grid where the height and width of the tiles differ from each other by a maximum of one LCU. In HEVC, a slice is defined as an integer number of coding tree units included in an independent slice segment and all subsequent dependent slice segments (if any) before the next independent slice segment (if any) within the same access unit. In HEVC, a segment is defined as an integer number of coding tree units that are ordered consecutively in a tile scan and included in a single NAL unit. The division of each picture into segments is a partitioning. In HEVC, an independent slice segment is defined as a slice segment for which the values ​​of the syntax elements of the slice segment header are not inferred from the values ​​of the previous slice segment, and a dependent slice segment is defined as a slice segment for which the values ​​of some syntax elements of the slice segment header are inferred from the values ​​of the previous independent slice segment in decoding order. In HEVC, a slice header is defined as a slice segment header of an independent slice segment, which is either the current slice segment or an independent slice segment preceding the current dependent slice segment, and a slice segment header is defined as the portion of the coded slice segment that includes data elements belonging to the first or all coding tree units represented in the slice segment. If tiles are not used, CUs are scanned in the raster scan order of LCUs within tiles or within a picture. Within an LCU, CUs have a specific scan order.

[0090] The decoder reconstructs the output video by applying a prediction component similar to the encoder to form a predicted representation of pixel blocks (using motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (the inverse operation of prediction error encoding to recover quantized prediction error signaling in the spatial pixel domain). After applying the prediction and prediction error decoding components, the decoder adds the prediction and prediction error signaling (pixel values) to form the output video frame. The decoder (and encoder) may also apply additional filtering components to improve the quality of the output video before passing it for display and / or storing it as a prediction reference for upcoming frames in the video sequence.

[0091] The filtering may, for example, include one or more of: deblocking, sample adaptive offset (SAO), and / or adaptive loop filtering (ALF).H.264 / AVC includes deblocking, while HEVC includes both deblocking and SAO.

[0092] In a typical video codec, motion information, such as a prediction unit, is indicated by a motion vector associated with each motion compensated picture block. Each of these motion vectors represents the displacement of a picture block in a picture to be encoded (on the encoder side) or decoded (on the decoder side) and a prediction source block in one of the previously encoded or decoded pictures. In order to efficiently represent motion vectors, those motion vectors are usually differentially encoded relative to block-specific predicted motion vectors. In a typical video codec, predicted motion vectors are created in a predetermined manner, such as calculating the median of the encoded or decoded motion vectors of adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in a temporal reference picture, and signal the selected candidates as motion vector predictors. In addition to predicting motion vector values, it is also possible to predict which (multiple) reference pictures are used for motion compensated prediction, and this prediction information can be represented, for example, by reference indices of previously encoded / decoded pictures. Reference indices are usually predicted based on adjacent blocks and / or co-located blocks in a temporal reference picture. In addition, a typical high-efficiency video codec adopts an additional motion information encoding / decoding mechanism generally referred to as merge / merge mode, in which all motion field information including motion vectors and corresponding reference picture indexes for each available reference picture list is predicted and used without any modification / correction. Similarly, prediction of motion field information is performed using motion field information of neighboring blocks and / or co-located blocks in a temporal reference picture, and the used motion field information is signaled in a list of motion field candidate lists populated with motion field information of available neighboring / co-located blocks.

[0093] In a typical video codec, the motion compensated prediction residual is first transformed with a transform kernel (such as DCT) and then encoded. The reason for this is that there is usually still some correlation in the residual, and the transform can help reduce this correlation and provide more efficient coding in many cases.

[0094] Video coding standards and specifications may allow the encoder to divide a coded picture into coded slices, etc. Intra-picture prediction is typically disabled across slice boundaries. Slices may therefore be considered a way to partition a coded picture into independently decodable segments. In H.264 / AVC and HEVC, intra-picture prediction across slice boundaries may be disabled. Slices may therefore be considered a way to partition a coded picture into independently decodable slices, and slices are therefore typically considered a basic unit for transmission. In many cases, the encoder may indicate in the bitstream which type of intra-picture prediction is turned off across slice boundaries, and the decoder operation may take this information into account, for example, when determining which prediction sources are available. For example, if adjacent CUs reside in different slices, samples from adjacent CUs may be considered unavailable for intra prediction.

[0095] The basic unit for the output of the H.264 / AVC or HEVC encoder and the input of the H.264 / AVC or HEVC decoder is the network abstraction layer (NAL) unit, respectively. For transmission over a packet-oriented network or storage in a structured file, the NAL unit can be encapsulated into a packet or similar structure. In H.264 / AVC and HEVC, a byte stream format has been specified for a transmission or storage environment that does not provide a framing structure. The byte stream format separates NAL units from each other by appending a start code before each NAL unit. In order to avoid erroneous detection of NAL unit boundaries, the encoder runs a byte-oriented start code emulation prevention algorithm that adds emulation prevention bytes to the NAL unit payload if the start code has already appeared. In order to enable direct gateway operation between packet-oriented systems and stream-oriented systems, prevention of start code emulation can always be performed, regardless of whether the byte stream format is used. The NAL unit can be defined as a grammatical structure including an indication of the data type to be followed and bytes including data in the form of RBSP, and the RBSP is interspersed with emulation prevention bytes as needed. A Raw Byte Sequence Payload (RBSP) may be defined as a syntax structure comprising an integer number of bytes encapsulated in a NAL unit. An RBSP is either empty or has the form of a data bit string comprising syntax elements, followed by an RBSP stop bit, followed by zero or more subsequent bits equal to 0.

[0096] The NAL unit consists of a header and a payload. In H.264 / AVC and HEVC, the NAL unit header indicates the type of the NAL unit.

[0097] In HEVC, a 2-byte NAL unit header is used for all specified NAL unit types. The NAL unit header includes a reserved bit, a six-bit NAL unit type indication, a three-bit nuh_temporal_id_plus1 indication for the temporal level (which may need to be greater than or equal to 1), and a six-bit nuh_layer_id syntax element. The temporal_id_plus1 syntax element can be regarded as a temporal identifier for the NAL unit, and the zero-based TemporalId variable can be derived as follows: TemporalId = temporal_id_plus1–1. The abbreviation TID can be used interchangeably with the TemporalId variable. A TemporalId equal to 0 corresponds to the lowest temporal level. The value of temporal_id_plus1 is required to be non-zero to avoid start code emulation involving two NAL unit header bytes. The bitstream created by excluding all VCL NAL units with a TemporalId greater than or equal to the selected value and including all other VCL NAL units is consistent. Therefore, a picture with TemporalId equal to tid_value does not use any picture with a TemporalId greater than tid_value as an inter-frame prediction reference. A sublayer or temporal sublayer can be defined as a temporal scalable layer (or temporal layer TL) of a temporal scalable bitstream, which consists of VCL NAL units with a specific value of the TemporalId variable and associated non-VCL NAL units. nuh_layer_id can be understood as a scalability layer identifier.

[0098] NAL units can be classified into video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are usually coded slice NAL units. In HEVC, VCL NAL units include syntax elements representing one or more CUs.

[0099] A non-VCL NAL unit may be, for example, one of the following types: a sequence parameter set, a picture parameter set, a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end-of-sequence NAL unit, an end-of-bitstream NAL unit, or a filler data NAL unit. Parameter sets may be required to reconstruct a deconstructed picture, while many other non-VCL NAL units are not required to reconstruct decoded sample values.

[0100] Parameters that remain unchanged in the encoded video sequence may be included in a sequence parameter set. In addition to parameters that may be required for the decoding process, the sequence parameter set may optionally include video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation. In HEVC, the sequence parameter set RBSP includes parameters that can be referenced by one or more picture parameter set RBSPs or one or more SEI NAL units containing a buffering period SEI message. The picture parameter set includes these parameters that may not change in several coded pictures. The picture parameter set RBSP may include parameters that can be referenced by the coded slice NAL units of one or more coded pictures.

[0101] In HEVC, a video parameter set (VPS) can be defined as a syntax structure comprising syntax elements that apply to zero or more complete coded video sequences, where the syntax elements are determined by the content of the syntax elements found in the SPS, which are referenced by the syntax elements found in the PPS, which are referenced by the syntax elements found in each slice segment header.

[0102] A video parameter set RBSP may include parameters that may be referenced by one or more sequence parameter set RBSPs.

[0103] The relationship and hierarchy between video parameter sets (VPS), sequence parameter sets (SPS) and picture parameter sets (PPS) can be described as follows. The VPS is located one level above the SPS in the parameter set hierarchy and in the context of scalability and / or 3D video. The VPS may include parameters common to all slices across all (scalability or view) layers in the entire coded video sequence. The SPS includes parameters common to all slices in a specific (scalability or view) layer in the entire coded video sequence, and may be shared by multiple (scalability or view) layers. The PPS includes parameters common to all slices in a specific layer representation (a representation of a scalability or view layer in an access unit) and may be shared by all slices in multiple layer representations.

[0104] The VPS can provide information about the dependency relationships of the layers in the bitstream, as well as a lot of other information that can apply to all slices on all (scalability or view) layers in the entire coded video sequence. The VPS can be considered to consist of two parts, a base VPS and a VPS extension, where the VPS extension can optionally be present.

[0105] Out-of-band transmission, signaling or storage may additionally or alternatively be used for other purposes besides tolerating transmission errors, such as ease of access or session negotiation. For example, a sample entry for a track in a file conforming to the ISO base media file format may include a parameter set, while the coded data in the bitstream is stored elsewhere in the file or in another file. Phrases along a bitstream (e.g., indicated along a bitstream) or along a coded unit of a bitstream (e.g., indicated along a coded tile) may be used in the claims and embodiments to refer to out-of-band transmission, signaling or storage in a manner such that the out-of-band data is associated with the bitstream or the coded unit, respectively. Phrases such as decoding along a bitstream or along a coding unit of a bitstream may refer to decoding the mentioned out-of-band data (which may be obtained from an out-of-band transmission, signaling or storage device) associated with the bitstream or coding unit, respectively.

[0106] A SEI NAL unit may contain one or more SEI messages, which are not necessary for decoding of output pictures, but may assist with related processes such as picture output timing, rendering, error detection, error concealment, and resource conservation.

[0107] A coded picture is an encoded representation of a picture.

[0108] In HEVC, a coded picture can be defined as the coded representation of a picture including all coding tree units of the picture. In HEVC, an access unit (AU) can be defined as a set of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and contain at most one picture with any specific value of nuh_layer_id. In addition to VCL NAL units including coded pictures, access units can also include non-VCL NAL units. The specified classification rule can, for example, associate pictures with the same output time or picture output count value to the same access unit.

[0109] A bitstream can be defined as a sequence of bits in the form of a stream of NAL units or a stream of bytes, which forms a representation of a coded picture and related data forming one or more coded video sequences. The first bitstream may be followed by a second bitstream in the same logical channel, for example in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) can be defined as a sequence of one or more bitstreams. The end of the first bitstream can be indicated by a specific NAL unit, which can be called an end-of-bitstream (EOB) NAL unit and is the last NAL unit of the bitstream. In HEVC and its current draft extensions, the EOB NAL unit is required to have nuh_layer_id equal to 0.

[0110] In H.264 / AVC, a coded video sequence is defined as a contiguous sequence of access units in decoding order from an IDR access unit (inclusive) to the next IDR access unit (exclusive) or to the end of the bitstream, whichever comes first.

[0111] In HEVC, a coded video sequence (CVS) may be defined as, for example, a sequence of access units consisting, in decoding order, of an IRAP access unit with NoRaslOutputFlag equal to 1, followed by 0 or more access units that are not IRAP access units with NoRaslOutputFlag equal to 1, including all subsequent access units, up to but not including any subsequent access unit that is an IRAP access unit with NoRaslOutputFlag equal to 1. An IRAP access unit may be defined as an access unit in which a base layer picture is an IRAP picture. For each IDR picture, each BLA picture, and each IRAP picture (the first picture in that particular layer in the bitstream in decoding order), the value of NoRaslOutputFlag is equal to 1, the first IRAP picture in decoding order after the end NAL unit of the sequence, which has the same nuh_layer_id value. There may be a component that provides the value of HandleCraAsBlaFlag to the decoder from an external entity, such as a player or receiver, that may control the decoder. For example, HandleCraAsBlaFlag may be set to 1 by a player that seeks to a new position in the bitstream or tunes to a broadcast and starts decoding and then starts decoding from a CRA picture. When HandleCraAsBlaFlag is equal to 1 for a CRA picture, the CRA picture is processed and decoded as if it were a BLA picture.

[0112] In HEVC, a coded video sequence may additionally or alternatively (according to the above specification) be designated as ending when a specific NAL unit, called the End of Sequence (EOS) NAL unit, is present in the bitstream and has nuh_layer_id equal to 0.

[0113] A group of pictures (GOP) and its characteristics can be defined as follows. A GOP can be decoded regardless of whether the previous picture is decoded. An open GOP is a group of pictures in which the pictures before the initial intra picture in the output order may not be correctly decoded when decoding starts from the initial intra picture of the open GOP. In other words, the pictures of the open GOP may refer to pictures belonging to the previous GOP (in inter prediction). The HEVC decoder can identify the intra picture that starts the open GOP because a specific NAL unit type, the CRA NAL unit type, may be used for its coded slice. A closed GOP is a group of pictures in which all pictures can be correctly decoded when decoding starts from the initial intra picture of the closed GOP. In other words, no picture in the closed GOP refers to a picture in the previous GOP. In H.264 / AVC and HEVC, a closed GOP can start with an IDR picture. In HEVC, a closed GOP can also start with a BLA_W_RADL or BLA_N_LP picture. Due to the greater flexibility in selecting reference pictures, an open GOP coding architecture may be more efficient in compression than a closed GOP coding structure.

[0114] The decoded picture buffer (DPB) can be used in the encoder and / or decoder. There are two reasons for buffering decoded pictures for reference in inter-frame prediction and for reordering decoded pictures to output order. Since H.264 / AVC and HEVC provide a lot of flexibility for both reference picture marking and output reordering, separate buffers for reference picture buffering and output picture buffering may waste memory resources. Therefore, the DPB can include a unified decoded picture buffering process for reference pictures and output reordering. When the decoded picture is no longer used as a reference and does not need to be output, the decoded picture can be removed from the DPB.

[0115] In multiple coding modes of H.264 / AVC and HEVC, reference pictures used for inter prediction are indicated by indices to reference picture lists. The indices may be encoded using variable length coding, which generally results in smaller indices having shorter values ​​for the corresponding syntax elements. In H.264 / AVC and HEVC, two reference picture lists (reference picture list 0 and reference picture list 1) are generated for each bidirectionally predicted (B) slice, and one reference picture list (reference picture list 0) is formed for each inter-coded (P) slice.

[0116] Multiple coding standards, including H.264 / AVC and HEVC, may have a decoding process to derive a reference picture index to a reference picture list, which may be used to indicate which of multiple reference pictures is used for inter-frame prediction of a particular block. The reference picture index may be encoded into the bitstream by the encoder in some inter-frame coding modes, or may be derived (by the encoder and decoder) using neighboring blocks, for example in some other inter-frame coding modes.

[0117] The motion parameter type or motion information may include but is not limited to one or more of the following types:

[0118] - an indication of the prediction type (e.g. intra prediction, uni prediction, bi prediction) and / or the number of reference pictures;

[0119] - an indication of a prediction direction, such as inter (also called temporal) prediction, inter-layer prediction, inter-view prediction, view synthesis prediction (VSP) and inter-component prediction (which may be indicated per reference picture and / or per prediction type, and in some embodiments inter-view and view synthesis prediction may be jointly considered as one prediction direction) and / or

[0120] - an indication of the reference picture type, such as short-term reference picture and / or long-term reference picture and / or

[0121] or inter-layer reference pictures (which may be indicated, for example, per reference picture)

[0122] - a reference index to a reference picture list and / or any other identifier of a reference picture (which may be indicated, for example, per reference picture, whose type may depend on the prediction direction and / or the referenced picture type, and may be accompanied by other relevant information, such as the reference picture list to which the reference index applies, etc.);

[0123] - horizontal motion vector component (which may be indicated, for example, by prediction block or by reference index, etc.);

[0124] - vertical motion vector component (which may be indicated, for example, by prediction block or by reference index, etc.);

[0125] - one or more parameters, such as a picture order count difference and / or a relative camera spacing between a picture including or associated with a motion parameter and its reference picture, which may be used for scaling a horizontal motion vector component and / or a vertical motion vector component in one or more motion vector prediction processes (wherein the one or more parameters may be indicated, for example, per reference picture or per reference index, etc.);

[0126] - the coordinates of the block to which the motion parameters and / or motion information apply, e.g. the coordinates of the top left sample of the block in units of luma samples;

[0127] - The extent (eg width and height) of the block to which the motion parameters and / or motion information apply.

[0128] Compared with previous video coding standards, the General Video Codec (H.26 / VVC) introduces a variety of new coding tools, such as:

[0129] Intra-frame prediction

[0130] - 67 fps intra-frame mode with wide-angle mode expansion

[0131] - Block size and mode dependent 4-tap interpolation filter

[0132] -Position Dependent Intra Prediction Combination (PDPC)

[0133] -Cross-Component Linear Model Intra Prediction (CCLM)

[0134] -Multiple reference line intra prediction

[0135] - Intra-frame sub-partition

[0136] -Weighted intra prediction with matrix multiplication

[0137] Inter-picture prediction

[0138] - Block motion replication using spatial, temporal, history-based and pairwise average merging and candidate

[0139] -Affine motion inter-frame prediction

[0140] - Sub-block based temporal motion vector prediction

[0141] -Adaptive motion vector solution

[0142] -8x8 block-based motion compression for temporal motion prediction

[0143] - High-precision (1 / 16 pixel) motion vector storage and motion compensation, 8-tap interpolation filter for luminance component and 4-tap interpolation filter for chrominance component

[0144] -Triangulation

[0145] - Combined intra and inter prediction

[0146] -Merged with MVD (MMVD)

[0147] -Symmetrical MVD encoding

[0148] - Bidirectional optical flow

[0149] -Decoder side motion vector refinement

[0150] - Use CU level weights for dual prediction

[0151] Transform, quantization and coefficient coding

[0152] - Multiple primary transform options with DCT2, DST7 and DCT8

[0153] -Secondary transformation in low frequency area

[0154] -Sub-block transform of inter prediction residual

[0155] - Increased max QP from 51 to 63 for quantization dependent

[0156] - Transform coefficient coding with sign data hiding

[0157] -Transform skip residual coding

[0158] Entropy coding

[0159] -Arithmetic coding engine with adaptive dual-window probability update

[0160] Loop filter

[0161] - Cycle shaping

[0162] - Deblocking filter with strong longer filter

[0163] - Sample Adaptive Offset

[0164] - Adaptive loop filter

[0165] Screen content encoding:

[0166] - Current picture reference with reference area restrictions

[0167] 360-degree video encoding

[0168] -Horizontal surround motion compensation

[0169] Advanced syntax and parallel processing

[0170] - Reference picture management using direct reference picture list signaling

[0171] - Tile groups with rectangular tile groups

[0172] Although the new coding tools listed above lack decoder-side intra mode derivation (DIMD), it has been considered for adoption in the VVC / H.266 video codec. DIMD techniques have been shown to have a beneficial effect on prior art video codecs. These methods typically rely on an inference process that operates on a support region formed by reconstructed samples around the current block. Gradient estimation techniques can be used to predict the directionality and strength of edges in the support region. These are then used to infer the intra prediction direction and ultimately derive the directional intra prediction mode. These modes are used to predict the current block. The use of DIMD is typically signaled for a given block. When DIMD is signaled to be used, no other information is needed to perform intra prediction of the current block, but rather the intra prediction mode is inferred.

[0173] However, while DIMD can successfully reduce the overhead required to signal a given intra-prediction mode, the resulting prediction is often suboptimal.

[0174] An improved method for performing decoder-side intra-mode derivation is now presented.

[0175] Figure 5 , wherein the method comprises decoding (500) encoded samples of a block of video sample data; determining (502) two or more regions of support for the block of video sample data, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; performing (504) a decoder-side intra mode derivation process for each region of support to obtain decoder-side intra mode derivation parameters specific to each region of support; and

[0176] A prediction of samples of the block of video sample data is computed (506) based on decoder-side intra-mode derivation parameters specific to each region of support.

[0177] Thus, the method generates intra-prediction for a given block, wherein at least two or more decoder-side intra mode derivation (DIMD) processes are performed, wherein each process operates on reconstructed samples extracted from a specific support region formed by a set of samples at specific positions relative to the current block, wherein each process outputs DIMD parameters specific to the support region.

[0178] Thus, at least two support regions around the current block are identified. The support regions include previously reconstructed decoded samples (or pixels). Each support region may be formed by samples at a specific position relative to the current block. The support regions may overlap each other.

[0179] Figure 6An example of a support region is given in which three support regions for a current coding unit (CU) are identified: support region 1 corresponds to samples directly above the current block, support region 3 corresponds to samples directly to the left of the current block, and support region 2 corresponds to samples in the upper left region of the current block. In this example, a first support region (support region 1) may be formed by a three-pixel high region of reconstructed samples directly above the current block; a second support region (support region 2) may be formed by a 3×3 pixel region of reconstructed samples located to the upper left of the current block; and a third support region (support region 3) may be formed by a three-pixel wide region of reconstructed samples located directly to the left of the current block.

[0180] For each supported region, a decoder-side intra mode derivation (DIMD) process is performed to derive specific DIMD parameters for each region.

[0181] According to an embodiment, the method comprises performing, as part of a decoder-side intra-mode derivation process, an analysis of the directionality of samples belonging to at least one region of support.

[0182] Therefore, an analysis of the directionality of pixels belonging to a specific support region may be performed as part of the DIMD processing and thereafter specific DIMD parameters for each region derived.

[0183] According to an embodiment, the method comprises deriving, as part of a decoder-side intra-mode derivation process, a gradient histogram of samples of at least one region of support.

[0184] As an example of directionality analysis, 3×3 Sobel gradient filters (horizontal and vertical) can be used. They are convolved with the samples in the support region to obtain a given gradient histogram.

[0185] According to an embodiment, the method comprises obtaining one or more intra prediction modes from a frame intra mode derivation process at the decoder side.

[0186] The histogram has a number of bins corresponding to the number of possible intra prediction modes M. The magnitude of a given bin in the histogram at a particular index m=0...M-1 represents the cumulative magnitude of gradients estimated to have the same direction as intra prediction mode m. Figure 7 In the example shown, the shaded (grey) area shows the pixels used as the center of the sliding window for support region 1, where the dashed line shows the 3×3 sliding window of samples convolved with the Sobel kernel.

[0187] According to an embodiment, the method comprises inferring whether one or more intra prediction modes are position-dependent on a specific region of support based on a decoder-side intra mode derivation parameter specific to each region of support.

[0188] Thus, the DIMD parameters obtained for each support region are used to determine whether a given derived DIMD pattern is position-dependent, and if so, which position it is dependent on. As an example, a given DIMD pattern that can be obtained using any conventional method can be classified as position-dependent, and its position can be identified by considering and possibly comparing the DIMD parameters output from different support regions.

[0189] For example, we can consider two support regions, one formed by the samples directly above the current block and the other formed by the samples directly to the left of the current block, respectively Figure 6 The support region 1 and the support region 3 in . Two gradient histograms can be obtained, H above Specific to the support area above, H left is specific to the left support region. For a given DIMD pattern m, the amplitudes of the two histograms acquired in the two support regions can be used to determine whether the pattern m is position dependent. For example, if H left (m)==0 and H above (m)≠0, it can be considered that the position of mode m depends on the indication of the support region mentioned above. On the contrary, if H left (m)≠0 and H above (m) == 0, then m can be classified as position dependent on the left support region.

[0190] This can be summarized as follows: If we consider N support regions, we get N histograms H0, H1, ... H N-1 , then a given mode m can be determined to be dependent on the location of region i if: H i (m)≠0, and H j (m)==0, j=0, 1, ..., N-1 and j≠i

[0191] The height of a bar of a particular histogram of the upper and left support regions can also be used to determine the positional dependency of a given pattern m. For example, if the factor K∈

[01] is considered, then the dependence of pattern m on the position of region i can be determined:

[0192]

[0193] According to an embodiment, the method includes normalizing the DIMD parameter outputted by each support region according to the number of samples belonging to each support region. For example, considering N support regions, N histograms are obtained. And will s i Denote M as the number of pixels in each histogram and consider the cumulative number of pixels in all support regions.

[0194] Then the normalized histogram Can be obtained as:

[0195] According to an embodiment, the method comprises inferring the strength of positional dependency of one or more intra prediction modes over a specific region of support based on decoder-side intra mode derivation parameters specific to each region of support.

[0196] Additionally or alternatively, in order to determine whether a given DIMD pattern is classified as position-dependent on a given support region, the strength of the position dependence of the given pattern on the given support region may be determined based on the DIMD parameters output for each support region. For example, considering N support regions, N histograms H0, H1, ... H N-1 , which may or may not be normalized. Consider a given mode m that is classified as position dependent on region i. The strength of the position dependence of mode m on support region i can be obtained from the calculation of the ratio as follows:

[0197]

[0198] According to an embodiment, the method comprises deriving a decoder-side intra-mode derivation mode specific to each support region based on decoder-side intra-mode derivation parameters specific to each support region.

[0199] According to this method, which can be used alone or in combination with other embodiments, the DIMD parameters obtained for each support region are used to determine a specific DIMD pattern for each support region. For example, a given DIMD pattern can be obtained using any conventional method, where the reconstructed pixels used to calculate the DIMD parameters are limited to pixels belonging to a specific support region. For example, considering N support regions, N histograms H0, H1, ... H N-1 , and consider M as the number of bins in each histogram, then for a given support region i, a specific DIMD pattern m can be obtained i As H i The mode with the highest peak, or: m i =max m=0,...,M-1 H i (m).

[0200] According to an embodiment, the method comprises computing a prediction for the current block based on a decoder-side intra mode derivation mode specific to each region of support.

[0201] A specific DIMD mode for each support region may be used to compute a prediction for the current block. As an example, an index signaling may be issued in the bitstream identifying a specific support region used to determine the DIMD mode for the current block. As another example, a flag signaling may be issued for each support region to identify whether to compute a prediction for the current block as a result of a hybrid determined specific DIMD mode.

[0202] This method can also be used in combination with other embodiments, for example, a specific DIMD mode m can be obtained for a given support area. i , and the strength of its position dependence on the support region i can be obtained as follows:

[0203]

[0204] According to an embodiment, which can be implemented independently of or in combination with other embodiments, the method comprises computing at least two predictors for a block of video sample data based on decoder-side intra mode derivation parameters specific to each region of support; and combining the at least two predictors together using blending to form a prediction for the block, wherein the blending is performed using sample based weights.

[0205] If implemented independently, Figure 8 The method shown in the flowchart of can be performed, wherein the method includes decoding (800) encoded samples of a block of video sample data; performing (802) a decoder-side intra-mode derivation process to obtain decoder-side intra-mode derivation parameters; obtaining (804) one or more intra-prediction modes from the decoder-side intra-mode derivation parameters; calculating (806) at least two predictors for the video block sample data based on the intra-prediction modes; and combining (808) at least the two predictors together using blending to form a prediction for the block, wherein the blending is performed using sample-based weights.

[0206] Thus, according to one approach, sample-based blending of multiple DIMD modes may be performed to obtain a prediction for a current block. Sample-based blending may operate based on determining specific weights for each predictor and each sample within the current block. The weights may be determined based on the position of each sample within the block, where different weights may be used to blend different samples within the block. However, blending does not necessarily require determining DIMD parameters specific to each region of support.

[0207] According to an embodiment, the method comprises determining two or more regions of support for a block of video sample data, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; and performing a decoder-side Intra mode derivation procedure to derive decoder-side Intra mode derivation parameters specific to each region of support.

[0208] Therefore, although it is not limited to using only these parameters to determine the intra prediction mode, DIMD parameters specific to each support region can still be derived. Instead, the parameters can be used to determine other things, such as the weights used to combine predictors, or to infer whether to use sample-based weighting at all or not.

[0209] According to an embodiment, the method includes obtaining the at least two predictors based on the two or more support regions; performing a decoder-side intra-frame mode derivation process to derive a decoder-side intra-frame mode derivation mode specific to each support region; and calculating at least two predictors based on the decoder-side intra-frame mode derivation mode specific to the two or more support regions.

[0210] For example, suppose Figure 6 As shown, three support regions are considered. Assume that two DIMD modes m0 and m1 are considered. These modes can be determined according to any method described in this article. Assume that two predictors P0 and P1 are obtained by performing intra-frame prediction processing according to modes m0 and m1 respectively. Then, the final predictor P of the current block can be obtained by the following formula:

[0211] P(x,y)=w1(x,y)P1(x,y)+w0(x,y)P0(x,y)

[0212] Where P i (x,y) refers to the predictor P i The pixel at position (x,y) in .

[0213] According to an embodiment, a method comprises inferring the use of sample-based weighting based on characteristics of a block of video data.

[0214] According to an embodiment, the sample-based weight depends on the position of the sample in the block.

[0215] According to an embodiment, the sample-based weights depend on the positional dependency of a given intra prediction mode over a specific support region.

[0216] Therefore, the use of sample-based mixtures can be inferred, for example, depending on whether the DIMD patterns depend on the position of a particular support region, and / or the strength of the position dependence of each DIMD pattern. i The weight of (x,y) can depend on the position (x,y). The weight wi(x,y) can also depend on m i Whether it is classified as being dependent on a particular support region. In the above example, assume that pattern m0 is classified as being dependent on the above support region, and m1 is classified as being dependent on the left support region, such as Fig. 9Assume a block of size H x W, where W is the width and H is the height. Then, the weight of the predictor can be determined for each sample as follows:

[0217]

[0218] w1(x,y)=1-w0(x,y)

[0219] We can also consider integer precision representation of weights. Assuming a 6-bit representation of weights, these can be defined as:

[0220]

[0221] w1(x, y) = 64 - w0(x, y)

[0222] It is also possible to use exemption operations to determine weights, such as by scaling and shifting.

[0223] The weights may also depend on the strength of the positional dependence of a given pattern on a given support region. In the above example, the strength of the positional dependence of pattern m0 from the above support region is denoted as F above (m0), and similarly denote the position-dependent strength of m1 from the left support region as F left (m1). Then, we can use F above (m0) and / or F left (m1) to determine the weight w i (x,y). For example, two intermediate parameters can be derived, called Δ x and Δ y . These parameters can be derived as:

[0224]

[0225] Other ways of deriving intermediate parameters can be used. The intermediate parameters can then be used to calculate weights, such as:

[0226]

[0227] w1(x, y) = 1 - w0(x, y)

[0228] The weights may be clipped within a predetermined range. For example, the weights may be clipped between 0 and 1.

[0229] The weights may also be pre-computed and stored, for example, in a lookup table, for reuse during the decoding process.

[0230] According to an embodiment, the method comprises combining one or more intra prediction modes derived using a decoder-side intra mode derivation process in combination with one or more predetermined intra prediction modes to compute a prediction for a current block.

[0231] Therefore, additional predetermined predictors may also be considered and used. For example, in addition to the two DIMD-derived predictors P0 and P1, an additional predictor P obtained by performing a predetermined planar mode on the current block may also be considered. pln When using these pre-derived predictors, the two DIMD derived predictors and the additional predictors may be mixed using appropriate weights. These weights may be derived based on the positional dependency of each mode on a given support region and / or the strength of such positional dependency and / or the DIMD parameters output during the DIMD process.

[0232] According to an embodiment, the method comprises computing a prediction of a current block using sample-based weights, wherein the weights for one or more predetermined intra prediction modes depend on decoder-side intra mode derivation parameters specific to each region of support.

[0233] Thus, if the combined DIMD uses a predetermined (eg, planar) mode, sample-based weights may be used, where the weights used on the planar mode also depend on the DIMD parameters.

[0234] An apparatus according to an aspect comprises means for decoding encoded samples of a block of video sample data; means for determining two or more regions of support for the block of video sample data, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; means for performing a decoder-side intra mode derivation process for each region of support to obtain decoder-side intra mode derivation parameters specific to each region of support; and means for calculating a prediction of samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each region of support.

[0235] According to an embodiment, the apparatus comprises means for performing a directionality analysis of samples belonging to at least one region of support as part of an intra mode derivation process at a decoder side.

[0236] According to an embodiment, an apparatus comprises means for deriving one or more intra prediction modes from a decoder-side intra mode derivation process.

[0237] According to an embodiment, the apparatus comprises means for inferring whether one or more intra prediction modes depend on the position of a specific region of support based on decoder-side intra mode derivation parameters specific to each region of support.

[0238] According to an embodiment, the apparatus comprises means for inferring the strength of positional dependency of one or more intra prediction modes over a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.

[0239] According to an embodiment, the apparatus comprises means for deriving a gradient histogram of samples of at least one region of support as part of a decoder-side intra mode derivation process.

[0240] According to an embodiment, the apparatus comprises means for deriving a decoder-side intra-mode derivation mode specific to each support region based on decoder-side intra-mode derivation parameters specific to each support region.

[0241] According to an embodiment, the apparatus comprises means for computing a prediction of a current block based on a decoder-side intra mode derivation mode specific to each region of support.

[0242] The apparatus according to the second aspect comprises means for decoding coded samples of a block of video sample data; means for performing a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters; means for deriving one or more intra prediction modes from the decoder-side intra mode derivation parameters; means for computing at least two predictors for the block of video sample data based on the decoder-side intra mode derivation parameters specific to each support region; and means for combining at least the two predictors together using blending to form a prediction for the block, wherein the blending is performed using sample based weights.

[0243] According to an embodiment, the apparatus comprises means for determining two or more regions of support for a block of video sample data, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data, and means for performing a decoder-side Intra mode derivation procedure to derive decoder-side Intra mode derivation parameters specific to each region of support.

[0244] According to an embodiment, the apparatus comprises means for inferring whether one or more intra prediction modes depend on the position of a specific region of support based on decoder-side intra mode derivation parameters specific to each region of support.

[0245] According to an embodiment, the sample-based weight depends on the position of the sample in the block.

[0246] According to an embodiment, the sample-based weights depend on the positional dependency of a given intra prediction mode over a specific support region.

[0247] According to an embodiment, an apparatus comprises means for inferring the use of sample-based weightings based on features of a block of video data.

[0248] According to an embodiment, the apparatus comprises means for inferring whether one or more intra prediction modes depend on the position of a specific region of support based on decoder-side intra mode derivation parameters specific to each region of support.

[0249] According to an embodiment, the sample-based weighting is inferred depending on the strength of the position dependence of the decoder-side intra-mode derivation mode over a certain support region.

[0250] According to an embodiment, the apparatus comprises means for combining one or more intra prediction modes derived using a decoder-side intra mode derivation process with computing a prediction for a current block using one or more predetermined intra prediction modes.

[0251] According to an embodiment, the apparatus comprises means for computing a prediction for a current block using sample-based weights, wherein the weights for one or more predetermined intra prediction modes depend on decoder-side intra mode derivation parameters specific to each region of support.

[0252] As a further aspect, there is provided an apparatus comprising: at least one processor and at least one memory having code stored thereon, the code when executed by the at least one processor causing the apparatus to at least perform: decoding coded samples of a block of video sample data; determining two or more support regions for the block of video sample data, wherein each support region comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; performing a decoder-side intra mode derivation process for each support region to obtain decoder-side intra mode derivation parameters specific to each support region; and calculating a prediction for samples of the block of video sample data based on the decoder-side intra mode derivation parameters specific to each support region.

[0253] According to an embodiment, an apparatus comprises code configured to cause the apparatus to perform an analysis of directionality of samples belonging to at least one region of support as part of a decoder-side intra-mode derivation process.

[0254] According to an embodiment, the apparatus comprises code configured to cause the apparatus to obtain one or more intra prediction modes from a decoder-side intra mode derivation process.

[0255] According to an embodiment, the apparatus comprises code configured to cause the apparatus to infer whether one or more intra prediction modes are position-dependent on a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.

[0256] According to an embodiment, the apparatus comprises code configured to cause the apparatus to infer strength of one or more intra-prediction mode position dependencies over a particular region of support based on decoder-side intra-mode derivation parameters specific to each region of support.

[0257] According to an embodiment, the apparatus comprises code configured to cause the apparatus to derive a gradient histogram of samples of at least one region of support as part of a decoder-side intra-mode derivation process.

[0258] According to an embodiment, the apparatus comprises code configured to cause the apparatus to derive a decoder-side intra-mode derivation mode specific to each support region based on decoder-side intra-mode derivation parameters specific to each support region.

[0259] According to an embodiment, the apparatus comprises code configured to cause the apparatus to compute a prediction for a current block based on a decoder-side intra mode derivation mode specific to each region of support.

[0260] The apparatus according to the fourth aspect comprises at least one processor and at least one memory having code stored thereon, which when executed by the at least one processor causes the apparatus to at least: decode coded samples of a block of video sample data; perform a decoder-side intra mode derivation process to obtain decoder-side intra mode derivation parameters; obtain one or more intra prediction modes from the decoder-side intra mode derivation parameters; calculate at least two predictors for the block of video sample data based on the decoder-side intra mode derivation parameters specific to each support region; and combine at least the two predictors together using blending to form a prediction of the block, wherein the blending is performed using sample-based weights.

[0261] According to an embodiment, the apparatus comprises code configured to cause the apparatus to determine, for a block of video sample data, two or more regions of support, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the video sample data; and code configured to cause the apparatus to perform a decoder-side intra-mode derivation process to derive decoder-side intra-mode derivation parameters specific to each region of support.

[0262] According to an embodiment, the apparatus comprises code configured to cause the apparatus to infer whether one or more intra prediction modes are position-dependent on a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.

[0263] According to an embodiment, the sample-based weights depend on the positions of the samples in the block.

[0264] According to an embodiment, the sample-based weights depend on the positional dependency of a given intra prediction mode over a specific support region.

[0265] According to an embodiment, the apparatus comprises code configured to cause the apparatus to infer usage of sample-based weighting based on features of a block of video data.

[0266] According to an embodiment, the apparatus comprises code configured to cause the apparatus to infer whether one or more intra prediction modes are position-dependent on a particular region of support based on decoder-side intra mode derivation parameters specific to each region of support.

[0267] According to an embodiment, the sample-based weighting is inferred from the strength of the positional dependency of the decoder-side intra-mode derivation mode over a specific support region.

[0268] According to an embodiment, the apparatus comprises code configured to cause the apparatus to combine one or more intra prediction modes derived using a decoder-side intra mode derivation process with one or more predetermined intra prediction modes to calculate a prediction for a current block.

[0269] According to an embodiment, the apparatus comprises code configured to cause the apparatus to compute a prediction of a current block using sample-based weights, wherein the weights for one or more predetermined intra-prediction modes depend on decoder-side intra-mode derivation parameters specific to each support region.

[0270] Such means may include, for example Figure 1 , Figure 2 , Figure 4a and Figure 4b The functional units disclosed in the for implementing the embodiments.

[0271] Such an apparatus further comprises code stored in the at least one memory, which, when executed by the at least another processor, causes the apparatus to perform one or more embodiments disclosed herein.

[0272] Fig.1015 is a graphical representation of an example multimedia communication system in which various embodiments can be implemented. Data source 1510 provides a source signal in analog, uncompressed digital or compressed digital format or any combination of these formats. Encoder 1520 may include or be connected to preprocessing, such as data format conversion and / or filtering of source signal. Encoder 1520 encodes source signal into coded media bit stream. It should be noted that the bit stream to be decoded can be received directly or indirectly from a remote device located in almost any type of network. Additionally, the bit stream can be received from local hardware or software. Encoder 1520 can encode more than one media type, such as audio and video, or more than one encoder 1520 may be needed to encode different media types of source signals. Encoder 1520 can also obtain synthetically generated input, such as graphics and text, or it can produce a coded bit stream of synthetic media. In the following, only the processing of a coded media bit stream of a media type is considered to simplify the description. However, it should be noted that usually real-time broadcast services include several streams (usually at least one audio stream, video stream and text subtitle stream). It should also be noted that the system may include many encoders, but only one encoder 1520 is shown in the figure to simplify the description without lack of generality. It should also be understood that although the text and examples included herein may specifically describe the encoding process, those skilled in the art will understand that the same concepts and principles also apply to the corresponding decoding process, and vice versa.

[0273] The coded media bitstream may be transmitted to a storage device 1530. The storage device 1530 may include any type of mass storage to store the coded media bitstream. The format of the coded media bitstream in the storage device 1530 may be a basic self-contained bitstream format, or one or more coded media bitstreams may be encapsulated into a container file, or the coded media bitstream may be encapsulated into a segmented format suitable for DASH (or a similar streaming system) and stored as a segmented sequence. If one or more media bitstreams are encapsulated in a container file, a file generator (not shown in the figure) may be used to store the one or more media bitstreams in a file and create file format metadata, which may also be stored in the file. The encoder 1520 or the storage device 1530 may include a file generator, or the file generator may be operatively connected to the encoder 1520 or the storage device 1530. Some systems operate in "real time", i.e., storage is omitted and the coded media bitstream is transmitted directly from the encoder 1520 to the transmitter 1540. The coded media bitstream may then be transmitted to the transmitter 1540, also referred to as a server, as needed. The format used in the transmission may be a basic self-contained bitstream format, a packetized stream format, a segmented format suitable for DASH (or a similar streaming system), or one or more coded media bitstreams may be encapsulated into a container file. The encoder 1520, storage device 1530, and server 1540 may reside in the same physical device, or they may be included in separate devices. The encoder 1520 and server 1540 may operate with live real-time content, in which case the coded media bitstreams are typically not stored permanently, but are buffered for a short period of time in the content encoder 1520 and / or server 1540 to smooth out variations in processing delays, transmission delays, and coded media bitrates.

[0274] The server 1540 sends the coded media bitstream using a communication protocol stack. The stack may include, but is not limited to, one or more of a real-time transport protocol (RTP), a user datagram protocol (UDP), a hypertext transfer protocol (HTTP), a transmission control protocol (TCP), and an Internet protocol (IP). When the communication protocol stack is packet-oriented, the server 1540 encapsulates the coded media bitstream into packets. For example, when using RTP, the server 1540 encapsulates the coded media bitstream into RTP packets according to the RTP payload format. Typically, each media type has a dedicated RTP payload format. It should be noted again that the system may include multiple servers 1540, but for simplicity, the following description only considers one server 1540.

[0275] If the media content is encapsulated in a container file for storage device 1530 or for inputting data into transmitter 1540, transmitter 1540 may include or may be operably attached to a "sending file parser" (not shown in the figure). In particular, if the container file is not sent as such, but at least one of the included coded media bitstreams is encapsulated for transmission via a communication protocol, the sending file parser locates the appropriate portion of the coded media bitstream for transmission via the communication protocol. The sending file parser can also help create the correct format for the communication protocol, such as packet headers and payloads. The multimedia container file may include encapsulation instructions, such as the recommended track in ISOBMFF, for encapsulating at least one of the media bitstreams included on the communication protocol.

[0276] The server 1540 may or may not be connected to the gateway 1550 through a communication network, which may be, for example, a combination of a CDN, the Internet, and / or one or more access networks. The gateway may also or alternatively be referred to as a middlebox. For DASH, the gateway may be an edge server (of a CDN) or a web proxy. It should be noted that the system may generally include any number of gateways or the like, but for simplicity, the following description only considers one gateway 1550. The gateway 1550 may perform different types of functions, such as converting a packet stream according to one communication protocol stack to another communication protocol stack, merging and forking data streams, and manipulating data streams according to downlink and / or receiver capabilities, such as controlling the bit rate of the forwarded stream according to the prevailing downlink network conditions. In various embodiments, the gateway 1550 may be a server entity.

[0277] The system includes one or more receivers 1560, which are generally capable of receiving, demodulating and decapsulating the transmitted signal into a coded media bitstream. The coded media bitstream can be transmitted to a recording storage device 1570. The recording storage device 1570 can include any type of mass storage to store the coded media bitstream. The recording storage device 1570 can alternatively or additionally include a computing memory, such as a random access memory. The format of the coded media bitstream in the recording storage device 1570 can be a basic self-contained bitstream format, or one or more coded media bitstreams can be encapsulated into a container file. If there are multiple coded media bitstreams associated with each other, such as audio streams and video streams, a container file is generally used, and the receiver 1560 includes or is attached to a container file generator that generates the container file from the input stream. Some systems operate in "real time", that is, the recording storage device 1570 is omitted and the coded media bitstream is transmitted directly from the receiver 1560 to the decoder 1580. In some systems, only the most recent portion of the recorded stream, such as an excerpt of the most recent 10 minutes of the recorded stream, is retained in the recording storage device 1570, and any earlier recorded data is discarded from the recording storage device 1570.

[0278] The coded media bitstream may be transferred from the recording storage device 1570 to the decoder 1580. If there are many coded media bitstreams, such as audio streams and video streams, which are associated with each other and encapsulated into a container file, or a single media bitstream is encapsulated in a container file, for example, for easier access, a file parser (not shown in the figure) is used to decapsulate each coded media bitstream from the container file. The recording storage device 1570 or the decoder 1580 may include the file parser, or the file parser is attached to the recording storage device 1570 or the decoder 1580. It should also be noted that the system may include many decoders, but only one decoder 1570 is discussed here to simplify the description without loss of generality.

[0279] The coded media bitstream may also be processed by a decoder 1570, the output of which is one or more uncompressed media streams. Finally, a renderer 1590 may reproduce the uncompressed media streams using, for example, a speaker or a display. The receiver 1560, the recording storage 1570, the decoder 1570, and the renderer 1590 may reside in the same physical device, or they may be included in separate devices.

[0280] The transmitter 1540 and / or the gateway 1550 may be configured to perform switching between different representations, such as switching between different viewports for 360-degree video content, view switching, bitrate adaptation and / or fast start, and / or the transmitter 1540 and / or the gateway 1550 may be configured to select the representation sent. The switching between different representations may occur for a variety of reasons, such as in response to a request from the receiver 1560 or the prevailing conditions of the network through which the bitstream is transmitted, such as throughput. In other words, the receiver 1560 may initiate switching between representations. The request from the receiver may be, for example, a request for a segment or sub-segment from a different representation than previously, a request for a change in the scalability layer and / or sub-layer sent, or a request for a change in a rendering device with different capabilities than previously. The request for the segment may be an HTTP GET request. The request for the sub-segment may be an HTTP GET request with a byte range. Additionally or alternatively, bitrate adjustment or bitrate adaptation may be used, for example, to provide so-called fast start in streaming services, where after starting or randomly accessing a stream, the bitrate of the stream sent is lower than the channel bitrate in order to start playback immediately and achieve a buffer occupancy level that tolerates occasional packet delays and / or retransmissions. The bitrate adaptation may include multiple representation or layer up-switching and representation or layer down-switching operations occurring in various orders.

[0281] Decoder 1580 may be configured to perform switching between different representations, such as switching between different viewports of 360-degree video content, view switching, bit rate adaptation and / or quick start, and / or decoder 1580 may be configured to select the (multiple) representations sent. Switching between different representations may occur for a variety of reasons, such as to achieve faster decoding operations or to adapt the transmitted bitstream (e.g., according to the bitrate) to the main conditions (e.g., throughput) of the network that transmits the bitstream. For example, if the device including decoder 1580 is multitasking and uses computing resources for other purposes other than decoding video bitstreams, faster decoding operations may be required. In another example, when replaying content at a speed faster than normal playback speed (e.g., twice or three times faster than conventional real-time playback speed), faster decoding operations may be required.

[0282] In the above, some embodiments have been described with reference to and / or using the terminology of HEVC and / or VVC. It should be understood that the embodiments can be similarly implemented with any video encoder and / or video decoder.

[0283] Where example embodiments have been described above with reference to an encoder, it will be appreciated that the resulting bitstream and decoder may have corresponding elements therein. Likewise, where example embodiments have been described with reference to a decoder, it will be appreciated that the encoder may have a structure and / or computer program for generating a bitstream to be decoded by a decoder. For example, some embodiments have been described relating to generating prediction blocks as part of encoding. Embodiments may be similarly implemented by generating prediction blocks as part of decoding, except that encoding parameters, such as horizontal offset and vertical offset, are decoded from the bitstream rather than being determined by the encoder.

[0284] The embodiments of the present invention described above are described in terms of separate encoder and decoder devices to help understand the processes involved. However, it should be appreciated that the apparatus, structure, and operation may be implemented as a single encoder-decoder apparatus / structure / operation. In addition, the encoder and decoder may share some or all common elements.

[0285] Although the above examples describe embodiments of the present invention operating within a codec within an electronic device, it should be understood that the present invention as defined in the claims may be implemented as part of any video codec. Thus, for example, embodiments of the present invention may be implemented in a video codec that may implement video encoding over a fixed or wired communication path.

[0286] Hence, the user equipment may comprise a video codec such as described above in embodiments of the invention.It will be appreciated that the term user equipment is intended to cover any suitable type of wireless user equipment, such as a mobile phone, a portable data processing device or a portable web browser.

[0287] Furthermore, elements of a public land mobile network (PLMN) may also include a video codec as described above.

[0288] Generally, various embodiments of the present invention may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although various aspects of the present invention may be shown and described as block diagrams, flow charts, or using some other graphical representations, it is well understood that these block diagrams, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.

[0289] Embodiments of the present invention may be implemented by computer software that can be executed by a data processor of a mobile device, for example in a processor entity, or by hardware, or by a combination of software and hardware. In addition, in this regard, it should be noted that any block of the logic flow shown in the figure can represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on a physical medium such as a memory chip or a memory block implemented in a processor, on a magnetic medium such as a hard disk or a floppy disk, and on an optical medium such as a DVD and its data variants, a CD.

[0290] The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. As non-limiting examples, the data processor may be of any type suitable for the local technical environment and may include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and a processor based on a multi-core processor architecture.

[0291] Embodiments of the present invention may be implemented in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0292] Programs such as those offered by Synopsys Inc. of Mountain View, Calif., and Cadence Design of San Jose, Calif., automatically route conductors and position components on a semiconductor chip using generally accepted design rules and a library of pre-stored design modules. Once the design of a semiconductor circuit is complete, the final design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be transmitted to a semiconductor manufacturing facility or "fab" for fabrication.

[0293] The foregoing description has provided a complete and informative description of exemplary embodiments of the present invention by way of exemplary and non-limiting examples. However, various modifications and adaptations may be apparent to those skilled in the relevant art in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the present teachings will still fall within the scope of the present invention.

Claims

1. A device comprising: means for decoding encoded samples of a block of video sample data; means for determining two or more regions of support for a block of video sample data, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; means for performing a decoder-side intra-mode derivation process for each support region to obtain decoder-side intra-mode derivation parameters specific to each support region; as well as Means for computing a prediction of said samples of said block of video sample data based on said decoder-side intra mode derivation parameters specific to each region of support.

2. The device according to claim 1, comprising: Means for performing a directionality analysis of said samples belonging to at least one region of support as part of an intra mode derivation process at a decoder side.

3. The device according to claim 1 or 2, comprising: Means for deriving one or more intra prediction modes from said decoder-side intra mode derivation process.

4. The device according to claim 3, comprising: means for inferring whether said one or more intra prediction modes are position-dependent on a particular region of support based on said decoder-side intra mode derivation parameters specific to each region of support.

5. The device according to claim 3 or claim 4, comprising: means for inferring a strength of positional dependency of the one or more intra-prediction modes on each region of support based on the decoder-side intra-mode derivation parameters specific to the particular region of support.

6. The device according to any one of the preceding claims, comprising: Means for deriving a gradient histogram of said samples of at least one region of support as part of said decoder side intra mode derivation process.

7. The device according to any one of the preceding claims, comprising: Means for deriving a decoder side intra-mode derivation mode specific to each region of support based on said decoder side intra-mode derivation parameters specific to each region of support.

8. The device according to claim 7, comprising: Means for computing a prediction for a current block based on said decoder-side intra mode derivation mode specific to each region of support.

9. An apparatus comprising: means for decoding encoded samples of a block of video sample data; A component for performing a decoder-side intra-frame mode derivation process to obtain a decoder-side intra-frame mode derivation parameter; means for obtaining one or more intra prediction modes from said decoder-side intra mode derivation parameters; means for computing at least two predictors for said block of video sample data based on said decoder-side intra mode derivation parameters specific to each region of support; as well as Means for combining at least the two predictors together using blending to form a prediction for the block, wherein the blending is performed using sample based weights.

10. The device according to claim 9, comprising: means for determining two or more regions of support for a block of video sample data, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; as well as Means for performing said decoder-side intra-mode derivation process to derive decoder-side intra-mode derivation parameters specific to each region of support.

11. The apparatus according to claim 9, comprising: means for inferring whether said one or more intra prediction modes are position-dependent on a particular region of support based on said decoder-side intra mode derivation parameters specific to each region of support.

12. The apparatus of claim 9, wherein the sample-based weights are dependent on the positions of samples in the blocks.

13. The apparatus of claim 9, wherein the sample-based weights depend on positional dependency of a given intra-prediction mode over a particular region of support.

14. The device according to any one of claims 9 to 13, comprising: Means for inferring usage of sample-based weighting based on characteristics of the block of video data.

15. The apparatus according to claim 14, comprising: means for inferring whether said one or more intra prediction modes are position-dependent on a particular region of support based on said decoder-side intra mode derivation parameters specific to each region of support.

16. The apparatus of claim 14, wherein the sample-based weighting is inferred depending on the strength of the position dependency of a decoder-side intra-mode derivation mode over a specific region of support.

17. Apparatus according to any preceding claim, comprising: Means for combining one or more intra prediction modes derived using the decoder-side intra mode derivation process with computing a prediction for the current block using one or more predetermined intra prediction modes.

18. The apparatus according to claim 17, comprising: Means for using sample-based weights to compute the prediction for the current block, wherein the weights for the one or more predetermined intra-prediction modes depend on the decoder-side intra-mode derivation parameters specific to each region of support.

19. A method comprising: decoding encoded samples of a block of video sample data; determining two or more regions of support for a block of video sample data, wherein each region of support comprises a set of reconstructed samples at predetermined positions relative to the block of video sample data; performing a decoder-side intra mode derivation process for each support region to obtain decoder-side intra mode derivation parameters specific to each support region; as well as A prediction of samples of the block of video sample data is computed based on the decoder-side intra mode derivation parameters specific to each region of support.

20. A method comprising: decoding encoded samples of a block of video sample data; Performing a decoder-side intra-frame mode derivation process to obtain a decoder-side intra-frame mode derivation parameter; Obtain one or more intra prediction modes from intra mode derivation parameters at the decoder side; calculating at least two predictors for the block of video sample data based on the intra prediction mode; as well as At least the two predictors are combined together using blending to form a prediction for the block, wherein the blending is performed using sample-based weights.