Apparatus, method and computer program for video encoding and decoding
By reducing the data bit depth and arithmetic operation complexity in video encoding, and adjusting the parameters of the cross component prediction model with fixed values, the problems of data bit depth and arithmetic operation complexity in video encoding are solved, and a video encoding process that is more in line with standard specifications and is easy to implement by hardware.
Patent Information
- Application Number
- CN202380070911.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-07
- Filing Date
- 2023-08-28
- Publication Date
- 2025-05-13
AI Technical Summary
In video encoding, the cross component linear model and the convolutional cross component model require high-bit depth data and complex arithmetic operations during the calculation process, which are difficult to implement in 32-bit operations defined in the video encoding standard specification, and are not easily obtained in the hardware environment.
By reducing the bit depth of the data and complex arithmetic operations, a method of subtracting fixed values from the sample values in the reference area before determining the cross component prediction model parameters is used, prediction is performed using the determined parameters, and the fixed values are added to the value of the predicted target sample.
It effectively reduces the bit depth of the data and the complexity of arithmetic operations, making the video encoding process more in line with the video encoding standard specifications and is easier to implement in the hardware environment.
Smart Images

Figure CN119999204A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to apparatus, methods and computer programs for video encoding and decoding. Background Art
[0002] In video coding, video and image samples are often encoded using a color representation such as YUV or YCbCr consisting of one luminance (luma) and two chrominance (chroma) channels. In these cases, the luminance channel, which primarily represents the scene lighting, is usually encoded at a certain resolution, while the chrominance channels, which typically represent the difference between certain color components, are usually encoded at a second resolution that is lower than the resolution of the luminance signal. The purpose of this differential representation is to decorrelate the color components and enable more efficient compression of the data.
[0003] In the general video coding (VVC / H.266) standard, a cross-component linear model (CCLM) or a convolutional cross-component model (CCCM) is used as a linear model for predicting samples in chroma channels (e.g., Cb and Cr). Model parameters are derived based on reconstructed samples in the neighborhood of chroma blocks, co-located neighboring samples in luma blocks, and reconstructed samples within co-located luma blocks.
[0004] Both CCLM and CCCM use a so-called solver process to derive model parameters. In the solver process, a variety of computational operations such as addition, subtraction, multiplication and division are usually used. However, the data required in the different stages of the calculation usually have large values, so the intermediate and arithmetic operations require a high bit depth, which is not suitable for the 32-bit operations defined in the video coding standard specifications and is not easily available in many hardware environments used to implement video codecs. Summary of the invention
[0005] Now, to at least alleviate the above problems, this paper introduces an enhanced method which reduces the bit depth of data and reduces complex arithmetic operations.
[0006] The independent claims define the scope of protection of various embodiments of the present invention. Embodiments and features described in this specification that do not fall within the scope of the independent claims, if any, are to be construed as examples that aid in understanding the various embodiments of the present invention.
[0007] The method according to the first aspect includes receiving an image block unit of a frame, the image block unit including samples in a color channel, wherein the color channel includes at least one chrominance channel and one luminance channel; reconstructing samples of the luminance channel of the image block unit; determining a reference area for predicting a target sample of at least one color channel of the image block unit, wherein the reference area includes one or more reference samples of reference samples in neighboring blocks in a current color channel / frame, the reference samples being in the neighboring blocks of a co-located block in a reference color channel / frame and / or in a co-located block in a reference color channel / frame; determining at least one fixed value to be subtracted from sample values of samples in the reference area before determining parameters of a cross-component prediction model; predicting the target sample of at least one color channel of the image block unit using the determined parameters of the cross-component prediction model; and adding the at least one fixed value to the value of the predicted target sample of at least one color channel of the image block.
[0008] The apparatus according to the second aspect comprises a component for receiving an image block unit of a frame, the image block unit comprising samples in a color channel, wherein the color channel comprises at least one chrominance channel and one luminance channel; a component for reconstructing samples of the luminance channel of the image block unit; a component for determining a reference area for predicting a target sample of at least one color channel of the image block unit, wherein the reference area comprises one or more reference samples of reference samples in neighboring blocks in a current color channel / frame, the reference samples being in the neighboring blocks of a co-located block in a reference color channel / frame and / or in a co-located block in a reference color channel / frame; a component for determining at least one fixed value to be subtracted from sample values of samples in the reference area before determining parameters of a cross-component prediction model; a component for predicting the target sample of at least one color channel of the image block unit using the determined parameters of the cross-component prediction model; and a component for adding the at least one fixed value to the value of the predicted target sample of at least one color channel of the image block.
[0009] According to an embodiment, the apparatus comprises means for separately calculating said at least one fixed value for at least one chrominance channel and said luma channel, wherein an average value of neighboring reconstructed samples used for training a cross-component prediction model is a fixed value.
[0010] According to an embodiment, the apparatus comprises means for calculating said at least one fixed value using a subset of said neighboring reconstructed samples.
[0011] According to an embodiment, the subset of neighboring reconstructed samples comprises a predefined sample from a reference region.
[0012] According to an embodiment, the apparatus comprises means for sub-sampling the reference samples; and means for determining at least one fixed value using one or more sub-sampled reference samples.
[0013] According to an embodiment, the apparatus comprises means for signaling the at least one fixed value in or along the bitstream per frame, slice or coding unit.
[0014] According to an embodiment, the device comprises means for subtracting a block-based fixed value from the values of said at least one chroma channel and said luma channel.
[0015] According to an embodiment, the apparatus comprises means for subtracting the at least one fixed value from the sample values before applying the autocorrelation matrix and the cross-correlation vector used in the cross-component prediction model.
[0016] According to an embodiment, the device comprises means for calculating a term based on fixed values of said at least one chrominance channel and said luma channel; and means for subtracting said term from the values of the autocorrelation matrix and the cross-correlation vector used in the cross-component prediction model.
[0017] According to an embodiment, the device comprises means for subtracting at least one fixed value of the luma channel during a downsampling process; and means for subtracting at least one fixed value of the at least one chroma channel during parameter derivation of a cross-component prediction model.
[0018] According to an embodiment, the apparatus comprises means for predicting said target sample of at least one color channel of an image block unit using intermediate sample values obtained by subtracting said at least one fixed value from sample values of samples in said reference area.
[0019] As described above, the apparatus and the computer-readable storage medium having code stored thereon are thus arranged to perform one or more of the above-described methods and embodiments related thereto. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] For a better understanding of the present invention, reference will now be made by way of example to the accompanying drawings, in which:
[0021] Figure 1 An electronic device using an embodiment of the present invention is schematically shown;
[0022] Figure 2 Schematically illustrates a user equipment suitable for adopting an embodiment of the present invention;
[0023] Figure 3 Further schematically illustrated are electronic devices employing embodiments of the present invention connected using wireless and wired network connections;
[0024] Figure 4a and Figure 4b Schematically illustrates an encoder and a decoder suitable for implementing an embodiment of the present invention;
[0025] Figure 5 The locations of samples used to derive the cross-component linear model (CCLM) parameters are shown;
[0026] Figure 6a and Figure 6b Examples of brightness samples being classified into two categories in the sample domain and the spatial domain are shown respectively;
[0027] Figure 7 An example of a co-located reference sample region consisting of reconstructed luma and chroma samples defined for both luma and chroma in a convolutional cross-component model (CCCM) is shown;
[0028] Figure 8 Various examples of filter kernel dimensions in CCCM are shown;
[0029] Fig. 9 An example of four reference lines adjacent to a prediction block is shown;
[0030] Fig.10 The matrix-weighted intra prediction process is shown;
[0031] Fig.11 An example of a low frequency non-separable transform (LFNST) process is shown;
[0032] Fig.12 A flowchart of a method according to an embodiment of the present invention is shown; and
[0033] Fig.13 A schematic diagram of an example multimedia communication system is shown in which various embodiments may be implemented. DETAILED DESCRIPTION
[0034] Suitable means and possible mechanisms for predicting chrominance samples are described in further detail below. Figure 1 and Figure 2 ,in Figure 1 A block diagram of a video encoding system according to an example embodiment is shown as a schematic block diagram of an exemplary apparatus or electronic device 50 that may incorporate a codec according to an embodiment of the present invention. Figure 2 The layout of the device according to the exemplary embodiment is shown. Figure 1 and Figure 2 components.
[0035] The electronic device 50 may be, for example, a mobile terminal or user equipment of a wireless communication system. However, it should be understood that the embodiments of the present invention may be implemented in any electronic device or apparatus that may need to encode and decode or encode or decode video images.
[0036] The device 50 may include a housing 30 for housing and protecting the device. The device 50 may further include a display 32 in the form of a liquid crystal display. In other embodiments of the present invention, the display may be any suitable display technology suitable for displaying images or videos. The device 50 may further include a keypad 34. In other embodiments of the present invention, any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data input system as part of a touch-sensitive display.
[0037] The device may include a microphone 36 or any suitable audio input, which may be a digital or analog signal input. The device 50 may further include an audio output device, which in embodiments of the present invention may be any of the following: headphones 38, speakers, or analog audio or digital audio output connections. The device 50 may further include a battery (or in other embodiments of the present invention, the device may be powered by any suitable mobile energy device, such as a solar cell, a fuel cell, or a clockwork generator). The device may further include a camera capable of recording or capturing images and / or video. The device 50 may further include an infrared port for short-range line-of-sight communication with other devices. In other embodiments, the device 50 may further include any suitable short-range communication solution, such as, for example, a Bluetooth wireless connection or a USB / Firewire wired connection.
[0038] The apparatus 50 may include a controller 56, a processor or processor circuitry for controlling the apparatus 50. The controller 56 may be connected to a memory 58 which, in embodiments of the invention, may store data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to a codec circuitry 54 which is adapted to perform encoding and decoding of audio and / or video data or to assist in encoding and decoding performed by the controller.
[0039] The apparatus 50 may further include a card reader 48 and a smart card 46, such as a UICC and a UICC reader, for providing user information and suitable for providing authentication information for authenticating and authorizing the user on the network.
[0040] The device 50 may include a radio interface circuit device 52, which is connected to the controller and is suitable for generating wireless communication signals, such as communicating with a cellular communication network, a wireless communication system or a wireless local area network. The device 50 may further include an antenna 44, which is connected to the radio interface circuit device 52, for transmitting the radio frequency signals generated at the radio interface circuit device 52 to (multiple) other devices, and for receiving radio frequency signals from (multiple) other devices.
[0041] The apparatus 50 may include a camera capable of recording or detecting individual frames, which are then passed to a codec 54 or controller for processing. The apparatus may receive video image data from another device for processing before transmission and / or storage. The apparatus 50 may also receive images wirelessly or via a wired connection for encoding / decoding. The structural elements of the apparatus 50 described above represent examples of components for performing corresponding functions.
[0042] about Figure 3 , showing an example of a system in which embodiments of the present invention may be utilized. System 10 includes a plurality of communication devices that may communicate via one or more networks. System 10 may include any combination of wired or wireless networks, including but not limited to wireless cellular telephone networks (such as GSM, UMTS, CDMA networks, etc.), wireless local area networks (WLANs) such as defined by any IEEE 802.x standard, Bluetooth personal area networks, Ethernet local area networks, token ring local area networks, wide area networks, and the Internet.
[0043] System 10 may include devices and / or apparatus 50 suitable for implementing both wired and wireless communications of embodiments of the present invention.
[0044] For example, Figure 3 The illustrated system shows a representation of the mobile telephone network 11 and the Internet 28. Connections to the Internet 28 may include, but are not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communications paths.
[0045] The example communication devices shown in system 10 may include, but are not limited to, electronic devices or apparatuses 50, a combination of a personal digital assistant (PDA) and mobile phone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22. The apparatus 50 may be stationary or mobile when carried by an individual who is moving. The apparatus 50 may also be located in a mode of transportation, including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle, or any similar suitable mode of transportation.
[0046] Embodiments may also be implemented in a set-top box; i.e., in a tablet or (laptop) personal computer (PC) with a combination of hardware or software or encoder / decoder implementations, in various operating systems, and in a chipset, processor, DSP and / or embedded system providing hardware / software based encoding, possibly with / without a display or wireless capabilities.
[0047] Some or additional devices may send and receive calls and messages and communicate with service providers via a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the Internet 28. The system may include additional communication devices and various types of communication devices.
[0048] The communication devices may communicate using various transmission technologies, including but not limited to code division multiple access (CDMA), global system for mobile communications (GSM), universal mobile telecommunications system (UMTS), time division multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-Internet protocol (TCP-IP), short message service (SMS), multimedia message service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11 and any similar wireless communication technology. The communication devices involved in implementing various embodiments of the present invention may communicate using various media, including but not limited to radio, infrared, laser, cable connection and any suitable connection.
[0049] In telecommunications and data networks, a channel can refer to a physical channel or a logical channel. A physical channel can refer to a physical transmission medium such as a wire, while a logical channel can refer to a logical connection on a multiplexed medium capable of transmitting several logical channels. A channel can be used to transmit an information signal (e.g., a bit stream) from one or several senders (or transmitters) to one or several receivers.
[0050] The MPEG-2 Transport Stream (TS), as specified in ISO / IEC 13818-1 or equivalently in ITU-T Recommendation H.222.0, is a format for carrying audio, video and other media, as well as program metadata or other metadata, in a multiplexed stream. A packet identifier (PID) is used to identify elementary streams (also called packetized elementary streams) within a TS. Therefore, logical channels within an MPEG-2 TS can be considered to correspond to specific PID values.
[0051] Available media file format standards include ISO Base Media File Format (ISO / IEC 14496-12, which may be abbreviated as ISOBMFF) and a file format for NAL unit structured video derived from ISOBMFF (ISO / IEC 14496-15).
[0052] A video codec consists of an encoder that converts the input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a visual form. The video encoder and / or the video decoder can also be separate from each other, i.e. they do not need to form a codec. Typically, the encoder discards some information in the original video sequence in order to represent the video in a more compact form (i.e. at a lower bit rate).
[0053] A typical hybrid video encoder, such as many encoder implementations of ITU-T H.263 and H.264, encodes video information in two stages. First, the pixel values in a certain picture area (or "block") are predicted, for example by a motion compensation component (finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being encoded) or by a spatial component (using the pixel values around the block to be encoded in a specified manner). Second, the prediction error, that is, the difference between the predicted pixel block and the original pixel block, is encoded. This is usually achieved by transforming the difference in pixel values using a specified transform (such as the discrete cosine transform (DCT) or a variant thereof), quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (image quality) and the size of the final encoded video representation (file size or transmission bit rate).
[0054] In temporal prediction, the prediction source is a previously decoded picture (also called a reference picture). In intra-block copy (IBC; also called intra-block copy prediction), the application of prediction is similar to temporal prediction, but the reference picture is the current picture, and only previously decoded samples can be referenced during the prediction process. Inter-layer or inter-view prediction can be applied similarly to temporal prediction, but the reference picture is a decoded picture from another scalable layer or another view, respectively. In some cases, inter-frame prediction may refer only to temporal prediction, while in other cases, inter-frame prediction may be collectively referred to as any temporal prediction and intra-block copy, inter-layer prediction, and inter-view prediction, as long as they are performed with the same or similar process as temporal prediction. Inter-frame prediction or temporal prediction may sometimes be referred to as motion compensation or motion compensated prediction.
[0055] Motion compensation can be performed with full sample or sub-sample accuracy. In the case of full sample accurate motion compensation, the motion can be represented as a motion vector with integer values for horizontal and vertical displacements, and the motion compensation process uses these displacements to effectively copy samples from a reference picture. In the case of sub-sample accurate motion compensation, the motion vector is represented by fractional or decimal values for the horizontal and vertical components of the motion vector. In the case where the motion vector refers to a non-integer position in the reference picture, a sub-sample interpolation process is usually called to calculate the predicted sample value based on the reference sample and the selected sub-sample position. The sub-sample interpolation process usually consists of horizontal filtering compensation for the horizontal offset relative to the full sample position, followed by vertical filtering compensation for the vertical offset relative to the full sample position. However, in some environments, vertical processing can also be completed before horizontal processing.
[0056] Inter prediction (also called temporal prediction, motion compensation or motion compensated prediction) reduces temporal redundancy. In inter prediction, the prediction source is a previously decoded picture. Intra prediction exploits the fact that adjacent pixels within the same picture may be correlated. Intra prediction can be performed in the spatial domain or in the transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically used for intra-frame coding, where inter prediction is not applied.
[0057] One result of the encoding process is a set of coded parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy coded more efficiently if they are first predicted from spatially or temporally adjacent parameters. For example, a motion vector can be predicted from spatially adjacent motion vectors, and only the difference with respect to the motion vector predictor can be encoded. Prediction of coding parameters and intra-frame prediction can be collectively referred to as intra-picture prediction.
[0058] Figure 4a and Figure 4b An encoder and decoder suitable for use with embodiments of the present invention are shown. A video codec consists of an encoder that converts input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically, an encoder discards and / or loses some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bit rate). An example of the encoding process is shown in FIG. Figure 4a Shown in. Figure 4a The image to be encoded (I n ); The predicted representation of the image block (P' n ); prediction error signal (D n ); reconstructed prediction error signal (D' n ); Preliminary reconstructed image (I' n ); Final reconstructed image (R' n ); Transform (T) and Inverse Transform (T -1), quantization (Q) and inverse quantization (Q -1 ); Entropy coding (E); Reference frame memory (RFM); Inter-frame prediction (P inter ); intra prediction (P intra ); mode selection (MS) and filtering (F).
[0059] An example of the decoding process is shown in Figure 4b Shown in. Figure 4b The predicted representation of the image block (P' n ); reconstructed prediction error signal (D' n ); Preliminary reconstructed image (I' n ); Final reconstructed image (R' n ); inverse transform (T -1 ); inverse quantization (Q -1 ); Entropy decoding (E -1 ); reference frame memory (RFM); prediction (inter or intra) (P); and filtering (F).
[0060] Many hybrid video encoders encode video information in two stages. First, the pixel values in a certain picture area (or "block") are predicted, for example, by a motion compensation component (finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being encoded) or by a spatial component (using the pixel values around the block to be encoded in a specified manner). Second, the prediction error, that is, the difference between the predicted pixel block and the original pixel block, is encoded. This is usually achieved by transforming the difference in pixel values using a specified transform (such as the discrete cosine transform (DCT) or a variant thereof), quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and the size of the final encoded video representation (file size or transmission bit rate). Video codecs may also provide a transform skip mode, which the encoder may choose to use. In transform skip mode, the prediction error is encoded in the sample domain, for example by deriving sample-by-sample differences relative to certain neighboring samples, and encoding the sample-by-sample differences with an entropy encoder.
[0061] Entropy coding / decoding can be performed in a variety of ways. For example, context-based coding / decoding can be applied, in which both the encoder and the decoder modify the context state of the coding parameters based on the coding parameters of the previous coding / decoding. Context-based coding can be, for example, context-adaptive binary arithmetic coding (CABAC) or context-based variable length coding (CAVLC) or any similar entropy coding. Entropy coding / decoding can alternatively or additionally be performed using a variable length coding scheme, such as Huffman coding / decoding or exponential Golomb coding / decoding. Decoding the coding parameters from an entropy-coded bit stream or codeword can be referred to as parsing.
[0062] The phrase along the bitstream (e.g., indication along the bitstream) may be defined to refer to out-of-band transmission, signaling, or storage in a manner that out-of-band data is associated with the bitstream. The phrase decoding along the bitstream etc. may refer to decoding the referenced out-of-band data associated with the bitstream (which may be obtained from the out-of-band transmission, signaling, or storage). For example, indication along the bitstream may refer to metadata in a container file that encapsulates the bitstream.
[0063] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunication Standardization Sector of the International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by two parent standardization organizations, and it is known as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There are multiple versions of the H.264 / AVC standard, integrating new extensions or features into the specification. These extensions include Scalable Video Coding (SVC) and Multi-view Video Coding (MVC).
[0064] Version 1 of the High Efficiency Video Coding (H.265 / HEVC, also known as HEVC) standard was developed by the Joint Collaboration Team on Video Coding (JCT-VC) of VCEG and MPEG. The standard was published by two parent standardization organizations and is known as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Subsequent versions of H.265 / HEVC include Scalable, Multi-view, Fidelity Range, 3D, and Screen Content Coding extensions, which can be abbreviated as SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.
[0065] Versatile Video Coding (VVC) (MPEG-I Part 3), also known as ITU-T H.266, is a video compression standard developed by the Joint Video Experts Team (JVET) of the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union (ITU) (formally known as ISO / IEC JTC1 SC29 WG11), as the successor to HEVC / H.265.
[0066] This section describes some key definitions, bitstream and coding structures, and concepts of H.264 / AVC and HEVC as examples of video encoders, decoders, encoding methods, decoding methods, and bitstream structures in which these embodiments may be implemented. Some of the key definitions, bitstream and coding structures, and concepts of H.264 / AVC are the same as those in HEVC, and therefore, they are described jointly below. Aspects of the present invention are not limited to H.264 / AVC or HEVC, but rather are described for one possible basis on which the present invention may be partially or fully implemented.
[0067] Similar to many earlier video coding standards, the bitstream syntax and semantics and the decoding process for error-free bitstreams are specified in H.264 / AVC and HEVC. The encoding process is not specified, but the encoder must produce a conforming bitstream. The conformance of the bitstream and the decoder can be verified using a hypothetical reference decoder (HRD). These standards contain coding tools that help cope with transmission errors and losses, but the use of these tools in encoding is optional, and the decoding process for erroneous bitstreams is not specified.
[0068] The basic unit for the input of an H.264 / AVC or HEVC encoder and the output of an H.264 / AVC or HEV decoder, respectively, is a picture. A picture given as an encoder input may also be referred to as a source picture, and a picture decoded by a decoder may be referred to as a decoded picture.
[0069] The source picture and the decoded picture each consist of one or more sample arrays, such as one of the following sets of sample arrays:
[0070] - Luminance (Y) only (monochrome).
[0071] - Luma and dual chroma (YCbCr or YCgCo).
[0072] - Green, Blue and Red (GBR, also known as RGB).
[0073] - Arrays representing other unspecified monochromatic or tristimulus color sampling (e.g., YZX, also known as XYZ).
[0074] In H.264 / AVC and HEVC, a picture can be a frame or a field. A frame consists of a matrix of luma samples and possibly corresponding chroma samples. A field is a set of alternating sample rows of a frame and can be used as encoder input when the source signal is interlaced. The chroma sample array may not be present (and therefore monochrome sampling may be used), or the chroma sample array may be subsampled compared to the luma sample array. The chroma format can be summarized as follows:
[0075] - In monochrome sampling, there is only one sample array, which can nominally be considered the brightness array.
[0076] - In 4:2:0 sampling, the width and height of each of the two chroma arrays is half the height and width of the luma array.
[0077] - In 4:2:2 sampling, each of the two chroma arrays has the same height as the luma array and a width that is half the width of the luma array.
[0078] - In 4:4:4 sampling, when separate color planes are not used, the width and height of each of the two chroma arrays are the same as the height and width of the luma array.
[0079] In H.264 / AVC and HEVC, sample arrays can be encoded into the bitstream as separate color planes, and the separately encoded color planes can be decoded from the bitstream separately. When separate color planes are used, each of them is processed separately (by the encoder and / or decoder) as a picture with monochrome sampling.
[0080] Partitioning can be defined as dividing a set into subsets such that every element of the set is in exactly one of the subsets.
[0081] When describing the operation of HEVC encoding and / or decoding, the following terms may be used. A coding block may be defined as a block of N x N samples for some value of N, such that partitioning a coding tree block into coding blocks is a partition. A coding tree block (CTB) may be defined as a block of N x N samples for some value of N, such that partitioning a component into coding tree blocks is a partition. A coding tree unit (CTU) may be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples of a picture with three sample arrays, or a coding tree block of samples of a monochrome picture or a picture encoded using three separate color planes and a syntax structure for encoding samples. A coding unit (CU) may be defined as a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture with three sample arrays, or a coding block of samples of a monochrome picture or a picture encoded using three separate color planes and a syntax structure for encoding samples. A CU with the maximum allowed size may be named an LCU (maximum coding unit) or a coding tree unit (CTU), and a video image is divided into non-overlapping LCUs.
[0082] A CU consists of one or more prediction units (PUs) that define the prediction process for samples within the CU and one or more transform units (TUs) that define the prediction error coding process for samples in the CU. Typically, a CU consists of a square block of samples, whose size can be selected from a predetermined set of possible CU sizes. To increase the granularity of the prediction and prediction error coding processes, respectively, each PU and TU can be further split into smaller PUs and TUs. Each PU has prediction information associated with it, defining which prediction is applied to the pixels within the PU (e.g., motion vector information for inter-predicted PUs and intra-prediction directionality information for intra-predicted PUs).
[0083] Each TU may be associated with information describing the prediction error decoding process for the samples within the TU (including, for example, DCT coefficient information). It is usually signaled at the CU level whether prediction error coding is applied to each CU. If there is no prediction error residual associated with a CU, it can be considered that there is no TU for the CU. The partitioning of an image into CUs, and the partitioning of a CU into PUs and TUs is usually signaled in the bitstream, allowing a decoder to reproduce the intended structure of these units.
[0084] In HEVC, a picture can be partitioned into tiles, which are rectangular and contain an integer number of LCUs. In HEVC, the partitioning of tiles forms a regular grid where the height and width of the tiles differ by a maximum of one LCU. In HEVC, a slice is defined as an integer number of coding tree units contained in an independent slice segment, and all subsequent dependent slice segments (if any) before the next independent slice segment (if any) in the same access unit. In HEVC, a slice segment is defined as an integer number of coding tree units that are ordered consecutively in a tile scan and contained in a single NAL unit. The division of each picture into slice segments is a partitioning. In HEVC, an independent slice segment is defined as a slice segment whose values of the syntax elements of the slice segment header are not inferred from the values of the previous slice segment, and a dependent slice segment is defined as a slice segment whose values of some syntax elements of the slice segment header are inferred from the values of the previous independent slice segment in decoding order. In HEVC, a slice header is defined as a slice segment header of an independent slice segment, which is the current slice segment or an independent slice segment preceding the current dependent slice segment, and a slice segment header is defined as a portion of a coded slice segment that contains data elements related to the first coding tree unit or all coding tree units represented in the slice segment. If tiles are not used, CUs are scanned in the raster scan order of LCUs within tiles or within pictures. In an LCU, CUs have a specific scanning order.
[0085] The decoder reconstructs the output video by applying a prediction component similar to the encoder to form a predicted representation of pixel blocks (using motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (the inverse operation of prediction error encoding, recovering the quantized prediction error signal in the spatial pixel domain). After applying the prediction and prediction error decoding components, the decoder adds the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) may also apply additional filtering components to improve the quality of the output video before passing the output video for display and / or storing it as a prediction reference for upcoming frames in the video sequence.
[0086] The filtering may, for example, include one or more of: deblocking, sample adaptive offset (SAO), and / or adaptive loop filtering (ALF).H.264 / AVC includes deblocking, while HEVC includes both deblocking and SAO.
[0087] In a typical video codec, motion information is represented by a motion vector associated with each motion compensated image block, such as a prediction unit. Each of these motion vectors represents the displacement of an image block in a picture to be encoded (in the encoder side) or decoded (in the decoder side) and a prediction source block in one of the previously encoded or decoded pictures. In order to efficiently represent the motion vectors, those motion vectors are usually differentially encoded relative to the block-specific predicted motion vector. In a typical video codec, the predicted motion vector is created in a predefined manner, such as calculating the median of the encoded or decoded motion vectors of neighboring blocks. Another way to create a motion vector prediction is to generate a list of candidate predictions from neighboring blocks and / or co-located blocks in a temporal reference picture, and signal the selected candidate as a motion vector predictor. In addition to predicting the motion vector value, it is possible to predict which reference pictures are used for motion compensated prediction, and this prediction information can be represented, for example, by a reference index of a previously encoded / decoded picture. The reference index is usually predicted based on neighboring blocks and / or co-located blocks in a temporal reference picture. Furthermore, typical high efficiency video codecs employ an additional motion information encoding / decoding mechanism, generally referred to as merge / merge mode, in which all motion field information including motion vectors and corresponding reference picture indexes for each available reference picture list is predicted and used without any modification / correction. Similarly, prediction of the motion field information is performed using motion field information of neighboring blocks and / or co-located blocks in temporal reference pictures, and the used motion field information is signaled in a list of motion field candidate lists populated with motion field information of available neighboring / co-located blocks.
[0088] In a typical video codec, the prediction residuals after motion compensation are first transformed with a transform kernel (such as DCT) and then encoded. The reason for this is that there is usually still some correlation between the residuals, and in many cases, the transform can help reduce this correlation and provide more efficient coding.
[0089] Video coding standards and specifications may allow the encoder to divide a coded picture into coded slices, etc. Intra-picture prediction across slice boundaries is typically disabled. Therefore, slices can be viewed as a way to partition a coded picture into independently decodable fragments. In H.264 / AVC and HEVC, intra-picture prediction can be disabled across slice boundaries. Therefore, slices can be viewed as a way to partition a coded picture into independently decodable fragments, and slices are therefore typically viewed as the basic unit of transmission. In many cases, the encoder can indicate in the bitstream which types of intra-picture prediction are turned off across slice boundaries, and the decoder operation takes this information into account, such as when deriving which prediction sources are available. For example, if adjacent CUs reside in different slices, samples from adjacent CUs may be considered unavailable for intra prediction.
[0090] The basic unit for the output of the H.264 / AVC or HEVC encoder and the input of the H.264 / AVC or HEV decoder is the network abstraction layer (NAL) unit, respectively. For transmission over a packet-oriented network or storage in a structured file, the NAL unit can be encapsulated into a packet or similar structure. A byte stream format is specified in H.264 / AVC and HEVC for a transmission or storage environment that does not provide a frame structure. The byte stream format separates NAL units from each other by appending a start code in front of each NAL unit. In order to avoid erroneous detection of NAL unit boundaries, the encoder runs a byte-oriented start code emulation prevention algorithm that adds emulation prevention bytes to the NAL unit payload if the start code would otherwise occur. In order to achieve direct gateway operation between packet-oriented and stream-oriented systems, start code emulation prevention can always be performed regardless of whether the byte stream format is used. The NAL unit can be defined as a syntax structure that contains an indication of the type of data to be followed, and bytes containing the data in the form of RBSP, interspersed with emulation prevention bytes when necessary. The original byte sequence payload (RBSP) can be defined as a syntax structure containing an integer byte encapsulated in a NAL unit. The RBSP is either empty or has the form of a string of data bits containing a syntax element, followed by the RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0091] The NAL unit consists of a header and a payload. In H.264 / AVC and HEVC, the NAL unit header indicates the type of the NAL unit.
[0092] In HEVC, a two-byte NAL unit header is used for all specified NAL unit types. The NAL unit header contains a reserved bit, a six-bit NAL unit type indication, a three-bit nuh_temporal_id_plus1 indication for the temporal level (which may need to be greater than or equal to 1), and a six-bit nuh_layer_id syntax element. The temporal_id_plus1 syntax element can be regarded as a temporal identifier for the NAL unit, and the zero-based TemporalId variable can be derived as follows: TemporalId = temporal_id_plus1-1. The abbreviation TID can be used interchangeably with the TemporalId variable. TemporalId equal to 0 corresponds to the lowest temporal level. The value of temporal_id_plus1 must be non-zero to avoid start code emulation involving two NAL unit header bytes. The bitstream created by excluding all VCL NAL units with TemporalId greater than or equal to the selected value and including all other VCL NALs units is consistent. Therefore, a picture with TemporalId equal to tid_value does not use any picture with TemporalId greater than tid_value as inter prediction reference. A sublayer or temporal sublayer can be defined as a temporal scalable layer (or temporal layer TL) of a temporal scalable bitstream, consisting of VCL NAL units with a specific value of the TemporalId variable and associated non-VCL NAL units. nuh_layer_id can be understood as a scalability layer identifier.
[0093] NAL units can be divided into video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are usually coded slice NAL units. In HEVC, VCL NAL units contain syntax elements representing one or more CUs.
[0094] A non-VCL NAL unit may be, for example, one of the following types: a sequence parameter set, a picture parameter set, a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end-of-sequence NAL unit, an end-of-bitstream NAL unit, or a filler data NAL unit. Parameter sets may be required to reconstruct a decoded picture, while many other non-VCL NAL units are not required to reconstruct decoded sample values.
[0095] Parameters that remain unchanged in the coded video sequence may be included in the sequence parameter set. In addition to the parameters that may be required for the decoding process, the sequence parameter set may optionally include video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, presentation, and resource reservation. In HEVC, the sequence parameter set RBSP includes parameters that can be referenced by one or more picture parameter sets RBS or one or more SEI NAL units containing a buffering period SEI message. The picture parameter set contains such parameters that may not change in several coded pictures. The picture parameter set RBSP may include parameters that can be referenced by the coded slice NAL units of one or more coded pictures.
[0096] In HEVC, a video parameter set (VPS) can be defined as a syntax structure containing syntax elements that apply to zero or more complete coded video sequences, as determined by the content of the syntax elements found in the SPS, which are referenced by syntax elements found in the PPS, which are referenced by syntax elements found in each slice segment header.
[0097] A video parameter set RBSP may include parameters that may be referenced by one or more sequence parameter sets RBS.
[0098] The relationship and hierarchy between the video parameter set (VPS), the sequence parameter set (SPS) and the picture parameter set (PPS) can be described as follows. The VPS is one level above the SPS in the parameter set hierarchy and in the context of scalability and / or 3D video. The VPS may include parameters common to all slices across all (scalability or view) layers in the entire coded video sequence. The SPS includes parameters common to all slices in a specific (scalability or view) layer in the entire coded video sequence, which may be shared by multiple (scalability or view) layers. The PPS includes parameters common to all slices in a specific layer representation (a representation of a scalability or view layer in an access unit), and these parameters may be shared by all slices in a multi-layer representation.
[0099] The VPS may provide information about dependencies of various layers in the bitstream, as well as a lot of other information applicable to all slices across all (scalability or view) layers in the entire coded video sequence. The VPS may be considered to consist of two parts, a base VPS and a VPS extension, where the VPS extension may optionally be present.
[0100] Out-of-band transmission, signaling or storage may additionally or alternatively be used for other purposes besides tolerating transmission errors, such as ease of access or session negotiation. For example, a sample entry for a track in a file conforming to the ISO base media file format may include a parameter set, while the coded data in the bitstream is stored elsewhere in the file or in another file. In the claims and described embodiments, phrases along the bitstream (e.g., indicated along the bitstream) or along the bitstream coding unit (e.g., indicated along the coding tile) may be used to refer to out-of-band transmission, signaling or storage in such a way that the out-of-band data is associated with the bitstream or coding unit, respectively. Phrases decoding along the bitstream or along the coding unit of the bitstream, etc. may refer to decoding the referenced out-of-band data (which may be obtained from the out-of-band transmission, signaling or storage) associated with the bitstream or coding unit, respectively.
[0101] A SEI NAL unit may contain one or more SEI messages that are not required for decoding of output pictures but may help with related processes such as picture output timing, presentation, error detection, error concealment, and resource reservation.
[0102] A coded picture is an encoded representation of a picture.
[0103] In HEVC, a coded picture can be defined as a coded representation of a picture that contains all coding tree units of the picture. In HEVC, an access unit (AU) can be defined as a set of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and contain at most one picture with any specific value of nuh_layer_id. In addition to VCL NAL units containing coded pictures, access units can also contain non-VCL NAL units. The specified classification rule can, for example, associate pictures with the same output time or picture output count value into the same access unit.
[0104] A bitstream can be defined as a sequence of bits in the form of a NAL unit stream or a byte stream, which forms a representation of coded pictures and associated data, which form one or more coded video sequences. The first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or the same connection of a communication protocol. An elementary stream (in the context of video coding) can be defined as a sequence of one or more bitstreams. The end of the first bitstream can be indicated by a specific NAL unit, which can be called the end of bitstream (EOB) NAL unit and is the last NAL unit of the bitstream. In HEVC and its current draft extensions, the EOB NAL unit requires nuh_layer_id to be equal to 0.
[0105] In H.264 / AVC, a coded video sequence is defined as a contiguous sequence of access units from an IDR access unit (inclusive) to the next IDR access unit (exclusive) or to the end of the bitstream, whichever is earlier, in decoding order.
[0106] In HEVC, a coded video sequence (CVS) may be defined, for example, as a sequence of access units consisting, in decoding order, of an IRAP access unit with NoRaslOutputFlag equal to 1, followed by zero or more non-IRAP access units with NoRaslOputFlag equal to 1, including all subsequent access units, up to but not including any subsequent access unit that is an IRAP access unit with NoRaslOputFlag equal to 1. An IRAP access unit may be defined as an access unit in which a base layer picture is an IRAP picture. For each IDR picture, each BLA picture, and each IRAP picture, the value of NoRaslOutputFlag is equal to 1, the IRAP picture is the first picture in that particular layer in the decoding order in the bitstream, and is the first IRAP picture after the end of a sequence of NAL units with the same nuh_layer_id value in the decoding order. There may be a component that provides the value of HandleCraAsBlaFlag to the decoder from an external entity that can control the decoder, such as a player or receiver. HandleCraAsBlaFlag may be set to 1, for example by a player seeking a new position in the bitstream or tuning to a broadcast and starting decoding, and subsequently starting decoding from a CRA picture. When HandleCraAsBlaFlag of a CRA picture is equal to 1, the CRA picture will be handled and decoded as if it were a BLA picture.
[0107] In HEVC, when a specific NAL unit (which may be referred to as an end of sequence (EOS) NAL unit) is present in the bitstream and nuh_layer_id is equal to 0, it may additionally or alternatively (according to the above specification) specify the end of a coded video sequence.
[0108] A group of pictures (GOP) and its characteristics can be defined as follows. A GOP can be decoded regardless of whether any previous picture is decoded. An open GOP is a group of pictures in which, when decoding starts from the initial intra picture of an open GOP, pictures that precede the initial intra picture in output order may not be decoded correctly. In other words, pictures of an open GOP may refer to pictures belonging to a previous GOP (in inter prediction). The HEVC decoder can identify the intra picture that starts an open GOP because a specific NAL unit type, a CRA NAL unit type, may be used for its coded slice. A closed GOP is a group of pictures in which all pictures can be correctly decoded when decoding starts from the initial intra picture of a closed GOP. In other words, no picture in a closed GOP refers to any picture in a previous GOP. In H.264 / AVC and HEVC, a closed GOP can start from an IDR picture. In HEVC, a closed GOP can also start from a BLA_W_RADL or BLA_N_LP picture. Due to the greater flexibility in selecting reference pictures, an open GOP coding architecture may be more efficient in compression than a closed GOP coding structure.
[0109] The decoded picture buffer (DPB) can be used for the encoder and / or decoder. There are two reasons for buffering decoded pictures, namely for reference in inter-frame prediction and for reordering decoded pictures to output order. Since H.264 / AVC and HEVC provide great flexibility for both reference picture marking and output reordering, separate buffers for reference picture buffering and output picture buffering may waste memory resources. Therefore, the DPB can include a unified decoded picture buffering process for reference pictures and output reordering. When the decoded picture is no longer used as a reference and does not need to be output, the decoded picture can be removed from the DPB.
[0110] In many coding modes of H.264 / AVC and HEVC, reference pictures used for inter prediction are represented by indices into reference picture lists. The index can be encoded using variable length coding, which typically results in smaller indices having shorter values for the corresponding syntax elements. In H.264 / AVC and HEVC, two reference picture lists (reference picture list 0 and reference picture list 1) are generated for each bi-predicted (B) slice, and one reference picture list (reference picture list 0) is formed for each inter-coded (P) slice.
[0111] A number of coding standards, including H.264 / AVC and HEVC, may have a decoding process to derive a reference picture index of a reference picture list, which may be used to indicate which of a plurality of reference pictures is used for inter prediction of a particular block. The reference picture index may be encoded into the bitstream by the encoder in some inter coding modes, or may be derived (by the encoder and decoder) using neighboring blocks, for example, in some other inter coding schemes.
[0112] The motion parameter type or motion information may include but is not limited to one or more of the following types:
[0113] - an indication of the prediction type (e.g. intra prediction, uni prediction, bi prediction) and / or the number of reference pictures;
[0114] - an indication of a prediction direction, such as inter (also called temporal) prediction, inter-layer prediction, inter-view prediction, view synthesis prediction (VSP) and inter-component prediction (which may be indicated per reference picture and / or per prediction type, and in some embodiments, inter-view and view synthesis prediction may be jointly considered as one prediction direction) and / or
[0115] - an indication of the reference picture type, such as a short-term reference picture and / or a long-term reference picture and / or an inter-layer reference picture (which may be indicated, for example, per reference picture)
[0116] - a reference index to a reference picture list and / or any other identifier of a reference picture (which may be indicated, for example, per reference picture, and its type may depend on the prediction direction and / or the reference picture type, and it may be accompanied by other pieces of relevant information, such as the reference picture list to which the reference index applies, etc.);
[0117] - horizontal motion vector component (which may be indicated, for example, by prediction block or by reference index, etc.);
[0118] - vertical motion vector component (which may be indicated, for example, per prediction block or per reference index, etc.);
[0119] - one or more parameters, such as a picture order count difference and / or a relative camera spacing between a picture comprising or associated with a motion parameter and its reference picture, which may be used for scaling a horizontal motion vector component and / or a vertical motion vector component in one or more motion vector prediction processes (wherein the one or more parameters may be indicated, for example, per reference picture or per reference index, etc.);
[0120] - the coordinates of the block to which the motion parameters and / or motion information apply, e.g. the coordinates of the top left sample of the block in units of luma samples;
[0121] - The extent (eg width and height) of the block to which the motion parameters and / or motion information apply.
[0122] - Compared with previous video coding standards, the General Video Codec (H.26 / VVC) introduces several new coding tools, such as the following:
[0123] Intra-frame prediction
[0124] -67 Intra-frame modes with wide-angle mode expansion,
[0125] - Block size and mode dependent 4-tap interpolation filter
[0126] -Position-dependent intra prediction combination (PDPC)
[0127] -Cross-Component Linear Model Intra Prediction (CCLM)
[0128] -Multiple reference line intra prediction
[0129] - Intra-frame sub-partition
[0130] -Weighted intra prediction with matrix multiplication
[0131] Inter-picture prediction
[0132] - Block motion replication with spatial, temporal, history-based and pairwise average merge candidates
[0133] -Affine motion inter-frame prediction
[0134] - Sub-block based temporal motion vector prediction
[0135] -Adaptive motion vector solution
[0136] - 8x8 block-based motion compression for temporal motion prediction
[0137] - High-precision (1 / 16 pixel) motion vector storage and motion compensation, where the luminance component uses 8-tap interpolation filtering and the chrominance component uses 4-tap interpolation filtering
[0138] -Triangular partition
[0139] - Combined intra and inter prediction
[0140] -Merged with MVD (MMVD)
[0141] -Symmetrical MVD encoding
[0142] - Bidirectional optical flow
[0143] -Decoder side motion vector refinement
[0144] -Dual prediction based on CU level weight
[0145] Transform, quantization and coefficient coding
[0146] - Multiple primary transform options with DCT2, DST7 and DCT8
[0147] - Secondary transformation in low frequency area
[0148] -Sub-block transform of inter prediction residual
[0149] - Increased max QP from 51 to 63 for quantization dependent
[0150] - Transform coefficient coding with sign data hiding
[0151] -Transform skip residual coding
[0152] Entropy coding
[0153] -Arithmetic coding engine with adaptive dual-window probability update
[0154] In-loop filtering
[0155] - In-loop shaping
[0156] - Deblocking filter with strong longer filter
[0157] - Sample Adaptive Offset
[0158] - Adaptive loop filtering
[0159] Screen content encoding:
[0160] - Current picture references with reference area restrictions
[0161] 360-degree video encoding
[0162] -Horizontal surround motion compensation
[0163] Advanced syntax and parallel processing
[0164] - Reference picture management with direct reference picture list signaling
[0165] - Tile groups with rectangular tile groups
[0166] Partitioning in VVC is performed similarly to HEVC, i.e., each picture is divided into coding tree units (CTUs). Pictures can also be divided into slices, tiles, bricks, and sub-pictures. CTUs can be partitioned into smaller CUs using a quadtree structure. Each CU can be partitioned using quadtrees and nested multi-type trees, including ternary and binary partitioning. However, there are specific rules to infer partitions in picture boundaries, and redundant partitioning patterns are not allowed in nested multi-type partitions.
[0167] Among the new coding tools listed above, VVC uses the cross-component linear model (CCLM) prediction mode to reduce cross-component redundancy. The chrominance samples are predicted based on the reconstructed luminance samples of the same CU using the following linear model:
[0168] pred C (i, j) = α·rec L ′(i,j)+β (Equation 1a)
[0169] where pred C (i, j) represents the predicted chroma sample in the CU, and rec L ′(i, j) represents the downsampled reconstructed luma sample of the same CU.
[0170] Alternatively, the following equation may be used for CCLM:
[0171]
[0172] The >> operation represents a bit shift to the right by value k.
[0173] The CCLM parameters (α and β) are derived using up to four adjacent chroma samples and their corresponding downsampled luma samples. Assuming the current chroma block size is W×H, W' and H' are set to
[0174] - When LM mode is applied, W'=W, H'=H;
[0175] - When LM-A mode is applied, W'=W+H;
[0176] - When LM-L mode is applied, H'=H+W;
[0177] In this paper, LM-A mode refers to linear model_above, in which only the upper template (i.e., the sample values from the adjacent position above the CU) is used to calculate the linear model coefficients. In order to obtain more samples, the upper template is expanded to (W+H). LM-L mode then refers to linear model_left, in which only the left template (i.e., the sample values from the adjacent position on the left side of the CU) is used to calculate the linear model coefficients. In order to obtain more samples, the left template is expanded to (H+W). For non-square blocks, the upper template is expanded to W+W and the left template is expanded to H+H.
[0178] The upper adjacent position is represented as S[0,-1]…S[W'-1,-1], and the left adjacent position is represented as S[-1,0]…S[-1,H'-1]. Then four samples are selected as
[0179] - When LM mode is applied and both upper and left neighboring samples are available, S[W' / 4,-1], S[3*W' / 4,-1], S[-1,H' / 4], S[-1,3*H' / 4];
[0180] - When LM-A mode is applied or only upper adjacent samples are available, S[W' / 8,-1], S[3*W' / 8,-1], S[5*W' / 8,-1], S[7*W' / 8,-1];
[0181] - When LM-L mode is applied or only left neighbor samples are available, S[-1,H' / 8],S[-1,3*H' / 8],S[-1,5*H' / 8],S[-1,7*H' / 8];
[0182] The four adjacent luma samples at the selected position are downsampled and compared four times to find the two smaller values: x0A and x1A, and the two larger values: x0B and x1B. Their corresponding chroma sample values are denoted as y0A, y1A, y0B, and y1B. Then, xA, xB, yA, and yB can be derived as:
[0183] X a =(x 0 A +x 1 A +1)>>1;X b =(x 0 B +x 1 B +1)>>1;Y a =(y 0 A +y 1 A +1)>>1;Y b =(y 0 B +y 1 B +1)>>1(Equation 2)
[0184] Finally, the linear model parameters α and β are obtained according to the following equations:
[0185]
[0186] β=Y b -α·X b (Equation 4)
[0187] Figure 5 An example of the positions of the left and upper samples and the current block samples involved in the CCLM mode is shown.
[0188] The division operation for calculating the parameter α is implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented in exponential notation. For example, diff is approximated using a 4-bit significant part and an exponent. Therefore, for 16 values of the significant part, the 1 / diff table is reduced to 16 elements as follows:
[0189] DivTable[] = {0,7,6,5,5,4,4,3,2,2,1,1,1,1,1,1,1,0} (Equation 5)
[0190] This reduces both the computational complexity and the memory size required to store the required tables.
[0191] In order to match the chroma sample positions of the 4:2:0 video sequence, two types of downsampling filters are applied to the luma samples to achieve a 2 to 1 downsampling ratio in both horizontal and vertical directions. The choice of downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "type-0" and "type-2" content respectively:
[0192]
[0193] Note that when the upper reference line is at a CTU boundary, only one luma line (common line buffer in intra prediction) is used to make the downsampled luma samples.
[0194] This parameter calculation is performed as part of the decoding process and not just as an encoder search operation. Therefore, no syntax is used to communicate the α and β values to the decoder.
[0195] For chroma intra mode coding, a total of 8 intra modes are allowed for chroma intra mode coding. Those modes include five traditional intra modes and three cross-component linear model modes (CCLM, LM_A and LM_L). The chroma mode signaling and derivation process are shown in Table 1. Chroma mode coding depends directly on the intra prediction mode of the corresponding luminance block. Since the separate block partition structure of luminance and chrominance components is enabled in the I slice, a chroma block can correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited.
[0196]
[0197] Table 1
[0198] As shown in Table 2, a single binarization table is used regardless of the value of sps_cclm_enabled_flag.
[0199] Value of intra_chroma_pred_mode Binary String 4 00 0 0100 1 0101 2 0110 3 0111 5 10 6 110 7 111
[0200] Table 2
[0201] In Table 2, the first binary number indicates whether it is normal (0) or LM mode (1). If it is LM mode, the next binary number indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next 1 binary number indicates whether it is LM_L (0) or LM_A (1). For this case, when sps_ccl_enabled_flag is 0, the first binary number of the binarization table corresponding to intra_chroma_pred_mode can be discarded before entropy coding. Or, in other words, the first binary number is inferred to be 0 and therefore not encoded. This single binarization table is used for the case where sps_cclm_enabled_flag is equal to 0 and 1. The first two bins in Table 2 are context encoded with their own context model, and the remaining bins are bypassed for encoding.
[0202] In addition, to reduce luma-chroma latency in dual trees, when the 64x64 luma coding tree node is partitioned with NotSplit (and intra sub-partitioning (ISP) is not used for 64x64 CUs) or QT, the chroma CUs in the 32x32 / 32x16 chroma coding tree nodes are allowed to use CCLM as follows:
[0203] - If a 32x 32 chroma node is not split or partitioned QT split, all chroma CUs in the 32x 32 node can use CCLM.
[0204] - If a 32x32 chroma node is partitioned with horizontal BT, and the 32x16 child node is not partitioned or uses vertical BT partitioning, all chroma CUs in the 32x16 chroma node can use CCLM.
[0205] Under all other luma and chroma coding tree split conditions, CCLM is not allowed for chroma CUs.
[0206] Multi-Model LM (MMLM)
[0207] The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, the reconstructed neighboring samples are divided into two categories using a threshold value, which is the average value of the brightness reconstructed neighboring samples. The linear model for each category is derived using the least mean square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. Figure 6a and Figure 6bTwo luminance-to-chrominance models are shown in the sample domain and the spatial domain respectively when the luminance (Y) threshold is 17. Each luminance-to-chrominance model has its own linear model parameters α and β. Figure 6b As can be seen in , each luminance-to-chrominance model corresponds to a spatial segmentation of the content (ie, they correspond to different objects or textures in the scene).
[0208] Convolutional Cross-Component Model (CCCM)
[0209] An improved version of cross-component prediction, called CCCM, uses a 2D filter kernel to derive a luma-to-chroma model. The filter coefficients are derived at the decoder side using the reconstructed input data set and chroma samples. For filter coefficient derivation, a co-located reference sample region (consisting of reconstructed luma and chroma samples) is defined for both luma and chroma, as Figure 7 , where the commonly used 4:2:0 chroma downsampling is applied. Figure 7 As shown, the reference sample area for a given block can be the six lines above and to the left, but any number of reference lines (which can be implemented by both the encoder and the decoder) can be used. In general, the reference samples can contain any chrominance and luminance samples reconstructed by both the encoder and the decoder. Once the reference samples are determined, different types of linear regression tools can be used to derive the filter coefficients, such as using ordinary least squares estimation, orthogonal matching pursuit, optimized orthogonal matching pursuit, ridge regression, or least absolute shrinkage and selection operators.
[0210] The filter kernel can be of dimensions such as 1x3 (1D vertical), 3x1 (1D horizontal), 3x3, 7x7 or any dimension, and can be shaped as a cross or a diamond (e.g., by selecting only a subset of all possible kernel positions) Figure 8 When referring to samples within the filter kernel, the following notation is used: North (above), East (right), South (below), West (left), and Center, as shown in Figure 4, using the letters N, E, S, W, C.
[0211] The overall approach of reconstructing chrominance samples using convolution between filter kernels obtained at the decoder side and the input data set is referred to herein as the cross-convolutional component model (CCCM). The following steps may be applied to perform the CCCM operation:
[0212] 1) Define co-located reference regions on luma and chroma components.
[0213] 2) Downsample the luma samples to match the chroma grid (optional).
[0214] 3) Scan the luminance and chrominance samples of the reference area and collect available statistics (such as autocorrelation matrix and cross-correlation vector) based on the filter shape.
[0215] 4) Solve for the filter coefficients by minimizing the squared error (or any other metric) based on available statistics (such as the autocorrelation matrix and the cross-correlation vector).
[0216] 5) Compute the predicted chrominance block by convolving the downsampled luma samples with the filter kernel.
[0217] Let's define the (possibly downsampled) luma samples as a 2D array Y(x,y) indexed using the horizontal x-coordinate and the vertical Y-coordinate. Let's also define the co-located chroma samples as a 2D array C(x,y), and the filter kernel (i.e., coefficients) as a 3x3 array F(i,j). At the sample level, we define the convolution between Y and F as,
[0218]
[0219] When other data terms are used, such as the nonlinear square root term, the additional convolution becomes,
[0220]
[0221] in are the filter coefficients outside the 2D filter kernel, but have been obtained as part of the linear system of equations used to solve the 2D filter coefficients in step 4 above. Similarly, we can add a bias term to the convolution using the following equation,
[0222]
[0223] Multiple Reference Line (MRL) Intra Prediction
[0224] Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Fig. 9 In Figure 1, an example of 4 reference lines is depicted, where the samples of segments A and F are not extracted from the reconstructed neighboring samples, but are filled with the closest samples from segments B and E, respectively. HEVC intra-image prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines are used (reference line 1 and reference line 3).
[0225] The index of the selected reference line (mrl_idx) is signaled and used to generate the intra predictor. For reference line idx greater than 0, only additional reference line modes are included in the MPM list, and only the MPM index is signaled without signaling the remaining modes. The reference line index is signaled before the intra prediction mode, and in case a non-zero reference line index is signaled, the planar mode is excluded from the intra prediction mode.
[0226] MRL is disabled for the first row of blocks within a CTU to prevent the use of extended reference samples outside the current CTU line. In addition, PDPC is disabled when additional lines are used. For MRL mode, the derivation of DC values in DC intra prediction mode for non-zero reference line index is aligned with the derivation of reference line index 0. MRL requires storage of 3 adjacent luma reference lines with a CTU to generate the prediction. The CCLM tool also requires 3 adjacent luma reference lines for its downsampling filtering. The definition of MLR using the same 3 lines is aligned with CCLM to reduce the storage requirements of the decoder.
[0227] Intra-frame sub-partitioning (ISP)
[0228] Intra subpartitioning (ISP) divides the luma intra prediction block vertically or horizontally into 2 or 4 subpartitions depending on the block size. For example, the minimum block size of ISP is 4x8 (or 8x4). If the block size is larger than 4x8 (or 8x4), the corresponding block is divided into 4 subpartitions. It has been noted that M×128 (where M≤64) and 128×N (where N≤64) ISP blocks may generate potential problems for 64×64 VDPU. For example, an M×128 CU in a single tree case has an M×128 luma TB and two corresponding M / 2×64 chroma TBs. If the CU uses ISP, the luma TB will be divided into four M×32 TBs (only horizontal partitioning is possible), each of which is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are not divided. Therefore, the size of both chroma components will be larger than a 32×32 block. Similarly, a 128×N CU using ISP may also create a similar situation. Therefore, these two cases are problematic for a 64×64 decoder pipeline. To this end, the CU size of the ISP can be limited to a maximum of 64×64. All sub-partitions meet the condition of having at least 16 samples.
[0229] Matrix Weighted Intra Prediction (MIP)
[0230] The matrix-weighted intra prediction (MIP) method is a new intra prediction technique in VVC. In order to predict the samples of a rectangular block of width W and height H, the matrix-weighted intra prediction (MIP) takes a row of H reconstructed adjacent boundary samples on the left side of the block and a row of W reconstructed adjacent boundary samples above the block as input. If the reconstructed samples are not available, they are generated in the traditional intra prediction manner. The generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication and linear interpolation, as shown in Fig.10 shown.
[0231] Decoder-side intra mode derivation (DIMD)
[0232] When DIMD is applied, two intra modes are derived from the reconstructed neighboring samples and these two predictors are combined with the planar mode predictor, whose weights are derived from the gradients as described in JVET-O0449. The division operation in the weight derivation is performed using the same lookup table (LUT) based integerization scheme used by CCLM. For example, the directional calculation
[0233] Orient=G y / G x
[0234] The partitioning operation in is calculated by the following LUT-based scheme:
[0235] x = Floor(Log2(Gx))
[0236] normDiff=((Gx<<4)>>x)&15
[0237] x+=(3+(normDiff!=0)?1:0)
[0238] Orient=(Gy*(DivSigTable[normDiff]|8)+(1<<(x-1)))>>x
[0239] in
[0240] DivSigTable
[16] ={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}.
[0241] The derived intra modes are included in the preliminary list of intra most probable modes (MPMs), so the DIMD process is performed before building the MPM list. The preliminary derived intra modes of a DIMD block are stored with the block and used for MPM list construction for neighboring blocks.
[0242] Template-based Intra Mode Derivation (TIMD) Fusion
[0243] For each intra prediction mode in MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The first two intra prediction modes with the smallest SATD are selected as TIMD modes. These two TIMD modes are fused with weights after applying the PDPC process, and the current CU is encoded using this weighted intra prediction. Position-dependent intra prediction combination (PDPC) is included in the derivation of TIMD modes.
[0244] The costs of the two selected modes are compared to a threshold, and in the test a cost factor of 2 is applied as follows: costMode2<2*costMode1.
[0245] If this condition is true, then fusion is applied, otherwise only mode 1 is used.
[0246] The weight of a mode is calculated based on its SATD cost as follows:
[0247] weight1=costMode2 / (costMode1+costMode2)
[0248] weight2=1-weight1
[0249] The division operation is performed using the lookup table (LUT) based integerization scheme used by CCLM.
[0250] Low Frequency Non-separable Transform (LFNST)
[0251] In VVC, LFNST is applied between the forward primary transform and quantization (at the encoder) and between dequantization and inverse primary transform (at the decoder side), as Fig.11 As shown in Figure 2. In LFNST, either 4x4 non-separable transform or 8x8 non-separable transform is applied depending on the block size. For example, 4x4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and 8x8 LFNST is applied to large blocks (i.e., min(width, height) > 4).
[0252] The following uses the input as an example to describe the application of the non-separable transformation used in LFNST. To apply 4x 4LFNST, the 4x4 input block X
[0253]
[0254] First, we represent it as a vector
[0255]
[0256] The inseparable transformation is calculated as in indicates the transform coefficient vector, and T is the 16x 16 transform matrix. 16x 1 coefficient vector The blocks are then reorganized into 4x 4 blocks using their scan order (horizontally, vertically or diagonally). The coefficients with smaller indices will be placed in a 4x4 coefficient block with smaller scan indices.
[0257] Reduced Inseparable Transformation
[0258] The Low-Frequency Non-Separable Transform (LFNST) applies a non-separable transform based on a direct matrix multiplication method, enabling its implementation in a single iteration without multiple iterations. However, it is necessary to reduce the dimension of the non-separable transform matrix to minimize the computational complexity and the memory space for storing the transform coefficients. Therefore, the Reduced Non-Separable Transform (or RST) method is used in LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (for an 8x8 NSST, N is usually equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, the RST matrix becomes an R×N matrix instead of an NxN matrix, as follows:
[0259]
[0260] where the R rows of the transform are the R bases of the N-dimensional space.
[0261] The inverse transform matrix of RT is the transpose of its forward transform. For an 8x8 LFNST, a reduction factor of 4 is applied, and the 64x64 direct matrix of the traditional 8x8 non-separable transform matrix size is reduced to a 16x48 direct matrix. Thus, a 48×16 inverse RST matrix is used at the decoder side to generate the core (primary) transform coefficients in the 8×8 upper-left region. When applying a 16x48 matrix instead of a 16x64 matrix with the same transform set configuration, each matrix obtains 48 input data from three 4x4 blocks (excluding the lower-right 4x4 block) in the upper-left 8x8 block. With the help of the reduced dimension, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, and the performance degradation is reasonable. To reduce the complexity, LFNST is only applicable when all coefficients outside the first coefficient subgroup are unimportant. Thus, when applying LFNST, all only primary transform coefficients must be zero. This allows for the adjustment of the LFNST index signaling at the last valid position and thus avoids the additional coefficient scanning in the current LFNST design, which is only required to check for valid coefficients at specific positions. The worst-case scenario for LFNST (in terms of multiplications per pixel) limits the non-separable transforms of 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In those cases, for other sizes less than 16, when applying LFNST, the last valid scan position must be less than 8. For blocks with shapes of 4xN, Nx4, and N > 8, the proposed limitation means that LFNST is now applied only once and only to the upper-left 4x4 region. Since all only primary coefficients are zero when applying LFNST, the number of operations required for the primary transform is reduced in such cases. From the encoder's perspective, when testing the LFNST transform, the quantization of the coefficients is significantly simplified. For the first 16 coefficients (in scan order), rate-distortion optimized quantization must be maximized, and the remaining coefficients are forced to zero.
[0262] LFNST Transform Selection
[0263] There are 4 transform sets in LFNST, and each transform set in LFNST uses 2 inseparable transform matrices (kernels). The mapping from intra prediction mode to transform set is predefined, as shown in the following table. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81<=predModeIntra<=83), transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable secondary transform candidate is further specified by the LFNST index sent with an explicit signal. After the transform coefficients, each intra CU sends the index with a signal in the bitstream.
[0264]
[0265]
[0266] Change selection mode
[0267] LFNST index signaling and interaction with other tools
[0268] Since LFNST is only applicable when all coefficients outside the first coefficient subgroup are insignificant, the LFNST index encoding depends on the position of the last significant coefficient. In addition, the LFNST index is context-coded, but does not depend on the intra prediction mode, and only the first binary number is context-coded. In addition, LFNST is applicable to intra CUs in both intra and inter slices, and for both luma and chroma. If dual tree is enabled, the LFNST indexes for luma and chroma will be signaled separately. For inter slices (dual tree is disabled), a single LFNST index is signaled and used for both luma and chroma.
[0269] Considering that large CUs larger than 64x 64 are implicitly split (TU tiling) due to the existing maximum transform size limit (64x 64), LFNST index search can increase the data buffering of a certain number of decoding pipeline stages by a factor of four. Therefore, the maximum size allowed by LFNST is limited to 64x 64. Note that LFNST is enabled only in the case of DCT2. LFNST index signaling is placed before MTS index signaling.
[0270] It is not obvious that using the scaling matrix for perceptual quantization, the scaling matrix specified for the primary matrix may be useful for the LFNST coefficients. Therefore, using the scaling matrix for LFNST coefficients is not allowed. For single-tree partitioning mode, chroma LFNST is not applied.
[0271] Derivation of CCCM (solver) parameters
[0272] In the above-mentioned process related to CCCM, the parameters of the filter can be represented by a vector x, and it can be convolved with the input vector z to produce predicted sample values for video or image coding purposes, for example. The vector x can consist of n parameters and can be given as follows:
[0273] x=[x0...x n-1 ] T
[0274] The input vector z can also consist of n input values and can be represented as:
[0275] z=[z0...z n-1 ] T
[0276] The predicted sample value p can then be calculated by convolving or multiplying the input z with the filter x as follows:
[0277]
[0278] The input vector z can be configured to include, for example, a luma sample value (e.g., a central luma sample corresponding to adjacent samples of a chroma sample and a central luma sample), a function of a luma sample value, or a constant, or a combination thereof. Including a constant in the input vector z corresponds to adding a constant to the output of the filter p. Such a constant can be referred to as a bias term or bias parameter and can be used to represent an offset between input and output values.
[0279] The filter parameters x can usually be calculated by finding solutions to a set of equations that can be expressed in matrix form as:
[0280] Ax=y
[0281] Where A represents the autocorrelation matrix of the determined input reference or "training" samples used in the process, and y represents the cross-correlation vector between the input training samples and the corresponding output training samples. The entries in the nxn matrix A and the vector y with n values can be calculated as follows:
[0282]
[0283] Where N is the number of training vectors included in the process, the R matrix contains the input training vectors (e.g., z) as its rows, and s represents the vector with output training samples. For example, in the CCCM case, R contains luma (or resampled luma) samples in the training region, and s contains the corresponding chroma samples in the training region.
[0284] In order to solve the filter coefficients xi in the vector x in the equation Ax=y, a method based on matrix decomposition is used. Different decompositions can be selected for this process. For example, LDL decomposition, Cholesky decomposition or QR decomposition can be used. For example, the implementation described in the following pseudocode can be used, with matrix A as input and upper triangular matrix U and vector d as output:
[0285]
[0286] By performing this decomposition, the upper triangular output matrix U corresponds to the transpose of the lower triangular L matrix of the LDL decomposition, and the output vector d contains the values of the diagonal elements of the diagonal matrix d of the LDL decomposition. The values of the output vector d can be referred to as scaling values, because these values are used to scale the values of the triangular matrix, and are also used to scale intermediate values when solving the decomposition system. The values of the vector d can be represented as the diagonal elements of, for example, a vector, an array, a list, or a matrix. In order to save memory in software or hardware implementations, the values of the vector d can be stored as the diagonal elements of the output matrix U, otherwise according to the above pseudocode example, there are unspecified values in the implementation.
[0287] As an alternative example, a lower triangular matrix L may be generated. Therefore, the filter coefficient vector x of Eq.
[0288] Ax=y
[0289] Now, after replacing A with its LDL decomposition, it can be solved by a simple back-substitution and scaling operation:
[0290] LDL T x=y
[0291] Or the U matrix calculated using the pseudocode example:
[0292] U T DUx=y
[0293] In practice, the vector x can be solved in three steps. In the first step, the DUx term can be labeled as a vector z, which can be solved by a simple back substitution:
[0294] DUx=z
[0295] U T z=y
[0296] In the second step, we can remove D by dividing the elements of z by the elements of vector D:
[0297]
[0298] In the third step, since the above equation is again in the form of an upper / lower triangular matrix multiplied by a vector x equal to another vector, the vector x can be solved directly using back substitution. The entire process of solving the filter coefficients can therefore be configured to have three stages: a first back substitution process, a scaling process, and a second back substitution process. Advantageously, the scaling process between the two back substitutions is performed using a vector generated as a product of the matrix decomposition.
[0299] The multiplication operation MULT and the division operation DIV can be implemented in different ways. For example, a floating point or fixed point implementation can be used. Since a fixed point implementation can provide faster execution on some computing architectures, it may be beneficial to use fixed point arithmetic in general. In order to achieve a favorable balance between the numerical stability of the computation process and the accuracy of the fixed point representation, the MULT operation can be rounded to the nearest integer, while the DIV operation can be rounded to zero. For example, the function can be defined in pseudo code as follows:
[0300]
[0301] Alternatively, the DIV operation can be implemented as a combination of, for example, a table lookup operation and a bit shift operation. It can also include rounding terms, just like the MULT function in the above example. The DECIM_BITS parameter determines the number of decimal digits in the fixed-point representation and can be set to different values depending on the desired accuracy of the operation.
[0302] The reverse substitution can be done in different ways. For example, the following pseudo code can be used:
[0303]
[0304] There are different ways to reduce the computational complexity of the bias term. For example, if the input to the filter is luma values and the output is predicted luma, and the filter is generated using a set of reference luma and chroma samples, then the average reference luma value y can be calculated mean and the average reference chrominance sample value c mean When calculating the filter coefficients, y can be deducted from the reference brightness value before or when generating the autocorrelation matrix A and the cross-correlation vector y. mean , and calculate c from the reference chrominance samples mean Similarly, when performing a convolution operation to compute the filter output, y can be subtracted from the input luma samples. mean Then, by taking the average reference chromaticity value c mean Adding to the output of the filter, the deviation between luma and chroma samples can be recovered as follows:
[0305]
[0306] Therefore, the multi-model cross-component intra prediction methods described above, such as the cross-component linear model (CCLM) and the convolutional cross-component model (CCCM), calculate some autocorrelation and cross-correlation data, which are fed to the solver step to derive the model parameters. The solver usually uses addition, subtraction, multiplication and division operations.
[0307] However, the data required in the different stages of the calculation often have large values, so intermediate and arithmetic operations require a high bit depth, which does not fit into the 32-bit operations defined in the video coding standard specifications and is not readily available in many hardware environments used to implement video codecs.
[0308] Improved methods are now introduced for achieving reduced data bit depth and less complex arithmetic operations in the solver.
[0309] Fig.12 A method according to one aspect is shown in the figure, wherein the method includes receiving (1200) an image block unit of a frame, the image block unit including samples in color channels, wherein the color channels include at least one chrominance channel and one luma channel; reconstructing (1202) samples of the luma channel of the image block unit; determining (1204) a reference area for predicting a target sample of at least one color channel of the image block unit, wherein the reference area includes one or more reference samples of reference samples in neighboring blocks in a current color channel / frame, the reference samples being within neighboring blocks of a co-located block in a reference color channel / frame and / or within a co-located block in a reference color channel / frame; determining (1206) at least one fixed value to be subtracted from sample values of samples in the reference area before determining parameters of a cross-component prediction model; predicting (1208) the target sample of at least one color channel of the image block unit using the determined parameters of the cross-component prediction model; and adding (1210) the at least one fixed value to the value of the predicted target sample of at least one color channel of the image block.
[0310] Therefore, in the method, fixed values can be subtracted from luma and / or chroma sample values used to determine or train parameters of a cross-component model. As a result, the cross-component model estimates the relationship between samples in the input (e.g., luma) and target (e.g., chroma) color planes, where the offset corresponds to the subtracted sample value. When performing cross-component prediction, the same fixed sample value subtracted from the reference input (luma) sample is also subtracted from the input sample used for output (chroma) sample prediction. Therefore, the bit depth of the data used in cross-component prediction is reduced. In addition, the fixed sample value subtracted from the target output (chroma) sample during training will be added to the intermediate output of the cross-component prediction to restore the complete predicted output (color difference) sample value.
[0311] In the following, various embodiments are mainly described in the context of multi-model CCLM and CCCM methods. It needs to be understood that these two models are used only as examples, and the method and its embodiments can be applied to any image and video coding tool with similar concepts as the cross-component prediction method.
[0312] According to an embodiment, the method comprises predicting the target sample of at least one color channel of the image block unit using an intermediate sample value obtained by subtracting the at least one fixed value from the sample value of the sample in the reference area.
[0313] Thus, prediction of a target sample for at least one color channel (e.g., chrominance) may be performed using sample values from which at least one fixed value has been subtracted. Note that this is an optional processing step, as the effect of the removed sample values may already be integrated into the model itself.
[0314] According to an embodiment, the method comprises separately calculating said at least one fixed value for at least one chrominance channel and said luma channel, wherein an average value of neighboring reconstructed samples used for training the cross-component prediction model is the fixed value.
[0315] The fixed values can be calculated using different methods. One alternative is to calculate the average of neighboring reconstruction samples, which are used to train the model and collect data separately for luminance and chrominance.
[0316] According to an embodiment, the method comprises using a subset of said adjacent reconstructed samples for calculating said at least one fixed value.Thus, a limited number of adjacent reconstructed samples, for example, every other sample, or two corner samples out of four corner samples in a neighboring region.
[0317] According to an embodiment, the subset of neighboring reconstructed samples comprises one predefined sample from the reference region.Thus, the value of a single sample in a specified position in the reference region, eg the top left neighboring sample of the current block, may be used as a fixed value.
[0318] According to an embodiment, the method comprises sub-sampling the reference samples; and using one or more sub-sampled reference samples for determining at least one fixed value.
[0319] Therefore, a sub-sampling process may also be used to determine the sample value or set of sample values used to derive the fixed sample value. For example, the same sub-sampling process used to derive the reference samples of the autocorrelation matrix and the cross-component vector may be used.
[0320] Another method is to use all or part of the reconstructed samples in the current luma block to calculate the average value of the luma value. For example, the fixed value can be set to the average value of the four corner samples of the current block.
[0321] According to an embodiment, the method comprises signaling the at least one fixed value in or along the bitstream for each frame, slice or coding unit.
[0322] The fixed value may be signaled in or along the bitstream for each frame, slice, or CTU. The fixed value may be calculated at the encoder side as the average of the sample values in each frame, slice, or CTU. As another example, the fixed value may be calculated at the encoder side as the average of the raw values in the selected region, or a histogram of the raw values in the region may be analyzed.
[0323] The selection of the at least one fixed value may depend on the block size. For example, for narrow blocks, it may be beneficial to select the fixed value from one or more of the reference samples in the larger side of the block, because those samples may have a better correlation with the samples within the block. In an alternative example, the at least one fixed value may be selected from one or more of the reference samples in the smaller side of the block. The selection of the sample position may also be a default setting, for example, defined in a coding standard. The at least one fixed value may also be calculated by analyzing a histogram of the reference samples or a selected subset of the reference samples.
[0324] According to an embodiment, the method comprises subtracting a block-based fixed value from the values of said at least one chroma channel and said luma channel.
[0325] For example, a first fixed value may be subtracted from the luma, a second fixed value may be subtracted from the chroma value of the first chroma component (e.g., Cb), and a third fixed value may be subtracted from the chroma value of the second chroma component (e.g., Cr). During the training and data collection steps, the luma and chroma fixed values may be subtracted from all luma and chroma neighboring reconstructed samples of the current block to calculate the average value of the removed training samples (i.e., R(i,c)=z(i,c)–z0, and s(i)=p(i)–p0). Subsequently, autocorrelation and cross-correlation data are collected. Similarly, a fixed luma value is subtracted from the reconstructed luma sample values within the current block, and the chroma prediction samples are calculated using the CCCM parameters and the luma samples from which the fixed values have been removed. The fixed chroma value may be added to the chroma prediction sample according to the following equation, where z and z0 are the reconstructed luma value and the luma fixed value, respectively, and p0 is the fixed chroma value.
[0326]
[0327] According to an embodiment, the method comprises subtracting the at least one fixed value from the sample values before applying the autocorrelation matrix and the cross-correlation vector used in the cross-component prediction model.
[0328] Therefore, during the autocorrelation and cross-correlation data, fixed value removal can be performed during the pre-processing step of modifying the reconstructed luma and chroma samples instead of subtracting at least one fixed value from the sample values. For example, in the case of non-4:4:4 image content (e.g., 4:2:0) content, the luma samples need to be downsampled to the chroma sampling grid. In this case, fixed value removal can be combined with this downsampling process. As another example, fixed value removal can be integrated into applying filtering (e.g., low pass filtering) to the neighboring and current block sample values.
[0329] As an example, a typical downsampling process for luma samples can be defined as a 6-tap filter with the reconstructed luma samples r(x,y) as input and the downsampled sample values z(i) as output, as follows:
[0330] z(i)=(r(x-1,y)+2*r(x,y)+r(x+1,y)+r(x-1,y+1)+2*r(x,y+1)+r(x+1,y+1)+4) / 8
[0331] where i may be the index of a training sample in the training region (eg, in raster order).
[0332] This can be modified to include the removal of the fixed value f, for example, as follows:
[0333] z(i)=(r(x-1,y)+2*r(x,y)+r(x+1,y)+r(x-1,y+1)+2*r(x,y+1)+r(x+1,y+1)+4) / 8-f
[0334] Advantageously, the fixed value can be combined with the rounding term to form a new rounding term f'. In this way, the computational complexity of the downsampling operation remains unchanged, but the fixed value f is removed from the downsampling operation result. For example, if the rounding value is 4 and the scale is 8, as shown in the above example, the new rounding value f' can be determined as:
[0335] f'=4–8*f
[0336] And the downsampling function can be determined as:
[0337] z(i)=(r(x-1,y)+2*r(x,y)+r(x+1,y)+r(x-1,y+1)+2*r(x,y+1)+r(x+1,y+1)+f') / 8
[0338] In a more general form, the rounding term f' integrated with the sample removal function can be given as a function of the sum of the filter taps S as follows:
[0339] f'=S / 2–S*f
[0340] Having removed the fixed values from the reconstructed luma and chroma values, the predicted sample values can be calculated as follows:
[0341]
[0342] According to an embodiment, the method comprises calculating a term based on fixed values of said at least one chrominance channel and said luma channel; and subtracting said term from the values of the autocorrelation matrix and the cross-correlation vector used in the cross-component prediction model.
[0343] Therefore, removing the effects of fixed luminance and / or chrominance values from the reference samples can also be achieved by updating the initial autocorrelation matrix a and the cross-correlation vector y. This can be performed by subtracting the terms calculated based on the fixed luminance and chrominance values from the autocorrelation matrix a and the cross-correlation vector y.
[0344] It is expected that setting the fixed value to the mean value of the training samples during the parameter derivation process gives the best performance in terms of reducing data bit depth and arithmetic operations. On the other hand, calculating the mean value of the training samples requires an additional processing stage before calculating the autocorrelation and cross-correlation data. In addition, calculating the mean value may require a large number of calculations, especially for large blocks with a large number of training samples.
[0345] In an embodiment, after collecting data and before the solver process, the autocorrelation and cross-correlation data can be modified according to fixed brightness and / or chromaticity values. Therefore, the bit depth of the data is reduced before starting the solver process. First, fixed values of brightness and chromaticity are selected using any of the above methods. One of the input samples (e.g., z[constIdx]) can be set to a constant or bias value, such as CONST, which can be set to a mid-range value (e.g., 512 for 10-bit samples), a power of 1, 2, or any integer value. Another input sample (e.g., z[nonLinearIdx]) can be set to the square of the center brightness sample. Therefore, A[c][constIdx] / CONST is the sum of z[c] (i.e., the center or adjacent brightness training sample). For example, if z[centerIdx] represents the center luma point, then A[centerIdex][constIdx] / CONST is the sum of the luma center samples, and if z[nonLinearIdx] represents the square of the center luma samples, then A[nonLinearIdx][constIdx / CONST is the sum of the square of the luma center samples. Similarly, y[constIdx] / CONST is the sum of the p (i.e., chroma training samples) values. The autocorrelation and cross-correlation data can then be modified as follows.
[0346] Update cross-correlation data:
[0347]
[0348] Update the autocorrelation data:
[0349]
[0350]
[0351] In the above-mentioned update process, autocorrelation data with constant value (i.e. A[c][constIdx]) is used in both update processes. If the update process is performed in a platform with serial execution (such as a CPU), and the autocorrelation data stored in the memory is reused as described above, it is preferred to first perform the update cross-correlation data, and then perform the update autocorrelation data. In addition, the component indexed as constIdxis is preferably processed as the last component. This means that the index "constIdx" can be set to M-1.
[0352] The fixedLuma of the center and neighboring luma samples can be set to the same value calculated using the following equation:
[0353] For c = {0..M-3}, 0 and M-3 are included, fixedLuma[c] = A[centerIdx][constIdx] / (CONST*sampleNum).
[0354] In this equation, it is assumed that the center brightness sample value is labeled with the index "center", and for implementation purposes, "center" may be set to 0. As described above, the division operation may be performed using DIV(x, y).
[0355] The fixedLuma value of the nonlinear term can be calculated by applying the nonlinear function to fixedLuma[centerIdx]. For example, when the nonlinear function is a square or square root, the fixed value of the nonlinear term can be set to the square or square root of fixedLuma[centerIdx], respectively. Alternatively, the fixedLuma of the nonlinear function can be calculated as the average of the nonlinear term of the training data, as shown below:
[0356] FixedLuma[nonLinearIdx] =
[0357] A[nonLinearIdx][constIdx] / (CONST*sampleNum)
[0358] The fixedLuma value of the constant term can be set to 0.
[0359] In the case where CONST (e.g., 512) is a power of 2, the division by CONST (i.e., " / CONST") in the above equation can be implemented by a right shift operation as follows, where the CONST_BITS logarithm of CONST is based on 2 (e.g., when CONST is 512, CONST_BITS is 9):
[0360] FixedLuma[nonLinearIdx]=(A[nonLinearIdx][constIdx] / sampleNum)>>CONST_BITS
[0361] According to an embodiment, a method comprises subtracting at least one fixed value of said luma channel during a downsampling process; and subtracting at least one fixed value of said at least one chroma channel during parameter derivation of a cross-component prediction model.
[0362] Particularly in the case of non-4:4:4 color formats, it may be beneficial to remove fixed luma values from luma samples during the downsampling process, and to remove fixed chroma values from chroma samples during parameter derivation. In this way, luma value subtraction may be performed without additional sample-level operations and may be performed without additional operations during parameter derivation. Furthermore, chroma value subtraction may then be performed without sample-level operations and with minimal additional operations during parameter derivation.
[0363] As another example, luma sample subtraction can be performed during the downsampling process, and chroma sample removal can be omitted entirely. This can reduce the computational overhead to almost zero compared to full bit depth operation, while still providing a reasonable bit depth reduction as the magnitude of luma-related variables is reduced in the process.
[0364] As another example, luma sample subtraction may be performed during the downsampling process, and chroma sample removal may be performed during the parameter derivation process, which only requires updating cross-correlation parameters (no autocorrelation parameters need to be updated), which has lower computational complexity, as shown below:
[0365]
[0366] In some cases, it may be beneficial to limit the range of fixed values. For example, if the fixed value is close to zero or close to the maximum sample value defined by the bit depth of the sample (e.g., 1023 for 10-bit samples), then after the removal is performed, removing such fixed values has little effect on reducing the maximum amplitude of the sample. Therefore, it is advantageous to limit or clip the fixed value to a specific range of values. For example, if the sample is represented by N bit values, the fixed value can be limited so that its value is not less than 1 / 4 of 2^N and not greater than 3 / 4 of 2^N. As another example, the fixed value can be limited so that its value is not less than 1 / 8 of 2^N and not greater than 7 / 8 of 2^N. That is, for a 10-bit sample with a maximum sample value of 1023, the fixed value can be limited to always be between 256 and 768, or between 128 and 896.
[0367] In an embodiment, the fixed value is limited or clipped so that its value is within a range determined based on the bit depth of the sample value.
[0368] In an embodiment, the fixed value is limited or clipped so that its value is within a range determined based on the bit depth of the sample value, wherein the range does not include values below a first threshold and above a second threshold.
[0369] As another example, the effect of removing sample values can be embedded in the coefficients of the convolution model. For example, the convolution prediction operation can be determined as the filter coefficients x i , input sample z i The sum of the products of the sample value z0 removed is as follows:
[0370]
[0371] Now, the equation can be rewritten in a form that separates the term that depends on the input samples from the part (-q0+p0) that is a constant for the block to be predicted:
[0372]
[0373] Advantageously, the constant part p0-q0 can be pre-computed before predicting the block and sample-level addition operations can be avoided. If the convolution model also includes a constant offset or so-called bias term, the constant part can be further integrated in the filter coefficients associated with the bias term. For example, if the bias term is related to the coefficient x n-1 Related to, and related to input parameter z n-1 is a constant, then the coefficient x n-1 It can be updated as follows:
[0374] x n-i =x n-i +(p0-q0) / z n-i
[0375] Or in the case where q0 is zero (for example, when luma samples have been removed from the input samples before convolution, but the effect of the removed chroma samples needs to be added back to the output of the prediction process):
[0376] x n-i =x n-i +p0 / z n-i
[0377] In a typical implementation, with the bias term z n-1 The relevant input parameters are chosen to be powers of two so that the division operation can be conveniently implemented using bit shift operations.
[0378] The bit depth issues of data and arithmetic operations mainly occur in large block sizes with a large number of training samples and are therefore less likely to occur in small block sizes. Also, more clock cycles or time can be allocated to encode and decode larger block sizes.
[0379] In an embodiment, fixed value removal may be applied only to blocks larger than a certain size, or blocks with a number of training samples greater than a predefined threshold. Similarly, fixed value removal may be skipped in smaller block sizes where execution time and allocated processing power are limited. By adopting this mechanism, the worst-case execution time of the method can be controlled. Block size limits (height and / or width) may have default values, such as specified in the specification of a coding standard. There may be signaling mechanisms of various granularities (e.g., by sequence, by frame, by CTU, etc.) that indicate block size limits for applying fixed value removal. The limits may be defined according to a codec profile or at the decoder level. Such a signaling mechanism may provide flexibility to encoders and decoders with higher processing power to use or not use such methods in encoding and decoding pipelines.
[0380] According to one aspect, an apparatus includes a component for receiving an image block unit of a frame, the image block unit including samples in color channels including at least one chrominance channel and one luminance channel; a component for reconstructing samples of the luminance channel of the image block unit; a component for determining a reference area for predicting a target sample of at least one color channel of the image block unit, wherein the reference area includes one or more reference samples of reference samples in neighboring blocks in a current color channel / frame, the reference samples being in neighboring blocks of a reference color channel / frame and / or being within a co-located block in a reference color channel / frame; a component for determining at least one fixed value to be subtracted from sample values of samples in the reference area before determining parameters of a cross-component prediction model; a component for predicting the target sample of at least one color channel of the image block unit using the determined parameters of the cross-component prediction model; and a component for adding the at least one fixed value to the value of the predicted target sample of at least one color channel of the image block.
[0381] According to an embodiment, the apparatus comprises means for separately calculating said at least one fixed value for at least one chrominance channel and said luma channel, wherein an average value of neighboring reconstructed samples used for training the cross-component prediction model is the fixed value.
[0382] According to an embodiment, the apparatus comprises means for calculating said at least one fixed value using a subset of said neighboring reconstructed samples.
[0383] According to an embodiment, the subset of neighboring reconstructed samples comprises a predefined sample from a reference region.
[0384] According to an embodiment, the apparatus comprises means for sub-sampling the reference samples; and means for determining at least one fixed value using one or more sub-sampled reference samples.
[0385] According to an embodiment, the apparatus comprises means for signaling the at least one fixed value in or along the bitstream per frame, slice or coding unit.
[0386] According to an embodiment, the device comprises means for subtracting a block-based fixed value from the values of said at least one chroma channel and said luma channel.
[0387] According to an embodiment, the apparatus comprises means for subtracting the at least one fixed value from the sample values before applying the autocorrelation matrix and the cross-correlation vector used in the cross-component prediction model.
[0388] According to an embodiment, the device comprises means for calculating a term based on fixed values of said at least one chrominance channel and said luma channel; and means for subtracting said term from the values of the autocorrelation matrix and the cross-correlation vector used in the cross-component prediction model.
[0389] According to an embodiment, the device comprises means for subtracting at least one fixed value of the luma channel during a downsampling process; and means for subtracting at least one fixed value of the at least one chroma channel during parameter derivation of a cross-component prediction model.
[0390] According to an embodiment, the apparatus comprises means for predicting said target sample of at least one color channel of an image block unit using intermediate sample values obtained by subtracting said at least one fixed value from sample values of samples in said reference area.
[0391] As a further aspect, a device is provided, comprising: at least one processor and at least one memory, the at least one memory having code stored thereon, which, when executed by the at least one processor, causes the device to at least perform: receiving an image block unit of a frame, the image block unit comprising samples in a color channel, wherein the color channel comprises at least one chrominance channel and one luminance channel; reconstructing samples of the luminance channel of the image block unit; determining a reference area for predicting a target sample of at least one color channel of the image block unit, wherein the reference area comprises one or more reference samples of reference samples in neighboring blocks in a current color channel / frame, the reference samples being in neighboring blocks of a reference color channel / frame and / or in a co-located block in a reference color channel / frame; determining at least one fixed value to be subtracted from sample values of samples in the reference area before determining parameters of a cross-component prediction model; predicting the target sample of at least one color channel of the image block unit using the determined parameters of the cross-component prediction model; and adding the at least one fixed value to the value of the predicted target sample of at least one color channel of the image block.
[0392] According to an embodiment, the apparatus comprises code causing the apparatus to calculate the at least one fixed value separately for at least one chrominance channel and the luma channel, wherein an average value of neighboring reconstructed samples used for training a cross-component prediction model is the fixed value.
[0393] According to an embodiment, the apparatus comprises code causing the apparatus to calculate the at least one fixed value using a subset of the adjacent reconstructed samples.
[0394] According to an embodiment, the subset of neighboring reconstructed samples comprises a predefined sample from a reference region.
[0395] According to an embodiment, an apparatus comprises code causing the apparatus to subsample the reference samples; and means for determining at least one fixed value using one or more subsampled reference samples.
[0396] According to an embodiment, the apparatus comprises code causing the apparatus to signal the at least one fixed value in or along the bitstream for each frame, slice or coding unit.
[0397] According to an embodiment, the apparatus comprises code causing the apparatus to subtract a block-based fixed value from values of the at least one chroma channel and the luma channel.
[0398] According to an embodiment, the apparatus comprises code causing the apparatus to subtract the at least one fixed value from the sample values before applying the autocorrelation matrix and the cross-correlation vector used in the cross-component prediction model.
[0399] According to an embodiment, the device comprises code causing the device to calculate a term based on fixed values of the at least one chrominance channel and the luma channel; and means for subtracting the term from the values of the autocorrelation matrix and the cross-correlation vector used in the cross-component prediction model.
[0400] According to an embodiment, the device comprises code causing the device to subtract at least one fixed value of the luma channel during a downsampling process; and means for subtracting at least one fixed value of the at least one chroma channel during parameter derivation of a cross-component prediction model.
[0401] According to an embodiment, the apparatus comprises code causing the apparatus to predict the target sample of at least one color channel of the image block unit using an intermediate sample value obtained by subtracting the at least one fixed value from the sample values of the samples in the reference area.
[0402] Such devices may include, for example Figure 1 , Figure 2 , Figure 4a and Figure 4b The functional units disclosed in the for implementing the embodiments.
[0403] Such an apparatus further comprises code stored in the at least one memory, which, when executed by the at least one processor, causes the apparatus to perform one or more of the embodiments disclosed herein.
[0404] Fig.1315 is a graphical representation of an example multimedia communication system, in which various embodiments can be implemented. Data source 1510 provides source signal in analog, uncompressed digital or compressed digital format or any combination of these formats. Encoder 1520 may include pre-processing, such as data format conversion and / or filtering of source signal, or be connected with pre-processing. Encoder 1520 encodes source signal into coded media bit stream. It should be noted that the bit stream to be decoded can be directly or indirectly received from a remote device located in almost any type of network. In addition, the bit stream can be received from local hardware or software. Encoder 1520 can encode more than one media type (such as audio and video), or more than one encoder 1520 may be needed to encode different media types of source signal. Encoder 1520 can also obtain synthetically generated input, such as graphics and text, or it can produce a coded bit stream of synthetic media. In the following, in order to simplify the description, only consider a coded media bit stream for processing a media type. However, it should be noted that usually real-time broadcast services include several streams (usually at least one audio, video and text subtitle stream). It should also be noted that the system may include many encoders, but only one encoder 1520 is shown in the figure to simplify the description without lack of generality. It should be further understood that although the text and examples contained herein may specifically describe the encoding process, those skilled in the art will understand that the same concepts and principles also apply to the corresponding decoding process, and vice versa.
[0405] The coded media bitstream can be delivered to the storage device 1530. The storage device 1530 may include any type of mass storage to store the coded media bitstream. The format of the coded media bitstream in the storage device 1530 may be a basic self-contained bitstream format, or one or more coded media bitstreams may be encapsulated into a container file, or the coded media bitstream may be encapsulated into a segmented format suitable for DASH (or a similar streaming system) and stored as a segmented sequence. If one or more media bitstreams are encapsulated in a container file, a file generator (not shown in the figure) may be used to store one or more media bitstreams in a file and create file format metadata, which may also be stored in the file. The encoder 1520 or the storage device 1530 may include a file generator, or the file generator may be operably attached to the encoder 1520 or the storage 1530. Some systems operate in real time, i.e., storage is omitted, and the coded media bitstream is transmitted directly from the encoder 1520 to the transmitter 1540. The coded media bitstream can then be transmitted to the transmitter 1540, also referred to as a server, as needed. The format used in the transmission may be a basic self-contained bitstream format, a packetized stream format, a segmented format suitable for DASH (or similar streaming systems), or one or more coded media bitstreams may be encapsulated into a container file. The encoder 1520, storage device 1530, and server 1540 may reside in the same physical device, or they may be included in separate devices. The encoder 1520 and server 1540 may operate with real-time live content, in which case the coded media bitstreams are typically not stored permanently, but are cached for a short period of time in the content encoder 1520 and / or server 1540 to smooth out processing delays, delivery delays, and variations in the coded media bitrate.
[0406] The server 1540 sends the coded media bitstream using a communication protocol stack. The stack may include, but is not limited to, one or more of the Real-time Transport Protocol (RTP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Transmission Control Protocol (TCP), and Internet Protocol (IP). When the communication protocol stack is packet-oriented, the server 1540 encapsulates the coded media bitstream into packets. For example, when RTP is used, the server 1540 encapsulates the coded media bitstream into RTP packets according to the RTP payload format. Typically, each media type has a dedicated RTP payload format. It should be noted again that the system may include more than one server 1540, but for simplicity, the following description only considers one server 1540.
[0407] If the media content is encapsulated in a container file for storage device 1530 or for inputting data to transmitter 1540, transmitter 1540 may include or be operably attached to a "sending file parser" (not shown in the figure). In particular, if the container file is not transmitted in this way, but at least one of the contained coded media bitstreams is encapsulated for transmission via a communication protocol, the sending file parser locates the appropriate portion of the coded media codestream to be transmitted via the communication protocol. The sending file parser can also help create the correct format for the communication protocol, such as data packet headers and payloads. The multimedia container file may contain encapsulation instructions, such as a hint track in ISOBMFF, for encapsulating at least one of the contained media bitstreams on the communication protocol.
[0408] Server 1540 may or may not be connected to gateway 1550 via a communication network, which may be, for example, a CDN, the Internet, and / or a combination of one or more access networks. A gateway may also or alternatively be referred to as a middlebox. For DASH, a gateway may be an edge server (of a CDN) or a network proxy. Note that the system may typically include any number of gateways, etc., but for simplicity, the following description only considers one gateway 1550. Gateway 1550 may perform different types of functions, such as converting a packet stream from one communication protocol stack to another communication protocol stack, merging and forking data streams, and manipulating data streams according to downlink and / or receiver capabilities, such as controlling the bit rate of the forwarded stream according to the prevailing downlink network conditions. In various embodiments, gateway 1550 may be a server entity.
[0409] The system includes one or more receivers 1560, which are generally capable of receiving, demodulating and decapsulating the transmitted signal into a coded media bitstream. The coded media bitstream can be delivered to a recording storage device 1570. The recording storage device 1570 can include any type of mass storage to store the coded media bitstream. The recording storage device 1570 can alternatively or additionally include a computing memory, such as a random access memory. The format of the coded media bitstream in the recording storage device 1570 can be a basic self-contained bitstream format, or one or more coded media bitstreams can be encapsulated into a container file. If there are multiple coded media bitstreams associated with each other, such as audio streams and video streams, a container file is generally used, and the receiver 1560 includes or is attached to a container file generator that generates the container file from the input stream. Some systems operate in "real time", that is, the recording storage device 1570 is omitted, and the coded media bitstream is delivered directly from the receiver 1560 to the decoder 1580. In some systems, only the most recent portion of the recorded stream, such as an excerpt of the most recent 10 minutes of the recorded stream, is saved in the recorded storage device 1570, while any earlier recorded data is discarded from the recorded storage device 1570.
[0410] The coded media bitstream may be passed from the recording storage device 1570 to the decoder 1580. If there are many coded media bitstreams, such as audio and video streams, which are associated with each other and encapsulated in a container file, or a single media bitstream is encapsulated in a container file (for example, for easier access), a file parser (not shown in the figure) is used to decapsulate each coded media bitstream from the container file. The recording storage device 1570 or the decoder 1580 may include a file parser, or the file parser is attached to either the recording storage device 1570 or the decoder 1580. It should also be noted that the system may include many decoders, but only one decoder 1570 is discussed here to simplify the description without loss of generality.
[0411] The coded media bitstream may be further processed by a decoder 1570, the output of which is one or more uncompressed media streams. Finally, a renderer 1590 may reproduce the uncompressed media streams, for example, using a speaker or display. The receiver 1560, the recording storage 1570, the decoder 1570, and the renderer 1590 may reside in the same physical device, or they may be included in separate devices.
[0412] The transmitter 1540 and / or the gateway 1550 may be configured to perform switching between different representations, for example, for switching between different viewing areas of 360-degree video content, view switching, bit rate adaptation and / or fast start, and / or the transmitter 1540 and / or the gateway 1550 may be configured to select a representation that has been transmitted. The switching between different representations may be for a variety of reasons, such as in response to a request from the receiver 1560 or the prevailing conditions of the network that transmits the bitstream, such as throughput. In other words, the receiver 1560 may initiate switching between representations. The request from the receiver may be, for example, a request for a segment or sub-segment from a different representation than previously, a request to turn off a scalability layer and / or sub-layer that has been transmitted, or a change to a rendering device with different capabilities than previously. The request for the segment may be an HTTP GET request. The request for the sub-segment may be an HTTP GET request with a byte range. Additionally or alternatively, bit rate adjustment or bit rate adaptation may be used, for example, in streaming services to provide so-called fast start, where upon startup or random access to a stream, the bit rate of the transport stream is lower than the channel bit rate in order to start playback immediately and achieve a buffer occupancy level that tolerates occasional packet delays and / or retransmissions. The bit rate adaptation may include multiple representation or layer up-switching and representation or layer down-switching operations occurring in various orders.
[0413] Decoder 1580 can be configured to perform switching between different representations, such as for switching between different viewports of 360-degree video content, view switching, bit rate adaptation and / or fast start, and / or decoder 1580 can be configured to select the representation sent. Switching between different representations may be for a variety of reasons, such as to achieve faster decoding operations, or to adapt the transmitted bitstream (e.g., in terms of bit rate) to the prevailing conditions of the network transmitting the bitstream, such as throughput. For example, if a device including decoder 1580 is multitasking and uses computing resources for purposes other than decoding video bitstreams, faster decoding operations may be required. In another example, when content is played at a speed faster than normal playback speed, such as two to three times faster than conventional real-time playback rates, faster decoding operations may be required.
[0414] In the above, some embodiments have been described with reference to and / or using the terminology of HEVC and / or VVC. It should be understood that the embodiments can be similarly implemented with any video encoder and / or video decoder.
[0415] In the above, where example embodiments are described with reference to an encoder, it is to be understood that the resulting bitstream and decoder may have corresponding elements therein. Similarly, where example embodiments are described with reference to a decoder, it is to be understood that the encoder may have a structure and / or computer program for generating a bitstream to be decoded by a decoder. For example, some embodiments have been described in connection with generating prediction blocks as part of encoding. Embodiments may be similarly implemented by generating prediction blocks as part of decoding, except that encoding parameters, such as horizontal offsets and vertical offsets, are decoded from the bitstream rather than determined by the encoder.
[0416] The embodiments of the present invention described above describe the codec in terms of separate encoder and decoder devices to help understand the processes involved. However, it should be understood that the device, structure and operation can be implemented as a single encoder-decoder device / structure / operation. In addition, the encoder and decoder may share some or all common elements.
[0417] Although the above examples describe embodiments of the present invention operating within a codec within an electronic device, it should be understood that the present invention defined in the claims can be implemented as part of any video codec. Thus, for example, embodiments of the present invention can be implemented in a video codec that can implement video encoding via a fixed or wired communication path.
[0418] Thus, the user equipment may comprise a video codec such as those described in the embodiments of the invention above.It will be appreciated that the term user equipment is intended to cover any suitable type of wireless user equipment such as a mobile phone, a portable data processing device or a portable web browser.
[0419] Furthermore, elements of a public land mobile network (PLMN) may also include a video codec as described above.
[0420] In general, various embodiments of the present invention may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be executed by a controller, microprocessor, or other computing device, although the present invention is not limited thereto. Although various aspects of the present invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it is well understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof.
[0421] Embodiments of the present invention can be implemented by computer software that can be executed by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. In addition, in this regard, it should be noted that any block of the logic flow shown in the figure can represent a program step, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software can be stored on a physical medium such as a memory chip or a memory block implemented in a processor, on a magnetic medium such as a hard disk or a floppy disk, and on an optical medium such as a DVD and its data variants, a CD.
[0422] The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and a processor based on a multi-core processor architecture, as non-limiting examples.
[0423] Embodiments of the present invention may be practiced in various components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0424] Programs, such as those offered by Synopsys Inc. of Mountain View, Calif., and Cadence Design of San Jose, Calif., automatically route conductors and position components on a semiconductor chip using generally accepted design rules along with a library of pre-stored design blocks. Once the design of a semiconductor circuit is complete, the final design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be transmitted to a semiconductor manufacturing facility or "fab" for fabrication.
[0425] The above description provides a complete and informative description of exemplary embodiments of the present invention by way of exemplary and non-limiting examples. However, various modifications and adaptations may be apparent to those skilled in the relevant art in view of the above description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the present invention's teachings will still fall within the scope of the present invention.
Claims
1. A device comprising: means for receiving an image block unit of a frame, the image block unit comprising samples in color channels, wherein the color channels include at least one chrominance channel and one luma channel; means for reconstructing samples of said luminance channel of said image block unit; means for determining a reference region of target samples for predicting at least one color channel of the image block unit, wherein the reference region comprises one or more reference samples of reference samples in neighboring blocks in a current color channel / frame, the reference samples being in the neighboring blocks of a co-located block in a reference color channel / frame and / or being within the co-located block in the reference color channel / frame; means for determining at least one fixed value to be subtracted from sample values of samples in said reference region prior to determining parameters of a cross-component prediction model; means for predicting the target sample of at least one color channel of the image block unit using the determined parameters of the cross-component prediction model; as well as Means for adding said at least one fixed value to a predicted value of said target sample of at least one color channel of said image block.
2. The device according to claim 1, comprising: Means for separately calculating the at least one fixed value for at least one chrominance channel and the luma channel, wherein an average value of adjacent reconstructed samples used for training the cross-component prediction model is the fixed value.
3. The device according to claim 1 or 2, comprising: Means for calculating said at least one fixed value using a subset of said adjacent reconstructed samples. 4 . The apparatus of claim 3 , wherein the subset of the adjacent reconstructed samples comprises a predefined sample from the reference region.
5. An apparatus according to any preceding claim, comprising: means for sub-sampling said reference sample; as well as Means for determining said at least one fixed value using one or more sub-sampled said reference samples.
6. Apparatus according to any preceding claim, comprising: Means for signaling the at least one fixed value in or along the bitstream for each frame, slice or coding unit.
7. Apparatus according to any preceding claim, comprising: Means for subtracting a block-based fixed value from values of said at least one chroma channel and said luma channel.
8. Apparatus according to any preceding claim, comprising: Means for subtracting said at least one fixed value from said sample values prior to applying the autocorrelation matrix and cross-correlation vector used in said cross-component prediction model.
9. The device according to any one of claims 2 to 7, comprising: means for calculating a term based on said fixed values for said at least one chroma channel and said luma channel; as well as Means for subtracting said terms from values of said autocorrelation matrix and said cross-correlation vector used in said cross-component prediction model.
10. The device according to any one of claims 2 to 7, comprising: means for subtracting said at least one fixed value of said luma channel during a downsampling process; as well as Means for subtracting said at least one fixed value of said at least one chroma channel during derivation of parameters of said cross-component prediction model.
11. Apparatus according to any preceding claim, comprising: A means for predicting the target sample of at least one color channel of the image block unit using an intermediate sample value, the intermediate sample value being obtained by subtracting the at least one fixed value from the sample value of the sample in the reference area.
12. A method comprising: Receiving an image block unit of a frame, the image block unit comprising samples in a color channel, wherein the color channel comprises at least one chroma channel and a luma channel; Reconstructing samples of the brightness channel of the image block unit; Determine a reference area for predicting target samples of at least one color channel of the image block unit, wherein the reference area includes one or more reference samples of reference samples in neighboring blocks in a current color channel / frame, the reference samples being in the neighboring blocks of a co-located block in a reference color channel / frame and / or being within the co-located block in the reference color channel / frame; determining at least one fixed value to be subtracted from sample values of samples in the reference region before determining parameters of the cross-component prediction model; predicting the target sample of at least one color channel of the image block unit using the determined parameters of the cross-component prediction model; as well as The at least one fixed value is added to a predicted value of the target sample of at least one color channel of the image block.
13. The method according to claim 12, comprising: The at least one fixed value is calculated separately for at least one chrominance channel and the luma channel, wherein an average value of adjacent reconstructed samples used for training the cross-component prediction model is the fixed value.
14. The method according to claim 12 or 13, comprising: The at least one fixed value is calculated using a subset of the adjacent reconstructed samples.
15. The method according to any one of claims 12 to 14, comprising: The target sample of at least one color channel of the image block unit is predicted using an intermediate sample value, wherein the intermediate sample value is obtained by subtracting the at least one fixed value from the sample value of the sample in the reference area.