An apparatus, a method and a computer program for video coding and decoding
By integrating inverse luma mapping within the video coding loop and applying it to reconstructed pictures before storage, the method addresses sub-optimal prediction and bit rate issues in existing video coding technologies, achieving improved accuracy and reduced bit rates.
Patent Information
- Application Number
- PCT/EP2024/077327
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-09-27
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video coding methods suffer from sub-optimal prediction and higher bit rates due to the reconstruction of object masks being filtered outside the video coding loop, leading to loss of compression accuracy.
The proposed method involves an apparatus and method for video coding and decoding that includes receiving luma mapping parameters, decoding pictures with indications for luma mapping, and applying inverse luma mapping functions to reconstructed pictures before storing them in a reference picture buffer, ensuring accurate prediction and reduced bit rates.
This approach enhances reference picture accuracy, improves inter-prediction, reduces bit rate requirements, and enhances reconstruction quality by integrating inverse luma mapping within the video coding loop.
Smart Images

Figure EP2024077327_30052025_PF_FP_ABST
Abstract
Description
AN APPARATUS, A METHOD AND A COMPUTER PROGRAM FOR VIDEO CODING AND DECODINGTECHNICAL FIELD
[0001] The present invention relates to an apparatus, a method and a computer program for video coding and decoding.BACKGROUND
[0002] Object masks are used in video encoding and decoding for providing the decoder information about objects detected upon encoding, such as information about object bounding boxes and object labels, so as to reduce the workload and power consumption of the decoder. Object masks can be coded as grayscale images, with each code value identifying one object. Lossless coding of such maps is costly in terms of bit rate.
[0003] Lossy compression would reduce the data amount, however at the cost of mask reconstruction accuracy. Various post-filtering proposals for such mask data nevertheless exist.
[0004] However, the proposed post-filtering methods are performed outside the video coding loop, which thus brings about a major drawback: As a reconstructed picture is filtered outside the coding loop, after the decoding process, it is stored as unfiltered in the reference picture buffer. Consequently, any prediction from this unfiltered picture suffers from the loss introduced due to compression. The result is a sub-optimal prediction and thus higher bit rate and lower reconstruction accuracy.SUMMARY
[0005] Now, improved methods and technical equipment implementing the methods have been invented, by which the above problems are alleviated. Various aspects include methods, apparatuses and a computer readable medium comprising a computer program, or a signal stored therein, which are characterized by what is stated in the independent claims. Various details of the embodiments are disclosed in the dependent claims and in the corresponding images and description.
[0006] The scope of protection sought for various embodiments of the invention is set out by the independent claims. The embodiments and features, if any, described in this specification thatdo not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the invention.
[0007] An apparatus according to a first aspect comprises: means for receiving luma mapping parameters; means for receiving a first coded picture with a first indication about luma mapping being enabled for said first coded picture and a second indication that the luma mapping parameters are used for said first coded picture; means for receiving a third indication about applying only an inverse luma mapping for said first coded picture; means for deriving an inverse luma mapping function based on the luma mapping parameters; means for decoding the first coded picture to a reconstruction of a first picture; means for applying at least the inverse luma mapping function for the reconstruction of the first picture to obtain a reconstructed first picture; means for storing the reconstructed first picture in a reference picture buffer; means for receiving a second coded picture with a fourth indication about luma mapping being enabled for said second coded picture and a fifth indication that the luma mapping parameters are used for said second coded picture; means for decoding a sixth indication about an identify forward luma mapping function being used in inter prediction; and means for decoding the second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for the inter prediction by applying the identity forward luma mapping function.
[0008] According to an embodiment, the apparatus comprises means for applying the identity forward mapping function to any picture subsequent to the first reconstructed picture in the reference picture buffer used as a reference for the inter prediction and applying the inverse luma mapping function for a reconstruction of any picture subsequent to the reconstruction of the first picture before storing said picture in the reference picture buffer.
[0009] According to an embodiment, the first coded picture is an intra-coded picture and the second coded picture is an inter-coded picture.
[0010] According to an embodiment, a luma mapping with chroma scaling (LMCS) adaptation parameter set comprises the luma mapping parameters.
[0011] According to an embodiment, luma mapping parameters comprise valid luma sample values and the apparatus comprises means for deriving the inverse luma mapping function from the valid luma sample values.
[0012] According to an embodiment, said indication about applying only an inverse luma mapping is received in a flag included in a slice header syntax structure.
[0013] According to an embodiment, said indication about applying only an inverse luma mapping is received in a flag included in a picture header syntax structure.
[0014] According to an embodiment, said indication about applying only an inverse luma mapping is received in a flag included in a sequence parameter set syntax structure.
[0015] An apparatus according to a second aspect comprises: at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive luma mapping parameters; receive a first coded picture with a first indication about luma mapping being enabled for said first coded picture and a second indication that the luma mapping parameters are used for said first coded picture; receive a third indication about applying only an inverse luma mapping for said first coded picture; derive an inverse luma mapping function based on the luma mapping parameters; decode the first coded picture to a reconstruction of a first picture; apply at least the inverse luma mapping function for the reconstruction of the first picture to obtain a reconstructed first picture; store the reconstructed first picture in a reference picture buffer; receive a second coded picture with a fourth indication about luma mapping being enabled for said second coded picture and a fifth indication that the luma mapping parameters are used for said second coded picture; decode a sixth indication about an identify forward luma mapping function being used in inter prediction; and decode the second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for the inter prediction by applying the identity forward luma mapping function.
[0016] A method according to a third aspect comprises receiving luma mapping parameters; receiving a first coded picture with a first indication about luma mapping being enabled for said first coded picture and a second indication that the luma mapping parameters are used for said first coded picture; receiving a third indication about applying only an inverse luma mapping for said first coded picture; deriving an inverse luma mapping function based on the luma mapping parameters; decoding the first coded picture to a reconstruction of a first picture; applying at least the inverse luma mapping function for the reconstruction of the first picture to obtain a reconstructed first picture; storing the reconstructed first picture in a reference picture buffer; receiving a second coded picture with a fourth indication about luma mapping being enabled for said second coded picture and a fifth indication that the luma mapping parameters are used forsaid second coded picture; decoding a sixth indication about an identify forward luma mapping function being used in inter prediction; and decoding the second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for the inter prediction by applying the identity forward luma mapping function.
[0017] An apparatus according to a fourth aspect comprises: means for receiving a first picture; means for deriving an inverse luma mapping function; means for encoding luma mapping parameters based on said inverse luma mapping function; means for encoding the first picture to a first coded picture; means for encoding an indication in or along the first coded picture that the luma mapping parameters are applied; means for applying the inverse luma mapping function for a reconstruction of the first picture; means for storing the reconstructed first picture in a reference picture buffer; means for receiving a second picture; means for encoding the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for inter prediction by applying an identity forward mapping function; and means for encoding an indication in or along the second coded picture that the identify forward mapping function being used in the inter prediction.
[0018] According to an embodiment, the apparatus comprises means for deriving a forward luma mapping function based on the inverse luma mapping function; and means for encoding a luma mapping with chroma scaling (LMCS) adaptation parameter set based on said forward luma mapping function.
[0019] According to an embodiment, the first coded picture is an intra-coded picture and the second coded picture is an inter-coded picture.
[0020] According to an embodiment, the apparatus comprises means for deriving the luma mapping parameters from the inverse luma mapping function.
[0021] According to an embodiment, the luma mapping parameters comprise a luma mapping with chroma scaling (LMCS) adaptation parameter set.
[0022] According to an embodiment, the apparatus comprises means for encoding LMCS pivot points in a manner that a valid luma sample value is represented by two LMCS pivot points, which specify a piece in the piece-wise linear inverse luma mapping function in such a manner that a luma sample given as input and mapped to this piece returns approximately or exactly the valid luma sample value.
[0023] According to an embodiment, the apparatus comprises means for deriving the inverse luma mapping function from the valid luma samples.
[0024] According to an embodiment, the apparatus comprises means for sending said indication about applying only an inverse luma mapping in a flag included in a slice header syntax structure.
[0025] According to an embodiment, the apparatus comprises means for sending said indication about applying only an inverse luma mapping in a flag included in a picture header syntax structure.
[0026] According to an embodiment, the apparatus comprises means for sending said indication about applying only an inverse luma mapping in a flag included in a sequence parameter set syntax structure.
[0027] An apparatus according to a fifth aspect comprises: at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive a first picture; derive an inverse luma mapping function; encode luma mapping parameters based on said inverse luma mapping function; encode the first picture to a first coded picture; encode an indication in or along the first coded picture that the luma mapping parameters are applied; apply the inverse luma mapping function for a reconstruction of the first picture; store the reconstructed first picture in a reference picture buffer; receive a second picture; encode the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for inter prediction by applying an identity forward mapping function; and encode an indication in or along the second coded picture that the identify forward mapping function being used in the inter prediction.
[0028] A method according to a sixth aspect comprises receiving a first picture; deriving an inverse luma mapping function; encoding luma mapping parameters based on said inverse luma mapping function; encoding the first picture to a first coded picture; encoding an indication in or along the first coded picture that the luma mapping parameters are applied; applying the inverse luma mapping function for a reconstruction of the first picture; storing the reconstructed first picture in a reference picture buffer; receiving a second picture; encoding the second picture to a second coded picture using the reconstruction of the first picture in the reference picture bufferas a reference for inter prediction by applying an identity forward mapping function; and encoding an indication in or along the second coded picture that the identify forward mapping function being used in the inter prediction.
[0029] The further aspects relate to apparatuses and computer readable storage media stored with code thereon, which are arranged to carry out the above methods and one or more of the embodiments related thereto.BRIEF DESCRIPTION OF THE DRAWINGS
[0030] For better understanding of the present invention, reference will now be made by way of example to the accompanying drawings in which:
[0031] Figures la and lb show schematically an encoder and a decoder suitable for implementing embodiments of the invention;
[0032] Figure 2 shows in-loop filter implementation in versatile Video Coding (WC);
[0033] Figure 3 shows a general illustration of the pipeline of Video Coding for Machines;
[0034] Figures 4a - 4c show examples of object mask data;
[0035] Figures 5a and 5b show examples of using the sample values of an auxiliary picture as the ID of the mask;
[0036] Figure 6 shows a flow chart of encoding object mask auxiliary pictures as layered coding;
[0037] Figures 7a and 7b show examples of single-mask and multiple-mask sequences, respectively;
[0038] Figure 8 shows an example of decoding architecture of luma mapping with chroma scaling (LMCS);
[0039] Figure 9 shows a flow chart of a decoding method according to an embodiment of the invention;
[0040] Figures 10a - 10 c show an example of a result of a mask generation process;
[0041] Figures I la and 11b show flow charts of an encoding and a decoding method according to another embodiment of the invention;
[0042] Figures 12a and 12b show flow charts of an encoding and a decoding method according to yet another embodiment of the invention;
[0043] Figure 13 shows schematically an electronic device suitable for employing embodiments of the invention;
[0044] Figure 14 shows schematically a user equipment suitable for employing embodiments of the invention; and
[0045] Figure 15 shows a schematic diagram of an example multimedia communication system within which various embodiments may be implemented.DETAILED DESCRIPTON OF SOME EXAMPLE EMBODIMENTS
[0046] A video codec consists of an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can uncompress the compressed video representation back into a viewable form. A video encoder and / or a video decoder may also be separate from each other, i.e. need not form a codec. Typically encoder discards some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).
[0047] Figures la and lb show an encoder and decoder for encoding and decoding the 2D pictures. A video codec consists of an encoder that transforms an input video into a compressed representation suited for storage / transmission and a decoder that can uncompress the compressed video representation back into a viewable form. Typically, the encoder discards and / or loses some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).
[0048] An example of an encoding process is illustrated in Figure la. Figure la illustrates an image to be encoded (In); a predicted representation of an image block (P'n); a prediction error signal (Dn); a reconstructed prediction error signal (D'n); a preliminary reconstructed image (I'n); a final reconstructed image (R'n); a transform (T) and inverse transform (T-l); a quantization (Q) and inverse quantization (Q-l); entropy encoding (E); a reference frame memory (RFM); inter prediction (Pinter); intra prediction (Pintra); mode selection (MS) and filtering (F).
[0049] An example of a decoding process is illustrated in Figure lb. Figure lb illustrates a predicted representation of an image block (P'n); a reconstructed prediction error signal (D'n); a preliminary reconstructed image (I'n); a final reconstructed image (R'n); an inverse transform (T-1); an inverse quantization (Q-l); an entropy decoding (E-l); a reference frame memory (RFM); a prediction (either inter or intra) (P); and filtering (F).
[0050] Thus, the decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence. The filtering may for example include one more of the following: deblocking, sample adaptive offset (SAO), and / or adaptive loop filtering (ALF).
[0051] Many hybrid video encoders, such as H.264 / AVC encoders, High Efficiency Video Coding (H.265 / HEVC a.k.a. HEVC) and Versatile Video Coding (H.266 / WC a.k.a. WC) encoders, encode the video information in two phases. Firstly pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate). Video codecs may also provide a transform skip mode, which the encoders may choose to use. In the transform skip mode, the prediction error is coded in a sample domain, for example by deriving a sample-wise difference value relative to certain adjacent samples and coding the sample-wise difference value with an entropy coder.
[0052] In temporal prediction, the sources of prediction are previously decoded pictures (a.k.a. reference pictures). In intra block copy (IBC; a.k.a. intra-block-copy prediction),prediction is applied similarly to temporal prediction but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process. Interlayer or inter-view prediction may be applied similarly to temporal prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal prediction only, while in other cases inter prediction may refer collectively to temporal prediction and any of intra block copy, inter-layer prediction, and inter-view prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0053] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, reduces temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0054] In many video codecs, including H.264 / AVC, HEVC and WC, motion information is indicated by motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder) or decoded (at the decoder) and the prediction source block in one of the previously coded or decoded images (or picture).
[0055] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0056] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organisation for Standardization (ISO) / International ElectrotechnicalCommission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). Extensions of the H.264 / AVC include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0057] High Efficiency Video Coding (H.265 / HEVC a.k.a. HEVC) standard was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard was published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Later versions of H.265 / HEVC included scalable, multiview, fidelity range, three-dimensional, and screen content coding extensions which may be abbreviated SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.
[0058] Versatile Video Coding (H.266 a.k.a. WC), defined in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, (also referred to as MPEG-I Part 3) is a video compression standard developed as the successor to HEVC.
[0059] A specification of the AVI bitstream format and decoding process were developed by the Alliance for Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.
[0060] Some key definitions, bitstream and coding structures, and concepts of WC, are described in this section as an example of a platform, where the embodiments may be implemented. Some of the key definitions, bitstream and coding structures, and concepts of H.266 / WC are the same as in H.264 / AVC and H.265 / HEVC. The embodiments are not limited to H.266 / WC, but rather the description is given for one possible basis on top of which the aspects of the invention and the related embodiments may be partly or fully realized.
[0061] An elementary unit for the input to an encoder and the output of a decoder, respectively, in most cases is a picture. A picture given as an input to an encoder may also be referred to as a source picture, and a picture decoded by a decoder may be referred to as a decoded picture or a reconstructed picture.
[0062] The source and decoded pictures are each comprises of one or more sample arrays, such as one of the following sets of sample arrays:Luma (Y) only (monochrome),Luma and two chroma (YCbCr or YCgCo),Green, Blue and Red (GBR, also known as RGB), Arrays representing other unspecified monochrome or tri-stimulus color samplings (for example, YZX, also known as XYZ).
[0063] In the following, these arrays may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr; regardless of the actual color representation method in use. The actual color representation method in use can be indicated e.g. in a coded bitstream e.g. using the Video Usability Information (VUI) syntax of HEVC or alike. A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma) that compose a picture e.g. in 4:2:0, 4:2:2 or 4:4:4 chroma format or the array or a single sample of the array that compose a picture in monochrome format.
[0064] A picture may be defined to be either frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, when the source signal is interlaced. Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays.
[0065] Some chroma formats (a.k.a color formats) may be summarized as follows:In monochrome sampling there is only one sample array, which may be nominally considered the luma array.In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array.In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array.In 4:4:4 sampling when no separate color planes are in use, each of the two chroma arrays has the same height and width as the luma array.
[0066] Partitioning may be defined as a division of a set into subsets such that each element of the set is in exactly one of the subsets.
[0067] In the following, partitioning a picture into subpictures, slices, and tiles according to H.266 / WC is described more in detail.
[0068] A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of coding tree units (CTU) that covers a rectangular region of a picture. The CTUs in a tile are scanned in raster scan order within that tile.
[0069] A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture. Consequently, each vertical slice boundary is always also a vertical tile boundary. It is possible that a horizontal boundary of a slice is not a tile boundary but consists of horizontal CTU boundaries within a tile; this occurs when a tile is split into multiple rectangular slices, each of which consists of an integer number of consecutive complete CTU rows within the tile.
[0070] Two modes of slices are supported, namely the raster-scan slice mode and the rectangular slice mode. In the raster-scan slice mode, a slice contains a sequence of complete tiles in a tile raster scan of a picture. In the rectangular slice mode, a slice contains either a number of complete tiles that collectively form a rectangular region of the picture or a number of consecutive complete CTU rows of one tile that collectively form a rectangular region of the picture. Tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to that slice.
[0071] A subpicture may be defined as a rectangular region of one or more slices within a picture, wherein the one or more slices are complete. Thus, a subpicture consists of one or more slices that collectively cover a rectangular region of a picture. Consequently, each subpicture boundary is also always a slice boundary, and each vertical subpicture boundary is always also a vertical tile boundary. The slices of a subpicture may be required to be rectangular slices.
[0072] One or both of the following conditions shall be fulfilled for each subpicture and tile:- All CTUs in a subpicture belong to the same tile.- All CTUs in a tile belong to the same subpicture.
[0073] The samples are processed in units of coding tree blocks (CTB). The array size for each luma CTB in both width and height is CtbSizeY in units of samples. The width and height of the array for each chroma CTB are CtbWidthC and CtbHeightC, respectively, in units of samples.
[0074] CTU may be split into smaller CUs using quaternary tree structure. Each CU may be divided using quad-tree and nested multi-type tree including ternary and binary split. There are specific rules to infer partitioning in in picture boundaries. The redundant split patterns are disallowed in nested multi-type partitioning.
[0075] In comparison to the previous video coding standards, such as H.264 / AVC and H.265 / HEVC, the H.266 / VCC introduces some new coding tools, such as:Intra prediction o 67 intra mode with wide angles mode extension o Block size and mode dependent 4 tap interpolation filter o Position dependent intra prediction combination (PDPC) o Cross component linear model intra prediction (CCLM) o Multi-reference line intra prediction o Intra sub-partitions o Weighted intra prediction with matrix multiplicationInter-picture prediction o Block motion copy with spatial, temporal, history-based, and pairwise average merging candidates o Affine motion inter prediction o sub-block based temporal motion vector prediction o Adaptive motion vector resolution o 8x8 block-based motion compression for temporal motion prediction o High precision (1 / 16 pel) motion vector storage and motion compensation with 8- tap interpolation filter for luma component and 4-tap interpolation filter for chroma component o Triangular partitions o Combined intra and inter prediction o Merge with MVD (MMVD) o Symmetrical MVD coding o Bi-directional optical flow o Decoder side motion vector refinement o Bi-prediction with CU-level weightTransform, quantization and coefficients coding o Multiple primary transform selection with DCT2, DST7 and DCT8 o Secondary transform for low frequency zone o Sub-block transform for inter predicted residual o Dependent quantization with max QP increased from 51 to 63 o Transform coefficient coding with sign data hidingo Transform skip residual codingEntropy Coding o Arithmetic coding engine with adaptive double windows probability update In loop filter o In-loop reshaping o Deblocking filter with strong longer filter o Sample adaptive offset o Adaptive Loop FilterScreen content coding: o Current picture referencing with reference region restriction 360-degree video coding o Horizontal wrap-around motion compensation High-level syntax and parallel processing o Reference picture management with direct reference picture list signalling o Tile groups with rectangular shape tile groups
[0076] Some general concepts and definitions applicable to most video coding approaches are given in the following.
[0077] A bitstream may be defined as a sequence of bits.
[0078] A bitstream format may comprise a sequence of syntax structures.
[0079] A syntax element may be defined as an element of data represented in the bitstream. A syntax structure may be defined as zero or more syntax elements present together in the bitstream in a specified order.
[0080] In some coding formats or standards, a bitstream may be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences (CVS).
[0081] In some coding formats, such as AVI, a bitstream may comprise a sequence of open bitstream units (OBUs). An OBU comprises a header and a payload, wherein the header identifies a type of the OBU. Furthermore, the header may comprise a size of the payload in bytes.
[0082] A NAL unit may be defined as a syntax structure containing an indication of the type of data to follow and bytes containing that data in the form of an RBSP interspersed as necessarywith start code emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure containing an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits containing syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0083] NAL units consist of a header and payload. The NAL unit header indicates the type of the NAL unit among other things.
[0084] NAL units can be categorized into Video Coding Layer (VCL) NAL units and non- VCL NAL units. VCL NAL units are typically coded slice NAL units.
[0085] A non-VCL NAL unit may be for example one of the following types: a sequence parameter set, a picture parameter set, a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit. Parameter sets may be needed for the reconstruction of decoded pictures, whereas many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.
[0086] Some coding formats specify parameter sets that may carry parameter values needed for the decoding or reconstruction of decoded pictures. A parameter may be defined as a syntax element of a parameter set. A parameter set may be defined as a syntax structure that contains parameters and that can be referred to from or activated by another syntax structure for example using an identifier.
[0087] Some types of parameter sets are briefly described in the following but it needs to be understood that other types of parameter sets may exist and that embodiments may be applied but are not limited to the described types of parameter sets. A video parameter set (VPS) may include parameters that are common across multiple layers in a coded video sequence or describe relations between layers. Parameters that remain unchanged through a coded video sequence (in a single-layer bitstream) or in a coded layer video sequence may be included in a sequence parameter set (SPS). In addition to the parameters that may be needed by the decoding process, the sequence parameter set may optionally contain video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation. A picture parameter set (PPS) contains such parameters that are likely to be unchanged in several coded pictures. A picture parameter set may include parameters that can be referred to by the coded image segments of one or more coded pictures. A header parameter set(HPS) has been proposed to contain such parameters that may change on picture basis. In VVC, an Adaptation Parameter Set (APS) may comprise parameters for decoding processes of different types, such as adaptive loop filtering or luma mapping with chroma scaling.
[0088] A parameter set may be activated when it is referenced e.g., through its identifier. For example, a header of an image segment, such as a slice header, may contain an identifier of the PPS that is activated for decoding the coded picture containing the image segment. A PPS may contain an identifier of the SPS that is activated, when the PPS is activated. An activation of a parameter set of a particular type may cause the deactivation of the previously active parameter set of the same type. Instead of explicitly activating and deactivating parameter sets, the syntax element values of a parameter set may be used in the (de)coding process when the parameter set is referenced e.g. through its identifier, similarly as explained above regarding parameter set activation.
[0089] An adaptation parameter set (APS) may be defined as a syntax structure that applies to zero or more slices. There may be different types of adaptation parameter sets. An adaptation parameter set may for example contain filtering parameters for a particular type of a filter. In WC, three types of APSs are specified carrying parameters for one of: adaptive loop filter (ALF), luma mapping with chroma scaling (LMCS), and scaling lists. A scaling list may be defined as a list that associates each frequency index with a scale factor for the scaling process, which multiplies transform coefficient levels by a scaling factor, resulting in transform coefficients. In WC, an APS is referenced through its type (e.g. ALF, LMCS, or scaling list) and an identifier. In other words, different types of APSs have their own identifier value ranges.
[0090] Instead of or in addition to parameter sets at different hierarchy levels (e.g., sequence and picture), video coding formats may include header syntax structures, such as a sequence header or a picture header. A sequence header may precede any other data of the coded video sequence in the bitstream order. A picture header may precede any coded video data for the picture in the bitstream order.
[0091] In WC, a picture header (PH) may be defined as a syntax structure containing syntax elements that apply to all slices of a coded picture. In other words, contains information that is common for all slices of the coded picture associated with the PH. A picture header syntax structure is specified as an RBSP and is contained in a NAL unit.
[0092] A first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams. In some coding formats or standards, the end of the first bitstream may be indicated by a specific NAL unit, which may be referred to as the end of bitstream (EOB) NAL unit and which is the last NAL unit of the bitstream.
[0093] The phrase along the bitstream (e.g. indicating along the bitstream) or along a coded unit of a bitstream (e.g. indicating along a coded tile) may be used in claims and described embodiments to refer to transmission, signaling, or storage in a manner that the "out-of-band" data is associated with but not included within the bitstream or the coded unit, respectively. The phrase decoding along the bitstream or along a coded unit of a bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream or the coded unit, respectively. For example, the phrase along the bitstream may be used when the bitstream is contained in a container file, such as a file conforming to the ISO Base Media File Format, and certain file metadata is stored in the file in a manner that associates the metadata to the bitstream, such as boxes in the sample entry for a track containing the bitstream, a sample group for the track containing the bitstream, or a timed metadata track associated with the track containing the bitstream.
[0094] A coded video sequence (CVS) may be defined as such a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream. A coded video sequence may additionally or alternatively be specified to end, when a specific NAL unit, which may be referred to as an end of sequence (EOS) NAL unit, appears in the bitstream.
[0095] Media coding standards may specify “profiles” and “levels.” A profile may be defined as a subset of algorithmic features of the standard (of the encoding algorithm or the equivalent decoding algorithm). In another definition, a profile is a specified subset of the syntax of the standard (and hence implies that the encoder may only use features that result into a bitstream conforming to that specified subset and the decoder may only support features that are enabled by that specified subset).
[0096] A level may be defined as a set of limits to the coding parameters that impose a set of constraints in decoder resource consumption. In another definition, a level is a defined set of constraints on the values that may be taken by the syntax elements and variables of the standard. These constraints may be simple limits on values. Alternatively, or in addition, they may take the form of constraints on arithmetic combinations of values (e.g., picture width multiplied by picture height multiplied by number of pictures decoded per second). Other means for specifying constraints for levels may also be used. Some of the constraints specified in a level may for example relate to the maximum picture size, maximum bitrate and maximum data rate in terms of coding units, such as macroblocks, per a time period, such as a second. The same set of levels may be defined for all profiles. It may be preferable for example to increase interoperability of terminals implementing different profiles that most or all aspects of the definition of each level may be common across different profiles.
[0097] Video encoders may utilize Lagrangian cost functions to find coding modes, e.g. the desired block partitioning and / or motion vectors. This kind of cost function uses a weighting factor X to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel value in an image area:C = D + XR (Eq. 1) where C is the Lagrangian cost to be minimized, D is the image distortion (e.g Mean Squared Error), e.g. with the block partitioning and / or motion vectors considered, and R is the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the block partitioning and / or the motion vectors). A Lagrangian cost may be a way to quantify rate distortion (RD) performance. Using a Lagrangian cost in coding mode selection may be referred to as rate-distortion optimization or ratedistortion-optimized encoding.
[0098] The purpose of in-loop filtering is to reduce artifacts and distortions that can occur during the compression process. Compression techniques such as block-based motion compensation and discrete cosine transform (DCT) can introduce artifacts such as blocking, ringing, and blurring in the decoded video. In-loop filtering is designed to reduce these artifacts and improve the perceived visual quality of the video. In-loop filters play a critical role in themaintenance of compressed video quality, since they can not only improve the quality of the current frame but can also provide a higher quality reference for subsequent frames.
[0099] In HEVC, one or more (e.g., two) in-loop filters, such as a deblocking filter (DBF) followed by a sample adaptive offset (SAG) filter, may be applied to one or more reconstructed samples.
[0100] The in-loop filters in WC are depicted in Figure 2. Four processing steps, namely a luma mapping with chroma scaling (LMCS) process, followed by a deblocking filter (DBF), a sample adaptive offset (SAG) filter, and an adaptive loop filter (ALF) are applied to the reconstructed samples before writing them into the decoded picture buffer. The DBF and SAG are similar to that of the HEVC standard, whereas LMCS and ALF are newly introduced in WC.
[0101] A block-based ALF is used in WC, which comprises luma ALF, chroma ALF and cross-component ALF (CC-ALF). The ALF filter coefficients are either pre-defined and fixed in both encoder and decoder or adaptively signaled on a picture basis using adaptation parameter set (APS).
[0102] ALF filter parameters are signalled in Adaptation Parameter Set (APS). In one APS, up to 25 sets of luma filter coefficients and clipping value indices, and up to eight sets of chroma filter coefficients and clipping value indices could be signalled. To reduce the overhead, filter coefficients of different classification for luma component can be merged. In slice header, the indices of the APSs used for the current slice are signaled.
[0103] Clipping value indices, which are decoded from the APS, allow determining clipping values using a table of clipping values for both luma and chroma components. These clipping values are dependent of the internal bit depth. More precisely, the clipping values may be obtained by the following formula:AlfClip={round(2p’a*n) for n£[0..N-l]} where 0 is equal to the internal bit depth, a is a pre-defined constant value, and N equal to the number of allowed clipping values. In WC, for example the values of a = 2.35 and N=4 may be used. The AlfClip is then rounded to the nearest value with the format of power of 2.
[0104] In slice header, up to 7 ALF APS indices can be signaled to specify the luma filter sets that are used for the current slice. The filtering process can be further controlled at CTB level. A flag is signalled to indicate whether ALF is applied to a luma CTB. A luma CTB may choose afilter set among 16 fixed filter sets and the filter sets from APSs. A filter set index is signaled for a luma CTB to indicate which filter set is applied. The 16 fixed filter sets are pre-defined in the WC standard and hard-coded in both the encoder and the decoder. The 16 fixed filter sets may be referred to as the pre-defined ALFs.
[0105] A slice header may contain sh alf cb enabled flag and sh alf cr enabled flag to enable (when equal to 1) or disable (when equal to 0) ALF for Cb and Cr components, respectively. When ALF is enabled for either or both chroma components, an ALF APS index is signaled in slice header to indicate the chroma filter sets being used for the current slice. At CTB level, a filter index is signaled for each chroma CTB if there is more than one chroma filter set in the ALF APS.
[0106] The filter coefficients are quantized with norm equal to 128. In order to restrict the multiplication complexity, a bitstream conformance is applied so that the coefficient value of the non-central position shall be in the range of -27 to 27 - 1, inclusive. The central position coefficient is not signalled in the bitstream and is considered as being equal to 128.
[0107] For each ALF APS, in the current draft of WC there areMaximum of 25 luma ALF filtersMaximum of 8 chroma ALF filtersMaximum of 4 cross-component ALF filters for Cb componentMaximum of 4 cross-component ALF filters for Cr component
[0108] To limit the required memory for storing ALF APSs at decoder side, the WC standard limits the value range for ALS APS ID values so that the storage of up to 8 ALF APSs is needed.
[0109] Table 1 shows the syntax of APS in WC. Since the APS ID syntax element (referred to as aps_adaptation_parameter_set_id) is a 5-bit unsigned integer (denoted as u(5) in the syntax), it may have a value between 0 to 31, inclusive. However, WC has semantic restrictions that for ALF APS and scaling list APS, the value of APS ID shall be in the range of 0 to 7, inclusive, and for LMCS APS, the value of APS ID shall be in the range of 0 to 3, inclusive.Table 1.
[0110] The descriptors specifying the parsing process of each syntax element are defined in the VCC specifications as follows: ae(v): context-adaptive arithmetic entropy-coded syntax element. b(8): byte having any pattern of bit string (8 bits). The parsing process for this descriptor is specified by the return value of the function read_bits( 8 ). f(n): fixed-pattern bit string using n bits written (from left to right) with the left bit first. The parsing process for this descriptor is specified by the return value of the function read_bits( n ). i(n): signed integer using n bits. When n is "v" in the syntax table, the number of bits varies in a manner dependent on the value of other syntax elements. The parsing process for this descriptor is specified by the return value of the function read_bits( n ) interpreted as a two's complement integer representation with most significant bit written first. se(v): signed integer O-th order Exp-Golomb-coded syntax element with the left bit first. u(n): unsigned integer using n bits. When n is "v" in the syntax table, the number of bits varies in a manner dependent on the value of other syntax elements. The parsing process for this descriptor is specified by the return value of the function read_bits( n ) interpreted as a binary representation of an unsigned integer with most significant bit written first. ue(v): unsigned integer O-th order Exp-Golomb-coded syntax element with the left bit first.
[0111] Table 2 shows the syntax of indicating active / used LMCS APS in the picture header in WCTable 2.
[0112] Scalable video coding refers to coding structure where one bitstream can contain multiple representations of the content e.g. at different bitrates, resolutions or frame rates. In these cases, the receiver can extract the desired representation depending on its characteristics (e.g. resolution that matches best the display device). Alternatively, a server or a network element can extract the portions of the bitstream to be transmitted to the receiver depending on e.g. the network characteristics or processing capabilities of the receiver.
[0113] Scalable video coding may be realized through multi-layered coding. Multi-layered coding is a concept wherein an un-encoded visual representation of a scene is, by processes such as transformation and filtering, mapped into multiple dependent or independent representations (called layers). One or more encoders are used to encode a layered visual representation. When the layers contain redundancies, the use of a single encoder can, by using inter-layer prediction techniques, encode with a significant gain in coding efficiency. Layered video coding is typically used to provide some form of scalability in services - e.g. quality scalability, spatial scalability, temporal scalability, and view scalability.
[0114] A portion of a scalable video bitstream that provides a certain decoded representation, such as a base quality video or a depth map video for a bitstream that also contains texture video, and is independently decodable from other portions of the scalable video bitstream, may be referred to as an independent layer. A scalable video bitstream may comprise multiple independent layers, e.g. a texture video layer, a depth video layer, and an alpha map video layer. A portion of a scalable video bitstream that provides a certain decoded representation or enhancement, such as a quality enhancement to a particular fidelity or a resolution enhancementto a certain picture width and height in samples, and requires decoding of one or more other layers (a.k.a. reference layers) in the scalable video bitstream due to inter-layer prediction may be referred to as a dependent layer or a predicted layer.
[0115] In some scenarios, a scalable bitstream includes a "base layer", which may provide a basic representation, such as the lowest quality video available, and one or more enhancement layers. In order to improve coding efficiency for an enhancement layer, the coded representation of that layer may depend on one or more of the lower layers, i.e. inter-layer prediction may be applied. E.g. the motion and mode information of the enhancement layer can be predicted from lower layers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer. The term enhancement layer may refer to enhancing one or more aspects of reference layer(s), such as quality or resolution. A portion of the bitstream that remains after removal of all enhancement layers may be referred to as the base layer.
[0116] It needs to be understood that the term layer may be conceptual, i.e. the bitstream syntax might not include signaling of layers or the signaling of layers is not in use in a scalable bitstream that conceptually comprises several layers. The term scalability layer may be used interchangeably with the term layer.
[0117] Scalability modes or scalability dimensions may include but are not limited to the following:Quality scalability: Base layer pictures are coded at a lower quality than enhancement layer pictures, which may be achieved for example using a greater quantization parameter value (i.e., a greater quantization step size for transform coefficient quantization) in the base layer than in the enhancement layer.Spatial scalability: Base layer pictures are coded at a lower resolution (i.e. have fewer samples) than enhancement layer pictures. Spatial scalability and quality scalability, particularly its coarse-grain scalability type, may sometimes be considered the same type of scalability.Bit-depth scalability: Base layer pictures are coded at lower bit-depth (e.g. 8 bits) than enhancement layer pictures (e.g. 10 or 12 bits).Dynamic range scalability: Scalable layers represent a different dynamic range and / or images obtained using a different tone mapping function and / or a different optical transfer function.Chroma format scalability: Base layer pictures provide lower spatial resolution in chroma sample arrays (e.g. coded in 4:2:0 chroma format) than enhancement layer pictures (e.g. 4:4:4 format).Color gamut scalability: enhancement layer pictures have a richer / broader color representation range than that of the base layer pictures - for example the enhancement layer may have UHDTV (ITU-R BT.2020) color gamut and the base layer may have the ITU-R BT.709 color gamut.Region-of-interest (ROI) scalability: An enhancement layer represents of spatial subset of the base layer. ROI scalability may be used together with other types of scalability, e.g. quality or spatial scalability so that the enhancement layer provides higher subjective quality for the spatial subset.View scalability, which may also be referred to as multiview coding. The base layer represents a first view, whereas an enhancement layer represents a second view. A view may be defined as a sequence of pictures representing one camera or viewpoint. It may be considered that in stereoscopic or two-view video, one video sequence or view is presented for the left eye while a parallel view is presented for the right eye.Depth scalability, which may also be referred to as depth-enhanced coding. A layer or some layers of a bitstream may represent texture view(s), while other layer or layers may represent depth view(s).
[0118] In modern video coding standards, such as High Efficiency Video Coding (HEVC, ISO / IEC 23008-2) and Versatile Video Coding (WC, ISO / IEC 23090-3), layers can be grouped into layer-sets in which the combination of layers in a layer-set can produce a valid decodable bitstream. When layers are coded using inter-layer prediction techniques, the coded layers have a decoding hierarchy.
[0119] WC supports temporal scalability by including a temporal ID in the NAL unit header. A hierarchical coding structure is defined since pictures of a particular temporal sub-layer cannot be used as reference for inter-prediction by pictures of a lower temporal sub-layer.
[0120] In addition to temporal scalability, WC also supports multi-layered scalability for various types of scalability including spatial scalability, SNR scalability, and view scalability.For spatial scalability, an enhancement layers can be predicted using single-layer intra-prediction and inter-prediction. In addition to inter-layer prediction, where resampled reconstructed videocontent from sone or more reference layer can be used for prediction. Furthermore, SNR scalability is achieved similarly as spatial scalability, but without changes in the spatial resolution. With WC it may be possible to extract a sub-bitstream extraction process, and the requirement that each sub-bitstream extraction output be a conforming bitstream.
[0121] Similar to WC, HEVC also supports temporal scalability through temporal layering. In HEVC, scalability is provided by extensions, e.g. scalable HEVC (SHVC), Multiview HEVC (MV-HEVC) and 3D-HEVC. In MV-HEVC and 3D-HEVC, a layer can represent texture, depth, etc. There can be dependencies between layers that are exploited through inter-layer prediction. As in WC, it is also possible in MV-HEVC and 3D-HEVC to extract sub-bitstreams. SHVC, MV-HEVC, and 3D-HEVC use a common basis specification, specified in Annex F of the version 2 of the HEVC standard. This common basis comprises for example high-level syntax and semantics e.g. specifying some of the characteristics of the layers of the bitstream, such as inter-layer dependencies, as well as decoding processes, such as reference picture list construction including inter-layer reference pictures and picture order count derivation for multilayer bitstream. Annex F may also be used in potential subsequent multi-layer extensions of HEVC. It is to be understood that even though a video encoder, a video decoder, encoding methods, decoding methods, bitstream structures, and / or embodiments may be described in the following with reference to specific extensions, such as SHVC and / or MV-HEVC, they are generally applicable to any multi-layer extensions of HEVC, and even more generally to any multi-layer video coding scheme.
[0122] In HEVC and WC, the NAL unit header syntax comprises a layer identifier (ID) syntax element, also referred to as nuh layer id. Coded video NAL units of the same layer have the same nuh layer id value. A dependent layer may be required to have a greater nuh layer id value than the nuh layer id values of its reference layers.
[0123] An auxiliary picture may be defined as a picture that has no normative effect on the decoding process of primary pictures. A primary picture may be defined as a picture regarded as the primary output of the decoding process. A sequence of the primary pictures may be, e.g., the texture video that is displayed, whereas the associated sequences of auxiliary pictures may supplement the primary video but may not be used directly for displaying. An auxiliary picture or a primary picture may either be a coded picture or a decoded picture, depending on the context that the term is used.
[0124] In some video coding specifications, such as HEVC or WC, a sequence of auxiliary pictures may be included in an independent layer, which may be referred to as an auxiliary picture layer. Characteristics, such as the type (which may be indicative of, but may not be limited to, e.g., depth map or alpha plane) of an auxiliary picture layer may be included in or along a bitstream, such as in an SEI message or in a VPS.
[0125] Video Coding for Machines (VCM)
[0126] Reducing the distortion in image and video compression is often intended to increase human perceptual quality, as humans are considered to be the end users, i.e. consuming / watching the decoded image. Recently, with the advent of machine learning, especially deep learning, there is a rising number of machines (i.e., autonomous agents) that analyze data independently from humans and that may even take decisions based on the analysis results without human intervention. Examples of such analysis are object detection, scene classification, semantic segmentation, video event detection, anomaly detection, pedestrian tracking, etc. Example use cases and applications are self-driving cars, video surveillance cameras and public safety, smart sensor networks, smart TV and smart advertisement, person re-identification, smart traffic monitoring, drones, etc. When decoded data is consumed by machines, different quality metric, i.e. other than human perceptual quality, may be advantageous when considering media compression in inter-machine communications. Also, dedicated algorithms for compressing and decompressing data for machine consumption are likely to be different than those for compressing and decompressing data for human consumption. The set of tools and concepts for compressing and decompressing data for machine consumption is referred to here as Video Coding for Machines.
[0127] It is likely that the receiver-side device has multiple “machines” or neural networks (NNs). These multiple machines may be used in a certain combination which is for example determined by an orchestrator sub-system. The multiple machines may be used for example in succession, based on the output of the previously used machine, and / or in parallel. For example, a video which was compressed and then decompressed may be analyzed by one machine (NN) for detecting pedestrians, by another machine (another NN) for detecting cars, and by another machine (another NN) for estimating the depth of all the pixels in the frames.
[0128] It is noted that the terms “receiver-side” or “decoder-side” are used to refer to the physical or abstract entity or device which contains one or more machines, and runs these one ormore machines on some encoded and eventually decoded video representation which is encoded by another physical or abstract entity or device, the “encoder-side device”.
[0129] The encoded video data may be stored into a memory device, for example as a file. The stored file may later be provided to another device.
[0130] Alternatively, the encoded video data may be streamed from one device to another.
[0131] Figure 3 provides a general illustration of the pipeline of Video Coding for Machines. A VCM encoder encodes the input video into a bitstream. A bitrate may be computed from the bitstream in order to evaluate the size of the bitstream. A VCM decoder decodes the bitstream output by the VCM encoder. The output of the VCM decoder is referred in the figure as “Decoded data for machines”. This data may be considered as the decoded or reconstructed video. However, in some implementations of this pipeline, this data may not have same or similar characteristics as the original video which was input to the VCM encoder. For example, this data may not be easily understandable by a human by simply rendering the data onto a screen. The output of VCM decoder is then input to one or more task neural network. In the figure, for the sake of illustrating that there may be any number of task-NNs, there are three example task-NNs, and a non-specified one (Task-NN X). The goal of VCM is to obtain a low bitrate while guaranteeing that the task-NNs still perform well in terms of the evaluation metric associated to each task.
[0132] When a conventional video encoder, such as a H.266 / VVC encoder, is used as a VCM encoder, one or more of the following approaches may be used to adapt the encoding to be suitable to machine analysis tasks:One or more regions of interest (ROIs) may be detected. An ROI detection method may be used. For example, ROI detection may be performed using a task NN, such as an object detection NN. In some cases, ROI boundaries of a group of pictures or an intra period may be spatially overlaid and rectangular areas may be formed to cover the ROI boundaries. The detected ROIs (or rectangular areas, likewise) may be used in one or more of the following ways: o The quantization parameter (QP) may be adjusted spatially in a manner that ROIs are encoded using finer quantization step size(s) than other regions. For example, QP may be adjusted CTU-wise.o The video is preprocessed to contain only the ROIs, while the other areas are replaced by one or more constant sample values or removed. o A grid is formed in a manner that a single grid cell covers a ROI. Grid rows or grid columns that contain no ROIs are downsampled as preprocessing to encoding.Quantization parameter of the highest temporal sublay er(s) is increased (i.e. coarser quantization is used) when compared to practices for human watchable video.The original video is temporally downsampled as preprocessing prior to encoding. A frame rate upsampling method may be used as postprocessing subsequent to decoding, if machine analysis at the original frame rate is desired.A filter is used to preprocess the input to the conventional encoder. The filter may be a machine learning based filter, such as a convolutional neural network.
[0133] Object mask information in video coding
[0134] Sending object masks as video data allows results of video analysis performed by the encoder, such as information about object bounding boxes and object labels, to be sent to the decoder. This could reduce the workload and power consumption of the decoder.
[0135] Figures 4a - 4c show examples of such mask data. Figure 4a shows a texture video frame, from where an object mask with 8 identified objects (persons, phones) according to Figure 4b is derived. The details of Figure 4b magnified in Figure 4c.
[0136] Object masks can be coded as grayscale images, with each code value identifying one object. Lossless coding of such maps is costly in terms of bit rate. Lossy coding of objects maps affects the reconstruction accuracy negatively and consequently affects any subsequent tasks performed at the decoder relying on such mask data, i.e. person identification, tracking, etc.
[0137] Object mask video may be regarded as one type of mask information video. Mask information video may be defined as a video sequence where the source sample values are known by an encoder to occupy a subset of all the available sample values for the bit-depth of the video sequence. In a typical case, the number of sample values used in the mask information video is significantly smaller than what could be expressed with the bit-depth of the video sequence. The luma sample values that may be present in the mask information video may be referred to valid luma sample values. For example, if the mask information video comprises background represented by luma sample value equal to 0 and one object represented by the lumasample value equal to 128, the valid luma sample values comprise values 0 and 128. Mask information video may be monochrome but may likewise have any other chroma format.
[0138] Forms of mask information video may include but are not limited to:- Object mask- Region of Interest (ROI) masks- ROI masks for VCM- Alpha mask (a.k.a. alpha plane)- Occupancy mask- Filtering mask- Depth map with a limited number of possible depth values
[0139] A region-of-interest mask may indicate one or more regions of interest in an associated image. A region of interest may be perceptually more important to human viewers or may be more important to improve a computer vision task accuracy when compared to areas outside of the regions of interest.
[0140] An alpha mask may be used to provide transparency information for an associated image. A first value of an alpha mask may represent a fully opaque pixel, and a second value may represent a fully transparent pixel. Values between the first and second values may represent different levels of transparency between fully opaque and fully transparent.
[0141] An occupancy mask may be used in patch-based volumetric video coding to indicate which sample locations of an associated texture image (a.k.a. texture atlas) are occupied by pixel values to be used in volumetric video reconstruction and which sample locations are unoccupied, i.e., should not be used in volumetric video reconstruction. A first value of an occupancy mask may indicate an occupied sample location, and a second value may indicate an unoccupied sample location.
[0142] A filtering mask may be used to indicate the areas that are subject to filtering and / or are not subject to filtering, and / or a filtering strength. The type of filtering may be, but may not be limited to, one or more of the following: film grain noise synthesis, adaptive loop filtering (ALF), post-filtering. A first value of a filtering mask may indicate that no filtering is applied for the collocated pixel in an associated image, and a second value may indicate that filtering is applied for the collocated pixel in an associated image. Values between the first and secondvalues may represent different levels of filtering strength between no filtering and filtering at full strength.
[0143] All of these forms of mask have the same problem when it comes to video coding. Lossless compression is costly in terms of bit rate, lossy compression is affecting the accuracy of the reconstructed masks.
[0144] It is to be understood that even though some embodiments are described with reference to object mask video, embodiments may be applied to any type of mask information video.
[0145] It is to be understood that even though some embodiments are described with reference to auxiliary pictures, embodiments may be applied in a case where object mask video is represented as primary pictures in a first bitstream, and the primary video associated with the object mask video may be represented as primary pictures in a second bitstream.
[0146] Object mask video coding
[0147] Usually, an object mask is a binary matrix wherein “0” represents background and “1” represents the foreground. However, in an auxiliary picture with luma sample bit-depth equal to BitDepthY, there are l«BitDepthY different values for each sample, where « indicates a bitshift operation to the left. To distinguish different masks within one picture, JVET-AD0175 uses the sample value of the auxiliary picture as the ID of the mask. I.e, in the auxiliary picture, the samples with the same value form a mask and the regions with different sample values represent different masks. For example, as shown in Figure 5a, there is a 16x8 sample auxiliary picture, and numbers in the samples denote the sample values. There are four different masks in the auxiliary picture of Figure 5a: the samples with value 5 form a mask with ID equal to 5; the samples with value 10 form a mask with ID equal to 10; the samples with value 20 form a mask with ID equal to 20; and the samples with value 0 (no shading) form a mask with ID equal 0. The special mask with ID equal 0 may be labelled as “background”.
[0148] As the masks can be overlapping, multiple object mask auxiliary pictures (one auxiliary picture in one layer) can used for one primary picture to handle overlapping case. In that case, the samples with same position but in the different mask picture could belong to different masks overlapped with each other. For example, as shown in Figure 5b, there are two 4x4 object mask auxiliary pictures. In the picture 0, the top-left 2x2 block is a mask with ID 5,and in the picture 1 , the center 2x2 block is a mask with ID 10. When a decoder receives these two auxiliary pictures, it is evident that there are two masks being overlapped at position (1,1).
[0149] JVET-AF0087 discloses an experiment where the object mask auxiliary pictures are coded using VTM-21.0 layered coding, as illustrated in the flow chart of Figure 6. In the experiment, only objects that are labelled as “person” were considered for simplicity. Two types of mask sequences were generated: 1) single-mask sequences and 2) multiple-mask sequences. For single-mask sequences, all the detected persons in the picture are treated as one mask (with the same mask ID), and mask ID (i.e., the sample value in the mask area) is set to 255. For multiple-mask sequences, at most four persons in a picture are kept, and the mask IDs (i.e., the sample value in the mask area) are set to 64, 128, 192 and 255. The background samples are set to 0.
[0150] Examples of single-mask and multiple-mask sequences are given in Figures 7a and 7b, respectively. In the single-mask sequence of Figure 7a, all persons are provided with same mask ID (similar shading). In the multiple-mask sequence of Figure 7b, different persons are provided with different mask IDs (different shadings). The bit-depth of the mask sequences is 8-bit and the color format is YUV4:0:0.
[0151] In an example, VCC multi-layer coding is used. The normal sequence was coded as primary pictures in layer 0 and the accompanying mask sequence was coded as auxiliary pictures in layer 1.
[0152] A “mask recovery” process is added after the mask sequences are decoded. It outputs the recovered mask sequences.
[0153] For single-mask sequences, the samples with decoded value less than or equal to 128 are decided as background area and the sample value is set to 0; the samples with decoded value larger than 128 are decided as mask area and the sample value is set to 255.
[0154] For multiple-mask sequences, the samples with decoded value less than or equal to 32 are decided as background area and the sample value is set to 0; the samples with decoded value larger than 32 and less than or equal to 96 are decided as maskO area and the sample value is set to 64; the samples with decoded value larger than 96 and less than or equal to 160 are decided as maskl area and the sample value is set to 128; the samples with decoded value larger than 160 and less than or equal to 223 are decided as mask2 area and the sample value is set to 192; thesamples with decoded value larger than 223 are decided as mask3 area and the sample value is set to 255.
[0155] LMCS filter
[0156] Luma mapping with chroma scaling (LMCS) was originally proposed to improve the subjective coding performance of HEVC for high dynamic range (HDR) and / or wide color gamut (WCG) and later adopted as in-loop filter into WC. LMCS contains two components: luma mapping (LM) and luma-dependent chroma residue scaling (CS).
[0157] The basic idea behind luma mapping is to make better use of the range of luma code values allowed at a specified bit depth. It is commonplace that not all allowed luma code values are used in a video signal. For example, as specified by ITU-R BT.2100-2, only luma code values from 64 to 940 are allowed for a 10-bit narrow range video signal. The luma code values from 0 to 63 and from 941 to 1023 are not allowed for the video signal but may be used within the coding process.
[0158] The luma mapping component of LMCS provides a process of reallocating the signaldomain luma code values within all or part of the full range of luma code values allowed in the coding domain. The chroma residue scaling component of LMCS is designed to compensate for the interaction between the luma signal and its corresponding chroma signals. In WC, the quantization parameter (QP) applied to the chroma residue signal depends on the value of the corresponding luma signal. When the luma mapping component of LMCS is enabled, the luma value in the coding domain may be different from the luma value in the original signal domain. As a result, the value of chroma QP may not be optimal. The chroma residue scaling component of LMCS aims at compensating this defect by making luma-dependent adjustments to the values of the chroma residue signals within a chroma coding block.
[0159] The decoding architecture of LMCS is illustrated in Figure 8. The upper part of Figure 8 illustrates the chroma residue scaling component of LMCS. The lower part of the figure illustrates the luma mapping component of LMCS. LMCS introduces the concepts of signals being in either the “mapped domain” or the “original (non-mapped signal) domain”. The luma code values of video signal in mapped domain which is processed by LMCS may be different from those of original video signals which is in original domain. The functional blocks in Figure 8 where the processing is performed on signals in the mapped domain when LMCS is enabled include inverse quantization (Q1) & inverse transform (T1) (800), luma intra prediction (IntraPrediction, 804), and summing the luma prediction with the luma residue values (Reconstruction, 802). LMCS also introduces new functional blocks, such as Inverse and Forward luma mapping (806, 814) on the luma mapping side and Chroma scaling (816) on the chroma mapping side.
[0160] The LMCS functional blocks are the following. Forward Mapping (814) maps luma code values in the original domain to luma code values in the mapped domain. The original domain is used for reference pictures and for decoded output pictures, whereas the mapped domain is used in the (de)coding process. Inverse Mapping (806) maps luma code values in the mapped domain used in the (de)coding process to the code values in the original domain. Chroma Scaling (816) determines a chroma scaling factor and applies the factor to chroma residue values. The blocks where the processing is applied in the original (non-mapped) domain include loop filters (808, 822) such as deblocking, adaptive loop filter (ALF), and sample adaptive offset (SAG), motion compensated (inter) prediction (812), chroma intra prediction (826), summing chroma prediction (818) with the chroma residue values, and storage of pictures in a decoded picture buffer (DPB; 810, 824).
[0161] Luma mapping makes use of a forward mapping function, FwdMap, and a corresponding inverse mapping function, InvMap. In WC, the FwdMap function is signaled using a piecewise linear model. InvMap function does not need to be signaled and is instead derived in the decoder from the FwdMap function.
[0162] In WC, the FwdMap piecewise linear model is determined as follows. The range of code values supported by a particular signal bit depth is partitioned into 16 equal pieces. For example, each of the 16 pieces for a 10-bit input signal would have 64 codewords assigned to it. The number of code words assigned to each piece is denoted by the variable OrgCW. The variable InputPivot[i], with i = 0..16, indicating the pivot point of each piece in original domain, is derived as InputPivot[i] = i * OrgCW. During the encoding process, values for mapped pivot points, denoted here by the variable MappedPivot[i], are determined. The difference MappedPivot[i+l] - MappedPivot[i] is the number of mapped luma code word values of the i-th piece of the piecewise linear model. The values of InputPivot[i] and MappedPivot[i] completely specify the FwdMap function.
[0163] In WC, the parameter values required for determining the values of InputPivot[i] and MappedPivot[i] at the decoder are signaled in the adaptation parameter set (APS) syntax structure with aps_params_type set equal to 1 (LMCS APS). The value range for an adaptationparameter set identifier (aps_adaptation_parameter_set_id) is from 0 to 3, inclusive, for LMCS APSs. While up to 4 LMCS APSs are available for any picture to be encoded, only up to 1 LMCS APS may be used for a picture. At the picture header, an LMCS enable flag is signaled to indicate if the LMCS process as depicted in Figure 8 is applied to the current picture. If LMCS is enabled for the current picture, an aps id is signaled in the picture header in the ph lmcs aps id syntax element to identify the APS that carries the luma mapping parameters. Thus, the same LMCS parameters are used for entire picture. However, when LMCS is enabled in the picture header and the picture header is a separate NAL unit and not included in the slice header, LMCS may still be enabled or disabled on slice basis using the sh lmcs used flag syntax element. sh lmcs used flag equal to 1 specifies that luma mapping is used for the current slice and chroma scaling could be used for the current slice (depending on the value of ph chroma residual scale flag). sh lmcs used flag equal to 0 specifies that luma mapping with chroma scaling is not used for the current slice.
[0164] Table 3 shows the syntax of the Imcs data syntax structure, which is included in the LMCS APS of WC.Table 3.
[0165] The semantics of the luma mapping related syntax elements of Imcs data as specified in WC are included in the next paragraphs.
[0166] Imcs min bin idx specifies the minimum bin index used in the luma mapping with chroma scaling construction process. The value of Imcs min bin idx shall be in the range of 0 to 15, inclusive.
[0167] Imcs delta max bin idx specifies the delta value between 15 and the maximum bin index LmcsMaxBinldx used in the luma mapping with chroma scaling construction process. The value of Imcs delta max bin idx shall be in the range of 0 to 15, inclusive. The value of LmcsMaxBinldx is set equal to 15 - Imcs delta max bin idx. The value of LmcsMaxBinldx shall be greater than or equal to Imcs min bin idx.
[0168] lmcs_delta_cw_prec_minusl plus 1 specifies the number of bits used for the representation of the syntax lmcs_delta_abs_cw[ i ]. The value of lmcs_delta_cw_prec_minusl shall be in the range of 0 to 14, inclusive.
[0169] lmcs_delta_abs_cw[ i ] specifies the absolute delta codeword value for the ith bin.
[0170] lmcs_delta_sign_cw_flag[ i ] specifies the sign of the variable lmcsDeltaCW[ i ] as follows:- If lmcs_delta_sign_cw_flag[ i ] is equal to 0, lmcsDeltaCW[ i ] is a positive value.- Otherwise ( lmcs_delta_sign_cw_flag[ i ] is not equal to 0 ), lmcsDeltaCW[ i ] is a negative value.
[0171] When lmcs_delta_sign_cw_flag[ i ] is not present, it is inferred to be equal to 0.
[0172] The variable OrgCW is derived as follows:OrgCW = ( 1 « BitDepth ) / 16
[0173] The variable lmcsDeltaCW[ i ], with i = Imcs min bin idx.. LmcsMaxBinldx, is derived as follows: lmcsDeltaCW[ i ] = ( 1 - 2 * lmcs_delta_sign_cw_flag[ i ] ) * lmcs_delta_abs_cw[ i ]
[0174] The variable lmcsCW[ i ] is derived as follows:- For i = 0.. Imcs min bin idx - 1, lmcsCW[ i ] is set equal 0.- For i = Imcs min bin idx.. LmcsMaxBinldx, the following applies: lmcsCW[ i ] = OrgCW + lmcsDeltaCW[ i ]The value of lmcsCW[ i ] shall be in the range of OrgCW » 3 to ( OrgCW « 3 ) - 1, inclusive.- For i = LmcsMaxBinldx + 1..15, lmcsCW[ i ] is set equal 0.
[0175] The variable InputPivot[ i ], with i = 0..15, is derived as follows:InputPivot[ i ] = i * OrgCW
[0176] The variable LmcsPivot[ i ] with i = 0..16, the variables ScaleCoeff[ i ] and InvScaleCoeff[ i ] with i = 0..15, are derived as follows: LmcsPivot
[0000] = 0 for( i = 0; i <= 15; i++ ) {LmcsPivot[ i + 1 ] = LmcsPivot[ i ] + lmcsCW[ i ]ScaleCoeff[ i ] = ( lmcsCW[ i ] * (1 « 11 ) + ( 1 « ( Log2( OrgCW ) - 1 ) ) ) » ( Log2( OrgCW ) ) if( lmcsCW[ i ] = = 0 )InvScaleCoeff[ i ] = 0 elseInvScaleCoeff[ i ] = OrgCW * ( 1 « 11 ) / lmcsCW[ i ]}
[0177] It should also be noted that when the luma mapping with chroma scaling is enabled in a picture header and a chroma format including the chroma components is in use, the chroma scaling part can be enabled or disabled in the picture header through ph chroma residual scale flag. When a picture has multiple slices, the luma mapping with chroma scaling is further enabled or disabled in the slice header for each slice.
[0178] The picture inverse mapping process 806 for luma samples may be specified as follows. Input to this process is a reconstructed picture luma sample array SL. Output of this process is a modified reconstructed picture luma sample array SL. For each coordinate value pair of x and y within the reconstructed picture luma sample array SL, the inverse mapping process for a luma sample SL[ X ][ y ] is invoked with the variable lumaSample set equal to SL[ x ][ y ] as the input and the output is assigned to the luma sample SL[ X ][ y ].
[0179] The inverse mapping process for a luma sample may be specified as follows. Input to this process is a luma sample lumaSample. Output of this process is a modified luma sample invLumaSample. The value of invLumaSample is derived as follows:- If sh lmcs used flag of the slice that contains the luma sample lumaSample is equal to 1, the following ordered steps apply:1. The variable idxYInv is derived by invoking the identification of piece-wise function index process for a luma sample with lumaSample as the input and idxYInv as the output.2. The variable invSample is derived as follows:
[0180] invSample = InputPivot[ idxYInv ] + ( ( InvScaleCoeff[ idxYInv ] *( lumaSample - LmcsPivot[ idxYInv ] ) + ( 1 « 10 ) ) » 11 )3. The inverse mapped luma sample invLumaSample is derived as follows:
[0181] invLumaSample = Clip 1 ( invSample ) where the function Clipl is specified as follows:
[0182] Clipl( x ) = Clip3( 0, ( 1 « BitDepth ) - 1, x ) where BitDepth is the sample bit-depth and Clip3 is specified as follows: ; z < x ; z > y; otherwise- Otherwise, invLumaSample is set equal to lumaSample.
[0184] The identification of piecewise function index process for a luma sample may be specified as follows. Input to this process is a luma sample lumaSample. Output of this process is an index idxYInv identifying the piece to which the luma sample lumaSample belongs. The variable idxYInv is derived as follows:
[0185] for( idxYInv = Imcs min bin idx; idxYInv <= LmcsMaxBinldx; idxYInv++ ) if( lumaSample < LmcsPivot[ idxYInv + 1 ] ) break idxYInv = Min( idxYInv, 15 )
[0186] It may be concluded from the above that mask images and / or mask video are an important tool for various applications, such asObject masks / Annotated Region masks / Object identification / ROI masks for VCM, Region of Interest (ROI) masks in general,Alpha masks to represent transparency, andOccupancy mask, e.g. in V3C (Visual Volumetric Video-based Coding) videoFiltering mask to control post-processing (such as film grain synthesis), in-loop filtering, and / or post-filteringDepth map with a constrained number of depth levels
[0187] Mask images typically contain sharp object edges, which may in turn cause typical compression artefacts, such as ringing noise or softening of the sharp edges. However, many applications using mask images would greatly benefit from accurate and sharp object edges. For example, errors in the occupancy mask may cause "flying pixels" in the reconstructed volumetric image or video.
[0188] Coding such masks in a lossless fashion requires a significant amount of data. Lossy compression would reduce the data requirement at the cost of mask reconstruction accuracy. Various post-filtering proposals for such mask video exist, e.g. as described in JVET-AF0087 above or Enhanced Occupancy map reconstruction in V3C.
[0189] However, such steps are outside the video coding loop and thus brings one major drawback: As the reconstructed picture is filtered outside the coding loop, after the decoding process, it is stored “unfiltered” in the reference picture buffer. Consequently, any prediction from this unfiltered picture suffers from the loss introduced due to compression. The result is a sub-optimal prediction and thus higher bit rate and lower reconstruction accuracy. A further drawback in this approach is the lack of standardized mask post-filters. Consequently, every vendor may use its own reconstruction, thus limiting the possibility of a broad application of lossy mask compression.
[0190] Now improved methods for enabling higher reference picture accuracy are introduced.
[0191] The method according to a first aspect, as shown in Figure 9, comprises receiving (900) a first picture; deriving (902) an inverse luma mapping function; encoding (904) the first picture to a first coded picture; applying (906) the inverse luma mapping function for a reconstruction of the first picture; storing (908) the reconstructed first picture in a reference picture buffer; receiving (910) a second picture; and encoding (912) the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference without applying any forward mapping function.
[0192] According to an embodiment, the inverse luma mapping function is a piece- wise linear function.
[0193] According to an embodiment, the inverse luma mapping function is represented by a lookup table, where the table index may be regarded as the input argument to the function, and the value of the table entry for the table index may be regarded as the value returned by the function.
[0194] According to an embodiment, the inverse luma mapping function is represented by InvScaleCoeff, InputPivot, and LmcsPivot arrays of WC.
[0195] According to an embodiment, the first coded picture is an intra-coded picture and the second coded picture is an inter-coded picture.
[0196] According to an embodiment, the method further comprises encoding a luma mapping with chroma scaling (LMCS) adaptation parameter set based on said inverse luma mapping function; and encoding an indication in or along the first coded picture that the LMCS adaptation parameter set is applied.
[0197] According to an embodiment, the method further comprises deriving the forward luma mapping function based on the inverse luma mapping function; encoding a luma mapping with chroma scaling (LMCS) adaptation parameter set based on said forward luma mapping function; and encoding an indication in or along the first coded picture that the LMCS adaptation parameter set is applied.
[0198] According to an embodiment, the method further comprises receiving information of valid luma sample values in the first picture; and deriving LMCS pivot points for the LMCS adaptation parameter set from said valid luma sample values, where the pivot points are set according to the total number of different valid luma samples and available bit depth.
[0199] According to an embodiment, the method further comprises deriving LMCS pivot points in a manner that a valid luma sample value is represented by two LMCS pivot points, which specify a piece in the piece-wise linear inverse luma mapping function in such a manner that a luma sample given as input and mapped to this piece returns approximately or exactly the valid luma sample value.
[0200] According to an embodiment, the method further comprises encoding an indication in or along the first coded picture that chroma residual scaling is disabled. For example, the method may comprise encoding ph chroma residual scale flag equal to 0 in the picture header of the first coded picture.
[0201] According to an embodiment, the method further comprises encoding an indication in or along the second coded picture that LMCS is disabled. For example, the method may comprise encoding ph lmcs enabled flag equal to 0 in the picture header of the second coded picture.
[0202] Accordingly, a decoder receives a first (intra) picture with an indication about luma mapping with chroma scaling being enabled, such as ph lmcs enabled flag = 1. The decoder also receives an LMCS adaptation parameter set (APS) reflecting a piecewise value distribution, e.g. pivot point values [ 0 0 50 50 50 100 100 100 150 150 150 200 200200250250] determined, for example, based on the number of masks used in the video sequence of the first picture. The decoder derives the inverse luma mapping function, i.e. an inverse look-up table (LUT). The decoder applies the derived inverse LUT to the first picture and stores the resulting picture in the reference picture buffer. The decoder receives a second picture (inter) with an indication about luma mapping with chroma scaling being disabled, such as ph lmcs enabled flag = 0. The decoder may then use the first picture for prediction without applying any LUT.
[0203] Thus, instead of carrying out the mask recovery as an out-of-loop application, as shown in Fig. 6 illustrating the solution according to JVET-AF0087, the mask recovery is shifted to be carried out as an in-loop filtering within the video coding loop. This results in higher reference picture accuracy and thus consequently better inter-prediction, reducing bit rate requirements and improving reconstruction quality.
[0204] In the following, the operation of an encoder is discussed further in an exemplified manner.
[0205] In this example, the encoder receives a mask image with a plurality of masks, such as up to 7 masks. The masks are coded as grayscale values in the range of 0 to (2A(bitdepth)-l), such as 0 to 255 for 8-bit video. The masks are carried only in the luma plane, and the chroma planes either may be set to a constant value or are not present (i.e. YUV 4:0:0).
[0206] The code word CW of a mask M(n) is set based on the available bit depth B and the total number of individual masks N per video, following the constraint of being a multiple of 1 « (B-5), as follows:Minstep = (1 « (B-5))Steps = floor(32 / N) / / B / ( 1 « (B-5) ) is always 32, divided by MasksStepsize = Steps * Minstep CW(n) = n* Stepsize
[0207] The background (BG) CW is set to 0. , the maximum CW value is capped to 31*(1(B-5)), as the maximum value 2AB-1 is not a multiple of 1 « (B-5)
[0208] The Table 4 below shows an example of the mask allocation for a 8-bit video:Table 4.
[0209] It is noted that the CW values in the table are just examples, and other values may be used, as well. The mask values are preferably distributed over the available bit depth to allow for some guard interval between mask values, as well as background. The background (BG) CW may also set to another value than 0.
[0210] Figures 10a - 10c show an example of a result of the mask generation process, where from the original texture picture shown in Figure 10a, an original mask identifying 7 different objects (6 elephants and background) is derived according to Figure 10b. Figure 10c shows the grayscale mask used for video coding.
[0211] In another example, the encoder receives information on how many masks are present in the video (with total number of masks N < 8) and creates the 16 LMCS forward mapping pivot points according to the process, where B is the available internal bit depth and i is the index of the LMCS forward mapping pivot point. The process may be illustrated by the following pseudocode: pS = floor(16 / (N+l)) / / pivot point steps size for 16 steps for N masks plus BG fwdPivot(0:pS) = 0 / / set background pivot points n = 0 / / set mask index for i=pS:pS: 16 / / loop through all pivot points with step size pSfwdPivot(i:i+pS-l) = M(n) / / set all pivot points within step size to mask value n = n+1 / / iterate through masks fwdPivot(16)=248 / / fill last pivot point in case it is still empty
[0212] As the pivot points require constant reconstruction for the mask values M(n), it is necessary to have at least two pivot points for each mask, wherein the maximum number of supported masks with this approach is 7.
[0213] In WC, it is a requirement of bitstream conformance that, for i =Imcs min bin idx. XmcsMaxBinldx, when the value of LmcsPivot[ i ] is not a multiple of 1 « ( BitDepth - 5 ), the value of ( LmcsPivot[ i ] » ( BitDepth - 5 ) ) shall not be equal to the value of ( LmcsPivot[ i + 1 ] » ( BitDepth - 5 ) ). Therefore, mask values M(n) shall be multiples of 1 « ( BitDepth - 5 ), e.g. 8 for B = 8.
[0214] The forward mapping points for this approach for 8-bit video may be, for example, as follows:1 Mask: [0 0 0 0 0 0 0 0 248 248 248 248 248 248 248 248]2 Masks: [0 0 0 0 0 128 128 128 128 128 248 248 248 248 248 248]5 Masks: [0 0 4848 48 96 96 96 144 144 144 192 192 192248 248]
[0215] It is noted that the above pivot point values are just examples, and other values may be used, as well. Likewise, it is not necessary to use the exact process for generating the pivot point values, but another process may be used, as well. Irrespective of the underlying process or the resulting values, it is more important to distribute the mask values over the available pivot points to allow for reconstruction of the mask intervals by applying the inverse LMCS LUT.
[0216] According to an embodiment, the encoder receives information on how many masks are present in the video (with total number of masks N < 7) and creates the N*2+2 LMCS forward mapping pivot points according to the process, where B is the available internal bit depth, I is the number of LMCS pivot points, and i is the index of the LMCS forward mapping pivot point. The process may be illustrated by the following pseudocode:I = N*2+2 / / 2 pivot points per mask + 2 points for background fwdPivot(0:2) = 0 / / set background pivot points n = 0 / / set mask index for i=2:2:I / / loop through all pivot points with step size 2 fwdPivot(i:i+l) = M(n) / / set all pivot points within step size to mask valuen = n+1 / / iterate through masks
[0217] The forward mapping points for this approach for 8-bit video may be, for example, as follows:1 Mask: [0 0248 248]2 Masks: [0 0 128 128 248 248]5 Masks: [0 048 48 96 96 144 144 192 192248 248 ]
[0218] According to an embodiment, the encoder creates an inverse mapping LUT according to these pivot points and omits creation of a forward mapping LUT. The encoder encodes intra pictures using the LMCS and disables LMCS for any inter-predicted pictures.
[0219] According to an embodiment, the encoder creates an inverse mapping LUT according to these pivot points and a transparent (identity) forward mapping LUT. The encoder encodes intra pictures using the LMCS and disables LMCS for any inter-predicted pictures.
[0220] As a conclusion, the encoding process described above enables to utilize the encoding process of Figure 9 via the signalling and decoding process as defined in WC. The encoder signals the modified LMCS pivot points to the decoder for an intra picture, but then disables LMCS for any pictures using the intra picture as a reference. Thus, the inverse LUT is calculated and applied at the decoder similarly as for the encoder (only for intra pictures). For interpredicted pictures with LMCS, a forward LUT would have to be applied to the reference picture before using it in the decoding process. According to WC, it is not possible to signal a transparent (identity) LUT for LMCS. However, by disabling LMCS for any inter-predicted pictures, the inverse LM process can be used for intra pictures of mask video without any normative changes to WC.
[0221] It has been estimated that the presence of a filtered reference picture in the reference picture buffer may reduce the error rate for predicted images by 60-70% at the same bit rate. Thus, clear benefits are obtained for object mask video coding.
[0222] The encoding method according to a second aspect, as shown in Figure I la, comprises receiving (1100) a first picture; deriving (1102) an inverse luma mapping function; encoding (1104) luma mapping parameters based on said inverse luma mapping function; encoding (1106) the first picture to a first coded picture; encoding (1108) an indication in or along the first coded picture that the luma mapping parameters are applied; applying (1110) the inverse luma mapping function for a reconstruction of the first picture; storing (1112) the reconstructed first picture in areference picture buffer; receiving (1114) a second picture; encoding (1116) the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for inter prediction by applying an identity forward mapping function; and encoding (1118) an indication in or along the second coded picture that the identity forward mapping function being used in the inter prediction.
[0223] According to an embodiment, the method comprises encoding a luma mapping with chroma scaling (LMCS) adaptation parameter set based on said luma mapping parameters.
[0224] According to an embodiment, the method comprises encoding a luma mapping with chroma scaling (LMCS) adaptation parameter set based on said inverse luma mapping function.
[0225] According to an embodiment, the method comprises receiving or obtaining valid luma sample values; and deriving the luma mapping parameters from the valid luma sample values.
[0226] For example, valid luma sample values may be known by the application that controls the encoding method. In another example, the encoding method comprises analyzing the pictures received as input to encoding to obtain the valid luma sample values.
[0227] According to an embodiment, the method comprises deriving the luma mapping parameters from the inverse luma mapping function.
[0228] According to an embodiment, the method comprises deriving a forward luma mapping function based on the inverse luma mapping function; and deriving the luma mapping parameters from the forward luma mapping function.
[0229] According to an embodiment, the method comprises encoding a luma mapping with chroma scaling (LMCS) adaptation parameter set comprising luma mapping parameters.
[0230] The decoding method according to the second aspect, as shown in Figure 11b, comprises receiving (1130) luma mapping parameters; receiving (1132) a first coded picture with a first indication about luma mapping being enabled for said first coded picture and a second indication that the luma mapping parameters are used for said first coded picture; receiving (1134) a third indication about applying only an inverse luma mapping for said first coded picture; deriving (1136) an inverse luma mapping function based on the luma mapping parameters; decoding (1138) the first coded picture to a reconstruction of a first picture; applying (1140) at least the inverse luma mapping function for the reconstruction of the first picture to obtain a reconstructed first picture; storing (1142) the reconstructed first picture in a reference picture buffer; receiving (1144) a second coded picture with a fourth indication about lumamapping being enabled for said second coded picture and a fifth indication that the luma mapping parameters are used for said second coded picture; decoding (1146) a sixth indication about an identity forward luma mapping function being used in inter prediction; and decoding (1148) the second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for the inter prediction by applying the identity forward luma mapping function.
[0231] According to an embodiment, the luma mapping parameters comprise valid luma sample values.
[0232] According to an embodiment, the luma mapping parameters comprise a representation of the inverse luma mapping function.
[0233] According to an embodiment, the luma mapping parameters comprise a representation of a forward luma mapping function from which the inverse luma mapping function may be derived.
[0234] According to an embodiment, the method comprises receiving a luma mapping with chroma scaling (LMCS) adaptation parameter set comprising the luma mapping parameters.
[0235] According to an embodiment, the method comprises applying the identity forward mapping function to any picture subsequent to the first reconstructed picture in the reference picture buffer used as a reference for the inter prediction and applying the inverse luma mapping function for a reconstruction of any picture subsequent to the reconstruction of the first picture before storing said reconstruction of the picture in the reference picture buffer.
[0236] It is to be understood that in addition to inverse luma mapping, other loop filters may be applied to the reconstruction of a picture (e.g., to the reconstruction of the first picture) and / or to the picture resulting from the inverse luma mapping before storing the reconstructed picture resulting from all the loop filter operations into the reference picture buffer. For example, the deblocking filter, the sample adaptive offset filter, and the adaptive loop filter may be enabled for a picture and may be applied successively after the inverse luma mapping.
[0237] According to an embodiment, the first coded picture is an intra-coded picture and the second coded picture is an inter-coded picture.
[0238] Accordingly, a decoder receives a first (intra) picture with an indication about luma mapping with chroma scaling being enabled, such as ph lmcs enabled flag = 1. The decoder also receives an LMCS adaptation parameter set (APS) reflecting a piecewise value distribution,e.g. pivot point values [ 0 0 50 50 50 100 100 100 150 150 150 200 200200250250] determined based on the number of masks used in the video sequence of the first picture. Additionally, the decoder receives an indication about applying only an inverse luma mapping for said first picture, such as a flag indicating the presence of “inverse luma mapping only”. The decoder derives the inverse luma mapping function, i.e. an inverse LUT. The decoder applies the derived inverse LUT to the first picture and stores the resulting picture in the reference picture buffer. The decoder receives a second picture (inter) with an indication about luma mapping with chroma scaling being enabled, such as ph lmcs enabled flag = 1. Then based on the flag indicating the presence of “inverse luma mapping only”, the decoder (a) applies a transparent (identity) Forward LUT to any picture in the reference picture buffer, and (b) skips the Forward LUT application to any pictures to be stored in the reference picture buffer, but rather applies the inverse LUT to pictures to be stored in the reference picture buffer.
[0239] Again, instead of carrying out the mask recovery as an out-of-loop application, as shown in Fig. 6 illustrating the solution according to JVET-AF0087, the mask recovery is shifted to be carried out as an in-loop filtering within the video coding loop. The same benefits as described above may be achieved: higher reference picture accuracy and thus consequently better inter-prediction, reducing bit rate requirements and improving reconstruction quality.
[0240] According to an embodiment, the inverse luma mapping function is represented in syntax by an indication of mask information video, an indication of the number of masks, and the valid luma sample values. For example, Imcs dataQ or any similar syntax structure may contain a flag indicating whether syntax elements used in VVC for luma mapping syntax elements or alike follow, or whether an indication of the number of masks and syntax element(s) indicative of the valid luma sample values follows. A decoding process to convert the valid luma samples to an inverse luma mapping function may be realized according to any embodiment or example in this disclosure.
[0241] In the following, the operation of an encoder upon encoding pictures suitable for the encoding process of Figure I la is discussed further in an exemplified manner. Herein, instead of relying on the WC, changes are proposed in the signalling and decoding process.
[0242] Also in this example, the encoder receives a mask image with a plurality of masks, such as up to 7 masks. The encoder encodes the masks and creates forward mapping pivot pointssimilarly to what is described above. The encoder creates an inverse mapping LUT according to the pivot points and a transparent (identity) forward mapping LUT.
[0243] The encoder encodes all pictures (intra and inter-predicted) using LMCS and sends a flag to indicate to the decoder to use transparent (identity) forward LUT. The flag may be provided, for example, in the slice header syntax as shown below.
[0244] Herein, sh_fwd_lmcs_used_flag equal to 1 specifies that forward luma mapping is used for the current slice, sh fwd lmcs used flag equal to 0 specifies that forward luma mapping is not used for the current slice. When sh fwd lmcs used flag is not present, it is inferred to be equal to 0.
[0245] Alternatively, the flag can be signalled in the picture header syntax, for example as shown below.
[0246] Herein, ph_fwd_lmcs_enabled_flag equal to 1 specifies that forward LM is enabled for the current picture, ph fwd lmcs enabled flag equal to 0 specifies that forward LM is disabled for the current picture. When not present, the value of ph fwd lmcs enabled flag is inferred to be equal to 0.
[0247] Yet alternatively, the flag can be signalled in the sequence parameter set syntax, for example as shown below.
[0248] Herein, sps_fwd_lmcs_enabled_flag equal to 1 specifies that forward LM is enabled for the CLVS. sps fwd lmcs enabled flag equal to 0 specifies that inverse LM is disabled for the CLVS.
[0249] It is to be understood that embodiments could be similarly realized with other syntax structures carrying the flag(s), such as a picture parameter set or sequence header. It is also possible to have a combination of the above flags, and / or derive flag values from flags higher up in the signalling hierarchy. For example, if sps fwd lmcs enabled flag = 0, then ph fwd lmcs enabled flag, if not present, can be derived as 0.
[0250] According to an embodiment, the decoder receives a bitstream including a flag indicating the use of a transparent (identity) forward LUT for LM in-loop filtering, e.g. ph fwd lmcs enabled flag = 0. The decoder reconstructs the inverse LUT based on the received pivot points in the LMCS APS (as in the approach of Figure 9), but instead of reconstructing also the forward mapping LUT from the received pivot points, it uses a transparent (identity) LUT as forward mapping LUT.
[0251] It is to be understood that embodiments could be similarly realized with other syntax elements indicative of enabling / disabling inverse luma mapping and enabling / disabling forward luma mapping. For example, a forward luma mapping enabled flag (fwd lm enabled flag) could be indicated. If fwd lm enabled flag is equal to 0 (disabled), an inverse luma mapping enabled flag (inv lm enabled flag) could be indicated. Otherwise, if fwd lm enabled flag is equal to 1 (enabled), inv lm enabled flag could be absent and inferred to be equal to 1. A variable LmEnabledFlag could be set equal to fwd lm enabled flag | inv lm enabled flag (where | is a bitwise XOR operation).
[0252] According to an embodiment, a decoder decodes an indication of mask information video. In response to the indication indicating mask information video, the decoder infers the value of the flag indicating forward luma mapping to be equal to 0 (disabled).
[0253] As a conclusion, some embodiments of the encoding process described above introduce a new flag to disable forward luma mapping, and the decoding processes described above decode the new flag from the bitstream or infer its value, which enables the decoder to use inverse LM filtering of all pictures, also inter-predicted pictures, but at the same time, removing the effect of forward LMCS LUT application on any reference pictures. Forward and inverse LUT is set and applied at the decoder same as for the encoder. Thus, better reconstruction quality can be achieved for sub-sequent (reference) pictures, greatly improving coding performance with only a minor change in the signalling.
[0254] The encoding method according to a third aspect, as shown in Figure 12a, comprises receiving (1200) a first picture; selecting (1202) a predefined inverse luma mapping function; deriving (1204) a forward luma mapping function based on the inverse luma mapping function; encoding (1206) the first picture to a first coded picture; applying (1208) the inverse luma mapping function for a reconstruction of the first picture; storing (1210) the reconstructed first picture in a reference picture buffer; receiving (1212) a second picture; and encoding (1214) the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference by applying an identity forward mapping function.
[0255] In an embodiment, the encoding method further comprises indicating in or along the first coded picture the selected predefined inverse luma mapping function. For example, the predefined inverse luma mapping functions may be given predefined indices, and the encodingmethod may comprise including the index of the selected predefined inverse luma mapping function in or along the first coded picture.
[0256] The decoding method according to the third aspect, as shown in Figure 12b, comprises receiving (1220) a first coded picture with an indication about luma mapping being enabled for said first coded picture and an indication about applying only an inverse luma mapping for said first coded picture; decoding (1222) the first coded picture, applying (1224) a predefined inverse luma mapping function for a reconstruction of the first coded picture; storing (1226) the reconstructed first picture in a reference picture buffer; receiving (1228) a second coded picture; and decoding (1230) the second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for the inter prediction by applying an identity forward luma mapping function.
[0257] In an embodiment, the decoding method further comprises decoding from or along the first coded picture an identification of the predefined inverse luma mapping function. For example, the predefined inverse luma mapping functions may be given predefined indices, and the decoding method may comprise decoding an index of the predefined inverse luma mapping function from or along the first coded picture, wherein the index identifies the predefined inverse luma mapping function to be used for decoding the first coded picture.
[0258] According to an embodiment, the method comprises applying the identity forward mapping function to any input picture in the reference picture buffer used as a reference for the inter prediction; and applying the predefined inverse luma mapping function for a reconstruction of any picture before storing said reconstruction of the picture in the reference picture buffer.
[0259] According to an embodiment, the first coded picture is an intra-coded picture and the second coded picture is an inter-coded picture.
[0260] Accordingly, a decoder receives a first (intra) picture with an indication about luma mapping with chroma scaling being enabled, such as ph lmcs enabled flag = 1. The decoder also receives an indication about applying only an inverse luma mapping for said first picture, such as a flag indicating the presence of “inverse luma mapping only”. The decoder applies a pre-stored (hard-coded) inverse LUT to the first picture and stores the resulting picture in the reference picture buffer. The decoder receives a second picture (inter) with an indication about luma mapping with chroma scaling being enabled, such as ph lmcs enabled flag = 1. Then based on the flag indicating the presence of “inverse luma mapping only”, the decoder (a) appliesa transparent (identity) Forward LUT to any picture in the reference picture buffer, and (b) skips the Forward LUT application to any pictures in the reference picture buffer.
[0261] In the following, the operation of an encoder upon encoding pictures is discussed further in an exemplified manner. It is noted that the previous approaches reuse significant parts of the LMCS signalling. Thus, they are limited to a maximum of 7 masks per image (maximum number of flat areas within 16 LMCS bins). Herein, this limitation is broken by introducing a modified inverse LUT available at the decoder, while not changing the application of LMCS at the decoder itself.
[0262] As a result, several signalling embodiments for such an inverse LUT are possible: a) In one embodiment, only one hard-coded inverse LUT is available at the decoder, for example, covering 15 masks maximum. Indicating the use of inverse LM automatically indicates the use of this inverse LUT. b) In one embodiment, the encoder signals the use of inverse LM plus the number of masks present. The decoder picks the appropriate inverse LUT from a pool of available LUT based on the number of masks in the video. c) In one embodiment, the encoder signals the use of inverse LM plus the number of masks present. The decoder recreates the inverse LUT based on the number of masks and available bit depth following the same process as at the encoder. d) In one embodiment, the inverse LUT is signalled directly to the decoder.
[0263] The embodiments a) and b), while being straightforward, are briefly described below.
[0264] a) In one embodiment, an encoder receives an object mask video with up to 15 masks. The mask values are in steps of (2AB) / 16, e.g. 16 for 8-bit video (BG: 0, Maskl: 16, ... , Maskl5: 255). The encoder uses a pre-set inverse mapping LUT and a transparent (identity) forward mapping LUT. The encoder encodes all pictures (intra and inter-predicted) using LMCS and sends a normative flag to indicate to the decoder to use inverse LM, e.g. as shown in the approach of Figure 11b.
[0265] The decoder receives a bitstream including a flag indicating the use of a inverse LM, e.g. ph fwd lmcs enabled flag = 0. The decoder uses a predefined inverse LUT and a transparent (identity) LUT as forward mapping LUT.
[0266] b) In one embodiment, an encoder receives an object mask video with up to 15 masks. The values of the masks are set as described in the approach of Figure 9, but limited to amaximum of 15 masks. The encoder receives information on how many objects are represented by masks in the video. The encoder selects a pre-set inverse mapping LUT, based on the number of masks represented in the video, and a transparent (identity) forward mapping LUT. The encoder encodes all pictures (intra and inter-predicted) using LMCS and sends a normative flag to indicate to the decoder to use inverse LM, e.g. as shown in the approach of Figure 11, as well as signalling the number of masks to the decoder.
[0267] The decoder receives a bitstream including a flag indicating the use of an inverse LM, e.g. ph fwd lmcs enabled flag = 0, as well as information on how many masks are present. The decoder selects a predefined inverse mapping LUT, based on the number of masks and a transparent LUT as forward mapping LUT.
[0268] The embodiment c) is described more in detail below.
[0269] c) An encoder receives a mask image with up to N masks. The encoder encodes the masks similarly to what is described in the approach of Figure 9, without the limitation to 7 masks. The encoder creates an inverse LUT based on the number of masks N and the available bit depth B. The process may be illustrated by the following pseudocode: bd=2AB-l / / max CW value sZ=bd / (N+l) / / step size n = 0 / / set mask index invLUT(0:sZ-l) = 0 / / set LUT for BG values for i=sZ:sZ:bd / / loop through LUT for all masks invLUT(i) = M(n) n = n+1 / / iterate mask index
[0270] The encoder encodes all pictures (intra and inter-predicted) using LMCS and sends a normative flag to indicate to the use of inverse LM. No LMCS pivot points are signalled as the inverse LUT can be derived at the decoder with the same process.
[0271] The decoder receives a bitstream including a flag indicating the use of a inverse LM, e.g. ph fwd lmcs enabled flag = 0, as well as information on how many masks are present. The decoder creates the inverse mapping LUT following the same process as the encoder, and uses a transparent LUT as forward mapping LUT.
[0272] The signalling of the number of masks for the embodiments b) and c) may be carried out in a similar way.
[0273] According to an embodiment, a LMCS mode flag is included the Imcs dataQ syntax structure within an LMCS APS indicating if conventional LMCS parameters or the number of masks are signalled. For example, the following syntax structure may be used for a maximum of 15 masks:
[0274] Herein, lmcs_masks_minusl plus 1 specifies the number of masks present in the video. The value of Imcs masks minusl shall be in the range of 0 to 14, inclusive.
[0275] According to an embodiment, the following syntax structure may be used for the Imcs dataQ:
[0276] It is noted that the existing WC bitstreams conform to this syntax. Thus, this syntax may be introduced in encoders without modifications to the current LMCS APS writing.
[0277] Furthermore, the constraints of Imcs min bin idx and Imcs delta max bin idx are relaxed so that they may express the number of masks as indicated in the following:
[0278] lmcs_min_bin_idx, when less than or equal to LmcsMaxBinldx derived below, specifies the minimum bin index used in the luma mapping with chroma scaling construction process. The value of Imcs min bin idx shall be in the range of 0 to 15, inclusive.
[0279] lmcs_delta_max_bin_idx is used to derive LmcsMaxBinldx.
[0280] The value of Imcs delta max bin idx shall be in the range of 0 to 15, inclusive. The value of LmcsMaxBinldx is set equal to 15 - Imcs delta max bin idx.
[0281] When LmcsMaxBinldx is greater than or equal to Imcs min bin idx, it specifies the maximum bin index used in the luma mapping with chroma scaling construction process.
[0282] When LmcsMaxBinldx is less than Imcs min bin idx, the number of masks is derived to be equal to Imcs min bin idx - LmcsMaxBinldx.
[0283] The embodiment d) is similar to the embodiment c), except that the LUT is signalled directly to the decoder. This allows for flexible ways of creating the inverse mapping LUT at the encoder.
[0284] Signalling of the LUT may be carried out, for example, in the LMCS data as follows:
[0285] Herein, lmcs_inv_LUT_delta_minusl[n] plus 1 specifies the delta between the inverse LUT value at position n and position n-1. lmcs_inv_LUT_delta[O] is inferred to be equal to 0.
[0286] As a conclusion, the encoding process described above enables to utilize the decoding process of Figure 12b, wherein in addition to the benefits gained for the decoding process of Figure 11b, the support for more than 7 masks per picture is allowed.
[0287] It is noted that each of the decoding methods and the related embodiments according to the various aspects may be implemented independently of each other.
[0288] According to an embodiment, which may be combined with any of the above embodiments, the number of masks may differ from one picture to another and / or the mask values change from one picture to another.
[0289] In this embodiment, rather than using a transparent (identity) forward mapping LUT, a forward mapping LUT is specified to map the mask values of a reference picture to respective mask values of a current picture being encoded or decoded. For example, if a reference picture is a three-level mask with values 0, 127, and 255 only, and the current picture being encoded or decoded adds a fourth mask value, the mask values 0, 127, and 255 in the reference picture may respectively correspond to masks value 0, 85, and 170 in the current picture, and the new mask value in the current gets value 255. The forward mapping LUT is specified to map values 0, 127, and 255 to the values 0, 85, and 170, respectively.
[0290] This embodiment may be realized with any signaling means, such as the following:
[0291] According to an embodiment, the applied LMCS APS identifier is indicated separately for the forward mapping LUT and the inverse mapping LUT.
[0292] According to an embodiment, the number of masks for the inverse mapping LUT of the current picture is indicated or inferred, and the forward mapping LUT is derived from the number of masks in the current picture relative to the that of the reference picture.
[0293] Herein, if the number of masks in both is the same, a transparent (identity) forward mapping LUT is inferred. If the number of masks differ, the forward mapping LUT is derived, e.g., so that it linearly scales the values proportionally to the number of intervals between mask values in the reference picture divided by the number of intervals between mask values in the current picture.
[0294] The encoding method according to the fourth aspect comprises receiving a first picture; receiving valid luma sample values; encoding luma mapping parameters based on said valid luma sample values; deriving an inverse luma mapping function based on said valid luma sample values; encoding the first picture to a first coded picture; encoding an indication in or along the first coded picture that the luma mapping parameters are applied; applying the inverse luma mapping function for a reconstruction of the first picture; storing the reconstructed first picture in a reference picture buffer; receiving a second picture; encoding the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for inter prediction by applying an identity forward mapping function; and encoding an indication in or along the second coded picture that the identity forward mapping function being used in the inter prediction.
[0295] The embodiments relating to the decoding aspects may be implemented in an apparatus comprising means for receiving luma mapping parameters; means for receiving a first coded picture with a first indication about luma mapping being enabled for said first coded picture and a second indication that the luma mapping parameters are used for said first coded picture; means for receiving a third indication about applying only an inverse luma mapping for said first coded picture; means for deriving an inverse luma mapping function based on the luma mapping parameters; means for decoding the first coded picture to a reconstruction of a first picture; means for applying at least the inverse luma mapping function for the reconstruction of the first picture to obtain a reconstructed first picture; means for storing the reconstructed first picture in a reference picture buffer; means for receiving a second coded picture with a fourth indication about luma mapping being enabled for said second coded picture and a fifth indication that the luma mapping parameters are used for said second coded picture; means for decoding asixth indication about an identity forward luma mapping function being used in inter prediction; and means for decoding the second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for the inter prediction by applying the identity forward luma mapping function.
[0296] According to an embodiment, the apparatus comprises means for applying the identity forward mapping function to any picture subsequent to the first reconstructed picture in the reference picture buffer used as a reference for the inter prediction and applying the inverse luma mapping function for a reconstruction of any picture subsequent to the reconstruction of the first picture before storing said picture in the reference picture buffer.
[0297] According to an embodiment, the first coded picture is an intra-coded picture and the second coded picture is an inter-coded picture.
[0298] According to an embodiment, a luma mapping with chroma scaling (LMCS) adaptation parameter set comprises the luma mapping parameters.
[0299] According to an embodiment, luma mapping parameters comprise valid luma sample values and the apparatus comprises means for deriving the inverse luma mapping function from the valid luma sample values.
[0300] According to an embodiment, the adaptation parameter set reflects a piecewise value distribution determined based on a number of object masks used in a video sequence of the first picture.
[0301] According to an embodiment, said indication about applying only an inverse luma mapping is received in a flag included in a slice header syntax structure.
[0302] According to an embodiment, said indication about applying only an inverse luma mapping is received in a flag included in a picture header syntax structure.
[0303] According to an embodiment, said indication about applying only an inverse luma mapping is received in a flag included in a sequence parameter set syntax structure.
[0304] The embodiments relating to the decoding aspects may likewise be implemented in an apparatus comprising at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive luma mapping parameters; receive a first coded picture with a first indication about luma mapping being enabled for said first coded picture and a second indication that the lumamapping parameters are used for said first coded picture; receive a third indication about applying only an inverse luma mapping for said first coded picture; derive an inverse luma mapping function based on the luma mapping parameters; decode the first coded picture to a reconstruction of a first picture; apply the inverse luma mapping function for the reconstruction of the first picture to obtain a reconstructed first picture; store the reconstructed first picture in a reference picture buffer; receive a second coded picture with a fourth indication about luma mapping being enabled for said second coded picture and a fifth indication that the luma mapping parameters are used for said second coded picture; decode a sixth indication about an identify forward luma mapping function being used in inter prediction; and decode the second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for the inter prediction by applying the identity forward luma mapping function.
[0305] The embodiments relating to the encoding aspects may be implemented in an apparatus comprising means for receiving a first picture; means for deriving an inverse luma mapping function; means for encoding luma mapping parameters based on said inverse luma mapping function; means for encoding the first picture to a first coded picture; means for encoding an indication in or along the first coded picture that the luma mapping parameters are applied; means for applying the inverse luma mapping function for a reconstruction of the first picture; means for storing the reconstructed first picture in a reference picture buffer; means for receiving a second picture; means for encoding the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for inter prediction by applying an identity forward mapping function; and means for encoding an indication in or along the second coded picture that the identify forward mapping function being used in the inter prediction.
[0306] According to an embodiment, the apparatus comprises means for deriving the luma mapping parameters from the inverse luma mapping function.
[0307] According to an embodiment, the apparatus comprises means for deriving a forward luma mapping function based on the inverse luma mapping function; and means for encoding a luma mapping with chroma scaling (LMCS) adaptation parameter set based on said forward luma mapping function.
[0308] According to an embodiment, the first coded picture is an intra-coded picture and the second coded picture is an inter-coded picture.
[0309] According to an embodiment, the luma mapping parameters comprise a luma mapping with chroma scaling (LMCS) adaptation parameter set.
[0310] According to an embodiment, the apparatus comprises means for encoding LMCS pivot points in a manner that a valid luma sample value is represented by two LMCS pivot points, which specify a piece in the piece-wise linear inverse luma mapping function in such a manner that a luma sample given as input and mapped to this piece returns approximately or exactly the valid luma sample value.
[0311] According to an embodiment, the apparatus comprises means for deriving the inverse luma mapping function from the valid luma samples.
[0312] According to an embodiment, the apparatus comprises means for sending said indication about applying only an inverse luma mapping in a flag included in a slice header syntax structure.
[0313] According to an embodiment, the apparatus comprises means for sending said indication about applying only an inverse luma mapping in a flag included in a picture header syntax structure.
[0314] According to an embodiment, the apparatus comprises means for sending said indication about applying only an inverse luma mapping in a flag included in a sequence parameter set syntax structure.
[0315] The embodiments relating to the encoding aspects may likewise be implemented in an apparatus comprising at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive a first picture; derive an inverse luma mapping function; encode luma mapping parameters based on said inverse luma mapping function; encode the first picture to a first coded picture; encode an indication in or along the first coded picture that the luma mapping parameters are applied; apply the inverse luma mapping function for a reconstruction of the first picture; store the reconstructed first picture in a reference picture buffer; receive a second picture; encode the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for inter prediction by applying an identity forward mapping function; and encode an indication in or along the second coded picture that the identify forward mapping function being used in the inter prediction.
[0316] Such apparatuses may comprise e.g. the functional units disclosed in any of the Figures la, lb, 2, 6 and 8 for implementing the embodiments. The embodiments may also be implemented in apparatus according Figures 13 and 14, where Figure 13 shows a schematic block diagram of an exemplary apparatus or electronic device 50, which may incorporate a codec according to an embodiment of the invention. Figure 14 shows a layout of an apparatus according to an example embodiment. The elements of Figs. 13 and 14 will be explained next.
[0317] The electronic device 50 may for example be a mobile terminal or user equipment of a wireless communication system. However, it would be appreciated that embodiments of the invention may be implemented within any electronic device or apparatus which may require encoding and decoding or encoding or decoding video images.
[0318] The apparatus 50 may comprise a housing 30 for incorporating and protecting the device. The apparatus 50 further may comprise a display 32 in the form of a liquid crystal display. In other embodiments of the invention the display may be any suitable display technology suitable to display an image or video. The apparatus 50 may further comprise a keypad 34. In other embodiments of the invention any suitable data or user interface mechanism may be employed. For example the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0319] The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The apparatus 50 may further comprise an audio output device which in embodiments of the invention may be any one of: an earpiece 38, speaker, or an analogue audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other embodiments of the invention the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera capable of recording or capturing images and / or video. The apparatus 50 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 50 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.
[0320] The apparatus 50 may comprise a controller 56, processor or processor circuitry for controlling the apparatus 50. The controller 56 may be connected to memory 58 which in embodiments of the invention may store both data in the form of image and audio data and / ormay also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and decoding of audio and / or video data or assisting in coding and decoding carried out by the controller.
[0321] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
[0322] The apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and for receiving radio frequency signals from other apparatus(es).
[0323] The apparatus 50 may comprise a camera capable of recording or detecting individual frames which are then passed to the codec 54 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / or storage. The apparatus 50 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.
[0324] Figure 15 is a graphical representation of an example multimedia communication system within which various embodiments may be implemented. A data source 1510 provides a source signal in an analog, uncompressed digital, or compressed digital format, or any combination of these formats. An encoder 1520 may include or be connected with a preprocessing, such as data format conversion and / or filtering of the source signal. The encoder 1520 encodes the source signal into a coded media bitstream. It should be noted that a bitstream to be decoded may be received directly or indirectly from a remote device located within virtually any type of network. Additionally, the bitstream may be received from local hardware or software. The encoder 1520 may be capable of encoding more than one media type, such as audio and video, or more than one encoder 1520 may be required to code different media types of the source signal. The encoder 1520 may also get synthetically produced input, such as graphics and text, or it may be capable of producing coded bitstreams of synthetic media. In thefollowing, only processing of one coded media bitstream of one media type is considered to simplify the description. It should be noted, however, that typically real-time broadcast services comprise several streams (typically at least one audio, video and text sub-titling stream). It should also be noted that the system may include many encoders, but in the figure only one encoder 1520 is represented to simplify the description without a lack of generality. It should be further understood that, although text and examples contained herein may specifically describe an encoding process, one skilled in the art would understand that the same concepts and principles also apply to the corresponding decoding process and vice versa.
[0325] The coded media bitstream may be transferred to a storage 1530. The storage 1530 may comprise any type of mass memory to store the coded media bitstream. The format of the coded media bitstream in the storage 1530 may be an elementary self-contained bitstream format, or one or more coded media bitstreams may be encapsulated into a container file, or the coded media bitstream may be encapsulated into a Segment format suitable for DASH (or a similar streaming system) and stored as a sequence of Segments. If one or more media bitstreams are encapsulated in a container file, a file generator (not shown in the figure) may be used to store the one more media bitstreams in the file and create file format metadata, which may also be stored in the file. The encoder 1520 or the storage 1530 may comprise the file generator, or the file generator is operationally attached to either the encoder 1520 or the storage 1530. Some systems operate “live”, i.e. omit storage and transfer coded media bitstream from the encoder 1520 directly to the sender 1540. The coded media bitstream may then be transferred to the sender 1540, also referred to as the server, on a need basis. The format used in the transmission may be an elementary self-contained bitstream format, a packet stream format, a Segment format suitable for DASH (or a similar streaming system), or one or more coded media bitstreams may be encapsulated into a container file. The encoder 1520, the storage 1530, and the server 1540 may reside in the same physical device or they may be included in separate devices. The encoder 1520 and server 1540 may operate with live real-time content, in which case the coded media bitstream is typically not stored permanently, but rather buffered for small periods of time in the content encoder 1520 and / or in the server 1540 to smooth out variations in processing delay, transfer delay, and coded media bitrate.
[0326] The server 1540 sends the coded media bitstream using a communication protocol stack. The stack may include but is not limited to one or more of Real-Time Transport Protocol(RTP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Transmission Control Protocol (TCP), and Internet Protocol (IP). When the communication protocol stack is packet-oriented, the server 1540 encapsulates the coded media bitstream into packets. For example, when RTP is used, the server 1540 encapsulates the coded media bitstream into RTP packets according to an RTP payload format. Typically, each media type has a dedicated RTP payload format. It should be again noted that a system may contain more than one server 1540, but for the sake of simplicity, the following description only considers one server 1540.
[0327] If the media content is encapsulated in a container file for the storage 1530 or for inputting the data to the sender 1540, the sender 1540 may comprise or be operationally attached to a "sending file parser" (not shown in the figure). In particular, if the container file is not transmitted as such but at least one of the contained coded media bitstream is encapsulated for transport over a communication protocol, a sending file parser locates appropriate parts of the coded media bitstream to be conveyed over the communication protocol. The sending file parser may also help in creating the correct format for the communication protocol, such as packet headers and payloads. The multimedia container file may contain encapsulation instructions, such as hint tracks in the ISOBMFF, for encapsulation of the at least one of the contained media bitstream on the communication protocol.
[0328] The server 1540 may or may not be connected to a gateway 1550 through a communication network, which may e.g. be a combination of a CDN, the Internet and / or one or more access networks. The gateway may also or alternatively be referred to as a middle-box. For DASH, the gateway may be an edge server (of a CDN) or a web proxy. It is noted that the system may generally comprise any number gateways or alike, but for the sake of simplicity, the following description only considers one gateway 1550. The gateway 1550 may perform different types of functions, such as translation of a packet stream according to one communication protocol stack to another communication protocol stack, merging and forking of data streams, and manipulation of data stream according to the downlink and / or receiver capabilities, such as controlling the bit rate of the forwarded stream according to prevailing downlink network conditions. The gateway 1550 may be a server entity in various embodiments.
[0329] The system includes one or more receivers 1560, typically capable of receiving, demodulating, and de-capsulating the transmitted signal into a coded media bitstream. The coded media bitstream may be transferred to a recording storage 1570. The recording storage 1570 maycomprise any type of mass memory to store the coded media bitstream. The recording storage 1570 may alternatively or additively comprise computation memory, such as random access memory. The format of the coded media bitstream in the recording storage 1570 may be an elementary self-contained bitstream format, or one or more coded media bitstreams may be encapsulated into a container file. If there are multiple coded media bitstreams, such as an audio stream and a video stream, associated with each other, a container file is typically used and the receiver 1560 comprises or is attached to a container file generator producing a container file from input streams. Some systems operate “live,” i.e. omit the recording storage 1570 and transfer coded media bitstream from the receiver 1560 directly to the decoder 1580. In some systems, only the most recent part of the recorded stream, e.g., the most recent 10-minute excerption of the recorded stream, is maintained in the recording storage 1570, while any earlier recorded data is discarded from the recording storage 1570.
[0330] The coded media bitstream may be transferred from the recording storage 1570 to the decoder 1580. If there are many coded media bitstreams, such as an audio stream and a video stream, associated with each other and encapsulated into a container file or a single media bitstream is encapsulated in a container file e.g. for easier access, a file parser (not shown in the figure) is used to decapsulate each coded media bitstream from the container file. The recording storage 1570 or a decoder 1580 may comprise the file parser, or the file parser is attached to either recording storage 1570 or the decoder 1580. It should also be noted that the system may include many decoders, but here only one decoder 1580 is discussed to simplify the description without a lack of generality
[0331] The coded media bitstream may be processed further by a decoder 1580, whose output is one or more uncompressed media streams. Finally, a Tenderer 1590 may reproduce the uncompressed media streams with a loudspeaker or a display, for example. The receiver 1560, recording storage 1570, decoder 1580, and Tenderer 1590 may reside in the same physical device or they may be included in separate devices.
[0332] A sender 1540 and / or a gateway 1550 may be configured to perform switching between different representations e.g. for switching between different viewports of 360-degree video content, view switching, bitrate adaptation and / or fast start-up, and / or a sender 1540 and / or a gateway 1550 may be configured to select the transmitted representation(s). Switching between different representations may take place for multiple reasons, such as to respond to requests ofthe receiver 1560 or prevailing conditions, such as throughput, of the network over which the bitstream is conveyed. In other words, the receiver 1560 may initiate switching between representations. A request from the receiver can be, e.g., a request for a Segment or a Subsegment from a different representation than earlier, a request for a change of transmitted scalability layers and / or sub-layers, or a change of a rendering device having different capabilities compared to the previous one. A request for a Segment may be an HTTP GET request. A request for a Subsegment may be an HTTP GET request with a byte range. Additionally or alternatively, bitrate adjustment or bitrate adaptation may be used for example for providing so-called fast start-up in streaming services, where the bitrate of the transmitted stream is lower than the channel bitrate after starting or random-accessing the streaming in order to start playback immediately and to achieve a buffer occupancy level that tolerates occasional packet delays and / or retransmissions. Bitrate adaptation may include multiple representation or layer up-switching and representation or layer down-switching operations taking place in various orders.
[0333] A decoder 1580 may be configured to perform switching between different representations e.g. for switching between different viewports of 360-degree video content, view switching, bitrate adaptation and / or fast start-up, and / or a decoder 1580 may be configured to select the transmitted representation(s). Switching between different representations may take place for multiple reasons, such as to achieve faster decoding operation or to adapt the transmitted bitstream, e.g. in terms of bitrate, to prevailing conditions, such as throughput, of the network over which the bitstream is conveyed. Faster decoding operation might be needed for example if the device including the decoder 1580 is multi-tasking and uses computing resources for other purposes than decoding the video bitstream. In another example, faster decoding operation might be needed when content is played back at a faster pace than the normal playback speed, e.g. twice or three times faster than conventional real-time playback rate.
[0334] In the above, any embodiment is not limited to be applied only under the first, second or third aspect that is the last aspect preceding the description of the embodiment, but can be applied to any aspect. For example, embodiments relating to the derivation of the inverse luma mapping function from the valid luma sample values apply to each of the first, second and third aspect.
[0335] In the above, some embodiments have been described with reference to LMCS. It is to be understood that embodiments may be realized by applying only the luma mapping part of LMCS, even if not explicitly mentioned.
[0336] In the above, some embodiments have been described with reference to encoding luma mapping related parameters into an adaptation parameter set or decoding luma mapping related parameters from an adaptation parameter set. It is to be understood that embodiments may be realized by encoding luma mapping related parameters to any syntax structure in or along a bitstream, such as a picture parameter set, a picture header, or a slice header or by decoding luma mapping related parameters from any syntax structure in or along a bitstream, such as a picture parameter set, a picture header, or a slice header.
[0337] In the above, some embodiments have been described with reference to and / or using terminology of VVC. It needs to be understood that embodiments may be similarly realized with any video encoder and / or video decoder with respective terms of other codecs, such as AVI, AV2, HEVC, or H.264 / AVC.
[0338] In the above, where the example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder may have corresponding elements in them. Likewise, where the example embodiments have been described with reference to a decoder, it needs to be understood that the encoder may have structure and / or computer program for generating the bitstream to be decoded by the decoder.
[0339] The embodiments of the invention described above describe the codec in terms of separate encoder and decoder apparatus in order to assist the understanding of the processes involved. However, it would be appreciated that the apparatus, structures and operations may be implemented as a single encoder-decoder apparatus / structure / operation. Furthermore, it is possible that the coder and decoder may share some or all common elements.
[0340] Although the above examples describe embodiments of the invention operating within a codec within an electronic device, it would be appreciated that the invention as defined in the claims may be implemented as part of any video codec. Thus, for example, embodiments of the invention may be implemented in a video codec which may implement video coding over fixed or wired communication paths.
[0341] Thus, user equipment may comprise a video codec such as those described in embodiments of the invention above. It shall be appreciated that the term user equipment isintended to cover any suitable type of wireless user equipment, such as mobile telephones, portable data processing devices or portable web browsers.
[0342] Furthermore elements of a public land mobile network (PLMN) may also comprise video codecs as described above.
[0343] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0344] The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.
[0345] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on multi-core processor architecture, as non-limiting examples.
[0346] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automatedprocess. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0347] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
[0348] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention.
Claims
CLAIMS:
1. An apparatus comprising: means for receiving luma mapping parameters; means for receiving a first coded picture with a first indication about luma mapping being enabled for said first coded picture and a second indication that the luma mapping parameters are used for said first coded picture; means for receiving a third indication about applying only an inverse luma mapping for said first coded picture; means for deriving an inverse luma mapping function based on the luma mapping parameters; means for decoding the first coded picture to a reconstruction of a first picture; means for applying at least the inverse luma mapping function for the reconstruction of the first picture to obtain a reconstructed first picture; means for storing the reconstructed first picture in a reference picture buffer; means for receiving a second coded picture with a fourth indication about luma mapping being enabled for said second coded picture and a fifth indication that the luma mapping parameters are used for said second coded picture; means for decoding a sixth indication about an identity forward luma mapping function being used in inter prediction; and means for decoding the second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for the inter prediction by applying the identity forward luma mapping function.
2. The apparatus according to claim 1, comprising means for applying the identity forward mapping function to any picture subsequent to the first reconstructed picture in the reference picture buffer used as a reference for the inter prediction and applying the inverse luma mapping function for a reconstruction of any picture subsequent to the reconstruction of the first picture before storing said picture in the reference picture buffer.
3. The apparatus according to claim 1 or 2, wherein the first coded picture is an intra-coded picture and the second coded picture is an inter-coded picture.
4. The apparatus according to any preceding claim, wherein a luma mapping with chroma scaling (LMCS) adaptation parameter set comprises the luma mapping parameters.
5. The apparatus according to any preceding claim, wherein luma mapping parameters comprise valid luma sample values and the apparatus comprises means for deriving the inverse luma mapping function from the valid luma sample values.
6. The apparatus according to any preceding claim, wherein said indication about applying only an inverse luma mapping is received in a flag included in a slice header syntax structure, a picture header syntax structure or a sequence parameter set syntax structure.
7. A method comprising: receiving luma mapping parameters; receiving a first coded picture with a first indication about luma mapping being enabled for said first coded picture and a second indication that the luma mapping parameters are used for said first coded picture; receiving a third indication about applying only an inverse luma mapping for said first coded picture; deriving an inverse luma mapping function based on the luma mapping parameters; decoding the first coded picture to a reconstruction of a first picture; applying at least the inverse luma mapping function for the reconstruction of the first picture to obtain a reconstructed first picture; storing the reconstructed first picture in a reference picture buffer; receiving a second picture with a fourth indication about luma mapping being enabled for said second coded picture and a fifth indication that the luma mapping parameters are used for said second coded picture;decoding a sixth indication about an identity forward luma mapping function being used in inter prediction; and decoding the second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for the inter prediction by applying the identity forward luma mapping function.
8. An apparatus comprising: means for receiving a first picture; means for deriving an inverse luma mapping function; means for encoding luma mapping parameters based on said inverse luma mapping function; means for encoding the first picture to a first coded picture; means for encoding an indication in or along the first coded picture that the luma mapping parameters are applied; means for applying the inverse luma mapping function for a reconstruction of the first picture; means for storing the reconstructed first picture in a reference picture buffer; means for receiving a second picture; means for encoding the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for inter prediction by applying an identity forward mapping function; and means for encoding an indication in or along the second coded picture that the identify forward mapping function being used in the inter prediction.
9. The apparatus according to claim 8, comprising means for deriving the inverse luma mapping function from valid luma samples.
10. The apparatus according to claim 8 or 9, comprising means for deriving the luma mapping parameters from the inverse luma mapping function.
11. The apparatus according to any of claims 8 or 10, comprising means for deriving a forward luma mapping function based on the inverse luma mapping function; and means for encoding a luma mapping with chroma scaling (LMCS) adaptation parameter set based on said forward luma mapping function.
12. The apparatus according to any of claims 8 - 11, wherein the first coded picture is an intracoded picture and the second coded picture is an inter-coded picture.
13. The apparatus according to any of claims 8 - 12, wherein the luma mapping parameters comprise a luma mapping with chroma scaling (LMCS) adaptation parameter set.
14. The apparatus according to claim 12, comprising means for encoding LMCS pivot points in a manner that a valid luma sample value is represented by two LMCS pivot points, which specify a piece in the piece-wise linear inverse luma mapping function in such a manner that a luma sample given as input and mapped to this piece returns approximately or exactly the valid luma sample value.
15. The apparatus according to any of claims 8 - 14, comprising means for sending said indication about applying only an inverse luma mapping in a flag included in a slice header syntax structure, a picture header syntax structure or a sequence parameter set syntax structure.
16. A method comprising: receiving a first picture; deriving an inverse luma mapping function; encoding luma mapping parameters based on said inverse luma mapping function; encoding the first picture to a first coded picture; encoding an indication in or along the first coded picture that the luma mapping parameters are applied; applying the inverse luma mapping function for a reconstruction of the first picture;storing the reconstructed first picture in a reference picture buffer; receiving a second picture; encoding the second picture to a second coded picture using the reconstruction of the first picture in the reference picture buffer as a reference for inter prediction by applying an identity forward mapping function; and encoding an indication in or along the second coded picture that the identity forward mapping function being used in the inter prediction.
Citation Information
Patent Citations
Configuring luma-dependent chroma residue scaling for video coding
US20220030267A1