Method and apparatus for cross-component prediction for video coding
By decoding and reconstructing the brightness sample points of video data blocks and applying a linear prediction model based on edge information classification, the problem of low cross-component prediction efficiency in the prior art is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202380072452.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-10-11
- Publication Date
- 2025-05-16
AI Technical Summary
When existing video encoding and decoding technologies apply cross-component prediction, it is difficult to effectively improve the encoding and decoding efficiency of image/video blocks.
By receiving the encoded block of the brightness sample points of the video data block, the brightness sample points are decoded and reconstructed, and classified them into multiple sample groups based on the edge information of the brightness sample points, and a corresponding linear prediction model is applied to predict the chrominance sample points.
Improve the encoding and decoding efficiency of video data blocks, reduce the bit rate, and maintain video quality.
Smart Images

Figure CN120019640A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based on and claims the benefit of provisional application No. 63 / 415,532 filed on October 12, 2022 and provisional application No. 63 / 416,220 filed on October 14, 2022. The entire contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present application relates to video coding and compression. More specifically, the present application relates to a method and apparatus for improving coding efficiency of image / video blocks applying cross-component prediction techniques. Background Art
[0004] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data across a communication network, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of a communication network and the limited memory resources of a storage device, video codecs can be used to compress video data according to one or more video codec standards before the video data is transmitted or stored. For example, video codec standards include general video codecs (VVC), joint exploration test models (JEM), high efficiency video codecs (HEVC / H.265), advanced video codecs (AVC / H.264), moving picture experts group (MPEG) codecs, etc. Video codecs typically utilize prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), which utilize the redundancy inherent in video data. Video codecs are intended to compress video data into a form using a lower bit rate while avoiding or minimizing the degradation of video quality. Summary of the invention
[0005] Embodiments of the present disclosure provide methods and apparatus for improving encoding and decoding efficiency of image / video blocks to which cross-component prediction techniques are applied.
[0006] The following presents a simplified overview of one or more aspects of the present disclosure in order to provide a basic understanding of such aspects. This overview is not an extensive overview of all contemplated aspects, and is neither intended to identify the primary or critical elements of all aspects, nor to describe the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to a more detailed description presented later.
[0007] According to one aspect of the present disclosure, a method for decoding video data is provided. The method includes: receiving an encoded block of luma samples for a block of the video data; decoding the encoded block of luma samples to obtain reconstructed luma samples of the block; classifying the luma samples of the block into one of a plurality of sample groups based on edge information of the luma samples, wherein the luma samples are obtained from one or more of the reconstructed luma samples to correspond to chroma samples of the block; and predicting the chroma samples by applying one of a plurality of linear prediction models corresponding to the classified sample groups to the luma samples.
[0008] According to one aspect of the present disclosure, a computer system is provided, comprising one or more processors and one or more storage devices storing computer-executable instructions, wherein when the computer-executable instructions are executed, the one or more processors perform operations including: receiving an encoded block of luma samples for a block of video data; decoding the encoded block of luma samples to obtain reconstructed luma samples for the block; classifying the luma samples for the block into one of a plurality of sample groups based on edge information of the luma samples, wherein the luma sample is obtained from one or more of the reconstructed luma samples to correspond to chroma samples of the block; and predicting the chroma samples by applying one of a plurality of linear prediction models corresponding to the classified sample group to the luma samples.
[0009] According to one aspect of the present disclosure, a method for video decoding using an edge-classified linear model (ELM) is provided. The method may include: receiving an encoded block of luma samples for a first block of a video signal; decoding the encoded block of luma samples to obtain reconstructed luma samples; classifying the reconstructed luma samples into a plurality of sample groups based on the direction and strength of edge information; applying different linear prediction models to the reconstructed luma samples in different sample groups; and predicting chroma samples for the first block of the video signal based on the applied linear prediction models.
[0010] According to one aspect of the present disclosure, a method for video decoding using a filter-based linear model (FLM) is provided. The method may include: receiving an encoded block of luma samples for a first block of a video signal; decoding the encoded block of luma samples to obtain reconstructed luma samples; determining a luma sample region and a chroma sample region to derive a multivariate linear regression (MLR) model; deriving the MLR model by pseudo-inverse matrix calculation; applying the MLR model to the reconstructed luma samples; and predicting chroma samples for the first block of the video signal based on the applied MLR model.
[0011] According to one aspect of the present disclosure, a method for video decoding using a gradient linear model (GLM) is provided. The method may include: receiving an encoded block of luma samples for a first block of a video signal; decoding the encoded block of luma samples to obtain reconstructed luma samples; using the sample gradient to develop a correlation between luma AC information and chroma intensity; determining a luma sample area and a chroma sample area to derive a multivariate linear regression (MLR) model; deriving the MLR model by pseudo-inverse matrix calculation; applying the MLR model to the reconstructed luma samples; and predicting chroma samples for the first block of the video signal based on the applied MLR model.
[0012] According to one aspect of the present disclosure, a method for video coding without utilizing a downsampling process in a convolutional cross-component model (CCCM) is provided. The method may include: receiving an encoded block of luma samples for a first block of a video signal; decoding the encoded block of luma samples to obtain reconstructed luma samples; utilizing a non-downsampled luma reference value and / or a different selection of a non-downsampled luma reference; determining a luma sample region and a chroma sample region to derive a convolutional cross-component model (CCCM); deriving the CCCM parameters by LDL decomposition; applying the CCCM to the reconstructed luma samples; and predicting the chroma samples for the first block of the video signal based on the applied CCCM.
[0013] According to one aspect of the present disclosure, a method for video encoding and decoding using LDL decomposition in a cross-component linear model (CCLM) / multi-model LM (MMLM) is provided. The method may include: receiving a coded block of luma samples for a first block of a video signal; decoding the coded block of luma samples to obtain reconstructed luma samples; determining a luma sample area and a chroma sample area to derive a cross-component linear model (CCLM) / multi-model LM (MMLM); deriving the CCLM / MMLM parameters by LDL decomposition; applying the CCLM / MMLM to the reconstructed luma samples; and predicting the chroma samples for the first block of the video signal based on the applied CCLM / MMLM.
[0014] According to one aspect of the present disclosure, a method for video coding and decoding with minimum sample restriction in FLM / GLM / ELM / CCCM is provided. The method may include: determining whether to apply the FLM / GLM / ELM / CCCM scheme in the intra prediction, wherein the number of samples is greater than or equal to a predefined number in the encoded block.
[0015] According to one aspect of the present disclosure, a method for performing video coding and decoding using a non-subsampled luma reference value and a subsampled luma reference value in CCCM is provided.
[0016] According to one aspect within the present disclosure, a method for performing video encoding and decoding using a plurality of modes of a combination of FLM / GLM / ELM / CCCM / CCLM is provided.
[0017] According to one aspect of the present disclosure, a method for decoding video data is provided, comprising: obtaining a video block from a bitstream; obtaining an internal luma sample value of the video block, an external luma sample value of an external region of the video block, and an external chroma sample value of the external region; determining a set of weighting coefficients corresponding to a filter shape based on the external luma sample value and the external chroma sample value, wherein the filter shape and the set of weighting coefficients are configured to predict each of the internal chroma sample values based on a plurality of corresponding luma sample values, wherein the plurality of corresponding luma sample values include: one or more non-subsampled luma sample values and one or more non-linear luma sample values associated with the filter shape; predicting the internal chroma sample values based on the internal luma sample values using the filter shape and the set of weighting coefficients; and obtaining a predicted video block using the predicted internal chroma sample values.
[0018] According to one aspect of the present disclosure, a method for encoding video data is provided, comprising: obtaining a video block; obtaining an internal luma sample value of the video block, an external luma sample value of an external region of the video block, and an external chroma sample value of the external region; determining a set of weighting coefficients corresponding to a filter shape based on the external luma sample value and the external chroma sample value, wherein the filter shape and the set of weighting coefficients are configured to predict each of the internal chroma sample values based on a plurality of corresponding luma sample values, wherein the plurality of corresponding luma sample values include: one or more non-subsampled luma sample values and one or more non-linear luma sample values associated with the filter shape; predicting the internal chroma sample values based on the internal luma sample values using the filter shape and the set of weighting coefficients; and generating a bitstream including an encoded video block by using the predicted internal chroma sample values.
[0019] According to one aspect of the present disclosure, a computer system is provided, comprising: one or more processors; and one or more storage devices storing computer executable instructions, which, when executed, cause the one or more processors to perform the operations of the method of the present disclosure.
[0020] According to one aspect of the present disclosure, there is provided a computer program product storing computer executable instructions that, when executed, cause one or more processors to perform the operations of the method of the present disclosure.
[0021] According to one aspect of the present disclosure, a computer-readable storage medium storing instructions is provided, which, when executed by a computing device having one or more processors, causes the one or more processors to perform the decoding method of the present disclosure and store a bit stream to be decoded by the decoding method of the present disclosure.
[0022] According to one aspect of the present disclosure, a computer-readable storage medium storing instructions is provided, which, when executed by a computing device having one or more processors, causes the one or more processors to perform the encoding method of the present disclosure and store a bit stream generated by the encoding method of the present disclosure.
[0023] According to one aspect of the present disclosure, a computer-readable medium storing a bit stream is provided, wherein the bit stream is to be decoded by performing the operations of the method of the present disclosure.
[0024] According to one aspect of the present disclosure, a computer-readable medium storing a bit stream is provided, wherein the bit stream is obtained by executing the operation of the method of the present disclosure.
[0025] It is to be understood that both the foregoing general description and the following detailed description are examples only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0027] Figure 1 is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to some implementations of the present disclosure.
[0028] Figure 2 is a block diagram illustrating an exemplary video encoder according to some implementations of the present disclosure.
[0029] Figure 3 is a block diagram illustrating an exemplary video decoder according to some implementations of the present disclosure.
[0030] Figures 4A to 4E is a block diagram illustrating how a frame may be recursively partitioned into multiple video blocks of different sizes and shapes according to some implementations of the present disclosure.
[0031] Figure 5 An overall diagram of a block-based video encoder for VVC is shown.
[0032] Figures 6A to 6E is a schematic diagram showing block partitioning in VVC.
[0033] Figure 7 An overall diagram of a video decoder for VVC is shown.
[0034] Figure 8 is a schematic diagram showing the positions of sample points used to derive α and β.
[0035] Figures 9A to 9C is a schematic diagram showing examples of MDLM, MDLM_L, and MDLM_T.
[0036] Fig.10 is a schematic diagram showing an example of classifying adjacent sample points into two groups.
[0037] Fig.11 is a schematic diagram showing the inflection point T.
[0038] Figures 12A to 12B is a schematic diagram showing the effect of the slope adjustment parameter "u".
[0039] Fig.13 is a schematic diagram showing the co-located reconstructed luminance samples used.
[0040] Fig.14 is a schematic diagram showing the used adjacent reconstruction samples.
[0041] Figures 15A to 15D is a schematic diagram showing the process of intra mode derivation at the decoder side.
[0042] Fig.16 is a diagram showing an example of four reference lines adjacent to a prediction block.
[0043] Fig.17 is a schematic diagram showing the spatial part of a convolutional filter.
[0044] Fig.18 is a schematic diagram showing the reference area (with its filling) used for deriving filter coefficients.
[0045] Figures 19A to 19B is a schematic diagram showing that chrominance samples can be simultaneously associated with multiple luma samples.
[0046] Fig. 20 is a schematic diagram showing that coefficients / offsets of multiple (eg, 6) luma samples relative to one chroma sample are trained to linearly predict chroma samples within a CU.
[0047] Fig.21is a schematic diagram showing that different chroma types / color formats may have different predefined filter shapes / taps.
[0048] Fig. 22 is a schematic diagram showing that FLM can use only the top or left luma / chroma samples (extended) for parameter derivation.
[0049] Fig.23 is a schematic diagram showing that FLM can use different lines for parameter derivation.
[0050] Figures 24A to 24D is a schematic diagram showing a pre-operation before applying the MLR model (GLM 1-tap / 2-tap).
[0051] Fig.25 is a schematic diagram showing examples of different shapes / numbers of filter taps.
[0052] Fig.26 is a schematic diagram showing examples of different shapes / numbers of filter taps.
[0053] Figures 27A to 27B is a schematic diagram showing examples of different shapes / numbers of filter taps.
[0054] Figures 28A to 28G is a schematic diagram showing examples of different sets of filter taps.
[0055] Figures 29A to 29B is a schematic diagram showing 2-fold training for implicit filter shape derivation.
[0056] Fig.30 A workflow of a method for decoding video data according to one or more aspects of the present disclosure is shown.
[0057] Fig.31 A workflow of a method for encoding video data according to one or more aspects of the present disclosure is shown.
[0058] Fig.32 is a schematic diagram illustrating a computing environment coupled with a user interface according to some implementations of the present disclosure. DETAILED DESCRIPTION
[0059] Reference will now be made in detail to specific implementations, examples of which are shown in the accompanying drawings. In the following detailed description, many non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.
[0060] It should be noted that the terms "first", "second", etc. used in the description, claims, and drawings of the present disclosure are used to distinguish objects and are not used to describe any specific order or sequence. It should be understood that the terms used in this manner can be interchanged under appropriate conditions, so that the embodiments of the present disclosure described herein can be implemented in an order other than those shown in the drawings or described in the present disclosure.
[0061] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. Figure 1 As shown in , system 10 includes a source device 12 that generates and encodes video data to be decoded at a later time by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0062] In some implementations, the destination device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may include a communication medium so that the source device 12 can send the encoded video data directly to the destination device 14 in real time. The encoded video data may be modulated and sent to the destination device 14 according to a communication standard such as a wireless communication protocol. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be used to facilitate communication from the source device 12 to the destination device 14.
[0063] In some other implementations, the encoded video data may be sent from the output interface 22 to the storage device 32. Subsequently, the encoded video data in the storage device 32 may be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device 32 may correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing encoded video data and sending the encoded video data to the destination device 14. Exemplary file servers include a network server (e.g., for a website), a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Destination device 14 may access the encoded video data through any standard data connection suitable for accessing encoded video data stored on a file server, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both. The transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
[0064] like Figure 1 As shown in , source device 12 includes video source 18, video encoder 20 and output interface 22. Video source 18 may include a source such as a video capture device, for example, a camera, a video archive containing previously captured video, a video feed interface receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, if video source 18 is a camera of a security surveillance system, source device 12 and destination device 14 may form a camera phone or a video phone. However, the implementation described in this application may be generally applicable to video encoding and decoding, and may be applied to wireless and / or wired applications.
[0065] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be sent directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored on storage device 32 for later access by destination device 14 or other devices for decoding and / or playback. Output interface 22 may also include a modem and / or a transmitter.
[0066] Destination device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 may include a receiver and / or a modem, and receives encoded video data over link 16. The encoded video data transmitted over link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data sent over a communication medium, stored on a storage medium, or stored on a file server.
[0067] In some implementations, the destination device 14 may include a display device 34, which may be an integrated display device or an external display device, configured to communicate with the destination device 14. The display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0068] The video encoder 20 and the video decoder 30 may operate according to proprietary or industry standards such as VVC, HEVC, MPEG-4 Part 10, AVC, or extensions of such standards. It should be understood that the present application is not limited to a specific video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that the video encoder 20 of the source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is also generally contemplated that the video decoder 30 of the destination device 14 may be configured to decode video data according to any of these current or future standards.
[0069] The video encoder 20 and the video decoder 30 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders, any of which can be integrated in the corresponding device as part of a combined encoder / decoder (CODEC).
[0070] In some implementations, at least a portion of the components of source device 12 (e.g., video source 18, as described below with reference to Figure 2 The video encoder 20 described herein or components included in the video encoder 20 and the output interface 22) and / or at least a portion of the components of the destination device 14 (eg, the input interface 28, as described below with reference to Figure 3 The video decoder 30 described herein or the components included in the video decoder 30 and the display device 34) may be operated in a cloud computing service network, which may provide software, platform and / or infrastructure, such as software as a service (SaaS), platform as a service (PaaS) or infrastructure as a service (IaaS). In some implementations, one or more components in the source device 12 and / or the destination device 14 that are not included in the cloud computing service network may be provided in one or more client devices, and the one or more client devices may communicate with a server computer in the cloud computing service network through a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In an embodiment, at least a portion of the operations described herein may be implemented as a cloud-based service provided by one or more server computers, which are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the destination device 14 in the cloud computing service network; and one or more other operations described herein may be implemented by one or more client devices. In some implementations, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of the present disclosure, terms such as "cloud," "cloud computing," "cloud-based," etc., may be used interchangeably herein when appropriate. It should be understood that the present disclosure is not limited to implementation in the above-described cloud computing service network. Rather, the present disclosure may also be implemented in any other type of computing environment currently known or developed in the future.
[0071] Figure 2 is a block diagram illustrating an exemplary video encoder 20 according to some implementations described in the present application. The video encoder 20 can perform intra-frame and inter-frame predictive coding of video blocks within a video frame. Intra-frame predictive coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-frame predictive coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. It should be noted that the term "frame" can be used as a synonym for the term "image" or "picture" in the field of video coding and decoding.
[0072] like Figure 2As shown in FIG. 1 , the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (BC) unit 48. In some implementations, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62 for video block reconstruction. A loop filter 63 (such as a deblocking filter) can be positioned between the summer 62 and the DPB 64 to filter the block boundaries to remove block artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter such as a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF) can be used to filter the output of the summer 62. It should be noted that for the CCSAO technique, the present application is not limited to the embodiments described herein, and alternatively, the present application may be applied to a case where an offset for any other of the luminance component, the Cb chrominance component, and the Cr chrominance component is selected according to any of the luminance component, the Cb chrominance component, and the Cr chrominance component to modify the any other component based on the selected offset. In addition, it should also be noted that the first component mentioned herein may be any one of the luminance component, the Cb chrominance component, and the Cr chrominance component, the second component mentioned herein may be any other of the luminance component, the Cb chrominance component, and the Cr chrominance component, and the third component mentioned herein may be the remaining component of the luminance component, the Cb chrominance component, and the Cr chrominance component. In some examples, the loop filter may be omitted, and the decoded video block may be provided directly to the DPB 64 by the summer 62. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided between one or more of the fixed or programmable hardware units shown.
[0073] Video data memory 40 may store video data to be encoded by components of video encoder 20. Video data in video data memory 40 may be obtained, for example, from Figure 1 6 is obtained from the video source 18 shown in . DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) used by the video encoder 20 (e.g., in intra-frame or inter-frame predictive coding mode) when encoding video data. Video data memory 40 and DPB 64 can be formed by any of a variety of memory devices. In various examples, video data memory 40 can be on-chip with other components of the video encoder 20, or off-chip relative to those components.
[0074] like Figure 2 As shown in , after receiving the video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. The segmentation may also include segmenting the video frame into slices, tiles (e.g., video block sets) or other larger coding units (CUs) according to a predefined segmentation structure (such as a quadtree (QT) structure associated with the video data). A video frame is or may be considered as a two-dimensional array or matrix of samples with sample values. The samples in the array may also be referred to as the basic unit of a pixel or image. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. The video frame may be divided into a plurality of video blocks by, for example, using QT segmentation. A video block is again or may be considered as a two-dimensional array or matrix of samples with sample values, but has a smaller dimension than a video frame. The number of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. A video block may also be segmented into one or more block partitions or sub-blocks (which may again form blocks) by, for example, iteratively using QT segmentation, binary tree (BT) segmentation, or ternary tree (TT) segmentation, or any combination thereof. It should be noted that the term "block" or "video block" as used herein may be a portion of a frame or picture, in particular a rectangular (square or non-square) portion. For example, with reference to HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU) or a transform unit (TU), and / or may be or correspond to a corresponding block, such as a coding tree block (CTB), a coding block (CB), a prediction block (PB) or a transform block (TB) and / or a sub-block.
[0075] Prediction processing unit 41 may select one of a plurality of possible prediction coding modes for the current video block based on the error results (e.g., coding rate and distortion level), such as one of a plurality of intra-frame predictive coding modes or one of a plurality of inter-frame predictive coding modes. Prediction processing unit 41 may provide the resulting intra-frame or inter-frame predictive coded block to summer 50 to generate a residual block, and to summer 62 to reconstruct the coded block for subsequent use as part of a reference frame. Prediction processing unit 41 also provides syntax elements such as motion vectors, intra-frame mode indicators, partition information, and other such syntax information to entropy encoding unit 56.
[0076] To select an appropriate intra-predictive coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-predictive coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-predictive coding of the current video block relative to one or more predictive blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple encoding passes, e.g., to select an appropriate coding mode for each block in the video data.
[0077] In some implementations, motion estimation unit 42 determines the inter-prediction mode for a current video frame by generating motion vectors according to a predetermined pattern within a sequence of video frames, the motion vector indicating the displacement of a video block within the current video frame relative to a predictive block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors, which estimate the motion of video blocks. For example, a motion vector may indicate the displacement of a video block within a current video frame or picture relative to a predictive block within a reference frame relative to a current block being encoded within the current frame. The predetermined pattern may designate a video frame in a sequence as a P frame or a B frame. Intra BC unit 48 may determine a vector (e.g., a block vector) for intra BC coding in a manner similar to the motion vectors determined by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine the block vector.
[0078] The predictive block for the video block may be or may correspond to a block or reference block of a reference frame that is considered to closely match the video block to be encoded in terms of pixel difference, which may be determined by a sum of absolute differences (SAD), a sum of squared differences (SSD), or other difference metrics. In some implementations, video encoder 20 may calculate values for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, motion estimation unit 42 may perform a motion search relative to full pixel positions and fractional pixel positions, and output a motion vector with fractional pixel precision.
[0079] Motion estimation unit 42 calculates a motion vector for a video block in an inter-prediction encoded frame by comparing the position of the video block to the position of a predictive block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.
[0080] The motion compensation performed by the motion compensation unit 44 may involve extracting or generating a predictive block based on the motion vector determined by the motion estimation unit 42. Upon receiving the motion vector for the current video block, the motion compensation unit 44 may locate the predictive block pointed to by the motion vector in one of the reference frame lists, retrieve the predictive block from the DPB 64, and forward the predictive block to the summer 50. Summer 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the predictive block provided by the motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel difference values forming the residual video block may include luma component differences or chroma component differences or both. The motion compensation unit 44 may also generate syntax elements associated with the video block of the video frame for use by the video decoder 30 when decoding the video block of the video frame. The syntax elements may include, for example, syntax elements defining the motion vector used to identify the predictive block, any flag indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are illustrated separately for conceptual purposes.
[0081] In some implementations, the intra BC unit 48 may generate a vector and extract a predictive block in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44, but the predictive block is in the same frame as the current block being encoded and the vector is referred to as a block vector rather than a motion vector. In particular, the intra BC unit 48 may determine an intra prediction mode for encoding the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, during a separate encoding pass, and test its performance through rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may use rate-distortion analysis for various tested intra prediction modes to calculate rate-distortion values, and select an intra prediction mode with the best rate-distortion characteristic among the tested modes as the appropriate intra prediction mode to be used. The rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original, unencoded block that was encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. Intra BC unit 48 may calculate ratios based on the distortion and rates of the various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.
[0082] In other examples, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction in accordance with implementations described herein. In either case, for intra block copying, a predictive block may be a block that is deemed to closely match the block to be encoded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identification of the predictive block may include calculation of values for sub-integer pixel positions.
[0083] Regardless of whether the predictive block is from the same frame according to intra-frame prediction or a different frame according to inter-frame prediction, the video encoder 20 can form a residual video block by subtracting the pixel values of the predictive block from the pixel values of the current video block being encoded, thereby forming pixel difference values. The pixel difference values forming the residual video block may include both luma component differences and chroma component differences.
[0084] The intra-prediction processing unit 46 may intra-predict the current video block as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44 or the intra-block copy prediction performed by the intra BC unit 48, as described above. In particular, the intra-prediction processing unit 46 may determine an intra-prediction mode to encode the current block. To this end, the intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, for example, during separate encoding passes, and the intra-prediction processing unit 46 (or in some examples, the mode selection unit) may select an appropriate intra-prediction mode to be used from the tested intra-prediction modes. The intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra-prediction mode into the bitstream.
[0085] After prediction processing unit 41 determines the predictive block for the current video block via inter-prediction or intra-prediction, summer 50 forms a residual video block by subtracting the predictive block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0086] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of the matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 may perform the scan.
[0087] After quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream may then be encoded as Figure 1 is sent to the video decoder 30 as shown in Figure 1 3 is archived in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 may also entropy encode motion vectors and other syntax elements for the current video frame being encoded.
[0088] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain to generate a reference block for prediction of other video blocks. As mentioned above, motion compensation unit 44 may generate a motion compensated predictive block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the predictive block to calculate sub-integer pixel values used in motion estimation.
[0089] Summer 62 adds the reconstructed residual block to the motion compensated predictive block produced by motion compensation unit 44 to produce a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a predictive block to inter-predict another video block in a subsequent video frame.
[0090] Figure 3 is a block diagram illustrating an exemplary video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, and a DPB 92. The prediction processing unit 81 also includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 may perform operations generally similar to those described above with respect to combining Figure 2 The decoding process is the inverse of the encoding process described for the video encoder 20. For example, the motion compensation unit 82 may generate prediction data based on the motion vector received from the entropy decoding unit 80, and the intra-prediction unit 84 may generate prediction data based on the intra-prediction mode indicator received from the entropy decoding unit 80.
[0091] In some examples, units of the video decoder 30 may be tasked to perform implementations of the present application. Furthermore, in some examples, implementations of the present disclosure may be divided between one or more of the units of the video decoder 30. For example, the intra BC unit 85 may perform implementations of the present application alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81, such as the motion compensation unit 82.
[0092] The video data memory 79 may store video data (such as an encoded video bitstream) to be decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from a storage device 32, via a wired or wireless network communication of video data, or from a local video source (such as a camera) by accessing a physical data storage medium (e.g., a flash drive or hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from an encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data for use when decoding video data by the video decoder 30 (e.g., in an intra-frame or inter-frame predictive coding mode). The video data memory 79 and the DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 is shown in FIG. 1 . Figure 3 9 depicts video data memory 79 and DPB 92 as two separate components of video decoder 30. However, it will be appreciated by those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, video data memory 79 may be on-chip with other components of video decoder 30, or off-chip relative to those components.
[0093] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. Video decoder 30 may receive syntax elements at the video frame level and / or the video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra-frame prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors or intra-frame prediction mode indicators and other syntax elements to prediction processing unit 81.
[0094] When a video frame is encoded as an intra-predictively coded (I) frame or for intra-coded predictive blocks in other types of frames, intra-prediction unit 84 of prediction processing unit 81 may generate prediction data for a video block of the current video frame based on a signaled intra-prediction mode and reference data from a previously decoded block of the current frame.
[0095] When the video frame is encoded as an inter-frame predictive coded (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 generates one or more predictive blocks for a video block of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the predictive blocks may be generated from a reference frame within one of the reference frame lists. Video decoder 30 may construct reference frame lists (list 0 and list 1) using a default construction technique based on the reference frames stored in DPB 92.
[0096] In some examples, when the video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a predictive block for the current video block based on the block vector and other syntax elements received from entropy decoding unit 80. The predictive block may be within a reconstructed region of the same picture as the current video block defined by video encoder 20.
[0097] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for a video block of the current video frame by parsing the motion vector and other syntax elements, and then uses the prediction information to generate a predictive block for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra or inter prediction) used to encode the video block of the video frame, the inter prediction frame type (e.g., B or P), construction information for one or more of the reference frame lists for the frame, the motion vector for each inter-predictively encoded video block of the frame, the inter prediction state for each inter-predictively encoded video block of the frame, and other information used to decode the video block in the current video frame.
[0098] Similarly, the intra BC unit 85 may use some of the received syntax elements (e.g., flags) to determine that the current video block is predicted using the intra BC mode, construction information of which video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, block vectors for each intra BC predicted video block of the frame, the intra BC prediction state of each intra BC predicted video block of the frame, and other information used to decode the video block in the current video frame.
[0099] Motion compensation unit 82 may also perform interpolation to calculate interpolated values for sub-integer pixels of reference blocks using interpolation filters as used by video encoder 20 during encoding of the video blocks. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 from the received syntax elements and use the interpolation filters to produce the predictive blocks.
[0100] The quantized transform coefficients provided in the bitstream are inverse quantized by inverse quantization unit 86 and entropy decoded by entropy decoding unit 80 using the same quantization parameters calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct residual blocks in the pixel domain.
[0101] After the motion compensation unit 82 or the intra BC unit 85 generates a predictive block for the current video block based on the vector and other syntax elements, the summer 90 reconstructs the decoded video block for the current video block by summing the residual block from the inverse transform processing unit 88 and the corresponding predictive block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter 91 (such as a deblocking filter, an SAO filter, a CCSAO filter, and / or an ALF) may be positioned between the summer 90 and the DPB 92 to further process the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be provided directly to the DPB 92 by the summer 90. The decoded video blocks in a given frame are then stored in the DPB 92, which stores reference frames used for subsequent motion compensation of the next video block. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later display on a display device (such as Figure 1 is presented on a display device 34).
[0102] In a typical video encoding and decoding process, a video sequence usually includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other cases, a frame may be monochrome and therefore include only a two-dimensional array of luma samples.
[0103] like Figure 4AAs shown in , the video encoder 20 (or more specifically, the partition unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs sequentially ordered in a raster scan order from left to right and from top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size of one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As Figure 4B As shown in , each CTU may include one CTB of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements used to encode the samples of the coding tree blocks. The syntax elements describe the properties of different types of units of the encoded pixel blocks and how the video sequence can be reconstructed at the video decoder 30, including inter-frame or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements used to encode the samples of the coding tree block. The coding tree block may be an N×N sample block.
[0104] To achieve better performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof, on the coding tree block of the CTU and divide the CTU into smaller CUs. Figure 4C As depicted in FIG, a 64x64 CTU 400 is first divided into four smaller CUs, each CU having a block size of 32x32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs of a block size of 16x16. Two 16x16 CUs 430 and 440 are each further divided into four CUs of a block size of 8x8. Figure 4D Depicts Figure 4C The quadtree data structure of the final result of the partitioning process of CTU 400 is depicted in FIG. 4 , where each leaf node of the quadtree corresponds to a CU of a corresponding size ranging from 32x32 to 8x8. Figure 4B In the CTU depicted in FIG, each CU may include a CB of luma samples and two corresponding coding blocks of chroma samples of the same size frame, and syntax elements used to encode the samples of the coding blocks. In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and syntax structures used to encode the samples of the coding block. It should be noted that Figure 4C and Figure 4DThe quadtree partitioning depicted in FIG is for illustrative purposes only, and a CTU can be partitioned into CUs to adapt to varying local characteristics based on quadtree / ternary / binary tree partitioning. In multiple types of tree structures, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU can be further partitioned by a binary tree and a ternary tree structure. Figure 4E As shown in , there are five possible partitioning types of a coding block with width W and height H, namely, quadruple partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.
[0105] In some implementations, the video encoder 20 may further partition the coding block of a CU into one or more MxN PBs. A PB is a rectangular (square or non-square) block of samples on which the same prediction (inter or intra) is applied. The PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements used to predict the PBs. In a monochrome picture or a picture with three separate color planes, a PU may include a single PB and a syntax structure used to predict the PB. The video encoder 20 may generate predictive luma, Cb, and Cr blocks for the luma, Cb, and Cr PBs of each PU of the CU.
[0106] Video encoder 20 may use intra prediction or inter prediction to generate the predictive blocks of a PU. If video encoder 20 uses intra prediction to generate the predictive blocks of a PU, video encoder 20 may generate the predictive blocks of the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate the predictive blocks of a PU, video encoder 20 may generate the predictive blocks of the PU based on decoded samples of one or more frames different from the frame associated with the PU.
[0107] After the video encoder 20 generates the predictive luma, Cb, and Cr blocks of one or more PUs of a CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predictive luma block of the CU from its original luma coding block, so that each sample point in the luma residual block of the CU indicates the difference between a luma sample in one of the predictive luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, so that each sample point in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predictive Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample point in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predictive Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0108] In addition, if Figure 4CAs shown in , the video encoder 20 can use quadtree partitioning to decompose the luma, Cb and Cr residual blocks of a CU into one or more luma, Cb and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block on which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements used to transform the transform block samples. Therefore, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure used to transform the samples of the transform block.
[0109] The video encoder 20 may apply one or more transforms to the luma transform block of the TU to generate a luma coefficient block of the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficient may be a scalar. The video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block of the TU. The video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block of the TU.
[0110] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized to possibly reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream, which includes a sequence of bits forming a representation of an encoded frame and associated data, which is stored in the storage device 32 or sent to the destination device 14.
[0111] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally inverse to the encoding process performed by the video encoder 20. For example, the video decoder 30 may perform an inverse transform on a coefficient block associated with a TU of the current CU to reconstruct a residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coding block of the current CU by adding samples of the predictive block for the PU of the current CU to corresponding samples of the transform block of the TU of the current CU. After reconstructing the coding block of each CU of the frame, the video decoder 30 may reconstruct the frame.
[0112] As described above, video codecs mainly use two modes, i.e., intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction) to achieve video compression. It should be noted that IBC can be regarded as intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to codec efficiency than intra-frame prediction because a motion vector is used to predict a current video block from a reference video block.
[0113] However, with the ever-improving video data capture technology and finer video block sizes for maintaining details in video data, the amount of data required for the motion vector representing the current frame has also increased significantly. One way to overcome this challenge is to benefit from the fact that not only do neighboring CU groups in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, it is possible to use the motion information of spatially neighboring CUs and / or temporally co-located CUs as an approximation of the motion information (e.g., motion vector) of the current CU by exploring the spatial and temporal correlation of the current CU, which is also referred to as the "motion vector predictor (MVP)" of the current CU.
[0114] Instead, it will be combined as above Figure 2 The actual motion vector of the current CU determined by the motion estimation unit 42 is encoded into the video bitstream, and the motion vector prediction value of the current CU is subtracted from the actual motion vector of the current CU to generate the motion vector difference (MVD) of the current CU. By doing so, the motion vector determined by the motion estimation unit 42 for each CU of the frame does not need to be encoded into the video bitstream, and the amount of data used to represent the motion information in the video bitstream can be significantly reduced.
[0115] Similar to the process of selecting a predictive block in a reference frame during inter-frame prediction of a code block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules to construct a motion vector candidate list (also referred to as a "merge list") for the current CU using those potential candidate motion vectors associated with the spatial neighboring CUs and / or the temporally co-located CUs of the current CU, and then select one member from the motion vector candidate list as the motion vector predictor of the current CU. By doing so, the motion vector candidate list itself does not need to be sent from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to encode and decode the current CU using the same motion vector predictor within the motion vector candidate list.
[0116] introduce
[0117] Various video codec techniques can be used to compress video data. Video coding is performed according to one or more video codec standards. For example, video codec standards include general video codec (VVC), high efficiency video codec (H.265 / HEVC), advanced video codec (H.264 / AVC), moving picture experts group (MPEG) codec, etc. Video codecs usually use prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that use redundancy in video images or sequences. An important goal of video coding technology is to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of video quality.
[0118] The first version of the VVC standard was completed in July 2020, which provides approximately 50% bitrate savings or equivalent perceptual quality compared to the existing generation video codec standard HEVC. Although the VVC standard provides significant codec improvements over its predecessor, there is evidence that additional codec tools can be used to achieve excellent codec efficiency. Recently, the Joint Video Exploration Team (JVET) in collaboration with ITU-T VCEG and ISO / IEC MPEG began exploring advanced technologies that can achieve substantial enhancements in codec efficiency over VVC. In April 2021, a software code base called the Enhanced Compression Model (ECM) was established for future video codec exploration work. The ECM reference software is based on the VVC Test Model (VTM) developed by JVET for VVC, in which several existing modules (e.g., intra / inter prediction, transform, loop filtering, etc.) are further extended and / or improved. In the future, any new codec tools beyond the VVC standard can be integrated into the ECM platform and tested using the JVET Common Test Conditions (CTC).
[0119] Similar to all the aforementioned video codec standards, ECM is built on a block-based hybrid video codec framework. Figure 5 A block diagram of a generic block-based hybrid video coding system is shown. The input video signal is processed block by block (referred to as a coding unit (CU)). In ECM-1.0, a CU can be up to 128x128 pixels. However, like VVC, a coding tree unit (CTU) is split into CUs to accommodate varying local characteristics based on quadtree / binarytree / ternarytree. In the multi-type tree structure, a CTU is first partitioned by a quadtree structure. Then, each quadtree leaf node can be further partitioned by a binary tree and a ternary tree structure. As Fig. 6A , Figure 6B , Figure 6C , Fig.6D and Fig. 6E As shown, there are five types of splits: quadruple split, vertical binary split, horizontal binary split, vertical extended quadruple split, and horizontal extended quadruple split.
[0120] exist Figure 5In the video code, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels of samples from coded neighboring blocks (which are called reference samples) in the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion compensated prediction") uses reconstructed pixels from coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal of a given CU is usually signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. In addition, if multiple reference pictures are supported, a reference picture index is additionally sent, which is used to identify which reference picture in the reference picture storage stores the temporal prediction signal. After spatial and / or temporal prediction, the mode decision block in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block; and the prediction residual is decorrelated and quantized using a transform. The quantized residual coefficients are dequantized and inverse transformed to form a reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Further loop filtering (such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF)) can be applied to the reconstructed CU before it is placed in the reference picture storage and used to encode future video blocks. In order to form an output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit to be further compressed and packaged to form a bitstream. It should be noted that the term "block" or "video block" as used herein can be a part of a frame or picture, especially a rectangular (square or non-square) part. For example, with reference to HEVC and VVC, a block or video block can be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU), or a transform unit (TU), and / or can be or correspond to a corresponding block, such as a coding tree block (CTB), a coding block (CB), a prediction block (PB), or a transform block (TB) and / or a sub-block.
[0121] Figure 7 A general block diagram of a block-based video decoder is shown. The video bitstream is first entropy decoded at the entropy decoding unit. The coding mode and prediction information are sent to the spatial prediction unit (if intra-frame coding) or the temporal prediction unit (if inter-frame coding) to form a prediction block. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual block. The prediction block and the residual block are then added together. The reconstructed block can be further subjected to loop filtering before it is stored in the reference picture storage. The reconstructed video in the reference picture storage is then sent out to drive the display device and is used to predict future video blocks.
[0122] The main focus of the present disclosure is to further enhance the codec efficiency of codec tools applied to cross-component prediction, cross-component linear model (CCLM) in ECM. In the following, some related codec tools in ECM are briefly reviewed. After that, some defects in the existing design of CCLM are discussed. Finally, solutions are provided to improve the existing CCLM prediction design.
[0123] Cross-component linear model prediction
[0124] To reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode is used in VVC, for which chroma samples are predicted based on the reconstructed luma samples of the same CU by using the following linear model:
[0125] pred C (i,j)=α·rec L ′(i,j)+β (1)
[0126] where pred C (i, j) represents the predicted chroma sample in the CU, and rec L ′(i,j) represents the reconstructed luminance sample rec of the same CU. L The downsampled reconstructed luma samples obtained by downsampling (i, j). The above α and β are linear model parameters, which are derived from at most four adjacent chroma samples and their corresponding downsampled luma samples, which can be called adjacent luma-chroma sample pairs. Assuming that the current chroma block has a size of W×H, W' and H' are obtained as follows:
[0127] – When LM mode is applied, W'=W, H'=H;
[0128] – When LM-A mode is applied, W'=W+H;
[0129] – When LM-L mode is applied, H'=H+W;
[0130] In LM mode, both the upper and left samples of the CU are used to calculate the linear model coefficients; in LM_A mode, only the upper samples of the CU are used to calculate the linear model coefficients; and in LM_L mode, only the left samples of the CU are used to calculate the linear model coefficients.
[0131] If the positions of the upper neighboring samples of the chroma block are represented as S[0, -1] ... S[W'-1, -1], and the positions of the left neighboring samples of the chroma block are represented as S[-1, 0] ... S[-1, H'-1], the positions of the four neighboring chroma samples are selected as follows:
[0132] – When LM mode is applied and both the upper and left neighboring samples are available, S[W' / 4, -1], S[3*W' / 4, -1], S[-1, H' / 4], S[-1, 3*H' / 4] are selected as the positions of the four adjacent chroma samples;
[0133] – When LM-A mode is applied or only upper adjacent samples are available, S[W' / 8, -1], S[3*W' / 8, -1], S[5*W' / 8, -1], S[7*W' / 8, -1] are selected as the positions of the four adjacent chroma samples;
[0134] – When LM-L mode is applied or only left adjacent samples are available, S[-1, H' / 8], S[-1, 3*H' / 8], S[-1, 5*H' / 8], S[-1, 7*H' / 8] are selected as the positions of the four adjacent chroma samples.
[0135] Four adjacent brightness samples corresponding to the selected position are obtained by downsampling operation, and the obtained four adjacent brightness samples are compared four times to find two larger values: x 0 A and x 1 A , and two smaller values: x 0 B and x 1 B The chrominance sample values corresponding to the two larger values and the two smaller values are represented as y 0 A ,y 1 A ,y 0 B and 1 B Then, X a , X b , Y a and Y b is derived as:
[0136] X a =(x 0 A +x 1 A +1)>>1;
[0137] X b =(x 0 B +x 1 B +1)>>1;
[0138] Y a =(y 0A +y 1 A +1)>>1;
[0139] Y b =(y 0 B +y 1 B +1)>>1 (2)
[0140] Finally, the linear model parameters α and β are obtained according to the following equations.
[0141]
[0142] β=Y b -α·X b (4)
[0143] Figure 8 Examples of the positions of the left and top samples involved in the CCLM mode and the samples of the current block are shown, including the positions of the left and top samples of the N×N chroma block in the CU and the positions of the left and top samples of the 2N×2N luminance block in the CU.
[0144] The division operation for calculating the parameter α is implemented using a lookup table. In order to reduce the memory required for storing the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented by exponential notation. For example, diff is approximated as a 4-bit significant part and an exponent. Therefore, for 16 values of the significant number as follows, the table for 1 / diff is reduced to 16 elements:
[0145] DivTable[]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}(5)
[0146] This will have the benefit of reducing both the complexity of the calculations and the memory size required to store the required tables.
[0147] In addition to the upper and left templates being used together to calculate the linear model coefficients, they can also be used alternatively in the other 2 LM modes (referred to as LM_A and LM_L modes).
[0148] In LM_T mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to (W+H) samples. In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to (H+W) samples.
[0149] In LM_LT mode, the linear model coefficients are calculated using the left and upper templates.
[0150] To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to the luma samples to achieve a 2 to 1 downsampling ratio in both the horizontal and vertical directions. The choice of downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "type 0" and "type 2" content, respectively.
[0151]
[0152] Note that when the above reference line is located at a CTU boundary, only one luma line (common line buffer in intra prediction) is used to get the downsampled luma samples.
[0153] This parameter calculation is performed as part of the decoding process, not just as an encoder search operation. Therefore, no syntax is used to convey the values of α and β to the decoder.
[0154] For chroma intra mode coding, a total of 8 intra modes are allowed for chroma intra mode coding. These modes include five traditional intra modes and three cross-component linear model modes (CCLM, LM_A and LM_L). The chroma mode signaling and derivation process are shown in Table 1. Chroma mode coding depends directly on the intra prediction mode of the corresponding luminance block. Since separate block partitioning structures for luminance components and chrominance components are enabled in I slices, one chroma block can correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited.
[0155] Table 1 - Chroma prediction modes derived from luma mode when CCLM is enabled
[0156]
[0157] Regardless of the value of sps_cclm_enabled_flag, a single binarization table is used, as shown in Table 2.
[0158] Table 2 - Unified binarization table for chroma prediction mode
[0159] Value of intra_chroma_pred_mode Binary bit string 4 00 0 0100 1 0101 2 0110 3 0111 5 10 6 110 7 111
[0160] In Table 2, the first binary bit indicates whether it is a normal mode (0) or a LM mode (1). If it is a LM mode, then the next binary bit indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, then the next binary bit indicates whether it is LM_L (0) or LM_A (1). For this case, when sps_cclm_enabled_flag is 0, the first binary bit of the binarization table for the corresponding intra_chroma_pred_mode can be discarded before entropy coding and decoding. Or, in other words, the first binary bit is inferred to be 0 and is therefore not coded and decoded. This single binarization table is used for both the cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two binary bits in Table 2 are context coded using their own context model, and the remaining binary bits are bypass coded.
[0161] In addition, to reduce luma-chroma latency in dual trees, when a 64x64 luma coding tree node is split using no split (and ISP is not used for 64x64 CUs) or QT, chroma CUs in 32x32 / 32x16 chroma coding tree nodes are allowed to use CCLM in the following manner:
[0162] – If a 32x32 chroma node is not split or is split by a partitioned QT, all chroma CUs in the 32x32 node can use CCLM
[0163] – If a 32x32 chroma node is split using horizontal BT and the 32x16 child nodes are not split or use vertical BT split, all chroma CUs in the 32x16 chroma node can use CCLM.
[0164] Under all other luma and chroma coding tree split conditions, CCLM is not allowed for chroma CUs.
[0165] During the development of the ECM, the simplified derivation of α and β (min-max approximation) was removed. Instead, the model parameters α and β were derived using a linear least squares solution between the causal reconstructed data of the downsampled luma samples and the causal chroma samples.
[0166]
[0167] Among them, Rec C (i) and Rec' L (i) indicates the reconstructed chroma samples and down-sampled reconstructed luminance samples around the target block, and I indicates the total number of samples of neighboring data.
[0168] The LM_A and LM_L models are also called multi-directional linear models (MDLM). Fig. 9AAn example of the operation of the MDLM is shown when the block content cannot be predicted from the L-shaped reconstructed region. Fig. 9B MDLM_L is shown using only the left reconstructed samples to derive CCLM parameters. Fig. 9C MDLM_T is shown using only the top reconstruction samples to derive CCLM parameters.
[0169] JCTVC-C206: Integration
[0170] The integration of the least mean square (LMS) discussed above (see equations (8)-(9)) has been proposed as an improvement to CCLM. The initial integration design of the LMS CCLM was first proposed in JCTVC-C206. The method was then improved through a series of simplifications, including the addition of α to the precision n α JCTVC-F0233 / I0178 which reduced the maximum multiplier bit width from 13 to 7, JCTVC-I0151 which reduced the maximum multiplier bit width, and JCTVC-H0490 / I0166 which reduced the division LUT entries from 64 to 32, finally leading to the ECM LMS version.
[0171] Basic Algorithm
[0172] As discussed in equation (1), the integrated design utilizes a linear relationship to model the correlation of the luma and chroma signals. The chroma values are predicted from the reconstructed luma values of the co-located blocks.
[0173] The luminance component and the chrominance component have different sampling rates in YUV420 sampling. The sampling rate of the chrominance component is half of the sampling rate of the luminance component, and has a 0.5 pixel phase difference in the vertical direction. The reconstructed luminance needs to be downsampled in the vertical direction and subsampled in the horizontal direction to match the size of the chrominance signal. For example, downsampling can be achieved by:
[0174] Rec L ′(i,j)=(rec L (2i,2j)+rec L (2i,2j+1))>>1 (10)
[0175] Integer Implementation
[0176] Floating point operations are required to calculate the linear model parameter α in equation (8) to maintain high data accuracy. And when α is represented by a floating point value, floating point multiplication is involved in equation (1). In this section, an integer implementation of the algorithm is designed. Specifically, using n α The fractional part of parameter α is quantized with a bit data accuracy. The value of parameter α is obtained by expanding and rounding the integer value α' and a'=a×(1<<n α). Then, the linear model of equation (1) is changed to:
[0177] pred C [x,y]=(α'·Rec L '[x,y]>>n α )+β' (11)
[0178] where β′ is the rounded value of the floating point β, and α′ can be calculated as follows.
[0179]
[0180] It is proposed to replace the division operation of equation (12) by table lookup and multiplication. First, A2 is reduced to reduce the table size. A1 is also reduced to avoid product overflow. Then, in A2, only the The most significant bit of the value is defined, while the other bits are set to zero. The approximate value A2' can be calculated as:
[0181]
[0182] where [...] means rounding operation, and can be calculated as:
[0183]
[0184] where bdepth(A2) means the bit depth of value A2.
[0185] Perform the same operation on A1 as follows:
[0186]
[0187] Considering the quantized representations of A1 and A2, equation (12) can be rewritten as follows.
[0188]
[0189] in is represented by a length of to avoid division.
[0190] In the simulation, the constant parameters are set as:
[0191] ·n α Equal to 13, this value is a compromise between data accuracy and computational cost.
[0192] · Equal to 6, resulting in a lookup table size of 64, when bdepth(A2)<6 (for example, A2<32), the table size can be further reduced to 32 by expansion.
[0193] ·n table Equal to 15, resulting in a 16-bit data representation of the table elements.
[0194] · Set to 15 to avoid product overflow and maintain 16-bit multiplication.
[0195] Finally, α' is clipped to [-2 -15 ,2 15 -1] to maintain 16-bit multiplication in equation (11). With this limit, when n α When it is equal to 13, the actual a value is limited to [-4,4), which is useful for preventing error amplification.
[0196] Using the calculated parameter α', the parameter β' is calculated as follows:
[0197]
[0198] Here, the division in the above equation can be simply replaced by a shift since the value I is a power of 2.
[0199] JCTVC-I0166: Simplified parameter calculation
[0200] Similar to the discussion above about equation (1), in HM6.0, an intra prediction mode called LM is applied to predict the chroma PU based on a linear model using the reconstruction of the co-located luma PU. The parameters of the linear model consist of the slope (a>>k) and y-intercept (b) derived from neighboring luma and chroma pixels using a least mean square solution. The value of the prediction sample predSamples[x,y] is derived as follows, where x,y=0...nS-1, where nS specifies the block size of the current chroma PU:
[0201] predSamples[x,y]=Clip1 C (((p Y '[x,y]*a)>>k)+b), where, x,y=0...nS-1 (17)
[0202] Where P Y '[x,y] is the reconstructed pixel from the corresponding luminance component. When the coordinates x and y are equal to or greater than 0, P Y ' is the reconstructed pixel from the same luma PU. When x or y is less than 0, P Y ' is the reconstructed neighboring pixels of the same-position brightness PU.
[0203] Some intermediate variables L, C, LL, LC, k2 and k3 in the derivation process are derived as follows:
[0204]
[0205] k2=Log2((2*nS)>>k3) (18-5)
[0206] k3=Max(0,BitDepth C +Log2(nS)-14) (18-6)
[0207] Therefore, the variables a, b, and k can be derived as:
[0208] a1=(LC< <k2)–L*C (19-1)
[0209] a2=(LL< <k2)–L*L (19-2)
[0210] k1=Max(0,Log2(abs(a2))-5)–Max(0,Log2(abs(a1))-14)+2 (19-3)
[0211] a1s=a1>>Max(0,Log2(abs(a1))-14) (19-4)
[0212] a2s=abs(a2>>Max(0,Log2(abs(a2))-5)) (19-5)
[0213] a3=a2s<1?0:Clip3(-2 15 ,2 15 -1,a1s*lmDiv+(1<<(k1-1))>>k1) (19-6)
[0214] a=a3>>Max(0,Log2(abs(a3))-6) (19-7)
[0215] k=13–Max(0,Log2(abs(a))-6) (19-8)
[0216] b=(L–((a*C)>>k1)+(1<<(k2-1)))>>k2 (19-9) where lmDiv is specified in a 63-entry lookup table (i.e., Table 3), which is generated online as follows:
[0217] lmDiv(a2s)=((1<<15)+a2s / 2) / a2s (20)
[0218] Table 3 - lmDiv Specifications
[0219]
[0220]
[0221] In equation (19-6), a1s is a 16-bit signed integer, and lmDiv is a 16-bit unsigned integer. Therefore, a 16-bit multiplier and 16-bit storage are required. It is proposed to reduce the bit depth of the multiplier to the internal bit depth and reduce the size of the lookup table, as described in detail below.
[0222] Reduced multiplier bit depth
[0223] The bit depth of a1s is reduced to the internal bit depth by changing equation (19-4) to the following:
[0224] a1s=a1>>Max(0,Log2(abs(a1))–(BitDepth C –2)) (21)
[0225] The value of lmDiv with internal bit depth is implemented using the following equation (22) and stored in a lookup table:
[0226] lmDiv(a2s)=((1<<(BitDepth C -1))+a2s / 2) / a2s (22)
[0227] Table 4 shows an example of an internal bit depth of 10.
[0228] Table 4 - Specification of lmDiv with internal bit depth equal to 10
[0229] a2s 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 lmDiv 512 256 171 128 102 85 73 64 57 51 47 43 39 37 34 32 a2s 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 lmDiv 30 28 27 26 24 23 22 21 20 20 19 18 18 17 17 16 a2s 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 lmDiv 16 15 15 14 14 13 13 13 12 12 12 12 11 11 11 11 a2s 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 lmDiv 10 10 10 10 10 9 9 9 9 9 9 9 8 8 8
[0230] Equations (19-3) and (19-8) are also modified as follows:
[0231] k1=Max(0,Log2(abs(a2))-5)–Max(0,Log2(abs(a1))–(BitDepth C –2)) (23-1)
[0232] k = BitDepth C –1–Max(0,Log2(abs(a))-6) (23-2)
[0233] Reduced lookup table entries
[0234] It is also proposed to reduce the number of entries from 63 to 32, and the number of bits per entry from 16 to 10, as shown in Table 5. By doing so, a memory saving of almost 70% can be achieved. The corresponding changes of equation (19-6), equation (20) and equation (19-8) are as follows:
[0235] a3=a2s<32?0:Clip3(-2 15 ,2 15 -1,a1s*lmDiv+(1<<(k1-1))>>k1) (24-1)
[0236] lmDiv(a2s)=((1<<(BitDepth C +4))+a2s / 2) / a2s (24-2)
[0237] k = BitDepth C +4–Max(0,Log2(abs(a))-6) (24-3)
[0238] Table 5 - Specification of lmDiv with internal bit depth equal to 10
[0239] a2s 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 lmDiv 512 496 482 468 455 443 431 420 410 400 390 381 372 364 356 349 a2s 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 lmDiv 341 334 328 321 315 309 303 298 293 287 282 278 273 269 264 260
[0240] Multi-model linear model prediction
[0241] In ECM-1.0, a multi-model LM (MMLM) prediction mode is proposed, for which chrominance samples are predicted based on the reconstructed luma samples of the same CU by using the following two linear models:
[0242]
[0243] where pred C (i, j) represents the predicted chroma sample in the CU, and rec L ′(i,j) represents the downsampled reconstructed luma sample of the same CU. Threshold is calculated as the average of adjacent reconstructed luma samples. Fig.10 An example of classifying adjacent samples into two groups based on the value Threshold is shown. For each group, the parameters αi and βi (where i is equal to 1 and 2, respectively) are obtained from two samples from the inside of the group (which is the minimum brightness sample A(X A , Y A ) and the maximum brightness sample B(X B , Y B )) is derived from the linear relationship between the brightness value and the chromaticity value. Here, X A , Y A is the x-coordinate (i.e., luminance value) and y-coordinate (i.e., chrominance value) value of sample point A, and X B , Y B are the x-coordinate and y-coordinate values of sample point B. The linear model parameters α and β are obtained according to the following equations.
[0244]
[0245] β=y A -αx A (26)
[0247] Such a method is also called the min-max method. The division in the above equation can be avoided and replaced by multiplication and shift.
[0248] For coding blocks with a square shape, the above two equations apply directly.For non-square coding blocks, the neighboring samples of the longer boundary are first subsampled to have the same number of samples as for the shorter boundary.
[0249] In addition to the scenario in which the upper template and the left template are used together to calculate the linear model coefficients, the two templates may also be used alternatively in the other two MMLM modes (referred to as MMLM_A and MMLM_L modes).
[0250] In MMLM_A mode, only pixel samples in the upper template are used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to a size of (W+W). In MMLM_L mode, only pixel samples in the left template are used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to a size of (H+H).
[0251] Note that when the above reference line is at a CTU boundary, only one luma line (which is stored in the line buffer for intra prediction) is used to get the downsampled luma samples.
[0252] For chroma intra mode coding, a total of 11 intra modes are allowed for chroma intra mode coding. These modes include five traditional intra modes and six cross-component linear model modes (CCLM, LM_A, LM_L, MMLM, MMLM_A and MMLM_L). The chroma mode signaling and derivation process are shown in Table 6. Chroma mode coding directly depends on the intra prediction mode of the corresponding luminance block. Since the separate block partitioning structures for luminance components and chrominance components are enabled in I slices, one chroma block can correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited.
[0253] Table 6 - Chroma prediction modes derived from luma mode when MMLM_ is enabled
[0254]
[0255] Adaptive Enabling of LM and MMLM for Forecasting
[0256] MMLM and LM modes can also be used together in an adaptive manner. For MMLM, the two linear models are as follows:
[0257]
[0258] where pred C (i, j) represents the predicted chroma sample in the CU, and rec L ′(i,j) represents the downsampled reconstructed luma sample of the same CU. Threshold can be simply determined based on the luma and chroma average values and their minimum and maximum values. Fig.11 An example of classifying adjacent samples into two groups based on the inflection point T indicated by the arrow is shown. The linear model parameters α1 and β1 are obtained from two samples (which are the minimum brightness sample A (X A , Y A )) between the brightness and chromaticity values and the Threshold (X T , Y T ) is derived. The linear model parameters α2 and β2 are derived from two samples (which are the maximum brightness samples B (X B , Y B )) between the brightness and chromaticity values and the Threshold (X T , Y T ) is derived. Here, X A , Y A is the x-coordinate (i.e., luminance value) and y-coordinate (i.e., chrominance value) value of sample point A, and X B , Y B is the x-coordinate and y-coordinate value of sample point B. The linear model parameter α of each group is obtained according to the following equation i and β i , where i is equal to 1 and 2 respectively.
[0259]
[0260] β1=Y A -α1X A
[0261]
[0262] β2=Y T -α2X T (28)
[0263] For coding blocks with a square shape, the above equations apply directly. For non-square coding blocks, the neighboring samples of the longer boundary are first subsampled to have the same number of samples as for the shorter boundary.
[0264] In addition to the scenario in which the upper template and the left template are used together to determine the linear model coefficients, the two templates may alternatively be used in the other two MMLM modes (referred to as MMLM_A and MMLM_L modes, respectively).
[0265] In MMLM_A mode, only pixel samples in the upper template are used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to a size of (W+W). In MMLM_L mode, only pixel samples in the left template are used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to a size of (H+H).
[0266] Note that when the above reference line is at a CTU boundary, only one luma line (which is stored in the line buffer for intra prediction) is used to get the downsampled luma samples.
[0267] For chroma intra mode coding, there is a condition check that is used to select LM mode (CCLM, LM_A and LM_L) or multi-model LM mode (MMLM, MMLM_A and MMLM_L). The condition check is as follows:
[0268]
[0269] Where BlkSizeThres LM Indicates the minimum block size in LM mode, and BlkSizeThres MM Indicates the minimum block size of the MMLM mode. The symbol d represents a predetermined threshold value. In one example, d can be taken as a value of 0. In another example, d can be taken as a value of 8.
[0270] For chroma intra mode coding, a total of 8 intra modes are allowed for chroma intra mode coding. These modes include five traditional intra modes and three cross-component linear model modes. The chroma mode signaling and derivation process are shown in Table 1. It is worth noting that for a given CU, if it is encoded in the linear model mode, whether it is a conventional single model LM mode or an MMLM mode is determined based on the above conditional check. Unlike the case shown in Table 6, there is no separate MMLM mode to be signaled. Chroma mode coding directly depends on the intra prediction mode of the corresponding luminance block. Since separate block partitioning structures for luminance components and chrominance components are enabled in I slices, one chroma block can correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited.
[0271] Slope adjustment for CCLM
[0272] During the development of the ECM, scale (slope) adjustment for the CCLM was proposed as a further improvement, for example, as described in JVET-Y0055 / Z0049.
[0273] As discussed above, CCLM uses a model with 2 parameters to map luma values to chroma values. The scale parameter "a" and the bias parameter "b" define the following mapping:
[0274] chromaVal=a*lumaVal+b (30)
[0275] It is proposed to signal an adjustment of the scale parameter “u” to update the model to the following form:
[0276] chromaVal=a'*lumaVal+b' (31)
[0277] Where a' = a + u, and b' = bu * y r .
[0278] With this choice, the mapping function is centered around the brightness value y r It is recommended to use the average value of the reference brightness samples used in model creation as the y r , in order to provide meaningful modifications to the model. Figures 12A to 12B shows the effect of the rescaling parameter "u", where Fig. 12A shows the model created without the scaling parameter "u", and Fig. 12B The model created with the scaling parameter "u" is shown.
[0279] In one example, the rescaling parameter is provided as an integer between -4 and 4 (inclusive) and signaled in the bitstream. The unit of the rescaling parameter is 1 / 8 of the chroma sample value per luma sample value (for 10-bit content).
[0280] In one example, the CCLM model that is available for reference samples both above and to the left of the block being used ("LM_CHROMA_IDX" and "MMLM_CHROMA_IDX") but not for "one-sided" mode is adjusted. This choice is based on codec efficiency versus complexity tradeoff considerations.
[0281] When applying rescaling for a multi-mode CCLM model, both models may be rescaled and thus up to two scale updates may be signaled for a single chroma block.
[0282] To implement rescaling at the encoder, the encoder may perform a SATD-based search for the best value for the scale update for Cr and a similar SATD-based search for Cb. If either result is a non-zero rescaling parameter, the combined rescaling pair (SATD-based update for Cr, SATD-based update for Cb) is included in the RD check list of the TU.
[0283] Fusion of Chroma Intra Prediction Modes
[0284] During the development of ECM, JVET-Y0092 / Z0051 proposed the fusion of chroma intra mode.
[0285] The intra prediction modes enabled for the chroma components in ECM-4.0 are six cross-component linear model (LM) modes (including CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L and MMLM_T modes), direct mode (DM) and four default chroma intra prediction modes. The four default modes are given by a list {0, 50, 18, 1}, and if the DM mode already belongs to the list, the mode in the list will be replaced with mode 66.
[0286] The decoder-side intra mode derivation (DIMD) method for luma intra prediction is included in ECM-4.0. First, the horizontal and vertical gradients are calculated for each reconstructed luma sample of the L-shaped template of the second adjacent row and column of the current block to construct a gradient histogram (HoG). Then, the two intra prediction modes with the largest and second largest histogram amplitude values are mixed with the planar mode to generate the final prediction value of the current luma block.
[0287] In order to improve the coding efficiency of chroma intra prediction, two methods are proposed, including the chroma intra prediction mode derived at the decoder side (DIMD chroma) and the fusion of non-LM mode and MMLM_LT mode.
[0288] DIMD Chroma Mode
[0289] In the first embodiment, a DIMD chroma mode is proposed. The proposed DIMD chroma mode uses the DIMD derivation method to derive the chroma intra prediction mode of the current block based on the co-located reconstructed luminance samples. Specifically, Fig.13 As shown in , the horizontal gradient and the vertical gradient are calculated for each co-located reconstructed luma sample of the current chroma block to construct the HoG. Then, the intra prediction mode with the maximum histogram magnitude value is used to perform chroma intra prediction of the current chroma block.
[0290] When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode having the second largest histogram magnitude value is used as the DIMD chroma mode.
[0291] A CU level flag is signaled to indicate whether the proposed DIMD chroma mode is applied, as shown in Table 7.
[0292] Table 7. Binarization process of intra_chroma_pred_mode used in the proposed method
[0293] intra_chroma_pred_mode Binary bit string Chroma Intra Mode 0 1100 List[0] 1 1101 List[1] 2 1110 List [2] 3 1111 List[3] 4 10 DIMD Chroma 5 0 DM
[0294] Fusion of Chroma Intra Prediction Modes
[0295] In a second embodiment, a fusion of chroma intra prediction modes is proposed, wherein the DM mode and four default modes can be fused with the MMLM_LT mode as follows:
[0296] pred=(w0*pred0+w1*pred1+(1<<(shift-1)))>>shift
[0297] Where pred0 is the prediction value obtained by applying the non-LM mode, pred1 is the prediction value obtained by applying the MMLM_LT mode, and pred is the final prediction value of the current chroma block. The two weights w0 and w1 are determined by the intra-frame prediction mode of the adjacent chroma block, and shift is set to be equal to 2. Specifically, when both the upper adjacent block and the left adjacent block are encoded in the LM mode, {w0, w1} = {1, 3}; when both the upper adjacent block and the left adjacent block are encoded in the non-LM mode, {w0, w1} = {3, 1}; otherwise, {w0, w1} = {2, 2}.
[0298] For syntax design, if non-LM mode is selected, a flag is signaled to indicate whether fusion is applied. And the proposed fusion is only applied to I slices.
[0299] Combination of DIMD chroma mode and chroma intra prediction mode fusion
[0300] In the third embodiment, the DIMD chroma mode is combined with the fusion of the chroma intra prediction mode. Specifically, the DIMD chroma mode described in the first embodiment is applied, and for I slices, the DM mode, the four default modes, and the DIMD chroma mode can be fused with the MMLM_LT mode using the weights described in the second embodiment, while for non-I slices, only the DIMD chroma mode can be fused with the MMLM_LT mode using equal weights.
[0301] Combination of DIMD chroma mode and chroma intra prediction mode fusion with reduced processing
[0302] In a fourth embodiment, the DIMD chroma mode with reduced processing is combined with a fusion of chroma intra prediction modes. Specifically, the DIMD chroma mode with reduced processing derives an intra mode based on adjacent reconstructed Y, Cb, and Cr samples in the second adjacent row and column, such as Fig.14 The other parts are the same as those of the third embodiment.
[0303] Decoder-side intra mode derivation (DIMD)
[0304] In one embodiment, when DIMD is applied, two intra modes are derived from reconstructed neighboring samples and these two predictions are combined with the planar mode prediction using weights derived from gradients as described in JVET-00449, as FIG. 15A to FIG. 15D The division operation in the weight derivation is performed using the same LUT-based integration scheme used by CCLM. For example, the division operation in the orientation calculation Orient=G y / G x It is calculated by the following LUT-based scheme:
[0305] x=Floor(Log2(Gx))
[0306] normDiff=((Gx<<4)>>x)&15
[0307] x+=(3+(normDiff!=0)?1:0)
[0308] Orient=(Gy*(DivSigTable[normDiff]|8)+(1<<(x-1)))>>x
[0309] in
[0310] DivSigTable
[16] ={0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}.
[0311] The derived intra modes are included in the main list of intra most probable modes (MPMs), so the DIMD process is performed before building the MPM list. The main derived intra modes of a DIMD block are stored with the block and are used for MPM list construction for neighboring blocks.
[0312] Figures 15A to 15D The steps of intra mode derivation at the decoder side are shown, where the intra prediction direction is estimated without intra mode signaling. Fig.15A The first step shown in consists of estimating for each sample point (for Fig.15A The gradient of the light gray sample points shown in Fig. 15B The second step shown in consists in mapping the gradient value to the closest prediction direction within [2, 66]. Fig. 15C The third step shown in comprises selecting 2 prediction directions, wherein, for each prediction direction, all absolute gradients Gx and Gy of neighboring pixels having this direction are added, and the first 2 directions are selected. Fig.15D The fourth step shown in comprises implementing weighted intra prediction using the selected direction.
[0313] Multiple Reference Line (MRL) Intra Prediction
[0314] Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Fig.16 In Figure 1, an example of 4 reference lines is depicted, where the samples of segments A and F are not extracted from the reconstructed neighboring samples, but are filled with the closest samples from segments B and E, respectively. HEVC intra-picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used.
[0315] The index of the selected reference line (mrl_idx) is signaled and used to generate intra prediction values. For reference line indices greater than 0, only additional reference line modes in the MPM list are included, and only the mpm index is signaled if there are no remaining modes. The reference line index is signaled before the intra prediction mode, and planar mode is excluded from the intra prediction mode if a non-zero reference line index is signaled.
[0316] MRL is disabled for the first line of blocks inside a CTU to prevent the use of extended reference samples outside the current CTU line. In addition, PDPC is disabled when additional lines are used. For MRL mode, the derivation of the DC value in the DC intra prediction mode for non-zero reference line index is consistent with the derivation of the DC value for reference line index 0. MRL requires storage of 3 adjacent luma reference lines for a CTU to generate predictions. The Cross Component Linear Model (CCLM) tool also requires 3 adjacent luma reference lines for its downsampling filter. The definition of MRL using the same 3 lines is aligned with CCLM to reduce storage requirements for the decoder.
[0317] Convolutional Cross Component Model (CCCM) for Intra Prediction
[0318] During the development of ECM, the Convolutional Cross-Component Model (CCCM) for Chroma Intra Mode was proposed.
[0319] It is proposed to apply a convolutional cross-component model (CCCM) to predict chroma samples from reconstructed luma samples in a similar spirit as done by the current CCLM mode. As with CCLM, when chroma subsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid.
[0320] Furthermore, similar to CCLM, there is an option to use a single model or a multi-model variant of CCCM. The multi-model variant uses two models, one derived for samples above the average luminance reference value, and another for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.
[0321] Convolutional Filters
[0322] The proposed convolutional 7-tap filter consists of a 5-tap plus sign-shaped spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter consists of the following terms: the center (C) luma sample (which is co-located with the chroma sample to be predicted) and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, as Fig.17 as shown in .
[0323] The nonlinear term P is expressed as a power of 2 of the center luma sample C and scaled to the sample value range of the content:
[0324] P=(C*C+midVal)>>bitDepth
[0325] That is, for 10 bits of content, it is calculated as:
[0326] P=(C*C+512)>>10
[0327] The bias term B represents a scalar offset between input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content).
[0328] The output of the filter is calculated as the filter coefficient c i Convolution with the input value and clipped to the range of valid chroma samples:
[0329] predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B
[0330] Calculation of filter coefficients
[0331] The filter coefficients c are calculated by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region i . Fig.18 A reference region consisting of 6 rows of chroma samples above and to the left of the PU is shown. The reference region extends one PU width to the right below the PU boundary and one PU height. The region is adjusted to include only available samples. The extension of the region shown in blue is needed to support the "side samples" of the plus-shaped spatial filter and is filled in when in an unavailable region.
[0332] MSE minimization is performed by computing the autocorrelation matrix of the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL decomposed and back substitution is used to compute the final filter coefficients. The process roughly follows the calculation of the ALF filter coefficients in ECM, however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. The proposed method uses only integer operations.
[0333] Bitstream signaling
[0334] The use of the mode is signaled using the CABAC coded PU level flag. A new CABAC context is included to support this. CCCM is considered a sub-mode of CCLM when it comes to signaling. That is, the CCCM flag is signaled only when the intra prediction mode is LM_CHROMA_IDX (to enable single-mode CCCM) or MMLM_CHROMA_IDX (to enable multi-mode CCCM).
[0335] Encoder Operation
[0336] The encoder performs two new RD checks in the chroma prediction mode loop, one for checking single-model CCCM mode and one for checking multi-model CCCM mode.
[0337] In the existing CCLM or MMLM design, adjacent reconstructed luma-chroma samples are classified into one or more sample groups based on the value Threshold that only considers the luma DC value. That is, luma-chroma sample pairs are classified by only considering the intensity of the luma samples. However, the luma component usually retains a lot of texture, and the current luma sample can be highly correlated with the adjacent luma samples. Such inter-sample correlation (AC correlation) can be beneficial to the classification of luma-chroma sample pairs and can bring additional coding efficiency.
[0338] like Fig.19A As shown in , CCLM assumes that a given chroma sample is only related to the corresponding luma sample (L0.5, which can be regarded as a fractional luma sample position) and uses simple linear regression (SLR) with ordinary least squares (OLS) estimation to predict the given chroma sample. Fig.19BAs shown in , in some video contents, one chrominance sample may be simultaneously correlated with multiple luminance samples (AC or DC correlated), so a multiple linear regression (MLR) model can further improve the prediction accuracy.
[0339] Although the CCCM mode can enhance intra-frame prediction efficiency, there is room for further improvement of its performance. At the same time, some parts of the existing CCCM mode also need to be simplified for efficient codec hardware implementation or improved for better codec efficiency. In addition, it is necessary to further improve the trade-off between its implementation complexity and its codec efficiency benefits.
[0340] Marginal Linear Model (ELM)
[0341] In order to improve the coding efficiency of the luma component and the chroma component, a classifier that considers luma edge or AC information is introduced, in contrast to the above-mentioned implementation in which only the luma DC value is considered. In addition to the existing MMLM with classification, the present disclosure provides an exemplary classifier. The process of generating a linear prediction model for different groups of samples can be similar to CCLM or MMLM (e.g., via least squares or a simplified minimum-maximum method, etc.), but with different metrics for classification. Different classifiers can be used to classify adjacent luma samples (e.g., adjacent luma-chroma sample pairs) and / or luma samples corresponding to the chroma samples to be predicted. Luma samples corresponding to chroma samples can be obtained by a downsampling operation to match the position of corresponding chroma samples of a 4:2:0 video sequence. For example, luma samples corresponding to chroma samples can be obtained by performing a downsampling operation on more than one (e.g., 4) reconstructed luma samples corresponding to chroma samples (e.g., located around chroma samples) to obtain luma samples corresponding to chroma samples. Alternatively, for example, in the case of a 4:4:4 video sequence, the luma samples may be obtained directly from the reconstructed luma samples. Alternatively, the luma samples may be obtained from corresponding ones of the reconstructed luma samples at corresponding co-located positions of the corresponding chroma samples. For example, the luma sample to be classified may be obtained from one of the four reconstructed luma samples corresponding to the chroma sample at the upper left position of the four reconstructed luma samples, which may be regarded as the co-located position for the chroma samples.
[0342] The first classifier can classify the luma samples according to the edge strength of the luma samples. For example, a direction (e.g., 0 degrees, 45 degrees, or 90 degrees, etc.) can be selected to calculate the edge strength. The direction can be formed by the current sample and the adjacent sample along the direction (e.g., the adjacent sample located to the upper right of the current sample at 45 degrees). The edge strength can be calculated by subtracting the neighboring sample from the current sample. The edge strength can be quantized to one of the M segments by M-1 thresholds, and the first classifier can use M categories to classify the current sample. Alternatively or additionally, N directions can be formed by the current sample and N adjacent sample along N directions. N edge strengths can be calculated by subtracting N adjacent sample from the current sample, respectively. Similarly, if each of the N edge strengths can be quantized to one of the M segments by M-1 thresholds, the first classifier can use MN categories to classify the current sample.
[0343] The second classifier can be used to classify according to the local pattern. For example, the current luma sample Y0 can be compared with its N adjacent luma samples Yi. If the value of Y0 is greater than the value of Yi, the score can be increased by 1, otherwise, the score can be reduced by 1. The score can be quantized to form K categories. The second classifier can classify the current sample into one of the K categories. For example, the adjacent luma samples can be obtained from four neighbors located above, to the left, to the right, and below (i.e., without diagonal neighbors) the current luma sample.
[0344] It is contemplated that multiple first classifiers, second classifiers, or different instances of the first classifiers or second classifiers or other classifiers described herein may be combined. For example, the first classifier may be combined with an existing MMLM threshold-based classifier. For another example, an instance A of a first classifier may be combined with another instance B of the first classifier, wherein instances A and B are in different directions (e.g., vertical and horizontal directions, respectively).
[0345] It will be appreciated by those skilled in the art that, although the existing CCLM design in the VVC standard is used as the basic CCLM method in the description, the proposed cross-component method described in the present disclosure can also be applied to other predictive codecs with similar design spirits. For example, for chroma from luma (CfL) in the AV1 standard, the proposed method can also be applied by dividing the luma-chroma sample pairs into multiple sample groups.
[0346] The skilled person will understand that Y / Cb / Cr can also be represented as Y / U / V in the field of video coding. If the video data is in RGB format, the proposed method can also be applied by, for example, simply mapping the YUV representation to GBR.
[0347] Filter-based Linear Model (FLM)
[0348] A filter-based linear model (FLM) utilizing the MLR model is introduced as follows to account for the possibility that one chrominance sample may be simultaneously correlated with multiple luma samples.
[0349] For chroma samples to be predicted, the reconstructed co-located and adjacent luma samples can be used to predict the chroma samples to capture the inter-sample correlation between co-located luma samples, adjacent luma samples, and chroma samples. The reconstructed luma samples are linearly weighted and combined with an "offset" to generate predicted chroma samples (C: predicted chroma samples, L i : The i-th reconstructed co-located or adjacent brightness sample, α i : filter coefficient, β: offset, N: filter tap), as shown in the following equation (32-1). Note that the linear weighting plus the offset value directly forms the predicted chroma samples (which can be low-pass, high-pass adaptively according to the video content), and then it is added to the residual to form the reconstructed chroma samples.
[0350]
[0351] In some implementations like CCCM, the offset term may also be implemented as the intermediate chrominance value B (512 for 10-bit content) multiplied by another coefficient, as shown in equation (32-2) below.
[0352]
[0353] For a given CU, the top and left reconstructed luma and chroma samples can be used to derive or train the FLM parameters (α i ,,β). As with CCLM, α i and β can be derived via OLS. The top and left training samples are collected and a pseudo-inverse matrix is calculated at both the encoder and decoder sides to derive the parameters, which are then used to predict the chroma samples in a given CU. Let N denote the number of filter taps applied to the luma samples, M denote the total top and left reconstructed luma-chroma sample pairs used for training the parameters, represents the luma sample with the i-th sample pair and the j-th filter tap, C i Denotes the chrominance sample with the i-th sample pair, the following equation shows the pseudo-inverse matrix A + The derivation and parameters. Fig. 20 An example is shown where N is 6 (6-tap), M is 8, the top 2 rows and left 3 columns of luma samples and the top 1 row and left 1 column of chroma samples are used to derive or train parameters.
[0354]
[0355] b = Ax
[0356] x=(A T A) -1 A T b=A + b (33)
[0357] Please note that it is possible to i Instead, there is no offset β to predict the chrominance samples, which can be a subset of the proposed method.
[0358] It should be noted that although the existing CCLM design in the VVC standard is used as the basic CCLM method in the following description, for those skilled in the art of video coding, the proposed cross-component method described in the present disclosure can also be applied to other predictive coding tools with similar design spirit. For example, for the chroma from luma (CfL) in the AV1 standard, the proposed FLM can also be applied by including multiple luma samples into the MLR model.
[0359] The proposed ELM / FLM / GLM (as discussed below) can be directly extended to the CfL design in the AV1 standard, which explicitly sends the model parameters (α, β). For example, (1-tap case) α and / or β are derived at the encoder at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level and signaled to the decoder for CfL mode.
[0360] Filter shape
[0361] To further improve the encoding and decoding performance, additional designs can be used in FLM prediction. Fig. 20 As shown in and discussed above, a 6-tap luma filter is used for FLM prediction. However, although the multi-tap filter can fit the training data well (e.g., top and left adjacent reconstructed luma and chroma samples), in some cases, the training data does not capture all the characteristics of the test data, but it may lead to overfitting and may not predict the test data well (i.e., the chroma block samples to be predicted). Moreover, different filter shapes can adapt well to different video block contents, resulting in more accurate predictions.
[0362] To address this issue, the filter shape and the number of filter taps can be predefined, or signaled or switched in a sequence parameter set (SPS), adaptive parameter set (APS), picture parameter set (PPS), picture header (PH), slice header (SH), region, CTU, CU, sub-block, or sampling level. The filter shape candidate set can be predefined, and the selection of the filter shape candidate set can be signaled or switched in the SPS, APS, PPS, PH, SH, region, CTU, CU, sub-block, or sampling level. Different components (e.g., U and V) can have different filter switching controls. For example, a filter shape candidate set (e.g., indicated by indices 0 to 5) can be predefined, and a filter shape (1, 2) can represent a 2-tap luminance filter, a filter shape (1, 2, 4) can represent a 3-tap luminance filter, and so on, as in Fig. 20 As shown in . The filter shape selection for U and V components can be switched in PH or in CU or CTU level. Note that N taps can represent N taps with or without offset β as described herein. An example is given in Table 8 as follows.
[0363] Table 8 - Example signaling and switching for different filter shapes
[0364]
[0365] The FLM or CCCM filter shape may include nonlinear terms. For example, for a CCCM filter, the following equations may be used to predict the chroma sample values:
[0366] predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B,P=(C*C+midVal)>>bitDepth
[0367] Wherein, the filter corresponds to the weighting coefficients c0, c1, ..., c6 of the center (C) luma sample value (which is co-located with the chroma sample to be predicted), its upper / north (N), lower / south (S), right / east (E) and left / west (W) neighbors, the nonlinear term P and the bias term B. The value used to derive the nonlinear term P can be a combination of the current and neighboring luma samples, but is not limited to C*C. For example, P can be derived as follows:
[0368] P=(Q*R+midVal)>>bitDepth
[0369] Where Q and R represent the values used to derive the nonlinear term P.
[0370] Q and R can be linear combinations of current and neighboring luma samples in the downsampled domain (eg, Q and R are pre-operation luma samples obtained by a weighted average operation) or without any downsampling process.
[0371] For example, each of Q and R can be selected from one of N, S, E, W and C brightness sample values, for example, Q*R=C*N, C*S, C*E, C*W, S*N or N*N, etc.; or Q and R can be equal to the average value of N, S, E and W brightness sample values, that is, Q=R=(N+S+E+W) / 4; or Q is equal to the C brightness sample value, and R is equal to the average value of N, S, E and W brightness sample values, that is, Q=C, and R=(N+S+E+W) / 4.
[0372] Different values (Q and R) used to derive the non-linear terms are considered different filter shapes and may be predefined or signaled / switched at SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. The filter shape candidate set may be predefined or signaled / switched at SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0373] Different chroma types and / or color formats may have different predefined filter shapes and / or taps. For example, a predefined filter shape (1, 2, 4, 5) may be used for 4:2:0 type 0, a predefined filter shape (0, 1, 2, 4, 7) may be used for 4:2:0 type 2, and a predefined filter shape (1, 4) may be used for 4:2:2, and a predefined filter shape (0, 1, 2, 3, 4, 5) may be used for 4:4:4, as shown in FIG. Fig.21 as shown in .
[0374] In another aspect of the present disclosure, unavailable luma and chroma samples used to derive the MLR model can be filled in from available reconstruction samples. Fig.21 For the CU located at the left picture boundary, the left column including sample (0, 3) is unavailable (outside the picture boundary), so sample (0, 3) is repeated and padded from sample (1, 4) to apply the 6-tap filter. Note that the padding process can be applied to both training data (top and left adjacent reconstructed luma and chroma samples) and test data (luma and chroma samples in the CU).
[0375] One or more shapes / number of filter taps can be used for FLM prediction, for example in Fig.25 , Fig.26 and FIG. 27A to FIG. 27B One or more sets of filter taps may be used for FLM prediction, examples being FIG. 28A to FIG. 28G Shown in.
[0376] Implicit filter shape derivation
[0377] The filter shape candidate may be derived implicitly without explicit signaling bits. For example, the filter shape candidate may be a filter shape candidate for FLM or GLM (as discussed below). In another example, the filter shape candidate may be a cross-shaped filter for CCCM, Fig.25 , Fig.26 , Fig.27A , Fig.27B as well as Figures 28A to 28G Any of the filters shown in or other filters mentioned in this disclosure. Since longer filter taps theoretically always fit better in the training data (template area), but may overfit, the "N-fold cross validation" technique known in the field of machine learning can be used to train the filter coefficients. This technique divides the available training data into N sets, and uses some sets for training and other sets for validation.
[0378] The following example involves implicit filter shape derivation for FLM prediction:
[0379] Step 1: Determine M filter shape candidates for predicting the chroma sample values of the current CU;
[0380] Step 2: Divide the available L-shaped template area outside the CU into N regions, denoted as R0, R1, ... R N-1 , that is, the training data is divided into N sets for N-fold training, where the luminance sample values and chrominance sample values of the available template area are known values;
[0381] Step 3: Apply each of the M filter shape candidates to a portion of the available template region, i.e., all N regions R0, R1, ... R N-1 one or more areas in
[0382] Step 4: Derive M sets of filter coefficients corresponding to the M filter shape candidates, denoted as F0, F1, ... F M- 1;
[0383] Step 5: Substitute the derived F0, F1, ... F M-1 Applying the set of filter coefficients to another portion of the available template area to predict chrominance sample values based on corresponding luma sample values, wherein the other portion of the available template area is different from the portion of the available template area mentioned in step 3;
[0384] Step 6: Accumulate the errors (denoted as E0, E1, ..., E2) between the predicted chroma sample values and the known chroma sample values in another part of the available template area by sum of absolute differences (SAD), sum of squared differences (SSD), or sum of absolute transformed differences (SATD) for each of the M filter shapes, respectively. M-1 );
[0385] Step 7: Sort and select the K smallest errors (denoted as E'0, E'1, ... E' K-1 ), which corresponds to K filter shapes and K sets of filter coefficients; and
[0386] Step 8: Select a filter shape candidate from the K filter shape candidates to apply to the current CU for chroma prediction. If K is greater than 1, the decoder can still receive a signal from the encoder indicating the filter to be applied. However, if K is 1, the signaling can be omitted and the filter with the smallest accumulated error is determined as the applied filter.
[0387] Fig.29A and Fig.29B An example of 2-fold training for implicit filter shape derivation is shown. For the current chroma CU prediction (blue area), Fig.29A It is shown that in the template area, the even-numbered row area R0 (yellow) is used for training / deriving 4 sets of filter coefficients, and the odd-numbered row area R1 (red) is used for verification / comparison of the 4 sets of filter coefficients and sorting the costs; Fig.29B Shown in the template region, R0 (yellow) and R1 (red) are interleaved. It should be understood that R0 and R1 can be swapped in these examples.
[0388] In one example, to select one of 4 filter shape candidates as the applied filter, the L-shaped template area is divided into even-numbered and odd-numbered rows or columns. The steps include:
[0389] Four filter shape candidates are predefined for the current CU (e.g., from Fig.25 4 filter shape candidates of );
[0390] The available L-shaped template area (e.g., 6 chroma rows and columns for CCCM, note that in CCCM design, each chroma sample refers to 6 luma samples for downsampling) is divided into 2 regions denoted as R0, R1, where, for example, R0 consists of even rows or columns, and R1 consists of odd rows or columns ( Fig.29AThe following example is shown: in the template area, the even-numbered row area R0 is used to train / derive 4 sets of filter coefficients, and the odd-numbered row area R1 is used to verify / compare and sort the costs of the 4 sets of filter coefficients);
[0391] Apply the 4 filter shape candidates independently to a portion of the available template area (e.g., area R0);
[0392] Four sets of filter coefficients are derived for each of the four filter shapes, denoted as F0, F1, ... F3;
[0393] Applying the derived F0, F1, ... F3 filter coefficient sets to another portion of the available template area (e.g., area R1) to predict corresponding chrominance sample values;
[0394] Accumulating the errors (denoted as E0, E1, ... E3) between the predicted chroma sample values and the known chroma sample values in another part of the available template area (e.g., area R1) by SAD, SSD or SATD for each of the four filters, respectively; and
[0395] A filter shape candidate (denoted as E'0) with the minimum accumulated error among E0, E1, ... E3 of the four filter shape candidates is selected, which corresponds to a filter shape and a set of filter coefficients. In this example, only the filter shape candidate with the minimum accumulated error will be determined as the applied filter, and then the decoder does not need to receive a signal indicating the applied filter.
[0396] In one example, the L-shaped template region can be divided into staggered parts when K = 2. The steps include:
[0397] Determine 4 filter shape candidates for the current CU;
[0398] The available L-shaped template area (e.g., 6 chroma rows or columns for CCCM) is divided into 2 regions denoted as R0, R1, where the luma samples in R0 and R1 are interleaved, for example, as shown in the following table:
[0399] <![CDATA[R0]]> <![CDATA[R1]]> <![CDATA[R1]]> <![CDATA[R0]]>
[0400] or
[0401] <![CDATA[R1]]> <![CDATA[R0]]> <![CDATA[R0]]> <![CDATA[R1]]>
[0402] as well as Fig.29B The following example is shown: in the template area, R0 and R1 are interleaved;
[0403] Apply the 4 filter shape candidates independently to a portion of the available template area (e.g., area R0);
[0404] derive 4 sets of filter coefficients (denoted as F0, F1, ... F3) for the 4 filter shapes respectively;
[0405] Applying the derived F0, F1, ... F3 filter coefficient sets to another portion of the available template area (e.g., area R1) to predict corresponding chrominance sample values;
[0406] Accumulating the errors (denoted as E0, E1, ... E3) between the predicted chroma sample values and the known chroma sample values in another part of the available template area (e.g., area R1) by SAD, SSD or SATD for each of the 4 filter shapes, respectively;
[0407] sorting and selecting the 2 smallest accumulated errors (denoted as E'0, E'1) among E0, E1, ... E3, which correspond to the 2 filter shapes and the 2 sets of filter coefficients; and
[0408] Based on the signal received from the encoder, one of the two filter shape candidates is selected to be applied to the current CU for chroma prediction.
[0409] Note that the implicit filter shape derivation method can also be used to determine whether to introduce nonlinear terms in the CCCM filter coefficients (using / not using nonlinear terms is processed as different filter shapes).
[0410] Although the examples above are shown for CCCM filters, it should be understood that the nonlinear term P may also be included in FLM filters (e.g., Fig. 20 ) and is derived in a similar manner as discussed above.
[0411] In one example, the step of dividing the available L-shaped template area can be omitted. In this example, M sets of filter coefficients can be derived based on sample values from the available template area, and then applied back to the available template area to predict corresponding chrominance sample values for accumulating errors.
[0412] Matrix derivation
[0413] As mentioned above, the MLR model (linear equation) must be derived at both the encoder and the decoder. According to one or more aspects of the present disclosure, several methods are proposed to derive the pseudo-inverse matrix A + , or directly solve the linear equation. Other known methods include Newton's method, Cayley-Hamilton method, and eigendecomposition (e.g. https: / / en.wikipedia.org / wiki / Invertible_matrix ) can also be applied.
[0414] In this disclosure, for simplicity, A + It can be represented as A -1 . The linear equation can be solved as follows:
[0415] 1. Solve A by adjoint matrix (adjA), closed form, analytical solution -1 :
[0416] The following shows an nxn general form, a 2x2 and a 3x3 case. If FLM uses 3x3, then 2 scalers plus an offset need to be solved.
[0417] b=Ax,x=(A T A) -1 A T b=A + b, represented by A -1 b
[0418]
[0419] By removing the (n-1)x(n-1) submatrix with the jth row and ith column
[0420]
[0421] 2. Gauss-Jordan elimination method
[0422] We can use Gauss-Jordan elimination, by augmenting the matrix [AI n ] and a series of basic row operations to solve the linear equation to obtain the reduced row echelon form matrix [I|X]. 2x2 and 3x3 examples are shown below.
[0423]
[0424] 3. Cholesky decomposition
[0425] To solve Ax=b, A can first be decomposed by the Cholesky-Crout algorithm, resulting in an upper triangular matrix and a lower triangular matrix, and a forward substitution followed by a backward substitution can be applied serially to obtain the solution. A 3x3 example is shown below.
[0426]
[0427] In addition to the above examples, some conditions require special handling. For example, if some conditions result in the inability to solve a linear equation, default values can be used to fill the chroma prediction values. For example, when 1<<(bitDepth-1), meanC, meanL, or mean C-meanL (average current chroma or other chroma, from available luminance values, or a subset of the FLM-reconstructed adjacent region) is predefined, the default value can be predefined, or signaled or switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0428] The following example shows the case when matrix A cannot be solved, where the default prediction value can be assigned to the entire current block:
[0429] 1. Solve by closed form (analytical, adjoint matrix), but A is singular (i.e., detA = 0);
[0430] 2. Solve by Cholesky decomposition, but A cannot be Cholesky decomposed, g jj <REG_SQR, where REG_SQR is a small value, which can be predefined, or signaled or switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0431] Application Area
[0432] Fig. 20 Illustrate a typical case of deriving FLM parameters using the top 2 luminance lines and / or the left 3 luminance lines and the top 1 chroma line and / or the left 1 chroma line. However, as mentioned above, due to different block contents and different reconstruction qualities of adjacent samples, using different regions for parameter derivation can bring coding and decoding benefits. The following presents several ways to select the regions applied for parameter derivation:
[0433] 1. Similar to MDLM, FLM derivation can use only the top or left luminance and / or chroma samples to derive parameters. Whether to use FLM, FLM_L, or FLM_T can be predefined, or signaled or switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Assuming the current chroma block has a size of W×H, then W' and H' are obtained as follows:
[0434] – When the FLM mode is applied, W’ = W, H’ = H;
[0435] – When FLM_T mode is applied, W'=W+We, where We represents the extended top luma / chroma sample;
[0436] – When FLM_L mode is applied, H′=H+He; where He represents the extended left luma / chroma samples.
[0437] The number of extended luma / chroma samples (We, He) may be predefined, or signaled or switched in SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0438] For example, (We, He) = (H, W) is predefined as a VVC CCLM, or (W, H) is predefined as an ECM CCLM. Unavailable (We, He) luma / chroma samples can be repeatedly filled from the nearest (horizontal, vertical) luma / chroma samples.
[0439] Fig. 22 The specification of FLM_L and FLM_T is shown (eg, less than 4 taps). When FLM_L or FLM_T is applied, only H' or W' luma / chroma samples, respectively, are used for parameter derivation.
[0440] 2. Similar to MRL, different line indices can be predefined, or signaled or switched in SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to indicate the selected luma-chroma sample pair lines. This can benefit from different reconstruction qualities of different line samples.
[0441] Fig.23 It is shown that similar to MRL, FLM can use different lines for parameter derivation (e.g., lower than 4 taps). For example, FLM can use light blue / yellow luma and / or chroma samples of index 1.
[0442] 3. Expand the CCLM region and employ all top N rows and / or left M rows for parameter derivation. Fig.23 It is shown that all the dark blue and light blue and dark yellow and light yellow areas can be used at once. Training using larger areas (data) can lead to more robust MLR models.
[0443] It should be understood that luma sample values of an outer region of a video block to be decoded may be referred to as “outer luma sample values” and chroma sample values of an outer region may be referred to as “outer chroma sample values” throughout the disclosure.
[0444] grammar
[0445] The corresponding syntax may be defined as follows in Table 9 for FLM prediction. Wherein, FLC represents fixed length code, TU represents truncated unary code, EGk represents exponential Golomb code with order k, where k may be fixed or signaled / switched in SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level, SVLC represents signed EG0, and UVLC represents unsigned EG0.
[0446] Table 9 - Example of FLM syntax
[0447]
[0448]
[0449] Note that the binarization of each syntax element may be varied.
[0450] Gradient Linear Model (GLM)
[0451] A new method for cross-component prediction is proposed based on the existing linear model design to further improve the encoding and decoding accuracy and efficiency. The main aspects of the proposed method are detailed as follows.
[0452] Although the FLM discussed above provides the best flexibility (leading to the best performance), if the number of filter taps goes up, many unknown parameters need to be solved. When the inverse matrix is larger than 3×3, the closed form derivation is not suitable (too many multipliers) and an iterative method like Cholesky's is required, which will burden the decoder processing cycle. In this section, a pre-operation before applying the linear model is proposed, including exploiting the correlation between the luminance AC information and the chrominance intensity using the sample gradients. With the help of the gradients, the number of filter taps can be effectively reduced.
[0453] Please note that the methods / examples in this section can be combined / reused from any of the designs discussed above, including but not limited to classification, filter shapes, matrix derivation (with special processing), application areas, syntax. In addition, the methods / examples listed in this section can also be applied to any of the designs discussed above to have better performance and certain complexity tradeoffs.
[0454] Please note that the reference samples / training templates / reconstructed neighboring regions used in this article generally refer to luma samples used to derive MLR model parameters, which are then applied to internal luma samples in a CU to predict chroma samples in the CU.
[0455] Filter shape
[0456] According to the proposed method, instead of directly using luma sample intensity values as input to the linear model, pre-operations (e.g., pre-linear weighting, sign, scaling / abs, thresholding, ReLU) can be applied to reduce the dimensionality of unknown parameters. In one example, the pre-operations can include calculating sample differences based on luma sample values. As understood by those skilled in the art, sample differences can be characterized as gradients, and therefore the new method is also referred to as a gradient linear model (GLM) in some embodiments.
[0457] Please note that the following detailed description discusses scenarios in which the proposed pre-operation can be reused for the SLR model (also known as the 1-tap case) / combined with the SLR model, and can be reused for the MLR model (also known as the multi-tap case, e.g., 2-tap) / combined with the MLR model.
[0458] For example, instead of applying a 2-tap to the 2 luma samples, the 2 luma samples may be pre-operated on and then a simpler 1-tap may be applied to reduce complexity. Figures 24A to 24D Some examples for 1-tap / 2-tap (with offset) pre-operation are shown, where the 2-tap coefficients are represented as (a, b). Note that Figures 24A to 24D Each circle shown in represents an illustrative chroma position for the YUV 4:2:0 format. As discussed above, in the YUV 4:2:0 format, a luma sample corresponding to a chroma sample may be obtained by performing a downsampling operation on more than one (e.g., 4) reconstructed luma samples corresponding to (e.g., located around) the chroma sample. In other words, a chroma position may correspond to one or more luma samples including a co-located luma sample. Different 1-tap patterns are designed for gradient calculations for different gradient directions and using different "interpolated" luma samples (weighted to different luma positions). For example, in Fig.24A , 24C A typical filter [1, 0, -1; 1, 0, -1] is shown in FIG. 24D, which represents the following operation:
[0459]
[0460] Among them, rec L Represents the reconstructed brightness sample value and Rec L ″(i,j) represents the pre-operation brightness sample value. Also note that, as in Fig.24A , Fig.24C and Fig.24D The 1-tap filter shown in can be understood as an alternative to the downsampling filter with changed filter coefficients as used in CCLM (see equations (6)-(7)).
[0461] The pre-operation can be based on gradient, edge direction (detection), pixel intensity, pixel change, pixel variance, Roberts / Prewitt / compass / Sobel / Laplacian operator, high pass filter (by calculating gradient or other related operators), low pass filter (by performing weighted average operation), etc. The edge direction detector listed in the example can be extended to different edge directions. For example, 1 tap (1, -1) or 2 taps (a, b) are applied along different directions to detect different edge gradients. The filter shape / coefficient can be symmetric about the chrominance position, such as Figures 24A to 24D Example (420 Type-0 case).
[0462] Pre-operation parameters (coefficients, sign, scaling / absolute value, thresholding, ReLU) can be fixed or signaled / switched at SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Note that in the example, if multiple coefficients apply to one sample (e.g., -1, 4), they can be merged (e.g., 3) to reduce operations.
[0463] In one example, the pre-operation may involve calculating sample differences of the luma sample values. Alternatively, the pre-operation may include performing downsampling by a weighted average operation. In some cases, the pre-operation may be applied repeatedly. For example, a template filter may be applied to the template using a low-pass smoothing FIR filter [1, 2, 1] / 4 or [1, 2, 1; 1, 2, 1] / 8 (i.e., downsampling) to remove outliers, and then a 1-tap GLM filter may be applied to calculate sample differences to derive a linear model. It is conceivable that sample differences may also be calculated and then downsampling may be implemented.
[0464] In one example, pre-operation coefficients (either ultimately applied (eg, 3) or intermediately applied (eg, -1, 4) to each luma sample) may be limited to power-of-2 values to save multipliers.
[0465] In one aspect of the present disclosure, the proposed new method can be reused for / combined with the CCLM discussed above, which utilizes a simple linear regression (SLR) model and uses one corresponding luma sample value to predict the chroma sample value. This is also referred to as the 1-tap case. In this case, deriving the linear model further includes deriving a scale parameter α and an offset parameter β by using pre-operated adjacent luma sample values and adjacent chroma sample values. Alternatively, the linear model can be rewritten as:
[0466] C=α·L+β (35) Where L here represents the "pre-operation" luminance sample. The parameter derivation of the 1-tap GLM can reuse the CCLM design, but consider the directional gradient (possibly with a high-pass filter). In one example, the scale parameter α can be derived by using a division lookup table as detailed below to achieve simplification.
[0467] In one example, when the GLM is combined with the SLR model, the scale parameter α and the offset parameter β can be derived by utilizing the minimum-maximum method discussed above. Specifically, the scale parameter α and the offset parameter β can be derived by comparing the adjacent brightness sample values of the pre-operation to determine the minimum brightness sample value Y A and the maximum brightness sample value Y B ; For the minimum brightness sample value Y A and the maximum brightness sample value Y B Determine the corresponding chroma sample value X A and X B ; and based on the minimum brightness sample value Y A , maximum brightness sample value Y B and the corresponding chrominance sample value X A and X B , the scale parameter α and the offset parameter β are derived according to the following equations:
[0468]
[0469] β=Y A -αX A (36)
[0470] In one example, when the GLM is combined with the SLR model, the rescaling discussed above can be reused. In this case, the encoder can determine the rescaling value (e.g., "u") to be signaled in the bitstream and add the rescaling value to the derived scale parameter α. The decoder can determine the rescaling value (e.g., "u") from the bitstream and add the rescaling value to the derived scale parameter α. The added value is ultimately used to predict the intra chroma sample values.
[0471] In one aspect of the present disclosure, the proposed new method can be reused / combined with FLM, which utilizes a multiple linear regression (MLR) model and uses multiple luma sample values to predict chroma sample values. This is also called a multi-tap case, e.g., 2-tap. In this case, the linear model can be rewritten as:
[0472]
[0473] In this case, multiple scale parameters α and offset parameters β may be derived by using adjacent luma sample values and adjacent chroma sample values of the pre-operation. In one example, the offset parameter β is optional. In one example, at least one of the multiple scale parameters α may be derived by using sample differences. In addition, another of the multiple scale parameters α may be derived by using downsampled luma sample values. In one example, at least one of the multiple scale parameters α may be derived by using horizontal or vertical sample differences calculated based on downsampled adjacent luma sample values. In other words, the linear model may combine multiple scale parameters α associated with different pre-operations.
[0474] Implicit filter shape derivation
[0475] In one example, instead of explicitly signaling the selected filter shape index, the directional orientation filter shape used can be derived at the decoder to save bit overhead. For example, at the decoder, several directional gradient filters can be applied to each reconstructed luminance sample of the L-shaped template of the i-th adjacent row and column of the current block. Then, the filtered values (gradients) can be accumulated for each direction of the several directional gradient filters respectively. In the example, the accumulated values are the accumulated values of the absolute values of the corresponding filtered values. After accumulation, the direction of the directional gradient filter for which the accumulated value is the largest can be determined as the derived (luminance) gradient direction. For example, a gradient histogram (HoG) can be constructed to determine the maximum value. The derived direction can be further applied as the direction for predicting the chrominance samples in the current block.
[0476] The following example involves reusing the decoder-side intra mode derivation (DIMD) method for luma intra prediction included in ECM-4.0:
[0477] Step 1: Apply 2 directional gradient filters (3x3 horizontal / vertical Sobel) to each reconstructed luminance sample of the L-shaped template of the second adjacent row and column of the current block;
[0478] Step 2: Accumulate the filtered values (gradients) of each of the directional gradient filters by SAD (Sum of Absolute Difference);
[0479] Step 3: constructing a Histogram of Gradients (HoG) based on the accumulated filtered values; and
[0480] Step 4: The maximum value in HoG is determined as the derived (brightness) gradient direction, based on which the GLM filter can be determined.
[0481] In one example, if the shape candidates are [-1, 0, 1; -1, 0, 1] (horizontal) and [1, 2, 1; -1, -2, -1] (vertical), when the maximum value is associated with the horizontal shape, the shape [-1, 0, 1; -1, 0, 1] is used for GLM-based chrominance prediction.
[0482] The gradient filter used to derive the gradient direction can be the same or different in shape from the GLM filter. For example, the two filters can be horizontal [-1, 0, 1; -1, 0, 1], or the two filters can have different shapes, and the GLM filter can be determined based on the gradient filter.
[0483] Classification
[0484] The proposed GLM can be combined with the MMLM or ELM discussed above. When combined with classification, each group can share or have its own filter shape, where the syntax indicates the shape of each group. For example, as an exemplary classifier, the horizontal gradient grad_hor can be classified into a first group corresponding to a first linear model, and the vertical gradient grad_ver can be classified into a second group corresponding to a second linear model. In one example, the horizontal brightness pattern can be generated only once.
[0485] Further possible classifiers are also provided as follows. Using the classifier, adjacent and internal luma-chroma sample pairs of the current video block can be classified into a plurality of groups based on one or more thresholds. Note that, as discussed above, each adjacent / internal chroma sample and its corresponding luma sample can be referred to as a luma-chroma sample pair. The one or more thresholds are associated with the intensity of the adjacent / internal luma samples. In this case, each of the plurality of groups corresponds to a respective one of the plurality of linear models.
[0486] When combined with the MMLM classifier, the following operations may be performed: adjacent reconstructed luminance-chrominance sample pairs of the current video block are classified into two groups based on Threshold; different linear models for different groups are derived, wherein the derivation process may be GLM simplified, i.e., the number of taps is reduced using the above-mentioned pre-operation; similarly, luminance-chrominance sample pairs within the CU (intra-luminance-chrominance sample pairs, wherein each intra-luminance-chrominance sample pair in the intra-luminance-chrominance sample pairs includes an intra-chrominance sample value to be predicted using the derived linear model) are classified into two groups based on Threshold; different linear models are applied to the reconstructed luminance samples in different groups; and chrominance samples in the CU are predicted based on different classified linear models.
[0487]
[0488] Among them, rec L ′(i,j) can be the down-sampled reconstructed brightness sample; rec C (i,j) may be reconstructed chroma samples (note: only neighbors are available); Threshold may be the average of neighboring reconstructed luma samples. Note that the number of categories (2) can be extended to multiple categories by increasing the number of thresholds (e.g., equally divided based on the min / max values of neighboring reconstructed (downsampled) luma samples, fixed, or signaled / switched at SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level).
[0489] In one example, instead of the MMLM luma DC intensity, the filtered values of the FLM / GLM applied to the adjacent luma samples are used for classification. For example, if a 1-tap (1, -1) GLM is applied, the average AC value (physical meaning) is used. The processing can be: classifying adjacent reconstructed luma-chroma sample pairs into K groups based on one or more filter shapes, one or more filtered values, and K-1 thresholds Ti; deriving different MLR models for different groups, wherein the derivation process can be GLM simplified, that is, using the above-mentioned pre-operation to reduce the number of taps; similarly classifying the luma-chroma sample pairs within the CU (internal luma-chroma sample pairs, wherein each internal luma-chroma sample pair in the internal luma-chroma sample pair includes an internal chroma sample value to be predicted using the derived linear model) into K groups based on one or more filter shapes, one or more filtered values, and K-1 thresholds Ti; applying different linear models to the reconstructed luma samples in different groups; predicting the chroma samples in the CU based on different classified linear models. Among them, Threshold can be predefined (e.g., 0, or can be a table), or signaled / switched in SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, Threshold can be the average AC value (filtered value) of adjacent reconstructed (can be downsampled) luma samples (2 groups), or equally divided based on minimum / maximum AC (K groups).
[0490] It is also proposed to combine GLM with ELM classifiers. Figures 24A to 24D As shown in , a filter shape (e.g., 1 tap) can be selected to calculate the edge strength. The direction is determined as the direction along which the sample difference between the current sample and the N adjacent sample points (e.g., all 6 luma samples) is calculated. For example, Fig.24AThe filter in the upper middle of the filter (shape [1, 0, -1; 1, 0, -1]) indicates the horizontal direction, because the sample differences between the samples in the horizontal direction can be calculated, and the filter below it (shape [1, 2, 1; -1, -2, -1]) indicates the vertical direction, because the sample differences between the samples in the vertical direction can be calculated. The positive coefficients and negative coefficients in each of the filters implement the calculation of the sample differences. Then, the processing may include: calculating an edge strength through the filtered value (e.g., equivalent); quantizing the edge strength into M segments through M-1 thresholds Ti; classifying the current sample using K categories (e.g., K == M); deriving different MLR models for different groups, wherein the derivation process may be GLM simplified, that is, using the above-mentioned pre-operation to reduce the number of taps; classifying the luminance-chrominance sample pairs within the CU into K groups; applying different MLR models to the reconstructed luminance samples in different groups; and predicting the chrominance samples in the CU based on the differently classified MLR models. Note that the filter shape used for classification may be the same or different from the filter shape used for MLR prediction. Both the number of thresholds M-1 and the threshold value Ti may be fixed or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. In addition, other classifiers / combinations of classifiers as discussed in ELM may also be used for FLM and / or GLM.
[0491] If the number of classification samples in a group is less than a certain number (e.g., a predefined 4), the default values mentioned when discussing the matrix derivation of the MLR model can be applied to the group parameters (α i , β). If the corresponding adjacent reconstructed samples are not available for the selected LM mode, then a default value may be applied, for example, when the MMLM_L mode is selected but the left samples are invalid.
[0492] Simplify and unify
[0493] In order to further improve the encoding and decoding efficiency, several methods related to the simplification of GLM are introduced as follows.
[0494] Matrix / parameter derivation in FLM requires floating-point operations (e.g., closed-form division), which is expensive for decoder hardware, thus requiring a fixed-point design. For the 1-tap GLM case, it can be viewed as a modified luminance reconstruction sample generation of CCLM (e.g., horizontal gradient direction, from CCLM [1, 2, 1; 1, 2, 1] / 8 to GLM [-1, 0, 1; -1, 0, 1]), and the original CCLM process can be reused for GLM, including fixed-point operations, MDLM downsampling, division table, application size restrictions, min-max approximation, and scale adjustment. The 1-tap GLM can have its own configuration or share the same design as CCLM for all entries. For example, a simplified min-max method is used to derive parameters (instead of LMS) and combined with scale adjustment after deriving the GLM model. In this case, the center point (luminance value y r ) becomes the average value of the reference luma sample “gradient”. Another example, when GLM is turned on for this CU, CCLM slope adjustment is inferred to be off and no syntax related to slope adjustment needs to be signaled.
[0495] For example, this section uses typical reference points (top row and left column). Fig.23 As shown in , the extended reconstruction region can also use simplifications in the same spirit and can have a syntax indicating a specific region (such as MDLM, MRL).
[0496] Please note that the following aspects can be combined and applied jointly: For example, combining reference sample downsampling and a division table to perform a division process.
[0497] When applying classification (MMLM / ELM), each group can apply the same or different simplification operations. For example, fill the samples of each group to the target number of samples before applying the right shift, and then apply the same derivation process and the same division table.
[0498] Please note that the implicit filter shape derivation method can also be used to determine whether to disable the downsampling process in the CCCM filter coefficients (using / not using the downsampling process is processed as different filter shapes).
[0499] Fixed-point implementation
[0500] The CCLM design can be reused for the 1-tap case, division by n can be implemented by right shift, and division by A2 can be implemented by LUT. α , n tableThe integration parameters for deriving the intermediate parameters of the linear model (Equations (19)-(20)) can be the same as CCLM or have different values to have higher accuracy. The integration parameters can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level and can be conditioned on the sequence bit depth. For example, n table =bitdepth+4.
[0501] MDLM downsampling
[0502] When GLM is combined with MDLM, the total number of samples used for parameter derivation may not be a power of 2 value and needs to be filled to a power of 2 to replace the division with a right shift operation. For example, for an 8x4 chroma CU, MDLM requires W+H=12 samples, of which MDLM_T has only 8 samples available (reconstructed), and then the downsampled 4 samples (0, 2, 4, 6) can be equally filled. The code for implementing such an operation is shown below:
[0503]
[0504] Other filling methods may also be applied, such as repeated / mirror filling relative to the last neighboring sample point (rightmost / lowest).
[0505] The filling method used for GLM can be the same or different from the filling method used for CCLM.
[0506] Note that in the ECM version, 8x4 chroma CU MDLM_T / MDLM_L requires 2T / 2L=16 / 8 samples respectively, in which case the same padding method can be applied to meet the target power-of-2 number of samples.
[0507] Division LUT
[0508] The division LUT proposed for CCLM / LIC (local illumination compensation) in known standard developments like AVC / HEVC / AV1 / VVC / AVS can be used for GLM division. For example, the LUT in JCTVC-I0166 is reused for the case of bit depth = 10 (Table 4). The division LUT may be different from CCLM. For example, CCLM uses a min-max method with a division table as in Equation 5, but GLM uses a 32-entry LMS division LUT as in Table 5.
[0509] When GLM is combined with MMLM, meanL values may not always be positive (e.g., using filtered / gradient values to classify groups), so sgn(meanL) needs to be extracted and abs(meanL) used to look up the division LUT. Note that the division LUT used for MMLM classification and parameter derivation can be different. For example, use a lower precision LUT (such as the LUT in the min-max method) for mean classification and a higher precision LUT (such as in LMS) for parameter derivation.
[0510] Size Limits and Latency Constraints
[0511] Similar to the CCLM design, some size restrictions may be applied to ELM / FLM / GLM. For example, the same constraints may be applied for luma-chroma delays in a dual tree.
[0512] Size limits can be based on CU area / width / height / depth. Thresholds can be predefined or signaled at SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, for chroma CU area, the predefined threshold can be 128.
[0513] In one example, at least one pre-operation is performed in response to determining that a video block satisfies an enablement threshold, wherein the enablement threshold is associated with an area, width, height, or partition depth of the video block. Specifically, the enablement threshold may define a minimum or maximum area, width, height, or partition depth of the video block. As will be appreciated by those skilled in the art, a video block may include a current chroma block and its co-located luminance block. It is also proposed that the above-mentioned enablement thresholds for the current chroma block and its co-located luminance block be applied jointly. For example, in response to determining that an enablement threshold is satisfied for both the current chroma block and its co-located luminance block, at least one pre-operation is performed.
[0514] Line buffer reduction
[0515] Similar to the CCLM design, if the co-located luma region of the current chroma CU contains the first line inside a CTU, the top template sample generation can be limited to 1 line to reduce the CTU line buffer storage. Note that when the above reference line is at the CTU boundary, only one luma line (common line buffer in intra prediction) is used to get the downsampled luma samples.
[0516] For example, in Fig. 22In , if the co-located luma region of the current chroma CU contains the first row inside a CTU, the top template can be restricted to using only 1 row (but not 2) for parameter derivation (other CUs can still use 2 rows). This saves luma sample line buffer storage when processing CTUs line by line at the decoder hardware. Several methods can be used to achieve line buffer reduction. Note that the example of limited "1" row can be extended to N rows with similar operations. Similarly, 2-tap or multi-tap can also apply such operations. When multi-tap is applied, chroma samples may also need to apply operations.
[0517] For example, using Fig.24A The 1-tap filter [1, 0, -1; 1, 0, -1] shown in is explained as an example. This filter can be reduced to [0, 0, 0; 1, 0, -1], that is, only the coefficients of the lower row are used. Alternatively, the limited upper row of luma samples can be filled from the luma samples of the lower row (repeated, mirrored, 0, meanL, meanC, etc.).
[0518] Take N=4 as an example, that is, the video block is located at the top boundary of the current CTU, and the adjacent luminance sample values and corresponding chrominance sample values of the top 4 rows are used to derive the linear model. Please note that the corresponding chrominance sample values can refer to the corresponding adjacent chrominance sample values of the top 4 rows (for example, for YUV 4:4:4 format). Alternatively, the corresponding chrominance sample values can refer to the corresponding adjacent chrominance sample values of the top 2 rows (for example, for YUV 4:2:0 format). In this case, the adjacent luminance sample values and corresponding chrominance sample values of the top 4 rows can be divided into two regions: a first region including valid sample values (for example, the luminance sample values and corresponding chrominance sample values of the nearest row) and a second region including invalid sample values (for example, the luminance sample values and corresponding chrominance sample values of the other three rows). The coefficients of the filter corresponding to the sample positions that do not belong to the first region can then be set to zero, so that only the sample values from the first region are used to calculate the sample difference. For example, as discussed above, in this case, the filter [1, 0, -1; 1, 0, -1] can be reduced to [0, 0, 0; 1, 0, -1]. Alternatively, the nearest sample values in the first region can be padded to the second region so that the padded sample values can be used to calculate the sample differences.
[0519] Fusion of Chroma Intra Prediction Modes
[0520] In one example, since GLM can be considered as a special CCLM mode, the fusion design can be reused or have its own way. Multiple (two or more) weights can be applied to generate the final prediction value. For example,
[0521] pred=(w0*pred0+w1*pred1+(1<<(shift-1)))>>shift
[0522] where pred0 is the prediction based on the non-LM model and pred1 is the prediction based on the GLM, or
[0523] pred0 is the prediction based on one of the CCLMs (including all MDLMs / MMLMs), and pred1 is the prediction based on the GLM, or
[0524] pred0 is the predicted value based on GLM, and pred1 is the predicted value based on GLM.
[0525] Different I / P / B slices may have different designs for weights w0 and w1 depending on whether the neighboring blocks are coded with CCLM / GLM / other coding modes or block size / width / height.
[0526] For example, the design for the weight can be determined by the intra prediction mode of the adjacent chroma block, and the shift is set to be equal to 2. Specifically, when both the upper and left adjacent blocks are encoded using the LM mode, {w0, w1} = {1, 3}; when both the upper and left adjacent blocks are encoded using the non-LM mode, {w0, w1} = {3, 1}; otherwise {w0, w1} = {2, 2}. For non-I slices, both w0 and w1 can be set to be equal to 2.
[0527] For grammar design, if non-LM mode is selected, a flag is signaled to indicate whether fusion is applied.
[0528] 1-tap linear model
[0529] As described above, the 1-tap GLM has a good gain-complexity tradeoff because it can reuse the existing CCLM module without introducing additional derivations. According to one or more aspects of the present disclosure, such a 1-tap design can be further extended or generalized.
[0530] In one aspect of the present disclosure, for a chroma sample to be predicted, a single corresponding luma sample L may be generated by combining the co-located luma sample and the adjacent luma sample. For example, the combination may be a combination of different linear filters, such as a high-pass gradient filter (GLM) and a low-pass smoothing filter (e.g., a [1, 2, 1; 1, 2, 1] / 8 FIR downsampling filter that may be commonly used in CCLM); and / or a combination of a linear filter and a nonlinear filter (e.g., with nth power, e.g., Ln, where n may be positive, negative, or +-fractional (e.g., +1 / 2, square root or +3, cube, which may be rounded and rescaled to the bit depth dynamic range)).
[0531] In one aspect of the present disclosure, the combination may be applied repeatedly. For example, a combination of GLM and [1, 2, 1; 1, 2, 1] / 8 FIR may be applied to the reconstructed luma samples, and then a nonlinear power of 1 / 2 may be applied. For example, a nonlinear filter may be implemented as a LUT (lookup table), e.g., for bit depth = 10, nth power, n = 1 / 2, LUT [i] = (int) (sqrt (i) + 0.5) << 5, i = 0 to 1023, where 5 will be scaled to the dynamic range of bit depth = 10. When a linear filter cannot efficiently handle the luma-chroma relationship, a nonlinear filter may provide an option. Whether to use a nonlinear term may be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0532] In one or more aspects of the present disclosure, GLM may refer to a generalized linear model (which may be used to linearly or nonlinearly generate a single luminance sample, and the generated single luminance sample may be fed into a CCLM linear model to derive parameters of the CCLM linear model), and the linear / nonlinear generation may be referred to as a general pattern. Different gradients or general patterns may be combined to form another pattern. For example, a gradient pattern may be combined with a CCLM downsampled value; a gradient pattern may be combined with a nonlinear L 2 Value combination; a gradient style can be combined with another gradient style. The two gradient styles to be combined can have different directions or the same direction. For example, [1, 1, 1; -1, -1, -1] and [1, 2, 1; -1, -2, -1], which both have vertical directions, can be combined. [1, 1, 1; -1, -1, -1] and [1, 0, -1; 1, 0, -1], which have vertical and horizontal directions, can be combined. Figures 24A to 24D The combination may include additive, subtractive or linear weighting.
[0533] GLM applied on the downsampled domain
[0534] As described above, the pre-operation may be repeatedly applied, and the GLM may be applied on the pre-linearly weighted / pre-operated samples. For example, as with CCLM, a template filter may be applied to the luma samples to remove outliers and generate downsampled luma samples (one downsampled luma sample corresponding to one chroma sample) using a low-pass smoothing FIR filter [1, 2, 1; 1, 2, 1] / 8 (i.e., a CCLM downsample smoothing filter). And thereafter, a 1-tap GLM may be applied to the smoothed downsampled luma samples to derive the MLR model.
[0535] Some gradient filter patterns, such as 3x3 Sobel or Privitt operator, can be applied to downsample the luma samples. The following table shows some of the gradient filter patterns.
[0536]
[0537]
[0538] The gradient filter pattern can be combined with other gradient / generic filter patterns in the downsampled luma domain. In one example, the combined filter pattern can be applied to the downsampled luma samples. For example, the combined filter pattern can be derived by performing addition or subtraction operations on the corresponding coefficients of the gradient filter pattern and a DC / lowpass based filter pattern such as a filter pattern [0, 0, 0; 0, 1, 0; 0, 0, 0] or [1, 2, 1; 2, 4, 1; 1, 2, 1]. In another example, the combined filter pattern can be derived by performing addition or subtraction operations on the coefficients of the gradient filter pattern and a filter pattern such as L 2 The combined filter pattern is derived by performing addition or subtraction operations on the nonlinear values of the gradient filter pattern and the corresponding coefficients of another gradient filter pattern with different or the same direction. In another example, the combined filter pattern is derived by performing linear weighting operations on the coefficients of the gradient filter pattern.
[0539] A GLM applied on the downsampled domain can fit into the CCCM framework, but may sacrifice high frequency accuracy because a low-pass smoothing is applied before applying the GLM.
[0540] GLM used as input to CCCM
[0541] As shown above, just like CCLM, CCCM applies luma downsampling before convolution. When chroma subsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid. Since the 1-tap GLM can also be viewed as changing the CCLM downsampling filter coefficients (e.g., from [1, 2, 1; 1, 2, 1] / 8 to [1, 2, 1, -1, -2, -1], i.e., from low pass to high pass), the GLM can be used as the input of the CCCM. Specifically, the gradient filter of the GLM replaces the luma downsampling filter ([1, 2, 1; 1, 2, 1] / 8) with gradient-based coefficients (e.g., [1, 2, 1, -1, -2, -1)). In this case, the CCCM operation becomes a "linear / nonlinear combination of gradients", as shown by the following equation:
[0542] predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B
[0543] Where C, N, S, E, W, P are the gradients of the current or neighboring samples (compared to the original downsampled values of CCCM). The related GLM methods described in this disclosure can be applied in the same way before entering the CCCM convolution, for example, classification, separate Cb / Cr control, syntax, pattern combination, PU size restriction, etc.
[0544] Gradient-based coefficient replacement can be applied to specific CCCM taps. Furthermore, not only high-pass coefficient replacement can be used, but also low-pass / band-pass / all-pass coefficient replacement. Replacement can be combined with the FLM / CCCM shape switch discussed above (resulting in a different number of taps). For example, Figures 24A to 24D The gradient style in can be used for replacement. In one example, the operation for applying GLM as an input to CCCM includes: predefining one or more coefficient candidates for CCCM / FLM downsampling; determining the CCCM / FLM filter shape and the number of filter taps for the CU; applying different CCLM downsampling coefficients to different filter taps, where the coefficients can be a high-pass filter (GLM) or a low-pass / band-pass / all-pass filter; generating downsampled luma samples for CCCM input samples (using the applied coefficients); and feeding the generated downsampled luma samples into the CCCM process.
[0545] Some examples for changing the CCLM downsampling filter coefficients are shown below:
[0546] Example 1:
[0547] The candidate filters are [1, 2, 1; 1, 2, 1] / 8 and [1, 0, -1; 1, 0, -1]; predChromaVal = c0C+c1N+c2S+c3E+c4W+c5P+c6B, using the typical CCCM cross shape, 7 taps; N, S, W, E use the filter [1, 2, 1; 1, 2, 1] / 8, keeping the original CCLM downsampling filter; and C, P use the filter [1, 0, -1; 1, 0, -1], i.e., a horizontal gradient filter, and then P physically means gradient^2.
[0548] Example 2:
[0549] The candidate filters are [1, 2, 1; 1, 2, 1] / 8, [1, 0, -1; 1, 0, -1], [1, 2, 1; -1, -2, -1], [2, 1, -1; 1, 1, -1, -2] and [-1, 1, 2; -2, -1, 1]; predChromaVal = c0C0+c1C1+c2C2+c3C3+c4C4+c5P+c6B; C0 uses the filter [1, 2, 1; 1, 2, 1] / 8, keeping the original CCLM downsampling filter; C1 uses the filter [1, 0, -1; 1, 0, -1]; C2 uses filter [1, 2, 1; -1, -2, -1]; C3 uses filter [2, 1, -1; 1, -1, -2]; C4 uses filter [-1, 1, 2; -2, -1, 1]; C5 uses filter [1, 2, 1; 1, 2, 1] / 8, keeping the original CCLM downsampling filter; C0 to C5, and P have the same downsampled brightness position (in the typical CCCM cross shape == C); and C1 to C4 are generated by Sobel-based gradient filters in different directions (such as FIG. 24A to FIG. 24D ).
[0550] Example 3:
[0551] The candidate filters are [1, 2, 1; 1, 2, 1] / 8, [1, 0, -1; 1, 0, -1], [1, 2, 1; -1, -2, -1], [0, 1, 1; 0, 1, 1] and [1, 1, 0; 1, 1, 0]; redChromaVal = c0C0+c1C1+c2C2+c3C3+c4C4+c5P+c6B; C0 uses the filter [1, 2, 1; 1, 2, 1] / 8, keeping the original CCLM downsampling filter; C1 uses the filter [1, 0, -1; 1, 0, -1]; C2 uses filter [1, 2, 1; -1, -2, -1]; C3 uses filter [0, 1, 1; 0, 1, 1]; C4 uses filter [1, 1, 0; 1, 1, 0]; C5 uses filter [1, 2, 1; 1, 2, 1] / 8, keeping the original CCLM downsampling filter; C0 to C5, and P have the same downsampled brightness position (in the typical CCCM cross shape == C); C1 to C2 are generated by gradient filters based on different directions of Sobel (such as Figures 24A to 24D ); and C3 to C4 are generated by a low-pass filter.
[0552] Which CCCM / FLM taps to which coefficient replacement is applied may be predefined (as in the above example), or signaled / switched in the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0553] For each CCCM / FLM tap, coefficient candidates for CCCM / FLM downsampling may be predefined or signaled / switched in SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0554] Example 4:
[0555] The candidate filters are [1, 2, 1; 1, 2, 1] / 8, [1, 0, -1; 1, 0, -1], [1, 2, 1; -1, -2, -1], [2, 1, -1; 1, -1, -2] and [-1, 1, 2; -2, -1, 1]; predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B, using the typical CCCM cross shape, 7 taps; C uses the switched downsampling filter among the 5 candidate filters; N, S, W, E use the filter [1, 2, 1; 1, 2, 1] / 8, keeping the original CCLM downsampling filter; and P uses the switched downsampling filter among the 5 candidate filters.
[0556] Example 5:
[0557] The candidate filters are: [1, 2, 1; 1, 2, 1] / 8, [1, 0, -1; 1, 0, -1], [1, 2, 1; -1, -2, -1], [0, 1, 1; 0, 1, 1] and [1, 1, 0; 1, 1, 0]; predChromaVal = c0C + c1W + c2E + c3P + c4B, i.e., horizontal minus shape, 5 taps; C uses the switched downsampling filter among the 5 candidate filters; W and E use the switched downsampling filters among the 3 candidate filters: [1, 2, 1; 1, 2, 1] / 8, [0, 1, 1; 0, 1, 1] and [1, 1, 0; 1, 1, 0]; and P uses the filter [1, 2, 1; 1, 2, 1] / 8, maintaining the original CCLM downsampling filter.
[0558] grammar
[0559] In one or more aspects of the present disclosure, one or more syntaxes may be introduced to indicate information about a GLM. An example of a GLM syntax is shown in Table 10 below.
[0560] Table 10
[0561]
[0562]
[0563] FLC: Fixed Length Code
[0564] TU: Truncated Unary Code
[0565] EGk: Exponential Golomb code with order k, where k can be fixed or signaled / switched in SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0566] SVLC: Signed EG0
[0567] UVLC: Unsigned EG0
[0568] Note that the binarization of each syntax element can be changed.
[0569] In one aspect of the present disclosure, the GLM on / off control for the Cb / Cr components can be joint or separate. For example, at the CU level, 1 flag can be used to indicate whether the GLM is active for the CU. If it is active, 1 flag can be used to indicate whether both Cb / Cr are active. If it is not active, the 1 flag indicating Cb or Cr is active. When Cb and / or Cr are active, the filter index / gradient (generic) pattern can be signaled separately. All flags can have their own context models or be bypass-coded.
[0570] In another aspect of the present disclosure, whether to signal the GLM on / off flag can depend on the luminance / chrominance coding mode and / or CU size. For example, in the ECM5 chrominance intra-frame mode syntax, when applying MMLM or MMLM_L or MMLM_T, the GLM can be inferred as off; when the CU area < A, where A can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level; if combined with CCCM, when CCCM is on, the GLM can be inferred as off.
[0571] Note that when the GLM is combined with MMLM, different models can share the same gradient / generic pattern or have their own gradient / generic pattern.
[0572]
[0573] When the GLM is combined with CCCM / FLM, if the current CU is implemented as CCCM / FLM, then the CU-level GLM enable flag can be inferred as off.
[0574] hasGlmFlag &=!pu.cccmFlag;
[0575] CCCM without downsampling
[0576] CCCM needs to process the downsampled luminance reference values before calculating the model parameters and applying the CCCM model, which burdens the decoder processing cycle. In this section, a CCCM that does not utilize the downsampling process is proposed, including different options that utilize non-downsampled luminance references and / or non-downsampled luminance. One or more filter shapes can be used for the purposes described below.
[0577] In one example, a convolutional 7-tap filter may include a 5-tap plus-shaped spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter includes the center (C) non-subsampled luma sample (which is co-located with the chroma sample to be predicted) and its non-subsampled upper or north (N), lower or south (S), left or west (W), and right or east (E) neighbors, such as Fig.17 as shown in .
[0578] The nonlinear term P is expressed as a power of 2 of the center luma sample C and scaled to the sample value range of the content:
[0579] P=(C*C+midVal)>>bitDepth
[0580] That is, for 10 bits of content, it is calculated as:
[0581] P=(C*C+512)>>10
[0582] The bias term B represents a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content).
[0583] The output of the filter is calculated as the convolution between the filter coefficients ci and the input value and is clipped to the range of valid chroma samples:
[0584] predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B
[0585] In another example, a convolutional 7-tap filter may include a 6-tap rectangular shape spatial component and a bias term. The input to the spatial 6-tap component of the filter includes the center (b) non-downsampled luma sample (which is co-located with the chroma sample to be predicted) and its downsampled lower left or southwest (d), lower right or southeast (f), lower or south (e), left or west (a), and right or east (c) neighbors, as shown in FIG. Fig.25 The shape is shown in 1.
[0586] The bias term B represents a scalar offset between input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content).
[0587] The output of the filter is calculated as the convolution between the filter coefficients ci and the input value and is clipped to the range of valid chroma samples:
[0588] predChromaVal=c0a+c1b+c2c+c3d+c4e+c5f+c6B
[0589] In yet another example, the convolutional 8-tap filter may consist of a 6-tap rectangular shape spatial component, a nonlinear term, and a bias term. The input to the spatial 6-tap component of the filter consists of the following terms: the center (b) non-subsampled luma sample (which is co-located with the chroma sample to be predicted) and its non-subsampled lower left or south-west (d), lower right or southeast (f), lower or south (e), left or west (a), and right or east (c) neighbors, as Fig.25 The shape is shown in 1.
[0590] The nonlinear term P is expressed as a power of 2 of the center luma sample (b) and scaled to the sample value range of the content:
[0591] P=(b*b+midVal)>>bitDepth
[0592] That is, for 10 bits of content, it is calculated as:
[0593] P=(b*b+512)>>10
[0594] The bias term B represents a scalar offset between input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content).
[0595] The output of the filter is calculated as the convolution between the filter coefficients ci and the input value and is clipped to the range of valid chroma samples:
[0596] predChromaVal=c0a+c1b+c2c+c3d+c4e+c5f+c6P+c7B
[0597] It should be appreciated that the examples shown above are merely sample examples, and other implementations may be possible without departing from the present disclosure, such as having more or fewer taps and having Fig.25 any of the shapes shown in (where the chroma samples to be predicted are represented as circles).
[0598] In yet another example, a convolutional 9-tap filter may consist of a 6-tap rectangular shape spatial component, two nonlinear terms, and a bias term. The input to the spatial 6-tap component of the filter consists of the following terms: the center (b) non-subsampled luma sample (which is co-located with the chroma sample to be predicted and is substantially centered in the filter shape) and its non-subsampled lower left / southwest (d), lower right / southeast (f), lower / south (e), left / west (a), and right / east (c) neighbors, as Fig.25 The shape is shown in 1.
[0599] The non-linear terms P and Q are two non-linear luma sample values expressed as a power of the luma sample value of the center (b) luma sample and the luma sample value of the bottom / south (e) luma sample, respectively, and then scaled to the sample value range of the content:
[0600] P = (b*b+midVal)>>bitDepth;
[0601] Q=(e*e+midVal)>>bitDepth.
[0602] That is, for 10 bits of content, it is calculated as:
[0603] P = (b*b+512)>>10;
[0604] Q=(e*e+512)>>10.
[0605] The bias term B is a bias value representing a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content).
[0606] The output of the filter is calculated as the convolution between the filter coefficients ci and the input value and is clipped to the range of valid chroma samples:
[0607] predChromaVal=c0a+c1b+c2c+c3d+c4e+c5f+c6P+c7Q+c8B
[0608] It should be understood that the nonlinear terms P and Q can be expressed as powers of any luma sample value of the non-downsampled luma sample of the filter. The two nonlinear terms P and Q are exemplary only, and the corresponding chroma sample values can be calculated based on one or more nonlinear values.
[0609] Please note that the methods / examples in this section can be combined / reused with the methods mentioned above, including but not limited to methods related to classification, filter shape, matrix derivation (with special processing), application area and syntax. In addition, the methods / examples listed in this section can also be applied together with the above methods / examples (more taps) to have better performance and some complexity tradeoffs.
[0610] In the present disclosure, reference samples / training templates / reconstructed neighboring regions generally refer to luma samples used to derive MLR model parameters, which are then applied to internal luma samples in a CU to predict chroma samples in the CU.
[0611] According to one or more embodiments of the present disclosure, the reference sample / training template / reconstructed neighboring region may be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, Fig. 9A An example of an L-shaped reconstruction region, the left / top reconstruction region, is shown to derive parameters.
[0612] Filter shape
[0613] One or more shapes / number of filter taps may be used for CCCM prediction, such as Fig.25 , Fig.26 and FIG. 27A to FIG. 27B One or more sets of filter taps may be used for FLM prediction, for example in FIG. 28A to FIG. 28G The selected luma reference value is non-downsampled. One or more predefined shapes / number of filter taps may be used for CCCM prediction based on previously decoded information at TB / CB / slice / picture / sequence level.
[0614] Although the multi-tap filter can fit the training data (i.e., top / left adjacent reconstructed luminance / chrominance samples) well, in some cases, the training data does not capture all the characteristics of the test data, and it may lead to overfitting and may not predict the test data (i.e., the chrominance block samples to be predicted) well. Moreover, different filter shapes can adapt well to different video block contents, resulting in more accurate predictions. To address this issue, the filter shape / number of filter taps can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. The filter shape candidate set can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different components (U / V) can have different filter switch controls. For example, as shown in the following table, the predefined filter shape candidate set (index = 0 to 5) and the filter shape (1, 2) represent a 2-tap luminance filter, while the filter shape (1, 2, 4) represents Fig. 20 3-tap luma filter shown in etc. The filter shape selection for U / V components can be switched in PH or in CU / CTU level. Note that, as described above, N-tap can represent N-tap with or without offset β.
[0615]
[0616] Different chroma types / color formats may have different predefined filter shapes / taps. For example, using predefined filter shapes for 420 type 0: (1, 2, 4, 5), 420 type 2: (0, 1, 2, 4, 7), 422: (1, 4), 444: (0, 1, 2, 3, 4, 5), such as Fig.21 as shown in .
[0617] The unavailable luminance / chrominance samples used to derive the MLR model can be filled in from the available reconstructed samples. Fig.21 , for the CU located at the left picture boundary, the left column including (0, 3) is unavailable (outside the picture boundary), so (0, 3) is repeated padding from (1, 4) to apply the 6-tap filter. Note that the padding process is applied in both the training data (top / left adjacent reconstructed luma / chroma samples) and the test data (luma / chroma samples in the CU).
[0618] According to one or more embodiments of the present disclosure, unavailable luma / chroma samples for deriving the MLR model may be skipped and not used. Then, no filling process is required for the unavailable luma / chroma samples.
[0619] CCLM / MMLM with LDL decomposition
[0620] CCCM needs to process LDL decomposition to calculate the model parameters of CCCM model, avoiding the use of square root operations and only requiring integer operations. In this section, CCLM / MMLM with LDL decomposition is proposed. As mentioned above, LDL decomposition can also be used in ELM / FLM / GLM.
[0621] Please note that the methods / examples in this section can be combined / reused with the methods mentioned above, including but not limited to methods related to classification, filter shape, matrix derivation (with special processing), application area and syntax. In addition, the methods / examples listed in this section can also be applied together with the methods / examples above to have better performance and certain complexity tradeoffs.
[0622] In the present disclosure, reference samples / training templates / reconstructed neighboring regions generally refer to luma samples used to derive MLR model parameters, which are then applied to internal luma samples in a CU to predict chroma samples in the CU.
[0623] CCLM / MMLM with extended range
[0624] One or more reference samples can be used for CCLM / MMLM prediction, i.e. Fig.18As shown in , the reference region can be the same as the reference region in CCCM. Based on the previous decoding information at TB / CB / slice / picture / sequence level, different reference regions can be used for CCLM / MMLM prediction.
[0625] Although training data with multiple reference regions can fit the calculation of model parameters well, in some cases, the training data does not capture all the characteristics of the test data, but it may lead to overfitting and may not predict the test data (i.e., the chroma block samples to be predicted) well. In addition, different reference regions can adapt well to different video block contents, resulting in more accurate predictions. To address this issue, the reference shape / number of reference regions can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. The reference region candidate set can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different components (U / V) can have different reference region switch controls. For example, the predefined reference region candidate set (index = 0 to 4) is shown in the following table. The reference region selection for the U / V component can be switched in the PH or in the CU / CTU level. Different colorimetric types / color formats may have different predefined reference areas.
[0626]
[0627] Unavailable luma / chroma samples used to derive the MLR model can be filled from available reconstructed samples, and the filling process is applied in both training data (top / left neighboring reconstructed luma / chroma samples) and test data (luma / chroma samples in the CU).
[0628] According to one or more embodiments of the present disclosure, unavailable luma / chroma samples for deriving the MLR model may be skipped and not used. Then, no filling process is required for the unavailable luma / chroma samples.
[0629] FLM / GLM / ELM / CCCM with minimum sample point restriction
[0630] FLM needs to process downsampled luma reference values and calculate model parameters, which burdens the decoder processing cycle, especially for small blocks. In this section, FLM with minimum sample restriction is proposed, for example, FLM is only used for samples greater than a predefined number (such as 64, 128). One or more different restrictions can be used for this purpose, for example, FLM is only used for samples greater than a predefined number (such as 256) in a single model, and FLM is only used for samples greater than a predefined number (such as 128) in a multi-model.
[0631] According to one or more embodiments of the present disclosure, the predefined minimum number of samples for a single model may be greater than or equal to the predefined minimum number of samples for a multi-model. For example, FLM / GLM / ELM / CCCM is only used in a single model for samples greater than or equal to a predefined number (such as 128), and FLM / GLM / ELM / CCCM is only used in a multi-model for samples greater than or equal to a predefined number (such as 256).
[0632] According to one or more embodiments of the present disclosure, the predefined minimum number of samples for FLM / GLM / ELM may be greater than or equal to the predefined minimum number of samples for CCCM. For example, CCCM is used only in a single model for samples greater than or equal to a predefined number (such as 0), and CCCM is used only in a multi-model for samples greater than or equal to a predefined number (such as 128). FLM is used only in a single model for samples greater than or equal to a predefined number (such as 128), and FLM is used only in a multi-model for samples greater than or equal to a predefined number (such as 256).
[0633] Please note that the methods / examples in this section can be combined / reused with the methods mentioned above, including but not limited to methods related to classification, filter shape, matrix derivation (with special processing), application area and syntax. In addition, the methods / examples listed in this section can also be applied together with the above methods / examples (more taps) to have better performance and some complexity tradeoffs.
[0634] Multi-mode combination of FLM / GLM / ELM / CCCM / CCLM
[0635] According to one or more embodiments of the present disclosure, two models in the multiple modes of FLM / GLM / ELM / CCCM / CCLM can be further combined to bring additional coding efficiency. For example, the parameters of CCCM (ci) and GLM (a, b) are first derived separately, and then the weight (w) between CCCM and GLM is derived by linear regression. i ), and finally use the weighted CCCM and GLM to predict the chrominance samples from the reconstructed luminance samples.
[0636] GLMpredChromaVal=a*lumaVal+b
[0637] CCCMpredChromaVal=c0*C+c1*N+c2*S+c3*E+c4*W+c5*P+c6*B
[0638] FinalpredChromaVal=w0*GLMPredChromaVal+w1*CCCMPredChromaVal
[0639] According to one or more embodiments of the present disclosure, there is a flag signaled / switched in the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to indicate whether the combined mode is used.
[0640] According to one or more embodiments of the present disclosure, instead of explicitly signaling the selected mode flag, the mode flag may be derived at the decoder to save bit overhead.
[0641] (1) Determine M combination candidates for the current CU
[0642] (2) Divide the available L-shaped template area into N regions, denoted as R0, R1, ... R N-1
[0643] (Divide the training data into N sets, N-fold training)
[0644] (3) Independently apply the M combined candidates to a portion of the available template area (which can be R0, R1, ... R N-1 Single or multiple regions in
[0645] (4) Derive M sets of filter coefficients (based on the M filter shapes), denoted as F0, F1, ... F M-1
[0646] (5) The derived F0, F1, ... F M-1 The set of filter coefficients is applied to other parts of the available template area, other than the part of the available template in (3)
[0647] (6) The error is accumulated by SAD, SSD or SATD and is expressed as E0, E1, ... E M-1
[0648] (7) Sort and select the K smallest errors, denoted as E'0, E'1, ... E' K-1 , which corresponds to K combinations of filters / K sets of filter coefficients
[0649] (8) Signal and select 1 of the K combined filters to apply to the current CU for chroma prediction; if K is 1, no signaling is required (the filter shape with the smallest error is the applied combined filter)
[0650] For example,
[0651] (1) Predefine 3 filter candidates for the current CU, such as CCCM, GLM, and combined CCCM and GLM
[0652] (2) Divide the available L-shaped template area (CCCM 6 chroma rows / columns, note that in the CCCM design, each chroma sample involves 6 luma samples for downsampling) into 2 regions, denoted as R0 and R1
[0653] For example, even rows / columns: R0, odd rows / columns: R1
[0654] Fig.29A The following example is shown: in the template area, the even-numbered row area R0 is used to train / derive 3 sets of filter coefficients, and the odd-numbered row area R1 is used to verify / compare and sort the costs of the 3 sets of filter coefficients.
[0655] (3) Independently apply the 3 filter candidates to a portion of the available template area (single R0)
[0656] (4) Three sets of filter coefficients are derived (based on the four filter shapes), denoted as F0, F1, and F2
[0657] (5) Applying the derived set of F0, F1, F2 filter coefficients to other parts (R1) of the available template area other than the part of the available template in (3).
[0658] (6) The error is accumulated by SAD, SSD or SATD and is expressed as E0, E1, E2
[0659] (7) Sort and select 1 minimum error, denoted as E'0, which corresponds to 1 filter shape / 1 filter coefficient set.
[0660] (8) If K is 1, no signaling is required (the filter with the smallest error is the filter shape applied)
[0661] Please note that the methods / examples in this section can be combined / reused from the methods mentioned in all sections, including but not limited to classification, filter shape, matrix derivation (with special processing), application area, syntax. In addition, the methods / examples listed in this section can also be applied to all sections to have better performance and certain complexity tradeoffs.
[0662] CCCM with non-subsampled and subsampled luminance reference values
[0663] Only downsampled luma reference values may be used in CCCM to calculate model parameters and apply the CCCM model. In this section, non-downsampled luma reference values are also used in CCCM to calculate model parameters and apply the CCCM model, including using non-downsampled luma reference values and downsampled luma reference values in different or the same positions. As described above, one or more filter shapes may be used for this purpose.
[0664] In one example, a convolutional 8-tap filter may include a 6-tap rectangular shape spatial component, a downsampled luma sample, and a bias term. The input to the spatial 6-tap component of the filter includes the center (b) non-downsampled luma sample (which is co-located with the chroma sample to be predicted) and its non-downsampled lower left or southwest (d), lower right or southeast (f), lower or south (e), left or west (a), and right or east (c) neighbors, as shown in Fig.25 , and the center (C) downsampled luma sample (which is co-located with the chroma samples to be predicted).
[0665] The bias term B represents a scalar offset between input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content).
[0666] The output of the filter is calculated as the convolution between the filter coefficients ci and the input value and is clipped to the range of valid chroma samples:
[0667] predChromaVal=c0a+c1b+c2c+c3d+c4e+c5f+c6C+c7B
[0668] In another example, a convolutional 9-tap filter may consist of a 6-tap rectangular shape spatial component, two nonlinear terms, and a bias term. The input to the spatial 6-tap component of the filter consists of the following terms: the center (b) non-subsampled luma sample (which is co-located with the chroma sample to be predicted and is substantially centered in the filter shape) and its non-subsampled lower left / southwest (d), lower right / southeast (f), lower / south (e), left / west (a), and right / east (c) neighbors, as Fig.25 As shown in shape 1 in , and the center (C) downsampled luma sample (which is co-located with the chroma sample to be predicted (e.g., the center (C) downsampled luma sample value is determined by a weighted average operation)).
[0669] The non-linear terms P and Q are two non-linear luma sample values expressed as the power of the luma sample value of the center (b) non-subsampled luma sample and the power of the luma sample value of the center (c) subsampled luma sample, respectively, and then scaled to the sample value range of the content:
[0670] P = (b*b+midVal)>>bitDepth;
[0671] Q=(C*C+midVal)>>bitDepth.
[0672] That is, for 10 bits of content, it is calculated as:
[0673] P = (b*b+512)>>10;
[0674] Q=(C*C+512)>>10.
[0675] The bias term B is a bias value representing a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content).
[0676] The output of the filter is calculated as the convolution between the filter coefficients ci and the input value and is clipped to the range of valid chroma samples:
[0677] predChromaVal=c0a+c1b+c2c+c3d+c4e+c5f+c6P+c7Q+c8B
[0678] It should be understood that the non-linear terms P and Q can be expressed as powers of any luma sample value for predicting downsampled / non-downsampled luma samples of corresponding chroma sample values. The two non-linear terms P and Q are only exemplary, and corresponding chroma sample values can be calculated based on one or more non-linear values.
[0679] As mentioned above, the non-linear term Q can also be expressed as a power of the non-subsampled luma samples of the filter. In such an example, the convolutional 9-tap filter can consist of a 6-tap rectangular shape spatial component, two non-linear terms and a bias term. The input to the spatial 6-tap component of the filter consists of the following terms: the center (b) non-subsampled luma sample (which is co-located with the chroma sample to be predicted and is essentially located in the center of the filter shape) and its non-subsampled left bottom / southwest (d), right bottom / southeast (f), bottom / south (e), left / west (a) and right / east (c) neighbors, as shown Fig.25 The shape is shown in 1.
[0680] The non-linear terms P and Q are two non-linear luma sample values expressed as a power of the luma sample value of the center (b) luma sample and the luma sample value of the bottom / south (e) luma sample, respectively, and then scaled to the sample value range of the content:
[0681] P = (b*b+midVal)>>bitDepth;
[0682] Q=(e*e+midVal)>>bitDepth.
[0683] That is, for 10 bits of content, it is calculated as:
[0684] P = (b*b+512)>>10;
[0685] Q=(e*e+512)>>10.
[0686] The bias term B is a bias value representing a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content).
[0687] The output of the filter is calculated as the convolution between the filter coefficients ci and the input value and is clipped to the range of valid chroma samples:
[0688] predChromaVal=c0a+c1b+c2c+c3d+c4e+c5f+c6P+c7Q+c8B
[0689] Please note that the methods / examples in this section can be combined / reused from the methods mentioned above, including but not limited to classification, filter shape, matrix derivation (with special processing), application area, syntax. In addition, the methods / examples listed in this section can also be applied to the methods mentioned above (more taps) to have better performance and some complexity tradeoffs.
[0690] In the present disclosure, reference samples / training templates / reconstructed neighboring regions generally refer to luma samples used to derive MLR model parameters, which are then applied to internal luma samples in a CU to predict chroma samples in the CU.
[0691] According to one or more embodiments of the present disclosure, CCCM without a downsampling process and CCCM with a downsampling process may be used by different taps or shapes. For example, the number of filter shapes / filter taps with and without downsampling values may be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0692] According to one or more embodiments of the present disclosure, the reference sample / training template / reconstructed neighboring region may be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, Fig. 9A An example of an L-shaped reconstruction region, the left / top reconstruction region, is shown to derive parameters.
[0693] Fig.30 A workflow of a method 3000 for encoding video data according to one or more aspects of the present disclosure is shown.
[0694] At step 3010, method 3000 includes obtaining a video block from a bitstream.
[0695] At step 3020, method 3000 includes obtaining an internal luma sample value of the video block, an external luma sample value of an external region of the video block, and an external chroma sample value of an external region.
[0696] At step 3030, method 3000 includes determining a set of weighting coefficients corresponding to a filter shape based on the outer luma sample values and the outer chroma sample values, wherein the filter shape and the set of weighting coefficients are configured to predict each of the inner chroma sample values based on a plurality of corresponding luma sample values, wherein the plurality of corresponding luma sample values include: one or more non-subsampled luma sample values associated with the filter shape, and one or more non-linear luma sample values.
[0697] At step 3040, method 3000 includes predicting intra chroma sample values based on the intra luma sample values using a filter shape and a set of weighting coefficients.
[0698] At step 3050 , method 3000 includes obtaining a predicted video block using the predicted intra chroma sample values.
[0699] In one example, the one or more non-linear luma sample values include: a first non-linear luma sample value derived by scaling a square of a first luma sample value to a luma sample value range of a video block, wherein the first luma sample value is a non-downsampled luma sample value associated with the filter shape.
[0700] In one example, the first luma sample value is a non-subsampled luma sample value of a first luma sample located at the center of the filter shape.
[0701] In one example, the one or more non-linear luma sample values further include: a second non-linear luma sample value derived by scaling the square of the second luma sample value to a luma sample value range of the video block.
[0702] In one example, the second luma sample value is a non-downsampled luma sample value of a second luma sample below the first luma sample.
[0703] In one example, the second luma sample value is a downsampled luma sample value determined based on a luma sample value of a co-located luma sample corresponding to a chroma sample to be predicted.
[0704] In one example, the downsampled luma sample values are determined by a weighted averaging operation.
[0705] In one example, the filter shape is determined based on six luma samples including: a first luma sample located at the center of the filter shape; a second luma sample located below the first luma sample; a third luma sample to the left of the first luma sample; a fourth luma sample to the right of the first luma sample; a fifth luma sample below and to the left of the first luma sample; and a sixth luma sample below and to the right of the first luma sample.
[0706] In one example, the set of weighting coefficients also includes a weighting coefficient for a bias value.
[0707] In one example, whether the second luma sample value is a downsampled luma sample value or a non-downsampled luma sample value is predefined or signaled at an SPS, DPS, VPS, SEI, APS, PPS, PH, SH, region, CTU, CU, sub-block, or sample level.
[0708] Fig.31 A workflow of a method 3100 for encoding video data according to one or more aspects of the present disclosure is shown.
[0709] At step 3110 , method 3100 includes obtaining a video block.
[0710] At step 3120 , method 3100 includes obtaining an internal luma sample value of the video block, an external luma sample value of an external region of the video block, and an external chroma sample value of an external region.
[0711] At step 3130, method 3100 includes determining a set of weighting coefficients corresponding to a filter shape based on external luma sample values and external chroma sample values, wherein the filter shape and the set of weighting coefficients are configured to predict each of the internal chroma sample values based on multiple corresponding luma sample values, wherein the multiple corresponding luma sample values include: one or more non-downsampled luma sample values associated with the filter shape and one or more non-linear luma sample values.
[0712] At step 3140, method 3100 includes predicting intra chroma sample values based on the intra luma sample values using a filter shape and a set of weighting coefficients.
[0713] At step 3150 , the method 3100 includes generating a bitstream including the encoded video block by using the predicted intra chroma sample values.
[0714] In one example, the one or more non-linear luma sample values include: a first non-linear luma sample value derived by scaling a square of a first luma sample value to a luma sample value range of a video block, wherein the first luma sample value is a non-downsampled luma sample value associated with the filter shape.
[0715] In one example, the first luma sample value is a non-subsampled luma sample value of a first luma sample located at the center of the filter shape.
[0716] In one example, the one or more non-linear luma sample values further include: a second non-linear luma sample value derived by scaling the square of the second luma sample value to a luma sample value range of the video block.
[0717] In one example, the second luma sample value is a non-downsampled luma sample value of a second luma sample below the first luma sample.
[0718] In one example, the second luma sample value is a downsampled luma sample value determined based on a luma sample value of a co-located luma sample corresponding to a chroma sample to be predicted.
[0719] In one example, the downsampled luma sample values are determined by a weighted averaging operation.
[0720] In one example, the filter shape is determined based on six luma samples including: a first luma sample at the center of the filter shape; a second luma sample below the first luma sample; a third luma sample to the left of the first luma sample; a fourth luma sample to the right of the first luma sample; a fifth luma sample below and to the left of the first luma sample; and a sixth luma sample below and to the right of the first luma sample.
[0721] In one example, the set of weighting coefficients also includes a weighting coefficient for a bias value.
[0722] In one example, whether the second luma sample value is a downsampled luma sample value or a non-downsampled luma sample value is predefined or signaled at an SPS, DPS, VPS, SEI, APS, PPS, PH, SH, region, CTU, CU, sub-block, or sample level.
[0723] Fig.32A computing environment 3210 is shown coupled to a user interface 3250. The computing environment 3210 may be part of a data processing server. The computing environment 3210 includes a processor 3220, a memory 3230, and an input / output (I / O) interface 3240.
[0724] The processor 3220 generally controls the overall operation of the computing environment 3210, such as operations associated with display, data acquisition, data communication, and image processing. The processor 3220 may include one or more processors to execute instructions to perform all or some steps in the above method. In addition, the processor 3220 may include one or more modules that facilitate the interaction between the processor 3220 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a graphics processing unit (GPU), etc.
[0725] The memory 3230 is configured to store various types of data to support the operation of the computing environment 3210. The memory 3230 may include predetermined software 3232. Examples of such data include instructions for any application or method operating on the computing environment 3210, video data sets, image data, etc. The memory 3230 may be implemented by using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0726] The I / O interface 3240 provides an interface between the processor 3220 and peripheral interface modules such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 3240 may be coupled to an encoder and a decoder.
[0727] In an embodiment, a non-transitory computer-readable storage medium is also provided, which includes multiple programs in the memory 3230, which can be executed by the processor 3220 in the computing environment 3210 to perform the above method and / or store the bit stream generated by the above encoding method or the bit stream decoded by the above decoding method. In one example, the multiple programs can be executed by the processor 3220 in the computing environment 3210 to (for example, from Figure 2The video encoder 20 in the computing environment 3210 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements, etc.), and can also be executed by the processor 3220 in the computing environment 3210 to perform the decoding method described above according to the received bitstream or data stream. In another example, multiple programs can be executed by the processor 3220 in the computing environment 3210 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 3220 in the computing environment 3210 to send the bitstream or data stream (e.g., to Figure 3 Alternatively, a non-transitory computer-readable storage medium may have stored therein a bitstream or data stream including the video decoder 30 in the encoder (e.g., Figure 2 The video encoder 20 in the example generates a video signal for the decoder (e.g., Figure 3 The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0728] In an embodiment, a bitstream generated by the above encoding method or a bitstream to be decoded by the above decoding method is provided. In an embodiment, a bitstream is provided, which includes the encoded video information generated by the above encoding method or the encoded video information to be decoded by the above decoding method.
[0729] In an embodiment, a computing device is also provided, which includes one or more processors (e.g., processor 3220) and a non-temporary computer-readable storage medium or memory 3230 storing therein a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the above-mentioned method when executing the plurality of programs.
[0730] In an embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, the bitstream including the encoded video information generated by the above encoding method or the encoded video information to be decoded by the above decoding method. In an embodiment, a computer program product is also provided, which includes a plurality of programs executable by the processor 3220 in the computing environment 3210 in, for example, a memory 3230, for performing the above method. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0731] In an embodiment, the computing environment 3210 may be implemented with one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0732] In an embodiment, a method for storing a bitstream is also provided, comprising storing the bitstream on a digital storage medium, wherein the bitstream comprises encoded video information generated by the encoding method or encoded video information to be decoded by the decoding method.
[0733] In an embodiment, a method for transmitting a bit stream generated by the above encoder is also provided. In an embodiment, a method for receiving a bit stream to be decoded by the above decoder is also provided.
[0734] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or to limit the present disclosure. Many modifications, variations, and alternative implementations will be apparent to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.
[0735] Unless otherwise specifically stated, the order of the steps of the method according to the present disclosure is intended to be illustrative only, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but can be changed according to actual conditions. In addition, at least one step of the method according to the present disclosure can be adjusted, combined or deleted according to actual requirements.
[0736] The examples are selected and described in order to explain the principles of the present disclosure and to enable others skilled in the art to understand the present disclosure for various implementations and to best utilize the basic principles and various implementations with various modifications suitable for the specific use contemplated. Therefore, it should be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed implementations, and modifications and other implementations are intended to be included within the scope of the present disclosure.
Claims
1. A method for decoding video data, comprising: Get video chunks from a bitstream; Obtaining an internal luminance sample value of the video block, an external luminance sample value of an external area of the video block, and an external chrominance sample value of the external area; Determining a set of weighting coefficients corresponding to a filter shape based on the outer luma sample values and the outer chroma sample values, wherein the filter shape and the set of weighting coefficients are configured to predict each of the inner chroma sample values based on a plurality of corresponding luma sample values, wherein the plurality of corresponding luma sample values include: one or more non-subsampled luma sample values associated with the filter shape, and One or more non-linear brightness sample values; predicting the intra chroma sample values based on the intra luma sample values using the filter shape and the set of weighting coefficients; and The predicted intra chroma sample values are used to obtain a predicted video block.
2. The method according to claim 1, wherein: The one or more non-linear brightness sample values include: A first non-linear luma sample value is derived by scaling a square of a first luma sample value to a range of luma sample values for the video block, wherein the first luma sample value is a non-subsampled luma sample value associated with the filter shape.
3. The method according to claim 2, wherein: The first luma sample value is the non-subsampled luma sample value of a first luma sample located at the center of the filter shape.
4. The method according to claim 3, wherein: The one or more non-linear brightness sample values also include: A second non-linear luma sample value is derived by scaling the square of a second luma sample value to the luma sample value range of the video block.
5. The method according to claim 4, wherein: The second luma sample value is a non-subsampled luma sample value of a second luma sample below the first luma sample.
6. The method according to claim 4, wherein: The second luma sample value is a downsampled luma sample value determined based on a luma sample value of a co-located luma sample corresponding to the chroma sample to be predicted.
7. The method according to claim 6, wherein: The downsampled luminance sample values are determined by a weighted averaging operation.
8. The method according to claim 1, wherein: The filter shape is determined based on six luma samples including: a first luminance sample located at the center of the filter shape; a second luminance sample below the first luminance sample; a third luminance sample point to the left of the first luminance sample point; a fourth luminance sample point to the right of the first luminance sample point; a fifth brightness sample point below and to the left of the first brightness sample point; as well as A sixth luma sample is below and to the right of the first luma sample.
9. The method according to claim 1, wherein: The set of weighting coefficients also includes a weighting coefficient for a bias value.
10. The method according to claim 4, wherein: Whether the second luma sample value is a downsampled luma sample value or a non-downsampled luma sample value is predefined or signaled at an SPS, DPS, VPS, SEI, APS, PPS, PH, SH, region, CTU, CU, sub-block or sample level.
11. A method for encoding video data, comprising: Get the video block; Obtaining an internal luminance sample value of the video block, an external luminance sample value of an external area of the video block, and an external chrominance sample value of the external area; Determining a set of weighting coefficients corresponding to a filter shape based on the outer luma sample values and the outer chroma sample values, wherein the filter shape and the set of weighting coefficients are configured to predict each of the inner chroma sample values based on a plurality of corresponding luma sample values, wherein the plurality of corresponding luma sample values include: one or more non-subsampled luma sample values associated with the filter shape, and One or more non-linear brightness sample values; predicting the intra chroma sample values based on the intra luma sample values using the filter shape and the set of weighting coefficients; and A bitstream including an encoded video block is generated by using the predicted intra chroma sample values.
12. The method according to claim 11, wherein: The one or more non-linear brightness sample values include: A first non-linear luma sample value is derived by scaling a square of a first luma sample value to a range of luma sample values for the video block, wherein the first luma sample value is a non-subsampled luma sample value associated with the filter shape.
13. The method according to claim 12, wherein: The first luma sample value is the non-subsampled luma sample value of a first luma sample located at the center of the filter shape.
14. The method according to claim 13, wherein: The one or more non-linear brightness sample values also include: A second non-linear luma sample value is derived by scaling the square of a second luma sample value to the luma sample value range of the video block.
15. The method according to claim 14, wherein: The second luma sample value is a non-subsampled luma sample value of a second luma sample below the first luma sample.
16. The method according to claim 14, wherein: The second luma sample value is a downsampled luma sample value determined based on a luma sample value of a co-located luma sample corresponding to the chroma sample to be predicted.
17. The method according to claim 16, wherein: The downsampled luminance sample values are determined by a weighted averaging operation.
18. The method according to claim 11, wherein: The filter shape is determined based on six luma samples including: a first luminance sample located at the center of the filter shape; a second luminance sample below the first luminance sample; a third luminance sample point to the left of the first luminance sample point; a fourth luminance sample point to the right of the first luminance sample point; a fifth brightness sample point below and to the left of the first brightness sample point; as well as A sixth luma sample is below and to the right of the first luma sample.
19. The method according to claim 11, wherein: The set of weighting coefficients also includes a weighting coefficient for a bias value.
20. The method according to claim 14, wherein: Whether the second luma sample value is a downsampled luma sample value or a non-downsampled luma sample value is predefined or signaled at an SPS, DPS, VPS, SEI, APS, PPS, PH, SH, region, CTU, CU, sub-block or sample level.
21. A computer system comprising: one or more processors; as well as One or more storage devices storing computer executable instructions which, when executed, cause the one or more processors to perform the operations of the method according to any one of claims 1-20.
22. A computer program product storing computer executable instructions which, when executed, cause one or more processors to perform the operations of the method according to any one of claims 1-20.
23. A computer-readable storage medium storing instructions which, when executed by a computing device having one or more processors, cause the one or more processors to perform a decoding method according to any one of claims 1-10, and store a bit stream to be decoded by the decoding method according to any one of claims 1-10.
24. A computer-readable storage medium storing instructions which, when executed by a computing device having one or more processors, cause the one or more processors to perform the encoding method according to any one of claims 11-20, and store a bit stream generated by the encoding method according to any one of claims 11-20.
25. A computer-readable medium storing a bit stream, wherein: The bit stream is to be decoded by performing the operations of the method according to any one of claims 1-10.
26. A computer-readable medium storing a bit stream, wherein: The bit stream is obtained by performing the operations of the method according to any one of claims 11-20.