Cross-component sample adaptive offset

By utilizing the cross-component relationship in the video encoding process, and modifying the sample point value through the decoder and encoder selection offset, the problem of low coding efficiency of brightness and chromaticity components is solved, and the compression performance of video data is improved.

CN120513627APending Publication Date: 2025-08-19BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480007609.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-11
Filing Date
2024-01-11
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing video encoding technology has low encoding efficiency when processing brightness and chrominance components, making it difficult to fully utilize the cross-component relationship between the two.

Method used

The video signal is received by the decoder and the encoder respectively and determine a plurality of offsets associated with the second component, and the classifier selects the offset from it to modify the sample point value of the second component, thereby achieving an adaptive offset between the brightness and chrominance components.

Benefits of technology

Improves the encoding efficiency of brightness and chromaticity components, enhances the compression performance of video data, and reduces the bit rate requirement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120513627A_ABST
    Figure CN120513627A_ABST
Patent Text Reader

Abstract

Methods and apparatus for video coding and decoding are provided. In one method, a decoder may receive a video signal including a first component and a second component, and receive a plurality of offsets associated with the second component. Further, the decoder may obtain a classifier associated with the second component based on the residual sample value of the first component, and select an offset from a plurality of offsets of the second component based on the classifier. Further, the decoder may obtain a modified sample value for the second component based on the selected offset.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is based upon and claims the benefit of U.S. Provisional Patent Application No. 63 / 438,531, filed on January 11, 2023, entitled “CROSS-COMPONENT SAMPLEADAPTIVE OFFSET,” the contents of which are incorporated herein by reference in their entirety for all purposes. Technical Field

[0002] The present disclosure relates generally to video coding and compression, and more particularly to methods and apparatus for improving luma and chroma coding efficiency. Background Art

[0003] Various electronic devices (such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc.) support digital video. Electronic devices send and receive or otherwise transmit digital video data through a communication network, and / or store digital video data on a storage device. Since the bandwidth capacity of the communication network is limited and the storage resources of the storage device are limited, before the video data is transmitted or stored, a video codec can be used to compress the video data according to one or more video codec standards. For example, video codec standards include Versatile Video Codec (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. AOMedia Video 1 (AV1) was developed as a successor to its predecessor VP9. Audio and Video Codec (AVS) refers to a compression standard for digital audio and digital video, and is another series of video compression standards. Video codecs typically employ prediction methods (eg, inter-frame prediction, intra-frame prediction, etc.) that exploit the redundancy inherent in video data. Video codecs aim to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation in video quality. Summary of the Invention

[0004] The present disclosure describes embodiments related to video data encoding and decoding, and more particularly, to methods and apparatus for improving coding efficiency of both luma and chroma components, including improving coding efficiency by exploring cross-component relationships between luma and chroma components.

[0005] According to a first aspect of the present application, a method for video decoding is provided. The method may include: a decoder receiving a video signal including a first component and a second component, and receiving a plurality of offsets associated with the second component. Furthermore, the method may include: the decoder obtaining a classifier associated with the second component based on residual sample values of the first component, and selecting an offset from the plurality of offsets of the second component based on the classifier. Furthermore, the method may include: the decoder obtaining a modified sample value of the second component based on the selected offset.

[0006] According to a second aspect of the present application, a method for video encoding is provided. The method may include: an encoder determining a video signal including a first component and a second component, and determining a plurality of offsets associated with the second component. Furthermore, the method may include: the encoder obtaining a classifier associated with the second component based on residual sample values of the first component, and selecting an offset from a plurality of offsets of the second component based on the classifier. Furthermore, the method may include the encoder obtaining a modified sample value of the second component based on the selected offset.

[0007] According to a third aspect of the present application, a device for video decoding is provided. The device may include: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. When executing the instructions, the one or more processors are configured to perform the method according to the first aspect.

[0008] According to a fourth aspect of the present application, a device for video encoding is provided. The device may include: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. When executing the instructions, the one or more processors are configured to perform the method according to the second aspect.

[0009] According to a fifth aspect of the present application, a non-transitory computer-readable storage medium stores computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the first aspect.

[0010] According to a sixth aspect of the present application, a non-transitory computer-readable storage medium is used to store computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the second aspect and send a bit stream.

[0011] According to a seventh aspect of the present application, a non-transitory computer-readable storage medium is provided for storing a bit stream to be decoded by the method according to the first aspect.

[0012] According to an eighth aspect of the present application, a non-transitory computer-readable storage medium is used to store a bit stream generated by the method according to the second aspect.

[0013] It is to be understood that both the foregoing general description and the following detailed description are examples only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples according to the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0015] Figure 1 is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to some embodiments of the present disclosure.

[0016] Figure 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0017] Figure 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0018] Figures 4A to 4E is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure.

[0019] Figure 4F is a block diagram illustrating intra mode as defined in VVC.

[0020] Figure 4G is a block diagram illustrating multiple reference lines used for intra prediction.

[0021] Figure 5A is a block diagram illustrating four gradient modes used in Sample Adaptive Offset (SAO) according to some embodiments of the present disclosure.

[0022] Figure 5B is a block diagram illustrating a decoder for a deblocking filter (DBF) combined with the proposed SAO filtering SAOV and SAOH according to some embodiments of the present disclosure.

[0023] Figure 6 is a block diagram illustrating that the proposed bilateral filter (BIF) and SAO both use samples from the deblocking stage as input according to some embodiments of the present disclosure.

[0024] Figure 7 is a block diagram illustrating a naming convention for samples around a center sample according to some embodiments of the present disclosure.

[0025] Figure 8A is a block diagram illustrating a 5x5 diamond ALF filter applied to chroma components according to some embodiments of the present disclosure.

[0026] Figure 8B is a block diagram illustrating a 7x7 diamond ALF filter applied to the luma component according to some embodiments of the present disclosure.

[0027] 9A to 9D The sub-sample Laplacian calculation according to some embodiments of the present disclosure is shown.

[0028] Figure 10A is a block diagram illustrating a system-level diagram of a CC-ALF process with respect to SAO, luma ALF process, and chroma ALF process according to some embodiments of the present disclosure.

[0029] Figure 10B It is shown that filtering in CC-ALF is achieved by applying a linear diamond filter to the luma channel according to some embodiments of the present disclosure.

[0030] Figure 11 Modified block classification at virtual boundaries is shown in accordance with some embodiments of the present disclosure.

[0031] Figure 12 Modified ALF filtering for luma components at virtual boundaries is shown in accordance with some embodiments of the present disclosure.

[0032] Figure 13A CCSAO applied to chroma samples and using DBF Y as input is shown, according to some embodiments of the present disclosure.

[0033] Figure 13B CCSAO applied to luma and chroma samples and using DBF Y / Cb / Cr as input is shown according to some embodiments of the present disclosure.

[0034] Figure 13C A CCSAO is shown operating independently according to some embodiments of the present disclosure.

[0035] Figure 13D A recursively applied CCSAO is shown according to some embodiments of the present disclosure.

[0036] Figure 13E The parallel application of SAO and BIF according to some embodiments of the present disclosure is shown.

[0037] Figure 13F An alternative SAO according to some embodiments of the present disclosure is shown and applied in parallel with the BIF.

[0038] Figure 14 It is shown that CCSAO is applied in parallel with other coding tools according to some embodiments of the present disclosure.

[0039] Figure 15A The position of CCSAO according to some embodiments of the present disclosure is shown after SAO.

[0040] Figure 15B CCSAO is shown working independently without CCALF, according to some embodiments of the present disclosure.

[0041] Figure 15C CCSAO used as a post-reconstruction filter according to some embodiments of the present disclosure is shown.

[0042] Figure 16 CCSAO applied in parallel with CCALF is shown according to some embodiments of the present disclosure.

[0043] Figure 17 It is shown that different brightness sample positions are used as another classifier for C0 classification according to some embodiments of the present disclosure.

[0044] Figures 18A to 18G Different candidate shapes are shown to which constraints may be applied according to some embodiments of the present disclosure.

[0045] Figure 19 It is shown that in addition to luma, other cross-component co-located and neighboring chroma samples may be fed into the CCSAO classification according to some embodiments of the present disclosure.

[0046] FIG. 20A to FIG. 20B Replacing co-located luma sample values with phase correction values by weighting adjacent luma samples according to some embodiments of the present disclosure is shown.

[0047] Figures 21A to 21B Replacing co-located luma sample values with phase correction values by weighting adjacent luma samples according to some embodiments of the present disclosure is shown.

[0048] FIG. 22A to FIG. 22B An example of classifying c using edge strength according to some embodiments of the present disclosure is shown.

[0049] FIG. 23A to FIG. 23B It is shown that according to some embodiments of the present disclosure, CCSAO is not applied to the current chroma sample if any of the co-located and neighboring luma samples used for classification is outside the current picture.

[0050] FIG. 24A to FIG. 24BIt is shown that according to some embodiments of the present disclosure, if any of the co-located and adjacent luma samples for classification is outside the current picture, the missing sample is used for repeating or mirroring to create the sample for classification.

[0051] Figure 25 It is shown that according to some embodiments of the present disclosure, 9 luma candidate CCSAOs in AVS can add 2 additional luma line buffers.

[0052] Figure 26A It is shown that according to some embodiments of the present disclosure, 9 luma candidate CCSAOs in VVC can add 1 additional luma line buffer.

[0053] Figure 26B It is shown that if co-located or adjacent chroma samples are used to classify the current luma sample, the selected chroma candidate may span VBs and require an additional chroma line buffer, according to some embodiments of the present disclosure.

[0054] Figures 27A to 27C It is shown that in AVS and VVC, if any of the luma candidates for the chroma sample spans VB (outside the current chroma sample VB), CCSAO is disabled for the chroma sample in accordance with some embodiments of the present disclosure.

[0055] FIG. 28A to FIG. 28B An example of a virtual boundary of C0 with 9 luminance position candidates is shown, according to some embodiments of the present disclosure.

[0056] Figures 29A to 29C It is shown that in AVS and VVC, if any of the luma candidates of the chroma samples spans VB (outside the current chroma sample VB), CCSAO is enabled using repeated padding of chroma samples according to some embodiments of the present disclosure.

[0057] Figures 30A to 30C It is shown that in AVS and VVC, if any of the luma candidates of the chroma samples spans VB (outside the current chroma sample VB), CCSAO is enabled using mirror padding of the chroma samples according to some embodiments of the present disclosure.

[0058] Figures 31A to 31B It is shown that in AVS and VVC, if one side is outside the VB, bilateral symmetric padding is used to enable CCSAO according to some embodiments of the present disclosure.

[0059] FIG. 32A to FIG. 32B It is shown that repeating or mirroring padding may be applied to luma samples outside of the virtual boundary according to some embodiments of the present disclosure.

[0060] Figures 33A to 33BLimitations for reducing the line buffer required for CCSAO and simplifying boundary processing condition checks according to some embodiments of the present disclosure are shown.

[0061] Figure 34 Regions where CCSAO is applied that are not aligned with CTB boundaries according to some embodiments of the present disclosure are shown.

[0062] Figure 35 FIG. 4 illustrates a region frame segmentation to which CCSAO can be fixedly applied according to some embodiments of the present disclosure.

[0063] Figure 36 It is shown that the region segmentation to which CCSAO is applied according to some embodiments of the present disclosure may be dynamic and switched at a picture level.

[0064] Figure 37 It shows how to apply a classifier set index that can be switched at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block levels if multiple classifiers are used in one frame according to some embodiments of the present disclosure.

[0065] Figure 38 It is shown that according to some embodiments of the present disclosure, the area where CCSAO is applied may be BT / QT / TT separated from the frame / slice / CTB level.

[0066] Figure 39 A CCSAO classifier that considers current or cross-component coding information according to some embodiments of the present disclosure is shown.

[0067] Figure 40A A block diagram illustrating the SAO classification method disclosed in this disclosure used as a prediction post-filter according to some embodiments of the present disclosure.

[0068] Figures 40B to 40D A block diagram illustrating that for a post-prediction SAO filter, each component may be classified using current and neighboring samples according to some embodiments of the present disclosure.

[0069] Figure 41 is a diagram illustrating a computing environment coupled with a user interface according to some embodiments of the present disclosure.

[0070] Figure 42 is a flowchart illustrating a method for video decoding according to an example of the present disclosure.

[0071] Figure 43 is a flowchart illustrating a method for video encoding according to an example of the present disclosure. DETAILED DESCRIPTION

[0072] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to facilitate understanding of the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein may be implemented on many types of electronic devices with digital video capabilities.

[0073] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a," "an," "the," and "the" in this disclosure and the appended claims are intended to include the plural forms as well, unless otherwise expressly indicated throughout the disclosure. It should also be understood that the term "and / or" as used in this disclosure refers to and includes one or any or all possible combinations of the listed items.

[0074] Reference throughout this specification to "one embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language means that the particular feature, structure, or characteristic being described is included in at least one embodiment or example. Unless expressly stated otherwise, a feature, structure, element, or characteristic described in conjunction with one or some embodiments may also apply to other embodiments.

[0075] Throughout the disclosure, unless otherwise expressly stated, the terms "first," "second," "third," etc. are used merely as references to related elements (e.g., devices, components, compositions, steps, etc.) without implying any spatial or temporal order. For example, "first device" and "second device" may refer to two separately formed devices, or two parts, components, or operating states of the same device, and may be named arbitrarily.

[0076] The terms "module," "sub-module," "circuit," "sub-circuit," "circuitry," "sub-circuitry," "unit," or "sub-unit" may include memory (shared, dedicated, or group) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits with or without stored code or instructions. A module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be physically attached to or located adjacent to each other.

[0077] As used herein, the terms "if" or "when" may be understood to mean "at the time of" or "in response to" depending on the context. These terms, if they appear in a claim, may not indicate that the associated limitation or feature is conditional or optional. For example, a method may include the following steps: i) when condition X exists or if condition X exists, perform function or action X', and ii) when condition Y exists or if condition Y exists, perform function or action Y'. The method may be implemented with the ability to perform function or action X' as well as the ability to perform function or action Y'. Thus, both functions X' and Y' may be performed at different times during multiple executions of the method.

[0078] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, a unit or module may include functionally related code blocks or software components that are linked together directly or indirectly to perform a specific function.

[0079] The first-generation AVS standard includes the Chinese national standards "Information Technology, Advanced Audio and Video Coding, Part 2: Video" (AVS1) and "Information Technology, Advanced Audio and Video Coding, Part 16: Broadcast and Television Video" (AVS+). Compared to the MPEG-2 standard, it can achieve approximately 50% bitrate savings while maintaining the same perceptual quality. The second-generation AVS standard includes a series of Chinese national standards, "Information Technology, High-Efficiency Multimedia Coding" (AVS2), primarily targeting the transmission of ultra-high-definition television programs. AVS2 boasts twice the coding efficiency of AVS+. The video portion of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) for international application. The AVS3 standard is a next-generation video coding standard for UHD video applications, aiming to surpass the coding efficiency of the latest international standard, HEVC, and offering approximately 30% bitrate savings compared to HEVC. The AVS3-P2 baseline, finalized at the 68th AVS meeting in March 2019, offers approximately 30% bitrate savings compared to HEVC. Currently, a reference software called the High Performance Model (HPM) is maintained by the AVS group to demonstrate a reference implementation of the AVS3 standard. Similar to HEVC, the AVS3 standard is built on a block-based hybrid video coding framework.

[0080] Figure 1 FIG. 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. Figure 1As shown in , system 10 includes a source device 12 that generates and encodes video data to be later decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including a cloud server, a server computer, a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0081] In some embodiments, target device 14 can receive the encoded video data to be decoded via link 16. Link 16 can include any type of communication medium or device capable of moving the encoded video data from source device 12 to target device 14. In one example, link 16 can include a communication medium that enables source device 12 to send the encoded video data directly to target device 14 in real time. The encoded video data can be modulated according to a communication standard (e.g., a wireless communication protocol) and sent to target device 14. The communication medium can include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form a portion of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium can include a router, a switch, a base station, or any other device that can facilitate communication from source device 12 to target device 14.

[0082] In some other embodiments, the encoded video data can be sent from the output interface 22 to a storage device 32. The encoded video data in the storage device 32 can then be accessed by the target device 14 via the input interface 28. The storage device 32 can include any of a variety of distributed or locally accessible data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The target device 14 can access the stored video data from the storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing and sending the encoded video data to the target device 14. Exemplary file servers include a network server (e.g., for a website), a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Target device 14 may access the encoded video data through any standard data connection suitable for accessing encoded video data stored on a file server, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both. The transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both streaming and download transmissions.

[0083] like Figure 1 As shown in , source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include a source such as a video capture device (e.g., a video camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources. As an example, if video source 18 is a camera of a security monitoring system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application may be generally applicable to video encoding and decoding, and may be applied to wireless and / or wired applications.

[0084] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be sent directly to target device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored on storage device 32 for later access by target device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.

[0085] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.

[0086] In some implementations, target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with target device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0087] The video encoder 20 and the video decoder 30 can operate according to a proprietary standard or an industry standard (e.g., VVC, HEVC, Part 10 of MPEG-4, AVC, AVS) or an extension of such a standard. It should be understood that the present application is not limited to a specific video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally believed that the video encoder 20 of the source device 12 can be configured to encode the video data according to any of these current standards or future standards. Similarly, it is also generally believed that the video decoder 30 of the target device 14 can be configured to decode the video data according to any of these current standards or future standards.

[0088] The video encoder 20 and the video decoder 30 can be implemented as any of a variety of suitable encoder and / or decoder circuits, respectively, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in the hardware to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders, and either encoder or decoder can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.

[0089] In some embodiments, components of source device 12 (e.g., video source 18, video encoder 20, or the like) may be configured to: Figure 2 The components included in the video encoder 20 and the output interface 22) and / or the components of the target device 14 (for example, the input interface 28, the video decoder 30 or the following reference Figure 3The components included in the video decoder 30 and at least a portion of the components in the display device 34 may be operated in a cloud computing service network such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), wherein the cloud computing service network can provide software, platform, and / or infrastructure. In some embodiments, one or more components of the source device 12 and / or the target device 14 that are not included in the cloud computing service network may be set in one or more client devices, and the one or more client devices may communicate with a server computer in the cloud computing service network via a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In one embodiment, at least a portion of the operations described herein may be implemented as a cloud-based service provided by one or more server computers, wherein the one or more server computers are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in the cloud computing service network; and one or more other operations described herein may be implemented by one or more client devices. In some embodiments, the cloud computing service network can be a private cloud, a public cloud, or a hybrid cloud. Terms such as "cloud," "cloud computing," and "cloud-based" may be used interchangeably herein without departing from the scope of this disclosure. It should be understood that this disclosure is not limited to implementation within the aforementioned cloud computing service network. Instead, this disclosure may also be implemented within any other type of computing environment currently known or developed in the future.

[0090] Figure 2 is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described herein. Video encoder 20 can perform intra-frame prediction coding and inter-frame prediction coding on video blocks within a video frame. Intra-frame prediction coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-frame prediction coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. It should be noted that in the field of video coding, the term "frame" can be used as a synonym for the term "image" or "picture."

[0091] like Figure 2As shown in FIG, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58 for video block reconstruction, an inverse transform processing unit 60, and an adder 62. A loop filter 63, such as a deblocking filter, can be located between the adder 62 and the DPB 64 to filter block boundaries to remove blocking artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can also be used to filter the output of the adder 62. It should be noted that with respect to the CCSAO technique, the present application is not limited to the embodiments described herein, but may also be applied to the case where an offset is selected for any other of the luma component, the Cb chroma component, and the Cr chroma component based on any of the luma component, the Cb chroma component, and the Cr chroma component to modify the other component based on the selected offset. Furthermore, it should be noted that the first component mentioned herein may be any one of the luma component, the Cb chroma component, and the Cr chroma component, the second component mentioned herein may be any other one of the luma component, the Cb chroma component, and the Cr chroma component, and the third component mentioned herein may be the remaining component of the luma component, the Cb chroma component, and the Cr chroma component. In some examples, the loop filter may be omitted, and the decoded video block may be provided directly by the adder 62 to the DPB 64. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be distributed among one or more of the fixed or programmable hardware units described.

[0092] Video data memory 40 may store video data to be encoded by the components of video encoder 20. Figure 1 The video source 18 shown obtains video data from the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use by the video encoder 20 (e.g., in intra-frame or inter-frame prediction coding mode) when encoding the video data. The video data memory 40 and the DPB 64 can be formed by any of a variety of memory devices. In various examples, the video data memory 40 can be on-chip with other components of the video encoder 20, or off-chip relative to those components.

[0093] like Figure 2As shown in , after receiving the video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a set of video blocks), or other larger coding units (CUs) according to a predefined splitting structure associated with the video data, such as a quadtree (QT) structure. A video frame is or can be viewed as a two-dimensional array or matrix of sample values. The samples in the array may also be referred to as pixels or picture elements (pel). The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. For example, a video frame may be divided into multiple video blocks by using QT segmentation. A video block is again or can be viewed as a two-dimensional array or matrix of sample values, but its dimensions are smaller than the dimensions of the video frame. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. The video block may be further partitioned into one or more block partitions or sub-blocks (which may again form blocks) by, for example, iteratively using QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof. It should be noted that the terms "block" or "video block" as used herein may be a portion of a frame or picture, in particular a rectangular (square or non-square) portion. With reference to, for example, HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU), or a transform unit (TU) and / or may be or correspond to a corresponding block (e.g., a coding tree block (CTB), a coding block (CB), a prediction block (PB), or a transform block (TB)) and / or a sub-block.

[0094] The prediction processing unit 41 may select one of a plurality of possible prediction coding modes for the current video block based on the error results (e.g., coding rate and distortion level), such as one of one or more inter-frame prediction coding modes among a plurality of intra-frame prediction coding modes. The prediction processing unit 41 may provide the resulting intra-frame prediction coding block or inter-frame prediction coding block to the adder 50 to generate a residual block and to the adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, partition information, and other such syntax information) to the entropy coding unit 56.

[0095] To select an appropriate intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-prediction coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple encoding passes, for example, to select an appropriate coding mode for each block of video data.

[0096] In some embodiments, motion estimation unit 42 determines the inter-prediction mode for the current video frame by generating motion vectors according to a predetermined pattern within a sequence of video frames, where the motion vectors indicate the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of a video block. For example, a motion vector may indicate the displacement of a video block within the current video frame or picture relative to a prediction block within a reference frame associated with the current block being encoded within the current frame. The predetermined pattern may designate the video frames in the sequence as P-frames or B-frames. Intra BC unit 48 may determine vectors (e.g., block vectors) for intra BC coding in a manner similar to the motion vectors determined by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine the block vectors.

[0097] In terms of pixel differences, the prediction block for a video block may be or may correspond to a block or reference block of a reference frame that is considered to closely match the video block to be encoded, and the pixel differences may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some embodiments, video encoder 20 may calculate values for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, motion estimation unit 42 may perform motion searches relative to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.

[0098] Motion estimation unit 42 calculates a motion vector for a video block in an inter-prediction coded frame by comparing the position of the video block to the position of a prediction block of a reference frame selected from either a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.

[0099] Motion compensation performed by motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by motion estimation unit 42. After receiving the motion vector for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel difference values forming the residual video block may include luma component differences, chroma component differences, or both. Motion compensation unit 44 may also generate syntax elements associated with the video block of the video frame for use by video decoder 30 when decoding the video block of the video frame. The syntax elements may include, for example, syntax elements defining a motion vector for identifying the prediction block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated but are described separately for conceptual purposes.

[0100] In some embodiments, the intra BC unit 48 may generate vectors and obtain prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44, but these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra BC unit 48 may determine the intra prediction mode to be used to encode the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example during separate encoding passes, and test their performance using rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode to use from the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values for the various tested intra prediction modes using rate-distortion analysis and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original, uncoded block that was coded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra BC unit 48 may calculate ratios based on the distortion and rate for various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.

[0101] In other examples, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction in accordance with embodiments described herein. In either case, for intra block copying, the prediction block may be a block that is considered to closely match the block to be encoded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identifying the prediction block may include calculating values for sub-integer pixel positions.

[0102] Regardless of whether the prediction block is from the same frame according to intra-frame prediction or from a different frame according to inter-frame prediction, video encoder 20 can form pixel difference values by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded, thereby forming a residual video block. The pixel difference values forming the residual video block may include both luma component differences and chroma component differences.

[0103] As an alternative to the inter-frame prediction performed by motion estimation unit 42 and motion compensation unit 44 or the intra-frame block copy prediction performed by intra BC unit 48 as described above, intra-frame prediction processing unit 46 can perform intra-frame prediction on the current video block. Specifically, intra-frame prediction processing unit 46 can determine an intra-frame prediction mode to use for encoding the current block. To do so, intra-frame prediction processing unit 46 can use various intra-frame prediction modes to encode the current block, for example, during separate encoding passes, and intra-frame prediction processing unit 46 (or in some examples, mode selection unit) can select an appropriate intra-frame prediction mode to use from the tested intra-frame prediction modes. Intra-frame prediction processing unit 46 can provide information indicating the intra-frame prediction mode selected for the block to entropy coding unit 56. Entropy coding unit 56 can encode the information indicating the selected intra-frame prediction mode into the bitstream.

[0104] After prediction processing unit 41 determines a prediction block for the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0105] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the quantization unit 54 may then perform a scan on the matrix comprising the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.

[0106] After quantization, entropy coding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream may then be sent to a video bitstream such as Figure 1 The video decoder 30 shown, or archived as Figure 1 The video frame is shown in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 may also entropy encode motion vectors and other syntax elements for the current video frame being encoded.

[0107] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain for use in generating a reference block for predicting other video blocks. As noted above, motion compensation unit 44 may generate a motion compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.

[0108] Adder 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block to inter-predict another video block in a subsequent video frame.

[0109] Figure 3 3 is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 may perform the operations described above in conjunction with the above. Figure 2 The encoding process is essentially the inverse of the decoding process described with respect to video encoder 20. For example, motion compensation unit 82 may generate prediction data based on motion vectors received from entropy decoding unit 80, and intra-prediction unit 84 may generate prediction data based on intra-prediction mode indicators received from entropy decoding unit 80.

[0110] In some examples, units of the video decoder 30 may be tasked with performing embodiments of the present application. Furthermore, in some examples, embodiments of the present disclosure may be dispersed across one or more of the units of the video decoder 30. For example, the intra BC unit 85 may perform embodiments of the present application alone or in combination with other units of the video decoder 30 (e.g., the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra BC unit 85, and the functionality of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., the motion compensation unit 82).

[0111] The video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, from a local video source (e.g., a camera), via a wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data for use by the video decoder 30 (e.g., in intra-frame or inter-frame prediction coding mode) when decoding the video data. The video data memory 79 and the DPB 92 may be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are stored in Figure 3 92 as two distinct components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or by separate memory devices. In some examples, video data memory 79 may be on-chip with the other components of video decoder 30, or off-chip relative to those components.

[0112] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes the bitstream to generate quantization coefficients, motion vectors or intra-frame prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-frame prediction mode indicators, and other syntax elements to the prediction processing unit 81.

[0113] When a video frame is encoded as an intra-frame prediction coded (I) frame or for intra-frame coded prediction blocks in other types of frames, intra-frame prediction unit 84 of prediction processing unit 81 can generate prediction data for a video block of the current video frame based on the intra-frame prediction mode transmitted by the signal and reference data from a previously decoded block of the current frame.

[0114] When the video frame is encoded as an inter-frame prediction coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the video block of the current video frame based on the motion vector and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame in one of the reference frame lists. The video decoder 30 can use a default construction technique to construct the reference frame lists, i.e., List 0 and List 1, based on the reference frames stored in the DPB 92.

[0115] In some examples, when a video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from entropy decoding unit 80. The prediction block may be within a reconstructed region of the same picture as the current video block, as defined by video encoder 20.

[0116] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for a video block of the current video frame by parsing the motion vectors and other syntax elements, and then uses the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used to encode the video block of the video frame, the inter-prediction frame type (e.g., B or P), construction information for one or more of the reference frame lists for the frame, the motion vector for each inter-prediction-encoded video block of the frame, the inter-prediction state for each inter-prediction-encoded video block of the frame, and other information used to decode the video block in the current video frame.

[0117] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as flags, to determine whether the current video block is predicted using intra BC mode, construction information of which video blocks of the frame are within the reconstruction region and should be stored in the DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding video blocks in the current video frame.

[0118] Motion compensation unit 82 may also perform interpolation to calculate interpolated values for sub-integer pixels of a reference block using interpolation filters, such as those used by video encoder 20 during encoding of the video block. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 from received syntax elements and use these interpolation filters to produce the prediction block.

[0119] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80, using the same quantization parameters that were calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to reconstruct the residual block in the pixel domain.

[0120] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter 91 (e.g., a deblocking filter, an SAO filter, a CCSAO filter, and / or an ALF) may be located between the adder 90 and the DPB 92 to further process the decoded video block. The loop filter 91 may be applied to the reconstructed CU before it is placed in the reference picture storage. In some examples, the loop filter 91 may be omitted, and the decoded video block may be provided directly to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92, or a memory device separate from the DPB 92, may also store the decoded video for later presentation on a display device (e.g., Figure 1 on the display device 34).

[0121] In a typical video encoding process, a video sequence typically consists of an ordered set of frames or pictures. Each frame can include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other examples, a frame can be monochrome and therefore only include a two-dimensional array of luma samples.

[0122] Similar to HEVC, the AVS3 standard is built on a block-based hybrid video coding framework. The input video signal is processed block by block (called coding unit (CU)). Unlike HEVC, which is only based on quadtree partitioning blocks, in AVS3, a coding tree unit (CTU) is partitioned into CUs based on quadtree / binary tree / extended quadtree to adapt to changing local characteristics. In addition, the concept of multiple partitioning unit types in HEVC is removed, that is, in AVS3, there is no separation of CU, prediction unit (PU) and transform unit (TU). Instead, each CU is always used as a basic unit for both prediction and transformation without further partitioning. In the tree partitioning structure of AVS3, a CTU is first partitioned based on the quadtree structure. Then, each quadtree leaf node can be further partitioned based on the binary tree and extended quadtree structure.

[0123] like Figure 4A As shown in , the video encoder 20 (or more specifically, the segmentation unit 45) generates an encoded representation of a frame by first segmenting the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and from top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a particular size. As Figure 4B As shown in , each CTU may include one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units of coding pixel blocks and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding the samples of the coding tree block. The coding tree block may be an N×N block of samples.

[0124] To achieve better performance, the video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof, on the coding tree block of the CTU and divide the CTU into smaller CUs. Figure 4C As depicted in FIG, a 64×64 CTU 400 is first divided into four smaller CUs, each having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. Two 16×16 CUs 430 and CU 440 are each further divided into four CUs with a block size of 8×8. Figure 4D Depicted is a diagram showing Figure 4C The quadtree data structure is the final result of the partitioning process of the CTU 400 depicted in FIG. 4 , with each leaf node of the quadtree corresponding to a CU of a corresponding size ranging from 32×32 to 8×8. Figure 4B Each CU may include a CB of luma samples and two corresponding coding blocks of chroma samples of the same size frame, and syntax elements for encoding the samples of the coding blocks. In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and syntax structures for encoding the samples of the coding block. It should be noted that Figure 4C and Figure 4D The quadtree partitioning depicted in FIG is for illustrative purposes only, and one CTU can be split into multiple CUs based on quadtree partitioning / ternary tree partitioning / binary tree partitioning to adapt to varying local characteristics. In the multi-type tree structure, one CTU is partitioned according to the quadtree structure, and each quadtree leaf CU can be further partitioned according to the binary and ternary tree structures. Figure 4E As shown, there are five possible partition types for a coding block with width W and height H, namely, quadruple partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning. In AVS3, there are five possible partition types, namely, quadruple partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal extended quadtree partitioning, and vertical extended quadtree partitioning.

[0125] In some embodiments, the video encoder 20 may further partition the coding block of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. The PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In a monochrome picture or a picture with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PB. The video encoder 20 may generate a predicted luma block, a predicted Cb block, and a predicted Cr block for the luma PB, Cb PB, and Cr PB of each PU of the CU.

[0126] Video encoder 20 may use intra prediction or inter prediction to generate a prediction block for a PU. If video encoder 20 uses intra prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0127] After the video encoder 20 generates the predicted luma block, the predicted Cb block, and the predicted Cr block for one or more PUs of a CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predicted luma block of the CU from the original luma coding block of the CU, such that each sample in the luma residual block of the CU indicates the difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0128] In addition, if Figure 4C As shown in , the video encoder 20 can use quadtree partitioning to decompose the luma residual block, Cb residual block and Cr residual block of a CU into one or more luma transform blocks, Cb transform blocks and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples and syntax elements for transforming the transform block samples. Therefore, each TU of a CU may be associated with a luma transform block, a Cb transform block and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.

[0129] Video encoder 20 may apply one or more transforms to the luma transform block of a TU to generate a luma coefficient block for the TU. A coefficient block may be a two-dimensional array of transform coefficients. A transform coefficient may be a scalar. Video encoder 20 may apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to the Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0130] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to the process by which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream comprising a sequence of bits forming a representation of the encoded frame and associated data, which is stored in the storage device 32 or sent to the target device 14.

[0131] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 can reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient blocks associated with the TUs of the current CU to reconstruct the residual blocks associated with the TUs of the current CU. The video decoder 30 also reconstructs the coding blocks of the current CU by adding samples of the prediction blocks for the PUs of the current CU to corresponding samples of the transform blocks of the TUs of the current CU. After reconstructing the coding blocks for each CU of the frame, the video decoder 30 can reconstruct the frame.

[0132] As mentioned above, video coding mainly uses two modes: intra-frame prediction (or intra prediction) and inter-frame prediction (or inter prediction) to achieve video compression. It should be noted that IBC can be regarded as intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block based on the reference video block.

[0133] However, with ever-improving video data capture technologies and finer video block sizes for preserving details in video data, the amount of data required to represent the motion vector for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only do a group of neighboring CUs in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU (which is also called the "motion vector predictor" (MVP) of the current CU) by exploiting their spatial and temporal correlations.

[0134] Instead of combining as above Figure 2As described, the actual motion vector of the current CU determined by the motion estimation unit 42 is encoded into the video bitstream, and the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to produce a motion vector difference (MVD) for the current CU. By doing so, the motion vector determined by the motion estimation unit 42 for each CU of the frame does not need to be encoded into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.

[0135] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coding block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules for constructing a motion vector candidate list (also called a "merge list") for the current CU using those potential candidate motion vectors associated with the spatially neighboring CUs and / or temporally co-located CUs of the current CU, and then selecting one member from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, the motion vector candidate list itself does not need to be sent from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector predictor within the motion vector candidate list to encode and decode the current CU.

[0136] Generally, the basic intra prediction scheme applied in VVC remains almost the same as that of HEVC, except that several prediction tools are further extended, added and / or improved, such as extended intra prediction with wide-angle intra mode, multiple reference line (MRL) intra prediction, position-dependent intra prediction combination (PDPC), intra subpartition (ISP) prediction, cross-component linear model (CCLM) prediction, and matrix-weighted intra prediction (MIP).

[0137] Similar to HEVC, VVC uses a set of reference samples adjacent to the current CU (i.e., above or to the left of the current CU) to predict the samples of the current CU. However, in order to capture the finer edge directions present in natural video (especially high-resolution (e.g., 4K) video content), the number of angular intra modes is expanded from 33 in HEVC to 93 in VVC. Figure 1 FIG. 1 shows a schematic diagram of the intra mode defined in VVC. Figure 1 As shown, among the 93 angle intra modes, Mode 2 to Mode 66 are traditional angle intra modes, Mode -1 to Mode -14 and Mode 67 to Mode 80 are wide angle intra modes. In addition to the angle intra mode, the planar mode ( Figure 1 Mode 0) and DC mode ( Figure 1 Mode 1) in is also applied to VVC.

[0138] like Figure 4E As shown in , since the partition structure of quadtree / binarytree / ternarytree is applied in VVC, for intra prediction in VVC, in addition to square video blocks, there are also rectangular video blocks. Since a given video block has unequal width and height, various angular intra mode sets can be selected for different block shapes from 93 angular intra modes. More specifically, for square video blocks and rectangular video blocks, in addition to planar mode and DC mode, 65 angular intra modes out of 93 angular intra modes are supported for each block shape. When the rectangular block shape of the video block meets specific conditions, the index of the wide-angle intra mode of the video block can be adaptively determined by the video decoder 30 according to the index of the traditional angular intra mode received from the video encoder 20 using the mapping relationship shown in Table 1 below. That is, for non-square blocks, the wide-angle intra mode is signaled by the video encoder 20 using the index of the conventional angular intra mode, wherein the index of the conventional angular intra mode is mapped to the index of the wide-angle intra mode after being parsed by the video decoder 30, thereby ensuring that the total number (i.e., 67) of intra modes (i.e., planar mode, DC mode, and 65 of the 93 angular intra modes) remains unchanged and the intra mode encoding method remains unchanged. Therefore, good signaling efficiency of the intra mode is achieved while providing a consistent design across different block sizes.

[0139] Table 1-0 shows the mapping relationship between the index of the traditional angle intra mode and the index of the wide angle intra mode for intra prediction of different block shapes in VCC, where W represents the width of the video block and H represents the height of the video block. Table 1-0

[0140] Similar to intra prediction in HEVC, all intra modes in VVC (i.e., planar mode, DC mode, and angular intra mode) use the reference sample sets above and to the left of the current video block for intra prediction. However, unlike using only the nearest reference sample row / column (i.e., Figure 2 Unlike HEVC (row 0 201 in VVC), MRL intra prediction is introduced in VVC. In MRL intra prediction, in addition to the nearest reference sample row / column, two additional reference sample rows / columns (i.e., Figure 2 The index of the selected reference sample row / column is signaled from the video encoder 20 to the video decoder 30. When a non-nearest reference sample row / column is selected (e.g., Figure 2When the first row 203 or the third row 205 in the current CTU is selected, the planar mode is excluded from the set of intra modes that can be used to predict the current video block. For the first row / column video block in the current CTU, MRL intra prediction is disabled to prevent the use of extended reference samples outside the current CTU.

[0141] Sample Adaptive Offset (SAO)

[0142] Sample Adaptive Offset (SAO) is a process that modifies decoded samples by conditionally adding an offset value to each sample after the deblocking filter is applied, based on values in a lookup table sent by the encoder. SAO filtering is performed on a region-by-region basis, based on the filter type selected per CTB via the syntax element sao-type-idx. A sao-type-idx value of 0 indicates that the SAO filter is not applied to the CTB, while values of 1 and 2 signal the use of band offset and edge offset filter types, respectively. In band offset mode, specified by sao-type-idx equal to 1, the selected offset value depends directly on the sample amplitude. In this mode, the full sample amplitude range is evenly divided into 32 segments called bands, and the sample values belonging to four of these bands (which are contiguous within the 32 bands) are modified by adding a signaled value denoted as a band offset, which can be positive or negative. The main reason for using four contiguous bands is that the sample amplitudes in a CTB tend to be concentrated in only a few bands in smooth areas where banding artifacts can occur. Furthermore, the design choice of using four offsets is consistent with the edge offset mode of operation which also uses four offset values. In the edge offset mode specified by sao-type-idx equal to 2, the syntax element sao-eo-class has a value from 0 to 3, indicating whether the edge offset classification in the CTB uses horizontal, vertical, or one of the two diagonal gradient directions.

[0143] Figure 5A is a block diagram depicting four gradient modes used in SAO according to some embodiments of the present disclosure. The four gradient modes 502, 504, 506, and 508 are used for corresponding sao-eo-class in edge offset mode. The sample point labeled "p" indicates the center sample point to be considered. The two samples labeled "n0" and "n1" specify two adjacent sample points along the (a) horizontal (sao-eo-class=0), (b) vertical (sao-eo-class=1), (c) 135° diagonal (sao-eo-class=2), and (d) 45° (sao-eo-class=3) gradient modes. Each sample point in the CTB is classified into one of the five EdgeIdx categories by comparing the sample value p at a certain position with the values n0 and n1 of two samples at adjacent positions, such as Figure 5AThis classification is performed for each sample based on the decoded sample value, so no additional signaling is required for EdgeIdx classification. Depending on the EdgeIdx category at the sample position, for EdgeIdx categories 1 to 4, an offset value from the transmitted lookup table is added to the sample value. The offset value is always positive for categories 1 and 2, and always negative for categories 3 and 4. Therefore, the filter generally has a smoothing effect in edge offset mode. Table 1-1 below shows the EdgeIdx categories of samples in the SAO edge class. Table 1-1

[0144] For SAO types 1 and 2, a total of four amplitude offset values are sent to the decoder for each CTB. For type 1, the sign is also encoded. The offset values and related syntax elements (such as sao-type-idx and sao-eo-class) are determined by the encoder, typically using a criterion that optimizes rate-distortion performance. A merge flag can be used to indicate that the SAO parameters are inherited from the left or above CTB to make signaling efficient. In summary, SAO is a nonlinear filtering operation that allows additional refinement of the reconstructed signal and can enhance signal representation both in smooth areas and around edges.

[0145] Pre-sample adaptive offset (Pre-SAO)

[0146] In some examples, pre-sample adaptive offset (Pre-SAO) is implemented. Pre-SAO codec performance with low complexity has broad application prospects in the development of future video codec standards. In some examples, Pre-SAO is applied only to luma component samples classified using luma samples. Pre-SAO operates by applying two SAO-like filtering operations (called SAOV and SAOH) and applying them in conjunction with a deblocking filter (DBF) before applying an existing (legacy) SAO. The first SAO-like filter, SAOV, operates to apply SAO to the input picture Y2 after applying a deblocking filter (DBFV) for vertical edges.

[0147] where T is a predetermined positive constant, and d1 and d2 are offset coefficients associated with two classes based on the sample-by-sample difference between Y1(i) and Y2(i), which is given by: f(i)=Y1(i)-Y2(i).

[0148] The first category for d1 is given by taking all sample positions i such that f(i)>T, while the second category for d2 is given by f(i)<-T. The offset coefficients d1 and d2 are calculated at the encoder such that the mean square error between the output picture Y3 of the SAOV and the original picture X is minimized in the same way as in the existing SAO process. After applying SAOV, a second SAO-like filter SAOH operates by applying SAO to Y4 after SAOV has been applied, where the classification is performed based on the sample-by-sample difference between the output pictures Y3(i) and Y4(i) of the deblocking filter based on horizontal edges (DBFH), as Figure 5B As shown in Figure 2, the same process as SAOV is applied to SAOH, where Y3(i)-Y4(i) is used instead of Y1(i)-Y2(i) for its classification. Two offset coefficients, a predetermined threshold T, and an enable flag for each of SAOH and SAOV are signaled at the slice level. SAOH and SAOV are applied independently for the luma and two chroma components.

[0149] In some examples, both SAOV and SAOH operate only on the picture samples affected by the corresponding deblocking (DBFV or DBFH). Thus, unlike existing SAO processing, Pre-SAO processes only a subset of all samples in a given spatial region (a picture or a CTU in the case of traditional SAO), which keeps the resulting increase in decoder-side averaging operations per picture sample low (according to preliminary estimates, in the worst case, two or three comparisons and two additions per sample). Pre-SAO only requires the samples used by the deblocking filter and does not require storing additional samples at the decoder.

[0150] Bilateral filter (BIF)

[0151] In some embodiments, a bilateral filter (BIF) is implemented to explore compression efficiency beyond VVC. The BIF is performed in the sample adaptive offset (SAO) loop filter stage. Both the bilateral filter (BIF) and SAO use samples from deblocking as input. Each filter creates an offset for each sample, and these offsets are added to the input samples and then clipped before entering the ALF.

[0152] In detail, the output sample I OUT Obtained as I OUT =clip3(I C +ΔI BIF +ΔI SAO ), Among them I C is the input sample from deblocking, ΔI BIF is the offset from the bilateral filter, and ΔI SAO is the offset from SAO.

[0153] In some embodiments, this embodiment provides the encoder with the possibility to enable or disable filtering at CTU and slice level.The encoder makes this decision by evaluating the rate-distortion optimization (RDO) cost.

[0154] The following syntax elements are introduced in the PPS in Table 1-2 showing the picture parameter set RBSP syntax. Table 1-2

[0155] pps_bilateral_filter_enabled_flag equal to 0 specifies that the bilateral loop filter is disabled for slices referring to the PPS. pps_bilateral_filter_flag equal to 1 specifies that the bilateral loop filter is enabled for slices referencing the PPS.

[0156] bilateral_filter_strength specifies the bilateral loop filter strength value used in the bilateral transform block filtering process. The value of bilateral_iltet_strength should be in the range of 0 to 2, inclusive.

[0157] bilateral_filter_qp_offset specifies the offset used in the derivation of the bilateral filter lookup table LUT(x) for the slice referencing the PPS. bilateral_filter_qp_offset should be in the range of -12 to +12, inclusive.

[0158] The following syntax elements are introduced in Table 1-3 showing the slice header syntax and Table 1-4 showing the coding tree unit syntax. Table 1-3 Table 1-4

[0159] The semantics are as follows: slice_bilateral_filter_all_ctb_enabled_flag equal to 1 specifies that the bilateral filter is enabled and applied to all CTBs in the current slice. When slice_bilateral_filter_all_ctb_enabled_flag is not present, it is inferred to be equal to 0.

[0160] slice_bilateral_filter_enabled_flag equal to 1 specifies that the bilateral filter is enabled and can be applied to the CTBs of the current slice. When slice_bilateral_filter_enabled_flag is not present, it is inferred to be equal to slice_bilateral_filter_all_ctb_enabled_flag.

[0161] When bilateral_filter_ctb_flag[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 1, it specifies that the bilateral filter is applied to the luma coding tree block of the coding tree unit at the luma position (xCtb, yCtb). When bilateral_filter_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, it specifies that the bilateral filter is not applied to the luma coding tree block of the coding tree unit at the luma position (xCtb, yCtb). When bilateral_filter_ctb_flag is not present, it is inferred to be equal to (slice_bilateral_filter_all_ctb_enabled_flag & slice_bilateral_filter_enabled_flag).

[0162] In some examples, for a filtered CTU, the filtering process proceeds as follows. At picture boundaries where samples are unavailable, the bilateral filter uses expansion (sample repetition) to fill in the unavailable samples. For virtual boundaries, the behavior is the same as for SAO, i.e., no filtering occurs. When crossing horizontal CTU boundaries, the bilateral filter can access the same samples as SAO is accessing. Figure 7 is a block diagram depicting a naming convention for sample points around a center sample point according to some embodiments of the present disclosure. As an example, if the center sample point I C Located in the top row of CTU, read I from the top of CTU NW , I A and I NE , just like SAO does, but filled with I AA , so no additional line buffer is required. Center sample point I C The surrounding samples are based on Figure 7 is represented by A, B, L, and R, where A, B, L, and R stand for up, down, left, and right, and where NW, NE, SW, SE stand for northwest, etc. Similarly, AA stands for up-up, BB stands for down-down, etc. This diamond shape is different from another method that uses square filter support and does not use I AA , IBB , I LL or I RR .

[0163] Each surrounding sample point I A , I R etc. will contribute the corresponding modification value Etc. These values are calculated as follows: From the right sample point I R Starting with the contribution of , the difference is calculated as: ΔI R =(|I R -I C |+4)>>3, Where |·| represents the absolute value. For data other than 10 bits, use ΔI R =(|I R -I C |+2 n-6 )>>(n-7) instead, where n=8 for 8-bit data, etc. The result value is now clipped Cut it so that it is less than 16: sI R =min(15,ΔI R ).

[0164] The modified value is now calculated as Among them LUT ROW [] is an array of 16 values determined by the value of qpb=clip(0,25,QP+bilateral_filter_qp_offset-17): {0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,},if qpb=0 {0,1,1,1,1,0,0,0,0,0,0,0,0,0,0,0,},if qpb=1 {0,2,2,2,1,1,0,1,0,0,0,0,0,0,0,0,},if qpb=2 {0,2,2,2,2,1,1,1,1,1,1,1,0,1,1,-1,},if qpb=3 {0,3,3,3,2,2,1,2,1,1,1,1,0,1,1,-1,},if qpb=4 {0,4,4,4,3,2,1,2,1,1,1,1,0,1,1,-1,},if qpb=5 {0,5,5,5,4,3,2,2,2,2,2,1,0,1,1,-1,},if qpb=6 {0,6,7,7,5,3,3,3,3,2,2,1,1,1,1,-1,},if qpb=7 {0,6,8,8,5,4,3,3,3,3,3,2,1,2,2,-2,},if qpb=8 {0,7,10,10,6,4,4,4,4,3,3,2,2,2,2,-2,},if qpb=9 {0,8,11,11,7,5,5,4,5,4,4,2,2,2,2,-2,},if qpb=10 {0,8,12,13,10,8,8,6,6,6,5,3,3,3,3,-2,},if qpb=11 {0,8,13,14,13,12,11,8,8,7,7,5,5,4,4,-2,},if qpb=12 {0,9,14,16,16,15,14,11,9,9,8,6,6,5,6,-3,},if qpb=13 {0,9,15,17,19,19,17,13,11,10,10,8,8,6,7,-3,},if qpb=14 {0,9,16,19,22,22,20,15,12,12,11,9,9,7,8,-3,},if qpb=15 {0,10,17,21,24,25,24,20,18,17,15,12,11,9,9,-3,},if qpb=16 {0,10,18,23,26,28,28,25,23,22,18,14,13,11,11,-3,},if qpb=17 {0,11,19,24,29,30,32,30,29,26,22,17,15,13,12,-3,},if qpb=18 {0,11,20,26,31,33,36,35,34,31,25,19,17,15,14,-3,},if qpb=19 {0,12,21,28,33,36,40,40,40,36,29,22,19,17,15,-3,},if qpb=20 {0,13,21,29,34,37,41,41,41,38,32,23,20,17,15,-3,},if qpb=21 {0,14,22,30,35,38,42,42,42,39,34,24,20,17,15,-3,},if qpb=22 {0,15,22,31,35,39,42,42,43,41,37,25,21,17,15,-3,},if qpb=23 {0,16,23,32,36,40,43,43,44,42,39,26,21,17,15,-3,},if qpb=24 {0,17,23,33,37,41,44,44,45,44,42,27,22,17,15,-3,},if qpb=25

[0165] These values can be stored using six bits per entry, resulting in 26*16*6 / 8 = 312 bytes, or 300 bytes (if the first row of all zeros is excluded). Modify Value and In the same way from I L , I A and I B Calculation. For diagonal sample point I NW , I NE , I SE , I SW , and the sample point I two steps away AA , I BB , I RR and I LL , the calculation also follows equations 2 and 3, but the values used are shifted by 1. SE For example, And the other diagonal samples and samples beyond two steps are calculated similarly.

[0166] Modified values are summed together

[0167] In some examples, for the previous sample point, equal Similarly, for the above sample points, equal And similar symmetries can be found for the modified values on the diagonal and two steps away. This means that in hardware implementation, the calculation and These six values are sufficient, and the remaining six values can be obtained from previously calculated values.

[0168] m sumThe value is now multiplied by c = 1, 2 or 3, which can be done using a single adder and a logic AND gate in the following way: c v =k1&(m sum <<1)+k2&m sum , Where & represents logical AND, k1 is the most significant bit of the multiplier c, and k2 is the least significant bit. The value to be multiplied is obtained using the minimum block size D = min (width, height), as shown in Table 1-5, which shows that the c parameter is obtained from the minimum block size D = min (width, height). Table 1-5 Block Type D≤4 4<D<16 D≥16 Intraframe 3 2 1 Interframe 2 2 1

[0169] Finally, calculate the bilateral filter offset ΔI BIF For full-intensity filtering, use the following: ΔI BIF =(c v +16)>>5, And for half-intensity filtering, use the following: ΔI BIF =(c v +32)>>6.

[0170] The general formula for n-bit data is to use r add =2 14-n-bilatera_filter_strength r shift =15-n-bilateal_filter_strength ΔI BIF =(c v +r add )>>r shi , The bilateral_filter_strength can be 0 or 1 and is signaled in the pps.

[0171] Adaptive loop filter (ALF)

[0172] In VVC, an adaptive loop filter (ALF) with block-based filter adaptation is applied. For the luma component, one of 25 filters is selected for each 4×4 block based on the direction and activity of the local gradient.

[0173] Using two diamond filter shapes (such as Figures 8A to 8B ). A 7×7 diamond shape is applied to the luma component, and a 5×5 diamond shape is applied to the chroma components.

[0174] For the luminance component, each 4×4 block is classified into one of 25 categories. The category index C is a quantized value based on its directionality D and activity. The exported ones are as follows:

[0175] To calculate D and First, use the one-dimensional Laplace method to calculate the gradient in the horizontal, vertical and two diagonal directions: where indices i and j refer to the coordinates of the top left sample point within the 4×4 block, and R(i,j) indicates the coordinate (i,j) The reconstructed sample points at .

[0176] In order to reduce the complexity of block classification, a sub-sampled one-dimensional Laplace calculation is applied. 9A to 9D As shown, the gradient calculations in all directions use the same subsampling position.

[0177] Then set the maximum and minimum values of the horizontal and vertical gradients D to: The maximum and minimum values of the gradients in the two diagonal directions are set to:

[0178] To derive the value of the directionality D, these values are compared with each other and with two thresholds t1 and t2: Step 1. If and If both are true, set D to 0. Step 2. If Then continue with step 3; otherwise, continue with step 4. Step 3. If Then set D to 2; otherwise, set D to 1. Step 4. If Then set D to 4; otherwise, set D to 3. The activity value A is calculated as: A is further quantized to the range of 0 to 4 (inclusive), and the quantized value is expressed as

[0179] For the chrominance components of the image, no classification method is applied.

[0180] Geometric transformation of filter coefficients and clipping values

[0181] Before filtering each 4×4 luma block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f(k,l) and the corresponding filter clipping values c(k,l), depending on the gradient values calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make different blocks to which the ALF is applied more similar by aligning their directionality.

[0182] The following three geometric transformations are introduced, including diagonal, vertical flip and rotation: Diagonal: f D (k,l)=f(l,k),c D (k,l)=c(l,k), Flip vertically: f V (k,l)=f(k,Kl-1),c V (k, l) = c(k, Kl-1) Rotation: f R (k,l)=f(Kl-1,k),c R (k,l)=c(Kl-1,k)

[0183] Where K is the size of the filter, and 0≤k, l≤K-1 are the coefficient coordinates, so that the position (0,0) is in the upper left corner and the position (K-1,K-1) is in the lower right corner. These transforms are applied to the filter coefficients f(k,l) and the clipping value c(k,l) depending on the gradient value calculated for the block. The relationship between the transforms and the four gradients in the four directions is summarized in Tables 1-6 below, which show the mapping of the gradient values calculated for a block to these transforms. Table 1-6 Gradient value Transform <![CDATA[g d2 <g d1 And g h <g v ]]> No transformation <![CDATA[g d2 <g d1 And g v <g h ]]> diagonal <![CDATA[g d1 <g d2 And g h <g v ]]> Vertical Flip <![CDATA[g d1 <g d2 And g v <g h ]]> Rotation

[0184] Filtering

[0185] At the decoder side, when ALF is enabled for CTB, each sample point R(i,j) in the CU is filtered to obtain the sample value R′(i,j) as shown below, Where f(k,l) represents the decoded filter coefficients, K(x,y) is the clipping function, and c(k,l) represents the decoded clipping parameters. and where L represents the filter length. The clipping function K(x,y)=min(y,max(-y,x)) corresponds to the function Clip3(-y,y,x). Clipping introduces nonlinearity to make the ALF more efficient by reducing the influence of neighboring sample values that differ significantly from the current sample value.

[0186] Cross-Component Adaptive Loop Filter (CC-ALF)

[0187] CC-ALF uses luma sample values to refine each chroma component by applying an adaptive linear filter to the luma channel and then using the output of this filtering operation for chroma refinement. Figure 10A A system-level diagram of CC-ALF processing relative to SAO processing, luma ALF processing, and chroma ALF processing is provided.

[0188] The filtering in CC-ALF is achieved by applying a linear diamond filter ( Figure 10B ) is applied to the luminance channel. A filter is used for each chrominance channel, and the operation is expressed as Among them, (x, y) is the position of the chrominance component i to be refined, (x Y ,y Y ) is the brightness position based on (x,y), S i is the filter support area in the luminance component, c i (s0,y0) represents the filter coefficient.

[0189] like Figure 10B As shown, the luma filter support is the area that is co-located with the current chroma sample after taking into account the spatial scaling factor between the luma plane and the chroma plane.

[0190] In the VVC reference software, CC-ALF filter coefficients are calculated by minimizing the mean squared error (MSE) of each chroma channel relative to the original chroma content. To achieve this, the VTM algorithm uses a coefficient derivation process similar to that used for chroma ALF. Specifically, a correlation matrix is derived, and coefficients are calculated using a Cholesky decomposition solver to attempt to minimize the mean squared error metric. During filter design, up to eight CC-ALF filters can be designed and transmitted per picture. The resulting filters are then indicated for each of the two chroma channels on a CTU-by-CTU basis.

[0191] Additional features of CC-ALF include: The design uses a 3×4 diamond with 8 taps. Seven filter coefficients are transmitted in APS; Each transmitted coefficient has a dynamic range of 6 bits and is restricted to values that are powers of 2; Deriving the eighth filter coefficient at the decoder such that the sum of the filter coefficients is equal to 0; APS can be referenced in the stripe header; Control CC-ALF filter selection at the CTU level for each chroma component; Border fill for horizontal virtual borders uses the same memory access pattern as luma ALF.

[0192] As an additional feature, the reference encoder can be configured to enable some basic subjective adjustments via a profile. When enabled, VTM reduces the application of CC-ALF in areas encoded at high QP and close to mid-gray or containing a lot of luminance high frequencies. Algorithmically, this is achieved by disabling the application of CC-ALF in a CTU if any of the following conditions are met: The slice QP value minus 1 is less than or equal to the base QP value; The number of chroma samples with a local contrast greater than (1<<(bitDepth–2))–1 exceeds the CTU height, where the local contrast is the difference between the maximum and minimum luminance sample values within the filter support area; More than a quarter of the chroma samples are in the range between (1<<(bitDepth–1))–16 and (1<<(bitDepth–1))+16.

[0193] The motivation for this feature is to provide some guarantee that CC-ALF does not amplify artifacts introduced earlier in the decoding path (largely due to the fact that VTM is not currently explicitly optimized for chroma subjective quality). It is expected that alternative encoder implementations will either not use this feature or incorporate alternative strategies appropriate to their coding characteristics.

[0194] Filter parameter signaling

[0195] ALF filter parameters are signaled in an Adaptive Parameter Set (APS). In one APS, up to 25 sets of luma filter coefficients and clipping value indices, and up to 8 sets of chroma filter coefficients and clipping value indices can be signaled. To reduce bit overhead, filter coefficients for different classifications of luma components can be merged. The index of the APS for the current slice is signaled in the slice header.

[0196] The clipping value index decoded from the APS allows the determination of the clipping value using a table of clipping values for both luma and chroma components. These clipping values depend on the internal bit depth. More precisely, the clipping value is obtained by the following formula: AlfClip={round(2 B-α*n )forn∈[0..N-1]} Where B is equal to the internal bit depth, α is a predefined constant value equal to 2.35, and N is equal to 4, which is the number of cropping values allowed in VVC. AlfClip is then rounded to the nearest value using the power of 2 format.

[0197] In the slice header, up to seven APS indices can be signaled to specify the luma filter set for the current slice. Filtering can be further controlled at the CTB level. A flag is always signaled to indicate whether the ALF is applied to the luma CTB. The luma CTB can select a filter set from 16 fixed filter sets and a filter set from the APS. A filter set index is signaled for the luma CTB to indicate which filter set to apply. The 16 fixed filter sets are predefined and hard-coded in both the encoder and decoder.

[0198] For chroma components, the APS index is signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if there is more than one chroma filter set in the APS, the filter index is signaled for each chroma CTB.

[0199] The filter coefficients are quantized with a norm equal to 128. To limit the multiplication complexity, bitstream consistency is applied so that the coefficient values in non-central positions are within -2 7 to 2 7 The center position coefficient is not signaled in the bitstream and is considered equal to 128.

[0200] Virtual boundary filtering for line buffer reduction

[0201] In VVC, in order to reduce the line buffer requirements of ALF, a modified block classification and filtering is applied to samples near the horizontal CTU boundary. For this purpose, the virtual boundary is defined as follows Figure 11 The horizontal CTU boundary is shifted by "N" rows of samples as shown, where N is equal to 4 for the luma component and 2 for the chroma components.

[0202] The modified block classification is applied to the luminance component, e.g. Figure 11 As shown in Figure 2 . For the 1D Laplacian gradient calculation of the 4×4 block above the virtual boundary, only the samples above the virtual boundary are used. Similarly, for the 1D Laplacian gradient calculation of the 4×4 block below the virtual boundary, only the samples below the virtual boundary are used. By taking into account the reduced number of samples used in the 1D Laplacian gradient calculation, the quantization of the activity value A is thus scaled.

[0203] For the filtering process, symmetric padding operations at the virtual boundaries are used for both the luma and chroma components. Figure 12As shown in FIG, when a filtered sample is below a virtual boundary, the adjacent sample above the virtual boundary is filled in. At the same time, the corresponding sample on the other side is also filled in symmetrically.

[0204] Unlike the symmetrical padding used at horizontal CTU boundaries, a simple padding process is applied to slice, tile, and sub-picture boundaries when filters across the boundaries are disabled. A simple padding process is also applied at picture boundaries. The padded samples are used for both classification and filtering. To compensate for extreme padding when the filtered samples are directly above or below the virtual boundary, the filter strengths for both luma and chroma are reduced for these cases by increasing the right shift by 3 in the equation that obtains the sample value R′(i,j).

[0205] For existing SAO designs in the HEVC, VVC, AVS2, and AVS3 standards, luma Y, chroma Cb, and chroma Cr sample offset values are determined independently. That is, for example, the current chroma sample offset is determined only by the current and neighboring chroma sample values, without considering co-located or neighboring luma samples. However, luma samples retain more original picture detail information than chroma samples, and they can benefit the determination of the current chroma sample offset. In addition, since chroma samples typically lose high-frequency details after color conversion from RGB to YCbCr or after quantization and deblocking filtering, introducing luma samples that preserve high-frequency details for chroma offset determination can benefit chroma sample reconstruction. Therefore, further benefits can be expected by exploiting cross-component correlations (e.g., by using methods and systems for cross-component sample adaptive offset (CCSAO)). In some embodiments, the correlation here includes not only cross-component sample values, but also picture / codec information such as prediction / residual codec modes, transform types, and quantization / deblocking / SAO / ALF parameters from across components.

[0206] Another example is that for SAO, the luma sample offset is determined only by the luma samples. However, for example, luma samples with the same band offset (BO) classification can be further classified according to their co-located and adjacent chroma samples, which may lead to more efficient classification. SAO classification can be used as a shortcut to compensate for the sample difference between the original image and the reconstructed image. Therefore, an efficient classification is desired.

[0207] Cross-Component Sample Adaptive Offset (CCSAO)

[0208] The existing SAO designs in the HEVC, VVC, AVS2 and AVS3 standards are used as the basic SAO method in the following description. For those skilled in the art of video coding and decoding, the proposed cross-component method described in this disclosure can also be applied to other loop filter designs or other codec tools with similar design spirit. For example, in the AVS3 standard, SAO is replaced by a codec tool called Enhanced Sample Adaptive Offset (ESAO), however, the proposed CCSAO can also be applied in parallel with ESAO. In another example, CCSAO can be applied in parallel with the Constrained Directional Enhancement Filter (CDEF) in the AV1 standard.

[0209] 13A to 13F The figure shows the proposed method. Figure 13A In

[13] , the luma samples after the luma deblocking filter (DBF Y) are used to determine the additional offsets for the chroma Cb and Cr after SAO Cb and SAO Cr. For example, the current chroma sample (1302) is first classified using the co-located luma sample (1304) and the neighboring luma samples (1306), and the CCSAO offset of the corresponding category is added to the current chroma sample. Figure 13B In , CCSAO is applied to luma samples and chroma samples, and uses DBF Y / Cb / Cr as input. Figure 13C In , CCSAO can work independently. Figure 13D In

[15] , CCSAO can be applied recursively (2 or N times) with the same or different offsets in the same codec level or repeatedly in different levels. Figure 13E In , CCSAO is applied in parallel with SAO and BIF. Figure 13F In

[15] , CCSAO replaces SAO and is applied in parallel with BIF.

[0210] Therefore, in order to classify the current luma sample, the information of the current and adjacent luma samples, the co-located and adjacent chroma samples (Cb and Cr) can be used. In addition, in order to classify the current chroma sample (Cb or Cr), the information of the co-located and adjacent luma samples, the co-located and adjacent cross chroma samples, and the current and adjacent chroma samples can be used.

[0211] Figure 14 It shows that CCSAO can also be applied in parallel with other codec tools, such as ESAO in the AVS standard or CDEF in the AV1 standard. Figure 15A It shows that the position of CCSAO can be after SAO, that is, the position of CCALF in the VVC standard. Figure 15B In , CCSAO can work independently without CCALF. Figure 15CIn

[15] , CCSAO can be used as a post-reconstruction filter, i.e., the reconstructed samples are used as input for classification, compensating the luma / chroma samples before entering the neighboring intra prediction. Figure 16 It is shown that CCSAO can also be applied in parallel with CCALF. Figure 16 In , the positions of CCALF and CCSAO can be switched. Figures 13A to 16 In the above or other sections of this disclosure, SAO Y / Cb / Cr blocks can be replaced by ESAO Y / Cb / Cr (in AVS3) or by CDEF (in AV1). Note that Y / Cb / Cr can also be represented as Y / U / V in the video codec area.

[0212] In some examples, if the video is in RGB format, the proposed CCSAO can also be applied by simply mapping the YUV notation to GBR in the following paragraphs.

[0213] Note that the drawings in this disclosure can be combined with all examples mentioned in this disclosure.

[0214] Classification

[0215] 13A to 13F and Figure 19 The input to the CCSAO classification is shown. 13A to 13F and Figure 19 It is also shown that all co-located and adjacent luma samples / chroma samples can be fed into CCSAO classification. Please note that the classifier mentioned in this disclosure can be used not only for cross-component classification (e.g., using luma to classify chroma and vice versa), but also for single-component classification (e.g., using luma to classify luma or using chroma to classify chroma), because the newly proposed classifier in this disclosure can also benefit the original SAO classification.

[0216] The classifier example (C0) uses the same luminance or chrominance sample value ( Figure 13A Y0)( Figure 13B-13C Let band_num be the number of equal bands of luminance or chrominance dynamic range, bit_depth be the sequence bit depth, and an example of the class index of the current chrominance sample is Class(C0)=(Y0*band_num)>>bit_depth

[0217] Table 2-2 below lists some examples of band_num and bit_depth. Table 2-2 shows three classification examples when the number of bands is different for each classification example. The classification may take into account rounding. Class(C0)=((Y0*band_num)+(1<<bit_depth))> >bit_depth Some band_num and bit_depth examples are listed in Table 2-1 below. Table 2-1

[0218] In some examples, the classifier uses different luminance (or chrominance) sample locations for C0 classification. For example, using the adjacent Y7 instead of Y0 for C0 classification, such as Figure 17 As shown in the figure. Different classifiers can be switched at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, Figure 17 In the example, Y0 is used for POC0 and Y7 is used for POC1, as shown in Table 2-2 below. Table 2-2

[0219] Figures 18A to 18G Some examples of brightness candidates of different shapes are shown. A constraint can be applied to the shape: the total number of candidates must be a power of 2, such as 18B to 18D As shown in . Constraints can be applied to the shape: the number of luma candidates must be horizontally and vertically symmetric to the number of chroma samples, as Figure 18A 、 Figures 18C to 18E As shown in . The power-of-2 constraint and symmetry constraint can also be applied to chrominance candidates. Figures 13B to 13C In , the U / V part shows an example of symmetry constraint.

[0220] In some examples, different color formats may have different classifier "constraints." For example, 420 uses Figures 13B to 13C Luma / chroma candidate selection (selecting a candidate from a 3x3 shape), but 444 uses Figure 18F To perform luminance and chrominance candidate selection, 422 uses Figure 18G For luma (2 chroma samples share 4 luma candidates), and Figure 18F For chroma candidates.

[0221] The C0 position and C0 band_num can be combined and switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different combinations can result in different classifiers, as shown in Table 2-3 below. Table 2-3

[0222] In some examples, the co-located luminance sample value (Y0) is replaced with a value (Yp) by weighting the co-located and neighboring luminance samples. FIG. 20A to FIG. 20B Two examples are shown. Different Yps can be different classifiers. Different Yps can be applied to different chroma formats. For example, Figure 20A the Yp in Figure 20B is used for the 420 case,

[0223] the Yp in is used for the 422 case, and Y0 is used for the 444 case. In some examples, another classifier example (C1) is the comparison score [-8, 8] of the co-located luminance sample (Y0) and the neighboring 8 luminance samples, which altogether produces 17 classes. Initial class (C1) = 0, loop over the neighboring 8 luminance samples (Yi, i = 1 to 8)

[0224] In some examples, the C1 example is equal to the following function with a threshold th of 0. ClassIdx = Index2ClassTable(f(C, P1)+f(C, P2)+…+f(C, P8)) [[ID=2X]] f(x, y) = 1 if x - y > th; f(x, y) = 0 if x - y = th; f(x, y) = -1 if x - y < th

[0225] In some examples, similar to the C4 classifier, one or more thresholds can be predefined (e.g., saved in a LUT), or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to assist in classifying (quantifying) the differences.

[0226] In some examples, the variant (C1’) only calculates the comparison score [0, 8], which produces 8 classes. (C1, C1’) is a classifier group and the PH / SH level flag can be signaled to switch between C1 and C1’. Initial class (C1’) = 0, loop over the neighboring 8 luminance samples (Yi, i = 1 to 8) If Y0 > Yi, then class += 1

[0227] In some examples, a variant (C1s) is to selectively use N neighboring samples out of M neighboring samples to calculate the comparison score. An M-bit bit mask can be signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to indicate which neighboring samples are selected for statistical comparison scores. Figure 13B As an example of a luma classifier: 8 adjacent luma samples are candidates, and an 8-bit bitmask (01111110) is signaled at PH, indicating that 6 samples of Y1 to Y6 are selected, so the comparison scores are in [-6, 6], which produces 13 offsets. The selective classifier C1s provides the encoder with more options to trade off the offset signaling overhead with the classification granularity.

[0228] In some examples, similar to C1s, the variant (C1's) only computes comparison scores [0,+N], the previous bit mask 01111110 example gives comparison scores in [0,6], which produces 7 offsets.

[0229] Different classifiers can be combined to produce a general classifier. For example, for different pictures, for different pictures (different POC values), different classifiers are applied, as shown in Table 2-4 below. Table 2-4

[0230] In some examples, another classifier instance (C2) uses the difference (Yn) of co-located and neighboring luminance samples. Figures 21A to 21B An example of Yn whose dynamic range is [-1024, 1023] when the bit depth is 10 is shown. Let C2 band_num be the number of equally divided bands of the dynamic range of Yn, Class(C2)=(Yn+(1<<bit_depth)*band_num)> >(bit_depth+1) C0 and C2 can be combined to produce a general classifier. For example, as shown in Table 2-5 below: Table 2-5

[0231] In some examples, another classifier example (C3) uses a bit mask for classification, as shown in Table 2-6. Table 2-6 shows an example of a classifier that uses a bit mask for classification (the bit mask positions are underlined). A 10-bit bit mask is signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to indicate the classifier. For example, the bit mask 11 1100 0000 means that for a given 10-bit luma sample value, only the MSB 4 bits are used for classification, and this yields a total of 16 categories. Another exemplary bit mask 10 0100 0001 means that only 3 bits are used for classification, and this yields a total of 8 categories. The bit mask length (N) can be fixed or can be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, for a 10-bit sequence, a 4-bit bitmask 1110 is signaled in the picture PH, and the MSB 3 bits b9, b8, and b7 are used for classification. Another example is a 4-bit bitmask 0011 on the LSB. b0 and b1 are used for classification. The bitmask classifier can be applied to luma or chroma classification. Whether the MSB or LSB is used for the bitmask N can be fixed or can be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level.

[0232] In some examples, the luma position and C3 bitmask can be combined and switched at SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. Different combinations can be different classifiers.

[0233] In some examples, a bit mask limit of "maximum number of 1s" can be applied to limit the corresponding number of offsets. For example, limiting the "maximum number of 1s" of the bit mask to 4 in the SPS will result in a maximum offset of 16 in the sequence. The bit mask in different POCs can be different, but the "maximum number of 1s" should not exceed 4 (the total number of categories should not exceed 16). The value of "maximum number of 1s" can be signaled and switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Table 2-6

[0234] As in Figure 19In the example above, other cross-component chroma samples can also be fed into the CCSAO classification. The classifier for the cross-component chroma samples can be the same as the luma cross-component classifier, or have its own classifier as mentioned in this disclosure. The two classifiers can be combined to form a joint classifier to classify the current chroma sample. For example, combining the joint classifier for cross-component luma and chroma samples produces a total of 16 categories, as shown in Table 2-7 below. Table 2-7 shows an example of a classifier using a joint classifier that combines cross-component luma and chroma samples (the bit mask positions are underlined). Table 2-7

[0235] All the above-mentioned classifications (C0, C1, C1', C2, C3) can be combined. For example, see Table 2-8 below. Table 2-8 shows that different classifiers are combined. Table 2-8

[0236] In some examples, another classifier example (C4) uses the difference between the CCSAO input and the sample value to be compensated for classification. For example, if CCSAO is applied to the ALF stage, the difference between the sample value before and after ALF of the current component is used for classification. One or more thresholds can be predefined (e.g., maintained in the LUT) or sent by signal at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to help classify (quantize) the difference. The C4 classifier can be combined with C0Y / U / VbandNum to form a joint classifier (e.g., as shown in the POC1 example in Table 2-9). Table 2-9 shows a classifier example that uses the difference between the CCSAO input value and the sample value to be compensated for classification. Table 2-9

[0237] In some embodiments, the classifier example (C5) uses "codec information" to help sub-block classification, because different codec modes can introduce different distortion statistics in the reconstructed image. CCSAO samples are classified according to their previous codec information, and the combination of codec information can form a classifier, for example, as shown in Table 2-10 below. Figure 39 Another example of different stages of C5 codec information is shown. Table 2-10 shows that CCSAO samples are classified according to their previous codec information, and the combination of codec information can form a classifier. Table 2-10

[0238] In some examples, the classifier example (C6) uses YUV color conversion values for classification. For example, to classify the current Y component, 1 / 1 / 1 co-located or adjacent Y / U / V samples are selected to convert their colors to RGB, and the R value is quantized to the current Y component classifier using C3bandNum.

[0239] In some examples, the classifier example (C7) can be used as a generalized version of C0 / C3 and C6. To derive the current component C0 / C3 bandNum classification, all 3 color components are used. For example, to classify the current U sample, the co-located and adjacent Y / V samples, as Figure 13B As shown in , using the current and neighboring U samples, it can be formulated as Where S is the intermediate sample point to be used for C0 / C3 bandNum classification, R ij is the jth co-located / adjacent / current sample of the i-th component, where the i-th component can be a Y / U / V component, c ij is a weighting coefficient, which can be predefined or signaled at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.

[0240] In some embodiments, a special subset case of C7 can use only 1 / 1 / 1 co-located or adjacent Y / U / V samples to derive the intermediate sample S, which can also be regarded as a special case of C6 (using 3-component color transform). S can be further fed into the C0 / C3 bandNum classifier. classIdx=bandS=(S*bandNumS)>>BitDepth;

[0241] In some embodiments, similar to the C0 / C3 bandNum classifiers, C7 can also be combined with other classifiers to form a joint classifier. In some examples, C7 can be different from the following example of jointly using co-located and adjacent Y / U / V samples for classification (joint bandNum classification for each Y / U / V component).

[0242] In some embodiments, a constraint may be applied: c ij The sum = 1, to reduce c ij Signaling overhead and limit the value of S to the bit depth range. For example, let c00 = (1 – other c ij The sum of ). Which c ij(c00 in this example) is forced (derived from other coefficients) and can be predefined or signaled at SPS / APS / PPS / PH / SH region / CTU / CU / sub-block / sample level.

[0243] In some embodiments, another classifier example (C8) uses cross-component / current component spatial activity information as a classifier. Similar to the block activity classifier described above, a sample at (k, l) can obtain the sample activity by the following operation: (1) Calculate N directional gradients (Laplacian or forward / backward) (2) Add N directional gradients to obtain the activity A (3) Quantize (or map) A to obtain the category index

[0244] In some embodiments, for example, the two-directional Laplace gradient for obtaining A and the two-directional Laplace gradient for obtaining The predefined mapping {Q n} g v =V k,l =|2R(k,l)-R(k,l-1)-R(k,l+1)| g h =H k,l =|2R(k,l)-R(k-1,l)-R(k+1,l)| A=(V k,l +H k,l )>>(BD-6) where (BD-6), or denoted as B, is a predefined normalization term associated with the bit depth.

[0245] In some embodiments, A can then be mapped to the range [0, 4]: {Q n}={0,1,2,2,2,2,2,3,3,3,3,3,3,3,3,4}} B and Qn can be predefined at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level or sent through signals.

[0246] In some embodiments, another classifier example (C9) can use spatial gradient information across components / current components as a classifier. Similar to the block gradient classifier described above, a sample point at (k, l) can be obtained by the following operation: (1) Calculate N directional gradients (Laplacian or forward / backward); (2) Calculate the maximum and minimum values of the gradients in M grouping directions (M <= N); (3) By comparing N values with each other and with m thresholds t1 to t m Compare to calculate the directionality D; (4) Apply a geometric transformation based on the relative gradient magnitude (optional).

[0247] For example, as an ALF block classifier, but applied at the sample level for sample classification, (1) Calculate the four directional gradients (Laplace) (2) Calculate the maximum and minimum values of the gradients in the two grouping directions (H / V and D / A) (3) By comparing N values with each other and with two thresholds t1 to t m Compare this to calculate the directivity D: (4) Apply geometric transformation according to the relative gradient size as shown in Table 1-2.

[0248] In some examples, C8 and C9 can be combined to form a joint classifier.

[0249] In some examples, another classifier example (C10) can use the edge information of cross-component / current component for current component classification. By extending the original SAO classifier, C10 can more efficiently extract the edge information of cross-component / current component in the following way: (1) selecting a direction to calculate two edge strengths, where one direction is formed by the current sample point and two adjacent sample points, and where an edge strength is calculated by subtracting the current sample point from one adjacent sample point; (2) Quantize each edge strength into M segments using M-1 thresholds Ti; (3) Use M*M categories to classify the current component sample.

[0250] FIG. 22A to FIG. 22B An example of using edge information of a cross component / current component for current component classification according to some embodiments of the present disclosure is shown. The current sample point is represented by c, and the two neighboring sample points of the current component / cross component are represented by a and b. In this example, (1) Select a diagonal direction from the four direction candidates. The differences (ca) and (cb) are two edge intensities in the range of -1023 to 1023 (for example, for a 10b sequence); (2) quantize each edge strength into 4 segments using a common threshold [-T, 0, T]; (3) Use 16 categories to classify the current component sample points.

[0251] like FIG. 22A to FIG. 22BAs shown, a diagonal direction is selected and the differences (ca) and (cb) are quantized into 4 and 4 segments with thresholds [-T, 0, T], which form 16 edge segments. The position of (a, b) can be indicated by signaling 2 syntaxes edgeDir and edgeStep.

[0252] In some examples, the directional mode can be 0, 45, 90, 135 degrees (45 degrees between directions), or extending to 22.5 degrees between directions, or a predefined set of directions, or sent through signals at SPS / APS / PPS / PH / SH / region (set) / CTU / CU / sub-block / sample level.

[0253] In some examples, the edge strength can also be defined as (ba), which simplifies the calculation but sacrifices accuracy.

[0254] In some examples, the M-1 threshold may be signaled or predefined at SPS / APS / PPS / PH / SH / region(set) / CTU / CU / sub-block / sample level.

[0255] In some examples, the M-1 threshold can be used for different sets of edge strength calculations, for example, different sets of (ca) and (cb). If different sets are used, the total number of categories may be different. For example, when [-T, 0, T] is used to calculate (ca) and [-T, T] is used to calculate (cb), the total number of categories is 4*3.

[0256] In some examples, the M-1 threshold can use a "symmetric" property to reduce signaling overhead. For example, the predefined pattern [-T, 0, T] can be used, but not [T0, T1, T2], which requires signaling three thresholds. Another example is [-T, T].

[0257] In some examples, the thresholds may only include values that are powers of 2, which not only efficiently obtains the edge intensity distribution but also reduces the comparison complexity (only the MSB N bits need to be compared).

[0258] In some examples, the positions of a and b can be indicated by signaling two syntaxes: (1) edgeDir indicating the selected direction, and (2) edgeStep indicating the distance between the sample points used to calculate the edge strength, such as FIG. 22A to FIG. 22B shown.

[0259] In some examples, edgeDir / edgeStep may be signaled or predefined at SPS / APS / PPS / PH / SH / region(set) / CTU / CU / subblock / sample level.

[0260] In some examples, edgeDir / edgeStep can be encoded using fixed length coding (FLC) or other methods, such as truncated unary code (TU), exponential golomb code of order k (EGk), signed EGO (SVLC), or unsigned EGO (UVLC).

[0261] In some examples, C10 can be combined with bandNumY / U / V or other classifiers to form a joint classifier. For example, combining 16 edge strengths with up to 4 bandNumY bands produces 64 classes.

[0262] Another classifier example (C11) is to use the residual samples of the cross-component / current component as a classifier. C11 is performed in the same way as the other disclosed classifiers in this disclosure, but uses the residual Y / U / V samples as input (instead of the reconstructed Y / U / V samples). The classifiers in this disclosure can be reused as C11 classifiers by simply changing the input to the residual Y / U / V samples, as well as their variants, syntax indicators, APS classifier indicator storage, combination with other classifiers, etc.

[0263] For example, a Y / U / V component sample is classified using a band segment of a Y residual sample, similar to C0. band_num can be switched at SPS / APS / PPS / PH / SH / region (set) / CTU / CU / subblock / sample level. Class(C11)=(Y0*band_num)>>bit_depth, Where Y0 is the co-located / current Y residual sample point

[0264] In some examples, the residual samples may also be (1) residual samples after chroma scaling in LMCS (which are added to the prediction samples to generate the reconstructed samples); or (2) residual samples before chroma scaling in LMCS (residual values present in the syntax). In some other examples, (1) and (2) may be switched at the SPS / APS / PPS / PH / SH / region (set) / CTU / CU / subblock / sample level.

[0265] Some preprocessing can be applied to the residual samples before deriving the class index. For example, taking the absolute value, clipping to bit depth, clipping to a specific range, linear / non-linear quantization. Note that the preprocessing can be applied independently or in combination in a predefined order. For example, taking the absolute value and then clipping to bit depth, Class(C11)=(Clip1(|Y0|)*band_num)>>bit_depth

[0266] The signs of the residual samples can also be considered to form a joint classifier. For example, bandR=(Clip1(|Y0|)*band_num)>>bit_depth signY=Y0>0?0:1 classIdx=signY*2 +bandR. In this example, band_num is different from the band_num in the above classifier C0 (ie, category (C0)). The band_num in this example is different from the value of the residual sample (ie, the residual sample value) Related.

[0267] As with C0, other variations of C0' can be applied. For example, using the current / co-located / currently adjacent / co-located residual samples for bandNum classification, or the above linear weighting, such as Figure 13B Specifically, the following values may be used when obtaining the residual sample value: the value of the co-located residual sample of one component relative to the residual sample of another component, the value of the adjacent residual sample of one component relative to the residual sample of another component, a value obtained by linearly weighting the co-located residual sample and the adjacent residual sample of one component relative to the residual sample of another component, or a comparison value of the co-located residual sample and the adjacent residual sample of one component relative to the residual sample of another component.

[0268] The C11 classifier can be combined with other classifiers to form a joint classifier. For example, let candR be the residual sample of the current component for C11 classification, bandR is the band index (quantized value of candR), and it is combined with C0 (jointly using the same and adjacent Y / U / V samples for classification) bandY=(candY*bandNumY)>>BitDepth; bandU=(candU*bandNumU)>>BitDepth; bandV=(candV*bandNumV)>>BitDepth; bandR=(candR*bandNumR)>>BitDepth; classIdx=bandY*bandNumU*bandNumV*bandNumR +bandU*bandNumV*bandNumR +bandV*bandNumR +bandR;

[0269] For example, combining C0, C10, and C11 (16 edge intensities, maximum 2bandNumY bands, and maximum 2bandNumR bands) yields a maximum of 64 classes.

[0270] In some embodiments, other classifier examples that only use current component information for current component classification can be used for cross-component classification. Figure 5A As shown in Table 1-1, luma sample information and eo-class are used to derive EdgeIdx and classify the current chroma sample. Other non-cross-component classifiers that can also be used as cross-component classifiers include edge direction, pixel intensity, pixel variance, pixel Laplacian sum, Sobel operator, compass operator, high-pass filter value, low-pass filter value, etc.

[0271] In some embodiments, multiple classifiers are used in the same POC. The current frame is divided into several regions, and the same classifier is used in each region. For example, three different classifiers are used in POC0, and which classifier (0, 1, or 2) is used is signaled at the CTU level, as shown in Table 2-11 below, which shows that different general classifiers are applied to different regions in the same picture. Table 2-11 POC Classifier C0 band_num area 0 C0 uses Y0 position 16 0 0 C0 uses Y0 position 8 1 0 C0 uses Y1 position 8 2

[0272] In some embodiments, the maximum number of multiple classifiers (multiple classifiers may also be referred to as alternative offset sets) may be fixed or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. In one example, the fixed (predefined) maximum number of multiple classifiers is 4. In that case, 4 different classifiers are used in POC0, and which classifier to use (0, 1, or 2) is signaled at the CTU level. A truncated unary (TU) code may be used to indicate the classifier used for each luma or chroma CTB. For example, as shown in Table 2-12 below, when TU code is 0: CCSAO is not applied; when TU code is 10: set 0 is applied; when TU code is 110, set 1 is applied; when TU code is 1110: set 2 is applied; when TU code is 1111: set 3 is applied. Fixed-length codes, Golomb-Rice codes, and exponential Golomb codes may also be used to indicate the classifier (offset set index) used for a CTB. Three different classifiers were used in POC1. Table 2-12

[0273] Examples of Cb and Cr CTB offset set indices are given for the 1280x720 sequence POC0 (if the CTU size is 128x128, then the number of CTUs in the frame is 10x6). POC0 Cb uses 4 offset sets and Cr uses 1 offset set. As shown in Table 2-13 below, when the offset set index is 0: CCSAO is not applied; when the offset set index is 1: set 0 is applied; when the offset set index is 2: set 1 is applied; when the offset set index is 3: set 2 is applied; when the offset set index is 4: set 3 is applied. Type means the position of the selected co-located luma sample (Yi). Different offset sets can have different types, band_nums, and corresponding offsets. Table 2-13 shows examples of Cb and Cr CTB offset set indices given for the 1280x720 sequence POC0 (if the CTU size is 128x128, then the number of CTUs in the frame is 10x6). Table 2-13

[0274] In some embodiments, examples of joint classification using co-located / current and neighboring Y / U / V samples are listed in Table 2-14 below (3-component joint bandNum classification for each Y / U / V component). Table 2-14 shows an example of joint classification using co-located / current and neighboring Y / U / V samples. In POC0, {2, 4, 1} offset sets are used for {Y, U, V} respectively. Each offset set can be adaptively switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different offset sets can have different classifiers. For example, as Figure 13B and Figure 13C In order to classify the current Y4 luma sample, Y set0 selects {current Y4, co-located U4, co-located V4} as candidates, each with a different bandNum {Y, U, V} = {16, 1, 2}. In the case of {candY, candU, candV} as the sample value of the selected {Y, U, V} candidate, the total number of categories is 32, and the category index derivation can be shown as: bandY=(candY*bandNumY)>>BitDepth; bandU=(candU*bandNumU)>>BitDepth; bandV=(candV*bandNumV)>>BitDepth; classIdx=bandY*bandNumU*bandNumV +bandU*bandNumV +bandV.

[0275] In some embodiments, the classIdx derivation of the joint classifier can be expressed as an "or-shift" form to simplify the derivation process. For example, max bandNum = {16, 4, 4} classIdx=(bandY<<4)|(bandU<<2)|bandV

[0276] Another example is in the POC1 component V set 1 classification. In this example, candPos = {neighboring Y8, neighboring U3, neighboring V0} with bandNum = {4, 1, 2} is used, which results in 8 classes. Table 2-14

[0277] In some embodiments, examples of using a combination of co-located and neighboring Y / U / V samples for classification of the current Y / U / V sample are listed (joint edgeNum (C1s) and bandNum classification for the 3-components of each Y / U / V component), for example, as shown in Table 2-15 below. The edge candidate position (edge CandPos) is the center position for the C1s classifier, the edge bit mask (edgebitMask) is the C1s neighboring sample activation indicator, and edgeNum is the corresponding number of C1s classes. In this example, C1s is only applied to the Y classifier (so edgeNum is equal to edgeNumY), where edge candPos is always Y4 (the current / co-located sample position). However, C1s can be applied to the Y / U / V classifier with edge candPos as the neighboring sample position.

[0278] In the case where diff represents the comparison score of Y C1s, the classIdx derivation can be bandY=(candY*bandNumY)>>BitDepth; bandU=(candU*bandNumU)>>BitDepth; bandV=(candV*bandNumV)>>BitDepth; edgeIdx=diff+(edgeNum>>1); bandIdx=bandY*bandNumU*bandNumV +bandU*bandNumV +bandV; classIdx=bandIdx*edgeNum+edgeIdx; Table 2-15 (Part 1) Table 2-15 (Part 2) Table 2-15 (Part 3)

[0279] In some embodiments, as described above, for a single component, multiple C0 classifiers (different positions or weight combinations, bandNum) can be combined to form a joint classifier. This joint classifier can be combined with other components to form another joint classifier, for example, using 2 Y samples (candY / candX and bandNumY / bandNumX), 1 U sample (candU and bandNumU) and 1 V sample (candV and bandNumV) to classify one U sample (Y / V can have the same concept). Class index derivation can be shown as: bandY=(candY*bandNumY)>>BitDepth; bandX=(candX*bandNumX)>>BitDepth; bandU=(candU*bandNumU)>>BitDepth; bandV=(candV*bandNumV)>>BitDepth; classIdx=bandY*bandNumX*bandNumU*bandNumV +bandX*bandNumU*bandNumV +bandU*bandNumV +bandV;

[0280] In some embodiments, if multiple C0s are used for a single component, some decoder-specific or encoder-consistent constraints may be applied. The constraints include: (1) the selected C0 candidates must be different from each other (e.g., candX!=candY), and / or (2) the newly added bandNum must be smaller than other bandNums (e.g., bandNumX<=bandNumY). By applying intuitive constraints within a single component (Y), redundant cases can be removed to save bit cost and complexity.

[0281] In some embodiments, the maximum band_num (bandNumY, bandNumU, or bandNumV) can be fixed or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. For example, max band_num = 16 is fixed in the decoder and for each frame, 4 bits are signaled to indicate the C0 band_num in the frame. Some other maximum band_num examples are listed in Table 2-16 below. Table 2-16

[0282] In some embodiments, the maximum number of classes or offsets (combination of multiple classifiers used together, e.g., C1s edgeNum*C1 bandNumY*bandNumU*bandNumV) for each set (or all sets added) can be fixed or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. For example, max is fixed for all added sets, class_num=256*4, and encoder consistency check or decoder compliance check can be used to check the constraints.

[0283] In some embodiments, restrictions may be applied to the C0 classification, for example, by setting band_num (bandNumY, bandNumU, or bandNumV) is restricted to only power-of-two values. Instead of explicitly signaling band_num, the syntax band_num_shift is signaled. The decoder can use the shift operation to avoid multiplication. Different band_num_shifts can be used for different components. Class(C0)=(Y0>>band_num_shift)>>bit_depth

[0284] Another example of operation is to consider rounding to reduce errors. Class(C0)=((Y0+(1<<(band_num_shift-1)))>>band_num_shift)>>bit_depth

[0285] For example, if band_num_max (Y, U, or V) is 16, the possible band_num_shift candidates are 0, 1, 2, 3, 4, corresponding to band_num=1, 2, 4, 8, 16, as shown in Table 2-17. Table 2-17 POC Classifier C0 band_num_shift C0 band_num General Category 0 C0 uses Y0 position 4 16 16 1 C0 uses Y7 position 3 8 8

[0286] Offset signaling

[0287] In some embodiments, the classifiers applied to Cb and Cr are different. The Cb and Cr offsets for all categories can be signaled separately. For example, different offsets signaled are applied to different chrominance components, as shown in Table 2-18 below. Table 2-18

[0288] In some embodiments, the maximum offset value is fixed or signaled at the sequence parameter set (SPS) / adaptation parameter set (APS) / picture parameter set (PPS) / picture header (PH) / slice header (SH) / region / CTU / CU / sub-block / sample level. For example, the maximum offset is between [-15, 15]. Different components can have different maximum offset values.

[0289] In some embodiments, the offset signaling may use differential pulse-code modulation (DPCM). For example, the offset {3, 3, 2, 1, -1} may be signaled as {3, 0, -1, -1, -2}.

[0290] In some embodiments, the offsets may be stored in an APS or memory buffer for multiplexing with the next picture / slice. An index may be signaled to indicate which previous frame offsets are stored for the current picture.

[0291] In some embodiments, the classifiers for Cb and Cr are the same. The Cb and Cr offsets for all classes may be jointly signaled, for example, as shown in Table 2-19 below. Table 2-19

[0292] In some embodiments, the classifiers for Cb and Cr can be the same. The Cb and Cr offsets for all categories can be jointly signaled along with the sign difference, as shown in Table 16 below. According to Table 2-20, when the Cb offset is (3, 3, 2, -1), the derived Cr offset is (-3, -3, -2, 1). Table 2-20

[0293] In some embodiments, a sign flag may be signaled for each category, as shown in Table 2-21 below. According to Table 2-21, when the Cb offset is (3, 3, 2, -1), the Cr offset derived from the corresponding sign flag is (-3, 3, 2, 1). Table 2-21

[0294] In some embodiments, the classifiers for Cb and Cr can be the same. The Cb and Cr offsets for all classes can be jointly signaled together with the weight differences, for example, as shown in Table 2-22 below, which shows that the Cb and Cr offsets for all classes can be jointly signaled together with the weight differences. The weights (w) can be selected from a finite table, for example, +-1 / 4, +-1 / 2, 0, +-1, +-2, +-4, etc., where |w| only includes values that are powers of 2. According to Table 18, when the Cb offsets are (3, 3, 2, -1), the Cr offsets derived from the corresponding sign flags are (-6, -6, -4, 2). Table 2-22

[0295] In some embodiments, the weight for each class can be signaled, for example, as shown in Table 2-23 below, which shows that the Cb and Cr offsets for each class can be jointly signaled and the weight for each class can be signaled. According to Table 2-22, when the Cb offset is (3, 3, 2, -1), the Cr offset derived from the corresponding sign flag is (-6, 12, 0, -1). Table 2-22

[0296] In some embodiments, if multiple classifiers are used in the same POC, different sets of offsets are signaled individually or jointly.

[0297] In some embodiments, previously decoded offsets can be stored for use in future frames. An index can be signaled to indicate which previously decoded offset set to use for the current frame to reduce offset signaling overhead. For example, POC2 can reuse POC0 offsets and signal an offset set idx=0 as shown in Table 2-23 below, which shows that an index can be signaled to indicate which previously decoded offset set to use for the current frame. Table 2-23

[0298] In some embodiments, the multiplexing offset set indices for Cb and Cr may be different, for example, as shown in Table 2-24 below, which shows that an index may be signaled to indicate which previously decoded offset set is used for the current frame, and the index may be different for the Cb and Cr components. Table 2-24

[0299] In some embodiments, offset signaling can use additional syntax including start and length to reduce signaling overhead. For example, when band_num = 256, only the offsets of band_idx = 37 to 44 are signaled. In the example in Table 2-25 below, the syntax for start and length are both 8-bit fixed-length codes that should match the band_num bits. Table 2-25

[0300] In some embodiments, if CCSAO is applied to all YUV 3 components, co-located and adjacent YUV samples can be used jointly for classification, and all the offset signaling methods described above for Cb / Cr can be extended to Y / Cb / Cr. In some embodiments, different component offset sets can be stored and used separately (each component has its own storage set) or stored and used jointly (each component shares / reuses the same storage). Examples of separate sets are shown in Table 2-26 below, which shows examples of how different component offset sets can be stored and used separately (each component has its own storage set) or stored and used jointly (each component shares / reuses the same storage). Table 2-26

[0301] In some embodiments, if the sequence bit depth is higher than 10 (or a specific bit depth), the offset can be quantized before being sent through the signal. At the decoder side, the decoded offset is dequantized before being applied, as shown in Table 2-27 below. For example, for a 12-bit sequence, the decoded offset is shifted left by 2 bits (dequantized). Table 2-27 Offset sent by signal Dequantization and applied offset 0 0 1 4 2 8 3 12 … 14 56 15 60

[0302] In some embodiments, the offset may be calculated as CcSaoOffsetVal = (1-2*ccsao_offset_sign_flag)*(ccsao_offset_abs<<(BitDepth-Min(10,BitDepth)))

[0303] In some embodiments, offset quantization can be encoder-selective (programmable). Whether offset quantization is enabled (on / off control) and the indicated quantization step size can be sent by signal or predefined at the SPS / APS / PPS / PH / SH / region (set) / CTU / CU / sub-block / sample level. For example, the quantization step size is predefined according to the bit depth or resolution and switched in the PH. The on / off control flag and quantization step size can be stored in the APS for future frame multiplexing. The step size range supported in the sequence can be sent by signal or predefined at the SPS / APS / PPS / PH / SH / region (set) / CTU / CU / sub-block / sample level. Due to the offset precision, the offset quantization mechanism enables the encoder to make a trade-off between bit cost and image quality improvement.

[0304] In some embodiments, the offset binarization method may depend on the quantization step size. The offset binarization method may be signaled or predefined at the SPS / APS / PPS / PH / SH / region (set) / CTU / CU / sub-block / sample level. Different components may have different or share the same {on / off control, quantization step size, offset binarization method}. For example, U / V uses the same method and Y uses a different method. Different sequence bit depths may have different predefined quantization step sizes / offset binarization methods. For example, different EGk orders are used for different quantization step sizes.

[0305] For example, step size = {0, 0, 2, 4, 6} is predefined for the {8b, 10b, 12b, 14b, 16b} sequence, and the step size / binarization method is switched at different levels.

[0306] For example, for the 8b sequence, 1 SPS flag to enable offset quantization, predefined step size = 0 (offset = 0, +-1, +-2, ...) 1 PH syntax to adaptively change the step size to 2 (offset = 0, +-4, +-8, ...) 1 region (set) level syntax for adaptively changing the step size used for each set (Set0=0, Set1=1...) 1 region (set) level syntax for switching between predefined binarization methods, Set0: TU, Set1: EG1, Set2: FLC...

[0307] For example, for a 10b sequence, 1 SPS flag to enable offset quantization, predefined step size = 1 (offset = 0, +-2, +-4, ...) Predefined quantization step size to binarization mapping: 0->EG0, 1->EG1, 2->EG2 · 1 APS syntax is used to store the previously used quantization step size (q) / according to EGk order. New pictures can add new APS indices. Index 0: Set 0: q = 0, Set 1: q = 2, Set 2: q = 1, Set 3: q = 0 Index 1: Set 0: q = 1, Set 1: q = 0, Set 2: q = 0, Set 3: q = 2 … Each region (set) in a picture can reuse the quantization step size (q) / EGk order according to the stored APS.

[0308] For example, step size = {0, 0, 2, 4, 6} is predefined for the {<480p, 720p, 1080p, 4K, > = 8K} sequence, and the step size / binarization method is switched at different levels.

[0309] In some embodiments, the concept of filter strength is further introduced herein. For example, the classifier offset can be further weighted before being applied to the sample. The weight (w) can be selected from a table of power-of-2 values. For example, +-1 / 4, +-1 / 2, 0, +-1, +-2, +-4... and so on, where |w| only includes power-of-2 values. The weight index can be signaled at the SPS / APS / PPS / PH / SH / region (set) / CTU / CU / sub-block / sample level. The quantized offset signaling can be used as a subset of the weight application. If Figure 13D As shown, recursive CCSAO is applied, and a similar weight indexing mechanism can be applied between the first and second levels.

[0310] In some examples, different classifiers are weighted: the offsets of multiple classifiers can be applied to the same instance with a weighted combination. A similar weight indexing mechanism can be signaled as described above. For example, offset_final=w*offset_1+(1-w)*offset_2, or offset_final=w1*offset_1+w2*offset_2+…

[0311] Adaptive Parameter Set (APS)

[0312] In some embodiments, instead of directly signaling CCSAO parameters in PH / SH, previously used parameters / offsets can be stored in an Adaptive Parameter Set (APS) or memory buffer for reuse by subsequent pictures / slices. An index can be signaled in PH / SH to indicate which stored previous frame offsets are used for the current picture / slice. A new APS ID can be created to maintain CCSAO history offsets. The following table shows the usage of Figure 13E For example, candPos and bandNum {Y, U, V} = {16, 4, 4}. In some examples, candPos, bandNum, and the offset signaling method can be a fixed length code (FLC) or other methods such as truncated unary (TU) code, exponential-golomb code with order k (EGk), signed EG0 (SVLC), or unsigned EG0 (UVLC). In this case, sao_cc_y_class_num (or cb, cr) is equal to sao_cc_y_band_num_y*sao_cc_y_band_num_u*sao_cc_y_band_num_v (or cb, cr). ph_sao_cc_y_aps_id is the parameter index used in this picture / slice. Note that the cb and cr components can follow the same signaling logic. Table 2-28

[0313] aps_adaptation_parameter_set_id provides an identifier of the APS for reference by other syntax elements. When aps_params_type is equal to CCSAO_APS, the value of aps_adaptation_parameter_set_id should be in the range of 0-7, inclusive.

[0314] ph_sao_cc_y_aps_id specifies the aps_adaptation_parameter_set_id of the CCSAO APS referenced by the Y color component of the slice in the current picture. When ph_sao_cc_y_aps_id is present, the following apply: the value of sao_cc_y_set_signal_flag of the APS NAL unit with aps_params_type equal to CCSAO_APS and aps_adaptation_parameter_set_id equal to ph_sao_cc_y_aps_id shall be equal to 1; the TemporalId of the APS Network Abstraction Layer (NAL) unit with aps_params_type equal to CCSAO_APS and aps_adaptation_parameter_set_id equal to ph_sao_cc_y_aps_id shall be less than or equal to the TemporalId of the current picture.

[0315] In some embodiments, an APS update mechanism is described herein. The maximum number of APS offset sets may be predefined or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different components may have different maximum number limits. If the APS offset set is full, the newly added offset set may replace an existing stored offset using a first-in-first-out (FIFO), last-in-first-out (LIFO), or least recently used (LRU) mechanism, or an index value is received indicating which APS offset set should be replaced. In some examples, if the selected classifier consists of candPos / edge info / coding info, etc., all classifier information may be part of the APS offset set and may also be stored in the APS offset set together with its offset value. In some cases, the above-mentioned update mechanism may also be predefined or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.

[0316] In some embodiments, constraints may be applied, referred to as “pruning.” For example, newly received classifier information and offsets cannot be identical to any of the stored APS offset sets (of the same component, or across different components).

[0317] In some examples, if the C0 candPos / bandNum classifier is used, the maximum number of APS offset sets is 4 per Y / U / V, and the FIFO update is for Y / V, the idx indicating the update is for U. Table 2-29 shows CCSAO offset set updates using FIFO. Table 2-29

[0318] In some embodiments, the pruning criteria can be relaxed to provide a more flexible way for the encoder to make trade-offs: for example, N offsets are allowed to be different when applying the pruning operation (e.g., N=4); in another example, the value of each offset is allowed to be different (denoted as "thr") when applying the pruning operation (e.g., +-2).

[0319] In some embodiments, the two criteria may be applied simultaneously or separately. Whether each criterion is applied is predefined or switched at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.

[0320] In some embodiments, N / thr may be predefined or switched at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.

[0321] In some embodiments, the FIFO update may be (1) a cyclic update from the previously left set index (if all are updated, start again from set 0), as in the above example, or (2) an update each time from set 0. In some examples, when a new offset set is received, the update may be at the PH (as in the example) or SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.

[0322] For LRU updates, the decoder maintains a counter table that counts the "total offset set usage count", which can be updated at the SPS / APS / per-picture-group (GOP) structure / PPS / PH / SH / region / CTU / CU / sub-block / sample level. A newly received offset set replaces the least recently used offset set in the APS. If two stored offset sets have the same count, FIFO / LIFO can be used. For example, see component Y in Table 2-30 below. Table 2-30

[0323] In some embodiments, different components may have different update mechanisms.

[0324] In some embodiments, different components (eg, U / V) may share the same classifier (same candPos / edge information / codec information / offset, possibly with additional weights with modifiers).

[0325] In some embodiments, since the offset sets used by different pictures / slices may have only slight differences in offset values, a "patch" implementation can be used in the offset replacement mechanism. In some embodiments, the "patch" implementation is differential pulse-code modulation (DPCM). For example, when a new offset set (OffsetNew) is sent by signal, the offset value can be on top of the offset set (OffsetOld) stored by the existing APS. The encoder only sends the delta value to update the old offset set (DPCM: OffsetNew = OffsetOld + delta). In the following examples shown in Table 2-31, options other than FIFO updates can also be used (LRU, LIFO, or signaling an index indicating which set to update). The YUV components can have the same update mechanism or use different update mechanisms. Although the classifier candPos / bandNum is not changed in the example in Table 2-31, the overlay set classifier can be indicated by sending an additional flag by signaling (flag = 0: only update the set offset, flag = 1: update both the set classifier and the set offset). Table 2-31

[0326] In some embodiments, the DPCM incremental offset value can be signaled in the FLC / TU / Egk (order=0, 1, ...) code. A flag can be signaled for each offset set, indicating whether DPCM signaling is enabled. The DPCM incremental offset value or the newly added offset value (when APSDPCM is enabled=0, it is directly signaled without DPCM) (ccsao_offset_abs) may be dequantized / mapped before being applied to the target offset (CcSaoOffsetVal). The offset quantization step may be signaled or predefined at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. For example, one approach is to directly signal an offset with a quantization step size of 2: CcSaoOffsetVal=(1-2*ccsao_offset_sign_flag)*(ccsao_offset_abs<<1)

[0327] Another approach is to use a DPCM signaling offset with a quantization step size of 2: CcSaoOffsetVal=CcSaoOffsetVal+(1-2*ccsao_offset_sign_flag)*(ccsao_offset_abs<<1)

[0328] In some embodiments, a constraint can be applied to reduce direct offset signaling overhead, for example, the updated offset value must have the same sign as the old offset value. By using such an inferred offset sign, the new updated offset does not need to send the sign flag again (ccsao_offset_sign_flag is inferred to be the same as the old offset).

[0329] In some embodiments, the sample processing is described as follows. Let R(x,y) be the input luma or chroma sample value before CCSAO, and R'(x,y) be the output luma or chroma sample value after CCSAO: offset=ccsao_offset[class_index of R(x,y)] R'(x,y)=Clip3(0,(1< <bit_depth)–1,R(x,y)+offset)

[0330] Sample processing

[0331] According to the above equation, each luma or chroma sample value R(x,y) is classified using the indicated classifier of the current picture and / or the current offset set index. The corresponding offset of the derived class index is added to each luma or chroma sample value R(x,y). A clipping function Clip3 is applied to (R(x,y)+offset) so that the output luma or chroma sample value R'(x,y) is within the bit depth dynamic range, for example, the range 0 to (1 < <bit_depth)–1。

[0332] For each luma or chroma sample, first it is classified using the indicated classifier of the current picture / current offset set index; second it is increased by the corresponding offset of the derived class index; and third it is cropped to the bit depth dynamic range.

[0333] Figure 6 is a block diagram illustrating that the proposed bilateral filter (BIF) and SAO both use samples from the deblocking stage as input according to some embodiments of the present disclosure.

[0334] In some embodiments, when CCSAO is operated with other loop filters, the clipping operation can be: (1) Adding post-clipping. The following equations show examples of (a) CCSAO operating with SAO and BIF, or (b) CCSAO replacing SAO but still used with BIF. (a)IOUT =clip1(I C +ΔI SAO +ΔI BIF ++ΔI CCSAO ) (b)I OUT =clip1(I C +ΔI CCSAO +ΔI BIF ) (2) Add pre-clipping, using BIF operations. In some embodiments, the clipping order can be switched. (a)I OUT =clip1(I C +ΔI SAO ) I′ OUT =clip1(I OUT +ΔI BIF ) I″ OUT =clip1(I″ OUT +ΔI CCSAO ) (b)I OUT =clip1(I C +ΔI BIF ) I′ OUT =clip1(I′ OUT +ΔI CCSAO ) (3) Partial addition and trimming (a)I OUT =clip1(I C +ΔI SAO +ΔI BIF ) I′ OUT =clip1(I OUT +ΔI CCSAO )

[0335] In some embodiments, different trimming combinations give different tradeoffs between correction accuracy and hardware temporary buffer size (register or SRAM bit width).

[0336] Figure 6 SAO / BIF offset clipping is shown. More specifically, for example, Figure 6The current design when BIF interacts with SAO is shown. The offsets of SAO and BIF are added to the input samples, and then a bit depth clipping is performed. However, when CCSAO is also added to the SAO level, two possible clipping designs can be chosen: (1) adding an additional bit depth clipping for CCSAO, and (2) a coordinated design that performs joint clipping after adding SAO / BIF / CCSAO offsets to the input samples. In some embodiments, the above clipping designs differ only on luma samples, because BIF is only applied to them.

[0337] Boundary Processing

[0338] In some embodiments, boundary handling is described below. CCSAO is not applied to the current chroma (chroma) sample if any of the co-located and adjacent luma (chroma) samples used for classification are outside the current picture. Figures 23A-23B is a block diagram illustrating that CCSAO is not applied to the current chrominance (chroma) sample if any of the co-located and adjacent luma (chroma) samples used for classification is outside the current picture, according to some embodiments of the present disclosure. Figure 23A In , if the classifier is used, CCSAO is not applied to the left column chroma components of the current picture. For example, if C1' is used, CCSAO is not applied to the left column and the first row chroma components of the current picture, such as Figure 23B shown.

[0339] Figures 24A-24B A block diagram illustrating applying CCSAO to a current luma or chroma sample if any of the co-located and adjacent luma or chroma samples used for classification is outside the current picture according to some embodiments of the present disclosure. In some embodiments, a variation is that if any of the co-located and adjacent luma or chroma samples used for classification is outside the current picture, then Figure 24A The missing samples are reused as shown, or as Figure 24B The mirror fills the missing samples to create samples for classification, and CCSAO can be applied to the current luma or chroma sample. In some embodiments, if any of the co-located and adjacent luma (chroma) samples used for classification is outside the current sub-picture / slice / tile / patch / CTU / 360 virtual boundary, the disabling / repeating / mirroring picture boundary processing method disclosed herein can also be applied to the sub-picture / slice / tile / CTU / 360 virtual edge.

[0340] For example, a picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that covers a rectangular area of the picture.

[0341] A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a picture tile.

[0342] A sub-image consists of one or more strips that together cover a rectangular area of the image.

[0343] In some embodiments, 360-degree video is captured on a sphere and has no "boundaries" in nature; reference samples outside the reference picture boundaries in the projection domain can always be obtained from neighboring samples in the spherical domain. For projection formats consisting of multiple planes, discontinuities will occur between two or more adjacent planes in a frame-packed picture, regardless of the compact frame packing arrangement used. In VVC, vertical and / or horizontal virtual boundaries are introduced that prohibit in-loop filtering operations, and the positions of these boundaries are signaled in the SPS or picture header. Compared to using two tiles (one for each set of consecutive planes), the use of 360 virtual boundaries is more flexible because it does not require the plane size to be a multiple of the CTU size. In some embodiments, the maximum number of vertical 360 virtual boundaries is 3, and the maximum number of horizontal 360 virtual boundaries is also 3. In some embodiments, the distance between two virtual boundaries is greater than or equal to the CTU size, and the virtual boundary granularity is 8 luma samples, for example, an 8x8 sample grid.

[0344] Figures 28A-28B is a block diagram illustrating that CCSAO is not applied to a current chroma sample if the corresponding co-located or adjacent luma sample selected for classification is outside the virtual space defined by a virtual boundary, according to some embodiments of the present disclosure. In some embodiments, the virtual boundary (VB) is a virtual line that separates the space within a picture frame. In some embodiments, if the virtual boundary (VB) is applied in the current frame, CCSAO is not applied to chroma samples whose corresponding luma positions have been selected outside the virtual space defined by the virtual boundary. Figures 28A-28B An example of a virtual boundary for the C0 classifier with 9 luma position candidates is shown. For each CTU, CCSAO is not applied to the chroma samples for which the corresponding selected luma position is outside the virtual space enclosed by the virtual boundary. For example, Figure 28A In , CCSAO is not applied to chroma samples 2802 when the selected Y7 luma sample position is on the other side of the horizontal virtual boundary 2806, which is located 4 pixel rows from the bottom side of the frame. Figure 28B , CCSAO is not applied to chroma samples 2804 when the selected Y5 luma sample location is on the other side of a vertical virtual boundary 2808, which is located y pixel rows from the right side of the frame.

[0345] Figure 32A-Figure 32BIt is shown that repeating or mirroring padding may be applied to luma samples outside the virtual boundary according to some embodiments of the present disclosure. Figure 32A An example of repeated padding is shown. If the original Y7 is selected as the classifier located at the bottom side of VB3202, the Y4 luma sample value is used for classification (copied to the Y7 position) instead of the original Y7 luma sample value. Figure 32B An example of mirror padding is shown. If Y7 is selected as the classifier located at the bottom side of VB 3204, the Y1 luma sample value, which is symmetrical to the Y7 value relative to the Y0 luma sample, is used for classification instead of the original Y7 luma sample value. The padding method provides the possibility of applying CCSAO with more chroma samples, thereby obtaining more coding gain.

[0346] In some embodiments, restrictions may be applied to reduce the line buffer required for CCSAO and to simplify boundary handling condition checking. Figure 26A It is shown that according to some embodiments of the present disclosure, if all 9 co-located adjacent luma samples are used for classification, an additional 1 luma line buffer may be required, ie, a full line of luma samples above line -5 of the current VB 1602. Figures 18A-18G shows an example of classification using only 6 brightness candidates, which reduces Figures 23A-23B and Figures 24A-24B The row buffer in and does not require any additional bounds checking.

[0347] In some embodiments, using luma samples for CCSAO classification may increase the luma line buffer, thereby increasing the decoder hardware implementation cost. Figure 25This figure illustrates an example of an AVS decoder according to some embodiments of the present disclosure, whereby two additional luma line buffers can be added for nine luma CCSAO candidates spanning VB 1702. For luma and chroma samples above virtual boundary (VB) 1702, DBF / SAO / ALF processing is performed on the current CTU line. For luma and chroma samples below VB 1702, DBF / SAO / ALF processing is performed on the next CTU line. In the AVS decoder hardware design, luma lines -4 to -1 DBF front samples, line -5 SAO front samples, and chroma lines -3 to -1 DBF front samples, line -4 SAO front samples are stored in line buffers for DBF / SAO / ALF processing on the next CTU line. When processing the next CTU line, luma and chroma samples not in the line buffers are unavailable. However, for example, at chroma line -3(b), chroma samples are processed on the next CTU line, but CCSAO requires SAO front luma sample lines -7, -6, and -5 for classification. SAO pre-luminance sample lines -7 and -6 are not in the line buffer, so they are not available. Adding SAO pre-luminance sample lines -7 and -6 to the line buffer would increase the decoder hardware implementation cost. In some examples, the luma VB (line -4) and chroma VB (line -3) may be different (misaligned).

[0348] and Figure 25 similar, Figure 26A This figure shows an illustration of VVC according to some embodiments of the present disclosure, whereby nine luma CCSAO candidates spanning VB 1802 can add one additional luma line buffer. VB can be different in different standards. In VVC, the luma VB is line -4 and the chroma VB is line -2, so nine luma CCSAO candidates can add one additional luma line buffer.

[0349] In some embodiments, in a first solution, if any of the luma candidates for the chroma sample spans the VB (outside the current chroma sample VB), CCSAO is disabled for the chroma sample. Figures 27A-27C It is shown that in AVS and VVC, according to some embodiments of the present disclosure, if any of the luma candidates of the chroma sample spans VB 2702 (outside the current chroma sample VB), then CCSAO is disabled for the chroma sample. Figures 28A-28B Some examples of this embodiment are also shown.

[0350] In some embodiments, in a second solution, repetition padding is used for CCSAO from a luma row that is close to and on the other side of VB, e.g., luma row -4, for "across VB" luma candidates. In some embodiments, repetition padding is implemented for "across VB" chroma candidates from the luma nearest neighbor below VB. Figures 29A-29CIt is shown that in AVS and VVC, according to some embodiments of the present disclosure, if any of the luma candidates of the chroma samples spans VB 2902 (outside the current chroma sample VB), CCSAO is enabled using repeated padding for the chroma samples. Figure 28A Some examples of this embodiment are also shown.

[0351] In some embodiments, in a third solution, for "cross-VB" luma candidates, mirror fill is used for CCSAO from luma VB down. Figures 30A-30C It is shown that in AVS and VVC, according to some embodiments of the present disclosure, if any of the luma candidates for the chroma samples spans VB 3002 (outside the current chroma sample VB), CCSAO is enabled using mirror padding for the chroma samples. Figure 28B and Figure 24B Some examples of this embodiment are also shown.In some embodiments, in a fourth solution, "double-sided symmetric filling" is used to apply CCSAO. Figure 31A-Figure 31B Some examples of different CCSAO shapes (e.g., 9 brightness candidates ( Figure 31A ) and 8 brightness candidates ( Figure 31B )), use bilateral symmetric padding to enable CCSAO. For a luma sample set with a co-located central luma sample with chroma samples, if one side of the luma sample set is outside VB 3102, bilateral symmetric padding is applied to both sides of the luma sample set. For example, in Figure 31A In , the luminance samples Y0, Y1 and Y2 are outside VB 3102, so Y3, Y4, Y5 are used to fill Y0, Y2 and Y6, Y7, Y8. Figure 31B In , luma sample Y0 is outside VB 3102, so Y0 is filled with Y2, and Y7 is filled with Y5.

[0352] Figure 26B This figure shows that according to some embodiments of the present disclosure, when co-located or adjacent chroma samples are used to classify the current luma sample, the selected chroma candidate may span VB and require an additional chroma line buffer. Similar solutions 1 to 4 as described above can be applied to handle this problem.

[0353] Solution 1 is to disable CCSAO for luma samples when any chroma candidate for luma samples may span VB.

[0354] Solution 2 is to use repeated padding of the nearest neighbors of the chroma below VB for the “cross-VB” chroma candidates.

[0355] Solution 3 is to use mirror padding lower than the chroma VB for the "cross-VB" chroma candidates.

[0356] Solution 4 is to use "double-sided symmetric padding". For a candidate set centered on the CCSAO co-located chroma sample, if one side of the candidate set is outside the VB, double-sided symmetric padding is applied to both sides.

[0357] The padding method provides the possibility of applying more luma or chroma samples for CCSAO, thus obtaining more coding gain.

[0358] In some embodiments, at the bottom picture (or slice, tile, brick) boundary CTU row, the samples below VB are processed in the current CTU row, so the above special processing (solutions 1, 2, 3, 4) are not applied to the bottom picture (or slice, tile, brick) boundary CTU row. For example, a 1920x1080 frame is divided into 128x128 CTUs. A frame contains 15x9 CTUs (rounded up). The bottom row of CTU is the 15th CTU row. The decoding process is CTU row by row, and each CTU row is CTU row by row. Deblocking needs to be applied along the horizontal CTU boundary between the current CTU row and the next CTU row. CTB VB is applied to each CTU row because within a CTU, at the bottom 4 / 2 luminance / chrominance rows, DBF samples (VVC case) are processed in the next CTU row and are not available for CCSAO at the current CTU row. However, in the bottom CTU row of the image frame, the bottom 4 / 2 luma / chroma row DBF samples are available in the current CTU row because there is no next CTU row left and they are DBF processed in the current CTU row.

[0359] In some embodiments, the VB shown in Figures 13 to 22 can be replaced by the boundaries of the sub-picture / slice / tile / block / CTU / 360 virtual boundaries. In some embodiments, the positions of the chroma and luma samples in Figures 13 to 22 can be switched. In some embodiments, Figure 6 、 Figures 23A-32B The positions of the chroma and luma samples in the ALF VB may be replaced by the positions of the first chroma sample and the second chroma sample. In some embodiments, the ALF VB within a CTU may generally be horizontal. In some embodiments, the boundaries of the sub-picture / slice / tile / block / CTU / 360 virtual boundaries may be horizontal or vertical.

[0360] In some embodiments, restrictions can be applied to reduce the line buffer required for CCSAO and simplify boundary processing condition checks. Figure 26 shows that if all 9 co-located adjacent luma samples are used for classification, an additional luma line buffer may be required (the entire row of luma samples at line: -5). Figure 33A-Figure 33BThe limitation of using a limited number of brightness candidates for classification according to some embodiments of the present disclosure is shown. Figure 33A The limitation of using only 6 brightness candidates for classification is shown. Figure 33B The limitation of using only 4 brightness candidates for classification is shown.

[0361] Applied area

[0362] In some embodiments, the application area is implemented. The area unit of CCSAO application can be based on CTB. That is, in a CTB, the on / off control and CCSAO parameters (offset for classification, luma candidate position, band_num, bit mask, etc., offset set index) are the same.

[0363] In some embodiments, the applied region may not be aligned with a CTB boundary. For example, the applied region may not be aligned with a chroma CTB boundary, but may be offset. Syntax (on / off control, CCSAO parameters) is still signaled for each CTB, but the actual applied region may not be aligned with a CTB boundary. Figure 34 It is shown that the CCSAO application area according to some embodiments of the present disclosure is not aligned with the CTB / CTU boundary 3406. For example, the applied area is not aligned with the chroma CTB / CTU boundary 3406, but is offset to the upper left by (4, 4) samples to the VB 3408. This non-aligned CTB boundary design is beneficial for deblocking because the same deblocking parameters are used for each 8x8 deblocking area.

[0364] In some embodiments, the area unit (mask size) of CCSAO application can be variable (larger or smaller than the CTB size), as shown in Table 2-32. The mask size may be different for different components. The mask size can be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, in PH, a series of mask on / off flags and offset set indices are signaled to indicate each CCSAO area information. Table 2-32 POC Quantity CTB size Mask size 0 Cb 64x64 128x128 0 Cr 64x64 32x32 1 Cb 64x64 16x16 1 Cr 64x64 256x256

[0365] In some embodiments, the region frame segmentation applied by CCSAO may be fixed, for example, the frame is segmented into N regions. Figure 35 It is shown that the region frame segmentation of the CCSAO application can be fixed using CCSAO parameters according to some embodiments of the present disclosure.

[0366] In some embodiments, each region may have its own region on / off control flag and CCSAO parameters. In addition, if the region size is larger than the CTB size, it may have a CTB on / off control flag and a region on / off control flag. Figure 35 (a) and (b) show some examples of segmenting a frame into N regions. Figure 35 (a) shows the vertical segmentation into four regions. Figure 35 (b) shows a square partition of four regions. In some embodiments, similar to the picture-level CTB full-on control flag (ph_cc_sao_cb_ctb_control_flag / ph_cc_sao_cr_ctb_control_flag), if the region on / off control flag is off, a CTB on / off flag can be further signaled. Otherwise, CCSAO is applied to all CTBs in the region without further signaling of the CTB flag.

[0367] In some embodiments, different CCSAO application areas can share the same area on / off control and CCSAO parameters. Figure 35 In (c), regions 0 to 2 share the same parameters, and regions 3 to 15 share the same parameters. Figure 35 (c) also shows the region on / off control flag, and the CCSAO parameters can be signaled in Hilbert scan order.

[0368] In some embodiments, the area unit for applying CCSAO can be a quadtree / binarytree / ternarytree separated from the picture / slice / CTB level. Similar to CTB partitioning, a series of partitioning flags are signaled to indicate the area partitioning for CCSAO application. Figure 36 It is shown that the CCSAO application area according to some embodiments of the present disclosure may be a binary tree (BT) / quad tree (QT) / ternary tree (TT) separated from the frame / slice / CTB level.

[0369] Figure 37 This is a block diagram illustrating multiple classifiers used and switched at different levels within a picture frame according to some embodiments of the present disclosure. In some embodiments, if multiple classifiers are used within a frame, the method for applying the classifier set index can be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, four classifier sets are used within a frame, and these sets are switched within the PH, as shown in Table 2-33 below. Figure 37 (a) and (c) show the default fixed region classifier. Figure 37(b) shows the classifier set index is signaled at the mask / CTB level, where 0 indicates CCSAO is off for that CTB, and 1 to 4 represent the set index. Table 2-33 POC 0 The square is divided into 4 regions (same as the frame QT divided into a maximum depth of 1) (a) 1 CTB level switching classifier (b) 2 Split into 4 regions vertically (c) 3 Split into frames with a maximum depth of 2 QT

[0370] In some embodiments, for the default region case, if the CTB in the region does not use the default set index (e.g., the region level flag is 0) but uses another classifier set in the frame, the region level flag may be signaled. For example, if the default set index is used, the region level flag is 1. For example, in a square partition of 4 regions, the following classifier set is used as shown in Table 2-34 below, which shows that a region level flag may be signaled to indicate whether the CTB in the region does not use the default set index. Table 2-34 POC area Logo Using the default collection index 0 1 1 Use default collection: 1 2 1 Using the default collection: 2 3 1 Using the default collection: 3 4 0 CTB switch sets 1 to 4

[0371] Figure 38 is a block diagram illustrating that the region segmentation of a CCSAO application according to some embodiments of the present disclosure may be dynamic and switched at the picture level. For example, Figure 38 (a) shows that 3 CCSAO offset sets (set_num=3) are used in this POC, thus dividing the picture frame vertically into 3 regions. Figure 38 (b) shows that 4 CCSAO offset sets (set_num=4) are used in this POC, thus dividing the picture frame horizontally into 4 regions. Figure 38 (c) shows that 3 CCSAO offset sets (set_num=3) are used in this POC, thus raster-dividing the picture frame into 3 regions. Each region can have its own region-wide on flag to store each CTB on / off control bit. The number of regions depends on the set_num of the picture sent by signal.

[0372] The CCSAO application area can be a specific area based on the coding information within the block (sample position, sample coding mode, loop filter parameters, etc.). For example, 1) the CCSAO application area can be applied only when the sample is skip-coded, or 2) the area to which CCSAO is applied only contains N samples along the CTU boundary, or 3) the area to which CCSAO is applied only includes samples on the 8x8 grid in the frame, or 4) the area to which CCSAO is applied only contains DBF-filtered samples, or 5) the area to which CCSAO is applied only contains the top M rows and left N rows in the CU, or (6) the area to which CCSAO is applied only contains intra-coded samples, or (7) the area to which CCSAO is applied only contains samples in cbf=0 blocks, or (8) the area to which CCSAO is applied is only on blocks with a block QP of [N,M], where (N,M) can be predefined or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Cross-component coding information can also be considered. (9) The area where CCSAO is applied is on the chroma samples, where the co-located luma samples are in the cbf=0 block.

[0373] In some embodiments, whether to introduce coding information application area restriction can be predefined, or a control flag can be sent by signaling at the SPS / APS / PPS / PH / SH / region (each alternative set) / CTU / CU / sub-block / sample level to indicate whether the specified coding information is included / excluded in the CCSAO application. The decoder skips CCSAO processing of these areas according to the predefined conditions or control flags. For example, YUV uses different predefined / flag control conditions that are switched at the region (set) level. The CCSAO application judgment can be at the CU / TU / PU or sample level. Table 2-35 shows that YUV uses different predefined / flag control conditions that are switched at the region (set) level. Table 2-35

[0374] Another example is to reuse all or part of the bilateral enablement constraints (predefined)

[0375] In some embodiments, excluding certain areas can facilitate the collection of CCSAO statistics. Offset derivation can be more accurate or tailored to areas that truly require correction. For example, a block with cbf=0 typically indicates a perfectly predicted block, which may not require further correction. Excluding these blocks can facilitate offset derivation for other areas.

[0376] Different application areas can use different classifiers. For example, in a CTU, skip mode uses C1, 8x8 grid uses C2, and skip mode and 8x8 grid use C3. For example, in a CTU, skip mode coded samples use C1, CU center samples use C2, and skip mode coded samples in the CU center use C3. Figure 39 FIG29 is a schematic diagram illustrating that a CCSAO classifier according to some embodiments of the present disclosure can take into account current or cross-component coding information. For example, different codec modes / parameters / sample locations can form different classifiers. Different codec information can be combined to form a joint classifier. Different classifiers can be used for different regions. FIG29 also shows another example of an application region.

[0377] In some embodiments, a predefined or flag-controlled "codec information exclusion area" mechanism may be used in DBF / SAO pre / SAO / BIF / CCSAO / ALF / CCALF / NN loop filter (NNLF) or other loop filters.

[0378] grammar

[0379] In some embodiments, the implemented CCSAO syntax is shown in Table 2-36 below. In some examples, the binarization of each syntax element can be changed. In AVS3, the term block is similar to slice, and the block header is similar to the slice header. FLC stands for fixed length code. TU stands for truncated unary code. EGk stands for exponential golomb code of order k, where k can be fixed. SVLC stands for signaled EG0. UVLC stands for unsigned EG0. Table 2-36

[0380] If a higher level flag is off, the lower level flags can be inferred from the off state of the flag and do not need to be signaled. For example, if ph_cc_sao_cb_flag is false in this picture, then ph_cc_sao_cb_band_num_minus1, ph_cc_sao_cb_luma_type, cc_sao_cb_offset_sign_flag, cc_sao_cb_offset_abs, ctb_cc_sao_cb_flag, cc_sao_cb_merge_left_flag, and cc_sao_cb_merge_up_flag are not present and are inferred to be false.

[0381] In some embodiments, the SPS ccsao_enabled_flag is conditional on the SPS SAO enabled flag shown in Table 2-37 below. Table 2-37

[0382] In some embodiments, ph_cc_sao_cb_ctb_control_flag and ph_cc_sao_cr_ctb_control_flag indicate whether Cb / Cr CTB on / off control granularity is enabled. If ph_cc_sao_cb_ctb_control_flag and ph_cc_sao_cr_ctb_control_flag are enabled, ctb_cc_sao_cb_flag and ctb_cc_sao_cr_flag may be further signaled. Otherwise, whether CCSAO is applied in the current picture depends on ph_cc_sao_cb_flag and ph_cc_sao_cr_flag, without further signaling ctb_cc_sao_cb_flag and ctb_cc_sao_cr_flag at the CTB level.

[0383] In some embodiments, for ph_cc_sao_cb_type and ph_cc_sao_cr_type, a flag may be further signaled to distinguish whether to use the center-co-located luma position ( Figures 18A-18GThe chroma samples are classified using the Y0 position in the CTB to reduce bit overhead. Similarly, if cc_sao_cb_type and cc_sao_cr_type are signaled at the CTB level, a flag can be further signaled using the same mechanism. For example, if the number of C0 luma position candidates is 9, cc_sao_cb_type0_flag is further signaled to distinguish whether the center-collocated luma position is used, as shown in Table 2-38 below. If the center-collocated luma position is not used, cc_sao_cb_type_iddc is used to indicate which of the remaining 8 adjacent luma positions is used. Table 2-38

[0384] Table 2-39 below shows an example in AVS where a single (set_num=1) or multiple (set_num>1) classifiers are used in the frame. Note that the syntax tokens can be mapped to the tokens used previously. Table 2-39

[0385] If each region has its own set of Figure 35 or Figure 37 In combination, the syntax example may include the area on / off control flag (picture_ccsao_lcu_control_flag[compIdx][setIdx]) shown in the following Table 2-40. Table 2-40

[0386] In some embodiments, for high-level syntax, pps_ccsao_info_in_ph_flag and gci_no_sao_constraint_flag may be added.

[0387] In some embodiments, pps_ccsao_info_in_ph_flag equal to 1 specifies that CCSAO filter information may be present in the PH syntax structure, but not in slice headers that reference a PPS that does not contain a PH syntax structure. pps_ccsao_info_in_ph_flag equal to 0 specifies that CCSAO filter information is not present in the PH syntax structure, and may be present in slice headers that reference a PPS. When not present, the value of pps_ccsao_info_in_ph_flag is inferred to be equal to 0.

[0388] In some embodiments, gci_no_ccsao_constraint_flag equal to 1 specifies that sps_ccsao_enabled_flag for all pictures in OlsInScope should be equal to 0. gci_no_ccsao_constraint_flag equal to 0 does not impose such constraints. In some embodiments, the bitstream of the video includes one or more output layer sets (OLSs) according to a rule. In the examples herein, OlsInScope refers to one or more OLSs in scope. In some examples, the profile_tier_level() syntax structure provides level information, and optionally the profile, layer, sub-profile, and general constraint information followed by OlsInScope. When the profile_tier_level() syntax structure is included in a VPS, OlsInScope is one or more OLSs specified by the VPS. When the profile_tier_level() syntax structure is included in an SPS, OlsInScope is an OLS that includes only the layer that is the lowest layer among the layers referencing the SPS, and the lowest layer is an independent layer.

[0389] In some embodiments, separate signaling of band_num_y_minus1, band_num_u_minus1, band_num_v_minus1 may introduce syntax redundancy. For example, as shown in Table 2-41, U1 / V1 are the same as Y1 because no segmentation is applied on U / V (redundant). Table 2-41

[0390] In some embodiments, if the classifier is a band classifier or a joint classifier including band classifiers, the encoder may predefine an indicator or signal the indicator. In some examples, the band classifier may be determined by: utilizing one or more samples based on co-located and / or neighboring samples from the Y component and sample values of the current and neighboring samples of the U / V component relative to corresponding samples of the U / V component; dividing the range of the sample values into a number of bands; and selecting a band from the number of bands.

[0391] In some examples, if a classifier or combined classifier (e.g., C0+C10) consists of bandNum segments for one or more components, the indicator can be a bandNumber indicator (bandIdc, bandNum mapping table) signaled or predefined at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock level to indicate the bandNum segment for one or more components. Different components may have different bandNum indicators. Different components may share the same bandNum indicator. For example, U / V share the same bandNum mapping table.

[0392] For example, bandNum indicators can be predefined for different components, where the bandIdc of U / V contains bandNum segments for more than one component. Table 2-42 shows an example of a bandNum mapping table. As shown in Table 2-42 below, if the current sample is of the U component, a bandNum indicator of 2 can be signaled to indicate the use of the Y3 segment, that is, the bandNumber segment of the Y component. In another example, if the current sample is of the V component, a bandNum indicator of 6 can be signaled to indicate the use of the U2 segment and the V2 segment, while a bandNum indicator of 7 can be signaled to indicate the use of the Y2 segment, the U2 segment, and the V2 segment. Table 2-42 bandIdc 0 1 2 3 4 5 6 7 Y Y1 Y2 Y3 Y4 U2 U3 V2 V3 U Y1 Y2 Y3 Y4 Y2U2 U2 Y2V2 V2 V Y1 Y2 Y3 Y4 U2 V2 U2V2 Y2U2V2

[0393] In some embodiments, C (current chroma) and CT (chroma transposed, using another chroma component) can be used to represent bandIdc. Table 2-43 shows an example of a bandNum mapping table. For example, as shown in Table 2-43, when C represents the current chroma component, CT represents another chroma component. For example, when C represents the U component, CT represents the V component; and when C represents the V component, CT represents the U component. Table 2-43 bandIdc 0 1 2 3 4 5 6 7 Y Y1 Y2 Y3 Y4 U2 U3 V2 V3 U Y1 Y2 Y3 Y4 C2 C3 CT2 CT3 V Y1 Y2 Y3 Y4 Y2C2 Y2C3 C2CT2 Y3CT3

[0394] In some embodiments, the bandNum mapping table can be adjusted / changed or predefined by the encoder at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block levels according to different granularity requirements.

[0395] Extensions to SAO filters after intra and inter prediction

[0396] In some embodiments, extensions to intra- and inter-prediction post-SAO filters are further described below. In some embodiments, the SAO classification methods disclosed in this disclosure (including cross-component sample / coding information classification) can be used as prediction post-filters, and the prediction can be intra, inter, or other prediction tools such as intra block copy. Figure 40A is a block diagram illustrating the use of the SAO classification method disclosed in the present disclosure as a post-prediction filter according to some embodiments of the present disclosure.

[0397] In some embodiments, a corresponding classifier is selected for each Y, U, and V component. For each component prediction sample, it is first classified and the corresponding offset is added. For example, each component can be classified using the current sample and neighboring samples. Y is classified using the current Y and neighboring Y samples, and U / V is classified using the current U / V sample, as shown in Table 2-44 below. Figures 40B to 40D is a block diagram illustrating that for a post-prediction SAO filter, each component may be classified using current and neighboring samples according to some embodiments of the present disclosure. Table 2-44

[0398] In some embodiments, the refined prediction samples (Ypred', Upred', Vpred') are updated by adding the corresponding class offsets and are thereafter used for intra, inter, or other predictions. Ypred'=clip3(0,(1< <bit_depth)-1,Ypred+h_Y[i]) Upred'=clip3(0,(1< <bit_depth)-1,Upred+h_U[i]) Vpred'=clip3(0,(1< <bit_depth)-1,Vpred+h_V[i])

[0399] In some embodiments, for chroma U and V components, in addition to the current chroma component, a cross component (Y) can be used for further offset classification. Additional cross component offsets (h'_U, h'_V) can be added to the current component offsets (h_U, h_V), for example, as shown in Table 2-45 below. Table 2-45

[0400] In some embodiments, the refined prediction samples (Upred", Vpred") are updated by adding corresponding class offsets and are thereafter used for intra, inter, or other predictions. Upred"=clip3(0,(1< <bit_depth)-1,Upred’+h’_U[i]) Vpred"=clip3(0,(1< <bit_depth)-1,Vpred’+h’_V[i])

[0401] In some embodiments, intra and inter prediction may use different SAO filter offsets.

[0402] Extension to post-reconstruction filter

[0403] Figure 15C is a block diagram illustrating the use of the SAO classification method disclosed in this disclosure as a post-reconstruction filter according to some embodiments of the present disclosure.

[0404] In some embodiments, the SAO / CCSAO classification methods disclosed herein (including cross-component sample / coding information classification) can be used as filters applied to the reconstructed samples of a tree unit (TU). Figure 15C As shown, CCSAO can be used as a post-reconstruction filter, that is, using the reconstructed samples (after adding the prediction / residual samples and before deblocking) as the input of the classification, compensating the luma / chroma samples before entering the adjacent intra / inter prediction. The CCSAO post-reconstruction filter can reduce the distortion of the current TU samples and provide better prediction for the adjacent intra / inter blocks. Better compression efficiency can be expected through more accurate prediction.

[0405] Encoding algorithm

[0406] In some embodiments, in order to efficiently determine the optimal CCSAO parameters in an image, a hierarchical rate-distortion (RD) optimization algorithm is designed, including 1) a progressive scheme for searching for the best single classifier; 2) a training process for refining the offset value of a classifier; 3) a robust algorithm for effectively assigning appropriate classifiers to different local areas. Typical CCSAO classifiers are: band Y =(Y col ·N Y )>>BD band U =(U col ·N U )>>BD band Y =(V col ·N V )>>BD i=band Y ·(N U ·N V )+band U·N V +band V C′ rec =Clip1(C rec +σ CCSAO [i]) Among them, {Y col ,U col ,V col} are three co-located samples used to classify the current sample; {N Y ,N U ,N V} is the number of bands applied to the Y, U, and V components respectively; BD is the encoding bit depth; C rec and C′ rec are the reconstructed samples before and after applying CCSAO; σ CCSAO [i] is the value of the CCSAO offset applied to the i-th category; Clip1(·) is the clipping function that clips the input to the depth range, i.e., [0,2 BD -1]; >> represents a right shift operation. In the example, the same luminance sample can be selected from 9 candidate positions, while the same chrominance sample is fixed.

[0407] Incremental search scheme

[0408] In some embodiments, to search for N categories (N Y ·N U ·N V ), a multi-stage early termination method is applied. When the classifier with fewer categories cannot improve the RD cost, the classifier with more categories is skipped. Depending on the configuration, multiple breakpoints are set for N categories for early termination. For example, AI: every 4 categories (N Y ·N U ·N V <4,8,12…). RA / LB: Every 16 categories (N Y ·N U ·N V <16,32,48,64…).

[0409] In addition, if N Y Less than N U or N V , or the total number of categories N is greater than the threshold, the classifier will be skipped. The progressive scheme not only adjusts the total bit cost, but also significantly reduces the encoding time. col This process is repeated for each position to determine the best single classifier.

[0410] Offset value refinement

[0411] In some embodiments, for a given classifier, the reconstructed samples in the picture are first classified according to Equation (1). The SAO fast distortion estimate is used to derive the initial offset for each class. The RD cost is further estimated iteratively with smaller offset values until the value is 0. Then, the CTBs with no RD cost improvement are disabled, and the remaining CTBs are retrained to obtain refined offset values. The CTB on-off process is repeated until there is no RD cost improvement for the picture or a threshold count is reached. ΔD=Nh 2 -2hE ΔJ=ΔD+λR

[0412] In some embodiments, for a class, k, s(k), x(k) are the sample positions, original samples and samples before CCSAO, E is the sum of the differences between s(k) and x(k), N is the sample count, ΔD is the delta distortion estimated by applying the offset h, ΔJ is the RD cost, λ is the Lagrange multiplier, and R is the bit cost.

[0413] In some embodiments, the original samples can be true original samples (raw image samples without pre-processing) or motion compensated temporal filter (MCTF) original samples, a classic coding algorithm that pre-processes the original samples before encoding. λ can be the same as that of SAO / ALF or weighted by a factor (depending on the configuration / resolution).

[0414] In some embodiments, the encoder optimizes CCSAO by weighing the total RD cost of all classes.

[0415] In some embodiments, statistics E and N for each category are stored for each CTB for further determining multiple region classifiers.

[0416] Robust multi-classifier assignment

[0417] In some embodiments, to investigate whether the second classifier contributes to the overall image quality, CTBs with enabled CCSAO are sorted in ascending order according to distortion (or according to RD cost, including bit cost).

[0418] In some embodiments, half of the CTBs with less distortion (or a predefined / related ratio, e.g., (setNum-1) / setNum-1) retain the same classifier, while the other half are trained with a new, second classifier. Simultaneously, during CTB switch offset refinement, each CTB can select its best classifier, allowing good classifiers to be propagated to more CTBs. This strategy, characterized by randomness and diffusion, combines the randomness and robustness of parameter decisions. If the current number of classifiers does not further increase the RD cost, further classifiers are skipped.

[0419] Figure 41 A computing environment 4110 is shown coupled to a user interface 4150. The computing environment 4110 may be part of a data processing server. The computing environment 4110 includes a processor 4120, a memory 4130, and an input / output (I / O) interface 4140.

[0420] The processor 4120 generally controls the overall operation of the computing environment 4110, such as operations associated with display, data acquisition, data communication, and image processing. The processor 4120 may include one or more processors for executing instructions to perform all or some of the steps in the above-described method. In addition, the processor 4120 may include one or more modules that facilitate interaction between the processor 4120 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, a graphics processing unit (GPU), etc.

[0421] The memory 4130 is configured to store various types of data to support the operation of the computing environment 4110. The memory 4130 may include predetermined software 4132. Examples of such data include instructions for any application or method operating on the computing environment 1610, video data sets, image data, etc. The memory 4130 may be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0422] The I / O interface 4140 provides an interface between the processor 4120 and peripheral interface modules (e.g., a keyboard, a click wheel, buttons, etc.). The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 4140 may be coupled to an encoder and a decoder.

[0423] Figure 42FIG. 4 is a flow chart illustrating a method for video decoding according to an example of the present disclosure. The method illustrates steps for implementing the above-mentioned classifier C11.

[0424] In step 4201 , the processor 4120 may receive a video signal including a first component and a second component from a video decoder side.

[0425] In some examples, the first component may include one of the following components: a luma component, a first chroma component, or a second chroma component, and the second component may include one of the following components: a luma component, a first chroma component, or a second chroma component. The luma component may be a Y component, the first chroma component may be a U component, and the second chroma component may be a V component. The first chroma component and the second chroma component are interchangeable.

[0426] In step 4202, processor 4120 may receive a plurality of offsets associated with a second component.

[0427] In step 4203, the processor 4120 may obtain a classifier associated with the second component according to the residual sample value of the first component.

[0428] In some examples, the residual sample value may be obtained based on one of the following values: a value of a co-located residual sample of the first component relative to a residual sample of the second component; a value of a neighboring residual sample of the first component relative to a residual sample of the second component; a value obtained by linearly weighting the co-located residual sample and the neighboring residual sample of the first component relative to the residual sample of the second component; or a comparison value of the co-located residual sample and the neighboring residual sample of the first component relative to the residual sample of the second component.

[0429] In some examples, the residual sample value of the first component may include one of the following values: the first residual sample value of the first component after chroma scaling in luma mapping with chroma scaling (LMCS); or the second residual sample value of the first component before LMCS. In addition, the switching syntax element may indicate whether to use the first residual sample value or the second residual sample value, and the switching syntax element may be fixed or signaled at one or more of a sequence parameter set (SPS), an adaptation parameter set (APS), a picture parameter set (PPS), a picture header (PH), a slice header (SH), a region, a coding tree unit (CTU), a coding unit (CU), a subblock, and a sample level.

[0430] In some examples, step 4203 may further include obtaining a classifier associated with the second component according to the residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component.

[0431] In some examples, step 4203 may further include determining a band based on the number of bands partitioned from the range of residual sample values of the first component using the residual sample value of the first component; and obtaining a classifier associated with the second component based on the band. The band may be determined based on the residual sample value of the first component, the number of bands partitioned from the range of residual sample values of the first component, and the sequence bit depth. For example, Class(C11) = (Y0*band_num)>>bit_depth, where Y0 is the co-located / current Y residual sample.

[0432] In some other examples, step 4203 may further include obtaining preprocessed residual sample values by preprocessing the residual sample values, and determining the bands using the preprocessed residual sample values according to the number of bands divided from the range of the residual sample values of the first component. Preprocessing the residual sample values may include one or more of the following steps: obtaining absolute values of the residual sample values; applying clipping to a bit depth to the residual sample values; applying clipping to a specific range to the residual sample values; or applying linear or nonlinear quantization to the residual sample values. For example, preprocessing the residual sample values may include performing an absolute operation on the residual sample values, followed by clipping to a bit depth, and the classifier may be determined as Class(C11)=(Clip1(Y0)*band_num)>>bit_depth.

[0433] In some other examples, step 4203 may further include obtaining preprocessed residual sample values by preprocessing the residual sample values; obtaining a sign of the residual sample values; and obtaining a classifier associated with the second component based on the sign of the residual sample value, the preprocessed residual sample value of the first component, and the number of bands divided from the range of the residual sample values of the first component. In addition, preprocessing the residual sample values includes one or more of the following steps: obtaining an absolute value of the residual sample value; applying clipping to a bit depth to the residual sample values; applying clipping to a specific range to the residual sample values; or applying linear or nonlinear quantization to the residual sample values. For example, preprocessing the residual sample values may include an absolute value operation followed by clipping to a bit depth, and the classifier may be determined using the following method: bandR=(Clip1(|Y0|)*band_num)>>bit_depth signY=Y0>0?0:1 classIdx=signY*2 +bandR. In this example, band_num is different from the band_num in the above classifier C0 (ie, category (C0)). The band_num in this example is different from the value of the residual sample (ie, the residual sample value) Related.

[0434] In step 4204, the processor 4120 may select an offset from a plurality of offsets of the second component according to the classifier.

[0435] In step 4205, the processor 4120 may obtain a modified sample value of the second component based on the selected offset.

[0436] In some examples, the classifier may be a joint classifier determined based on a combination of residual samples and at least one of the reconstructed samples or edge strength. For example, the C11 classifier may be combined with other classifiers (C0 and / or C10) to form a joint classifier. Steps 4204 and 4205 may include selecting an offset from a plurality of offsets of the second component based on the joint classifier; and obtaining a modified sample value of the second component based on the sample offsets, the sample offsets including the selected offset and at least one of the band offset (C0) or the edge offset (C10).

[0437] In some examples, the processor 4120 may store the classifier and the selected offset in a memory or an adaptation parameter set (APS) for future use in subsequent video frames. The offsets / classifier parameters using the C11 residual classifier may be stored in a memory or AP for future reuse. A mechanism may be designed to store the CCSAO offsets of a previous frame for future reuse to save bits sent by signaling for future frames. In some examples, if the selected classifier includes candPos / edge information / coding information, etc., all classifier information may be considered part of the APS offset set and may also be stored therein along with its offset values.

[0438] Figure 43 FIG. 4 is a flow chart illustrating a method for video encoding according to an example of the present disclosure. The method illustrates steps for implementing the above-mentioned classifier C11.

[0439] In step 4301 , the processor 4120 from the video encoder side may determine a video signal including a first component and a second component.

[0440] In some examples, the first component may include one of the following components: a luma component, a first chroma component, or a second chroma component, and the second component may include one of the following components: a luma component, a first chroma component, or a second chroma component. The luma component may be a Y component, the first chroma component may be a U component, and the second chroma component may be a V component. The first chroma component and the second chroma component are interchangeable.

[0441] In step 4302, processor 4120 may determine a plurality of offsets associated with the second component.

[0442] In step 4303, the processor 4120 may obtain a classifier associated with the second component according to the residual sample value of the first component.

[0443] In some examples, the residual sample value may be obtained based on one of the following values: a value of a co-located residual sample of the first component relative to a residual sample of the second component; a value of a neighboring residual sample of the first component relative to a residual sample of the second component; a value obtained by linearly weighting the co-located residual sample and the neighboring residual sample of the first component relative to the residual sample of the second component; or a comparison value of the co-located residual sample and the neighboring residual sample of the first component relative to the residual sample of the second component.

[0444] In some examples, the residual sample value of the first component may include one of the following values: the first residual sample value of the first component after chroma scaling in a luma map with chroma scaling (LMCS); or the second residual sample value of the first component before LMCS. In addition, the switching syntax element may indicate whether to use the first residual sample value or the second residual sample value, and the switching syntax element may be fixed or signaled at one or more of a sequence parameter set (SPS), an adaptation parameter set (APS), a picture parameter set (PPS), a picture header (PH), a slice header (SH), a region, a coding tree unit (CTU), a coding unit (CU), a subblock, and a sample level.

[0445] In some examples, step 4303 may further include determining a band based on the number of bands partitioned from the range of residual sample values of the first component using the residual sample value of the first component; and obtaining a classifier associated with the second component based on the band. The band may be determined based on the residual sample value of the first component, the number of bands partitioned from the range of residual sample values of the first component, and the sequence bit depth. For example, Class(C11) = (Y0*band_num)>>bit_depth, where Y0 is the co-located / current Y residual sample.

[0446] In some other examples, step 4303 may further include obtaining preprocessed residual sample values by preprocessing the residual sample values, and determining the bands using the preprocessed residual sample values according to the number of bands divided from the range of the residual sample values of the first component. Preprocessing the residual sample values may include one or more of the following steps: obtaining an absolute value of the residual sample values; applying clipping to a bit depth to the residual sample values; applying clipping to a specific range to the residual sample values; or applying linear or nonlinear quantization to the residual sample values. For example, preprocessing the residual sample values may include an absolute value operation followed by clipping to a bit depth, and the classifier may be determined as Class(C11)=(Clip1(|Y0|)*band_num)>>bit_depth.

[0447] In some other examples, step 4303 may further include obtaining preprocessed residual sample values by preprocessing the residual sample values; obtaining a sign of the residual sample values; and obtaining a classifier associated with the second component based on the sign of the residual sample value, the preprocessed residual sample value of the first component, and the number of bands divided from the range of the residual sample values of the first component. In addition, preprocessing the residual sample values includes one or more of the following steps: obtaining an absolute value of the residual sample value; applying clipping to a bit depth to the residual sample values; applying clipping to a specific range to the residual sample values; or applying linear or nonlinear quantization to the residual sample values. For example, preprocessing the residual sample values may include an absolute value operation followed by clipping to a bit depth, and the classifier may be determined using the following method: bandR=(Clip1(|Y0|)*band_num)>>bit_depth signY=Y0>0?0:1 classIdx=signY*2 +bandR. In this example, band_num is different from the band_num in the above classifier C0 (ie, category (C0)). The band_num in this example is different from the value of the residual sample (ie, the residual sample value) Related.

[0448] In step 4304, the processor 4120 may select an offset from a plurality of offsets of the second component according to the classifier.

[0449] In step 4305, the processor 4120 may obtain a modified sample value of the second component based on the selected offset.

[0450] In some examples, the classifier may be a joint classifier determined based on a combination of residual samples and at least one of the reconstructed samples or edge strength. For example, the C11 classifier may be combined with other classifiers (C0 and / or C10) to form a joint classifier. Steps 4204 and 4205 may include selecting an offset from a plurality of offsets of the second component based on the joint classifier; and obtaining a modified sample value of the second component based on the sample offsets, the sample offsets including the selected offset and at least one of the band offset (C0) or the edge offset (C10).

[0451] In some examples, the processor 4120 may store the classifier and the offset selected in a memory or adaptive parameter set (APS) for future use in subsequent video frames. The offsets / classifier parameters using the C11 residual classifier may be stored in a memory or AP for future reuse. A mechanism may be designed to store the CCSAO offsets of a previous frame for future reuse to save bits sent by signal for future frames. In some examples, if the selected classifier includes candPos / edge information / coding information, etc., all classifier information may be considered part of the APS offset set and may also be stored therein along with its offset values.

[0452] In an embodiment, a non-transitory computer-readable storage medium including, for example, a plurality of programs in a memory 4130 and / or storing a bit stream generated by the above encoding method or a bit stream to be decoded by the above decoding method is also provided. The plurality of programs can be executed by a processor 4120 in a computing environment 4110 to perform the above method. In one example, the plurality of programs can be executed by a processor 4120 in a computing environment 4110 to (for example, from Figure 2 The video encoder 20 in the computing environment 4110 receives a bit stream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 4120 in the computing environment 4110 to perform the above-mentioned decoding method according to the received bit stream or data stream. In another example, the multiple programs can be executed by the processor 4120 in the computing environment 4110 to perform the above-mentioned encoding method to encode the video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bit stream or data stream, and can also be executed by the processor 4120 in the computing environment 4110 to (e.g., to Figure 3 Alternatively, a non-transitory computer readable storage medium may store the bit stream or data stream generated by the encoder (e.g., Figure 2 The video encoder 20 in FIG. 1 generates a video signal for use by a decoder (eg, Figure 3 A bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or associated one or more syntax elements, etc.) used by the video decoder 30 in the video decoder 30 when decoding video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0453] In an embodiment, a bit stream generated by the above encoding method or a bit stream to be decoded by the above decoding method is provided. In an embodiment, a bit stream including coded video information generated by the above encoding method or coded video information to be decoded by the above decoding method is provided.

[0454] In an embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 4120); and a non-transitory computer-readable storage medium or memory 4130 having stored therein a plurality of programs that can be executed by the one or more processors, wherein the one or more processors are configured to perform the above-mentioned method when executing the plurality of programs.

[0455] In an embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, wherein the bitstream includes encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method. In an embodiment, a computer program product is also provided, including, for example, a plurality of programs in a memory 4130, which can be executed by a processor 4120 in a computing environment 4110 to perform the above method. For example, the computer program product can include a non-transitory computer-readable storage medium.

[0456] In an embodiment, the computing environment 4110 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.

[0457] In an embodiment, a method for storing a bitstream is further provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream comprises encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method.

[0458] In an embodiment, a method for transmitting a bit stream generated by the above encoder is also provided. In an embodiment, a method for receiving a bit stream to be decoded by the above decoder is also provided.

[0459] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.

[0460] Unless otherwise specifically stated, the order of steps of the method according to the present disclosure is intended to be illustrative only, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but can be changed according to actual circumstances. In addition, at least one of the steps of the method according to the present disclosure can be adjusted, combined, or deleted according to actual needs.

[0461] The examples are chosen and described in order to explain the principles of the present disclosure and to enable others skilled in the art to understand the various embodiments of the present disclosure and to best utilize the basic principles and various embodiments with various modifications as are suited to the particular use contemplated. Therefore, it will be understood that the scope of the present disclosure is not limited to the specific examples of the embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the present disclosure.

Claims

1. A method for video decoding, comprising: receiving, by a decoder, a video signal comprising a first component and a second component; receiving, by the decoder, a plurality of offsets associated with the second component; Obtaining, by the decoder, a classifier associated with the second component based on the residual sample value of the first component; selecting, by the decoder, an offset from the plurality of offsets of the second component based on the classifier; as well as A modified sample value of the second component is obtained by the decoder based on the selected offset.

2. The method according to claim 1, wherein The residual sample value is obtained according to one of the following values: the value of the co-located residual sample of the first component relative to the residual sample of the second component; The values of the neighboring residual samples of the first component relative to the residual samples of the second component; a value obtained by linearly weighting the co-located residual sample and the adjacent residual samples of the first component with respect to the residual sample of the second component; or Comparison values of the co-located residual samples and adjacent residual samples of the first component relative to the residual samples of the second component.

3. The method according to claim 1, wherein The residual sample value of the first component includes one of the following values: a first residual sample value of the first component after chroma scaling in a luma mapping with chroma scaling (LMCS); or The second residual sample value of the first component before LMCS, The switching syntax element indicates whether to use the first residual sample value or the second residual sample value, and the switching syntax element is fixed or signaled at one or more of a sequence parameter set (SPS), an adaptation parameter set (APS), a picture parameter set (PPS), a picture header (PH), a slice header (SH), a region, a coding tree unit (CTU), a coding unit (CU), a sub-block, and a sample level.

4. The method according to claim 1, wherein Obtaining the classifier associated with the second component according to the residual sample value of the first component includes: The classifier associated with the second component is obtained according to the residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component.

5. The method according to claim 4, further comprising: Obtaining preprocessed residual sample values by preprocessing the residual sample values; The step of obtaining the classifier associated with the second component according to the residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component comprises: The classifier associated with the second component is obtained according to the preprocessed residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component.

6. The method according to claim 5, wherein: Preprocessing the residual sample values includes one or more of the following steps: Obtaining the absolute value of the residual sample value; Applying clipping to the depth to the residual sample value; Applying clipping to a specific range to the residual sample values; or Linear or non-linear quantization is applied to the residual sample values.

7. The method according to claim 4, wherein: Obtaining the classifier associated with the second component according to the residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component includes: determining a band according to the number of bands divided from a range of the residual sample values of the first component using the residual sample value of the first component; and The classifier associated with the second component is obtained based on the band.

8. The method according to claim 7, wherein: The bands are determined according to the residual sample values of the first component, the number of bands divided from a range of the residual sample values of the first component, and a sequence bit depth.

9. The method according to claim 4, further comprising: Obtaining preprocessed residual sample values by preprocessing the residual sample values; Obtaining the sign of the residual sample value; as well as The step of obtaining the classifier associated with the second component according to the residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component comprises: The classifier associated with the second component is obtained according to the sign of the residual sample value, the preprocessed residual sample value of the first component, and the number of bands divided from the range of the residual sample value of the first component.

10. The method according to claim 9, wherein: Preprocessing the residual sample values includes one or more of the following steps: Obtaining the absolute value of the residual sample value; Applying clipping to the depth to the residual sample value; Applying clipping to a specific range to the residual sample values; or Linear or non-linear quantization is applied to the residual sample values.

11. The method according to claim 1, wherein The classifier is a joint classifier determined according to a combination of residual samples and at least one of reconstructed samples or edge strength.

12. The method according to claim 11, wherein selecting the offset from the plurality of offsets of the second component according to the classifier; And obtaining a modified sample value of the second component based on the selected offset includes: selecting the offset from the plurality of offsets of the second component according to the joint classifier; as well as A modified sample value of the second component is obtained based on a sample offset comprising the selected offset and at least one of a band offset or an edge offset.

13. The method according to claim 1, further comprising: The classifier and the selected offset are stored by the decoder in a memory or an Adaptation Parameter Set (APS) for future use in subsequent video frames.

14. A method for video encoding, comprising: determining, by an encoder, a video signal comprising a first component and a second component; determining, by the encoder, a plurality of offsets associated with the second component; Obtaining, by the encoder, a classifier associated with the second component according to the residual sample value of the first component; selecting, by the encoder, an offset from the plurality of offsets of the second component based on the classifier; as well as A modified sample value of the second component is obtained by the encoder based on the selected offset.

15. The method according to claim 14, wherein The residual sample value is obtained according to one of the following values: the value of the co-located residual sample of the first component relative to the residual sample of the second component; The values of the neighboring residual samples of the first component relative to the residual samples of the second component; a value obtained by linearly weighting the co-located residual sample and the adjacent residual samples of the first component with respect to the residual sample of the second component; or Comparison values of the co-located residual samples and adjacent residual samples of the first component relative to the residual samples of the second component.

16. The method according to claim 14, wherein The residual sample value of the first component includes one of the following values: a first residual sample value of the first component after chroma scaling in a luma mapping with chroma scaling (LMCS); or The second residual sample value of the first component before LMCS, The first residual sample value or the second residual sample value is fixed, or is sent through a signal at one or more of a sequence parameter set (SPS), an adaptation parameter set (APS), a picture parameter set (PPS), a picture header (PH), a slice header (SH), a region, a coding tree unit (CTU), a coding unit (CU), a sub-block, and a sample level.

17. The method according to claim 14, wherein: Obtaining the classifier associated with the second component according to the residual sample value of the first component includes: The classifier associated with the second component is obtained according to the residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component.

18. The method according to claim 17, further comprising: Obtaining preprocessed residual sample values by preprocessing the residual sample values; The step of obtaining the classifier associated with the second component according to the residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component comprises: The classification associated with the second component is obtained according to the preprocessed residual sample value of the first component and the number of bands divided from a range of the residual sample value of the first component.

19. The method according to claim 18, wherein Preprocessing the residual sample values includes one or more of the following steps: Obtaining the absolute value of the residual sample value; Applying clipping to the depth to the residual sample value; Applying clipping to a specific range to the residual sample values; or Linear or non-linear quantization is applied to the residual sample values.

20. The method according to claim 17, wherein Obtaining the classifier associated with the second component according to the residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component includes: determining a band according to the number of bands divided from a range of the residual sample values of the first component using the residual sample value of the first component; and The classifier associated with the second component is obtained based on the band.

21. The method according to claim 20, wherein The bands are determined according to the residual sample values of the first component, the number of bands divided from a range of the residual sample values of the first component, and a sequence bit depth.

22. The method of claim 17, further comprising: Obtaining preprocessed residual sample values by preprocessing the residual sample values; Obtaining the sign of the residual sample value; as well as The step of obtaining the classifier associated with the second component according to the residual sample value of the first component and the number of bands divided from the range of the residual sample value of the first component comprises: The classifier associated with the second component is obtained according to the sign of the residual sample value, the preprocessed residual sample value of the first component, and the number of bands divided from the range of the residual sample value of the first component.

23. The method according to claim 22, wherein Preprocessing the residual sample values includes one or more of the following steps: Obtaining the absolute value of the residual sample value; Applying clipping to the depth to the residual sample value; Applying clipping to a specific range to the residual sample values; or Linear or non-linear quantization is applied to the residual sample values.

24. The method according to claim 14, wherein The classifier is a joint classifier determined according to a combination of residual samples and at least one of reconstructed samples or edge strength.

25. The method according to claim 24, wherein selecting the offset from the plurality of offsets of the second component according to the classifier; And obtaining a modified sample value of the second component based on the selected offset includes: selecting the offset from the plurality of offsets of the second component according to the joint classifier; as well as A modified sample value of the second component is obtained based on a sample offset comprising the selected offset and at least one of a band offset or an edge offset.

26. The method of claim 14, further comprising: The classifier and the selected offset are stored by the encoder in a memory or adaptive parameter set (APS) for future use in subsequent video frames.

27. An apparatus for video decoding, comprising: one or more processors; as well as a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, Wherein, the one or more processors are configured to perform the method according to any one of claims 1 to 13 when executing the instructions.

28. An apparatus for video encoding, comprising: one or more processors; as well as a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, Wherein, when executing the instructions, the one or more processors are configured to perform the method according to any one of claims 14 to 26.

29. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 1 to 26.

30. A non-transitory computer-readable storage medium for storing a bit stream to be decoded by the method of any one of claims 1 to 13.

31. A non-transitory computer-readable storage medium for storing a bitstream generated by the method of any one of claims 14 to 26.