Enhanced coding in inter-component sample adaptive offsets.

CCSAO methods enhance video coding efficiency by modifying chroma samples based on weighted luma and chroma relationships, addressing inefficiencies in existing technologies and improving bit rate utilization and video quality.

JP7827829B2Active Publication Date: 2026-03-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-17
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing luma and chroma components, leading to suboptimal bit rate usage and video quality degradation.

Method used

Implementing cross-component sample adaptive offset (CCSAO) methods that utilize weighted sample values from adjacent luma and chroma components to determine sample offsets and modify chroma samples, enhancing coding efficiency by exploiting cross-component relationships.

Benefits of technology

Improves coding efficiency by reducing bit rates and maintaining video quality through optimized compression of luma and chroma components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007827829000107
    Figure 0007827829000107
  • Figure 0007827829000108
    Figure 0007827829000108
  • Figure 0007827829000109
    Figure 0007827829000109
Patent Text Reader

Abstract

An electronic device performs a method for decoding video data, the method including: receiving a picture frame from a video signal, the picture frame including a first component and a second component; determining a classifier for each sample of the second component using a set of weighted sample values ​​from a first set of samples of the first component associated with the respective sample of the second component and a second set of samples of the second component associated with the respective sample of the second component, the first set of samples and the second set of samples being sequential, adjacent, and current samples for the respective sample of the second component; determining a sample offset for the respective sample of the second component in accordance with the classifier; and modifying the respective sample of the second component based on the determined sample offset.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 235,090, entitled "CROSS-COMPONENT SAMPLE ADAPTIVE OFFSET," filed August 19, 2021, the entire contents of which are incorporated herein by reference.

[0002] This application relates generally to video coding and compression, and more particularly to methods and apparatus for improving both luma and chroma coding efficiency. [Background technology]

[0003] Digital video is supported by a variety of electronic devices, including digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video gaming consoles, smartphones, video teleconferencing devices, and video streaming devices. Electronic devices transmit and receive digital video data over communication networks and / or store it on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding may be used to compress video data according to one or more video coding standards before it is communicated or stored. For example, video coding standards include Versatile Video Coding (VVC), Joint Exploration test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Expert Group (MPEG) coding. AOMei Video 1 (AV1) was developed as a successor to the predecessor standard VP9. Audio Video Coding (AVS) refers to digital audio and video compression standards, and is another series of video compression standards. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. The goal of video coding is to compress video data into a form that avoids or minimizes degradation of video quality while using lower bit rates. Summary of the Invention

[0004] This application describes implementation examples relating to the encoding and decoding of video data, and more particularly, to methods and apparatus for improving the coding efficiency of both luma and chroma components, including improving coding efficiency by exploring cross-component relationships between the luma and chroma components.

[0005] According to a first aspect of the present application, a method for decoding a video signal includes receiving a picture frame from the video signal, the picture frame including a first component and a second component; determining a classifier for each sample of the second component using a set of weighted sample values ​​from a first set of samples of the first component associated with the respective sample of the second component and a second set of samples of the second component associated with the respective sample of the second component, wherein the first set of samples of the first component includes an array sample of the first component for the respective sample of the second component and an adjacent sample of the array sample of the first component, and the second set of samples of the second component includes a current sample of the second component for the respective sample of the second component and an adjacent sample of the current sample of the second component; determining a sample offset for each sample of the second component according to the classifier; and modifying each sample of the second component based on the determined sample offset.

[0006] According to a second aspect of the present application, a method for decoding a video signal includes receiving a picture frame from the video signal, the picture frame including a first component, a second component, and a third component; determining a classifier for each sample of the second component using a set of weighted sample values ​​from a first set of samples of the first component associated with each sample of the second component and a third set of samples of the third component associated with each sample of the second component, wherein the first set of samples of the first component includes an array sample of the first component for each sample of the second component and an adjacent sample of the array sample of the first component, and the third set of samples of the third component includes an array sample of the third component for each sample of the second component and an adjacent sample of the array sample of the third component; determining a sample offset for each sample of the second component according to the classifier; and modifying each sample of the second component based on the determined sample offset.

[0007] According to a third aspect of the present application, an electronic device includes one or more processing units, a memory, and a plurality of programs stored in the memory, the programs, when executed by the one or more processing units, causing the electronic device to perform the method for encoding a video signal described above.

[0008] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by an electronic device having one or more processing units, the programs, when executed by the one or more processing units, causing the electronic device to perform the method for encoding a video signal as described above.

[0009] According to a fifth aspect of the present application, a computer-readable storage medium stores a bitstream including instructions that, when executed, cause a decoding device to perform the method for decoding a video signal described above.

[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure.

[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram illustrating an example system for video block encoding and decoding according to some implementations of the present disclosure.

[0013] [Figure 2] FIG. 1 is a block diagram illustrating an example video encoder according to some implementations of the present disclosure.

[0014] [Figure 3] FIG. 1 is a block diagram illustrating an example video decoder according to some implementations of the present disclosure.

[0015] [Figures 4A-4E] A block diagram illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes in accordance with some implementations of the present disclosure.

[0016] [Figure 4F] FIG. 1 is a block diagram showing the intra-modes defined in VVC.

[0017] [Figure 4G] FIG. 1 is a block diagram illustrating multiple reference lines for intra prediction.

[0018] [Figure 5A] FIG. 1 is a block diagram illustrating four gradient patterns used in sample adaptive offset (SAO) according to some implementations of the present disclosure.

[0019] [Figure 5B] FIG. 10 is a block diagram illustrating the naming convention for samples surrounding a central sample according to some implementations of the present disclosure.

[0020] [Figure 6A] FIG. 1 is a block diagram illustrating a CCSAO system and process applied to chroma samples by some example implementations of the present disclosure and using DBF Y as input.

[0021] [Figure 6B] FIG. 1 is a block diagram illustrating a CCSAO system and process applied to luma and chroma samples by some example implementations of the present disclosure, using DBF Y / Cb / Cr as input.

[0022] [Figure 6C] FIG. 1 is a block diagram illustrating a CCSAO system and process that can function independently according to some implementations of the present disclosure.

[0023] [Figure 6D] FIG. 1 is a block diagram illustrating a CCSAO system and process that may be applied recursively (2 or N times) with the same or different offsets by some implementations of the present disclosure.

[0024] [Figure 6E] FIG. 1 is a block diagram illustrating a CCSAO system and process applied in parallel to Enhanced Sample Adaptive Offset (ESAO) in the AVS standard according to some implementations of the present disclosure.

[0025] [Figure 6F] FIG. 1 is a block diagram illustrating a CCSAO system and process applied after SAO according to some implementations of the present disclosure.

[0026] [Figure 6G]FIG. 10 is a block diagram illustrating that some implementations of the present disclosure allow the CCSAO system and process to function independently without a CCALF.

[0027] [Figure 6H] FIG. 1 is a block diagram illustrating a CCSAO system and process applied in parallel to a cross-component adaptive loop filter (CCALF) according to some implementations of the present disclosure.

[0028] [Figure 6I] FIG. 1 is a block diagram illustrating a system and process of a CCSAO applied in parallel to an SAO and a BIF according to some implementations of the present disclosure.

[0029] [Figure 6J] FIG. 10 is a block diagram illustrating a system and process of a CCSAO applied in parallel to a BIF by exchanging an SAO according to some implementations of the present disclosure.

[0030] [Figure 7] FIG. 10 is a block diagram illustrating a sample process for using a CCSAO according to some implementations of the present disclosure.

[0031] [Figure 8] FIG. 10 is a block diagram illustrating how the CCSAO process is interleaved with vertical and horizontal deblocking filters (DBFs) according to some implementations of the present disclosure.

[0032] [Figure 9] 1 is a flow diagram illustrating an example process for decoding a video signal using cross-component correlation according to some implementations of the present disclosure.

[0033] [Figure 10A] FIG. 10 is a block diagram illustrating a classifier that uses different luma (or chroma) sample positions for C0 classification according to some implementations of the present disclosure.

[0034] [Figure 10B] 1A and 1B are diagrams illustrating several examples of different shapes for luma candidates according to some implementations of the present disclosure.

[0035] [Figure 11] FIG. 10 is a block diagram of a sample process illustrating that all sequence and adjacent luma / chroma samples may be provided to CCSAO classification according to some implementations of the present disclosure.

[0036] [Figure 12A] 10A-10C illustrate an exemplary classifier by replacing an array luma sample value with a value obtained by weighting the array and neighboring luma samples, according to some implementations of the present disclosure.

[0037] [Figure 12B] FIG. 1 illustrates a subsampled Laplacian calculation according to some example implementations of the present disclosure.

[0038] [Figure 13] 10 is a block diagram illustrating how CCSAO is applied and other in-loop filters have different clipping combinations, according to some implementations of the present disclosure.

[0039] [Figure 14A] A block diagram showing that, according to some implementation examples of the present disclosure, CCSAO is not applied to the current chroma (luma) sample if the array used for classification and any of the neighboring luma (chroma) samples are outside the current picture.

[0040] [Figure 14B] FIG. 10 is a block diagram illustrating that, according to some example implementations of the present disclosure, CCSAO is applied to a current luma or chroma sample when either the array used for classification and the neighboring luma or chroma sample are outside the current picture.

[0041] [Figure 14C] A block diagram illustrating that, in some implementations of the present disclosure, CCSAO is not applied to the current chroma sample if the corresponding selected array or neighboring luma sample used for classification is outside the virtual space defined by the virtual boundary (VB).

[0042] [Figure 15] 10A and 10B illustrate how repetition or mirror padding is applied to luma samples that fall outside the virtual boundary, according to some implementations of the present disclosure.

[0043] [Figure 16] FIG. 10 illustrates that, according to some implementations of the present disclosure, an additional luma line buffer is required when all nine arrays and adjacent luma samples are used for classification.

[0044] [Figure 17] FIG. 10 illustrates that in some implementations of the present disclosure, in AVS, the nine luma candidate CCSAOs intersecting the VB can be increased by two additional luma line buffers.

[0045] [Figure 18A] FIG. 10 illustrates that in VVC, nine luma candidate CCSAOs intersecting a VB can be increased by one additional luma line buffer according to some implementation examples of the present disclosure.

[0046] [Figure 18B] FIG. 10 illustrates that, according to some implementations of the present disclosure, when an array or neighboring chroma samples is used to classify the current luma sample, the selected chroma candidate may cross the VB and require additional chroma line buffers.

[0047] [Figures 19A-19C]FIG. 10 illustrates that, in some implementation examples of the present disclosure, in AVS and VVC, CCSAO is disabled for a chroma sample if any of the luma candidates for the chroma sample crosses VB (is outside the current chroma sample VB).

[0048] [Figures 20A-20C] FIG. 10 illustrates that, in some implementation examples of the present disclosure, in AVS and VVC, CCSAO is enabled using repeated padding on chroma samples when any of the luma candidates for a chroma sample crosses the VB (is outside the current chroma sample VB).

[0049] [Figures 21A-21C] FIG. 10 illustrates that, in some implementation examples of the present disclosure, in AVS and VVC, CCSAO is enabled using mirror padding on chroma samples when any of the luma candidates for the chroma samples crosses the VB (is outside the current chroma sample VB).

[0050] [Figures 22A-22B] FIG. 10 illustrates how CCSAO is enabled using two-sided symmetric padding for different CCSAO sample shapes according to some implementations of the present disclosure.

[0051] [Figure 23] 10A and 10B illustrate the limitation of using a limited number of luma candidates for classification in accordance with some implementations of the present disclosure.

[0052] [Figure 24] A diagram illustrating that, in some implementations of the present disclosure, the CCSAO application region is not aligned to a coding tree block (CTB) / coding tree unit (CTU) boundary.

[0053] [Figure 25]A diagram showing that the CCSAO application area frame partition can be fixed by CCSAO parameters according to some implementation examples of the present disclosure.

[0054] [Figure 26] A diagram showing that, according to some implementation examples of the present disclosure, the CCSAO application area may be a binary tree (BT) / quadtree (QT) / ternary tree (TT) split from the frame / slice / CTB level.

[0055] [Figure 27] A block diagram showing multiple classifiers used and switched at different levels within a picture frame in accordance with some implementations of the present disclosure.

[0056] [Figure 28] A block diagram showing that CCSAO application area partitions are dynamic and can be switched at the picture level according to some implementation examples of the present disclosure.

[0057] [Figure 29] FIG. 10 illustrates that some implementations of the present disclosure allow the CCSAO classifier to consider current or inter-component coding information.

[0058] [Figure 30] FIG. 10 is a block diagram illustrating the SAO classification method disclosed in this disclosure acting as a post prediction filter, according to some implementations of the present disclosure.

[0059] [Figure 31] FIG. 10 is a block diagram illustrating that, for a post prediction SAO filter, each component can use the current sample and neighboring samples for classification, according to some implementations of the present disclosure.

[0060] [Figure 32]FIG. 10 is a block diagram illustrating the SAO classification method disclosed in the present disclosure acting as a post reconstruction filter, according to some implementations of the present disclosure.

[0061] [Figure 33] 1 is a flow diagram illustrating an example process for decoding a video signal using inter-component correlation according to some implementations of the present disclosure.

[0062] [Figure 34] FIG. 1 illustrates a computing environment coupled to a user interface in accordance with some implementations of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0063] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used without departing from the scope of the claims, and that the subject matter may be practiced without these specific details. For example, it will be apparent to those skilled in the art that the subject matter presented herein may be implemented in many types of electronic devices having digital video capabilities.

[0064] It should be noted that terms such as "first," "second," and the like, used in this description, the claims of this disclosure, and the accompanying drawings, are used to distinguish between objects and are not intended to describe any particular order or sequence. It should be understood that the data used in this manner may be interchanged under appropriate conditions, and that the embodiments of the disclosure described herein may therefore be implemented in orders other than those illustrated in the accompanying drawings or described in this disclosure.

[0065] The first generation of AVS standards includes the Republic of China national standard "Information Technology, Advanced Audio Video Coding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video" (known as AVS+). Compared with the MPEG-2 standard, this can provide approximately 50% bit rate savings with the same perceptual quality. The second generation of AVS standards includes a series of Republic of China national standards "Information Technology, Efficient Multimedia Coding" (known as AVS2), which primarily targets the transmission of additional HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. Meanwhile, the video part of the AVS2 standard has been submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for application. The AVS3 standard is a new generation video coding standard for UHD video that aims to surpass the coding efficiency of the latest international standard, HEVC, and provides approximately 30% bitrate savings compared to the HEVC standard. The AVS3-P2 baseline was completed at the 68th AVS Conference in March 2019, providing approximately 30% bitrate savings compared to the HEVC standard. Currently, a single reference software called the High Performance Model (HPM) is maintained by the AVS group to demonstrate a reference implementation of the AVS3 standard. Like HEVC, the AVS3 standard is built on a block-based hybrid video coding framework.

[0066] 1 is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12, which generates and encodes video data for subsequent decoding by a destination device 14. Source device 12 and destination device 14 may comprise any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0067] In some implementations, destination device 14 may receive encoded video data to be decoded via link 16. Link 16 may comprise any type of communication medium or device capable of moving encoded video data from source device 12 to destination device 14. In one example, link 16 may comprise a communication medium to enable source device 12 to transmit encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network, e.g., the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from source device 12 to destination device 14.

[0068] In some other implementations, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by destination device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessible data storage media, such as a hard drive, Blu-ray disc, digital versatile disc (DVD), compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In a further example, storage device 32 may correspond to a file server or another intermediate storage device capable of holding the encoded video data generated by source device 12. Destination device 14 may access the stored video data from storage device 32 via streaming or download. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to destination device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network-attached storage (NAS) devices, or local disk drives. Destination device 14 can access the encoded video data over any standard data connection suitable for accessing the encoded video data stored on the file server, including wireless channels (e.g., Wireless Fidelity (Wi-Fi) connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or a combination of both. Transmission of the encoded video data from storage device 32 can be a streaming transmission, a download transmission, or a combination of both.

[0069] As shown in FIG. 1 , source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 may include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, if video source 18 is a video camera in a security surveillance system, source device 12 and destination device 14 may form a camera phone or video phone. However, implementations described herein may apply to video coding in general and may be applicable to wireless and / or wired applications.

[0070] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored on storage device 32 for later access by destination device 14 or another device for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.

[0071] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem and may receive encoded video data over link 16. The encoded video data communicated over link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.

[0072] In some implementations, destination device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may comprise any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0073] Video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as VVC, HEVC, MPEG-4 Part 10, AVC, AVS, or extensions to such standards. It should be understood that the present application is not limited to any particular video encoding / decoding standard and may apply to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data in accordance with any of these current or future standards. Similarly, it is generally contemplated that video decoder 30 of destination device 14 may be configured to decode video data in accordance with any of these current or future standards.

[0074] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, an electronic device can store instructions for the software on a suitable non-transitory computer-readable medium and execute these instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.

[0075] FIG. 2 is a block diagram illustrating an example video encoder 20 according to some implementations described herein. Video encoder 20 can perform intra- and inter-predictive coding of video blocks within a video frame. Intra-predictive coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-predictive coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. Note that the term "frame" may be used synonymously with the terms "image" or "picture" in the field of video coding.

[0076] As shown in FIG. 2, video encoder 20 includes video data memory 40, prediction processing unit 41, decoded picture buffer (DPB) 64, adder 50, transform processing unit 52, quantization unit 54, and entropy coding unit 56. Prediction processing unit 41 further includes motion estimation unit 42, motion compensation unit 44, partition unit 45, intra-prediction processing unit 46, and intra-block copy (BC) unit 48. In some implementations, video encoder 20 also includes inverse quantization unit 58 for video block reconstruction, inverse transform processing unit 60, and adder 62. An in-loop filter 63, such as a deblocking filter, may be placed between adder 62 and DPB 64 to filter block boundaries and remove blockiness artifacts from the reconstructed video. In addition to the deblocking filter for filtering the output of adder 62, another in-loop filter, such as a sample adaptive offset (SAO) filter and / or an adaptive in-loop filter (ALF), may also be used. In some examples, the in-loop filter may be omitted, and the decoded video blocks may be provided directly to DPB 64 by summer 62. Video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided among one or more of the fixed or programmable hardware units shown.

[0077] Video data memory 40 may store video data to be encoded by components of video encoder 20. As shown in FIG. 1, the video data in video data memory 40 may be obtained, for example, from video source 18. DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use by video encoder 20 (e.g., in intra- or inter-predictive coding modes) in encoding the video data. Video data memory 40 and DPB 64 may be formed by any of a variety of memory devices. In various examples, video data memory 40 may be on-chip with or off-chip relative to other components of video encoder 20.

[0078] As shown in FIG. 2, after receiving the video data, partition unit 45 in prediction processing unit 41 divides the video data into video blocks. This division may also include dividing the video frame into slices, tiles (e.g., sets of video blocks), or other larger coding units (CUs) according to a predefined division structure, such as a quadtree (QT) structure associated with the video data. A video frame may be, or may be viewed as, a two-dimensional array or matrix of samples having sample values. The samples in the array may also be referred to as pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. The video frame may be divided into multiple video blocks, for example, by using QT division. Again, a video block may be, or may be viewed as, a two-dimensional array or matrix of samples having sample values, but with smaller dimensions than a video frame. The number of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. A video block may be further divided into one or more block partitions or sub-blocks (which can again form blocks), e.g., by iteratively using QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof. Note that, as used herein, the term “block” or “video block” can refer to a portion of a frame or picture, particularly a rectangular (square or non-square) portion. With reference to HEVC and VVC, for example, a block or video block can be or correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU), and / or can be or correspond to a corresponding block, e.g., a coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB), and / or can correspond to a sub-block.

[0079] Prediction processing unit 41 may select one of multiple possible predictive coding modes, such as one of multiple intra-predictive coding modes or one of multiple inter-predictive coding modes, for the current video block based on the error result (e.g., coding rate and distortion level). Prediction processing unit 41 may provide the resulting intra- or inter-predictive coded block to summer 50 to generate a residual block, or to summer 62 to reconstruct the coded block for later use as part of a reference frame. Prediction processing unit 41 also provides syntax elements such as motion vectors, intra-mode indicators, partition information, and other such syntax information to entropy coding unit 56.

[0080] To select an appropriate intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-predictive coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-predictive coding of the current video block relative to one or more predictive blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple coding passes, e.g., to select an appropriate coding mode for each block of video data.

[0081] In some implementations, motion estimation unit 42 determines the inter-prediction mode for a current video frame by generating motion vectors that indicate the displacement of video blocks in the current video frame relative to predictive blocks in a reference video frame according to a predetermined pattern in a sequence of video frames. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors to estimate motion for video blocks. The motion vectors may indicate, for example, the displacement of video blocks in a current video frame or picture relative to predictive blocks in a reference frame relative to the current block being coded in the current frame. The predetermined pattern may designate a video frame in the sequence as a P-frame or a B-frame. Intra BC unit 48 may determine, or utilize motion estimation unit 42 to determine, vectors, e.g., block vectors, for intra BC coding, similar to the determination of motion vectors by motion estimation unit 42 for inter prediction.

[0082] A prediction block for a video block may be or correspond to a block or reference block in a reference frame that is deemed to closely match the video block to be coded in terms of pixel difference, which may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metric. In some implementations, video encoder 20 may calculate values ​​for sub-integer pixel locations of the reference frame stored in DPB64. For example, video encoder 20 may interpolate values ​​for quarter-pixel locations, eighth-pixel locations, or other fractional pixel locations of the reference frame. Thus, motion estimation unit 42 may perform motion searches for full pixel locations and fractional pixel locations and output motion vectors with fractional pixel accuracy.

[0083] Motion estimation unit 42 calculates motion vectors for video blocks in inter-predictive coded frames by comparing the position of the video block with the position of a predictive block in a reference frame selected from a first reference frame list (List0) or a second reference frame list (List1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vectors to motion compensation unit 44 and then to entropy coding unit 56.

[0084] Motion compensation performed by motion compensation unit 44 may involve fetching or generating a predictive block based on the motion vector determined by motion estimation unit 42. Upon receiving the motion vector for the current video block, motion compensation unit 44 may locate the predictive block pointed to by the motion vector in one of the reference frame lists, retrieve the predictive block from DPB 64, and forward the predictive block to summer 50. Summer 50 then forms a residual video block of pixel difference values ​​by subtracting pixel values ​​of the predictive block provided by motion compensation unit 44 from pixel values ​​of the current video block being coded. The pixel difference values ​​forming the residual video block may include luma or chroma difference components, or both. Motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use by video decoder 30 in decoding the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vector used to identify the predictive block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes.

[0085] In some implementations, the intra BC unit 48 can generate a vector and fetch a predictive block, similar to that described above with reference to the motion estimation unit 42 and the motion compensation unit 44, except that the predictive block is within the same frame as the current block being coded, and the vector is referred to as a block vector rather than a motion vector. In particular, the intra BC unit 48 can determine an intra prediction mode to use to encode the current block. In some examples, the intra BC unit 48 can encode the current block using various intra prediction modes, e.g., during separate coding passes, and test their performance through rate-distortion analysis. The intra BC unit 48 can then select a suitable intra prediction mode from among the various tested intra prediction modes and use and generate an intra mode indicator accordingly. For example, the intra BC unit 48 can calculate rate-distortion values ​​for the various tested intra prediction modes using rate-distortion analysis, and select and use the intra prediction mode with the best rate-distortion characteristics from among the tested modes as the suitable intra prediction mode. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that was coded to result in the coded block, as well as the bit rate (i.e., number of bits) used to result in the coded block. Intra BC unit 48 can calculate ratios from the distortion and rate for various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for that block.

[0086] In other examples, intra BC unit 48 may perform such functions for intra BC prediction, in whole or in part, using motion estimation unit 42 and motion compensation unit 44, depending on the implementation described herein. In either case, for intra block copying, the predictive block may be a block deemed to closely match the block to be coded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metric, and identification of the predictive block may include calculation of values ​​for sub-integer pixel locations.

[0087] Regardless of whether the predictive block is from the same frame according to intra prediction or from a different frame according to inter prediction, video encoder 20 may form a residual video block by subtracting pixel values ​​of the predictive block from pixel values ​​of the current video block being coded to form pixel difference values. The pixel difference values ​​that form the residual video block may include both luma and chroma component differences.

[0088] Intra-prediction processing unit 46 may intra-predict the current video block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44 or intra-block copy prediction performed by intra BC unit 48, as described above. In particular, intra-prediction processing unit 46 may determine an intra-prediction mode to use to encode the current block. In doing so, intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, e.g., during separate encoding passes, and intra-prediction processing unit 46 (or, in some examples, a mode selection unit) may select and use an appropriate intra-prediction mode from the tested intra-prediction modes. Intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to entropy coding unit 56. Entropy coding unit 56 may encode the information indicating the selected intra-prediction mode in the bitstream.

[0089] After prediction processing unit 41 determines a predictive block for a current video block via inter- or intra-prediction, adder 50 forms a residual video block by subtracting the predictive block from the current video block. The residual video data in the residual block, which may be included in one or more TUs, is provided to transform processing unit 52. Transform processing unit 52 converts the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0090] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of a matrix containing the quantized transform coefficients. Alternatively, entropy coding unit 56 may perform the scan.

[0091] Following quantization, entropy coding unit 56 entropy codes the quantized transform coefficients into a video bitstream, using, for example, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The coded bitstream may then be transmitted to video decoder 30, or may be stored in storage device 32 for later transmission to or retrieval by video decoder 30, as shown in FIG. 1. Entropy coding unit 56 may also entropy code motion vectors and other syntax elements for the current video frame being coded.

[0092] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain and generate reference blocks for prediction of other video blocks. As described above, motion compensation unit 44 may generate a motion-compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values ​​for use in motion estimation.

[0093] Adder 62 adds the reconstructed residual block to the motion compensated prediction block provided by motion compensation unit 44 to provide a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block for inter predicting another video block in a subsequent video frame.

[0094] 3 is a block diagram illustrating an exemplary video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process that is generally reverse to the encoding process described above for the video encoder 20 in connection with FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on a motion vector received from the entropy decoding unit 80, and the intra prediction unit 84 may generate prediction data based on an intra prediction mode indicator received from the entropy decoding unit 80.

[0095] In some examples, units of video decoder 30 may be tasked with performing implementations of the present disclosure. Also, in some examples, implementations of the present disclosure may be divided among one or more of the units of video decoder 30. For example, intra BC unit 85 may perform implementations of the present disclosure alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some examples, video decoder 30 may not include intra BC unit 85, and the functionality of intra BC unit 85 may be performed by other components of prediction processing unit 81, such as motion compensation unit 82.

[0096] Video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of video decoder 30. The video data stored in video data memory 79 may be obtained, for example, from storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include a coded picture buffer (CPB) that stores coded video data from the coded video bitstream. DPB 92 of video decoder 30 stores reference video data for use in decoding video data by video decoder 30 (e.g., in intra- or inter-predictive coding modes). Video data memory 79 and DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, video data memory 79 and DPB 92 are depicted in Figure 3 as two separate components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or by separate memory devices. In some examples, video data memory 79 may be on-chip with other components of video decoder 30 or off-chip with respect to those components.

[0097] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks and associated syntax elements of encoded video frames. Video decoder 30 may receive the syntax elements at the video frame level and / or the video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors, or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators and other syntax elements to prediction processing unit 81.

[0098] When a video frame is coded as an intra-predictive coded (I) frame, or coded for intra-coded predictive blocks within other types of frames, intra prediction unit 84 of prediction processing unit 81 may generate predictive data for video blocks of the current video frame based on the signaled intra-prediction mode and reference data from previously decoded blocks of the current frame.

[0099] When a video frame is coded as an inter-predictively coded (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 produces one or more prediction blocks for video blocks of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks may be produced from one reference frame in the reference frame lists. Video decoder 30 may construct reference frame lists List0 and List1 using a default construction technique based on the reference frames stored in DPB 92.

[0100] In some examples, when a video block is coded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 produces a predictive block for the current video block based on the block vectors and other syntax elements received from entropy decoding unit 80. The predictive block may be located within the same reconstructed region of the picture as the current video block as defined by video encoder 20.

[0101] Motion compensation unit 82 and / or intra BC unit 85 determine prediction information for video blocks of the current video frame by analyzing the motion vectors and other syntax elements, and then use the prediction information to yield a prediction block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra- or inter-prediction) used to code the video blocks of the video frame, the inter-prediction frame type (e.g., B or P), construction information for one or more of the reference frame lists for the frame, the motion vector for each inter-predictive coded video block of the frame, the inter-prediction state for each inter-predictive coded video block of the frame, and other information for decoding video blocks in the current video frame.

[0102] Similarly, intra BC unit 85 may use some of the received syntax elements, such as flags, to determine that the current video block was predicted using an intra BC mode, construction information for the video blocks of the frame that are located within the reconstructed region and that should be stored in DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction states for each intra BC predicted video block of the frame, and other information for decoding video blocks in the current video frame.

[0103] Motion compensation unit 82 may also perform interpolation using interpolation filters used by video encoder 20 during encoding of the video block to calculate interpolated values ​​for sub-integer pixels of the reference block. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 from the received syntax elements and may use those interpolation filters to result in a predictive block.

[0104] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients entropy decoded by entropy decoding unit 80 provided in the bitstream, using the same quantization parameters calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients to reconstruct the residual block in the pixel domain.

[0105] After motion compensation unit 82 or intra BC unit 85 generates a predictive block for the current video block based on the vectors and other syntax elements, adder 90 reconstructs a decoded video block for the current video block by adding a residual block from inverse transform processing unit 88 and the corresponding predictive block generated by motion compensation unit 82 and intra BC unit 85. To further process the decoded video block, an in-loop filter 91, such as a deblocking filter, an SAO filter, and / or an ALF, may be placed between adder 90 and DPB 92. The in-loop filter 91 may be applied on the reconstructed CU before being placed in the reference picture store. In some examples, the in-loop filter 91 may be omitted, and the decoded video block may be provided directly to DPB 92 by adder 90. The decoded video block in a given frame is then stored in DPB 92, which stores the reference frame used for motion compensation after the next video block. DPB 92, or a memory device separate from DPB 92, may also store the decoded video for later presentation on a display device, such as display device 34 of FIG.

[0106] In a typical video coding process, a video sequence typically includes an ordered set of frames or pictures. Each frame can include three sample arrays, denoted SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other cases, a frame may be monochromatic and therefore include only one two-dimensional array of luma samples.

[0107] Like HEVC, the AVS3 standard is built on a block-based hybrid video coding framework. The input video signal is processed block by block (called a coding unit (CU)). Unlike HEVC, which divides blocks solely based on a quadtree, AVS3 divides a single coding tree unit (CTU) into CUs based on a quadtree, binary tree, or extended quadtree to adapt to varying local characteristics. In addition, the concept of multiple partition unit types in HEVC has been eliminated; that is, AVS3 does not separate CUs, prediction units (PUs), and transform units (TUs). Instead, each CU is always used as the basic unit for both prediction and transformation without further partitioning. In the AVS3 tree partitioning structure, a CTU is first partitioned based on a quadtree structure. Then, each quadtree leaf node can be further partitioned based on binary tree and extended quadtree structures.

[0108] As shown in FIG. 4A, video encoder 20 (or more specifically, partition unit 45) generates a coded representation of a frame by first dividing the frame into a set of CTUs. A video frame may contain an integer number of CTUs, ordered consecutively from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by video encoder 20 in a sequence parameter set. All CTUs in a video sequence have the same size, which may be one of 128x128, 64x64, 32x32, or 16x16. However, it should be noted that the present application is not necessarily limited to a particular size. As shown in FIG. 4B, each CTU may comprise one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements used to code the samples in these coding tree blocks. The syntax elements describe the characteristics of different types of units of coding blocks of pixels and how a video sequence may be reconstructed at video decoder 30, including inter or intra prediction, intra prediction mode, motion vectors, and other parameters. For a monochrome picture or a picture with three distinct color planes, a CTU may comprise a single coding tree block and the syntax elements used to code the samples of this coding tree block. A coding tree block may be an NxN block of samples.

[0109] To achieve better performance, video encoder 20 can recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination thereof, on the coding tree blocks of a CTU to partition the CTU into smaller CUs. As illustrated in FIG. 4C, a 64×64 CTU 400 is first partitioned into four smaller CUs, each with a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each partitioned into four CUs with a block size of 16×16. Two 16×16 CUs, 430 and 440, are each further partitioned into four CUs with a block size of 8×8. FIG. 4D illustrates a quad tree data structure resulting from the partitioning process of CTU 400 illustrated in FIG. 4C, where each leaf node of the quad tree corresponds to one CU, each with a size ranging from 32×32 to 8×8. Similar to the CTU depicted in Figure 4B, each CU may comprise a CB of luma samples, two corresponding coded blocks of chroma samples for the same frame size, and the syntax elements used to code the samples of these coded blocks. In a monochrome picture or a picture with three distinct color planes, a CU may comprise a single coded block and the syntax structure used to code the samples of this coded block. Note that the quadtree partitioning depicted in Figures 4C and 4D is for illustrative purposes only; a CTU may be partitioned into multiple CUs to accommodate varying local characteristics based on quadtree / ternary / binary tree partitioning. In multiple types of tree structures, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU may be further partitioned by binary and ternary tree structures. As shown in Figure 4E, a coding block has a width W and a height H, and there are five possible partition types: 4-way partition, horizontally bi-way partition, vertically bi-way partition, horizontally tri-way partition, and vertically tri-way partition. In AVS3, there are five possible partition types: 4-partition, horizontal 2-partition, vertical 2-partition, horizontal extended quadtree partition (not shown in Figure 4E), and vertical extended quadtree partition (not shown in Figure 4E).

[0110] In some implementations, video encoder 20 may further divide a coding block of a CU into one or more M×N PBs. A PB is a rectangular (square or non-square) block of samples to which the same inter or intra prediction is applied. A PU of a CU may comprise a PB of luma samples and two corresponding PBs of chroma samples, along with syntax elements used to predict these PBs. In a monochromatic picture or a picture with three distinct color planes, a PU may comprise a single PB and syntax structures used to predict this PB. Video encoder 20 may generate predicted luma, Cb, and Cr blocks for the luma, Cb, and CrPBs of each PU of a CU.

[0111] Video encoder 20 may generate predictive blocks for a PU using intra prediction or inter prediction. If video encoder 20 uses intra prediction to generate predictive blocks for a PU, video encoder 20 may generate predictive blocks for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate predictive blocks for the PU, video encoder 20 may generate predictive blocks for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0112] After video encoder 20 generates predictive luma, Cb, and Cr blocks for one or more PUs of a CU, video encoder 20 may generate a luma residual block for the CU by subtracting the predictive luma block of the CU from its original luma-coded block, where each sample in the luma residual block of the CU indicates a difference between a luma sample in one of the predictive luma blocks of the CU and a corresponding sample in the original luma-coded block of the CU. Similarly, video encoder 20 may generate Cb residual blocks and Cr residual blocks for the CU, where each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predictive Cb blocks of the CU and a corresponding sample in the original Cb-coded block of the CU, and each sample in the Cr residual block of the CU indicates a difference between a Cr sample in one of the predictive Cr blocks of the CU and a corresponding sample in the original Cr-coded block of the CU.

[0113] Further, as shown in FIG. 4C , video encoder 20 may use quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of a CU may comprise a transform block of luma samples and two corresponding transform blocks of chroma samples, as well as syntax elements used to transform the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may comprise a single transform block and syntax structures used to transform the samples of the transform block.

[0114] Video encoder 20 may apply one or more transforms to a luma transform block of a TU to generate a luma coefficient block for the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 may apply one or more transforms to a Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to a Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0115] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to a process by which transform coefficients are quantized, possibly to reduce the amount of data used to represent the transform coefficients and provide further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 may output a bitstream including a sequence of bits and associated data that form a representation of a coded frame, which may be stored in storage device 32 or transmitted to destination device 14.

[0116] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct frames of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally reverse to the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on coefficient blocks associated with TUs of a current CU to reconstruct residual blocks associated with the TUs of the current CU. Video decoder 30 also reconstructs coded blocks of the current CU by adding samples of predictive blocks for PUs of the current CU to corresponding samples of transform blocks of TUs of the current CU. After reconstructing the coded blocks for each CU of a frame, video decoder 30 may reconstruct the frame.

[0117] As mentioned above, video coding achieves video compression using two main modes: intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). Note that IBC can be considered intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes to coding efficiency more than intra-frame prediction because it uses motion vectors to predict the current video block from a reference video block.

[0118] However, with the increasing advances in video data capture technology and the further refinement of video block sizes to preserve details in video data, the amount of data required to represent motion vectors for a current frame also increases substantially. One way to overcome this challenge is to take advantage of the fact that a group of neighboring CUs in both the spatial and temporal domains not only have similar video data for prediction purposes, but also have similar motion vectors between these neighboring CUs. Therefore, by exploring the spatial and temporal correlations of the current CU, also known as a "motion vector predictor (MVP)," it is possible to use the motion information of spatially neighboring CUs and / or temporally aligned CUs as an approximation of the motion information (e.g., motion vector) of the current CU.

[0119] Instead of encoding into the video bitstream the actual motion vector of the current CU determined by motion estimation unit 42 as described above in connection with Figure 2, the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to obtain a Motion Vector Difference (MVD) for the current CU. By doing so, the motion vector determined by motion estimation unit 42 for each CU of a frame does not need to be encoded into the video bitstream, and the amount of data used to represent motion information in the video bitstream may be significantly reduced.

[0120] Similar to the process of choosing a predictive block in a reference frame during inter-frame prediction of a code block, a set of rules needs to be adopted by both video encoder 20 and video decoder 30 to construct a motion vector candidate list (also known as a "merge list") for the current CU using those potential motion vector candidates associated with spatially neighboring CUs and / or temporally aligned CUs of the current CU, and then select one member from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, the motion vector candidate list itself does not need to be transmitted from video encoder 20 to video decoder 30; the index of the selected motion vector predictor in the motion vector candidate list is sufficient for video encoder 20 and video decoder 30 to encode and decode the current CU using the same motion vector predictor in the motion vector candidate list.

[0121] Generally, the basic intra prediction schemes applied in VVC remain largely the same as those in HEVC, except that some prediction tools have been further extended, added, and / or improved, such as enhanced intra prediction with wide-angle intra mode, multiple baseline (MRL) intra prediction, composite position-dependent intra prediction (PDPC), intra subdivision (ISP) prediction, component-to-component linear model (CCLM) prediction, and matrix weighted intra prediction (MIP).

[0122] Like HEVC, VVC uses a set of reference samples adjacent to the current CU (i.e., above or to the left of the current CU) to predict samples in the current CU. However, to capture the finer edge orientations present in raw video (especially in high-resolution, e.g., 4K, video content), the number of angular intra modes is expanded from 33 in HEVC to 93 in VVC. Figure 4F is a block diagram illustrating the intra modes defined in VVC. As shown in Figure 4F, among the 93 angular intra modes, modes 2 through 66 are traditional angular intra modes, while modes -1 through -14 and modes 67 through 80 are wide-angle intra modes. In addition to the angular intra modes, HEVC's planar mode (mode 0 in Figure 1) and direct current (DC) mode (mode 1 in Figure 1) are also applied in VVC.

[0123] As shown in FIG. 4E , because a quadtree / binary / ternary tree partitioning structure is applied in VVC, in addition to square video blocks, rectangular video blocks also exist for intra prediction in VVC. Because the width and height of a given video block are not equal, various sets of angular intra modes may be selected from 93 angular intra modes for different block shapes. More specifically, for both square and rectangular video blocks, in addition to planar and DC modes, 65 angular intra modes out of the 93 angular intra modes are also supported for each block shape. When a rectangular block of video blocks meets certain conditions, video decoder 30 may adaptively determine the index of the wide-angle intra mode of the video block according to the index of the conventional angular intra mode received from video encoder 20 using the mapping relationship shown in Table 1 below. That is, for non-square blocks, wide-angle intra modes are signaled by video encoder 20 using the index of the conventional angular intra modes, which are then analyzed and mapped to the index of the wide-angle intra modes by video decoder 30, thus ensuring that the total number (i.e., 67) of intra modes (i.e., planar mode, DC mode, and 65 angular intra modes out of 93 angular intra modes) remains unchanged, and the intra-prediction mode coding method remains unchanged. As a result, good signaling efficiency of intra-prediction modes is achieved while providing a consistent design across different block sizes.

[0124] Table 1-0 shows the mapping relationship between the indices of conventional angle intra modes and wide angle intra modes for intra prediction of different block shapes in VCC, where W represents the width of the video block and H represents the height of the video block. [Table 1]

[0125] Similar to intra prediction in HEVC, all intra modes in VVC (i.e., planar, DC, and angular intra modes) utilize a set of reference samples above and to the left of the current video block for intra prediction. However, unlike HEVC, in which only the nearest row / column of reference samples (i.e., the zeroth line 201 in FIG. 4G ) is used, VVC introduces MRL intra prediction, in which, in addition to the nearest row / column of reference samples, two additional rows / columns of reference samples (i.e., the first line 203 and the third line 205 in FIG. 4G ) may also be used for intra prediction. The index of the selected row / column of reference samples is signaled and sent from video encoder 20 to video decoder 30. When a non-nearest row / column of reference samples (i.e., the first line 203 or the third line 205 in FIG. 4G ) is selected, planar modes are excluded from the set of intra modes that may be used to predict the current video block. To prevent the use of extended reference samples outside the current CTU, MRL intra prediction is disabled for the first row / column of video blocks within the current CTU.

[0126] Sample Adaptive Offset (SAO) is a process that modifies decoded samples after the application of a deblocking filter by conditionally adding an offset value to each sample based on values ​​in a lookup table transmitted by the encoder. SAO filtering is performed region-by-region based on the filtering type selected for each CTB by the syntax element sao-type-idx. A value of 0 for sao-type-idx indicates that no SAO filter is applied to the CTB, while values ​​1 and 2 signal the use of the band-offset and edge-offset filtering types, respectively. In band-offset mode, specified by sao-type-idx equal to 1, the selected offset value depends directly on the sample amplitude. In this mode, the complete sample amplitude range is uniformly divided into 32 segments called bands, and sample values ​​belonging to four of these bands (consecutive within the 32 bands) are modified by adding a transmitted value, denoted as band offset, which can be positive or negative. The main reason for using four consecutive bands is that in smooth areas where banding artifacts may appear, the sample amplitudes of the CTB tend to be concentrated in only a few of these bands. Additionally, the design choice of using four offsets unifies with the edge offset operation mode, which also uses four offset values. In the edge offset mode specified by sao-type-idx equal to 2, the syntax element sao-eo-class, with a value between 0 and 3, signals whether horizontal, vertical, or one of the two diagonal gradient directions is used for edge offset classification in the CTB.

[0127] FIG. 5A is a block diagram illustrating four gradient patterns used in SAO according to some implementations of the present disclosure. Four gradient patterns 502, 504, 506, and 508 are for each sao-eo-class in edge offset mode. The sample labeled "p" indicates the center sample to be considered. Two samples labeled "n0" and "n1" designate two adjacent samples along the following gradient patterns: (a) horizontal (sao-eo-class=0), (b) vertical (sao-eo-class=1), (c) 135° diagonal (sao-eo-class=2), and (d) 45° diagonal (sao-eo-class=3). Each sample in the CTB is classified into one of five EdgeIdx categories by comparing the sample value p located at a certain position with the values ​​n0 and n1 of the two adjacent samples located at the same position, as shown in FIG. 5A. This classification is done for each sample based on the decoded sample value, so no additional signaling for EdgeIdx classification is required. For EdgeIdx categories 1-4, an offset value from the transmitted lookup table is added to the sample value depending on the EdgeIdx category at the sample position. The offset value is always positive for categories 1 and 2 and negative for categories 3 and 4. Therefore, the filter generally has a smoothing effect in edge offset mode. Table 1-1 below shows sample EdgeIdx categories for SAO edge classes. [Table 2]

[0128] For SAO types 1 and 2, a total of four amplitude offset values ​​are transmitted to the decoder for each CTB. For type 1, a sign is also encoded. The offset values ​​and associated syntax elements, such as sao-type-idx and sao-eo-class, are determined by the encoder, typically using criteria that optimize rate-distortion performance. For efficient signaling, SAO parameters can be indicated as inherited from the left or top CTB using a merge flag. In summary, SAO is a nonlinear filtering operation that allows additional refinement of the reconstructed signal, enhancing the signal representation both in smooth areas and around edges.

[0129] In some embodiments, a sample adaptive offset pre-SAO (pre-SAO) is implemented. The low-complexity pre-SAO coding performance is promising for future video coding standards development. In some examples, the pre-SAO is applied only to luma component samples, using luma samples for classification. The pre-SAO operates by applying two SAO-like filtering operations, called SAOV and SAOH, which are applied jointly with a deblocking filter (DBF) before applying the existing (legacy) SAO. The first SAO-like filter, SAOV, operates to apply SAO to input picture Y2 after the deblocking filter for vertical edges (DBFV) has been applied. Y3(i)=Clip1(Y2(i)+d1·(f(i)>T?1:0)-d2·(f(i)<-T?1:0))

[0130] where T is a predetermined positive constant, and d1 and d2 are f(i)=Y1(i)-Y2(i) is the offset coefficient associated with the two classes based on the sample-by-sample difference between Y1(i) and Y2(i) given by

[0131] The first class for d1 is given by f(i) > T for all sample locations i, and the second class for d2 is given by f(i) < -T. Similar to the existing SAO process, offset coefficients d1 and d2 are calculated in the encoder to minimize the mean squared error between the SAOV output picture Y3 and the original picture X. After SAOV is applied, a second SAO-like filter, SAOH, operates to apply SAO to Y4 after SAOV is applied by classifying it based on the sample-by-sample difference between Y3(i) and Y4(i). This is the output picture of the deblocking filter for horizontal edges (DBFH). The same procedure as for SAOV is applied to SAOH, except that Y3(i)-Y4(i) are used for classification instead of Y1(i)-Y2(i). Two offset coefficients, a predetermined threshold T, and an enable flag for each of SAOH and SAOV are signaled at the slice level. SAOH and SAOV are applied independently to the luma and two chroma components.

[0132] In some cases, both SAOV and SAOH operate only on picture samples affected by their respective deblocking (DBFV or DBFH) filters. Thus, unlike existing SAO processes, only a fraction of all samples in a given spatial region (picture, or CTU in the case of legacy SAO) are processed by pre-SAO, keeping the resulting average decoder-side work per picture sample low (two or three comparisons and two additions per sample in the worst-case scenario, based on preliminary estimates). Pre-SAO only requires samples used by the deblocking filter and does not store any additional samples at the decoder.

[0133] In some embodiments, to explore compression efficiency beyond VVC, a Bi-Directional Filter (BIF) is implemented. The BIF is implemented with a Sample Adaptive Offset (SAO) loop filter stage. Both the Bi-Directional Filter (BIF) and the SAO use samples from the deblocking filter as input. Each filter creates a per-sample offset, which is added to the input sample and then clipped before proceeding to the ALF.

[0134] For details, see the output sample I OUT teeth, I OUT =clip3(I C +ΔI BIF +ΔI SAO ) where I C is the input sample from deblocking, and ΔI BIF is the offset from the two-way filter, and ΔI SAO is the offset from SAO.

[0135] In some embodiments, the implementation provides the possibility for the encoder to enable or disable filtering at the CTU and slice level. The encoder makes the decision by evaluating the rate-distortion optimization (RDO) cost.

[0136] The following syntax elements are introduced in PPS: [Table 3]

[0137] pps_bilateral_filter_enabled_flag equal to 0 specifies that the bilateral loop filter is disabled for slices that reference the PPS. pps_bilateral_filter_flag equal to 1 specifies that the bilateral loop filter is enabled for slices that reference the PPS.

[0138] bilateral_filter_strength specifies the bilateral loop filter strength value used in the bilateral transform block filter process. The value of bilateral_filter_strength shall be in the range of 0 to 2, inclusive.

[0139] bilateral_filter_qp_offset specifies the offset used in deriving the bilateral filter lookup table LUT(x) for the slice referencing the PPS. bilateral_filter_qp_offset shall be in the range of -12 to +12, inclusive.

[0140] The following syntax elements are introduced: [Table 4] [Table 5]

[0141] The meaning is as follows: slice_bilateral_filter_all_ctb_enabled_flag equal to 1 specifies that the bilateral filter is enabled and applies to all CTBs in the current slice. When slice_bilateral_filter_all_ctb_enabled_flag is not present, it is inferred to be equal to 0.

[0142] slice_bilateral_filter_enabled_flag equal to 1 specifies that the bilateral filter is enabled and may be applied to the CTB of the current slice. When slice_bilateral_filter_enabled_flag is not present, it is inferred to be equal to slice_bilateral_filter_all_ctb_enabled_flag.

[0143] bilateral_filter_ctb_flag[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] equal to 1 specifies that a bilateral filter is applied to the luma coding tree block of the coding tree unit at luma location (xCtb,yCtb). bilateral_filter_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] equal to 0 specifies that a bilateral filter is not applied to the luma coding tree block of the coding tree unit at luma location (xCtb,yCtb). When bilateral_filter_ctb_flag is not present, it is inferred to be equal to (slice_bilateral_filter_all_ctb_enabled_flag&slice_bilateral_filter_enabled_flag).

[0144] In some examples, for a filtered CTU, the filtering process proceeds as follows: At picture boundaries where samples are not available, the two-way filter uses dilation (sample repetition) to fill in the unavailable samples. For virtual boundaries, the behavior is the same as for SAO, i.e., no filtering occurs. When crossing a horizontal CTU boundary, the two-way filter can access the same samples that the SAO has access to. Figure 5B is a block diagram illustrating the naming convention for samples surrounding a central sample according to some implementation examples of the present disclosure. As an example, the central sample I C If is located at the top line of the CTU, I NW , I A , and I NE is read from the CTU above, just like SAO, but I AA is padded, so no extra line buffer is needed. CThe sample surrounding is shown in Figure 5B, where A, B, L, and R represent up, down, left, and right, and NW, NE, SW, SE represent northwest, etc. Similarly, AA represents up-up, BB represents down-down, etc. This diamond shape is I AA , I BB , I LL , or I RR Unlike the alternative method which uses a square filter, which does not use a square filter.

[0145] Each peripheral sample I A , I R etc. are the corresponding modifier values

number

number

[0146] Then the modifier value is

number

[0147] These values ​​can be stored using 6 bits per entry, resulting in 26*16*6 / 8=312 bytes or 300 bytes if you exclude the first row which is all 0's.

number

number

number

number

[0148] The modifier values ​​are summed together.

number

[0149] In some instances,

number

number

number

number

number

number

number

number

number

number

[0150] Finally, the two-way filter offset ΔI BIF is calculated. For full strength filtering, the following is used: ΔI BIF =(c v +16)>>5 For half strength filtering, the following is used: ΔI BIF =(c v +32)>>6

[0151] The general formula for n bits of data is r add =2 14-n-bilateral_filter_strength rshift =15-n-bilateal_filter_strength ΔI BIF =(c v +r add )>>r shift where bilateral_filter_strength can be 0 or 1 and is signaled in pps.

[0152] In some embodiments, methods and systems are disclosed herein for improving coding efficiency or reducing the complexity of sample adaptive offset (SAO) by introducing inter-component information. SAO is used in the HEVC, VVC, AVS2, and AVS3 standards. In the following description, the existing SAO design in the HEVC, VVC, AVS2, and AVS3 standards is used as the basic SAO method for those skilled in the art of video coding. However, the inter-component method described in this disclosure may also be applied to other loop filter designs or other coding tools with similar design spirit. For example, in the AVS3 standard, SAO is replaced by a coding tool called enhanced sample adaptive offset (ESAO). However, the CCSAO disclosed herein may be applied in parallel to the ESAO. In another example, the CCSAO may be applied in parallel to the Constrained Directional Enhancement Filter (CDEF) in the AV1 standard.

[0153] In the existing SAO designs in the HEVC, VVC, AVS2, and AVS3 standards, the luma Y, chroma Cb, and chroma Cr sample offset values ​​are determined independently. That is, for example, the current chroma sample offset is determined only by the current chroma sample value and adjacent chroma sample values, without considering the alignment or adjacent luma samples. However, luma samples retain more original picture detail than chroma samples and can benefit from the current chroma sample offset determination. Furthermore, because chroma samples typically lose high-frequency detail after RGB-to-YCbCr color conversion or after quantization and deblocking filters, introducing luma samples with preserved high-frequency detail for chroma offset determination can benefit from chroma sample reconstruction. Therefore, further gains can be expected by exploring inter-component correlations, for example, by using Cross-Component Sample Adaptive Offset (CCSAO) methods and systems. In some embodiments, this correlation not only includes inter-component sample values, but also picture / coding information such as prediction / residual coding mode, transform type, and quantization / deblocking / SAO / ALF parameters from the inter-component sample values.

[0154] Another example is that in the case of SAO, the luma sample offset is determined by the luma sample only. However, for example, luma samples with the same band offset (BO) classification may be further classified by their alignment and neighboring chroma samples, thereby obtaining a more effective classification. SAO classification may be obtained as a shortcut to compensate for sample differences between the original picture and the reconstructed picture. Therefore, an effective classification is desirable.

[0155] FIG. 6A is a block diagram illustrating a CCSAO system and process applied to chroma samples by some implementations of the present disclosure and using DBF Y as an input. The luma sample after the luma deblocking filter (DBF Y) is used to determine additional offsets for chroma Cb and Cr after SAO Cb and SAO Cr. For example, a current chroma sample 602 is first classified using an array luma sample 604 and an adjacent (white) luma sample 606, and the corresponding CCSAO offset value of the corresponding class is added to the current chroma sample value. FIG. 6B is a block diagram illustrating a CCSAO system and process applied to luma samples and chroma samples by some implementations of the present disclosure and using DBF Y / Cb / Cr as an input. FIG. 6C is a block diagram illustrating a CCSAO system and process that can function independently by some implementations of the present disclosure. FIG. 6D is a block diagram illustrating a CCSAO system and process, which may be applied recursively (2 or N times) with the same or different offsets at the same codec stage or repeated at different stages according to some implementations of the present disclosure. In summary, in some embodiments, information about the current luma sample and neighboring luma samples, and information about the alignment and neighboring chroma samples (Cb and Cr) may be used to classify the current luma sample. In some embodiments, alignment and neighboring luma samples, alignment and neighboring cross-chroma samples, and the current chroma sample and neighboring chroma samples may be used to classify the current chroma sample (Cb or Cr). In some embodiments, CCSAO may be cascaded (1) after DBF Y / Cb / Cr, (2) after pre-DBF reconstructed image Y / Cb / Cr, (3) after SAO Y / Cb / Cr, or (4) after ALF Y / Cb / Cr.

[0156] In some embodiments, the CCSAO may also be applied in parallel to other coding tools, such as the ESAO in the AVS standard, or the CDEF in the AV1 standard, or the Neural Network Loop Filter (NNLF). Figure 6E is a block diagram illustrating a CCSAO system and process applied in parallel to the ESAO in the AVS standard by some implementations of the present disclosure.

[0157] FIG. 6F is a block diagram illustrating a CCSAO system and process applied after the SAO according to some implementations of the present disclosure. In some embodiments, FIG. 6F illustrates that the location of the CCSAO can be after the SAO, i.e., in the place of the Inter-Component Adaptive Loop Filter (CCALF) in the VVC standard. FIG. 6G is a block diagram illustrating that the CCSAO system and process can function independently without the CCALF according to some implementations of the present disclosure. In some embodiments, the SAO Y / Cb / Cr can be replaced with ESAO, for example, in the AVS3 standard.

[0158] FIG. 6H is a block diagram illustrating a CCSAO system and process applied in parallel to CCALF according to some implementations of the present disclosure. In some embodiments, CCSAO may be applied in parallel to CCALF. In some embodiments, the locations of CCALF and CCSAO may be switched, as shown in FIG. 6H. In some embodiments, the SAO Y / Cb / Cr blocks may be swapped for ESAO Y / Cb / Cr (AVS3) or CDEF (AV1), as shown in FIGS. 6A-6H or throughout this disclosure. Note that Y / Cb / Cr may also be denoted as Y / U / V in the video coding domain. In some embodiments, if the video is in RGB format, CCSAO can also be applied by simply mapping the respective YUV representations to GBR in this disclosure.

[0159] FIG. 6I is a block diagram illustrating a system and process of CCSAO applied in parallel to an SAO and a BIF according to some implementations of the present disclosure. FIG. 6J is a block diagram illustrating a system and process of CCSAO applied in parallel to a BIF by replacing an SAO according to some implementations of the present disclosure. In some embodiments, the current chroma sample classification reuses the SAO type (edge ​​offset (EO) or BO), class, and category of the ordered luma samples. The corresponding CCSAO offsets may be signaled or derived from the decoder itself. For example, let h_Y be the ordered luma SAO offset, and let h_Cb and h_Cr be the CCSAO Cb and Cr offsets, respectively. h_Cb (or h_Cr) = w*h_Y, where w may be selected from a limited table, such as ±1 / 4, ±1 / 2, 0, ±1, ±2, ±4, etc., where |w| includes only values ​​that are powers of two.

[0160] In some embodiments, the comparison score of the sequence luma sample (Y0) and the eight adjacent luma samples [-8, 8] is used, resulting in a total of 17 classes. Initial Class=0 Loop through 8 adjacent luma samples (Yi, i=1~8) if Y0>Yi Class+=1 else if Y0 <Yi Class-=1

[0161] In some embodiments, the classification methods described above may be combined. For example, to increase diversity, the comparison scores combined with SAO BO (32-band classification) are used, resulting in a total of 17*32 classes. In some embodiments, Cb and Cr may use the same class to reduce complexity or save bits.

[0162] FIG. 7 is a block diagram illustrating a sample process using the CCSAO according to some implementations of the present disclosure. Specifically, FIG. 7 illustrates that the input of the CCSAO can incorporate vertical and horizontal DBF inputs to simplify class determination or increase flexibility. For example, let Y0_DBF_V, Y0_DBF_H, and Y0 be the array luma samples at the inputs of DBF_V, DBF_H, and SAO, respectively. Yi_DBF_V, Yi_DBF_H, and Yi are the eight adjacent luma samples at the inputs of DBF_V, DBF_H, and SAO, respectively, where i = 1 to 8. Max Y0=max(Y0_DBF_V,Y0_DBF_H,Y0_DBF) Max Yi=max(Yi_DBF_V,Yi_DBF_H,Yi_DBF) max Y0 and max Yi are given as CCSAO classifications.

[0163] Figure 8 is a block diagram illustrating how the CCSAO process is interleaved across vertical and horizontal DBFs according to some example implementations of the present disclosure. In some embodiments, the CCSAO blocks of Figures 6, 7, and 8 may be optional. For example, the DBF_V luma sample input is used as the CCSAO input, while the Y0_DBF_V and Yi_DBF_V are used for the first CCSAO_V, which applies the same sample processing as in Figure 6.

[0164] In some embodiments, the CCSAO syntax implemented is shown in Table 2 below. [Table 7]

[0165] In some embodiments, when one additional chroma offset is signaled to signal the CCSAO Cb and Cr offset values, the other chroma component offset can be derived by a positive or negative sign or by weighting to save bit overhead. For example, let h_Cb and h_Cr be the offsets of the CCSAO Cb and Cr, respectively. By explicitly signaling w, when w=±|w| and the |w| candidates are limited, h_Cr can be derived from h_Cb without explicitly signaling h_Cr itself. h_Cr=w*h_Cb

[0166] Figure 7 is a block diagram illustrating an example process using CCSAO according to some implementations of the present disclosure. Figure 8 is a block diagram illustrating how the CCSAO process is interleaved with vertical and horizontal deblocking filters (DBFs) according to some implementations of the present disclosure.

[0167] FIG. 9 is a flow diagram illustrating an example process 900 for decoding a video signal using inter-component correlation according to some implementations of this disclosure.

[0168] Video decoder 30 receives a video signal including a first component and a second component (910). In some embodiments, the first component is a luma component and the second component is a chroma component of the video signal.

[0169] Video decoder 30 also receives a plurality of offsets associated with the second component (920).

[0170] Video decoder 30 then utilizes the characteristic measure of the first component to obtain a classification category associated with the second component (930). For example, in FIG. 6, current chroma sample 602 is first classified using aligned luma sample 604 and adjacent (white) luma sample 606, and the corresponding CCSAO offset value is added to the current chroma sample.

[0171] Video decoder 30 further selects a first offset from a plurality of offsets for the second component according to the classification category (940).

[0172] Video decoder 30 additionally modifies the second component based on the selected first offset (950).

[0173] In some embodiments, obtaining a classification category associated with a second component using characteristic measurements of the first component (930) includes obtaining a classification category for each sample of the second component using each sample of the first component, where each sample of the first component is a corresponding sample of the first component for each sample of the second component. For example, the current chroma sample classification reuses the SAO type (EO or BO), class, and category of the corresponding luma sample.

[0174] In some embodiments, obtaining a classification category associated with the second component using characteristic measurements of the first component (930) includes obtaining a classification category for each sample of the second component using each sample of the first component, where each sample of the first component is reconstructed before being deblocked or reconstructed after being deblocked. In some embodiments, the first component is deblocked with a deblocking filter (DBF). In some embodiments, the first component is deblocked with a luma deblocking filter (DBF Y). For example, as an alternative to FIG. 6 or FIG. 7, the CCSAO input may be before DBF Y.

[0175] In some embodiments, the characteristic measure is derived by dividing the range of sample values ​​of the first component into bands and selecting bands based on the intensity values ​​of the samples in the first component. In some embodiments, the characteristic measure is derived from a band offset (BO).

[0176] In some embodiments, the characteristic measure is derived based on the direction and intensity of edge information of the samples in the first component. In some embodiments, the characteristic measure is derived from edge offset (EO).

[0177] In some embodiments, modifying (950) the second component includes adding the selected first offset directly to the second component, e.g., adding a corresponding CCSAO offset value to the current chroma component sample.

[0178] In some embodiments, modifying (950) the second component includes mapping the selected first offset to a second offset and adding the mapped second offset to the second component. For example, if one additional chroma offset is signaled to signal the CCSAO Cb and Cr offset values, the other chroma component offset can be derived by using a positive or negative sign or by weighting to save bit overhead.

[0179] In some embodiments, receiving 910 the video signal includes receiving a syntax element indicating whether decoding the video signal using CCSAO is enabled for the video signal in a sequence parameter set (SPS). In some embodiments, cc_sao_enabled_flag indicates whether CCSAO is enabled at the sequence level.

[0180] In some embodiments, receiving (910) the video signal includes receiving a syntax element indicating whether decoding the video signal using CCSAO is enabled for the second component at the slice level. In some embodiments, slice_cc_sao_cb_flag or slice_cc_sao_cr_flag indicates whether CCSAO is enabled within the slice for Cb or Cr, respectively.

[0181] In some embodiments, receiving 920 multiple offsets associated with the second component includes receiving different offsets for different coding tree units (CTUs), where for a CTU, cc_sao_offset_sign_flag indicates a sign for the offset and cc_sao_offset_abs indicates CCSAO Cb and Cr offset values ​​for the current CTU.

[0182] In some embodiments, receiving 920 multiple offsets associated with the second component includes receiving a syntax element indicating whether the received offset of the CTU is the same as one of the CTU's neighboring CTUs, where the neighboring CTU is a left or upper neighboring CTU. For example, cc_sao_merge_up_flag indicates whether the CCSAO offset is merged from a left or upper CTU.

[0183] In some embodiments, the video signal further includes a third component, and the method of decoding the video signal using the CCSAO further includes receiving a second plurality of offsets associated with the third component, utilizing characteristic measurements of the first component to obtain a second classification category associated with the third component, selecting a third offset from the second plurality of offsets for the third component according to the second classification category, and modifying the third component based on the selected third offset.

[0184] Figure 11 is a block diagram of a sample process that illustrates how all alignment and adjacent (white) luma / chroma samples can be fed to the CCSAO classification in some implementations of the present disclosure. Figures 6A, 6B, and 11 show the inputs of the CCSAO classification. In Figure 11, the current chroma sample is 1104, the inter-component alignment chroma sample is 1102, and the alignment luma sample is 1106.

[0185] In some embodiments, an exemplary classifier (C0) uses the following sequence of luma or chroma sample values ​​(Y0) (Y4 / U4 / V4 in FIGS. 6B and 6C) in FIG. 12 for classification: If band_num is the number of equally divided bands of the luma or chroma dynamic range and bit_depth is the sequence bit depth, then an example class index for the current chroma sample is: class(C0)=(Y0*band_num)>>bit_depth

[0186] In some embodiments, this classification takes into account rounding, for example: class(C0)=((Y0*band_num)+(1<<bit_depth))> >bit_depth

[0187] Some band_num and bit_depth examples are listed below in Table 3. Table 3 shows three classification examples when the number of bands for each classification example is different. [Table 8]

[0188] In some embodiments, the classifier uses a different luma sample location for the C0 classification. Figure 10A is a block diagram illustrating a classifier that uses a different luma (or chroma) sample location for the C0 classification, for example, using neighboring Y7 instead of Y0 for the C0 classification, according to some implementations of this disclosure.

[0189] In some embodiments, different classifiers can be switched at the Sequence Parameter Set (SPS) / Adaptation Parameter Set (APS) / Picture Parameter Set (PPS) / Picture Header (PH) / Slice Header (SH) / Region / Coding Tree Unit (CTU) / Coding Unit (CU) / Sub-block / Sample level. For example, in Figure 10, we use Y0 for POC0 but Y7 for POC1, as shown in Table 4 below. [Table 9]

[0190] In some embodiments, FIG. 10B shows some examples of different shapes for luma candidates according to some implementations of the present disclosure. For example, constraints may be applied to the shapes. In some cases, the total number of luma candidates must be a power of two, as shown in FIG. 10B(b), (c), and (d). In some cases, the number of luma candidates must be horizontally and vertically symmetric about the chroma sample (center), as shown in FIG. 10B(a), (c), (d), and (e). In some embodiments, the power of two constraint and the symmetry constraint may also be applied to the chroma candidates. The U / V portions of FIG. 6B and FIG. 6C show an example for the symmetry constraint. In some embodiments, different color formats may have different classifier "constraints." For example, the 420 color format uses the luma / chroma candidate selection shown in Figures 6B and 6C (one candidate is selected from a 3x3 shape), while the 444 color format uses Figure 10B(f) for luma and chroma candidate selection, and the 422 color format uses Figure 10B(g) for luma (two chroma samples share four luma candidates) and Figure 10B(f) for chroma candidates.

[0191] In some embodiments, C0 position and C0 band_num may be combined and switched at SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. Different combinations may result in different classifiers, as shown in Table 5 below. [Table 10]

[0192] In some embodiments, the array luma sample value (Y0) is replaced with a value (Yp) obtained by weighting the array and adjacent luma samples. FIG. 12A shows an example classifier by replacing the array luma sample value with a value obtained by weighting the array and adjacent luma samples, according to some implementations of the present disclosure. The array luma sample value (Y0) may be replaced with a phase correction value (Yp) obtained by weighting adjacent luma samples. Different Yp values ​​may result in different classifiers.

[0193] In some embodiments, different Yp are applied to different chroma formats. For example, as shown in Figure 12A, Yp in Figure 12A(a) is used for 420 chroma format, Yp in Figure 12A(b) is used for 422 chroma format, and Y0 is used for 444 chroma format.

[0194] In some embodiments, another classifier (C1) compares the scores of the sequence luma sample (Y0) and the eight adjacent luma samples [-8, 8], resulting in a total of 17 classes, as shown below: Initial class (C1) = 0, loop through 8 adjacent luma samples (Yi,i = 1 to 8) if Y0>Yi Class+=1 else if Y0 <Yi Class-=1

[0195] In some embodiments, an example of C1 is equal to the following function, where the threshold th is 0: ClassIdx=Index2ClassTable(f(C,P1)+f(C,P2)+...+f(C,P8)) if xy>th,f(x,y)=1;if xy=th,f(x,y)=0;if xy <th,f(x,y)=-1 In the above formula, Index2ClassTable is a look-up table (LUT), C is the current sample or array sample, and P1~P8 are neighboring samples.

[0196] In some embodiments, similar to the C4 classifier, one or more thresholds may be predefined (e.g., maintained in a LUT) or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level to aid in classification (quantization) of the differences.

[0197] In some embodiments, the variation (C1') counts only the comparison score [0, 8], which results in 8 classes. (C1, C1') is the classifier group, and the PH / SH level flag can be signaled to switch between C1 and C1'.

[0198] Initial class (C1') = 0, loop through 8 adjacent luma samples (Yi,i = 1 to 8) if Y0>Yi Class+=1

[0199] In some embodiments, the variation (C1) selectively uses N neighbors out of M neighboring samples to calculate the comparison score. An M-bit bitmask may be signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level to indicate which neighboring samples were selected to calculate the comparison score. Using FIG. 6B as an example for the luma classifier, eight neighboring luma samples are candidates, and an 8-bit bitmask (01111110) is signaled at PH, indicating that six samples Y1 through Y6 are selected. Therefore, the comparison score is [-6, 6], resulting in an offset of 13. The selective classifier C1 gives the encoder more options for trading off offset signaling overhead and classification granularity.

[0200] Like C1, variation (C1') counts only the comparison score [0,+N], so the previous example bitmask 01111110 gives a comparison score of [0,6], resulting in an offset of 7.

[0201] In some embodiments, different classifiers are combined to result in a generic classifier, for example, different classifiers are applied to different pictures (different POC values), as shown in Table 6-1 below. [Table 11]

[0202] In some embodiments, another exemplary classifier (C3) uses a bit mask for classification, as shown in Table 6-2. To indicate this classifier, a 10-bit bit mask is signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, a bit mask of 11 1100 0000 means that for a given 10-bit luma sample value, only the most significant bits (MSBs): 4 bits are used for classification, resulting in a total of 16 classes. Another exemplary bit mask of 10 0100 0001 means that only 3 bits are used for classification, resulting in a total of 8 classes.

[0203] In some embodiments, the bit mask length (N) may be fixed or may be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, for a 10-bit sequence, a 4-bit bit mask 1110 is signaled in the PH in the picture, with the three most significant bits b9, b8, and b7 used for classification. Another example is a 4-bit bit mask 0011 on the least significant bits, with b0 and b1 used for classification. The bit mask classifier may correspond to luma or chroma classification. Whether the most significant bit or least significant bit is used for the bit mask N may be fixed or may be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.

[0204] In some embodiments, the luma position and C3 bitmask may be combined and switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different combinations may result in different classifiers.

[0205] In some embodiments, a "max number of ones" bit mask restriction may be applied to limit the number of corresponding offsets. For example, limiting the "max number of ones" bit mask to 4 in an SPS results in a maximum offset of 16 in the sequence. The bit masks for different POCs may be different, but the "max number of ones" shall not exceed 4 (and shall not exceed 16 for all classes). The "max number of ones" value may be signaled and can be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. [Table 12]

[0206] In some embodiments, as shown in FIG. 11 , other inter-component chroma samples, such as chroma sample 1102 and its neighboring samples, may also be provided to the CCSAO classification, e.g., for the current chroma sample 1104. For example, the Cr chroma sample may be provided to the CCSAO Cb classification. The Cb chroma sample may be provided to the CCSAO Cr classification. The classifier for the inter-component chroma sample may be the same as the luma inter-component classifier or may have its own classifier, as described in this disclosure. The two classifiers may be combined to form a joint classifier for classifying the current chroma sample. For example, a joint classifier combining the inter-component luma sample and the chroma sample results in a total of 16 classes, as shown in Table 6-3 below. [Table 13]

[0207] All the above mentioned classifiers (C0, C1, C1', C2, C3) can be combined. See for example Table 6-4 below. [Table 14]

[0208] In some embodiments, an exemplary classifier (C2) uses the sequence and the difference of adjacent luma samples (Yn). Figure 12A(c) shows an example of Yn, which has a dynamic range of [-1024, 1023] when the bit depth is 10. Let band_num in C2 be the number of equally divided bands of the Yn dynamic range. Class(C2)=(Yn+(1<<bit_depth)*band_num)> >(bit_depth+1).

[0209] In some embodiments, C0 and C2 are combined to result in a general classifier, with different classifiers applied to different pictures (different POCs), for example, as shown in Table 7 below. [Table 15]

[0210] In some embodiments, all the above-mentioned classifiers (C0, C1, C1', C2) are combined, for example, different classifiers are applied to different pictures (different POCs), as shown in Table 8-1 below. [Table 16]

[0211] In some embodiments, an exemplary classifier (C4) uses the difference between the CCSAO input value and the sample value to be compensated for classification, as shown in Table 8-2 below. For example, if CCSAO is applied in the ALF stage, the difference between the pre-ALF and post-ALF sample values ​​of the current component is used for classification. To aid in the classification (quantization) of the difference, one or more thresholds may be predefined (e.g., maintained in a look-up table (LUT)) or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. The C4 classifier can be combined with the Y / U / VbandNum of C0 to form a joint classifier (e.g., exemplary POC1, as shown in Table 8-2). [Table 17]

[0212] In some embodiments, because different coding modes may introduce different distortion statistics into the reconstructed image, an exemplary classifier (C5) uses "coding information" to aid in sub-block classification. For example, as shown in Table 8-3 below, a CCSAO sample is classified by its previous coding information, and the combination of coding information can form a classifier. Figure 30 below shows another example of different stages of coding information for C5. [Table 18]

[0213] In some embodiments, an exemplary classifier (C6) uses YUV color transform values ​​for classification. For example, to classify the current Y component, a 1 / 1 / 1 array or adjacent Y / U / V samples is selected to be color transformed to RGB, and the bandNum of C3 is used to quantize the R value to become the current Y component classifier.

[0214] In some embodiments, an example classifier (C7) may be obtained as a simplified variant of C0 / C3 and C6. To derive the bandNum classification for the current component C0 / C3, the sequence / current and adjacent samples of all three color components are used. For example, to classify the current U sample, sequence and adjacent Y / V, the current and adjacent U samples are used, similar to FIG. 6B, which is

number

[0215] In some embodiments, one special case of C7 may derive intermediate sample S using only 1 / 1 / 1 sequences or adjacent Y / U / V samples, which may also be obtained as a special case of C6 (color transformation using three components). S may be further fed into the bandNum classifier of C0 / C3. classIdx=bandS=(S*bandNumS)>>BitDepth

[0216] In some embodiments, similar to the C0 / C3 bandNum classifiers, C7 may also be combined with other classifiers to form a joint classifier. In some examples, C7 may not be the same as the later example (three-component joint bandNum classification for each Y / U / V component) that jointly uses the sequence and adjacent Y / U / V samples for classification.

[0217] In some embodiments, to reduce the cij signaling overhead and limit the value of S within the bit depth range, one constraint may be applied: sum of cij = 1. For example, enforcing c00 = (1 - sum of other cij). Which cij (c00 in this example) is enforced (derived by other coefficients) may be predefined or signaled at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level.

[0218] In some embodiments, a block activity classifier is implemented. For the luma component, each 4x4 block is classified into one of 25 classes. The classification index C is determined by its directionality D and quantized activity value

number

number

[0219] In some embodiments, D and

number

number

[0220] In some embodiments, a subsampled 1-D Laplacian calculation is applied to reduce the complexity of block classification. Figure 12B illustrates the subsampled Laplacian calculation according to some implementations of the present disclosure. As shown in Figure 12B, the same subsample position is used for gradient calculations in all directions.

[0221] Then, the maximum and minimum values ​​of the horizontal and vertical gradients D are

number

[0222] In some embodiments, the maximum and minimum values ​​of the two diagonal gradients are:

number

[0223] In some embodiments, these values ​​are compared with each other and with two thresholds t1 and t2 to derive the value of the directionality D, Step 1.

number

number

number

number

number

[0224] In some embodiments, the activity value A is

number

[0225] In some embodiments, A is further quantized to the range of 0 to 4, inclusive, and the quantized value is:

number

[0226] In some embodiments, the classification method is not applied to the chroma components in the picture.

[0227] In some embodiments, before filtering each 4x4 luma block, a geometric transformation such as a rotation or a diagonal and vertical flip is applied to the filter coefficients f(k,l) and the corresponding filter clipping values ​​c(k,l) according to the gradient values ​​calculated for that block. This is equivalent to applying these transformations to samples within the filter-enabled region. The idea is to make different blocks to which ALF is applied more similar by aligning their orientations.

[0228] In some embodiments, three geometric transformations are introduced, including skew, vertical flip and rotation. Diagonal: f D (k,l)=f(l,k), c D (k,l)=c(l,k) Vertical flip:fV (k,l)=f(k,Kl-1), c V (k,l)=c(k,Kl-1) Rotation:f R (k,l)=f(Kl-1,k), c R (k,l)=c(Kl-1,k) where K is the filter size and 0≦k,l≦K-1 are the coefficient coordinates, so location (0,0) is the upper-left corner and location (K-1,K-1) is the lower-right corner. These transformations are applied to the filter coefficients f(k,l) and clipping values ​​c(k,l) depending on the gradient values ​​calculated for that block. The relationship between the transformations and the four gradients in the four directions is summarized in Table 8-4 below. [Table 19]

[0229] In some embodiments, at the decoder side, when ALF is enabled for a CTB, each sample R(i,j) in that CU is filtered, resulting in a sample value R'(i,j), as shown below: R'(i,j)=R(i,j)+((Σ k≠0 Σ l≠0 f(k,l)×K(R(i+k,j+l)-R(i,j),c(k,l))+64)>>7) where f(k,l) denotes the decoded filter coefficients, K(x,y) is the clipping function, and c(k,l) denotes the decoded clipping parameters. The variables k and l are

number

number

[0230] In some embodiments, another classifier example (C8) uses inter-component / current component spatial activity information as a classifier. Similar to the block activity classifier above, one sample located at (k,l) may obtain the sample activity by: (1) Compute N directional gradients (Laplacian or forward / backward) (2) Sum the N directional gradients to obtain the activation A. (3) Quantize (or map) A to obtain a class index.

number

[0231] In some embodiments, for example, two directional Laplacian gradients to obtain A and

number

[0232] In some embodiments, A may then be further mapped to the range [0,4].

number

[0233] In some embodiments, other exemplary classifiers that use only current component information for the current component classification may be used as inter-component classifiers. For example, as shown in FIG. 5A and Table 1, luma sample information and eo-class are used to derive EdgeIdx and classify the current chroma sample. Other "non-inter-component" classifiers that may also be used as inter-component classifiers include edge direction, pixel intensity, pixel variance, pixel variance, sum of pixel Laplacian, Sobel operator, Compass operator, high-pass filter value, low-pass filter value, etc.

[0234] In some embodiments, multiple classifiers are used in the same POC. The current frame is divided into several regions, and each region uses the same classifier. For example, three different classifiers are used in POC0, and which classifier (0, 1, or 2) is used is signaled at the CTU level, as shown in Table 9 below. [Table 20]

[0235] In some embodiments, the maximum number of multiple classifiers (which may also be referred to as alternative offset sets) may be fixed or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. In one example, the fixed (predefined) maximum number of multiple classifiers is four. In that case, four different classifiers are used in POC0, and which classifier (0, 1, or 2) is used is signaled at the CTU level. Truncated-unary (TU) codes may be used to indicate the classifier used for each luma or chroma CTB. For example, as shown in Table 10 below, when the TU code is 0, no CCSAO is applied; when the TU code is 10, set 0 is applied; when the TU code is 110, set 1 is applied; when the TU code is 1110, set 2 is applied; and when the TU code is 1111, set 3 is applied. To indicate the classifier for the CTB (offset set index), fixed-length codes, Golomb-Rice codes, and exponential-Golomb codes may be used. Three different classifiers are used in POC1. [Table 21]

[0236] An example of CTB offset set indices for Cb and Cr is given in the 1280x720 sequence POC0 (if the CTU size is 128x128, the number of CTUs in a frame is 10x6). Cb in POC0 uses four offset sets, and Cr uses one offset set. As shown in Table 11-1 below, when the offset set index is 0, CCSAO is not applied; when the offset set index is 1, set 0 is applied; when the offset set index is 2, set 1 is applied; when the offset set index is 3, set 2 is applied; and when the offset set index is 4, set 3 is applied. The type indicates the position of the selected array luma sample (Yi). Different offset sets can have different types, band_nums, and corresponding offsets. [Table 22]

[0237] In some embodiments, an example of jointly using array / current and adjacent Y / U / V samples for classification is shown in Table 11-2 below (three-component joint bandNum classification for each Y / U / V component). At POC0, a {2, 4, 1} offset set is used for {Y, U, V}, respectively. Each offset set can be adaptively switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different offset sets can have different classifiers. For example, to classify the current Y4 luma sample as the candidate position (candPos) shown in Figures 6B and 6C, Yset0 selects {current Y4, array U4, array V4} as candidates, with different bandNum{Y, U, V} = {16, 1, 2}, respectively. When using {candY,candU,candV} as the sample values ​​of the selected {Y,U,V} candidates, the total number of classes is 32, and the class index derivation can be shown as follows: bandY=(candY*bandNumY)>>BitDepth; bandU=(candU*bandNumU)>>BitDepth; bandV=(candV*bandNumV)>>BitDepth; classIdx=bandY*bandNumU*bandNumV +bandU*bandNumV +bandV;

[0238] In some embodiments, the derivation of classIdx for the joint classifier can be expressed as an "or-shift" format to simplify the derivation process. For example, max bandNum={16,4,4}, classIdx=(bandY<<4)|(bandU<<2)|bandV.

[0239] Another example is the component V set1 classification in POC1, where candPos={neighbor Y8, neighbor U3, neighbor V0} is used with bandNum={4,1,2}, resulting in 8 classes. [Table 23]

[0240] In some embodiments, an example is provided in which the array and adjacent Y / U / V samples are jointly used for current Y / U / V sample classification (three-component joint edgeNum(C1) and bandNum classification for each Y / U / V component), as shown in Table 11-3 below. Edge CandPos is the center position used for the C1 classifier, edge bitMask is the C1 adjacent sample activation indicator, and edgeNum is the number of the corresponding C1 class. In this example, C1 is only applied to the Y classifier (thus edgeNum is equal to edgeNumY), and edge candPos is always Y4 (current / array sample position). However, C1 may also be applied to the Y / U / V classifier with edge candPos as the adjacent sample position.

[0241] If diff denotes the C1 comparison score of Y, then the classIdx derivation can be: bandY=(candY*bandNumY)>>BitDepth; bandU=(candU*bandNumU)>>BitDepth; bandV=(candV*bandNumV)>>BitDepth; edgeIdx=diff+(edgeNum>>1); bandIdx=bandY*bandNumU*bandNumV +bandU*bandNumV +bandV; classIdx=bandIdx*edgeNum+edgeIdx; [Table 24] [Table 25] [Table 26]

[0242] In some embodiments, max band_num (bandNumY, bandNumU, or bandNumV) may be fixed or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, if the decoder fixes max band_num=16 for each frame, 4 bits are signaled to indicate the band_num of C0 in the frame. Some other max band_num examples are listed in Table 12 below. [Table 27]

[0243] In some embodiments, the maximum number of classes or offsets (combinations of multiple classifiers used jointly, e.g., C1 edgeNum*C1 bandNumY*bandNumU*bandNumV) for each set (or all added sets) may be fixed or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. For example, max may be fixed at class_num=256*4 for all added sets, and an encoder conformance check or decoder norm check may be used to verify the constraint.

[0244] In some embodiments, restrictions may be applied to the C0 classification, for example restricting band_num (bandNumY, bandNumU, or bandNumV) to only power-of-two values. Instead of explicitly signaling band_num, the syntax band_num_shift is signaled. The decoder can use shift operations to avoid multiplications. Different band_num_shifts can be used for different components. class(C0)=(Y0>>band_num_shift)>>bit_depth

[0245] Another example operation considers rounding to reduce error. class(C0)=((Y0+(1<<(band_num_shift-1)))>>band_num_shift)>>bit_depth

[0246] For example, if band_num_max (Y, U, or V) is 16, then the possible band_num_shift candidates are 0, 1, 2, 3, 4, corresponding to band_num=1, 2, 4, 8, 16, as shown in Table 13. [Table 28] [Table 29]

[0247] In some embodiments, different classifiers are applied to Cb and Cr. The Cb and Cr offsets for all classes may be signaled separately. For example, different offsets are signaled and applied to different chroma components, as shown in Table 14 below. [Table 30]

[0248] In some embodiments, the maximum offset value is fixed or signaled at the sequence parameter set (SPS), adaptive parameter set (APS), picture parameter set (PPS), picture header (PH), slice header (SH), region, CTU, CU, sub-block, or sample level. For example, the maximum offset is [-15, 15]. Different components can have different maximum offset values.

[0249] In some embodiments, the offset signaling may use differential pulse code modulation (DPCM), for example, an offset of {3, 3, 2, 1, -1} may be signaled as {3, 0, -1, -1, -2}.

[0250] In some embodiments, the offsets may be stored in an APS or memory buffer for reuse for the next picture / slice, and an index may be signaled to indicate which stored previous frame offset is to be used for the current picture.

[0251] In some embodiments, the classifiers for Cb and Cr are the same. The Cb and Cr offsets for all classes may be jointly signaled, for example as shown in Table 15 below. [Table 31]

[0252] In some embodiments, the classifiers for Cb and Cr may be the same. For example, the Cb and Cr offsets for all classes may be jointly signaled by a sign flag difference, as shown in Table 16 below. According to Table 16, when the Cb offset is (3, 3, 2, -1), the derived Cr offset is (-3, -3, -2, 1). [Table 32]

[0253] In some embodiments, a sign flag may be signaled for each class, for example, as shown in Table 17 below: According to Table 17, when the Cb offset is (3, 3, 2, -1), the derived Cr offset is (-3, 3, 2, 1) according to the respective sign flags. [Table 33]

[0254] In some embodiments, the classifiers for Cb and Cr may be the same. For example, as shown in Table 18 below, the Cb and Cr offsets for all classes may be jointly signaled by weight differences. The weights (w) may be selected in a restricted table, e.g., ±¼, ±½, 0, ±1, ±2, ±4..., where |w| includes only values ​​that are powers of 2. According to Table 18, when the Cb offset is (3, 3, 2, -1), the derived Cr offset is (-6, -6, -4, 2) depending on the respective sign flags. [Table 34]

[0255] In some embodiments, weights may be signaled for each class, for example, as shown in Table 19 below: According to Table 19, when the Cb offset is (3, 3, 2, -1), the derived Cr offset is (-6, 12, 0, -1) according to the respective sign flags. [Table 35]

[0256] In some embodiments, when multiple classifiers are used in the same POC, different offset sets are signaled separately or jointly.

[0257] In some embodiments, previously decoded offsets may be stored for use in future frames. To reduce the overhead of offset signaling, an index may be signaled to indicate which previously decoded offset set is used for the current frame. For example, as shown in Table 20 below, the POC0 offset may be reused by POC2 by signaling offset set idx=0. [Table 36]

[0258] In some embodiments, the reuse offsets set idx for Cb and Cr may be different, for example as shown in Table 21 below. [Table 37]

[0259] In some embodiments, offset signaling can use additional syntax including start and length to reduce signaling overhead. For example, when band_num=256, only offsets for band_idx=37-44 are signaled. In the example in Table 22-1 below, the syntax for start and length are both 8-bit fixed length codes that should align to the band_num bits. [Table 38]

[0260] In some embodiments, if CCSAO is applied to all YUV3 components, the alignment and adjacent YUV samples may be used jointly for classification, and all of the offset signaling methods described above for Cb / Cr may be extended to Y / Cb / Cr. In some embodiments, different component offset sets may be stored and used separately (each component has its own stored set) or jointly (each component shares / reuses the same stored set). Examples of separate sets are shown in Table 22-2 below. [Table 39] TIFF0007827829000077.tif90170

[0261] In some embodiments, if the sequence bit depth is greater than 10 (or a specific bit depth), the offset may be quantized before signaling. At the decoder side, the decoded offset is dequantized before application, as shown in Table 23-1 below. For example, for a 12-bit sequence, the decoded offset is left-shifted (dequantized) by 2. [Table 40]

[0262] In some embodiments, the offset may be calculated as CcSaoOffsetVal=(1-2*ccsao_offset_sign_flag)*(ccsao_offset_abs<<(BitDepth-Min(10,BitDepth))).

[0263] In some embodiments, the concept of filter strength is further introduced herein. For example, the classifier offsets can be further weighted before being applied to the samples. The weights (w) can be selected from a table of power-of-two values, such as ±¼, ±½, 0, ±1, ±2, ±4..., where |w| includes only power-of-two values. The weight index can be signaled at the SPS / APS / PPS / PH / SH / Region (Set) / CTU / CU / Subblock / Sample level. The quantized offset signaling can be obtained as part of this weight application. As shown in Figure 6D, when a recursive CCSAO is applied, a similar weight indexing mechanism can be applied between the first and second stages.

[0264] In some instances, weighting different classifiers allows offsets of multiple classifiers to be applied to the same sample with a combination of weights. Similar weight indexing mechanisms can be signaled as described above. For example, offset_final=w*offset_1+(1-w)*offset_2, or offset_final=w1*offset_1+w2*offset_2+...

[0265] In some embodiments, instead of directly signaling the CCSAO parameters in the PH / SH, previously used parameters / offsets can be stored in an adaptive parameter set (APS) or memory buffer for reuse for the next picture / slice. An index can be signaled in the PH / SH to indicate which stored previous frame offset is to be used for the current picture / slice. A new APS ID can be created to maintain the CCSAO history offsets. The table below shows an example using Figure 6I, candPos, and bandNum{Y,U,V}={16,4,4}. In some examples, the candPos, bandNum, and offset signaling method can be fixed-length code (FLC) or other methods such as truncated unary (TU) code, exponential-Golomb code with degree k (EGk), signed EG0 (SVLC), or unsigned EG0 (UVLC). In this case, sao_cc_y_class_num (or cb, cr) is equal to sao_cc_y_band_num_y*sao_cc_y_band_num_u*sao_cc_y_band_num_v (or cb, cr). ph_sao_cc_y_aps_id is the parameter index used within this picture / slice. Note that the cb and cr components can follow the same signaling logic. [Table 41] TIFF0007827829000080.tif234170

[0266] aps_adaptation_parameter_set_id provides an identifier for the APS for reference by other syntax elements. When aps_params_type is equal to CCSAO_APS, the value of aps_adaptation_parameter_set_id shall be (for example) in the range 0 to 7, inclusive.

[0267] ph_sao_cc_y_aps_id specifies the aps_adaptation_parameter_set_id of the CCSAO APS to which the Y color component of the slice in the current picture refers. When ph_sao_cc_y_aps_id is present, the following applies: The value of sao_cc_y_set_signal_flag of an APS NAL unit with aps_params_type equal to CCSAO_APS and aps_adaptation_parameter_set_id equal to ph_sao_cc_y_aps_id shall be equal to 1, and the TemporalId of an APS Network Abstraction Layer (NAL) unit with aps_params_type equal to CCSAO_APS and aps_adaptation_parameter_set_id equal to ph_sao_cc_y_aps_id shall be less than or equal to the TemporalId of the current picture.

[0268] In some embodiments, an APS update mechanism is described herein. The maximum number of APS offset sets can be predefined or signaled at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level. Different components can have different maximum number limits. If an APS offset set is full, a newly added offset set can replace an existing stored offset in a first-in-first-out (FIFO), last-in-first-out (LIFO), or least recently used (LRU) mechanism, or an index value is received indicating which APS offset set should be replaced. In some examples, if the selected classifier consists of candPos / edge info / coding info..., etc., all classifier information can be obtained as part of the APS offset set and can be stored within the APS offset set with its offset value. In some cases, the update mechanism described above may be predefined or signaled at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level.

[0269] In some embodiments, a constraint called "pruning" can be applied, e.g., newly received classifier information and offsets cannot be the same as any of the stored APS offset sets (same or different components).

[0270] In some examples, when the C0 candPos / bandNum classifier is used, the maximum number of APS offset sets is 4 per Y / U / V, and FIFO updates are used for Y / V, and idx indicating updates are used for U. [Table 42] TIFF0007827829000082.tif112170

[0271] In some embodiments, the pruning criteria may be relaxed to provide a more flexible method for encoder tradeoffs, for example, allowing N offsets to vary when applying the pruning operation (e.g., N=4), and in another example, allowing a difference (represented as “thr”) in the value of each offset when applying the pruning operation (e.g., ±2).

[0272] In some embodiments, the two criteria may be applied simultaneously or separately, and the application of each criterion may be predefined or toggled at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level.

[0273] In some embodiments, N / thr can be predefined or can be switched at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level.

[0274] In some embodiments, the FIFO update can (1) circularly update the previously leftover set idx (if all have been updated, start again with set 0), as in the example above, or (2) update from set 0 each time. In some examples, when a new offset set is received, the update can be done at the PH (as in the example), or at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level.

[0275] For LRU updates, the decoder maintains a count table that counts the "total number of used offset sets," which can be refreshed at the SPS / APS / Group of Pictures (GOP) structure / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level. A newly received offset set replaces the most recently used offset set in an APS. If two stored offset sets have the same number, FIFO / LIFO can be used. For example, see component Y in Table 23-4 below. [Table 43] TIFF0007827829000084.tif32170

[0276] In some embodiments, different components may have different update mechanisms.

[0277] In some embodiments, different components (e.g., U / V) can share the same classifier (the same candPos / edge info / coding info / offset can further have weights with modifiers).

[0278] In some embodiments, a "patch" implementation may be used in the offset exchange mechanism because offset sets used by different pictures / slices may only have small offset value differences. In some embodiments, a "patch" implementation is Differential Pulse Code Modulation (DPCM). For example, when signaling a new offset set (OffsetNew), the offset value may be located on top of the existing APS-stored offset set (OffsetOld). The encoder signals only a delta value to update the old offset set (DPCM: OffsetNew = OffsetOld + Delta). In the following example shown in Table 23-5, options other than FIFO update (LRU, LIFO, or signaling an index indicating which set should be updated) may be used. YUV components may have the same update mechanism or may use different update mechanisms. In the example in Table 23-5, the classifier candPos / bandNum remains unchanged, but an additional flag may be signaled to indicate overwriting the set classifier (flag = 0: update only the set offset; flag = 1: update both the set classifier and the set offset). [Table 44] TIFF0007827829000086.tif92170

[0279] In some embodiments, the DPCM delta offset value may be signaled with FLC / TU / EGk (order = 0, 1, ...) code. For each offset set, one flag may be signaled indicating whether DPCM signaling is enabled. The DPCM delta offset value, or the newly added offset value (signaled directly without DPCM when enabling APS DPCM = 0) (ccsao_offset_abs), may be dequantized / mapped before being applied to the target offset (CcSaoOffsetVal). The offset quantization step can be predefined or signaled at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level. For example, one method is to directly signal the offset with a quantization step = 2. CcSaoOffsetVal=(1-2*ccsao_offset_sign_flag)*(ccsao_offset_abs<<1) Another method is to use a DPCM signaling offset with quantization step=2. CcSaoOffsetVal=CcSaoOffsetVal+(1-2*ccsao_offset_sign_flag)*(ccsao_offset_abs<<1)

[0280] In some embodiments, to reduce direct offset signaling overhead, a constraint may be applied, for example, the updated offset value must have the same sign as the old offset value. By using such an inferred offset sign, the newly updated offset does not need to carry the sign flag again (ccsao_offset_sign_flag is inferred to be the same as the old offset).

[0281] In some embodiments, sample processing is described as follows: Let R(x,y) be the input luma or chroma sample value before CCSAO, and R'(x,y) be the output luma or chroma sample value after CCSAO. offset = ccsao_offset[class_index of R(x, y)] R'(x, y) = Clip3(0, (1 << bit_depth) - 1, R(x, y) + offset)

[0282] According to the above equation, each luma or chroma sample value R(x, y) is classified using the classifier indicated by the current picture and / or the current offset set idx. The corresponding offset of the derived class index is added to each luma or chroma sample value R(x, y). The clip function Clip3 is applied to (R(x, y) + offset) to create an output luma or chroma sample value R'(x, y) within the bit depth dynamic range, for example, in the range from 0 to (1 << bit_depth) - 1.

[0283] FIG. 13 is a block diagram showing that CCSAO is applied and other in-loop filters have different combinations of clipping according to some implementation examples of the present disclosure.

[0284] In some embodiments, when CCSAO is operated by other in-loop filters, the clip operation can include the following. (1) Clipping after addition The following equations show (a) an example where CCSAO is operated by SAO and BIF, or (b) an example where CCSAO replaces SAO but is still operated by BIF. (a) I OUT = clip1(I C + ΔI SAO + ΔI BIF ++ ΔI CCSAO ) (b) I OUT = clip1(I C + ΔI CCSAO + ΔI BIF ) (2) Clipping after addition operated by BIF In some embodiments, the clip order can be switched. (a)I OUT =clip1(I C +ΔI SAO ) I' OUT =clip1(I OUT +ΔI BIF ) I” OUT =clip1(I” OUT +ΔI CCSAO ) (b)I OUT =clip1(I C +ΔI BIF ) I' OUT =clip1(I' OUT +ΔI CCSAO ) (3) Clipping after partial addition (a)I OUT =clip1(I C +ΔI SAO +ΔI BIF ) I' OUT =clip1(I OUT +ΔI CCSAO )

[0285] In some embodiments, different clipping combinations provide different trade-offs between correction accuracy and hardware temporary buffer size (register or SRAM bit width).

[0286] FIG. 13(a) shows SAO / BIF offset clipping. FIG. 13(b) shows one additional bit-depth clipping for CCSAO. FIG. 13(c) shows joint clipping after adding SAO / BIF / CCSAO offsets to the input samples. More specifically, for example, FIG. 13(a) shows the current BIF design when interacting with SAO. Offsets from the SAO and BIF are added to the input samples, followed by one bit-depth clipping. However, as shown in FIG. 13(b) and FIG. 13(c), when a CCSAO is also added at the SAO stage, two possible clipping designs can be selected: (1) adding one additional bit-depth clipping for the CCSAO, and (2) a harmonized design that performs joint clipping after adding SAO / BIF / CCSAO offsets to the input samples. In some embodiments, the clipping designs described above differ only for luma samples because the BIF is applied only to luma samples.

[0287] In some embodiments, boundary processing is described below. If the array used for classification and any of the neighboring luma (chroma) samples are outside the current picture, the CCSAO is not applied to the current chroma (luma) sample. Figure 14A is a block diagram illustrating that, according to some implementation examples of the present disclosure, if the array used for classification and any of the neighboring luma (chroma) samples are outside the current picture, the CCSAO is not applied to the current chroma (luma) sample. For example, in Figure 14A(a), when a classifier is used, the CCSAO is not applied to the chroma components in the leftmost column of the current picture. For example, when C1' is used, the CCSAO is not applied to the chroma components in the leftmost column and the topmost row of the current picture, as shown in Figure 14A(b).

[0288] Figure 14B is a block diagram illustrating that, according to some example implementations of the present disclosure, if the array used for classification and any of the adjacent luma or chroma samples are located outside the current picture, a CCSAO is applied to the current luma or chroma sample. In some embodiments, if the array used for classification and any of the adjacent luma or chroma samples are located outside the current picture, as shown in Figure 14B(b), the missing sample may be repeated or mirror-padded to create a sample for classification, as shown in Figure 14B(a), and a CCSAO may be applied to the current luma or chroma sample. In some embodiments, the invalidation / repetition / mirror picture boundary processing methods disclosed herein may also be applied to the subpicture / slice / tile / CTU / 360 virtual boundary if the array used for classification and any of the adjacent luma (chroma) samples are located outside the current subpicture / slice / tile / patch / CTU / 360 virtual boundary.

[0289] For example, a picture is divided into one or more tile rows and one or more tile columns, where a tile is a series of CTUs that cover a rectangular area of ​​the picture.

[0290] A slice consists of an integer number of complete tiles or an integer number of contiguous complete CTU rows within a tile of a picture.

[0291] A subpicture contains one or more slices that collectively cover a rectangular area of ​​the picture.

[0292] In some embodiments, 360-degree video is captured on a sphere, which inherently has no "boundaries," and reference samples located outside the boundaries of the reference picture in the projected domain can always be obtained from neighboring samples in the spherical domain. For projection formats consisting of multiple surfaces, discontinuities appear between two or more adjacent surfaces in a frame-packed picture, regardless of the type of compact frame-packing arrangement used. VVC introduces vertical and / or horizontal virtual boundaries where in-loop filtering operations are disabled, and the location of these boundaries is signaled in the SPS or picture header. Compared to using two tiles, one for each pair of consecutive surfaces, the use of 360 virtual boundaries is more flexible because it does not require the surface size to be a multiple of the CTU size. In some embodiments, the maximum number of vertical 360 virtual boundaries is three, and the maximum number of horizontal 360 virtual boundaries is also three. In some embodiments, the distance between two virtual boundaries is greater than or equal to the CTU size, and the granularity of the virtual boundaries is 8 luma samples, e.g., an 8x8 sample grid.

[0293] Figure 14C is a block diagram illustrating that, according to some example implementations of the present disclosure, CCSAO is not applied to a current chroma sample if the corresponding selected array or neighboring luma sample used for classification is outside the virtual space defined by the virtual boundary. In some embodiments, the virtual boundary (VB) is a virtual line separating spaces within a picture frame. In some embodiments, if a virtual boundary (VB) is applied within the current frame, CCSAO is not applied to chroma samples selected for corresponding luma positions that are outside the virtual space defined by the virtual boundary. Figure 14C shows an example of virtual boundaries for a C0 classifier with nine luma position candidates. For each CTU, CCSAO is not applied to chroma samples whose corresponding selected luma positions are outside the virtual space encompassed by the virtual boundary. For example, in Figure 14C(a), when the selected Y7 luma sample position is on the opposite side of the horizontal virtual boundary 1406, which is located four pixels from the bottom of the frame, CCSAO is not applied to chroma sample 1402. For example, in Figure 14C(b), when the selected Y5 luma sample position is located on the opposite side of the vertical virtual boundary 1408, which is located y pixels from the right edge of the frame, CCSAO is not applied to the chroma sample 1404.

[0294] FIG. 15 illustrates that, in some implementations of the present disclosure, repetition or mirror padding may be applied to luma samples outside the virtual boundary. FIG. 15(a) illustrates an example of repetition padding. If the original Y7 is selected to be the classifier located at the bottom of VB1502, the Y4 luma sample value is used for classification (copied to the Y7 position) rather than the original Y7 luma sample value. FIG. 15(b) illustrates an example of mirror padding. If Y7 is selected to be the classifier located at the bottom of VB1504, the Y1 luma sample value, which is symmetrical to the Y7 value relative to the Y0 luma sample, is used for classification rather than the original Y7 luma sample value. This padding method allows the possibility of applying CCSAO to more chroma samples, thereby realizing greater coding gain.

[0295] In some embodiments, restrictions may be applied to reduce the line buffer required by the CCSAO and simplify boundary processing condition checks. Figure 16 shows that, according to some example implementations of the present disclosure, if all nine array and adjacent luma samples are used for classification, an additional luma line buffer may be required, namely, the full line luma samples of line -5 above the current VB1602. Figure 10B(a) shows an example in which only six luma candidates are used for classification, which reduces the line buffer and does not require any of the additional boundary checks in Figures 14A and 14B.

[0296] In some embodiments, using luma samples for CCSAO classification may increase the luma line buffer and therefore decoder hardware implementation costs. Figure 17 illustrates how, in some example implementations of this disclosure, nine luma candidate CCSAOs crossing VB 1702 in AVS may increase two additional luma line buffers. For luma and chroma samples above the virtual boundary (VB) 1702, DBF / SAO / ALF are processed in the current CTU row. For luma and chroma samples below VB 1702, DBF / SAO / ALF are processed in the next CTU row. In the AVS decoder hardware design, the pre-DBF samples for luma lines -4 to -1, the pre-SAO samples for line -5, and the pre-DBF samples for chroma lines -3 to -1, and the pre-SAO samples for line -4 are stored as line buffers for DBF / SAO / ALF processing of the next CTU row. When processing the next CTU row, luma and chroma samples that are not in the line buffer are unavailable. However, for example, at chroma line -3(b), chroma samples are processed in the next CTU row, but CCSAO requires pre-SAO luma sample lines -7, -6, and -5 for classification. Pre-SAO luma sample lines -7 and -6 are not in the line buffer and therefore unavailable. Adding pre-SAO luma sample lines -7 and -6 to the line buffer increases the decoder hardware implementation cost. In some examples, luma VB (line -4) and chroma VB (line -3) may be different (not aligned).

[0297] 17, FIG. 18A illustrates that in VVC, nine luma candidate CCSAOs intersecting VB 1802 may increase one additional luma line buffer according to some implementations of the present disclosure. VB may be different in different standards. In VVC, luma VB is line -4 and chroma VB is line -2, and therefore nine candidate CCSAOs may increase one luma line buffer.

[0298] In some embodiments, in the first solution, if any of the luma candidates for a chroma sample crosses VB (is outside the current chroma sample VB), the CCSAO is disabled for that chroma sample. Figures 19A-19C show that, in some implementation examples of the present disclosure, if any of the luma candidates for a chroma sample crosses VB1902 (is outside the current chroma sample VB), the CCSAO is disabled for that chroma sample in AVS and VVC. Figure 14C also shows some examples of this implementation.

[0299] In some embodiments, in the second solution, for "cross-VB" luma candidates, repeat padding is used in the CCSAO from the luma line close to and opposite the VB, e.g., luma line -4. In some embodiments, repeat padding from the luma closest to the adjacent luma below the VB is implemented for the "cross-VB" chroma candidates. Figures 20A-20C show that, in some implementation examples of the present disclosure, in AVS and VVC, CCSAO is enabled using repeat padding for chroma samples if any of the luma candidates for the chroma samples crosses the VB 2002 (is outside the current chroma sample VB). Figure 14C(a) also shows some examples of this implementation example.

[0300] In some embodiments, in the third solution, for "cross-VB" luma candidates, mirror padding is used for CCSAO from below the luma VB. Figures 21A-21C show that some implementation examples of the present disclosure enable CCSAO using mirror padding for chroma samples in AVS and VVC when any of the luma candidates for the chroma samples crosses VB 2102 (is outside the current chroma sample VB). Figures 14C(b) and 14B(b) also show some examples of this implementation. In some embodiments, in the fourth solution, "double-sided symmetric padding" is used to apply CCSAO. Figures 22A-22B show that some implementation examples of the present disclosure enable CCSAO using double-sided symmetric padding for several examples of different CCSAO shapes (e.g., 9 luma candidates (Figure 22A) and 8 luma candidates (Figure 22B)). For luma sample sets with a center luma sample aligned with the chroma samples, if one side of the luma sample set is outside VB2 202, double-sided symmetric padding is applied to both sides of the luma sample set. For example, in Figure 22A, luma samples Y0, Y1, and Y2 are outside VB2 202, so Y0, Y1, Y2 and Y6, Y7, Y8 are both padded using Y3, Y4, and Y5. For example, in Figure 22B, luma sample Y0 is outside VB2 202, so Y0 is padded using Y2 and Y7 is padded using Y5.

[0301] 18B illustrates how, in some implementations of the present disclosure, an array or neighboring chroma samples are used to classify the current luma sample, and the selected chroma candidate may cross VB and require additional chroma line buffers. Solutions 1-4 similar to those described above may be applied to address this issue.

[0302] Solution 1 is to disable CCSAO for luma samples when any of their chroma candidates may cross VB.

[0303] Solution 2 is to use repeated padding from the chroma closest to the adjacent chroma below VB for the "cross VB" chroma candidate.

[0304] Solution 3 is to use mirror padding from below chroma VB for the "cross VB" chroma candidate.

[0305] Solution 4 is to use "double-sided symmetric padding". For a candidate set located at the center of the CCSAO array chroma sample, if one side of the candidate set is outside the VB, double-sided symmetric padding is applied to both sides.

[0306] The padding method gives the possibility to apply CCSAO to more luma or chroma samples, and therefore more coding gain can be realized.

[0307] In some embodiments, in a bottom picture (or slice, tile, brick) boundary CTU row, samples below the VB are processed in the current CTU row, and therefore the above special treatment (solutions 1, 2, 3, and 4) does not apply to this bottom picture (or slice, tile, brick) boundary CTU row. For example, a 1920x1080 frame is divided into 128x128 CTUs. The frame contains 15x9 CTUs (rounded up). The bottom CTU row is the 15th CTU row. The decoding process is performed CTU-by-CTU for each CTU row. Deblocking needs to be applied along the horizontal CTU boundary between the current CTU row and the next CTU row. Inside one CTU, in the bottom 4 / 2 luma / chroma lines, the VB of the CTB is applied to each CTU row because the DBF samples (in the case of VVC) are processed in the next CTU row and are not available for the CCSAO of the current CTU row. However, for the bottom CTU row of the picture frame, there are no next CTU rows remaining, so the bottom 4 / 2 luma / chroma line DBF samples are available in the current CTU row and are DBF processed in the current CTU row.

[0308] In some embodiments, the VBs shown in Figures 13-22 may be swapped with the boundaries of the subpicture / slice / tile / patch / CTU / 360 virtual boundary. In some embodiments, the positions of the chroma and luma samples in Figures 13-22 may be switched. In some embodiments, the positions of the chroma and luma samples in Figures 13-22 may be swapped with the positions of the first chroma sample and the second chroma sample. In some embodiments, the VBs of the ALF within the CTU may be generally horizontal. In some embodiments, the boundaries of the subpicture / slice / tile / patch / CTU / 360 virtual boundary may be horizontal or vertical.

[0309] In some embodiments, restrictions may be applied to reduce the line buffer required by the CCSAO and simplify boundary processing condition checks, as illustrated in Figure 16. Figure 23 illustrates restrictions to use a limited number of luma candidates for classification according to some implementations of the present disclosure. Figure 23(a) illustrates restrictions to use only six luma candidates for classification. Figure 23(b) illustrates restrictions to use only four luma candidates for classification.

[0310] In some embodiments, an application region is implemented. The CCSAO application region unit can be based on a CTB, i.e., the on / off control, CCSAO parameters (offsets used for classification offset set index, luma candidate position, band_num, bitmask, etc.) are the same within one CTB.

[0311] In some embodiments, the application region may not be aligned to the CTB boundary. For example, the application region may be shifted rather than aligned to the chroma CTB boundary. While the syntax (on / off control, CCSAO parameters) is still shown for each CTB, the true application region is not aligned to the CTB boundary. Figure 24 illustrates that, according to some implementations of the present disclosure, the CCSAO application region is not aligned to the CTB / CTU boundary 2406. For example, the application region is not aligned to the chroma CTB / CTU boundary 2406, but is shifted (4,4) samples up and left relative to the VB 2408. This unaligned CTB boundary design benefits the deblocking process because the same deblocking parameters are used for each 8x8 deblocking process region.

[0312] In some embodiments, the CCSAO application region unit (mask size) may be variable (larger or smaller than the CTB size), as shown in Table 24. The mask size may be different for different components. The mask size may be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, in the PH, a series of mask on / off flags and offset set indexes are shown to indicate each CCSAO region information. [Table 45]

[0313] In some embodiments, the CCSAO application area frame partition may be fixed, for example, by dividing the frame into N regions. Figure 25 illustrates that, according to some implementations of the present disclosure, the CCSAO application area frame partition may be fixed by a CCSAO parameter.

[0314] In some embodiments, each region can have its own region on / off control flag and CCSAO parameters. Also, if the region size is larger than the CTB size, it can have both a CTB on / off control flag and a region on / off control flag. Figures 25(a) and 25(b) show some examples of dividing a frame into N regions. Figure 25(a) shows a vertical division consisting of four regions. Figure 25(b) shows a square division consisting of four regions. In some embodiments, similar to the picture-level CTB on all-on control flag (ph_cc_sao_cb_ctb_control_flag / ph_cc_sao_cr_ctb_control_flag), if the region on / off control flag is off, a CTB on / off flag can be further signaled. Otherwise, CCSAO is applied to all CTBs in this region without further signaling of the CTB flag.

[0315] In some embodiments, different CCSAO application regions can share the same region on / off control and CCSAO parameters. For example, in Figure 25(c), regions 0-2 share the same parameters, and regions 3-15 share the same parameters. Figure 25(c) also shows that the region on / off control flags and CCSAO parameters can be signaled in Hilbert scan order.

[0316] In some embodiments, the CCSAO application area unit can be a quadtree / binarytree / ternarytree split from the picture / slice / CTB level. Similar to CTB splitting, a series of split flags are signaled to indicate the CCSAO application area partition. Figure 26 shows that in some implementations of the present disclosure, the CCSAO application area can be a binary tree (BT) / quadtree (QT) / ternary tree (TT) split from the frame / slice / CTB level.

[0317] Figure 27 is a block diagram illustrating multiple classifiers used and switched at different levels within a picture frame according to some implementations of this disclosure. In some embodiments, when multiple classifiers are used in a frame, the method of applying the classifier set index may be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, as shown in Table 25 below, four classifier sets are used within a frame and switched at the PH. Figures 27(a) and 27(c) show the default fixed region classifiers. Figure 27(b) shows the classifier set index signaled at the mask / CTB level, where 0 means CCSAO off for this CTB and 1-4 means set index. [Table 46]

[0318] In some embodiments, for a default region, the region-level flag may be signaled if the CTB in this region does not use the default set index (e.g., the region-level flag is 0), but uses other classifier sets in this frame. For example, if the default set index is used, the region-level flag will be 1. For example, in four regions of a square plot, the following classifier sets are used, as shown in Table 26-1 below: [Table 47]

[0319] Figure 28 is a block diagram illustrating that, in some implementations of the present disclosure, the CCSAO application region partition can be dynamic and can be switched at the picture level. For example, Figure 28(a) shows that three CCSAO offset sets are used in this POC (set_num=3), thus dividing the picture frame vertically into three regions. Figure 28(b) shows that four CCSAO offset sets are used in this POC (set_num=4), thus dividing the picture frame horizontally into four regions. Figure 28(c) shows that three CCSAO offset sets are used in this POC (set_num=3), thus dividing the picture frame horizontally into three regions. Each region can have its own region all-on flag to save CTB on / off control bits. The number of regions depends on the picture set_num being signaled. The CCSAO application area can be a specific area according to the coding information in the block (sample position, sample coding mode, loop filter parameters, etc.). For example, (1) the CCSAO application region may be applied only when the sample is coded in skip mode, or (2) the CCSAO application region may include only N samples along a CTU boundary, or (3) the CCSAO application region may include only samples on an 8x8 grid within a frame, or (4) the CCSAO application region may include only DBF-filtered samples, or (5) the CCSAO application region may include only the top M and left N rows within a CU, or (6) the CCSAO application region may include only intra-coded samples, or (7) the CCSAO application region may include only samples within cbf=0 blocks, or (8) the CCSAO application region may be located only on blocks with a block QP of [N,M], where (N,M) may be predefined or may be signaled at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level. Inter-component coding information may be taken into account, and (9) the CCSAO application region is on the chroma samples whose constellation luma samples are within the cbf=0 block.

[0320] In some embodiments, whether to introduce coding information application region restrictions can be predefined or signaled in a control flag at the SPS / APS / PPS / PH / SH / Region (per alternative set) / CTU / CU / Subblock / Sample level to indicate whether the specified coding information is included / excluded in the CCSAO application. The decoder omits CCSAO processing for those areas according to the predefined condition or control flag. For example, YUV uses different predefined / flag control conditions that switch at the region (set) level. The CCSAO application decision can be made at the CU / TU / PU or sample level. [Table 48]

[0321] Another example is to reuse all or part of the two-way enabling constraints (which are predefined). bool isInter=(currCU.predMode==MODE_INTER)?true:false; if(ccSaoParams.ctuOn[ctuRsAddr] &&((TU::getCbf(currTU,COMPONENT_Y)||isInter==false)&&(currTU.cu->qp>17)) &&(128>std::max(currTU.lumaSize().width,currTU.lumaSize().height)) &&((isInter==false)||(32>std::min(currTU.lumaSize().width,currTU.lumaSize().height))))

[0322] In some embodiments, excluding specific areas may favor CCSAO statistics collection. Offset derivation may be more precise or suitable for those areas that really need to be corrected. For example, blocks with cbf=0 usually mean that blocks that do not need further correction are perfectly predicted. Excluding those blocks may favor offset derivation for other areas.

[0323] Different application areas can use different classifiers. For example, in a CTU, skip mode uses C1, an 8x8 lattice uses C2, and skip mode and an 8x8 lattice use C3. For example, in a CTU, skip mode coded samples use C1, CU-centered samples use C2, and CU-centered skip mode coded samples use C3. Figure 29 illustrates how some implementations of the present disclosure allow a CCSAO classifier to consider current or inter-component coding information. For example, different coding modes / parameters / sample positions can form different classifiers. Different coding information can be combined to form a joint classifier. Different regions can use different classifiers. Figure 29 also illustrates another example of application areas.

[0324] In some embodiments, a predefined or flag-controlled "coding information exclusion zone" mechanism may be used in DBF / Pre-SAO / SAO / BIF / CCSAO / ALF / CCALF / NN Loop Filter (NNLF), or other loop filters.

[0325] In some embodiments, the CCSAO syntax implemented is shown in Table 27 below. In some instances, the binarization of each syntax element may be changed. In AVS3, the term patch is similar to slice, and patch header is similar to slice header. FLC stands for fixed length code. TU stands for shortened unary code. EGk stands for exponential-Golomb code with degree k, where k may be fixed. SVLC stands for signed EG0. UVLC stands for unsigned EG0. [Table 49] TIFF0007827829000092.tif254161TIFF0007827829000093.tif34170

[0326] If a high-order flag is off, the low-order flags may be inferred from the off state of the flags and do not need to be signaled. For example, if ph_cc_sao_cb_flag is false in this picture, then ph_cc_sao_cb_band_num_minus1, ph_cc_sao_cb_luma_type, cc_sao_cb_offset_sign_flag, cc_sao_cb_offset_abs, ctb_cc_sao_cb_flag, cc_sao_cb_merge_left_flag, and cc_sao_cb_merge_up_flag are inferred to be absent and false.

[0327] In some embodiments, the ccsao_enabled_flag of the SPS is adjusted with the SAO enable flag of the SPS, as shown in Table 28 below. [Table 50]

[0328] In some embodiments, ph_cc_sao_cb_ctb_control_flag, ph_cc_sao_cr_ctb_control_flag indicate whether to enable the granularity of CTB on / off control for Cb / Cr. If ph_cc_sao_cb_ctb_control_flag and ph_cc_sao_cr_ctb_control_flag are enabled, ctb_cc_sao_cb_flag and ctb_cc_sao_cr_flag may be further signaled. Otherwise, whether CCSAO is applied to the current picture depends on ph_cc_sao_cb_flag, ph_cc_sao_cr_flag, and ctb_cc_sao_cb_flag and ctb_cc_sao_cr_flag are not further signaled at the CTB level.

[0329] In some embodiments, to reduce bit overhead, for ph_cc_sao_cb_type and ph_cc_sao_cr_type, a flag may be further signaled to distinguish whether the center array luma position (Y0 position in FIG. 10) is used for classification for chroma samples. Similarly, when cc_sao_cb_type and cc_sao_cr_type are signaled at the CTB level, a flag may be further signaled by the same mechanism. For example, if the number of C0 luma position candidates is 9, cc_sao_cb_type0_flag is further signaled to distinguish whether the center array luma position is used, as shown in Table 29 below. If the center array luma position is not used, cc_sao_cb_type_idc is used to indicate which of the remaining eight adjacent luma positions is used. [Table 51]

[0330] Table 30 below shows an example where a single (set_num=1) or multiple (set_num>1) classifiers are used within a frame in AVS. The syntax notation can be mapped to the notation used above. [Table 52] [Table 53]

[0331] When combined with Figure 25 or Figure 27, where each region has its own set, an example syntax may include region on / off control flags (picture_ccsao_lcu_control_flag[compIdx][setIdx]), as shown in Table 31 below. [Table 54]

[0332] In some embodiments, for high level syntax, pps_ccsao_info_in_ph_flag and gci_no_sao_constraint_flag may be added.

[0333] In some embodiments, pps_ccsao_info_in_ph_flag equal to 1 specifies that CCSAO filter information may be present in a PH syntax structure but not in a slice header that references a PPS that does not contain a PH syntax structure. pps_ccsao_info_in_ph_flag equal to 0 specifies that CCSAO filter information may be present in a slice header that references a PPS but not in a PH syntax structure. When not present, the value of pps_ccsao_info_in_ph_flag is inferred to be equal to 0.

[0334] In some embodiments, gci_no_ccsao_constraint_flag equal to 1 specifies that sps_ccsao_enabled_flag is equal to 0 for all pictures in OlsInScope. gci_no_ccsao_constraint_flag equal to 0 imposes no such constraint. In some embodiments, a video bitstream includes one or more Output Layer Sets (OLSs) according to a rule. In the examples herein, OlsInScope references one or more OLSs that are in scope. In some examples, the profile_tier_level( ) syntax structure provides level information, and optionally provides the profile, layer, subprofile, and general constraint information to which OlsInScope adheres. When the profile_tier_level( ) syntax structure is included within a VPS, OlsInScope is one or more OLSs specified by the VPS. When the profile_tier_level( ) syntax structure is included within an SPS, OlsInScope is an OLS that includes only the lowest layer among the layers that reference the SPS, where this lowest layer is an independent layer.

[0335] In some embodiments, extensions to intra- and inter-prediction post-SAO filters are further illustrated below. In some embodiments, the SAO classification method disclosed in this disclosure (including classification of inter-component sample / coding information) can act as a post-prediction filter, and the prediction can be intra, inter, or other prediction tools, such as intra-block copy. Figure 30 is a block diagram illustrating the SAO classification method disclosed in this disclosure acting as a post-prediction filter according to some implementation examples of this disclosure.

[0336] In some embodiments, a corresponding classifier is selected for each Y, U, and V component. For each component prediction sample, the corresponding classifier is first classified and a corresponding offset is added. For example, each component can use the current sample and neighboring samples for classification. As shown in Table 32 below, Y uses the current Y sample and neighboring Y samples, and U / V uses the current U / V sample for classification. Figure 31 is a block diagram illustrating that, for a post-prediction SAO filter, each component can use the current sample and neighboring samples for classification, according to some implementations of the present disclosure. [Table 55]

[0337] In some embodiments, the refined prediction samples (Ypred', Upred', Vpred') are updated by adding the corresponding class offsets and are then used for intra, inter, or other prediction.

[0338] Ypred'=clip3(0,(1< <bit_depth)-1,Ypred+h_Y[i])

[0339] Upred'=clip3(0,(1< <bit_depth)-1,Upred+h_U[i])

[0340] Vpred'=clip3(0,(1< <bit_depth)-1,Vpred+h_V[i])

[0341] In some embodiments, in addition to the current chroma component, for the chroma U and V components, an inter-component offset (Y) may be used for further offset classification. For example, as shown below in Table 33, an additional inter-component offset (h'_U, h'_V) may be added with the current component offset (h_U, h_V). [Table 56]

[0342] In some embodiments, the improved prediction samples (Upred”, Vpred”) are updated by adding the corresponding class offset and are then used for intra, inter, or other prediction.

[0343] Upred”=clip3(0,(1< <bit_depth)-1,Upred’+h’_U[i])

[0344] Vpred”=clip3(0,(1< <bit_depth)-1,Vpred’+h’_V[i])

[0345] In some embodiments, intra and inter predictions may use different SAO filter offsets.

[0346] FIG. 32 is a block diagram illustrating the SAO classification method disclosed in this disclosure acting as a post-reconstruction filter according to some implementations of the present disclosure.

[0347] In some embodiments, the SAO / CCSAO classification method disclosed herein (including classification of inter-component samples / coding information) can act as a filter applied to the reconstructed samples of a tree unit (TU). As shown in Figure 32, the CCSAO can act as a post-reconstruction filter, i.e., using the reconstructed samples (after addition of prediction / residual samples, but before deblocking) as input for classification and compensating luma / chroma samples before entering neighboring intra / inter prediction. The CCSAO post-reconstruction filter can reduce distortion of the current TU sample and provide better prediction for neighboring intra / inter blocks. With more accurate prediction, better compression efficiency can be expected.

[0348] FIG. 33 is a flow diagram illustrating an example process 3300 for decoding a video signal using inter-component correlation according to some implementations of this disclosure.

[0349] In one aspect, video decoder 30 (shown in FIG. 3) receives a picture frame from a video signal, the picture frame including a first component and a second component (3310A).

[0350] Video decoder 30 determines a classifier for each sample of the second component using a set of weighted sample values ​​from a first set of samples of the first component associated with each sample of the second component and a second set of samples of the second component associated with each sample of the second component (3320A). In some embodiments, the first set of samples of the first component includes an array sample of the first component for each sample of the second component and neighboring samples of the array sample of the first component, and the second set of samples of the second component includes a current sample of the second component for each sample of the second component and neighboring samples of the current sample of the second component (3320A-1).

[0351] Video decoder 30 determines a sample offset for each sample of the second component according to the classifier (3330A).

[0352] Video decoder 30 modifies each sample of the second component based on the determined sample offset (3340A).

[0353] In some embodiments, the first component is a component selected from the group consisting of a luma component, a first chroma component, and a second chroma component, and the second component is a component selected from the group consisting of a luma component, a first chroma component, and a second chroma component, and the first component is different from the second component.

[0354] In some embodiments, the picture frame further includes a third component, and determining a classifier for each sample of the second component (3320A) further includes determining the classifier for each sample of the second component using a set of weighted sample values ​​and an additional set of weighted sample values ​​from a third set of samples of the third component associated with the each sample of the second component, wherein the third set of samples of the third component includes an array sample of the third component for the each sample of the second component and adjacent samples of the array sample of the third component.

[0355] In another aspect, video decoder 30 receives a picture frame from a video signal, the picture frame including a first component, a second component, and a third component (3310B).

[0356] Video decoder 30 determines a classifier for each sample of the second component using a set of weighted sample values ​​from the first set of samples of the first component associated with each sample of the second component and the third set of samples of the third component associated with each sample of the second component (3320B). In some embodiments, the first set of samples of the first component includes an array sample of the first component for each sample of the second component and neighboring samples of the array sample of the first component, and the third set of samples of the third component includes an array sample of the third component for each sample of the second component and neighboring samples of the array sample of the third component (3320B-1).

[0357] Video decoder 30 determines a sample offset for each sample of the second component according to the classifier (3330B).

[0358] Video decoder 30 modifies each sample of the second component based on the determined sample offset (3340B).

[0359] In some embodiments, the first component is a component selected from the group consisting of a luma component, a first chroma component, and a second chroma component; the second component is a component selected from the group consisting of a luma component, a first chroma component, and a second chroma component; the third component is a component selected from the group consisting of a luma component, a first chroma component, and a second chroma component; and the first component, the second component, and the third component are different components.

[0360] In some embodiments, determining a classifier for each sample of the second component (3320B) further includes determining a classifier for each sample of the second component using a set of weighted sample values ​​and an additional set of weighted sample values ​​from a second set of samples of the second component associated with the each sample of the second component, wherein the second set of samples of the second component includes a current sample of the second component for the each sample of the second component and neighboring samples of the current sample of the second component.

[0361] In some embodiments, determining a classifier for each sample of the second component (3320A or 3320B) includes determining a classifier for each sample of the second component using an addition of a set of weighted sample values.

[0362] In some embodiments, determining a classifier for each sample of the second component (3320A or 3320B) includes determining a classifier for each sample of the second component using an addition of one set of weighted sample values ​​and an additional set of weighted sample values.

[0363] In some embodiments, determining a classifier for each sample of the second component using summation (3320A or 3320B) may be performed by dividing the classified sample value S by:

number

[0364] In some embodiments, addition

number

[0365] In some embodiments, the classification sample value S is within the dynamic bit depth of the video signal.

[0366] In some embodiments, the classifier for each sample of the second component includes a first classifier that uses summation of a set of weighted sample values ​​jointly combined with a second classifier determined from edge directions and intensities of array samples of the first component for each sample of the second component, and the first classifier is different from the second classifier.

[0367] In some embodiments, one of the weighting coefficients for the current sample of the second component is derived from the other weighting coefficients.

[0368] In some embodiments, the classification sample value S is used to determine a class index for the classifier as classIdx=(S*bandNumS)>>BitDepth, where bandNumS is the band number associated with the sample value of S and BitDepth is the dynamic bit depth of the video signal.

[0369] 34 illustrates a computing environment 3410 coupled to a user interface 3450. The computing environment 3410 may be part of a data processing server. The computing environment 3410 includes a processor 3420, a memory 3430, and an input / output (I / O) interface 3440.

[0370] The processor 3420 typically controls the overall operation of the computing environment 3410, such as operations associated with display, data acquisition, data communication, and image processing. The processor 3420 may include one or more processors to execute instructions for performing all or some of the steps in the methods described above. Additionally, the processor 3420 may include one or more modules that facilitate interaction between the processor 3420 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single chip machine, a graphical processing unit (GPU), etc.

[0371] The memory 3430 is configured to store various types of data to support the operation of the computing environment 3410. The memory 3430 may include predefined software 3432. Examples of such data include instructions for any application or method operated on the computing environment 3410, video data sets, image data, etc. The memory 3430 may be implemented using any type of volatile or non-volatile memory device, or combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disks.

[0372] The I / O interface 3440 provides an interface between the processor 3320 and a peripheral interface module, such as a keyboard, a click wheel, and buttons. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 3440 may be coupled to an encoder and a decoder.

[0373] In one embodiment, a non-transitory computer-readable storage medium is also provided, comprising, for example, in memory 3430, a plurality of programs executable by processor 3420 in computing environment 3410 to perform the aforementioned methods. Alternatively, the non-transitory computer-readable storage medium may store a bitstream or datastream including encoded video information (e.g., video information including one or more syntax elements), the bitstream or datastream generated by an encoder (e.g., video encoder 20 of FIG. 2) using, for example, the encoding method described above for use by a decoder (e.g., video decoder 30 of FIG. 3) in decoding the video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0374] In one embodiment, a computing device is also provided that includes one or more processors (e.g., processor 3420) and a non-transitory computer-readable storage medium or memory 3430 having stored thereon a plurality of programs executable by the one or more processors, the one or more processors being configured to, upon execution of the plurality of programs, perform the aforementioned method.

[0375] In one embodiment, a computer program product is also provided that includes a plurality of programs executable by a processor 3420 in the computing environment 3410 to perform the methods described above, e.g., in memory 3430. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0376] In one embodiment, the computing environment 3410 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0377] Further embodiments also include various subsets of the above embodiments combined or otherwise rearranged in various other embodiments.

[0378] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communications protocol. In this manner, computer-readable media may generally correspond to (1) tangible computer-readable storage media that is non-transitory, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the implementations described herein. A computer program product may include computer-readable media.

[0379] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the present disclosure. Many modifications, variations, and alternative implementations will be apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.

[0380] Unless otherwise specifically stated, the order of steps of the method according to the present disclosure is intended for illustrative purposes only, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but may be changed according to practical conditions. In addition, at least one of the steps of the method according to the present disclosure may be adjusted, combined, or deleted according to practical requirements.

[0381] The examples have been chosen and described to explain the principles of the disclosure and to enable those skilled in the art to understand the disclosure with respect to various implementations and to best utilize the underlying principles and various implementations, along with various modifications suited to the particular applications contemplated. Accordingly, it should be understood that the scope of the disclosure is not limited to the particular implementations disclosed, but that modifications and other implementations are also intended to be included within the scope of the disclosure.

Claims

1. 1. A method for decoding a video signal, comprising: receiving a picture frame from the video signal, the picture frame including a first component and a second component; determining a classifier for each sample of the second component using a set of weighted sample values ​​obtained from a first set of samples of the first component associated with the respective sample of the second component and a second set of samples of the second component associated with the respective sample of the second component, wherein the first set of samples of the first component includes an array sample of the first component for the respective sample of the second component and an adjacent sample of the array sample of the first component, and the second set of samples of the second component includes a current sample of the second component for the respective sample of the second component and an adjacent sample of the current sample of the second component; determining a sample offset for each sample of the second component according to the classifier; and modifying the respective samples of the second component based on the determined sample offset.

2. determining the classifier for the respective sample of the second component; The method of claim 1 , comprising determining the classifier for the respective sample of the second component using a summation of the set of weighted sample values.

3. the first component is a component selected from the group consisting of a luma component, a first chroma component, and a second chroma component; the second component is a component selected from the group consisting of the luma component, the first chroma component, and the second chroma component; the first component is different from the second component; The method of claim 1.

4. the picture frame further includes a third component, and determining the classifier for the respective samples of the second component comprises:

2. The method of claim 1, further comprising: determining the classifier for the respective samples of the second component using the set of weighted sample values ​​from the first set of samples and the second set of samples, and an additional set of weighted sample values ​​obtained from a third set of samples of the third component associated with the respective samples of the second component, wherein the third set of samples of the third component includes an array sample of the third component for the respective sample of the second component and adjacent samples of the array sample of the third component.

5. determining the classifier for the respective sample of the second component; The method of claim 4 , comprising determining the classifier for the respective sample of the second component using a sum of the set of weighted sample values ​​and the additional set of weighted sample values.

6. Determining the classifier for each sample of the second component using the summation may include summing a classified sample value S to [Equation 1] The purpose is to judge as R ij is the jth sample of the i-th component of the sequence or the current and adjacent samples for said respective sample of said second component, where i is equal to 1, 2, and 3 representing each of said first, second, and third components, respectively, and j is equal to 0 to N-1, where i and j are both integers; ij However, the R ij and N is the total number of the sequence or current sample and the adjacent samples of the i component for the respective sample of the second component; determining the classifier for each sample of the second component based on the classification sample value S. The method of claim 5.

7. Addition [Equation 2] or the classification sample value S is within the dynamic bit depth range of the video signal; The method of claim 6.

8. the classifier for the respective samples of the second component includes a first classifier using the summation of the set of weighted sample values ​​jointly combined with a second classifier determined from edge directions and intensities of the sequence samples of the first component for the respective samples of the second component, the first classifier being different from the second classifier; The method of claim 2.

9. one of the weighting coefficients for the current sample of the second component is derived from the other weighting coefficients, or The classification sample value S is used to determine a class index for the classifier as classIdx = (S * bandNumS) >> BitDepth, where bandNumS is the band number associated with the sample value of S and BitDepth is the dynamic bit depth of the video signal. The method of claim 6.

10. The method of claim 1, wherein the bit depth of the video signal is equal to 12.

11. 1. A method for decoding a video signal, comprising: receiving a picture frame from the video signal, the picture frame including a first component, a second component, and a third component; determining a classifier for each sample of the second component using a set of weighted sample values ​​obtained from a first set of samples of the first component associated with the respective sample of the second component and a third set of samples of the third component associated with the respective sample of the second component, wherein the first set of samples of the first component includes a sequence sample of the first component for the respective sample of the second component and adjacent samples of the sequence sample of the first component, and the third set of samples of the third component includes a sequence sample of the third component for the respective sample of the second component and adjacent samples of the sequence sample of the third component; determining a sample offset for each sample of the second component according to the classifier; and modifying the respective samples of the second component based on the determined sample offset.

12. determining the classifier for the respective sample of the second component; The method of claim 11 , comprising determining the classifier for the respective sample of the second component using a summation of the set of weighted sample values.

13. the first component is a component selected from the group consisting of a luma component, a first chroma component, and a second chroma component; the second component is a component selected from the group consisting of the luma component, the first chroma component, and the second chroma component; the third component is a component selected from the group consisting of the luma component, the first chroma component, and the second chroma component; the first component, the second component, and the third component are different components; The method of claim 11.

14. determining the classifier for the respective sample of the second component; 12. The method of claim 11, further comprising: determining the classifier for the respective sample of the second component using the set of weighted sample values ​​from the first set of samples and the third set of samples, and an additional set of weighted sample values ​​obtained from a second set of samples of the second component associated with the respective sample of the second component, wherein the second set of samples of the second component includes a current sample of the second component for the respective sample of the second component and neighboring samples of the current sample of the second component.

15. determining the classifier for the respective sample of the second component; The method of claim 14 , comprising determining the classifier for the each sample of the second component using a sum of the set of weighted sample values ​​and the additional set of weighted sample values.

16. determining the classifier for each sample of the second component using the summation; The classification sample value S is [Equation 3] R ij is the jth sample of the i-th component of the sequence or the current and adjacent samples for said respective sample of said second component, where i is equal to 1, 2, and 3 representing each of said first, second, and third components, respectively, and j is equal to 0 to N-1, where i and j are both integers; ij However, the R ij and N is the total number of the sequence or current sample and the adjacent samples of the i component for the respective sample of the second component; and determining the classifier for the respective sample of the second component based on the classification sample value S.

17. Addition [Equation 4] or The method of claim 16, wherein the classification sample value S is within a dynamic bit depth range of the video signal.

18. 13. The method of claim 12, wherein the classifier for the respective sample of the second component comprises a first classifier using the summation of the set of weighted sample values ​​jointly combined with a second classifier determined from edge directions and intensities of the array samples of the first component for the respective sample of the second component, and the first classifier is different from the second classifier.

19. one of the weighting coefficients for the current sample of the second component is derived from the other weighting coefficients, or The classification sample value S is used to determine a class index for the classifier as classIdx = (S * bandNumS) >> BitDepth, where bandNumS is the band number associated with the sample value of S and BitDepth is the dynamic bit depth of the video signal.

17. The method of claim 16.

20. The method of claim 11, wherein the bit depth of the video signal is equal to 12.

21. one or more processing units; a memory coupled to the one or more processing units; a plurality of programs stored in said memory, said plurality of programs, when executed by said one or more processing units, causing said electronic device to perform the method of any one of claims 1 to 20. electronic equipment.

22. 1. A method for storing a bitstream, comprising: performing the following encoding method steps to generate a bitstream: obtaining a picture frame from the video signal, the picture frame including a first component and a second component; determining a classifier for each sample of the second component using a set of weighted sample values ​​obtained from a first set of samples of the first component associated with the respective sample of the second component and a second set of samples of the second component associated with the respective sample of the second component, wherein the first set of samples of the first component includes an array sample of the first component for the respective sample of the second component and an adjacent sample of the array sample of the first component, and the second set of samples of the second component includes a current sample of the second component for the respective sample of the second component and an adjacent sample of the current sample of the second component; determining a sample offset for each sample of the second component according to the classifier; modifying the respective samples of the second component based on the determined sample offsets; and storing the bitstream; The method, wherein the bitstream is decoded by the video decoding method of claim 1 .

23. The method of claim 22, wherein the bit depth of the video signal is equal to 12.

24. A method for storing a bitstream, comprising: performing the following encoding method steps to generate a bitstream: obtaining a picture frame from the video signal, the picture frame including a first component, a second component, and a third component; determining a classifier for each sample of the second component using a set of weighted sample values ​​obtained from a first set of samples of the first component associated with the respective sample of the second component and a third set of samples of the third component associated with the respective sample of the second component, wherein the first set of samples of the first component includes a sequence sample of the first component for the respective sample of the second component and adjacent samples of the sequence sample of the first component, and the third set of samples of the third component includes a sequence sample of the third component for the respective sample of the second component and adjacent samples of the sequence sample of the third component; determining a sample offset for each sample of the second component according to the classifier; modifying the respective samples of the second component based on the determined sample offsets; and storing the bitstream; The method, wherein the bitstream is decoded by the video decoding method of claim 11.

25. The method of claim 24, wherein the bit depth of the video signal is equal to 12.

26. A computer program product for causing a computer to carry out the method according to any one of claims 1 to 20.

27. 1. A method for transmitting a bitstream, comprising: performing the following encoding method steps to generate a bitstream: obtaining a picture frame from the video signal, the picture frame including a first component and a second component; determining a classifier for each sample of the second component using a set of weighted sample values ​​obtained from a first set of samples of the first component associated with the respective sample of the second component and a second set of samples of the second component associated with the respective sample of the second component, wherein the first set of samples of the first component includes an array sample of the first component for the respective sample of the second component and an adjacent sample of the array sample of the first component, and the second set of samples of the second component includes a current sample of the second component for the respective sample of the second component and an adjacent sample of the current sample of the second component; determining a sample offset for each sample of the second component according to the classifier; modifying the respective samples of the second component based on the determined sample offsets; and transmitting the bitstream; The method, wherein the bitstream is decoded by the video decoding method of claim 1 .

28. A method for transmitting a bitstream, comprising: performing the following encoding method steps to generate a bitstream: obtaining a picture frame from the video signal, the picture frame including a first component, a second component, and a third component; determining a classifier for each sample of the second component using a set of weighted sample values ​​obtained from a first set of samples of the first component associated with the respective sample of the second component and a third set of samples of the third component associated with the respective sample of the second component, wherein the first set of samples of the first component includes a sequence sample of the first component for the respective sample of the second component and adjacent samples of the sequence sample of the first component, and the third set of samples of the third component includes a sequence sample of the third component for the respective sample of the second component and adjacent samples of the sequence sample of the third component; determining a sample offset for each sample of the second component according to the classifier; determining a sample offset for each sample of the second component according to the classifier; modifying the respective samples of the second component based on the determined sample offsets; and transmitting the bitstream; The method, wherein the bitstream is decoded by the video decoding method of claim 11.

Citation Information

Patent Citations

  • Chroma block prediction method and apparatus

    CA3125500A1

  • Method and device for decoding of compensation offsets for set of reconstructed samples of image

    JP2019110587A

  • Method of Sample Adaptive Offset Processing for Video Coding and Inter-Layer Scalable Coding

    US20140348222A1

  • Sample adaptive offset systems and methods

    US20170230656A1

  • Methods and devices for improvement in obtaining linear component sample prediction parameters

    WO2019162117A1