Enhancement of Chroma Encoding / Decoding in Cross-Component Sampling Adaptive Offset

By employing a Cross-Component Sampling Adaptive Offset (CCSAO) method that adjusts sample offsets based on cross-component relationships, the encoding and decoding efficiency of video data is improved, addressing the challenges of high-resolution video data management.

JP7684411B2Active Publication Date: 2025-05-27BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023546286
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-04
Filing Date
2022-01-24
Publication Date
2025-05-27
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face challenges in efficiently encoding and decoding luminance and chrominance components, particularly as video resolutions increase, leading to exponential growth in video data.

Method used

The method involves improving encoding and decoding efficiency by exploring the cross-component relationship between luminance and chrominance components, specifically through a Cross-Component Sampling Adaptive Offset (CCSAO) process that adjusts sample offsets based on classifiers determined from sets of samples.

Benefits of technology

This approach enhances encoding and decoding efficiency by utilizing cross-component correlations, allowing for more effective compression and decompression of video data while maintaining image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007684411000042
    Figure 0007684411000042
  • Figure 0007684411000043
    Figure 0007684411000043
  • Figure 0007684411000044
    Figure 0007684411000044
Patent Text Reader

Abstract

An electronic device performs a method for decoding a video signal, the method comprising: receiving an image frame from a video signal, the image frame including a first component and a second component; determining a classifier for the first component based on a first set of one or more samples of the second component associated with each sample of the first component; determining a sample offset for each sample of the first component according to the classifier; and modifying a value of each sample of the first component based on the determined sample offset, the first component being a luma component and the second component being a first chroma component. In one embodiment, the classifier for the first component is further determined based on a second set of one or more samples of the first component associated with each sample of the first component.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 144,414, filed on February 1, 2021, with the title "Cross - Component Sampling Adaptive Offset" and U.S. Provisional Application No. 63 / 145,940, filed on February 4, 2021, with the title "Cross - Component Sampling Adaptive Offset", and the entire specifications of these patent applications are incorporated herein by reference.

Technical Field

[0002] This application generally relates to video encoding, decoding, and compression, and more particularly, to methods and apparatuses for improving the encoding and decoding efficiency of luminance and chrominance.

Background Art

[0003] Various electronic devices such as digital TVs, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing devices, and video streaming devices support digital video. The electronic devices receive and transmit, encode, decode, and store digital video data by implementing video compression / decompression standards. Known video encoding / decoding standards include Versatile Video Coding (VVC), jointly developed by ISO / IEC MPEG and ITU-T VCEG, High Efficiency Video Coding (also known as HEVC or H.265 or MPEG-H Part 2), and Advanced Video Coding (also known as AVC or H.264 or MPEG-4 Part 10). AOMedia Video (AV1) was developed by the Alliance for Open Media (AOM) as a successor to the previous standard VP9. Audio Video Coding Standard (AVS), which is a digital audio and digital video compression standard, is another video compression standard system formulated by the China Audio Video Coding Standard Working Group.

[0004] Video compression typically involves performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or eliminate redundancy inherent in video data. In block-based video coding, a video frame is partitioned into one or more slices, each containing a plurality of video blocks called coding tree units (CTUs). Each CTU contains one coding unit (CU) or may be recursively partitioned into smaller CUs until a syntactically defined minimum CU size is reached. Each CU (also called a leaf CU) contains one or more transform units (TUs) and one or more prediction units (PUs). Each CU can be coded in either intra, inter, or IBC mode. Video blocks within an intra-coded (I) slice in a video frame are coded using spatial prediction with respect to reference samples in adjacent blocks within the same video frame. Video blocks within an inter-coded (P or B) slice in a video frame use spatial prediction with respect to reference samples in adjacent blocks within the same video frame or temporal prediction with respect to reference samples in other previous and / or future reference video frames.

[0005] In previous symbolized reference blocks, for example, in spatial prediction or temporal prediction based on adjacent blocks, a prediction block of the current video block to be coded is obtained. The process of expressing the reference block can be realized by a block matching algorithm. Residual data indicating the pixel difference between the current block to be coded and the prediction block is called a residual block or prediction error. An inter-coded block is coded according to a motion vector pointing to a reference block in the reference frame where the prediction block is generated and the residual block. The process of determining the motion vector is usually called motion estimation. An intra-coded block is coded according to an intra prediction mode and the residual block. For further compression, the residual block is converted from a pixel domain to a transform domain, for example, a frequency domain, and as a result, residual transform coefficients to be quantized later are obtained. Then, the transform coefficients initially arranged in a two-dimensional matrix and quantized are scanned to generate a one-dimensional vector of the transform coefficients, and then entropy coded into a video bit stream to achieve further compression.

[0006] Then, the coded video bit stream is stored in a computer-readable storage medium (e.g., flash memory), accessed by another electronic device with digital video capabilities, or directly transmitted to this electronic device via wire or wirelessly. And this electronic device, for example, analyzes the coded video bit stream to obtain syntax elements from this bit stream, and based on at least a part of the syntax elements obtained from this bit stream, reconstructs digital video data from this coded video stream into the original format, thereby performing video decompression (a process opposite to the above-mentioned video compression), and reproduces this reconstructed digital video data on the display of this electronic device.

[0007] As the quality of digital video progresses from high definition to 4K×2K and even 8K×4K, the amount of video data to be encoded / decoded increases exponentially. It has always been a challenge to encode / decrypt video data more efficiently while maintaining the image quality of the decoded video data.

Summary of the Invention

[0008] This application describes an implementation related to a method and apparatus for improving the encoding / decoding efficiency of a luminance component and a chrominance component, including encoding and decoding video data, and more specifically, improving the encoding / decoding efficiency by exploring the cross-component relationship between the luminance component and the chrominance component.

[0009] According to a first aspect of the present application, a method for decoding a video signal includes receiving, from the video signal, an image frame including a first component and a second component; determining, based on a first set of one or more samples of the second component associated with each sample of the first component, a classifier for the first component; determining, according to the classifier, a sample offset for each sample of the first component; and changing the value of each sample of the first component based on the determined sample offset, where the first component is a luminance component and the second component is a first chrominance component.

[0010] In some embodiments, the classifier for the first component is further determined based on a second set of one or more samples of the first component associated with each sample of the first component.

[0011] In one embodiment, the image frame further includes a third component, and the classifier for the first component is further determined based on a third set of one or more samples of the third component associated with each sample of the first component, and the third component is a second chroma component.

[0012] According to a second aspect of the present application, an electronic device includes one or more processing units, a memory connected to the one or more processing units, and a plurality of programs stored in the memory. When the plurality of programs are executed by the one or more processing units, the electronic device is caused to execute the method described above.

[0013] According to a third aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs to be executed by an electronic device having one or more processing units. When the plurality of programs are executed by the one or more processing units, the electronic device is caused to execute the method described above.

Brief Description of the Drawings

[0014] The accompanying drawings, which are incorporated herein as part of this specification, provide a further understanding of the implementation of the present invention, show the above-described implementation, and together with the description, serve to explain the basic principles. The same reference numerals denote the same or corresponding parts.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 6D

Figure 6E

Figure 6F

Figure 6G

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In the following detailed description, a plurality of non-limiting specific details are described in order to easily understand the gist described in this specification. However, it is obvious to those skilled in the art that the present invention can be implemented by various modifications without departing from the scope of the claims and their gist. For example, it is obvious to those skilled in the art that the gist described in this specification can be implemented in many types of electronic devices having digital video functions.

[0016] The first-generation AVS standard includes the Chinese national standards "Information Technology, Advanced Audio and Video Coding and Decoding, Part 2: Video" (referred to as AVS1) and "Information Technology, Advanced Audio and Video Coding and Decoding, Part 16: TV Video Broadcasting" (referred to as AVS+). Compared with the MPEG-2 standard, it can save about 50% of the bitrate with the same visual image quality. The second-generation AVS standard mainly includes the Chinese national standard "Information Technology, High-Efficiency Multimedia Coding and Decoding" (referred to as AVS2) series for the transmission of ultra-high HD TV programs. The coding and decoding efficiency of AVS2 is twice that of AVS+. At the same time, the video part of the AVS2 standard has been submitted as an international application standard by the Institute of Electrical and Electronics Engineers (IEEE). The AVS3 standard is a next-generation video coding and decoding standard for UHD video applications that aims to reduce the bitrate by about 30% compared to the HEVC standard and exceed the coding and decoding efficiency of the latest international standard HEVC. At the 68th AVS meeting in March 2019, the AVS3-P2 baseline that achieves a bitrate reduction of about 30% compared to the HEVC standard was completed. Currently, the AVS group maintains reference software called the High Performance Model (HPM) and demonstrates the reference implementation of the AVS3 standard. The AVS3 standard, like HEVC, is established on a block-based hybrid video coding and decoding framework.

[0017] FIG. 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to an embodiment of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data to be decoded in the future by a destination device 14. The source device 12 and the destination device 14 may include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some embodiments, the source device 12 and the destination device 14 have wireless communication capabilities.

[0018] In some embodiments, the destination device 14 receives the encoded video data to be decoded via a link 16. The link 16 can include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may include a communication medium that can directly transmit the encoded video data from the source device 12 to the destination device 14 in real time. The encoded video data is modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium can include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may be configured as part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, and any other device useful for communication from the source device 12 to the destination device 14.

[0019] In some other embodiments, the encoded video data is transmitted from the output interface 22 to the storage device 32. Thereafter, the encoded video data in the storage device 32 is accessed by the target device 14 via the input interface 28. The storage device 32 can include any of a variety of distributed or locally accessible data storage media such as a hard drive, a Blu-ray disk, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or other suitable digital storage media for storing the encoded video data. In another example, the storage device 32 may correspond to a file server or another intermediate storage device capable of holding the encoded video data generated by the source device 12. The target device 14 can access the video data stored in the storage device 32 by streaming or downloading. The file server can be any type of computer capable of storing the encoded video data and transmitting this encoded video data to the target device 14. Exemplary file servers include a web server (e.g., for a website), an FTP server, a network-attached storage (NAS) device, or a local disk drive. The target device 14 can access the encoded video data via any standard data connection including a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination thereof suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 32 can be a streaming transmission, a download transmission, or a combination thereof.

[0020] As shown in FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 can include sources such as a video capture device (e.g., a video camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or combinations thereof. As an example, when the video source 18 is a video camera of a security monitoring system, the source device 12 and the target device 14 can constitute a camera-equipped mobile phone or a video phone. However, the embodiments described in the present application are generally applicable to video encoding and are applicable to wireless and / or wired applications.

[0021] The video encoder 20 can encode the video to be captured, previously captured video, or video generated by a computer. The encoded video data can be directly transmitted to the target device 14 via the output interface 22 of the source device 12. In addition (or alternatively), the encoded video data may be stored in the storage device 32 so that it can then be accessed by the target device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.

[0022] The target device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 includes a receiver and / or a modem and receives encoded video data via link 16. The encoded video data communicated via link 16 or provided to storage device 32 may include many syntax elements generated by video encoder 20 and used for decoding the video data by video decoder 30. These encoded video data may include such syntax elements regardless of whether they are transmitted over a communication medium, stored in a storage medium, or stored in a file server.

[0023] In some embodiments, the target device 14 may include a display device 34 that is an integrated display device or an external display device configured to communicate with the target device 14. The display device 34 displays the decoded video data to the user and may include any of various display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0024] The video encoder 20 and the video decoder 30 operate according to specialized or industry standards such as VVC, HEVC, MPEG-4, Part10, Advanced Video Coding (AVC), AVS, or extensions of such standards. It should be understood that the present application is not limited to specific video coding / decoding standards and is also applicable to other video coding / decoding standards. The video encoder 20 of the source device 12 is configured to encode video data according to any of these current or future standards. Similarly, the video decoder 30 of the target device 14 is configured to decode video data according to any of these current or future standards.

[0025] The video encoder 20 and the video decoder 30 can each be implemented by any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented in part by software, the electronic device may store software instructions in a suitable non-transitory computer-readable medium and execute the instructions in hardware by one or more processors to perform the video encoding / decoding operations described in this disclosure. The video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders integrated as part of a combined encoder / decoder (CODEC) in each device.

[0026] FIG. 2 is a block diagram illustrating a video encoder 20 according to an embodiment described in the present application. The video encoder 20 can perform intra prediction encoding and inter prediction encoding on video blocks within a video frame. Intra prediction encoding relies on spatial prediction and reduces or eliminates spatial redundancy of video data within a given video frame or image. Inter prediction encoding relies on temporal prediction and reduces or eliminates temporal redundancy of video data within adjacent video frames or images of a video sequence.

[0027] As shown in FIG. 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a conversion processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra prediction processing unit 46, and an intra block copy (BC) unit 48. In one embodiment, the video encoder 20 also further includes an inverse quantization unit 58, an inverse conversion processing unit 60, and an adder 62 for video block reconstruction. Between the adder 62 and the DPB 64, it is possible to arrange, for example, a deblocking filter, an in-loop filter 63 that filters the boundaries between blocks from the reconstructed video to remove block artifacts. Also, in addition to the deblocking filter, another in-loop filter 63 may be used to filter the output of the adder 62. Before the reconstructed CU is placed in the reference image memory and used as a reference for encoding / decoding future video blocks, for example, sampling adaptive offset (SAO), adaptive in-loop filter (ALF), etc. of i The in-loop filter 63 may be applied to the reconstructed CU. Further The video encoder 20 may be formed in the form of a fixed or programmable hardware unit, or may be partitioned within one or more of the illustrated fixed or programmable hardware units.

[0028] The video data memory 40 stores the video data to be encoded by components in the video encoder 20. The video data in the video data memory 40 is obtained, for example, from the video source 18. The DPB 64 is a buffer that stores reference video data used when the video encoder 20 encodes video data (e.g., in an intra prediction or inter prediction coding mode). The video data memory 40 and the DPB 64 can be formed of any of various memory devices. In various examples, the video data memory 40 may be on-chip with other components in the video encoder 20, or may be off-chip with respect to those components.

[0029] As shown in FIG. 2, after receiving the video data, the partitioning unit 45 in the prediction processing unit 41 partitions this video data into video blocks. This partitioning may include partitioning the video frame into slices, tiles, or other larger coding units (CUs) according to a predetermined partitioning structure such as a quad-tree structure for this video data. The video frame can be partitioned into a plurality of video blocks (or a set of video blocks called tiles). The prediction processing unit 41 selects one of a plurality of possible prediction coding modes so as to select one of one or more intra prediction coding modes or one of one or more inter prediction coding modes for the current video block based on error results (e.g., coding rate and distortion level). Then, the prediction processing unit 41 provides the obtained intra or inter prediction coded block to the adder 50 to generate a residual block, and reconstructs the coded block to be used as a part of a subsequent reference frame. Further, the prediction processing unit 41 provides syntax elements such as motion vectors, intra mode indicators, partitioning information, and other syntax information to the entropy coding unit 56.

[0030] The intra prediction processing unit 46 in the prediction processing unit 41 can perform spatial prediction by performing intra prediction encoding of the current video block in relation to one or more adjacent blocks within the same frame as the current block to be encoded, in order to select an intra prediction encoding mode suitable for the current video block. The motion estimation unit 42 and the motion compensation unit 44 in the prediction processing unit 41 perform temporal prediction by performing inter prediction encoding of the current video block in relation to one or more prediction blocks within one or more reference frames. The video encoder 20 may perform encoding processing for a plurality of paths to select, for example, an appropriate encoding mode for each block in the video data.

[0031] In one embodiment, the motion estimation unit 42 determines the inter prediction mode by generating a motion vector indicating the displacement of the prediction unit (PU) of the video block within the current video frame with respect to the prediction block within the reference video frame for the current video frame according to a predetermined pattern of the sequence of video frames. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector for estimating the motion of the video block. The motion vector can indicate, for example, the displacement of the PU of the video block within the current video frame (or other encoding unit) with respect to the prediction block within the reference frame (or other encoding unit) for the currently encoded video block within the current video frame or image. The predetermined pattern of the sequence can specify the video frames in this sequence as P frames or B frames. The intra BC unit 48 may determine a vector for intra BC encoding, such as a block vector, in a manner similar to the determination of the motion vector for inter prediction by the motion estimation unit 42, or may use the motion estimation unit 42 to determine this block vector.

[0032] The prediction block is a block in the reference frame that exactly matches the PU of the video block to be coded with respect to the pixel difference that can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In one embodiment, the video encoder 20 can calculate the values of the sub-pixel positions of the reference frames stored in the DPB64. For example, the video encoder 20 may interpolate the values of the 1 / 4 pixel position, 1 / 8 pixel position, or other fractional pixel positions of the reference frame. Accordingly, the motion estimation device 42 can execute a motion search process for all pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.

[0033] The motion estimation unit 42 calculates a motion vector for this PU by comparing the position of the PU of the video block in the inter-predicted coded frame with the position of the prediction block of the reference frame selected from the first reference frame list (List0) or the second reference frame list (List1) that identifies one or more reference frames stored in the DPB64, respectively. The motion estimation unit 42 transmits the calculated motion vector to the motion compensation unit 44 and also transmits it to the entropy coding unit 56.

[0034] The motion compensation executed by the motion compensation unit 44 may include obtaining or generating a prediction block based on the motion vector determined by the motion estimation unit 42. When the motion compensation unit 44 receives a motion vector for the PU of the current video block, it locates the prediction block pointed to by this motion vector in one of the reference frame lists, searches for this prediction block from the DPB 64, and transfers this prediction block to the adder 50. Then, the adder 50 forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by the motion compensation unit 44 from the pixel values of the encoded current video block. The pixel difference values forming the residual video block may include a luminance difference component or a chrominance difference component, or both. Also, the motion compensation unit 44 can further generate syntax elements related to the video blocks of the video frame, and these syntax elements are used when the video decoder 30 decodes the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vector for identifying this prediction block, any flag indicating the prediction mode, or any other syntax information described herein. Note that although the motion estimation unit 42 and the motion compensation unit 44 are shown separately for conceptual purposes, they may be highly integrated.

[0035] In one embodiment, the intra BC unit 48 can generate vectors and obtain prediction blocks in the same manner as the methods described above for the motion estimation unit 42 and the motion compensation unit 44. Here, the prediction block is in the same frame as the currently encoded block, and the vector is called a block vector rather than a motion vector. In particular, the intra BC unit 48 can determine the intra prediction mode used to encode the current block. In one example, the intra BC unit 48 can encode the current block using various intra prediction modes, for example, in the encoding of individual paths, and test their performance by rate-distortion analysis. Next, the intra BC unit 48 selects and uses one appropriate intra prediction from the various tested intra prediction modes to generate the corresponding intra mode indicator. For example, the intra BC unit 48 can calculate the rate-distortion values of the various tested intra prediction modes by rate-distortion analysis, and select and use the intra prediction mode with the optimal rate-distortion characteristics from the tested modes as the appropriate intra prediction mode. In rate-distortion analysis, generally, the amount of distortion (or error) between the encoded block and the original unencoded block from which the encoded block was generated, and the bit rate (i.e., the number of bits) used to generate this encoded block are determined. The intra BC unit 48 can calculate the ratio from the distortion and rate for various blocks to be encoded to determine which intra prediction mode shows the optimal rate-distortion value for this block.

[0036] In another example, the intra BC unit 48 may use all or part of the motion estimation unit 42 and the motion compensation unit 44 to perform the functions related to intra BC prediction according to the embodiments described herein. In any case, for intra-block copy, the prediction block is considered to exactly match the block to be encoded in terms of the pixel difference that can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics, and the identification of the prediction block may include the calculation of the values at sub-integer pixel positions.

[0037] Video encoder 20 can generate a residual video block by subtracting the pixel values of a prediction block from the pixel values of the currently encoded video block to generate a pixel difference value, regardless of whether the prediction block is from the same frame based on intra prediction or from different frames based on inter prediction. The pixel difference values forming the residual video block may include both a luminance component difference and a chrominance component difference.

[0038] The intra prediction processing unit 46 can perform intra prediction on the current video block instead of the inter prediction executed by the motion estimation unit 42 and the motion compensation unit 44 described above, or the intra-block copy prediction executed by the intra BC unit 48. In particular, the intra prediction processing unit 46 can determine one intra prediction mode and encode the current block. To achieve this, the intra prediction processing unit 46, for example, encodes the current block using various intra prediction modes in the encoding process of an individual path, and the intra prediction processing unit 46 (or in some examples, the mode selection unit) may select and use one appropriate intra prediction mode from the tested intra prediction modes. The intra prediction processing unit 46 may provide information indicating the intra prediction mode selected for this block to the entropy encoding unit 56. The entropy encoding unit 56 can encode the information indicating the selected intra prediction mode into a bitstream.

[0039] After the prediction processing unit 41 determines the prediction block of the current video block by inter prediction or intra prediction, the adder 50 generates a residual video block by subtracting this prediction block from the current video block. The residual video data in the residual block is included in one or more transform units (TUs) and provided to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients by, for example, discrete cosine transform (DCT) or a conceptually similar transform.

[0040] The conversion processing unit 52 transmits the obtained conversion coefficients to the quantification unit 54. The quantification unit 54 quantifies these conversion coefficients to further reduce the bit rate. The quantification process can reduce the bit depth associated with some or all of these coefficients. The degree of quantification can be changed by adjusting the quantization parameter. And in one example, the quantification unit 54 can perform a scan on the matrix including the quantified conversion coefficients. This scan may be performed by the entropy encoding unit 56.

[0041] Following quantification, the entropy encoding unit 56 entropy-encodes the quantified conversion coefficients into a video bitstream by, for example, context-adaptive variable length coding / decoding (CAVLC), context-adaptive binary arithmetic coding / decoding (CABAC), syntax-based context-adaptive binary arithmetic coding / decoding (SBAC), probability interval partitioning entropy (PIPE) coding / decoding, or another entropy coding method or technique. And the encoded bitstream may be transmitted to the video decoder 30, or may be transmitted to the video decoder 30 later, or may be archived in the storage device 32 for retrieval by the video decoder 30. Also, the entropy encoding unit 56 may entropy-encode the motion vectors and other syntax elements for the current video frame being encoded.

[0042] The inverse quantization unit 58 and the inverse conversion processing unit 60 respectively reconstruct the residual video block within the pixel region for generating a reference block used for prediction of other video blocks by inverse quantization and inverse conversion. As described above, the motion compensation unit 44 can generate a motion compensation prediction block from one or more reference blocks of the frames stored in the DPB 64. Also, the motion compensation unit 44 may apply one or more interpolation filters to this prediction block to calculate sub-pixel values used for motion estimation.

[0043] The adder 62 adds the reconstructed residual block to the motion compensation prediction block generated by the motion compensation unit 44 to generate a reference block that is stored in the DPB 64. Then, this reference block can be used as a prediction block by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 to inter-predict another video block in a subsequent video frame.

[0044] FIG. 3 is a block diagram showing an exemplary video decoder 30 according to an embodiment of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction processing unit 84, and an intra BC unit 85. The video decoder 30 can execute a decoding process that is approximately the reverse of the encoding process described above with respect to the video encoder 20 with reference to FIG. 2. For example, the motion compensation unit 82 can generate prediction data based on the motion vector received from the entropy decoding unit 80, and the intra prediction unit 84 can generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.

[0045] In one example, one component in the video decoder 30 may be responsible for the task of implementing the present application. Also, in one example, the implementation of the present disclosure may be partitioned among one or more components in the video decoder 30. For example, the intra BC unit 85 may implement the present application alone, or in combination with other components in the video decoder 30 such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80. In one example, the video decoder 30 may not include the intra BC unit 85, and the function of the intra BC unit 85 may be implemented by other components in the prediction processing unit 81 such as the motion compensation unit 82.

[0046] The video data memory 79 can store video data such as an encoded video bit stream decoded by other components in the video decoder 30. The video data stored in the video data memory 79 is obtained from a local video source such as a storage device 32 or a camera, for example, by accessing a wired or wireless network communication of the video data or a physical data storage medium (e.g., a flash drive or a hard disk). The video data memory 79 may include an encoded picture buffer (CPB) that stores encoded video data from the encoded video bit stream. The decoded picture buffer (DPB) 92 in the video decoder 30 stores reference video data used for decoding video data by the video decoder 30 (e.g., in an intra prediction or inter prediction coding / decoding mode). The video data memory 79 and the DPB 92 can be formed by any of various memory devices such as synchronous DRAM (SDRAM), magnetic resistance RAM (MRAM), resistive RAM (RRAM), dynamic random access memory (DRAM), or other types of memory devices. For convenience of explanation, the video data memory 79 and the DPB 92 are shown as two separate components in the video decoder 30 in FIG. 3. However, it is obvious to those skilled in the art that the video data memory 79 and the DPB 92 may be provided by the same memory device or individual memory devices. In one example, the video data memory 79 may be on-chip with other components in the video decoder 30 or off-chip with respect to those components.

[0047] Video decoder 30 receives an encoded video bitstream indicating video blocks of an encoded video frame and associated syntax elements in a decoding process. The video decoder 30 may receive syntax elements at a video frame level and / or a video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes this bitstream to generate quantized coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Then, the entropy decoding unit 80 transfers the motion vectors and other syntax elements to the prediction processing unit 81.

[0048] When the video frame is encoded as an intra prediction encoded (I) frame or is used for an intra encoded prediction block in another type of frame, the intra prediction processing unit 84 in the prediction processing unit 81 can generate prediction data for the video blocks of the current video frame based on the signaled intra prediction mode and reference data from previously decoded blocks of the current frame.

[0049] When the video frame is encoded as an inter prediction encoded (i.e., B or P) frame, the motion compensation unit 82 in the prediction processing unit 81 can generate one or more prediction blocks for the video blocks of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each prediction block is generated from a reference frame within one of the reference frame lists. The video decoder 30 can configure these reference frame lists, List0 and List1, by default configuration techniques based on the reference frames stored in the DPB 92.

[0050] In one example, when a video block is encoded according to the intra BC mode described herein, the intra BC unit 85 in the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. This prediction block may be in the reconstruction area of the same image as the current video block determined by the video encoder 20.

[0051] The motion compensation unit 82 and / or the intra BC unit 85 analyzes the motion vector and other syntax elements to determine prediction information for the video blocks of the current video frame, and generates a prediction block for the current video block being decoded using this prediction information. For example, the motion compensation unit 82 uses a part of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) for encoding the video blocks of this video frame, an inter prediction frame type (e.g., B or P), structural information of one or more reference frame lists for this frame, the motion vectors of each inter prediction encoded video block of this frame, the inter prediction state of each inter prediction encoded video block of this frame, and other information for decoding the video blocks in the current video frame.

[0052] Similarly, the intra BC unit 85 can use a part of the received syntax elements, such as a flag, to determine that the current video block is predicted in the intra BC mode, structural information regarding which video blocks in this frame are in the reconstruction area and should be stored in the DPB 92, the block vectors of each intra BC predicted video block in this frame, the intra BC prediction state of each intra BC predicted video block in this frame, and other information for decoding the video blocks in the current video frame.

[0053] Further, the motion compensation unit 82 can also perform interpolation using the interpolation filter used by the video encoder 20 in encoding the video block, and calculate the interpolation value of the sub-pixel of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the received syntax elements, and generate a prediction block using this interpolation filter.

[0054] The inverse quantization unit 86 uses the same quantization parameter calculated for each video block in this video frame to determine the degree of quantization by the video encoder 20, and inverse quantizes the quantization conversion coefficients provided to the bitstream and entropy decoded by the entropy decoding unit 80. The inverse transform processing unit 88 applies an inverse transform, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to these transform coefficients so as to reconstruct the residual block in the pixel domain.

[0055] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 adds the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85 to reconstruct the decoded video block for the current video block. An in-loop filter 91 can be arranged between the adder 90 and the DPB 92 to further process this decoded video block. Before the reconstructed CU is put into the reference picture memory, the in-loop filter 91, such as a deblocking filter, a sampling adaptive offset (SAO), or an adaptive in-loop filter (ALF), may be applied to the reconstructed CU. Then, these decoded video blocks within a predetermined frame are stored in the DPB 92 that stores the reference frame used for future motion compensation of the next video block. Also, the decoded video can be stored in the DPB 92, or in a memory device separate from the DPB 92, so as to be subsequently displayed on a display device such as the display device 34 in FIG. 1.

[0056] In a typical video encoding / decoding process, one video sequence usually includes a set of ordered frames or images. Each frame can include three sample matrices denoted as SL, SCb, and SCr. SL is a two-dimensional matrix of luminance samples. SCb is a two-dimensional matrix of Cb chrominance samples. SCr is a two-dimensional matrix of Cr chrominance samples. In another example, the frame may be monochrome, in which case only one two-dimensional matrix of luminance samples is included.

[0057] Similar to HEVC, the AVS3 standard is built on a block-based hybrid video encoding / decoding framework. The input video signal is processed in block units (referred to as coding / decoding units (CUs)). Different from HEVC that partitions blocks based only on quadtree, in AVS3, one coding tree unit (CTU) is divided into CUs based on quadtree / binary tree / extended quadtree to accommodate varying local characteristics. Also, in AVS3, the concept of multi-partition unit types in HEVC is removed, that is, there is no distinction between CUs, prediction units (PUs), and transform units (TUs). Instead, each CU is always used as the basic unit for prediction and transform without further division. In the tree partitioning structure of AVS3, first, one CTU is divided based on the quadtree structure. Next, each leaf node of the quadtree may be further partitioned based on the binary tree and extended quadtree structures.

[0058] As shown in FIG. 4A, the video encoder 20 (or, more specifically, the partitioning unit 45) first generates an encoded representation of a frame by partitioning the frame into a set of coding tree units. A video frame contains an integer number of CTUs that are sequentially ordered from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height are notified in the sequence parameter set by the video encoder 20 such that all CTUs in the video sequence have the same size, which is one of 128×128, 64×64, 32×32, and 16×16. Note that the present application is not necessarily limited to a specific size. As shown in FIG. 4B, each CTU may include one coding tree block (CTB) of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntax elements used to encode the samples of the coding tree blocks. The syntax elements describe the attributes of different types of units of the coded block of pixels and how the video sequence is reconstructed at the video decoder 30, and include, for example, inter prediction or intra prediction, intra prediction mode, motion vectors, and other parameters. In a monochrome image or an image having three separate color planes, one CTU may include a single coding tree block and syntax elements used to encode the samples of this coding tree block. The coding tree block can be an N×N sample block.

[0059] To achieve better performance, the video encoder 20 can recursively perform tree partitioning, such as binary tree partitioning, quadtree partitioning, or a combination thereof, on the encoding tree block of the CTU to partition this CTU into smaller coding units (CUs). To achieve better performance, the video encoder 20 can recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof, on the encoding tree block of the CTU to partition this CTU into smaller coding units (CUs). As shown in FIG. 4C, the 64×64 CTU 400 is first partitioned into four smaller CUs of 32×32 block size. Among these four smaller CUs, CU 410 and CU 420 are each partitioned into four CUs of 16×16 block size. The two CUs 430 and 440 of 16×16 block size are each further partitioned into four CUs of 8×8 block size. FIG. 4D shows a quadtree data structure representing the final result of the partitioning process of the CTU 400 shown in FIG. 4C, and each leaf node of the quadtree corresponds to one CU of each size from 32×32 to 8×8. Similar to the CTU shown in FIG. 4B, each CU may include one coding block (CB) of luminance samples of the same size in the frame, two corresponding coding blocks of chroma samples, and syntax elements used to encode the samples of these coding blocks. For a monochrome image or an image having three separate color planes, one CU may include a single coding block and the syntax structure used to encode the samples of this coding block. Note that the quadtree partitioning shown in FIGS. 4C and 4D is only exemplary, and one CTU can be divided into CUs suitable for various local characteristics based on quaternary / ternary / binary tree partitioning. In a multi-type tree structure, one CTU can be divided according to the quadtree structure, and each quadtree leaf CU can be further divided according to the binary tree and ternary tree structures. As shown in FIG. 4E, there are five types of partitioning / types in AVS3, namely, quaternary partitioning, horizontal binary partitioning, vertical binary partitioning, horizontally extended quadtree partitioning, and vertically extended quadtree partitioning.

[0060] In one embodiment, the video encoder 20 can further partition the coded block of the CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) sample block to which the same prediction (inter prediction or intra prediction) is applied. The prediction unit (PU) of the CU can include a prediction block of one luma sample, two corresponding prediction blocks of chroma samples, and the syntax elements used to predict these prediction blocks. In a monochrome image or an image having three individual color planes, the PU can include a single prediction block and the syntax structure used to predict this prediction block. The video encoder 20 can generate a predicted luma block, a predicted Cb block, and a predicted Cr block for the luma prediction block, the Cb prediction block, and the Cr prediction block of each PU of the CU.

[0061] The video encoder 20 can generate these prediction blocks for the PU by intra prediction or inter prediction. When the video encoder 20 generates the prediction block of the PU by intra prediction, it can generate the predicted block of this PU based on the decoded samples of the frame related to this PU. When the video encoder 20 generates the predicted block of the PU by inter prediction, it can generate the predicted block of this PU based on the decoded samples of one or more frames other than the frame related to this PU.

[0062] After the video encoder 20 generates one or more predictive luminance blocks, predictive Cb blocks, and predictive Cr blocks of a CU, the video encoder 20 subtracts the predictive luminance block of the CU from the original luminance encoding block of the CU to generate a luminance residual block of the CU. Here, each sample in the luminance residual block of the CU indicates the difference between the luminance sample in one of the predictive luminance blocks of the predictive luminance block of the CU and the corresponding sample in the original luminance encoding block of the CU. Similarly, the video encoder 20 generates a Cb residual block and a Cr residual block of the CU respectively. Here, each sample in the Cb residual block of the CU indicates the difference between the Cb sample in one of the predictive Cb blocks of the predictive Cb block of the CU and the corresponding sample in the original Cb encoding block of the CU, and each sample in the Cr residual block of the CU indicates the difference between the Cr sample in one of the predictive Cr blocks of the predictive Cr block of the CU and the corresponding sample in the original Cr encoding block of the CU.

[0063] Furthermore, as shown in FIG. 4C, video encoder 20 can decompose the luminance residual block, Cb residual block, and Cr residual block of a CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks by quadtree partitioning. The transform blocks are rectangular (square or non-square) sample blocks to which the same transform is applied. The transform unit (TU) of a CU can include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements used to transform the transform block samples. Thus, each TU of a CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In one example, the luminance transform block associated with a TU can be a sub-block of the luminance residual block of the CU. The Cb transform block can be a sub-block of the Cb residual block of the CU. The Cr transform block can be a sub-block of the Cr residual block of the CU. In a monochrome image or an image having three separate color planes, a TU can include a single transform block and a syntax structure used to transform the samples of this transform block.

[0064] Video encoder 20 can apply one or more transforms to the luminance transform block of a TU to generate the luminance coefficient block of this TU. The coefficient block can be a two-dimensional matrix of transform coefficients. The transform coefficients can be scalar values. Video encoder 20 can apply one or more transforms to the Cb transform block of a TU to generate the Cb coefficient block of this TU. Video encoder 20 can apply one or more transforms to the Cr transform block of a TU to generate the Cr coefficient block of this TU.

[0065] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally means quantizing transform coefficients to reduce as much as possible the amount of data representing these transform coefficients and achieve further compression. After quantizing the coefficient block, the video encoder 20 can entropy encode the syntax elements representing the quantized transform coefficients. For example, the video encoder 20 may perform context-adaptive binary arithmetic coding / decoding (CABAC) on the syntax elements representing the quantized transform coefficients. Finally, the video encoder 20 outputs a bitstream including a bit sequence constituting the encoded frame and the representation of related data, and saves it in the storage device 32 or transmits it to the target device 14.

[0066] After receiving the bitstream generated by the video encoder 20, the video decoder 30 analyzes this bitstream to obtain syntax elements from the bitstream. Based on at least a part of the syntax elements obtained from the bitstream, the video decoder 30 can reconstruct a frame of video data. The process of reconstructing video data is generally the reverse of the encoding process executed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block related to the TU of the current CU to reconstruct the residual block related to the TU of the current CU. Also, the video decoder 30 reconstructs the encoded block of the current CU by adding the samples of the prediction block for the PU of the current CU and the corresponding samples of the transform block of the TU of the current CU. After the encoded blocks of each CU in the frame are reconstructed, the video decoder 30 can reconstruct this frame.

[0067] SAO is a process that modifies the samples decoded based on the values in the lookup table transmitted by the encoder by conditionally adding offset values to each sample after applying the deblocking filter. SAO filtering is performed on a zone basis based on the filter type selected for each CTB by the syntax element sao-type-idx. A value of 0 for sao-type-idx indicates that the SAO filter is not applied to the CTB, and values of 1 and 2 indicate that the band offset and edge offset filter types are used, respectively. In the band offset mode specified by sao-type-idx being equal to 1, the selected offset value directly depends on the sample amplitude. In this mode, the entire sample amplitude range is evenly divided into 32 segments called bands, and the sample values belonging to 4 of these bands (consecutive within the 32 bands) are represented as a band offset that is either positive or negative and changed by adding the transmitted value. The reason 4 consecutive bands are used is mainly because in smoothed regions where strip artifacts may occur, the sample amplitudes in the CTB tend to concentrate in only a few bands. Furthermore, the design choice of using 4 offsets is unified with the edge offset operation mode that also uses 4 offset values. In the edge offset mode specified by sao-type-idx being equal to 2, the syntax element sao-eo-class with values from 0 to 3 indicates whether any of the horizontal, vertical, or two diagonal gradient directions are used for edge offset classification in the CTB.

[0068] FIG. 5 is a block diagram showing four gradient patterns used in SAO according to an embodiment of the present disclosure. The four gradient patterns 502, 504, 506, 508 are used for each sao-eo-class in the edge offset mode. Samples denoted as "p" indicate the central sample to be considered. Two samples denoted as "n0" and "n1" specify two adjacent samples along the (a) horizontal (sao-eo-class = 0), (b) vertical (sao-eo-class = 1), (c) 135° diagonal (sao-eo-class = 2), and (d) 45° (sao-eo-class = 3) gradient patterns. As shown in FIG. 5, by comparing the sample value p at a certain position with the values n0 and n1 of two adjacent samples at adjacent positions, each sample in the CTB is classified into one of five EdgeIdx categories. Since each sample is thus classified based on the decoded sample value, the EdgeIdx category does not require additional signaling. Depending on the EdgeIdx category of the sample position, offset values from the transmitted look-up table are added to the sample value for EdgeIdx categories from 1 to 4. The offset values for categories 1 and 2 are always positive, and the offset values for categories 3 and 4 are always negative. Therefore, the filter usually has a smoothing effect in the edge offset mode. Table 1 below exemplifies the sample EdgeIdx categories in SAO edge classification. JPEG0007684411000001.jpg44166

[0069] In the case of SAO types 1 and 2, a total of four amplitude offset values are sent to the decoder for each CTB. In the case of type 1, the symbols are also encoded. For example, offset values and correlation syntax elements such as sao-type-idx and sao-eo-class are usually determined by the encoder using criteria that optimize the distortion rate performance. The SAO parameters can be indicated by a merge flag to inherit from the left or upper CTB to enable signaling. In short, SAO is a non-linear filtering operation that enables further refinement of the reconstructed signal and can enhance the signal representation in the smoothing region and around the edges.

[0070] In one embodiment, methods and systems are disclosed herein for improving encoding / decoding efficiency or reducing the complexity of sample adaptive offset (SAO) by introducing cross-component information. SAO is used in the HEVC, VVC, AVS2, and AVS3 standards. In the following description, the existing SAO design in the HEVC, VVC, AVS2, and AVS3 standards is used as the basic SAO method. However, for those skilled in the art of video encoding / decoding, the cross-component method described in this disclosure is also applicable to other loop filter designs or other encoding / decoding tools having similar design concepts. For example, in the AVS3 standard, SAO is replaced by an encoding / decoding tool called extended sample adaptive offset (ESAO). However, the CCSAO disclosed herein can also be applied in parallel with ESAO. In another example, CCSAO may be applied in parallel with the constraint-directed enhancement filter (CDEF) in the AV1 standard.

[0071] In existing SAO designs in the HEVC, VVC, AVS2, and AVS3 standards, the luminance Y, chrominance Cb, and chrominance Cr sample offset values are determined independently. That is, for example, the current chrominance sample offset is determined only by the current chrominance sample value and adjacent chrominance sample values, regardless of the juxtaposed or adjacent luminance samples. However, luminance samples retain more detailed information of the original image than chrominance samples, and can facilitate the determination of the offset of the current chrominance sample. Furthermore, after the color conversion from RGB to YCbCr, or after quantization and deblocking filtering, chrominance samples often lose high-frequency details. Therefore, introducing luminance samples that retain high-frequency details for chrominance offset determination can facilitate chrominance sample reconstruction. Thus, for example, further gains can be expected by exploring cross-component correlation by using methods and systems of cross-component sample adaptive offset (CCSAO). In another example of SAO, the luminance sample offset is determined only by luminance samples. However, for example, luminance samples having the same frequency band offset (BO) classification can be further classified by their juxtaposed and adjacent chrominance samples, which can result in more efficient classification. SAO classification can be used as a shortcut to compensate for the difference in samples between the original image and the reconstructed image. Therefore, effective classification is required.

[0072] FIG. 6A is a block diagram showing a CCSAO system and process applied to a chroma sample according to an embodiment of the present disclosure, taking DBF Y as an input. The luminance samples after the luminance block filter (DBF Y) are for determining additional offsets of chroma Cb and Cr after SAO Cb and SAO Cr. For example, first, the current chroma sample 602 is classified using the juxtaposed 604 and adjacent (white) luminance samples 606, and the corresponding CCSAO offset values for each classification are added to the current chroma sample value. FIG. 6B is a block diagram showing a CCSAO system and process applied to luminance and chroma samples according to an embodiment of the present disclosure, taking DBF Y / Cb / Cr as an input. FIG. 6C is a block diagram showing a CCSAO system and process that can operate independently according to an embodiment of the present disclosure. In summary, in one embodiment, the current luminance sample and adjacent luminance samples, juxtaposed and adjacent chroma samples (Cb and Cr) can be used to classify the current luminance sample. In one embodiment, the juxtaposed and adjacent luminance samples, juxtaposed and adjacent cross-chroma samples, and current and adjacent chroma samples can be used to classify the current chroma sample (Cb or Cr). In one embodiment, CCSAO can be cascaded (1) after DBF Y / Cb / Cr, (2) after the reconstructed image Y / Cb / Cr before DBF, (3) after SAO Y / Cb / Cr, or (4) after ALF Y / Cb / Cr.

[0073] In one embodiment, CCSAO can also be applied in parallel with other encoding / decoding tools such as ESAO in the AVS standard or CDEF in the AV1 standard. FIG. 6D is a block diagram showing a CCSAO system and process applied in parallel with ESAO in the AVS standard according to an embodiment of the present disclosure.

[0074] FIG. 6E is a block diagram showing a CCSAO system and process applied after SAO according to an embodiment of the present disclosure. In one embodiment, FIG. 6E shows that the position of CCSAO is after SAO, i.e., at the position of the cross-component adaptive loop filter (CCALF) in the VVC standard. FIG. 6F is a block diagram showing that the CCSAO system and process according to an embodiment of the present disclosure can operate independently without CCALF. In one embodiment, SAO Y / Cb / Cr may be replaced by, for example, ESAO in the AVS 3 standard.

[0075] FIG. 6G is a block diagram showing a CCSAO system and process applied in parallel with CCALF according to an embodiment of the present disclosure. In one embodiment, FIG. 6G shows that CCSAO can be applied in parallel with CCALF. In one embodiment, in FIG. 6G, the positions of CCALF and CCSAO can be switched. In one embodiment, in FIGS. 6A-6G, or throughout the present disclosure, the SAO Y / Cb / Cr block may be replaced by ESAO Y / Cb / Cr (in AVS 3) or CDEF (in AV1). Note that Y / Cb / Cr can also be represented as Y / U / V in the video encoding / decoding region.

[0076] In one embodiment, the current chroma sample classification is to reuse the SAO type (edge offset (EO) or BO), classification, and category of the collocated luminance samples. The corresponding CCSAO offset can be signaled or derived by the decoder itself. For example, let h_Y be the collocated luminance SAO offset, and h_Cb and h_Cr be the CCSAO Cb and Cr offsets respectively. h_Cb (or h_Cr) = w * h_Y, where w can be selected from a limited table. For example, +-1 / 4, +-1 / 2, 0, +-1, +-2, +-4... etc., where |w| includes only powers of 2 values.

[0077] In one embodiment, a comparison score [-8, 8] between a collocated luminance sample (Y0) and eight adjacent luminance samples is used, thereby generating a total of 17 classifications. Initial classification = 0 Circulate on eight adjacent luminance samples (Yi, i = 1~8) If Y0 > Yi, classification += 1 Otherwise, if Y0 < Yi, classification -= 1

[0078] In one embodiment, the above classification methods can be combined. For example, the comparison score can be combined with SAO BO (32-band classification) to increase diversity, generating a total of 17 * 32 classifications. In one embodiment, Cb and Cr can use the same classification to reduce complexity or save bits.

[0079] FIG. 7 is a block diagram showing a sample process using CCSAO according to an embodiment of the present disclosure. Specifically, FIG. 7 shows that the inputs of vertical and horizontal DBFs can be introduced as inputs of CCSAO to simplify classification determination or increase flexibility. For example, Y0_DBF_V, Y0_DPF_H, and Y0 are the collocated luminance samples in the inputs of DBF_V, DBF_H, and SAO, respectively. Yi_DBF_V, Yi_DBF_H, and Yi are the eight adjacent luminance samples in the inputs of DBF_V, DBF_H, and SAO, respectively, where i = 1~8. Max Y0 = max (Y0_DBF_V, Y0_DBF_H, Y0_DBF) Max Yi = max (Yi_DBF_V, Yi_DBF_H, Yi_DBF) Then, the maximum Y0 and the maximum Yi are input into the CCSAO classification.

[0080] FIG. 8 is a block diagram showing that the CCSAO process is interleaved into vertical and horizontal DBFs according to an embodiment of the present disclosure. In some embodiments, the CCSAO blocks of FIGS. 6, 7, and 8 may be optional. For example, for the first CCSAO_V, the input of the DBF_V luminance samples is used as the CCSAO input while using Y0_DBF_V and Yi_DBF_V that apply a sample process similar to FIG. 6.

[0081] In some embodiments, the implemented CCSAO syntax is shown in Table 2 below. JPEG0007684411000002.jpg94166

[0082] In some embodiments, when one additional chroma offset is signaled to signal the CCSAO Cb and Cr offset values, another chroma component offset can be derived by positive sign, negative sign, or weighting to save bit overhead. For example, let h_Cb and h_Cr be the offset amounts of CCSAO Cb and Cr respectively. In the case of explicit signaling w with limited |w| candidates and w = +-|w|, h_Cr can be derived from h_Cb without explicit signaling of h_Cr itself. h_Cr = w * h_Cb

[0083] FIG. 9 is a flowchart showing an exemplary process 900 for decoding a video signal using cross-component correlation according to an embodiment of the present disclosure.

[0084] The video decoder 30 receives a video signal including a first component and a second component (910). In some embodiments, the first component is the luminance component of the video signal and the second component is the chroma component of the video signal.

[0085] The video decoder 30 also receives a plurality of offsets associated with the second component (920).

[0086] Next, the video decoder 30 obtains a classification category related to the second component by using the characteristic measurement of the first component (930). For example, in FIG. 6, first, the current chroma sample 602 is classified using the juxtaposed 604 and adjacent (white) luminance sample 606, and the corresponding CCSAO offset value is added to the current chroma sample.

[0087] Furthermore, the video decoder 30 selects a first offset from among a plurality of offsets for the second component according to the classification category (940).

[0088] The video decoder 30 additionally modifies the second component based on the selected first offset (950).

[0089] In an embodiment, obtaining a classification category related to the second component by using the characteristic measurement of the first component (930) includes obtaining the classification category of each sample of the second component by using each sample that is a juxtaposed sample of the first component corresponding to each sample of the second component. For example, the current chroma sample classification reuses the SAO type (EO or BO), classification, and category of the juxtaposed luminance sample.

[0090] In an embodiment, obtaining a classification category related to the second component by using the characteristic measurement of the first component (930) includes obtaining the classification category of each sample of the second component by using each sample of the first component that is reconstructed before being deblocked or reconstructed after being deblocked. In an embodiment, the first component is deblocked by a deblocking filter (DBF). In an embodiment, the first component is deblocked by a luminance deblocking filter (DBF Y). For example, instead of FIG. 6 or FIG. 7, the CCSAO input may be before the DBF Y.

[0091] In one embodiment, the characteristic measurement is derived by dividing the sample value range of the first component into a plurality of bands and selecting a band based on the intensity value of the sample in the first component. In one embodiment, the characteristic measurement is derived from a band offset (BO).

[0092] In one embodiment, the characteristic measurement is derived based on the direction and intensity of the edge information of the sample in the first component. In one embodiment, the characteristic measurement is derived from an edge offset (EO).

[0093] In one embodiment, changing the second component (950) includes directly adding the selected first offset to the second component. For example, adding the corresponding CCSAO offset value to the current chroma component sample.

[0094] In one embodiment, changing the second component (950) includes mapping the selected first offset to the second offset and adding this mapped second offset to the second component. For example, when signaling one additional chroma offset to signal the CCSAO Cb and Cr offset values, another chroma offset can be derived using positive or negative signs or weighting to save bit overhead.

[0095] In one embodiment, receiving a video signal (910) includes receiving a syntax element indicating whether a video signal decoding method using CCSAO for the video signal is valid in a sequence parameter set (SPS). In one embodiment, the cc_sao_enabled_flag indicates whether CCSAO is valid at the sequence level.

[0096] In one embodiment, receiving (910) a video signal includes receiving a syntax element indicating whether a video signal decoding method using CCSAO for a second component is effective at a slice level. In one embodiment, slice_cc_sao_cb_flag or slice_cc_sao_cr_flag indicates whether CCSAO is effective for respective slices for Cb or Cr.

[0097] In one embodiment, receiving (920) a plurality of offsets related to a second component includes receiving different offsets of different coding tree units (CTUs). In one embodiment, cc_sao_offset_sign_flag indicates the sign of the offset for a CTU, and cc_sao_offset_abs indicates the CCSAO Cb and Cr offset values of the current CTU.

[0098] In one embodiment, receiving (920) a plurality of offsets related to a second component includes receiving a syntax element indicating whether the offset of the received CTU is the same as the offset of one of the adjacent CTUs that is the left adjacent CTU or the upper adjacent CTU of this CTU. For example, cc_sao_merge_up_flag indicates whether the CCSAO offset is merged from the left CTU or the upper CTU.

[0099] In one embodiment, the video signal further includes a third component, and the method for decoding the video signal using CCSAO includes receiving a second plurality of offsets related to the third component, obtaining a second classification category related to the third component using the characteristic measurement of the first component, selecting a third offset from the second plurality of offsets of the third component according to the second classification category, and changing the third component based on the selected third offset.

[0100] FIG. 11 is a block diagram showing a sample process in which all juxtaposed and adjacent (white) luminance / chroma samples according to an embodiment of the present disclosure can be fed into the CCSAO classification. FIGS. 6A, 6B and FIG. 11 show the input of the CCSAO classification. In FIG. 11, the current chroma sample is 1104, the cross-component juxtaposed chroma sample is 1102, and the juxtaposed luminance sample is 1106.

[0101] In one embodiment, classifier example (C0) uses the juxtaposed luminance or chroma sample value (Y0) (Y4 / U4 / V4 in FIGS. 6B and 6C) in FIG. 12 below for classification. Let band_num be the number of equally divided bands of the dynamic range of luminance or chroma, and bit_depth be the sequence bit depth. An example of the classification index of the current chroma sample is as follows. Class (C0) = (Y0 * band_num) >> bit_depth

[0102] In one embodiment, the classification takes rounding into account, for example, as follows. Class (C0) = ((Y0 * band_num) + (1 << bit_depth)) >> bit_depth

[0103] Table 3 shows some examples of band_num and bit_depth. Table 3 shows three classification examples when the number of bands is different for each classification example. JPEG0007684411000003.jpg101166JPEG0007684411000004.jpg8157

[0104] In one embodiment, the classifier uses different luminance sample positions for the C0 classification. FIG. 10A is a block diagram showing a classifier that uses different luminance (or chroma) sample positions for the C0 classification according to an embodiment of the present disclosure. For example, for the C0 classification, adjacent Y7 is used instead of Y0.

[0105] In one embodiment, different classifiers can be switched at the sequence parameter set (SPS) / adaptive parameter set (APS) / picture parameter set (PPS) / picture header (PH) / slice header (SH) / coded tree unit (CTU) / coded unit (CU) level. For example, in FIG. 10, as shown in Table 4 below, Y0 is used for POC0, while Y7 is used for POC1. JPEG0007684411000005.jpg39157In one embodiment, FIG. 10B shows some examples of different shapes of luminance candidates according to an embodiment of the present disclosure. For example, constraints may be applied to the shape. As shown in FIGS. 10B(b)(c)(d), the total number of luminance candidates may have to be a power of 2. As shown in FIGS. 10B(a)(c)(d)(e), the number of luminance candidates may have to be symmetric horizontally and vertically with respect to the chroma sample (at the center). In one embodiment, both the power-of-2 constraint and the symmetry constraint may be applied to the chroma candidates. The U / V portions of FIGS. 6B and 6C show examples of the symmetry constraint. In one embodiment, different color formats can have different classifier "constraints". For example, as shown in FIGS. 6B and 6C, the 420 color format uses luminance / chroma candidate selection (one candidate selected from a 3×3 shape), the 444 color format uses FIG. 10B(f) for luminance and chroma candidate selection, the 422 color format uses FIG. 10B(g) for luminance candidates (2 chroma samples share 4 luminance candidates), and FIG. 10B(f) for chroma candidates.

[0106] In one embodiment, the C0 position and C0 band_num can be combined and switched at the SPS / APS / PPS / PH / SH / CTU level. Different combinations may be different classifiers, as shown in Table 5 below. JPEG0007684411000006.jpg39157

[0107] In one embodiment, the juxtaposed luminance sample value (Y0) is replaced with a value (Yp) obtained by weighting the juxtaposed luminance sample and the adjacent luminance sample. FIG. 12 shows an exemplary classifier according to an embodiment of the present disclosure for replacing the juxtaposed luminance sample value with a value obtained by weighting the juxtaposed luminance sample and the adjacent luminance sample. The juxtaposed luminance sample value (Y0) can be replaced with a phase correction value (Yp) obtained by weighting the adjacent luminance sample. Different Yps may be different classifiers.

[0108] In one embodiment, different Yps are applied to different chroma formats. For example, in FIG. 12, the Yp in (a) is used for the 420 chroma format, the Yp in (b) is used for the 422 chroma format, and Y0 is used for the 444 chroma format.

[0109] In one embodiment, another classifier (C1) is a comparison score [-8, 8] between the juxtaposed luminance sample (Y0) that generates a total of 17 classifications as shown below and 8 adjacent luminance samples. Initial Class (C1) = 0, circulating through 8 adjacent luminance samples (Yi, i = 1 to 8) If Y0 > Yi, then Class += 1 Otherwise, if Y0 < Yi, then Class -= 1

[0110] In one embodiment, the variable (C1’) calculates only the comparison score [0, 8] and generates 8 classifications. (C1, C1’) is a classifier group, and the PH / SH level flag can be signaled to switch between C1 and C1’. Initial Class (C1) = 0, circulating through 8 adjacent luminance samples (Yi, i = 1 to 8) If Y0 > Yi, then Class += 1

[0111] In one embodiment, different classifiers are combined to generate a common classifier. For example, for different images (different POC values), different classifiers are applied as shown in Table 6-1 below. JPEG0007684411000007.jpg51156

[0112] In one embodiment, another classifier example (C3) classifies using a bit mask as shown in Table 6-2. A 10-bit bit mask is signaled at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock levels to instruct the classifier. For example, the bit mask 11 1100 0000 means that for a given 10-bit luminance sample value, only the most significant bits (MSBs) are used for classification, generating a total of 16 classifications. Another example bit mask 10 0100 0001 means that only 3 bits are used for classification, generating a total of 8 classifications.

[0113] In one embodiment, the bit mask length (N) may be fixed or switched at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock levels. For example, in the case of a 10-bit sequence, the 4-bit mask 1110 is signaled for PH in the image, and the MSB 3 bits b9, b8, b7 are used for classification. Another example is the 4-bit mask 0011 at the LSB, and b0, b1 are used for classification. The bit mask classifier can be applied to luminance or chroma classification. Whether to use the MSB or the LSB for the bit mask N can be fixed or switched at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock levels.

[0114] In one embodiment, the luminance position and the C3 bit mask can be combined and switched at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock levels. Different combinations may be different classifiers.

[0115] In one embodiment, the "maximum number of 1s" of the bitmask restriction can be applied to limit the corresponding offset number. For example, in SPS, the "maximum number of 1s" of the bitmask is restricted to 4, so that the maximum offset in the sequence is 16. Although the bitmask varies depending on the POC, the "maximum number of 1s" does not exceed 4 (the total classification does not exceed 16). The value of the "maximum number of Ss" is signaled and can be switched at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock levels. JPEG0007684411000008.jpg159166

[0116] In one embodiment, as shown in FIG. 11, other cross-component chroma samples such as the chroma sample 1102 and its adjacent samples may also be supplied to the CCSAO classification of the current chroma sample 1104. For example, the Cr chroma sample may be supplied to the CCSAO Cb classification. The Cb chroma sample may be supplied to the CCSAO Cr classification. The classifier for the cross-component chroma samples may be the same as the luminance cross-component classifier, or may have a unique classifier as described in the present disclosure. Two classifiers can be combined to form a combined classifier for classifying the current chroma sample. For example, as shown in Table 6-3 below, the combined classifier combining the cross-component luminance and chroma samples generates a total of 16 classifications. JPEG0007684411000009.jpg80170

[0117] All of the above classifications (C0, C1, C1′, C2, C3) may be combined. For example, refer to Table 6-4 below. JPEG0007684411000010.jpg79170

[0118] In one embodiment, classifier example (C2) uses the differences (Yn) between juxtaposed and adjacent luminance samples. FIG. 12(c) shows an example of Yn having a dynamic range of [-1024, 1023] when the bit depth is 10. Let C2 band_num be the number of equally divided bands of the Yn dynamic range, Class (C2) = (Yn + (1 << bit_depth) * band_num) >> (bit_depth + 1).

[0119] In one embodiment, C0 and C2 are combined to generate a general-purpose classifier. For example, for different images (different POCs), different classifiers are applied as shown in Table 7 below. JPEG0007684411000011.jpg44170

[0120] In one embodiment, all of the above-described classifiers (C0, C1, C1′, C2) are combined. For example, for different images (different POCs), different classifiers are applied as shown in Table 8 below. JPEG0007684411000012.jpg48170

[0121] In one embodiment, multiple classifiers are used for the same POC. The current frame is divided into multiple regions, and each region uses the same classifier. For example, as shown in Table 9 below, three different classifiers are used for POC 0, and it is signaled which classifier (0, 1, or 2) is being used at the CTU level. JPEG0007684411000013.jpg56170

[0122] In one embodiment, the maximum number of multiple classifiers (the multiple classifiers are also referred to as alternative offset sets) can be fixed or signaled at the SPS / APS / PPS / PH / SH / Region / CTU / CU / Subblock levels. In one example, the fixed (predetermined) maximum number of multiple classifiers is 4. In this case, at POC 0, four different classifiers are used, and at the CTU level, which classifier (0, 1, or 2) is being used is signaled. The cut-off unit (TU) code can indicate whether the classifier is applied to each luminance or chrominance CTB. For example, as shown in Table 10 below, when the TU code is 0, CCSAO is not applied; when the TU code is 10, set 0 is applied; when the TU code is 110, set 1 is applied; when the TU code is 1110, set 2 is applied; and when the TU code is 1111, set 3 is applied. Fixed-length codes, golom-rice codes, and exponential-golomb codes can also be used for the classifier (offset set index) with respect to the CTB. At POC 1, three different classifiers are used. JPEG0007684411000014.jpg90170

[0123] Examples of Cb and Cr CTB offset set indices for 1280×720 sequence POC 0 (when the CTU size is 128×128, the number of CTUs in a frame is 10×6) are provided. POC 0 Cb uses 4 offset sets and Cr uses 1 offset set. As shown in Table 11-1 below, when the offset set index is 0, CCSAO is not applied; when the offset set index is 1, set 0 is applied; when the offset set index is 2, set 1 is applied; when the offset set index is 3, set 2 is applied; when the offset set index is 4, set 3 is applied. The type refers to the position of the selected collocated luminance sample (Yi). Different offset sets may have different types, band_num, and corresponding offsets. JPEG0007684411000015.jpg105167

[0124] In one embodiment, examples of combining collocated / current and adjacent Y / U / V samples for application to classification (3-component combined bandNum classification for each Y / U / V component) are shown in Table 11-2 below. For POC 0, {2, 4, 1} offset sets are applied to {Y, U, V} respectively. Each offset set can be adaptively switched at the SPS / APS / PPS / PH / SH / CTU / CU / Subblock levels. Different offset sets can have different classifiers. For example, as the candidate positions (candPos) shown in FIGS. 6B and 6C, to classify the current Y4 luminance sample, Y set 0 selects {current Y4, collocated U4, collocated V4} as candidates, each having different bandNum {Y, U, V} = {16, 1, 2}. The sample values of the selected {Y, U, V} candidates {candY, candU, candV} are used, and the total number of classifications is 32. The classification index derivation can be expressed as follows. bandY = (candY * bandNumY) >> BitDepth; bandU = (candU * bandNumU) >> BitDepth; bandV = (candV * bandNumV) >> BitDepth; classIdx = bandY * bandNumU * bandNumV + bandU * bandNumV + bandV

[0125] Another example is the POC1 component V set1 classification. In this example, candPos = {neighboring Y8, neighboring U3, neighboring V0} with bandNum = {4, 1, 2} is used to generate 8 classifications. JPEG0007684411000016.jpg83167

[0126] In some embodiments, the maximum band_num (bandNumY, bandNumU, or bandNumV) may be fixed or signaled at the SPS / APS / PPS / PH / SH / CTU / CU level. For example, a maximum band_num = 16 is fixed in the decoder, and 4 bits are signaled for each frame to indicate the C0 band_num in the frame. Table 12 below shows some examples of other maximum band_nums. JPEG0007684411000017.jpg106168

[0127] In some embodiments, restrictions may be applied to the C0 classification. For example, band_num (bandNumY, bandNumU, or bandNumV) may be restricted to only powers of 2. Instead of explicitly signaling band_num, the syntax band_num_shift is signaled. The decoder may use shift operations to avoid multiplication. Different band_num_shifts may be used for different components. Class (C0) = (Y0 >> band_num_shift) >> bit_depth

[0128] In another operation example, rounding is considered to reduce errors. Class (C0) = ((Y0 + (1 << (band_num_shift - 1))) >> band_num_shift) >> bit_depth

[0129] For example, when band_num_max (Y, U, or V) is 16, the possible band_num_shift candidates are 0, 1, 2, 3, 4 corresponding to band_num = 1, 2, 4, 8, 16 as shown in Table 13. JPEG0007684411000018.jpg96168

[0130] In one embodiment, the classifiers applied to Cb and Cr are different. All Cb and Cr offsets for all classifications may be signaled individually. For example, as shown in Table 14 below, different offsets signaled are applied to different chroma components. JPEG0007684411000019.jpg43168

[0131] In one embodiment, the maximum offset value is fixed or signaled in the Sequence Parameter Set (SPS) / Adaptive Parameter Set (APS) / Picture Parameter Set (PPS) / Picture Header (PH) / Slice Header (SH). For example, the maximum offset is between [-15, 15]. Different components may have different maximum offset values.

[0132] In one embodiment, differential pulse code modulation (DPCM) may be used for offset signaling. For example, the offsets {3, 3, 2, 1, -1} may be signaled as {3, 0, -1, -1, -2}.

[0133] In one embodiment, the offset may be stored in the APS or memory buffer for the next image / slice reuse. The index may be signaled to indicate whether the stored previous frame offset is used for the current image.

[0134] In one embodiment, the classifiers for Cb and Cr are the same. The Cb and Cr offsets for all classifications may be signaled in combination, for example, as shown in Table 15 below. JPEG0007684411000020.jpg30165

[0135] In one embodiment, the classifiers for Cb or Cr may be the same. The Cb or Cr offsets for all classifications may be signaled in combination by the sign flag difference, for example, as shown in Table 16 below. According to Table 16, when the Cb offset is (3, 3, 2, -1), the derived Cr offset is (-3, -3, -2, 1). JPEG0007684411000021.jpg43165

[0136] In one embodiment, a sign flag may be signaled for each classification. For example, it is shown in Table 17 below. According to Table 17, when the Cb offset is (3, 3, 2, -1), the Cr offset derived based on each signed flag is (-3, 3, 2, 1). JPEG0007684411000022.jpg44165

[0137] In one embodiment, the classifiers for Cb and Cr may be the same. The Cb and Cr offsets for all classifications may be signaled in combination by a weight difference, for example, as shown in Table 18 below. The weight (w) can be selected within a limited table such as +-1 / 4, +-1 / 2, 0, +-1, +-2, +-4..., where |w| includes only powers of two values. According to Table 18, when the Cb offset is (3, 3, 2, -1), the Cr offset derived based on each signed flag is (-6, -6, -4, 2). JPEG0007684411000023.jpg42166

[0138] In one embodiment, the weights for each classification may be signaled. For example, as shown in Table 19 below. According to Table 19, when the Cb offset is (3, 3, 2, -1), the Cr offset derived based on each signed flag is (-6, 12, 0, -1). JPEG0007684411000024.jpg48166

[0139] In one embodiment, when multiple classifiers are used with the same POC, different offset sets are signaled individually or in combination.

[0140] In one embodiment, previously decoded offsets may be stored for future frames. An index may be signaled for the current frame to indicate which previously decoded offset set is being used to reduce the signaling overhead of the offset. For example, as shown in Table 20 below, the POC0 offset can be reused by POC2 with the signaling offset set idx = 0. JPEG0007684411000025.jpg135167

[0141] In one embodiment, the reuse offset sets idx for Cb and Cr may be different, for example, as shown in Table 21 below. JPEG0007684411000026.jpg137165

[0142] In certain embodiments, offset signaling may use additional syntax including start and length so as to reduce signaling overhead. For example, when band_num = 256, only the offsets of band_idx = 37 to 44 are signaled. In the example of Table 22-1 below, both the start and length syntax are 8-bit fixed-length encoded and decoded to match the band_num bits. JPEG0007684411000027.jpg84153JPEG0007684411000028.jpg9164

[0143] In certain embodiments, when CCSAO is applied to all YUV3 components, adjacent and neighboring YUV samples may be combined for classification, and all the above offset signaling methods for Cb / Cr may be extended to Y / Cb / Cr. In certain embodiments, different component offset sets may be stored and used individually (each component has its own stored save set), or stored and used in combination (each component shares / reuses the same stored one). Table 22-2 below shows an example of an individual set. JPEG0007684411000029.jpg169169

[0144] In certain embodiments, when the sequence bit depth is higher than 10 (or a certain bit depth), the offset may be quantized before being signaled. On the decoder side, as shown in Table 23 below, inverse quantization is performed before applying the decoded offset. For example, for a 12-bit sequence, the decoded offset is shifted 2 bits to the left (inverse quantization). JPEG0007684411000030.jpg52169

[0145] In one embodiment, the offset amount may be calculated as CcSaoOffsetVal=( 1 - 2 * ccsao_offset_sign_flag ) * (ccsao_offset_abs << ( BitDepth - Min( 10, BitDepth ) ) ).

[0146] In one embodiment, the sample processing is described below. Let R(x, y) be the input luminance or chroma sample value before CCSAO, and R’(x, y) be the output luminance or chroma sample value after CCSAO. Then it is as follows. offset = ccsao_offset [class_index of R(x, y)] R’(x, y) = Clip3( 0, (1 << bit_depth) - 1, R(x, y) + offset )

[0147] According to the above formula, each luminance or chroma sample value R(x, y) is classified by the current image and / or the indicated classifier of the current offset set idx. The corresponding offset of the derived classification index is added to each luminance or chroma sample value R(x, y). The clip function Clip 3 is applied to (R(x, y)+offset) so that the output luminance or chroma sample value R’(x, y) is within the bit depth dynamic range, for example, in the range 1 to (1 << bit_depth) - 1.

[0148] In one embodiment, the boundary processing is described below. When either the juxtaposed and adjacent luminance (chroma) samples for classification are outside the current image, CCSAO is not applied to the current chroma (luminance) sample. FIG. 13A is a block diagram showing that when either the juxtaposed and adjacent luminance (chroma) samples for classification are outside the current image, CCSAO is not applied to the current chroma (luminance) sample according to an embodiment of the present disclosure. For example, in FIG. 13A(a), when applying the classifier, CCSAO is not applied to the chroma components in the leftmost column of the current image. For example, when using C1’, as shown in FIG. 13A(b), CCSAO should not be applied to the chroma components in the leftmost column and the uppermost row of the current image.

[0149] FIG. 13B is a block diagram showing that when either the juxtaposed and adjacent luminance or chroma samples for classification are outside the current image, CCSAO is applied to the current luminance or chroma sample according to an embodiment of the present disclosure. In one embodiment, one change is that when either the juxtaposed and adjacent luminance or chroma samples for classification are outside the current image, the lost samples can be repeatedly used as shown in FIG. 13B(a), or the lost samples can be mirror padded to create samples for classification as shown in FIG. 13B(b), so that CCSAO can be applied to the current luminance or chroma sample.

[0150] FIG. 14 is a block diagram showing that when corresponding selected juxtaposed or adjacent luminance samples for classification are outside the virtual space defined by a virtual boundary according to an embodiment of the present disclosure, CCSAO is not applied to the current chroma sample. In one embodiment, the virtual boundary (VB) is a virtual line that separates the space within the image frame. In one embodiment, when a virtual boundary (VB) is applied to the current frame, CCSAO shall not be applied to chroma samples having corresponding selected luminance positions outside the virtual space defined by the virtual boundary. FIG. 14 shows an example of a virtual boundary for a C0 classifier having nine luminance position candidates. For each CTU, CCSAO is not applied to chroma samples where the corresponding selected luminance position is outside the virtual space surrounded by the virtual boundary. For example, in FIG. 14(a), when the selected Y7 luminance sample position is on the other side of the horizontal virtual boundary 1406 located 4 pixel rows from the bottom of the frame, CCSAO is not applied to the chroma sample 1402. For example, in FIG. 14(b), when the selected Y5 luminance sample position is on the other side of the vertical virtual boundary 1408 located y pixel rows from the right side of the frame, CCSAO is not applied to the chroma sample 1404.

[0151] FIG. 15 shows applying duplicate or mirror padding to luminance samples outside the virtual boundary according to an embodiment of the present disclosure. FIG. 15(a) shows an example of duplicate padding. When the original Y7 is selected as the classifier located on the bottom side of VB 1502, instead of the original Y7 luminance sample value, the Y4 luminance sample value is applied to the classification (copied to the Y7 position). FIG. 15(b) shows an example of mirror padding. When Y7 is selected as the classifier located on the bottom side of VB 1504, instead of the original Y7 luminance sample value, the Y1 luminance sample value symmetric to the Y7 value with respect to the Y0 luminance sample is applied to the classification. The padding method provides the possibility of applying CCSAO to more chroma samples and can obtain more coding and decoding gains.

[0152] In certain embodiments, restrictions can be applied to reduce the line buffer required for CCSAO and simplify the boundary processing condition check. FIG. 16 shows that when all nine juxtaposed adjacent luminance samples according to an embodiment of the present disclosure are used for classification, one additional luminance line buffer, i.e., all the line luminance samples of line -5 above the current VB 1602, is required. (a) of FIG. 10B shows an example where only six luminance candidates are used for classification to reduce the line buffer and eliminate the need for the additional boundary checks of FIGS. 13A and 13B.

[0153] In certain embodiments, using luminance samples for CCSAO classification may increase the implementation cost of the luminance line buffer and further increase the implementation cost of the decoder hardware. FIG. 17 shows that in the AVS according to an embodiment of the present disclosure, the intersection of nine luminance candidate CCSAO and VB 1702 may increase two additional luminance line buffers. For the luminance and chrominance samples above the virtual boundary (VB) 1702, DBF / SAO / ALF is processed in the current CTU row. For the luminance and chrominance samples below VB 1702, DBF / SAO / ALF is processed in the next CTU row. In the AVS decoder hardware design, the luminance line -4~-1 samples before DBF, the line -5 sample before SAO, the chrominance line -3~-1 samples before DBF, and the line -4 sample before SAO are stored as line buffers for the next CTU row DBF / SAO / ALF processing. When processing the next CTU row, luminance and chrominance samples that do not exist in the line buffer cannot be used. However, for example, at the chrominance line -3(b) position, chrominance samples are processed in the next CTU row, but CCSAO requires the luminance sample lines -7, -6, -5 before SAO for classification. Since the luminance sample lines -7, -6 before SAO are not in the line buffer, they cannot be used. Also, increasing the luminance sample lines -7 and -6 before SAO in the line buffer increases the implementation cost of the decoder hardware. In one example, the luminance VB (line -4) and the chrominance VB (line -3) may be different (not aligned).

[0154] FIG. 18, similar to FIG. 17, shows that in a certain embodiment of the present disclosure, in VVC, the intersection of nine luminance candidate CCSAOs and VB 1802 may increase one additional luminance line buffer. VB may vary according to different standards. In VVC, since the luminance VB is line - 4 and the chroma VB is line - 2, the nine candidate CCSAOs may increase one luminance line buffer.

[0155] In a certain embodiment, in the first countermeasure, when any luminance candidate of the chroma samples crosses VB (outside the current chroma sample VB), CCSAO is disabled for the chroma samples. FIGS. 19A - 19C show that in AVS and VVC according to a certain embodiment of the present disclosure, when any of the luminance candidates of the chroma samples crosses VB 1902 (outside the current chroma sample VB), CCSAO is disabled for the chroma samples. FIG. 14 also shows an example of this embodiment.

[0156] In a certain embodiment, in the second countermeasure, for the luminance candidate that "crosses VB", for example, from a luminance line near VB, such as luminance line - 4, and located on the opposite side of VB, duplicate padding is used for CCSAO. FIGS. 20A - 20C show that in AVS and VVS according to a certain embodiment of the present disclosure, when any of the luminance candidates of the chroma samples crosses VB 2002 (outside the current chroma sample VB), CCSAO is enabled to use duplicate padding for the chroma samples. FIG. 14(a) also shows an example of this embodiment.

[0157] In a certain embodiment, in the third countermeasure, for the luminance candidate that "crosses VB", mirror padding is used for CCSAO from below the luminance VB. FIGS. 21A - 21C show that in AVS and VVC according to a certain embodiment of the present disclosure, when any of the luminance candidates of the chroma samples crosses VB 2102 (outside the current chroma sample VB), CCSAO is enabled to use mirror padding for the chroma samples. FIGS. 14(b) and 13B(b) also show examples of this embodiment.

[0158] In one embodiment, in the fourth countermeasure, "bilateral symmetric padding" is used to apply CCSAO. FIGS. 22A - 22B show that for some examples of different CCSAO shapes (e.g., 9 luminance candidates (FIG. 22A) and 8 luminance candidates (FIG. 22B)) according to an embodiment of the present disclosure, CCSAO enables the use of bilateral symmetric padding. For a luminance sample set having a juxtaposed center luminance sample of chroma samples, when one side of the luminance sample set is outside VB 2202, bilateral symmetric padding is applied to both sides of this luminance sample set. For example, in FIG. 22A, since the luminance samples Y0, Y1, Y2 are outside VB 2202, Y0, Y1, Y2 and Y6, Y7, Y8 are padded by Y3, Y4, Y5. For example, in FIG. 22B, since the luminance sample Y0 is outside VB 2202, Y0 is padded by Y2, and Y7 is padded by Y5.

[0159] The padding method provides the possibility of applying CCSAO to more chroma samples and can obtain more coding / decoding gains.

[0160] In one embodiment, in the bottom image (or slice, tile, brick) boundary CUT row, since the samples below VB are processed in the current CTU row, the above special processing (countermeasures 1, 2, 3, 4) is not applied in the bottom image (or slice, tile, brick) boundary CTU row. For example, a 1920×1080 frame is divided by 128×128 CTUs. One frame contains 15×9 CTUs (rounded). The bottom row of CTUs is the 15th row of CTUs. The decoding process is executed for each CTU row, and for each CTU row, CTUs are processed one by one. Deblocking needs to be applied along the horizontal CTU boundary between the current CTU row and the next CTU row. Within one CTU, for the bottom 4 / 2 luminance / chroma lines, the DBF samples (in the case of VVC) are processed in the next CTU row and cannot be used for CCSAO in the current CTU row, so CTB VB is applied to each CTU row. However, in the bottom CTU row of the image frame, there is no remaining next CTU row, and the bottom 4 / 2 luminance / chroma line DBF samples are available in the current CTU row and are processed by DBF in the current CTU row.

[0161] In one embodiment, restrictions are applied to reduce the line buffer required for CCSAO and simplify the boundary processing condition check shown in FIG. 16. FIG. 23 shows the restrictions using a limited number of luminance candidates for classification according to an embodiment of the present disclosure. (a) of FIG. 23 shows the restrictions using only 6 luminance candidates for classification. (b) of FIG. 23 shows the restrictions using only 4 luminance candidates for classification.

[0162] In one embodiment, an application area is realized. The CCSAO application area unit may be CTB-based. That is, within one CTB, the on / off control, CCSAO parameters (offsets for classification, luminance candidate positions, band_num, bit masks, etc., offset set indexes) are the same.

[0163] In some embodiments, the application area may not be aligned with the CTB boundary. For example, the application area is not aligned with the chroma CTB boundary and is offset. The syntax (on / off control, CCSAO parameters) is still signaled for each CTB, but the actual application area is not aligned with the CTB boundary. FIG. 24 shows that the CCSAO application area according to an embodiment of the present disclosure is not aligned with the CTB / CTU boundary 2406. For example, the application area is not aligned with the chroma CTB / CTU boundary 2406 and is shifted 4 samples to the upper left in the VB 2408 only by (4, 4). Such a CTB boundary design that is not aligned is advantageous for the block release process because the same block release parameters are used for each 8×8 block release process area.

[0164] In some embodiments, the CCSAO application area unit (mask size) may be variable (larger or smaller than the CTB size) as shown in Table 24. The mask size may vary depending on the component. The mask size may be switched at the SPS / APS / PPS / PH / SH / area / CTU / CU / sub-block level. For example, in the PH, a series of mask on / off flags and offset set indexes indicating each CCSAO area information are signaled. JPEG0007684411000031.jpg59170

[0165] In some embodiments, the CCSAO application area frame partition may be fixed. For example, the frame is partitioned into N areas. FIG. 25 shows fixing the CCSAO application area frame partition using the CCSAO parameters according to an embodiment of the present disclosure.

[0166] In one embodiment, each region may have its own region on / off control flag and CCSAO parameter. Also, when the region size is larger than the CTB size, it may have both a CTB on / off control flag and a region on / off control flag. FIGS. 25(a) and (b) are diagrams showing an example of dividing a frame into N regions. FIG. 25(a) shows a vertical division of four regions. FIG. 25(a) shows a square division of four regions. In one embodiment, similar to the image-level CTB full-on control flag (ph_cc_sao_cb_ctb_control_flag / ph_cc_sao_cr_ctb_control_flag), when the region on / off control flag is off, the CTB on / off flag may be further signaled. Otherwise, CCSAO is applied to all CTBs within that region without signaling the CTB flag.

[0167] In one embodiment, different CCSAO application regions may share the same region on / off control and CCSAO parameters. For example, in FIG. 25(c), regions 1 to 2 share the same parameters, and regions 3 to 15 share the same parameters. FIG. 25(c) also shows that the region on / off control flag and CCSAO parameters are signaled in Hilbert scan order.

[0168] In one embodiment, the CCSAO application region unit may be a quadtree / binary tree / trinary tree divided from the image / slice / CTB level. Similar to CTB division, a series of division flags are signaled to indicate the CCSAO application region partition. FIG. 26 is a diagram showing that the CCSAO application region according to an embodiment of the present disclosure is divided into a binary tree (BT) / quadtree (QT) / trinary tree (TT) from the frame / slice / CTB level.

[0169] FIG. 27 is a block diagram showing a plurality of classifiers that are used and switched at different levels within an image frame according to an embodiment of the present disclosure. In one embodiment, when multiple classifiers are used in one frame, the method of how to apply the classifier set index can be switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / Subblock level. For example, use four sets of classifiers in one frame and switch at the PH as shown in Table 25 below. In FIGS. 27(a) and (c), the default fixed region classifier is shown. FIG. 27(b) shows that the classifier set index is signaled at the mask / CTB level, where 0 indicates CCSAO off for the CTB, and 1-4 indicate the set index. JPEG0007684411000032.jpg58170

[0170] In one embodiment, for the default region, when the CTB within the region does not use the default set index (for example, the region level flag is 0) and uses other classifier sets in the frame, the region level flag may be signaled. For example, when using the default set index, the region level flag is 1. For example, in the 4 regions of the square section, the classifier sets as shown in Table 26 below are used. JPEG0007684411000033.jpg43170

[0171] FIG. 28 is a block diagram showing that the CCSAO application area division according to an embodiment of the present disclosure is dynamic and can be switched at the image level. For example, in FIG. 28(a), since three CCSAO offset sets (set_num = 3) are used in this POC, it shows that the image frame is divided into three regions vertically. In FIG. 28(b), since four CCSAO offset sets (set_num = 4) are used in this POC, it shows that the image frame is divided into four regions horizontally. In FIG. 28(c), since three CCSAO offset sets (set_num = 3) are used in this POC, it shows that the image frame is divided into three regions in raster. Each region may have its own region full-on flag to save each CTB on / off control bit. The number of regions depends on the image set_num notified by the signal.

[0172] In an embodiment, the implemented CCSAO syntax is shown in Table 27 below. In AVS3, the term patch is similar to a slice, and the patch header is similar to the slice header. FLC represents a fixed-length code. TU represents a truncated unary code. EGk represents the k-th exponential Golomb code, where k may be fixed. JPEG0007684411000034.jpg243165JPEG0007684411000035.jpg101167

[0173] When the upper-level flag is off, the lower-level flag can be estimated from the off state of the flag, and there is no need to signal it. For example, when ph_cc_sao_cb_flag in this image is false, ph_cc_sao_cb_band_num_minus1, ph_cc_sao_cb_luma_type, cc_sao_cb_offset_sign_flag, cc_sao_cb_offset_abs, ctb_cc_sao_cb_flag, cc_sao_cb_merge_left_flag, and cc_sao_cb_merge_up_flag do not exist and are presumed to be false.

[0174] In one embodiment, sps_ccsao_enabled_flag is conditional on the SPS SAO enable flag as shown in Table 28 below. JPEG0007684411000036.jpg42167

[0175] In one embodiment, ph_cc_sao_cb_ctb_control_flag and ph_cc_sao_cr_ctb_control_flag indicate whether to enable the Cb / Cr CTB on / off control granularity. When ph_cc_sao_cb_ctb_control_flag and ph_cc_sao_cr_ctb_control_flag are enabled, ctb_cc_sao_cb_flag and ctb_cc_sao_cr_flag may be further signaled. Otherwise, whether to apply CCSAO to the current image depends on ph_cc_sao_cb_flag and ph_cc_sao_cr_flag without further signaling ctb_cc_sao_cb_flag and ctb_cc_sao_cr_flag at the CTB level.

[0176] In one embodiment, for ph_cc_sao_cb_type and ph_cc_sao_cr_type, a flag is further signaled to indicate whether the central co-located luminance position (Y0 position in FIG. 10) is used for the classification of chroma samples, so as to reduce the bit overhead. Similarly, when cc_sao_cb_type and cc_sao_cr_type are signaled at the CTB level, the flag may be further signaled by the same mechanism. For example, when the number of C0 luminance position candidates is 9, as shown in Table 29 below, cc_sao_cb_type0_flag is further signaled to determine whether the central co-located luminance position is used. If the central co-located luminance position is not used, cc_sao_cb_type_idc is used to indicate which of the remaining 8 adjacent luminance positions is used. JPEG0007684411000037.jpg85167

[0177] The following Table 30 shows examples of using single (set_num = 1) or multiple (set_num > 1) classifiers within a frame in AVS. Note that the syntax notation may be mapped to the notation used above. JPEG0007684411000038.jpg250170

[0178] When each region is combined with FIG. 25 or FIG. 27 that has its own set, the syntax example may include a region on / off control flag (picture_ccsao_lcu_control_flag[compIdx][setIdx]) as shown in Table 31 below. JPEG0007684411000039.jpg91169

[0179] In one embodiment, hereinafter, the extension of the intra and inter post-prediction SAO filters will be further described. In one embodiment, the SAO classification method disclosed in the present disclosure can be used as a post-prediction filter, and the prediction can be used as other prediction tools such as intra, inter, or intra-block copy. FIG. 29 is a block diagram showing the use of the SAO classification method disclosed in the present disclosure according to an embodiment of the present disclosure as a post-prediction filter.

[0180] In one embodiment, each classifier is selected for each of the Y, U, and V components. For each component prediction sample, it is first classified and the corresponding offset is added. For example, each component may use the current sample and adjacent samples for classification. As shown in Table 32 below, Y uses the current Y sample and adjacent Y samples, and U / V uses the current U / V sample for classification. FIG. 30 is a block diagram showing that for the post-prediction SAO filter according to an embodiment of the present disclosure, each component uses the current sample and adjacent samples for classification. JPEG0007684411000040.jpg51169

[0181] In one embodiment, the subdivided prediction samples (Ypred’, Upred’, Vpred’) are updated by adding the corresponding classification offsets, and then used for intra, inter, or other predictions.

[0182] Ypred’ = clip3(0, (1 << bit_depth)-1, Ypred + h_Y[i])

[0183] Upred’ = clip3(0, (1 << bit_depth)-1, Upred + h_U[i])

[0184] Vpred’ = clip3(0, (1 << bit_depth)-1, Vpred + h_V[i])

[0185] In one embodiment, for the saturation U and V components, in addition to the current saturation components, the cross-component (Y) can be used for further offset classification. Additional cross-component offsets (h’_U, h’_V) can be added to the current component offsets (h_U, h_V), for example, as shown in Table 33 below. JPEG0007684411000041.jpg67169

[0186] In one embodiment, the subdivided prediction samples (Upred’’, Vpred’’) are updated by adding the corresponding class offsets and then used for intra, inter, or other predictions.

[0187] Upred’’ = clip3(0, (1 << bit_depth)-1, Upred’ + h’_U[i])

[0188] Vpred’’ = clip3(0, (1 << bit_depth)-1, Vpred’ + h’_V[i])

[0189] In one embodiment, intra prediction and inter prediction may use different SAO filter offsets.

[0190] FIG. 31 is a flowchart showing an exemplary process 3100 for decoding a video signal using cross-component correlation according to an embodiment of the present disclosure.

[0191] A video decoder 30 (as shown in FIG. 3) receives (3110) an image frame including a first component and a second component from the video signal.

[0192] Video decoder 30 determines a classifier for a first component based on a first set of one or more samples of a second component associated with each sample of the first component, where the first component is a luminance component and the second component is a first chrominance component (3120).

[0193] Video decoder 30 determines a sample offset for each sample of the first component according to this classifier (3130).

[0194] Video decoder 30 changes the value of each sample of the first component based on the determined sample offset (3140).

[0195] In some embodiments, the classifier for the first component is further determined based on a second set of one or more samples of the first component associated with each sample of the first component (3150).

[0196] In some embodiments, the image frame further includes a third component, and the classifier for the first component is further determined based on a third set of one or more samples of the third component associated with each sample of the first component, where the third component is a second chrominance component (3160).

[0197] In some embodiments, before determining the classifier for the first component, each sample of the first component is reconstructed by an in-loop filter and / or a first set of one or more samples of the second component is reconstructed by an in-loop filter, where the in-loop filter is a deblocking filter (DBF) or a sample adaptive offset (SAO).

[0198] In one embodiment, a first set of one or more samples of a second component associated with each sample of a first component is selected from one or more juxtapositions and adjacent samples of the second component relative to each sample of the first component.

[0199] In one embodiment, a second set of one or more samples of a first component associated with each sample of a first component is selected from the current sample of the first component and one or more of the adjacent samples relative to each sample of the first component.

[0200] In one embodiment, the class index of a classifier for each sample of a first component is derived by dividing a first dynamic range of values of a first set of one or more samples of a second component according to a first sub-classifier of the classifier into a first number of bands, dividing a second dynamic range of values of a second set of one or more samples of the first component into a second number of bands according to a second sub-classifier of the classifier, dividing a third dynamic range of values of a third set of one or more samples of the third component into a third number of bands according to a third sub-classifier of the classifier, and combining the first sub-classifier, the second sub-classifier, and the third sub-classifier.

[0201] In one embodiment, the class index of a classifier for each sample of a first component is derived as follows. bandY = (candY * bandNumY) >> BitDepth; bandU = (candU * bandNumU) >> BitDepth; bandV = (candV * bandNumV) >> BitDepth; classIdx = bandY * bandNumU * bandNumV+ bandU * bandNumV+ bandV; Here, classIdx is the class index of the classifier for each sample of the first component, bandNumY is the number of divided bands of the dynamic range of the first component, bandNumU is the number of divided bands of the dynamic range of the second component, bandNumV is the number of divided bands of the dynamic range of the third component, candY is a value based on one or more samples of the 1 first set of the 2 component, candU is a value based on one or more samples of the 2 first set of the 1 component, candV is a value based on one or more samples of the third set of the third component, and BitDepth is the bit depth of the video signal.

[0202] In one embodiment, the layout of the first set of one or more samples of the second component is symmetric with respect to the corresponding samples of the first component.

[0203] In one embodiment, the definition of the first set of one or more samples of the second component for determining the classifier and the second set of one or more samples of the first component associated with each sample of the first component are switched at one or more levels among a sequence parameter set (SPS), an adaptive parameter set (APS), an image parameter set (PPS), an image header (PH), a slice header (SH), a region, a coding tree unit (CTU), a coding unit (CU), and a sub-block level. The definition of the first set of one or more samples of the second component and the second set of one or more samples of the first component associated with each sample of the first component are the defined relative positions of the selected samples of the second component and the first component, and including juxtaposed positions and / or adjacent positions the including classification using a single classifier or classification using a combination of multiple classifiersIncludes one or more of a classification method and a bitmask definition for values based on a first set and a second set of one or more samples.

[0204] In certain embodiments, a first band number offset is applied to a first value based on a first set of one or more samples of a second component to obtain a first classification index for each sample of the first component, and a second band number offset is applied to a second value based on a second set of one or more samples of the first component to obtain a second classification index for each sample of the first component.

[0205] In certain embodiments, video decoder 30 further determines a sample offset for each sample of the second component. In certain embodiments, the sample offset for each sample of the first component is determined within a first maximum offset range, the sample offset for each sample of the second component is determined within a second maximum offset range, and the first maximum offset range and the second maximum offset range are fixed or signaled at one or more levels of a sequence parameter set (SPS), an adaptive parameter set (APS), a picture parameter set (PPS), a picture header (PH), and a slice header (SH).

[0206] In certain embodiments, a first sample shape constraint is applied to a classifier based on a first set of one or more samples of a second component associated with each sample of a first component based on a first video color format. And a second sample shape constraint is applied to a classifier based on a first set of one or more samples of a second component associated with each sample of the first component based on a second video color format.

[0207] In one embodiment, if samples of the first set of one or more samples of the second component and the second set of one or more samples of the first component are within the image frame of the corresponding sample of the first component, the value of the corresponding sample of the first component is changed based on the determined sample offset.

[0208] In one embodiment, if samples of the first set of one or more samples of the second component and the second set of one or more samples of the first component are outside the image frame of the corresponding sample of the first component, samples outside the image frame are derived by duplication or mirror padding by copying samples within the image frame from the first set of one or more samples of the second component and the second set of one or more samples of the first component.

[0209] FIG. 32 shows a computing environment 3210 connected to a user interface 3250. The computing environment 3210 may be part of a data processing server. The computing environment 3210 includes a processor 3220, a memory 3230, and an input / output (I / O) interface 3240.

[0210] The processor 3220 typically controls the overall operation of the computing environment 3210, such as operations related to display, data collection, data communication, and image processing. The processor 3220 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Further, the processor 3220 may include one or more modules for facilitating interaction between the processor 3220 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip device, a graphics processing unit (GPU), or the like.

[0211] Memory 3230 is configured to store various types of data to support the operation of computing environment 3210. Memory 3230 may include predetermined software 3232. Examples of such data include instructions for any application or method for operating on computing environment 3210, video data sets, image data, and the like. Memory 3230 can be realized by using any type of volatile or non-volatile memory device such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk, or a combination thereof.

[0212] I / O interface 3240 provides an interface between processor 3220 and peripheral interface modules such as a keyboard, click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a scan start button, and a scan stop button. I / O interface 3240 may be connected to an encoder and a decoder.

[0213] In one embodiment, there is also provided a non-transitory computer-readable storage medium including a plurality of programs executable by a processor 3220 in a computing environment 3210, such as a memory 3230, to execute the above-described method. Alternatively, the non-transitory computer-readable storage medium can store a bitstream or data stream including encoded video information (e.g., video information including one or more syntax elements) generated by an encoder (e.g., video encoder 20 of FIG. 2) using, for example, the above-described encoding method, for use by a decoder (e.g., video decoder 30 of FIG. 3) when decoding video data. The non-transitory computer-readable storage medium can be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0214] In one embodiment, there is also provided a computing device including one or more processors (e.g., processor 3220), and a non-transitory computer-readable storage medium or memory 3230 stores a plurality of programs executable by the one or more processors, and the one or more processors are configured to implement the above-described method when executing the plurality of programs.

[0215] In one embodiment, there is also provided a computer program product including a plurality of programs executable by a processor 3220 in a computer environment 3210, such as in a memory 3230, to execute the above-described method. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0216] In one embodiment, the computing environment 3210 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to execute the above-described method.

[0217] Further embodiments include various subsets of the above-described embodiments that are combined or rearranged in various other embodiments.

[0218] In one or more examples, the functions described above are implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions are stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a processing unit in hardware. A computer-readable medium can include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, a computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium.

[0219] The terms used herein to describe the embodiments are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "one", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used herein means and includes any and all possible combinations of one or more of the associated listed items. The term "comprising" used herein indicates the presence of the described features, elements, and / or components, but does not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0220] Here, it should also be understood that the use of terms such as first, second, etc. to describe various elements does not limit these elements by these terms. These terms are used only to distinguish one element from another. For example, unless departing from the scope of the embodiments, the first electrode may be referred to as the second electrode, and similarly, the second electrode may be referred to as the first electrode. The first electrode and the second electrode are both electrodes, but not the same electrode.

[0221] In this specification, references in the singular or plural forms of "one example", "an example", "exemplary example", etc. mean that one or more specific features, structures, or characteristics described in relation to the example are included in at least one example of the present disclosure. Thus, appearances in the singular or plural forms of terms such as "in one example" or "in an example" throughout this specification do not necessarily refer to the same example. Furthermore, the specific features, structures, or characteristics in one or more examples can include being combined in any suitable manner.

[0222] The description of the present application is presented for purposes of illustration and explanation and is not limited to the invention in its comprehensive or disclosed form. Various changes, modifications, and alternative implementations will be apparent to those skilled in the art having obtained the teachings presented in the foregoing description and the associated drawings. The embodiments are selected and described in order to best explain the principles of the invention, its practical application, to enable those skilled in the art to understand the invention for various implementations, and to best utilize the principles and various implementations on which the basis for various modifications suitable for a particular use. Therefore, it should be understood that the claims are not limited to the specific examples of the disclosed implementations, and that changes and other implementations are also included within the scope of the appended claims.

Claims

1. Receiving an image frame including a first component and a second component from a video signal; Determining a classifier for the first component based on a measurement of characteristics of a first set of one or more samples of the second component juxtaposed or adjacent to each sample of the first component; Determining a sample offset for each sample of the first component according to the classifier; Changing the value of each sample of the first component based on the determined sample offset; comprising; A method for decoding a video signal, wherein the first component is a luminance component and the second component is a first chrominance component.

2. The method according to claim 1, wherein the classifier for the first component is further determined based on a second set of one or more samples of the first component associated with each sample of the first component.

3. The image frame further includes a third component, The classifier for the first component is further determined based on a third set of one or more samples of the third component associated with each sample of the first component, The method according to claim 1, wherein the third component is a second chrominance component.

4. Before determining the classifier for the first component, each sample of the first component is reconstructed by an in-loop filter, and a first set of one or more samples of the second component is reconstructed by an in-loop filter, The method according to claim 1, wherein the in-loop filter is a deblocking filter (DBF) or a sample adaptive offset (SAO).

5. The method according to claim 2, wherein the second set of one or more samples of the first component associated with each sample of the first component is selected from among the current sample of the first component and one or more of the adjacent samples for each sample of the first component.

6. The image frame further includes a third component, The classifier for the first component is further determined based on one or more samples of a third set of a third component associated with each sample of the first component, the third component is a second chroma component, the class index of the classifier for each sample of the first component is dividing a first dynamic range of values of a first set of one or more samples of the second component into a first number of bands according to a first sub-classifier of the classifier, dividing a second dynamic range of values of a second set of one or more samples of the first component into a second number of bands according to a second sub-classifier of the classifier, dividing a third dynamic range of values of a third set of one or more samples of the third component into a third number of bands according to a third sub-classifier of the classifier, and combining the first sub-classifier, the second sub-classifier, and the third sub-classifier, the method according to claim 2 derived thereby. **Claim 7** the image frame further includes a third component, the classifier for the first component is further determined based on one or more samples of a third set of a third component associated with each sample of the first component, the third component is a second chroma component, the classifier index of the classifier for each sample of the first component is bandY = (candY * bandNumY) >> BitDepth; bandU = (candU * bandNumU) >> BitDepth; bandV = (candV * bandNumV) >> BitDepth; classIdx = bandY * bandNumU * bandNumV+ bandU * bandNumV+ bandV derived thereby, Here, classIdx is a classifier for each sample of the first component, bandNumY is the number of divided bands of the dynamic range of the first component, bandNumU is the number of divided bands of the dynamic range of the second component, bandNumV is the number of divided bands of the dynamic range of the third component, candY is a value based on a second set of one or more samples of the first component, candU is a value based on a first set of one or more samples of the second component, candV is a value based on a third set of one or more samples of the third component, and BitDepth is the bit depth of the video signal. The method according to claim 2.

8. The layout of the first set of one or more samples of the second component is symmetric with respect to the corresponding samples of the first component. The method according to claim 1.

9. The definition of the first set of one or more samples of the second component and the second set of one or more samples of the first component related to each sample of the first component for determining the classifier is switched at one or more levels among a sequence parameter set (SPS), an adaptive parameter set (APS), an image parameter set (PPS), an image header (PH), a slice header (SH), a region, a coding tree unit (CTU), a coding unit (CU), and a sub-block level. The definition of the first set of one or more samples of the second component and the second set of one or more samples of the first component related to each sample of the first component includes the second component and the selected samples of the first component at relative positions including defined juxtaposed positions and / or adjacent positions related to each sample of the first component, a classification method using a single classifier or a combination of multiple classifiers, and a bit mask definition for values based on the first set and the second set of one or more samples. The method according to claim 2.

10. The first band number offset is applied to a first value based on a first set of one or more samples of the second component to obtain a first classification index for each sample of the first component, and a second band number offset is applied to a second value based on a second set of one or more samples of the first component to obtain a second classification index for each sample of the first component. The method according to claim 2.

11. Further comprising determining a sample offset for each sample of the second component, The sample offset for each sample of the first component is determined within a first maximum offset range, The sample offset for each sample of the second component is determined within a second maximum offset range, The first maximum offset range and the second maximum offset range are fixed or signaled at one or more levels of a sequence parameter set (SPS), an adaptive parameter set (APS), an image parameter set (PPS), an image header (PH), and a slice header (SH). The method according to claim 2.

12. A first sample shape constraint is applied to a classifier based on a first set of one or more samples of the second component associated with each sample of the first component based on a first video color format, and a second sample shape constraint is applied to a classifier based on a first set of one or more samples of the second component associated with each sample of the first component based on a second video color format. The method according to claim 2.

13. When samples of the first set of one or more samples of the second component and the second set of one or more samples of the first component are within the image frame of the corresponding sample of the first component, changing the value of the corresponding sample of the first component based on the determined sample offset. The method according to claim 2.

14. If one or more samples of the first set of one or more samples of the second component and one or more samples of the second set of one or more samples of the first component are outside the image frame of the corresponding samples of the first component, the samples within the image frame from the first set of one or more samples of the second component and the second set of one or more samples of the first component are copied to derive the samples outside the image frame by overlapping or mirror padding. The method according to claim 2.

15. One or more processing units, A memory connected to the one or more processing units, A plurality of programs stored in the memory, Including, When the plurality of programs are executed by the one or more processing units, the electronic device is caused to execute the method according to any one of claims 1 to 14. Electronic device.

16. Performing a method for video encoding to generate a bitstream, Storing the bitstream, Including, The bitstream is decoded by the method according to any one of claims 1 to 14, The method for video encoding is, Receiving an image frame including a first component and a second component, Determining a classifier for the first component based on the measurement of the characteristics of one or more samples of the second component juxtaposed or adjacent to each sample of the first component, Determining a sample offset for each sample of the first component according to the classifier, Changing the value of each sample of the first component based on the determined sample offset, Including, A method for storing a bitstream, wherein the first component is a luminance component and the second component is a first chrominance component.

17. Performing a method for video encoding to generate a bitstream, Transmitting the bitstream, Including, The bitstream is decoded by the method according to any one of claims 1 to 14, The method for video encoding is, Receiving an image frame including a first component and a second component; Determining a classifier for the first component based on measurements of characteristics of a first set of one or more samples of the second component juxtaposed or adjacent to each sample of the first component; Determining a sample offset for each sample of the first component according to the classifier; Changing the value of each sample of the first component based on the determined sample offset; comprising; A method for transmitting a bitstream, wherein the first component is a luminance component and the second component is a first chrominance component. **Claim 18** A computer program storing instructions, wherein the instructions, when executed by a processor, cause the processor to execute the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Method and apparatus for optimizing the encoding / decoding of compensation offsets for a set of reconstructed image samples.

    JP2014534762A

  • Cross Component Filter

    JP2019525679A

  • Cross plane filtering for chroma signal enhancement in video coding

    JP2020036353A

  • Systems and methods for reducing a reconstruction error in video coding based on a cross-component correlation

    WO2020262396A1

  • Chroma coding enhancement in cross-component sample adaptive offset with virtual boundary

    WO2022093992A1