Coding enhancements in cross-component sample adaptive offsets
By employing Cross-Component Sample Adaptive Offset (CCSAO) to leverage cross-component relationships, the method addresses the challenge of efficiently encoding high-resolution video data, thereby enhancing coding efficiency and image quality.
Patent Information
- Application Number
- JP2023557382
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-22
- Filing Date
- 2022-03-16
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2042-03-16
AI Technical Summary
Existing video encoding technologies face challenges in efficiently encoding and decoding high-resolution video data, such as 4Kx2K or 8Kx4K, due to the exponential increase in video data, which affects image quality and encoding efficiency.
The implementation of methods and apparatus for improving coding efficiency by searching for cross-component relationships between luma and chroma components, specifically through a Cross-Component Sample Adaptive Offset (CCSAO) process that modifies chroma sample offsets based on luma component samples.
This approach enhances coding efficiency by effectively modifying chroma sample offsets, leading to improved compression performance and maintained image quality, even at high resolutions.
Smart Images

Figure 0007673225000046 
Figure 0007673225000047 
Figure 0007673225000048
Abstract
Description
[Technical field]
[0001] Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 200,626, entitled "Cross-component Sample Adaptive Offset," filed March 18, 2021, and U.S. Provisional Patent Application No. 63 / 164,459, entitled "Cross-component Sample Adaptive Offset," filed March 22, 2021, which are incorporated by reference in their entireties.
[0002] This application relates generally to video encoding and compression, and more specifically to methods and apparatus for improving both luma and chroma encoding efficiency. [Background technology]
[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Electronic devices transmit, receive, encode, decode, and / or store digital video data by implementing video compression / decompression standards. Some well-known video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), jointly developed by ISO / IEC MPEG and ITU-TVCEG. AOMedia Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to the earlier VP9 standard. Audio Video Coding (AVS), which refers to digital audio and video compression standards, is another series of video compression standards developed by the Audio and Video Coding Standard workgroup.
[0004] Video compression typically involves performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in video data. For block-based video coding, a video frame is divided into one or more slices, each having multiple video blocks, which may also be referred to as a coding tree unit (CTU). Each CTU may contain one coding unit (CU) or may be recursively divided into smaller CUs until a predetermined minimum CU size is reached. Each CU (also referred to as a leaf CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU may be coded in either intra mode, inter mode, or IBC mode. Video blocks in an intra-coded (I) slice of a video frame are encoded using spatial prediction relative to reference samples in neighboring blocks in the same video frame. Video blocks within an inter-coded (P or B) slice of a video frame may use spatial prediction relative to reference samples in adjacent blocks in the same video frame, or temporal prediction relative to reference samples in other previous and / or future reference video frames.
[0005] Spatial or temporal prediction based on previously encoded reference blocks, e.g., neighboring blocks, results in a prediction block for the current video block to be encoded. The process of finding the reference block may be accomplished by a block matching algorithm. Residual data representing pixel differences between the current block to be encoded and the prediction block is referred to as a residual block or prediction error. Inter-coded blocks are encoded according to a motion vector that points to a reference block in a reference frame that forms the prediction block, and the residual block. The process of determining the motion vector is typically referred to as motion estimation. Intra-coded blocks are encoded according to an intra prediction mode and a residual block. For further compression, the residual block may be transformed from the pixel domain to a transform domain, e.g., to the frequency domain, resulting in residual transform coefficients, which may then be quantized. The quantized transform coefficients, initially arranged in a two-dimensional array, are scanned to produce a one-dimensional vector of transform coefficients, which may then be entropy encoded into a video bitstream to achieve even further compression.
[0006] The encoded video bitstream is then stored in a computer-readable storage medium (e.g., flash memory) where it can be accessed by another electronic device with digital video capabilities or transmitted directly to the electronic device via wired or wireless connection. The electronic device then performs video decompression (which is the opposite process to video compression described above), for example, by parsing the encoded video bitstream to obtain syntax elements from the bitstream, reconstructing the digital video data from the encoded video bitstream back to its original format based at least in part on the syntax elements obtained from the bitstream, and rendering the reconstructed digital video data on a display of the electronic device.
[0007] As digital video quality progresses from high definition to 4Kx2K or even 8Kx4K, the amount of video data to be encoded / decoded increases exponentially. This is a constant challenge as to how to more efficiently encode / decode video data while maintaining the image quality of the decoded video data. Summary of the Invention
[0008] This application relates to encoding and decoding video data, and more particularly, describes implementations of methods and apparatus for improving the encoding efficiency of both luma and chroma components, including improving the encoding efficiency by exploring cross-component relationships between the luma and chroma components.
[0009] According to a first aspect of the present application, a method of decoding a video signal includes: receiving from the video signal a picture frame comprising a first component and a second component, determining a classifier for the second component from a set of one or more samples of the first component associated with respective samples of the second component, determining whether to modify values of respective samples of the second component within a region of the picture frame according to the classifier, determining sample offsets of respective samples of the second component according to the classifier in response to determining to modify values of respective samples of the second component within the region according to the classifier, and modifying values of respective samples of the second component based on the determined sample offsets. In some embodiments, the region is formed by dividing the picture frame.
[0010] According to a second aspect of the present application, an electronic device includes one or more processing units, a memory, and a number of programs stored in the memory, the programs, when executed by the one or more processing units, causing the electronic device to perform a method for encoding a video signal as described above.
[0011] According to a third aspect of the present application, a non-transitory computer-readable storage medium stores a number of programs for execution by an electronic device having one or more processing units, the programs, when executed by the one or more processing units, causing the electronic device to perform a method for encoding a video signal as described above.
[0012] According to a fourth aspect of the present application, a computer-readable storage medium stores a bitstream comprising video information generated by a method for video encoding as described above.
[0013] It is to be understood that both the foregoing general description and the following detailed description are merely exemplary and are not restrictive of the present disclosure.
[0014] The accompanying drawings, which are included to provide a further understanding of the implementations, and which are incorporated in and constitute a part of this specification, illustrate the described implementations and, together with the description, serve to explain the underlying principles. Like reference numerals refer to corresponding parts. [Brief description of the drawings]
[0015] [Figure 1] FIG. 1 is a block diagram illustrating an example video encoding and decoding system according to some implementations of the present disclosure. [Diagram 2] FIG. 2 is a block diagram illustrating an example video encoder according to some implementations of this disclosure. [Diagram 3] FIG. 3 is a block diagram illustrating an example video decoder according to some implementations of the disclosure. [Figure 4A]FIG. 4A is a block diagram illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes in accordance with some implementations of this disclosure. [Figure 4B] FIG. 4B is a block diagram illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes in accordance with some implementations of this disclosure. [Figure 4C] FIG. 4C is a block diagram illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes in accordance with some implementations of this disclosure. [Figure 4D] FIG. 4D is a block diagram illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes in accordance with some implementations of this disclosure. [Figure 4E] FIG. 4E is a block diagram illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes in accordance with some implementations of this disclosure. [Diagram 5] FIG. 5 is a block diagram illustrating four gradient patterns used in Sample Adaptive Offset (SAO) in accordance with some implementations of the present disclosure. [Figure 6A] FIG. 6A is a block diagram illustrating a system and process of a CCSAO applied to chroma samples and using DBF Y as input according to some implementations of the present disclosure. [Figure 6B] FIG. 6B is a block diagram illustrating a system and process of CCSAO applied to luma and chroma samples and using DBF Y / Cb / Cr as input according to some implementations of the present disclosure. [Figure 6C] FIG. 6C is a block diagram illustrating systems and processes of a CCSAO that can operate independently according to some implementations of the present disclosure. [Figure 6D]FIG. 6D is a block diagram illustrating a system and process of a CCSAO that can be applied recursively (2 or N times) with the same offset or different offsets according to some implementations of the present disclosure. [Figure 6E] FIG. 6E is a block diagram illustrating a system and process of a CCSAO applied in parallel with an Enhanced Sample Adaptive Offset (ESAO) in the AVS standard according to some implementations of the present disclosure. [Figure 6F] FIG. 6F is a block diagram illustrating a system and process of a CCSAO applied after an SAO according to some implementations of the present disclosure. [Figure 6G] FIG. 6G is a block diagram illustrating that the systems and processes of the CCSAO can operate independently without the CCALF, according to some implementations of the present disclosure. [Figure 6H] FIG. 6H is a block diagram illustrating a system and process of a CCSAO applied in parallel with a Cross-Component Adaptive Loop Filter (CCALF) according to some implementations of the present disclosure. [Figure 7] FIG. 7 is a block diagram illustrating a sample process using a CCSAO according to some implementations of the present disclosure. [Figure 8] FIG. 8 is a block diagram illustrating the CCSAO process being interleaved into vertical and horizontal deblocking filters (DBFs) in accordance with some implementations of the present disclosure. [Figure 9] FIG. 9 is a flowchart illustrating an example process for decoding a video signal using cross-component correlation according to some implementations of the disclosure. [Figure 10A] FIG. 10A is a block diagram illustrating a classifier that uses different luma (or chroma) sample positions for C0 classification, according to some implementations of this disclosure. [Figure 10B] FIG. 10B illustrates some examples of different shapes of luma candidates according to some implementations of this disclosure. [Figure 11] FIG. 11 is a block diagram of a sample process illustrating that all collocated and adjacent luma / chroma samples may be fed into CCSAO classification according to some implementations of the present disclosure. [Figure 12] FIG. 12 illustrates an example classifier by replacing collocated luma sample values with values obtained by weighting collocated and adjacent luma samples according to some implementations of this disclosure. [Figure 13A] FIG. 13A is a block diagram illustrating that, in accordance with some implementations of the present disclosure, CCSAO is not applied to the current chroma (luma) sample if any of the co-located and adjacent luma (chroma) samples used in classification are outside the current picture. [Figure 13B] FIG. 13B is a block diagram illustrating that CCSAO is applied to the current luma or chroma sample if any of the co-located and adjacent luma or chroma samples used in classification are outside the current picture, according to some implementations of the present disclosure. [Figure 14] FIG. 14 is a block diagram illustrating that, in accordance with some implementations of the present disclosure, CCSAO is not applied to a current chroma sample if the corresponding selected co-located or adjacent luma sample used for classification is outside a virtual space defined by a virtual boundary (VB). [Figure 15] FIG. 15 illustrates repetitive or mirror padding applied to luma samples that are outside the virtual boundary according to some implementations of this disclosure. [Figure 16] FIG. 16 shows that, according to some implementations of this disclosure, if all nine collocated adjacent luma samples are used for classification, then one additional luma line buffer is needed. [Figure 17]FIG. 17 shows a diagram of an AVS where a nine luma candidate CCSAO across a VB may increase two additional luma line buffers in accordance with some implementations of the present disclosure. [Figure 18A] FIG. 18A shows a diagram of a VVC in which a nine luma candidate CCSAO across a VB may increase one additional luma line buffer in accordance with some implementations of the present disclosure. [Figure 18B] FIG. 18B shows a diagram when collocated and adjacent chroma samples are used to classify the current luma sample, and the selected chroma candidate may span VB, requiring additional chroma line buffers according to some implementations of the present disclosure. [Figure 19A] FIG. 19A shows that in AVS and VVC, in some implementations of this disclosure, if any of the luma candidates of a chroma sample crosses VB (is outside the current chroma sample VB), the CCSAO is disabled for the chroma sample. [Figure 19B] FIG. 19B shows that in AVS and VVC, in some implementations of this disclosure, if any of the luma candidates of a chroma sample crosses VB (is outside the current chroma sample VB), the CCSAO is disabled for the chroma sample. [Figure 19C] FIG. 19C shows that in AVS and VVC, in some implementations of this disclosure, if any of the luma candidates of a chroma sample crosses VB (is outside the current chroma sample VB), the CCSAO is disabled for the chroma sample. [Figure 20A] FIG. 20A shows that in AVS and VVC, if any of the luma candidates of a chroma sample crosses VB (is outside the current chroma sample VB), CCSAO is enabled with repeated padding of the chroma sample, according to some implementations of the present disclosure. [Figure 20B]FIG. 20B shows that in AVS and VVC, if any of the luma candidates of a chroma sample crosses VB (is outside the current chroma sample VB), CCSAO is enabled with repeated padding of the chroma sample, according to some implementations of this disclosure. [Figure 20C] FIG. 20C shows that in AVS and VVC, if any of the luma candidates of a chroma sample crosses VB (is outside the current chroma sample VB), CCSAO is enabled with repeated padding of the chroma sample, according to some implementations of this disclosure. [Figure 21A] FIG. 21A shows that in AVS and VVC, if any of the luma candidates of a chroma sample crosses VB (is outside the current chroma sample VB), CCSAO is enabled with mirror padding of the chroma sample, according to some implementations of this disclosure. [Figure 21B] FIG. 21B shows that in AVS and VVC, if any of the luma candidates of a chroma sample crosses VB (is outside the current chroma sample VB), CCSAO is enabled with mirror padding of the chroma sample, according to some implementations of this disclosure. [Figure 21C] FIG. 21C shows that in AVS and VVC, if any of the luma candidates of a chroma sample crosses VB (is outside the current chroma sample VB), CCSAO is enabled with mirror padding of the chroma sample, according to some implementations of this disclosure. [Figure 22A] FIG. 22A illustrates CCSAO being enabled with two-sided symmetric padding for different CCSAO sample shapes according to some implementations of the present disclosure. [Figure 22B] FIG. 22B illustrates that CCSAO is enabled with two-sided symmetric padding for different CCSAO sample shapes according to some implementations of the present disclosure. [Figure 23]FIG. 23 illustrates the limitation of using a limited number of luma candidates for classification, according to some implementations of this disclosure. [Figure 24] FIG. 24 illustrates that the CCSAO application region is not aligned to a coding tree block (CTB) / coding tree unit (CTU) boundary in accordance with some implementations of the present disclosure. [Diagram 25] FIG. 25 illustrates that the CCSAO application region frame partition may be fixed using CCSAO parameters according to some implementations of the present disclosure. [Figure 26] FIG. 26 illustrates that in some implementations of the present disclosure, the CCSAO application region may be a binary-tree (BT) / quad-tree (QT) / ternary-tree (TT) split from the frame / slice / CTB level. [Figure 27] FIG. 27 is a block diagram illustrating multiple classifiers used and switched at different levels within a picture frame in accordance with some implementations of the present disclosure. [Figure 28] FIG. 28 is a block diagram illustrating that CCSAO application region partitions are dynamic and can be switched at the picture level, according to some implementations of the present disclosure. [Figure 29] FIG. 29 is a diagram illustrating that a CCSAO classifier can take into account current component encoding information or cross-component encoding information, according to some implementations of the present disclosure. [Diagram 30] FIG. 30 is a block diagram illustrating the SAO classification method disclosed in this disclosure functioning as a post-prediction filter according to some implementations of the present disclosure. [Diagram 31] FIG. 31 is a block diagram illustrating that for a post-predictive SAO filter, each component can use the current sample and neighboring samples for classification, according to some implementations of the present disclosure. [Diagram 32] FIG. 32 is a flowchart illustrating an example process for decoding a video signal using cross-component correlation according to some implementations of the disclosure. [Diagram 33] FIG. 33 illustrates a computing environment coupled with a user interface in accordance with some implementations of the disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives can be used without departing from the scope of the claims, and that the subject matter can be practiced without these specific details. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.
[0017] The first generation of AVS standards includes the People's Republic of China national standard "Information Technology, Advanced Audio Video Coding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video" (known as AVS+), which can bring about 50% bitrate savings with the same perceptual quality compared to the MPEG-2 standard. The second generation of AVS standards includes a series of People's Republic of China national standards "Information Technology, Efficient Multimedia Coding" (known as AVS2), which mainly targets the transmission of further HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. Meanwhile, the AVS2 standard video part has been submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for applications. The AVS3 standard is a new generation video coding standard for UHD video applications that aims to exceed the coding efficiency of the latest international standard HEVC, which saves about 30% bitrate than the HEVC standard. In March 2019, at the 68th AVS Conference, the AVS3-P2 baseline was completed, which achieved approximately 30% bitrate savings over the HEVC standard. Currently, one reference software, called the high performance model (HPM), is maintained by the AVS group to demonstrate the reference implementation of the AVS3 standard. Like HEVC, the AVS3 standard is built on a block-based hybrid video coding framework.
[0018] 1 is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that is subsequently decoded by a destination device 14. Source device 12 and destination device 14 may comprise any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0019] In some implementations, the destination device 14 may receive the encoded video data to be decoded via a link 16. The link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may include a communication medium that allows the source device 12 to transmit the encoded video data directly to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device 14. The communication medium may include any wireless communication medium or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from the source device 12 to the destination device 14.
[0020] In some other implementations, the encoded video data may be sent from the output interface 22 to the storage device 32. The encoded video data in the storage device 32 may then be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that may hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or download. The file server may be any type of computer that may store the encoded video data and transmit the encoded video data to the destination device 14. Exemplary file servers include a web server (e.g., for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. Destination device 14 may access the encoded video data over any standard data connection suitable for accessing encoded video data stored on a file server, including a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both. The transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
[0021] As shown in FIG. 1, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include sources such as a video capture device, such as a video camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources. As an example, if video source 18 is a video camera of a security surveillance system, source device 12 and destination device 14 may form a camera phone or video phone. However, the implementations described herein may be applicable to video encoding in general and may be applied to wireless and / or wired applications.
[0022] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored in storage device 32 for later access by destination device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.
[0023] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or modem and may receive encoded video data over link 16. The encoded video data communicated over link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data that is transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0024] In some implementations, destination device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0025] Video encoder 20 and video decoder 30 may operate according to a proprietary or industry standard, such as, for example, VVC, HEVC, MPEG-4, Part 10, High Efficiency Video Coding (AVC), AVS, or an extension of such a standard. It should be understood that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is also generally contemplated that video decoder 30 of destination device 14 may be configured to decode video data according to any of these current or future standards.
[0026] Each of the video encoder 20 and the video decoder 30 may be implemented as any of a variety of suitable encoder circuits, such as, for example, one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic elements, software, hardware, firmware, or any combination thereof. If implemented partially in software, the electronic device may store instructions for the software on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.
[0027] 2 is a block diagram illustrating an example video encoder 20 according to some implementations described herein. Video encoder 20 may perform intra-predictive and inter-predictive coding of video blocks within a video frame. Intra-predictive coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-predictive coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence.
[0028] As shown in FIG. 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partition unit 45, an intra-prediction processing unit 46, and a block copy (BC) unit 48. In some implementations, the video encoder 20 also includes an inverse quantization unit 58 for video block reconstruction, an inverse transform processing unit 60, and an adder 62. An in-loop filter 63, such as a deblocking filter, may be disposed between the adder 62 and the DPB 64 to filter block boundaries and remove blocky artifacts from the reconstructed video. In addition to the deblocking filter, another in-loop filter 63 may also be used to filter the output of the adder 62. Further in-loop filtering 63, such as a sample adaptive offset (SAO) and an adaptive in-loop filter (ALF), may be applied to the reconstructed CU before it is placed in a reference picture store and used as a reference for encoding future video blocks. Video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided among one or more of the illustrated fixed or programmable hardware units.
[0029] Video data memory 40 may store video data to be encoded by components of video encoder 20. The video data in video data memory 40 may be obtained, for example, from video source 18. DPB 64 is a buffer that stores reference video data used in encoding the video data by video encoder 20 (e.g., in intra-predictive or inter-predictive coding modes). Video data memory 40 and DPB 64 may be formed by any of a variety of memory devices. In various examples, video data memory 40 may be on-chip with other components of video encoder 20 or off-chip relative to those components.
[0030] As shown in FIG. 2, after receiving the video data, partition unit 45 in prediction processing unit 41 partitions the video data into video blocks. This partitioning may also include partitioning the video frame into slices, tiles, or other larger coding units (CUs) according to a predetermined partitioning structure, such as a quadtree structure, associated with the video data. The video frame may be divided into a number of video blocks (or sets of video blocks referred to as tiles). Prediction processing unit 41 may select one of a number of possible predictive coding modes, such as one of a number of intra-predictive coding modes or one of a number of inter-predictive coding modes, for the current video block based on the error result (e.g., the coding rate and the level of distortion). Prediction processing unit 41 may provide the resulting intra-predictive or inter-predictive coded block to adder 50 to generate a residual block, and to adder 62 to reconstruct the encoded block, which may then be used as part of the reference frame. Prediction processing unit 41 also provides syntax elements, such as motion vectors, intra mode indicators, partition information, and other such syntax information to entropy encoding unit 56.
[0031] To select an appropriate intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-prediction coding of the current video block relative to one or more predictive blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple encoding passes to select an appropriate coding mode for each block of video data, for example.
[0032] In some implementations, motion estimation unit 42 determines an inter prediction mode for a current video frame by generating a motion vector that indicates the displacement of a prediction unit (PU) of a video block in a current video frame relative to a predictive block in a reference video frame according to a predetermined pattern in a sequence of video frames. Motion estimation performed by motion estimation unit 42 is a process of generating motion vectors that estimate the motion of a video block. The motion vector may indicate, for example, the displacement of a PU of a video block in a current video frame or picture relative to a predictive block in a reference frame (or other coding unit) relative to a current block being coded in the current frame (or other coding unit). The predetermined pattern may designate a video frame in the sequence as a P frame or a B frame. Intra BC unit 48 may determine vectors, e.g., block vectors, for intra BC coding in a manner similar to the determination of motion vectors by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine the block vectors.
[0033] A prediction block is a block of a reference frame that is deemed to closely match the PU of the video block to be encoded in terms of pixel differences, which may be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metric. In some implementations, video encoder 20 may calculate values for sub-integer pixel locations of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values for quarter-pixel locations, eighth-pixel locations, or other fractional pixel locations of the reference frame. Thus, motion estimation unit 42 may perform motion search for whole pixel locations and fractional pixel locations and output motion vectors with fractional pixel accuracy.
[0034] Motion estimation unit 42 calculates a motion vector for a PU of a video block in an inter-predictively coded frame by comparing the position of the PU to the position of a predictive block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.
[0035] The motion compensation performed by motion compensation unit 44 may include fetching or generating a predictive block based on the motion vector determined by motion estimation unit 42. Upon receiving the motion vector for the PU of the current video block, motion compensation unit 44 may search for the predictive block pointed to by the motion vector in one of the reference frame lists, obtain the predictive block from DPB 64, and forward the predictive block to summer 50. Summer 50 then forms a residual video block of pixel difference values by subtracting pixel values of the predictive block provided by motion compensation unit 44 from pixel values of the current video block being coded. The pixel difference values forming the residual video block may include a luma difference component or a chroma difference component, or both. Motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frames for use by video decoder 30 in decoding the video blocks of the video frames. The syntax elements may include, for example, syntax elements defining the motion vector used to identify the predictive block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are illustrated separately for conceptual purposes.
[0036] In some implementations, the intra BC unit 48 may generate vectors and fetch predictive blocks in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, except that the predictive block is in the same frame as the current block being coded, and the vectors are referred to as block vectors as opposed to motion vectors. In particular, the intra BC unit 48 may determine an intra prediction mode to use to encode the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, e.g., during separate encoding passes, and test their performance by rate-distortion analysis. The intra BC unit 48 may then select an appropriate intra prediction mode from among the various tested intra prediction modes and use and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate a rate-distortion value using a rate-distortion analysis of the various tested intra prediction modes, and may select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block that was encoded to produce the encoded block, and the bitrate (i.e., number of bits) used to produce the encoded block. Intra BC unit 48 may calculate ratios from the distortions and rates of the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.
[0037] In other examples, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functionality for intra BC prediction according to implementations described herein. In either case, for intra block copying, the predictive block may be a block that is deemed to closely match the block to be coded in terms of pixel differences, which may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metric, and identification of the predictive block may include calculation of values at sub-integer pixel positions.
[0038] Regardless of whether the predictive block is from the same frame via intra prediction or from a different frame via inter prediction, video encoder 20 may form a residual video block by subtracting pixel values of the predictive block from pixel values of the current video block being encoded to form pixel difference values. The pixel difference values that form the residual video block may include both luma component difference and chroma component difference.
[0039] Intra-prediction processing unit 46 may intra-predict the current video block instead of the inter prediction performed by motion estimation unit 42 and motion compensation unit 44, or instead of the intra block copy prediction performed by intra BC unit 48, as described above. In particular, intra-prediction processing unit 46 may determine an intra-prediction mode to use to encode the current block. To that end, intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, e.g., during separate encoding passes, and intra-prediction processing unit 46 (or, in some examples, a mode selection unit) may select an appropriate intra-prediction mode to use from the tested intra-prediction modes. Intra-prediction processing unit 46 may provide information indicative of the selected intra-prediction mode for the block to entropy encoding unit 56. Entropy encoding unit 56 may encode the information indicative of the selected intra-prediction mode into the bitstream.
[0040] After prediction processing unit 41 determines a predictive block for the current video block via either inter- or intra-prediction, summer 50 forms a residual video block by subtracting the predictive block from the current video block. The residual video data in the residual block, which may be included in one or more transform units (TUs), is provided to transform processing unit 52. Transform processing unit 52 converts the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0041] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54, which quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of a matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 may perform the scan.
[0042] Following quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream, e.g., using context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding method or technique. The encoded bitstream may then be transmitted to video decoder 30 or archived to storage device 32 for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 may also entropy encode motion vectors and other syntax elements for the current video frame being encoded.
[0043] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain for generating reference blocks for prediction of other video blocks. As mentioned above, motion compensation unit 44 may generate motion-compensated prediction blocks from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0044] Summer 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block for inter predicting another video block in a subsequent video frame.
[0045] 3 is a block diagram illustrating an example video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-prediction processing unit 84, and an intra-BC unit 85. The video decoder 30 may perform a decoding process that is generally reciprocal to the encoding process described above for the video encoder 20 in connection with FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on a motion vector received from the entropy decoding unit 80, while the intra-prediction unit 84 may generate prediction data based on an intra-prediction mode indicator received from the entropy decoding unit 80.
[0046] In some examples, the units of the video decoder 30 may be assigned to perform implementations of the present disclosure. Also, in some examples, implementations of the present disclosure may be divided among one or more of the units of the video decoder 30. For example, the intra BC unit 85 may perform implementations of the present disclosure alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decode unit 80. In some examples, the video decoder 30 does not include the intra BC unit 85, and the functionality of the intra BC unit 85 may be performed by other components of the prediction processing unit 81, such as the motion compensation unit 82.
[0047] Video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of video decoder 30. The video data stored in video data memory 79 may be obtained, for example, from storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. A decoded picture buffer (DPB) 92 of video decoder 30 stores reference video data for use in decoding of video data by video decoder 30 (e.g., in intra-predictive or inter-predictive coding modes). Video data memory 79 and DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magneto-resistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, video data memory 79 and DPB 92 are shown as two separate components of video decoder 30 in FIG. 3 . However, it will be apparent to one skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, video data memory 79 may be on-chip with other components of video decoder 30 or may be off-chip with respect to those components.
[0048] During the decoding process, video decoder 30 receives an encoded video bitstream, which represents video blocks and associated syntax elements of encoded video frames. Video decoder 30 may receive the syntax elements at the video frame level and / or the video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors and other syntax elements to prediction processing unit 81.
[0049] If the video frame is coded as an intra-predictive (I) frame, or for intra-coded predictive blocks within other types of frames, intra-prediction processing unit 84 of prediction processing unit 81 may generate predictive data for video blocks of the current video frame based on the signaled intra-prediction mode and reference data from previously decoded blocks of the current frame.
[0050] If the video frame is coded as an inter-predictive (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 produces one or more prediction blocks for video blocks of the current video frame based on the motion vectors and other syntax elements received from entropy decode unit 80. Each of the prediction blocks may be produced from a reference frame in one of the reference frame lists. Video decoder 30 may construct the reference frame lists, List 0, and List 1, using a default construction technique based on the reference frames stored in DPB 92.
[0051] In some examples, if the video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 produces a prediction block of the current video block based on the block vectors and other syntax elements received from entropy decoding unit 80. The prediction block may be within the same reconstructed region of the picture as the current video block defined by video encoder 20.
[0052] Motion compensation unit 82 and / or intra BC unit 85 determine prediction information for video blocks of the current video frame by parsing the motion vectors and other syntax elements, and then produce predictive blocks for the current video block that are decoded using the prediction information. For example, motion compensation unit 82 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) used to encode the video blocks of the video frame, an inter prediction frame type (e.g., B or P), configuration information for one or more of the reference frame lists for the frame, a motion vector for each inter predictive encoded video block of the frame, an inter prediction state for each inter predictive encoded video block of the frame, and other information for decoding video blocks in the current video frame.
[0053] Similarly, intra BC unit 85 may use some of the received syntax elements, such as flags, to determine that the current video block was predicted using intra BC mode, which configuration information of the video blocks of the frame are within the reconstruction region and should be stored in DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction states for each intra BC predicted video block of the frame, and other information for decoding video blocks in the current video frame.
[0054] Motion compensation unit 82 may also perform the interpolation using an interpolation filter as used by video encoder 20 during encoding of the video block to calculate sub-integer pixel interpolated values of the reference block. In this case, motion compensation unit 82 may determine the interpolation filter used by video encoder 20 from the received syntax element and may determine the interpolation filter to produce the predictive block.
[0055] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80 using the same quantization parameters calculated by video encoder 20 for each video block in a video frame to determine the degree of quantization. Inverse transform processing unit 88 performs an inverse transform, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, on the transform coefficients to reconstruct residual blocks in the pixel domain.
[0056] After motion compensation unit 82 or intra BC unit 85 generates a predictive block for the current video block based on the vectors and other syntax elements, adder 90 reconstructs a decoded video block for the current video block by adding the residual block from inverse transform processing unit 88 and the corresponding predictive block generated by motion compensation unit 82 and intra BC unit 85. To further process the decoded video block, an in-loop filter 91 may be disposed between adder 90 and DPB 92. In-loop filtering 91, such as a deblocking filter, sample adaptive offset (SAO), and adaptive in-loop filter (ALF), may be applied to the reconstructed CU before being placed in the reference picture store. The decoded video blocks in a given frame are then stored in DPB 92, which stores reference frames used for subsequent motion compensation of the video block. DPB 92, or a memory device separate from DPB 92, may also store decoded video for later presentation on a display device, such as display device 34 of FIG. 1.
[0057] In a typical video encoding process, a video sequence usually includes a set of ordered frames or pictures. Each frame may include three sample arrays, denoted SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. As another example, a frame may be black and white and therefore includes only one two-dimensional array of luma samples.
[0058] Similar to HEVC, the AVS3 standard is structured based on a block-based hybrid video coding framework. An input video signal is processed in blocks (called coding units (CUs)). Unlike HEVC, which divides blocks based only on quadtrees, in AVS3, one coding tree unit (CTU) is divided into CUs to adapt to changing local characteristics based on quadtrees / binary trees / extended quadtrees. In addition, the concept of multiple partition unit types in HEVC is removed, i.e., the separation of CUs, prediction units (PUs), and transform units (TUs) does not exist in AVS3. Instead, each CU is always used as a basic unit for both prediction and transformation without further partitioning. In the tree partition structure of AVS3, one CTU is first divided based on a quadtree structure. Then, each quadtree leaf node may be further divided based on a binary tree structure and an extended quadtree structure.
[0059] As shown in FIG. 4A, video encoder 20 (or, more specifically, partition unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of coding tree units (CTUs). A video frame may include an integer number of CTUs ordered consecutively in raster scan order, from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTUs are signaled by video encoder 20 in the sequence parameter set such that all CTUs in a video sequence have the same size, which is one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a particular size. As shown in FIG. 4B, each CTU may comprise one coding tree block (CTB) of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements used to code the samples of the coding tree blocks. The syntax elements describe the properties of different types of units of coded pixel blocks and how a video sequence may be reconstructed at video decoder 30, including inter or intra prediction, intra prediction mode, motion vectors, and other parameters. For a monochrome image or an image with three separate color planes, a CTU may include a single coding tree block and syntax elements used to code samples of the coding tree block. A coding tree block may be an N×N block of samples.
[0060] To achieve better performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination of both, on the coding tree block of the CTU to divide the CTU into smaller coding units (CUs). As shown in FIG. 4C, 64x64 CTU 400 is first divided into four smaller CUs, each with a block size of 32x32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16x16. Two 16x16 CUs 430 and 440 are each further divided into four CUs with a block size of 8x8. FIG. 4D shows a quad tree data structure illustrating the final result of the division process of CTU 400 as shown in FIG. 4C, where each leaf node of the quad tree corresponds to one CU with a respective size ranging from 32x32 to 8x8. Similar to the CTU shown in FIG. 4B, each CU may include a coding block (CB) of luma samples and two corresponding coding blocks of chroma samples of the same size frame, and syntax elements used to code the samples of the coding block. In a monochrome image, or an image with three separate color planes, a CU may include a single coding block and syntax structures used to code the samples of the coding block. It should be noted that the quadtree partitioning shown in FIG. 4C and FIG. 4D is for illustrative purposes only, and one CTU can be partitioned into CUs to accommodate different local characteristics based on quadtree / ternary tree / binary tree partitioning. In a multi-type tree structure, one CTU is partitioned by a quadtree structure, and each quadtree leaf CU may be further partitioned by a binary tree structure and a ternary tree structure. As shown in FIG. 4E, there are five partitioning / partitioning types in AVS3, namely, quaternary partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal extended quadtree partitioning, and vertical extended quadtree partitioning.
[0061] In some implementations, video encoder 20 may further partition the coding blocks of a CU into one or more MxN prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. A prediction unit (PU) of a CU may comprise a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and syntax elements used to predict the prediction block. For monochrome images, or images with three separate color planes, a PU may comprise a single prediction block and syntax structures used to predict the prediction block. Video encoder 20 may generate predicted luma, Cb, and Cr blocks for luma, Cb, and Cr prediction blocks for each PU of the CU.
[0062] Video encoder 20 may use intra prediction or inter prediction to generate the predictive blocks for a PU. If video encoder 20 uses intra prediction to generate the predictive blocks for a PU, video encoder 20 may generate the predictive blocks for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate the predictive blocks for a PU, video encoder 20 may generate the predictive blocks for the PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0063] After video encoder 20 generates the predictive luma, Cb, and Cr blocks for one or more PUs of a CU, video encoder 20 may generate a luma residual block of the CU by subtracting the predictive luma block of the CU from its original luma coding block, such that each sample in the luma residual block of the CU indicates a difference between a luma sample in one of the predictive Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU. Similarly, video encoder 20 may generate a Cb residual block and a Cr residual block of the CU, respectively, such that each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predictive Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates a difference between a Cr sample in one of the predictive Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0064] Further, as illustrated in FIG. 4C, video encoder 20 may use quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A transform unit (TU) of a CU may comprise a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements used to transform the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome image, or an image with three separate color planes, a TU may comprise a single transform block and syntax structures used to transform the samples of the transform block.
[0065] Video encoder 20 may apply one or more transforms to a luma transform block of a TU to generate a luma coefficient block of the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 may apply one or more transforms to a Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to a Cr transform block of a TU to generate a Cr coefficient block for the TU.
[0066] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients and provide further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform context-adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream including a sequence of bits forming a representation of the encoded frame and associated data, which is either stored in the storage device 32 or transmitted to the destination device 14.
[0067] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct frames of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally reciprocal to the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on coefficient blocks associated with TUs of the current CU to reconstruct residual blocks associated with the TUs of the current CU. Video decoder 30 also reconstructs coding blocks of the current CU by adding samples of predictive blocks for PUs of the current CU to corresponding samples of transform blocks of the TUs of the current CU. After reconstructing the coding blocks for each CU of the frame, video decoder 30 may reconstruct the frame.
[0068] SAO is a process of modifying decoded samples by conditionally adding an offset value to each sample after the application of a deblocking filter based on values in a lookup table transmitted by the encoder. SAO filtering is performed on a per-region basis based on the filtering type selected per CTB by the syntax element sao-type-idx. A value of 0 in sao-type-idx indicates that no SAO filter is applied to the CTB, while values 1 and 2 signal the use of band-offset and edge-offset filtering types, respectively. In band-offset mode, specified by sao-type-idx equal to 1, the offset value selected depends directly on the sample amplitude. In this mode, the total sample amplitude range is uniformly divided into 32 segments called bands, and sample values belonging to four of these bands (that are consecutive within the 32 bands) are modified by adding a transmitted value, denoted as band offset, which can be positive or negative. The main reason for using four consecutive bands is that in smooth areas where banding artifacts may appear, the sample amplitudes in the CTB tend to be concentrated in very few bands. In addition, the design choice of using four offsets is unified with the edge-offset computation mode that also uses four offset values. In the edge-offset mode specified by sao-type-idx equal to 2, the syntax element sao-eo-class with a value between 0 and 3 signals whether horizontal, vertical, or one of the two diagonal gradient directions is used for edge-offset classification in the CTB.
[0069] FIG. 5 is a block diagram illustrating four gradient patterns used in sample-adaptive offset SAO according to some implementations of the present disclosure. The four gradient patterns 502, 504, 506, and 508 are for each sao-eo-class in edge-offset mode. The sample labeled with "p" indicates the center sample to be considered. The two samples labeled with "n0" and "n1" specify two adjacent samples along the (a) horizontal (sao-eo-class=0), (b) vertical (sao-eo-class=1), (c) 135° diagonal (sao-eo-class=2), and (d) 45° (sao-eo-class=3) gradient patterns. Each sample in the CTB is classified into one of five EdgeIdx categories by comparing the sample value p located at a certain position with the values n0 and n1 of the two samples located at the adjacent positions, as shown in FIG. 5. No additional signaling is required for EdgeIdx classification, since this classification is done for each sample based on the decoded sample value. Depending on the EdgeIdx category at the sample position, an offset value from the transmitted lookup table is added to the sample value for EdgeIdx categories from 1 to 4. The offset value is always positive for categories 1 and 2, and always negative for categories 3 and 4. Thus, the filter generally has a smoothing effect in edge offset mode. Table 1 below illustrates sample EdgeIdx categories in SAO edge classification. [Table 1]
[0070] For SAO types 1 and 2, a total of four amplitude offset values are transmitted to the decoder for each CTB. For type 1, the sign is also encoded. The offset values and related syntax elements such as sao-type-idx and sao-eo-class are determined by the encoder, typically using a criterion that optimizes rate-distortion performance. SAO parameters can be indicated to be inherited from the left or top of the CTB using a merge flag to make the signaling efficient. In summary, SAO is a nonlinear filtering operation that allows further refinement of the reconstructed signal, which may enhance the signal representation both in smooth areas and around edges.
[0071] In some embodiments, methods and systems are disclosed herein for improving coding efficiency or reducing the complexity of sample adaptive offset (SAO) by introducing cross-component information. SAO is used in the HEVC, VVC, AVS2, and AVS3 standards. In the following description, the existing SAO design in the HEVC, VVC, AVS2, and AVS3 standards is used as the basic SAO method, but those skilled in the art of video coding may also understand that the cross-component method described in this disclosure may also be applied to other loop filter designs, or other coding tools with similar design ideas. For example, in the AVS3 standard, SAO is replaced by a coding tool called Improved Sample Adaptive Offset (ESAO). However, the CCSAO disclosed herein may also be applied in parallel with the ESAO. In another example, the CCSAO may be applied in parallel with the Constrained Directional Enhancement Filter (CDEF) of the AV1 standard.
[0072] In the existing SAO designs of the HEVC, VVC, AVS2, and AVS3 standards, the luma Y, chroma Cb, and chroma Cr sample offset values are determined independently. That is, for example, the current chroma sample offset is determined only by the current chroma sample value and the adjacent chroma sample value without considering the co-located or adjacent luma samples. However, luma samples preserve more original picture detail information than chroma samples, and they can benefit from the current chroma sample offset determination. Furthermore, since chroma samples usually lose high frequency details after RGB to YCbCr color conversion or after quantization and deblocking filters, introducing luma samples with preserved high frequency details for chroma offset determination can benefit from chroma sample reconstruction. Therefore, further gains can be expected by exploring cross-component correlations, for example, by using a Cross-Component Sample Adaptive Offset (CCSAO) method and system. In some embodiments, the correlation here includes not only cross-component sample values, but also picture / coding information such as prediction / residual coding mode, transform type, and quantization / deblocking / SAO / ALF parameters from the cross-component.
[0073] Another example is the case of SAO, where the luma sample offset is determined by the luma sample only. However, for example, luma samples with the same band offset (BO) classification can be further classified by their collocated and adjacent chroma samples, which may result in a more effective classification. SAO classification may be obtained as a shortcut to correct sample differences between the original and reconstructed pictures. Therefore, an effective classification is desired.
[0074] FIG. 6A is a block diagram illustrating a system and process of CCSAO applied to chroma samples and using DBF Y as input according to some implementations of the present disclosure. The luma sample after the luma deblocking filter (DBF Y) is used to determine additional offsets for chroma Cb and chroma Cr after SAO Cb and SAO Cr. For example, the current chroma sample 602 is first sorted using the collocated 604 and adjacent (white) luma samples 606, and the corresponding CCSAO offset value of the corresponding sort is added to the current chroma sample value. FIG. 6B is a block diagram illustrating a system and process of CCSAO applied to luma samples and chroma samples and using DBF Y / Cb / Cr as input according to some implementations of the present disclosure. FIG. 6C is a block diagram illustrating a system and process of CCSAO that can operate independently according to some implementations of the present disclosure. FIG. 6D is a block diagram illustrating a system and process of CCSAO, which may be applied recursively (2 or N times) with the same or different offsets in the same codec stage, or may be repeated in different stages, according to some implementations of the present disclosure. In summary, in some embodiments, information of the current luma sample and adjacent luma samples, information of the co-located and adjacent chroma samples (Cb and Cr) may be used to classify the current luma sample. In some embodiments, co-located and adjacent luma samples, co-located and adjacent cross chroma samples, and current and adjacent chroma samples may be used to classify the current chroma sample (Cb or Cr). In some embodiments, CCSAO may be cascaded (1) after DBF Y / Cb / Cr, (2) after pre-DBF reconstructed image Y / Cb / Cr, or (3) after SAO Y / Cb / Cr, or (4) after ALF Y / Cb / Cr.
[0075] In some embodiments, the CCSAO may also be applied in parallel with other encoding tools, such as the ESAO in the AVS standard or the CDEF in the AV1 standard. Figure 6E is a block diagram illustrating a system and process of the CCSAO applied in parallel with the improved sample adaptive offset ESAO in the AVS standard according to some implementations of the present disclosure.
[0076] FIG. 6F is a block diagram illustrating a system and process of a CCSAO applied after the SAO according to some implementations of the present disclosure. In some embodiments, FIG. 6F shows that the position of the CCSAO may be after the SAO, i.e., the position of the cross-component adaptive loop filter (CCALF) in the VVC standard. FIG. 6G is a block diagram illustrating that the system and process of the CCSAO according to some implementations of the present disclosure can operate independently without the CCALF. In some embodiments, the SAO Y / Cb / Cr may be replaced with ESAO, for example in the AVS3 standard.
[0077] FIG. 6H is a block diagram illustrating a system and process of CCSAO applied in parallel with CCALF according to some implementations of the present disclosure. In some embodiments, FIG. 6H shows that CCSAO may be applied in parallel with CCALF. In some embodiments, in FIG. 6H, the positions of CCALF and CCSAO may be switched. In some embodiments, in FIG. 6A-FIG. 6H or throughout this disclosure, the SAO Y / Cb / Cr block may be replaced with ESAO Y / Cb / Cr (in AVS3) or CDEF (in AV1). Note that Y / Cb / Cr may also be represented as Y / U / V in the video coding area.
[0078] In some embodiments, current chroma sample classification is to reuse the SAO type (edge offset (EO) or BO), classification, and category of the collocated luma sample. The corresponding CCSAO offset can be signaled or derived from the decoder itself. For example, let h_Y be the collocated luma SAO offset, and h_Cb and h_Cr be the CCSAO Cb and Cr offsets respectively. h_Cb (or h_Cr) = w * h_Y, where w can be selected from a limited table. For example, it can be +-1 / 4, +-1 / 2, 0, +-1, +-2, +-4, etc., where |w| only includes values that are powers of 2.
[0079] In some embodiments, the comparison scores [-8, 8] of the collocated luma sample (Y0) and eight adjacent luma samples are used, where a total of 17 classifications are produced. Initial classification = 0 Loop through the eight adjacent luma samples (Yi, i = 1 to 8) If Y0 > Yi, classification += 1 Otherwise, if Y0 < Yi, classification = 1
[0080] In some embodiments, the above classification methods may be combined. For example, the comparison scores combined with SAO BO (32-band classification) are used to increase diversity, which produces a total of 17×32 classifications. In some embodiments, Cb and Cr may use the same classification to reduce complexity or save bits.
[0081] FIG. 7 is a block diagram illustrating a sample process using CCSAO according to some implementations of the present disclosure. Specifically, FIG. 7 shows that the input of CCSAO may introduce vertical DBF and horizontal DBF inputs to simplify classification decisions or to increase flexibility. For example, let Y0_DBF_V, Y0_DBF_H, and Y0 be the collocated luma samples at the inputs of DBF_V, DBF_H, and SAO, respectively. Yi_DBF_V, Yi_DBF_H, and Yi are the adjacent eight luma samples at the inputs of DBF_V, DBF_H, and SAO, respectively, where i=1 to 8. Max Y0 = Max (Y0_DBF_V, Y0_DBF_H, Y0_DBF) Max Yi = Max (Yi_DBF_V, Yi_DBF_H, Yi_DBF) Then, maxY0 and maxYi are fed to the CCSAO classification.
[0082] 8 is a block diagram illustrating the CCSAO process being interleaved into vertical and horizontal deblocking filters DBF according to some implementations of the present disclosure. In some embodiments, the CCSAO blocks of Figures 6, 7, and 8 may be optional. For example, using the input of DBF_V luma samples as the CCSAO input, while using Y0_DBF_V and Yi_DBF_V for the first CCSAO_V applying the same sample processing as in Figure 6.
[0083] In some embodiments, the implemented CCSAO syntax is shown in Table 2 below. [Table 2]
[0084] In some embodiments, when one additional chroma offset is signaled to signal the CCSAO Cb and Cr offset values, other chroma component offsets can be derived by plus or minus sign, or weighting to save bit overhead. For example, let h_Cb and h_Cr be the offsets of CCSAO Cb and Cr, respectively. With explicit signaling w (where w=+-|w|, with limited |w| candidates), h_Cr can be derived from h_Cb without explicitly signaling h_Cr itself. h_Cr=w*h_Cb
[0085] FIG. 9 is a flowchart illustrating an example process 900 for decoding a video signal using cross-component correlation according to some implementations of the disclosure.
[0086] Video decoder 30 receives a video signal including a first component and a second component (910). In some embodiments, the first component is a luma component and the second component is a chroma component of the video signal.
[0087] Video decoder 30 also receives a number of offsets associated with the second component (920).
[0088] Video decoder 30 then utilizes the characteristic measurements of the first component to obtain a classification category associated with the second component (930). For example, in FIG. 6, the current chroma sample 602 is first classified using the collocated 604 and adjacent (white) luma sample 606, and the corresponding CCSAO offset value is added to the current chroma sample value.
[0089] Video decoder 30 further selects the first offset from the multiple offsets of the second component according to the classification category (940).
[0090] Video decoder 30 additionally modifies the second component based on the selected first offset (950).
[0091] In some embodiments, utilizing the characteristic measurements of the first component to obtain a classification category associated with the second component (930) includes utilizing each sample of the first component to obtain a classification category for each respective sample of the second component, where the each sample of the first component is a collocated sample of the first component with the each respective sample of the second component. For example, the current chroma sample classification is to reuse the SAO type (EO or BO), classification, and category of the collocated luma sample.
[0092] In some embodiments, utilizing the characteristic measurements of the first component to obtain a classification category associated with the second component (930) includes utilizing each sample of the first component to obtain a classification category for each sample of the second component, where each sample of the first component is reconstructed before being deblocked or reconstructed after being deblocked. In some embodiments, the first component is deblocked with a deblocking filter (DBF). In some embodiments, the first component is deblocked with a luma deblocking filter (DBF Y). For example, instead of FIG. 6 or FIG. 7, the CCSAO input may also be before DBF Y.
[0093] In some embodiments, the characteristic measure is derived by dividing the range of sample values of the first component into bands and selecting the bands based on the intensity values of the samples in the first component, hi some embodiments, the characteristic measure is derived from a band offset (BO).
[0094] In some embodiments, the characteristic measure is derived based on a direction and intensity of edge information of the samples in the first component, hi some embodiments, the characteristic measure is derived from an edge offset (EO).
[0095] In some embodiments, modifying the second component (950) includes adding the selected first offset directly to the second component, e.g., a corresponding CCSAO offset value is added to the current chroma component sample.
[0096] In some embodiments, modifying the second component (950) includes mapping the selected first offset to a second offset and adding the mapped second offset to the second component. If one additional chroma offset is signaled, e.g., to signal CCSAO Cb and Cr offset values, the other chroma component offsets can be derived by using a plus or minus sign, or a weighting to save bit overhead.
[0097] In some embodiments, receiving the video signal (910) includes receiving a syntax element indicating whether a method for decoding the video signal using CCSAO is enabled for the video signal in a sequence parameter set (SPS). In some embodiments, cc_sao_enabled_flag indicates whether CCSAO is enabled at the sequence level.
[0098] In some embodiments, receiving the video signal (910) includes receiving a syntax element indicating whether a method for decoding the video signal using CCSAO is enabled for a slice level second component. In some embodiments, slice_cc_sao_cb_flag or slice_cc_sao_cr_flag indicates whether CCSAO is enabled for the Cb or Cr slice, respectively.
[0099] In some embodiments, receiving multiple offsets associated with the second component (920) includes receiving different offsets for different coding tree units (CTUs), where cc_sao_offset_sign_flag indicates the sign of the offset and cc_sao_offset_abs indicates the CCSAO Cb and Cr offset values for the current CTU.
[0100] In some embodiments, receiving the multiple offsets associated with the second component (920) includes receiving a syntax element indicating whether the received offset of the CTU is the same as one of the CTU's neighboring CTUs, where the neighboring CTU is either the left or the top neighboring CTU. For example, cc_sao_merge_up_flag indicates whether the CCSAO offset is merged from the left or the top CTU.
[0101] In some embodiments, the video signal further includes a third component, and a method of decoding the video signal using CCSAO includes: receiving a second plurality of offsets associated with the third component; utilizing characteristic measurements of the first component to obtain a second classification category associated with the third component; selecting a third offset from the second plurality of offsets of the third component according to the second classification category; and modifying the third component based on the selected third offset.
[0102] 11 is a block diagram of a sample process illustrating that all co-located and adjacent (white) luma / chroma samples may be fed into a CCSAO classification, according to some implementations of the present disclosure. FIGS. 6A, 6B, and 11 show the inputs of the CCSAO classification. In FIG. 11, the current chroma sample is 1104, the cross-component co-located chroma sample is 1102, and the co-located luma sample is 1106.
[0103] In some embodiments, an example classifier (C0) uses the following co-located luma or chroma sample values (Y0) of FIG. 12 for classification (Y4 / U4 / V4 in FIGS. 6B and 6C). If band_num is the number of equal bands of the luma or chroma dynamic range and bit_depth is the sequence bit depth, an example classification index for the current chroma sample is as follows: Classification (C0)=(Y0*band_num)>>bit_depth
[0104] In some embodiments, the classification takes into account rounding, for example, as follows: Classification (C0)=((Y0*band_num)+(1<<bit_depth))> >bit_depth
[0105] Some examples of band_num and bit_depth are listed below in Table 3. Table 3 shows three classification examples where the number of bands in each classification example is different. [Table 3]
[0106] In some embodiments, the classifier uses a different luma sample location for the C0 classification. Figure 10A is a block diagram showing a classifier that uses a different luma (or chroma) sample location for the C0 classification, for example, using neighboring Y7 instead of Y0 for the C0 classification, according to some implementations of the present disclosure.
[0107] In some embodiments, different classifiers may be switched at sequence parameter set (SPS) / adaptation parameter set (APS) / picture parameter set (PPS) / picture header (PH) / slice header (SH) / region / coding tree unit (CTU) / coding unit (CU) / sub-block / sample level. For example, in Figure 10, Y0 is used for POC0, but Y7 is used for POC1, as shown in Table 4 below. [Table 4]
[0108] In some embodiments, FIG. 10B illustrates some examples of different shapes of luma candidates according to some implementations of the present disclosure. For example, the shapes may be constrained. In some cases, the total number of luma candidates must be a power of two, as shown in FIG. 10B(b), FIG. 10B(c), FIG. 10B(d). In some cases, the number of luma candidates must be horizontally and vertically symmetric about the (center) chroma sample, as shown in FIG. 10B(a), FIG. 10B(c), FIG. 10B(d), FIG. 10B(e). In some embodiments, power of two constraints and symmetry constraints may also be applied to the chroma candidates. The U / V parts of FIG. 6B and FIG. 6C show an example of a symmetry constraint. In some embodiments, different color formats may have different classifier "constraints." For example, as shown in Figures 6B and 6C, the 420 color format uses luma / chroma candidate selection (one candidate selected from a 3x3 shape), while the 444 color format uses Figure 10B(f) for luma candidate selection and chroma candidate selection, and the 422 color format uses Figure 10B(g) for luma (two chroma samples share four luma candidates) and Figure 10B(f) for chroma candidates.
[0109] In some embodiments, the C0 position and C0 band_num may be switched in combination at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different combinations may be different classifiers as shown in Table 5 below.
Table 5
[0110] In some embodiments, the collocated luma sample value (Y0) is replaced by a value (Yp) obtained by weighting the collocated and adjacent luma samples. FIG. 12 illustrates an exemplary classifier according to some implementations of the present disclosure, where the collocated luma sample value is replaced by a value obtained by weighting the collocated and adjacent luma samples. The collocated luma sample value (Y0) may be replaced by a phase correction value (Yp) obtained by weighting adjacent luma samples. Different Yps may be different classifiers.
[0111] In some embodiments, different Yps are applied to different chroma formats. For example, in FIG. 12, the Yp in (a) is used for the 420 chroma format, the Yp in (b) is used for the 422 chroma format, and Y0 is used for the 444 chroma format.
[0112] In some embodiments, another classifier (C1) is the comparison score [-8,8] of the collocated luma sample (Y0) and eight adjacent luma samples, which produces a total of 17 classifications as shown below. Initial classification (C1) = 0, loop through eight adjacent luma samples (Yi, i = 1~8) If Y0>Yi, classification += 1 Otherwise, Y0<Yi, classification = 1
[0113] In some embodiments, an example of C1 is equal to the following function where the threshold th is 0. ClassIdx = Index2ClassTable(f(C, P1) + f(C, P2) + … + f(C, P8)) When x - y > th, f(x, y) = 1; when x - y = th, f(x, y) = 0; when x - y < th, f(x, y) = -1 In the formula, Index2ClassTable is a look-up table (LUT), C is the current or collocated sample, and P1 to P8 are adjacent samples.
[0114] In some embodiments, similar to the C4 classifier, one or more thresholds may be predefined (e.g., held in the LUT) or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to help classify (quantize) the differences.
[0115] In some embodiments, only the variation (C1’) counts the comparison scores [0, 8], thereby producing eight classifications. (C1, C1’) is a classifier group, and the PH / SH level flag may be signaled to switch between C1 and C1’.
[0116] Initial classification (C1’) = 0, loop through eight adjacent lumasamples (Yi, i = 1 to 8) If Y0 > Yi, classification += 1
[0117] In some embodiments, the variation (C1s) selectively uses the neighboring N of the M neighboring samples to count the comparison score. An M-bit bitmask may be signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level to indicate which neighboring samples are selected to count the comparison score. Using FIG. 6B as an example of a luma classifier, 8 neighboring luma samples are candidates and an 8-bit bitmask (01111110) is signaled at the PH to indicate that 6 samples Y1 to Y6 are selected, and thus the comparison score is [-6, 6], yielding 13 offsets. The selective classifier C1s gives the encoder more options to trade off between offset signaling overhead and classification granularity.
[0118] Like C1s, variation(C1's) only counts comparison scores [0,+N], so the previous example bitmask 01111110 would give a comparison score in [0,6], yielding an offset of 7.
[0119] In some embodiments, different classifiers are combined to produce a general classifier, for example for different pictures (different POC values), different classifiers are applied as shown in Table 6-1 below. [Table 6-1]
[0120] In some embodiments, another classifier example (C3) uses a bitmask for classification, as shown in Table 6-2. A 10-bit bitmask is signaled at SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level to indicate the classifier. For example, a bitmask of 11 1100 0000 means that for a given 10-bit luma sample value, only the most significant bit (MSB): 4 bits are used for classification, and a total of 16 classifications are calculated. Another exemplary bitmask of 10 0100 0001 means that only 3 bits are used for classification, and a total of 8 classifications are calculated.
[0121] In some embodiments, the bitmask length (N) may be fixed or switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. For example, for a 10-bit sequence, a 4-bit bitmask 1110 is signaled at the PH in the picture, and the MSB 3 bits b9, b8, b7 are used for classification. Another example is a 4-bit bitmask 0011 on the LSB, and b0, b1 are used for classification. The bitmask classifier may apply to luma classification or chroma classification. Whether to use the MSB or LSB for the bitmask N may be fixed or switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level.
[0122] In some embodiments, the luma position and C3 bitmask can be combined and switched at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. Different combinations can be different classifiers.
[0123] In some embodiments, a bitmask limit "max number of sec" may be applied to limit the corresponding offset number. For example, the bitmask "max number of sec" may be limited to 4 in an SPS, yielding a maximum offset in a sequence of 16. The bitmasks of different POCs may be different, but the "max number of sec" may not exceed 4 (total classification may not exceed 16). The "max number of sec" value may be signaled and switched at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. [Table 6-2]
[0124] In some embodiments, as shown in FIG. 11, other cross-component chroma samples, e.g., chroma sample 1102 and its neighbors, may also be fed into the CCSAO classification, e.g., for the current chroma sample 1104. For example, the Cr chroma sample may be fed into the CCSAO Cb classification. The Cb chroma sample may be fed into the CCSAO Cr classification. The classifier for the cross-component chroma sample may be the same as the luma cross-component classifier, or may have its own classifier as described in this disclosure. The two classifiers may be combined to form a combined classifier for classifying the current chroma sample. For example, the combined classifier combining the cross-component luma and chroma samples produces a total of 16 classifications, as shown in Table 6-3 below. [Table 6-3]
[0125] All of the above categories (C0, C1, C1', C2, C3) may be combined, see for example Table 6-4 below. [Table 6-4]
[0126] In some embodiments, an example classifier (C2) uses the difference between the co-located luma sample and the adjacent luma sample (Yn). Figure 12(c) shows an example of Yn with a dynamic range of [-1024, 1023] for a bit depth of 10. Let C2 band_num be the number of equal bands of the Yn dynamic range, Classification (C2)=(Yn+(1<<bit_depth)*band_num)> > (bit_depth+1).
[0127] In some embodiments, C0 and C2 are combined to produce a general classifier, for example for different pictures (different POCs), different classifiers are applied, as shown in Table 7 below. [Table 7]
[0128] In some embodiments, all the above classifiers (C0, C1, C1', C2) are combined. For example, for different pictures (different POCs), different classifiers are applied, as shown in Table 8-1 below. [Table 8-1]
[0129] In some embodiments, an example classifier (C4) uses the difference between the CCSAO input value and the compensated sample value for classification, as shown in Table 8-2 below. For example, if CCSAO is applied at the ALF stage, the difference between the pre-ALF and post-ALF sample values of the current component is used for classification. One or more thresholds may be predefined (e.g., stored in a look-up table (LUT)) or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level to help classify (quantize) the difference. The C4 classifier may be combined with C0 Y / U / V bandNum to form a combined classifier (e.g., the example POC1 shown in Table 8-2). [Table 8-2]
[0130] In some embodiments, an example classifier (C5) uses "encoding information" to aid sub-block classification, since different encoding modes may introduce different distortion statistics into the reconstructed image. CCSAO samples are classified by their sample pre-encoding information, and the combination of encoding information may form a classifier, for example, as shown in Table 8-3 below. Figure 30 below shows another example of different stages of encoding information in C5. [Table 8-3]
[0131] In some embodiments, the classifier example (C6) uses YUV color transform values for classification, for example, to classify the current Y component, select 1 / 1 / 1 collocated or adjacent Y / U / V samples to color transform to RGB, and quantize the R value using C3 bandNum to become the current Y component classifier.
[0132] In some embodiments, other example classifiers that only use current component information for current component classification may be used as cross-component classification. For example, luma sample information and eo-class are used to derive EdgeIdx and classify the current chroma sample as shown in FIG. 5 and Table 1. Other "non-cross-component" classifiers that can also be used as cross-component classifiers include edge direction, pixel intensity, pixel variation, pixel variance, Laplacian sum of pixels, Sobel operator, compass operator, high pass filtering value, low pass filtering value, etc.
[0133] In some embodiments, multiple classifiers are used in the same POC. The current frame is divided by several regions, and each region uses the same classifier. For example, three different classifiers are used in POC0, and which classifier (0, 1, or 2) is used is signaled at the CTU level, as shown in Table 9 below. [Table 9]
[0134] In some embodiments, the maximum number of classifiers (classifiers may also be referred to as alternative offset sets) may be fixed or signaled at the SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. In one example, the fixed (predefined) maximum number of classifiers is 4. In that case, 4 different classifiers are used in POC0, and which classifier (0, 1, or 2) is used is signaled at the CTU level. A truncated-unary (TU) code may be used to indicate the classifier used for each luma CTB or each chroma CTB. For example, as shown in Table 10 below, if the TU code is 0, no CCSAO is applied; if the TU code is 10, set 0 is applied; if the TU code is 110, set 1 is applied; if the TU code is 1110, set 2 is applied; if the TU code is 1111, set 3 is applied. It is also possible to indicate the classifiers (offset set indexes) of the CTB using fixed length codes, Golomb-Rice codes, and exponential-Golomb codes. Three different classifiers are used in POC1. [Table 10]
[0135] Examples of Cb and Cr CTB offset set indexes are given for a 1280x720 sequence POC0 (if the CTU size is 128x128, the number of CTUs in a frame is 10x6). POC0 Cb uses 4 offset sets and Cr uses 1 offset set. As shown in Table 11-1 below, if the offset set index is 0, no CCSAO is applied; if the offset set index is 1, set 0 is applied; if the offset set index is 2, set 1 is applied; if the offset set index is 3, set 2 is applied; if the offset set index is 4, set 3 is applied. Type means the position of the selected co-located luma sample (Yi). Different offset sets may have different types, band_num, and corresponding offsets. [Table 11-1]
[0136] In some embodiments, an example of using collocated / current and adjacent Y / U / V samples together for classification is listed in Table 11-2 below (three-component joint bandNum classification for each Y / U / V component). In POC0, {Y,U,V} have {2,4,1} offset sets respectively. Each offset set may be adaptively switched at SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. Different offset sets may have different classifiers. For example, to classify the current Y4 luma sample as the candidate position (candPos) shown in FIG. 6B and FIG. 6C, Y set 0 selects {current Y4, collocated U4, collocated V4} as candidates, each of which has different bandNum{Y,U,V}={16,1,2}. Using {candY, candU, candV} as the sample values of the selected {Y,U,V} candidates, the total number of classifications is 32, and the classification index derivation can be shown as follows: bandY=(candY*bandNumY)>>BitDepth; bandU=(candU*bandNumU)>>BitDepth; bandV=(candV*bandNumV)>>BitDepth; classIdx=bandY*bandNumU*bandNumV +bandU*bandNumV +bandV;
[0137] Another example is the POC1 component V set 1 classification: in this example, candPos={Y8 adjacent, U3 adjacent, V0 adjacent} with bandNum={4,1,2} is used, which produces 8 classifications. [Table 11-2]
[0138] In some embodiments, examples are enumerated using collocated and adjacent Y / U / V samples together for the current Y / U / V sample classification (three component combined edgeNum(C1s) and bandNum classification for each Y / U / V component), for example as shown in Table 11-3 below. Edge CandPos is the center position used for C1s classifier, edge bitMask is the C1s adjacent sample activation indicator, and edgeNum is the corresponding number of C1s classification. In this example, C1s is only applied to the Y classifier (hence edgeNum is equal to edgeNumY), where edge candPos is always Y4 (current / collocated sample position). However, C1s may be applied to the Y / U / V classifier with edge candPos as the adjacent sample position.
[0139] With the difference representing the comparison score of Y C1s, the classIdx derivation can be as follows: bandY=(candY*bandNumY)>>BitDepth; bandU=(candU*bandNumU)>>BitDepth; bandV=(candV*bandNumV)>>BitDepth; edgeIdx = diff + (edgeNum>>1); bandIdx=bandY*bandNumU*bandNumV +bandU*bandNumV +bandV; classIdx=bandIdx*edgeNum+edgeIdx; [Table 11-3(1)] [Table 11-3(2)] [Table 11-3(3)]
[0140] In some embodiments, max band_num (bandNumY, bandNumU, or bandNumV) may be fixed or signaled at SPS / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level. For example, max band_num=16 is fixed at the decoder, and for each frame, 4 bits are signaled to indicate the C0 band_num in the frame. Some other max band_num examples are listed in Table 12 below. [Table 12]
[0141] In some embodiments, the maximum number of classifications or offsets (combinations of multiple classifiers used together, e.g., C1s edgeNum*C1 bandNumY*bandNumU*bandNumV) for each set (or all sets added) may be fixed or signaled at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, the maximum is fixed for all added sets class_num=256*4, and an encoder conformance check or a decoder norm check can be used to check the constraint.
[0142] In some embodiments, restrictions may be applied to the C0 classification, for example band_num (bandNumY, bandNumU, or bandNumV) may be restricted to only power-of-two values. Instead of explicitly signaling band_num, the syntax band_num_shift is signaled. The decoder may use shift operations to avoid multiplications. Different band_num_shift may be used for different components. Classification (C0)=(Y0>>band_num_shift)>>bit_depth
[0143] Another example operation is to take into account rounding to reduce error. Classification (C0)=((Y0+(1<<(band_num_shift-1)))>>band_num_shift)>>bit_depth
[0144] For example, as shown in Table 13, if band_num_max (Y, U, or V) is 16, the possible band_num_shift candidates are 0, 1, 2, 3, 4, corresponding to band_num=1, 2, 4, 8, 16. [Table 13(1)] [Table 13(2)]
[0145] In some embodiments, the classifiers applied to Cb and Cr are different. The Cb and Cr offsets for every classification may be signaled separately. For example, different signaled offsets are applied to different chroma components, as shown in Table 14 below. [Table 14]
[0146] In some embodiments, the maximum offset value is fixed or signaled at the sequence parameter set (SPS) / adaptive parameter set (APS) / picture parameter set (PPS) / picture header (PH) / slice header (SH) / region / CTU / CU / subblock / sample level. For example, the maximum offset is between [-15,15]. Different components may have different maximum offset values.
[0147] In some embodiments, the offset signaling may use Differential Pulse-Code Modulation (DPCM), e.g., the offset {3, 3, 2, 1, -1} may be signaled as {3, 0, -1, -1, -2}.
[0148] In some embodiments, the offsets may be stored in an APS or memory buffer for next picture / slice reuse, and an index may be signaled to indicate which stored previous frame offset is used for the current picture.
[0149] In some embodiments, the classifiers for Cb and Cr are the same, for example, the Cb and Cr offsets for all classes may be signaled together, as shown in Table 15 below. [Table 15]
[0150] In some embodiments, the classifiers for Cb and Cr may be the same. The Cb and Cr offsets for all classifications may be signaled together using a sign flag differential, for example, as shown in Table 16 below. According to Table 16, if the Cb offset is (3, 3, 2, -1), the derived Cr offset is (-3, -3, -2, 1). [Table 16]
[0151] In some embodiments, a signed flag may be signaled for each classification, for example as shown in Table 17 below: According to Table 17, if the Cb offset is (3,3,2,-1), the derived Cr offset is (-3,3,2,1) according to the respective signed flag. [Table 17]
[0152] In some embodiments, the classifiers for Cb and Cr may be the same. The Cb and Cr offsets of all classifications may be signaled together using weighted differentials, for example, as shown in Table 18 below. The weights (w) may be selected from a limited table, for example, +-1 / 4, +-1 / 2, 0, +-1, +-2, +-4..., etc., where |w| includes only values that are powers of 2. According to Table 18, if the Cb offset is (3,3,2,-1), the derived Cr offset is (-6,-6,-4,2) according to the respective signed flags. [Table 18]
[0153] In some embodiments, weights may be signaled for each classification, for example as shown in Table 19 below: According to Table 19, if the Cb offset is (3,3,2,-1), the derived Cr offset is (-6,12,0,-1) according to the respective signed flags. [Table 19]
[0154] In some embodiments, when multiple classifiers are used in the same POC, different offset sets are signaled separately or together.
[0155] In some embodiments, previously decoded offsets may be stored for use in future frames. To reduce offset signaling overhead, an index may be signaled to indicate which previously decoded offset set is used for the current frame. For example, the POC0 offset may be reused by POC2 with signaling offset set to idx=0, as shown in Table 20 below. [Table 20]
[0156] In some embodiments, the reuse offset sets idx for Cb and Cr may be different, for example as shown in Table 21 below. [Table 21]
[0157] In some embodiments, offset signaling may use additional syntax including start and length to reduce signaling overhead. For example, if band_num=256, only offsets from band_idx=37 to 44 are signaled. In the example in Table 22-1 below, both start and length syntaxes are 8-bit fixed length encodings that should match the band_num bits. [Table 22-1]
[0158] In some embodiments, when CCSAO is applied to all YUV 3 components, collocated and adjacent YUV samples may be used together for classification, and all offset signaling methods described above for Cb / Cr may be extended to Y / Cb / Cr. In some embodiments, different component offset sets may be stored and used separately (each component has its own stored set) or may be used together (each component shares / reuses its stored one). Another example set is shown in Table 22-2 below. [Table 22-2]
[0159] In some embodiments, if the sequence bit depth is higher than 10 (or a specific bit depth), the offset may be quantized before signaling. At the decoder side, the decoded offset is dequantized before applying it, as shown in Table 23 below. For example, for a 12-bit sequence, the decoded offset is left shifted (dequantized) by 2. [Table 23]
[0160] In some embodiments, the offset may be calculated as CcSaoOffsetVal=(1-2*ccsao_offset_sign_flag)*(ccsao_offset_abs<<(BitDepth-Min(10,BitDepth))).
[0161] In some embodiments, the sample processing is described below. Let R(x,y) be the input luma or chroma sample value before CCSAO, and R’(x,y) be the output luma or chroma sample value after CCSAO: offset = ccsao_offset[class_index of R(x,y)] R’(x,y) = Clip3(0, (1<<bit_depth)-1, R(x,y) + offset)
[0162] According to the above formula, each luma or chroma sample value R(x,y) is classified using the classifier indicated by the current picture and / or the current offset set idx. The corresponding offset of the derived classification index is added to each luma or chroma sample value R(x,y). The clip function Clip3 is applied to (R(x,y) + offset) to ensure that the output luma or chroma sample value R’(x,y) is within the dynamic range of the bit depth, for example, in the range of 0 to (1<<bit_depth)-1.
[0163] In some embodiments, the boundary processing is described below. If any of the co-located and adjacent luma (chroma) samples used for classification are outside the current picture, the CCSAO is not applied to the current chroma (luma) sample. FIG. 13A is a block diagram illustrating that if any of the co-located and adjacent luma (chroma) samples used for classification are outside the current picture, the CCSAO is not applied to the current chroma (luma) sample according to some implementations of the present disclosure. For example, in FIG. 13A(a), if a classifier is used, the CCSAO is not applied to the chroma components in the left one column of the current picture. For example, if C1' is used, the CCSAO is not applied to the chroma components in the left one column and the top one row of the current picture, as shown in FIG. 13A(b).
[0164] FIG. 13B is a block diagram illustrating that a CCSAO is applied to a current luma or chroma sample if any of the collocated and adjacent luma or chroma samples used for classification are outside the current picture, according to some implementations of the present disclosure. In some embodiments, if any of the collocated and adjacent luma or chroma samples used for classification are outside the current picture, the missing samples may be repeated as shown in FIG. 13B(a) or the missing samples may be mirror padded to create samples for classification as shown in FIG. 13B(b), and a CCSAO may be applied to the current luma or chroma sample. In some embodiments, if any of the collocated and adjacent luma (chroma) samples used for classification are outside the current subpicture / slice / tile / patch / CTU / 360 virtual boundary, the invalidation / repeat / mirror picture boundary processing methods disclosed herein may also be applied to the subpicture / slice / tile / CTU / 360 virtual boundary.
[0165] For example, a picture may be divided into one or more tile rows and one or more tile columns, where a tile is a sequence of CTUs that covers a rectangular region of the picture.
[0166] A slice consists of an integer number of complete tiles, or an integer number of consecutive complete CTU rows within a tile of a picture.
[0167] A subpicture contains one or more slices that collectively cover a rectangular region of the picture.
[0168] In some embodiments, 360-degree video is captured on a sphere and does not have inherently "borders", and reference samples that fall outside the boundaries of the reference picture in the projection domain may always be obtained from neighboring samples in the sphere domain. For projection formats that consist of multiple faces, discontinuities between two or more adjacent faces will appear in the frame-packed picture, no matter how compact the frame-packing arrangement is used. In VVC, vertical and / or horizontal virtual boundaries are introduced where in-loop filtering operations are disabled, and the location of those boundaries is signaled in either the SPS or the picture header. Compared to using two tiles, one for each set of consecutive faces, the use of 360 virtual boundaries is more flexible, since the face size does not need to be a multiple of the CTU size. In some embodiments, the maximum number of vertical 360 virtual boundaries is 3, and the maximum number of horizontal 360 virtual boundaries is also 3. In some embodiments, the distance between two virtual boundaries is equal to or greater than the CTU size, and the virtual boundary granularity is 8 luma samples, e.g., an 8x8 sample grid.
[0169] FIG. 14 is a block diagram illustrating that the CCSAO is not applied to a current chroma sample if the corresponding selected collocated or adjacent luma sample used for classification is outside the virtual space defined by the virtual boundary, according to some implementations of the disclosure. In some embodiments, the virtual boundary (VB) is a virtual line separating spaces in a picture frame. In some embodiments, when the virtual boundary (VB) is applied to the current frame, the CCSAO is not applied to chroma samples that selected corresponding luma positions outside the virtual space defined by the virtual boundary. FIG. 14 shows an example with a virtual boundary for a C0 classifier with nine luma position candidates. For each CTU, the CCSAO is not applied to chroma samples whose corresponding selected luma positions are outside the virtual space bounded by the virtual boundary. For example, in FIG. 14(a), the CCSAO is not applied to chroma sample 1402 if the selected Y7 luma sample position is on the other side of the horizontal virtual boundary 1406 located four pixel lines from the bottom of the frame. For example, in FIG. 14(b), if the selected Y5 luma sample position is located on the other side of the vertical virtual boundary 1408, which is y pixel lines from the right side of the frame, then no CCSAO is applied to the chroma sample 1404.
[0170] FIG. 15 illustrates repetitive or mirror padding that may be applied to luma samples outside the virtual boundary according to some implementations of the present disclosure. FIG. 15(a) illustrates an example of repetitive padding. If the original Y7 is selected to be the classifier located below VB 1502, the Y4 luma sample value is used for classification (copied to the Y7 position) instead of the original Y7 luma sample value. FIG. 15(b) illustrates an example of mirror padding. If Y7 is selected to be the classifier located below VB 1504, the Y1 luma sample value, which is symmetrical to the Y7 value with respect to the Y0 luma sample, is used for classification instead of the original Y7 luma sample value. The padding method may achieve more coding gain because more chroma samples give the possibility to apply CCSAO.
[0171] In some embodiments, restrictions may be applied to reduce the CCSAO required line buffer and simplify boundary processing condition checks. Figure 16 shows that if all nine collocated adjacent luma samples are used for classification according to some implementations of the present disclosure, an additional luma line buffer may be required, namely, all line luma samples of line-5 above the current VB 1602. Figure 10B(a) shows an example using only six luma candidates for classification, which reduces the line buffer and does not require any of the additional boundary checks of Figures 13A and 13B.
[0172] In some embodiments, using luma samples for CCSAO classification may increase the luma line buffer and therefore increase the implementation cost of the decoder hardware. FIG. 17 shows a diagram of an AVS where nine luma candidate CCSAOs across VB 1702 may increase two additional luma line buffers according to some implementations of the present disclosure. For luma and chroma samples that are beyond the virtual boundary (VB) 1702, DBF / SAO / ALF are processed in the current CTU row. For luma and chroma samples that are below VB 1702, DBF / SAO / ALF are processed in the next CTU row. In the AVS decoder hardware design, the pre-DBF samples of luma lines -4 to -1, the pre-SAO samples of line -5, and the pre-DBF samples of chroma lines -3 to -1, the pre-SAO samples of line -4 are stored as line buffers for the next CTU row DBF / SAO / ALF processing. When processing the next CTU row, luma and chroma samples that are not in the line buffer cannot be used. However, for example, in the chroma line-3(b) position, the chroma samples are processed in the next CTU row, but the CCSAO needs pre-SAO luma sample lines-7,-6, and-5 for classification. Pre-SAO luma sample lines-7 and-6 cannot be used because they are not in the line buffer. Also, adding pre-SAO luma sample lines-7 and-6 to the line buffer increases the implementation cost of the decoder hardware. In some cases, luma VB (line-4) and chroma VB (line-3) may be different (not aligned).
[0173] Similar to FIG. 17, FIG. 18A shows a diagram of VVC where nine luma candidate CCSAOs across VB 1802 may increase one additional luma line buffer, according to some implementations of this disclosure. VB may be different in different standards. In VVC, luma VB is line -4 and chroma VB is line -2, so nine candidate CCSAOs may increase one luma line buffer.
[0174] In some embodiments, in a first solution, the CCSAO is disabled for a chroma sample if any of the luma candidates for the chroma sample crosses VB (outside the current chroma sample VB). Figures 19A-19C show that in AVS and VVC, the CCSAO is disabled for a chroma sample if any of the luma candidates for the chroma sample crosses VB 1902 (outside the current chroma sample VB) according to some implementations of the present disclosure. Figure 14 also shows some examples of this implementation.
[0175] In some embodiments, in the second solution, for "crossed VB" luma candidates, repeated padding is used for CCSAO from the luma line opposite VB, e.g., the luma line close to luma line -4. In some embodiments, repeated padding from the luma nearest neighbor below VB is implemented for "crossed VB" chroma candidates. Figures 20A-20C show that in AVS and VVC, if any of the luma candidates of a chroma sample crosses VB 2002 (is outside the current chroma sample VB), CCSAO is enabled with repeated padding of the chroma sample according to some implementations of this disclosure. Figure 14(a) also shows some examples of this implementation.
[0176] In some embodiments, in the third solution, mirror padding is used for CCSAO from below the luma VB for "cross-VB" luma candidates. Figures 21A-21C show that in AVS and VVC, CCSAO is enabled with mirror padding of chroma samples if any of the luma candidates of the chroma samples crosses the VB 2102 (is outside the current chroma sample VB) according to some implementations of the present disclosure. Figures 14(b) and 13B(b) also show some examples of this implementation. In some embodiments, in the fourth solution, "two-sided symmetric padding" is used to apply the CCSAO. Figures 22A-22B show that CCSAO is enabled with two-sided symmetric padding for some examples of different CCSAO shapes (e.g., 9 luma candidates (Figure 22A) and 8 luma candidates (Figure 22B)) according to some implementations of the present disclosure. For luma sample sets with a central luma sample collocated with a chroma sample, bilateral symmetric padding is applied to both sides of the luma sample set if one side of the luma sample set is outside of VB 2202. For example, in Figure 22A, luma samples Y0, Y1, and Y2 are outside of VB 2202, so both Y0, Y1, Y2, and Y6, Y7, Y8 are padded with Y3, Y4, Y5. For example, in Figure 22B, luma sample Y0 is outside of VB 2202, so Y0 is padded with Y2, and Y7 is padded with Y5.
[0177] 18B shows a diagram when collocated and adjacent chroma samples are used to classify the current luma sample, and the selected chroma candidate may span VB, requiring additional chroma line buffers according to some implementations of the present disclosure. Solutions 1 to 4 similar to those described above may be applied to handle this problem.
[0178] Solution 1 is to disable CCSAO for luma samples if any of their chroma candidates may cross VB.
[0179] Solution 2 is to use iterative padding from chroma nearest neighbors below VB for the "cross VB" chroma candidates.
[0180] Solution 3 is to use mirror padding from below chroma VB for the "cross VB" chroma candidate.
[0181] Solution 4 is to use "two-sided symmetric padding". For a candidate set centered on a CCSAO collocated chroma sample, if one side of the candidate set is outside the VB, then two-sided symmetric padding is applied to both sides.
[0182] The padding method may achieve more coding gain since it provides more luma or chroma samples the possibility to apply CCSAO.
[0183] In some embodiments, for bottom picture (or slice, tile, brick) boundary CTU rows, the samples below VB are processed in the current CTU row, so the above special handling (Solution 1, Solution 2, Solution 3, Solution 4) is not applied to the bottom picture (or slice, tile, brick) boundary CTU row. For example, a 1920x1080 frame is divided by 128x128 CTUs. The frame contains 15x9 CTUs (rounded up). The bottom CTU row is the 15th CTU row. The decoding process is performed CTU row by CTU row, and CTU by CTU in each CTU row. Deblocking needs to be applied along the horizontal CTU boundary between the current CTU row and the next CTU row. Inside one CTU, in the bottom 4 / 2 luma / chroma line, CTB VB is applied to each CTU row, because the DBF samples (VVC case) are processed in the next CTU row and are not available in the CCSAO of the current CTU row. However, for the bottom CTU row of the picture frame, since there is no next CTU row remaining and DBF processing is performed on the current CTU row, the bottom 4 / 2 luma / chroma line DBF samples are available for the current CTU row.
[0184] In some embodiments, the VBs displayed in Figures 13-22 may be replaced with the boundaries of the subpicture / slice / tile / patch / CTU / 360 virtual boundary. In some embodiments, the positions of the chroma and luma samples in Figures 13-22 may be switched. In some embodiments, the positions of the chroma and luma samples in Figures 13-22 may be replaced with the positions of the first and second chroma samples. In some embodiments, the ALF VBs in the CTU may be generally horizontal. In some embodiments, the boundaries of the subpicture / slice / tile / patch / CTU / 360 virtual boundary may be horizontal or vertical.
[0185] In some embodiments, restrictions may be applied to reduce the CCSAO required line buffer and simplify boundary processing condition checks, as described in Figure 16. Figure 23 illustrates a restriction to use a limited number of luma candidates for classification, according to some implementations of the present disclosure. Figure 23(a) illustrates a restriction to use only six luma candidates for classification. Figure 23(b) illustrates a restriction to use only four luma candidates for classification.
[0186] In some embodiments, an adaptation region is implemented. The CCSAO adaptation region unit may be CTB-based, i.e., on / off control, CCSAO parameters (offsets used for classification, luma candidate positions, band_num, bitmask, etc., offset set index) are the same in one CTB.
[0187] In some embodiments, the application region may not be aligned to the CTB boundary. For example, the application region is not aligned to the chroma CTB boundary, but is shifted. The syntax (on / off control, CCSAO parameters) is still signaled for each CTB, but the truly applied region is not aligned to the CTB boundary. Figure 24 shows that the CCSAO application region is not aligned to the coding tree block CTB / CTU boundary 2406 according to some implementations of the present disclosure. For example, the application region is not aligned to the chroma CTB / CTU boundary 2406, but is shifted (4,4) samples up and left with respect to the VB 2408. This non-aligned CTB boundary design benefits the deblocking process because the same deblocking parameters are used for each 8x8 deblocking region.
[0188] In some embodiments, the CCSAO application region unit (mask size) can vary (larger or smaller than CTB size) as shown in Table 24. Mask size can vary per component. Mask size may be switched at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, in PH, a set of mask on / off flags and offset set index are signaled to indicate each CCSAO region information. [Table 24]
[0189] In some embodiments, the CCSAO application region frame partition may be fixed, e.g., partitioning the frame into N regions. Figure 25 illustrates that the CCSAO application region frame partition may be fixed using CCSAO parameters in accordance with some implementations of the present disclosure.
[0190] In some embodiments, each region may have its own region on / off control flag and CCSAO parameters. Also, if the region size is larger than the CTB size, it may have both a CTB on / off control flag and a region on / off control flag. Figures 25(a) and 25(b) show some examples of partitioning a frame into N regions. Figure 25(a) shows a vertical partitioning of four regions. Figure 25(b) shows a square partitioning of four regions. In some embodiments, the CTB on / off flag may be further signaled if the region on / off control flag is off, as well as all picture level CTB on control flags (ph_cc_sao_cb_ctb_control_flag / ph_cc_sao_cr_ctb_control_flag). Otherwise, the CCSAO applies to all CTBs in this region without further signaling of the CTB flag.
[0191] In some embodiments, different CCSAO application regions may share the same region on / off control and CCSAO parameters. For example, in Figure 25(c), Region 0 to Region 2 share the same parameters, and Region 3 to Region 15 share the same parameters. Figure 25(c) also shows the region on / off control flags, and the CCSAO parameters may be signaled in Hilbert scan order.
[0192] In some embodiments, the CCSAO application region unit may be quad-tree / binary-tree / ternary-tree partitioned from the picture / slice / CTB level. Similar to the CTB partition, a series of partition flags are signaled to indicate the CCSAO application region partitioning. Figure 26 shows that the CCSAO application region may be binary-tree (BT) / quad-tree (QT) / ternary-tree (TT) partitioned from the frame / slice / CTB level according to some implementations of the present disclosure.
[0193] FIG. 27 is a block diagram illustrating multiple classifiers used and switched at different levels in a picture frame according to some implementations of the present disclosure. In some embodiments, when multiple classifiers are used in one frame, the method of applying the classifier set index may be switched at SPS / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level. For example, four sets of classifiers are used in a frame, which are switched at PH, as shown in Table 25 below. FIG. 27(a) and FIG. 27(c) show the default fixed region classifiers. FIG. 27(b) shows that the classifier set index is signaled at the mask / CTB level, where 0 means CCSAO off for this CTB, and 1-4 means set index. [Table 25]
[0194] In some embodiments, for a default region, a region level flag may be signaled if the CTB in this region does not use the default set index (e.g., the region level flag is 0) but uses other classifier sets in this frame. As an example, if the default set index is used, the region level flag is 1. For example, in a square partition 4 region, the following classifier sets are used, as shown in Table 26 below: [Table 26]
[0195] FIG. 28 is a block diagram illustrating that the CCSAO applied region partitions according to some implementations of the present disclosure are dynamic and can be switched at the picture level. For example, FIG. 28(a) shows that three CCSAO offset sets are used in this POC (set_num=3), so the picture frame is divided vertically into three regions. FIG. 28(b) shows that four CCSAO offset sets are used in this POC (set_num=4), so the picture frame is divided horizontally into four regions. FIG. 28(c) shows that three CCSAO offset sets are used in this POC (set_num=3), so the picture frame is raster partitioned into three regions. Each region may have its own region all-on flag to save the on / off control bits per CTB. The number of regions depends on the signaled picture set_num.
[0196] The CCSAO application region may be a specific area depending on the coding information in the block (such as sample position, sample coding mode, loop filter parameters, etc.). For example, (1) the CCSAO application region may apply only if the sample is skip mode coded, or (2) the CCSAO application region includes only N samples along the CTU boundary, or (3) the CCSAO application region includes only samples on an 8×8 grid in a frame, or (4) the CCSAO application region includes only DBF filtered samples, or (5) the CCSAO application region includes only the top M rows and left N rows in a CU. Different application regions may use different classifiers. Different application regions may use different classifiers. For example, in a CTU, skip mode uses C1, 8×8 grid uses C2, and skip mode and 8×8 grid use C3. For example, in a CTU, skip mode coded samples use C1, CU-centered samples use C2, and CU-centered skip mode coded samples use C3. FIG. 29 is a diagram illustrating that the CCSAO classifier can take into account current component coding information or cross-component coding information, according to some implementations of the present disclosure. For example, different coding modes / parameters / sample positions may form different classifiers. Different coding information may be combined to form a joint classifier. Different areas may use different classifiers. Also shown in FIG. 29 is another example of a coating region.
[0197] In some embodiments, the CCSAO syntax implemented is shown in Table 27 below. In some examples, the binarization of each syntax element may be changed. In AVS3, the term patch is similar to slice, and patch header is similar to slice header. FLC stands for fixed length code. TU stands for truncated unary code. EGk stands for exponential-Golomb code with order k, where k may be fixed. SVLC stands for signed EG0. UVLC stands for unsigned EG0. [Table 27]
[0198] If the higher level flags are off, the lower level flags may be inferred from the off state of the flags and do not need to be signaled. For example, if ph_cc_sao_cb_flag is false in this picture, then ph_cc_sao_cb_band_num_minus1, ph_cc_sao_cb_luma_type, cc_sao_cb_offset_sign_flag, cc_sao_cb_offset_abs, ctb_cc_sao_cb_flag, cc_sao_cb_merge_left_flag, and cc_sao_cb_merge_up_flag are inferred to be not present and false.
[0199] In some embodiments, the SPS ccsao_enabled_flag is conditioned on the SPS SAO enabled flag as shown in Table 28 below. [Table 28]
[0200] In some embodiments, ph_cc_sao_cb_ctb_control_flag, ph_cc_sao_cr_ctb_control_flag indicate whether to enable Cb / Cr CTB on / off control granularity. If ph_cc_sao_cb_ctb_control_flag and ph_cc_sao_cr_ctb_control_flag are enabled, ctb_cc_sao_cb_flag and ctb_cc_sao_cr_flag may be further signaled. Otherwise, whether CCSAO is applied in the current picture depends on ph_cc_sao_cb_flag, ph_cc_sao_cr_flag without further signaling ctb_cc_sao_cb_flag and ctb_cc_sao_cr_flag at the CTB level.
[0201] In some embodiments, for ph_cc_sao_cb_type and ph_cc_sao_cr_type, a flag may be further signaled to distinguish whether the centrally collocated luma position is used for classification of chroma samples (Y0 position in FIG. 10) in order to reduce bit overhead. Similarly, when cc_sao_cb_type and cc_sao_cr_type are signaled at the CTB level, a flag may be further signaled with the same mechanism. For example, when the number of C0 luma position candidates is 9, cc_sao_cb_type0_flag is further signaled to distinguish whether the centrally collocated luma position is used, as shown in Table 29 below. If the centrally collocated luma position is not used, cc_sao_cb_type_idc is used to indicate which of the remaining 8 adjacent luma positions is used. [Table 29]
[0202] Table 30 below shows an example of an AVS where a single (set_num=1) or multiple (set_num>1) classifiers are used within a frame. Note that the syntax notation may be mapped to the notation used above. [Table 30(1)] [Table 30(2)]
[0203] When combined with FIG. 25 or FIG. 27 where each region has its own set, the example syntax may include region on / off control flags (picture_ccsao_lcu_control_flag[compIdx][setIdx]) as shown in Table 31 below. [Table 31]
[0204] In some embodiments, the extension to intra and inter post-prediction SAO filters is further illustrated below. In some embodiments, the SAO classification method disclosed in this disclosure can function as a post-prediction filter, and the prediction can be intra, inter, or other prediction tools such as Intra Block Copy. Figure 30 is a block diagram illustrating the SAO classification method disclosed in this disclosure functioning as a post-prediction filter according to some implementations of the present disclosure.
[0205] In some embodiments, for each Y, U, and V component, a corresponding classifier is selected. Then, for each component prediction sample, it is first classified and a corresponding offset is added. For example, each component may use the current sample and adjacent samples for classification. As shown in Table 32 below, Y uses the current Y sample and adjacent Y samples, and U / V uses the current U / V sample for classification. Figure 31 is a block diagram illustrating that for a post-prediction SAO filter according to some implementations of the present disclosure, each component can use the current sample and adjacent samples for classification. [Table 32]
[0206] In some embodiments, the refined prediction samples (Ypred', Upred', Vpred') are updated by adding the corresponding classification offsets and are then used for intra, inter, or other prediction.
[0207] Ypred'=clip3(0,(1< <bit_depth)-1,Ypred+h_Y[i])
[0208] Upred'=clip3(0,(1< <bit_depth)-1,Upred+h_U[i])
[0209] Vpred'=clip3(0,(1< <bit_depth)-1,Vpred+h_V[i])
[0210] In some embodiments, for chroma U and V components, in addition to the current chroma component, the cross component (Y) may be used for further offset classification. An additional cross component offset (h'_U, h'_V) may be added to the current component offset (h_U, h_V), for example, as shown in Table 33 below. [Table 33]
[0211] In some embodiments, the refined prediction samples (Upred", Vpred") are updated by adding the corresponding classification offsets and are then used for intra, inter, or other prediction.
[0212] Upred”=clip3(0,(1< <bit_depth)-1,Upred’+h’_U[i])
[0213] Vpred”=clip3(0,(1< <bit_depth)-1,Vpred’+h’_V[i])
[0214] In some embodiments, intra prediction and inter prediction may use different SAO filter offsets.
[0215] FIG. 32 is a flowchart illustrating an example process 3200 for decoding a video signal using cross-component correlation according to some implementations of the disclosure.
[0216] Video decoder 30 (as shown in FIG. 3) receives (3210) a picture frame from a video signal, the picture frame including a first component and a second component.
[0217] Video decoder 30 determines (3220) a classifier for the second component from a set of one or more samples of the first component associated with each sample of the second component.
[0218] Video decoder 30 determines whether to modify values of each sample of the second component within the region of the picture frame according to the classifier (3230). In some embodiments, the region is formed by dividing the picture frame.
[0219] In response to determining to modify values of each sample of the second component in the region according to the classifier, video decoder 30 determines sample offsets for each sample of the second component according to the classifier (3240).
[0220] Video decoder 30 modifies (3250) the values of each sample of the second component based on the determined sample offset.
[0221] In some embodiments, the first component is a luma component and the second component is a first chroma component, the first component is a first chroma component and the second component is a luma component, or the first component is a first chroma component and the second component is a second chroma component.
[0222] In some embodiments, the region is a region selected from the group consisting of a subpicture, a slice, a tile, a patch, a coding tree unit (CTU), and a 360 virtual boundary.
[0223] In some embodiments, the boundaries of the regions are aligned with the CTU boundaries.
[0224] In some embodiments, the boundaries of the region are aligned to the boundaries of an 8x8 sample grid in the picture frame.
[0225] In some embodiments, the region includes a region boundary that includes a first boundary of a first component and a second boundary of a second component.
[0226] In some embodiments, the first boundary is not aligned with the second boundary.
[0227] In some embodiments, the first boundary is aligned with the second boundary.
[0228] In some embodiments, determining whether to modify the value of the respective sample of the second component within the region of the picture frame in accordance with the classifier (3230) includes: determining to leave the value of the respective sample of the second component within the region of the picture frame unchanged in accordance with the classifier in accordance with a determination that any one sample of the set of one or more samples of the first component associated with the respective sample of the second component is located on a different side of the region boundary relative to the respective sample of the second component.
[0229] In some embodiments, determining whether to modify values of each sample of the second component within the region of the picture frame in accordance with the classifier (3230) includes: replacing the first subset with a second subset from the remaining subset of the set of one or more samples of the first component in accordance with a determination that a first subset of the set of one or more samples of the first component associated with the respective sample of the second component is located on a different side of the region boundary with respect to the respective sample of the second component, and a remaining subset of the set of one or more samples of the first component associated with the respective sample of the second component is located on the same side of the region boundary with respect to the respective sample of the second component; and determining to modify values of each sample of the second component within the region of the picture frame in accordance with the classifier.
[0230] In some embodiments, the second subset from the remaining subset of the set of one or more samples of the first component is from the nearest row or column of the samples of the first component in the remaining subset to the first subset.
[0231] In some embodiments, the second subset from the remaining subsets of the set of one or more samples of the first component is located at a region boundary of the second component or at a symmetrical position of the respective samples relative to the first subset.
[0232] In some embodiments, determining whether to modify values of respective samples of the second component within the region of the picture frame in accordance with the classifier (3230) includes: replacing the set of one or more samples of the first component with a second set of samples of the first component on the same side of the region boundary relative to the respective samples of the second component in accordance with a determination that a set of one or more samples of the first component associated with the respective samples of the second component are located on a different side of the region boundary relative to the respective samples of the second component; and determining to modify values of respective samples of the second component within the region of the picture frame in accordance with the classifier.
[0233] In some embodiments, the second set of samples on the same side of the region boundary to each sample of the second component is from the nearest row or column of samples of the first component that are on the same side of the region boundary to each sample of the second component to a set of one or more samples of the first component.
[0234] In some embodiments, the second set of samples on the same side of the region boundary as each sample of the second component is located at a symmetrical position of the region boundary or each sample of the second component relative to the set of one or more samples of the first component.
[0235] In some embodiments, determining whether to modify values of the respective samples of the second component within the region of the picture frame in accordance with the classifier (3230) includes: replacing the first subset with a second subset of one or more central subsets from the remaining subsets of the set of one or more samples of the first component in accordance with a determination that a first subset of the set of one or more samples of the first component associated with the respective samples of the second component are located on different sides of the region boundary with respect to the respective samples of the second component and a remaining subset of the set of one or more samples of the first component associated with the respective samples of the second component are located on the same side of the region boundary with respect to the respective samples of the second component; replacing a third subset within the remaining subsets located at the boundary position of the set of one or more samples of the first component with the second subset or a fourth subset of the one or more central subsets from the remaining subsets of the set of one or more samples of the first component; and determining to modify values of the respective samples of the second component within the region of the picture frame in accordance with the classifier.
[0236] In some embodiments, the third subset and the first subset are symmetrically located within the set of one or more samples of the first component.
[0237] In some embodiments, determining whether to modify values of each sample of the second component within the region of the picture frame in accordance with the classifier (3230) includes: determining not to modify values of each sample of the second component within the region of the picture frame in accordance with the classifier in response to determining that the chroma format of the video signal is 4:0:0.
[0238] FIG. 33 illustrates a computing environment 3310 coupled with a user interface 3350 . The computing environment 3310 may be part of a data processing server. The computing environment 3310 includes a processor 3320, a memory 3330, and an input / output (I / O) interface 3340.
[0239] The processor 3320 typically controls the overall operation of the computing environment 3310, such as operations related to display, data collection, data communication, and image processing. The processor 3320 may include one or more processors for executing instructions for performing all or a portion of the steps in the methods described above. Additionally, the processor 3320 may include one or more modules that facilitate interaction between the processor 3320 and other components. The processor may be a Central Processing Unit (CPU), a microprocessor, a single chip machine, a Graphical Processing Unit (GPU), etc.
[0240] The memory 3330 is configured to store various types of data to support the operation of the computing environment 3310 . The memory 3330 may include certain software 3332 . Examples of such data include instructions for any application or method operating on the computing environment 3310, video data sets, image data, etc. The memory 3330 may be implemented using any type of volatile or non-volatile memory device, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk, or a combination thereof.
[0241] The I / O interface 3340 provides an interface between the processor 3320 and a peripheral interface module such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 3340 may be coupled to an encoder and a decoder.
[0242] In one embodiment, a non-transitory computer readable storage medium is also provided, comprising a number of programs, e.g., in memory 3330, executable by processor 3320 in computing environment 3310 to perform the above-described methods. Alternatively, the non-transitory computer readable storage medium may store a bitstream or data stream comprising encoded video information (e.g., video information including one or more syntax elements) generated by an encoder (e.g., video encoder 20 of FIG. 2) using the encoding method described above for use by a decoder (e.g., video decoder 30 of FIG. 3) in decoding video data. The non-transitory computer readable storage medium may be, for example, a ROM, a Random Access Memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.
[0243] In one embodiment, there is also provided a computing device including one or more processors (e.g., processor 3320) and a non-transitory computer readable storage medium or memory 3330 having stored thereon a number of programs executable by the one or more processors, where the one or more processors are configured to perform the above-described methods upon execution of the number of programs.
[0244] In one embodiment, a computer program product is also provided that includes a number of programs, e.g., in memory 3330, executable by the processor 3320 in the computing environment 3310, to perform the above-described methods. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0245] In one embodiment, the computing environment 3310 may be implemented with one or more ASICs, DSPs, Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0246] Further embodiments also include various subsets of the above embodiments that are combined or otherwise rearranged in various other embodiments.
[0247] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted as one or more instructions or code to a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium, or a communication medium, which includes any medium that facilitates transfer of a computer program from one place to another, for example according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the implementations described herein. A computer program product may include a computer-readable medium.
[0248] The terms used in the description of the embodiments herein are intended only to describe particular implementations and are not intended to limit the scope of the claims. As used in the description of the implementations and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "comprises" and / or "comprising" as used herein specify the presence of the stated features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or groups thereof.
[0249] It will also be understood that, although terms such as first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first electrode can be referred to as a second electrode, and similarly, a second electrode can be referred to as a first electrode, without departing from the scope of the implementation. Although a first electrode and a second electrode are both electrodes, they are not the same electrode.
[0250] Throughout this specification, reference to the singular or plural "one example," "an example," "exemplary example," etc. means that one or more particular features, structures, or characteristics described in connection with the example are included in at least one example of the present disclosure. Thus, the appearances of phrases such as "in one example," "in an example," "in an exemplary example," etc. in the singular or plural in various places throughout this specification are not necessarily all referring to the same example. Furthermore, particular features, structures, or characteristics in one or more examples may be combined in any suitable manner.
[0251] The description in this application has been presented for purposes of illustration and description and is not intended to be exhaustive or limited to the invention in the disclosed form. Numerous modifications, variations, and alternative implementations will be apparent to those skilled in the art having the benefit of the teachings presented in the foregoing description and the associated drawings. The embodiments have been chosen and described in order to best explain the principles, practical applications of the invention and to enable others skilled in the art to understand the invention in its various implementations and to best utilize the basic principles and various implementations with various modifications suited to the particular use envisaged. It is therefore to be understood that the claims are not limited to the particular examples of implementations disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims.
Claims
1. 1. A method for decoding a video signal, comprising: receiving a picture frame from the video signal, the picture frame including a first component and a second component; determining a classifier for the second component based on characteristic measurements of a set of one or more samples of the first component juxtaposed or adjacent to each sample of the second component; determining whether to modify values of the respective samples of the second component within the region of the picture frame according to the classifier; in response to the determination of modifying values of the respective samples of the second component in the region according to the classifier, determining a sample offset for the respective samples of the second component according to the classifier; modifying the values of the respective samples of the second component based on the determined sample offsets; Including, The regions are formed by dividing the picture frame. method.
2. the first component is a luma component and the second component is a first chroma component; the first component is a first chroma component and the second component is a luma component; or the first component is a first chroma component and the second component is a second chroma component; The method of claim 1.
3. The method of claim 1 , wherein the region is a region selected from the group consisting of a subpicture, a slice, a tile, a patch, a coding tree unit (CTU), and a 360 virtual boundary.
4. The method of claim 3 , wherein boundaries of the regions are aligned with CTU boundaries.
5. The method of claim 3 , wherein boundaries of the region are aligned with boundaries of an 8×8 sample grid in the picture frame.
6. The method of claim 1 , wherein the region includes a region boundary that includes a first boundary of the first component and a second boundary of the second component.
7. The method of claim 6 , wherein the first boundary is not aligned with the second boundary or the first boundary is aligned with the second boundary.
8. Determining whether to modify the values of the respective samples of the second component in the region of the picture frame according to the classifier includes: in response to a determination that any one sample of the set of one or more samples of the first component collocated or adjacent to the respective sample of the second component is located on a different side of the region boundary relative to the respective sample of the second component; determining to leave the values of the respective samples of the second component within the region of the picture frame unchanged according to the classifier; The method of claim 6, comprising:
9. Determining whether to modify the values of the respective samples of the second component in the region of the picture frame according to the classifier includes: a first subset of the set of one or more samples of the first component that are collocated or adjacent to the respective samples of the second component are located on different sides of the region boundary relative to the respective samples of the second component; and A remaining subset of the set of one or more samples of the first component that are collocated or adjacent to each sample of the second component is located on the same side of a region boundary as the respective sample of the second component. According to the judgment, replacing the first subset with a second subset from the remaining subset of the set of one or more samples of the first component; determining to modify the values of the respective samples of the second component in the region of the picture frame according to the classifier; The method of claim 6, comprising:
10. 10. The method of claim 9, wherein the second subset from the remaining subset of the set of one or more samples of the first component is from a nearest row or column of a sample of the first component in the remaining subset to the first subset, or the second subset from the remaining subset of the set of one or more samples of the first component is located at a symmetric position of the region boundary or the respective sample of the second component relative to the first subset.
11. Determining whether to modify the values of the respective samples of the second component in the region of the picture frame according to the classifier includes: in response to determining that the set of one or more samples of the first component collocated or adjacent to the respective sample of the second component is located on a different side of the region boundary relative to the respective sample of the second component; replacing the set of one or more samples of the first component with a second set of samples of the first component on the same side of the region boundary relative to the respective samples of the second component; determining to modify the values of the respective samples of the second component in the region of the picture frame according to the classifier; The method of claim 6, comprising:
12. 12. The method of claim 11 , wherein the second set of samples on the same side of the region boundary to the respective samples of the second component is from a nearest row or column of samples of the first component on the same side of the region boundary to the respective samples of the second component to the set of one or more samples of the first component.
13. 12. The method of claim 11 , wherein the second set of samples on the same side of the region boundary as the respective samples of the second component are located at symmetric positions of the region boundary or the respective samples of the second component relative to the set of one or more samples of the first component.
14. Determining whether to modify the values of the respective samples of the second component in the region of the picture frame according to the classifier includes: in response to a determination that a first subset of the set of one or more samples of the first component collocated or adjacent to the respective sample of the second component are located on different sides of the region boundary relative to the respective sample of the second component, and a remaining subset of the set of one or more samples of the first component collocated or adjacent to the respective sample of the second component are located on the same side of the region boundary relative to the respective sample of the second component; replacing the first subset with a second subset of one or more central subsets from the remaining subsets of the set of one or more samples of the first component; replacing a third subset within the remaining subset at a boundary position of the set of one or more samples of the first component with the second subset or a fourth subset of the one or more central subsets from the remaining subset of the set of one or more samples of the first component; determining to modify the values of the respective samples of the second component in the region of the picture frame according to the classifier; The method of claim 6, comprising:
15. The method of claim 14 , wherein the third subset and the first subset are symmetrically located within the set of one or more samples of the first component.
16. Determining whether to modify the values of the respective samples of the second component in the region of the picture frame according to the classifier includes: in response to determining that the chroma format of the video signal is 4:0:0, determining not to modify the values of the respective samples of the second component within the region of the picture frame according to the classifier; The method of claim 1 , comprising:
17. 1. An electronic device comprising: One or more processing units; a memory coupled to the one or more processing units; a plurality of programs stored in said memory, which, when executed by said one or more processing units, cause said electronic device to perform a method according to any one of claims 1 to 16; Including electronic devices.
18. A method for storing a bitstream, the bitstream being used by a method for decoding a video, the method for decoding a video comprising: receiving a picture frame from a video signal, the picture frame including a first component and a second component; determining a classifier for the second component based on characteristic measurements of a set of one or more samples of the first component juxtaposed or adjacent to each sample of the second component; determining whether to modify values of the respective samples of the second component within the region of the picture frame according to the classifier; in response to the determination of modifying values of the respective samples of the second component in the region according to the classifier, determining a sample offset for the respective samples of the second component according to the classifier; modifying the values of the respective samples of the second component based on the determined sample offsets; Including, The regions are formed by dividing the picture frame. method.
19. A non-transitory computer-readable storage medium storing a plurality of programs executed by an electronic device having one or more processing units, the plurality of programs, when executed by the one or more processing units, causing the electronic device to perform a method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Method and apparatus for optimizing the encoding / decoding of compensation offsets for a set of reconstructed image samples.
JP2014534762A
Cross Component Filter
JP2019525679A
Cross plane filtering for chroma signal enhancement in video coding
JP2020036353A
Systems and methods for reducing a reconstruction error in video coding based on a cross-component correlation
WO2020262396A1
Chroma coding enhancement in cross-component sample adaptive offset with virtual boundary
WO2022093992A1