CCSO with advanced flag

By introducing a cross-component sample offset (CCSO) filter in video encoding and decoding technology, and controlling the filtering process using advanced syntax elements, the problem of large reconstruction errors in the prior art is solved, and video decoding with higher quality and efficiency is achieved.

CN120202665APending Publication Date: 2025-06-24TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004777.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-09
Filing Date
2024-05-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are difficult to effectively reduce reconstruction errors when compressing video data, especially in cross-component sample offsets.

Method used

Using a cross-component sample offset (CCSO) filter, CCSO filter is controlled by high-level syntax elements of the frame and component levels, and the sample offset value of the second color component is calculated using the reconstruction samples from the first color component.

Benefits of technology

By adjusting the reconstructed image samples, the reconstruction error is significantly reduced and the quality and efficiency of video decoding are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120202665A_ABST
    Figure CN120202665A_ABST
Patent Text Reader

Abstract

Various implementations described herein include methods and systems for encoding and decoding video. In one aspect, a video bitstream includes a high-level syntax element for a frame-level sample offset (CCSO) flag indicating whether to apply cross-component (CCSO) filtering to a current image frame. When the frame-level CCSO flag indicates that CCSO filtering is applied to the current image frame, the electronic device identifies, in the video bitstream, a first syntax element of a first component CCSO flag indicating whether to apply CCSO filtering to a first color component of the current image frame, and determines whether to apply CCSO filtering to a first color sample of the first color component based on the first syntax element. A current image frame comprising a first color sample of a first color component is reconstructed.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims the priority of U.S. Provisional Patent Application No. 63 / 543,272, titled "CCSO with High Level Flags", filed on October 9, 2023, and this application is a continuation of and claims the priority of U.S. Patent Application No. 18 / 660,074, titled "CCSO with High Level Flags", filed on May 9, 2024. The entire contents of these applications are incorporated herein by reference. Technical Field

[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for loop filtering (e.g., cross - component offset (CCSO) filtering) of video data. Background Art

[0003] A variety of electronic devices support digital video, such as digital televisions, laptop or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transfer digital video data via a communication network and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited storage resources of the storage device, video coding can be used to compress video data according to one or more video coding standards before the video data is transferred or stored. Video coding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.

[0004] Video coding typically uses prediction methods that exploit the redundancy inherent in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing a degradation in video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4). Versatile Video Coding (VVC / H.266) aims to be a video compression standard that succeeds HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (Version 1) and 2022 (Version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verified version 1.0.0 of the specification containing Errata 1 was released. Summary of the Invention

[0005] As described above, encoding (compression) reduces the bandwidth and / or storage space requirements. As will be described in detail later, both lossless compression and lossy compression can be employed. Lossless compression refers to a technique where an exact copy of the original signal can be reconstructed from the compressed signal via a decoding process. Lossy compression refers to an encoding / decoding process where the original video information is not fully retained during the encoding process and is not fully recovered during the decoding process. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is made small enough such that the reconstructed signal is useful for the intended application. The degree of tolerable distortion depends on the application. For example, users of some consumer video streaming applications may be more tolerant of higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect various distortion tolerances. Generally, higher tolerable distortion allows the use of coding algorithms that incur higher losses but have higher compression ratios.

[0006] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying a loop filter (e.g., a CCSO filter) to video (image) compression. A video codec includes multiple functional modules for one or more of the following operations: intra / inter prediction, transform coding, quantization, entropy coding, and in-loop filtering. In-loop filtering techniques are applied to adjust the reconstructed picture samples to further reduce the reconstruction error. In various embodiments of the present application, a video bitstream is transmitted between a video encoder and a video decoder, and the video bitstream includes a high-level syntax element for a frame-level cross-component sample offset (CCSO) flag. The frame-level CCSO flag indicates whether CCSO filtering is applied to the current picture frame. When the frame-level CCSO flag indicates that CCSO filtering is enabled, the video bitstream may further include another syntax element for a corresponding component CCSO flag. The component CCSO flag indicates whether CCSO filtering is applied to the corresponding color component of the current picture frame. Thus, cross-component offset filtering can be controlled at both the frame level and the component level for the corresponding color components of the current picture frame.

[0007] In some embodiments, a cross-component sample offset (CCSO) filter uses co-located reconstructed samples from a first color component (e.g., luminance samples) and their adjacent reconstructed samples to derive a sample offset value to be added to the current samples of a second color component (e.g., luminance or chrominance samples). In some embodiments, the CCSO filter may include an edge-preserving loop filter that depends on the values of the reconstructed samples to determine the sample offset values for luminance and / or chrominance samples. The CCSO flag for luminance / chrominance sample offset is written separately for each component at the frame level. Additional high-level flags may further be introduced to control the CCSO filtering parameters for different color components.

[0008] According to some embodiments, a video decoding method is provided. The method includes: receiving a video bitstream including a current picture frame. The video bitstream includes a high-level syntax element for a frame-level CCSO flag indicating whether CCSO filtering is applied to the current picture frame. The method further includes determining, based on the high-level syntax element, whether CCSO filtering is applied to the current picture frame. The method further includes: when CCSO filtering is applied to the current picture frame, identifying, in the video bitstream, a first syntax element for a first component CCSO flag, the first component CCSO flag indicating whether CCSO filtering is applied to a first color component of the current picture frame. The method further includes determining, based on the first syntax element, whether CCSO filtering is applied to the first color samples of the first color component of the current picture frame, and reconstructing the current picture frame including the first color samples of the first color component. In some embodiments, the high-level syntax element includes a frame-level syntax element.

[0009] According to some embodiments, a method for video encoding is provided. The method includes: receiving video data including a current picture frame; encoding the current picture frame; transmitting the encoded current picture frame via a video bitstream; and writing, via the video bitstream, a high-level syntax element for a frame-level CCSO flag indicating whether to apply CCSO filtering to the current picture frame. When the frame-level CCSO flag indicates that CCSO filtering is applied to the current picture frame, a first syntax element for a first component CCSO flag is written in the video bitstream, and the first component CCSO flag indicates whether to apply CCSO filtering to a first color component of the current picture frame.

[0010] According to some embodiments, a method for bitstream conversion is provided. The method includes: obtaining a source video sequence including a current picture frame, and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes the current picture frame and a high-level syntax element for a frame-level CCSO flag indicating whether to apply CCSO filtering to the current picture frame. When the frame-level CCSO flag indicates that the CCSO filtering is applied to the current picture frame, a first syntax element for a first component CCSO flag is written in the video bitstream, and the first component CCSO flag indicates whether to apply CCSO filtering to a first color component of the current picture frame.

[0011] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic devices. The computing system includes a control circuit and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.

[0012] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions executed by a computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.

[0013] Thus, devices and systems for video decoding methods are disclosed. Such methods, devices, and systems may supplement or replace conventional methods, devices, and systems for video decoding.

[0014] The features and advantages described in this specification are not necessarily all included. In particular, in view of the drawings, the specification, and the claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. In addition, it should be noted that the language used in the specification is mainly selected for readability and guidance purposes, and is not selected to describe or limit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] For a more detailed understanding of the present disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are shown in the accompanying drawings. However, the accompanying drawings only show the relevant features of the present disclosure and are therefore not necessarily considered restrictive; those skilled in the art will understand that the specification may include other effective features when reading the present disclosure.

[0016] Figure 1 A block diagram showing an example communication system according to some embodiments.

[0017] Figure 2A A block diagram showing example elements of an encoder component according to some embodiments.

[0018] Figure 2B A block diagram showing example elements of a decoder component according to some embodiments.

[0019] Figure 3 A block diagram showing an example server system according to some embodiments.

[0020] Figure 4 A flowchart showing an example process of decoding a video stream using CCSO filtering based on flags written at different levels according to some embodiments.

[0021] Figure 5 A flowchart showing an example process of applying CCSO filtering in in-loop filtering according to some embodiments.

[0022] Figure 6 A flowchart showing a method of decoding a video according to some embodiments.

[0023] By convention, the various features shown in the drawings are not necessarily drawn to scale, and the same reference numerals are used throughout the specification and drawings to denote the same features. Detailed Description

[0024] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying a loop filter to video (image) compression. Intra-loop filtering techniques are applied to adjust the reconstructed picture samples to further reduce the reconstruction error. In various embodiments of the present application, a video bitstream is transmitted between a video encoder and a video decoder, and the video bitstream includes an advanced syntax element for a frame-level cross-component sample offset (CCSO) flag. The frame-level CCSO flag indicates whether CCSO filtering is applied to the current picture frame. When the frame-level CCSO flag indicates that CCSO filtering is enabled, the video bitstream may further include another syntax element for a corresponding component CCSO flag. The component CCSO flag indicates whether CCSO filtering is applied to the corresponding color component of the current picture frame. Thus, cross-component offset filtering can be controlled both at the frame level and the component level for the corresponding color components of the current picture frame.

[0025] In some embodiments, a cross-component sample offset (CCSO) filter uses co-located reconstructed samples from a first color component (e.g., luminance samples) and their neighboring reconstructed samples to derive a sample offset value to be added to the current sample of a second color component (e.g., luminance or chrominance samples). In some embodiments, the CCSO filter may include an edge-preserving loop filter that depends on the values of the reconstructed samples to determine the sample offset values for luminance and / or chrominance samples. The CCSO flag for luminance / chrominance sample offset is written separately for each component at the frame level. Additional advanced flags may further be introduced to control the CCSO filtering parameters for different color components.

[0026] More specifically, in some embodiments, a video decoder identifies a set of luminance samples including a first luminance sample and one or more neighboring luminance samples of the first luminance sample. For example, the luminance samples are quantized using a scalar quantizer to generate one or more quantization values. The scalar quantizer may be specified by a quantization interval (e.g., a range of values assigned to the same integer) and a quantization level (e.g., an integer value assigned to the quantization interval). For example, a classifier classifies the first color sample based on one or more quantization values to determine a first sample offset of the first color sample. The first color sample is adjusted based on the first sample offset of the first color sample such that the current picture frame can be reconstructed.

[0027] Figure 1 A block diagram of a communication system 100 is shown in accordance with some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m), which are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0028] The source device 102 includes a video source 104 (e.g., a camera assembly or a media memory) and an encoder assembly 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder assembly 106 generates one or more encoded video bitstreams based on the video stream. The video stream from the video source 104 may have a higher data volume compared to the encoded video bitstreams 108 generated by the encoder assembly 106. Since the encoded video bitstreams 108 have a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less bandwidth to transmit and less storage space to store compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder assembly 106 (e.g., configured to send uncompressed video to the network 110).

[0029] One or more networks 110 represent any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0030] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as the encoded video stream from the source device 102). The server system 112 includes a codec assembly 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec assembly 114 includes an encoder assembly and / or a decoder assembly. In various embodiments, the codec assembly 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the codec assembly 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to produce encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings for the encoded video bitstream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim the encoded video bitstream 108 for customizing potentially different bitstreams for one or more of the multiple electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0031] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to produce an output video stream that can be presented on a display or other type of presentation device. In some embodiments, one or more of the plurality of electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0032] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the plurality of electronic devices 120 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0033] In an example operation of the communication system 100, the source device 102 sends the encoded video stream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video stream 108 and may use the codec component 114 to decode and / or encode the encoded video stream 108. For example, the server system 112 may apply an encoding that is optimal for network transmission and / or storage to the video data. The server system 112 may send the encoded video data 116 (e.g., one or more encoded video streams) to one or more of the plurality of electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.

[0034] Figure 2ABlock diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source of a component of a device different from the encoder component 106). The video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device storing previously acquired / prepared video. In some embodiments, the video source 104 is a camera that acquires local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial array of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can easily understand the relationship between pixels and samples.

[0035] The encoder component 106 is configured to encode and / or compress pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder component 106 is configured to perform a conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to these units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can easily identify other functions of the controller 204 as they may be related to the encoder component 106 optimized for a specific system design.

[0036] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. As a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and (a) reference picture(s)) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact corresponding between the local encoder and the remote encoder. In this way, the reference picture samples interpreted by the prediction part of the encoder are the same as the sample values that the decoder will interpret when using prediction during decoding. This principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is known to those of ordinary skill in the art.

[0037] The operation of the decoder 210 can be the same as that of the remote decoder (such as the decoder component 122) described in detail below in connection with Figure 2B However, briefly referring to Figure 2B , when the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implementable in the local decoder 210.

[0038] Except for parsing / entropy decoding, the decoder techniques described herein can exist in a corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder techniques can be simplified because encoder techniques can be inverse to decoder techniques.

[0039] As part of its operation, the source encoder 202 can perform motion-compensated predictive coding. Referring to one or more previously encoded frames in the video sequence designated as reference frames, this motion-compensated predictive coding performs predictive coding on the input frame. In this way, the coding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the reference frame, and the reference picture can be selected as the prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.

[0040] The decoder 210 decodes the encoded video data of the frame that can be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the coding engine 212 can be a lossy process. When the encoded video data is in the video decoder (Figure 2A When decoded at (not shown in the figure), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process, which can be performed by the remote video decoder on the reference frames, and can cause the reconstructed reference frames to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frames, which has the same content (without transmission errors) as the reconstructed reference frames to be obtained by the remote video decoder.

[0041] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search in the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata that can be used as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor 206 can operate on a per-pixel-block basis of the sample blocks to find a suitable prediction reference. Based on the search results obtained by the predictor 206, it can be determined that the input picture can have a prediction reference taken from multiple reference pictures stored in the reference picture memory 208.

[0042] The outputs of all the above functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by various functional units according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding), thereby converting the symbols into an encoded video sequence.

[0043] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer the encoded video sequence created by the entropy encoder 214, so as to prepare for transmission through the communication channel 218, which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter can be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter can transmit additional data when transmitting the encoded video. The source encoder 202 can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and redundant slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0044] The controller 204 can manage the operation of the encoder component 106. During encoding, the controller 204 can assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics, and thus will not be elaborated here. Predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0045] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictive encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictive encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non-predictively encoded through spatial prediction or through temporal prediction with reference to one previously encoded reference picture. Blocks of a B picture can be non-predictively encoded through spatial prediction or through temporal prediction with reference to one or two previously encoded reference pictures.

[0046] The captured video can be a plurality of source pictures (video pictures) in a time series. Intra picture prediction (often simplified to intra prediction) utilizes the spatial correlation within a given picture, while inter picture prediction utilizes the (temporal or other) correlation between pictures. In an embodiment, the particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0047] The encoder component 106 may perform encoding operations according to a predetermined video encoding technique or standard, such as any of the standards described herein. In operation, the encoder component 106 may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancies in the input video sequence. Accordingly, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.

[0048] Figure 2B FIG. is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256, the transmitter being configured to send data to the display 124 (e.g., via a wired or wireless connection).

[0049] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218, the receiver being configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from the channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not shown). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data when receiving the encoded video. This additional data may be part of the encoded video sequence. The decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0050] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The decoder component 122 may be implemented at least in part in software.

[0051] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to prevent network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 inside the decoder component 122 (e.g., configured to handle playback timing), a separate buffer memory is provided outside the decoder component 122 (e.g., to prevent network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous synchronous network, it may also not be necessary to configure the buffer memory 252, or the buffer memory can be made smaller. For use on a traffic packet network such as the Internet, the buffer memory 252 may also be required, and the buffer memory can be relatively large and / or have an adaptable size, and can be implemented at least partially in the operating system or a similar element outside the decoder component 122.

[0052] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols can include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a presentation device (e.g., the display 124). The control information for the presentation device can be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set segment (not labeled). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be performed according to a video coding technology or standard, and can follow various principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroup can include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser 254 can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0053] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter - picture and intra - picture, inter - block and intra - block) and other factors, the reconstruction of symbol 270 may involve multiple different units. Which units are involved and the way they are involved can be controlled by subgroup control information parsed by parser 254 from the encoded video sequence. For clarity, such subgroup control information flows between parser 254 and the multiple units below are not described.

[0054] Decoder component 122 can be conceptually divided into multiple functional units, and in some embodiments, these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptually divided functional units are retained here.

[0055] Scaler / inverse transform unit 258 receives the quantized transform coefficients as symbol 270 and control information from parser 254, including which transform mode to use, block size, quantization factor, and / or quantization scaling matrix, etc. Scaler / inverse transform unit 258 can output a block including sample values, which can be input into aggregator 268.

[0056] In some cases, the output samples of scaler / inverse transform unit 258 belong to intra - coded blocks; that is: blocks that do not use predictive information from previously reconstructed pictures, but may use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by intra - picture prediction unit 262. Intra - picture prediction unit 262 can generate surrounding blocks with the same size and shape as the block being reconstructed using the reconstructed information extracted from the current (partially reconstructed) picture in current picture memory 264. Aggregator 268 can add the predictive information generated by intra - picture prediction unit 262 to the output sample information provided by scaler / inverse transform unit 258 based on each sample.

[0057] In other cases, the output samples of scaler / inverse transform unit 258 belong to inter - coded and potentially motion - compensated blocks. In this case, motion - compensation prediction unit 260 can access reference picture memory 266 to extract samples for prediction. After motion - compensating the extracted samples according to symbol 270 related to the block, these samples can be added by aggregator 268 to the output of scaler / inverse transform unit 258 (which is called residual samples or residual signal in this case), thus generating output sample information. The acquisition of prediction samples by motion - compensation prediction unit 260 from addresses within reference picture memory 266 can be controlled by a motion vector. The motion vector can be provided to motion - compensation prediction unit 260 in the form of said symbol 270, which is, for example, including X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from reference picture memory 266, motion vector prediction mechanisms, etc. when using sub - sample accurate motion vectors.

[0058] The output samples of aggregator 268 can be adopted by various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in the encoded video bitstream, and the parameters can be used for loop filter unit 256 as symbols 270 from parser 254. However, video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) parts of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. The output of loop filter unit 256 can be a sample stream, which can be output to a rendering device such as display 124 and stored in reference picture memory 266 for subsequent inter-picture prediction.

[0059] Once reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is reconstructed and the encoded picture (by parser 254, for example) is identified as a reference picture, the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0060] Decoder component 122 can perform decoding operations according to predetermined video compression techniques recorded in a standard (such as any of the standards described herein). The encoded video sequence can conform to the syntax of the video compression technique or standard specified in the sense that the encoded video sequence follows the video compression technique or standard (especially the profile thereof) specified in the video compression technique or standard. In addition, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of HRD buffer management signaled in the encoded video sequence.

[0061] Figure 3A block diagram of a server system 112 according to some embodiments is shown. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).

[0062] The network interface 304 may be configured to interact with one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication network may be a local area network, a wide area network, a metropolitan area network, a vehicular and industrial network, a real-time network, a delay tolerant network, etc. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial networks including CANBus, etc. Such communication may be unidirectional receive only (e.g., broadcast TV), unidirectional transmit only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., to other computer systems using a local area network or a wide area digital network). Such communication may include communication to one or more cloud computing networks.

[0063] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 may include one or more of the following devices: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 may include one or more of the following devices: audio output devices (e.g., speakers), visual output devices (e.g., displays or viewers), etc.

[0064] The memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices located remotely from the control circuit 302. The memory 314 or the non-volatile solid-state memory device within the memory 314 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures or subsets or supersets thereof: · An operating system 316, including programs for handling various basic system services and for performing hardware-related tasks; · A network communication module 318, for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); · A codec module 320, for performing various functions regarding encoding and / or decoding data (such as video data). In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: ○ A decoding module 322, for performing various functions regarding decoding encoded data, such as those functions previously described regarding the decoder component 122; and ○ An encoding module 340, for performing various functions regarding encoding data, such as those functions previously described regarding the encoder component 106; and · A picture memory 352, for storing pictures and picture data, e.g., for use with the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0065] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described regarding the parser 254), a transform module 326 (e.g., configured to perform various functions previously described regarding the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions previously described regarding the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described regarding the loop filter 256).

[0066] In some embodiments, the encoding module 340 includes an encoder module 342 (e.g., configured to perform various functions previously described regarding the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described regarding the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the modules shown. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0067] Each of the above-identified modules stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but instead uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores subsets of the above-identified modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0068] Although Figure 3 FIG. 112 shows a server system 112 according to some embodiments, Figure 3 it is more intended as a functional description of the various features that may exist in one or more server systems rather than a structural schematic of the embodiments described herein. In practice, and as will be appreciated by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in FIG. 112 may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among them will vary depending on the implementation and optionally partly depend on the amount of data traffic processed by the server system during peak usage periods as well as during average usage periods.

[0069] Figure 4 FIG. 13 is a flowchart of an example process 400 for decoding a video bitstream 116 using CCSO filtering 414 based on flags written at different levels according to some embodiments. The GOP includes a sequence of picture frames, which further includes a current picture frame 406C. The current picture frame 406C includes a color picture, i.e., a non-monochrome picture frame, which has a plurality of co-located color samples (e.g., chroma samples 402 and luma samples 404). After reconstructing the plurality of color samples of the current picture frame 406C, in-loop filtering is applied to adjust a subset of the color samples to improve the picture quality of the current picture frame 406C. An example of in-loop filtering is CCSO filtering. In some embodiments associated with CCSO filtering, the reconstructed samples of the luma component and / or its adjacent reconstructed samples (e.g., luma samples 404) are combined to derive an offset value 422 for a first color component 410 (e.g., luma samples 404, chroma samples 402), and the reconstructed samples of the first color component 410 are co-located with the reconstructed samples of the luma component and are adjusted by the offset value 422. The first color component 410 is optionally the same as or different from the luma component.

[0070] In some embodiments, decoder 122 receives video bitstream 116 including current picture frame 406C. Video bitstream 116 includes high-level syntax element 408 for frame-level cross-component sample offset (CCSO) flag 412, and the frame-level CCSO flag indicates whether CCSO filtering 414 is applied to current picture frame 406C. Based on high-level syntax element 408, decoder 122 determines whether to apply CCSO filtering 414 to current picture frame 406C. When frame-level CCSO flag 412 indicates that CCSO filtering 414 is applied to current picture frame 406C, video decoder 122 identifies first syntax element 416 for first-component CCSO flag 418 in video bitstream 116. First-component CCSO flag 418 indicates whether CCSO filtering 414 is applied to first color component 410 of current picture frame 406C, and based on first syntax element 416, determines whether to apply CCSO filtering 414 to first color samples 410C of first color component 410 (e.g., to determine sample offset 422 for first color samples 410C). In other words, in some embodiments, video bitstream 116 may send first syntax element 416 as needed (e.g., when frame-level CCSO flag 412 is enabled). In some embodiments, first color component 410 is one of luminance component 404, blue-difference chrominance component 402B, and red-difference chrominance component 402R. Video decoder 122 reconstructs current picture frame 406C including first color samples 410C of first color component 410.

[0071] In some embodiments, when frame-level CCSO flag 412 indicates that CCSO filtering is disabled for first picture frame 406-1 (e.g., different from current picture frame 406C), video decoder 122 sets a distinct-component CCSO flag to indicate that CCSO filtering 414 is not applied to first color component 410 of first picture frame 406-1, independent of whether any syntax element for the distinct-component CCSO flag of first picture frame 406-1 is written into video bitstream 116. In some embodiments, the distinct-component CCSO flag may be written in first syntax element 416 for first picture frame 406-1. Conversely, in some embodiments, the distinct-component CCSO flag is not written into video bitstream 116, but is locally determined by video decoder 122 for first picture frame 406-1.

[0072] In some embodiments, after determining that the frame-level CCSO flag 412 indicates that CCSO filtering 414 is to be applied to the current image frame 406C, the video decoder 122 determines that the first-component CCSO flag 418 indicates that CCSO filtering 414 is not to be applied to the first color component 410. The video decoder 122 aborts generating the first sample offset 422 for the first color samples 410C of the first color component 410.

[0073] In some embodiments, the video bitstream 112 includes a first syntax element 416 for the first-component CCSO flag 418. When the first-component CCSO flag 418 indicates that CCSO filtering 414 is disabled for the first color component 410, the video decoder 122 sets the second-component CCSO flag 428 to indicate that CCSO filtering 414 is not to be applied to the second color component 420 of the current image frame 406C, independent of whether the second syntax element 426 is written to the video bitstream 112. In other words, CCSO filtering 414 is jointly disabled for the first color component 410 and the second color component 420, for example, under the control of the first-component CCSO flag 418. The first-component CCSO flag 418 may or may not be written to the video bitstream 116.

[0074] In some embodiments, the video bitstream 116 includes a first syntax element 416 for the first-component CCSO flag 418. When the first-component CCSO flag 418 indicates that CCSO filtering 414 is to be applied to the first color component 410, a second syntax element 426 is written to the video bitstream 116 to control the CCSO filtering 414 of the second color component 420, independent of the first syntax element 416. When the first-component CCSO flag 418 indicates that CCSO filtering 414 is to be applied to the first color component 410, the video decoder 122 identifies, in the video bitstream 116, a second syntax element 426 for the second-component CCSO flag 428, where the second-component CCSO flag indicates whether CCSO filtering 414 is to be applied to the second color component 420 of the current image frame 406C. The second-component CCSO flag 428 is used to determine whether CCSO filtering 414 is to be applied to the second color samples 420C of the second color component 420 of the current image frame 406C.

[0075] In addition, in some embodiments, the video decoder 122 determines that the first component CCSO flag 418 indicates that CCSO filtering 414 is applied to the first color component 410. When the second component CCSO flag 428 indicates that CCSO filtering 414 is not applied to the second color component 420, the video decoder 122 aborts generating the second sample offset 424 for the second color samples 420C of the second color component 420. Instead, in some embodiments, when the first component CCSO flag 418 indicates that CCSO filtering 414 is applied to the first color component 410 and the second component CCSO flag 428 indicates that CCSO filtering 414 is applied to the second color component 420, the second sample offset 424 for the second color samples 420C of the second color component 420 is generated based on one or more samples of the first color component 410.

[0076] In some embodiments, the first color component 410 includes a luminance component 404, and the second color component 420 includes a chrominance component 402. The two chrominance components 402 include a blue-difference chrominance component 402B and a red-difference chrominance component 402R. In some embodiments, the first color component 410 includes the first chrominance component of the two chrominance components 402R and 402B, and the second color component 420 includes the second chrominance component different from the first chrominance component of the two chrominance components 402R and 402B. In an example, the first color component 410 or the second color component 420 includes the luminance component 404, and according to the CCSO filtering 414, a set of one or more luminance samples 404 including at least the first luminance sample 404C is used to generate the sample offset 422 or 424 of the first luminance sample 404C. In another example, the first color component 410 or the second color component 420 includes the chrominance component 402, and according to the CCSO filtering 414, a set of one or more luminance samples 404 including at least the first luminance sample 404C is used to generate the sample offset of the first chrominance sample 402C.

[0077] In some embodiments, the video bitstream 116 includes a third syntax element 436 for a first CCSO filtering control parameter 438 that indicates whether to apply the CCSO setting 430 to a first color component 410 of the current picture frame 406C. For example, the CCSO setting 430 may include one or more of the following: a maximum number of bands 430A, a filter shape 430B, and a quantizer selection 430C. When the first CCSO filtering control parameter 438 indicates not to apply the CCSO setting 430 to the first color component 410, a second CCSO filtering control parameter 448 is set to indicate not to apply the CCSO setting to a second color component 420 of the current picture frame 406C, independent of whether a fourth syntax element 446 associated with the second CCSO filtering control parameter 448 is written into the video bitstream 116. That is, in some cases, the first CCSO filtering control parameter 438 and the second CCSO filtering control parameter 448 cannot be applied jointly.

[0078] In some embodiments, the video bitstream 116 includes a third syntax element 436 for a first CCSO filtering control parameter 438 that indicates whether to apply the CCSO setting 430, which includes a maximum number of bands 430A, a filter shape 430B, and / or a quantizer selection 430C, to a first color component 410 of the current picture frame 406C. When the first CCSO filtering control parameter 438 indicates to apply the CCSO setting 430 to the first color component 410, a fourth syntax element 446 for a second CCSO filtering control parameter 448 is identified in the video bitstream 116, and the second CCSO filtering control parameter indicates whether to apply the CCSO setting 430 to a second color component 420 of the current picture frame 406C. The video decoder 122 determines whether to apply the CCSO setting 430 to a second color sample 420C of the second color component 420 of the current picture frame 406C based on the fourth syntax element 446.

[0079] In some embodiments, video decoder 122 determines that first CCSO filtering control parameter 438 indicates applying CCSO setting 430 to first color component 410. When second CCSO filtering control parameter 448 indicates not applying CCSO setting 430 to second color component 420, video decoder 122 aborts applying CCSO setting 430 when generating second sample offset 424 for second color sample 420C of second color component 420. In contrast, in some embodiments, video decoder 122 determines that first CCSO filtering control parameter 438 indicates applying CCSO setting 430 to first color component 410. Video decoder 122 also determines that second CCSO filtering control parameter 448 indicates applying CCSO setting 430 to second color component 420. Based on CCSO setting 430, second sample offset 424 is generated for second color sample 420C of second color component 420.

[0080] Note that when applying CCSO filtering control parameter 438 or 448, CCSO setting 430 controlled by parameter 438 or 448 is selected from maximum number of bands 430A, filter shape 430B, and quantizer selection 430C.

[0081] In some embodiments, video bitstream 116 includes joint syntax element 442 for frame-level joint cross-component filtering (CCF) flag 440. Joint CCF flag 440 is configured to indicate whether multiple cross-component filtering operations 444 are jointly applied to current picture frame 406C. Examples of cross-component filtering operations 444 include but are not limited to CCSO filtering 414 and cross-component Wiener filtering. In each cross-component filtering operation 444, the sample value of a certain color sample is adjusted based on one or more sample values of color samples of the same or different color components. Further, in some embodiments, when frame-level joint CCF flag 440 indicates not applying multiple CCF operations 444 to current picture frame 406C including first color sample 410C of first color component 410, video decoder 122 aborts applying each CCF operation of multiple CCF operations 444 to current picture frame 406C.

[0082] Conversely, in some embodiments, when the frame-level combined CCF flag 440 indicates that multiple CCF operations 444 are to be applied to the current image frame 406C, the video decoder 122 applies at least one of the multiple CCF operations to the current image frame 406C that includes the first color samples 410C of the first color component 410. Alternatively, in some embodiments, when the frame-level combined CCF flag 440 indicates that multiple CCF operations 444 are to be applied to the current image frame 406C, the video decoder 122 identifies, in the video bitstream 116, multiple CCF syntax elements (e.g., the first syntax element 408) for multiple CCF flags (e.g., the CCSO flag 412), where the multiple CCF flags indicate whether the multiple CCF operations 444 (e.g., CCSO filtering 414) are to be applied to the current image frame 406C, respectively.

[0083] In some embodiments, the multiple CCF operations 444 include a last CCF operation 444L and one or more remaining CCF operations 444R (e.g., CCSO filtering 414), and the video bitstream 112 includes one or more remaining CCF syntax elements (e.g., the first syntax element 408) for one or more remaining CCF flags (e.g., the CCSO flag 412), where the one or more remaining CCF flags indicate whether the one or more remaining CCF operations 444R are to be applied to the current image frame 406C, respectively. For example, when the frame-level combined CCF flag 440 indicates that the multiple CCF operations 444 are to be applied to the current image frame 406C, the video decoder 122 identifies, in the video bitstream 116, one or more remaining CCF syntax elements for one or more remaining CCF flags. Based on the one or more remaining CCF syntax elements, the video decoder 122 determines to disable the one or more remaining CCF flags (e.g., the CCSO flag 412) and not apply the one or more remaining CCF operations 444R. For the last CCF operation 444L, no last CCF flag is written and identified, and the last CCF operation 444L is applied to the current image frame 406C.

[0084] Figure 5FIG. 0 is a flowchart of an example process of applying CCSO filtering 414 in in-loop filtering according to some embodiments. A decoder 122 receives a video bitstream 116 including a current picture frame 406C. The video bitstream 116 includes a high-level syntax element 408 for a picture-level cross-component sample offset (CCSO) flag 412, and the picture-level CCSO flag indicates whether to apply CCSO filtering 414 to the current picture frame 406C. Based on the high-level syntax element 408, the decoder 122 determines whether to apply CCSO filtering 414 to the current picture frame 406C. When the picture-level CCSO flag 412 indicates to apply CCSO filtering 414 to the current picture frame 406C, the video decoder 122 identifies a first syntax element 416 for a first-component CCSO flag 418 in the video bitstream 116. The first-component CCSO flag indicates whether to apply CCSO filtering 414 to a first color component 410 of the current picture frame 406C, and based on the first syntax element 416, determines whether to apply CCSO filtering 414 to first color samples 410C of the first color component 410. The video decoder 122 reconstructs the current picture frame 406C including the first color samples 410C of the first color component 410.

[0085] In some embodiments, when the picture-level CCSO flag 412 indicates to apply CCSO filtering 414 to the current picture frame 406C, the video decoder 122 further determines that the first-component CCSO flag 418 indicates to apply CCSO filtering 414 to the first color component 410, and generates a first sample offset 422 for the first color samples 410C of the first color component 410 based on one or more luma samples 404. Further, in some embodiments, the video decoder 122 identifies one or more luma samples 404 including a first luma sample 404C and one or more neighboring luma samples 404X. The first luma sample 404X is co-located with the first color samples 410C in the current picture frame 406C. The first sample offset 422 for the first color samples 410C is determined based on the first luma sample 404C and the one or more neighboring luma samples 404X. The first color samples 410C are adjusted based on the first sample offset 422 of the first color samples 410C.

[0086] In some embodiments associated with edge offset, one or more differences 404DX between one or more adjacent luminance samples 404X and a first luminance sample 404C are determined, where the first luminance sample 404C is co-located with a first color sample 410. The one or more differences are quantized to generate one or more quantization values 404Q. The first color sample 410C is classified based on the one or more quantization values 404Q to determine a first sample offset 422 of the first color sample 410. Alternatively, in some embodiments associated with band offset, a set of luminance samples 404 is quantized to generate one or more quantization values 404Q. The luminance values (not gradient values or differences) of the set of luminance samples 404 are quantized. The first color sample 410C is classified based on the one or more quantization values 404Q to determine a first sample offset 422 of the first color sample 410. The decoder 122 reconstructs the current image frame 406C at least by adjusting the first color sample 410 based on the first sample offset 422. In some embodiments, the first color sample 410C is one of the following: the first luminance sample 404C, the first chrominance red (Cr) sample, and the first chrominance blue (Cb) sample. The first luminance sample 404C, the first Cb sample, and the first Cr sample are co-located with each other.

[0087] For example, the classifier 512 classifies the first color sample 410C based on the quantization values 404Q to determine the first sample offset 422 of the first color sample 410C. In the example, the quantization values 404Q include quantization values 404QC, 404QN, 404QS, 404QW, and 404QE. The look-up table 514 maps multiple combinations of the quantization values 404QC, 404QN, 404QS, 404QW, and 404QE to different sample offset options SO (e.g., SO1 - SQ16). Based on the look-up table 514, the quantization values 404Q correspond to one of the combinations in the look-up table 514; and the corresponding sample offset option SO associated with the combination of the quantization values 404Q is identified, and thus the sample offset option SO is selected for the first sample offset 422. In other words, in some embodiments, the decoder 122 classifies the first color sample 410C by: identifying a combination of one or more quantization values 404Q in the look-up table 514 that associates multiple combinations of quantization values with multiple offset value options SO (e.g., SO1 - SO16); and determining the first sample offset 422 corresponding to the combination of one or more quantization values 404Q in the look-up table 514.

[0088] In some embodiments, a scalar quantizer 530 including a plurality of quantization intervals (QI) 518 and a plurality of quantization levels (QL) 520 quantizes the value (not the associated difference or gradient) of the luminance sample 404A into a plurality of integer values within the quantization range 516, and each quantization value in one or more quantization values 404Q includes a corresponding integer within the quantization range 516. For each integer value within the quantization range 516, the quantization interval 518 is defined as the range of values assigned to that corresponding integer value. The quantization level 520 corresponds to that corresponding integer value to which a difference range associated with the quantization interval 518 is assigned.

[0089] In some embodiments, in the band offset only mode, only the first luminance sample 404C is quantized, e.g., by the quantizer 530, and classified to determine a first sample offset 422 of the first color sample 410C. The first luminance sample 404C is determined to be associated with one of a plurality of bands. Each of the plurality of bands corresponds to a corresponding sample offset value. The first sample offset 422 is determined to be equal to the corresponding sample offset value corresponding to one of the plurality of bands.

[0090] The first color sample 410C is adjusted based on the first sample offset 422 of the first color sample 410C such that the current image frame 406C can be reconstructed. In some embodiments, the first color sample 410C includes a first chrominance sample 402C co-located with the first luminance sample 404C in the current image frame, and the first chrominance sample 402C is adjusted based on the first sample offset 422. Alternatively, in some embodiments, the first color sample 410C is the first luminance sample 404C, and the first luminance sample 404C is adjusted based on the first sample offset 422.

[0091] In some embodiments not shown, when both the frame-level CCSO flag 412 and the second component CCSO flag 428 are enabled, in CCSO filtering 414, a second sample offset 424 of the second color sample 420C of the second color component is determined, e.g., based on a set of luminance samples 404. For the sake of brevity, the details of the CCSO filtering 414 are not described further.

[0092] Figure 6 is a flowchart showing an example method 600 for decoding video according to some embodiments. The method 600 may be performed by a computing system including a control circuit and a memory storing instructions executed by the control circuit (e.g., Figure 1executed by the server system 112, source device 102, or electronic device 120) therein. In some embodiments, method 600 is applied in conjunction with one or more video codecs, including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In some embodiments, the CCSO filtering method 414( Figure 4 ) is implemented based on an edge-preserving loop filter that uses reconstructed samples to calculate the sample offsets for the luminance samples 404 or chrominance components 402. The control flags 418 and 428 for the luminance or chrominance sample offsets can be written individually for each component at the frame level. Additional advanced flags (e.g., CCSO filtering control parameters 438 and 448) can be further introduced to control the CCSO filtering parameters (e.g., CCSO settings 430) for different color components (e.g., luminance samples 404, chrominance samples 402).

[0093] In some embodiments, a frame-level CCSO filtering enable flag (e.g., frame-level CCSO flag 412) is written. If this flag (e.g., frame-level CCSO flag 412) is true, then the cross-component enable flags for each component (e.g., component CCSO flags 418 and 428) are further written. Otherwise, if this flag (e.g., frame-level CCSO flag 412) is false, then the cross-component enable flags for each component (e.g., component CCSO flags 418 and 428) are inferred to be false.

[0094] In some embodiments, the chrominance component CCSO filtering enable flag (e.g., component CCSO flag 428) depends on the luminance component flag (e.g., component CCSO flag 418). That is, if the luminance CCSO filtering enable flag (e.g., component CCSO flag 418) is true, then the chrominance component CCSO filtering enable flag (e.g., component CCSO flag 428) is further written. Otherwise, if the luminance CCSO filtering enable flag (e.g., component CCSO flag 418) is false, then the chrominance component cross-component sample offset enable flag (e.g., component CCSO flag 428) is inferred to be false.

[0095] In some embodiments, the chrominance component CCSO filtering enable flag (e.g., one of the two component CCSO flags 428) depends on another chrominance component flag (e.g., the other of the two component CCSO flags 428). If the Cb CCSO filtering enable flag is true, then the Cr CCSO filtering enable flag is further written. Otherwise, if the Cb CCSO filtering flag is false, then the Cr CCSO filtering flag is inferred to be false.

[0096] In some embodiments, dual control (e.g., frame-level control and component-level control, two component-level controls) is jointly applied to one of the CCSO filtering control flags or indicators different from CCSO flags 412, 418, and 428. Additionally, in some embodiments, the indicator for the maximum number of bands 430A can be shared among all different color components, or the indicator for the maximum number of bands 430A can be written independently for different color components. Alternatively, in some embodiments, the filter shape selection 430B can be shared among different color components 410 and 420. Alternatively, in some embodiments, the quantizer selection 430C can be shared among different color components 410 and 420. In other words, a frame-level indicator is applied to control CCSO settings 430A, 430B, or 430C, and a component-level indicator can be applied based on the frame-level indicator. Alternatively, two component-level indicators associated with CCSO settings 430A, 430B, or 430C can be jointly written and applied based on each other.

[0097] In some embodiments, the controls of CCSO flags 412, 418, and 428 can be combined. A frame-level cross-component enable flag (e.g., frame-level CCSO flag 412) is written. If the frame-level cross-component enable flag (e.g., frame-level CCSO flag 412) is true, dual control of two of CCSO flags 418 and 428 is applied. Otherwise, if the frame-level cross-component enable flag (e.g., frame-level CCSO flag 412) is false, all CCSO flags 418 and 428 are set to false. For example, if the frame-level CCSO flag 412 is true, CCSO flag 428 depends on CCSO flag 418.

[0098] In some embodiments, control flags (e.g., combined CCF flag 440) can be jointly written for multiple loop filtering methods 444. For example, the combined CCF flag 440 is applied to indicate whether any or all cross-component filtering methods 444 are applied; if the combined CCF flag 440 is false, no cross-component filtering method 444 (including CCSO filtering 414) is written or applied individually.

[0099] In some embodiments, a first flag (e.g., combined CCF flag 440) indicates whether any or all of the cross-component filtering methods are applied. If the combined CCF flag 440 is true, at least one of the cross-component filtering methods 444 (including CCSO filtering 414) is applied. Further, in some embodiments, if the combined CCF flag 440 is true, an enable flag for each cross-component filtering method 444 is further written. Alternatively, in some embodiments, the combined CCF flag 440 is true. The enable flags for all cross-component filtering methods 444 except the last cross-component filtering method 444L in the ordered sequence of written enable flags are written as false. The last cross-component filtering method 444L is enabled, however, the remaining CCF operations 444R are disabled. The enable flag for the last cross-component filtering method 444L is not written, but is determined to be true.

[0100] Although Figure 6 Although multiple logical stages are shown in a particular order, stages that are not order-dependent can be reordered and other stages can be combined or split. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, and thus the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.

[0101] Some example embodiments are described below.

[0102] (A1) In some implementations, method 600 is implemented for decoding video data. Method 600 includes receiving (operation 602) a video bitstream including a current picture frame, where the video bitstream includes (operation 604) a high-level syntax element indicating whether cross-component sample offset (CCSO) filtering is applied to the current picture frame; based on the high-level syntax element, determining (operation 606) whether to apply CCSO filtering to the current picture frame; when CCSO filtering is applied to the current picture frame: identifying (operation 610) in the video bitstream a first syntax element for a first component CCSO flag, the first component CCSO flag indicating whether CCSO filtering is applied to a first color component of the current picture frame; and determining (operation 612) whether to apply CCSO filtering to a first color sample of the first color component based on the first syntax element; and reconstructing (operation 614) the current picture frame including the first color sample of the first color component.

[0103] (A2) In some embodiments of A1, method 600 further includes: when the frame-level CCSO flag indicates that CCSO filtering is disabled for the first image frame, setting the independent-component CCSO flag to indicate that CCSO filtering is not applied to the first color component of the first image frame, regardless of whether any syntax elements of the independent-component CCSO flag of the first image frame are written to the video bitstream.

[0104] (A3) In some embodiments of any one of A1 or A2, method 600 further includes: when the frame-level CCSO flag indicates that CCSO filtering is applied to the current image frame and when the first-component CCSO flag indicates that CCSO filtering is not applied to the first color component, aborting the generation of the first sample offset for the first color sample of the first color component.

[0105] (A4) In some embodiments of A1 or A2, method 600 further includes: when the frame-level CCSO flag indicates that CCSO filtering is applied to the current image frame and when the first-component CCSO flag indicates that CCSO filtering is applied to the first color component, generating a first sample offset for the first color sample of the first color component based on one or more luma samples.

[0106] (A5) In some embodiments of A4, generating the first sample offset for the first color sample of the first color component further includes: identifying one or more luma samples including the first luma sample and one or more adjacent luma samples, wherein the first luma sample is co-located with the first color sample in the current image frame; and determining the first sample offset of the first color sample based on the first luma sample and the one or more adjacent luma samples, wherein the first color sample is adjusted based on the first sample offset of the first color sample.

[0107] (A6) In some embodiments of A5, generating the first sample offset for the first color sample of the first color component further includes: quantizing one or more luma samples to generate one or more quantization values; and classifying the first color sample based on the one or more quantization values to determine the first sample offset of the first color sample.

[0108] (A7) In some embodiments of A5, generating the first sample offset for the first color sample of the first color component further includes: determining one or more differences between the one or more adjacent luma samples and the first luma sample, wherein the first luma sample is co-located with the first color sample; quantizing the one or more differences to generate one or more quantization values; and classifying the first color sample based on the one or more quantization values to determine the first sample offset of the first color sample.

[0109] (A8)In some embodiments of any one of A1 to A7, the first color component is one of a luminance component, a blue-difference chrominance component, and a red-difference chrominance component.

[0110] (A9)In some embodiments of any one of A1 to A8, the video bitstream includes a first syntax element for a first component CCSO flag. Method 600 further includes: when the first component CCSO flag indicates that CCSO filtering is disabled for the first color component, setting a second component CCSO flag to indicate that CCSO filtering is not applied to the second color component of the current picture frame, independent of whether the second syntax element is written into the video bitstream.

[0111] (A10)In some embodiments of any one of A1 to A9, the video bitstream includes a first syntax element for a first component CCSO flag. Method 600 further includes: when the first component CCSO flag indicates that CCSO filtering is applied to the first color component: identifying, in the video bitstream, a second syntax element for a second component CCSO flag, the second component CCSO flag indicating whether CCSO filtering is applied to the second color component of the current picture frame; and determining, based on the second syntax element, whether to apply CCSO filtering to a second color sample of the second color component of the current picture frame.

[0112] (A11)In some embodiments of A10, Method 600 further includes: when the first component CCSO flag indicates that CCSO filtering is applied to the first color component and when the second component CCSO flag indicates that CCSO filtering is not applied to the second color component, aborting the generation of a second sample offset for a second color sample of the second color component.

[0113] (A12)In some embodiments of A10, Method 600 further includes: when the first component CCSO flag indicates that CCSO filtering is applied to the first color component and when the second component CCSO flag indicates that CCSO filtering is applied to the second color component, generating, based on one or more samples of the first color component, a second sample offset for a second color sample of the second color component.

[0114] (A13)In some embodiments of any one of A9 to A12, the first color component includes a luminance component, and the second color component includes a chrominance component.

[0115] (A13)In some embodiments of any one of A9 to A12, the two chrominance components include a blue-difference chrominance component and a red-difference chrominance component. The first color component includes the first chrominance component of the two chrominance components, and the second color component includes the second chrominance component different from the first chrominance component of the two chrominance components.

[0116] (A15)In some embodiments of any one of A1 to A14, the video bitstream includes a third syntax element for a first CCSO filtering control parameter, and the first CCSO filtering control parameter indicates whether to apply the CCSO setting to the first color component of the current picture frame. The method 600 further includes: when the first CCSO filtering control parameter indicates not to apply the CCSO setting to the first color component, setting a second CCSO filtering control parameter to indicate not to apply the CCSO setting to the second color component of the current picture frame, which is independent of whether the fourth syntax element is written into the video bitstream.

[0117] (A16)In some embodiments of any one of A1 to A15, the video bitstream includes a third syntax element for a first CCSO filtering control parameter, and the first CCSO filtering control parameter indicates whether to apply the CCSO setting including the maximum number of bands, filter shape, and quantizer selection to the first color component of the current picture frame. The method 600 further includes: when the first CCSO filtering control parameter indicates that the CCSO setting is applied to the first color component: identifying a fourth syntax element for a second CCSO filtering control parameter in the video bitstream, where the second CCSO filtering control parameter indicates whether to apply the CCSO setting to the second color component of the current picture frame; and determining whether to apply the CCSO setting to the second color samples of the second color component of the current picture frame based on the fourth syntax element.

[0118] (A17)In some embodiments of A16, the method 600 further includes: when the first CCSO filtering control parameter indicates applying the CCSO setting to the first color component and when the second CCSO filtering control parameter indicates not to apply the CCSO setting to the second color component, aborting the application of the CCSO setting when generating the second sample offset of the second color samples for the second color component.

[0119] (A18)In some embodiments of A16, the method 600 further includes: when the first CCSO filtering control parameter indicates applying the CCSO setting to the first color component and when the second CCSO filtering control parameter indicates applying the CCSO setting to the second color component, generating a second sample offset of the second color samples for the second color component based on the CCSO setting.

[0120] (A19)In some embodiments of any one of A15 to A18, the CCSO setting is selected from the maximum number of bands, filter shape, and quantizer selection.

[0121] (A20)In some embodiments of any one of A1 to A19, the video bitstream includes a joint syntax element for a frame-level joint CCF flag, and the joint CCF flag is configured to indicate whether to jointly apply multiple CCF operations to the current picture frame.

[0122] (A21) In some embodiments of A20, method 600 further includes: when the frame-level combined CCF flag indicates that multiple CCF operations are not applied to the current image frame, aborting the application of each of the multiple CCF operations to the current image frame.

[0123] (A22) In some embodiments of A20, method 600 further includes: when the frame-level combined CCF flag indicates that multiple CCF operations are applied to the current image frame, applying at least one of the multiple CCF operations to the current image frame including a first color sample of a first color component.

[0124] (A23) In some embodiments of A20, method 600 further includes: when the frame-level combined CCF flag indicates that multiple CCF operations are applied to the current image frame, identifying, in the video bitstream, multiple CCF syntax elements for multiple CCF flags, the multiple CCF flags indicating whether the multiple CCF operations are respectively applied to the current image frame.

[0125] (A24) In some embodiments of A20, the multiple CCF operations include a last CCF operation and one or more remaining CCF operations, and the video bitstream includes one or more remaining CCF syntax elements for one or more remaining CCF flags, the one or more remaining CCF flags indicating whether the one or more remaining CCF operations are respectively applied to the current image frame. Method 600 further includes: when the frame-level combined CCF flag indicates that multiple CCF operations are applied to the current image frame: identifying, in the video bitstream, one or more remaining CCF syntax elements for one or more remaining CCF flags; determining, based on the one or more remaining CCF syntax elements, to disable the one or more remaining CCF flags; and determining to apply the last CCF operation to the current image frame, wherein the last CCF flag is written and identified for the last CCF operation.

[0126] (A25) In some embodiments, a computing system includes a control circuit and a memory storing one or more programs configured to be executed by the control circuit. The one or more programs further include instructions for: receiving video data including a current image frame; encoding the current image frame; transmitting the encoded current image frame via a video bitstream; and writing, via the video bitstream, a high-level syntax element for a frame-level CCSO flag, the frame-level CCSO flag indicating whether CCSO filtering is applied to the current image frame. When the frame-level CCSO flag indicates that CCSO filtering is applied to the current image frame, writing a first syntax element for a first-component CCSO flag in the video bitstream, the first-component CCSO flag indicating whether CCSO filtering is applied to a first color component of the current image frame.

[0127] (A26) In some embodiments, a non-transitory computer-readable storage medium stores one or more programs executed by a control circuit of a computing system. The one or more programs include instructions for: obtaining a source video sequence including a current image frame and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes the current image frame and high-level syntax elements for a frame-level CCSO flag that indicates whether CCSO filtering is applied to the current image frame. When the frame-level CCSO flag indicates that CCSO filtering is applied to the current image frame, a first syntax element for a first-component CCSO flag is written in the video bitstream, and the first-component CCSO flag indicates whether CCSO filtering is applied to a first color component of the current image frame.

[0128] In another aspect, some embodiments include a computing system (e.g., server system 112). The computing system includes a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, and the memory stores one or more sets of instructions configured to be executed by the control circuit, and the one or more sets of instructions include instructions for performing any of the methods described herein (e.g., A1 - A26 above).

[0129] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium that stores one or more sets of instructions for being executed by a control circuit of a computing system, and the one or more sets of instructions include instructions for performing any of the methods described herein (e.g., A1 - A26 above).

[0130] The proposed methods can be used alone or in any order combination. Additionally, each of the methods (or embodiments), encoders, and decoders can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). For example, one or more processors execute programs stored in a non-transitory computer-readable medium. Hereinafter, the term block can be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU.

[0131] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0132] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. Further, it can be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0133] As used herein, depending on the context, the term "if" can be interpreted to mean "when", or "after", or "in response to determining", or "in accordance with determining", or "in response to detecting" that the prerequisite is true. Similarly, depending on the context, the phrases "if it is determined (that the prerequisite is true)", or "if (the prerequisite is true)", or "when (the conditional prerequisite is true)" can be interpreted to mean "after determining that the prerequisite is true", or "in response to determining that the prerequisite is true", or "in accordance with determining that the prerequisite is true", or "after detecting that the prerequisite is true", or "in response to detecting that the prerequisite is true".

[0134] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Given the above teachings, many modifications and variations are possible. The embodiments were chosen and described in order to best explain the operating principles and the practical application, so as to enable others skilled in the art to implement.

Claims

1. A method for decoding video data, comprising: Receiving a video code stream including a current image frame, wherein the video code stream includes a high-level syntax element indicating whether to apply a cross-component sample offset (CCSO) filter to the current image frame; Based on the high-level syntax element, determining whether to apply CCSO offset filtering to the current image frame; When CCSO filtering is applied to the current image frame: Identifying a first syntax element for a first component CCSO flag in the video bitstream, the first component CCSO flag indicating whether CCSO filtering is applied to a first color component of the current image frame; and Determining whether to apply CCSO filtering to a first color sample of the first color component based on the first syntax element; and reconstructing the current image frame including the first color sample of the first color component.

2. The method according to claim 1, further comprising: When the frame-level CCSO flag indicates that CCSO filtering is disabled for a first image frame, independent of whether any syntax element of the independent component CCSO flag for the first image frame is written into the video bitstream, the independent component CCSO flag is set to indicate that CCSO filtering is not applied to the first color component of the first image frame.

3. The method according to claim 1, further comprising: When the frame-level CCSO flag indicates that CCSO filtering is applied to the current image frame: When the first component CCSO flag indicates that CCSO filtering is not applied to the first color component, generating a first sample offset for a first color sample of the first color component is discontinued.

4. The method according to claim 1, further comprising: When the frame-level CCSO flag indicates that CCSO filtering is applied to the current image frame: When the first component CCSO flag indicates that CCSO filtering is applied to the first color component, a first sample offset for a first color sample of the first color component is generated based on one or more luma samples.

5. The method according to claim 4, wherein: Generating the first sample offset for the first color sample for the first color component further comprises: identifying the one or more luma samples comprising a first luma sample and one or more adjacent luma samples, wherein the first luma sample is co-located with the first color sample in the current image frame; and The first sample offset of the first color sample is determined based on the first luma sample and the one or more neighboring luma samples, wherein the first color sample is adjusted based on the first sample offset of the first color sample.

6. The method of claim 5, generating the first sample offset for the first color sample of the first color component, further comprising: quantizing the one or more luma samples to generate one or more quantized values; as well as The first color sample is classified based on the one or more quantized values ​​to determine the first sample offset of the first color sample.

7. The method of claim 5, generating the first sample offset for the first color sample of the first color component, further comprising: determining one or more differences between the one or more adjacent luma samples and the first luma sample, wherein the first luma sample is co-located with the first color sample; quantizing the one or more difference values ​​to generate one or more quantized values; and The first color sample is classified based on the one or more quantized values ​​to determine the first sample offset of the first color sample.

8. The method according to any one of claims 1, wherein: The first color component is one of a luminance component, a blue difference chrominance component, and a red difference chrominance component.

9. The method according to claim 1, wherein: The video code stream includes the first syntax element for the first component CCSO flag, and the method further includes: When the first component CCSO flag indicates that CCSO filtering is disabled for the first color component, independent of whether a second syntax element is written to the video bitstream, a second component CCSO flag is set to indicate that CCSO filtering is not applied to the second color component of the current image frame.

10. The method according to claim 1, wherein: The video code stream includes the first syntax element for the first component CCSO flag, and the method further includes: When the first component CCSO flag indicates that CCSO filtering is applied to the first color component: Identifying a second syntax element for a second component CCSO flag in the video bitstream, the second component CCSO flag indicating whether CCSO filtering is applied to a second color component of the current image frame; and Determine whether to apply CCSO filtering to a second color sample of a second color component of the current image frame based on the second syntax element.

11. The method according to claim 10, further comprising: When the first component CCSO flag indicates that CCSO filtering is applied to the first color component: When the second component CCSO flag indicates that CCSO filtering is not applied to the second color component, generating a second sample offset for a second color sample of the second color component is discontinued.

12. The method according to claim 10, further comprising: When the first component CCSO flag indicates that CCSO filtering is applied to the first color component: When the second component CCSO flag indicates that CCSO filtering is applied to the second color component, a second sample offset for second color samples of the second color component is generated based on one or more samples of the first color component.

13. The method according to any one of claims 9, wherein: The first color component includes a luminance component, and the second color component includes a chrominance component.

14. The method according to any one of claims 9, wherein: The two chrominance components include a blue difference chrominance component and a red difference chrominance component, and wherein the first color component includes a first chrominance component of the two chrominance components, and the second color component includes a second chrominance component of the two chrominance components that is different from the first chrominance component.

15. The method according to claim 1, wherein: The video bitstream includes a third syntax element for a first CCSO filter control parameter, the first CCSO filter control parameter indicating whether to apply a CCSO setting for the first color component of the current image frame, the method further comprising: When the first CCSO filtering control parameter indicates that the CCSO setting is not applied to the first color component, independent of whether the fourth syntax element is written to the video bitstream, a second CCSO filtering control parameter is set to indicate that the CCSO setting is not applied to the second color component of the current image frame.

16. The method according to claim 1, wherein: The video bitstream includes a third syntax element for a first CCSO filter control parameter, the first CCSO filter control parameter indicating whether to apply CCSO settings including a maximum number of frequency bands, a filter shape, and a quantizer selection to the first color component of the current image frame, the method further comprising: When the first CCSO filter control parameter indicates that the CCSO setting is to be applied to the first color component: identifying, in the video bitstream, a fourth syntax element for a second CCSO filter control parameter, the second CCSO filter control parameter indicating whether to apply the CCSO setting to a second color component of the current image frame; and Determining whether to apply the CCSO setting to a second color sample of the second color component of the current image frame based on the fourth syntax element.

17. The method according to claim 16, further comprising: When the first CCSO filter control parameter indicates that the CCSO setting is applied to the first color component: When the second CCSO filter control parameter indicates that the CCSO setting is not to be applied to the second color component, applying the CCSO setting is discontinued in generating a second sample offset for a second color sample of the second color component.

18. The method according to claim 16, further comprising: When the first CCSO filter control parameter indicates that the CCSO setting is applied to the first color component: When the second CCSO filter control parameter indicates that the CCSO setting is applied to the second color component, a second sample offset for a second color sample of the second color component is generated based on the CCSO setting.

19. The method according to any one of claims 15, wherein: The CCSO settings are selected from the maximum number of frequency bands, filter shape and quantizer selection.

20. The method according to claim 1, wherein: The video code stream includes a joint syntax element for a frame-level joint cross-component filtering (CCF) flag, wherein the frame-level joint CCF flag is configured to indicate whether multiple CCF operations are jointly applied to the current image frame, and the high-level syntax element includes a frame-level syntax element.

21. The method according to claim 20, further comprising: When the frame-level joint CCF flag indicates that the multiple CCF operations are not to be applied to the current image frame, applying each of the multiple CCF operations to the current image frame is suspended.

22. The method according to claim 20, further comprising: When the frame-level joint CCF flag indicates that the multiple CCF operations are applied to the current image frame, at least one CCF operation of the multiple CCF operations is applied to the current image frame including the first color sample of the first color component.

23. The method of claim 20, further comprising: When the frame-level joint CCF flag indicates that the multiple CCF operations are applied to the current image frame, multiple CCF syntax elements for multiple CCF flags are identified in the video code stream, and the multiple CCF flags indicate whether the multiple CCF operations are applied to the current image frame respectively.

24. The method according to claim 20, wherein: The multiple CCF operations include a last CCF operation and one or more remaining CCF operations, and the video code stream includes one or more remaining CCF syntax elements for one or more remaining CCF flags, and the one or more remaining CCF flags indicate whether the one or more remaining CCF operations are applied to the current image frame respectively, and the method also includes: When the frame-level joint CCF flag indicates that the multiple CCF operations are applied to the current image frame: Identifying the one or more remaining CCF syntax elements for the one or more remaining CCF flags in the video bitstream; Based on the one or more remaining CCF syntax elements, determining to disable the one or more remaining CCF flags; and Determine to apply the last CCF operation to the current image frame, wherein a last CCF flag is written and identified for the last CCF operation.

25. A computing system comprising: Control circuit; as well as A memory storing one or more programs configured to be executed by the control circuit, the one or more programs further comprising instructions for: Receiving video data including a current image frame; Encoding the current image frame; Sending the encoded current image frame via the video code stream; as well as Writing, via the video bitstream, a high-level syntax element for a frame-level cross-component sample offset (CCSO) flag, wherein the frame-level CCSO flag indicates whether CCSO filtering is applied to the current image frame; When the frame-level CCSO flag indicates that CCSO filtering is applied to the current image frame, a first syntax element for the first component CCSO flag is written in the video bitstream, and the first component CCSO flag indicates whether CCSO filtering is applied to the first color component of the current image frame.

26. A non-transitory computer-readable storage medium storing one or more programs executed by control circuitry of a computing system, the one or more programs comprising instructions for: Obtaining a source video sequence including a current image frame; as well as Performing conversion between the source video sequence and the video code stream, wherein: The video code stream includes: the current image frame; and a high-level syntax element for a frame-level cross-component sample offset (CCSO) flag, the frame-level CCSO flag indicating whether CCSO filtering is applied to the current image frame; When the frame-level CCSO flag indicates that CCSO filtering is applied to the current image frame, a first syntax element for the first component CCSO flag is written in the video bitstream, and the first component CCSO flag indicates whether CCSO filtering is applied to the first color component of the current image frame.