CCSO using downsampling filter and sampling position selection

By adopting a cross-component offset filtering method in video encoding and decoding technology, the offset value is calculated using the co-bit reconstruction sample of the first color component and the adjacent reconstruction sample, and adjusting the samples of the second color component, the problem of insufficient compression efficiency in the prior art is solved, and more efficient video compression and transmission is achieved.

CN120202668APending Publication Date: 2025-06-24TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004776.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-09
Filing Date
2024-05-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When existing video encoding and codec technology compresses and transmits video data, it is difficult to effectively utilize the redundancy in the video data, resulting in high bandwidth and storage requirements.

Method used

The cross-component offset filtering method is used to calculate the offset value through the co-bit reconstruction sample of the first color component and the adjacent reconstruction sample, and adjust the samples of the second color component to achieve more efficient video compression.

Benefits of technology

Through cross-component offset filtering technology, reconstruction errors can be further reduced, video compression rate can be improved, bandwidth and storage requirements can be reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120202668A_ABST
    Figure CN120202668A_ABST
Patent Text Reader

Abstract

Various implementations described herein include methods and systems for encoding and decoding video. In one aspect, a video bitstream includes a current image frame and a first syntax element. The electronic device determines that the first syntax element has a first predetermined value indicating that a cross-component sample offset (CCSO) mode is enabled, generates a set of adapted luminance samples based on a set of reconstructed luminance samples, the set of adapted luminance samples including an adapted first luminance sample and its adapted adjacent luminance samples. The reconstructed luminance sample includes a first luminance sample co-located with the first color sample. The electronic device determines a first sample offset for the first color sample based on the adapted first luminance sample and the one or more adapted adjacent luminance samples. The current image frame is reconstructed at least by adjusting the first color samples based on the first sample offset.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by reference

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 543,273, filed Oct. 9, 2023, entitled "CCSO with Downsampling Filters and Sample Position Selection", and this application is a continuation of, and claims priority to, U.S. Patent Application No. 18 / 660,067, filed May 9, 2024, entitled "CCSO with Downsampling Filters and Sample Position Selection", both of which are hereby incorporated by reference in their entirety. Technical field

[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for loop filtering (e.g., cross-component offset filtering) of video data. Background art

[0003] A variety of electronic devices support digital video, such as digital televisions, laptop or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise convey digital video data over a communication network and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited storage resources of the storage device, video coding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data. Video coding can be performed by software and / or hardware on an electronic device / client device or a server providing cloud services.

[0004] Video encoding typically uses prediction methods that exploit the redundancy inherent in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video encoding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing a degradation in video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4), respectively. Versatile Video Coding (VVC / H.266) is a video compression standard designed to succeed HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2), respectively. AOMedia Video 1 (AV1) is an open video coding format designed to replace HEVC. On January 8, 2019, the verified version 1.0.0 of the specification with Errata 1 was released. Summary of the Invention

[0005] As described above, encoding (compression) reduces the bandwidth and / or storage space requirements. As will be described in detail later, both lossless compression and lossy compression can be employed. Lossless compression refers to a technique by which an exact copy of the original signal can be reconstructed from the compressed signal through a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully retained during the encoding process and is not fully restored during the decoding process. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be useful for the intended application. The degree of tolerable distortion depends on the application. For example, users of some consumer video streaming applications may be more tolerant of higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: generally, a higher tolerable distortion allows the use of an encoding algorithm that incurs higher losses but has a higher compression ratio.

[0006] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying in-loop filters to video (image) compression. A video codec includes multiple functional modules for performing one or more of the following operations: intra / inter prediction, transform coding, quantization, entropy coding, and in-loop filtering. In-loop filtering techniques are used to adjust the reconstructed picture samples to further reduce the reconstruction error. A cross-component offset filtering method is implemented, which applies the co-located reconstructed samples of the first color component and the associated adjacent reconstructed samples to derive an offset value, which is added to the current sample of the second color component to adjust the reconstructed value of the current sample. An example of the first color component is the luminance color component, and an example of the second color component is the chrominance color component. In some implementations, the first color component and the second color component correspond to the same color component, such as luminance samples.

[0007] In various embodiments of the present application, the samples of the first color component are processed to generate adapted samples of the first color component, and then these adapted samples are processed by a cross-component offset filter to determine an offset value, which is added to the samples of the second color component. For example, the reconstructed luminance samples are downsampled, and the downsampled luminance samples are applied to generate an offset value for the first luminance sample or an offset value for the first chrominance sample co-located with the first luminance sample. In another example, a luminance filter is applied to the reconstructed luminance samples to generate adapted luminance samples, and these adapted luminance samples are further processed to generate an offset value.

[0008] According to some embodiments, a video decoding method is provided. The method includes receiving a video bitstream including a current image frame. The video bitstream includes a first syntax element that indicates whether a first sample offset of a first color sample of the current image frame is determined based on one or more luminance samples for a cross-component sample offset (CCSO) mode. The method further includes, when the CCSO mode is enabled, generating a set of adapted luminance samples based on a set of reconstructed luminance samples, the set of adapted luminance samples including an adapted first luminance sample and its adapted adjacent luminance samples. The set of reconstructed luminance samples includes a first luminance sample co-located with the first color sample of the current image frame. The method further includes determining a first sample offset of the first color sample based on the adapted first luminance sample and one or more adapted adjacent luminance samples, and reconstructing the current image frame at least by adjusting the first color sample based on the first sample offset.

[0009] According to some embodiments, a video encoding method is provided. The method includes receiving video data including a current picture frame, encoding the current picture frame, enabling a Cross-Component Sample Offset (CCSO) mode to generate a first sample offset of first color samples of the current picture frame based on one or more luma samples. In the CCSO mode, the first sample offset of the first color samples is determined based on a set of adapted luma samples, and the set of adapted luma samples is further generated based on a set of reconstructed luma samples including first luma samples co-located with the first color samples of the current picture frame. The method further includes transmitting the encoded current picture frame through a video bitstream and writing a first syntax element through the video bitstream to indicate the application of the CCSO mode to reconstruct the first color samples co-located with the first luma samples based on the first sample offset.

[0010] According to some embodiments, a bitstream conversion method is provided. The method includes obtaining a source video sequence including a current picture frame and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes the current picture frame and a first syntax element. The first syntax element indicates whether a first sample offset of first color samples of the current picture frame is generated based on one or more luma samples for a Cross-Component Sample Offset (CCSO) mode. The first sample offset of the first color samples is determined based on a set of adapted luma samples generated based on a first luma sample and one or more neighboring luma samples, and the first luma sample is co-located with the first color samples.

[0011] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic devices, is provided. The computing system includes a control circuit and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.

[0012] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions executable by a computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.

[0013] Therefore, apparatuses and systems having methods for encoding and decoding video are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding and decoding.

[0014] The features and advantages described in the specification are not necessarily all included. In particular, in view of the accompanying drawings, the specification, and the claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. In addition, it should be noted that the language used in the specification is mainly selected for readability and guidance purposes, and is not selected to describe or limit the subject matter described herein. Description of the Drawings

[0015] To understand the present disclosure in more detail, a more specific description can be made by referring to the features of various embodiments, some of which are shown in the accompanying drawings. However, the attached drawings only show the relevant features of the present disclosure and are not necessarily considered restrictive. This specification may recognize other effective features that those skilled in the art will understand when reading this disclosure.

[0016] Figure 1 is a block diagram showing an example communication system according to some embodiments.

[0017] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.

[0018] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.

[0019] Figure 3 is a block diagram showing an example server system according to some embodiments.

[0020] Figure 4 is a flowchart showing an example process of applying in-loop filtering in video decoding according to some embodiments.

[0021] Figure 5 is a schematic diagram showing information embedded in a video bitstream according to some embodiments.

[0022] Figure 6 is a flowchart showing a method of decoding video according to some embodiments.

[0023] As a matter of convention, the various features shown in the drawings are not necessarily drawn to scale, and the same reference numerals may be used throughout the specification and the drawings to denote the same features. Detailed Description of the Embodiments

[0024] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying in-loop filters for video (image) compression. In-loop filtering techniques are used to adjust reconstructed picture samples to further reduce reconstruction errors. A cross-component offset filtering method is implemented that applies co-located reconstructed samples of a first color component and associated neighboring reconstructed samples to derive an offset value that is added to a current sample of a second color component to adjust the reconstructed value of the current sample. In various embodiments of the present application, a decoder receives a video bitstream including a current image frame and a first syntax element from an encoder. Reconstructed samples of the first color component are used to generate adapted samples of the first color component, and these adapted samples are processed by a cross-component offset filter to determine an offset value that is added to a sample of the second color component. For example, adapted luma samples are used to generate an offset value for a first luma sample or a first chroma sample co-located with the first luma sample.

[0025] More specifically, in some embodiments, a video decoder identifies a first luma sample and one or more neighboring luma samples of the first luma sample based on a filter shape. The decoder processes these luma samples to generate adapted luma samples. The decoder may determine one or more adapted differences between one or more adapted neighboring luma samples and the adapted first luma sample. The adapted luma samples or one or more adapted differences are quantized, for example using a scalar quantizer, to generate one or more quantization values. The scalar quantizer may be specified by a quantization interval (e.g., a range of values assigned to the same integer) and a quantization level (e.g., an integer value assigned to the quantization interval). A first color sample is classified (e.g., by a classifier) based on one or more quantization values to determine a first sample offset of the first color sample. The first color sample is adjusted based on the first sample offset of the first color sample to effect reconstruction of the current image frame. Additionally, in some embodiments, the adapted luma samples include downsampled luma samples generated by a downsampling filter.

[0026] Figure 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) that are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, such as for use with video-enabled applications (e.g., video conferencing applications, digital television applications, and media storage and / or distribution applications).

[0027] The source device 102 includes a video source 104 (e.g., a camera assembly or a media storage) and an encoder assembly 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder assembly 106 generates one or more encoded video bitstreams based on the video stream. Compared with the video stream from the video source 104, the video stream generated by the encoder assembly 106 may have a higher data volume. Since the encoded video bitstream 108 has a lower data volume (less data) compared with the video stream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared with the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder assembly 106 (e.g., configured to transmit uncompressed video to the network 110).

[0028] One or more networks 110 represent any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0029] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content such as the encoded video stream from the source device 102). The server system 112 includes a codec assembly 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec assembly 114 includes an encoder assembly and / or a decoder assembly. In various embodiments, the codec assembly 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the codec assembly 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstream 108. In some embodiments, the server system 112 acts as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim the encoded video bitstream 108 to tailor potentially different bitstreams to one or more electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0030] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be presented on a display or other type of presentation device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0031] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smartphone, a tablet, or a laptop), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0032] In an example operation of the communication system 100, the source device 102 sends the encoded video stream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video stream 108 and may use the codec component 114 to decode and / or encode the encoded video stream 108. For example, the server system 112 may apply an encoding that is better suited for network transmission and / or storage to the video data. The server system 112 may send the encoded video data 116 (e.g., one or more encoded video streams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.

[0033] Figure 2AFIG. 0 is a block diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously acquired / prepared video. In some embodiments, the video source 104 is a camera that acquires local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial array of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples.

[0034] The encoder component 106 is configured to encode and / or compress pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder component 106 is configured to perform a conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to these units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. A person of ordinary skill in the art can easily identify other functions of the controller 204 as they may relate to optimizing the encoder component 106 for a particular system design.

[0035] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and (a) reference picture(s)) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder creates sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the contents in the reference picture memory 208 are also bit-exact corresponding between the local encoder and the remote encoder. In this way, the reference picture samples interpreted by the prediction part of the encoder are the same as the sample values that the decoder will interpret when using prediction during decoding. This principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well-known to those of ordinary skill in the art.

[0036] The operation of the decoder 210 can be the same as that of, for example, a remote decoder (such as the decoder component 122) described in detail below in connection with Figure 2B . Additionally, briefly referring to Figure 2B , however, when the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.

[0037] Except for parsing / entropy decoding, the decoder techniques described herein can exist in a corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder techniques can be simplified because encoder techniques can be reciprocal to decoder techniques.

[0038] As part of its operation, the source encoder 202 can perform motion compensated predictive coding, referring to one or more previously encoded frames in the video sequence designated as reference frames, and this motion compensated predictive coding performs predictive coding on the input frame. In this way, the coding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the reference frame, and the reference frame can be selected as the prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding the video data.

[0039] The decoder 210 decodes the encoded video data of the frames that can be specified as reference frames based on the symbols created by the source encoder 202. The operation of the encoding engine 212 can advantageously be a lossy process. When the encoded video data is decoded at the video decoder ( Figure 2A not shown), the reconstructed video sequence may be a copy of the source video sequence with some errors. The decoder 210 duplicates the decoding process that can be performed by a remote video decoder on the reference frames and enables the reconstructed reference frames to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frames, which has the same content (without transmission errors) as the reconstructed reference frames to be obtained by the remote video decoder.

[0040] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata that can serve as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor 206 can operate on a per-pixel block basis of the sample blocks to find a suitable prediction reference. According to the search results obtained by the predictor 206, it can be determined that the input picture may have a prediction reference taken from multiple reference pictures stored in the reference picture memory 208.

[0041] The outputs of all the above functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 converts the symbols into an encoded video sequence by losslessly compressing the symbols generated by various functional units according to techniques known to those of ordinary skill in the art, such as Huffman coding, variable length coding, and / or arithmetic coding.

[0042] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer the encoded video sequence created by the entropy encoder 214 to prepare for transmission over the communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter 440 can be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter can transmit additional data when transmitting the encoded video. The source encoder 202 can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and redundant slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0043] The controller 204 can manage the operation of the encoder component 106. During encoding, the controller 204 can assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi - predictive picture (B picture). An intra picture can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are aware of the variants of I pictures and their corresponding applications and characteristics, so they will not be repeated here. Predictive pictures can be encoded and decoded using intra - prediction or inter - prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi - predictive pictures can be encoded and decoded using intra - prediction or inter - prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0044] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block - by - block. These blocks can be predictively encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be non - predictively encoded, or the block can be predictively encoded with reference to already - encoded blocks of the same picture (spatial prediction or intra - prediction). Pixel blocks of a P picture can be non - predictively encoded with reference to a previously encoded reference picture through spatial prediction or through temporal prediction. Blocks of a B picture can be non - predictively encoded with reference to one or two previously encoded reference pictures through spatial prediction or through temporal prediction.

[0045] The captured video can be a plurality of source pictures (video pictures) in a time series. Intra - picture prediction (usually abbreviated as intra - prediction) utilizes the spatial correlation within a given picture, while inter - picture prediction utilizes the (temporal or other) correlation between pictures. In an embodiment, the specific picture being encoded / decoded is segmented into blocks, and the specific picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded with a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0046] The encoder component 106 may perform encoding operations according to any predetermined video coding technique or standard such as those described herein. In operation, the encoder component 106 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video coding technique or standard being used.

[0047] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to a loop filter 256, the transmitter being configured to send data to the display 124 (e.g., via a wired or wireless connection).

[0048] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218, the receiver being configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from the channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data when receiving the encoded video. This additional data may be part of the encoded video sequence. The additional data may be used by the decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0049] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The decoder component 122 is implemented at least in part in software.

[0050] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to prevent network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 inside the decoder component 122 (e.g., which is configured to handle playback timing), a separate buffer memory is provided outside the decoder component 122 (e.g., to prevent network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, it may also be unnecessary to configure the buffer memory 252, or the buffer memory can be made smaller. For use on a service packet network such as the Internet, the buffer memory 252 may also be required, which can be relatively large and / or can have an adaptive size, and can be implemented at least partially in the operating system or a similar element outside the decoder component 122.

[0051] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols can include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device can be, for example, in the form of a Supplementary Enhancement Information (SEI) message or a fragment of a Video Usability Information (VUI) parameter set (not shown). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be performed according to a video coding technology or standard, and can follow various principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser 254 can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups can include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser 254 can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0052] Depending on the type of the encoded video picture or a part thereof (e.g., inter - picture and intra - picture, inter - block and intra - block) and other factors, the reconstruction of symbol 270 may involve multiple different units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by parser 254 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between parser 254 and the multiple units below are not depicted.

[0053] Decoder component 122 can be conceptually subdivided into multiple functional units, and in some embodiments, these units interact closely with each other and can be at least partially integrated with each other. However, for the sake of clarity, the conceptually subdivided functional units are retained herein.

[0054] Scaler / inverse transform unit 258 receives the quantized transform coefficients as symbol 270 and control information (e.g., which transform mode to use, block size, quantization factor, and / or quantization scaling matrix, etc.) from parser 254. Scaler / inverse transform unit 258 may output a block including sample values, and the sample values can be input into aggregator 268.

[0055] In some cases, the output samples of scaler / inverse transform unit 258 belong to intra - coded blocks; that is, blocks that do not use predictive information from previously reconstructed pictures but may use predictive information from previously reconstructed parts of the current picture. Such predictive information may be provided by intra - picture prediction unit 262. Intra - picture prediction unit 262 may use the surrounding reconstructed information extracted from the current (partially reconstructed) picture in current picture memory 264 to generate a block having the same size and shape as the block being reconstructed. Aggregator 268 may add the prediction information generated by intra - picture prediction unit 262 to the output sample information provided by scaler / inverse transform unit 258 based on each sample.

[0056] In other cases, the output samples of scaler / inverse transform unit 258 belong to inter - coded and potentially motion - compensated blocks. In this case, motion - compensation prediction unit 260 may access reference picture memory 266 to extract samples for prediction. After motion - compensating the extracted samples according to symbol 270 related to the block, these samples may be added by aggregator 268 to the output of scaler / inverse transform unit 258 (referred to as residual samples or residual signal in this case), thereby generating output sample information. The extraction of prediction samples by motion - compensation prediction unit 260 from an address within reference image memory 266 may be controlled by a motion vector. The motion vector may be provided to motion - compensation prediction unit 260 in the form of symbol 270, and symbol 270 may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values extracted from reference picture memory 266 when using sub - sample accurate motion vectors, motion - vector prediction mechanisms, etc.

[0057] The output samples of aggregator 268 can be adopted by various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in the encoded video bitstream, and the parameters can be used for loop filter unit 256 as symbols 270 from parser 254. However, video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) parts of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values. The output of loop filter unit 256 can be a sample stream, which can be output to a rendering device (such as display 124), and stored in reference picture memory 266 for subsequent inter-picture prediction.

[0058] Once reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is reconstructed and the encoded picture (by, for example, parser 254) is identified as a reference picture, the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0059] Decoder component 122 can perform decoding operations according to a predetermined video compression technique recorded in any standard such as those described herein. An encoded video sequence can conform to the syntax specified by the video compression technique or standard used in the sense that the encoded video sequence follows the video compression technique or standard (especially the profile therein) specified in the video compression technique document or standard. Additionally, to conform to certain video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined at the video compression technique or standard level. In some cases, the layer limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the layer can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.

[0060] Figure 3is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).

[0061] The network interface 304 may be configured to interact with one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication network may be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay tolerant network, etc. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, and so on. Such communication may be only one-way receiving (e.g., broadcast television), only one-way transmitting (e.g., CANBus connected to certain CANbus devices), or two-way, e.g., using a local area network or a wide area network digital network to connect to other computer systems. Such communication may include communication to one or more cloud computing networks.

[0062] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 may include one or more of the following devices: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 may include one or more of the following devices: audio output devices (e.g., speakers), visual output devices (e.g., displays or viewers), etc.

[0063] The memory 314 may include high-speed random access memory (e.g., DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices located remotely from the control circuit 302. The memory 314, or alternatively, the non-volatile solid-state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: An operating system 316 that includes programs for handling various basic system services and for performing hardware-related tasks; A network communication module 318 for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); A codec module 320 for performing various functions regarding encoding and / or decoding data (e.g., video data). In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: A decoding module 322 for performing various functions regarding decoding encoded data, such as those functions previously described with respect to the decoder component 122; An encoding module 340 for performing various functions regarding encoding data, such as those functions previously described with respect to the encoder component 106; and A picture memory 352 for storing pictures and picture data, e.g., for use with the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0064] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform the various functions previously described with respect to the parser 254), a transform module 326 (e.g., configured to perform the various functions previously described with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described with respect to the loop filter 256).

[0065] In some embodiments, the encoding module 340 includes an encoder module 342 (e.g., configured to perform the various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 A subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.

[0066] Each of the above-identified modules stored in the memory 314 corresponds to a set of instructions for performing the functions described herein. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus, in various embodiments, subsets of these modules may be combined or otherwise rearranged. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but rather uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores subsets of the above modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0067] Although Figure 3 FIG. 112 shows a server system 112 according to some embodiments, Figure 3 it is more intended as a functional description of the various features that may exist in one or more server systems than as a structural diagram of the embodiments described herein. In fact, as will be appreciated by one of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in FIG. may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among them will vary depending on the implementation and, optionally, in part on the data traffic processed by the server system during peak usage as well as during average usage.

[0068] Figure 4 FIG. 13 is a flow chart showing an example process 400 of applying in-loop filtering in video decoding. A GOP includes a sequence of picture frames, which further includes a current picture frame. The current picture frame includes a color picture, i.e., a non-monochrome picture frame, which has a plurality of co-located color samples (e.g., chroma samples 402 and luma samples 404). After the plurality of color samples in the current picture frame are reconstructed, in-loop filtering is applied to adjust a subset of the color samples, thereby improving the picture quality of the current picture frame. In some embodiments, the reconstructed samples of a first color component and their adjacent reconstructed samples are combined to derive an offset value for a second color component; the reconstructed samples of the second color component are co-located with the reconstructed samples of the first color component, and the reconstructed samples of the second color component are adjusted by the offset value. Optionally, in some embodiments, the reconstructed samples of the first color component are processed to generate a set of adapted samples. The adapted samples of the first color component and their adapted adjacent samples (e.g., adapted luma samples 404A) are combined to derive an offset value for a second color component (e.g., luma samples 404, chroma samples 402). The samples of the second color component are adjusted by the offset value. Optionally, the first color component is the same as or different from the second color component.

[0069] For example, a set of luminance samples 404 is processed to generate a plurality of adapted luminance samples 404A, with a first luminance sample 404A co-located with an adapted first luminance sample 404AC. The adapted first luminance sample 404AC and its adapted neighboring luminance samples 404AX are combined to derive a sample offset 406, which is used to adjust the first luminance sample 404C itself. In another example, the adapted first luminance sample 404AC and its adapted neighboring luminance samples 404AX are combined to derive a sample offset 406, with a first chrominance sample 402C co-located with the adapted first luminance sample 404AC, and the first chrominance sample 402C is adjusted by the sample offset 406. To combine the adapted luminance samples 404AC and 404AX, a loop filter 256 is applied to determine one or more of the number, position, and weight of the adapted neighboring luminance samples 404AX for generating the sample offset 406.

[0070] More specifically, the decoder 122 receives a video bitstream 116 including a current image frame from the encoder 106. The video bitstream 116 includes a first syntax element. A cross-component sample offset (CCSO) mode 408 indicates whether a first sample offset 406 of a first color sample 410 of the current image frame is determined based on one or more luminance samples 404 for the CCSO mode. The first syntax element has a first predetermined value indicating that the CCSO mode is enabled. The decoder 122 identifies a set of reconstructed luminance samples 404 and generates a set of adapted luminance samples 404A based on the set of reconstructed luminance samples 404. The set of adapted luminance samples 404A includes an adapted first luminance sample 404AC and its adapted neighboring luminance samples 404AX. The set of reconstructed luminance samples 404 includes a first luminance sample 404C co-located with the first color sample 410 and one or more neighboring luminance samples 404X of the first luminance sample 404C. The first sample offset 406 of the first color sample 410 is determined based on the adapted first luminance sample 404AC and one or more adapted neighboring luminance samples 404AX. The decoder 122 reconstructs the current image frame at least by adjusting the first color sample 410 based on the first sample offset 406. In some embodiments, the first color sample 410 is one of the following: the first luminance sample 404C, the first blue-difference chrominance (Cb) sample 402Cb, and the first blue-difference chrominance (Cr) sample 402Cr. The first luminance sample 404C, the first Cb sample 402Cb, and the first Cr sample 402Cr are co-located with each other.

[0071] In some embodiments, a set of reconstructed luma samples 404 is downsampled by one or more downsampling filters 422 to generate adapted luma samples 404A for the CCSO mode 408, based on the resolution of the chroma samples 402 co-located with the set of reconstructed luma samples 404. Further, in some embodiments, one or more downsampling filters 422 are predefined and used in the encoder 106 and the decoder 122( Figure 1 ). For example, the downsampling filter 422 is selected from a plurality of predefined filters 424 stored in the encoder 106 and the decoder 122. A subset of the predefined filters 424 can be selected as the downsampling filter 422 by a syntax element. Optionally, a subset of the predefined filters 424 is selected based on other coded information.

[0072] In some embodiments, the downsampling filter 422 is selected from a plurality of predefined filters 424. The downsampling filter 422 is used based on a set of reconstructed luma samples 404 to generate a set of adapted luma samples 404A. The downsampling filter 422 is also applied to the cross-component intra prediction (CCIP) mode to determine a first chroma sample 402C co-located with a first luma sample 404C, based on a set of reconstructed luma samples 404. Further, in some embodiments,

[0073] In some embodiments, one or more luma filters 426 are applied to a set of reconstructed luma samples 404 to generate a set of adapted luma samples 404A, where the resolution of the reconstructed luma samples 404 is the same as that of the adapted luma samples 404A. For example, the filter type of the luma filter 426 has a cross shape and includes four taps. Adjacent luma samples 404X include: a north luma sample 404N (also referred to as an upper luma sample), a south luma sample 404S (also referred to as a lower luma sample), a west luma sample 404W (also referred to as a left luma sample), and an east luma sample 404E (also referred to as a right luma sample). The adapted first luma sample 404AC is co-located with the first luma sample 404C, and the adapted first luma sample 404AC is a weighted combination of the first luma sample 404C and the luma samples 404N, 404S, 404W, 404E.

[0074] In some embodiments, the decoder 122 generates a set of adapted luma samples 404A by applying a luma filter 426 to a set of reconstructed luma samples 404. The decoder 122 applies a cross-component Wiener filter 428 to process the set of reconstructed luma samples 404 when generating the set of adapted luma samples 404A, thereby reducing the noise level of the set of reconstructed luma samples 404. The Wiener filter 428 is a linear filter that is used to reduce the noise (e.g., mean square error) in the adapted luma samples 404A and enhance the quality of the luma information in the current image frame. The filter type of the Wiener filter 428 can be the same as the filter type of the luma filter 426 used for the CCSO mode 408.

[0075] In some embodiments, the resolution of the first color samples 410 (e.g., chroma samples 402) is lower than the resolution of the set of adapted luma samples 404A (e.g., the same as the resolution of the luma samples 404). The first color samples 410 are physically co-located with at least one subset of adapted luma samples 432 (e.g., a 2×2 luma sample array). One sample in the subset of adapted luma samples 432 (e.g., the left luma sample in the 2×2 luma sample array) is selected as the adapted first luma sample 404AC co-located with the first color sample 410 to determine the first sample offset 406 of the first color sample 410. Additionally, in some embodiments, the subset of adapted luma samples 432-1 includes a target luma sample C, a right luma sample R, a bottom luma sample B, and a bottom-right luma sample RB. Further, in some embodiments, the subset of adapted luma samples 432-2 further includes a top-left luma sample LT, a left luma sample L, a bottom-left luma sample LB, a top luma sample T, and a top-right luma sample RT in addition to the four luma samples in the subset of adapted luma samples 432-1. Optionally, in some embodiments, the subset of adapted luma samples 432 includes a target luma sample C and a set of N adapted adjacent luma samples 404A surrounding the target luma sample C, where the target luma sample L shares the upper left corner with the first color sample 410 and N is a positive integer. In one example, the subset of adapted luma samples 432-2 includes a target luma sample C (sharing the upper left corner with the first color sample 410) and eight adjacent luma samples 404AX. One of the eight adjacent luma samples 404AX is used as the center position to determine the sample offset 406.

[0076] In some embodiments, the CCSO mode 408 corresponds to a band offset classifier 412B. Based on the band offset classifier 412B, the decoder 122 determines that a set of adapted luminance samples 404A includes an adapted first luminance sample 404AC and one or more adapted adjacent luminance samples 404AX. The set of adapted luminance samples 404A is provided to the quantizer 430 and used to generate one or more quantization values 404Q, which are then used by the band offset classifier 412B to classify the first color sample 410. For example, in the CCSO mode 408, the filter type is cross-shaped and includes four taps. The set of adapted adjacent luminance samples 404AX includes a north-adapted luminance sample, a south-adapted luminance sample, a west-adapted luminance sample, and an east-adapted luminance sample, and these adapted luminance samples 404AX are respectively quantized into quantization values 404QN, 404QS, 404QW, and 404QE. The adapted first luminance sample 404AC is quantized into a quantization value 404QC.

[0077] In some embodiments, the CCSO mode 408 corresponds to at least one edge offset classifier 412E. Based on the edge offset classifier 412E, the decoder 122 determines that one or more adapted luminance samples 404A include an adapted first luminance sample 404AC and one or more adapted adjacent luminance samples 404AX, and further determines one or more adapted differences between the one or more adapted adjacent luminance samples and the adapted first luminance sample. One or more quantization values 404Q are generated based on the one or more adapted differences and are used by the edge offset classifier 412E to classify the first color sample 410. For example, in the CCSO mode 408, the filter type is cross-shaped and includes four taps. The one or more adapted adjacent luminance samples include a north-adapted luminance sample, a south-adapted luminance sample, a west-adapted luminance sample, and an east-adapted luminance sample. The decoder 122 determines one or more adapted differences between the one or more adapted adjacent luminance samples 404AX and the adapted first luminance sample 404AC. For example, the one or more adapted differences include one or more of the following: a north-adapted difference, a south-adapted difference, a west-adapted difference, and an east-adapted difference. Each adapted difference is respectively the difference between a corresponding adapted adjacent luminance sample in the adapted adjacent luminance samples 404AX and the first luminance sample 404AC. The one or more differences are quantized to generate one or more quantization values 404QX. For example, the one or more quantization values 404QX include one or more of the following: a north quantization value 404QN, a south quantization value 404QS, a west quantization value 404QW, and an east quantization value 404QE. Each adapted difference is provided to the quantizer 430 and is quantized to generate a corresponding one of the quantization values 404QN, 404QS, 404QW, and 404QE. The quantization value 404QC is equal to 0.

[0078] Classifier 412 classifies the first color sample 410 based on the quantization value 404Q to determine the first sample offset 406 of the first color sample 410. In one example, the quantization value 404Q includes quantization values 404QC, 404QN, 404QS, 404QW, and 404QE. The look-up table 414 maps multiple combinations of the quantization values 404QC, 404QN, 404QS, 404QW, and 404QE to different sample offset options SO (e.g., SO1 - SQ16). Based on the look-up table 414, the quantization value 404Q corresponds to one of the combinations in the look-up table 414, and the corresponding sample offset option SO is identified as corresponding to a combination of the quantization value 404Q and is thus selected for the first sample offset 406. In other words, in some embodiments, the decoder 122 classifies the first color sample 410 by: identifying one or more combinations of the quantization value 404Q in the look-up table 414 (the look-up table associates multiple quantization combinations with multiple offset value options SO (e.g., SO1 - SO16)); and determining the first sample offset 406 corresponding to the one or more combinations of the quantization value 404Q in the look-up table 414.

[0079] In some embodiments, a scalar quantizer 430 including multiple quantization intervals 418 (QI) and multiple quantization levels 420 (QL) quantizes the adapted luminance sample 404A or the adapted difference into multiple integer values within the quantization range 416, and each of the one or more quantization values 404Q includes a corresponding integer within the quantization range 416. For each integer value in the quantization range 416, the quantization interval 418 is defined as the value range assigned to the corresponding integer value. The quantization level 420 corresponds to the corresponding integer value to which a difference range associated with the quantization interval 518 is assigned.

[0080] The first color sample 410 is adjusted based on the first sample offset 406 of the first color sample 410, enabling the reconstruction of the current image frame. In some embodiments, the first color sample 410 includes a first chrominance sample 402C co-located with the first luminance sample 404C in the current image frame, and the first chrominance sample 402C is adjusted based on the first sample offset 406. Optionally, in some embodiments, the first color sample 410 is the first luminance sample 404C, and the first luminance sample 404C is adjusted based on the first sample offset 406.

[0081] Figure 5Schematic diagram showing information 500 embedded in video bitstream 116 according to some embodiments. Decoder 122 receives video bitstream 116 including the current image frame from encoder 106. Video bitstream 116 includes first syntax element 520. CCSO mode 408 indicates whether the first sample offset 406 of first color sample 410 of the current image frame is determined based on one or more luma samples 404 for cross-component sample offset (CCSO) mode 408. The first syntax element has a first predetermined value that indicates that the CCSO mode is enabled. Decoder 122 identifies a set of reconstructed luma samples 404 and generates a set of adapted luma samples 404A including adapted first luma sample 404AC and its adapted adjacent luma samples 404AX based on the set of reconstructed luma samples 404. The set of reconstructed luma samples 404 includes first luma sample 404C co-located with first color sample 410 and one or more adjacent luma samples 404X of first luma sample 404C. The first sample offset 406 of first color sample 410 is determined based on adapted first luma sample 404AC and one or more adapted adjacent luma samples 404AX. Decoder 122 reconstructs the current image frame at least by adjusting first color sample 410 based on first sample offset 406.

[0082] In some embodiments, video bitstream 116 further includes a first high-level flag 502 indicating whether loop filtering is enabled. According to determining that the first high-level flag 502 indicates that loop filtering is enabled, decoder 122 identifies a third high-level syntax element 504 (e.g., an index), and the third high-level syntax element 504 determines a luma filter 426 that is used to generate adapted luma samples 404A based on reconstructed luma samples 404. For example, decoder 122 selects luma filter 426 from a predefined set of filters stored in decoder-side memory using the index. In contrast, according to determining that the first high-level flag 502 indicates that loop filtering is disabled, decoder 122 uses the set of reconstructed luma samples 404 as the set of adapted luma samples 404A, e.g., without using luma filter 426. The first sample offset 406 of first color sample 410 is determined based on the set of reconstructed luma samples 404 (e.g., not based on the set of adapted luma samples 404A).

[0083] In some embodiments, a downsampling filter 422 is selected from a plurality of predefined filters 424. Using the downsampling filter 422, a set of adapted luminance samples 404A is generated based on a set of reconstructed luminance samples 404. The downsampling filter 422 is applied in the cross-component intra prediction (CCIP) mode to determine a first chrominance sample 402C collocated with a first luminance sample 404C based on a set of reconstructed luminance samples 404. Additionally, in some embodiments, the video bitstream 116 further includes a first high-level syntax element 506 for the type of the downsampling filter 422, which is configured to downsample a set of reconstructed luminance samples 404 in the CCIP mode to predict one or more chrominance samples 402. The first high-level syntax element 506 is written into one of the following: sequence header, picture header, sub-picture header, slice header, tile header, and super-block header. The downsampling filter 422 is selected based on the first high-level syntax element 506.

[0084] In some embodiments, the video bitstream 116 further includes a second high-level syntax element 508 for the type of the downsampling filter, which is configured to downsample a set of reconstructed luminance samples 404 in the CCSO mode 408 to generate a set of adapted luminance samples 404A. The second high-level syntax element is written into one of the following: sequence header, picture header, sub-picture header, slice header, tile header, and super-block header. The downsampling filter 422 is selected based on the second high-level syntax element 508. Although the type of the downsampling filter 422 is written together with the CCSO mode 408, the downsampling filter 422 is also used for chrominance prediction according to luminance samples in the CCIP mode.

[0085] In some embodiments, the video bitstream 116 further includes a fourth high-level syntax element 510 having a plurality of bits (e.g., an index). According to determining that the plurality of bits of the fourth high-level syntax element 510 correspond to a first predetermined value (e.g., "000"), the decoder 116 uses a set of reconstructed luminance samples 404 as a set of adapted luminance samples 404A. In other words, the CCSO mode 408 is disabled. A first sample offset 406 of the first color sample 410 is determined based on a set of reconstructed luminance samples 404. According to determining that the plurality of bits of the fourth high-level syntax element 510 correspond to a second predetermined value different from the first predetermined value (e.g., "100"), the decoder 116 determines a luminance filter 426 for generating adapted luminance samples 404A based on the second predetermined value according to the reconstructed luminance samples 404. For example, the plurality of bits has three bits. The first bit indicates whether the CCSO mode 408 is enabled, and if the CCSO mode 408 is enabled, the last two bits are used to select one from a plurality of predefined filters as the luminance filter 426. The second predetermined values equal to "100", "101", "110", and "111" respectively correspond to four different types of luminance filters.

[0086] In some embodiments, the first color sample 410 includes a first Cb sample 402Cb and a first Cr sample 402Cr that are co-located with the first luminance sample 404C of the current image frame. The video bitstream 116 also includes a second syntax element 522 for the CCSO mode 408, and the second syntax element 522 indicates whether a second sample offset 406R of the first Cr sample 402Cb is determined based on one or more luminance samples 404. The video bitstream 116 also includes two different high-level indexes 524R and 524B, and the two different high-level indexes 524R and 524B indicate whether two downsampling filters 422 are applied to generate the first Cb sample 402Cb and the first Cr sample 402Cr, respectively. Each of the two downsampling filters 422 is applied to cross-component intra prediction, loop filtering, or both of their respective chrominance samples. For example, the first downsampling filter is applied to one or more of the first Cb sample 402Cb predicted from the luminance sample 404, Wiener filtering of the first Cb sample 402Cb, and the CCSO mode 408 associated with the first Cb sample 402Cb.

[0087] In some embodiments, the first color sample 410 includes one of the first Cb sample 402Cb and the first Cr sample 402Cr, and the first Cb sample 420Cb and the first Cr sample 402Cr are co-located with the first luminance sample 404C. The video bitstream also includes a common high-level index 524R, and the common high-level index 524R indicates whether the downsampling filter 422 is applied to generate the first Cb sample 402Cb and the first Cr sample 402Cr. The downsampling filter 422 is applied to cross-component intra prediction and / or loop filtering of the first Cb sample 402Cb and the first Cr sample 402Cr.

[0088] In some embodiments, the resolution of the first color sample 410 (e.g., chrominance sample 402) is lower than the resolution of a set of adapted luminance samples 404A (e.g., having the same resolution as the luminance sample 404). The first color sample 410 is physically co-located with at least one subset of adapted luminance samples 432 (e.g., a 2×2 luminance sample array). One sample in the subset of adapted luminance samples 432 (e.g., the left luminance sample in the 2×2 luminance sample array) is selected as the adapted first luminance sample 404AC co-located with the first color sample 410 to determine the first sample offset 406 of the first color sample 410. Additionally, in some embodiments, the video bitstream 116 also includes a fifth high-level syntax element 512, and the fifth high-level syntax element 512 is used to select one sample in the subset of adapted luminance samples 432 (e.g., Figure 4Among LT, T, RT, L, C, R, LB, B, or RB), the selected sample is used to determine a first sample offset 406 of a first color sample 410 in the CCSO mode 408, and a fifth high-level syntax element 512 is written at a frame level or for a first color component corresponding to the first color sample 410.

[0089] In some embodiments, the video bitstream 116 further includes a second high-level flag 514 indicating whether a luminance filter 426 is applied in loop filtering (e.g., in the CSSO mode 408). Based on determining that the second high-level flag 514 indicates that the luminance filter 426 is applied in loop filtering (e.g., when the second high-level flag 514 is "1"), the decoder 122 identifies a dual-functional index 516 and determines, based on the dual-functional index 516, a luminance filter 426 for generating an adapted luminance sample 404A from the reconstructed luminance samples 404. For example, the luminance filter 426 is selected from a plurality of predefined filters, and different values of the dual-functional index 516 correspond to different luminance filter types 526. Conversely, based on determining that the second high-level flag 514 indicates that the luminance filter 426 is not applied in loop filtering (e.g., when the second high-level flag 514 is "0"), the decoder 122 identifies the dual-functional index 516 and selects, based on the dual-functional index 516, a sample from a set of adapted luminance samples 404A to determine a first sample offset 406 of a first color sample 410 in the CCSO mode 408. For example, a sample from a set of adapted luminance samples 404A is selected to be co-located with the first color sample 410, and the first sample offset 406 is jointly determined using the associated adapted neighboring luminance samples 404AX and a sample from the selected set of adapted luminance samples 404A. Different values of the dual-functional index 516 correspond to different selected adapted luminance samples 528 (e.g., LT, T, RT, L, C, R, LB, B, or RB).

[0090] Figure 6 is a flowchart illustrating an example video decoding method 600 according to some embodiments. The method 600 may be executed on a computing system (e.g., Figure 1 the server system 112, the source device 102, or the electronic device 120 in) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, the method 600 is jointly applied with one or more video codecs, including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In various embodiments of the present application, before inputting to a cross-component offset filter, one or more filters (e.g., the luminance filter 426) are first applied to samples of a first color component, and then the filtered samples (e.g., the adapted luminance samples 404A) are used to calculate an offset value 406.

[0091] In some embodiments, one or more filters (e.g., luminance filter 426) are applied to samples of a first color component (e.g., luminance samples 404) according to a chroma format. For example, for a 4:2:0 chroma format, downsampling filters 422 are applied both in the horizontal and vertical directions. The output samples of the filtering process (e.g., adapted luminance samples 404A) are used to calculate an offset value 406 applied to a second color component (e.g., luminance samples 404, chroma samples 402). In one example, the one or more filters include downsampling filters 422 for performing downsampling operations ( Figure 4 ). Conversely, in another example, one or more filters (e.g., Figure 4 luminance filter 426 therein) are used to filter reconstructed samples (e.g., luminance samples 404) without changing the resolution. In other words, the one or more filters are not downsampling filters. In some embodiments, the first color component is luminance and the second color component is a chroma component (e.g., blue-difference chroma (Cb) samples, chroma-difference chroma (Cb) samples). Optionally, in some embodiments, both the first and second color components are luminance. In some embodiments, one or more filters are used to perform downsampling operations, and these downsampling filters 422 are predefined (e.g., selected from predefined filters 424) and are used both in the encoder 106 and the decoder 122.

[0092] In some embodiments, one or more filters are used to perform downsampling operations, and the three downsampling filters used in a cross-component intra prediction mode (e.g., predicting chroma from luminance mode) are also used for this downsampling process. A high-level (frame / sequence) level filter selection indicator for the cross-component intra prediction mode indicates which filter to use in this downsampling process. That is, the downsampling filter selection is consistent and shared between the cross-component intra prediction module and the cross-component loop filtering module. The cross-component intra prediction module is configured to determine chroma samples 402 based on luminance samples 404. The cross-component loop filtering module is configured to generate a sample offset 406 in loop filtering.

[0093] Optionally, in some embodiments, one or more filters are used to perform downsampling operations, and the three downsampling filters used in a cross-component intra prediction mode (e.g., predicting chroma from luminance mode) are also used for this downsampling process. A dedicated high-level (frame / sequence) level indicator is written for the cross-component loop filter downsampling process.

[0094] In some embodiments, a high-level flag is written to indicate whether the filtering process is enabled. If the filtering process is enabled, a high-level index is further written to indicate which filter is used. If the downsampling process is disabled, the non-downsampled luma samples 404 are used to calculate the sample offset 406 for the chroma samples 402.

[0095] In some embodiments, the high-level index is written using N bits (e.g., 2 bits). If the index is equal to 0, the filtering process is disabled. Otherwise, the filtering process is enabled, and the index indicates which filter is used to generate the adapted luma samples 404A. The level of the high-level index is higher than the block level.

[0096] In some embodiments, the above high-level flag, index, indicator, or syntax element is written separately for each of the two chroma components (Cb and Cr). The filtering and downsampling processes between the two chroma components are controlled independently. In contrast, in some embodiments, the above high-level flag, index, indicator, or syntax element is written jointly for the two chroma components. The filtering and downsampling processes between the two chroma components are jointly controlled (e.g., using the same filter).

[0097] In some embodiments, the cross-component sample offset 406 and the cross-component Wiener filter 428 use the same filter type and are jointly controlled by the filtering and downsampling process decisions (e.g., based on the same syntax element). In other words, the Wiener filter is the same type of filter as the luma filter 426 (e.g., filter shape, position of adjacent samples).

[0098] In some embodiments, the non-downsampled luma samples 404 are used to determine the chroma offset value 406. The center position of the luma samples 404 is co-located with the first chroma sample 402C, and the center position of the luma samples 404 is selected from a subset of the luma samples 404 or 404A. In one example, the co-located luma sample 404C is selected from four positions, namely the co-located luma sample 404C or C, the co-located luma sample 404E or R on the right, the co-located luma sample 404S or B below, the co-located luma sample 404SE or RB in the lower right, and the co-located luma sample 404C is used to determine the chroma offset 406. The high-level index (for each frame or each component of the frame) is written to indicate which position is used. In another example, nine positions including the co-located luma position and the surrounding eight luma positions can be used as the center position to calculate the chroma offset 402. The high-level index (for each frame or each component of the frame) is written to indicate which position is used. In one example, N positions (e.g., eight positions) around the co-located luma position 404C can be used as the center position to calculate the chroma offset. The high-level index (for each frame or each component of the frame) is written to indicate which position is used.

[0099] In some embodiments, co-located luminance positions are selected for selecting adapted luminance samples from a set of adapted luminance samples 404A. An advanced flag (per frame or per frame component) is written to indicate whether downsampling or filtering has been applied. If downsampling or filtering has been applied, an index is written to indicate the type of filter used. Conversely, if downsampling or filtering has not been applied, an index is written to indicate which position is used as the luminance position for calculating the relevant chrominance offset 406.

[0100] Although Figure 6 Although multiple logical stages are shown in a particular order, stages that are not order-dependent can be reordered, and other stages can be combined or split. For those of ordinary skill in the art, some reorderings or other groupings not specifically mentioned will be obvious, so the orderings and groupings given here are not exhaustive. Additionally, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.

[0101] Now turning to some example embodiments.

[0102] (A1) In some implementations, method 600 is implemented for decoding video data. Method 600 includes: receiving (operation 602) a video bitstream including a current image frame, where the video bitstream includes (operation 604) a first syntax element that indicates, for a cross-component sample offset (CCSO) mode, whether a first sample offset of a first color sample of the current image frame is determined based on one or more luminance samples; when the CCSO mode is enabled, generating (operation 608) a set of adapted luminance samples based on a set of reconstructed luminance samples, the set of adapted luminance samples including an adapted first luminance sample and its adapted neighboring luminance samples, the set of reconstructed luminance samples including (operation 610) a first luminance sample co-located with the first color sample of the current image frame; determining (operation 612) the first sample offset of the first color sample based on the adapted first luminance sample and one or more adapted neighboring luminance samples; and reconstructing (operation 614) the current image frame at least by adjusting the first color sample based on the first sample offset.

[0103] (A2) In some embodiments of A1, one or more downsampling filters are applied (operation 616) to a set of reconstructed luminance samples to generate a set of adapted luminance samples based on the resolution of chrominance samples co-located with the set of reconstructed luminance samples.

[0104] (A3) In some embodiments of A2, one or more downsampling filters are predefined and used in both the encoder and the decoder.

[0105] (A4)In some embodiments of A1, one or more luminance filters are applied (operation 618) to a set of reconstructed luminance samples to generate a set of adapted luminance samples, where the resolution of the reconstructed luminance samples is the same as that of the adapted luminance samples.

[0106] (A5)In some embodiments of any one of A1 to A4, the first color sample is one of the following: a first luminance sample, a first blue chrominance difference (Cb) sample, and a first red chrominance difference (Cr) sample, where the first luminance sample, the first Cb sample, and the first Cr sample are co-located with each other.

[0107] (A6)In some embodiments of any one of A1 to A5, method 600 further includes: selecting a downsampling filter from a plurality of predefined filters; and applying the downsampling filter in a cross-component intra prediction (CCIP) mode to determine a first chrominance sample co-located with the first luminance sample based on a set of reconstructed luminance samples; where a set of adapted luminance samples is generated based on a set of reconstructed luminance samples using the downsampling filter.

[0108] (A7)In some embodiments of A6, the video bitstream further includes a first high-level syntax element for the type of the downsampling filter, where the downsampling filter is configured to downsample a set of reconstructed luminance samples in a cross-component intra prediction (CCIP) mode to predict one or more chrominance samples; the first high-level syntax element is written into one of the following: sequence header, picture header, sub-picture header, slice header, tile header, and super-block header; and the downsampling filter is selected based on the first high-level syntax element.

[0109] (A8)In some embodiments of A6 or A7, the video bitstream further includes a second high-level syntax element for the type of the downsampling filter, where the downsampling filter is configured to downsample a set of reconstructed luminance samples in a CCSO mode to generate a set of adapted luminance samples; the second high-level syntax element is written into one of the following: sequence header, picture header, sub-picture header, slice header, tile header, and super-block header; and the downsampling filter is selected based on the second high-level syntax element.

[0110] (A9)In some embodiments of any one of A1 to A8, the video bitstream further includes a first high-level flag indicating whether loop filtering is enabled. Method 600 further includes: identifying a third high-level index for determining a luminance filter for generating adapted luminance samples based on reconstructed luminance samples according to determining that the first high-level flag indicates that loop filtering is enabled; and using a set of reconstructed luminance samples as a set of adapted luminance samples according to determining that the first high-level flag indicates that loop filtering is disabled, where the first sample offset of the first color sample is determined based on a set of reconstructed luminance samples.

[0111] (A10)In some embodiments of any one of A1 to A9, the video bitstream further includes a fourth high-level syntax element having a plurality of bits. Method 600 further includes: using a set of reconstructed luma samples as a set of adapted luma samples according to determining that the plurality of bits of the fourth high-level syntax element correspond to a first predetermined value, wherein a first sample offset of the first color sample is determined based on the set of reconstructed luma samples; and determining a luma filter for generating adapted luma samples from the reconstructed luma samples based on a second predetermined value according to determining that the plurality of bits of the fourth high-level syntax element correspond to a second predetermined value different from the first predetermined value.

[0112] (A11)In some embodiments of any one of A1 to A10, the first color sample includes a first Cb sample co-located with a first luma sample and a first Cr sample of the current picture frame. The video bitstream further includes a second syntax element for the CCSO mode, the second syntax element indicating whether the second sample offset of the first Cr sample is determined based on one or more luma samples. The video bitstream further includes two different high-level indices, the two different high-level indices indicating whether two downsampling filters are applied to generate the first Cb sample and the first Cr sample respectively, wherein each of the two downsampling filters is applied to cross-component intra prediction, loop filtering, or both of their respective chroma samples.

[0113] (A12)In some embodiments of any one of A1 to A10, the first color sample includes one of a first Cb sample and a first Cr sample, the first Cb sample and the first Cr sample being co-located with the first luma sample. The video bitstream further includes a common high-level index indicating whether a downsampling filter is applied to generate the first Cb sample and the first Cr sample, wherein the downsampling filter is applied to cross-component intra prediction, loop filtering, or both of the first Cb sample and the first Cr sample.

[0114] (A13)In some embodiments of any one of A1 to A12, method 600 further includes: applying a cross-component Wiener filter to process a set of adapted luma samples generated according to a set of reconstructed luma samples for the CCSO mode.

[0115] (A14)In some embodiments of any one of A1 to A13, the resolution of the first color sample is lower than the resolution of a set of adapted luma samples, the first color sample is physically co-located with at least one subset of the adapted luma samples, and one sample in the set of adapted luma samples is selected as an adapted first luma sample to determine a first sample offset of the first color sample.

[0116] (A15)In some embodiments of A14, the video bitstream further includes a fifth high-level syntax element for selecting one sample from a set of adapted luma samples, the selected sample being used to determine a first sample offset of a first color sample in the CCSO mode, and the fifth high-level syntax element is written at the frame level or for a first color component corresponding to the first color sample.

[0117] (A16)In some embodiments of A14 or A15, the set of adapted luma samples includes a left luma sample, a right luma sample, a bottom luma sample, and a bottom-right luma sample.

[0118] (A17)In some embodiments of A16, the set of adapted luma samples further includes a top-left luma sample, a left luma sample, a bottom-left luma sample, a top luma sample, and a top-right luma sample.

[0119] (A18)In some embodiments of any one of A14 or A15, the set of adapted luma samples includes a top-left luma sample and a set of N adjacent luma samples surrounding the top-left luma sample, the top-left luma sample sharing the top-left corner with the first color sample, and N is a positive integer.

[0120] (A19)In some embodiments of any one of A1 to A18, the video bitstream further includes a second high-level flag indicating whether a luma filter is applied in loop filtering. Method 600 further includes: identifying a dual-functional index according to determining that the second high-level flag indicates that a luma filter is applied in loop filtering, determining a luma filter for generating adapted luma samples from reconstructed luma samples based on the dual-functional index; and identifying a dual-functional index according to determining that the second high-level flag indicates that a luma filter is not applied in loop filtering, and selecting one sample from a set of adapted luma samples based on the dual-functional index to determine a first sample offset of a first color sample in the CCSO mode.

[0121] (A20)In some embodiments of any one of A1 to A19, determining a first sample offset of a first color sample further includes: generating one or more quantization values based on adapted adjacent luma samples and an adapted first luma sample; and classifying the first color sample based on the one or more quantization values to determine a first sample offset of the first color sample.

[0122] (A21)In some embodiments of A20, generating one or more quantization values further includes: determining one or more differences between adapted adjacent luma samples and an adapted first luma sample, wherein the one or more differences are quantized to generate one or more quantization values.

[0123] (A22) In some embodiments, a method for encoding video data includes: receiving video data including a current picture frame; encoding the current picture frame; enabling a cross-component sample offset (CCSO) mode to generate a first sample offset of a first color sample of the current picture frame based on one or more luma samples, wherein, in the CCSO mode, the first sample offset of the first color sample is determined based on a set of adapted luma samples, and the set of adapted luma samples is further generated based on a set of reconstructed luma samples including a first luma sample collocated with the first color sample of the current picture frame; transmitting the encoded current picture frame through a video bitstream; and writing a first syntax element through the video bitstream to indicate that the CCSO mode is applied to reconstruct the first color sample collocated with the first luma sample based on the first sample offset.

[0124] (A23) In some embodiments, a method for generating a bitstream includes: obtaining a source video sequence including a current picture frame; and performing a conversion between the source video sequence and a video bitstream, wherein the video bitstream includes: the current picture frame; and a first syntax element for a cross-component sample offset (CCSO) mode, the first syntax element indicating whether to generate a first sample offset of a first color sample of the current picture frame based on one or more luma samples; wherein the first sample offset of the first color sample is determined based on a set of adapted luma samples generated based on a first luma sample and one or more neighboring luma samples, and the first luma sample is collocated with the first color sample.

[0125] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing one or more sets of instructions configured to be executed by the control circuit, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 - A23 above).

[0126] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by a control circuit of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 - A23 above).

[0127] The proposed methods can be used alone or in any combination in any order. Additionally, each method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). For example, one or more processors execute a program stored in a non-transitory computer-readable medium. Hereinafter, the term block may be interpreted as a prediction block, a coding block, or a coding unit (CU).

[0128] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0129] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It should be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated feature, integer, step, operation, element, and / or component, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0130] As used herein, depending on the context, the term "if" may be interpreted to mean "when...", or "after...", or "in response to determining...", or "in accordance with determining...", or "in response to detecting" that the prerequisite is true. Similarly, depending on the context, the phrases "if it is determined (that the prerequisite is true)", or "if (the prerequisite is true)", or "when (the conditional prerequisite is true)" may be interpreted to mean "after determining that the prerequisite is true", or "in response to determining that the prerequisite is true", or "in accordance with determining that the prerequisite is true", or "after detecting that the prerequisite is true", or "in response to detecting that the prerequisite is true".

[0131] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Given the above teachings, many modifications and variations are possible. The embodiments were chosen and described in order to best explain the operating principles and the practical application, so that others skilled in the art can implement them.

Claims

1. A method for decoding video data, comprising: Receiving a video code stream including a current image frame, wherein the video code stream includes a first syntax element, and the first syntax element indicates whether a first sample offset of a first color sample of the current image frame is determined based on one or more luma samples for a cross-component sample offset (CCSO) mode; When the CCSO mode is enabled, generating a set of adapted luma samples based on a set of reconstructed luma samples, the set of adapted luma samples comprising an adapted first luma sample and its adapted adjacent luma samples, the set of reconstructed luma samples comprising a first luma sample co-located with a first color sample of the current image frame; determining a first sample offset for the first color sample based on the adapted first luma sample and one or more adapted adjacent luma samples; and The current image frame is reconstructed by at least adjusting the first color samples based on the first sample offset.

2. The method according to claim 1, wherein: One or more downsampling filters are applied to the set of reconstructed luma samples to generate the set of adapted luma samples based on a resolution of chroma samples co-located with the set of reconstructed luma samples.

3. The method according to claim 2, wherein: The one or more down-sampling filters are predefined and used in the encoder and the decoder.

4. The method according to claim 1, wherein: One or more luma filters are applied to the set of reconstructed luma samples to generate the set of adapted luma samples, the reconstructed luma samples having a same resolution as the adapted luma samples.

5. The method according to claim 1, wherein: The first color sample is one of: the first luma sample, a first blue difference chroma (Cb) sample, a first blue difference chroma (Cb) sample, wherein the first luma sample, the first Cb sample, and the first Cr sample are co-located with each other.

6. The method according to claim 1, further comprising: selecting a downsampling filter from a plurality of predefined filters; as well as applying the downsampling filter in a cross-component intra prediction (CCIP) mode to determine a first chroma sample co-located with the first luma sample based on the set of reconstructed luma samples, Wherein, the set of adapted luma samples is generated based on the set of reconstructed luma samples using the downsampling filter.

7. The method according to claim 6, wherein: The video bitstream also includes a first high-level syntax element for a type of the downsampling filter, the downsampling filter being configured to downsample the set of reconstructed luma samples to predict one or more chroma samples in a cross-component intra prediction (CCIP) mode; The first high-level syntax element is written into one of the following: a sequence header, a picture header, a sub-picture header, a slice header, a tile header, and a super-block header; and The down-sampling filter is selected based on the first high-level syntax element.

8. The method according to claim 6, wherein: The video bitstream further includes a second high-level syntax element for the type of the downsampling filter, the downsampling filter being configured to downsample the set of reconstructed luma samples in the CCSO mode to generate the set of adapted luma samples; The second high-level syntax element is written into one of the following: a sequence header, a picture header, a sub-picture header, a slice header, a tile header, and a super-block header; and The down-sampling filter is selected based on the second high-level syntax element.

9. The method according to claim 1, wherein: The video code stream also includes a first advanced flag indicating whether loop filtering is enabled, and the method further includes: Based on determining that the first high-level flag indicates that loop filtering is enabled, identifying a third high-level index for determining a luma filter for generating the adapted luma samples from the reconstructed luma samples; and Based on determining that the first high level flag indicates that loop filtering is disabled, the set of reconstructed luma samples is used as the set of adapted luma samples, wherein a first sample offset for the first color sample is determined based on the set of reconstructed luma samples.

10. The method according to claim 1, wherein: The video code stream also includes a fourth high-level syntax element having a plurality of bits, and the method further includes: in accordance with determining that the plurality of bits of the fourth high-level syntax element corresponds to a first predetermined value, using the set of reconstructed luma samples as the set of adapted luma samples, wherein a first sample offset for the first color sample is determined based on the set of reconstructed luma samples; and In accordance with determining that the plurality of bits of the fourth high-level syntax element corresponds to a second predetermined value different than the first predetermined value, determining a luma filter for generating the adapted luma samples from the reconstructed luma samples based on the second predetermined value.

11. The method according to claim 1, wherein: The first color sample includes a first Cb sample co-located with the first brightness sample and a first Cr sample of the current image frame; The video code stream further includes a second syntax element for CCSO mode, wherein the second syntax element indicates whether to determine a second sample offset of the first Cr sample based on one or more luma samples; as well as The video code stream also includes two different advanced indexes, which indicate whether two downsampling filters are applied to generate the first Cb sample and the first Cr sample respectively, wherein each of the two downsampling filters is applied to cross-component intra-frame prediction, loop filtering, or both of the respective chroma samples.

12. The method according to claim 1, wherein: The first color sample includes one of a first Cb sample and a first Cr sample, and the first Cb sample and the first Cr sample are co-located with the first brightness sample; The video code stream also includes a common advanced index, which indicates whether a downsampling filter is applied to generate the first Cb sample and the first Cr sample, wherein the downsampling filter is applied to cross-component intra-frame prediction, loop filtering, or both of the first Cb sample and the first Cr sample.

13. The method according to claim 1, further comprising: A cross-component Wiener filter is applied to the adapted set of luma samples generated from the set of reconstructed luma samples for the CCSO mode.

14. The method according to claim 1, wherein: The first color sample has a lower resolution than the set of adapted luma samples, the first color sample is physically co-located with at least one of the adapted luma sample subsets, and one of the set of adapted luma samples is selected as the adapted first luma sample to determine a first sample offset for the first color sample.

15. The method according to claim 14, wherein: The video code stream also includes a fifth high-level syntax element for selecting one sample from the set of adapted luma samples, the selected sample being used to determine a first sample offset of the first color sample in the CCSO mode, the fifth high-level syntax element being written at a frame level or for a first color component corresponding to the first color sample.

16. The method according to claim 14, wherein: The set of adapted luma samples includes a left luma sample, a right luma sample, a bottom luma sample, and a bottom right luma sample.

17. The method according to claim 16, wherein: The set of adapted luma samples also includes an upper left luma sample, a left side luma sample, a lower left luma sample, an upper luma sample, and an upper right luma sample.

18. The method according to claim 14, wherein: The set of adapted brightness samples includes an upper left brightness sample and a set of N adjacent brightness samples surrounding the upper left brightness sample, the upper left brightness sample shares an upper left corner with the first color sample, and N is a positive integer.

19. The method according to claim 1, wherein: The video code stream also includes a second high-level flag indicating whether a luminance filter is applied in the loop filtering, and the method further includes: Based on determining that the second high-level flag indicates application of the luma filter in loop filtering, identifying a bifunctional index, and determining the luma filter for generating the adapted luma samples from the reconstructed luma samples based on the bifunctional index; and Based on determining that the second high-level flag indicates that the luma filter is not applied in loop filtering, identifying the dual-function index, and selecting a sample from the set of adapted luma samples based on the dual-function index to determine a first sample offset for the first color sample in the CCSO mode.

20. The method of claim 1, determining a first sample offset for the first color sample further comprising: generating one or more quantized values ​​based on the adapted adjacent luma samples and the adapted first luma sample; as well as The first color sample is classified based on the one or more quantized values ​​to determine a first sample offset for the first color sample.

21. The method according to claim 20, wherein: Generating the one or more quantized values ​​further comprises: One or more difference values ​​between the adapted adjacent luma samples and the adapted first luma sample are determined, wherein the one or more difference values ​​are quantized to generate the one or more quantized values.

22. A computing system comprising: Control circuit; as well as A memory storing one or more programs, the one or more programs configured to be executed by the control circuit, the one or more programs further comprising instructions for: Receiving video data including a current image frame; Encoding the current image frame; enabling a cross-component sample offset (CCSO) mode to generate a first sample offset for a first color sample of the current image frame based on one or more luma samples, wherein in the CCSO mode, the first sample offset for the first color sample is determined based on a set of adapted luma samples, the set of adapted luma samples being further generated based on a set of reconstructed luma samples, the set of reconstructed luma samples including a first luma sample co-located with the first color sample of the current image frame; The current image frame encoded by transmitting the video code stream; as well as A first syntax element is written through the video code stream to indicate application of the CCSO mode to reconstruct the first color sample co-located with the first luma sample based on the first sample offset.

23. A non-transitory computer-readable storage medium storing one or more programs executed by control circuitry of a computing system, the one or more programs comprising instructions for: Obtaining a source video sequence including a current image frame; as well as Performing conversion between the source video sequence and the video code stream, wherein: The video code stream includes: the current image frame; and a first syntax element indicating whether to generate a first sample offset for a first color sample of the current image frame based on one or more luma samples for a cross-component sample offset CCSO mode; The first sample offset of the first color sample is determined based on a set of adapted luma samples, the set of adapted luma samples is generated based on the first luma sample and one or more adjacent luma samples, and the first luma sample is co-located with the first color sample.