Cross-component sample offset (CCSO) with adaptive multi-tap filter classifier
By using loop filter technology and cross-component offset filtering method in video encoding, the problem of difficulty in effectively reducing reconstruction errors in the prior art is solved, and more efficient and higher quality video encoding is achieved.
Patent Information
- Application Number
- CN202380078571.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2023-10-31
- Publication Date
- 2025-06-20
AI Technical Summary
Existing video encoding technologies are difficult to effectively reduce reconstruction errors when compressing video data, especially during cross-component offset filtering.
Using loop filter technology, the offset value is calculated using the co-bit reconstruction sample of the first color component and the adjacent reconstruction sample through the cross-component offset filtering method, and added to the current sample of the second color component to adjust the reconstruction value.
By adjusting the reconstructed image samples, the reconstruction error is significantly reduced and the efficiency and quality of video encoding are improved.
Smart Images

Figure CN120188480A_ABST
Abstract
Description
Cross-Reference to Related Applications
[0001] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 535,468, filed on Aug. 30, 2023, entitled “CCSO with Adaptive Multi-Tap-Filter Classifier”, and this application is a continuation of, and claims the benefit of priority to, U.S. patent application Ser. No. 18 / 497,908, filed on Oct. 30, 2023, entitled “Cross-Component Sample Offset (CCSO) with Adaptive Multi-Tap-Filter Classifiers”. Technical Field
[0002] The disclosed embodiments generally relate to video coding, including but not limited to systems and methods for performing loop filtering (e.g., cross-component offset filtering) on video data. Background Art
[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise convey digital video data over a communication network and / or store the digital video data on a storage device. Since the bandwidth capacity of the communication network is limited and the memory resources of the storage device are limited, video coding may be used to compress the video data according to one or more video coding standards before transmitting or storing the video data.
[0004] A variety of video codec standards have been developed. For example, video coding standards include AOMedia Video1 (AV1), Versatile Video Coding (VVC), Joint Exploration test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Expert Group (MPEG) coding. Video coding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), such prediction methods exploiting the redundancy inherent in the video data. Video coding aims to compress the video data into a form that uses a lower bit rate while avoiding or minimizing degradation of the video quality.
[0005] High Efficiency Video Coding (HEVC), also known as H.265, is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4). Versatile Video Coding (VVC), also known as H.266, is intended as a successor video compression standard to HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (Version 1) and 2022 (Version 2). Alliance for Open Media (AOMedia) Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verified version 1.0.0 (with errata 1) of the specification was released. Summary of the Invention
[0006] As described above, encoding (compression) reduces the bandwidth and / or storage space requirements. As will be described in detail later, both lossless compression and lossy compression can be employed. Lossless compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully preserved during the encoding process and cannot be fully recovered during the decoding process. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be useful for the intended application. The amount of allowed distortion depends on the application. For example, users of some consumer video streaming applications can tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect the various allowed distortions, and a higher allowed distortion generally allows an encoding algorithm with higher loss and higher compression ratio.
[0007] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying loop filters to video (image) compression. A video codec includes multiple functional modules for one or more of the following operations: intra / inter prediction, transform coding, quantization, entropy coding, and in-loop filtering. In-loop filtering techniques are applied to adjust the reconstructed picture samples to further reduce the reconstruction error. A cross-component offset filtering method is implemented to apply the co-located reconstructed samples of a first color component and the associated adjacent reconstructed samples to obtain an offset value, which is added to the current sample of a second color component, thereby adjusting the reconstructed value of the current sample. An example of the first color component is the luminance color component, and an example of the second color component is the chrominance color component. In some embodiments, the first color component and the second color component correspond to the same color component, such as luminance samples.
[0008] According to some embodiments, a method of video decoding is provided. The method includes receiving a video bitstream that includes a current coded block of a current picture frame. The video bitstream includes (i) a first syntax element for a cross-component sample offset (CCSO) mode that indicates whether a first sample offset of a first color sample of the current coded block is determined based on one or more luma samples, and (ii) a second syntax element for a filter type index and a filter shape index. The filter type index indicates one of a plurality of filter types of a loop filter selected for the current coded block, and the filter shape index identifies a set of one or more neighboring positions used by the loop filter. The method further includes determining one or more differences between one or more neighboring luma samples and a first luma sample, quantizing the one or more differences to generate one or more quantized differences, classifying a first color sample based on the one or more quantized differences to determine a first sample offset of the first color sample, and reconstructing the current coded block by at least adjusting the first color sample based on the first sample offset of the first color sample.
[0009] In some embodiments, the first color sample includes a first chroma sample co-located with the first luma sample in the current coded block, and the first chroma sample is adjusted based on the first sample offset. Alternatively, in some embodiments, the first color sample is the first luma sample, and the first luma sample is adjusted based on the first sample offset.
[0010] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.
[0011] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets for execution by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.
[0012] Thus, the present application discloses apparatuses and systems for video coding methods. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video coding.
[0013] The features and advantages described in this specification are not necessarily all-inclusive, and in particular, in view of the accompanying drawings, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in the specification is mainly selected for readability and guidance purposes, and is not necessarily selected to depict or limit the subject matter described in this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To understand the present disclosure in more detail, a more specific description can be made by referring to the features of various embodiments, some of which are shown in the accompanying drawings. However, the drawings only show the relevant features of the present disclosure and are not necessarily considered restrictive, as those skilled in the art will understand when reading this disclosure that the specification of this application may include other valid features.
[0015] Figure 1 is a block diagram showing an example communication system according to some embodiments.
[0016] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.
[0017] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.
[0018] Figure 3 is a block diagram showing an example server system according to some embodiments.
[0019] Figure 4 is a flowchart of an example process for applying in-loop filtering in video decoding according to some embodiments.
[0020] Figure 5A and Figure 5B is a diagram showing the positions of adjacent luma samples of an example two-tap loop filter used in the CCSO mode according to some embodiments.
[0021] Figure 6A and Figure 6B is a diagram showing the positions of adjacent luma samples of an example three-tap loop filter used in the CCSO mode according to some embodiments.
[0022] Figure 7 is a diagram showing the positions of adjacent luma samples of an example five-tap loop filter used in the CCSO mode according to some embodiments.
[0023] Figure 8A is a flowchart of an example process for applying band-based in-loop filtering according to some embodiments, Figure 8BA flowchart of an example process for implementing classification based on a dynamic filter type during in-loop filtering according to some embodiments.
[0024] Figure 9 A flowchart showing a method of video coding according to some embodiments.
[0025] According to common practice, the various features shown in the drawings are not necessarily drawn to scale, and the same reference numerals may be used throughout the specification and drawings to indicate the same features. Detailed Description
[0026] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying in-loop filters to video (image) compression. In-loop filtering techniques are applied to adjust the reconstructed picture samples to further reduce the reconstruction error. A cross-component offset filtering method is implemented to obtain an offset value by applying the co-located reconstructed samples of the first color component and the associated adjacent reconstructed samples, and the offset value is added to the current sample of the second color component to adjust the reconstructed value of the current sample. Specifically, in some embodiments, the decoder receives a video bitstream from the encoder, and the video bitstream includes a current encoded block of a current image frame. The video bitstream includes a filter type index and a filter shape index of the in-loop filter. The filter type index is configured to select one of a plurality of filter types (e.g., 2-tap, 3-tap, and 5-tap) of the in-loop filter for the current encoded block. The filter shape index is configured to identify a set of one or more adjacent positions used by the in-loop filter. Based on the filter type index and the filter shape index of the in-loop filter, the decoder identifies one or more adjacent luminance samples of a first luminance sample. The decoder determines one or more differences between the one or more adjacent luminance samples and the first luminance sample. The one or more differences are quantized to generate one or more quantized differences. For example, a classifier classifies a first color sample based on the one or more quantized differences to determine a first sample offset of the first color sample. The first color sample is adjusted based on the first sample offset of the first color sample, so that the current encoded block can be reconstructed.
[0027] Figure 1 A block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, for use with video-enabled applications such as video conferencing applications, digital television applications, and media storage and / or distribution applications.
[0028] The source device 102 includes a video source 104 (e.g., a camera assembly or media storage) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 may have a high data volume compared to the encoded video bitstreams 108 generated by the encoder component 106. Since the encoded video bitstreams 108 have a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less bandwidth for transmission and less storage space for storage compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to send uncompressed video data to one or more networks 110).
[0029] One or more networks 110 represent any number of networks for conveying information between the source device 102, the server system 112, and / or the electronic device 120, including for example wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0030] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server, or the server system 112 includes a streaming server, which is for example configured to store and / or distribute video content, such as storing and / or distributing the encoded video stream from the source device 102. The server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the encoder component 114 is instantiated as hardware, software, or a combination of hardware and software. In some embodiments, the encoder component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video bitstream 108.
[0031] In some embodiments, the server system 112 functions as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim an encoded video stream 108 for customizing potentially different streams for one or more of the plurality of electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0032] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be displayed on a display or other type of display device. In some embodiments, one or more of the plurality of electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0033] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the plurality of electronic devices 120 are examples of server systems, personal computers, portable devices (e.g., smart phones, tablets, or laptop computers), wearable devices, video conferencing devices, and / or other types of electronic devices.
[0034] In an example operation of the communication system 100, the source device 102 sends the encoded video stream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video stream 108 and may decode and / or encode the encoded video stream 108 using the encoder component 114. For example, the server system 112 may apply the encoding that is optimal for network transmission and / or storage to the video data. The server system 112 may send the encoded video data 116 (e.g., one or more encoded video streams) to one or more of the plurality of electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.
[0035] In some embodiments, the transmission discussed above is unidirectional data transmission. Unidirectional data transmission is sometimes used in media service applications and the like. In some embodiments, the transmission discussed above is bidirectional data transmission. Bidirectional data transmission is sometimes used in video conferencing applications and the like. In some embodiments, the encoded video bitstream 108 and / or the encoded video data 116 are encoded and / or decoded according to any one of the video coding / compression standards described in the present disclosure, such as HEVC, VVC, and / or AV1.
[0036] Figure 2A FIG. is a block diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source of a component of a device different from the encoder component 106). The video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 YCrCb or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0 or YCrCb 4:4:4). In some embodiments, the video source 104 is a storage device storing previously acquired / prepared video. In some embodiments, the video source 104 is a camera that acquires local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0037] The encoder component 106 is configured to encode and / or compress the pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls and is functionally coupled to other functional units as described below. The parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can easily identify other functions of the controller 204, as they may be related to the encoder component 106 optimized for a specific system design.
[0038] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. As a simplified example, the encoding loop may include a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and (a) reference picture(s)) and a (local) decoder 210. The decoder 210 creates sample data in a manner similar to how a (remote) decoder creates sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact corresponding between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the same sample values as the reference picture samples that the decoder will interpret when using prediction during decoding. The principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is known to those of ordinary skill in the art.
[0039] The operation of the decoder 210 can be the same as that of a remote decoder (such as the decoder component 122) described in detail below, for example, in connection with Figure 2B However, briefly referring to Figure 2B , when the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.
[0040] It can be seen at this point that any decoder technology other than the parsing / entropy decoding present in the decoder must also necessarily exist in the corresponding encoder in substantially the same functional form. Therefore, the subject matter of this application focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is reciprocal to the fully described decoder technology. More detailed descriptions are only needed in certain areas and are provided below.
[0041] As part of its operation, the source encoder 202 may perform motion compensation predictive coding. The input frame is predictively encoded with reference to one or more previously encoded frames designated as reference image frames in the video sequence. In this way, the coding engine 212 encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference image frame, and the pixel blocks of the reference frame can be selected as the prediction reference for the input frame. The controller 204 may manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.
[0042] The decoder 210 decodes the encoded video data of a frame that can be specified as a reference image frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 can advantageously be a lossy process. When the encoded video data is decoded at the video decoder ( Figure 2A not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that can be performed by a remote video decoder on the reference image frame and can cause the reconstructed reference image frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference image frame that has the same content (in the absence of transmission errors) as the reconstructed reference image frame that will be obtained by the remote video decoder.
[0043] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor 206 can operate block by block based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 206, it can be determined that the input picture can have a prediction reference taken from multiple reference pictures stored in the reference picture memory 208.
[0044] The outputs of all the above functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding), thereby converting the symbols into an encoded video sequence.
[0045] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence created by the entropy encoder 214 in preparation for transmission over a communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such other data being, for example, encoded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may send additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / signal-to-noise ratio (SNR) enhancement layers, redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set segments, and the like.
[0046] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture may be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and characteristics, and thus will not be described in detail herein. A predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. A bi-predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0047] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the coding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively encoded, or the blocks of the I picture can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0048] The captured video can be a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (often simplified to intra-frame prediction) utilizes the spatial correlation in a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0049] The encoder component 106 can perform encoding operations according to a predetermined video coding technique or standard (such as any standard described herein). In operation, the encoder component 106 can perform various compression operations, including predictive coding operations that utilize the temporal and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0050] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to a loop filter 256, and the transmitter is configured to send data to the display 124 (e.g., via a wired or wireless connection).
[0051] In some embodiments, decoder component 122 includes a receiver coupled to channel 218, the receiver being configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not shown). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data when receiving the encoded video. The additional data may be part of the encoded video sequence. Decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0052] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, decoder component 122 is implemented at least partially in software.
[0053] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., configured to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to prevent network jitter). When receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may also be possible not to configure buffer memory 252, or buffer memory 252 may be made smaller. For use on a traffic packet network such as the Internet, buffer memory 252 may also be required, and the buffer memory may be relatively large and / or may have an advantageous adaptive size and may be implemented at least partially in the operating system or a similar element external to decoder component 122 (not shown).
[0054] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols may include information for managing the operation of the decoder components 122, and / or information for controlling a display device (e.g., the display 124). The control information for the display device may be, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not labeled). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technique or standard, and may follow various principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups may include a group of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0055] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols 270 may involve multiple different units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by the parser 254 from the encoded video sequence. For clarity, such subgroup control information flows between the parser 254 and the multiple units below are not described.
[0056] In addition to the functional blocks already mentioned, the decoder component 122 may be conceptually divided into multiple functional units as described below. In practical implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptually divided functional units below are retained.
[0057] The scaler / inverse transform unit 258 receives, from the parser 254, the quantized transform coefficients as symbols 270 and control information, including which transform mode to use, block size, quantization factor, and / or quantization scaling matrix, etc. The scaler / inverse transform unit 258 may output blocks including sample values, which may be input into the aggregator 268.
[0058] In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 may generate a block having the same size and shape as the block being reconstructed by using the surrounding reconstructed information extracted from the current picture (partially reconstructed) in the current picture memory 264. The aggregator 268 may add the predictive information that the intra-picture prediction unit 262 has generated to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.
[0059] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit 260 may access the reference picture memory 266 to extract samples for prediction. After the extracted samples are motion-compensated according to the symbol 270 related to the block, these samples may be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The address when the motion compensation prediction unit 260 extracts prediction samples from the reference picture memory 266 may be controlled by a motion vector. The motion vector is provided to the motion compensation prediction unit 260 in the form of the symbol 270, and the symbol 270 includes, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values extracted from the reference picture memory 266, a motion vector prediction mechanism, etc. when using sub-sample accurate motion vectors.
[0060] The output samples of the aggregator 268 may be adopted by various loop filtering techniques in the loop filter unit 256. Video compression techniques may include in-loop filter techniques, which are controlled by parameters included in the encoded video bitstream, and the parameters may be used for the loop filter unit 256 as the symbol 270 from the parser 254. However, video compression techniques may also respond to meta-information obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0061] The output of the loop filter unit 256 may be a sample stream, which may be output to a display device such as the display 124 and stored in the reference picture memory 266 for subsequent inter-picture prediction.
[0062] Once fully reconstructed, some of the encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture (by, for example, parser 254) is identified as a reference picture, the current reference picture can become part of the reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0063] The decoder component 122 can perform decoding operations according to a predetermined video compression technique recorded in a standard such as any of the standards described herein. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it complies with the syntax as specified in the video compression technique document or standard and specifically in the profile therein. Additionally, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within a range defined by the level of the video compression technique or standard. In some cases, the layer limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the layer can be further defined by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.
[0064] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).
[0065] The network interface 304 can be configured to interact with one or more communication networks (e.g., wireless, wired, and / or optical networks). The communication network can be local, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of communication networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, etc. Such communication can be unidirectional, receive-only (e.g., broadcast TV), send-only unidirectional (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., to other computer systems using local or wide area networks). Such communication can include communication to one or more cloud computing networks.
[0066] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 can include one or more of the following: keyboard, mouse, touchpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 can include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.
[0067] The memory 314 can include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices remote from the control circuit 302. The memory 314 or the non-volatile solid-state memory device within the memory 314 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures or subsets or supersets thereof: · An operating system 316, including procedures for handling various basic system services and for performing hardware-related tasks; · A network communication module 318 for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); · An encoding module 320 for performing various functions regarding encoding and / or decoding data (such as video data). In some embodiments, the encoding module 320 is an instance of the encoder component 114. The encoding module 320 includes, but is not limited to, one or more of the following: ο A decoding module 322, configured to perform various functions regarding decoding encoded data, such as those previously described with respect to decoder component 122; and ο An encoding module 340, configured to perform various functions regarding encoding data, such as those previously described with respect to encoder component 106; and · A picture memory 352, configured to store pictures and picture data, for example, for use with encoding module 320. In some embodiments, picture memory 352 includes one or more of the following: reference picture memory 208, buffer memory 252, current picture memory 264, and reference picture memory 266.
[0068] In some embodiments, decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to parser 254), a transform module 326 (e.g., configured to perform various functions previously described with respect to scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to motion compensation prediction unit 260 and / or intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to loop filter 256).
[0069] In some embodiments, encoding module 340 includes an encoding module 342 (e.g., configured to perform various functions previously described with respect to source encoder 202 and / or encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to predictor 206). In some embodiments, decoding module 322 and / or encoding module 340 includes Figure 3 a subset of the illustrated modules. For example, a shared prediction module is used by both decoding module 322 and encoding module 340.
[0070] Each of the above-identified modules stored in memory 314 corresponds to an instruction set for performing the functions described in this disclosure. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, processes, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, encoding module 320 optionally does not include separate decoding and encoding modules, but instead uses the same set of modules to perform both sets of functions. In some embodiments, memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, memory 314 stores additional modules and data structures not described above, such as an audio processing module.
[0071] In some embodiments, the server system 112 includes a web or Hypertext Transfer Protocol (HTTP) server, a File Transfer Protocol (FTP) server, and web pages and applications implemented using Common Gateway Interface (CGI) scripts, PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), Hypertext Markup Language (HTML), Extensible Markup Language (XML), Java, JavaScript, Asynchronous JavaScript and XML (AJAX), XHP, Javelin, Wireless Universal Resource File (WURFL), and the like.
[0072] Although Figure 3 a server system 112 according to some embodiments is shown, Figure 3 it is more intended as a functional description of the various features that may exist in one or more server systems rather than a structural schematic of the embodiments described in the present disclosure. In practice, and as will be recognized by those of ordinary skill in the art, the items shown separately may be combined, and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among them vary in different implementations and may optionally depend in part on the amount of data traffic processed by the server system during peak usage periods as well as during average usage periods.
[0073] Figure 4FIG. 400 is a flowchart of an example process 400 for applying in-loop filtering in video decoding according to some embodiments. A GOP includes a sequence of picture frames. The sequence of picture frames includes a current picture frame, which further includes a current coded block. The current picture frame includes a color picture, i.e., a non-monochrome picture frame, having a plurality of color samples (e.g., chrominance samples and luminance samples) co-located with each other. After reconstructing the plurality of color samples of the current picture frame, in-loop filtering is applied to adjust a subset of the color samples to improve the picture quality of the current picture frame. Reconstructed samples of a first color component and their neighboring reconstructed samples are combined to obtain an offset value for a second color component, and the reconstructed samples of the second color component are co-located with the reconstructed samples of the first color component and are adjusted by the offset value. The first color component is optionally the same as or different from the second color component. For example, a first luminance sample 404C and its neighboring luminance sample 404X are combined to obtain a sample offset 406, and a first chrominance sample 402C is co-located with the first luminance sample 404C and is adjusted by the sample offset 406. Alternatively, in another example, a first luminance sample 404C and its neighboring luminance sample 404X are combined to obtain a sample offset 406, and the sample offset 406 is applied to adjust the first luminance sample 404C itself. During the process of combining the luminance samples 404C and 404X, an in-loop filter 256 is applied to determine one or more of the following: the number, position, and weight of the neighboring luminance samples 404X, which are applied to generate the sample offset 406.
[0074] A decoder 122 receives a video bitstream 116 from an encoder 106, the video bitstream including a current coded block of a current picture frame. The video bitstream 116 includes (i) a cross-component sampling offset (CCSO) mode and (ii) a filter type index 408T and a filter shape index 408S of an in-loop filter 256. The CCSO mode indicates determining a first sample offset 406 for a first color sample 410 of the current coded block based on one or more luminance samples (e.g., a first luminance sample 404C, a neighboring luminance sample 404X). The filter type index 408T is configured to select one filter type from a plurality of filter types of the in-loop filter 256 for the current coded block. The filter shape index 408S is configured to identify a set of one or more neighboring positions used by the in-loop filter 256. The first color sample 410 is co-located with the first luminance sample 404C in the current coded block. In some embodiments, the filter type index 408T and the filter shape index 408S are written into and signaled in the video bitstream 116 at one of an image frame level and a key frame level of the current coded block.
[0075] Based on the filter type index 408T and the filter shape index 408S of the loop filter 256, the decoder 122 identifies one or more neighboring luma samples 404X of the first luma sample 404C. For example, one or more neighboring luma samples 404X include one or more of the following: a north neighboring luma sample (also referred to as an upper luma sample) 404N, a south neighboring luma sample (also referred to as a lower luma sample) 404S, a west neighboring luma sample (also referred to as a left luma sample) 404W, and an east neighboring luma sample (also referred to as a right luma sample) 404E.
[0076] The decoder determines one or more differences 404DX between one or more neighboring luma samples 404X and the first luma sample 404C. For example, one or more difference values 404DX include one or more of the following: a north difference 404DN, a south difference 404DS, a west difference 404DW, and an east difference 404DE. Each of the differences 404DN, 404DS, 404DW, and 404DE is respectively the difference between each of the luma samples 404W, 404S, 404W, and 404E and the first luma sample 404C. Quantize one or more differences 404DX to generate one or more quantized differences 404QX. For example, one or more quantized differences 404QX include one or more of the following: a quantized north difference 404QN, a quantized south difference 404QS, a quantized west difference 404QW, and a quantized east difference 404QE. Each of the differences 404DN, 404DS, 404DW, and 404DE is quantized to respectively generate the quantized differences 404QN, 404QS, 404QW, and 404QE.
[0077] For example, the classifier 412 classifies the first color sample 410 based on one or more quantized differences 404QX to determine the first sample offset 406 of the first color sample 410. For example, the quantized differences 404QX include quantized differences 404QN, 404QS, 404QW, and 404QE. The look-up table 414 maps multiple combinations of the quantized differences 404QN, 404QS, 404QW, and 404QE to different sample offset options SO (e.g., SO1 - SQ16). Based on the look-up table 414, the quantized difference 404QX determined for the first luminance sample 404C and the adjacent luminance sample 404X corresponds to one of the combinations in the look-up table 414, and the corresponding sample offset option SO is identified as corresponding to the combination of the quantized differences 404QX and is thus selected for the first sample offset 406. In other words, in some embodiments, the decoder 122 classifies the first color sample 410 by identifying one or more combinations of the quantized differences 404QX in the look-up table 414 and determining the first sample offset 406 corresponding to the one or more combinations of the quantized differences in the look-up table 414, in which multiple combinations of the quantized differences are associated with multiple offset value options SO (e.g., SO1 - SO16).
[0078] In addition, in some embodiments, the filter type index 408T selects one of the multiple filter types of the loop filter 256 for the current coding block. The multiple filter types of the loop filter include a first number of filter types, and each filter type has a different number of filter taps. The look-up table 414 includes a first number of sub-tables, each sub-table corresponding to a respective filter type. For example, the multiple filter types include three types (e.g., 2-tap, 3-tap, and 5-tap filters). The look-up table 414 includes 3 sub-tables corresponding to the 3 types of filters. The sub-tables corresponding to the 2-tap, 3-tap, or 5-tap filters have 2, 3, or 4 columns respectively. Additionally, in some embodiments, the filter shape index 408S selects one of the multiple filter shapes corresponding to the filter type index 408T of the loop filter 256. Each filter type 408S (e.g., 3-tap filter) has multiple corresponding filter shapes, each filter shape having a different combination of one or more adjacent positions. Each filter type corresponds to a respective sub-table that integrates the association of combinations of the quantized differences with multiple offset value options SO for the multiple corresponding filter shapes 408S.
[0079] Adjust the first color sample 410 based on the first sample offset 406 of the first color sample 410, so that the current coded block can be reconstructed. In some embodiments, the first color sample 410 includes a first chrominance sample 402C co-located with the first luma sample 404C in the current coded block, and the first chrominance sample 402C is adjusted based on the first sample offset 406. Alternatively, in some embodiments, the first color sample 410 is the first luma sample 404C, and the first luma sample 404C is adjusted based on the first sample offset 406.
[0080] In some embodiments, the filter type index 408T includes a type flag configured to select one of a two-tap filter ( Figure 5A and Figure 5B ) and a three-tap filter ( Figure 6A and Figure 6B ). Alternatively, in some embodiments, the filter type index 408T includes a type flag configured to select one of a five-tap filter ( Figure 7 ) and a three-tap filter ( Figure 6A and Figure 6B ). Alternatively, in some embodiments, the filter type index 408T includes a multi-bit (e.g., 2-bit) index configured to select at least one of a two-tap filter, a three-tap filter, and a five-tap filter.
[0081] In some embodiments, each of the plurality of filter types corresponds to a set of one or more candidate filter shape indices 408S. Each candidate filter shape index 408S is unique and different from any other filter type index 408S of the corresponding filter type and any other filter shape index 408S of any of the remaining filter types of the plurality of filter types. For example, the plurality of filter types includes three types (e.g., 2-tap, 3-tap, and 5-tap filters). The 2-tap filter corresponds to filter shape indices 408S in a first range (e.g., 0-11). The 3-tap filter corresponds to filter shape indices 408S in a second range (e.g., 12-17), and the 5-tap filter corresponds to filter shape indices 408S in a third range (e.g., 18-22). The first range, the second range, and the third range are mutually exclusive (i.e., non-overlapping).
[0082] Figure 5A and Figure 5BFIG. is a diagram showing positions 500 and 550 of adjacent luminance samples 404X of an exemplary two-tap loop filter 256 used in the CCSO mode according to some embodiments. The decoder 122 receives a video bitstream 116 from the encoder 106, and the video bitstream includes a current encoded block of a current image frame. The video bitstream 116 includes a filter type index 408T and a filter shape index 408S of the loop filter 256. The filter type index 408T is configured to select one of a plurality of filter types (e.g., two-tap, three-tap, and five-tap) of the loop filter 256 for the current encoded block. The filter shape index 408S is configured to identify a set of one or more adjacent positions used by the loop filter 256. Based on the filter type index 408T and the filter shape index 408S of the loop filter 256, the decoder 122 identifies one or more adjacent luminance samples 404X of the first luminance sample 404C. The decoder 122 determines one or more differences 404DX between the one or more adjacent luminance samples 404X and the first luminance sample 404C. The one or more differences 404DX are quantized to generate one or more quantized differences 404QX. For example, the classifier 412 classifies the first color sample 410 based on the one or more quantized differences 404QX to determine a first sample offset 406 of the first color sample 410. The first color sample 410 is adjusted based on the first sample offset 406 of the first color sample 410, so that the current encoded block can be reconstructed.
[0083] In some embodiments, the filter type index 408T selects a two-tap loop filter 256 having two taps. The two-tap loop filter 256 includes a first luminance sample 404C and a single adjacent luminance sample 404X, and the filter shape index 408S identifies a single adjacent position of the adjacent luminance sample 404X used by the loop filter 256. Refer to Figure 5A , in some embodiments, the adjacent luminance sample 404X is adjacent to the first luminance sample 404C. The filter shape index 408S is one of eight different values (e.g., consecutive integer values in the range of 0-7). Specifically, in some cases, the decoder 122 identifies one or more adjacent luminance samples 404X of the first luminance sample 404C by selecting a two-tap filter according to the determined filter type index 408T and by selecting one adjacent luminance sample from a set of eight adjacent luminance samples 404NW, 404N, 404NE, 404E, 404SE, 404S, 404SW, and 404W as the one or more adjacent luminance samples 404X based on the filter shape index 408S (e.g., a value in 0-7).
[0084] Refer to Figure 5B, in some embodiments, a single adjacent luminance sample 404X is adjacent to the first luminance sample 404C, or is separated from the first luminance sample 404C by at least one row or one column. The filter shape index 408S is one of twelve different values (e.g., twelve consecutive integer values in the range of 0 - 11). A filter shape index 408S having a value in the first value subset corresponds to a luminance position adjacent to the first luminance sample 404C, and a filter shape index 408S having a value in the second value subset corresponds to a luminance position separated from the first luminance sample 404C. For example, a filter shape index 408S having a value of 2, 6, 0, or 4 corresponds to adjacent luminance samples 404N, 404S, 404W, or 404E, respectively. Conversely, a filter shape index 408S having a value higher than 7 (e.g., 8, 9, 10, 11) corresponds to a luminance position in the same row as the position of the first luminance sample 404C but separated from the position of the first luminance sample 404C by one or more columns. Specifically, in some cases, the decoder 122 identifies one or more adjacent luminance samples 404X of the first luminance sample 404C by selecting a two - tap filter according to the determined filter type index 408T and selecting one adjacent luminance sample from a set of twelve adjacent luminance samples as one or more adjacent luminance samples 404X based on the filter shape index. The set of twelve adjacent luminance samples includes eight adjacent luminance samples 404NW, 404N, 404NE, 404E, 404SE, 404S, 404SW, and 404W, two luminance samples 404W1 and 404E1 in the same row and separated from the first luminance sample 404C by two luminance samples, and two luminance samples 404W2 and 404E2 in the same row and separated from the first luminance sample 404C by four luminance samples.
[0085] Figure 6A and Figure 6BFIG. is a diagram showing positions 600 and 650 of adjacent luminance samples 404X of an example three-tap loop filter 256 used in the CCSO mode according to some embodiments. In some embodiments, the filter type index 408T selects a three-tap loop filter 256 having three taps, the three-tap loop filter 256 includes a first luminance sample 404C and two adjacent luminance samples 404X, and the filter shape index 408S identifies a pair of adjacent positions (i.e., two adjacent positions) of the adjacent luminance samples 404X used by the loop filter 256. In some embodiments, the adjacent positions of the adjacent luminance samples 404X are indexed in pairs (indexed), for example, using the same filter shape index 408S to index two adjacent positions. Additionally, in some embodiments, the positions of two adjacent luminance samples 404X jointly indexed using the same filter shape index 408S are symmetric with respect to the first luminance sample 404C. Specifically, in some cases, the decoder 122 identifies one or more adjacent luminance samples 404X of the first luminance sample 404C by determining to select a three-tap filter according to the filter type index 408T and by selecting one combination from a plurality of different combinations of two adjacent positions (e.g., as one or more adjacent luminance samples 404X) based on the filter shape index 408S.
[0086] Reference Figure 6A , in some embodiments, two adjacent luminance samples 404X selected according to the filter shape index 408S are adjacent to the first luminance sample 404C. The filter shape index 408S is one of four different values (e.g., consecutive integer values in the range of 1-4). Specifically, in some cases, the decoder 122 identifies a plurality of different combinations of two adjacent positions, and the plurality of different combinations of two adjacent positions at least includes (1) a first combination of a north luminance sample 404N and a south luminance sample 404S, (2) a second combination of a west luminance sample 404W and an east luminance sample 404E, (3) a third combination of a northwest luminance sample 404NW and a southeast luminance sample 404SE, and (4) a fourth combination of a northeast luminance sample 404NE and a southwest luminance sample 404SW.
[0087] Reference Figure 6B, in some embodiments, two adjacent luma samples 404X identified by the same filter shape index 408S are adjacent to the first luma sample 404C or are separated from the first luma sample 404C by at least one row or one column. For example, the filter shape index 408S is one of six different values (e.g., six consecutive integer values in the range of 1 - 6). A filter shape index 408S having a value from a first subset of values (e.g., 1, 2, 3, 4) corresponds to a luma position adjacent to the first luma sample 404C, and a filter shape index 408S having a value from a second subset of values (e.g., 5, 6) corresponds to a luma position separated from the first luma sample 404C. For example, filter shape indices 408S having values 1, 2, 3, or 4 correspond respectively to adjacent luma sample pairs 404N and 404S, 404NW and 404SE, 404W and 404E, or 404NE and 404SW. Conversely, filter shape indices 408S having values 5 or 6 correspond respectively to adjacent luma sample pairs 404W1 and 404E1, or 404W2 and 404E2. Specifically, in some cases, the decoder 122 identifies multiple different combinations of two adjacent positions, which multiple different combinations at least include a first combination, a second combination, a third combination, and a fourth combination, a fifth combination of two luma samples 404W1 and 404E1 located in the same row and separated from the first luma sample by two luma samples, and a sixth combination of two luma samples 404W2 and 404E2 located in the same row and separated from the first luma sample by four luma samples.
[0088] In addition, in some embodiments, the filter shape index 408S selects one of multiple filter shapes corresponding to the filter type index 408T of the loop filter 256. The selected filter type 408S (e.g., a 3 - tap filter) has multiple corresponding filter shapes (e.g., Figure 5B the six shapes indexed by 1 - 6 in), and each filter shape has a different combination of one or more adjacent positions. In some embodiments, each filter shape has a corresponding look - up sub - table. For a 3 - tap filter, the 6 filter shapes ( Figure 5B ) correspond to 6 look - up sub - tables, and each sub - table associates a combination of quantized differences of the corresponding luma sample 404X with an offset value option SO. Alternatively, in some embodiments, each filter type corresponds to a single sub - table that integrates the association of combinations of quantized differences with multiple offset value options SO for multiple corresponding filter shapes 408S. In other words, for a 3 - tap filter, the 6 look - up sub - tables of the 6 filter shapes ( Figure 5B ) are integrated into a single sub - table of the corresponding filter type (e.g., a 3 - tap filter).
[0089] Figure 7 FIG. is a diagram showing the positions 700 of adjacent luminance samples 404X of an example five-tap loop filter 256 used in the CCSO mode according to some embodiments. In some embodiments, the filter type index 408T selects a five-tap loop filter 256 having five taps, the five-tap loop filter 256 including a first luminance sample 404C and four adjacent luminance samples 404X, and the filter shape index 408S identifies two pairs of adjacent positions (i.e., four adjacent positions) of the adjacent luminance samples 404X used by the loop filter 256. In some embodiments, the adjacent positions of the adjacent luminance samples 404X are indexed in pairs. For example, the filter shape index 408S having two index numbers is used to index the four adjacent positions. The two index numbers of the filter shape index 408S respectively select two pairs of adjacent luminance samples 404X. Additionally, in some embodiments, the positions of the two pairs of adjacent luminance samples 404X jointly indexed using the filter shape index 408S are symmetric with respect to the first luminance sample 404C. Specifically, in some cases, the decoder 122 identifies one or more adjacent luminance samples 404X of the first luminance sample 404C by determining to select a five-tap filter according to the filter type index 408T and by simultaneously selecting two combinations from multiple different combinations of two adjacent positions as one or more adjacent luminance samples 404X based on the filter shape index 408S.
[0090] In some embodiments, eight immediately adjacent luminance samples are divided into four pairs of luminance samples indexed by four different numbers. The four adjacent luminance samples 404X (i.e., two pairs of samples 404X) selected according to the filter shape index 408S are immediately adjacent to the first luminance sample 404C and are identified by two of the four different numbers. The filter shape index 408S includes two of the four different values, for example, two values from 1-4. Specifically, in some cases, the decoder 122 identifies multiple different combinations of two adjacent positions, the multiple different combinations of two adjacent positions at least including (1) a first combination of a north luminance sample 404N and a south luminance sample 404S, for example, corresponding to the number "1"; (2) a second combination of a west luminance sample 404W and an east luminance sample 404E, for example, corresponding to the number "2"; (3) a third combination of a northwest luminance sample 404NW and a southeast luminance sample 404SE, for example, corresponding to the number "3"; and (4) a fourth combination of a northeast luminance sample 404NE and a southwest luminance sample 404SW, for example, corresponding to the number "4". In the example, the filter shape index 408S includes "1" and "3", and adjacent luminance samples 404N, 404S, 404W, and 404S are selected to generate a first sample offset 406.
[0091] Figure 8AFIG. 800 is a flow chart of an example process 800 of applying band-based in-loop filtering according to some embodiments. As described above, decoder 122 receives video bitstream 116, in which filter type index 408T and filter shape index 408S of loop filter 256 are written. Filter type index 408T is configured to select one of multiple filter types (e.g., 2-tap, 3-tap, and 5-tap) of loop filter 256 for a current coding block. Filter shape index 408S is configured to identify a set of one or more neighboring positions used by loop filter 256. Based on filter type index 408T, decoder 122 selects one of multiple filter types (e.g., 2-tap, 3-tap, and 5-tap) of loop filter 256. The loop filter 256 of the selected filter type has a first number N of taps. First color sample 410 corresponds to multiple bands 802. Decoder 122 determines the number 804 of bands in the multiple bands based on the first number N. Each band corresponds to a respective value of filter shape index 408S', which identifies a respective set of one or more neighboring positions used by loop filter 256.
[0092] Decoder 122 identifies one of the multiple bands that includes first color sample 410 based on the value of first color sample 410. Based on the identified one of the multiple bands, decoder 122 identifies the respective value of filter shape index 408S' and the respective set of one or more neighboring positions, so as to select one or more neighboring luma samples 404X for generating first sample offset 406. In other words, filter shape index 408S' is determined by decoder 122 and applied in place of filter shape index 408S received in the video bitstream.
[0093] In the example, first color sample 410 is first chroma sample 402C. Filter type index 408T identifies 2-tap loop filter 256. The first number N is equal to 2, so the number of bands 804 is also 2. The chroma value is in the range of [0-1023], and corresponds to two bands of [0,511) and [512,1023], which further correspond to the first value and the second value of filter shape index 408S respectively. Refer to Figure 5A , the first value and the second value are equal to 8 and 4, and are used for the quantized difference 404QX ( Figure 4)The adjacent luminance samples 404X classified and generating the first sample offset 406 are luminance samples 404W and 404E of two different frequency bands respectively. In some cases, the value of the first chrominance sample 402C is 490, and the luminance sample 404W is applied to classify the quantized difference 404QW and generate the first sample offset 406. Alternatively, in some cases, the value of the first chrominance sample 402C is 890, and the luminance sample 404E is applied to classify the quantized difference 404QE to generate the first sample offset 406.
[0094] Figure 8B It is a flowchart of an example process 850 for implementing classification based on a dynamic filter type during in-loop filtering according to some embodiments. The decoder 122 receives the video bitstream 116, in which the filter type index 408T and the filter shape index 408S of the in-loop filter 256 are written. Based on the filter type index 408T and the filter shape index 408S of the in-loop filter 256, the decoder 122 identifies one or more adjacent luminance samples 404X of the first luminance sample 404C. The decoder 122 determines one or more differences 404DX between the one or more adjacent luminance samples 404X and the first luminance sample 404C. The one or more differences 404DX are quantized to generate one or more quantized differences 404QX. For example, the classifier 412 classifies the first color sample 410 based on the one or more quantized differences 404QX to determine the first sample offset 406 of the first color sample 410. The first color sample 410 is adjusted based on the first sample offset 406 of the first color sample 410, so that the current encoded block can be reconstructed.
[0095] In some embodiments, a scalar quantizer 422 including a plurality of quantization intervals (QI) 418 and a plurality of quantization levels (QL) 420 is used ( Figure 4 ) to quantize one or more differences 404DX into a plurality of integer values within the quantization range 416 ( Figure 4 ), where each of the one or more quantized differences 404DX includes a corresponding integer within the quantization range 416. For each integer value in the quantization range 416, the quantization interval 418 is defined as the range of differences 404DX assigned to the corresponding integer value. The quantization level 420 corresponds to the corresponding integer value to which the range of differences associated with the quantization interval 418 is assigned. Further, in some embodiments, one of the plurality of filter types of the in-loop filter selected by the filter type index 408T has a first number N of taps. The plurality of quantization intervals includes a second number M of quantization intervals. The first color sample 410 is classified by a classifier 412 ( Figure 4 ) having a third number L of classes, where L is equal to M N. For example, when applying a two-tap filter ( Figure 5A and Figure 5B ), the number of categories (e.g., the number of entries in the offset lookup table 414) is determined to be M 2 .
[0096] Additionally, in some embodiments, one of the multiple filter types of the loop filter 256 selected by the filter type index 408T has a first number N of taps, and the multiple quantization intervals 418 include a second number M of quantization intervals. Based on the first number N, the decoder 122 determines the second number M and the multiple quantization levels 420.
[0097] In some embodiments, one of the multiple filter types of the loop filter 256 selected by the filter type index 408T has a first number N of taps. The first color sample 410 is classified by a classifier 412 ( Figure 4 ) having a third number L of categories. The decoder 122 determines the third number L based on the first number N.
[0098] Figure 9 is a flowchart showing an example method 900 for encoding video according to some embodiments. The method 900 may be executed in a computing system (e.g., Figure 1 the server system 112, the source device 102, or the electronic device 120 in), the computing system having a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, the method 900 is applied in conjunction with one or more video codecs, one or more video codecs including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In some embodiments, the method 900 is executed by executing instructions stored in the memory of the computing system (e.g., the encoding module 320 of the memory 314). In some embodiments, the computing system transmits (902) the video bitstream 116 from the encoder 106 to the decoder 122, where the video bitstream 116 includes the current encoded block of the current image frame. The video bitstream 116 includes (904) (i) the cross-component sampling offset (CCSO) mode and (ii) the filter type index 408T and the filter shape index 408S of the loop filter 256.
[0099] The computing system identifies (906) a sample of a first color component (e.g., a first luminance sample 404C) of the current coding block and a first color sample 410 of a second color component, where the first color sample 410 of the second color component is co-located with the sample of the first color component in the current coding block, and the computing system identifies (908) one or more neighboring samples (e.g., one or more neighboring luminance samples 404X of the first luminance sample 404C) of the sample of the first color component based on the filter type index 408T and the filter shape index 408S of the loop filter 256. Determine (910) a difference 404DX between a co-located reconstructed luminance sample from the first color component and its neighboring reconstructed samples (e.g., between the first luminance sample 404C and the neighboring luminance samples 404X). Select the positions of the neighboring reconstructed samples (e.g., the neighboring luminance samples 404X) based on the given filter shape 408S.
[0100] Quantize (912) the difference 404DX using a scalar quantizer 422 ( Figure 4 ), where the scalar quantizer 422 is specified by a quantization interval (QI) 418 and a quantization level (QL) 420. The quantization interval (QI) 418 is defined as the range of values assigned to the same integer, and the quantization level (QL) 420 is defined as an integer value that assigns all the values within the corresponding quantization interval (QI) 418. Apply the quantized difference 404QX to classify the first color sample 410, thereby determining a first sample offset 406 of the first color sample 410. For example, a combination of the quantized difference 404QX is associated with an index in the selected look-up table 414 ( Figure 4 ), and the output of the selected look-up table 414 includes the value of the first sample offset 406. This cross-component sample offset filtering method is based on an edge-preserving loop filter 256, which uses reconstructed samples (e.g., 404C, 404X, 402C, 410) to calculate the sample offset 406 of the luminance component and / or the chrominance component. In some embodiments, samples around the current color sample ( Figure 6B "C" in) are used in a three-tap filter to classify the current color sample, and the current color sample is used in a band or edge offset. In some embodiments, an adaptive multi-tap filter classifier is applied.
[0101] In some embodiments, a multi-tap filter classifier 412 ( Figure 4 ) is used to classify the current central sample (e.g., the first luminance sample 404C) during the cross-component sample offset process. The predefined filter shape (i.e., the surrounding samples used) is optionally the same as or different from the design in Figure 6B .
[0102] In some embodiments, a two-tap filter 256 is used in the classification (Figure 5A and Figure 5B ). At the frame level (or any other level where sampling offset control information is written), if a newly introduced two-tap filter or the original three-tap filter is used, a write flag is set. If the flag has a value (e.g., true), the direction of the filter is further written to indicate which sample position is used. For example, if the flag indicates the use of a two-tap filter, the filter shape index is indicated as "1" (i.e., the sample marked as 1 in Figure 5B ). In another example, additional filter shape and direction are written for the two-tap filter. If the written filter shape index has a first value (e.g., "0"), the sample 404N above the current central sample (e.g., the first luminance sample 404C) is used in the two-tap filter. If the written filter shape index has a second value (e.g., "1"), the sample 404S below the current central sample (e.g., the first luminance sample 404C) is used in the two-tap filter.
[0103] Alternatively, in some embodiments, a two-tap filter is used in classification, and the number of filter shapes is explicitly written. An example of eight two-tap filters is shown in Figure 5A , and the index value of the filter shape index 408S ranging from 0 to 7 is written to indicate which filter tap is used. Figure 5B Another example of twelve two-tap filters 256 is shown in
[0104] In some embodiments, a five-tap filter ( Figure 7 ) is used for classification. At the frame level (or any other level where sampling offset control information is written), a flag is written to select one of the five-tap filter or the three-tap filter. If the five-tap filter is selected, the semantics of the filter shape index 408S are modified. For example, in the case of the three-tap filter, the filter shape index 408S (e.g., equal to one of 0 - 11) is used to indicate Figure 6B which two samples in Figure 7 are used for classification. In contrast, in the case of the proposed five-tap filter, the filter shape index 408S is used to indicate which set of samples in
[0105] In some embodiments, the encoder 106 adaptively selects multiple multi-tap filters and notifies the decoder 122 as a syntax element. For example, a two-tap filter ( Figure 5A and Figure 5B ) and a three-tap filter (Figure 6A and Figure 6B ) and a five-tap filter ( Figure 7 ). In an example, a syntax element having two bits is written at the frame level or any other level where CCSO control information is written. In another example, the syntax element is written at the key frame level.
[0106] In some embodiments, the described multi-tap filter may be combined with any other predefined filter size or quantizer involved in the cross-component sample offset process. That is, the applicability of the proposed method is not limited by changes in other parts of the cross-component sample offset process.
[0107] In some embodiments, the number of filter shapes (e.g., Figure 5A 8 in Figure 5B ), Figure 6A 11 in Figure 6B ) varies depending on the number of filter taps. For example, the number of 2-tap filter shapes is 1, 2, 3, 4, …, or 16, while the number of 3-tap filter shapes is 4 (
[0108] In some embodiments, the number of classes L given by the classifier 412 ( Figure 8B ) depends on the number of filter taps (N). For example, in some embodiments, the number of filter taps is represented by N, and the number of quantization intervals (QI) 418 of the difference between adjacent samples located at the filter taps is represented by M. The number of classes L ( Figure 8B ) given by the classifier 412 is determined by M N . In one example, when a two-tap filter is applied, the number of classes L (the number of entries in the offset look-up table 414) is determined to be M 2 .
[0109] In some embodiments, the number of applicable frequency bands 804 ( Figure 8A ) given by the classifier depends on the number of filter taps (N).
[0110] In some embodiments, the number of applicable quantization intervals (QI) 418 and / or candidate quantization step sizes for obtaining the classifier 412 depends on the number of filter taps (N).
[0111] Although Figure 9A number of logical stages are shown in a particular order, but stages that are not order-dependent can be reordered, and other stages can be combined or broken apart. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, so the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.
[0112] Some example embodiments are described below.
[0113] (A1) In some embodiments, an implementation method 900 is used to decode video data. Method 900 includes receiving (902) a video bitstream that includes a current encoded block of a current image frame, where the video bitstream includes (904) (i) a first syntax element for a cross-component sample offset (CCSO) mode that indicates whether a first sample offset of a first color sample of the current encoded block is determined based on one or more luma samples, and (ii) a second syntax element for a filter type index and a filter shape index, the first type index indicating one of a plurality of filter types of a loop filter selected for the current encoded block, and the filter shape index identifying a set of one or more neighboring positions used by the loop filter; determining (910) one or more differences between one or more neighboring luma samples and a first luma sample; quantizing (912) the one or more differences to generate one or more quantized differences; classifying (914) the first color sample based on the one or more quantized differences to determine the first sample offset of the first color sample; and reconstructing (916) the current encoded block at least by adjusting the first color sample based on the first sample offset of the first color sample.
[0114] (A2) In some embodiments of A1, wherein the first color sample includes a first chroma sample that is co-located with the first luma sample in the current encoded block, and the first chroma sample is adjusted based on the first sample offset.
[0115] (A3) In some embodiments of A1, wherein the first color sample is the first luma sample, and the first luma sample is adjusted based on the first sample offset.
[0116] (A4) In some embodiments of A1 - A3, wherein the filter type index includes a type flag that is configured to select one of a two-tap filter and a three-tap filter.
[0117] (A5) In some embodiments of A1 - A3, wherein the filter type index includes a type flag that is configured to select one of a four-tap filter and a three-tap filter.
[0118] (A6)In some embodiments of A1 - A3, wherein the filter type index includes a multi - bit index, the multi - bit index is configured to select at least one of a two - tap filter, a three - tap filter, and a four - tap filter.
[0119] (A7)In some embodiments of A1 - A4, method 900 further includes: selecting a two - tap filter according to the determined filter type index, and selecting one of the adjacent luma samples in a set of eight adjacent luma samples as one or more adjacent luma samples based on the filter shape index.
[0120] (A8)In some embodiments of A1 - A4, the method further includes: selecting a two - tap filter according to the determined filter type index, and selecting one of the adjacent luma samples in a set of twelve adjacent luma samples as one or more adjacent luma samples based on the filter shape index, wherein the set of twelve adjacent luma samples includes eight adjacent luma samples, two luma samples in the same row and separated from the first luma sample by two luma samples, and two luma samples in the same row and separated from the first luma sample by four luma samples.
[0121] (A9)In some embodiments of A1 - A6, the method further includes: selecting a three - tap filter according to the determined filter type index, and selecting one of multiple different combinations of two adjacent positions as one or more adjacent luma samples based on the filter shape index.
[0122] (A10)In some embodiments of A9, the multiple different combinations of two adjacent positions at least include (1) a first combination of a north luma sample and a south luma sample, (2) a second combination of a west luma sample and an east luma sample, (3) a third combination of a northwest luma sample and a southeast luma sample, and (4) a fourth combination of a northeast luma sample and a southwest luma sample.
[0123] (A11)In some embodiments of A10, the multiple different combinations of two adjacent positions further include a fifth combination of two luma samples in the same row and separated from the first luma sample by two luma samples and a sixth combination of two luma samples in the same row and separated from the first luma sample by four luma samples.
[0124] (A12)In some embodiments of A1 - A4, A5, and A6, the method further includes: selecting a four - tap filter according to the determined filter type index, and simultaneously selecting two of the multiple different combinations of two adjacent positions as one or more adjacent luma samples based on the filter shape index.
[0125] (A13)In some embodiments of A1 - A12, wherein the filter type index and the filter shape index are written into the video bitstream at one of the image frame level and the key frame level of the current coded block.
[0126] (A14)In some embodiments of A1 - A13, wherein each of the plurality of filter types corresponds to a set of one or more candidate filter shape indices, each candidate filter shape index being unique and different from any other filter shape index of the corresponding filter type and any other filter shape index of any of the remaining filter types of the plurality of filter types.
[0127] (A15)According to some embodiments of A1 - A14, wherein the one or more differences are quantized to a plurality of integer values within a quantization range using a scalar quantizer including a plurality of quantization intervals and a plurality of quantization levels, and each of the one or more quantized differences includes a corresponding integer within the quantization range.
[0128] (A16)In some embodiments of A15, wherein one of the plurality of filter types of the loop filter selected by the filter type index has a first number N of taps; the plurality of quantization intervals includes a second number M of quantization intervals; and the first color samples are classified by a classifier having a third number L of classes, where L is equal to MN.
[0129] (A17)According to some embodiments of A15, wherein one of the plurality of filter types of the loop filter selected by the filter type index has a first number N of taps, and the plurality of quantization intervals includes a second number M of quantization intervals, the method further comprising: determining the second number M and the plurality of quantization levels based on the first number N.
[0130] (A18)In some embodiments of A15, wherein one of the plurality of filter types of the loop filter selected by the filter type index has a first number N of taps, and the first color samples are classified by a classifier having a third number L of classes, the method further comprising: determining the third number L based on the first number N.
[0131] (A19)In some embodiments of A1 - A18, wherein one of the plurality of filter types of the loop filter selected by the filter type index has a first number N of taps, and the first color sample corresponds to a plurality of frequency bands, the method further includes: determining, based on the first number N, the number of frequency bands in the plurality of frequency bands, each frequency band corresponding to a respective value of the filter shape index that identifies a respective set of one or more adjacent positions used by the loop filter; identifying, based on the value of the first color sample, one of the plurality of frequency bands that includes the first color sample; and identifying, based on the identified one of the plurality of frequency bands, the respective value of the filter shape index and the respective set of the one or more adjacent positions, thereby selecting the one or more adjacent luminance samples for generating the first sample offset.
[0132] (A20)In some embodiments of A1 - A19, classifying the first color sample based on the one or more quantized differences further includes: identifying, in a look - up table, the combination of the one or more quantized differences, the look - up table associating combinations of a plurality of quantized differences with a plurality of offset value options; and determining a first sample offset corresponding to the combination of the one or more quantized differences in the look - up table.
[0133] (A21)In some embodiments of A20, wherein the plurality of filter types of the loop filter includes a first number of filter types, each filter type having a different number of filter taps, and the look - up table includes a first number of sub - tables, each sub - table corresponding to a respective filter type.
[0134] (A22)In some embodiments of A21, wherein the filter shape index selects one of a plurality of filter shapes corresponding to the filter type index of the loop filter; each filter type has a plurality of corresponding filter shapes, each filter shape having a different combination of a plurality of adjacent positions; and each filter type corresponds to a respective sub - table that integrates the association of combinations of the quantized differences with a plurality of offset value options for the plurality of corresponding filter shapes.
[0135] (A23)In some embodiments of A1 - A22, wherein method 900 further includes: identifying (906) a first luminance sample of the current coding block and a first color sample co - located with the first luminance sample in the current coding block; and identifying (908) one or more adjacent luminance samples of the first luminance sample based on the filter type index and the filter shape index of the loop filter.
[0136] In another aspect, some embodiments include a computing system (e.g., server system 112). The computing system includes control circuitry (e.g., control circuitry 302) and a memory coupled to the control circuitry (e.g., memory 314), the memory storing one or more instruction sets configured to be executed by the control circuitry, the one or more instruction sets including instructions for performing any of the methods described in this disclosure (e.g., A1 - A20 above).
[0137] In yet another aspect, some embodiments include a non - transitory computer - readable storage medium storing one or more instruction sets for execution by control circuitry of a computing system, the one or more instruction sets including instructions for performing any of the methods described in this disclosure (e.g., A1 - A22 above).
[0138] The proposed methods can be used alone or in any order of combination. Additionally, each of the methods (or embodiments), encoders, and decoders can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). For example, one or more processors execute a program stored in a non - transitory computer - readable medium. Hereinafter, the term block can be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU.
[0139] It should be understood that although terms such as "first", "second", etc. may be used in this disclosure to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.
[0140] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used in this disclosure refers to and encompasses any and all possible combinations of one or more of the associated listed items. Further, it can be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their combinations.
[0141] As used in this disclosure, depending on context, the term "if" can be interpreted to mean "when" or "in" or "in response to determining" or "in accordance with determining" or "in response to detecting" that the stated precondition is true. Similarly, depending on context, the phrases "if it is determined (that the stated precondition is true)" or "if (the stated precondition is true)" or "when (the stated precondition is true)" can be interpreted to mean "upon determining" or "in response to determining" or "in accordance with determining" or "upon detecting" or "in response to detecting", that the stated precondition is true.
[0142] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and the practical application, thereby enabling others skilled in the art to utilize the disclosure.
Claims
1. A method for decoding video data, comprising: Receive a video bitstream, where the video bitstream includes a current coding block of a current picture frame. Among them, the video bitstream includes (i) a first syntax element for a cross-component sample offset (CCSO) mode, and the CCSO mode indicates whether to determine a first sample offset of a first color sample of the current coding block based on one or more luma samples, and (ii) a second syntax element for a filter type index and a filter shape index, the filter type index indicates one of a plurality of filter types of a loop filter selected for the current coding block, and the filter shape index identifies a set of one or more neighboring positions used by the loop filter; Determine one or more differences between one or more neighboring luma samples and a first luma sample; Quantize the one or more differences to generate one or more quantized differences; Classify the first color sample based on the one or more quantized differences to determine a first sample offset of the first color sample; and Reconstruct the current coding block at least by adjusting the first color sample based on the first sample offset of the first color sample.
2. The method according to claim 1, further comprising: Identify the first luma sample of the current coding block and the first color sample co-located with the first luma sample in the current coding block; And Identify the one or more neighboring luma samples of the first luma sample based on the filter type index and the filter shape index of the loop filter.
3. The method according to claim 1, wherein, The first color sample includes a first chroma sample co-located with the first luma sample in the current coding block, and the first chroma sample is adjusted based on the first sample offset.
4. The method according to claim 1, wherein, The first color sample is the first luma sample, and the first luma sample is adjusted based on the first sample offset.
5. The method according to claim 1, wherein, The filter type index includes a type flag configured to select one of a two-tap filter and a three-tap filter.
6. The method according to claim 1, wherein, The filter type index includes a type flag configured to select one of a four-tap filter and a three-tap filter.
7. The method according to claim 1, wherein, The filter type index includes a multi-bit index configured to select at least one of a two-tap filter, a three-tap filter, and a four-tap filter.
8. The method according to claim 1, further comprising: Select a two-tap filter according to the determined filter type index, and select one of the neighboring luma samples in a set of eight neighboring luma samples as the one or more neighboring luma samples based on the filter shape index.
9. The method according to claim 1, further comprising: Select a two-tap filter according to the determined filter type index, and select one of the neighboring luma samples in a set of twelve neighboring luma samples as the one or more neighboring luma samples based on the filter shape index. Among them, the set of twelve neighboring luma samples includes eight neighboring luma samples, two luma samples in the same row and spaced two luma samples from the first luma sample, and two luma samples in the same row and spaced four luma samples from the first luma sample.
10. The method according to claim 1, further comprising: Select a three - tap filter according to the determined filter type index, and select one combination from multiple different combinations of two adjacent positions as the one or more adjacent luminance samples based on the filter shape index.
11. The method according to claim 10, wherein the plurality of different combinations of the two adjacent positions at least includes (1) a first combination of a north luminance sample and a south luminance sample, (2) a second combination of a west luminance sample and an east luminance sample, (3) a third combination of a northwest luminance sample and a southeast luminance sample, and (4) a fourth combination of a northeast luminance sample and a southwest luminance sample.
12. The method according to claim 11, wherein the plurality of different combinations of the two adjacent positions further includes a fifth combination of two luminance samples located in the same row and spaced two luminance samples from the first luminance sample, and a sixth combination of two luminance samples located in the same row and spaced four luminance samples from the first luminance sample.
13. The method according to claim 1, further comprising: Select a four - tap filter according to the determined filter type index, and simultaneously select two combinations from multiple different combinations of two adjacent positions as the one or more adjacent luminance samples based on the filter shape index.
14. A computing system, comprising: Control circuit; and A memory that stores one or more programs configured to be executed by the control circuit, and the one or more programs further include instructions for the following operations: Receive a video bitstream that includes a current encoded block of a current image frame, where the video bitstream includes (i) a first syntax element for a cross - component sample offset (CCSO) mode that indicates whether to determine a first sample offset of a first color sample of the current encoded block based on one or more luminance samples, and (ii) a second syntax element for a filter type index and a filter shape index, the filter type index indicating one filter type among multiple filter types of a loop filter selected for the current encoded block, and the filter shape index identifying a set of one or more adjacent positions used by the loop filter; Determine one or more differences between one or more adjacent luminance samples and a first luminance sample; Quantize the one or more differences to generate one or more quantized differences; Classify the first color sample based on the one or more quantized differences to determine a first sample offset of the first color sample; and Reconstruct the current encoded block at least by adjusting the first color sample based on the first sample offset of the first color sample.
15. The computing system according to claim 14, wherein, The filter type index and the filter shape index are written into the video bitstream at one of the image frame level and the key frame level of the current encoded block.
16. The computing system according to claim 14, wherein, Each of the multiple filter types corresponds to a set of one or more candidate filter shape indexes, and each candidate filter shape index is unique and different from any other filter shape index of the corresponding filter type and any other filter shape index of any remaining filter types among the multiple filter types.
17. The computing system according to claim 14, wherein, Quantize the one or more differences into multiple integer values within a quantization range using a scalar quantizer including multiple quantization intervals and multiple quantization levels, and each of the one or more quantized differences includes a corresponding integer within the quantization range.
18. The computing system according to claim 17, wherein, One of the multiple filter types of the loop filter selected by the filter type index has a first number N of taps; The multiple quantization intervals include a second number M of quantization intervals; and The first color sample is classified by a classifier having a third number L of categories, where L is equal to M N .
19. The computing system according to claim 17, wherein, One of the multiple filter types of the loop filter selected by the filter type index has a first number N of taps, and the multiple quantization intervals include a second number M of quantization intervals, and the one or more programs further include instructions for the following operations: Determine the second quantity M and the plurality of quantization levels according to the first quantity N.
20. The computing system according to claim 17, wherein, One of the plurality of filter types of the loop filter selected by the filter type index has a first quantity N of taps, and the first color sample is classified by a classifier having a third quantity L of categories. The one or more programs further include instructions for: Determine the third quantity L based on the first quantity N.
21. A non-transitory computer-readable storage medium storing one or more programs executed by a control circuit of a computing system, the one or more programs including instructions for the following operations: Receiving a video bitstream including a current encoded block of a current image frame, wherein, The video bitstream includes (i) a first syntax element for a cross-component sample offset (CCSO) mode that indicates whether a first sample offset of a first color sample of the current coding block is determined based on one or more luma samples, and (ii) a second syntax element for a filter type index and a filter shape index. The filter type index indicates one of the plurality of filter types of the loop filter selected for the current coding block, and the filter shape index identifies a set of one or more neighboring positions used by the loop filter; Determine one or more differences between one or more neighboring luma samples and a first luma sample; Quantize the one or more differences to generate one or more quantized differences; Classify the first color sample based on the one or more quantized differences to determine a first sample offset of the first color sample; And Reconstruct the current coding block at least by adjusting the first color sample based on the first sample offset of the first color sample.
22. The method according to claim 21, wherein, One of the plurality of filter types of the loop filter selected by the filter type index has a first quantity N of taps, and the first color sample corresponds to a plurality of bands. The one or more programs further include instructions for: Determine the number of bands in the plurality of bands based on the first quantity N, each band corresponding to a respective value of the filter shape index that identifies a respective set of one or more neighboring positions used by the loop filter; Identify one of the plurality of bands based on the value of the first color sample, the band including the first color sample; And Identify the respective value of the filter shape index and the respective set of the one or more neighboring positions based on the identified one of the plurality of bands, thereby selecting the one or more neighboring luma samples for generating the first sample offset.
23. The method according to claim 21, classifying the first color sample based on the one or more quantified differences, further comprising: Identify a combination of the one or more quantized differences in a look-up table that associates combinations of a plurality of quantized differences with a plurality of offset value options; And Determine the first sample offset corresponding to the combination of the one or more quantized differences in the look-up table.
24. The method according to claim 23, wherein The plurality of filter types of the loop filter includes a first quantity of filter types, each filter type having a different number of filter taps, and the look-up table includes the first quantity of sub-tables, each sub-table corresponding to a respective filter type.
25. The method according to claim 24, wherein Select one filter shape from a plurality of filter shapes corresponding to the filter type index of the loop filter according to the filter shape index; Each filter type has a plurality of corresponding filter shapes, and each filter shape has a plurality of different combinations of adjacent positions; And Each filter type corresponds to a respective sub-table that integrates the association of combinations of the quantized differences with the plurality of offset value options for the plurality of corresponding filter shapes.