Systems and methods for cross-component sample offset filter information signaling

By predicting filter parameters for CCSO filtering, the inefficiencies in signaling cross-component sample offset information are addressed, leading to improved coding efficiency and reduced bandwidth usage.

JP2025535629APending Publication Date: 2025-10-28TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024548549
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-10
Filing Date
2023-05-11
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

The signaling of cross-component sample offset filter information in video coding is costly, especially at small resolutions, limiting the coding performance of cross-component sample offset (CCSO) filtering techniques.

Method used

Predicting filter parameters for cross-component sample offset (CCSO) filtering based on previously predicted and/or selected parameter values, reducing the need for explicit signaling and improving coding efficiency.

Benefits of technology

Reduces signaling costs and enhances coding performance by implicitly predicting filter parameters for CCSO filtering, thereby optimizing video compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535629000001_ABST
    Figure 2025535629000001_ABST
Patent Text Reader

Abstract

Various embodiments described herein include methods and systems for encoding and decoding video. In one aspect, the method includes obtaining, from a video bitstream, video data having a plurality of blocks, including a first block. The method further includes predicting a parameter set for filtering the first block, the parameter set not being explicitly signaled in the video bitstream. The method also includes obtaining a filtered first block by applying cross-component sample offset (CCSO) to the first block, the CCSO being based on the parameter set. The method also includes reconstructing a picture from the video data using the filtered first block.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001]

[0001] Related Applications This application is a continuation of and claims priority to U.S. patent application Ser. No. 18 / 195,870, filed May 10, 2023, entitled "System and Method for Cross-Component Sample Offset Filter Information Signaling," and also claims priority to U.S. Provisional Patent Application Ser. No. 63 / 413,237, filed October 4, 2022, entitled "Improved Cross-Component Sample Offset Filter Information Signaling," all of which are incorporated herein by reference in their entireties.

[0002]

[0002] Technical field The disclosed embodiments relate generally to video coding and include, but are not limited to, systems and methods for signaling and predicting cross-component sample offset filter information. [Background technology]

[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, and video streaming devices. Electronic devices transmit, receive, or otherwise communicate digital video data over communication networks and / or store the digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can compress video data according to one or more video coding standards before the video data is communicated or stored.

[0004]

[0004] Multiple video codec standards have been developed. For example, video coding standards include AV1 (AOMedia Video 1), VVC (Versatile Video Coding), JEM (Joint Exploration test Model), HEVC / H.265 (High-Efficiency Video Coding), AVC / H.264 (Advanced Video Coding), and MPEG (Moving Picture Expert Group) coding. Video coding generally uses prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. The goal of video coding is to compress video data into a format that uses a lower bit rate while avoiding or minimizing degradation to video quality.

[0005]

[0005] HEVC, also known as H.265, is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4). VVC (Versatile Video Coding), also known as H.266, is a video compression standard intended as the successor to HEVC. ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (Version 1) and 2022 (Version 2). AV1 is an open video coding format designed as a replacement for HEVC. Validated version 1.0.0 was released on January 8, 2019, along with errata 1. Summary of the Invention

[0006] As mentioned above, encoding (compression) reduces bandwidth and / or storage space requirements. As will be explained in more detail later, both lossless and lossy compression can be used. Lossless compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal through a decoding process. Non-lossless compression refers to an encoding / decoding process in which the original video information is not fully preserved during encoding and is not fully recoverable during decoding. When using non-lossless compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough to make the reconstructed signal useful for the intended application. The amount of acceptable distortion depends on the application. For example, users of certain consumer video streaming applications may be able to tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect different distortion tolerances: higher distortion tolerance generally allows for coding algorithms that result in higher loss and higher compression ratios.

[0007]

[0007] Video encoders and / or decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding. A loop filtering technique called cross-component sample offset (CCSO) can sometimes be used to reduce distortion in reconstructed samples. In CCSO, given the processed input reconstructed samples of a first color component, a nonlinear mapping is used to derive an output offset, which is then added to the reconstructed samples of another color component in the CCSO filtering process. However, CCSO requires multiple syntax elements, such as quantization step size, filter shape index, band information, etc. This information is costly to signal, especially at small resolutions, and therefore limits the coding performance of CCSO.

[0008] According to some embodiments, there is provided a method for video decoding, comprising: (i) receiving, from a video bitstream, video data having a plurality of blocks including a first block; (ii) obtaining, from the video bitstream, syntax element values ​​indicating a difference between a parameter associated with the first block and a reference parameter associated with a second block, the parameter being one of a parameter set; (iii) deriving, based on the syntax element values, a value of a parameter associated with the first block, the parameter associated with the first block being not explicitly signaled in the video bitstream; (iv) obtaining a filtered first block by applying a cross-component sample offset (CCSO) to the first block, the CCSO being based on the parameter set; and (v) reconstructing a picture from the video data using the filtered first block.

[0009] According to some embodiments, there is provided a method for video encoding, the method including: (i) obtaining video data comprising a plurality of blocks including a first block and a second block; (ii) identifying a parameter set to use by applying a cross-component sample offset (CCSO) to the first block, where parameters in the parameter set are identified using a first reference parameter, the reference parameter corresponding to the second block; and (iii) signaling the identified parameter set in a video bitstream.

[0010] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes control circuitry and a memory that stores one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.

[0011] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets for execution by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.

[0012]

[0012] Accordingly, methods, as well as devices and systems, for encoding and decoding video are disclosed that can complement or replace conventional methods, devices and systems for video encoding / decoding.

[0013]

[0013] The features and advantages described in the specification are not necessarily exhaustive, and some additional features and advantages will be apparent to those skilled in the art, especially in view of the drawings, specification, and claims provided in this disclosure. Furthermore, it should be noted that the language used in this specification has been selected primarily for readability and instructional purposes, and not necessarily to delineate or limit the subject matter described herein. [Brief explanation of the drawings]

[0014]

[0014] In order that the present disclosure may be understood in more detail, a more particular description may be made by reference to features of various embodiments, some of which are shown in the accompanying drawings. However, the accompanying drawings merely illustrate relevant features of the present disclosure and are therefore not intended to be necessarily limiting to the description, since other useful features may be recognized by those skilled in the art upon reading the present disclosure. [Figure 1]

[0015] FIG. 1 is a block diagram illustrating an exemplary communication system in accordance with some embodiments. [Figure 2A]

[0016] FIG. 2A is a block diagram illustrating exemplary elements of an encoder component according to some embodiments. [Figure 2B]

[0017] FIG. 2B is a block diagram illustrating exemplary elements of a decoder component according to some embodiments. [Figure 3]

[0018] FIG. 3 is a block diagram illustrating an exemplary server system according to some embodiments. [Figure 4A]

[0019] 4A-4D illustrate exemplary coding tree structures according to some embodiments. [Figure 4B] 4A-4D illustrate exemplary coding tree structures according to some embodiments. [Figure 4C] 4A-4D illustrate exemplary coding tree structures according to some embodiments. [Figure 4D] 4A-4D illustrate exemplary coding tree structures according to some embodiments. [Figure 5A]

[0020] FIG. 5A illustrates an exemplary filter shape according to some embodiments. [Figure 5B]

[0021] FIG. 5B illustrates exemplary sub-sample locations for gradient calculations according to some embodiments. [Figure 5C]

[0022] FIG. 5C illustrates an exemplary quadtree division approach according to some embodiments. [Figure 5D]

[0023] FIG. 5D illustrates an exemplary quadtree split flag encoded in z-order according to some embodiments. [Figure 5E]

[0024] FIG. 5E illustrates an exemplary filter support area according to some embodiments. [Figure 6A]

[0025] FIG. 6A is a flow chart illustrating an example method for decoding video according to some embodiments. [Figure 6B]

[0026] FIG. 6B is a flow chart illustrating an example method for encoding video according to some embodiments.

[0027] According to common practice, the various features illustrated in the drawings are not necessarily drawn to scale and like reference numerals may be used to refer to like features throughout the specification and drawings. DETAILED DESCRIPTION OF THE INVENTION

[0015]

[0028] This disclosure describes, among other things, predicting, identifying, and signaling filtering parameters for block filtering. For example, filter parameters for cross-component sample offset (CCSO) filtering can be predicted based on previously predicted and / or selected parameter values. Predicting filter parameters, as opposed to explicitly signaling them, reduces signaling costs, which can improve coding efficiency and / or reduce bandwidth.

[0016]

[0029] Exemplary Systems and Devices 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 through 120-m) communicatively coupled to one another via one or more networks. In some embodiments, the communication system 100 is a streaming system for use with video-enabled applications, such as, for example, video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0017]

[0030] Source device 102 includes video source 104 (e.g., a camera component or media storage) and encoder component 106. In some embodiments, video source 104 is a digital camera (e.g., configured to generate an uncompressed video sample stream). Encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from video source 104 may be a large amount of data compared to encoded video bitstream 108 generated by encoder component 106. Because encoded video bitstream 108 is a smaller amount of data (less data) compared to the video stream from the video source, encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared to the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video data over network 110).

[0018]

[0031] The one or more networks 110 represent any number of networks that convey information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired and / or wireless communication networks. The one or more networks 110 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0019]

[0032] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as encoded video streams from source device 102). Server system 112 includes a coder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, coder component 114 includes an encoder component and / or a decoder component. In various embodiments, coder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, coder component 114 is configured to decode encoded video bitstream 108 and re-encode the video data using a different encoding standard and / or methodology to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video bitstream 108.

[0020]

[0033] In some embodiments, server system 112 functions as a Media-Aware Network Element (MANE). For example, server system 112 may be configured to prune encoded video bitstream 108 to accommodate potentially different bitstreams for one or more electronic devices 120. In some embodiments, a MANE is provided separate from server system 112.

[0021]

[0034] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, decoder component 122 is configured to decode encoded video data 116 to generate an output video stream capable of being rendered on a display or other type of rendering device. In some embodiments, one or more electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include media storage). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.

[0022]

[0035] The source device and / or the electronic devices 120 are sometimes referred to as “terminal devices” or “user devices.” In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smartphone, tablet, or laptop), a wearable device, a videoconferencing device, and / or other types of electronic devices.

[0023]

[0036] In an example operation of the communication system 100, the source device 102 transmits an encoded video bitstream 108 to the server system 112. For example, the source device 102 may code a stream of pictures captured by the source device. The server system 112 may receive the encoded video bitstream 108 and decode and / or encode the encoded video bitstream 108 using a coder component 114. For example, the server system 112 may apply coding to the video data that is more optimal for network transmission and / or storage. The server system 112 may transmit the encoded video data 116 (e.g., one or more coded video bitstreams) to one or more electronic devices 120. Each electronic device 120 may decode the encoded video data 116 to recover and optionally display video pictures.

[0024]

[0037] In some embodiments, the transmission is a one-way data transmission. One-way data transmission is sometimes used, such as in media serving applications. In some embodiments, the transmission is a two-way data transmission. Two-way data transmission is sometimes used, such as in video conferencing applications. In some embodiments, the coded video bitstream 108 and / or the coded video data 116 are coded and / or decoded according to any video coding / compression standard described herein, such as HEVC, VVC, and / or AV1.

[0025]

[0038] FIG. 2A is a block diagram illustrating exemplary elements of the encoder component 106, according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 can provide the source video sequence in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 YCrCb, or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0 or YCrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed sequentially. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., used. Those skilled in the art will readily understand the relationship between pixels and samples. The following discussion focuses on samples.

[0026]

[0039] The encoder component 106 is configured to code and / or compress pictures of a source video sequence into a coded video sequence 216 in real time or under other time constraints required by the application. Imposing an appropriate coding rate is one of the functions of the controller 204. In some embodiments, the controller 204 controls and is operatively coupled to other functional units, as described below. Parameters set by the controller 204 may include rate control-related parameters (e.g., picture skip, quantizer, and / or lambda value for rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art will readily be able to identify other functions of the controller 204, as they may be associated with the encoder component 106 to be optimized for a particular system design.

[0027]

[0040] In some embodiments, the encoder component 106 is configured to operate in a coding loop. In a simplified example, the coding loop includes a source coder 202 (e.g., responsible for generating a symbol stream based on an input picture to be coded and a reference picture) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to a (remote) decoder (if the compression between the symbols and the coded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory 208. Because decoding the symbol stream produces bit-exact results independent of the location of the decoder (local or remote), the contents in the reference picture memory 208 are also bit-exact between the local and remote encoders. In this way, the predictive portion of the encoder interprets the same sample values ​​as reference picture samples when using prediction during decoding. This principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, eg, due to channel errors) is known to those skilled in the art.

[0028]

[0041] The operation of decoder 210 may be identical to that of a remote decoder, such as decoder component 122, which is described in more detail below in connection with Figure 2B. However, with brief reference to Figure 2B, because symbols are available and the encoding / decoding of symbols into a coded video sequence by entropy coder 214 and parser 254 is lossless, the entropy decoding portion of decoder component 122, including buffer memory 252 and parser 254, may not be implemented entirely within local decoder 210.

[0029]

[0042] A possible insight at this point is that any decoder technique, with the exception of analysis / entropy decoding, that is present in the decoder must also be present in substantially identical functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder techniques can be omitted, as they are the reverse of the decoder techniques that are described generically. Only in certain areas is a more detailed description required and is provided below.

[0030]

[0043] As part of its operation, the source coder 202 may perform motion-compensated predictive coding (predictive coding of an input frame with reference to one or more previously coded frames from a video sequence, designated as reference frames). In this method, the coding engine 212 codes the differences between pixel blocks of the input frame and pixel blocks of reference frames that can be selected as prediction references for the input frame. The controller 204 may manage the coding operations of the source coder 202, including, for example, setting parameters and subgroup parameters used to encode the video data.

[0031]

[0044] The decoder 210 decodes the coded video data of frames that may be designated as reference frames based on symbols generated by the source coder 202. The operation of the coding engine 212 may advantageously be a non-lossless process. When the coded video data is decoded by a video decoder (not shown in FIG. 2A), the reconstructed video sequence may be a replica of the source video sequence with some errors. The decoder 210 may replicate the decoding process that may be performed by a remote video decoder with respect to the reference frames and cause the reconstructed reference frames to be stored in the reference picture memory 208. In this manner, the encoder component 106 locally stores copies of reconstructed reference frames that have common content as the reconstructed reference frames that would be obtained by the remote video decoder (assuming no transmission errors).

[0032]

[0045] The predictor 206 may perform the predictive search of the coding engine 212. That is, for a new frame to be coded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or specific metadata (e.g., reference picture motion vectors, block shapes, etc.) that can serve as suitable prediction references for the new picture. The predictor 206 may operate on a sample block-by-pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor 206, an input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory 208.

[0033]

[0046] The output of all the aforementioned functional units may be entropy coded in entropy coder 214. Entropy coder 214 converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).

[0034]

[0047] In some embodiments, the output of the entropy coder 214 is coupled to a transmitter. The transmitter may be configured to buffer coded video sequences created by the entropy coder 214 and prepare them for transmission over a communication channel 218 (which may be a hardware / software link to a storage device that stores the coded video data). The transmitter may be configured to merge the coded video data from the source coder 202 with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may transmit additional data along with the coded video. The source coder 202 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data (e.g., redundant pictures and slices), supplemental enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0035]

[0048] The controller 204 may manage the operation of the encoder component 106. During coding, the controller 204 may assign a specific coding picture type to each coded picture, which may affect the coding technique applied to the respective picture. For example, a picture may be designated as an intra picture (I picture), a predicted picture (P picture), or a bidirectionally predicted picture (B picture). An intra picture may be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will recognize these variations of I pictures, as well as their respective uses and characteristics, and therefore will not be repeated here. A predicted picture may be coded and decoded using intra prediction or inter prediction, and may use at most one motion vector and reference index to predict the sample values ​​of each block. Bi-directionally predicted pictures are coded and decoded using intra or inter prediction and can predict the sample values ​​of each block using at most two motion vectors and reference indices. Similarly, multi-predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0036]

[0049] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding specification applied to the block's respective picture. For example, blocks of an I-picture may be coded nonpredictively or predictively with reference to previously coded blocks of the same picture (spatial or intra prediction). Picture blocks of a P-picture may be nonpredictively coded with spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be nonpredictively coded with spatial or temporal prediction with reference to one or two previously coded reference pictures.

[0037]

[0050] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture being encoded / decoded (called the current picture) is divided into blocks. If a block in the current picture is similar to a reference block in a reference picture that was previously coded in the video and is still buffered, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0038]

[0051] The encoder component 106 may perform coding operations according to a given video coding technique or standard, as described elsewhere herein. In doing so, the encoder component 106 may perform various compression processes, including predictive coding processes that exploit temporal and spatial redundancies in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0039]

[0052] 2B is a block diagram illustrating exemplary elements of the decoder component 122 according to some embodiments. The decoder component 122 of FIG. 2B is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter unit 256 and configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0040]

[0053] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more coded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each coded video sequence is independent of the other coded video sequences. Each coded video sequence may be received from channel 218 (which may be a hardware / software link to a storage device that stores the coded video data). The receiver may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be transported using respective entities (not shown). The receiver may separate the coded video sequence from the other data. In some embodiments, the receiver receives additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data can be used by decoder component 122 to decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0041]

[0054] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra-picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. In some embodiments, the decoder component 122 is implemented at least partially in software.

[0042]

[0055] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to address network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., configured to handle playback timing), a separate buffer memory is provided external to decoder component 122 (e.g., to address network jitter). Buffer memory 252 may not be needed, or can be small, when receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isosynchronous network. For use over best effort packet networks such as the Internet, the buffer memory 252 may be required to be relatively large and may advantageously be adaptively sized and implemented, at least in part, in an operating system or similar element (not shown) external to the decoder component 122.

[0043]

[0056] The parser 254 is configured to reconstruct symbols 270 from the coded video sequence. The symbols may include, for example, information used to manage the operation of the decoder component 122 and / or information for controlling a rendering device such as the display 124. The rendering device control information may be in the form of, for example, a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser 254 parses (entropy decodes) the coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 may extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Parser 254 may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0044]

[0057] The reconstruction of symbol 270 may include several different units, depending on the type of coded video picture or portion thereof (inter and intra pictures, inter and intra blocks, etc.), and other factors. Which units are included and how they participate may be controlled by subgroup control information parsed from the coded video sequence by parser 254. The flow of such subgroup control information between parser 254 and subsequent units is not depicted for clarity.

[0045]

[0058] Beyond the functional blocks already described, the decoder component 122 may be conceptually subdivided into a number of functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units has been provided:

[0046]

[0059] The scaler / inverse transform unit 258 receives the quantized transform coefficients as well as control information (such as the transform to use, block size, quantization factor, and / or quantization scaling matrix) from the parser 254 as symbols 270. The scaler / inverse transform unit 258 can output blocks containing sample values ​​that can be input to the aggregator 268.

[0047]

[0060] In some cases, the output samples of the scaler / inverse transform unit 258 relate to intra-coded blocks; that is, blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 262, which may generate blocks of the same size and shape as the block being reconstructed using surrounding reconstructed information from the current (partially reconstructed) picture from the current picture memory 264. The aggregator 268 may add the prediction information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 on a sample-by-sample basis.

[0048]

[0061] In other cases, the output samples of the scalar / inverse transform unit 258 relate to an inter-coded, potentially motion-compensated, block. In such cases, the motion-compensated prediction unit 260 may access the reference picture memory 266 to retrieve samples for use in prediction. After motion-compensating the retrieved samples according to the symbols 270 associated with the block, these samples may be added by the aggregator 268 to the output of the scalar / inverse transform unit 258 (in this case, referred to as residual samples or a residual signal) to generate output sample information. The addresses in the reference picture memory 266 from which the motion-compensated prediction unit 260 retrieves the prediction samples may be controlled by a motion vector. The motion vector may be available to the motion-compensated prediction unit 260 in the form of a symbol 270, which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​as fetched from the reference picture memory 266, motion vector prediction mechanisms, etc., when sub-sample accurate motion vectors are used.

[0049]

[0062] The output samples of the aggregator 268 may be subjected to various loop filtering techniques in the loop filter unit 256. Video compression techniques may include in-loop filtering techniques controlled by parameters included in the coded video bitstream and made available to the loop filter unit 256 as symbols 270 from the parser 254, but may also be responsive to meta-information obtained during decoding of previous portions of the coded picture or coded video sequence (in decoding order), or to previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 may be a sample stream that may be output to a rendering device such as the display 124 or stored in the reference picture memory 266 for use in future inter-picture prediction.

[0050]

[0063] Certain coded pictures, once fully reconstructed, can be used as reference pictures for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 254), the current reference picture can become part of reference picture memory 266, and a fresh current picture memory can be reassigned before beginning reconstruction of a subsequent coded picture.

[0051]

[0064] Decoder component 122 may perform decoding operations according to a given video compression technique, which may be documented in a standard, such as any of the standards described herein. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that it conforms to the syntax of the video compression technique or standard as specified in the video compression technique document or standard, particularly the profile document therein. To comply with some video compression techniques or standards, the complexity of the coded video sequence may also be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited, in some cases, by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0052]

[0065] 3 is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., a CPU, a GPU, and / or a DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).

[0053]

[0066] The network interface 304 may be configured to interface with one or more communications networks (e.g., wireless, wired, and / or optical networks). Communications networks can be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of communications networks include Ethernet, wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, and the like), TV wired or wireless wide-area digital networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicular, and industrial networks including CANbus, and the like. Such communications can be unidirectional, receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus to a specific CANbus device), or bidirectional (e.g., to other computer systems using local or wide-area digital networks). Such communications can include communications to one or more cloud computing networks.

[0054]

[0067] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input devices 310 may include one or more of: a keyboard, a mouse, a trackpad, a touch screen, a data glove, a joystick, a microphone, a scanner, a camera, etc. The output devices 308 may include one or more of: an audio output device (e.g., a speaker), a visual output device (e.g., a display or monitor), etc.

[0055]

[0068] Memory 314 may include high-speed random-access memory (such as DRAM, SRAM, DDR RAM, and / or other random-access solid-state memory devices) and / or non-volatile memory (such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Memory 314 optionally includes one or more storage devices located remotely from control circuitry 302. Memory 314, or alternatively, a non-volatile solid-state memory device within memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, memory 314 or the non-transitory computer-readable storage medium of memory 314 stores the following programs, modules, instructions, and data structures, or a subset or superset thereof:

[0056] An operating system 316 that includes procedures for handling various basic system services and performing hardware-dependent tasks; · a network communications module 318 used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); A coding module 320 for performing various functions related to encoding and / or decoding data, such as video data; in some embodiments, the coding module 320 is an instance of the coder component 114. The coding module 320 may include, but is not limited to, one or more of the following: a decoding module 322 for performing various functions related to decoding the encoded data, such as those described above with respect to the decoder component 122; and an encoding module 340 for performing various functions on the encoded data, such as those described above with respect to the encoder component 106; Picture memory 352 for storing pictures and picture data, for example for use with coding module 320; in some embodiments, picture memory 352 includes one or more of reference picture memory 208, buffer memory 252, current picture memory 264, and reference picture memory 266.

[0057]

[0069] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions described above with respect to the parser 254), a transform module 326 (e.g., configured to perform various functions described above with respect to the scalar / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions described above with respect to the motion compensated prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions described above with respect to the loop filter unit 256).

[0058]

[0070] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions described above with respect to the source coder 202, the coding engine 212, and / or the entropy coder 214) and a prediction module 344 (e.g., configured to perform various functions described above with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include a subset of the modules shown in FIG. 3. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0059]

[0071] Each of the above-identified modules stored in memory 314 corresponds to a set of instructions for performing the functions described herein. The above-identified modules (e.g., sets of instructions) are not required to be implemented as separate software programs, procedures, or modules; thus, various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, coding module 320 optionally does not include separate decoding and encoding modules, but rather uses the same set of modules to perform both sets of functionality. In some embodiments, memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0060]

[0072] In some embodiments, server system 112 includes a web or Hypertext Transfer Protocol (HTTP) server, a File Transfer Protocol (FTP) server, and web pages and applications implemented using Common Gateway Interface (CGI) scripts, PHP Hyper-text Preprocessor (PHP), ASP (Active Server Pages), HTML (HyperText Markup Language), XML (Extensible Markup Language), Java, JavaScript, Asynchronous JavaScript and XML (AJAX), XHP, Javelin, Wireless Universal Resource File (WURFL), etc.

[0061]

[0073] While FIG. 3 illustrates a server system 112 according to some embodiments, FIG. 3 is intended as a functional illustration of various features that may be present in one or more server systems, rather than an architectural overview of the embodiments described herein. In practice, those skilled in the art will recognize that items shown separately may be combined and some items may be separated. For example, some items shown separately in FIG. 3 may be implemented on a single server, and a single item may be implemented by more than one server. The actual number of servers used to implement server system 112, and how functionality is allocated among them, will vary from implementation to implementation and, optionally, will depend in part on the amount of data traffic the server system handles during peak and average usage periods.

[0062]

[0074] Example Coding Techniques 4A-4D illustrate exemplary coding tree structures according to some embodiments. As shown in the first coding tree structure (400) of FIG. 4A, some coding schemes (e.g., VP9) use a four-way partition tree starting at the 64x64 level down to the 4x4 level, with some additional restrictions for 8x8 blocks. In FIG. 4A, the partitions shown as R can be referred to as recursive in that the same partition tree is repeated at lower scales until the lowest 4x4 level is reached.

[0075] As shown in the second coding tree structure (402) of Figure 4B, some coding techniques (e.g., AV1) extend the partition tree to a 10-way structure and increase the maximum size (e.g., called a superblock in VP9 / AV1 terminology) starting at 128x128. The second coding tree structure includes 4:1 / 1:4 rectangular partitions, which are not present in the first coding tree structure. The partition type with three subpartitions in the second row of Figure 4B is called a T-shaped partition. Rectangular partitions in this tree structure cannot be further subdivided. In addition to the coding block size, a coding tree depth can be defined to indicate the partition depth from the root node. For example, the coding tree depth of the root node, e.g., 128x128, is set to 0, and after the tree block is further divided, the coding tree depth is increased by 1. For example, instead of enforcing a fixed transform unit size as in VP9, ​​AV1 allows luma coding blocks to be divided into transform units of multiple sizes, which can be expressed by a recursive partition going down to a maximum of two levels. To incorporate AV1's expanded coding block partitioning, square, 2:1 / 1:2, and 4:1 / 1:4 transform sizes from 4x4 to 64x64 are supported. For chroma blocks, only the largest possible transform unit is allowed.

[0063]

[0076] As an example, a CTU can be divided into CUs (units) by using a quadtree structure, denoted as a coding tree, to adapt to various local characteristics, as in HEVC. In some embodiments, the decision of whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into one, two, or four PUs depending on the PU partition type. Within a PU, the same prediction process is applied, and related information is transmitted to the decoder on a PU-by-PU basis. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU can be partitioned into TUs according to another quadtree structure, such as the CU coding tree. One important feature of the HEVC structure is that it has multiple partition concepts, including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square, while a PU can be square or rectangular for inter-predicted blocks. In HEVC, a coding block is further divided into four square sub-blocks, and a transform is performed on each sub-block (TU). Each TU can be further divided recursively (using quad-tree partitioning) into smaller TUs, called residual quad-trees (RQTs). At picture boundaries, such as in HEVC, implicit quad-tree partitioning is used, allowing blocks to remain quad-tree partitioned until their size fits the picture boundary.

[0064]

[0077] A quad-tree with nested multi-type trees, using binary and ternary partition segmentation structures like those in VVC, can replace the concept of multiple partition unit types, eliminating the need for separate CU, PU, ​​and TU concepts except when necessary for CUs that are too large for the maximum transform length, and supporting more flexibility in CU partition shapes. In the coding tree structure, CUs can have either square or rectangular shapes. ACTUs are first partitioned using a quaternary tree (also called a quad-tree) structure. The quaternary tree leaf nodes can be further partitioned using a multi-type tree structure. As shown in the third coding tree structure (404) in Figure 4C, the multi-type tree structure includes four partition types. For example, the multi-type tree structure includes vertical binary splitting (SPLIT_BT_VER), horizontal binary splitting (SPLIT_BT_HOR), vertical ternary splitting (SPLIT_TT_VER), and horizontal ternary splitting (SPLIT_TT_HOR). The multi-type tree leaf nodes are called CUs, and this segmentation is used for prediction and transform processing without any further partitioning, unless the CU is too large for the maximum transform length. This means that in most cases, the CU, PU, ​​and TU have the same block size in a quad-tree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color component of the CU. An example of block partitioning for one CTU (406) is shown in Figure 4D, which shows an example of a quad-tree with a nested multi-type tree coding block structure.

[0065]

[0078] As in VVC, the maximum supported luma transform size may be 64x64, and the maximum supported chroma transform size may be 32x32. If the width or height of the CB is larger than the maximum transform width or height, the CB is automatically split horizontally and / or vertically to fit the transform size limit in that direction.

[0066]

[0079] The coding tree method supports the ability for luma and chroma to have separate block tree structures, as in VTM7. In some cases, for P and B slices, the luma and chroma coding tree blocks (CTBs) within one CTU share the same coding tree structure. However, for I slices, luma and chroma can have separate block tree structures. When the separate block tree mode is applied, the luma CTB is divided into CUs by one coding tree structure, and the chroma CTB is divided into chroma CUs by another coding tree structure. This means that a CU in an I slice can contain or consist of a coding block for the luma component or a coding block for two chroma components, and a CU in a P or B slice can always contain or consist of coding blocks for all three color components, unless the video is monochrome.

[0067]

[0080] To support extended coding block partitions, as in AV1, multiple transform sizes (e.g., ranging from 4 points to 64 points in each dimension) and transform shapes (e.g., square or rectangular with width / height ratios of 2:1 / 1:2 and 4:1 / 1:4) may be used.

[0068]

[0081] A loop filter can be applied to a block (e.g., CU) to reduce distortion introduced in the encoding process and thus improve the reconstructed quality. For example, an adaptive loop filter (ALF) with block-based filter adaptation can be applied. As an example, for the luma component, one of 25 filters is selected for each 4x4 block based on local gradient direction and activity. FIG. 5A illustrates exemplary ALF filter shapes according to some embodiments. Two diamond filter shapes shown in FIG. 5A can be used with the ALF filter. For example, a 7x7 diamond shape is applied to the luma component and a 5x5 diamond shape is applied to the chroma component.

[0069]

[0082] For the luma component, each 4x4 block can be classified into one of 25 classes. The classification index C is determined by its directionality D and the quantized activity value A. ^ Based on this, it can be derived as follows: Formula 1 - Classification Index

[0070]

number

[0083] D and A ^ To calculate , first the horizontal, vertical, and two diagonal gradients are calculated using the 1-D Laplacian, as shown in Equation 2-5 below: Equation 2 - Vertical Gradient

[0071]

number

[0072]

number

[0073]

number

[0074]

number

[0075]

[0084] The maximum and minimum values ​​of the horizontal and vertical gradients of D are set as shown in Equation 7, and the maximum and minimum values ​​of the gradients in the two diagonal directions are set as shown in Equation 8, as follows:

[0076] Equation 7 - Maximum and minimum horizontal and vertical gradients

[0077]

number

[0078]

number

[0085] These values ​​are compared with two thresholds t1 and t2 to derive the value of the directionality D. For example, in step 1:

[0079]

number

[0080]

number

[0081]

number

[0082]

number

[0083]

number

[0084]

[0086] For chroma components within a picture, no classification method needs to be applied, eg, a single set of ALF coefficients is applied for each chroma component.

[0085]

[0087] ALF filter parameters can be signaled in an adaptation parameter set (APS). For example, up to 25 pairs of luma filter coefficients and clipping value indices and up to 8 pairs of chroma filter coefficients and clipping value indices can be signaled in one APS. To reduce bit overhead, filter coefficients of different classifications for the luma component can be merged. In the slice header, the index of the APS used for the current slice is signaled. For example, ALF signaling is CTU-based in VVC (Draft 8).

[0086]

[0088] The clipping value index decoded from the APS allows the clipping values ​​to be determined using a table of luma and chroma clipping values. These clipping values ​​depend on the internal bit depth. The clipping value table can be obtained by applying Equation 10 below.

[0087] Equation 10 - Clipping Value

[0088]

number

[0089]

[0089] In the slice header, up to seven APS indices can be signaled to specify the luma filter set to be used for the current slice. The filtering process can be further controlled at the CTB level. For example, a flag can be signaled to indicate whether ALF is applied to the luma CTB. The luma CTB can select a filter set from an APS and a filter set from 16 fixed filter sets. A filter set index can be signaled to the luma CTB to indicate which filter set to apply. The 16 fixed filter sets may be predefined and hard-coded in both the encoder and the decoder.

[0090] For chroma components, an APS index may be signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if there is more than one chroma filter set in an APS, a filter index may be signaled for each chroma CTB.

[0091]

[0091] The filter coefficients can be quantized with a norm equal to 128. To limit the complexity of the multiplications, bitstream conformance is applied so that coefficient values ​​at non-central positions are within the limits of (-2 7 ) to (2 7 -1). The center position coefficient is not signaled in the bitstream and is assumed to be equal to 128.

[0092] AlfClip based on table 1-bitDepth and clipIdx

[0093] [Table 1]

[0092] On the decoder side, when ALF is enabled for a CTB, each sample R(i,j) in a CU is filtered to result in a sample value R' as shown in Equation 11 below.

[0094] Equation 11 - Sample Value Calculation

[0095]

number

[0096]

[0093] To improve coding efficiency, a coding unit synchronous picture quadtree-based adaptive loop filter can be applied. For example, a luma picture is divided into several multi-level quadtree partitions, and each partition boundary is aligned with the boundary of the largest coding unit (LCU). Each partition has its own filtering process and is therefore called a filter unit (FU).

[0097]

[0094] A two-pass encoding flow can be applied.

[0098] In the first pass, the quadtree division pattern and the best filter for each FU are determined. The filtering distortion can be estimated by fast filtering distortion estimation (FFDE) during the decision process. The reconstructed picture is filtered according to the determined quadtree division pattern and the selected filters for all FUs.

[0099] In the second pass, CU-synchronized ALF on / off control is performed. According to the result of ALF on / off, the original filtered picture is partially restored by the reconstructed picture.

[0100] A top-down partitioning strategy can be used to divide a picture into multi-level quadtree partitions by using a rate-distortion criterion. Each partition is called a filter unit. The partitioning process aligns the quadtree partitions to LCU boundaries. The coding order of the FUs follows the z-scan order. For example, in Figure 5C, the picture is divided into 10 FUs, and the coding order is FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, and FU9. To indicate the quadtree partitioning pattern of the picture, a partition flag can be coded and transmitted in z-order. Figure 5D shows the quadtree partitioning pattern corresponding to Figure 5C.

[0101]

[0096] The filter for each FU can be selected from two filter sets based on a rate-distortion criterion. The first set contains half-symmetric square and triangle-shaped filters newly derived for the current FU or current FU. The second set is derived from the time-delay filter buffer; the time-delay filter buffer stores filters previously derived for FUs of previous pictures. The filter with the smallest rate-distortion cost among these two sets is selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further divided into four child FUs, the rate-distortion costs of the four child FUs are calculated. By recursively comparing the rate-distortion costs for the division and non-division cases, a picture quadtree division pattern can be determined. If the maximum quadtree division level is 2, the maximum number of FUs is 16. During the quadtree division decision, correlation values ​​for deriving Wiener coefficients for the 16 FUs at the bottom quadtree level (smallest FU) can be reused. The remaining FUs can derive their Wiener filters from the correlation of the 16 FUs at the bottom quadtree level, so only one frame buffer access is required to derive the filter coefficients for all FUs.

[0102]

[0097] After the quadtree division pattern is determined, CU-synchronized ALF on / off control is performed to further reduce filtering distortion. By comparing filtering distortion with non-filtering distortion, leaf CUs can explicitly switch ALF on / off in their local regions. Coding efficiency can be further improved by redesigning the filter coefficients according to the ALF on / off result. However, the redesign process requires additional frame buffer accesses. In some encoder designs, there is no redesign process after CU-synchronized ALF on / off decision to minimize the number of frame buffer accesses.

[0103] A loop filtering technique called cross-component sample offset (CCSO) can reduce distortion in the reconstructed samples. In CCSO, given the processed input reconstructed samples of a first color component, a nonlinear mapping is used to derive an output offset, which is then added to the reconstructed samples of another color component in the CCSO filtering process.

[0104]

[0099] The input reconstructed samples are from the first color component located in the filter support area. As shown in Figure 5E, the filter support area includes four reconstructed samples: p0, p1, p2, and p3. The four input reconstructed samples follow a cross shape in the vertical and horizontal directions. The center sample c in the first color component and the filtering sample f in the second color component are co-located. In some embodiments, the center sample c corresponds to the luma pixel, and the sample f corresponds to the chroma pixel. When processing the input reconstructed samples, the following steps can be applied:

[0105] In step 1, the delta values ​​between p0-p3 and c are calculated and denoted as m0, m1, m2, and m3.

[0106] In step 2, the delta values ​​m0-m3 are further quantized, and the quantized values ​​are denoted as d0, d1, d2, and d3.

[0107] For example, the quantized values ​​may be calculated based on the following criteria: (i) If m<-N, then d=-1; (ii) -If N<=m<=N, then d=0; (iii) If m>N, then d=1; Based on the quantization step size, the possible values ​​are -1, 0, and 1. In this example, N is called the quantization step size, and example values ​​for N are 4, 8, 12, and 16.

[0108]

[0100] d0-d3 can be used to identify one of the nonlinear mapping combinations. CCSO has four filter taps d0-d3, and each filter tap can have one of three quantized values, resulting in a total of 3^4=81 combinations. Examples of offset values ​​are integers such as 0, 1, -1, 3, -3, 5, -5, and -7. The final filtering process of CCSO is applied using Equation 12 below.

[0109] Equation 12 - Filtered Sample Value f'=clip(f+s) where f is the reconstructed sample to filter and s is the output offset value looked up from the table of combinations identified by d0-d3. The filtered sample value f' is further clipped to be within the range associated with the bit depth.

[0110] As mentioned above, CCSO may require several syntax elements, such as quantization step size, filter shape index, band information, etc., to be signaled in the picture header or at the block level. However, this information can be costly to signal (especially for small resolutions) and limit the coding performance of CCSO. According to some embodiments, the quantization step size, number of bands, and / or filter shape values ​​are predicted from the quantization step size, number of bands, and / or filter shape index values ​​applied to different pictures, frames, planes, slices, or blocks.

[0111] 6A is a flow chart illustrating a method 600 for decoding video according to some embodiments. Method 600 may be performed on a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 600 is performed by executing instructions stored in a memory (e.g., memory 314) of the computing system.

[0112]

[0103] A system receives video data having a plurality of blocks including a first block from a video bitstream (602). The system determines a plurality of transform coefficients associated with the first block (604). In an embodiment, the system obtains syntax element values ​​from the video bitstream, the syntax element values ​​indicating differences between parameters associated with the first block and reference parameters associated with a second block, the parameters being one of a parameter set.

[0113] The system predicts (606) a parameter set for filtering the first block, where the parameter set is not explicitly signaled in the video bitstream. In some embodiments, the system derives values ​​of parameters associated with the first block based on syntax element values, where the parameters associated with the first block are not explicitly signaled in the video bitstream.

[0114] The system obtains a filtered first block by applying cross-component sample offset (CCSO) to the first block (608), where CCSO is based on the parameter set. The system reconstructs a picture from the video data using the filtered first block (610). Method 600 is optionally applied to luma and / or chroma blocks. The term "block" as used herein may be interpreted as a prediction block, a coding block, or a coding unit (CU). The term "block" may also refer to a transform block, a coding tree unit (CTU) block, or a CCSO superblock. For example, the CCSO filtering process uses reconstructed samples of a first color component (e.g., Y, Cb, or Cr) as input, and the output is applied to a second color component that is different from the first color component.

[0115] In some embodiments, the set of parameters includes a number of bands, a quantization step size, and / or a filter shape. In some embodiments, the quantization step size, number of bands, and / or filter shape values ​​are predicted from the quantization step size, number of bands, and filter shape index values ​​applied to different pictures, frames, planes, slices, or blocks. In some embodiments, instead of signaling the quantization step size, number of bands, and / or filter shape index values, the difference (delta) between these indexes and the predicted / previous index is signaled.

[0116] In some embodiments, the difference between the quantization step, number of bands, and / or filter shape value of the current and previous picture, frame, plane, slice, or block is signaled. In some embodiments, the difference between one or more of the quantization step, number of bands, and / or filter shape value of the current and previous picture, frame, plane, slice, or block is signaled. In some embodiments, the sign and magnitude of the delta value are signaled separately. In some embodiments, lookup table based signaling is used. For example, a lookup table index representing the delta value is signaled.

[0117] In some embodiments, whether the quantization step size, number of bands, and / or filter shape index value are predicted is signaled by a flag in a high level syntax (HLS), including, but not limited to, an APS, a slice header, a frame header, a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In some embodiments, a flag is signaled for one or more of the quantization step size, number of bands, and / or filter shape index value to indicate whether the value is predicted or explicitly signaled. In some embodiments, at least one of the quantization step size, number of bands, and filter shape index value is not signaled and is derived from the predicted index value.

[0118] In some embodiments, signaling of quantization step size, number of bands, and / or filter shape index values ​​is skipped entirely and predicted / previous values ​​are used instead. Prediction of quantization step size, number of bands, and / or filter shape index values ​​using previous picture, frame, plane, slice, or block values ​​may depend on coded information including, but not limited to, frame type, temporal layer, and / or quantization parameters.

[0119] 6B is a flow chart illustrating a method 650 for decoding video according to some embodiments. Method 650 may be performed on a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 650 is performed by executing instructions stored in a memory (e.g., memory 314) of the computing system.

[0120] The system obtains (652) video data having a plurality of blocks, including a first block and a second block. The system identifies (654) a parameter set to use by applying a cross-component sample offset (CCSO) to the first block, where parameters in the parameter set are identified using a first reference parameter, and the reference parameter corresponds to the second block. The system sends (656) a signal in the video bitstream signaling the identified parameter set. Method 650 is optionally applied to luma and / or chroma blocks.

[0121]

[0112] For example, the encoder search for the optimal quantization step, band, and / or filter shape value for the current picture, frame, plane, slice, or block may be optimized using previously selected / predicted quantization step, band, and / or filter shape values, reducing encoder complexity.

[0122] In some embodiments, the encoder search for at least one of the quantization step size, number of bands, and / or filter shape index value is skipped entirely, and predicted / previous values ​​are used instead. In some embodiments, the predicted / previously selected values ​​are used to reduce the encoder search for optimal quantization step size, number of bands, and filter shape index values.

[0123]

[0114] The use of previously selected / predicted quantization step size, number of bands, and filter shape index values ​​to optimize the encoder search for the current value selection may depend on coded information including, but not limited to, frame type, temporal layer, and / or quantization parameters.

[0124] 6A and 6B depict a number of logical steps in a particular order, steps that are not order-dependent may be reordered, and other steps may be combined or separated. Some reordering or other groupings not specifically mentioned will be apparent to those skilled in the art, and therefore the reordering and groupings presented herein are not exhaustive. Furthermore, it should be recognized that the various steps may be implemented in hardware, firmware, software, or any combination thereof.

[0125]

[0116] We note some examples.

[0126] (A1) In one aspect, some embodiments include a video decoding method (e.g., method 600). In some embodiments, the method is performed in a computing system (e.g., server system 112) having memory and control circuitry. In some embodiments, the method is performed in a coding module (e.g., coding module 320). In some embodiments, the method is performed in a parser (e.g., parser 254). The method includes: (i) obtaining, from a video bitstream, video data having a plurality of blocks including a first block; (ii) predicting a parameter set for filtering the first block, where the parameter set is not explicitly signaled in the video bitstream; (iii) obtaining a filtered first block by applying a cross-component sample offset (CCSO) to the first block, where the CCSO is based on the parameter set; and (iv) reconstructing a picture from the video data using the filtered first block. For example, a CCSO filtering process uses reconstructed samples of a first color component as input (e.g., Y, Cb, or Cr), and the output is applied to a second color component that is a different color component than the first color component. In some embodiments, the first block is a luma block or a chroma block. In some embodiments, the sets of parameters correspond to different types of filtering and / or decoding operations (e.g., non-CCSO filtering processes or other transform operations).

[0127]

[0118] In some embodiments, the method includes: (i) obtaining video data having a plurality of blocks including a first block from a video bitstream; (ii) predicting a set of parameters for filtering or transforming the first block, where the set of parameters is not explicitly signaled in the video bitstream; (iii) obtaining a filtered or transformed first block by applying a filter or transform to the first block, where the filter or transform is based on the set of parameters; and (iv) reconstructing a picture from the video data using the filtered or transformed first block.

[0128]

[0119] In some embodiments, the method includes: (i) receiving, from a video bitstream, video data having a plurality of blocks including a first block; (ii) obtaining a syntax element value from the video bitstream, the syntax element value indicating a difference between a parameter associated with the first block and a reference parameter associated with the second block, the parameter being one of a parameter set; (iii) deriving a value of a parameter associated with the first block based on the syntax element value, the parameter associated with the first block being not explicitly signaled in the video bitstream; (iv) obtaining a filtered first block by applying a cross-component sample offset (CCSO) to the first block, the CCSO being based on the parameter set; and (v) reconstructing a picture from the video data using the filtered first block.

[0129]

[0120] (A2) In some embodiments of A1, the parameter set includes the number of bands, the quantization step size, and the filter shape. For example, the band values ​​may be 0-127 or 128-255.

[0130] (A3) In some embodiments of A1 or A2, the individual values ​​of the parameter set are derived based on a reference parameter set, the reference parameter set being associated with the second block and including reference parameters. In some embodiments, the parameter set corresponds to a first picture, frame, plane, and / or slice, and the set of reference parameters corresponds to a different picture, frame, plane, and / or slice. In some embodiments, the filtered second block is obtained by applying cross-component sample offset (CCSO) to the second block, and the CCSO applied to the second block uses the set of reference parameters.

[0131] (A4) In some embodiments of A3, the method further includes obtaining, from the video bitstream, an indication of a difference between the parameter set and a reference parameter set, where each value of the parameter set is derived based on the indication of the difference. For example, differences between parameters (e.g., quantization step, number of bands, and / or filter shape values) of the current picture and a previous picture, frame, plane, slice, or block are signaled.

[0132] (A5) In some embodiments of A4, obtaining an indication of the difference from the video bitstream includes obtaining a sign of the difference and obtaining a magnitude of the difference. For example, the sign and magnitude of the delta value are signaled separately. Signaling the difference separately can improve entropy coding efficiency.

[0133] (A6) In some embodiments of any of A1-A5, the method further includes obtaining from the video bitstream an indication of individual differences between the parameter set and one or more reference parameter sets. For example, instead of signaling a quantization step size, a number of bands, and / or a filter shape index value, a delta between these indexes and a predicted / previous index is signaled. For example, a difference between one or more of the parameters (e.g., the quantization step, the number of bands, and the filter shape value) of a current and previous picture, frame, plane, slice, or block is signaled.

[0134] (A7) In some embodiments of A6, each delta indication includes an index into a lookup table. For example, a lookup table-based signaling method is used. In some embodiments, an index into the lookup table that represents the delta value is signaled.

[0135]

[0126] (A8) In some embodiments of any of A1-A7, the method further includes: (i) obtaining from the video bitstream an indication of whether the second parameter set should be predicted; (ii) refraining from predicting the second parameter set for filtering the first block in accordance with the indication that the second parameter set is not predicted; and (iii) predicting at least a subset of the second parameter set in accordance with the indication that at least a subset of the second parameter set is predicted.

[0136] For example, whether the quantization step size, number of bands, and / or filter shape index values ​​are predicted can be signaled by flags in the High Level Syntax (HLS), including, but not limited to, the adaptation parameter set (APS), slice header, frame header, picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS).

[0137] In some embodiments, the method includes: (i) obtaining an indication from the video bitstream of whether a second parameter set should be derived; (ii) refraining from deriving a second parameter set for filtering the first block in accordance with the indication indicating that the second parameter set is not to be derived; and (iii) deriving at least a subset of the second parameter set in accordance with the indication indicating that at least a subset of the second parameter set is to be derived.

[0138] (A9) In some embodiments of A8, the indication corresponds to one or more parameters of the parameter set, for example, a quantization step size, a number of bands, and / or a filter shape index value, and a flag may be signaled to indicate whether the value is predicted or explicitly signaled.

[0139] (A10) In some embodiments of A8 or A9, the method further includes obtaining the second parameter set from the video bitstream according to an indication that the second parameter set is not to be predicted (derived). For example, the non-predicted parameters are explicitly signaled in the video bitstream.

[0140] (A11) In some embodiments of any of A1-A10, at least one parameter of the parameter set is derived from predicted index values. For example, at least one of the quantization step size, the number of bands, and the filter shape index value is not signaled but is derived from the predicted index values. In some embodiments, others of the parameter set are explicitly signaled. For example, one parameter of the parameter set is predicted and another parameter of the parameter set is explicitly signaled.

[0141] (A12) In some embodiments of any of A1-A11, at least one parameter of the parameter set is predicted (derived) based on a value of at least one parameter of a previous block. For example, signaling of a quantization step size, a number of bands, and / or a filter shape index value is skipped. In that example, predicted and / or previous values ​​are used instead. For example, a value is predicted for at least one parameter based on context information (e.g., one or more aspects of the first block and / or a previous value used for at least one parameter (e.g., in a previous picture, frame, plane, slice, or block)). As another example, a previous value of at least one parameter is used (e.g., directly) for at least one parameter.

[0142] (A13) In some embodiments of any of A1-A12, the parameter set is predicted (derived) based on previously coded information. For example, the quantization step size, number of bands, and / or filter shape index value are predicted using previous picture, frame, plane, slice, or block values. In some embodiments, at least one parameter of the parameter set is predicted based on coded information such as frame type, temporal layer, and / or quantization parameter.

[0143] (B1) In another aspect, some embodiments include a video encoding method (e.g., method 650). In some embodiments, the method is implemented in a computing system (e.g., server system 112) having memory and control circuitry. In some embodiments, the method is implemented in a coding module (e.g., coding module 320). In some embodiments, the method is implemented in an entropy coder (e.g., entropy coder 214). The method includes: (i) obtaining video data having a plurality of blocks including a first block and a second block; (ii) identifying a parameter set to use by applying cross-component sample offset (CCSO) to the first block, where parameters in the parameter set are identified using a first reference parameter, the reference parameter corresponding to the second block; and (iii) signaling the identified parameter set in a video bitstream.

[0144] For example, the encoder's search for optimal quantization step, band, and filter shape values ​​for the current picture, frame, plane, slice, or block is optimized using previously selected / predicted quantization step, band, and / or filter shape values, resulting in reduced encoder / encoding complexity. In some embodiments, a parameter set to use in applying CCSO to the first block is identified, but the identified parameter set is not signaled in the video bitstream.

[0145] In some embodiments, a method includes: (i) obtaining video data having a plurality of blocks including a first block and a second block; (ii) identifying a parameter set to use in applying a transform or filter to the first block, wherein parameters in the parameter set are identified using a first reference parameter, the reference parameter corresponding to the second block; and (iii) signaling the identified parameter set in a video bitstream.

[0146] (B2) In some embodiments of B1, the first reference parameters are predicted parameters of the second block.

[0147] (B3) In some embodiments of B1 or B2, the first reference parameters are parameters selected for the second block, e.g., the first reference parameters are those used to apply the CCSO filter to the second block.

[0148]

[0135] (B4) In some embodiments of any of B1-B3, the parameter set includes a number of bands, a quantization step size, and / or a filter shape.

[0149] (B5) In some embodiments of any of B1-B4, identifying the parameters using the first reference parameters includes preceding the step of performing an optimization search for the parameters, e.g., the encoder search for at least one of the quantization step size, the number of bands, and the filter shape index value is skipped entirely, and a predicted / prior value is used instead.

[0150] (B6) In some embodiments of any of B1-B5, the method further includes identifying the first reference parameter based on previously coded information. For example, using previously selected / predicted quantization step size, number of bands, and filter shape index value to optimize the encoder search for selection of the current value depends on coded information such as frame type, temporal layer, and / or quantization parameter.

[0151]

[0138] (B7) In some embodiments of any of B1-B6, the parameters in the parameter set are identified using a reference parameter set, and the reference parameter set includes a first reference parameter.

[0152]

[0139] (B8) In some embodiments of B7, identifying the parameters includes performing an optimization search using a set of reference parameters and selecting a reference parameter from the reference parameter set that has the lowest associated cost.

[0153] The methods described herein may be used separately or combined in any order. Each method may be performed by a processing circuit (e.g., one or more processors or one or more integrated circuits). In some embodiments, the processing circuit executes a program stored on a non-transitory computer-readable medium. Although the above embodiments are described with respect to CCSO filtering, the methods described herein may be applied to other filtering techniques, such as those described above with reference to FIGS. 5A-5D.

[0154]

[0141] In another aspect, some embodiments include a computing system (e.g., server system 112) including control circuitry (e.g., control circuit 302) and memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more instruction sets configured to be executed by the control circuitry, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1-A13 and B1-B8 above).

[0155]

[0142] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more instruction sets for execution by control circuitry of a computing system, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1-A13 and B1-B8 above).

[0156]

[0143] Although terms such as "first," "second," etc. may be used herein to describe various elements, it will be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0157]

[0144] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in describing the embodiments and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or," as used herein, refers to and includes any and all possible combinations of one or more of the associated listed items. It will be further understood that, as used herein, the terms "comprises" and / or "comprising" specify the presence of referenced features, integers, steps, processes, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, processes, elements, components, and / or groups thereof.

[0158]

[0145] As used herein, the term "if" can be interpreted to mean "if" or "when" or "in response to determining that" or "in accordance with a determination that" or "in response to detecting that" a specified precondition is true, depending on the context. Similarly, the phrases "if it is determined that [the specified precondition is true]," "if [the specified precondition is true]," and "when [the specified precondition is true]" can be interpreted to mean "upon determining that" or "in response to determining that" or "by determining that" or "upon detecting a determination that" or "in response to detecting a determination that" that a specified precondition is true, depending on the context.

[0159]

[0146] The foregoing description has been described with reference to specific embodiments for explanatory purposes. However, the exemplary discussion above is not intended to be exhaustive or to limit the claims to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described to best explain the principles of operation and practical applications, thereby enabling those skilled in the art to implement them.

Claims

1. 1. A video decoding method executed on a computing system having a memory and one or more processors, the method comprising: receiving video data from a video bitstream, the video data having a plurality of blocks including a first block; obtaining syntax element values ​​from the video bitstream, the syntax element values ​​indicating differences in parameters associated with the first block and reference parameters associated with a second block, the parameters being one of a parameter set; deriving values ​​of parameters associated with the first block based on the syntax element values, the parameters associated with the first block not explicitly signaled in the video bitstream; obtaining a filtered first block by applying a cross-component sample offset (CCSO) to the first block, the CCSO being based on the parameter set; and reconstructing a picture from the video data using the filtered first blocks; A method having the following.

2. The method of claim 1 , wherein the set of parameters includes one or more of a number of bands, a quantization step size, and a filter shape.

3. 2. The method of claim 1, wherein each value of the parameter set is derived based on a reference parameter set, the reference parameter set being associated with the second block and including the reference parameters.

4. 4. The method of claim 3, further comprising obtaining from the video bitstream an indication of differences between the parameter set and the reference parameter set, and wherein individual values ​​of the parameter set are derived based on the indication of differences.

5. 5. The method of claim 4, wherein obtaining an indication of the difference from the video bitstream comprises obtaining a sign of the difference and obtaining a magnitude of the difference.

6. 10. The method of claim 1, further comprising obtaining from the video bitstream an indication of individual differences between the parameter set and one or more reference parameter sets.

7. 7. The method of claim 6, wherein the indication of each difference comprises an index into a lookup table.

8. 10. The method of claim 1, further comprising: obtaining from the video bitstream an indication of whether a second parameter set should be derived; refraining from deriving the second parameter set for filtering the first block in response to the indication indicating that the second parameter set is not to be derived; and deriving at least a subset of the second parameter set in response to the instructions indicating that at least a subset of the second parameter set is to be derived; A method comprising:

9. 9. The method of claim 8, wherein the indication corresponds to one or more parameters of the parameter set.

10. 9. The method of claim 8, further comprising: in response to the indication indicating that the second parameter set is not to be derived, obtaining the second parameter set from the video bitstream.

11. 10. The method of claim 1, wherein at least one parameter of the parameter set is derived from a predicted index value.

12. 10. The method of claim 1, wherein at least one parameter of the parameter set is derived based on a value of at least one parameter for a previous block.

13. 10. The method of claim 1, wherein the parameter set is derived based on previously coded information.

14. Control circuitry; memory; and one or more sets of instructions stored in the memory and configured to be executed by the control circuitry; 14. A computing system comprising: said one or more instruction sets causing said control circuitry to perform a method according to any one of claims 1 to 13.

15. A computer program product causing a computer of a computing device to carry out the method according to any one of claims 1 to 13.

16. 1. A video encoding method executed on a computing system having memory and control circuitry, comprising: obtaining video data having a plurality of blocks including a first block and a second block; identifying a parameter set to use by applying a cross-component sample offset (CCSO) to a first block, where parameters in the parameter set are identified using a first reference parameter, the reference parameter corresponding to a second block; and signaling the parameter sets identified in the identifying step in a video bitstream; A method comprising: