CCSO employing adaptive filtering unit size
By introducing loop filters, especially CCSO filters, into video encoding and decoding technology, the problem of large error in video data reconstruction in the prior art is solved, and more efficient video data compression and transmission are achieved.
Patent Information
- Application Number
- CN202480004843.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-09
- Filing Date
- 2024-05-20
- Publication Date
- 2025-06-20
AI Technical Summary
Existing video encoding and decoding technologies are difficult to effectively reduce reconstruction errors when compressing and transmitting video data, especially when bandwidth and storage resources are limited.
Loop filters, especially cross-component sample offset (CCSO) filters, reduce reconstruction errors by adjusting the reconstructed picture samples. The filter control parameters are streamed together with the image frame and are used to process the filter blocks of the current image frame.
Through loop filtering technology, the bit error of video data in transmission and storage is significantly reduced, the video quality is improved, and the bandwidth and storage resource utilization are optimized.
Smart Images

Figure CN120188477A_ABST
Abstract
Description
Incorporation by reference
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 544,409, entitled "CCSO with Adaptive Filter Unit Size", filed on October 16, 2023, and is a continuation of, and claims priority to, U.S. Patent Application No. 18 / 660,064, entitled "CCSO with Adaptive Filter Unit Size", filed on May 9, 2024, both of which are incorporated herein by reference in their entireties. Technical field
[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for loop filtering (e.g., cross-component offset filtering) of video data. Background art
[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video conferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data over a communication network and / or store digital video data on a storage device. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can be used to compress video data according to one or more video coding standards before the video data is transmitted or stored. Video coding can be performed by software and / or hardware on an electronic / client device or a server providing cloud services.
[0004] Video encoding typically utilizes prediction methods that exploit the redundancy inherent in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video encoding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing a degradation in video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended as a successor to HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the specification verification version 1.0.0 containing errata 1 was released. Summary of the Invention
[0005] As described above, encoding (compression) reduces the bandwidth and / or storage space requirements. As will be described in detail later, both lossless compression and lossy compression can be employed. Lossless compression refers to a technique where an exact copy of the original signal can be reconstructed from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process where the original video information is not fully retained during encoding and is not fully recovered during decoding. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal is useful for the intended application. The degree of tolerable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect the various distortion tolerances: generally, higher tolerable distortion allows the use of encoding algorithms that introduce higher losses but have a higher compression ratio.
[0006] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying loop filters to video (image) compression. A video codec includes multiple functional modules for one or more of the following: intra / inter prediction, transform coding, quantization, entropy coding, and in-loop filtering. In-loop filtering techniques are applied to adjust the reconstructed picture samples to further reduce the reconstruction error. In various embodiments of the present application, filter control parameters are streamed together with the current image frame for the loop filter to process a first filtering block of the current image frame. When the filter control parameters are received, the video decoder can identify the first filtering block, determine that the filter control parameters are enabled, and apply the loop filter to process the first filtering block.
[0007] In some embodiments, the loop filter includes a cross-component sample offset (CCSO) filter that uses co-located reconstructed samples from a first color component and its neighboring reconstructed samples (e.g., luminance samples) to derive a sample offset value to be added to the current sample (e.g., luminance or chrominance sample) of a second color component. In some embodiments, the CCSO filter may include an edge-preserving loop filter that depends on the values of the reconstructed samples to determine the sample offset values for luminance samples and / or chrominance samples. For each color component, corresponding filter control parameters can be written at the frame level for the CCSO filter, and a filter on / off control flag can also be written at the filtering unit level to control the application of the corresponding control parameters to each individual filtering unit. In an example, the filtering unit size is fixed at 256×256 luminance samples, which corresponds to 128×128 chrominance samples. Additionally, an adaptive filtering unit size can be used to achieve flexibility in cross-component sample offset.
[0008] According to some embodiments, a video decoding method is provided. The method includes: receiving a video bitstream that includes a current image frame and first filter control parameters for a loop filter to process a first filtering block of the current image frame; determining a filtering unit size for processing the current image frame by the loop filter, the first filtering block having the filtering unit size; identifying the first filtering block in the current image frame based on the filtering unit size; when the first filter control parameters are enabled, applying the loop filter to process one or more samples of the first filtering block; and reconstructing the current image frame including the first filtering block.
[0009] According to some embodiments, a video encoding method is provided. The method includes: receiving video data including a current picture frame; encoding the current picture frame; transmitting the encoded current picture frame via a video bitstream; and writing, via the video bitstream, first filter control parameters for a first filter block for loop filter processing of the current picture frame. A filter unit size is determined for processing the current picture frame by the loop filter, and the filter unit size is applied to identify the first filter block processed by the loop filter based on the first filter control parameters.
[0010] According to some embodiments, a bitstream conversion method is provided. The method includes: obtaining a source video sequence including a current picture frame; and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes: the current picture frame; and first filter control parameters for a first filter block for loop filter processing of the current picture frame. A filter unit size is determined for processing the current picture frame by the loop filter, and the filter unit size is applied to identify the first filter block processed by the loop filter based on the first filter control parameters.
[0011] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes a control circuit and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.
[0012] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions for execution by a computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.
[0013] Thus, apparatuses and systems having methods for video encoding and decoding are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding and decoding.
[0014] The features and advantages described in the specification are not necessarily all included, and in particular, considering the drawings, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in the specification is mainly selected for readability and guidance purposes and is not necessarily selected to depict or delimit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To better understand the present disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are shown in the accompanying drawings. However, the drawings only show the relevant features of the present disclosure and should not be considered restrictive, as those skilled in the art will understand after reading the present disclosure that the description may also have other valid features.
[0016] Figure 1 is a block diagram showing an example communication system according to some embodiments.
[0017] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments; and
[0018] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.
[0019] Figure 3 is a block diagram showing an example server system according to some embodiments.
[0020] Figure 4 is a flowchart of an example process 400 for decoding a video bitstream using in-loop filtering based on a first filter control parameter 408 according to some embodiments.
[0021] Figure 5 is a flowchart of an example process for applying cross-component sample offset in in-loop filtering according to some embodiments.
[0022] Figure 6 is a flowchart showing a method for decoding a video according to some embodiments.
[0023] By convention, the various features shown in the drawings are not necessarily drawn to scale, and the same reference numerals may be used throughout the specification and drawings to denote the same features. Detailed Description
[0024] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying a loop filter to video (image) compression. Intra-loop filtering techniques are applied to adjust the reconstructed picture samples to further reduce the reconstruction error. In various embodiments of the present application, filter control parameters are streamed to a video decoder along with the current image frame for loop filter processing of a first filter block of the current image frame. Upon receiving the filter control parameters, the video decoder can identify the first filter block, determine that the filter control parameters are enabled, and apply the loop filter to process the first filter block. In some embodiments, the loop filter includes a CCSO filter that uses reconstructed samples of a first color component (e.g., luminance samples) to derive a sample offset value to be added to samples of a second color component (e.g., luminance or chrominance samples). The CCSO filter can include an edge-preserving loop filter that depends on the values of the reconstructed samples to determine the sample offset values for the luminance samples and / or chrominance samples. For each color component, corresponding filter control parameters can be written at the frame level for the CCSO filter, and a filter on / off control flag can also be written at the filter unit level to control the application of the corresponding control parameters to each individual filter unit.
[0025] Implement a cross-component offset filtering method to apply co-located reconstructed color components of a first sample and associated adjacent reconstructed samples to derive an offset value to be added to a current sample of a second color component, thereby adjusting the reconstructed value of the current sample. In various embodiments of the present application, the decoder receives a video bitstream from an encoder, the video bitstream including a current image frame and first filter control parameters for a loop filter, and the first color sample is determined based on the values of one or more luminance samples (e.g., independent of any associated luminance gradient of the one or more luminance samples). The sample values of the first color component (e.g., unassociated gradient values) are used in the offset filtering to determine the offset value to be added to the samples of the second color component.
[0026] More specifically, in some embodiments, the video decoder identifies a set of luminance samples that includes a first luminance sample and one or more adjacent luminance samples of the first luminance sample. The luminance samples (e.g., using a scalar quantizer) are quantized to generate one or more quantization values. The scalar quantizer can be specified by a quantization interval (e.g., a range of values assigned to the same integer) and a quantization level (e.g., an integer value assigned to the quantization interval). The first color sample is classified based on the one or more quantization values (e.g., by a classifier) to determine a first sample offset of the first color sample. The first color sample is adjusted based on the first sample offset of the first color sample such that the current image frame can be reconstructed.
[0027] Figure 1is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m), which are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, e.g., for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0028] The source device 102 includes a video source 104 (e.g., a camera assembly or a media memory) and an encoder assembly 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder assembly 106 generates one or more encoded video bitstreams based on the video stream. The video stream from the video source 104 has a higher data volume compared to the encoded video bitstream 108 generated by the encoder assembly 106. Since the encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth for transmission and less storage space for storage. In some embodiments, the source device 102 does not include the encoder assembly 106 (e.g., configured to send uncompressed video to one or more networks 110).
[0029] One or more networks 110 represent any number of networks that convey information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0030] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as an encoded video stream from a source device 102). The server system 112 includes a codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec component 114 includes an encoder component and / or a decoder component. In various embodiments, the codec component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the codec component 114 is configured to decode an encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or multiple video encodings based on the encoded video bitstream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0031] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be presented on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0032] The source device and / or the multiple electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0033] In an example operation of communication system 100, source device 102 sends encoded video stream 108 to server system 112. For example, source device 102 may encode a picture stream captured by the source device. Server system 112 receives the encoded video stream 108 and may decode and / or encode the encoded video stream 108 using codec component 114. For example, server system 112 may apply an encoding that is more suitable for network transmission and / or better for storage to the video data. Server system 112 may send the encoded video data 116 (e.g., one or more encoded video streams) to one or more of electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display video pictures.
[0034] Figure 2A is a block diagram showing example elements of encoder component 106 according to some embodiments. Encoder component 106 receives video data (e.g., a source video sequence) from video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from encoder component 106). Video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that, when viewed in sequence, are given motion. The pictures themselves may be constructed as spatial pixel arrays, depending on the sampling structure, color space, etc. used, where each pixel may include one or more samples. Those of ordinary skill in the art can readily understand the relationship between pixels and samples.
[0035] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder component 106 is configured to perform a conversion between a source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may be related to the encoder component 106 optimized for a specific system design.
[0036] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols (e.g., a symbol stream) based on an input picture to be encoded and one or more reference pictures) and a (local) decoder 210. (When the compression between the symbols and the encoded video bitstream is lossless) the decoder 210 reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data. The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact corresponding between the local encoder and the remote encoder. In this way, the reference picture samples interpreted by the prediction part of the encoder are the same as the sample values interpreted by the decoder when using prediction during decoding. This principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well known to those of ordinary skill in the art.
[0037] The operation of the decoder 210 may be the same as that of a remote decoder such as the decoder component 122, which will be described in detail below in conjunction with Figure 2B However, briefly referring to Figure 2B , when the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols into the encoded video sequence losslessly, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.
[0038] Except for parsing / entropy decoding, the decoder techniques described herein may exist in corresponding encoders in substantially the same functional form. Thus, the disclosed subject matter focuses on decoder operations. Since encoder techniques are inverse to decoder techniques, the description of encoder techniques may be simplified.
[0039] As part of its operation, the source encoder 202 may perform motion-compensated predictive coding, referring to one or more previously encoded frames designated as reference frames in a video sequence, and the motion-compensated predictive coding performs predictive coding on an input frame. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the (one or more) reference frames, which may be selected as the (one or more) predictive references for the input frame. The controller 204 may manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.
[0040] The decoder 210 decodes the encoded video data of a frame that can be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 may advantageously be a lossy process. When the encoded video data is decoded in a video decoder ( Figure 2A not shown), the reconstructed video sequence may be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that can be performed by a remote video decoder on a reference frame and may cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame, which has the same content (in the absence of transmission errors) as the reconstructed reference frame that will be obtained by the remote video decoder.
[0041] The predictor 206 may perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor 206 may operate block by block based on sample blocks to find an appropriate prediction reference. As the search result obtained by the predictor 206, it can be determined that the input picture may have a prediction reference extracted from multiple reference pictures stored in the reference picture memory 208.
[0042] The outputs of all the above functional units may be entropy encoded in the entropy encoder 214. The entropy encoder 214 converts the symbols into an encoded video sequence by performing lossless compression on the symbols generated by various functional units according to techniques well known to those of ordinary skill in the art, such as Huffman coding, variable length coding, and / or arithmetic coding.
[0043] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the (one or more) encoded video sequences created by the entropy encoder 214 in preparation for transmission over a communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such other data being, for example, encoded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may transmit additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and redundant slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set segments, and the like.
[0044] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture may be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are aware of those variations of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. A predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. A bi-predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may be used to reconstruct a single block using more than two reference pictures and associated metadata.
[0045] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks determined by the coding assignments applied to the corresponding pictures of the blocks. For example, blocks of an I picture can be non-predictively encoded, or the block can be predictively encoded with reference to the already encoded block pairs of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0046] Video can be acquired as multiple source pictures (video pictures) in a time sequence. Intra picture prediction (usually abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded (referred to as the current picture) is partitioned into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0047] The encoder component 106 can perform encoding operations according to any predetermined video coding technique or standard such as those described herein. The encoder component 106 can perform various compression operations in its operation, including predictive coding operations that exploit the temporal redundancy and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.
[0048] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to the loop filter 256 and is configured to transmit data to the display 124 (e.g., via a wired or wireless connection).
[0049] In some embodiments, decoder component 122 includes a receiver that is coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data while receiving other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data when receiving the encoded video. The additional data may be part of the (one or more) encoded video sequences. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0050] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. Decoder component 122 may be implemented at least partially in software.
[0051] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 inside decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided outside decoder component 122 (e.g., to prevent network jitter). When receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may also be possible not to configure buffer memory 252, or buffer memory 252 may be made smaller. For use on a service packet network such as the Internet, buffer memory 252 may also be required, which may be relatively large and / or have an adaptive size and may be implemented at least partially in an operating system or a similar element outside decoder component 122.
[0052] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols can include, for example, information for managing the operation of the decoder components 122 and / or information for controlling a rendering device such as the display 124. The control information for the rendering device can be in the form of, for example, Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser 254 can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup can include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser 254 can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and the like.
[0053] Depending on the type of the encoded video picture or a portion of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols 270 may involve multiple different units. Which units are involved and the manner of involvement can be controlled by the subgroup control information parsed by the parser 254 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 254 and the multiple units below are not described.
[0054] The decoder components 122 can be conceptually subdivided into multiple functional units, and in some implementations, these units interact closely with each other and can be at least partially integrated with each other. However, for the sake of clarity, the conceptually subdivided functional units are retained here.
[0055] The scaler / inverse transform unit 258 receives the quantized transform coefficients as symbols 270 and control information (such as which transform mode to use, block size, quantization factor, and / or quantization scaling matrix, etc.) from the parser 254. The scaler / inverse transform unit 258 can output blocks including sample values, and the sample values can be input into the aggregator 268.
[0056] In some cases, the output samples of the scaler / inverse transform unit 258 belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but can use predictive information from a previously reconstructed portion of the current picture. Such predictive information can be provided by the intra picture prediction unit 262. The intra picture prediction unit 262 can generate a block of the same size and shape as the block being reconstructed by using the surrounding reconstructed information extracted from the current (partially reconstructed) picture in the current picture memory 264. The aggregator 268 can add the predictive information generated by the intra picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.
[0057] In other cases, the output samples of the scaler / inverse transform unit 258 can belong to an inter-coded and potentially motion-compensated block. In this case, the motion compensation prediction unit 260 can access the reference picture memory 266 to extract samples for prediction. After motion-compensating the extracted samples according to the symbol 270 associated with the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signal in this case), thereby generating output sample information. The address within the reference picture memory 266 from which the motion compensation prediction unit 260 extracts the prediction samples can be controlled by a motion vector. The motion vector can be in the form of a symbol 270 for use by the motion compensation prediction unit 260, which can have, for example, an X component, a Y component, and a reference picture component. Motion compensation can also include interpolation of sample values extracted from the reference picture memory 266 when using sub-sampled accurate motion vectors, a motion vector prediction mechanism, and so on.
[0058] The output samples of the aggregator 268 can be employed by various loop filtering techniques in the loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and that are available to the loop filter unit 256 as the symbol 270 from the parser 254. However, video compression techniques can also respond to meta-information obtained during the decoding of a previously (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 can be a sample stream that can be output to a rendering device such as the display 124 and can be stored in the reference picture memory 266 for subsequent inter-picture prediction.
[0059] Once reconstructed, some of the encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is reconstructed and the encoded picture is identified as a reference picture (e.g., by parser 254), the current reference picture can become part of the reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0060] The decoder component 122 can perform decoding operations according to, for example, a predetermined video compression technique documented in any of the standards described herein. The encoded video sequence can conform to the syntax specified by the video compression technique or standard used in the sense that the encoded video sequence follows the video compression technique or standard (especially the profile thereof) specified in the video compression technique document or standard. In addition, to conform to some video compression techniques or standards, the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (e.g., measured in megasamples per second), the maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management written into the encoded video sequence.
[0061] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one or more field-programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).
[0062] One or more network interfaces 304 may be configured to interact with one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication network may be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, and so on. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, and so on. Such communication may be unidirectional receive-only (e.g., broadcast television), unidirectional transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., using a local area digital network or a wide area digital network to connect to other computer systems). Such communication may include communication to one or more cloud computing networks.
[0063] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. One or more input devices 310 may include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. One or more output devices 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or viewers), etc.
[0064] The memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices located remotely from the control circuit 302. The memory 314, or alternatively the one or more non-volatile solid-state memory devices within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures or subsets or supersets thereof: · An operating system 316, which includes programs for handling various basic system services and for performing hardware-related tasks; · A network communication module 318, which is used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via a wired connection and / or a wireless connection); · A codec module 320 that is configured to perform various functions related to encoding and / or decoding data (e.g., video data). In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: ο A decoding module 322 that is configured to perform various functions related to decoding encoded data, such as those functions previously described with respect to the decoder component 122; and ο An encoding module 340 that is configured to perform various functions related to encoding data, such as those functions previously described with respect to the encoder component 106; and · A picture memory 352 that is configured to store pictures and picture data, e.g., for use with the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.
[0065] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to the parser 254), a transform module 326 (e.g., configured to perform various functions previously described with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter 256).
[0066] In some embodiments, the encoding module 340 includes an encoder module 342 (e.g., configured to perform various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the modules shown. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.
[0067] Each of the modules identified above stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The modules identified above (e.g., the instruction sets) need not be implemented as separate stand-alone software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but rather uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the modules and data structures identified above. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.
[0068] Although Figure 3 FIG. shows a server system 112 according to some embodiments, but Figure 3 is more intended as a functional description of the various features that may exist in one or more server systems rather than a structural schematic of the embodiments described herein. In practice, and as will be recognized by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in FIG. may be implemented on a single server, and a single item may also be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among them will vary depending on the implementation, and optionally, in part, on the data traffic processed by the server system during peak usage periods as well as during average usage periods.
[0069] Figure 4 is a flowchart of an example process 400 for decoding a video bitstream using in-loop filtering based on a first filter control parameter 408 according to some embodiments. The GOP includes a sequence of picture frames, which further includes a current picture frame 406. The current picture frame 406 includes a color picture, i.e., a non-monochrome picture frame, which has a plurality of co-located color samples (e.g., chrominance samples 402 and luminance samples 404). After reconstructing the plurality of color samples of the current picture frame 406, in-loop filtering is applied to adjust a subset of the color samples, thereby improving the picture quality of the current picture frame 406. An example of in-loop filtering is CCSO filtering. In some embodiments associated with CCSO filtering, the reconstructed samples of a first color component and their adjacent reconstructed samples are combined to derive an offset value for a second color component, and the reconstructed samples of the second color samples are co-located with the reconstructed samples of the first color component and adjusted by the offset value. The first color component is optionally the same as or different from the second color component.
[0070] In some embodiments, decoder 122 receives video bitstream 116, which includes current picture frame 406 and first filter control parameter 408 for first filter block 412 of loop filter 410 to process current picture frame 406. Decoder 122 determines filter unit size 414 for processing current picture frame 406 by loop filter 410. First filter block 412 has filter unit size 414. First filter block 412 is identified in current picture frame 406 based on filter unit size 414. In some embodiments, decoder 122 determines whether first filter control parameter 408 associated with first filter block 412 is enabled. When the first filter control parameter is enabled, decoder 122 applies loop filter 410 to process one or more samples of first filter block 412. Current picture frame 406 including first filter block 412 is reconstructed to generate reconstructed picture frame 416. In some embodiments, first filter block 412 is defined at the frame level and applied to multiple filter blocks including first filter block 412, and a block-level flag is used to indicate and determine whether first filter control parameter 408 is applied to first filter block 412.
[0071] In some embodiments, video bitstream 116 includes first syntax element 418 indicating filter unit size 414, which is used for loop filter 410 to process current picture frame 406. First syntax element 418 can be signaled at one of the following levels for the first color component (e.g., luminance component 404) among multiple color components 420: tile header of the current picture frame, sequence level, and frame level. Filter unit size 414 can be determined based on first syntax element 418 signaled through video bitstream 116. Additionally, in some embodiments, video bitstream 116 includes multiple syntax elements 430, which also include first syntax element 418, and the multiple syntax elements 430 can be signaled separately for multiple color components 420.
[0072] In some embodiments, the filter unit size 414 includes a first filter unit size, and a second color component (e.g., the chrominance component 402) among the plurality of color components is not written and has a second filter unit size. The second filter unit size may have a predefined relationship with the first filter unit size and is configured to be determined based on the first filter unit size. Additionally, in some embodiments, the plurality of color components 420 includes a luminance component 404 and chrominance components 402 (e.g., red-difference chrominance (Cr) component, blue-difference chrominance (Cb) component). The chrominance format of the current image frame 406 is 4:2:0. The second filter unit size is twice or half of the first filter unit size, depending on whether the first filter unit size corresponds to the chrominance component 402 or the luminance component 404. For example, based on a first syntax element 418 that defines the filter unit size 414, the first filter unit size of the luminance component 404 is determined to include 256×256 luminance samples. Based on a predefined relationship with the first filter unit size of the luminance samples 404, the second filter unit size of the chrominance samples 402 is determined to be 128×128 chrominance samples. In each dimension (e.g., width, length) of the filter unit, the second filter unit size of the chrominance samples 402 is half of the first filter unit size of the luminance samples 404. In another example, the first filter unit size of the luminance component 404 is 128×128 luminance samples, and the second filter unit size of the chrominance component 402 is 64×64 chrominance samples.
[0073] In some embodiments, the current image frame 406 has a superblock size that represents the maximum coding unit size applied to encode the current image frame 406, and the filter unit size 414 is equal to or greater than the superblock size. Additionally, in some embodiments, the superblock size is 128×128 pixels. The first filter unit size of the luminance component 404 of the current image frame 406 is 128×128 luminance samples, and the second filter unit size of the chrominance component 402 of the current image frame 406 is 64×64 chrominance samples corresponding to 128×128 pixels. In some embodiments, the video bitstream 116 includes a first syntax element 418 that indicates the filter unit size 414 for processing the current image frame 406 by the loop filter 410, and the first syntax element 418 is written based on the superblock size. In some embodiments, the video bitstream 116 includes a first syntax element 418 that indicates the filter unit size 414 for processing the current image frame 406 by the loop filter 410, and the first syntax element 418 includes a binary integer that is equal to the difference between the binary logarithm of the filter unit size 414 and the binary logarithm of the superblock size. For example, the superblock size is 128×128 pixels represented by a first binary logarithm value of 14, and the filter unit size 414 is 256×256 represented by a second binary logarithm value of 16. The first syntax element 418 is equal to 2 and is represented as "010".
[0074] In some embodiments, the first filter block 412 is divided into a plurality of sub-filter units 422 according to a tree structure, a control flag associated with the first filter control parameter 408 is written at the sub-filter unit level, and the control flag corresponds to the first sub-filter unit among the sub-filter units 422 of the first filter block 412. The first filter control parameter 408 is written at the frame level of the current image frame 406, and the filter unit size 414 corresponds to a fixed maximum filter unit size. Inside each filter unit (e.g., the first filter block 412), the filter unit is further (e.g., recursively) divided into sub-filter units 422 using a tree structure. In an example, the first filter control parameter 408 is associated with CCSO filtering. For each sub-filter unit 422, a CCSO on / off flag or other control parameters are written at the sub-filter unit level.
[0075] In some embodiments, the video bitstream 116 includes a first syntax element 418 that indicates a filter unit size 414 for processing the current picture frame by the loop filter 410. The first syntax element 418 is written at the picture level of the current picture frame 406. The filter unit size 414 includes a fixed filter unit size. The video bitstream 116 also includes a second syntax element 428 that indicates a block-level filter unit size 424 for a set of one or more superblocks of the current picture frame 406. The second syntax element 428 includes an index that defines a scaling factor between the fixed filter unit size 414 and the block-level filter unit size 424. Additionally, in some embodiments, the index is equal to one of 0, 1, 2 and corresponds respectively to a scaling factor equal to one of 1, 4, and 16, and the block-level filter unit size is equal to 1 times, 4 times, or 16 times the fixed filter unit size.
[0076] In some embodiments, the video bitstream 116 also includes a pointer index 426 associated with a second filter block 432 of the current picture frame 406, and the pointer index 426 identifies a first filter block 412 and controls the loop filter 410 to process the second filter block 432 based on a first filter control parameter 408 associated with the first filter block 412.
[0077] In some embodiments, the filter unit size 414 of the current picture frame 406 is determined based on encoded information 434. The encoded information 434 includes one or more of the following: quantization parameter, image resolution, tile size, whether the current picture frame is an intra-only picture, temporal layer, and whether the current picture frame is an interpolated picture. Additionally, in some embodiments, the current picture frame 406 is interpolated from two additional picture frames in a GOP. In one example, the filter unit size 414 is not written and is determined based on the encoded information 434. In another example, the filter unit size 414 is written in the first syntax element 418. The first syntax element 418 includes an index that selects the filter unit size 414 from a plurality of filter unit size options determined based on the encoded information 434.
[0078] In some embodiments, the filter unit size 414 includes a first filter unit size. The alternative picture frame corresponds to a second filter unit size for processing the alternative picture frame by the loop filter 410. The resolution of the current picture frame is greater than the resolution of the alternative picture frame, and the first filter unit size is greater than the second filter unit size.
[0079] In some embodiments, the loop filter 410 includes a CCSO mode 436, and the first filter control parameter 408 indicates whether the CCSO mode 436 is enabled to apply one or more samples of a first color component (e.g., the luminance samples 404) to determine a sample offset of samples of a second color component (e.g., the luminance samples 404, the chrominance samples 402). Additionally, in some embodiments, the loop filter 410 further includes one or more additional filtering operations 438 (e.g., Wiener filtering) different from the CCSO mode 436, and the filter unit size 414 is jointly written in the first syntax element 418 for the CCSO mode 436 and the one or more additional filtering operations 438. The one or more additional filtering operations may include at least one of Wiener filtering and cross-component Wiener filtering.
[0080] Figure 5 FIG. 500 is a flowchart of an example process for applying cross-component sample offsets in in-loop filtering according to some embodiments. The decoder 122 receives a video bitstream 116 that includes a current picture frame 406 and a first filter control parameter 408 for the loop filter 410 to process a first filtered block 412 of the current picture frame 406. The decoder 122 determines a filter unit size 414 for processing the current picture frame 406 by the loop filter 410. The first filtered block 412 has the filter unit size 414. The first filtered block 412 is identified in the current picture frame 406 based on the filter unit size 414. When the first filter control parameter is enabled, the decoder 122 applies the loop filter 410 to process one or more samples of the first filtered block 412. The current picture frame 406 including the first filtered block 412 is reconstructed to generate a reconstructed picture frame 416. In some embodiments, the loop filter 410 includes a CCSO mode 436, and the first filter control parameter 408 indicates whether the CCSO mode 436 is enabled to apply one or more samples of a first color component (e.g., the luminance samples 404) to determine a sample offset of samples of a second color component (e.g., the luminance samples 404, the chrominance samples 402), e.g., at the picture level of the current picture frame 406.
[0081] In some embodiments, a set of one or more luminance samples 404 is identified among one or more samples of a first color component. The set of one or more luminance samples 404 includes one or more luminance samples among a first luminance sample 404C and one or more adjacent luminance samples 404X. The first luminance sample 404C is co-located with a first color sample 510 of a second color component to be determined in the first filter block 412. A first sample offset 506 of the first color sample 510 is determined based on the set of one or more luminance samples 404. The first color sample 510 is adjusted based on the first sample offset 506 of the first color sample 510.
[0082] In some embodiments associated with edge offset, one or more differences between one or more adjacent luminance samples 404X and a first luminance sample 404C that is co-located with a first color sample 510 are determined. The one or more differences are quantized to generate one or more quantization values 404Q. The first color sample 510 is classified based on the one or more quantization values 404Q to determine a first sample offset 506 of the first color sample 510. Alternatively, in some embodiments associated with band offset, a set of luminance samples 404 (e.g., a first luminance sample 404C) is quantized to generate one or more quantization values 404Q. The luminance values (not gradient values or different values) of the set of luminance samples 404 are quantized. The first color sample 510 is classified based on the one or more quantization values 404Q to determine a first sample offset 506 of the first color sample 510. The decoder 122 reconstructs the current image frame 406 at least by adjusting the first color sample 510 based on the first sample offset 506.
[0083] In some embodiments, the first color sample 510 is one of the following samples: a first luminance sample 404C, a first chrominance red difference (Cr) sample 402Cr, and a first chrominance blue difference (Cb) sample 402Cb. The first luminance sample 404C, the first Cb sample 402Cb, and the first Cr sample 402Cr are co-located with each other.
[0084] For example, the first color sample 510 is classified based on the quantization values 404Q by a classifier 512 to determine a first sample offset 506 of the first color sample 510. In the example, the quantization values 404Q include quantization values 404QC, 404QN, 404QS, 404QW, and 404QE. A look-up table 514 maps multiple combinations of the quantization values 404QC, 404QN, 404QS, 404QW, and 404QE to different sample offset options SO (e.g., SO1 to SQ16). Based on the look-up table 514, the quantization values 404Q correspond to one of the combinations in the look-up table 514, and a corresponding sample offset option SO associated with the combination of the quantization values 404Q is identified and thus selected for the first sample offset 506. In other words, in some embodiments, the decoder 122 classifies the first color sample 510 by: identifying a combination of one or more of the quantization values 404Q in the look-up table 514 (the look-up table associates multiple quantization combinations with multiple offset value options SO (e.g., SO1 to SO16)); and determining the first sample offset 506 corresponding to the combination of one or more of the quantization values 404Q in the look-up table 514.
[0085] In some embodiments, a scalar quantizer 530 including a plurality of quantization intervals 518 (QI) and a plurality of quantization levels 520 (QL) quantizes the value (not the associated difference or gradient) of (one or more) luminance samples 404A to a plurality of integer values within a quantization range 516, and each of the one or more quantization values 404Q includes a corresponding integer within the quantization range 516. For each integer value within the quantization range 516, a quantization interval 518 is defined as the range of values assigned to the corresponding integer value. The quantization level 520 corresponds to the corresponding integer value to which a difference range associated with the quantization interval 518 is assigned.
[0086] In some embodiments, in the offset-only mode 440, only the first luminance sample 404C is quantized, e.g., by the quantizer 530, and classified to determine a first sample offset 506 of the first color sample 510. The first luminance sample 404C is determined to be associated with one of a plurality of bands. Each of the plurality of bands corresponds to a respective sample offset value. The first sample offset 506 is determined to be equal to the respective sample offset value corresponding to one of the plurality of bands.
[0087] The first color sample 510 is adjusted based on the first sample offset 506 of the first color sample 510, enabling reconstruction of the current image frame. In some embodiments, the first color sample 510 includes a first chrominance sample 402C co-located with the first luminance sample 404C in the current image frame, and the first chrominance sample 402C is adjusted based on the first sample offset 506. Alternatively, in some embodiments, the first color sample 510 is the first luminance sample 404C, and the first luminance sample 404C is adjusted based on the first sample offset 506.
[0088] Figure 6 is a flowchart showing an example method 600 of decoding a video. Method 600 may be implemented in a computing system having a control circuit and a memory (e.g., Figure 1executed in a server system 112, a source device 102, or an electronic device 120), and the memory stores instructions executed by the control circuit. In some embodiments, method 600 is applied in conjunction with one or more video codecs, including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. The filter unit size 414 is adaptively selected and written into the bitstream 116. The first filter control parameter 408 may be associated with a CCSO on / off decision flag, an alternative loop filter on / off decision, or other loop filter control parameters. The first filter control parameter 408 is written and applied according to the filter unit size 414. In some embodiments, the first filter unit size 414 is written separately at the frame level for each color component (e.g., the luminance component, the Cr component, the Cb component). In some embodiments, the filter unit size 414 is written at the sequence level or the tile level.
[0089] In some embodiments, the first filter unit size 414 is written using an advanced syntax (e.g., at the frame level, the sequence level, or the tile level), and the first filter unit size 414 is shared between two or more color components (e.g., the Cr and Cb components, the luminance and chrominance components) according to the color format of the current image frame 406. For example, for video data with a chroma format of 4:2:0, the first filter unit size 414 is written for one of the luminance components, and the filter unit size 414 for the chroma samples 402 is half of the written filter unit size 414. In other words, the written filter unit size 414 associated with the luminance component is twice the filter unit size of the chroma samples.
[0090] In some embodiments, the written filter unit size 414 cannot be smaller than the superblock size. In other words, the written filter unit size 414 is equal to or greater than the superblock size, i.e., the lower limit is the superblock size. For example, the superblock size is 128×128 pixels. The first filter unit size for the luminance component 404 of the current image frame 406 is 128×128 luminance samples, and the second filter unit size for the chrominance component 402 of the current image frame 406 is 64×64 chroma samples corresponding to 128×128 pixels. In some embodiments, the writing of the filter unit size 414 varies according to the superblock size. For example, the superblock size is 128×128 pixels, and the filter unit size 414 for the chroma samples 402 is 64×64. The filter unit size 414 is not written and is defaultly set to the superblock size. In another example,
[0091] In some embodiments, the difference between the write filter unit size 414 and the superblock size is determined, and this difference can be expressed as the difference between the base-2 logarithm of the filter unit size 414 and the base-2 logarithm of the superblock size. For example, the superblock size is 128×128 pixels represented by a first binary logarithm value 14, and the filter unit size 414 is 256×256 represented by a second binary logarithm value 16. The difference is equal to 2 and is represented as "010" by the first syntax element 418.
[0092] In some embodiments, the filter unit size 414 corresponds to a fixed maximum filter unit size. Inside each filter unit (e.g., the first filter block 412), the filter unit is further (e.g., recursively) divided into sub-filter units 422 using a tree structure. In an example, the first filter control parameter 408 is associated with CCSO filtering. For each sub-filter unit 422, a CCSO on / off flag or other control parameter is written at the level of the sub-filter unit. In some embodiments, the filter unit size 414 is written at the frame level. For a group of superblocks (e.g., a single superblock, a row of superblocks, a column of superblocks), a scaling factor between the frame-level filter unit size 414 and the block-level filter unit size 424 is written. For example, an index is written, and this index can be equal to one of 0, 1, 2, which respectively correspond to a scaling factor equal to one of 1, 4, and 16. The block-level filter unit size 424 is equal to 1 times, 4 times, or 16 times the frame-level filter unit size 414.
[0093] In some embodiments, e.g., at the frame level, the fixed or recursive filter unit size 414 is written using the first syntax element 418. A CCSO setting on / off flag or other control parameter is written at the filter unit level or the sub-filter unit level. These flags can be merged and shared between filter units (also referred to as filter blocks) to reduce signaling overhead. For example, the first filter block can be written with a flag / index to indicate where the control parameter is inherited from. In some embodiments, the filter unit size 414 can be written jointly for multiple loop filters (e.g., the CCSO filter 436 and the additional loop filter 438). In an example, the filter unit size 414 is written jointly for CCSO filtering 436 and Wiener filtering (e.g., cross-component Wiener filtering).
[0094] In some embodiments, the filter unit size options for multiple applications are not fixed and can be determined based on the encoded information 434, which includes but is not limited to one or more of the following: quantization parameter, picture resolution, tile size, whether it is an intra-only frame, the temporal layer of the current frame, and whether it is an interpolated frame. In an example, for video frames with different resolutions, the filter unit size options applied can be different. An image frame with a higher resolution may correspond to a larger filter unit size option 414. In another example, an interpolated frame refers to a frame reconstructed by interpolating between two encoded frames.
[0095] Although Figure 6 Multiple logical stages are shown in a particular order, but logical stages that are not order-dependent can be reordered, and other stages can be combined or split. For those of ordinary skill in the art, some reorderings or other groupings not specifically mentioned will be obvious, and thus the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.
[0096] Now turning to some example embodiments.
[0097] (A1) In some implementations, a method 600 for decoding video data is implemented. The method 600 includes: receiving (operation 602) a video bitstream that includes a current image frame and first filter control parameters for a first filter block for loop filter processing of the current image frame; determining (operation 604) a filter unit size for processing the current image frame by the loop filter, the first filter block having the filter unit size; identifying (operation 606) the first filter block in the current image frame based on the filter unit size; when the first filter control parameters are enabled, applying (operation 610) the loop filter to process one or more samples of the first filter block; and reconstructing (operation 612) the current image frame that includes the first filter block. In some embodiments, the method 600 further includes determining (operation 608) whether the first filter control parameters associated with the first filter block are enabled.
[0098] (A2) In some embodiments of A1, the video bitstream includes a first syntax element that indicates a filter unit size for processing the current image frame by the loop filter. For the first color component among multiple color components, the first syntax element is written at one of the following levels: the tile header of the current image frame, the sequence level, and the frame level. The filter unit size is determined based on the first syntax element written through the video bitstream.
[0099] (A3) In some embodiments of A2, the video bitstream includes a plurality of syntax elements, the plurality of syntax elements further includes the first syntax element, and the plurality of syntax elements are written for the plurality of color components respectively.
[0100] (A4) In some embodiments of A2 or A3, the filter unit size includes a first filter unit size, and the second color component among the plurality of color components is not written and has a second filter unit size. The second filter unit size has a predefined relationship with the first filter unit size and is configured to be determined based on the first filter unit size.
[0101] (A5) In some embodiments of A4, the plurality of color components includes a luminance component and a chrominance component, and the chrominance format of the current image frame is 4:2:0. The second filter unit size is twice or half of the first filter unit size.
[0102] (A6) In some embodiments of any one of A1 to A5, the current image frame has a superblock size, which represents the maximum coding unit size applied to encode the current image frame, and the filter unit size is equal to or greater than the superblock size.
[0103] (A7) In some embodiments of A6, the first filter unit size for the luminance component of the current image frame is 128×128, and the second filter unit size for the chrominance component of the current image frame is 64×64.
[0104] (A8) In some embodiments of A6 or A7, the video bitstream includes a first syntax element, which indicates the filter unit size for processing the current image frame by the loop filter, and the first syntax element is written based on the superblock size.
[0105] (A9) In some embodiments of A6 or A7, the video bitstream includes a first syntax element, the first syntax element indicates the filter unit size for processing the current image frame by the loop filter, and the first syntax element includes a binary integer, and the binary integer is equal to the difference between the binary logarithm of the filter unit size and the binary logarithm of the superblock size.
[0106] (A10) In some embodiments of any one of A1 to A9, the first filter block is divided into a plurality of sub-filter units according to a tree structure, a control flag associated with the first filter control parameter is written at the sub-filter unit level, and the control flag corresponds to the first sub-filter unit among the sub-filter units of the first filter block.
[0107] (A11) In some embodiments of any one of A1 to A10, the video bitstream includes a first syntax element that indicates a filter unit size for processing the current picture frame by the loop filter. The first syntax element is written at the picture level of the current picture frame, and the filter unit size includes a fixed filter unit size. The video bitstream further includes a second syntax element that indicates a block-level filter unit size for a set of one or more superblocks of the current picture frame. The second syntax element includes an index that defines a scaling factor between the fixed filter unit size and the block-level filter unit size.
[0108] (A12) In some embodiments of A11, the index is equal to one of 0, 1, 2 and corresponds to a scaling factor equal to one of 1, 4, and 16, and the block-level filter unit size is equal to 1 times, 4 times, or 16 times the fixed filter unit size.
[0109] (A13) In some embodiments of any one of A1 to A12, the video bitstream further includes a pointer index associated with a second filter block of the current picture frame. The pointer index identifies the first filter block and controls the loop filter to process the second filter block based on the first filter control parameter associated with the first filter block.
[0110] (A14) In some embodiments of any one of A1 to A13, the loop filter includes a cross-component sample offset (CCSO) mode, and the first filter control parameter indicates whether the CCSO mode is enabled to apply samples of one or more color components to determine a sample offset of samples of a second color component.
[0111] (A15) In some embodiments of A14, the loop filter further includes one or more additional filtering operations different from the CCSO mode, and the filter unit size is jointly written in the first syntax element for the CCSO mode and the one or more additional filtering operations.
[0112] (A16) In some embodiments of A15, the one or more additional filtering operations include at least one of Wiener filtering and cross-component Wiener filtering.
[0113] (A17)In some embodiments of any one of A14 to A16, the method 600 further includes identifying a set of luminance samples in the one or more samples. The set of luminance samples includes a first luminance sample and one or more luminance samples among one or more adjacent luminance samples, and the first luminance sample is co-located with a first color sample of the second color component to be determined in the CCSO mode in the first filter block. The method 600 further includes determining a first sample offset of the first color sample based on the set of luminance samples; and adjusting the first color sample based on the first sample offset of the first color sample.
[0114] (A18)In some embodiments of A17, the method 600 further includes: determining one or more differences between the one or more adjacent luminance samples and the first luminance sample, the first luminance sample being co-located with the first color sample; quantifying the one or more differences to generate one or more quantization values; and classifying the first color sample based on the one or more quantization values to determine the first sample offset of the first color sample.
[0115] (A19)In some embodiments of A17, the method 600 further includes: quantifying the set of luminance samples to generate one or more quantization values; and classifying the first color sample based on the one or more quantization values to determine the first sample offset of the first color sample.
[0116] (A20)In some embodiments of any one of A1 to A19, the filtering unit size of the current image frame is determined based on encoded information, the encoded information including one or more of the following: quantization parameter, image resolution, tile size, whether the current image frame is an intra-only frame, temporal layer, and whether the current image frame is an interpolated frame.
[0117] (A21)In some embodiments of A20, the current image frame is interpolated from two additional image frames.
[0118] (A22)In some embodiments of any one of A1 to A21, the filtering unit size includes a first filtering unit size. The alternate image frame corresponds to a second filtering unit size for processing the alternate image frame through the loop filter. The resolution of the current image frame is greater than the resolution of the alternate image frame, and the first filtering unit size is greater than the second filtering unit size.
[0119] (A23) In some embodiments, a computing system includes a control circuit and a memory that stores one or more programs executed by the control circuit. The one or more programs further include instructions for performing the following operations: receiving video data including a current image frame; encoding the current image frame; transmitting the encoded current image frame via a video bitstream; and writing, via the video bitstream, first filter control parameters for a first filter block that loop filter processes the current image frame. A filter unit size is determined for loop filter processing the current image frame, and the filter unit size is applied to identify the first filter block processed by the loop filter based on the first filter control parameters.
[0120] (A24) In some embodiments, a non-transitory computer-readable storage medium stores one or more programs for execution by a control circuit of a computing system. The one or more programs include instructions for performing the following operations: obtaining a source video sequence including a current image frame; and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes the current image frame and first filter control parameters for a first filter block that loop filter processes the current image frame. A filter unit size is determined for loop filter processing the current image frame, and the filter unit size is applied to identify the first filter block processed by the loop filter based on the first filter control parameters.
[0121] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit. The memory stores one or more sets of instructions configured to be executed by the control circuit, and the one or more sets of instructions include instructions for performing any of the various methods described herein (e.g., A1 to A24 above).
[0122] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium that stores one or more sets of instructions for execution by a control circuit of a computing system, and the one or more sets of instructions include instructions for performing any of the various methods described herein (e.g., A1 to A24 above).
[0123] The proposed methods can be used alone or in any combination in any order. Additionally, each of the various methods (or embodiments), the encoder, and the decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). For example, one or more processors execute a program stored in a non-transitory computer-readable medium. Hereinafter, the term "block" may be interpreted as a prediction block, a coding block, or a coding unit (i.e., a CU).
[0124] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.
[0125] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" are also intended to include the plural forms. It should also be understood that the term "and / or" as used herein refers to any and all possible combinations of one or more of the associated listed items, and encompasses any and all possible combinations of one or more of the associated listed items. It should be further understood that when used in this specification, the terms "comprises" and / or "comprising" expressly specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0126] As used herein, depending on the context, the term "if" can be interpreted, depending on the context, to mean "when" or "after" or "in response to determining..." or "in accordance with determining..." or "in response to detecting", i.e., the precondition is true. Similarly, depending on the context, the phrase "if it is determined that [the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" can be interpreted to mean "after determining that the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "when detecting that the precondition is true" or "in response to detecting that the precondition is true".
[0127] For purposes of explanation, the above description has been presented with reference to specific embodiments. However, the foregoing illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of operation and practical application, thereby enabling others skilled in the art to understand.
Claims
1. A method for decoding video data, comprising: Receiving a video code stream, the video code stream comprising a current image frame and a first filter control parameter of a first filter block for processing the current image frame by a loop filter; determining a filtering unit size for processing the current image frame by the loop filter, the first filtering block having the filtering unit size; identifying the first filter block in the current image frame based on the filter unit size; applying the loop filter to process one or more samples of the first filtering block when the first filter control parameter is enabled; and The current image frame including the first filter block is reconstructed.
2. The method according to claim 1, wherein: The video code stream includes a first syntax element, the first syntax element indicating the filtering unit size used to process the current image frame through the loop filter, and the first syntax element is written at one of the following levels for a first color component of multiple color components: a tile header, a sequence level, and a frame level of the current image frame, wherein the filtering unit size is determined based on the first syntax element written through the video code stream.
3. The method according to claim 2, wherein: The video code stream includes a plurality of syntax elements, the plurality of syntax elements also include the first syntax element, and the plurality of syntax elements are written respectively for the plurality of color components.
4. The method according to claim 2, wherein: The filter unit size includes a first filter unit size, and a second color component of the multiple color components is not written and has a second filter unit size, wherein the second filter unit size has a predefined relationship with the first filter unit size and is configured to be determined based on the first filter unit size.
5. The method according to claim 4, wherein: The plurality of color components include a luminance component and a chrominance component, and the chrominance format of the current image frame is 4:2:0; and The size of the second filtering unit is twice or half of the size of the first filtering unit.
6. The method according to claim 1, wherein: The current image frame has a super block size, the super block size indicating a maximum coding unit size applied to encode the current image frame, and the filtering unit size is equal to or greater than the super block size.
7. The method according to claim 6, wherein: The first filter unit size for the luma component of the current image frame is 128×128, and the second filter unit size for the chroma component of the current image frame is 64×64.
8. The method according to claim 6, wherein: The video code stream includes a first syntax element, which indicates the filtering unit size used to process the current image frame through the loop filter, and the first syntax element is written based on the super block size.
9. The method according to claim 6, wherein: The video code stream includes a first syntax element, the first syntax element indicates the filtering unit size used to process the current image frame through the loop filter, and the first syntax element includes a binary integer, which is equal to the difference between the binary logarithm of the filtering unit size and the binary logarithm of the super block size.
10. The method according to claim 1, wherein: The first filter block is divided into multiple sub-filter units according to a tree structure, and a control flag associated with the first filter control parameter is written at the sub-filter unit level, and the control flag corresponds to the first sub-filter unit among the sub-filter units of the first filter block.
11. The method according to claim 1, wherein: The video code stream includes a first syntax element, wherein the first syntax element indicates the filter unit size used for processing the current image frame through the loop filter; The first syntax element is written at a frame level of the current image frame, and the filtering unit size comprises a fixed filtering unit size; The video code stream also includes a second syntax element, which indicates a block-level filtering unit size for a set of one or more super blocks of the current image frame, and the second syntax element includes an index defining a scaling factor between the fixed filtering unit size and the block-level filtering unit size.
12. The method according to claim 11, wherein: The index is equal to one of 0, 1, 2 and corresponds to the scaling factor being equal to one of 1, 4 and 16 respectively, and the block-level filtering unit size is equal to 1 times, 4 times or 16 times the fixed filtering unit size.
13. The method according to claim 1, wherein: The video code stream also includes a pointer index associated with a second filter block of the current image frame, and the pointer index identifies the first filter block and controls the loop filter to process the second filter block based on the first filter control parameter associated with the first filter block.
14. The method according to claim 1, wherein: The loop filter includes a cross-component sample offset (CCSO) mode, and the first filter control parameter indicates whether the CCSO mode is enabled to apply one or more samples of a first color component to determine a sample offset of samples of a second color component.
15. The method according to claim 14, wherein: The loop filter also includes one or more additional filtering operations different from the CCSO mode, and the filtering unit size is jointly written in the first syntax element for the CCSO mode and the one or more additional filtering operations.
16. The method according to claim 15, wherein: The one or more additional filtering operations include at least one of Wiener filtering and cross-component Wiener filtering.
17. The method according to claim 14, further comprising: identifying a group of luma samples among the one or more samples, wherein the group of luma samples includes a first luma sample and one or more luma samples of one or more adjacent luma samples, and the first luma sample is co-located with a first color sample of the second color component to be determined in the CCSO mode in the first filter block; determining a first sample offset for the first color sample based on the set of luma samples; and The first color sample is adjusted based on a first sample offset of the first color sample.
18. The method according to claim 17, further comprising: determining one or more differences between the one or more adjacent luma samples and the first luma sample, the first luma sample being co-located with the first color sample; quantizing the one or more difference values to generate one or more quantized values; as well as The first color sample is classified based on the one or more quantized values to determine a first sample offset for the first color sample.
19. The method according to claim 17, further comprising: quantizing the set of luma samples to generate one or more quantized values; as well as The first color sample is classified based on the one or more quantized values to determine a first sample offset for the first color sample.
20. The method according to claim 1, wherein: The filtering unit size of the current image frame is determined based on encoded information, and the encoded information includes one or more of the following: quantization parameter, image resolution, tile size, whether the current image frame is an intra-frame only, temporal layer, and whether the current image frame is an interpolated frame. The method also includes determining whether the first filter control parameter associated with the first filtering block is enabled.
21. The method according to claim 20, wherein: The current image frame is obtained by interpolation based on two additional image frames.
22. The method of claim 1, wherein: The filtering unit size comprises a first filtering unit size; The substitute image frame corresponds to a second filtering unit size for processing the substitute image frame through the loop filter; as well as A resolution of the current image frame is greater than a resolution of the replacement image frame, and the first filtering unit size is greater than the second filtering unit size.
23. A computing system comprising: Control circuit; as well as A memory storing one or more programs, wherein the one or more programs are configured to be executed by the control circuit, and the one or more programs further include instructions for performing the following operations: Receiving video data including a current image frame; Encoding the current image frame; Transmitting the encoded current image frame via a video code stream; as well as Writing, via the video bitstream, a first filter control parameter of a first filter block used for loop filter processing of the current image frame; The filtering unit size is determined to be used for processing the current image frame by the loop filter, and the filtering unit size is applied to identify the first filter block processed by the loop filter based on the first filter control parameter.
24. A non-transitory computer-readable storage medium storing one or more programs executed by control circuitry of a computing system, the one or more programs comprising instructions for: Obtaining a source video sequence including a current image frame; as well as Execute conversion between the source video sequence and the video code stream, wherein: The video code stream includes: the current image frame; and a first filter control parameter for a first filter block for loop filter processing of the current image frame; The filtering unit size is determined to be used for processing the current image frame by the loop filter, and the filtering unit size is applied to identify the first filter block processed by the loop filter based on the first filter control parameter.