Cross-component sample offset filtering with asymmetric quantizer
By using an asymmetric quantizer and a cross-component sample offset filter, the problem of low video coding efficiency in existing technologies is solved, achieving more efficient video coding and decoding effects, reducing coding artifacts, and improving the quality of decoded images.
Patent Information
- Application Number
- CN202380100969.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2023-10-31
- Publication Date
- 2026-03-03
AI Technical Summary
Existing video coding techniques are inefficient in reducing redundancy in uncompressed video signals, especially when applying cross-component sample offset filtering, due to a lack of effective quantizer selection and signaling mechanisms.
An asymmetric quantizer is used to process the sample increment, and an appropriate quantizer is selected through a signaling mechanism. A cross-component sample offset filter is applied for filtering, including the correspondence between the symmetric and asymmetric quantizers. The quantizer selection and offset determination are performed using syntax elements.
It improves the efficiency and quality of video encoding, reduces encoding artifacts, and enhances the clarity and compression effect of decoded images.
Smart Images

Figure CN121605633A_ABST
Abstract
Description
[0001] By incorporating via reference
[0002] This disclosure is based on and claims priority to U.S. non-provisional patent application No. 18 / 497,662, filed October 30, 2023, entitled “CROSS COMPONENT SAMPLE OFFSET FILTERING WITH ASYMMETRIC QUANTIZER,” which is based on and claims priority to U.S. provisional patent application No. 63 / 530,571, filed August 3, 2023, entitled “CROSS COMPONENT SAMPLE OFFSET FILTERING WITH ASYMMETRIC QUANTIZER,” both of which are incorporated herein by reference in their entirety. Technical Field
[0003] This disclosure generally describes a range of advanced video coding techniques, and specifically relates to cross-component sample offset filtering. Background Technology
[0004] Uncompressed digital video can comprise a series of images and may have specific bitrate requirements for storage, data processing, and transmission bandwidth in streaming applications. One objective of video encoding and decoding is to reduce redundancy in the uncompressed input video signal through various compression techniques. Summary of the Invention
[0005] This disclosure describes various aspects of advanced video coding techniques, and particularly relates to Cross-Component Sample Offset (CCSO) filtering for reconstructed samples. For example, the quantizer used to quantize the sample increment to apply the CCSO filter can be constructed as an asymmetric quantizer relative to zero sample increments. Before applying the CCSO filter, either a symmetric or asymmetric quantizer can be selected to process the sample increment. The asymmetric quantizer can be correlated with the corresponding symmetric quantizer via a non-zero offset. The signaling for selecting the asymmetric quantizer can be independent of or dependent on the signaling provision of the symmetric quantizer.
[0006] In some example implementations, a method for video filtering is disclosed. The method may include: reconstructing frames from a video bitstream to generate reconstructed samples of at least a first color component and a second color component of the frame; obtaining a quantizer for quantizing a CCSO filter unit of a plurality of quantizers for the frame based on a first syntax element signaled in the video bitstream, the plurality of quantizers including at least one asymmetric quantizer, the at least one asymmetric quantizer being asymmetric with respect to values quantized on the positive and negative sides of zero; selecting a quantizer indicated by the first syntax element; applying the selected quantizer to a CCSO sample increment of the reconstructed samples of the first color component of the CCSO filter unit according to a CCSO filter to generate a quantized CCSO sample increment; determining a CCSO sample offset based on the quantized CCSO sample increment and the CCSO filter; and applying the CCSO sample offset to the reconstructed samples of the second color component in the CCSO filter unit to generate a filtered sample of the second color component in the CCSO filter unit.
[0007] In the example implementation described above, each of the multiple quantizers is represented by one or more quantization intervals.
[0008] In any of the above example implementations, at least one asymmetric quantizer corresponds to one of the symmetric quantizers among a plurality of quantizers via a CCSO sample increment offset, each of the symmetric quantizers being symmetric with respect to zero sample increment.
[0009] In any of the above example implementations, multiple quantizers are predefined and indexed, and the selected quantizer is signaled in the video bitstream by a first syntax element.
[0010] In any of the above example implementations, the multiple quantizers include a subset of symmetric quantizers with a one-to-one correspondence and a subset of asymmetric quantizers.
[0011] In any of the above example implementations, the symmetric quantizer subset and the asymmetric quantizer subset are correlated by a predefined sample increment offset.
[0012] In any of the above example implementations, the selection of an asymmetric quantizer is indicated in the video bitstream by a second syntax element and by a first syntax element, the second syntax element being used to signal the indication to use an asymmetric quantizer, and the first syntax element indicating a symmetric quantizer in a subset of symmetric quantizers corresponding to the selected asymmetric quantizer.
[0013] In any of the above example implementations, the first syntax element includes an index for a symmetric quantizer corresponding to the selected asymmetric quantizer.
[0014] In any of the above example implementations, the multiple quantizers include an unequal number of symmetric quantizer subsets and an asymmetric CCSO quantizer subset.
[0015] In any of the above example implementations, each symmetric quantizer in the subset of symmetric quantizers corresponds to zero or more asymmetric quantizers through one or more sample increment offsets from a predefined set of sample increment offsets.
[0016] In any of the above example implementations, the selection of an asymmetric quantizer is indicated in the video bitstream by a first syntax element and a second syntax element, the first syntax element indicating a symmetric quantizer corresponding to the selected asymmetric quantizer, and the second syntax element indicating an offset for applying the symmetric quantizer indicated by the first syntax element to generate the selected asymmetric quantizer.
[0017] In any of the above example implementations, the selection of the asymmetric quantizer is indicated in the video bitstream by a first syntax element and a second syntax element, the first syntax element indicating the symmetric quantizer corresponding to the selected asymmetric quantizer, and the second syntax element indicating the index among the asymmetric quantizers corresponding to the symmetric quantizer indicated by the first syntax element.
[0018] In any of the above example implementations, the number of asymmetric quantizers corresponding to the symmetric quantizer depends non-decreasingly on the size of the quantization interval of the symmetric quantizer.
[0019] In any of the above example implementations, at least one of the symmetric quantizer subsets does not correspond to any of the asymmetric quantizer subsets.
[0020] In any of the above example implementations, when the dead zone of a symmetric quantizer in the subset of symmetric quantizers is at or below a threshold, the symmetric quantizer does not correspond to any of the subset of asymmetric quantizers. The dead zone represents the quantization interval containing the zero-sample increment of the symmetric quantizer.
[0021] In any of the above example implementations, the first syntax element is signaled at the sequence header, picture header, tile header, or maximum coded block header level.
[0022] In any of the above example implementations, the first syntax element is used to select an asymmetric quantizer from a plurality of quantizers by: indicating a symmetric quantizer, deriving a sample increment offset based on the indicated symmetric quantizer, and generating the selected asymmetric quantizer by applying the resulting sample increment offset to the indicated symmetric quantizer.
[0023] In any of the above example implementations, the sample increment offset is derived based on the size of the dead zone of the indicated symmetric quantizer, where the dead zone represents the quantization interval containing the zero sample increment of the indicated symmetric quantizer.
[0024] In any of the above example implementations, determining the selected quantizer based on the first syntax element may include: determining the selected symmetric quantizer based on the first syntax element; determining the size of the dead zone of the selected symmetric quantizer, the dead zone representing the quantization interval containing the zero-sample increment of the selected symmetric quantizer; determining to apply the selected symmetric quantizer in response to the size of the dead zone of the selected symmetric quantizer being less than a predefined threshold; and determining whether to select an asymmetric quantizer based on a second syntax element in the video bitstream in response to the size of the dead zone of the selected symmetric quantizer being not less than the predefined threshold.
[0025] In some implementations, a video encoding or decoding apparatus is disclosed. This apparatus may include a circuit system configured to implement any of the methods described above.
[0026] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform methods for video decoding and / or encoding. Attached Figure Description
[0027] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0028] Figure 1 A schematic diagram of a simplified block diagram of a communication system (100) according to an example embodiment is shown.
[0029] Figure 2 A schematic diagram of a simplified block diagram of a communication system (200) according to an example embodiment is shown.
[0030] Figure 3 A schematic diagram of a simplified block diagram of a video decoder according to an example implementation is shown.
[0031] Figure 4 A schematic diagram of a simplified block diagram of a video encoder according to an example implementation is shown.
[0032] Figure 5 A block diagram of a video encoder according to another example implementation is shown.
[0033] Figure 6 A block diagram of a video decoder according to another example implementation is shown.
[0034] Figure 7An exemplary adaptive loop filter according to an embodiment of this disclosure is shown.
[0035] Figures 8A to 8D Examples of subsampling positions for calculating gradients in the vertical, horizontal and two diagonal directions, respectively, according to embodiments of this disclosure, are shown.
[0036] Figure 8E An example of how block directionality is determined based on various gradients for use by the Adaptive Loop Filter (ALF) is shown.
[0037] Figure 9A and Figure 9B A modified block classification at virtual boundaries is shown in an example implementation of this disclosure.
[0038] Figures 10A to 10F An exemplary adaptive loop filter with a fill operation at a corresponding virtual boundary is shown according to an embodiment of the present disclosure.
[0039] Figure 11 An example of image quadtree segmentation with maximum coding unit alignment according to an embodiment of this disclosure is shown.
[0040] Figure 12 An example implementation of the present disclosure is shown. Figure 11 The corresponding quadtree segmentation pattern.
[0041] Figure 13 A cross-component filter for generating chromaticity components is shown according to an example embodiment of the present disclosure.
[0042] Figure 14 An example of a cross-component ALF filter according to an embodiment of this disclosure is shown.
[0043] Figure 15 An exemplary position of a chromaticity sample relative to a luminance sample according to an embodiment of this disclosure is shown.
[0044] Figure 16 An example of a block-oriented search according to an embodiment of this disclosure is shown.
[0045] Figure 17 An example of subspace projection according to an embodiment of this disclosure is shown.
[0046] Figure 18 This shows an example location for Cross-Component Sample Offset (CCSO) filtering in a loop filter pipeline.
[0047] Figure 19 An example of a filter support region in a CCSO filter according to an embodiment of this disclosure is shown.
[0048] Figure 20 An example implementation of a 3-tap CCSO filter shape according to an embodiment of this disclosure is shown.
[0049] Figure 21 A flowchart outlining the process (2100) according to an embodiment of this disclosure is shown.
[0050] Figure 22 A schematic diagram of a computer system according to an example implementation is shown. Detailed Implementation
[0051] Throughout the specification and claims, terms may have subtle meanings implied or implied in the context beyond their expressly stated meanings. The phrases “in one embodiment” or “in some embodiments” as used herein do not necessarily refer to the same embodiment, and the phrases “in another embodiment” or “in other embodiments” as used herein do not necessarily refer to different embodiments. For example, it is intended that the claimed subject matter includes a combination of all or part of the exemplary embodiments.
[0052] Generally, terms can be understood, at least in part, from their usage in context. For example, terms such as “and,” “or,” or “and / or” as used herein can include a variety of context-dependent meanings. Generally, if “or” is used with an associative list such as A, B, or C, it is intended to mean: A, B, and C, used here in an inclusive sense; and A, B, or C, used here in an exclusive sense. Additionally, the terms “one or more,” “at least one,” “a,” “one,” or “described” as used herein, at least in part, depend on the context and can be used in a singular or plural sense. Furthermore, the terms “based on” or “determined by” can be understood as not necessarily intended to convey an exclusive set of factors, and again, at least in part, depend on the context, and may alternatively allow for additional factors that are not necessarily explicitly described.
[0053] Figure 1 A simplified block diagram of a communication system (100) according to an embodiment of the present disclosure is shown. The communication system (100) includes a plurality of terminal devices, such as terminal devices 110, 120, 130 and 140, that can communicate with each other via, for example, a network (150). Figure 1In the example, the first pair of terminal devices (110) and (120) can perform unidirectional data transmission. For example, terminal device (110) can encode video data (e.g., a video image stream captured by terminal device (110)) in the form of one or more encoded bitstreams for transmission via network (150). Terminal device (120) can receive the encoded video data from network (150), decode the encoded video data to recover the video images, and display the video images based on the recovered video images. Unidirectional data transmission can be implemented in media service applications, etc.
[0054] In another example, the second pair of terminal devices (130) and (140) can, for example, perform bidirectional transmission of encoded video data during a video conferencing application. For bidirectional data transmission, in the example, each of the terminal devices (130) and (140) can encode video data (e.g., a video image stream captured by the terminal device) for transmission to the other terminal device (130) and (140), and can also receive encoded video data from the other terminal device (130) and (140) to recover and display the video images.
[0055] exist Figure 1 In the examples, the terminal device can be implemented as a server, a personal computer, and a smartphone, but the applicability of the basic principles of this disclosure is not limited to these. Implementations of this disclosure can be carried out in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, etc. Network (150) refers to any number or type of network that transmits encoded video data between terminal devices, including, for example, wired (connected) and / or wireless communication networks. Communication network (150) can exchange data in circuit-switched channels, packet-switched channels, and / or other types of channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet.
[0056] As an example of the application of the disclosed topic, Figure 2 The placement of a video encoder and video decoder in a video streaming environment is illustrated. The disclosed subject matter is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, and storing compressed video on digital media including CDs, DVDs, Memory Sticks, etc.
[0057] like Figure 2As shown, a video streaming system may include a video capture subsystem (213), which may include a video source (201), such as a digital camera device, for creating an uncompressed video picture or image stream (202). In the example, the video picture stream (202) includes samples recorded by the digital camera device of the video source 201. The video picture stream (202) is depicted as a thick line to emphasize the high data volume when compared with encoded video data (204) (or encoded video bitstream), which may be processed by an electronic device (220) including a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter as described in more detail below. Encoded video data (204) (or encoded video bitstream (204)) is depicted as a thin line to emphasize its lower data volume when compared to a stream of uncompressed video images (202). The encoded video data (204) may be stored on a streaming server (205) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems, for example... Figure 2 The client subsystems (206) and (208) can access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) may include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes the incoming copy (207) of the encoded video data and creates an uncompressed outgoing video picture stream (211) that can be displayed on a display (212) (e.g., a screen) or other presentation device (not shown).
[0058] Figure 3 A block diagram of a video decoder (310) for an electronic device (330) according to any embodiment of the present disclosure described below is shown. The electronic device (330) may include a receiver (331) (e.g., a receiving circuitry system). A video decoder (310) may be used instead. Figure 2 The example video decoder (210).
[0059] like Figure 3As shown, the receiver (331) can receive one or more encoded video sequences from the channel (301). To combat network jitter and / or process playback timing, a buffer memory (315) can be placed between the receiver (331) and the entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) can reconstruct symbols (321) from the encoded video sequences. These symbols include information for managing the operation of the video decoder (310) and potential information for controlling presentation devices such as displays (312) (e.g., screens). The parser (320) can perform parsing / entropy decoding on the encoded video sequences. The parser (320) can extract a set of subgroup parameters from at least one subgroup of pixels in the encoded video sequences for use in the video decoder. Subgroups may include Group of Picture (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (320) can also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc., from the encoded video sequence. The reconstruction of the symbols (321) can involve multiple different processing or functional units. The units involved and how these units are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (320).
[0060] The first unit may include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive from the parser (320) quantization transform coefficients as symbols (321) and control information, including information indicating which type of inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block comprising sample values that can be input to the aggregator (355).
[0061] In some cases, the output samples of the scaler / inverse transform (351) may belong to intra-coded blocks, i.e., blocks that do not use predictive information from previously reconstructed images but can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) can use information from surrounding blocks that have been reconstructed and stored in the current picture buffer (358) to generate blocks of the same size and shape as the blocks in the reconstruction. For example, the current picture buffer (358) buffers the partially reconstructed current picture and / or the fully reconstructed current picture. In some implementations, the aggregator (355) may add the predictive information already generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.
[0062] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to inter-frame coded and possibly motion-compensated blocks. In such cases, the motion-compensated prediction unit (353) may access the reference image memory (357) based on motion vectors to obtain samples for inter-frame image prediction. After motion compensation of the obtained reference samples according to the symbols (321) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (351) (the output of unit 351 may be referred to as residual samples or residual signals) by the aggregator (355) to generate output sample information.
[0063] The output samples of the aggregator (355) can undergo various loop filtering techniques in a loop filter unit (356) that includes several types of loop filters. The output of the loop filter unit (356) can be a sample stream that can be output to a presentation device (312) and stored in a reference image memory (357) for future inter-frame image prediction.
[0064] Figure 4 A block diagram of a video encoder (403) according to an example embodiment of the present disclosure is shown. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may also include a transmitter (440) (e.g., a transmission circuitry system). The video encoder (403) may be used instead of Figure 4 The example video encoder (403).
[0065] The video encoder (403) can receive video samples from the video source (401). According to some example implementations, the video encoder (403) can encode and compress images of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate constitutes a function of the controller (450). In some implementations, the controller (450) can be functionally coupled to and control other functional units as described below. Parameters set by the controller (450) may include parameters related to rate control (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of images (GOP) layout, maximum motion vector search range, etc.
[0066] In some example implementations, the video encoder (403) can be configured to operate in a codec loop. The codec loop may include a source encoder (430) and a (local) decoder (433) embedded in the video encoder (403). Even if the embedded decoder 433 processes the encoded video stream encoded by the source encoder 430 without entropy coding, the decoder (433) reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder would create sample data (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bitstream in entropy coding can be lossless). It can be observed that any decoder technique other than parsing / entropy decoding, which may only exist in the decoder, may also need to exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operations related to the decoding portion of the encoder. Since the encoder technique is the inverse of the decoder technique, which has been fully described, the description of the encoder technique can be simplified. A more detailed description of the encoder is provided below only in certain areas or aspects.
[0067] During operation in some example implementations, the source encoder (430) may perform motion-compensated predictive coding that predictively codes the input image with reference to one or more previously coded images designated as “reference images” from the video sequence.
[0068] The local video decoder (433) can decode the encoded video data of an image that can be designated as a reference image. The local video decoder (433) replicates the decoding process that can be performed on the reference image by the video decoder, and can store the reconstructed reference image in a reference image cache (434). In this way, the video encoder (403) can store a copy of the reconstructed reference image locally, which has the same content as the reconstructed reference image to be obtained by the remote video decoder (without transmission errors).
[0069] The predictor (435) can perform a prediction search against the encoding engine (432). That is, for a new image to be encoded, the predictor (435) can search in the reference image memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image.
[0070] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.
[0071] The outputs of all the aforementioned functional units can undergo entropy encoding in the entropy encoder (445). The transmitter (440) can buffer one or more encoded video sequences created by the entropy encoder (445) in preparation for transmission via a communication channel (460), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (440) can combine the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0072] The controller (450) can manage the operation of the video encoder (403). During encoding, the controller (450) can assign a specific encoding picture type to each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures can typically be assigned to one of the following picture types: intra-frame picture (I picture), predictive picture (P picture), bidirectional predictive picture (B picture), or multi-predictive picture. As described in further detail below, source pictures can typically be spatially subdivided into multiple sample coding blocks.
[0073] Figure 5 A diagram of a video encoder (503) according to another example embodiment of this disclosure is shown. The video encoder (503) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video image in a video image sequence, and to encode the processing block into an encoded image that is part of an encoded video sequence. The example video encoder (503) can be used instead. Figure 4 The video encoder (403) in the example.
[0074] For example, the video encoder (503) receives a matrix of sample values of the processed block. The video encoder (503) then uses, for example, rate-distortion optimization (RDO) to determine whether to best encode the processed block using intra-frame mode, inter-frame mode, or dual prediction mode.
[0075] exist Figure 5 In the example, the video encoder (503) includes, for example, Figure 5 The example arrangement shows an inter-frame encoder (530), an intra-frame encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a master controller (521), and an entropy encoder (525) coupled together.
[0076] The inter-frame encoder (530) is configured to receive a sample of the current block (e.g., the processing block), compare the block with one or more reference blocks in the reference images (e.g., blocks in the previous and later images in display order), generate inter-frame prediction information (e.g., a description of redundancy information based on the inter-frame coding technique, motion vectors, merging mode information), and compute inter-frame prediction results (e.g., predicted blocks) based on the inter-frame prediction information using any suitable technique.
[0077] The intra encoder (522) is configured to receive samples of the current block (e.g., the processing block), compare the block with blocks that have already been encoded in the same picture, and generate transformed quantization coefficients, and in some cases also generate intra predictive information (e.g., intra predictive direction information based on one or more intra coding techniques).
[0078] The master controller (521) can be configured to determine master control data and, based on the master control data, control other components of the video encoder (503) to, for example, determine a prediction mode of a block and provide control signals to the switch (526) based on the prediction mode.
[0079] A residual calculator (523) can be configured to calculate the difference (residual data) between the received block and the prediction result of a block selected from an intra-encoder (522) or an inter-encoder (530). A residual encoder (524) can be configured to encode the residual data to generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients. In various example embodiments, the video encoder (503) also includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform and generate decoded residual data. An entropy encoder (525) can be configured to format the bitstream to include coded blocks and perform entropy encoding.
[0080] Figure 6 A diagram of an example video decoder (610) according to another embodiment of this disclosure is shown. The video decoder (610) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In the example, a video decoder (610) may be used instead of Figure 4The example video decoder (410).
[0081] exist Figure 6 In the example, the video decoder (610) includes, for example, Figure 6 The example arrangement shows an entropy decoder (671), an inter-frame decoder (680), a residual decoder (673), a reconstruction module (674), and an intra-frame decoder (672) coupled together.
[0082] An entropy decoder (671) can be configured to reconstruct certain symbols representing the syntax elements constituting the encoded picture from the encoded picture. An inter-frame decoder (680) can be configured to receive inter-frame prediction information and generate inter-frame prediction results based on the inter-frame prediction information. An intra-frame decoder (672) can be configured to receive intra-frame prediction information and generate prediction results based on the intra-frame prediction information. A residual decoder (673) can be configured to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. A reconstruction module (674) can be configured to combine the residual output by the residual decoder (673) with the prediction results (output by the inter-frame prediction module or the intra-frame prediction module, as appropriate) in the spatial domain to form a reconstruction block, which forms part of the reconstructed picture as part of the reconstructed video.
[0083] Note that any suitable technology can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In some example implementations, one or more integrated circuits can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In another implementation, one or more processors executing software instructions can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). And one or more processors executing software instructions can be used to implement (810).
[0084] In some example implementations, a loop filter can be included in both the encoder and decoder to reduce encoding artifacts and improve the quality of the decoded image. For example, a loop filter 356 could be included as... Figure 3 It is part of decoder 330. For another example, the loop filter could be... Figure 4These filters are part of the embedded decoder unit 433 in the encoder 420. These filters are called loop filters because they are included in the decoding loop used for video blocks in the decoder or encoder. Each loop filter can be associated with one or more filtering parameters. Such filtering parameters can be predefined or can be derived by the encoder during encoding processing. These filtering parameters (in the case of being derived by the encoder) or their indices (in the case of being predefined) can be included in the final bitstream in encoded form. The decoder can then parse these filtering parameters from the bitstream during decoding and perform loop filtering based on the parsed filtering parameters.
[0085] Various loop filters can be used to reduce coding artifacts and improve the quality of decoded video in different ways. Such loop filters can include, but are not limited to, one or more deblocking filters, adaptive loop filters (ALF), cross-component adaptive loop filters (CC-ALF), constrained directional enhancement filters (CDEF), sample adaptive offset (SAO) filters, cross-component sample offset (CCSO) filters, and local sample offset (LSO) filters. These filters may or may not depend on each other. They can be arranged in the decoding loop of the decoder or encoder in any suitable order compatible with their interdependencies (if any). These various loop filters are described in more detail in the disclosure below.
[0086] An Adaptive Loop Filter (ALF) with block-based filter adaptation can be applied by the encoder / decoder to reduce artifacts. An ALF is adaptive in the sense that the filter coefficients / parameters or their indices are signaled in the bitstream and can be designed based on image content and the distortion of the reconstructed image. An ALF can be applied to reduce distortion introduced by the encoding process and improve the quality of the reconstructed image.
[0087] For the luma component, one of several filters (e.g., 25 filters) can be selected for the luma block (e.g., a 4×4 luma block) based, for example, on the direction and activity of the local gradient. The filter coefficients of these filters can be derived by the encoder during the encoding process and communicated to the decoder via signaling in the bitstream.
[0088] ALFs can have any suitable shape and size. (See reference) Figure 7 For example, an ALF can have a rhombus shape, such as a 5×5 rhombus for ALF(710) and a 7×7 rhombus for ALF(711). In ALF(710), thirteen (13) elements can be used for filtering and form a rhombus shape. For these 13 elements, seven values (e.g., C0 to C6) can be used and arranged as shown in the example. In ALF(711), twenty-five (25) elements can be used for filtering and form a rhombus shape. Thirteen (13) values (e.g., C0 to C12) can be used for the 25 elements as shown in the example.
[0089] Reference Figure 7 In some examples, an ALF filter with one of two rhombus shapes (710) to (711) can be selected to process the luma block or chroma block. For example, a 5×5 rhombus shape filter (710) can be applied to the chroma component (e.g., chroma block, chroma CB), and a 7×7 rhombus shape filter (711) can be applied to the luma component (e.g., luma block, luma CB). Other suitable shapes and sizes can be used in the ALF. For example, a 9×9 rhombus shape filter can be used.
[0090] The filter coefficients at the locations indicated by values (e.g., C0 to C6 in (710) or C0 to C12 in (711)) can be nonzero. Furthermore, when the ALF includes a clipping function, the clipping value at those locations can be nonzero. The clipping function can be used to limit the upper limit of filter values in a luma or chroma block.
[0091] In some implementations, the specific ALF to be applied to a specific block with a luminance component can be based on the classification of the luminance block. For the block classification of the luminance component, a 4×4 block (or luminance block, luminance CB) can be classified or categorized into one of several (e.g., 25) categories, which correspond to, for example, 25 different ALFs (e.g., 25 ALFs with different filter coefficients in a 7×7 ALF). The classification index C can be derived using Equation (1) based on the quantization value of the directional parameter D and the activity value A.
[0092] Equation (1)
[0093] To calculate the directional parameter D and the quantization value Â, the gradient g in the vertical, horizontal, and two diagonal directions (e.g., d1 and d2) can be calculated using a 1-dimensional Laplacian as follows. v g h g d1 and gd2 .
[0094] Equation (2)
[0095] Equation (3)
[0096] Equation (4)
[0097] Equation (5)
[0098] Here, indices i and j refer to the coordinates of the top-left sample within the 4×4 block, and R(k, l) indicates the reconstructed sample at coordinates (k, 1). Directions (e.g., d1 and d2) refer to the two diagonal directions.
[0099] To reduce the complexity of the block classification described above, a 1-dimensional Laplace calculation based on subsampling can be applied. Figures 8A to 8D The following diagrams show the methods for calculating the vertical direction ( Figure 8A ), horizontal direction ( Figure 8B ) and the two diagonal directions d1 ( Figure 8C ) and d2 ( Figure 8D The gradient g of ) v g h g d1 and g d2 An example of a subsampling location. Figure 8A In the text, the label "V" indicates the symbol used to calculate the vertical gradient g. v The subsampling location. Figure 8B In the text, the label "H" indicates the symbol used to calculate the horizontal gradient g. h The subsampling location. Figure 8C In the text, the label "D1" indicates the value used to calculate the diagonal gradient g of d1. d1 The subsampling location. Figure 8D In the text, the label "D2" indicates the value used to calculate the diagonal gradient g of d2. d2 The sub-sampling position. Figure 8A and Figure 8B This demonstrates that the same subsampling location can be used for gradient calculation in different directions. In some other implementations, different subsampling schemes can be used in all directions. In yet another implementation, different subsampling schemes can be used in different directions.
[0100] The gradients g in the horizontal and vertical directions v and g h maximum value and minimum value It can be set to:
[0101] Equation (6)
[0102] gradient g in the two diagonal directions d1 and g d2 maximum value and minimum value It can be set to:
[0103] Equation (7)
[0104] The directional parameter D can be derived based on the above value and the following two thresholds t1 and t2.
[0105] Step 1. If (1) And (2) If true, then set D to 0.
[0106] Step 2. If If yes, proceed to step 3; otherwise, proceed to step 4.
[0107] Step 3. If If the result is positive, then set D to 2; otherwise, set D to 1.
[0108] Step 4. If If the result is positive, then set D to 4; otherwise, set D to 3.
[0109] In other words, such as Figure 8E As shown, the directional parameter D is represented by several discrete levels and is determined based on the gradient value distribution of the brightness block between the horizontal and vertical directions and between the two diagonal directions.
[0110] Activity value A can be calculated as:
[0111] Equation (8)
[0112] Therefore, the activity value A represents a combined metric of the horizontal and vertical 1-dimensional Laplacian. The activity value A for a luminance block can be further quantized, for example, to a range of 0 to 4 (inclusive), and the quantized value is represented as Â.
[0113] For the luminance component, the classification index C, calculated as above, can then be used to select one of several categories (e.g., 25 categories) of the diamond-shaped ALF filter. In some implementations, for the chrominance component in the image, block classification may not be applied, and therefore a single set of ALF coefficients can be applied for each chrominance component. In such an implementation, while multiple sets of ALF coefficients may exist for the chrominance components, the determination of the ALF coefficients may not depend on any classification of the chrominance blocks.
[0114] Geometric transformations can be applied to filter coefficients and the corresponding filter limiting values (also known as limiting values). Before filtering a block (e.g., a 4×4 brightness block), the limiting factor depends on the gradient values (e.g., g) calculated for the block. v g h g d1 and / or g d2 Geometric transformations, such as rotation or diagonal and vertical flips, can be applied to the filter coefficients f(k, l) and the corresponding filter threshold c(k, l). Applying geometric transformations to the filter coefficients f(k, l) and the corresponding filter threshold c(k, l) is equivalent to applying geometric transformations to samples within the region supported by the filter. Geometric transformations can make different blocks to which ALF is applied more similar by aligning their respective orientations.
[0115] The three geometric transformation options, including diagonal flip, vertical flip and rotation, can be performed as described in equations (9) to (11).
[0116] Equation (9)
[0117] Equation (10)
[0118] Equation (11)
[0119] Here, K represents the size of the ALF or filter, and 0 ≤ k, l ≤ K-1 are the coordinates of the coefficients. For example, position (0, 0) is at the top left corner of the filter f or the limiting matrix (or limiting matrix) c, while position (K-1, K-1) is at the bottom right corner. Depending on the gradient values computed for the block, transformations can be applied to the filter coefficients f(k, l) and the limiting value c(k, l). Table 1 summarizes examples of the relationship between the transformations and the four gradients.
[0120] Table 1: Mapping of gradients to transformations computed for blocks
[0121]
[0122] In some implementations, the ALF filter parameters derived by the encoder can be signaled in an Adaptation Parameter Set (APS) for the image. In the APS, one or more sets (e.g., up to 25 sets) of luminance filter coefficients and limiting indexes can be signaled. These can be indexed in the APS. In an example, one of the sets may include luminance filter coefficients and one or more limiting indexes. One or more sets (e.g., up to 8 sets) of chrominance filter coefficients and limiting indexes can be derived by the encoder and signaled. To reduce signaling overhead, filter coefficients for different classifications of the luminance components (e.g., with different classification indices) can be merged. The index of the APS used for the current slice can be signaled in the slice header. In another example, the ALF signaling can be CTU-based.
[0123] In an implementation, the limiting index (also called the limiting index) can be decoded according to the APS. The limiting index can be used, for example, to determine the corresponding limiting value based on the relationship between the limiting index and the corresponding limiting value. This relationship can be predefined and stored in the decoder. In an example, this relationship is described by one or more tables such as a table of limiting indexes and corresponding limiting values for the luma component (e.g., for luma CB) and a table of limiting indexes and corresponding limiting values for the chroma component (e.g., for chroma CB). The limiting value can depend on the bit depth B. The bit depth B can refer to the internal bit depth, the bit depth of the reconstructed sample in the CB to be filtered, etc. In some examples, Equation (12) can be used to obtain the table of limiting values (e.g., for luma and / or for chroma).
[0124] Equation (12)
[0125] Where AlfClip is the clipping value, B is the bit depth (e.g., bitDepth), N (e.g., N = 4) is the number of allowed clipping values, and α is a predefined constant value. In the example, α equals 2.35. n is the clipping value index (also called the clipping index or clipIdx). Table 2 shows an example of a table obtained using equation (12), where N = 4. The clipping index n can be 0, 1, 2, and 3 (up to N-1) in Table 2. Table 2 can be used for luma blocks or chroma blocks.
[0126] Table 2 – AlfClip can depend on bit depth B and clipIdx
[0127]
[0128] In the slice header of the current slice, one or more APS indices (e.g., up to 7 APS indices) can be signaled to specify the set of luma filters that can be used for the current slice. Filtering can be controlled at one or more suitable levels such as picture level, slice level, CTB level, etc. In the example implementation, filtering can be further controlled at the CTB level. A flag indicating whether an ALF is applied to the luma CTB can be signaled. The luma CTB can be selected from multiple fixed filter banks (e.g., 16 fixed filter banks) and filter banks(one or more) signaled in the APS (e.g., up to 25 filters, which are derived by the encoder as described above, and also referred to as one or more signaled filter banks). Filter bank indices can be signaled for the luma CTB to indicate the filter banks to be applied (e.g., filter banks among multiple fixed filter banks and one or more signaled filter banks). Multiple fixed filter banks can be predefined and hard-coded in the encoder and decoder, and can be referred to as predefined filter banks. Therefore, it is not necessary to signal predefined filter coefficients.
[0129] For chroma components, the APS index can be signaled in the slice header to indicate the chroma filter bank to be used for the current slice. At the CTB level, if there is more than one chroma filter bank in the APS, the filter bank index can be signaled for each chroma CTB.
[0130] The filter coefficients can be quantized using a norm equal to 128. To reduce multiplication complexity, bitstream consistency can be applied so that coefficient values at non-center positions can be in the range of -27 to 27-1 (inclusive). In the example, the center position coefficients are not signaled in the bitstream, and the center position coefficients can be assumed to be equal to 128.
[0131] In some implementations, the syntax and semantics of the clipping index and clipping value are defined as follows: `alf_luma_clip_idx[sfIdx][j]` can be used to specify the clipping index to be used before multiplying with the j-th coefficient of the luminance filter indicated by signaling, as specified by `sfIdx`. Bitstream consistency requirements may include that the value of `alf_luma_clip_idx[sfIdx][j]`, where `sfIdx` = 0 to `alf_luma_num_filters_signalled_minus1` and `j` = 0 to 11, should be within, for example, a range of 0 to 3, including both 0 and 3.
[0132] The luminance filter limiting value AlfClipL[adaptation_parameter_set_id] with the element AlfClipL[adaptation_parameter_set_id][filtIdx][j] can be obtained as specified in Table 2, based on bitDepth being set equal to BitDepthY and clipIdx being set equal to alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx][j].
[0133] The `alf_chroma_clip_idx[altIdx][j]` parameter can be used to specify the limiting index of the limiting value to be used before multiplying with the j-th coefficient of the alternative chroma filter with index `altIdx`. Bitstream consistency requirements may include that the value of `alf_chroma_clip_idx[altIdx][j]` should be in the range of 0 to 3 (inclusive), where `altIdx` = 0 to `alf_chroma_num_alt_filters_minus1` and `j` = 0 to 5.
[0134] The chroma filter limiting value AlfClipC[adaptation_parameter_set_id][altIdx] with the element AlfClipC[adaptation_parameter_set_id][altIdx][j] can be obtained as specified in Table 2, based on bitDepth being set equal to BitDepthC and clipIdx being set equal to alf_chroma_clip_idx[altIdx][j], where altIdx = 0 to alf_chroma_num_alt_filters_minus1 and j = 0 to 5.
[0135] In the implementation, the filtering process can be described as follows. On the decoder side, with ALF enabled for the CTB, the samples R(i,j) in the CU (or CB) of the CTB can be filtered to obtain the filtered sample value R'(i,j) as shown below using Equation (13). In the example, each sample in the CU is filtered.
[0136] Equation (13)
[0137] Where f(k,l) represents the decoded filter coefficients, K(x, y) is the limiting function, and c(k, l) represents the decoded limiting parameter (or limiting value). The variables k and l can vary between -L / 2 and L / 2, where L represents the filter length (e.g., for each of the following values). Figure 7 Example diamond filters 710 and 711 (L = 5 and 7) are used for the luminance and chrominance components. The clipping function K(x, y) = min(y, max(-y, x)) corresponds to the clipping function Clip3(-y, y, x). By including the clipping function K(x, y), the loop filtering method (e.g., ALF) becomes a nonlinear process and can be called nonlinear ALF.
[0138] The selected limiting value can be encoded in the "alf_data" syntax element as follows: A suitable encoding scheme (e.g., the Columbus encoding scheme) can be used to encode the limiting index corresponding to the selected limiting value, such as that shown in Table 2. The encoding scheme can be the same as that used to encode the filter bank index.
[0139] In this implementation, virtual boundary filtering can be used to reduce the line buffer requirements of the ALF. Therefore, modified block classification and filtering can be applied to samples near CTU boundaries (e.g., horizontal CTU boundaries). The virtual boundary (930) can be defined by shifting the horizontal CTU boundary (920) by "N". 样本 "A line generated from a sample, such as" Figure 9A As shown, where N 样本 It can be a positive integer. In the example, for the luminance component, N 样本 It equals 4, while for the chromaticity component, N 样本 It equals 2.
[0140] Reference Figure 9A Modified block classification can be applied to the luminance components. In the example, the 1D Laplacian gradient calculation for the 4×4 block (910) above the virtual boundary (930) uses only the samples above the virtual boundary (930). Similarly, refer to Figure 9B For the 1D Laplace gradient calculation of the 4×4 block (911) below the virtual boundary (931) shifted from the CTU boundary (921), only the samples below the virtual boundary (932) are used. The quantization of the active value A can be adjusted accordingly by taking into account the reduced number of samples used in the 1D Laplace gradient calculation.
[0141] For filtering, the symmetrical fill operation at the virtual boundary can be used for both the luminance and chrominance components. Figures 10A to 10FAn example of such a modified ALF filter for the luminance component at a virtual boundary is shown. When the filtered sample is below the virtual boundary, neighboring samples above the virtual boundary are filled. When the filtered sample is above the virtual boundary, neighboring samples below the virtual boundary are filled. (See reference...) Figure 10A Neighboring sample C0 can be filled with sample C2 located below the virtual boundary (1010). (See reference...) Figure 10B Neighboring sample C0 can be filled with sample C2 located above the virtual boundary (1020). (See reference...) Figure 10C Neighboring samples C1 to C3 can be filled with samples C5 to C7 located below the virtual boundary (1030), respectively. Sample C0 can be filled with sample C6. (See reference...) Figure 10D Neighboring samples C1 to C3 can be filled with samples C5 to C7 located above the virtual boundary (1040), respectively. Sample C0 can be filled with sample C6. (See reference...) Figure 10E Neighboring samples C4 to C8 can be filled with samples C10, C11, C12, C11, and C10 located below the virtual boundary (1050), respectively. Samples C1 to C3 can be filled with samples C11, C12, and C11. Sample C0 can be filled with sample C12. (See reference...) Figure 10F Neighboring samples C4 to C8 can be filled with samples C10, C11, C12, C11, and C10 located above the virtual boundary (1060), respectively. Samples C1 to C3 can be filled with samples C11, C12, and C11. Sample C0 can be filled with sample C12.
[0142] In some examples, the above description may be adjusted appropriately when (one or more) samples and (one or more) neighboring samples are located to the left (or right) and right (or left) of the virtual boundary.
[0143] Image quadtree segmentation can be performed using a maximum coding unit (LCU) aligned quadtree. To improve coding efficiency, an adaptive loop filter based on the LCU-synchronized image quadtree can be used in video coding. In the example, the luminance image can be segmented into multiple multi-level quadtree partitions, with each partition boundary aligned with the boundary of the maximum coding unit (LCU). Each partition can have filtering processing and can therefore be called a filter unit (or filtering unit, FU).
[0144] The example of the secondary encoding process is described below. In the first step, the quadtree segmentation pattern and the optimal filter (or best filter) for each FU can be determined. During the decision process, filtering distortion can be estimated using Fast Filtering Distortion Estimation (FFDE). Based on the determined quadtree segmentation pattern and the selected filters for the FUs (e.g., all FUs), the reconstructed image can be filtered. In the second step, CU-synchronized ALF on / off control can be performed. Based on the ALF on / off result, the first filtered image is partially recovered from the reconstructed image.
[0145] A top-down segmentation strategy can be employed to divide the image into multi-level quadtree partitions using a rate-distortion criterion. Each partition can be referred to as a FU. This segmentation process can align the quadtree partitions with the LCU boundaries, such as... Figure 11 As shown. Figure 11 An example of LCU-aligned image quadtree segmentation according to an embodiment of this disclosure is shown. In the example, the encoding order of the FUs follows the z-scan order. For example, refer to... Figure 11 The image is segmented into ten FUs (e.g., FU0-FU9, with a segmentation depth of 2, where FU0, FU1, and FU9 are first-level FUs, FU2, FU7, and FU8 are second-level FUs, and FU3 to FU6 are third-level FUs), and the encoding order is from FU0 to FU9, for example, FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, and FU9.
[0146] To indicate the quadtree segmentation mode of an image, segmentation flags ("1" indicates a quadtree segmentation, while "0" indicates no quadtree segmentation) can be encoded and transmitted in z-scan order. Figure 12 The embodiments according to this disclosure are shown in correspondence. Figure 11 The quadtree partitioning pattern. For example... Figure 12 As shown in the example, the quadtree partition flags are encoded in z-scan order.
[0147] The filters for each Functional Unit (FU) can be selected from two filter banks based on a rate-distortion criterion. The first bank can have newly derived 1 / 2 symmetric square and oblique square shape filters for the current FU. The second bank can come from a delay filter buffer. The delay filter buffer can store filters previously derived for FUs in an existing image. The filter with the minimum rate-distortion cost from the two filter banks can be selected for the current FU. Similarly, if the current FU is not the minimum FU and can be further segmented into four sub-FUs, the rate-distortion cost of the four sub-FUs can be calculated. By recursively comparing the rate-distortion costs of the segmented and unsegmented cases, the quadtree segmentation pattern of the image can be determined (in other words, whether quadtree segmentation of the current FU should stop).
[0148] In some examples, the maximum quadtree split level or depth can be limited to a predefined number. For example, the maximum quadtree split level or depth could be 2, and therefore the maximum number of FUs could be 16 (or a power of 4 of the maximum depth). During the quadtree split decision, the correlation values of the Wiener coefficients used to derive the 16 FUs (minimum FUs) at the bottom quadtree level can be reused. The Wiener filters for the remaining FUs can be derived from the correlations of the 16 FUs at the bottom quadtree level. Thus, in the example, there is only one frame buffer access to derive the filter coefficients for all FUs.
[0149] After determining the quadtree segmentation pattern, to further reduce filtering distortion, CU-synchronized ALF on / off control can be implemented. By comparing filtered and non-filtered distortion, the leaf CU can explicitly switch ALF on / off in the corresponding local region. Coding efficiency can be further improved by redesigning the filter coefficients based on the ALF on / off result. In the example, the redesign process requires additional frame buffer access. Therefore, in some examples, such as in the design of a Coding Unit Synchronous PictureQuadtree-Based Adaptive Loop Filter (CS-PQALF) encoder, no redesign process is needed after the CU-synchronized ALF on / off decision to minimize the number of frame buffer accesses.
[0150] Cross-component filtering can be applied using cross-component filters, such as the Cross-Component Adaptive Loop Filter (CC-ALF). A cross-component filter can use luminance sample values of a luminance component (e.g., luminance CB) to optimize a chrominance component (e.g., the chrominance CB corresponding to the luminance CB). In the example, both the luminance CB and chrominance CB are included in the CU.
[0151] Figure 13A cross-component filter (e.g., CC-ALF) for generating chromaticity components is shown according to an example embodiment of this disclosure. For example, Figure 13 Filtering processes for a first chromaticity component (e.g., first chromaticity CB), a second chromaticity component (e.g., second chromaticity CB), and a luminance component (e.g., luminance CB) are illustrated. The luminance component can be filtered by a Sample Adaptive Offset (SAO) filter (1310) to generate a SAO-filtered luminance component (1341). The SAO-filtered luminance component (1341) can be further filtered by an ALF luminance filter (1316) to become a filtered luminance CB (1361) (e.g., "Y").
[0152] The first chromaticity component can be filtered by a SAO filter (1312) and an ALF chromaticity filter (1318) to generate a first intermediate component (1352). Furthermore, the SAO-filtered luminance component (1341) can be filtered by a cross-component filter (e.g., CC-ALF) (1321) for the first chromaticity component to generate a second intermediate component (1342). Subsequently, a filtered first chromaticity component (1362) (e.g., 'Cb') can be generated based on at least one of the second intermediate component (1342) and the first intermediate component (1352). In the example, the filtered first chromaticity component (1362) (e.g., 'Cb') can be generated by combining the second intermediate component (1342) and the first intermediate component (1352) using an adder (1322). The example cross-component adaptive loop filtering process for the first chromaticity component can therefore include steps performed by the CC-ALF (1321) and steps performed by, for example, the adder (1322).
[0153] The above description can be applied to the second chromaticity component. The second chromaticity component can be filtered by a SAO filter (1314) and an ALF chromaticity filter (1318) to generate a third intermediate component (1353). Furthermore, the SAO-filtered luminance component (1341) can be filtered by a cross-component filter (e.g., CC-ALF) (1331) for the second chromaticity component to generate a fourth intermediate component (1343). Subsequently, a filtered second chromaticity component (1363) (e.g., 'Cr') can be generated based on at least one of the fourth intermediate component (1343) and the third intermediate component (1353). In the example, the filtered second chromaticity component (1363) (e.g., 'Cr') can be generated by combining the fourth intermediate component (1343) and the third intermediate component (1353) using an adder (1332). In the example, the cross-component adaptive loop filtering process for the second chromaticity component can therefore include steps performed by CC-ALF (1331) and steps performed by, for example, an adder (1332).
[0154] Cross-component filters (e.g., CC-ALF (1321), CC-ALF (1331)) can operate by applying a linear filter with any suitable filter shape to the luma component (or luma channel) to optimize each chroma component (e.g., first chroma component, second chroma component). CC-ALF utilizes the correlation between color components to reduce coding distortion in one color component based on samples from one color component.
[0155] Figure 14 An example of a CC-ALF filter (1400) according to an embodiment of the present disclosure is shown. The filter (1400) may include non-zero filter coefficients and zero filter coefficients. The filter (1400) has a diamond shape (1420) formed by filter coefficients (1410) (indicated by circles with black fill). In the example, the non-zero filter coefficients in the filter (1400) are included in the filter coefficients (1410), and the filter coefficients not included in the filter coefficients (1410) are zero. Therefore, the non-zero filter coefficients in the filter (1400) are included in the diamond shape (1420), and the filter coefficients not included in the diamond shape (1420) are zero. In the example, the number of filter coefficients of the filter (1400) is equal to the number of filter coefficients (1410), which is in Figure 14 The example shown is 18.
[0156] CC-ALF can include any suitable filter coefficients (also known as CC-ALF filter coefficients). Return to reference Figure 13 CC-ALF (1321) and CC-ALF (1331) can have the same filter shape, for example... Figure 14 The diamond shape (1420) and the same number of filter coefficients are shown. In the example, the values of the filter coefficients in CC-ALF (1321) are different from the values of the filter coefficients in CC-ALF (1331).
[0157] Typically, filter coefficients in CC-ALF (e.g., non-zero filter coefficients derived from the encoder) can be transmitted, for example, in the APS. In the example, the filter coefficients can be factored (e.g., 2). 10Scaling is performed, and rounding can be done for fixed-point representations. CC-ALF application can be controlled for variable block sizes and is signaled via context-encoded flags (e.g., CC-ALF enable flags) received for each sample block. Context-encoded flags such as the CC-ALF enable flag can be signaled at any suitable level, such as the block level. For each chroma component, the block size and CC-ALF enable flag can be received at the slice level. In some examples, block sizes of 16×16, 32×32, and 64×64 (in chroma samples) can be supported.
[0158] In the examples, the syntax changes of CC-ALF are described in Table 3 as follows.
[0159] Table 3: Syntactic Changes in CC-ALF
[0160]
[0161] The semantics of the CC-ALF related syntax in the above examples can be described as follows:
[0162] The value of alf_ctb_cross_component_cb_idc [xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] equal to 0 indicates that the cross component Cb filter was not applied to the Cb color component sample block at the luminance location (xCtb, yCtb).
[0163] The fact that alf_ctb_cross_component_cb_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] is not equal to 0 indicates that the alf_ctb_cross_component_cb_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] cross component Cb filter is applied to the Cb color component sample block at the luminance position (xCtb, yCtb).
[0164] The value of alf_ctb_cross_component_cr_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] equal to 0 indicates that the cross component Cr filter was not applied to the Cr color component sample block at the luminance location (xCtb, yCtb).
[0165] The fact that alf_ctb_cross_component_cr_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] is not equal to 0 indicates that the alf_ctb_cross_component_cr_idc[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] cross component Cr filter is applied to the Cr color component sample block at the luminance position (xCtb, yCtb).
[0166] The following describes an example of a chroma sample format. Typically, a luma block can correspond to one or more chroma blocks, such as two chroma blocks. The number of samples in each of the (one or more) chroma blocks can be less than the number of samples in the luma block. A chroma subsampling format (also known as, for example, a chroma subsampling format specified by chroma_format_idc) can indicate the chroma horizontal subsampling factor (e.g., SubWidthC) and chroma vertical subsampling factor (e.g., SubHeightC) between each of the (one or more) chroma blocks and the corresponding luma block. For a nominal 4 (horizontal) × 4 (vertical) block, the chroma subsampling scheme can be specified as a 4:x:y format, where x is the horizontal chroma subsampling factor (the number of chroma samples retained in the first row of the block), and y is the number of chroma samples retained in the second row of the block. In the example, the chroma subsampling format could be 4:2:0, indicating that both the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 2, as shown below. Figure 15 A to Figure 15 As shown in B. In another example, the chroma subsampling format could be 4:2:2, indicating a chroma horizontal subsampling factor (e.g., SubWidthC) of 2 and a chroma vertical subsampling factor (e.g., SubHeightC) of 1. In yet another example, the chroma subsampling format could be 4:4:4, indicating both a chroma horizontal subsampling factor (e.g., SubWidthC) and a chroma vertical subsampling factor (e.g., SubHeightC) of 1. Thus, the chroma sample format or type (also known as chroma sample position) can indicate the relative position of a chroma sample in a chroma block relative to at least one corresponding luminance sample in a luminance block.
[0167] Figure 15 A to Figure 15 B illustrates an exemplary position of a chromaticity sample relative to a luminance sample according to an embodiment of this disclosure. (Refer to...) Figure 15 A, the brightness sample (1501) is located in rows (1511) to (1518). Figure 15The luminance sample (1501) shown in A can represent a portion of the image. In the example, a luminance block (e.g., luminance CB) includes luminance sample (1501). A luminance block can correspond to two chroma blocks with a chroma subsampling format of 4:2:0. In the example, each chroma block includes chroma sample (1503). Each chroma sample (e.g., chroma sample (1503(1))) corresponds to four luminance samples (e.g., luminance samples (1501(1)) to (1501(4)). In the example, these four luminance samples are the top left sample (1501(1)), the top right sample (1501(2)), the bottom left sample (1501(3)), and the bottom right sample (1501(4)). A chroma sample (e.g., (1503(1))) can be located between the top left sample (1501(1)) and the bottom left sample (1501(4)). The chromaticity sample type of the chromaticity block with chromaticity sample (1503) located at the left center position between (1501(1)) can be referred to as chromaticity sample type 0. Chromaticity sample type 0 indicates the relative position 0 corresponding to the left center position between the upper left sample (1501(1)) and the lower left sample (1501(3)). Four luminance samples (e.g., (1501(1)) to (1501(4))) can be referred to as the adjacent luminance samples of chromaticity sample (1503) (1).
[0168] In the example, each chroma block may include a chroma sample (1504). The above description of the chroma sample (1503) can be applied to the chroma sample (1504), and therefore a detailed description may be omitted for brevity. Each chroma sample in the chroma sample (1504) may be located at the center of the four corresponding luminance samples, and the chroma sample type of the chroma block having chroma samples (1504) may be referred to as chroma sample type 1. Chroma sample type 1 indicates a relative position 1 corresponding to the center position of the four luminance samples (e.g., (1501(1)) to (1501(4))). For example, one of the chroma samples (1504) may be located at the center portion of the luminance samples (1501(1)) to (1501(4)).
[0169] In the example, each chroma block includes a chroma sample (1505). Each chroma sample in the chroma sample (1505) may be located at the top-left position co-located with the top-left sample of the four corresponding luminance samples (1501), and the chroma sample type of the chroma block having the chroma sample (1505) may be referred to as chroma sample type 2. Therefore, each chroma sample in the chroma sample (1505) co-located with the top-left sample of the four luminance samples (1501) corresponding to the chroma sample. Chroma sample type 2 indicates the relative position 2 corresponding to the top-left position of the four luminance samples (1501). For example, one of the chroma samples (1505) may be located at the top-left position of luminance samples (1501(1)) to (1501(4)).
[0170] In the example, each chroma block includes a chroma sample (1506). Each chroma sample in the chroma samples (1506) may be located at the top center position between the corresponding top-left sample and the corresponding top-right sample, and the chroma sample type of the chroma block having chroma samples (1506) may be referred to as chroma sample type 3. Chroma sample type 3 indicates a relative position 3 corresponding to the top center position between the top-left sample and the top-right sample. For example, one of the chroma samples (1506) may be located at the top center position of the luminance samples (1501(1)) to (1501(4)).
[0171] In the example, each chroma block includes a chroma sample (1507). Each chroma sample in the chroma sample (1507) may be located at a lower-left position co-located with the lower-left sample in the four corresponding luminance samples (1501), and the chroma sample type of the chroma block having the chroma sample (1507) may be referred to as chroma sample type 4. Thus, each chroma sample in the chroma sample (1507) co-located with the lower-left sample in the four luminance samples (1501) corresponding to the chroma sample. Chroma sample type 4 indicates the relative position 4 corresponding to the lower-left position in the four luminance samples (1501). For example, one of the chroma samples (1507) may be located at the lower-left position of the luminance samples (1501(1)) to (1501(4)).
[0172] In the example, each chroma block includes a chroma sample (1508). Each chroma sample in the chroma sample (1508) is located at the bottom center position between the bottom left sample and the bottom right sample, and the chroma sample type of the chroma block having chroma sample (150) can be referred to as chroma sample type 5. Chroma sample type 5 indicates the relative position 5 corresponding to the bottom center position between the bottom left sample and the bottom right sample in the four luminance samples (1501). For example, one of the chroma samples (1508) can be located between the bottom left sample and the bottom right sample of the luminance samples (1501(1)) to (1501(4)).
[0173] Generally, any suitable chroma sample type can be used for chroma subsampling formats. Chroma sample types 0 through 5 provide exemplary chroma sample types described using chroma subsampling format 4:2:0. Additional chroma sample types can be used for chroma subsampling format 4:2:0. Furthermore, other chroma sample types and / or variations of chroma sample types 0 through 5 can be used for other chroma subsampling formats, such as 4:2:2, 4:4:4, etc. In the example, a chroma sample type combining chroma samples (1505) and (1507) can be used for chroma subsampling format 4:2:2.
[0174] In another example, the luminance block is considered to have alternating rows, such as rows (1511) to (1512), which respectively include the top two samples (e.g., (1501(1)) to (1501(2))) of four luminance samples (e.g., (1501(1)) to (1501(4))) and the bottom two samples (e.g., (1501(3)) to (1501(4))) of four luminance samples (e.g., (1501(1)) to (1501(4))). Therefore, rows (1511), (1513), (1515), and (1517) can be referred to as the current row (also known as the top area), and rows (1512), (1514), (1516), and (1518) can be referred to as the next row (also known as the bottom area). Four luminance samples (e.g., (1501(1)) to (1501(4))) are located in the current row (e.g., (1511)) and the next row (e.g., (1512)). The relative chromaticity positions 2 to 3 above are located in the current row, the relative chromaticity positions 0 to 1 above are located between each current row and the corresponding next row, and the relative chromaticity positions 4 to 5 above are located in the next row.
[0175] The chroma samples (1503), (1504), (1505), (1506), (1507), or (1508) are located in rows (1551) through (1554) of each chroma block. The specific position of rows (1551) through (1554) can depend on the chroma sample type of the chroma sample. For example, for chroma samples (1503) through (1504) with corresponding chroma sample types 0 through 1, row (1551) is located between rows (1511) and (1512). For chroma samples (1505) through (1506) with corresponding chroma sample types 2 through 3, row (1551) is co-located with the current row (1511). For chroma samples (1507) through (1508) with corresponding chroma sample types 4 through 5, row (1551) is co-located with the next row (1512). The above description can be appropriately applied to lines (1552) to (1554), and detailed descriptions have been omitted for the sake of brevity.
[0176] The above can be displayed, stored, and / or transmitted using any suitable scanning method. Figure 15 A describes the luminance blocks and (one or more) corresponding chrominance blocks. In some example implementations, line-by-line scanning can be used.
[0177] Alternative locations can use interlaced scanning, such as... Figure 15 As shown in B. As mentioned above, the chroma subsampling format can be 4:2:0 (e.g., chroma_format_idc equals 1). In the example, the variable chroma position type (e.g., ChromaLocType) can indicate the current line (e.g., ChromaLocType is chroma_sample_loc_type_top_field) or the next line (e.g., ChromaLocType is chroma_sample_loc_type_bottom_field). The current lines (1511), (1513), (1515), and (1517) and the next lines (1512), (1514), (1516), and (1518) can be scanned individually. For example, the current lines (1511), (1513), (1515), and (1517) can be scanned first, and then the next lines (1512), (1514), (1516), and (1518) can be scanned. The current line may include a luminance sample (1501), and the next line may include a luminance sample (1502).
[0178] Similarly, the corresponding chroma blocks can be scanned in an interlaced manner. Lines (1551) and (1553) containing chroma samples (1503), (1504), (1505), (1506), (1507), or (1508) without fill can be referred to as the current line (or current chroma line), and lines (1552) and (1554) containing chroma samples (1503), (1504), (1505), (1506), (1507), or (1508) with grayscale fill can be referred to as the next line (or next chroma line). In the example, during interlaced scanning, lines (1551) and (1553) can be scanned first, followed by lines (1552) and (1554).
[0179] In addition to the ALF mentioned above, the Constrained Directional Enhancement Filter (CDEF) can also be used for loop filtering in video coding. In-loop CDEF can be used to filter out coding artifacts such as quantization ringing artifacts while preserving image details. In some coding techniques, a similar goal can be achieved by employing the Sample Adaptive Offset (SAO) algorithm to limit signal offsets for different categories of pixels. Unlike SAO, CDEF is a nonlinear spatial filter. In some examples, the design of CDEF filters is constrained to be easily vectorized (i.e., achievable through Single Instruction Multiple Data (SIMD) operation), which is not the case for other nonlinear filters such as median filters and bilateral filters.
[0180] The CDEF design stems from the following observation: In some cases, the amount of ringing artifacts in the encoded image may be approximately proportional to the quantization step size. The minimum detail preserved in the quantized image is also proportional to the quantization step size. Therefore, preserving image detail will require a smaller quantization step size, which will result in higher levels of undesirable quantization ringing artifacts. Fortunately, for a given quantization step size, the magnitude of ringing artifacts can be smaller than the magnitude of detail, thus providing an opportunity to design a CDEF that strikes a balance between preserving sufficient detail and filtering out ringing artifacts.
[0181] CDEF can first identify the orientation of each block. Then, CDEF can adaptively filter along the identified orientation, and filter to a lesser extent along a direction rotated 45° relative to the identified orientation. The filter strength can be explicitly signaled, allowing for a high degree of control over blur details. An efficient encoder search can be designed for the filter strength. CDEF can be based on two in-loop filters, and the combined filter can be used for video coding. In some example implementations, one or more CDEF filters can follow one or more deblocking filters for in-loop filtering.
[0182] like Figure 16As shown, orientation search can be performed on the reconstructed pixels (or samples), for example, after the deblocking filter. Since the reconstructed pixels are available to the decoder, orientation may not require signaling. Orientation search can be performed on blocks of suitable size (e.g., 8×8 blocks) that are small enough to adequately handle non-linear edges (making the edges appear straight enough within the filter block) and large enough to reliably estimate the orientation when applied to a quantized image. Having a constant orientation over an 8×8 region makes vectorization of the filter easier. For each block, the orientation that best matches the pattern in the block can be determined by minimizing a difference metric between the quantized block and each of the fully oriented blocks, such as the sum of squared differences (SSD), root mean square (RMS) error, etc. In the example, the fully oriented block (e.g., Figure 16 One of (1620) refers to a block in which all pixels along a line in one direction have the same value. Figure 16 An example of directional search for an 8×8 block (1610) according to an exemplary implementation of this disclosure is shown. Figure 16 In the example shown, the 45-degree direction (1623) out of a set of directions (1620) is selected because the 45-degree direction (1623) minimizes the error (1640). For example, the error for the 45-degree direction is 12, and it is the smallest of the errors in the range of 12 to 87 indicated by row (1640).
[0183] The example nonlinear low-pass directional filter is described in more detail below. Identifying the direction can help align the filter taps along the identified direction to reduce ringing artifacts while preserving directional edges or patterns. However, in some examples, directional filtering alone is not sufficient to reduce ringing artifacts. It is desirable to use additional filter taps on pixels that are not along the main direction (e.g., the identified direction). To reduce the risk of blurring, the additional filter taps can be handled more conservatively. Therefore, a CDEF can define primary and secondary taps. In some example implementations, the complete two-dimensional (2-D) CDEF filter can be represented as:
[0184] Equation (14)
[0185] In equation (14), D represents the damping parameter, and S (p) and S (s) These represent the strengths of the primary and secondary taps, respectively, and the function round(·) rounds the relation relative to zero. and Let f(d, S, D) represent the filter weights, and let f(d, S, D) represent the constraint function that operates on the difference d (e.g., d = x(m, n) - x(i, j)) between the filtered pixel (e.g., x(i, j)) and each of its neighboring pixels (e.g., x(m, n)). When the difference is small, f(d, S, D) can be equal to the difference d (e.g., f(d, S, D) = d), and therefore the filter can behave as a linear filter. When the difference is large, f(d, S, D) can be equal to 0 (e.g., f(d, S, D) = 0), which effectively ignores the filter taps.
[0186] As another in-loop processing component, a set of in-loop restoration schemes can be used in post-encoded deblocking to substantially denoise and improve edge quality outside of the deblocking operation. This set of in-loop restoration schemes can be switched within frames (or images) according to appropriately sized tiles. Some examples of in-loop restoration schemes are described below based on separable symmetric Wiener filters and dual self-guided filters with subspace projection. Since content statistics can vary significantly within a frame, filters can be integrated within a switchable frame, where different filters can be triggered in different regions of the frame.
[0187] The following describes an example of a separable symmetric Wiener filter. The Wiener filter can be used as one of the switchable filters. Each pixel (or sample) in the degraded frame can be reconstructed as a non-causal filtered version of the pixels within a w × w window surrounding that pixel, where w = 2r + 1, and is odd for integers r. The 2D filter taps can be generated by having w... 2 The vector F is represented in column vectorized form with × 1 elements, and direct linear minimum mean square error (LMMSE) optimization can cause the filter parameters to be derived from F = H. -1 M is given, where H equals E[XX] T And it is the autocovariance of x, where x is the w × w of the area around the pixel within the window. 2 A column vectorized version of each sample, where M equals E[YX]. T ] represents the cross-correlation between x and the scalar source sample y to be estimated. The encoder can be configured to estimate H and M based on the implementation in the deblocked frame and the source, and send the resulting filter F to the decoder. However, in some example implementations, during the transmission w 2Significant bit rate costs can occur when using multiple taps. Furthermore, non-separable filtering can make decoding overly complex. Therefore, several additional constraints can be imposed on the properties of F. For example, F can be constrained to be separable, allowing filtering to be implemented as separable horizontal w-tap convolutions and vertical w-tap convolutions. In the example, each of the horizontal and vertical filters is constrained to be symmetric. Additionally, in some example implementations, it can be assumed that the sum of the horizontal and vertical filter coefficients is 1.
[0188] Dual self-guided filtering with subspace projection can also be used as one of the switchable filters for in-loop recovery and is described below. In some example implementations, guided filtering can be used for image filtering, where a local linear model is used to compute the filtered output y from an unfiltered sample x. The local linear model can be written as...
[0189] Equation (15)
[0190] Wherein, F and G can be determined based on statistical results of the downgraded image and the guidance image (also called the guide image) near the filtered pixels. If the guide image is the same as the downgraded image, the resulting self-guided filtering can have the effect of preserving smooth edges. According to some aspects of this disclosure, the specific form of the self-guided filtering can depend on two parameters: radius r and noise parameter e, and is listed below:
[0191] 1. Obtain the mean μ and variance σ of the pixels within a (2r + 1) × (2r + 1) window surrounding each pixel. 2 For example, to obtain the mean μ and variance σ of the pixels. 2 This can be effectively achieved by using box filtering based on panoramic imaging.
[0192] 2. Calculate parameters f and g for each pixel based on equation (16).
[0193] Equation (16)
[0194] 3. Calculate F and G for each pixel as the average of the values of parameters f and g in the 3×3 window around that pixel for use.
[0195] Dual self-guided filtering can be controlled by the radius r and the noise parameter e. A larger radius r implies a higher spatial variance, while a higher noise parameter e implies a higher range variance.
[0196] Figure 17An example of subspace projection according to an exemplary implementation of this disclosure is shown. Figure 17 In the example shown, subspace projection can use low-cost repairs X1 and X2 to produce a final recovered X that is closer to the source Y. f Even if the low-cost repairs X1 and X2 are not close to the source Y, with appropriate multipliers {α,β}, the low-cost repairs X1 and X2 can be moved closer to the source Y if they are moved in the correct direction. For example, the final recovery X f It can be obtained based on the following equation (17).
[0197] Equation (17)
[0198] In addition to the deblocking filter, ALF, CDEF, and loop recovery mentioned above, a loop filtering method called Cross-Component Sample Offset (CCSO) filter or CCSO can be implemented in the loop filtering process to reduce distortion of the reconstructed samples (also known as reconstruction samples). The CCSO filter can be placed anywhere in the loop filtering stage. Figure 18 The image shows an example of a CCSO filter associated with deblocking, CDEF, and LR filters. In CCSO filtering, a nonlinear mapping can be used to determine the output offset based on the processed input reconstructed sample of the first color component. In CCSO filtering, the output offset can be added to the reconstructed sample of the second color component.
[0199] The input reconstructed samples can come from the first color component located in the filter's support region, such as... Figure 19 shown. Specifically, Figure 19 An example of a filter support region in a CCSO filter according to an embodiment of this disclosure is shown. The filter support region may include four reconstructed samples: p0 and p1. Figure 19 In the example, the two input reconstructed samples are located on either side of the center sample in the vertical direction. In the example, the center sample (denoted by rl) in the first color component (e.g., the luminance component) and the sample to be filtered in the second color component (e.g., the chroma component) are co-located. The following steps can be applied when processing the input reconstructed samples:
[0200] Step 1: Calculate the increment (e.g., difference) between the four reconstructed samples p0 and p1 and the center sample rl, and denote them as m0 and m1, respectively. For example, the increment between p0 and rl is m0.
[0201] Step 2: The increment values m0 to m1 can be further quantized into multiple (e.g., two) discrete values. For example, for m0 and m1, the quantized values can be represented as d0 and d1, respectively. In the example, based on the following quantization process, the quantized value of each of d0 and d1 can be -1, 0, or 1:
[0202] di = -1, if mi < -N; Equation (18)
[0203] di = 0, if -N <= mi <= N; Equation (19)
[0204] di = 1, if mi > N. Equation (20)
[0205] Where N is the quantization step size, and example values for N are 4, 8, 12, 16, etc., and di and mi refer to the quantization value and increment value, respectively, where i is 0, 1, 2 or 3.
[0206] Quantization values d0 to d3 can be used to identify combinations of nonlinear mappings. Figure 19 In the example shown, the CCSO filter has two filter inputs d0 to d1, and each filter input can have one of three quantization values (e.g., -1, 0, and 1), thus the total number of combinations is 9 (e.g., 3). 2 (The number of differences in the number of quantized values raised to the power of the number of values). Examples of the nine combinations with offset values are shown in Table 4 below as a lookup table (LUT).
[0207] Table 4 Example LUTs used in CCSO
[0208]
[0209] The last column can represent the output offset value for each combination that can be looked up based on the increment. The output offset value can be an integer, such as 0, 1, -1, 3, -3, 5, -5, -7, etc. The first column represents the index assigned to these combinations of quantized values d0 and d1. The middle column represents all possible combinations of quantized values (with three possible quantization levels) d0 and d1. The offset column can include the actual offset value. Alternatively, a finite number of allowed offset values can exist, and the offset column in Table 4 can include indexes to the allowed offset values. Therefore, the terms offset value and index can be used interchangeably.
[0210] The final filtering process of the CCSO filter can be applied as follows:
[0211] Equation (21)
[0212] Where f is the reconstructed sample to be filtered, and s is, for example, the output offset value retrieved from the LUT. In the example shown in equation (21), the filtered sample value f' of the reconstructed sample f to be filtered can be further constrained within a range associated with the bit depth.
[0213] like Figure 19 As shown, the example CCSO filtering performed on the reconstructed sample rc of the second color corresponding to the sample c of the first color using p0 and p1 of the first color can be referred to as a 3-tap CCSO filter design. Alternatively, other CCSO designs with different numbers of filter taps can be used. For example, two additional taps can be added horizontally to make it a 5-tap filter (4 differential inputs), and the number of incremental combinations can be 81 (3...) at the same quantization level as above. 4 ).
[0214] Figure 20 Example implementations of various 3-tap CCSO filter shapes according to embodiments of this disclosure are shown. The term CCSO filter shape is used to indicate the number and position of taps used for CCSO filtering. Multiple filter shapes for a specific number of taps can be predefined as CCSO filter options. For example, in Figure 20 In this context, for a 3-tap CCSO filter, any of six different example filter shapes can be defined. Each filter shape can define the positions of three reconstructed samples (also referred to as the three taps) within the first component (also referred to as the first color component). The three reconstructed samples can include a center sample (denoted as c) and two symmetrically positioned samples, such as... Figure 20 The same number (one from 1 to 6) is used to represent them. In the example, the reconstructed sample to be filtered in the second color component is co-located with the center sample c. For clarity, in Figure 20 The reconstructed sample to be filtered in the second color component is not shown.
[0215] For CCSO filtering, as mentioned above, each filter essentially corresponds to a LUT. The number of entries (rows) in the LUT is determined by the number of increment and quantization level combinations, as shown in Table 4 above. A CCSO filter can be associated with one or more CCSO filters or filter parameters, including but not limited to filter shape, quantization step size (and quantization levels), and number of bands. CCSO filters or filter parameters may alternatively be referred to as CCSO parameters.
[0216] As described above, a quantizer is first used to process the differences or increments between the CCSO filter taps with respect to the center samples in the first color component. The quantizer can be characterized by at least a number of quantization intervals (or multiple quantization levels) and one or more quantization steps. For an example quantizer with three quantization levels, the set of quantization intervals for the sample increments can be represented as (-∞, -T1), [-T1, T2], and (T2, +∞). The increment value falling within each of these quantization intervals can be quantized to a specific quantization increment level. For example, the three quantization increment levels corresponding to the three increment quantization intervals described above can be assigned values of -1, 0, and 1, or represented by -1, 0, and 1. The quantization interval associated with the zero quantization level value can be referred to as the dead zone of the quantizer. In other words, the dead zone is the range of values assigned to zero after quantization is applied. For example, the dead zone of the three-interval quantizer described above, where T1 and T2 are positive, is represented by [-T1, T2].
[0217] An example three-level quantizer is shown in equations (18) through (20) above. The example quantizers in equations (18) through (20) are shown as examples of symmetric quantizers, where T1 = T2 = T in the general three-level quantization partition above. In some example implementations, the quantizer can be asymmetric. For a general asymmetric three-level quantizer, T1 will not be equal to -T2.
[0218] While asymmetric quantizers with three or more levels can be arbitrarily determined by the encoder and signaled in the bitstream (e.g., including all information to indicate or derive each quantization interval), in some example implementations, they can be specified as deriveable by applying a non-zero offset to a predefined or signaled symmetric quantizer. For example, for the aforementioned three-interval symmetric quantizer with quantization intervals of (-∞, -T), [-T, T], and (T, +∞), a corresponding three-level asymmetric quantizer can be obtained, implemented, or derived by adding an offset F to the quantization intervals of the symmetric quantizer, such as (-∞, -T+F), [-T+F, T+F], and (T+F, +∞). The resulting quantizer is asymmetric because its quantization intervals are no longer symmetric relative to the zero-sample increment.
[0219] In some example implementations, multiple symmetric and asymmetric quantizers can be predefined. In other words, each of the multiple symmetric or asymmetric quantizers can have a predefined set of quantization intervals, which can be known to both the encoder and decoder. These predefined quantizers can be indexed or associated with quantizer identifiers, such that each of them can be appropriately identified or referenced in the encoder or decoder. In some example implementations, symmetric and asymmetric quantizers can be separately indexed as two distinct groups identified by distinguishable group identifiers. In some alternative implementations, symmetric and asymmetric quantizers can be indexed or identified as a single group.
[0220] In some example implementations, multiple symmetric quantizers can be associated with a variety of different numbers of quantization levels and / or quantization intervals. For example, the number of quantization levels (or quantization intervals) of a symmetric quantizer can be two, three, four, or five. The quantization intervals, except for the two end intervals, can have equal quantization steps.
[0221] In some example implementations, the encoder may determine a selection from multiple symmetric and asymmetric quantizers for use and include that selection (e.g., the index associated with the selected quantizer) in the bitstream. The encoder's selection of quantizers can be made at the encoding level, including but not limited to sequence level, picture level, slice level, tile level, superblock level, and code block level. Therefore, the selection can be signaled correspondingly in one or more high-level syntaxes at one or more locations such as sequence headers, picture headers, slice headers, tile headers, superblock headers, and code block headers. If a single set of identifier indices is used for both symmetric and asymmetric quantizers (in other words, when multiple predefined quantizers include a mixed set of symmetric and asymmetric quantizers), a single index or identifier can be signaled in the bitstream to indicate the selection. If the symmetric and asymmetric quantizers are indexed separately, the signaling for selection can first indicate whether a symmetric or asymmetric quantizer is selected, and then subsequently indicate the index of the selected symmetric or asymmetric quantizer. Alternatively, one can first signal the selection of an asymmetric quantizer, and then signal the index of the corresponding asymmetric quantizer.
[0222] In some example implementations, an asymmetric quantizer can be associated with a symmetric quantizer among a plurality of predefined quantizers. For example, as mentioned above, a predefined asymmetric quantizer can be associated with another predefined symmetric quantizer via a predefined offset.
[0223] For example, the sets of predefined asymmetric quantizers and predefined symmetric quantizers can have a one-to-one correlation with predefined offsets. Thus, the number of predefined symmetric quantizers can be equal to the number of predefined asymmetric quantizers, and the encoder's selection of a particular symmetric quantizer can be signaled by an index in a predefined index of the predefined symmetric quantizer set. The encoder's selection of a particular asymmetric quantizer can be signaled by including the index of the corresponding symmetric quantizer in the predefined index of the symmetric quantizer set in the bitstream, followed by or before an additional indication of the selection of the asymmetric quantizer. The decoder can then be able to determine the appropriate predefined asymmetric quantizer to apply (its offset relative to the corresponding symmetric quantizer with respect to the predefined offset).
[0224] For another example, a predefined set of asymmetric quantizers can be associated with predefined symmetric quantizers with predefined offsets, but this may not be a one-to-one correspondence. In other words, the number of predefined symmetric quantizers and the number of predefined asymmetric quantizers may not be equal. For example, a predefined symmetric quantizer can be associated with one or more corresponding predefined offsets and one or more asymmetric quantizers. In some example implementations, there can be N predefined offsets, and each predefined symmetric quantizer can correspond to N asymmetric quantizers. The predefined symmetric quantizers can be the encoder's selection of them and the indexes signaled to them. For example, the encoder's selection of asymmetric quantizers can be achieved by signaling the index of the corresponding symmetric quantizer, followed by or before the selection of offsets (which can be the index of an offset among the N predefined offsets, or it can be direct signaling of the offset amount).
[0225] In some of the other example implementations described above, two or more predefined symmetric quantizers can correspond to different numbers of asymmetric quantizers. For example, a symmetric quantizer with a large dead zone can correspond to a large number of predefined asymmetric quantizers. Thus, the selection of an asymmetric quantizer can be signaled by indicating its corresponding symmetric quantizer and then using an index for selection from its corresponding asymmetric quantizer. Such an index can be signaled in a manner dependent on the symmetric quantizer (because the number of bits used to signal the selection of the asymmetric quantizer can be different for different corresponding symmetric quantizers).
[0226] In some of the other implementation examples mentioned above, not every predefined symmetric quantizer corresponds to one or more asymmetric quantizers. In other words, some predefined symmetric quantizers may not correspond to any predefined asymmetric quantizers. Thus, when such a symmetric quantizer is signaled in the bitstream, it means that the symmetric quantizer has been selected, and no additional syntax for selecting asymmetric quantizers is included in the bitstream. In other words, asymmetric quantizers are signaled only for a subset of symmetric quantizers. This improves signaling efficiency. For a specific example, a predefined symmetric quantizer may not correspond to any asymmetric quantizer when its dead zone is smaller than a predefined threshold. In other words, one or more asymmetric quantizers corresponding to a symmetric quantizer may be predefined and can be additionally signaled only when the dead zone size of the symmetric quantizer is greater than a given threshold.
[0227] In some of the other example implementations described above, multiple symmetric and asymmetric quantizers are predefined, and the selection of one of them is signaled at various levels in the bitstream. The asymmetric quantizer is implemented by adding a non-zero offset value to the symmetric quantizer, and the offset values used to derive the asymmetric quantizer can depend on the corresponding symmetric quantizer. For example, a larger offset value can be used to derive the asymmetric quantizer for a symmetric quantizer with a larger dead-time size. In some other example implementations, the offset values used to derive the asymmetric quantizer can depend on the codec's internal bit depth or loop filtering. For example, a deeper bit depth results in a larger offset (e.g., a symmetric quantizer with different dead-time sizes or other properties can correspond to an asymmetric quantizer with different offsets).
[0228] In some example implementations, the number of possible offsets relative to a predefined symmetric quantizer that form an asymmetric quantizer can depend on the corresponding symmetric quantizer. For example, for a symmetric quantizer with a large dead zone size, a larger number of possible offset values can be used to derive the asymmetric quantizer.
[0229] In some of the other example implementations described above, an asymmetric quantizer can be derived on top of the selected symmetric quantizer. In some other example implementations, the selection of the symmetric quantizer is signaled before the selection of the asymmetric quantizer. In still other example implementations, the number of applicable offset values used to derive the asymmetric quantizer is different for the different symmetric quantizers used to derive the asymmetric quantizer.
[0230] In some example implementations, multiple symmetric quantizers can be predefined and known to both the encoder and decoder. The encoder can determine which symmetric quantizer to select for use at various coding levels and notify this via signaling in the bitstream. For example, as described above, multiple symmetric quantizers can be indexed, and the index of the selected symmetric quantizer can be signaled. If an asymmetric quantizer is to be applied, an indicator can be signaled after or before the signaling of the symmetric quantizer, and the offset can also be signaled. In some example implementations, multiple offsets can be predefined, and an index can be signaled to indicate which offset to apply to the signaled symmetric quantizer. In some other example implementations, the actual offset can be signaled. In some example implementations, the offset can be predefined, and therefore only an indicator indicating whether an asymmetric quantizer should be applied is signaled, without signaling the offset. In some other implementations, a predefined set of derivation rules can be provided, allowing an asymmetric quantizer to be derived from a signaled symmetric quantizer selected by the encoder. For example, the rule set can be derived based on the dead zone size or quantization interval size of the symmetric quantizer, as notified by signaling. For instance, a predefined rule set can specify the offset as a function of the dead zone size or quantization interval size of the symmetric quantizer.
[0231] In some example implementations, all available CCSO quantizers can be signaled using high-level syntax. For example, all allowed symmetric quantizers can be signaled, such as by signaling the identifiers of allowed symmetric quantizers in a predefined set of quantizers and indexing these identifiers for selection within the bitstream. Furthermore, when selecting an asymmetric quantizer in the bitstream, the allowed offsets for each allowed symmetric quantizer can be signaled and indexed for reference.
[0232] In the various implementations above, a CCSO filter unit refers to a segment of reconstructed video frames of any size that has undergone processing by a CCSO filter.
[0233] Figure 21Flowchart 2100 is shown. The logical flow 2100 begins at S2101. At S2110, a frame is reconstructed from the video bitstream to generate reconstructed samples of at least a first color component and a second color component of the frame. In S2120, a quantizer for quantizing the CCSO filter unit of the frame is obtained from among a plurality of quantizers based on a first syntax element notified by signaling in the video bitstream. The plurality of quantizers includes at least one asymmetric CCSO quantizer, which is asymmetric with respect to the values quantized on the positive and negative sides of zero. At S2130, a quantizer is selected according to at least the first syntax element. In S2140, the selected CCSO quantizer is applied to the CCSO sample increment of the reconstructed sample of the first color component of the CCSO filter unit according to the CCSO filter to generate a quantized CCSO sample increment. In S2150, a CCSO sample offset is determined based on the quantized CCSO sample increment and the CCSO filter. In S2160, the CCSO sample offset is applied to the reconstructed sample of the second color component in the CCSO filter unit to generate a filtered sample of the second color component in the CCSO filter unit. Logic flow 2100 ends at S2199.
[0234] The above operations can be combined or arranged in any number or order as needed. Two or more steps and / or operations can be performed in parallel. The embodiments and implementations in this disclosure can be used individually or in combination in any order. Furthermore, each of the method (or implementation), encoder, and decoder can be implemented by a processing circuit system (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments in this disclosure can be applied to luma blocks or chroma blocks. The term "block" can be interpreted as a prediction block, coding block, or coding unit, i.e., CU. The term "block" here can also be used to refer to a transform block. In the following, when referring to block size, it can refer to block width or block height, or the maximum value of width and height, or the minimum value of width and height, or area size (width * height), or the aspect ratio of the block (width:height, or height:width).
[0235] The above-mentioned techniques can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 22 A computer system (2200) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0236] Computer software can be coded using any suitable machine code or computer language. Machine code or computer language can be subjected to mechanisms such as assembly, compilation, and linking to create code that includes instructions. These instructions can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.
[0237] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0238] Figure 22 The components shown for the computer system (2200) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any one or a combination of components shown in the exemplary embodiments of the computer system (2200).
[0239] The computer system (2200) may include certain human-machine interface input devices. The input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard (2201), mouse (2202), touchpad (2203), touch screen (2210), data glove (not shown), joystick (2205), microphone (2206), scanner (2207), and camera device (2208).
[0240] The computer system (2200) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback via a touchscreen (2210), data gloves (not shown), or joystick (2205), but tactile feedback devices that are not used as input devices may also exist); audio output devices (e.g., speakers (2209), headphones (not depicted)); visual output devices (e.g., screens (2210), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability—some of which may be capable of outputting two-dimensional visual output or outputting more than three dimensions in a manner such as stereoscopic image output; virtual reality glasses (not depicted); holographic displays and ashtrays (not depicted)); and printers (not depicted).
[0241] The computer system (2200) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2220) having media such as CD / DVD (2221), thumb drives (2222), removable hard disk drives or solid-state drives (2223), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0242] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0243] The computer system (2200) may also include interfaces (2254) to one or more communication networks (2255). The networks may be, for example, wireless, wired, or optical. The networks may also be local area, wide area, metropolitan area, vehicular and industrial networks, real-time, latency-tolerant, etc. Examples of networks include: local area networks such as Ethernet; wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular networks including CAN bus and industrial networks, etc.
[0244] The human-machine interface device, human-accessible storage device and network interface mentioned above can be attached to the core (2240) of the computer system (2200).
[0245] The core (2240) may include one or more central processing units (CPUs) (2241), graphics processing units (GPUs) (2242), dedicated programmable processing units in the form of field-programmable gate arrays (FPGAs) (2243), hardware accelerators (2244) for certain tasks, graphics adapters (2250), etc. These devices, along with read-only memory (ROMs) (2245), random access memory (2246), and internal mass storage devices such as internal non-user-accessible hard disk drives, SSDs, etc. (2247), may be connected via a system bus (2248). In some computer systems, the system bus (2248) may be accessed as one or more physical connectors to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (2249) to the core's system bus (2248). In the example, a screen (2210) may be connected to a graphics adapter (2250). Peripheral bus architectures include PCI, USB, etc.
[0246] Computer-readable media may contain computer code for performing operations of various computer implementations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be media and computer code of a type known and available to those skilled in the art of computer software.
[0247] While this disclosure has described several exemplary embodiments, variations, substitutions, and various alternative equivalents fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art can conceive of many systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are thus within its spirit and scope.
Claims
1. A method for video decoding, comprising: Frames are reconstructed from the video bitstream to generate reconstructed samples of at least the first and second color components of the frames; A quantizer for quantizing the CCSO filter unit of the frame is obtained from a plurality of quantizers based on a first syntax element notified by signaling in the video bitstream, the plurality of quantizers including at least one asymmetric quantizer, the at least one asymmetric quantizer being asymmetric with respect to the values quantized on the positive and negative sides of zero. Select the quantizer indicated by the first syntax element; The CCSO sample increment is generated by applying the selected quantizer to the reconstructed sample of the first color component of the CCSO filter unit according to the CCSO filter. The CCSO sample offset is determined based on the quantized CCSO sample increment and the CCSO filter; as well as The CCSO sample offset is applied to the reconstructed sample of the second color component in the CCSO filtering unit to generate a filtered sample of the second color component in the CCSO filtering unit.
2. The method according to claim 1, wherein, Each of the plurality of quantizers is characterized by one or more quantization intervals.
3. The method according to claim 2, wherein, Each of the at least one asymmetric quantizer corresponds to one of the symmetric quantizers among the plurality of quantizers via a CCSO sample increment offset, and each of the symmetric quantizers is symmetric with respect to zero sample increment.
4. The method according to claim 3, wherein, The plurality of CCSO quantizers are predefined and indexed, and the selected CCSO quantizer is signaled by the first syntax element in the video bitstream.
5. The method according to claim 3, wherein, The plurality of quantizers includes a subset of symmetric quantizers with a one-to-one correspondence and a subset of asymmetric quantizers.
6. The method according to claim 5, wherein, The symmetric CCSO quantizer subset and the asymmetric CCSO quantizer subset are correlated by a predefined sample increment offset.
7. The method according to claim 6, wherein, The selection of an asymmetric quantizer is indicated in the video bitstream by a second syntax element and by a first syntax element, the second syntax element being used to signal the indication to use an asymmetric quantizer, and the first syntax element indicating the symmetric quantizer in the subset of symmetric quantizers corresponding to the selected asymmetric CCSO quantizer.
8. The method according to claim 7, wherein, The first syntax element includes an index for the symmetric quantizer corresponding to the selected asymmetric quantizer.
9. The method according to claim 3, wherein, The plurality of quantizers includes unequal subsets of symmetric quantizers and subsets of asymmetric quantizers.
10. The method according to claim 9, wherein, Each symmetric quantizer in the subset of symmetric quantizers corresponds to zero or more asymmetric quantizers through one or more sample increment offsets from a predefined set of sample increment offsets.
11. The method according to claim 10, wherein, The selection of an asymmetric quantizer is indicated in the video bitstream by a first syntax element and a second syntax element, the first syntax element indicating a symmetric quantizer corresponding to the selected asymmetric quantizer, and the second syntax element indicating an offset for applying the symmetric quantizer indicated by the first syntax element to generate the selected asymmetric quantizer.
12. The method according to claim 10, wherein, The selection of an asymmetric quantizer is indicated in the video bitstream by a first syntax element and a second syntax element, the first syntax element indicating a symmetric quantizer corresponding to the selected asymmetric quantizer, and the second syntax element indicating an index among the asymmetric quantizers corresponding to the symmetric quantizer indicated by the first syntax element.
13. The method according to claim 10, wherein, The number of asymmetric quantizers corresponding to symmetric quantizers depends non-decreasingly on the size of the quantization interval of the symmetric quantizer.
14. The method according to claim 9, wherein, At least one of the symmetric quantizer subsets does not correspond to any of the asymmetric quantizer subsets.
15. The method according to claim 9, wherein, When the dead zone of a symmetric quantizer in the subset of symmetric quantizers is at or below a threshold, the symmetric quantizer does not correspond to any of the asymmetric quantizers in the subset. The dead zone represents the quantization interval containing the zero-sample increment of the symmetric quantizer.
16. The method according to any one of claims 1 to 15, wherein, The first syntax element is signaled at the sequence header, picture header, tile header, or maximum coded block header level.
17. The method according to any one of claims 1 to 15, wherein, The first syntax element is used to select an asymmetric quantizer from the plurality of quantizers by: indicating a symmetric quantizer, deriving a sample increment offset based on the indicated symmetric quantizer, and generating the selected asymmetric quantizer by applying the resulting sample increment offset to the indicated symmetric quantizer.
18. The method according to claim 17, wherein, The sample increment offset is determined based on the size of the dead zone of the indicated symmetric quantizer, where the dead zone represents the quantization interval containing the zero sample increment of the indicated symmetric quantizer.
19. The method according to any one of claims 1 to 15, wherein, The quantizer selected based on the first syntax element includes: The selected symmetric quantizer is determined based on the first syntax element; Determine the size of the dead zone of the selected symmetric quantizer, where the dead zone represents the quantization interval containing the zero-sample increment of the selected symmetric quantizer; In response to the dead zone size of the selected symmetric quantizer being less than a predefined threshold, it is determined which symmetric quantizer to apply; and In response to the fact that the dead zone size of the selected symmetric quantizer is not less than the predefined threshold, it is determined whether to select an asymmetric quantizer based on the second syntax element in the video bitstream.
20. An apparatus for filtering a video bitstream, comprising a memory for storing instructions and a processor for executing the instructions to perform the following operations: Frames are reconstructed from the video bitstream to generate reconstructed samples of at least the first and second color components of the frames; A quantizer for quantizing the CCSO filter unit of the frame is obtained from a plurality of quantizers based on a first syntax element notified by signaling in the video bitstream, the plurality of quantizers including at least one asymmetric quantizer, the at least one asymmetric quantizer being asymmetric with respect to the values quantized on the positive and negative sides of zero. The quantizer is selected based on the first syntax element; The CCSO sample increment is generated by applying the selected quantizer to the reconstructed sample of the first color component of the CCSO filter unit according to the CCSO filter. The CCSO sample offset is determined based on the quantized CCSO sample increment and the CCSO filter; as well as The CCSO sample offset is applied to the reconstructed sample of the second color component in the CCSO filtering unit to generate a filtered sample of the second color component in the CCSO filtering unit.