Filtering method and apparatus in video decoding

By using filters with nonlinear mapping to switch filter shape configurations in video coding, intra-frame prediction and motion compensation are optimized, solving the problem of insufficient efficiency in reducing redundancy in existing video coding technologies, and achieving more efficient video data compression and decoding performance.

CN115104303BActive Publication Date: 2026-02-10TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180014369.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-06
Filing Date
2021-08-03
Publication Date
2026-02-10
Estimated Expiration
2041-08-03

AI Technical Summary

Technical Problem

Existing video coding technologies suffer from insufficient efficiency in reducing redundancy during intra-frame prediction and motion compensation, especially when processing complex video content. This makes it difficult to effectively compress video data, resulting in high bandwidth and storage requirements.

Method used

A filter based on nonlinear mapping is used to reconstruct video samples by switching different filter shape configurations. Cross-component sampling offset (CCSO) and local sampling offset (LSO) filters are used to optimize the intra-frame prediction and motion compensation process and reduce redundant information.

Benefits of technology

It improves the compression efficiency of video encoding, reduces bandwidth and storage requirements, and enhances the encoding quality and decoding performance of video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115104303B_ABST
    Figure CN115104303B_ABST
Patent Text Reader

Abstract

Aspects of the disclosure provide methods and apparatuses for video coding / decoding. In some examples, an apparatus for video decoding includes processing circuitry. The processing circuitry reconstructs first samples in a video carried in a coded video bitstream according to a non-linear mapping based filter having a first filter shape configuration. The processing circuitry then determines to switch from the first filter shape configuration to a second filter shape configuration, and reconstructs second samples in the video according to a non-linear mapping based filter having the second filter shape configuration.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application No. 17 / 368,734, filed July 6, 2021, entitled “METHOD AND APPARATUS FOR VIDEOFILTERING”, which claims priority to U.S. Provisional Application No. 63 / 122,780, filed December 8, 2020, entitled “IMPROVED FILTER SHAPE FOR SAMPLE OFFSET”. The entire disclosure of the earlier application is incorporated herein by reference. Technical Field

[0003] This disclosure describes embodiments that generally involve video coding. Background Technology

[0004] The background description provided herein is for the purpose of presenting the general content of this disclosure. Within the scope described in this background section, neither the work of the currently named inventor nor any aspect of this description that does not qualify as prior art at the time of submission is expressly or impliedly acknowledged as prior art to this disclosure.

[0005] Video encoding and decoding can be performed using inter-frame picture prediction with motion compensation. Uncompressed digital video can comprise a series of pictures, each with a spatial dimension of, for example, 1920x1080 luma samples and associated chroma samples. This series of pictures can have a fixed or variable picture rate (also informally referred to as the frame rate), such as 60 pictures per second or 60Hz. Uncompressed video has very high bitrate requirements. For example, at 8 bits per sample, 1080p60 4:2:0 video (with a 1920x1080 luma sample resolution at a 60Hz frame rate) requires close to 1.5 Gbit / s of bandwidth. One hour of such video would require more than 600 gigabytes of storage space.

[0006] One objective of video encoding and decoding is to reduce redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by two orders of magnitude or more. Lossless compression, lossy compression, and combinations thereof can be employed. Lossless compression refers to a technique that reconstructs an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may differ from the original signal, but the distortion between the original and reconstructed signals is small enough that the reconstructed signal can be used for the intended application. In the case of video, lossy compression is widely used. The tolerable amount of distortion depends on the application; for example, users of some consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can be reflected in the fact that higher permissible / acceptable distortion can result in a higher compression ratio.

[0007] Video encoders and decoders can employ techniques from several broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.

[0008] Video codec techniques can include a variety of techniques known as intra-frame coding. In intra-frame coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference images. In some video codecs, images are spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, the image can be an intra-frame image. Intra-frame images and their derived images (e.g., independent decoder refresh images) can be used to reset the decoder state and thus can be used as the first image in the encoded video bitstream and video session, or as a still image. Samples of an intra-frame block can be transformed, and the transform coefficients can be quantized before entropy coding. Intra-frame prediction can be a technique that minimizes the sample values ​​in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are needed to represent the entropy-coded block at a given quantization step size.

[0009] Traditional intra-frame coding (such as intra-frame coding known from techniques like MPEG-2 generation coding) does not use intra-frame prediction. However, some newer video compression techniques include attempts to use data based on, for example, surrounding sample data and / or metadata, obtained during the encoding / decoding of spatially adjacent and earlier-in-the-order data blocks. Such techniques are referred to below as "intra-frame prediction" techniques. It is noteworthy that, at least in some cases, intra-frame prediction uses only reference data from the current image being reconstructed, rather than from a reference image.

[0010] There can be many different forms of intra-prediction. When more than one such technique is used in a given video coding technique, the techniques used can be encoded in an intra-prediction mode. In some cases, a mode can have multiple sub-modes and / or multiple parameters, and these sub-modes and parameters can be encoded individually or included in the mode codeword. Which codeword is used for a given combination of mode / sub-mode / parameters can affect the coding efficiency gain through intra-prediction, and the same applies to entropy coding techniques used to convert codewords into bitstreams.

[0011] H.264 introduced an intra-prediction mode, which was improved in H.265 and further refined in newer coding techniques such as Joint Exploration Model (JEM), Universal Video Coding (VVC), and Baseline Matrix (BMS). Prediction blocks can be formed using neighboring sample values ​​that belong to already available samples. The sample values ​​of neighboring samples are copied into the prediction block according to the direction. The reference to the direction being used can be encoded in the bitstream or it can be predicted itself.

[0012] Referring to Figure 1A, a subset of nine known prediction directions from the 33 possible prediction directions of H.265 (corresponding to 33 angular modes of 35 intra-frame modes) is depicted in the lower right of Figure 1A. The point (101) where the arrows converge represents the sample being predicted. The arrows indicate the direction of the predicted sample. For example, arrow (102) indicates that sample (101) is predicted based on one or more samples to the upper right at a 45-degree angle to the horizontal direction. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more samples to the lower left of sample (101) at a 22.5-degree angle to the horizontal direction.

[0013] Referring again to Figure 1A, a 4×4 square block (104) of samples is depicted in the upper left of Figure 1A (represented by a bold dashed line). The square block (104) comprises 16 samples, each labeled “S”, with its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block size is 4×4 samples, sample S44 is located in the lower right. Reference samples following a similar numbering scheme are also shown. The reference samples are labeled R, with their Y position (e.g., row index) and X position (column index) relative to the block (104). In H.264 and H.265, the predicted samples are adjacent to the block being reconstructed; therefore, negative values ​​are not required.

[0014] Intra-frame image prediction works by copying reference sample values ​​from adjacent samples occupied by the prediction direction indicated by a signal. For example, suppose the encoded video bitstream includes signaling indicating a prediction direction consistent with arrow (102) for that block, i.e., predicting samples based on one or more prediction samples at a 45-degree angle to the upper right of the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Then sample S44 is predicted based on reference sample R08.

[0015] In some cases, especially when the predicted direction is not divisible by 45 degrees, the values ​​of multiple reference samples can be combined, for example, by interpolation, to calculate the reference sample.

[0016] With the development of video coding technology, the number of possible predicted directions has increased. In H.264 (2003), nine different predicted directions could be represented. This increased to 33 in H.265 (2013), while JEM / VVC / BMS, when publicly released, could support up to 65 directions. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those possible directions with a small number of bits, thus penalizing less likely directions. Furthermore, sometimes the direction itself can be predicted based on adjacent directions used in adjacent decoded blocks.

[0017] Figure 1B shows a schematic diagram (180) depicting 65 intra-frame predicted directions according to JEM, to illustrate the number of predicted directions increasing over time.

[0018] The mapping of intra-prediction direction bits representing direction in an encoded video bitstream can vary depending on the video coding technique, and this mapping can range from, for example, a simple direct mapping from prediction direction to intra-prediction mode, to codeword, to complex adaptive schemes involving the most probable mode (and similar techniques). However, in all cases, some directions are statistically less likely to appear in the video content compared to others. Since the goal of video compression is to reduce redundancy, in well-functioning video coding techniques, those less likely directions will require more bits to represent than the more likely directions.

[0019] Motion compensation can be a lossy compression technique and can involve using blocks of sample data from a previously reconstructed image or a portion thereof (the reference image) spatially offset along a direction indicated by a motion vector (hereinafter referred to as MV) to predict the newly reconstructed image or a portion thereof. In some cases, the reference image can be the same as the image currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference image being used (the latter could indirectly be a temporal dimension).

[0020] In some video compression techniques, the MV applicable to certain regions of sample data can be predicted based on other MVs (e.g., those MVs that are spatially adjacent to the region being reconstructed and whose decoding order precedes that MV). Doing so can significantly reduce the amount of data required to encode MVs, thereby eliminating redundancy and increasing compression ratio. MV prediction can work effectively, for example, because when encoding the input video signal obtained from the camera (called natural video), there is a statistical probability that a larger region than the region applicable to a single MV moves in similar directions. Therefore, in some cases, the larger region can be predicted using similar motion vectors derived from the MVs of neighboring regions. This makes the MV found for a given region similar to or the same as the MV predicted based on the surrounding MVs, and thus, after entropy coding, the MV found for the given region can be represented with fewer bits than when directly encoding the MV. In some cases, MV prediction can be an example of lossless compression of the signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when calculating predictions based on multiple surrounding MVs.

[0021] H.265 / HEVC (ITU-T H.265 Recommendation, “High Efficiency Video Coding”, December 2016) describes various MV prediction mechanisms. Among the various MV prediction mechanisms provided by H.265, this application describes the technique hereinafter referred to as “spatial combining”.

[0022] Referring to Figure 2, the current block (201) includes samples that have been discovered by the encoder during the motion search process, and these samples can be predicted based on previous blocks of the same size that have generated spatial offsets. Alternatively, the MV can be derived from metadata associated with one or more reference images, rather than being directly encoded. For example, using the MV associated with any of the five surrounding samples A0, A1, and B0, B1, B2 (corresponding to 202 to 206 respectively), the MV can be derived from metadata associated with one or more reference images (from the nearest reference image in decoding order). In H.265, MV prediction can use predictions from the same reference images that are also being used in adjacent blocks. Summary of the Invention

[0023] Various aspects of this disclosure provide methods and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes processing circuitry. The processing circuitry reconstructs a first sample of video carried in an encoded video bitstream based on a filter with a first filter shape configuration and a nonlinear mapping. The processing circuitry then determines to switch from the first filter shape configuration to a second filter shape configuration and reconstructs a second sample of video based on a filter with the second filter shape configuration and a nonlinear mapping.

[0024] In one embodiment, the filter based on the nonlinear mapping is a cross-component sample offset (CCSO) filter. In another embodiment, the filter based on the nonlinear mapping is a local sample offset (LSO) filter.

[0025] According to one aspect of this disclosure, the difference between the first filter shape configuration and the second filter shape configuration lies at least in: the geometry of the filter tap position; and the distance from the filter tap position to the center of the filter tap position.

[0026] In some examples, both the first filter shape configuration and the second filter shape configuration have at least one of a cross geometry at the filter tap location and a rectangular geometry at the filter tap location.

[0027] In one example, the first filter shape configuration and the second filter shape configuration have the same geometry, and the difference between the first filter shape configuration and the second filter shape configuration is the distance from the filter tap position to the center of the filter tap position.

[0028] In some examples, the processing circuitry decodes an index based on the encoded video bitstream carrying the video, the index indicating a second filter shape configuration. The processing circuitry then determines, based on the index, to switch from the first filter shape configuration to the second filter shape configuration. In one example, the processing circuitry determines the switch from the first filter shape configuration to the second filter shape configuration at the picture level. The first sample is located in the first picture of the video, and the second sample is located in the second picture of the video.

[0029] In another example, the processing circuitry determines to switch from a first filter shape configuration to a second filter shape configuration at the block level. The first sample is located in a first block of images of the video, and the second sample may be located in a second block of images of the video.

[0030] In some examples, the processing circuitry decodes the index based on syntax signaling from at least one of the block-level, video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), slice header, tile header, and frame header.

[0031] In some examples, to reconstruct a first sample in a video based on a filter with a first filter shape configuration using a nonlinear mapping, the processing circuitry may perform preprocessing operations on samples located at filter tap positions corresponding to the first filter shape configuration to generate preprocessed samples, and determine an offset to be applied to the first sample based on the preprocessed samples. In one example, the processing circuitry calculates the average sample value located at two or more filter tap positions as the preprocessed sample. In another example, the processing circuitry applies a filter to the samples located at the filter tap positions to generate filtered samples as the preprocessed sample.

[0032] Various aspects of this disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform any method for video decoding. Attached Figure Description

[0033] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0034] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, wherein:

[0035] Figure 1A is a schematic diagram of an exemplary subset of intra-prediction modes.

[0036] Figure 1B is a schematic diagram of an exemplary intra-frame prediction direction.

[0037] Figure 2 is a schematic diagram of the current block and its surrounding space merging candidates based on an example.

[0038] Figure 3 This is a simplified block diagram of a communication system according to one embodiment.

[0039] Figure 4 This is a simplified block diagram of a communication system according to another embodiment.

[0040] Figure 5 This is a simplified block diagram of a decoder according to one embodiment.

[0041] Figure 6 This is a simplified block diagram of an encoder according to one embodiment.

[0042] Figure 7 A block diagram of an encoder according to another embodiment is shown.

[0043] Figure 8 A block diagram of a decoder according to another embodiment is shown.

[0044] Figure 9 An example of a filter shape according to an embodiment of the present disclosure is shown.

[0045] Figures 10A to 10D An example of a subsampling location for calculating a gradient according to an embodiment of the present disclosure is shown.

[0046] Figure 11A and Figure 11B An example of a virtual boundary filtering process according to an embodiment of the present disclosure is shown.

[0047] Figures 12A to 12F An example of a symmetrical filling operation at a virtual boundary according to an embodiment of the present disclosure is shown.

[0048] Figure 13 Examples of image segmentation according to some embodiments of the present disclosure are shown.

[0049] Figure 14 The quadtree splitting pattern of the image is shown in some examples.

[0050] Figure 15 A cross-component filter according to an embodiment of the present disclosure is shown.

[0051] Figure 16 An example of a filter shape according to an embodiment of the present disclosure is shown.

[0052] Figure 17 Syntax examples of cross-component filters according to some embodiments of this disclosure are shown.

[0053] Figure 18A and Figure 18B An exemplary position of a chromaticity sample relative to a luminance sample according to an embodiment of the present disclosure is shown.

[0054] Figure 19 An example of directional search according to one embodiment of this disclosure is shown.

[0055] Figure 20 Examples illustrating subspace projection are shown in some examples.

[0056] Figure 21 A table of multiple sample adaptive offset (SAG) types according to one embodiment of the present disclosure is shown.

[0057] Figure 22 Examples of patterns for pixel classification in edge offsets are shown in some examples.

[0058] Figure 23 A table showing pixel classification rules for edge offsets is provided in some examples.

[0059] Figure 24 An example of a syntax that can be represented by signals is shown.

[0060] Figure 25 Examples of filter-supported regions according to some embodiments of this disclosure are shown.

[0061] Figure 26 An example of another filter support region according to some embodiments of this disclosure is shown.

[0062] Figures 27A to 27C A table with 81 combinations is shown according to one embodiment of the present disclosure.

[0063] Figure 28 An example of a filter shape configuration according to an embodiment of the present disclosure is shown.

[0064] Figure 29 Another example of a filter shape configuration according to an embodiment of the present disclosure is shown.

[0065] Figure 30 Another example of a filter shape configuration according to an embodiment of the present disclosure is shown.

[0066] Figure 31 Another example of a filter shape configuration according to an embodiment of the present disclosure is shown.

[0067] Figure 32 An example of three candidate filter shape configurations with intersecting geometry is shown.

[0068] Figure 33 An example of two candidate filter shape configurations with intersecting geometry is shown.

[0069] Figure 34 An example of two candidate filter shape configurations with rectangular geometry is shown.

[0070] Figure 35 Examples of four candidate filter shape configurations with a mixture of intersecting and rectangular geometries are shown.

[0071] Figure 36 An example of preprocessing according to an embodiment of this disclosure is shown.

[0072] Figure 37 Another example of preprocessing according to an embodiment of this disclosure is shown.

[0073] Figure 38 A flowchart outlining a process according to one embodiment of the present disclosure is shown.

[0074] Figure 39 This is a schematic diagram of a computer system according to one embodiment. Detailed Implementation

[0075] Figure 3 This is a simplified block diagram of a communication system (300) according to an embodiment disclosed in this application. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first terminal device (310) and a second terminal device (320) interconnected via a network (350). Figure 3 In this embodiment, the first terminal device (310) and the second terminal device (320) perform unidirectional data transmission. For example, the first terminal device (310) may encode video data (e.g., a video image stream captured by the terminal device (310)) for transmission over a network (350) to the second terminal device (320). The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to recover the video images, and display the video images based on the recovered video data. Unidirectional data transmission is common in applications such as media services.

[0076] In another embodiment, the communication system (300) includes a third terminal device (330) and a fourth terminal device (340) that perform bidirectional transmission of encoded video data, which may occur, for example, during a video conference. For bidirectional data transmission, in one example, each of the third terminal device (330) and the fourth terminal device (340) may encode video data (e.g., a stream of video images captured by the terminal device) for transmission over a network (350) to the other terminal device. Each of the third terminal device (330) and the fourth terminal device (340) may also receive encoded video data transmitted by the other terminal device and may decode the encoded video data to recover video images, and may display the video images on an accessible display device based on the recovered video data.

[0077] exist Figure 3 In the examples, the first terminal device (310), the second terminal device (320), the third terminal device (330), and the fourth terminal device (340) may be servers, personal computers, and smartphones, but the principles disclosed herein are not limited thereto. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (350) refers to any number of networks that transmit encoded video data between the first terminal device (310), the second terminal device (320), the third terminal device (330), and the fourth terminal device (340), including, for example, wired (connected) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switched and / or packet-switched channels. The network may include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of the network (350) may be irrelevant to the operation of this application.

[0078] As an example of the application of the disclosed subject matter Figure 4 The diagram illustrates the placement of a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0079] The streaming system may include an acquisition subsystem (413) that may include a video source (401) such as a digital camera, which creates an uncompressed video image stream (402). In an embodiment, the video image stream (402) includes samples captured by a digital camera. The video image stream (402) is depicted as a thick line to emphasize the high data volume compared to encoded video data (404) (or encoded video bitstream), and the video image stream (402) may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. Compared to the video image stream (402), the encoded video data (404) (or the encoded video bitstream (404)) is depicted as a thin line to emphasize the lower data volume. This encoded video data can be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as Figure 4 Client subsystems (406) and (408) can access a streaming server (405) to retrieve copies (407) and (409) of encoded video data (404). Client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and produces an output video picture stream (411) that can be displayed on a display (412) (e.g., a screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (404), video data (407), and video data (409) (e.g., video bitstream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In embodiments, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application can be used in the context of the VVC standard.

[0080] It should be noted that electronic devices (420) and (430) may include other components (not shown). For example, electronic device (420) may include a video decoder (not shown), and electronic device (430) may also include a video encoder (not shown).

[0081] Figure 5 This is a block diagram of a video decoder (510) according to an embodiment disclosed in this application. The video decoder (510) may be disposed in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., receiving circuitry). The video decoder (510) may be used in place of... Figure 3 The video decoder (410) in the embodiment.

[0082] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510); in the same embodiment or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not indicated). The receiver (531) may separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be located external to the video decoder (510) (not indicated). In other cases, an external buffer (not shown) may be provided for the video decoder (510) to prevent network jitter, for example, and another buffer (515) may be configured internally for, for example, handling broadcast timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer (515) may not be necessary, or it may be made smaller. Of course, for use on packet networks such as the Internet, a buffer (515) may be required; this buffer may be relatively large and adaptive in size, and may be at least partially implemented in the operating system or a similar component (not shown) external to the video decoder (510).

[0083] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (510) and potential information for controlling a display device (512) (e.g., a display screen), which is not part of the electronic device (530) but may be coupled to it, such as... Figure 5As shown in the diagram. The control information for the display device may be a parameter set fragment (not shown) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (520) may parse / decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set of at least one subgroup of pixels in the subgroup of pixels for use in the video decoder based on at least one parameter corresponding to a group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (520) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0084] The parser (520) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0085] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (521) may involve multiple different units. Which units are involved and how they are involved can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (520). For brevity, the flow of such subgroup control information between the parser (520) and the various units described below is not described.

[0086] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0087] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantization transform coefficients as symbols (521) from the parser (520) and control information, including the transform method used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block containing sample values, which can be input into the aggregator (555).

[0088] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images but can use predictive information from previously reconstructed portions of the current image. Such predictive information may be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses reconstructed information extracted from the current picture buffer (558) to generate surrounding blocks of the same size and shape as the block being reconstructed. For example, the current picture buffer (558) buffers partially reconstructed and / or fully reconstructed current images. In some cases, the aggregator (555) adds the predictive information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) based on each sample.

[0089] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (553) can access the reference image memory (557) to extract samples for prediction. After motion compensation is performed on the extracted samples according to the symbol (521), these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (referred to as residual samples or residual signals in this case) to generate output sample information. The motion compensation prediction unit (553) can obtain the predicted samples from the address in the reference image memory (557) under motion vector control, and the motion vector is available to the motion compensation prediction unit (553) in the form of the symbol (521), which, for example, includes X, Y and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory (557) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0090] The output samples of the aggregator (555) can be employed by various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video stream), and these parameters can be used as symbols (521) from the parser (520) in the loop filter unit (556). However, in other embodiments, the video compression techniques may also respond to metadata obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0091] The output of the loop filter unit (556) can be a sample stream, which can be output to the display device (512) and stored in the reference image memory (557) for subsequent inter-frame image prediction.

[0092] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded images corresponding to the current image have been fully reconstructed and the encoded images (by, for example, the parser (520)) are identified as reference images, the current image buffer (558) can become part of the reference image memory (557), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.

[0093] The video decoder (510) can perform decoding operations according to a predetermined video compression technique, such as that specified in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technique or standard as the only tools available under that configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the limits defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.

[0094] In this embodiment, the receiver (531) may receive supplemental (redundant) data along with the encoded video. This supplemental data may be a portion of the encoded video sequence. The supplemental data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The supplemental data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0095] Figure 6 This is a block diagram of a video encoder (603) according to an embodiment disclosed in this application. The video encoder (603) is disposed in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used to replace... Figure 4 The video encoder (403) in the embodiment.

[0096] The video encoder (603) can obtain data from the video source (601) (not) Figure 6 In one embodiment, a portion of the electronic device (620) receives video samples, the video source being capable of capturing video images to be encoded by a video encoder (603). In another embodiment, the video source (601) is a portion of the electronic device (620).

[0097] A video source (601) can provide a sequence of source video samples to be encoded by a video encoder (603) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) can be a storage device storing previously prepared video. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc., used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0098] According to an embodiment, the video encoder (603) can encode and compress images of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units described below. For simplicity, coupling is not shown in the figures. Parameters set by the controller (650) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be used with other suitable functions related to the video encoder (603) optimized for a particular system design.

[0099] In some embodiments, the video encoder (603) operates within an encoding loop. For simplicity, in an embodiment, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (633) embedded within the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (because in the video compression techniques considered in this application, any compression between the symbols and the encoded video stream is lossless). The reconstructed sample stream (sample data) is input to a reference image memory (634). Since decoding of the symbol stream produces bit-precise results independent of the decoder's location (local or remote), the contents of the reference image memory (634) also correspond bit-precisely between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction portion are exactly the same sample values ​​that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.

[0100] The operation of the “local” decoder (633) can be combined with, for example, the above-mentioned Figure 5 The video decoder (510) is described in detail as the same as the "remote" decoder. However, a further brief reference is provided. Figure 5 When symbols are available and the entropy encoder (645) and parser (520) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (510), including the buffer (515) and parser (520), may not be fully implemented in the local decoder (633).

[0101] It can be observed that any decoder technique other than parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in essentially the same functional form. For this reason, this application focuses on decoder operation. The description of encoder techniques can be simplified because encoder techniques are inverses of the fully described decoder techniques. More detailed descriptions are only required in certain areas, and are provided below.

[0102] During operation, in some embodiments, the source encoder (630) may perform motion-compensated predictive coding. This motion-compensated predictive coding predictively encodes the input image, referencing one or more previously encoded images from the video sequence designated as "reference images." In this manner, the encoding engine (632) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.

[0103] The local video decoder (633) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (630). The operation of the encoding engine (632) can be a lossy process. When the encoded video data can be decoded by the video decoder (633), Figure 6 When the source video sequence (not shown) is decoded, the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in a reference image cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0104] The predictor (635) can perform a prediction search against the encoding engine (632). That is, for a new image to be encoded, the predictor (635) can search in the reference image memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (635) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, based on the search results obtained by the predictor (635), it can be determined that the input image may have prediction references obtained from multiple reference images stored in the reference image memory (634).

[0105] The controller (650) can manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0106] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding, thereby converting the symbols into an encoded video sequence.

[0107] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission via a communication channel (660), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0108] The controller (650) manages the operation of the video encoder (603). During encoding, the controller (650) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types:

[0109] An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will understand variations of I-pictures and their corresponding applications and characteristics.

[0110] A predictive picture (P-picture) can be a picture that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block.

[0111] A bidirectional predictive picture (B-picture) can be a picture that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values ​​for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0112] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and each block is encoded sequentially. These blocks can be predictively coded with reference to other (already coded) blocks, determined by the coding assignment of the corresponding images applied to the block. For example, a block of an I-image can be non-predictively coded, or it can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. A pixel block of a P-image can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. A block of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.

[0113] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0114] In this embodiment, the transmitter (640) may transmit additional data while transmitting encoded video. The source encoder (630) may include such data as part of the encoded video sequence. Additional data may include temporal / spatial / SNR enhancement layers, redundant images and slices, other forms of redundant data, SEI messages, VUI parameter set fragments, etc.

[0115] The acquired video can be presented as multiple source images (video images) in a time-series format. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In an embodiment, a specific image being encoded / decoded is segmented into blocks, referred to as the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when multiple reference images are used, the motion vector may have a third dimension that identifies the reference image.

[0116] In some embodiments, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. Specifically, the block can be predicted using a combination of the first and second reference blocks.

[0117] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0118] According to some embodiments disclosed in this application, predictions such as inter-frame image prediction and intra-frame image prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a video image sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the images have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Furthermore, each CTU can be further subdivided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be subdivided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In embodiments, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. Furthermore, depending on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In embodiments, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Taking a luma prediction block as an example, a prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0119] Figure 7 This is a diagram of a video encoder (703) according to another embodiment disclosed in this application. The video encoder (703) is used to receive a processing block (e.g., a prediction block) of sample values ​​within a current video image in a video image sequence, and to encode the processing block into an encoded image that is part of an encoded video sequence. In this embodiment, the video encoder (703) is used instead of Figure 4The video encoder (403) in the embodiment.

[0120] In the HEVC embodiment, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as an 8×8 sample prediction block. The video encoder (703) uses, for example, rate-distortion (RD) optimization to determine whether to use intra-frame mode, inter-frame mode, or bidirectional prediction mode to encode the processing block. When encoding the processing block in intra-frame mode, the video encoder (703) can use intra-frame prediction techniques to encode the processing block into an already encoded picture; and when encoding the processing block in inter-frame mode or bidirectional prediction mode, the video encoder (703) can use inter-frame prediction or bidirectional prediction techniques to encode the processing block into an already encoded picture, respectively. In some video coding techniques, the merging mode can be an inter-frame picture prediction sub-mode, in which motion vectors are derived from one or more motion vector prediction values ​​without relying on already encoded motion vector components outside the prediction values. In some other video coding techniques, motion vector components applicable to the subject block may exist. In the embodiment, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the processing block mode.

[0121] exist Figure 7 In one embodiment, the video encoder (703) includes, as shown below: Figure 7 The inter-frame encoder (730), intra-frame encoder (722), residual calculator (723), switch (726), residual encoder (724), general controller (721) and entropy encoder (725) are shown coupled together.

[0122] The inter-frame encoder (730) is configured to receive samples of the current block (e.g., the processing block), compare the block with one or more reference blocks in a reference image (e.g., blocks in previous and later images), generate inter-frame prediction information (e.g., redundancy information description, motion vectors, and merging mode information based on the inter-frame prediction information), and calculate inter-frame prediction results (e.g., predicted blocks) using any suitable technique based on the inter-frame prediction information. In some embodiments, the reference image is a decoded reference image based on encoded video information.

[0123] The intra encoder (722) is used to receive samples of the current block (e.g., the processing block), in some cases compare the block with encoded blocks in the same image, generate quantization coefficients after transformation, and in some cases also (e.g., based on intra prediction direction information of one or more intra coding techniques) generate intra prediction information. In an embodiment, the intra encoder (722) also calculates intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same image.

[0124] A general controller (721) determines general control data and controls other components of the video encoder (703) based on this general control data. In an embodiment, the general controller (721) determines the mode of a block and provides control signals to a switch (726) based on this mode. For example, when the mode is an intra-frame mode, the general controller (721) controls the switch (726) to select an intra-frame mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select intra-frame prediction information and add the intra-frame prediction information to the bitstream; and when the mode is an inter-frame mode, the general controller (721) controls the switch (726) to select an inter-frame prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select inter-frame prediction information and add the inter-frame prediction information to the bitstream.

[0125] A residual calculator (723) is used to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). A residual encoder (724) is used to operate on the residual data to encode the residual data to generate transform coefficients. In an embodiment, the residual encoder (724) is used to transform the residual data from the time domain to the frequency domain and generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is used to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter-frame prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra-frame prediction information. The decoded blocks are processed appropriately to generate a decoded image, and in some embodiments, the decoded image may be buffered in a memory circuit (not shown) and used as a reference image.

[0126] An entropy encoder (725) is used to format the bitstream to produce encoded blocks. The entropy encoder (725) generates various information according to a suitable standard such as the HEVC standard. In an embodiment, the entropy encoder (725) is used to obtain general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. It should be noted that, according to the disclosed subject matter, residual information is not present when blocks are encoded in a merged sub-mode of inter-frame mode or bidirectional prediction mode.

[0127] Figure 8This is a diagram of a video decoder (810) according to another embodiment disclosed in this application. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In the embodiment, the video decoder (810) is used instead of Figure 4 The video decoder (410) in the embodiment.

[0128] exist Figure 8 In the embodiment, the video decoder (810) includes, as follows: Figure 7 The entropy decoder (871), inter-frame decoder (880), residual decoder (873), reconstruction module (874), and intra-frame decoder (872) are shown coupled together.

[0129] An entropy decoder (871) can be used to reconstruct certain symbols from an encoded image, these symbols representing the syntax elements constituting the encoded image. Such symbols may include, for example, a mode used to encode the block (e.g., intra-frame mode, inter-frame mode, bidirectional prediction mode, a combined sub-mode of the latter two, or another sub-mode), prediction information (e.g., intra-frame prediction information or inter-frame prediction information) that can respectively identify certain samples or metadata used by the intra-frame decoder (872) or the inter-frame decoder (880) for prediction, residual information in the form of, for example, quantized transform coefficients, and so on. In an embodiment, when the prediction mode is inter-frame or bidirectional prediction mode, inter-frame prediction information is provided to the inter-frame decoder (880); and when the prediction type is intra-frame prediction type, intra-frame prediction information is provided to the intra-frame decoder (872). Residual information may be provided to the residual decoder (873) via inverse quantization.

[0130] The inter-frame decoder (880) is used to receive inter-frame prediction information and generate inter-frame prediction results based on the inter-frame prediction information.

[0131] The intra-frame decoder (872) is used to receive intra-frame prediction information and generate prediction results based on the intra-frame prediction information.

[0132] The residual decoder (873) performs inverse quantization to extract the dequantized transform coefficients and processes the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require some control information (to obtain the quantizer parameter (QP)), which can be provided by the entropy decoder (871) (the data path is not indicated because this is only low-level control information).

[0133] The reconstruction module (874) combines the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, which may be part of a reconstructed image, which in turn may be part of a reconstructed video. It should be noted that other suitable operations, such as deblocking, may be performed to improve visual quality.

[0134] It should be noted that any suitable technology can be used to implement the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810). In one embodiment, one or more integrated circuits can be used to implement the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810). In another embodiment, one or more processors executing software instructions can be used to implement the video encoder (403), video encoder (603), and video encoder (603), as well as the video decoder (410), video decoder (510), and video decoder (810).

[0135] Various aspects of this disclosure provide filtering techniques for video encoding / decoding.

[0136] The encoder / decoder can apply an adaptive loop filter (ALF) with block-based filter adaptation to reduce artifacts. For the luma component, for example, based on the direction and activity of the local gradient, one of several filters (e.g., 25 filters) can be selected for a 4×4 luma block.

[0137] ALFs can have any suitable shape and size. (Reference) Figure 9 ALF(910)-(911) have a rhombus shape; for example, ALF(910) has a 5×5 rhombus shape, and ALF(911) has a 7×7 rhombus shape. In ALF(910), elements (920)-(932) form a rhombus shape and can be used in the filtering process. Seven values ​​(e.g., C0-C6) are available for elements (920)-(932). In ALF(911), elements (940)-(964) form a rhombus shape and can be used in the filtering process. Thirteen values ​​(e.g., C0-C12) are available for elements (940)-(964).

[0138] refer to Figure 9In some examples, two ALFs (910)-(911) with diamond filter shapes are used. A 5×5 diamond filter (910) can be applied to the chroma component (e.g., chroma block, chroma CB), and a 7×7 diamond filter (911) can be applied to the luma component (e.g., luma block, luma CB). Other suitable shapes and sizes can be used in the ALF. For example, a 9×9 diamond filter can be used.

[0139] Filter coefficients located at positions indicated by values ​​(e.g., C0-C6 in (910) or C0-C12 in (920)) can be nonzero. Furthermore, when the ALF includes a clipping function, the clipping values ​​at these positions can be nonzero.

[0140] For block classification of the luminance component, 4×4 blocks (or luminance blocks, luminance CB) can be classified or grouped into one of several (e.g., 25) categories. This can be based on the quantized values ​​of the orientation parameter D and the activity value A. Use equation (1) to derive the category index C.

[0141]

[0142] To calculate the direction parameter D and the quantization value The gradients g in the vertical direction, horizontal direction, and two diagonal directions (e.g., d1 and d2) can be calculated using the 1-D Laplace as shown below. v g h g d1 And gd2.

[0143]

[0144]

[0145]

[0146]

[0147] Here, indices i and j refer to the coordinates of the top-left sample within the 4×4 block, and R(k, l) indicates the reconstructed sample at coordinates (k, l). Directions (e.g., d1 and d2) can refer to two diagonal directions.

[0148] To reduce the complexity of the above block classification, subsampling 1-D Laplace calculation can be applied. Figures 10A to 10D The diagram shows the methods for calculating the vertical direction ( Figure 10A ), horizontal direction ( Figure 10B ) and two diagonal directions d1( Figure 10C ) and d2( Figure 10D The gradient g of ) v gh g d1 and g d2 Examples of subsampling locations. The same subsampling location can be used for gradient calculation in different directions. Figure 10A In the diagram, the 'V' symbol indicates the calculation of the vertical gradient g. v The subsampling location. Figure 10B In the diagram, the label 'H' indicates the calculation of the horizontal gradient g. h The subsampling location. Figure 10C In the diagram, the label 'D1' indicates the calculation of the diagonal gradient g of d1. d1 The subsampling location. Figure 10D In the diagram, the label 'D2' indicates the calculation of the diagonal gradient g of d2. d2 The sub-sampling position.

[0149] Horizontal gradient g h and the vertical gradient g v maximum value and minimum value It can be set to:

[0150]

[0151] Gradients g in two diagonal directions d1 and g d2 maximum value and minimum value It can be set to:

[0152]

[0153] The direction parameter D can be derived as follows based on the above values ​​and two thresholds t1 and t2.

[0154] Step 1: If (1) And (2) If true, then D is set to 0.

[0155] Step 2: If If yes, continue to step 3; otherwise, continue to step 4.

[0156] Step 3: If Then D is set to 2; otherwise, D is set to 1.

[0157] Step 4: If Then D is set to 4; otherwise, D is set to 3.

[0158] The activity value A can be calculated as:

[0159]

[0160] A can be further quantized to the range of 0 to 4 (inclusive), and the quantized value is represented as

[0161] Block classification is not applied to the chroma components in the image, so a single ALF coefficient set can be applied to each chroma component.

[0162] Geometric transformations can be applied to filter coefficients and corresponding filter clipping values ​​(also known as trim values). Before filtering a block (e.g., a 4×4 brightness block), these transformations are performed, for example, based on gradient values ​​calculated for the block (e.g., g). v g h g d1 and / or g d2 Geometric transformations, such as rotation or diagonal and vertical flipping, can be applied to the filter coefficients f(k, l) and the corresponding filter clipping values ​​c(k, l). Applying geometric transformations to the filter coefficients f(k, l) and the corresponding filter clipping values ​​c(k, l) is equivalent to applying the geometric transformation to samples within the region supported by the filter. Geometric transformations can make different blocks applied by ALF more similar by aligning the corresponding orientations.

[0163] The three geometric transformations, including diagonal flip, vertical flip, and rotation, can be performed as described by equations (9)-(11).

[0164] f D (k,l)=f(l,k),c D (k,l)=c(l,k) Equation (9)

[0165] f V (k,l)=f(k,Kl-1),c V (k,l)=c(k,Kl-1) Equation (10)

[0166] f R (k,l)=f(Kl-1,k),c R (k,l)=c(Kl-1,k) Equation (11)

[0167] Where K is the size of the ALF or filter, and 0 ≤ k, l ≤ K⁻¹ are the coordinates of the coefficients. For example, position (0, 0) is at the top left corner, and position (K⁻¹, K⁻¹) is at the bottom right corner of the filter f or the clipping matrix (or clipping matrix) c. The transformation can be applied to the filter coefficients f(k, l) and the clipping values ​​c(k, l) based on the gradient values ​​computed for the block. Examples of the relationship between the transformation and the four gradients are summarized in Table 1.

[0168] Table 1: Mapping of gradients and transformations for block computation

[0169] gradient value Transformation <![CDATA[g d2 <g d1 And g h <g v ]]> No transformation <![CDATA[g d2 <g d1 And g v <g h ]]> Flip diagonally <![CDATA[g d1 <g d2 And g h <g v ]]> Vertical flip <![CDATA[g d1 <g d2 And g v <g h ]]> Rotation

[0170] In some embodiments, ALF filter parameters are represented by signals in the Adaptive Parameter Set (APS) of the image. In the APS, one or more sets (e.g., up to 25 sets) of luminance filter coefficients and crop value indices can be represented by signals. In one example, one or more sets may include luminance filter coefficients and one or more crop value indices. One or more sets (e.g., up to 8 sets) of chrominance filter coefficients and crop value indices can be represented by signals. To reduce signaling overhead, filter coefficients for different classifications of the luminance component (e.g., with different classification indices) can be combined. In the slice header, the index of the APS used for the current slice can be represented by signals.

[0171] In one embodiment, the clipping value index (also referred to as the clipping index) can be decoded according to the APS. The clipping value index can be used, for example, to determine the corresponding clipping value based on the relationship between the clipping value index and the corresponding clipping value. The relationship can be predefined and stored in the decoder. In one example, the relationship is described by tables such as a luminance table (e.g., for luminance CB) of clipping value index and corresponding clipping value, and a chrominance table (e.g., for chrominance CB) of clipping value index and corresponding clipping value. The clipping value may depend on the bit depth B. The bit depth B may refer to the internal bit depth, the bit depth of the reconstructed sample in the CB to be filtered, etc. In some examples, the tables (e.g., luminance table, chrominance table) are obtained using equation (12).

[0172]

[0173] Where AlfClip is the clipping value, B is the bit depth (e.g., bitDepth), N (e.g., N=4) is the number of allowed clipping values, and (n-1) is the clipping value index (also called the clipping index or clipIdx). Table 2 shows an example of the table obtained using equation (12) when N=4. In Table 2, the clipping index (n-1) can be 0, 1, 2, and 3, and n can be 1, 2, 3, and 4. Table 2 can be used for luma blocks or chroma blocks.

[0174] Table 2 - AlfClip can depend on bit depth B and clipIdx

[0175]

[0176]

[0177] In the slice header of the current slice, one or more APS indices (e.g., up to seven APS indices) can be represented by signals to specify the set of luma filters available for the current slice. The filtering process can be controlled at one or more appropriate levels (e.g., picture level, slice level, CTB level, etc.). In one embodiment, the filtering process can be further controlled at the CTB level. Flags can be represented by signals to indicate whether the ALF is applied to the luma CTB. The luma CTB can select a filter set from multiple fixed filter sets (e.g., 16 fixed filter sets) and a filter set represented by signals in the APS (also referred to as a filter set represented by signals). For the luma CTB, filter set indices can be represented by signals to indicate the filter set to be applied (e.g., a filter set among multiple fixed filter sets and filter sets represented by signals). Multiple fixed filter sets can be predefined and hard-coded in the encoder and decoder, and can be referred to as predefined filter sets.

[0178] For chroma components, the APS index can be represented by a signal in the slice header to indicate the chroma filter set to be used for the current slice. At the CTB level, if there is more than one chroma filter set in the APS, the filter set index can be represented by a signal for each chroma CTB.

[0179] Filter coefficients can be quantized using a norm equal to 128. To reduce multiplication complexity, bitstream compliance can be applied, allowing coefficient values ​​at non-center positions to be in the range of -27 to 27-1 (inclusive). In one example, center position coefficients are not represented by signals in the bitstream, and the center position coefficients can be considered equal to 128.

[0180] In some embodiments, the syntax and semantics of pruning indexes and pruning values ​​are defined as follows:

[0181] `alf_luma_clip_idx[sfIdx][j]` can be used to specify the clipping index of the clipping value used before multiplying with the j-th coefficient of the luminance filter represented by the signal, which is represented by `sfIdx`. Bitstream compliance requirements may include that the value of `alf_luma_clip_idx[sfIdx][j]` should be in the range of 0 to 3 (inclusive) when `sfIdx` = 0 to `alf_luma_num_filter_signal_minus` 1 and `j` = 0 to 11.

[0182] Based on bitDepth set to equal BitDepthY and clipIdx set to equal alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx]][j], the luminance filter clipping value AlfClipL[adaption_parameter_set_id][filtIdx][j] can be derived as specified in Table 2, where filtIdx = 0 to NumAlfFilters-1 and j = 0 to 11. alf_chroma_clip_idx[altIdx][j] can be used to specify the clipping index of the clipping value used before multiplying with the j-th coefficient of the alternative chroma filter with index altIdx. Bitstream compliance requirements may include: the value of alf_chroma_clip_idx[altIdx][j] should be in the range of 0 to 3 (inclusive) when altIdx = 0 to alf_chroma_num_alt_filters_minus 1 and j = 0 to 5.

[0183] Based on bitDepth set to equal BitDepthC and clipIdx set to equal alf_chroma_clip_idx[altIdx]][j], the chroma filter clipping value AlfClipC[adaption_parameter_set_id][altIdx][j] can be derived as specified in Table 2, where altIdx = 0 to alf_chroma_num_alt_filters_minus 1 and j = 0 to 5.

[0184] In one embodiment, the filtering process may be described as follows. On the decoder side, when ALF is enabled for CTB, samples R(i,j) within the CU (or CB) may be filtered to produce filtered sample values ​​R'(i,j) as shown in equation (13). In one example, each sample in the CU is filtered.

[0185]

[0186] Where f(k, l) represents the decoded filter coefficients, K(x, y) is the clipping function, and c(k, l) indicates the decoded clipping parameters (or clipping values). The variables k and l can vary between -L / 2 and L / 2, where L indicates the filter length. The clipping function K(x, y) = min(y, max(-y, x)), corresponding to the clipping function Clip3(-y, y, x). By incorporating the clipping function K(x, y), the loop filtering method (e.g., ALF) becomes a nonlinear process and can be called nonlinear ALF.

[0187] In nonlinear ALF, multiple clipping value sets can be provided as shown in Table 3. In one example, the luma set includes four clipping values ​​{1024, 181, 32, 6}, and the chroma set includes four clipping values ​​{1024, 161, 25, 4}. The four clipping values ​​in the luma set can be selected from the full range (e.g., 1024) of sample values ​​(encoded in 10 bits) that are approximately equally divided in the logarithmic domain into luma blocks. For the chroma set, this range can be from 4 to 1024.

[0188] Table 3 - Examples of clipping values

[0189] Intra-frame / Inter-frame tile groups brightness {1024,181,32,6} chromaticity {1024,161,25,4}

[0190] The selected clipping values ​​can be encoded in the "alf_data" syntax element as follows: A suitable encoding scheme (e.g., the Golomb encoding scheme) can be used to encode the clipping indices corresponding to the clipping values ​​selected as shown in Table 3. The encoding scheme can be the same one used to encode the filter set indices.

[0191] In one embodiment, a virtual boundary filtering process can be used to reduce the line buffer requirements of the ALF. Therefore, modified block classification and filtering can be applied to samples close to the CTU boundary (e.g., horizontal CTU boundary). Figure 11A As shown, the virtual boundary (1130) can be moved by "N" by shifting the horizontal CTU boundary (1120). samples "N samples are used to define a line, where N" samples It can be a positive integer. In one example, for the luminance component, N samples It equals 4, and for the chromaticity components, N samples It equals 2.

[0192] refer to Figure 11A Modified block classification can be applied to the luminance components. In one example, for the 1D Laplacian gradient calculation of the 4×4 block (1110) above the virtual boundary (1130), only samples above the virtual boundary (1130) are used. Similarly, refer to... Figure 11BFor the 1D Laplace gradient calculation of the 4×4 block (1111) below the virtual boundary (1131) shifted from the CTU boundary (1121), only the samples below the virtual boundary (1131) are used. Therefore, the quantization size of the activity value A can be adjusted by considering reducing the number of samples used in the 1D Laplace gradient calculation.

[0193] For filtering, the symmetrical fill operation at the virtual boundary can be used for both the luminance and chrominance components. Figures 12A to 12F An example of this modified ALF filter being used for the luminance component at a virtual boundary is shown. When the filtered sample is below the virtual boundary, adjacent samples above the virtual boundary can be filled. When the filtered sample is above the virtual boundary, adjacent samples below the virtual boundary can be filled. (Reference) Figure 12A The adjacent sample C0 can be filled using sample C2 located below the virtual boundary (1210). (See reference...) Figure 12B The adjacent sample C0 can be filled using the sample C2 located above the virtual boundary (1220). (See reference...) Figure 12C The adjacent samples C1-C3 can be filled using samples C5-C7 located below the virtual boundary (1230). (See reference) Figure 12D The adjacent samples C1-C3 can be filled using samples C5-C7 located above the virtual boundary (1240). (See reference) Figure 12E The adjacent samples C4-C8 can be filled using samples C10, C11, C12, C11, and C10 located below the virtual boundary (1250). (See reference) Figure 12F The adjacent samples C4-C8 can be filled by using samples C10, C11, C12, C11 and C10 located above the virtual boundary (1260).

[0194] In some examples, the above description may be adjusted appropriately when the sample and its neighboring sample are located to the left (or right) and right (or left) of the virtual boundary.

[0195] According to one aspect of this disclosure, to improve coding efficiency, images can be segmented based on a filtering process. In some examples, the CTU is also referred to as the Maximum Coding Unit (LCU). In one example, the CTU or LCU may have a size of 64×64 pixels. In some embodiments, LCU-aligned image quadtree splitting can be used for filtering-based segmentation. In some examples, an adaptive loop filter based on synchronizing the image quadtree with the coding units can be used. For example, a luminance image can be split into several multi-level quadtree segments, with each segment boundary aligned with the boundary of the LCU. Each segment has its own filtering process and is therefore called a filter unit (FU).

[0196] In some examples, a two-pass coded stream can be used. At the first pass of the two-pass coded stream, the quadtree splitting pattern of the image and the optimal filter for each FU can be determined. In some embodiments, the determination of the quadtree splitting pattern of the image and the determination of the optimal filter for each FU are based on filtering distortion. During the determination process, filtering distortion can be estimated using the Fast Filter Distortion Estimation (FFDE) technique. The image is segmented using quadtree segmentation. Based on the determined quadtree splitting pattern and the selected filters for all FUs, the reconstructed image can be filtered.

[0197] At the second pass of the 2-pass coded stream, CU synchronous ALF on / off control is performed. Based on the ALF on / off result, the first filtered image is partially recovered from the reconstructed image.

[0198] Specifically, in some examples, a top-down splitting strategy is employed, dividing the image into multi-level quadtree segments based on a rate-distortion criterion. Each segment is called a filter unit (FU). The splitting process aligns the quadtree segments with the LCU boundaries. The encoding order of the FUs follows a z-scan order.

[0199] Figure 13 Examples of segmentation according to some embodiments of this disclosure are shown. Figure 13 In the example, the image (1300) is split into 10 FUs, and the encoding order is FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8 and FU9.

[0200] Figure 14 A quadtree splitting pattern (1400) is shown for image (1300). Figure 14 In the examples, the split flag is used to indicate the image segmentation mode. For example, "1" indicates that a quadtree split is performed on the block; "0" indicates that the block is not further segmented. In some examples, the smallest FU has an LCU size, and the smallest FU does not require a split flag. Figure 14 As shown, the splitting flags are encoded and transmitted in z-order.

[0201] In some examples, filters for each FU are selected from two filter sets based on a rate-distortion criterion. The first set contains 1 / 2-symmetric square and diamond filters derived for the current FU. The second set comes from a time-delayed filter buffer. The time-delayed filter buffer stores filters previously derived for FUs of previous images. The filter with the minimum rate-distortion cost from these two sets can be selected for the current FU. Similarly, if the current FU is not the minimum FU and can be further split into 4 sub-FUs, the rate-distortion cost of the 4 sub-FUs is calculated. The image quadtree splitting mode is determined by recursively comparing the rate-distortion costs with and without splitting.

[0202] In some examples, the maximum number of functional groups (FUs) can be limited using the maximum quadtree split level. In one example, when the maximum quadtree split level is 2, the maximum number of FUs is 16. Further, during the quadtree split determination, the correlation values ​​used to derive the Wiener coefficients of the 16 FUs (minimum FUs) at the bottom quadtree level can be reused. The remaining FUs can have their Wiener filters derived from the correlations of the 16 FUs at the bottom quadtree level. Therefore, in this example, only one framebuffer access is performed to derive the filter coefficients for all FUs.

[0203] After determining the quadtree splitting mode, to further reduce filtering distortion, CU-synchronized ALF on / off control can be implemented. By comparing the filtering distortion and non-filtering distortion at each leaf CU, the leaf CU can explicitly turn the ALF on / off in its local region. In some examples, coding efficiency can be further improved by redesigning the filter coefficients based on the ALF on / off results.

[0204] Cross-component filtering can be applied using cross-component filters, such as the cross-component adaptive loop filter (CC-ALF). A cross-component filter can use luminance sample values ​​of the luminance component (e.g., luminance CB) to refine the chrominance component (e.g., the chrominance CB corresponding to the luminance CB). In one example, the luminance CB and chrominance CB are included in the CU.

[0205] Figure 15 A cross-component filter (e.g., CC-ALF) for generating chromaticity components is illustrated according to one embodiment of this disclosure. In some examples, Figure 15 The filtering process for a first chromaticity component (e.g., first chromaticity CB), a second chromaticity component (e.g., second chromaticity CB), and a luminance component (e.g., luminance CB) is illustrated. The luminance component can be filtered by a Sample Adaptive Offset (SAO) filter (1510) to generate a SAO-filtered luminance component (1541). The SAO-filtered luminance component (1541) can be further filtered by an ALF luminance filter (1516) to become a filtered luminance CB (1561) (e.g., 'Y').

[0206] The first chromaticity component can be filtered by a SAO filter (1512) and an ALF chromaticity filter (1518) to generate a first intermediate component (1552). Furthermore, the SAO-filtered luminance component (1541) can be filtered by a cross-component filter (e.g., CC-ALF) (1521) for the first chromaticity component to generate a second intermediate component (1542). Subsequently, the filtered first chromaticity component (1562) (e.g., 'Cb') can be generated based on at least one of the first intermediate component 1552 and the second intermediate component 1542. In one example, the filtered first chromaticity component (1562) (e.g., 'Cb') can be generated by combining the first intermediate component (1552) and the second intermediate component (1542) with an adder (1522). The cross-component adaptive loop filtering process for the first chromaticity component may include steps performed by the CC-ALF (1521) and steps performed by, for example, the adder (1522).

[0207] The above description applies to the second chromaticity component. The second chromaticity component can be filtered by a SAO filter (1514) and an ALF chromaticity filter (1518) to generate a third intermediate component (1553). Furthermore, the SAO-filtered luminance component (1541) can be filtered by a cross-component filter (e.g., CC-ALF) (1531) for the second chromaticity component to generate a fourth intermediate component (1543). Subsequently, the filtered second chromaticity component (1563) (e.g., 'Cr') can be generated based on at least one of the third intermediate component (1553) and the fourth intermediate component (1543). In one example, the filtered second chromaticity component (1563) (e.g., 'Cr') can be generated by combining the third intermediate component (1553) and the fourth intermediate component (1543) with an adder (1532). In one example, the cross-component adaptive loop filtering process for the second chromaticity component may include steps performed by CC-ALF (1531) and steps performed by, for example, an adder (1532).

[0208] Cross-component filters (e.g., CC-ALF(1521), CC-ALF(1531)) can be operated by applying a linear filter with any suitable filter shape to the luminance component (or luminance channel) to refine each chrominance component (e.g., the first chrominance component, the second chrominance component).

[0209] Figure 16An example of a filter (1600) according to an embodiment of the present disclosure is shown. The filter (1600) may include non-zero filter coefficients and zero filter coefficients. The filter (1600) has a rhombus shape (1620) formed by filter coefficients (1610) (indicated by circles with black fill). In one example, the non-zero filter coefficients in the filter (1600) are included in the filter coefficients (1610), and the filter coefficients not included in the filter coefficients (1610) are zero. Therefore, the non-zero filter coefficients in the filter (1600) are included in the rhombus shape (1620), and the filter coefficients not included in the rhombus shape (1620) are zero. In one example, the number of filter coefficients of the filter (1600) is equal to the number of filter coefficients (1610), in Figure 16 In the example shown, the number of filter coefficients is 18.

[0210] CC-ALF can include any suitable filter coefficients (also known as CC-ALF filter coefficients). Return to Reference Figure 15 CC-ALF(1521) and CC-ALF(1531) can have the same filter shape, for example... Figure 16 The rhombus shape (1620) shown, as well as CC-ALF (1521) and CC-ALF (1531), can have the same number of filter coefficients. In one example, the values ​​of the filter coefficients in CC-ALF (1521) are different from the values ​​of the filter coefficients in CC-ALF (1531).

[0211] Typically, filter coefficients (e.g., non-zero filter coefficients) from the CC-ALF can be transmitted in the APS. In one example, the filter coefficients can be factored (e.g., 2). 10 Scaling is supported, and rounding is possible for fixed-point representations. CC-ALF application can be controlled with a variable block size, and its application is signaled by context-coded flags (e.g., CC-ALF enable flags) received for each block of a sample. Context-coded flags (e.g., CC-ALF enable flags) can be signaled at any suitable level (e.g., block level). For each chroma component, the block size and CC-ALF enable flag can be received together at the slice level. In some examples, block sizes of 16×16, 32×32, and 64×64 (in chroma samples) are supported.

[0212] Figure 17 Examples of syntax for CC-ALF according to some embodiments of this disclosure are shown. Figure 17In the example, `alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` is an index indicating whether a cross-component Cb filter is used, and if so, it indicates the index of the cross-component Cb filter. For example, when `alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` equals 0, the cross-component Cb filter is not applied to the block of Cb color component samples at the luminance location (xCtb, yCtb). When `alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` is not equal to 0, `alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` is the index of the filter to be applied. For example, the `alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]`-th cross-component Cb filter is applied to the block of Cb color component samples at the luminance position (xCtb, yCtb).

[0213] Furthermore, in Figure 17In the example, `alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` is used to indicate whether a cross-component Cr filter is used, and the index of whether the cross-component Cr filter is used. For example, when `alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` equals 0, the cross-component Cr filter is not applied to the block of Cr color component samples at the luminance location (xCtb, yCtb). When `alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` is not equal to 0, `alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` is the index of the cross-component Cr filter. For example, the `alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]`-th cross-component Cr filter can be applied to a block of Cr color component samples at the luminance location (xCtb, yCtb).

[0214] In some examples, chroma subsampling is used, so the number of samples in each chroma block can be less than the number of samples in the luma block. The chroma subsampling format (also known as, for example, the chroma subsampling format specified by chroma_format_idc) indicates the chroma horizontal subsampling factor (e.g., SubWidthC) and chroma vertical subsampling factor (e.g., SubHeightC) between each chroma block and its corresponding luma block. In one example, the chroma subsampling format is 4:2:0, so the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are both 2, as shown below. Figure 18A and Figure 18B As shown. In one example, the chroma subsampling format is 4:2:2, so the chroma horizontal subsampling factor (e.g., SubWidthC) is 2 and the chroma vertical subsampling factor (e.g., SubHeightC) is 1. In another example, the chroma subsampling format is 4:4:4, so both the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 1. The chroma sample type (also referred to as the chroma sample position) indicates the relative position of a chroma sample in a chroma block relative to at least one corresponding luminance sample in a luminance block.

[0215] Figure 18A and Figure 18B Exemplary positions of chromaticity samples relative to luminance samples according to embodiments of the present disclosure are shown. Reference Figure 18A The brightness sample (1801) is located in rows (1811)-(1818). Figure 18A The luminance sample (1801) shown may represent a portion of the image. In one example, a luminance block (e.g., luminance CB) includes luminance sample (1801). A luminance block may correspond to two chroma blocks with a chroma subsampling format of 4:2:0. In one example, each chroma block includes chroma sample (1803). Each chroma sample (e.g., chroma sample (1803(1))) corresponds to four luminance samples (e.g., luminance samples (1801(1))-(1801(4)). In one example, the four luminance samples are the top left sample (1801(1)), the top right sample (1801(2)), the bottom left sample (1801(3)), and the bottom right sample (1801(4)). The chroma sample (e.g., (1803(1)) is located at the left center position, which is located at the top left sample (1801(1)). Between 1801(1)) and the lower left sample (1801(3)), the chromaticity sample type of the chromaticity block with chromaticity sample (1803) can be called chromaticity sample type 0. Chromaticity sample type 0 indicates the relative position 0 corresponding to the left center position located between the upper left sample (1801(1)) and the lower left sample (1801(3)). Four luminance samples (e.g., (1801(1))-(1801(4))) can be called adjacent luminance samples of chromaticity sample (1803)(1).

[0216] In one example, each chroma block includes a chroma sample (1804). The above description of the chroma sample (1803) is applicable to the chroma sample (1804), so for brevity, a detailed description may be omitted. Each chroma sample (1804) may be located at the center of four corresponding luminance samples, and the chroma sample type of a chroma block having chroma samples (1804) may be called chroma sample type 1. Chroma sample type 1 indicates the relative position 1 corresponding to the center position of the four luminance samples (e.g., (1801(1))-(1801(4))). For example, a chroma sample (1804) may be located at the center portion of the luminance samples (1801(1))-(1801(4)).

[0217] In one example, each chroma block includes a chroma sample (1805). Each chroma sample (1805) may be located at the top-left position, which is the same position as the top-left sample among the four corresponding luminance samples (1801), and the chroma sample type of the chroma block with chroma sample (1805) may be called chroma sample type 2. Thus, each chroma sample (1805) is in the same position as the top-left sample among the four luminance samples (1801) corresponding to the corresponding chroma sample. Chroma sample type 2 indicates the relative position 2 of the top-left position corresponding to the four luminance samples (1801). For example, a chroma sample (1805) may be located at the top-left position of luminance samples (1801(1))-(1801(4)).

[0218] In one example, each chroma block includes a chroma sample (1806). Each chroma sample (1806) may be located at the top center position between the corresponding top-left sample and the corresponding top-right sample, and the chroma sample type of the chroma block with chroma sample (1806) may be referred to as chroma sample type 3. Chroma sample type 3 indicates the relative position 3 corresponding to the top center position between the top-left sample and the top-right sample. For example, a chroma sample (1806) may be located at the top center position of the luminance samples (1801(1))-(1801(4)).

[0219] In one example, each chroma block includes a chroma sample (1807). Each chroma sample (1807) may be located at the lower-left position, which is in the same position as the lower-left sample among the four corresponding luminance samples (1801), and the chroma sample type of the chroma block having the chroma sample (1807) may be called chroma sample type 4. Thus, each chroma sample (1807) is in the same position as the lower-left sample among the four luminance samples (1801) corresponding to the corresponding chroma sample. Chroma sample type 4 indicates the relative position 4 of the lower-left position corresponding to the four luminance samples (1801). For example, a chroma sample (1807) may be located at the lower-left position of luminance samples (1801(1))-(1801(4)).

[0220] In one example, each chroma block includes a chroma sample (1808). Each chroma sample (1808) is located at the bottom center position between the bottom left sample and the bottom right sample, and the chroma sample type of the chroma block with chroma sample (1808) can be referred to as chroma sample type 5. Chroma sample type 5 indicates the relative position 5 of the bottom center position between the bottom left and bottom right samples corresponding to the four luminance samples (1801). For example, a chroma sample (1808) can be located between the bottom left and bottom right samples of the luminance samples (1801(1))-(1801(4)).

[0221] Typically, any suitable chroma sample type can be used with a chroma subsampling format. Chroma sample types 0-5 are exemplary chroma sample types described with a chroma subsampling format of 4:2:0. Additional chroma sample types can be used with a chroma subsampling format of 4:2:0. Furthermore, variations of chroma sample types 0-5 and / or other chroma sample types can be used with other chroma subsampling formats, such as 4:2:2 and 4:4:4, etc. In one example, a chroma sample type combining chroma samples (1805) and (1807) is used with a chroma subsampling format of 4:2:2.

[0222] In one example, the luminance blocks are considered to have alternating rows, such as rows (1811)-(1812), where rows (1811)-(1812) consist of the top two samples (e.g., (1801(1))-(1801(2))) of four luminance samples (e.g., 1801(1))-(1801(4)) and the bottom two samples (e.g., 1801(3))-(1801(4))) of four luminance samples (e.g., 1801(1)-(1801(4))). Therefore, rows (1811), (1813), (1815), and (1817) can be called the current row (also called the top field), and rows (1812), (1814), (1816), and (1818) can be called the next row (also called the bottom field). Four brightness samples (e.g., (1801(1))-(1801(4))) are located in the current row (e.g., (1811)) and the next row (e.g., (1812)). Relative positions 2 and 3 are located in the current row, relative positions 0 and 1 are located between each current row and the corresponding next row, and relative positions 4 and 5 are located in the next row.

[0223] Within each chroma block, chroma samples (1803), (1804), (1805), (1806), (1807), or (1808) are located in rows (1851)-(1854). The specific position of rows (1851)-(1854) may depend on the chroma sample type. For example, for chroma samples (1803)-(1804) with chroma sample types 0 and 1 respectively, row (1851) is located between rows (1811)-(1812). For chroma samples (1805)-(1806) with chroma sample types 2 and 3 respectively, row (1851) is in the same position as the current row (1811). For chroma samples (1807)-(1808) with chroma sample types 4 and 5 respectively, row (1851) is in the same position as the next row (1812). The above description may be appropriately applied to lines (1852)-(1854), and for the sake of brevity, detailed descriptions are omitted.

[0224] Any suitable scanning method may be used to display, store and / or transmit the above text. Figure 18A The text describes the luminance blocks and their corresponding chrominance blocks. In one example, progressive scan is used.

[0225] Interlaced scanning can be used, such as Figure 18B As shown. As mentioned above, the chroma subsampling format is 4:2:0 (e.g., chroma_format_idc equals 1). In one example, the variable chroma position type (e.g., ChromaLocType) indicates the current line (e.g., ChromaLocType is chroma_sample_loc_type_top_field) or the next line (e.g., ChromaLocType is chroma_sample_loc_type_bottom_field). The current lines (1811), (1813), (1815), and (1817) and the next lines (1812), (1814), (1816), and (1818) can be scanned respectively. For example, the current lines (1811), (1813), (1815), and (1817) can be scanned first, and then the next lines (1812), (1814), (1816), and (1818) can be scanned. The current line may include a luminance sample (1801), and the next line may include a luminance sample (1802).

[0226] Similarly, corresponding chroma blocks can be scanned interlaced. Rows (1851) and (1853) containing unfilled chroma samples (1803), (1804), (1805), (1806), (1807), or (1808) can be referred to as the current row (or current chroma row), and rows (1852) and (1854) containing gray-filled chroma samples (1803), (1804), (1805), (1806), (1807), or (1808) can be referred to as the next row (or next chroma row). In one example, during interlaced scanning, rows (1851) and (1853) are scanned first, followed by rows (1852) and (1854).

[0227] In some examples, constrained directional enhancement filtering techniques can be used. The use of an in-loop constrained directional enhancement filter (CDEF) can filter out coded artifacts while preserving image details. In one example (e.g., HEVC), the Sample Adaptive Offset (SAO) algorithm can achieve a similar goal by defining signal offsets for different pixel categories. Unlike SAO, CDEF is a nonlinear spatial filter. In some examples, CDEF can be constrained to be easily vectorized (i.e., implemented using Single Instruction Multiple Data (SIMD) operations). It should be noted that other nonlinear filters (e.g., median filters, bilateral filters) cannot be handled in the same way.

[0228] In some cases, the amount of ringing artifacts in an encoded image tends to be roughly proportional to the quantization step size. The amount of detail is a property of the input image, but even the smallest details preserved in the quantized image tend to be proportional to the quantization step size. For a given quantization step size, the amplitude of ringing is typically smaller than the amplitude of detail.

[0229] CDEF can be used to identify the orientation of each block, then adaptively filter along the identified orientation, and filter to a lesser extent along the orientation rotated 45 degrees from the identified orientation. In some examples, the encoder can search for filter strength, and the filter strength can be explicitly represented as a signal, which allows for high control over the fuzziness.

[0230] Specifically, in some examples, orientation search is performed only on the reconstructed pixels after the deblocking filter. Since those pixels are available to the decoder, the orientation can be searched by the decoder, so in one example, orientation signaling is not required. In some examples, orientation search can operate on certain block sizes, such as 8×8 blocks, which are small enough to adequately handle non-straight edges, yet large enough to reliably estimate the orientation when applied to a quantized image. Furthermore, having a constant orientation over an 8×8 region makes vectorization of the filter easier. In some examples, each block (e.g., 8×8) can be compared to a perfectly oriented block to determine the differences. A perfectly oriented block is one such block where all pixels along a line in one direction have the same value. In one example, the difference measurement between this block and each of the perfectly oriented blocks, such as the sum of squared differences (SSD) or root mean square (RMS) errors, can be calculated. The perfectly oriented block with the minimum difference (e.g., minimum SSD, minimum RMS, etc.) can then be determined, and the orientation of the determined perfectly oriented block can be the orientation that best matches the pattern in the block.

[0231] Figure 19 An example of directional search according to an embodiment of this disclosure is shown. In one example, block (1910) is an 8×8 block reconstructed and output from the deblocking filter. Figure 19 In the example, direction search can determine one of the eight directions shown by (1920) for block (1910). Eight perfectly oriented blocks (1930) are formed, each corresponding to one of the eight directions (1920). A perfectly oriented block corresponding to a direction is a block that makes pixels along the line in that direction have the same value. Furthermore, the difference measurement, such as SSD, RMS error, etc., between each of block (1910) and perfectly oriented blocks (1930) can be calculated. Figure 19 In the example, the RMS error is shown by (1940). As shown in (1943), the RMS error of block (1910) and the perfectly oriented block (1933) is the smallest, so the orientation (1923) is the orientation that best matches the pattern in block (1910).

[0232] After identifying the orientation of the blocks, a nonlinear low-pass directional filter can be determined. For example, the filter taps of the nonlinear low-pass directional filter can be aligned along the identified orientation to reduce ringing while preserving directional edges or patterns. However, in some examples, directional filtering alone is sometimes insufficient to reduce ringing. In one example, additional filter taps are also used for pixels that are not aligned along the identified orientation. To reduce the risk of blurring, the additional filter taps are handled more conservatively. For this purpose, CDEF includes a primary filter tap and a secondary filter tap. In one example, the complete 2-D CDEF filter can be expressed as equation (14):

[0233]

[0234] Where D represents the damping parameter, S (p) S represents the intensity of the main filter tap. (s) The function represents the strength of the filter tap, `round(·)` indicates the operation to bypass constraints far from zero, `w` represents the filter weight, and `f(d,S,D)` is a constraint function that operates on the difference between the filtered pixel and each of its neighboring pixels. In one example, for small differences, the function `f(d,S,D)` equals `D`, which makes the filter behave like a linear filter; when the differences are large, the function `f(d,S,D)` equals `0`, which effectively ignores the filter tap.

[0235] In some examples, an intra-loop restoration scheme is used in post-encoded deblocking to substantially denoise and enhance edge quality, in addition to the deblocking operation. In one example, the intra-loop restoration scheme is toggleable within each frame for appropriately sized tiles. The intra-loop restoration scheme is based on a separable symmetric Wiener filter, a dual self-guided filter with subspace projection, and a domain transform recursive filter. Because content statistics can vary substantially within a frame, the intra-loop restoration scheme is integrated within a toggleable frame, where different schemes can be triggered in different regions of the frame.

[0236] Separable symmetric Wiener filters can be one of the in-loop recovery schemes. In some examples, each pixel in a degraded frame can be reconstructed as a noncausal filtered version of the pixels within a w×w window surrounding each pixel, where w = 2r + 1, and w is odd for integers r. If the 2D filter taps are in column vectorized form w... 2 The element vector F represents a ×1 element, and direct LMMSE optimization generates F = H. -1 M provides the filter parameters, where H = E[XX] T ] is the w in the w×w window around x and pixels. 2 The column vectorized version of the autocovariance of each sample, and M = E[YX] T] is the cross-correlation between x and the scalar source sample y to be estimated. In one example, the encoder can estimate H and M based on the implementation in the unblocked frame and the source, and can send the resulting filter F to the decoder. However, this is not only true when sending w 2 Each tap incurs a considerable bit rate cost, and non-separable filtering makes decoding extremely complex. In some embodiments, several additional constraints are imposed on the properties of F. For the first constraint, F is constrained to be separable, such that filtering can be implemented as a separable horizontal and vertical w-tap convolution. For the second constraint, each of the horizontal and vertical filters is constrained to be symmetric. For the third constraint, it is assumed that the sum of the horizontal and vertical filter coefficients is 1.

[0237] Dual self-guided filtering with subspace projection can be one of the in-loop recovery schemes. Guided filtering is an image filtering technique in which the local linear model is given by equation (15):

[0238] Equation (15) is y = Fx + G.

[0239] A local linear model is used to calculate the filtered output y based on the unfiltered sample x, where F and G are determined based on statistics of the guiding and degraded images near the filtered pixels. If the guiding image is identical to the degraded image, the resulting so-called self-guided filtering has the effect of preserving smooth edges. In one example, a specific form of self-guided filtering can be used. The specific form of self-guided filtering depends on two parameters: the radius r and the noise parameter e, and is listed as follows:

[0240] 1. Obtain the mean μ and variance σ of the pixels within a (2r+1)×(2r+1) window surrounding each pixel. 2 This step can be efficiently implemented using box filtering based on integral imaging.

[0241] 2. For each pixel, calculate: f = σ 2 / (σ 2 +e); g=(1-f)μ

[0242] 3. Calculate the F and G values ​​for each pixel as the average of the f and g values ​​in a 3×3 window around the pixel used.

[0243] The specific form of the self-guided filter is controlled by r and e, where a higher r means a higher spatial variance and a higher e means a higher range variance.

[0244] Figure 20 Examples illustrating subspace projection are shown in some instances. For example... Figure 20As shown, even if X1 and X2 are recovered, they are not close to the source Y. An appropriate multiplier {α, β} can make X1 and X2 closer to the source Y by simply moving X1 and X2 slightly to the right.

[0245] In some examples (e.g., HEVC), a filtering technique called Sample Adaptive Offset (SAO) can be used. In some examples, SAO is applied to the reconstructed signal after the deblocking filter. SAO can use the offset value given in the slice header. In some examples, for luminance samples, the encoder can decide whether to apply (enable) SAO on the slice. When SAO is enabled, the current image allows the coding unit to be recursively divided into four sub-regions, and each sub-region can select a SAO type from multiple SAO types based on features within the sub-region.

[0246] Figure 21 A table (2100) showing multiple SAO types according to one embodiment of the present disclosure is provided. In the table (2100), SAO types 0-6 are shown. It should be noted that SAO type 0 is used to indicate that SAO is not applied. Furthermore, each SAO type from SAO type 1 to SAO type 6 includes multiple categories. SAO can classify reconstructed pixels in a sub-region into multiple categories and reduce distortion by adding an offset to the pixels of each category in the sub-region. In some examples, edge attributes can be used for pixel classification in SAO types 1 to 4, and pixel intensity can be used for pixel classification in SAO types 5 and 6.

[0247] Specifically, in one embodiment, such as SAO types 5 and 6, band offset (BO) can be used to classify all pixels in a sub-region into multiple bands. Each of the multiple bands includes pixels within the same intensity range. In some examples, the intensity range is equally divided into multiple intervals, such as 32 intervals from zero to the maximum intensity value (e.g., 255 intervals for an 8-bit pixel), and each interval is associated with an offset. Furthermore, in one example, the 32 bands are divided into two groups, such as a first group and a second group. The first group includes the central 16 bands (e.g., the 16 intervals in the middle of the intensity range), while the second group includes the remaining 16 bands (e.g., 8 intervals on the lower side of the intensity range and 8 intervals on the higher side of the intensity range). In one example, only the offset of one of the two groups is transmitted. In some embodiments, when using pixel classification operations in BO, the five most valid bits of each pixel can be directly used as a band index.

[0248] Furthermore, in one embodiment, such as SAO types 1 to 4, edge offset (EO) can be used for pixel classification and offset determination. For example, edge orientation information can be considered to determine pixel classification based on a 1D 3-pixel pattern.

[0249] Figure 22Examples of 3-pixel patterns used for pixel classification of edge offsets are shown in some examples. Figure 22 In the example, the first pattern (2210) (as shown in 3 gray pixels) is called the 0-degree pattern (associated horizontally with the 0-degree pattern), the second pattern (2220) (as shown in 3 gray pixels) is called the 90-degree pattern (associated vertically with the 90-degree pattern), the third pattern (2230) (as shown in 3 gray pixels) is called the 135-degree pattern (associated diagonally with the 135-degree pattern), and the fourth pattern (2240) (as shown in 3 gray pixels) is called the 45-degree pattern (associated diagonally with the 45-degree pattern). In one example, edge orientation information of the sub-regions can be considered for selection. Figure 22 One of the four directional patterns shown is selected. In one example, a directional pattern is chosen that can be transmitted as auxiliary information in the encoded video bitstream. Pixels in a sub-region can then be classified into multiple categories by comparing each pixel with two adjacent pixels in the direction associated with the directional pattern.

[0250] Figure 23 Table (2300) shows pixel classification rules for edge offsets in some examples. Specifically, pixel c (also...) Figure 22 (shown in each pattern) with two adjacent pixels (also in) Figure 22 Each pattern is shown in gray for comparison, and pixel c can be based on the comparison, according to Figure 23 The pixel classification rules shown indicate that the pixels are classified into one of categories 0 to 4.

[0251] In some embodiments, the SAO on the decoder side can operate independently of the maximum coding unit (LCU) (e.g., CTU), thereby saving line buffers. In some examples, when 90-degree, 135-degree, and 45-degree classification patterns are selected, the pixels in the top and bottom rows of each LCU are not processed by SAO. When 0-degree, 135-degree, and 45-degree patterns are selected, the pixels in the leftmost and rightmost columns of each LCU are not processed by SAO.

[0252] Figure 24An example (2400) is shown where, if parameters are not merged from adjacent CTUs, it might be necessary to use signals to represent the syntax for the CTU. For example, the syntax element `sao_type_idx[cldx][rx][ry]` can be signal-represented to indicate the SAO type of a subregion. The SAO type can be BO (Band Offset) or EO (Edge Offset). When `sao_type_idx[cldx][rx][ry]` is 0, it indicates that the SAO is OFF; values ​​from 1 to 4 indicate the use of one of the four EO categories corresponding to 0°, 90°, 135°, and 45°; a value of 5 indicates the use of BO. Figure 24 In the example, each of the BO and EO types has four SAO offset values ​​(sao_offset[cIdx][rx][ry][0] to sao_offset[cIdx][rx][ry][3]) represented by signals.

[0253] Typically, the filtering process can use a reconstructed sample of the first color component as input (e.g., Y or Cb or Cr, or R or G or B) to generate an output, and the output of the filtering process is applied to a second color component, which may be the same as the first color component or may be another color component that is different from the first color component.

[0254] In examples of cross-component filtering (CCF), filter coefficients are derived based on mathematical equations. From the encoder to the decoder, these derived coefficients are represented as signals, and offsets are generated using a linear combination of these coefficients. The generated offsets are then added to the reconstructed samples as part of the filtering process. For example, an offset is generated based on a linear combination of the filter coefficients and the luma samples, and this offset is added to the reconstructed chroma samples. Examples of CCF rely on the assumption of a linear mapping between the reconstructed luma sample values ​​and the δ values ​​between the original and reconstructed chroma samples. However, the mapping between the reconstructed luma sample values ​​and the δ values ​​between the original and reconstructed chroma samples does not necessarily follow a linear mapping, thus limiting the coding performance of CCF under the assumption of a linear mapping.

[0255] In some examples, nonlinear mapping techniques can be used for cross-component filtering and / or same-color component filtering without significant signaling overhead. In one example, nonlinear mapping techniques can be used for cross-component filtering to generate cross-component sampling offsets. In another example, nonlinear mapping techniques can be used for same-color component filtering to generate local sampling offsets.

[0256] For convenience, the filtering process using nonlinear mapping techniques can be called Nonlinear Mapping Sample Offset (SO-NLM). In cross-component filtering, SO-NLM can be called Cross-Component Sample Offset (CCSO). In same-color component filtering, SO-NLM can be called Local Sample Offset (LSO). Filters using nonlinear mapping techniques can be called nonlinear mapping-based filters. Nonlinear mapping-based filters can include CCSO filters, LSO filters, etc.

[0257] In one example, CCSO and LSO can be used as loop filters to reduce distortion in the reconstructed samples. CCSO and LSO do not rely on the linear mapping assumption used in the relevant example CCF. For example, CCSO does not rely on the assumption of a linear mapping between the reconstructed luminance sample values ​​and the δ values ​​between the original chrominance samples and the reconstructed chrominance samples. Similarly, LSO does not rely on the assumption of a linear mapping between the reconstructed chrominance sample values ​​and the δ values ​​between the original chrominance samples and the reconstructed chrominance samples.

[0258] In the following description, an SO-NLM filtering process is described, which uses a reconstructed sample of a first color component as input (e.g., Y or Cb or Cr, or R or G or B) to generate an output, and the output of the filtering process is applied to a second color component. The description applies to LSO when the second color component is the same as the first color component; and to CCSO when the second color component is different from the first color component.

[0259] In SO-NLM, a nonlinear mapping is derived on the encoder side. This nonlinear mapping lies between the reconstructed samples of the first color component in the filter support region and the offset of the second color component added to the filter support region. When the second color component is the same as the first color component, the nonlinear mapping is used for LSO. When the second color component is different from the first color component, the nonlinear mapping is used for CCSO. The domain of the nonlinear mapping is determined by different combinations of the processed input reconstructed samples (also referred to as combinations of possible reconstructed sample values).

[0260] A concrete example can be used to illustrate the SO-NLM technique. In this example, a reconstructed sample from the first color component located within a filter support region (also referred to as the "filter support region") is determined. The filter support region is the area where a filter can be applied, and it can have any suitable shape.

[0261] Figure 25An example of a filter support region (2500) according to some embodiments of the present disclosure is shown. The filter support region (2500) includes four reconstructed samples of a first color component: P0, P1, P2, and P3. Figure 25 In the example, the four reconstructed samples can form an intersecting shape along the vertical and horizontal directions, and the center of the intersecting shape is the position of the sample to be filtered. The sample at the center position that has the same color component as P0-P3 is denoted by C. The sample at the center position that has a second color component is denoted by F. The second color component can be the same as or different from the first color component of P0-P3.

[0262] Figure 26 An example of another filter support region (2600) according to some embodiments of the present disclosure is shown. The filter support region (2600) includes four reconstructed samples P0, P1, P2, and P3 of a first color component, which form a square shape. Figure 26 In the example, the center of the square shape represents the location of the sample to be filtered. A sample at the center location that has the same color component as P0-P3 is denoted by C. A sample at the center location that has a second color component is denoted by F. The second color component may be the same as or different from the first color component of P0-P3.

[0263] The reconstructed samples are input to the SO-NLM filter and processed appropriately to form filter taps. In one example, the location of the reconstructed sample as input to the SO-NLM filter is called the filter tap location. In a specific example, the reconstructed samples are processed in the following two steps.

[0264] In the first step, the δ values ​​between P0-P3 and C are calculated respectively. For example, m0 represents the δ value between P0 and C; m1 represents the δ value between P1 and C; m2 represents the δ value between P2 and C; and m3 represents the δ value between P3 and C.

[0265] In the second step, the δ values ​​m0-m3 are further quantized, and the quantized values ​​are represented as d0, d1, d2, and d3. In one example, the quantized value can be one of -1, 0, or 1 based on the quantization process. For example, when m is less than -N (where N is a positive value and is called the quantization step size), the value m can be quantized to -1; when m is in the range [-N, N], the value m can be quantized to 0; and when m is greater than N, the value m can be quantized to 1. In some examples, the quantization step size N can be one of 4, 8, 12, 16, etc.

[0266] In some embodiments, the quantization values ​​d0-d3 are filter taps and can be used to identify a combination in the filter domain. For example, filter taps d0-d3 can form combinations in the filter domain. Each filter tap can have three quantization values, so when four filter taps are used, the filter domain includes 81 (3×3×3×3) combinations.

[0267] Figures 27A to 27C A table (2700) with 81 combinations is shown according to one embodiment of the present disclosure. The table (2700) comprises 81 rows corresponding to the 81 combinations. In each row corresponding to a combination, a first column includes the index of the combination; a second column includes the value of the filter tap d0 of the combination; a third column includes the value of the filter tap d1 of the combination; a fourth column includes the value of the filter tap d2 of the combination; a fifth column includes the value of the filter tap d3 of the combination; and a sixth column includes an offset value associated with the combination of nonlinear mappings. In one example, when filter taps d0-d3 are determined, the offset value (denoted by s) associated with the combination of d0-d3 can be determined according to the table (2700). In one example, the offset values ​​s0-s80 are integers, such as 0, 1, -1, 3, -3, 5, -5, -7, etc.

[0268] In some embodiments, the final filtering process of SO-NLM can be applied, as shown in equation (16):

[0269] f' = clip(f + s) Equation (16)

[0270] Where f is the reconstructed sample of the second color component to be filtered, and s is the offset value determined according to the filter tap, which is the result of processing the reconstructed sample of the first color component, for example, using the result processed using Table (2700). The sum of the reconstructed sample F and the offset value s is further constrained to a range associated with the bit depth to determine the final filtered sample f' of the second color component.

[0271] It should be noted that in the case of LSO, the second color component described above is the same as the first color component; in the case of CCSO, the second color component described above may be different from the first color component.

[0272] It should be noted that the above description may be adjusted for other embodiments of the present invention.

[0273] In some examples, on the encoder side, the encoding device can derive a mapping between reconstructed samples of a first color component in the filter support region and offsets of reconstructed samples to be added to the second color component. This mapping can be any suitable linear or non-linear mapping. Then, on the encoder side and / or the decoder side, a filtering process can be applied based on the mapping. For example, the decoder can be appropriately notified of the mapping (e.g., the mapping is included in the encoded video bitstream sent from the encoder side to the decoder side), and then the decoder can perform a filtering process based on the mapping.

[0274] According to some aspects of this disclosure, the performance of filters based on nonlinear mappings (e.g., CCSO filters, LSO filters, etc.) depends on the filter shape configuration. Using a fixed filter shape configuration can limit the performance of filters based on nonlinear mappings. Various aspects of the technology provide techniques for switching filter shape configurations for filters based on nonlinear mappings (e.g., CCSO filters, LSO filters, etc.).

[0275] According to some aspects of this disclosure, the filter shape configuration (also referred to as filter shape) of a filter can refer to the properties of a pattern formed by the positions of the filter taps. The pattern can be defined by various parameters, such as the geometry of the filter tap positions, the distance from the filter tap positions to the center of the pattern, etc.

[0276] In some embodiments, the filter shape configuration of a nonlinear mapping-based filter (e.g., a CCSO filter, an LSO filter, etc.) may have a cross geometry. Specifically, filter tap positions exist at the top, bottom, left, and right of the center position of the filter tap position. In one example, the distance from the filter tap position to the center of the filter tap position (denoted by n) can be any suitable positive integer in samples, such as 1, 2, 3, 4, 5, etc.

[0277] Figure 28 An example of a filter shape configuration (2800) according to an embodiment of the present disclosure is shown. The filter shape configuration (2800) has a cross geometry, and the distance from the filter tap position to the center of the filter shape configuration (2800) is 1 sample (n=1). Figure 28 In the diagram, each circle represents a sample. Figure 28In the diagram, the center position is indicated by C. The filter shape configuration (2800) includes four filter tap positions indicated by p0, p1, p2, and p3. As shown, filter tap position p0 is located at the top of the center position C; filter tap position p1 is located to the left of the center position C; filter tap position p2 is located at the bottom of the center position C; and filter tap position p3 is located to the right of the center position C. The distance from filter tap position p0 to the center position C is 1 sample; the distance from filter tap position p1 to the center position C is 1 sample; the distance from filter tap position p2 to the center position C is 1 sample; and the distance from filter tap position p3 to the center position C is 1 sample.

[0278] In one example, to apply the filter of the filter shape configuration (2800) to a sample, the sample to be filtered is located at the center position C; the reconstructed sample at filter tap position p0 is used to derive the first filter tap (d0). The reconstructed sample at filter tap position p1 is used to derive the second filter tap (d1). The reconstructed sample at filter tap position p2 is used to derive the third filter tap (d2). And the reconstructed sample at filter tap position p3 is used to derive the fourth filter tap (d1). Then, the filter taps d0-d3 are used to determine the sampling offset to be applied to the sample to be filtered.

[0279] Figure 29 Another example of a filter shape configuration (2900) according to an embodiment of the present disclosure is shown. The filter shape configuration (2900) has a cross geometry and the distance from the filter tap position to the center of the filter shape configuration (2900) is 4 samples (n=4). Figure 29 In the diagram, each circle represents a sample, with the center position indicated by C. The filter shape configuration (2900) includes four filter tap positions indicated by p0, p1, p2, and p3. Figure 29 As shown, filter tap position p0 is located at the top of center position C; filter tap position p1 is located to the left of center position C; filter tap position p2 is located at the bottom of center position C; and filter tap position p3 is located to the right of center position C. The distance from filter tap position p0 to center position C is 4 samples. The distance from filter tap position p1 to center position C is 4 samples; the distance from filter tap position p2 to center position C is 4 samples; and the distance from filter tap position p3 to center position C is 4 samples.

[0280] In one example, to apply the filter of the filter shape configuration (2900) to a sample, the sample to be filtered is located at the center position C; the reconstructed sample at filter tap position p0 is used to derive the first filter tap (d0); the reconstructed sample at filter tap position p1 is used to derive the second filter tap (d1); the reconstructed sample at filter tap position p2 is used to derive the third filter tap (d2); and the reconstructed sample at filter tap position p3 is used to derive the fourth filter tap (d1). Then, the filter taps d0-d3 are used to determine the sampling offset to be applied to the sample to be filtered.

[0281] Figure 30 An example of a filter shape configuration (3000) according to an embodiment of the present disclosure is shown. The filter shape configuration (3000) has a rectangular geometry, and the distance from the filter tap position to the center of the filter tap position is 1 sample (n=1). Specifically, in Figure 30 In the diagram, the center position is indicated by C. The filter shape configuration (3000) includes four filter tap positions, indicated by q0, q1, q2, and q3. As shown, filter tap position q0 is located to the upper left of the center position C; filter tap position q1 is located to the lower left of the center position C; filter tap position q2 is located to the lower right of the center position C; and filter tap position q3 is located to the upper right of the center position C. The distance from filter tap position q0 to the center position C is 1 sample; the distance from filter tap position q1 to the center position C is 1 sample; the distance from filter tap position q2 to the center position C is 1 sample; and the distance from filter tap position q3 to the center position C is 1 sample.

[0282] In one example, to apply a filter of filter shape configuration (3000) to a sample, the sample to be filtered is located at center position C. The reconstructed sample at filter tap position q0 is used to derive the first filter tap (d0). The reconstructed sample at filter tap position q1 is used to derive the second filter tap (d1). The reconstructed sample at filter tap position q2 is used to derive the third filter tap (d2). And the reconstructed sample at filter tap position q3 is used to derive the fourth filter tap (d1). The filter taps d0-d3 are then used to determine the sampling offset to be applied to the sample to be filtered.

[0283] Figure 31 An example of a filter shape configuration (3100) according to an embodiment of the present disclosure is shown. The filter shape configuration (3100) has a rectangular geometry and the distance from the filter tap position to the center of the filter tap position is 4 samples (n=4). Specifically, in Figure 31In the diagram, the center position is indicated by C. The filter shape configuration (3100) includes four filter tap positions indicated by q0, q1, q2, and q3. As shown, filter tap position q0 is located at the upper left of the center position C; filter tap position q1 is located at the lower left of the center position C; filter tap position q2 is located at the lower right of the center position C; and filter tap position q3 is located at the upper right of the center position C. The distance from filter tap position q0 to the center position C is 4 samples; the distance from filter tap position q1 to the center position C is 4 samples; the distance from filter tap position q2 to the center position C is 4 samples; and the distance from filter tap position q3 to the center position C is 4 samples.

[0284] In one example, to apply the filter of the filter shape configuration (3100) to a sample, the sample to be filtered is located at the center position C; the reconstructed sample at filter tap position q0 is used to derive the first filter tap (d0); the reconstructed sample at filter tap position q1 is used to derive the second filter tap (d1); the reconstructed sample at filter tap position q2 is used to derive the third filter tap (d2); and the reconstructed sample at filter tap position q3 is used to derive the fourth filter tap (d1). The filter taps d0-d3 are then used to determine the sampling offset to be applied to the sample to be filtered.

[0285] According to one aspect of this disclosure, the filter shape configuration of a nonlinear mapping-based filter (e.g., CCSO filter, LSO filter, etc.) can be switchable during video reconstruction from an encoded video bitstream. The nonlinear mapping-based filter can select one of several candidate filter shape configurations at appropriate levels, such as at the picture level, block level, slice level, tile level, etc.

[0286] In one embodiment, multiple candidate filter shape configurations may have the same geometry.

[0287] Figure 32 An example (3200) of three candidate filter shape configurations with intersecting geometry is shown. The distances from the filter tap positions to the center obtained by the three candidate filter shape configurations can be different.

[0288] Specifically, in Figure 32In the first candidate filter configuration, the center position C and filter tap positions p0, p1, p2, and p3 form the first candidate filter shape configuration. The distance from filter tap positions p0, p1, p2, and p3 to the center position C is 1 sample. The center position C and filter tap positions p0', p1', p2', and p3' form the second candidate filter shape configuration. The distance from filter tap positions p0', p1', p2', and p3' to the center position C is 4 samples. The center position C and filter tap positions p0", p1", p2", and p3" form the third candidate filter shape configuration. The distance from filter tap positions p0", p1", p2", and p3" to the center position C is 7 samples.

[0289] In some examples, one of the first candidate filter shape configuration, the second candidate filter shape configuration, and the third candidate filter shape configuration can be selected at an appropriate level, such as at the image level, block level, slice level, tile level, etc., to be used for sample reconstruction at the appropriate level.

[0290] Figure 33 An example (3300) of two candidate filter shape configurations with intersecting geometry is shown. The distances from the filter tap positions to the center obtained by the two candidate filter shape configurations can be different.

[0291] Specifically, in Figure 33 In the first candidate filter configuration, the center position C and filter tap positions p0, p1, p2, and p3 form the first candidate filter shape configuration. The distance from the filter tap positions p0, p1, p2, and p3 to the center position C is 1 sample. The center position C and filter tap positions p0', p1', p2', and p3' form the second candidate filter shape configuration. The distance from the filter tap positions p0', p1', p2', and p3' to the center position C is 4 samples.

[0292] In some examples, one of the first candidate filter shape configurations and the second candidate filter shape configurations can be selected at an appropriate level, such as at the image level, block level, slice level, tile level, etc., to be used for sample reconstruction at the appropriate level.

[0293] It should be noted that the candidate filter shape configuration can have other suitable geometries.

[0294] Figure 34 An example (3400) of two candidate filter shape configurations with rectangular geometry is shown. The distances from the filter tap positions to the center obtained by the two candidate filter shape configurations can be different.

[0295] Specifically, in Figure 34In the first candidate filter configuration, the center position C and filter tap positions q0, q1, q2, and q3 form the first candidate filter shape configuration. The distance from the filter tap positions q0, q1, q2, and q3 to the center position C is 1 sample. The second candidate filter shape configuration, consisting of the center position C and filter tap positions q0', q1', q2', and q3', forms the second candidate filter shape configuration. The distance from the filter tap positions q0', q1', q2', and q3' to the center position C is 4 samples.

[0296] In some examples, one of the first candidate filter shape configurations and the second candidate filter shape configurations can be selected at an appropriate level, such as at the image level, block level, slice level, tile level, etc., to be used for sample reconstruction at the appropriate level.

[0297] It should also be noted that candidate filter shape configurations can have different geometries.

[0298] Figure 35 An example (3500) of four candidate filter shape configurations with a mixture of intersecting and rectangular geometries is shown. The distances from the filter tap positions to the center obtained by the four candidate filter shape configurations can be different.

[0299] Specifically, in Figure 35 In the first candidate filter shape configuration, the center position C and filter tap positions p0, p1, p2, and p3 form the first candidate filter shape configuration. The distance from the filter tap positions p0, p1, p2, and p3 to the center position C is 1 sample. The center position C and filter tap positions p0', p1', p2', and p3' form the second candidate filter shape configuration. The distance from the filter tap positions p0', p1', p2', and p3' to the center position C is 4 samples. The first and second candidate filter shape configurations have an intersecting geometry.

[0300] Furthermore, the center position C and the filter tap positions q0, q1, q2, and q3 form the third candidate filter shape configuration. The distance from the filter tap positions q0, q1, q2, and q3 to the center position C is 1 sample. The center position C and the filter tap positions q0', q1', q2', and q3' form the fourth candidate filter shape configuration. The distance from the filter tap positions q0', q1', q2', and q3' to the center position C is 4 samples. Both the third and fourth candidate filter shape configurations have rectangular geometry.

[0301] In some examples, one of the first candidate filter shape configuration, the second candidate filter shape configuration, the third candidate filter shape configuration, and the fourth candidate filter shape configuration can be selected at an appropriate level, such as at the image level, block level, slice level, tile level, etc., to be used for sample reconstruction at the appropriate level.

[0302] According to one aspect of this disclosure, a filter shape configuration can be selected from a plurality of candidate filter shape configurations by means of a signal in the encoded video bitstream from encoder to decoder.

[0303] In one embodiment, the switching of filter shape configurations for nonlinear mapping-based filters (e.g., CCSO filters, LSO filters) is performed at the picture level. In one example, for each picture in the encoded video bitstream, a signal is used to indicate the index of the filter shape configuration selected from a plurality of candidate filter shape configurations.

[0304] In one embodiment, the switching of filter shape configurations for nonlinearly mapped filters (e.g., CCSO filters, LSO filters) is at the block level. A block can be interpreted as a prediction block, coding block, or coding unit, i.e., a CU, CTU block, or superblock, or a filtering unit (FU). In one example, for each block in the encoded video bitstream, a signal is used to indicate the index of the filter shape configuration selected from a plurality of candidate filter shape configurations.

[0305] It should be noted that in some embodiments, the index indicating the selected filter shape configuration can be represented by a signal in a high-level syntax, such as APS, slice header, frame header, PPS, SPS, VPS, etc.

[0306] According to one aspect of this disclosure, samples at filter tap locations can be preprocessed and then used as input to filters based on nonlinear mappings (e.g., CCSO filters, LSO filters, etc.).

[0307] In one embodiment, a weighted average of the sample values ​​at the filter taps can be calculated.

[0308] Figure 36 An example of preprocessing according to an embodiment of the present disclosure is shown (3600). Figure 36 As shown, example (3600) includes eight filter tap positions p0-p7. In one example, the average of the samples at filter tap positions p0 and p1 is calculated and denoted as p0', the average of the samples at filter tap positions p2 and p3 is calculated and denoted as p1', the average of the samples at filter tap positions p4 and p5 is calculated and denoted as p2', and the average of the samples at filter tap positions p6 and p7 is calculated and denoted as p3'. Then, p0'-p3' and c are used as inputs to a nonlinear mapping-based filter (e.g., CCSO filter, LSO filter, etc.).

[0309] It should be noted that the average value can be calculated as a weighted average. For example, when calculating the average value of the samples at filter tap positions p0 and p1, the samples at p0 and p1 can be weighted differently.

[0310] In another embodiment, the pre-filtering process can be applied to samples located at filter taps.

[0311] Figure 37 An example (3700) of preprocessing according to an embodiment of the present disclosure is shown. Example (3700) includes four filter tap positions p0, p1, p2, and p3. In one example, a filtering (pre-filtering) process is applied to samples at filter tap positions based on samples that are close to (e.g., neighboring or within K sample distances, where K is a positive integer) the filter tap positions.

[0312] For example, based on the first sample at adjacent positions q0-q3, a first filtering (pre-filtering) process is applied to the first sample at filter tap position p0; based on the second sample at adjacent positions r0-r3, a second filtering (pre-filtering) process is applied to the second sample at filter tap position p1; based on the third sample at adjacent positions s0-s3, a third filtering (pre-filtering) process is applied to the third sample at filter tap position p2; and based on the fourth sample at adjacent positions t0-t3, a fourth filtering (pre-filtering) process is applied to the fourth sample at filter tap position p3. Then, the filtered first, second, third, and fourth samples can be used as inputs to filters based on nonlinear mappings (e.g., CCSO filters and LSO filters).

[0313] It should be noted that pre-filtering can be performed by any suitable filter, whether linear or nonlinear.

[0314] Figure 38A flowchart outlining a process (3800) according to one embodiment of the present disclosure is shown. The process (3800) can be used to reconstruct the video carried in an encoded video bitstream. When using the term block, a block can be interpreted as a prediction block, coding unit, luma block, chroma block, etc. In various embodiments, the process (3800) is executed by processing circuitry such as: processing circuitry in terminal devices (310), (320), (330), and (340); processing circuitry performing the functions of a video encoder (403); processing circuitry performing the functions of a video decoder (410); processing circuitry performing the functions of a video decoder (510); processing circuitry performing the functions of a video encoder (603); etc. In some embodiments, the process (3800) is implemented as software instructions, so that when the processing circuitry executes these software instructions, the processing circuitry executes the process (3800). The process begins at (S3801) and proceeds to (S3810).

[0315] At (S3810), the first sample in the video carried in the encoded video bitstream is reconstructed based on a filter with a first filter shape configuration and a nonlinear mapping.

[0316] At (S3820), a switch from a first filter shape configuration to a second filter shape configuration is determined. The second filter shape configuration differs from the first filter shape configuration. In some examples, the difference between the first and second filter shape configurations may be the geometry of the filter tap position, and may be the distance from the filter tap position to the center of the filter tap position.

[0317] The geometry of the filter tap location can be an intersecting geometry or a rectangular geometry. The first filter shape configuration and the second filter shape configuration can have different or the same geometry at the filter tap locations. In some examples, the first and second filter shape configurations have the same geometry, but the distance from the filter tap location to its center differs between the two configurations.

[0318] In some examples, the index is decoded based on the encoded video bitstream carrying the video. The index indicates a second filter shape configuration. Then, based on the index, a switch from the first filter shape configuration to the second filter shape configuration is determined.

[0319] In one example, indices are represented by signals at the image level, and the switch from a first filter shape configuration to a second filter shape configuration is determined at the image level. The first sample is located in the first image of the video, and the second sample is located in the second image of the video.

[0320] In another example, indices are represented by signals at the block level, and the switch from a first filter shape configuration to a second filter shape configuration is determined at the block level. The first sample is located in a first block of images in the video, and the second sample is located in a second block of images in the video.

[0321] In some examples, the index can be decoded based on signaling in the high-level syntax, such as signaling in the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), slice header, tile header, frame header, etc.

[0322] At (S3830), the second sample in the video is reconstructed based on a filter with a second filter shape configuration and a nonlinear mapping.

[0323] The process (3800) proceeds to (S3899) and ends.

[0324] It should be noted that in some examples, the filter based on the nonlinear mapping is a cross-component sample offset (CCSO) filter, and in some other examples, the filter based on the nonlinear mapping is a local sample offset (LSO) filter.

[0325] It should also be noted that sample values ​​that serve as input to a nonlinear mapping-based filter can be preprocessed. For example, to reconstruct a first sample, a preprocessing operation can be performed on samples located at filter tap positions corresponding to the first filter shape configuration, generating preprocessed samples. These preprocessed samples are then used as input to the nonlinear mapping-based filter, and the offset applied to the first sample can be determined based on the preprocessed samples. In one example, the preprocessing operation is a weighted averaging operation. For example, the average sample values ​​located at two or more filter tap positions can be calculated to derive the preprocessed samples. In another example, the preprocessing operation is a filtering operation. For example, a filter is applied to samples located at filter tap positions to generate filtered samples as preprocessed samples. The filter can be any suitable filter, such as a linear filter, a nonlinear filter, etc.

[0326] The process (3800) can be adjusted as appropriate. Steps in the process (3800) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.

[0327] The embodiments of this disclosure can be used individually or in any combination in any order. Furthermore, each of the method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0328] The above-described technology can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 39 A computer system (3900) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0329] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.

[0330] The instructions can be executed on various types of computers or their components, including personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0331] Figure 39 The components of the computer system (3900) shown are exemplary in nature and are not intended to impose any limitation on the use or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependencies or requirements relating to any one or a combination of components shown in the exemplary embodiments of the computer system (3900).

[0332] The computer system (3900) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, images captured from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0333] Human-machine interface input devices may include one or more of the following (only one of each is shown): keyboard (3901), mouse (3902), touchpad (3903), touch screen (3910), data glove (not shown), joystick (3905), microphone (3906), scanner (3907), camera (3908).

[0334] The computer system (3900) may also include certain human-machine interface output devices. Such human-machine interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback from a touchscreen (3910), data gloves (not shown), or joystick (3905), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (3909), headphones (not depicted)), visual output devices (e.g., screens (3910) including CRT screens, LCD screens, plasma screens, OLED screens, each screen may or may not have touchscreen input functionality, each screen may or may not have tactile feedback functionality, some of which are capable of outputting two-dimensional visual output or more than three-dimensional output via devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), and printers (not depicted).

[0335] The computer system (3900) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (3920) having media such as CD / DVD (3921), finger drives (3922), removable hard disk drives or solid-state drives (3923), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0336] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0337] The computer system (3900) may also include an interface (3954) leading to one or more communication networks (3955). The network may be, for example, a wireless network, a wired network, or an optical network. The network may further be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (3900)) attached to some general-purpose data port or peripheral bus (3949). Other network interfaces are typically integrated into the core of the computer system (3900) by being attached to a system bus (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system (3900) can use any of these networks to communicate with other entities. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANbus connected to certain CANbus devices), or bidirectional, such as connecting to other computer systems using a local area network (LAN) or wide area network (WAN) digital network. As mentioned above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.

[0338] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel (3940) of the computer system (3900).

[0339] The core (3940) may include one or more central processing units (CPUs) (3941), graphics processing units (GPUs) (3942), dedicated programmable processing units in the form of field-programmable gate arrays (FPGAs) (3943), hardware accelerators (3944) for certain tasks, graphics adapters (3950), etc. These devices, along with read-only memory (ROM) (3945), random access memory (3946), and internal mass storage (3947) such as internal non-user-accessible hard disk drives, SSDs, etc., may be connected via a system bus (3948). In some computer systems, the system bus (3948) may be accessed via one or more physical connectors to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (3948) or attached via a peripheral bus (3949). In one example, a display (3910) may be connected to a graphics adapter (3950). Peripheral bus architectures include PCI, USB, etc.

[0340] The CPU (3941), GPU (3942), FPGA (3943), and accelerator (3944) can execute certain instructions that can be combined to form the aforementioned computer code. This computer code can be stored in ROM (3945) or RAM (3946). Transient data can also be stored in RAM (3946), while permanent data can be stored, for example, in internal mass storage (3947). Fast storage and retrieval to any storage device can be achieved using a cache, which can be closely associated with one or more CPUs (3941), GPUs (3942), mass storage (3947), ROM (3945), RAM (3946), etc.

[0341] Computer-readable media may have computer code thereon that performs various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0342] As a non-limiting example, a computer system having an architecture (3900), particularly a kernel (3940), can provide functionality by having one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, and memory of certain non-transitory kernels (3940), such as internal kernel mass storage (3947) or ROM (3945). Software implementing various embodiments of this disclosure can be stored in such devices and executed by the kernel (3940). Depending on specific needs, the computer-readable media may include one or more storage devices or chips. The software can cause the kernel (3940), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (3946) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system may provide functionality due to hard-wired or otherwise embodied logic in circuitry (e.g., the accelerator (3944)), which may replace or operate with the software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., an integrated circuit (IC)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0343] Appendix A: Acronyms

[0344] JEM: Joint Exploration Model

[0345] VVC: Universal Video Coding

[0346] BMS: Benchmark Set

[0347] MV: Motion Vector

[0348] HEVC: High-Efficiency Video Coding

[0349] MPM: Most Likely Pattern

[0350] WAIP: Wide-angle Intra-frame Prediction

[0351] SEI: Auxiliary Enhancement Information

[0352] VUI: Video Availability Information

[0353] GOP: Image Group

[0354] TU: Transformation Unit

[0355] PU: Prediction Unit

[0356] CTU: Coding Tree Unit

[0357] CTB: Coded Tree Block

[0358] PB: Prediction Block

[0359] HRD: Hypothetical Reference Decoder

[0360] SDR: Standard Dynamic Range

[0361] SNR: Signal-to-noise ratio

[0362] CPU: Central Processing Unit

[0363] GPU: Graphics Processing Unit

[0364] CRT: Cathode Ray Tube

[0365] LCD: Liquid Crystal Display

[0366] OLED: Organic Light Emitting Diode

[0367] CD: CD-ROM

[0368] DVD: Digital Video Disc

[0369] ROM: Read-Only Memory

[0370] RAM: Random Access Memory

[0371] ASIC: Application-Specific Integrated Circuit

[0372] PLD: Programmable Logic Device

[0373] LAN: Local Area Network

[0374] GSM: Global System for Mobile Communications

[0375] LTE: Long Term Evolution

[0376] CANBus: Controller Area Network Bus

[0377] USB: Universal Serial Bus

[0378] PCI: Interconnect Peripheral Devices

[0379] FPGA: Field Programmable Gate Array

[0380] SSD: Solid State Drive

[0381] IC: Integrated Circuit

[0382] CU: Encoding Unit

[0383] PDPC: Location-Related Prediction Combination

[0384] ISP: Intra-Frame Sub-Partition

[0385] SPS: Sequence Parameter Settings

[0386] Although several exemplary embodiments have been described in this disclosure, modifications, substitutions, and various equivalent alternatives that fall within the scope of this disclosure exist. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.

Claims

1. A filtering method in video decoding, comprising: The first sample in the video carried in the encoded video bitstream is reconstructed based on a filter with a first filter shape configuration and a nonlinear mapping. A switch is made from the first filter shape configuration to the second filter shape configuration, wherein the first filter shape configuration and the second filter shape configuration contain the same number of filter taps, and the distance from the filter tap position of the first filter shape configuration to the center of the first filter shape configuration is different from the distance from the filter tap position of the second filter shape configuration to the center of the second filter shape configuration. as well as The second sample in the video is reconstructed based on the filter based on the nonlinear mapping, which has the second filter shape configuration.

2. The method according to claim 1, wherein, The nonlinear mapping-based filter includes at least one of a cross-component sampling offset (CCSO) filter and a local sampling offset (LSO) filter.

3. The method according to claim 1, wherein, The difference between the first filter shape configuration and the second filter shape configuration is at least as follows: The geometry of the filter tap positions.

4. The method according to claim 1, wherein, Both the first filter shape configuration and the second filter shape configuration have at least one of a cross geometry at the filter tap position and a rectangular geometry at the filter tap position.

5. The method according to claim 1, wherein, The first filter shape configuration and the second filter shape configuration have the same geometry.

6. The method according to claim 1, further comprising: The index is decoded based on the encoded video bitstream carrying the video, the index indicating the second filter shape configuration; as well as Based on the index, a switch from the first filter shape configuration to the second filter shape configuration is determined.

7. The method according to claim 6, wherein, The first sample is located in a first image of the video, the second sample is located in a second image of the video, and the method further includes: Based on the index, the switch from the first filter shape configuration to the second filter shape configuration is determined at the image level.

8. The method according to claim 6, wherein the first sample is located in a first block of images in the video, and the second sample is located in a second block of images in the video, the method further comprising: Based on the index, a switch from the first filter shape configuration to the second filter shape configuration is determined at the block level.

9. The method according to claim 6, further comprising: The index is decoded based on the syntax signaling of at least one of the block-level, video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), slice header, tile header, and frame header.

10. The method according to any one of claims 1 to 9, wherein, Reconstructing the first sample in the video based on a nonlinear mapping using a filter with a first filter shape configuration also includes: Perform preprocessing operations on samples located at filter tap positions corresponding to the first filter shape configuration to generate preprocessed samples; and The offset applied to the first sample is determined based on the preprocessed sample.

11. The method according to claim 10, wherein, Performing a preprocessing operation on samples located at filter tap positions corresponding to the first filter shape configuration to generate preprocessed samples further includes at least one of the following: Calculate the average sample value located at two or more filter tap positions as the preprocessed sample; as well as Apply the filter to the sample located at the filter tap position to generate the filtered sample as the preprocessed sample.

12. A filtering device for video decoding, comprising a processing circuit, the processing circuit being configured to: The first sample in the video carried in the encoded video bitstream is reconstructed based on a filter with a first filter shape configuration and a nonlinear mapping. Determine to switch from the first filter shape configuration to the second filter shape configuration, wherein, The first filter shape configuration and the second filter shape configuration contain the same number of filter taps, and the distance from the filter tap position of the first filter shape configuration to the center of the first filter shape configuration is different from the distance from the filter tap position of the second filter shape configuration to the center of the second filter shape configuration. as well as The second sample in the video is reconstructed based on the filter based on the nonlinear mapping, which has the second filter shape configuration.

13. A computer-readable medium storing instructions that, when executed by a computer device for video decoding, cause the computer device to perform the method according to any one of claims 1 to 11.