Method and device for filtering in video coding and decoding, and storage medium

By adopting nonlinear mapping-based filters and constrained direction enhancement filters in video encoding and decoding technology, combining cross component sample offset and local sample offset technologies, the problems of low efficiency of video frame boundary filtering and difficulty in reducing encoding artifacts are solved, and efficient filtering processing and video quality improvement are achieved.

CN120111234APending Publication Date: 2025-06-06TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264599.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2021-09-22
Filing Date
2021-09-24
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has problems such as low filtering processing efficiency and difficulty in reducing encoding artifacts when processing video frame boundaries.

Method used

A filter based on nonlinear mapping and a constrained direction enhancement filter are used, and the process is carried out through a loop filter chain, combining cross component sample offset and local sample offset techniques to generate reconstructed samples and apply loop recovery filters.

Benefits of technology

It improves the filtering processing efficiency during video encoding and decoding, reduces the occurrence of encoding artifacts, and enhances video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111234A_ABST
    Figure CN120111234A_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods and apparatus for filtering in video coding. In some examples, an apparatus for video encoding and decoding includes processing circuitry. The processing circuitry buffers first boundary pixel values of a first reconstructed sample at a first node along a loop filter chain. The first node is associated with a non-linear mapping-based filter applied prior to a loop recovery filter in a loop filter chain. The first boundary pixel value is a value of a pixel at a frame boundary. The processing circuitry applies a loop recovery filter to the reconstructed sample to be filtered based on the buffered first boundary pixel values.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by reference

[0002] This application claims the benefit of priority to U.S. Patent Application No. 17 / 448,469, filed on September 22, 2021, entitled “METHOD AND APPARATUS FOR BOUNDARY HANDLING IN VIDEO CODING,” which claims the benefit of priority to U.S. Provisional Application No. 63 / 187,213, filed on May 11, 2021, entitled “ON LOOP RESTORATION BOUNDARY HANDLING.” The disclosures of the prior applications are incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure describes embodiments generally related to video coding. Background Art

[0004] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent described in this background section, the work of the presently named inventors and aspects of the description that may not be prior art at the time of filing are neither explicitly nor implicitly admitted to be prior art to the present disclosure.

[0005] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. An uncompressed digital video may include a series of pictures, each picture having a luminance sample (Sample(s), or sampling) and associated chrominance samples with a spatial dimension of, for example, 1920×1080. The series of pictures may have a fixed or variable image rate (also informally referred to as a frame rate), for example, 60 pictures per second or 60Hz. Uncompressed video has specific bit rate requirements. For example, 1080p60 4:2:0 video (1920x1080 luminance sample resolution at a 60Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5Gbit / s. One hour of such video requires more than 600G bytes of storage space.

[0006] One purpose of video encoding and decoding is to reduce redundancy in the input video signal through compression. Compression helps reduce the bandwidth and / or storage space requirements mentioned above, in some cases by two orders of magnitude or more. Lossless compression and lossy compression and their combinations can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal. When lossy compression is used, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. For video, lossy compression is widely used. The amount of distortion allowed depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that higher allowed / tolerable distortion can produce higher compression ratios.

[0007] Video encoders and decoders may utilize several broad categories of techniques including, for example, motion compensation, transforms, quantization, and entropy coding.

[0008] Video codec techniques may include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from a previously reconstructed reference picture. In some video codecs, pictures are spatially subdivided into blocks of samples. When all blocks of samples are encoded in intra-mode, the picture may be an intra-picture. Intra-pictures and their derivatives (e.g., independent decoder refresh pictures) may be used to reset the decoder state and may therefore be used as the first picture in a coded video bitstream and video session, or as a still image. The samples of the intra-block may be transformed, and the transform coefficients may be quantized prior to entropy encoding. Intra-prediction may be a technique for minimizing sample values ​​in a pre-transform domain. In some cases, the smaller the transformed DC value, the smaller the AC coefficient, and the fewer bits required to represent the entropy-coded block at a given quantization step size.

[0009] For example, conventional intra-frame coding known from MPEG-2 generation coding techniques does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to use surrounding sample data and / or metadata obtained during encoding / decoding of, for example, spatially adjacent and preceding data blocks in decoding order. These techniques are hereinafter referred to as "intra-frame prediction" techniques. Note that, at least in some cases, intra-frame prediction uses only reference data from the current picture being reconstructed, and not reference data from reference pictures.

[0010] There can be many different forms of intra prediction. When more than one such technique can be used in a given video coding technique, the technique used can be encoded in an intra prediction mode. In some cases, a mode can have sub-modes and / or parameters, which can be encoded separately or contained in a mode codeword. For a given mode / sub-mode / parameter combination, which codeword is used will have an impact on the coding efficiency gain through intra prediction and, therefore, the entropy coding technique used to convert the codeword into a bitstream.

[0011] H.264 introduced specific intra prediction modes, improved in H.265, and further improved in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). A prediction block can be formed using neighboring sample values ​​belonging to already available samples. Sample values ​​of neighboring samples are copied into the prediction block according to the direction. A reference to the direction in use can be encoded in the bitstream or can be predicted itself.

[0012] refer to Figure 1A , a subset of 9 known prediction directions from the 33 possible prediction directions of H.265 (corresponding to the 33 angular modes of the 35 intra modes) is depicted at the bottom right. The point (101) where the arrows converge represents the sample being predicted. The arrows indicate from which direction the sample is being predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45° to the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left of sample (101) at an angle of 22.5° to the horizontal.

[0013] Still reference Figure 1A , a square block (104) of 4×4 samples is depicted in the upper left (indicated by the thick dashed line). The square block (104) includes 16 samples, each sample labeled "S", its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the size of the block is 4×4 samples, S44 is located in the lower right corner. Reference samples are also shown following a similar numbering scheme. The reference samples are labeled with R, their Y position (e.g., row index) and X position (column index) relative to the block (104). In H.264 and H.265, the prediction samples are adjacent to the block being reconstructed; therefore, there is no need to use negative values.

[0014] Intra picture prediction can work by copying reference sample values ​​from neighboring samples using a signaled prediction direction. For example, assume that the coded video bitstream contains signaling that indicates a prediction direction consistent with arrow (102) for this block - that is, multiple samples are predicted from one or more prediction samples to the upper right at a 45° angle to the horizontal. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Sample S44 is then predicted from reference sample R08.

[0015] In some cases, the values ​​of multiple reference samples may be combined, for example, by interpolation, in order to calculate the reference sample; in particular, when the direction is not divisible by 45°.

[0016] As video coding techniques have evolved, the number of possible directions has increased. In H.264 (2003), nine different directions could be represented. This was increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of disclosure. Experiments have been done to identify the most likely directions, and certain techniques in entropy coding are used to represent those possible directions with a small number of bits, accepting some penalty for less likely directions. Additionally, the direction itself can sometimes be predicted from adjacent directions used in adjacent already decoded blocks.

[0017] Figure 1B A schematic diagram (180) is shown depicting 65 intra prediction directions according to JEM to illustrate that the number of prediction directions increases over time.

[0018] The mapping of the intra-prediction direction bits representing directions in the coded video bitstream can vary depending on the video coding technique; and can range from simple direct mappings, such as prediction directions to intra-prediction modes, or to codewords, to complex adaptive schemes involving most probable modes, and similar techniques. However, in all cases, some directions are statistically less likely to appear in the video content than some other directions. Since the goal of video compression is to reduce redundancy, in video coding techniques that work well, those less likely directions will be represented by more bits than more likely directions.

[0019] Motion compensation may be a lossy compression technique and may involve a technique in which a block of sample data from a previously reconstructed picture or part thereof (reference picture) is used to predict a newly reconstructed picture or part of a picture after being spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, the third dimension being an indication of the reference picture in use (the latter may indirectly be a temporal dimension).

[0020] In some video compression techniques, the MV applicable to a region of sample data can be predicted from other MVs, for example, from MVs associated with another region of sample data that is spatially adjacent to the region being reconstructed and precedes the MV in decoding order. Doing so can greatly reduce the amount of data required to encode the MV, thereby eliminating redundancy and improving compression. MV prediction can work effectively, for example, because when encoding an input video signal derived from a camera (called natural video), there is a statistical probability that regions larger than the region to which a single MV can be applied move in similar directions, so in some cases, similar motion vectors derived from MVs of neighboring regions can be used for prediction. This results in the MV found for a given region being similar or identical to the MV predicted from the surrounding MVs, and after entropy coding, this can in turn be represented with fewer bits than would be used if the MV were encoded directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself may be lossy, for example, due to rounding errors when the predicted value is calculated from several surrounding MVs.

[0021] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). In addition to the many MV prediction mechanisms provided by H.265, a technique referred to as "spatial merging" is described hereafter.

[0022] refer to Figure 2, the current block (201) includes samples that the encoder found during motion search to be predictable from a previous block of the same size that has been spatially shifted. Rather than encoding the MV directly, the MV may be derived from metadata associated with one or more reference pictures, e.g., from the most recent (in decoding order) reference picture, using the MV associated with any of the five surrounding samples, denoted A0, A1 and B0, B1, B2 (202 to 206, respectively). In H.265, MV prediction may use a predictor from the same reference picture being used by a neighboring block. Summary of the invention

[0023] Aspects of the present disclosure provide methods and devices for filtering in video codecs. In some examples, a device for video codecs includes a processing circuit. The processing circuit buffers a first boundary pixel value of a first reconstructed sample at a first node along a loop filter chain. The first node is associated with a nonlinear mapping-based filter applied before a loop recovery filter in the loop filter chain. In one example, the first boundary pixel value is the value of a pixel at a frame boundary. The processing circuit applies a loop recovery filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value.

[0024] In some examples, the filter based on non-linear mapping is a cross-component sample offset (CCSO) filter. In some examples, the filter based on non-linear mapping is a local sample offset (LSO) filter.

[0025] In some examples, the first reconstructed sample at the first node is an input to a filter based on a nonlinear mapping. In some examples, the processing circuit buffers a second boundary pixel value of the second reconstructed sample at a second node along the loop filter chain. The second reconstructed sample at the second node is generated after applying a sample offset generated by the filter based on the nonlinear mapping. The second boundary pixel value is the value of a pixel at a frame boundary. The processing circuit can then apply a loop recovery filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value.

[0026] In one example, the processing circuit combines the sample offsets generated by the filter based on the nonlinear mapping with the output of the constrained directional enhancement filter to generate a second reconstructed sample. In another example, the processing circuit combines the sample offsets generated by the filter based on the nonlinear mapping with the first reconstructed sample to generate an intermediate reconstructed sample, and applies the constrained directional enhancement filter to the intermediate reconstructed sample to generate a second reconstructed sample.

[0027] In some examples, the processing circuit buffers a second boundary pixel value of the second reconstructed sample at a second node along the loop filter chain, and combines the second reconstructed sample at the second node with the sample offset generated by the filter based on the nonlinear mapping to generate the reconstructed sample to be filtered. The second boundary pixel value is the value of the pixel at the frame boundary. The processing circuit then applies the loop restoration filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value.

[0028] In some examples, the processing circuit buffers a second boundary pixel value of a second reconstructed sample generated by the deblocking filter, and applies a constrained directional enhancement filter to the second reconstructed sample to generate an intermediate reconstructed sample. The second boundary pixel value is the value of a pixel at a frame boundary. The processing circuit combines the intermediate reconstructed sample with a sample offset generated by a filter based on a nonlinear mapping to generate a first reconstructed sample. Then, a loop recovery filter may be applied based on the buffered first boundary pixel value and the buffered second boundary pixel value.

[0029] In some examples, the processing circuitry clips the reconstructed samples to be filtered before applying the loop restoration filter. In some examples, the processing circuitry clips the intermediate reconstructed samples before combining with the sample offsets generated by the filter based on the nonlinear mapping.

[0030] Aspects of the present disclosure also provide a video encoding method, including: reconstructing pixel values ​​in a current image frame to obtain reconstructed samples; applying a deblocking filter to the reconstructed samples to obtain a first reconstructed sample; buffering a first boundary pixel value of a subset of the first reconstructed samples at a first node along a loop filter chain, the first boundary pixel value is the value of a pixel at a frame boundary; buffering a second boundary pixel value of a subset of the second reconstructed samples at a second node along the loop filter chain, the second reconstructed sample at the second node is generated after applying at least one of the first filter or the second filter to the first reconstructed sample, and the second boundary pixel value is the value of a pixel at a frame boundary; and applying a loop recovery filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value to obtain a reconstructed sample after filtering; and performing encoding processing based on the reconstructed sample after filtering.

[0031] Aspects of the present disclosure also provide a video decoding method, including: decoding a received encoded bit stream to obtain a decoded symbol; reconstructing pixels based on the decoded symbol to obtain a reconstructed sample; applying the reconstructed sample to a deblocking filter to obtain a first reconstructed sample; buffering a first boundary pixel value of a subset of the first reconstructed sample at a first node along a loop filter chain; buffering a second boundary pixel value of a subset of the second reconstructed sample at a second node along the loop filter chain, the second reconstructed sample at the second node being generated after applying at least one of the first filter or the second filter to the first reconstructed sample; applying a loop recovery filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value to obtain a filtered reconstructed sample; and outputting a decoded image frame based on the filtered reconstructed sample.

[0032] Aspects of the present disclosure also provide a method for generating a video bitstream, which generates a coded bitstream of a video according to the above video coding method.

[0033] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions, which, when executed by a computer, cause the computer to perform any of the above methods for video encoding / decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0035] Figure 1A is a schematic diagram of an exemplary subset of intra prediction modes;

[0036] Figure 1B is a diagram of exemplary intra prediction directions;

[0037] Figure 2 is a schematic diagram of a current block and its surrounding spatial merging candidates in an example;

[0038] Figure 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to one embodiment;

[0039] Figure 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment;

[0040] Figure 5 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment;

[0041] Figure 6 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment;

[0042] Figure 7shows a block diagram of an encoder according to another embodiment;

[0043] Figure 8 shows a block diagram of a decoder according to another embodiment;

[0044] Fig. 9 shows an example of a filter shape according to an embodiment of the present disclosure;

[0045] Figures 10A-10D An example of sub-sampling positions for calculating gradients according to an embodiment of the present disclosure is shown;

[0046] Figures 11A-11B An example of a virtual boundary filtering process according to an embodiment of the present disclosure is shown;

[0047] Figures 12A-12F An example of a symmetric fill operation at a virtual boundary according to an embodiment of the present disclosure is shown;

[0048] Fig.13 shows an example of segmentation of a picture according to some embodiments of the present disclosure;

[0049] Fig.14 shows the quadtree partitioning pattern of some example pictures;

[0050] Fig.15 A cross component filter according to an embodiment of the present disclosure is shown;

[0051] Fig.16 shows an example of a filter shape according to an embodiment of the present disclosure;

[0052] Fig.17 shows an example of syntax for a cross-component filter according to some embodiments of the present disclosure;

[0053] Figures 18A-18B shows exemplary positions of chroma samples relative to luma samples according to an embodiment of the present invention;

[0054] Fig.19 An example of direction search according to an embodiment of the present disclosure is shown;

[0055] Fig. 20 Examples illustrating subspace projections in some examples are shown;

[0056] Fig.21 A table showing multiple sample adaptive offset (SAO) types according to an embodiment of the present disclosure;

[0057] Fig. 22 Examples of patterns of pixel classification in edge offsets in some examples are shown;

[0058] Fig.23 A table showing pixel classification rules for edge offsets in some examples;

[0059] Fig.24 An example of a syntax that may be signaled is shown;

[0060] Fig.25 shows an example of a filter support region according to some embodiments of the present disclosure;

[0061] Fig.26 shows an example of another filter support region according to some embodiments of the present disclosure;

[0062] Figures 27A-27C A table with 81 combinations according to an embodiment of the present disclosure is shown;

[0063] Fig.28 7 filter shape configurations for 3 filter taps in one example are shown;

[0064] Fig.29 shows a block diagram of a loop filter chain in some examples;

[0065] Fig.30 An example of a loop filter chain including filters based on non-linear mapping is shown;

[0066] Fig.31 An example of another loop filter chain including filters based on non-linear mapping is shown;

[0067] Fig.32 An example of another loop filter chain including filters based on non-linear mapping is shown;

[0068] Fig.33 An example of another loop filter chain including filters based on non-linear mapping is shown;

[0069] Fig.34 An example of another loop filter chain including filters based on non-linear mapping is shown;

[0070] Fig.35 An example of another loop filter chain including filters based on non-linear mapping is shown;

[0071] Fig.36 A flowchart outlining a process according to an embodiment of the present disclosure is shown;

[0072] Fig.37 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION

[0073] Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present disclosure is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). Figure 3 In the example of , a first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (e.g., a video picture stream captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to restore the video picture, and display the video picture based on the restored video data. Unidirectional data transmission is common in media service applications and the like.

[0074] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, such as may occur during a video conference. For bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via a network (350). Each of the terminal devices (330) and (340) can also receive encoded video data transmitted by the other of the terminal devices (330) and (340), and can decode the encoded video data to restore the video picture, and can display the video picture on an accessible display device based on the restored video data.

[0075] exist Figure 3In the example of , terminal devices (310), (320), (330) and (340) can be shown as servers, personal computers and smart phones, but the principles of the present disclosure may not be limited to this. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players and / or dedicated video conferencing equipment. Network (350) represents any number of networks that transmit encoded video data between terminal devices (310), (320), (330) and (340), including, for example, lines (wired) and / or wireless communication networks. Communication network (350) can exchange data in circuit switching and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purpose of the current discussion, the architecture and topology of network (350) may be unimportant to the operation of the present disclosure unless explained below.

[0076] As examples of applications of the disclosed subject matter, Figure 4 The placement of a video encoder and a video decoder in a streaming environment is shown. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital television, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0077] The streaming system may include: a capture subsystem (413), which may include a video source (401), such as a digital camera; creating, for example, an uncompressed video picture stream (402). In one example, the video picture stream (402) includes samples captured by the digital camera. The video picture stream (402) is depicted as thick lines to emphasize the high amount of data when compared to the encoded video data (404) (or encoded video bitstream), and can be processed by an electronic device (420) including a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is depicted as thin lines to emphasize the lower amount of data when compared to the video picture stream (402), and can be stored on a streaming server (405) for future use. One or more streaming client subsystems (e.g., Figure 4The client subsystems (406) and (408) in the video server (405) can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the input copy (407) of the encoded video data and creates an output video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or other presentation device (not shown). In some streaming systems, the encoded video data (404), (407) and (409) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0078] Note that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0079] Figure 5 A block diagram of a video decoder (510) according to an embodiment of the present disclosure is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 A video decoder (410) is shown in an example.

[0080] A receiver (531) may receive one or more encoded video sequences to be decoded by a video decoder (510); in the same or another embodiment, one at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not shown). The receiver (531) may separate the encoded video sequence from the other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and an entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, it may be external to the video decoder (510) (not shown). In other cases, there may be a buffer memory (not shown) outside the video decoder (510), for example, to combat network jitter, and another buffer memory (515) inside the video decoder (510), for example, to handle playout timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be needed, or may be small. For use over a best-effort packet network such as the Internet, a buffer memory (515) may be required, which may be relatively large and may advantageously have an adaptive size and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).

[0081] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510) and potentially controlling a presentation device such as a presentation device (512) (e.g., a display screen) that is not part of the electronic device (530) but can be coupled to the electronic device (530), such as Figure 5As shown. The control information for the rendering device may be in the form of a supplemental enhancement information (SEI message) or a video usability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be based on a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a set of subgroup parameters of at least one pixel subgroup in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (Cu), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (520) may also extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0082] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0083] Depending on the type of coded video picture or part thereof (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors, the reconstruction of the symbol (521) may involve a number of different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For clarity, such subgroup control information flow between the parser (520) and the following multiple units is not described.

[0084] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into a number of functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.

[0085] The first unit is a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantized transform coefficients and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc., as symbols (521) from the parser (520). The scaler / inverse transform unit (551) can output blocks including sample values, which can be input into the aggregator (555).

[0086] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block; that is, a block that does not use prediction information from a previously reconstructed image, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information obtained from a current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, an aggregator (555) adds the prediction information already generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.

[0087] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-frame coded and possibly motion compensated block. In this case, the motion compensated prediction unit (553) may access the reference picture memory (557) to obtain samples for prediction. After the extracted samples are motion compensated according to the symbols (521) associated with the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case referred to as residual samples or residual signal) to generate output sample information. The addresses within the reference picture memory (557) from which the motion compensated prediction unit (553) obtains the predicted samples may be controlled by motion vectors, and the motion compensated prediction unit (553) may obtain these addresses in the form of symbols (521), which may have, for example, X, Y and reference picture components. When sub-sampled accurate motion vectors are used, motion compensation may also include interpolation of sample values ​​obtained from the reference picture memory (557), motion vector prediction mechanisms, etc.

[0088] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). The video compression techniques may include loop filtering techniques which are controlled by parameters contained in the coded video sequence (also referred to as the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of the coded picture or coded video sequence and to previously reconstructed and loop filtered sample values.

[0089] The output of the loop filter unit (556) may be a sample stream that may be output to a rendering device (512) and stored in a reference picture memory (557) for use in future inter-picture prediction.

[0090] Once fully reconstructed, certain coded pictures can be used as reference pictures for future predictions. For example, once a coded picture corresponding to a current picture is fully reconstructed, and the coded picture has been identified as a reference picture (e.g., by a parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting reconstruction of the next coded picture.

[0091] The video decoder (510) may perform decoding operations according to a predetermined video compression technique in a standard such as ITU-T Rec. H.265. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence conforms to the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available under the profile. Conformance to the standard also requires that the complexity of the encoded video sequence is within the range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the metadata of the hypothetical reference decoder (HRD) specification and HRD buffer management signaled in the encoded video sequence.

[0092] In one embodiment, the receiver (531) can receive additional (redundant) data with the encoded video. The additional data can be included as part of the encoded video sequence. The video decoder (510) can use the additional data to correctly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0093] Figure 6 A block diagram of a video encoder (603) according to an embodiment of the present disclosure is shown. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) can be used to replace Figure 4 The video encoder (403) in the example.

[0094] The video encoder (603) can be used to obtain the video source (601) (which is not Figure 6 In another example, the video source (601) is a part of the electronic device (620) to receive the video samples, and the video source can capture the video pictures to be encoded by the video encoder (603). In another example, the video source (601) is a part of the electronic device (620).

[0095] The video source (601) may provide a source video sequence to be encoded by the video encoder (603) in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. Video data may be provided as multiple individual pictures that impart motion when viewed sequentially. The pictures themselves may be organized as a spatial array of pixels, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. in use. The relationship between pixels and samples may be readily understood by those skilled in the art. The following description focuses on samples.

[0096] According to one embodiment, the video encoder (603) can encode and compress the pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing the appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to the other functional units. For clarity, the coupling is not described. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, ...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions related to the video encoder (603) optimized for a specific system design.

[0097] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an oversimplified description, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, e.g., a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols in a manner similar to what the (remote) decoder will also create to create sample data (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream results in a bit-accurate result independent of the decoder location (local or remote), the contents in the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, when prediction is used during decoding, the predicted part of the encoder "sees" as a reference picture sample exactly the same sample value as the sample value "seen" by the decoder. This basic principle of reference picture synchronization (and the resulting drift if the synchronization cannot be maintained, eg due to channel errors) is also used in some related techniques.

[0098] The operation of the "local" decoder (633) may be identical to the operation of a "remote" decoder such as the video decoder (510), described above in conjunction with Figure 5 However, a brief reference to Figure 5 , since symbols are available and the symbol encoding / decoding of the encoded video sequence by the entropy encoder (645) and the parser (520) can be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633).

[0099] It can be observed at this point that, in addition to the parsing / entropy decoding present in the decoder, any decoder technology must also be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because these technologies are the inverse of the fully described decoder technology. Only in certain areas is a more detailed description required and provided below.

[0100] During operation, in some examples, the source encoder (630) may perform motion compensated predictive coding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence that are designated as “reference pictures.” In this manner, the encoding engine (632) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.

[0101] The local video decoder (633) can decode the encoded video data of the picture that can be designated as the reference picture based on the symbols created by the source encoder (630). The operation of the encoding engine (632) can advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 6 When decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture, and may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) may locally store copies of the reconstructed reference pictures that have the same content (absent transmission errors) as the reconstructed reference pictures that will be obtained by the remote video decoder.

[0102] The predictor (635) may perform a prediction search on the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (635) may operate on a sample block-pixel block basis to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references extracted from multiple reference pictures stored in the reference picture memory (634).

[0103] The controller (650) may manage encoding operations of the source encoder (630), including, for example, the setting of parameters and sub-group parameters for encoding video data.

[0104] The outputs of all the aforementioned functional units may undergo entropy coding in an entropy encoder (645). The entropy encoder (645) converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0105] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission via a communication channel (660), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0106] The controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a specific coded picture type to each coded picture, which can affect the coding techniques that can be applied to the corresponding picture. For example, a picture can generally be assigned one of the following picture types:

[0107] An intra picture (I picture) may be a picture that is encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of those variants of I pictures and their corresponding applications and features.

[0108] A prediction picture (P picture) may be a picture that is encoded and decoded using intra prediction or inter prediction by predicting sample values ​​of each block using at most one motion vector and a reference index.

[0109] Bidirectional prediction pictures (B pictures) can be pictures that use up to two motion vectors and reference indices to predict the sample values ​​of each block, and are encoded and decoded using intra-frame prediction or inter-frame prediction. Similarly, multi-prediction pictures can use more than two reference pictures and related metadata to reconstruct a single block.

[0110] A source picture may typically be spatially subdivided into a plurality of blocks of samples (e.g., 4×4, 8×8, 4×8, or 16×16 blocks of samples each), and encoded on a block-by-block basis. Blocks may be predictively encoded with reference to other (already encoded) blocks determined by the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively encoded, or may be predictively encoded (spatial prediction or intra prediction) with reference to already encoded blocks of the same picture. Blocks of pixels of a P picture may be predictively encoded via spatial prediction or via temporal prediction with reference to one previously encoded reference picture. Blocks of a B picture may be predictively encoded via spatial prediction or via temporal prediction with reference to one or two previously encoded reference pictures.

[0111] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard (e.g., ITU-T Rec. H.265). In its operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.

[0112] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data (e.g., redundant pictures and slices), SEI messages, VUI parameter set fragments, etc.

[0113] Video may be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture in encoding / decoding, referred to as the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. In the case of using multiple reference pictures, the motion vector points to a reference block in a reference picture and may have a third dimension that identifies the reference picture.

[0114] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture, both of which are before the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0115] In addition, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0116] According to some embodiments of the present disclosure, prediction is performed in units of blocks, for example, inter-picture prediction and intra-picture prediction. For example, according to the HEVC standard, a picture in a video picture sequence is divided into coding tree units (CTUs) for compression, and the CTUs in the picture have the same size, for example, 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), namely a luminance CTB and two chrominance CTBs. Each CTU can be recursively quadtree-divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, for example, an inter-prediction type or an intra-prediction type. According to temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block contains a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0117] Figure 7 A diagram of a video encoder (703) according to another embodiment of the present disclosure is shown. The video encoder (703) is configured to receive a processed block (e.g., a prediction block) of sample values ​​within a current video picture in a sequence of video pictures and encode the processed block into an encoded image as part of an encoded video sequence. In one example, the video encoder (703) is used instead of Figure 4 The video encoder (403) in the example.

[0118] In the HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a prediction block of 8×8 samples, etc. The video encoder (703) determines whether to best encode the processing block using intra mode, inter mode, or a bidirectional prediction mode such as rate-distortion optimization. When the processing block is to be encoded in intra mode, the video encoder (703) may encode the processing block into a coded picture using intra prediction techniques; and when the processing block is to be encoded in inter mode or bidirectional prediction mode, the video encoder (703) may encode the processing block into a coded picture using inter prediction or bidirectional prediction techniques, respectively. In some video coding techniques, the merge mode may be an inter picture prediction submode, in which motion vectors are derived from one or more motion vector predictors without benefiting from a coded motion vector component outside the predictor. In some other video coding techniques, there may be a motion vector component applicable to the object block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0119] exist Figure 7 In the example of FIG. 7 , the video encoder ( 703 ) includes Figure 7 An inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) are shown coupled together.

[0120] The inter-frame encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-frame prediction information (e.g., description of redundant information according to an inter-frame coding technique, motion vectors, merge mode information), and calculate an inter-frame prediction result (e.g., a prediction block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0121] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with blocks already encoded in the same picture, generate quantization coefficients after transformation, and in some cases, generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). In one example, the intra encoder (722) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.

[0122] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, when the mode is intra-frame mode, the general controller (721) controls the switch (726) to select the intra-frame mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select the intra-frame prediction information and include the intra-frame prediction information in the bitstream; when the mode is inter-frame mode, the general controller (721) controls the switch (726) to select the inter-frame prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select the inter-frame prediction information and include the inter-frame prediction information in the bitstream.

[0123] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate a transform coefficient. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate a transform coefficient. The transform coefficient is then quantized to obtain a quantized transform coefficient. In various embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded blocks are processed appropriately to generate a decoded picture, and the decoded picture may be cached in memory circuitry (not shown) and used as a reference picture in some examples.

[0124] The entropy encoder (725) is configured to format the bitstream to include the coded blocks. The entropy encoder (725) is configured to include various information according to a suitable standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. Note that according to the disclosed subject matter, when the block is encoded in inter-frame mode or the merge sub-mode of the bidirectional prediction mode, there is no residual information.

[0125] Figure 8A diagram of a video decoder (810) according to another embodiment of the present disclosure is shown. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (810) is used to replace Figure 4 A video decoder (410) is shown in an example.

[0126] exist Figure 8 In the example of FIG. 8 , the video decoder ( 810 ) includes Figure 8 An entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874) and an intra-frame decoder (872) are shown coupled together.

[0127] The entropy decoder (871) can be configured to reconstruct certain symbols representing syntax elements constituting the coded image from the coded image. Such symbols may include, for example, a mode for encoding a block (e.g., intra mode, inter mode, bidirectional prediction mode, merge submode, or the latter two modes in another submode), prediction information (e.g., intra prediction information or inter prediction information) that can identify a specific sample or metadata used for prediction by an intra decoder (872) or an inter decoder (880), respectively, residual information such as quantized transform coefficients, etc. In one example, when the prediction mode is an inter or bidirectional prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information may be subjected to inverse quantization and provided to the residual decoder (873).

[0128] The inter-frame decoder (880) is configured to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.

[0129] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0130] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (QP)), and this information may be provided by the entropy decoder (871) (the data path is not shown because this may only be a small amount of control information).

[0131] The reconstruction module (874) is configured to combine the residual output by the residual decoder (873) and the prediction result (output by the inter-frame or intra-frame prediction module, as appropriate) in the spatial domain to form a reconstructed block, which can be part of a reconstructed picture, which in turn can be part of a reconstructed video. Note that other suitable operations such as deblocking operations can be performed to improve visual quality.

[0132] Note that the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using any suitable technology. In one embodiment, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603) and (603) and the video decoders (410), (510) and (810) may be implemented using one or more processors that execute software instructions.

[0133] Aspects of the present invention provide filtering techniques for video encoding / decoding.

[0134] The encoder / decoder may apply an adaptive loop filter (ALF) with block-based filter adaptation to reduce artifacts.For the luma component, for example, one of multiple filters (eg, 25 filters) may be selected for a 4×4 luma block based on the direction and activity of the local gradient.

[0135] The ALF may have any suitable shape and size. Fig. 9 , ALF (910)-(911) has a rhombus shape, for example, a 5×5 rhombus shape of ALF (910) and a 7×7 rhombus shape of ALF (911). In ALF (910), elements (920)-(932) form a rhombus and can be used in the filtering process. Seven values ​​(for example, C0-C6) can be used for elements (920)-(932). In ALF (911), elements (940)-(964) form a rhombus and can be used in the filtering process. Thirteen values ​​(for example, C0-C12) can be used for elements (940)-(964).

[0136] refer to Fig. 9 In some examples, two ALFs (910)-(911) with diamond filter shapes are used. A 5×5 diamond filter (910) may be applied to chroma components (e.g., chroma blocks, chroma CBs), while a 7×7 diamond filter (911) may be applied to luma components (e.g., luma blocks, luma CBs). Other suitable shapes and sizes may be used in the ALFs. For example, a 9×9 diamond filter may be used.

[0137] The filter coefficients at the positions indicated by the values ​​(e.g., C0-C6 in (910) or C0-C12 in (920)) may be non-zero. Furthermore, when the ALF includes a clipping function, the clipping values ​​at these positions may be non-zero.

[0138] For block classification of luma components, a 4×4 block (or luma block, luma CB) may be classified or categorized into one of a plurality of (eg, 25) categories. The quantized values ​​of the directionality parameter D and the activity value A may be used to classify the 4×4 block into one of a plurality of (eg, 25) categories. The classification index C is derived using Equation 1.

[0139]

[0140] In order to calculate the directivity parameter D and the quantization value The gradient g in vertical, horizontal, and two diagonal directions (e.g., d1 and d2) v , g h , g d1 and g d2 can be calculated using the 1-D Laplacian respectively as follows.

[0141]

[0142] Where the indices i and j refer to the coordinates of the top left sample within the 4×4 block, and R(k,l) represents the reconstructed sample at the coordinates (k,l). The directions (eg, d1 and d2) may refer to two diagonal directions.

[0143] To reduce the complexity of the above block classification, a sub-sampled 1-D Laplacian calculation can be applied. Figures 10A-10D The calculations for the vertical direction ( Fig. 10A ), horizontal direction( Fig. 10B ) and two diagonal directions d1( Fig. 10C ) and d2( Fig. 10D )’s gradient g v , g h , g d1 and g d2 The same subsampling position can be used for gradient calculations in different directions. Fig. 10A The label “V” shows the vertical gradient g v The sub-sampling position of Fig. 10B The label “H” shows the method used to calculate the horizontal gradient g. h The sub-sampling position of Fig. 10C The label "d1" shows the calculation method for the diagonal gradient g of d1. d1 The sub-sampling position of Fig. 10DIn the figure, label “d2” shows the number of diagonal gradients g used to calculate d2. d2 The subsampling position.

[0144] Horizontal and vertical gradients g v and g h The maximum value of and minimum value Can be set to:

[0145]

[0146] The maximum value of the gradients in the two diagonal directions gd1 and gd2 and minimum value Can be set to:

[0147]

[0148] It can be based on the above value and the following two thresholds t 1 and t 2 To derive the directivity parameter D.

[0149] Step 1: If (1) and (2) is true, D is set to 0.

[0150] Step 2: If Then proceed to step 3; otherwise, proceed to step 4.

[0151] Step 3: If Then D is set to 2; otherwise D is set to 1.

[0152] Step 4: If Then D is set to 4; otherwise D is set to 3.

[0153] The activity value A can be calculated as:

[0154]

[0155] A can be further quantized to a range of 0 to 4 (inclusive), and the quantized value is represented as

[0156] For the chroma components in a picture, block classification is not applied, so a single set of ALF coefficients can be applied for each chroma component.

[0157] The geometric transformation may be applied to the filter coefficients and the corresponding filter clipping values ​​(also referred to as clipping values). Before filtering a block (e.g., a 4×4 luma block), for example, based on the gradient values ​​(e.g., g) calculated for the block, v , g h , gd1 and / or d2 ), geometric transformations such as rotation or diagonal and vertical flipping may be applied to the filter coefficients f(k, l) and the corresponding filter cropping values ​​c(k, l). The geometric transformation applied to the filter coefficients f(k, l) and the corresponding filter cropping values ​​c(k, l) may be equivalent to applying the geometric transformation to samples in the region supported by the filter. The geometric transformation may make different blocks to which the ALF is applied more similar by aligning the corresponding directionality.

[0158] Three geometric transformations including diagonal flipping, vertical flipping and rotation can be performed as described in equations (9)-(11), respectively.

[0159] f D (k,l)=f(l,k),c D (k, l) = c(l, k) Equation (9)

[0160] f v (k,l)=f(k,Kl-1),c v (k, l) = c(l, Kl-1) Equation (10)

[0161] f R (k,l)=f(Kl-1,k),c R (k, l) = c(Kl-1, k) Equation (11)

[0162] Where K is the size of the ALF or filter, and 0≤k, l≤K-1 are the coordinates of the coefficients. For example, position (0, 0) is at the upper left corner of the filter f or the cropping value matrix (or cropping matrix) c, and position (K-1, K-1) is at the lower right corner. Depending on the gradient value calculated for the block, a transform can be applied to the filter coefficients f(k, l) and the cropping values ​​c(k, l). An example of the relationship between the transform and the four gradients is summarized in Table 1.

[0163] Table 1: Mapping of gradients and transformations computed for a block

[0164] Gradient Value Transform <![CDATA[g d2 <g d1 And g h <g v ]]> No transformation <![CDATA[g d2 <g d1 And g v <g h ]]> Flip diagonally <![CDATA[g d1 <g d2 And g h <g v ]]> Flip Vertically <![CDATA[g d1 <g d2 And g v <g h ]]> Rotation

[0165] In some embodiments, the ALF filter parameters are signaled in an adaptive parameter set (APS) of a picture. In the APS, one or more groups (e.g., up to 25 groups) of luminance filter coefficients and cropping value indices may be signaled. In one example, one of the one or more groups may include luminance filter coefficients and one or more cropping value indices. One or more groups (e.g., up to 8 groups) of chrominance filter coefficients and cropping value indices may be signaled. In order to reduce signaling overhead, filter coefficients of different classifications (e.g., with different classification indices) of luminance components may be merged. In a slice header, the index of the APS for the current slice may be signaled.

[0166] In one embodiment, a cropping value index (also referred to as a cropping index) may be decoded from the AP. For example, based on the relationship between the cropping value index and the corresponding cropping value, the cropping value index may be used to determine the corresponding cropping value. The relationship may be predefined and stored in the decoder. In one example, the relationship is described by a table, such as a brightness table of the cropping value index and the corresponding cropping value (e.g., for a brightness CB), a chrominance table of the cropping value index and the corresponding cropping value (e.g., for a chrominance CB). The cropping value may depend on the bit depth B. The bit depth B may refer to an internal bit depth, a bit depth of a reconstructed sample in a CB to be filtered, etc. In some examples, a table (e.g., a brightness table, a chrominance table) is obtained using equation (12).

[0167]

[0168] Wherein, AlfClip is the clipping value, B is the bit depth (e.g., bitDepth), N (e.g., N=4) is the number of allowed clipping values, and (n-1) is the clipping value index (also referred to as clipping index or clipIdx). Table 2 shows an example of a table obtained using equation (12), where N=4. In Table 2, the clipping index (n-1) can be 0, 1, 2, and 3, and n can be 1, 2, 3, and 4, respectively. Table 2 can be used for luma blocks or chroma blocks.

[0169] Table 2: AlfClip may depend on bit depth B and clipIdx

[0170]

[0171] In the slice header of the current slice, one or more APS indexes (e.g., up to 7 APS indexes) may be signaled to specify the luma filter group that may be used for the current slice. The filtering process may be controlled at one or more suitable levels, e.g., picture level, slice level, CTB level, etc. In one embodiment, the filtering process may be further controlled at the CTB level. A flag may be signaled to indicate whether the ALF is applied to the luma CTB. The luma CTB may select a filter group from a plurality of fixed filter groups (e.g., 16 fixed filter groups) and a filter group signaled in the APS (also referred to as a signaled filter group). The filter group index of the luma CTB may be signaled to indicate the filter group to be applied (e.g., a filter group in a plurality of fixed filter groups and a signaled filter group). A plurality of fixed filter groups may be predefined and hard-coded in the encoder and decoder and may be referred to as a predefined filter group.

[0172] For chroma components, the APS index may be signaled in the slice header to indicate the chroma filter set to be used for the current slice. At the CTB level, if more than one chroma filter set exists in the AP, the filter set index may be signaled for each chroma CTB.

[0173] The filter coefficients may be quantized with a norm equal to 128. To reduce multiplication complexity, bitstream consistency may be applied so that coefficient values ​​at non-center positions may be in the range of -27 to 27-1, inclusive. In one example, the center position coefficients are not signaled in the bitstream and may be considered equal to 128.

[0174] In some embodiments, the syntax and semantics of the clip index and clip value are defined as follows:

[0175] alf_luma_clip_idx[sfIdx][j] may be used to specify the clip index of the clip value, which is then multiplied by the jth coefficient of the signaled luma filter indicated by sfIdx. Bitstream conformance requirements may include that the value of alf_luma_clip_idx[sfIdx][j] for sfIdx=0 to alf_luma_num_filters_signalled_minus1 and j=0 to 11 shall be in the range of 0 to 3, inclusive.

[0176] The luma filter clipping values ​​AlfClipL[adaptation_parameter_set_id] with elements AlfClipL[adaptation_parameter_set_id][filtIdx][j] may be derived as specified in Table 2, where filtIdx = 0 to NumAlfFilters-1, and j = 0 to 11, depending on bitDepth being set equal to BitDepthY and clipIdx being set equal to alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx]][j].

[0177] alf_chroma_clip_idx[altIdx][j] may be used to specify the clipping index of the clipping value, which is then multiplied by the jth coefficient of the alternative chroma filter indexed by altIdx. Bitstream conformance requirements may include that the value of alf_chroma_clip_idx[altIdx][j] for altIdx=0 to alf_chroma_num_alt_filters_minus1 and j=0 to 5 shall be in the range of 0 to 3, inclusive.

[0178] The chroma filter clipping value AlfClipC[adaptation_parameter_set_id][altIdx] with elements AlfClipC[adaptation_parameter_set_id][altIdx][j] may be derived as specified in Table 2, where altIdx = 0 to alf_chroma_num_alt_filters_minus1, and j = 0 to 5, depending on bitDepth being set equal to BitDepthC, clipIdx being set equal to alf_chroma_clip_idx[altIdx][j].

[0179] In one embodiment, the filtering process can be described as follows. At the decoder side, when ALF is enabled for a CTB, the sample R(i, j) within the CU (or CB) can be filtered to produce a filtered sample value R'(i, j), as shown below using equation (13). In one example, each sample in the CU is filtered.

[0180]

[0181] Wherein, f(k, l) represents the decoded filter coefficient, K(x, y) is the clipping function, and c(k, l) represents the decoded clipping parameter (or clipping value). The variables k and l can vary between -L / 2 and L / 2, where L represents the filter length. The clipping function K(x, y) = min(y, max(-y, x)) corresponds to the clipping function Clip3(-y, y, x). By combining the clipping function K(x, y), the loop filtering method (e.g., ALF) becomes a nonlinear process and can be called a nonlinear ALF.

[0182] In a non-linear ALF, multiple groups of clipping values ​​may be provided in Table 3. In one example, the luma group includes four clipping values ​​{1024, 181, 32, 6} and the chroma group includes four clipping values ​​{1024, 161, 25, 4}. The four clipping values ​​in the luma group may be selected by approximately equally dividing the full range (e.g., 1024) of the sample values ​​of the luma block (encoded in 10 bits) in the logarithmic domain. The range of the chroma set may be from 4 to 1024.

[0183] Table 3: Examples of clipping values

[0184]

[0185] The selected cropping value may be encoded in the "alf_data" syntax element as follows: A suitable coding scheme (e.g., Golomb coding scheme) may be used to encode the cropping index corresponding to the selected cropping value, as shown in Table 3. The coding scheme may be the same as the coding scheme used to encode the filter bank index.

[0186] In one embodiment, a virtual boundary filtering process may be used to reduce the line buffer requirements of the ALF. Thus, modified block classification and filtering may be used for samples near a CTU boundary (e.g., a horizontal CTU boundary). The virtual boundary (1130) may be achieved by moving the horizontal CTU boundary (1120) by “N”. samples "A sample is defined as a line, such as Fig.11A As shown, where N samples Can be a positive integer. In one example, for the brightness component, N samples Equal to 4, for chrominance components, N samples Equals 2.

[0187] refer to Fig.11A , the modified block classification can be applied to the luma component. In one example, for the 1D Laplacian gradient calculation of the 4×4 block (1110) above the virtual boundary (1130), only the samples above the virtual boundary (1130) are used. Similarly, referring to Fig. 11B, for the 1D Laplacian gradient calculation of the 4×4 block (1111) under the virtual boundary (1131) offset from the CTU boundary (1121), only the samples under the virtual boundary (1131) are used. By considering the reduced number of samples used in the 1D Laplacian gradient calculation, the quantization of the activity value A can be scaled accordingly.

[0188] For the filtering process, symmetric padding operations at virtual boundaries can be used for luma and chroma components. Figures 12A-12F An example of such modified ALF filtering of the luma component at a virtual boundary is shown. When a filtered sample is below a virtual boundary, neighboring samples above the virtual boundary may be padded. When a filtered sample is above a virtual boundary, neighboring samples below the virtual boundary may be padded. Fig. 12A , the adjacent sample C0 can be filled with the sample C2 located below the virtual boundary (1210). Fig. 12B , the adjacent sample C0 can be filled with the sample C2 located above the virtual boundary (1220). Fig. 12C , adjacent samples C1-C3 can be filled with samples C5-C7 located below the virtual boundary (1230), respectively. Fig.12D , adjacent samples C1-C3 can be filled with samples C5-C7 located above the virtual boundary (1240), respectively. Fig.12E , adjacent samples C4-C8 can be filled with samples C10, C11, C12, C11, and C10 respectively located below the virtual boundary (1250). Fig.12F , adjacent samples C4-C8 may be filled with samples C10, C11, C12, C11, and C10, respectively, which are located above the virtual boundary (1260).

[0189] In some examples, when the sample and the adjacent sample are located on the left side (or right side) and the right side (or left side) of the virtual boundary, the above description may be appropriately modified.

[0190] According to one aspect of the present disclosure, in order to improve coding efficiency, pictures can be divided based on a filtering process. In some examples, a CTU is also referred to as a maximum coding unit (LCU). In one example, a CTU or LCU may have a size of 64×64 pixels. In some embodiments, LCU-aligned picture quadtree segmentation may be used for filtering-based partitioning. In some examples, an adaptive loop filter based on a coding unit synchronized picture quadtree may be used. For example, a luminance image may be divided into several multi-level quadtree partitions, with each partition boundary aligned with the boundary of the LCU. Each partition has its own filtering process and is therefore referred to as a filtering unit (FU).

[0191] In some examples, a two-pass encoding process can be used. In the first pass of the two-pass encoding process, the quadtree segmentation mode of the picture and the optimal filter for each FU can be determined. In some embodiments, the determination of the quadtree segmentation mode of the picture and the determination of the optimal filter for the FU are based on filtering distortion. In the determination process, the filtering distortion can be estimated by a fast filter distortion estimation (FFDE) technique. The picture is segmented using quadtree segmentation. Based on the determined quadtree segmentation mode and the selected filters of all FUs, the reconstructed picture can be filtered.

[0192] In the second pass of the two-pass encoding process, CU synchronous ALF on / off control is performed. According to the ALF on / off result, the first filtered picture is partially restored by the reconstructed picture.

[0193] Specifically, in some examples, a top-down segmentation strategy is adopted to divide the picture into multiple levels of quadtree partitions by using a rate-distortion criterion. Each partition is called a filter unit (FU). The segmentation process aligns the quadtree partitions with the LCU boundaries. The encoding order of the FU follows the z-scan order.

[0194] Fig.13 An example of partitioning according to some embodiments of the present disclosure is shown. Fig.13 In the example of , the picture (1300) is divided into 10 FUs, and the encoding order is FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8 and FU9.

[0195] Fig.14 The quadtree partitioning pattern (1400) of the picture (1300) is shown. Fig.14 In some examples, the split flag is used to indicate the picture split mode. For example, "1" indicates that quadtree partitioning is performed on the block; and "0" indicates that the block is not further split. In some examples, the minimum size FU has the LCU size, and the minimum size FU does not need a split flag. Fig.14 As shown, the segmentation flags are encoded and transmitted in z order.

[0196] In some examples, the filter for each FU is selected from two filter groups based on a rate-distortion criterion. The first group has 1 / 2 symmetric square and diamond filters derived for the current FU. The second group comes from a delay filter buffer; the delay filter buffer stores filters previously derived for FUs of previous pictures. The filter with the smallest rate-distortion cost in these two groups can be selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further divided into 4 sub-FUs, the rate-distortion costs of the 4 sub-FUs are calculated. By recursively comparing the rate-distortion costs in the segmented and non-segmented cases, the picture quadtree segmentation mode can be determined.

[0197] In some examples, the maximum quadtree split level can be used to limit the maximum number of FUs. In one example, when the maximum quadtree split level is 2, the maximum number of FUs is 16. In addition, during the quadtree split determination, the correlation values ​​used to derive the Wiener coefficients of the 16 FUs (the smallest FUs) at the bottom quadtree level can be reused. The remaining FUs can derive their Wiener filters from the correlations of the 16 FUs at the bottom quadtree level. Therefore, in this example, only one frame buffer access is performed to derive the filter coefficients of all FUs.

[0198] After deciding the quadtree partitioning mode, in order to further reduce the filtering distortion, CU synchronous ALF on / off control can be performed. By comparing the filtering distortion and non-filtering distortion at each leaf CU, the leaf CU can explicitly turn on / off ALF in its local area. In some examples, the coding efficiency can be further improved by redesigning the filter coefficients according to the ALF on / off results.

[0199] The cross-component filtering process may apply a cross-component filter, such as a cross-component adaptive loop filter (CC-ALF). The cross-component filter may refine a chroma component (e.g., a chroma CB corresponding to the luma CB) using luma sample values ​​of a luma component (e.g., a luma CB). In one example, the luma CB and the chroma CB are included in a CU.

[0200] Fig.15 A cross component filter (e.g., CC-ALF) for generating chrominance components according to an embodiment of the present invention is shown. In some examples, Fig.15 The filtering process for a first chroma component (e.g., a first chroma CB), a second chroma component (e.g., a second chroma CB), and a luma component (e.g., a luma CB) is shown. The luma component may be filtered by a sample adaptive offset (SAO) filter (1510) to generate an SAO filtered luma component (1541). The SAO filtered luma component (1541) may be further filtered by an ALF luma filter (1516) to become a filtered luma CB (1561) (e.g., "Y").

[0201] The first chroma component may be filtered by the SAO filter (1512) and the ALF chroma filter (1518) to generate a first intermediate component (1552). In addition, the SAO filtered luma component (1541) may be filtered by a cross component filter (e.g., CC-ALF) (1521) for the first chroma component to generate a second intermediate component (1542). Subsequently, a filtered first chroma component (1562) (e.g., “Cb”) may be generated based on at least one of the second intermediate component (1542) and the first intermediate component (1552). In one example, the filtered first chroma component (1562) (e.g., “Cb”) may be generated by combining the second intermediate component (1542) and the first intermediate component (1552) using an adder (1522). The cross component adaptive loop filtering process for the first chroma component may include steps performed by the CC-ALF (1521) and steps performed by, for example, the adder (1522).

[0202] The above description may be applied to the second chroma component. The second chroma component may be filtered by the SAO filter (1514) and the ALF chroma filter (1518) to generate a third intermediate component (1553). In addition, the SAO filtered luma component (1541) may be filtered by a cross component filter (e.g., CC-ALF) (1531) for the second chroma component to generate a fourth intermediate component (1543). Subsequently, a filtered second chroma component (1563) (e.g., "Cr") may be generated based on at least one of the fourth intermediate component (1543) and the third intermediate component (1553). In an example, the filtered second chroma component (1563) (e.g., "Cr") may be generated by combining the fourth intermediate component (1543) and the third intermediate component (1553) with an adder (1532). In an example, the cross component adaptive loop filtering process for the second chroma component may include steps performed by the CC-ALF (1531) and steps performed by, for example, the adder (1532).

[0203] The cross-component filter (e.g., CC-ALF (1521), CC-ALF (1531)) may operate by applying a linear filter having any suitable filter shape to the luma component (or luma channel) to refine each chroma component (e.g., first chroma component, second chroma component).

[0204] Fig.16An example of a filter (1600) according to an embodiment of the present disclosure is shown. The filter (1600) may include non-zero filter coefficients and zero filter coefficients. The filter (1600) has a diamond shape (1620) (indicated by a circle with black fill) formed by the filter coefficients (1610). In one example, the non-zero filter coefficients in the filter (1600) are included in the filter coefficients (1610), and the filter coefficients not included in the filter coefficients (1610) are zero. Therefore, the non-zero filter coefficients in the filter (1600) are included in the diamond (1620), and the filter coefficients not included in the diamond (1620) are zero. In one example, the number of filter coefficients of the filter (1600) is equal to the number of filter coefficients (1610), Fig.16 In the example shown, this is 18.

[0205] CC-ALF may include any suitable filter coefficients (also referred to as CC-ALF filter coefficients). Return to Reference Fig.15 , CC-ALF (1521) and CC-ALF (1531) may have the same filter shape, for example, Fig.16 The diamonds (1620) shown in FIG. 16A and have the same number of filter coefficients. In one example, the values ​​of the filter coefficients in CC-ALF (1521) are different from the values ​​of the filter coefficients in CC-ALF (1531).

[0206] Typically, filter coefficients in CC-ALF (e.g., non-zero filter coefficients) can be transmitted, for example, in the AP. In one example, the filter coefficients can be scaled by a factor (e.g., 210) and can be rounded for fixed-point representation. The application of CC-ALF can be controlled on a variable block size and signaled by a context coding flag (e.g., a CC-ALF enable flag) received for each sample block. The context coding flag (e.g., the CC-ALF enable flag) can be signaled at any suitable level (e.g., block level). For each chroma component, the block size and the CC-ALF enable flag can be received at the slice level. In some examples, block sizes of 16×16, 32×32, and 64×64 (in chroma samples) can be supported.

[0207] Fig.17 FIG. 2 shows an example of CC-ALF syntax according to some embodiments of the present disclosure. Fig.17In the example of , alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is an index indicating whether the cross component Cb filter is used and whether the cross component Cb filter is used. For example, when alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, the cross-component Cb filter is not applied to the block of Cb color component samples at the luma position (xCtb, yCtb); when alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not equal to 0, alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is the index of the filter to be applied. For example, the alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]th cross-component Cb filter is applied to a block of Cb color component samples at luma position (xCtb, yCtb).

[0208] In addition, Fig.17 In the example of , alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is used to indicate whether to use the cross component Cr filter and the index of the cross component Cr filter.

[0209] For example, when alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, the cross-component Cr filter is not applied to the block of Cr color component samples at the luma position (xCtb, yCtb); when alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not equal to 0, alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is the index of the cross-component Cr filter. For example, the alf_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]th cross-component Cr filter may be applied to a block of Cr color component samples at luma position (xCtb, yCtb).

[0210] In some examples, a chroma subsampling technique is used, so the number of samples in each chroma block can be less than the number of samples in the luma block. The chroma subsampling format (also referred to as the chroma subsampling format, e.g., specified by chroma_format_idc) can indicate a chroma horizontal subsampling factor (e.g., SubWidthC) and a chroma vertical subsampling factor (e.g., SubHeightC) between each chroma block and the corresponding luma block. In one example, the chroma subsampling format is 4:2:0, so the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 2, as shown in FIG. Figures 18A-18B In one example, the chroma subsampling format is 4:2:2, so the chroma horizontal subsampling factor (e.g., SubWidthC) is 2, and the chroma vertical subsampling factor (e.g., SubHeightC) is 1. In one example, the chroma subsampling format is 4:4:4, so the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 1. The chroma sample type (also called chroma sample position) may indicate the relative position of the chroma samples in the chroma block relative to at least one corresponding luma sample in the luma block.

[0211] Figures 18A-18B 2 shows an exemplary position of chroma samples relative to luma samples according to an embodiment of the present invention. Fig.18A , brightness samples (1801) are located in rows (1811)-(1818). Fig.18AThe luma sample (1801) shown in the figure may represent a portion of a picture. In an example, a luma block (e.g., luma CB) includes luma sample (1801). The luma block may correspond to two chroma blocks having a 4:2:0 chroma subsampling format. In an example, each chroma block includes a chroma sample (1803). Each chroma sample (e.g., chroma sample (1803(1)) corresponds to four luma samples (e.g., luma samples (1801(1))-(1801(4)). In an example, the four luma samples are an upper left sample (1801(1)), an upper right sample (1801(2)), a lower left sample (1801(3)), and a lower right sample (1801(4)). The chroma sample (e.g., (1803(1)) is located between the upper left sample (1801(1)) and the lower left sample (1801(1)). The chroma sample type of the chroma block having the chroma sample (1803) may be referred to as chroma sample type 0. The chroma sample type 0 indicates a relative position 0 corresponding to the left center position between the upper left sample (1801(1)) and the lower left sample (1801(3)). The four luma samples (e.g., (1801(1))-(1801(4)) may be referred to as adjacent luma samples of the chroma sample (1803)(1).

[0212] In one example, each chroma block includes a chroma sample (1804). The above description with reference to the chroma sample (1803) may be applicable to the chroma sample (1804), and thus a detailed description may be omitted for the sake of brevity. Each chroma sample (1804) may be located at a center position of four corresponding luma samples, and a chroma sample type of a chroma block having the chroma sample (1804) may be referred to as a chroma sample type 1. The chroma sample type 1 indicates a relative position 1 (e.g., (1801(1))-(1801(4)) corresponding to the center position of the four luma samples. For example, one chroma sample (1804) may be located at a center portion of luma samples (1801(1))-(1801(4)).

[0213] In one example, each chroma block includes a chroma sample (1805). Each chroma sample (1805) may be located at an upper left position co-located with an upper left sample of four corresponding luma samples (1801), and a chroma sample type of a chroma block having the chroma sample (1805) may be referred to as a chroma sample type 2. Thus, each chroma sample (1805) is co-located with an upper left sample of four luma samples (1801) corresponding to the corresponding chroma sample. The chroma sample type 2 indicates a relative position 2 corresponding to an upper left position of the four luma samples (1801). For example, one chroma sample (1805) may be located at an upper left position of luma samples (1801(1))-(1801(4)).

[0214] In one example, each chroma block includes a chroma sample (1806). Each chroma sample (1806) may be located at a top center position between a corresponding top left sample and a corresponding top right sample, and a chroma sample type of a chroma block having the chroma sample (1806) may be referred to as a chroma sample type 3. The chroma sample type 3 indicates a relative position 3 corresponding to a top center position between the top left sample (and the top right sample). For example, one chroma sample (1806) may be located at a top center position of luma samples (1801(1))-(1801(4)).

[0215] In one example, each chroma block includes a chroma sample (1807). Each chroma sample (1807) may be located at a lower left position co-located with a lower left sample of four corresponding luma samples (1801), and a chroma sample type of a chroma block having the chroma sample (1807) may be referred to as a chroma sample type 4. Thus, each chroma sample (1807) is co-located with a lower left sample of four luma samples (1801) corresponding to the corresponding chroma sample. The chroma sample type 4 indicates a relative position 4 corresponding to the lower left position of the four luma samples (1801). For example, one chroma sample (1807) may be located at a lower left position of luma samples (1801(1))-(1801(4)).

[0216] In one example, each chroma block includes a chroma sample (1808). Each chroma sample (1808) is located at a bottom center position between a bottom left sample and a bottom right sample, and a chroma sample type of a chroma block having the chroma sample (1808) may be referred to as a chroma sample type 5. The chroma sample type 5 indicates a relative position 5 corresponding to a bottom center position between a bottom left sample and a bottom right sample of four luma samples (1801). For example, one chroma sample (1808) may be located between a bottom left sample and a bottom right sample of luma samples (1801(1))-(1801(4)).

[0217] In general, any suitable chroma sampling type may be used for a chroma subsampling format. Chroma sample types 0-5 are exemplary chroma sample types described with the chroma subsampling format 4:2:0. Additional chroma sample types may be used with the chroma subsampling format 4:2:0. Furthermore, other chroma sample types and / or variants of chroma sample types 0-5 may be used with other chroma subsampling formats, e.g., 4:2:2, 4:4:4, etc. In one example, the chroma sample type of the combined chroma samples (1805) and (1807) is used for the chroma subsampling format 4:2:2.

[0218] In one example, a luma block is considered to have alternating rows, e.g., rows (1811)-(1812), which include the top two samples (e.g., (1801(1))-(1801(2))) of four luma samples (e.g., (1801(1))-(1801(4))) and the bottom two samples (e.g., (1801(3))-(1801(4))) of four luma samples (e.g., (1801(1))-(1801(4))). Thus, rows (1811), (1813), (1815), and (1817) may be referred to as current rows (also referred to as top fields), and rows (1812), (1814), (1816), and (1818) may be referred to as next rows (also referred to as bottom fields). Four brightness samples (e.g., (1801(1))-(1801(4))) are located in the current row (e.g., (1811)) and the next row (e.g., (1812)). Relative positions 2-3 are located in the current row, relative positions 0-1 are located between each current row and the corresponding next row, and relative positions 4-5 are located in the next row.

[0219] Chroma samples (1803), (1804), (1805), (1806), (1807), or (1808) are located in rows (1851)-(1854) of each chroma block. The specific location of rows (1851)-(1854) may depend on the chroma sample type of the chroma samples. For example, for chroma samples (1803)-(1804) having corresponding chroma sample types 0-1, row (1851) is located between rows (1811)-(1812). For chroma samples (1805)-(1806) having corresponding chroma sample types 2-3, row (1851) is located at the same position as the current row (1811). For chroma samples (1807)-(1808) having corresponding chroma sample types 4-5, row (1851) is located at the same position as the next row (1812). The above description can be appropriately applied to rows (1852)-(1854), and the detailed description is omitted for the sake of brevity.

[0220] Any suitable scanning method may be used to display, store and / or transmit the above Fig.18A In one example, a progressive scan is used.

[0221] You can use interlaced scanning, such as Fig.18BAs shown. As described above, the chroma subsampling format is 4:2:0 (e.g., chroma_format_idc is equal to 1). In one example, the variable chroma location type (e.g., ChromaLocType) indicates the current row (e.g., ChromaLocType is chroma_sample_loc_type_top_field) or the next row (e.g., ChromaLocType is chroma_sample_loc_type_bottom_field). The current row (1811), (1813), (1815), and (1817) and the next row (1812), (1814), (1816), and (1818) may be scanned respectively, for example, the current row (1811), (1813), (1815), and (1817) may be scanned first, and then the next row (1812), (1814), (1816), and (1818) may be scanned. The current row may include luma samples (1801) and the next row may include luma samples (1802).

[0222] Similarly, the corresponding chroma blocks can be scanned interlaced. The rows (1851) and (1853) including the chroma samples (1803), (1804), (1805), (1806), (1807), or (1808) without filling can be referred to as the current row (or current chroma row), and the rows (1852) and (1854) including the chroma samples (1803), (1804), (1805), (1806), (1807), or (1808) with gray filling can be referred to as the next row (or next chroma row). In one example, during interlaced scanning, rows (1851) and (1853) are scanned first, and then rows (1852) and (1854) are scanned.

[0223] In some examples, constrained directional enhancement filtering techniques can be used. Using an in-loop constrained directional enhancement filter (CDEF) can filter out coding artifacts while preserving image details. In one example (e.g., HEVC), the sample adaptive offset (SAO) algorithm can achieve similar goals by defining signal offsets for different categories of pixels. Unlike SAO, CDEF is a nonlinear spatial filter. In some examples, CDEF can be constrained to be easily vectorized (i.e., can be implemented with single instruction multiple data (SIMD) operations). Note that other nonlinear filters, e.g., median filters, bilateral filters, cannot be treated in the same way.

[0224] In some cases, the number of clear artifacts in the encoded image tends to be roughly proportional to the quantization step size. The amount of detail is a property of the input image, but the smallest detail preserved in the quantized image also tends to be proportional to the quantization step size. For a given quantization step size, the magnitude of clear artifacts is usually smaller than the magnitude of detail.

[0225] CDEF can be used to identify the direction of each block, and then adaptive filtering is performed along the identified direction, and to a lesser extent along the direction rotated 45° from the identified direction. In some examples, the encoder can search for filter strength and can explicitly signal the filter strength, which allows a high degree of control over the blur.

[0226] Specifically, in some examples, a direction search is performed on the reconstructed pixels after the deblocking filter. Because these pixels are available to the decoder, the decoder can search for the direction, so in one example, there is no need to signal the direction. In some examples, the direction search can operate on a specific block size, for example, an 8×8 block, which is small enough to fully handle non-straight edges and large enough to reliably estimate the direction when applied to a quantized image. In addition, maintaining a constant direction within an 8×8 area makes vectorization of the filter easier. In some examples, each block (e.g., 8×8) can be compared with a fully oriented block to determine the difference. A fully oriented block is a block in which all pixels along a line in one direction have the same value. In one example, a difference metric can be calculated for the block and each fully oriented block, such as a sum of squared differences (SSD), a root mean square (RMS) error. Then, the fully oriented block with the smallest difference (e.g., the smallest SSD, the smallest RMS, etc.) can be determined, and the direction of the determined fully oriented block can be the direction that best matches the pattern in the block.

[0227] Fig.19 An example of directional search according to an embodiment of the present disclosure is shown. In one example, the block (1910) is a reconstructed 8×8 block and is output from a deblocking filter. Fig.19 In the example of , the directional search can determine a direction for the block (1910) from the 8 directions shown in (1920). 8 fully oriented blocks (1930) are formed corresponding to the 8 directions (1920), respectively. A fully oriented block corresponding to a direction is a block where pixels along a line in the direction have the same value. In addition, a difference metric between the block (1910) and each fully oriented block (1930) can be calculated, such as SSD, RMS error, etc. Fig.19 In the example of , the RMS error is shown as (1940). As shown in (1943), the RMS error between block (1910) and the fully oriented block (1933) is the smallest, so direction (1923) is the direction that best matches the pattern in block (1910).

[0228] After identifying the direction of the block, a nonlinear low-pass directional filter can be determined. For example, the filter taps of the nonlinear low-pass directional filter can be aligned along the identified direction to reduce clear artifacts while retaining directional edges or patterns. However, in some examples, directional filtering alone is sometimes not enough to reduce clear artifacts. In one example, additional filter taps are also used for pixels that are not along the identified direction. In order to reduce the risk of blur, the additional filter taps are handled more conservatively. To this end, CDEF includes primary filter taps and secondary filter taps. In one example, the complete 2-DCDEF filter can be expressed as equation (14):

[0229]

[0230] Where D represents the damping parameter, S (p) represents the strength of the primary filter tap, S (s) represents the strength of the secondary filter tap, round(·) represents the operation of rounding ties away from zero, w represents the filter weight, and f(d,S,D) is a constraint function that operates on the difference between the filtered pixel and each neighboring pixel. In one example, for small differences, the function f(d,S,D) is equal to D, which can make the filter behave like a linear filter; when the difference is large, the function f(d,S,D) is equal to 0, which can effectively ignore the filter tap.

[0231] In some examples, in addition to the deblocking operation, an in-loop recovery scheme is used in the video encoding after deblocking to generally denoise and enhance the quality of edges. In one example, the in-loop recovery scheme can be switched within the frame of each appropriately sized tile. The in-loop recovery scheme is based on a separable symmetric Wiener filter, a dual self-guided filter with subspace projection, and a domain transform recursive filter. Because the content statistics may vary greatly within a frame, the in-loop recovery scheme is integrated in a switchable framework, where different schemes can be triggered in different regions of the frame.

[0232] According to one aspect of the present disclosure, in-loop recovery (LR) (also referred to as an LR filter) can use neighboring pixels during LR filtering, for example, a pixel window around the pixel to be filtered. In order to be able to apply the LR filter at the edge of the frame, in some examples, some boundary pixel values ​​of pixels at the frame boundary are copied to a boundary buffer before applying the LR filter. In some examples, two copies of the boundary pixels are stored in the boundary buffer. In one example, a first copy step is applied after the deblocking filter and before CDEF, and the copying of the boundary pixel values ​​of the pixels at the frame boundary by the first copy step is represented as COPY0; a second copy step is applied after CDEF and before the LR filter, and the copying of the boundary pixel values ​​of the pixels at the frame boundary by the second copy step is represented as COPY1. The boundary buffer is then used to fill the boundary pixels during the LR filtering process, for example, when the LR filter is applied at the edge of the frame. The process of copying the boundary pixel values ​​to the boundary buffer is called boundary processing.

[0233] A separable symmetric Wiener filter can be an in-loop restoration scheme. In some examples, each pixel in the degraded frame can be reconstructed as a non-causal filtered version of the pixels in a w×w window around it, where w=2r+1 for an integer r that is odd. If the 2D filter taps are represented by w in column vectorized form 2 × 1 element vector F, then direct LMMSE optimization results in the filter parameters being F = H -1 M is given, where H = E[XX T ] is the autocovariance of x, in the w×w window around the pixel w 2 Column vectorized version of the sample, M = E[YX T ] is the cross-correlation of x to be estimated with the scalar source sample y. In one example, the encoder can estimate H and M based on the implementation in the deblocked frame and the source, and the resulting filter F can be sent to the decoder. However, this will not only 2 The tapping results in a considerable bit rate cost, and the non-separable filtering makes decoding extremely complicated. In some embodiments, several additional constraints are imposed on the properties of F. For the first constraint, F is constrained to be separable so that the filtering can be implemented as separable horizontal and vertical w-tap convolutions. For the second constraint, each horizontal and vertical filter is constrained to be symmetric. For the third constraint, it is assumed that the sum of the horizontal and vertical filter coefficients is 1.

[0234] The double self-guided filtering of subspace projection can be used as an in-loop recovery scheme. Guided filtering is an image filtering technique in which the local linear model shown by equation (15) is:

[0235] y=F x+G equation (15)

[0236] For calculating the filtered output y from the unfiltered sample x, where F and G are determined based on the statistics of the degraded image and the guide image near the filtered pixel. If the guide image is the same as the degraded image, the resulting so-called self-guided filtering has the effect of edge-preserving smoothing. In one example, a specific form of self-guided filtering can be used. The specific form of the self-guided filtering depends on two parameters: the radius r and the noise parameter e, and the steps are listed as follows:

[0237] 1. Get the mean μ and variance σ of the pixels in the (2r+1)×(2r+1) window around each pixel 2 This step can be efficiently implemented using box filtering based on integral imaging.

[0238] 2. Calculate for each pixel: f = σ 2 / (σ 2 +e); g = (1-f) μ

[0239] 3. Calculate F and G for each pixel as the average of the f and g values ​​in a 3×3 window around that pixel.

[0240] The specific form of the self-guided filter is controlled by r and e, where higher r means higher spatial variance and higher e means higher range variance.

[0241] Fig. 20 An example of subspace projection is shown to illustrate some examples. Fig. 20 As shown, even if the restored X1 and X2 are not close to the source Y, as long as they are slightly moved to the right, the appropriate multipliers {α, β} can make them closer to the source Y.

[0242] In some examples, a technique called frame super-resolution (FSR) is used to improve the perceived quality of the decoded image. Typically, the FSR process is applied at a low bit rate and includes four steps. In the first step, on the encoder side, the source video is reduced as a non-standard process. In the second step, the reduced video is encoded, followed by a filtering process of a deblocking filter and CDEF. In the third step, a linear amplification process is applied as a standard process to restore the encoded video to its original spatial resolution. In the fourth step, a loop recovery filter is applied to resolve some high-frequency losses. In one example, the last two steps together can be referred to as a super-resolution process. Similarly, on the decoder side, the processes of decoding, deblocking filters, and CDEF can be applied at a lower spatial resolution. Then, the frame undergoes a super-resolution process. In some examples, in order to reduce the overhead in the line buffer implemented with respect to the hardware, the amplification and reduction processes are applied only to the horizontal dimension.

[0243] In some examples (e.g., HEVC), a filtering technique called sample adaptive offset (SAO) can be used. In some examples, SAO is applied to the reconstructed signal after the deblocking filter. SAO can use an offset value given in the slice header. In some examples, for luma samples, the encoder can decide whether to apply (enable) SAO to the slice. When SAO is enabled, the current picture allows the coding unit to be recursively split into four sub-regions, and each sub-region can select an SAO type from multiple SAO types based on the features in the sub-region.

[0244] Fig.21 A table (2100) of various SAO types according to an embodiment of the present disclosure is shown. In the table (2100), SAO types 0-6 are shown. Note that SAO type 0 is used to indicate that no SAO is applied. In addition, each SAO type in SAO type 1 to SAO type 6 includes multiple categories. SAO can classify reconstructed pixels of a sub-region and reduce distortion by adding offsets to pixels of each category in the sub-region. In some examples, edge attributes can be used for pixel classification in SAO types 1-4, and pixel intensity can be used for pixel classification in SAO types 5-6.

[0245] Specifically, in one embodiment, for example, SAO type 5-6, a band offset (BO) can be used to classify all pixels of a sub-region into multiple bands. Each band in the multiple bands includes pixels in the same intensity interval. In some examples, the intensity range is divided into multiple intervals, for example, 32 intervals from zero to the maximum intensity value (for example, 255 for 8-bit pixels), and each interval is associated with an offset. In addition, in one example, the 32 bands are divided into two groups, for example, a first group and a second group. The first group includes the central 16 bands (for example, 16 intervals in the middle of the intensity range), and the second group includes the remaining 16 bands (for example, 8 intervals at the low end of the intensity range and 8 intervals at the high end of the intensity range). In one example, only the offset of one of the two groups is transmitted. In some embodiments, when using the pixel classification operation in BO, the five most significant bits of each pixel can be directly used as a band index.

[0246] Furthermore, in one embodiment, for example, SAO types 1-4, edge offset (EO) may be used for pixel classification and offset determination. For example, pixel classification may be determined based on a 1-dimensional 3-pixel pattern that takes into account edge direction information.

[0247] Fig. 22 An example of a 3-pixel pattern used for pixel classification in edge offset in some examples is shown. Fig. 22In the example, the first mode (2210) (as shown by the three gray pixels) is called the 0° mode (the horizontal direction is associated with the 0° mode), the second mode (2220) (as shown by the three gray pixels) is called the 90° mode (the vertical direction is associated with the 90° mode), the third mode (2230) (as shown by the three gray pixels) is called the 135° mode (the 135° diagonal direction is associated with the 135° mode), and the fourth mode (2240) (as shown by the three gray pixels) is called the 45° mode (the 45° diagonal direction is associated with the 45° mode). In one example, the edge direction information of the sub-region may be considered to select Fig. 22 One of the four directional modes shown. In one example, this selection can be sent in the encoded video bitstream as side information. The pixels in the sub-region can then be classified into multiple categories by comparing each pixel with its two neighboring pixels in the direction associated with the directional map.

[0248] Fig.23 A table (2300) of pixel classification rules for edge offset in some examples is shown. Specifically, pixel c (also in Fig. 22 ) with two adjacent pixels (also shown in Fig. 22 The comparison is shown in gray in each mode of Fig.23 The pixel classification rule shown can classify pixel c into one of categories 0-4 based on the comparison.

[0249] In some embodiments, SAO on the decoder side can operate independently of the largest coding unit (LCU) (e.g., CTU), so that line buffers can be saved. In some examples, when the 90°, 135°, and 45° classification modes are selected, the pixels in the top and bottom rows of each LCU are not SAO processed; when the 0°, 135°, and 45° modes are selected, the pixels in the leftmost and rightmost columns of each LCU are not SAO processed.

[0250] Fig.24An example (2400) of syntax (syntax(es)) that may need to be signaled for a CTU if the parameters are not merged from adjacent CTUs is shown. For example, the syntax element sao_type_idx[cIdx][rx][ry] may be signaled to indicate the SAO type of the sub-region. The SAO type may be BO (bandoffset) or EO (edge ​​offset). When the value of sao_type_idx[cIdx][rx][ry] is 0, it indicates that SAO is off; values ​​1 to 4 indicate the use of one of the four EO categories corresponding to 0°, 90°, 135°, and 45°; a value of 5 indicates that BO is used. Fig.24 In the example shown in FIG. 4 , each of the BO and EO types has four SAO offset values ​​(sao_offset[cIdx][rx][ry][0] to sao_offset[cIdx][rx][ry][3]) signaled.

[0251] Typically, a filtering process may use reconstructed samples of a first color component as input (e.g., Y or Cb or Cr, or R or G or B) to generate an output, and the output of the filtering process is applied to a second color component, which may be the same as the first color component, or may be another color component different from the first color component.

[0252] In a related example of cross component filtering (CCF), filter coefficients are derived based on some mathematical equations. The derived filter coefficients are sent from the encoder side to the decoder side by signaling, and the derived filter coefficients are used to generate offsets using linear combinations. Then, as a filtering process, the generated offsets are added to the reconstructed samples. For example, offsets are generated based on a linear combination of filter coefficients and luminance samples, and the generated offsets are added to the reconstructed chrominance samples. The related example of CCF is based on the assumption of a linear mapping relationship between reconstructed luminance sample values ​​and the incremental values ​​between the original and reconstructed chrominance samples. However, the mapping between the reconstructed luminance sample values ​​and the incremental values ​​between the original and reconstructed chrominance samples does not necessarily follow a linear mapping process, and therefore, under the assumption of a linear mapping relationship, the coding performance of CCF may be limited.

[0253] In some examples, nonlinear mapping techniques can be used for cross component filtering and / or same color component filtering without significant signaling overhead. In one example, nonlinear mapping techniques can be used in cross component filtering to generate cross component sample offsets. In another example, nonlinear mapping techniques can be used in same color component filtering to generate local sample offsets.

[0254] For convenience, the filtering process using nonlinear mapping technology can be referred to as sample offset of nonlinear mapping (SO-NLM). SO-NLM in the cross-component filtering process can be referred to as cross-component sample offset (CCSO). SO-NLM in the same color component filtering can be referred to as local sample offset (LSO). The filter using nonlinear mapping technology can be referred to as a filter based on nonlinear mapping. Filters based on nonlinear mapping can include CCSO filters, LSO filters, etc.

[0255] In one example, CCSO and LSO can be used as loop filtering to reduce the distortion of reconstructed samples. CCSO and LSO do not rely on the linear mapping assumption used in the related example CCF. For example, CCSO does not rely on the assumption of a linear mapping relationship between the luma reconstructed sample value and the delta value between the original chroma sample and the chroma reconstructed sample. Similarly, LSO does not rely on the assumption of a linear mapping relationship between the reconstructed sample value of the color component and the delta value between the original sample of the color component and the reconstructed sample of the color component.

[0256] In the following description, a SO-NLM filtering process is described that uses a reconstructed sample of a first color component as input (e.g., Y or Cb or Cr, or R or G or B) to generate an output, and the output of the filtering process is applied to a second color component. When the second color component is the same color component as the first color component, the description applies to LSO; and when the second color component is different from the first color component, the description applies to CCSO.

[0257] In SO-NLM, a nonlinear mapping is derived at the encoder side. The nonlinear mapping is between the reconstructed samples of the first color component in the filter support region and the offset of the second color component to be added to the filter support region. Nonlinear mapping is used in LSO when the second color component is the same as the first color component and in CCSO when the second color component is different from the first color component. The domain of the nonlinear mapping is determined by the different combinations of processed input reconstructed samples (also called combinations of possible reconstructed sample values).

[0258] The technique of SO-NLM can be illustrated with a specific example. In the specific example, reconstructed samples from a first color component located in a filter support region (also referred to as a "filter support region") are determined. The filter support region is the region in which the filter can be applied, and the filter support region can have any suitable shape.

[0259] Fig.25An example of a filter support area (2500) according to some embodiments of the present disclosure is shown. The filter support area (2500) includes four reconstructed samples of the first color component: P0, P1, P2, and P3. Fig.25 In the example of , the four reconstructed samples may form a cross in the vertical and horizontal directions, and the center position of the cross is the position of the sample to be filtered. The sample of the same color component as P0-P3 at the center position is represented by C. The sample of the second color component at the center position is represented by F. The second color component may be the same as the first color component of P0-P3, or may be different from the first color component of P0-P3.

[0260] Fig.26 An example of another filter support region (2600) according to some embodiments of the present disclosure is shown. The filter support region (2600) includes four reconstructed samples P0, P1, P2 and P3 of the first color component, and the four reconstructed samples form a square. Fig.26 In the example of , the center position of the square is the position of the sample to be filtered. The sample of the same color component as P0-P3 at the center position is represented by C. The sample of the second color component at the center position is represented by F. The second color component can be the same as the first color component of P0-P3, or can be different from the first color component of P0-P3.

[0261] The reconstructed samples are input to the SO-NLM filter and are appropriately processed to form filter taps. In one example, the positions of the reconstructed samples as input to the SO-NLM filter are referred to as filter tap positions. In a specific example, the reconstructed samples are processed in the following two steps.

[0262] In the first step, the incremental values ​​between P0-P3 and C are calculated respectively. For example, m0 represents the incremental value between P0 and C; m1 represents the incremental value between P1 and C; m2 represents the incremental value between P2 and C; and m3 represents the incremental value between P3 and C.

[0263] In the second step, the incremental values ​​m0-m3 are further quantized, and the quantized values ​​are represented as d0, d1, d2, d3. In one example, based on the quantization process, the quantized value can be one of -1, 0, and 1. For example, when m is less than -N (N is a positive value, called a quantization step), the value m can be quantized to -1; when m is in the range of [-N, N], the value m can be quantized to 0; and when m is greater than N, the value m can be quantized to 1. In some examples, the quantization step N can be one of 4, 8, 12, 16, etc.

[0264] In some embodiments, the quantized values ​​d0-d3 are filter taps and can be used to identify a combination in the filter domain. For example, the filter taps d0-d3 can form a combination in the filter domain. Each filter tap can have three quantized values, so when four filter taps are used, the filter domain includes 81 (3×3×3×3) combinations.

[0265] Figures 27A-27C A table (2700) with 81 combinations according to an embodiment of the present disclosure is shown. The table (2700) includes 81 rows corresponding to the 81 combinations. In each row corresponding to a combination, the first column includes the index of the combination; the second column includes the value of the filter tap d0 of the combination; the third column includes the value of the filter tap d1 of the combination; the fourth column includes the value of the filter tap d2 of the combination; the fifth column includes the value of the filter tap d3 of the combination; and the sixth column includes an offset value associated with the combination of nonlinear mappings. In one example, when the filter taps d0-d3 are determined, the offset value (represented by s) associated with the combination of d0-d3 can be determined according to the table (2700). In one example, the offset values ​​s0-s80 are integers, for example, 0, 1, -1, 3, -3, 5, -5, -7, etc.

[0266] In some embodiments, the final filtering process of SO-NLM may be applied, as shown in equation (16):

[0267] f′=clip(f+s) Equation (16)

[0268] Wherein, f is the reconstructed sample of the second color component to be filtered, s is the offset value determined according to the filter tap, and the filter tap is the processing result of the reconstructed sample of the first color component, for example, using table (2700). The sum of the reconstructed sample F and the offset value s is further clipped to a range associated with the bit depth to determine the final filtered sample f' of the second color component.

[0269] Note that, in the case of LSO, the second color component in the above description is the same as the first color component; and, in the case of CCSO, the second color component in the above description may be different from the first color component.

[0270] Note that the above description may be adjusted for other embodiments of the present disclosure.

[0271] In some examples, at the encoder side, the encoding device may derive a mapping between a reconstructed sample of a first color component in a filter support region and an offset to be added to the reconstructed sample of a second color component. The mapping may be any suitable linear or nonlinear mapping. The filtering process may then be applied on the encoder side and / or decoder side based on the mapping. For example, the mapping is appropriately notified to the decoder (e.g., the mapping is included in a coded video bitstream transmitted from the encoder side to the decoder side), and the decoder may then perform filtering based on the mapping.

[0272] The performance of filters based on nonlinear mapping, such as CCSO filters, LSO filters, etc., depends on the filter shape configuration. The filter shape configuration (also referred to as filter shape) of a filter may refer to the properties of the pattern formed by the filter tap positions. The pattern may be defined by various parameters, such as the number of filter taps, the geometry of the filter tap positions, the distance of the filter tap positions to the center of the pattern, etc. Using a fixed filter shape configuration may limit the performance of filters based on nonlinear mapping.

[0273] like Fig.24 and Fig.25 as well as Figures 27A-27C As shown, some examples use a 5-tap filter design for a filter shape configuration of a filter based on a nonlinear mapping. The 5-tap filter design may use tap positions at P0, P1, P2, P3, and C. The 5-tap filter design for the filter shape configuration may produce a lookup table (LUT) with 81 entries, such as Figures 27A-27C As shown. The sample offset LUT needs to be signaled from the encoder side to the decoder side, and the signaling of the LUT contributes most of the signaling overhead and affects the coding efficiency using filters based on nonlinear mapping. According to some aspects of the present disclosure, the number of filter taps may be different from 5. In some examples, the number of filter taps may be reduced, information in the filter support region may still be captured, and coding efficiency may be improved.

[0274] In some examples, the filter shape configurations in the group of nonlinear mapping based filters each have 3 filter taps.

[0275] Fig.28Seven filter shape configurations for three filter taps in one example are shown. Specifically, the first filter shape configuration includes 3 filter taps at positions marked as "1" and "c", and position "c" is the center position of position "1"; the second filter shape configuration includes 3 filter taps at positions marked as "2" and position "c", and position "c" is the center position of position "2"; the third filter shape configuration includes 3 filter taps at positions marked as "3" and position "c", and position "c" is the center position of position "3"; the fourth filter shape configuration includes 3 filter taps at positions marked as "4" and position "c", and position "c" is the center position of position "4"; the fifth filter shape configuration includes 3 filter taps at positions marked as "5" and "c", and position "c" is the center position of position "5"; the sixth filter shape configuration includes 3 filter taps at positions marked as "6" and position "c", and position "c" is the center position of position "6"; the seventh filter shape configuration includes 3 filter taps at positions marked as "7" and position "c", and position "c" is the center position of position "7".

[0276] Various aspects of the present disclosure provide techniques for integrating video processing tools, such as filtering, boundary processing, and cropping tools. In some examples, a filter based on nonlinear mapping (e.g., CCSO, LSO) or other tools (e.g., a cropping module) is applied after a deblocking filter and before an LR filter, and a boundary processing process that stores two copies of boundary pixels is also applied after the deblocking filter and before the LR filter. The present disclosure provides various configurations for including a filter based on nonlinear mapping, a boundary processing module, and / or a cropping module after a deblocking filter and before an LR filter.

[0277] Fig.29 A block diagram of a loop filter chain (loop filter chain) (2900) in some examples is shown. The loop filter chain (2900) includes a plurality of filters connected in series in the filter chain. The loop filter chain (2900) can be used as a loop filtering unit, for example, the loop filter unit (loop filter unit) (556) in the example. The loop filter chain (2900) can be used in an encoding or decoding loop before storing the reconstructed image in a decoded image buffer (e.g., a reference image memory (557)). The loop filter chain (2900) receives input reconstructed samples from a previous processing module and applies filters to the reconstructed samples to generate output reconstructed samples.

[0278] The loop filter chain (2900) may include any suitable filters. Fig.29In the example of , the loop filter chain (2900) includes a deblocking filter (labeled as deblocking), a constrained directional enhancement filter (labeled as CDEF) and an in-loop restoration filter (labeled as LR) connected in a chain. The loop filter chain (2900) has an input node (2901), an output node (2909) and a plurality of intermediate nodes (2902)-(2903). The input node (2901) of the loop filter chain (2900) receives input reconstructed samples from a prior processing module, and the input reconstructed samples are provided to the deblocking filter. The intermediate node (2902) receives the reconstructed samples (after being processed by the deblocking filter) from the deblocking filter and provides the reconstructed samples to CDEF for further filtering processing. The intermediate node (2903) receives the reconstructed samples (after being processed by CDEF) from the CDEF and provides the reconstructed samples to the LR filter for further filtering processing. The output node (2909) receives the output reconstructed samples from the LR filter (after being processed by the LR filter). The output reconstructed samples may be provided to other processing modules, such as a post processing module, for further processing.

[0279] Note that the following description illustrates the technique of using filters based on non-linear mapping in the loop filter chain (2900). The technique of using filters based on non-linear mapping can be used in other suitable loop filter chains.

[0280] According to some aspects of the present invention, when a filter based on nonlinear mapping is used in a loop filter chain, at least one of the copies (COPY0 and COPY1) of the boundary pixel values ​​used for boundary processing in the LR filter is associated with the filter based on nonlinear mapping. In some examples, the reconstructed samples from which the boundary pixel values ​​are obtained and buffered can be the input of the filter based on nonlinear mapping. In some examples, the reconstructed samples from which the boundary pixel values ​​are obtained and buffered can be the result of applying the filter based on nonlinear mapping. In some examples, the reconstructed samples from which the boundary pixel values ​​are obtained and buffered can be combined with the sample offset generated by the filter based on nonlinear mapping.

[0281] In some examples, pixels after the deblocking filter and before CDEF or a filter based on non-linear mapping are used for COPY0, and pixels after applying CDEF or a filter based on non-linear mapping and before the LR filter are used for COPY1.

[0282] Fig.30An example of a loop filter chain (3000) is shown, which includes a nonlinear mapping based filter and a CDEF between the input and output of the nonlinear mapping based filter. The loop filter chain (3000) can be used in an encoding device or a decoding device instead of the loop filter chain (2900). In the loop filter chain (3000), the deblocking filter generates a first intermediate reconstructed sample at a first intermediate node (3011), and the first intermediate reconstructed sample is input to the CDEF and a nonlinear mapping based filter (non linear mapping based filter, marked as SO-NLM). The CDEF is applied to the first intermediate reconstructed sample and a second intermediate reconstructed sample (3012) at a second intermediate node is generated. The nonlinear mapping based filter generates a sample offset SO (sample offset(s)) based on the first intermediate reconstructed sample. The sample offset SO is combined with the second intermediate reconstructed sample to generate a third intermediate reconstructed sample (3013) at a third intermediate node. Then, the LR filter is applied to the third intermediate reconstructed sample to generate the output of the loop filter chain (3000). In addition, the first intermediate reconstructed sample at the first intermediate node (3011) is used to obtain a first copy COPY0 of the boundary pixel for boundary processing in the LR filter, and the third intermediate reconstructed sample at the third intermediate node (3013) is used to obtain a second copy COPY1 of the boundary pixel for boundary processing in the LR filter.

[0283] In some examples, pixels after the deblocking filter and before applying CDEF or a filter based on non-linear mapping are used for COPY0, and pixels after applying CDEF and before applying a filter based on non-linear mapping are used for COPY1.

[0284] Fig.31An example of a loop filter chain (3100) is shown, which includes a filter based on nonlinear mapping and a CDEF between the input and output of the filter based on nonlinear mapping. The loop filter chain (3100) can be used in an encoding device or a decoding device instead of the loop filter chain (2900). In the loop filter chain (3100), the deblocking filter generates a first intermediate reconstructed sample at a first intermediate node (3111), and the first intermediate reconstructed sample is input to the CDEF and the filter based on nonlinear mapping (labeled as SO-NLM). The CDEF is applied to the first intermediate reconstructed sample and a second intermediate reconstructed sample (3112) at a second intermediate node is generated. The filter based on nonlinear mapping generates a sample offset SO based on the first intermediate reconstructed sample. The sample offset SO is combined with the second intermediate reconstructed sample to generate a third intermediate reconstructed sample (3113) at an intermediate third node. Then, the LR filter is applied to the third intermediate reconstructed sample to generate the output of the loop filter chain (3100). In addition, the first intermediate reconstructed sample at the first intermediate node (3111) is used to obtain a first copy COPY0 of the boundary pixel for boundary processing in the LR filter, and the second intermediate reconstructed sample at the second intermediate node (3012) is used to obtain a second copy COPY1 of the boundary pixel for boundary processing in the LR filter.

[0285] In some examples, the nonlinear mapping based filter is connected in series with other filters in the loop filter chain, and there are no other filters between the input and output of the nonlinear mapping based filter. In one example, the nonlinear mapping based filter is applied after CDEF. The pixels after the deblocking filter and before CDEF are used for COPY0, and the pixels after the nonlinear mapping based filter and before the LR filter are used for COPY1.

[0286] Fig.32An example of a loop filter chain (3200) in an example is shown. The loop filter chain (3200) can be used in an encoding device or a decoding device instead of the loop filter chain (2900). In the loop filter chain (3200), CDEF is between a first intermediate node (3211) and a second intermediate node (3212), and a filter based on nonlinear mapping (labeled as SO-NLM) is applied to the second intermediate node (3212) of the loop filter chain (3200). Specifically, the deblocking filter generates a first intermediate reconstructed sample (3211) at the first intermediate node, CDEF is applied to the first intermediate reconstructed sample, and a second intermediate reconstructed sample (3212) at the second intermediate node is generated. The second intermediate reconstructed sample is an input to the filter based on nonlinear mapping. Based on the input, the filter based on nonlinear mapping generates a sample offset (SO). The sample offset is combined with the second intermediate reconstructed sample at the second intermediate node (3212) to generate a third intermediate reconstructed sample (3213) at the third intermediate node. The LR filter is applied to the third intermediate reconstructed sample to generate the output of the loop filter chain (3200). In addition, the first intermediate reconstructed sample at the first intermediate node (3211) is used to obtain a first copy COPY0 of the boundary pixel for boundary processing in the LR filter, and the third intermediate reconstructed sample at the third node (3213) is used to obtain a second copy COPY1 of the boundary pixel for boundary processing in the LR filter.

[0287] In one example, a non-linear mapping based filter is applied before CDEF. Pixels after the deblocking filter and before the non-linear mapping based filter are used for COPY0, and pixels after CDEF and before the LR filter are used for COPY1.

[0288] Fig.33An example of a loop filter chain (3300) in an example is shown. The loop filter chain (3300) can be used in an encoding device or a decoding device instead of the loop filter chain (2900). In the loop filter chain (3300), a filter based on nonlinear mapping (labeled as SO-NLM) is applied to a first intermediate node (3311) of the loop filter chain (3300), and CDEF is between a second intermediate node (3312) and a third intermediate node (3313). Specifically, the deblocking filter generates a first intermediate reconstructed sample (3311) at the first intermediate node. The first intermediate reconstructed sample is an input to the filter based on nonlinear mapping. Based on the input, the filter based on nonlinear mapping generates a sample offset (SO). The sample offset is combined with the first intermediate reconstructed sample (3311) at the first intermediate node to generate a second intermediate reconstructed sample (3312) at the second intermediate node. CDEF is applied to the second intermediate reconstructed sample and generates a third intermediate reconstructed sample (3313) at the third intermediate node. The LR filter is applied to the third intermediate reconstructed sample to generate the output of the loop filter chain (3300). In addition, the first intermediate reconstructed sample at the first intermediate node (3311) is used to obtain a first copy COPY0 of the boundary pixel for boundary processing in the LR filter, and the third intermediate reconstructed sample at the third intermediate node (3313) is used to obtain a second copy COPY1 of the boundary pixel for boundary processing in the LR filter.

[0289] According to aspects of the present invention, the enabling / disabling of boundary processing of the LR filter (e.g., obtaining COPY0 and COPY1 in the LR filter and using COPY0 and COPY1) depends on the enabling / disabling of other filtering tools, such as nonlinear mapping-based filters, CDEF and / or FSR in the loop filtering chain.

[0290] In some examples, the enabling / disabling of boundary processing depends on the enabling / disabling of the filter based on nonlinear mapping, and / or CDEF, and / or FSR. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing; the flag En-SO-NLM is used to indicate the enabling (e.g., the flag En-SO-NLM has a binary 1) or disabling (e.g., the flag En-SO-NLM has a binary 0) of applying the filter based on nonlinear mapping; the flag En-CDEF is used to indicate the enabling (e.g., the flag En-CDEF has a binary 1) or disabling (e.g., the flag En-CDEF has a binary 0) of applying CDEF; the flag En-FSR is used to indicate the enabling (e.g., the flag En-FSR has a binary 1) or disabling (e.g., the flag En-FSR has a binary 0) of applying FSR. Then, in one example, the flag En-Boundary is a logical combination of the flag En-SO-NLM, the flag En-CDEF, and the flag En-FSR. Note that any suitable logical operator may be used, eg, AND, OR, NOT, etc.

[0291] In some examples, the enabling / disabling of boundary processing depends on the enabling / disabling of the filter based on nonlinear mapping and / or CDEF. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing; the flag En-SO-NLM is used to indicate the enabling (e.g., the flag En-SO-NLM has a binary 1) or disabling (e.g., the flag En-SO-NLM has a binary 0) of applying the filter based on nonlinear mapping; the flag En-CDEF is used to indicate the enabling (e.g., the flag En-CDEF has a binary 1) or disabling (e.g., the flag En-CDEF has a binary 0) of applying CDEF. Then, in one example, the flag En-Boundary is a logical combination of the flag En-SO-NLM and the flag En-CDEF. Note that any suitable logical operator can be used, such as AND, OR, NOT, etc.

[0292] In some examples, the enabling / disabling of boundary processing depends on the enabling / disabling of the filter and / or FSR based on nonlinear mapping. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing; the flag En-SO-NLM is used to indicate the enabling (e.g., the flag En-SO-NLM has a binary 1) or disabling (e.g., the flag En-SO-NLM has a binary 0) of applying the filter based on nonlinear mapping; the flag En-FSR is used to indicate the enabling (e.g., the flag En-FSR has a binary 1) or disabling (e.g., the flag En-FSR has a binary 0) of applying the FSR. Then, in one example, the flag En-Boundary is a logical combination of the flag En-SO-NLM and the flag En-FSR. Note that any suitable logical operator can be used, such as AND, OR, NOT, etc.

[0293] In some examples, the enabling / disabling of the boundary processing depends on the enabling / disabling of the filter based on the nonlinear mapping. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of the boundary processing; the flag En-SO-NLM is used to indicate the enabling (e.g., the flag En-SO-NLM has a binary 1) or disabling (e.g., the flag En-SO-NLM has a binary 0) of applying the filter based on the nonlinear mapping. Then, in one example, the flag En-Boundary can be the flag En-SO-NLM, or can be the logical NOT of the flag En-SO-NLM.

[0294] In one embodiment, the enabling / disabling of boundary processing depends on the enabling / disabling of CDEF and / or FSR.

[0295] In some examples, the enabling / disabling of boundary processing depends on the enabling / disabling of CDEF and / or FSR. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing; the flag En-CDEF is used to indicate the enabling (e.g., the flag En-CDEF has a binary 1) or disabling (e.g., the flag En-CDEF has a binary 0) of applying CDEF; the flag En-FSR is used to indicate the enabling (e.g., the flag En-FSR has a binary 1) or disabling (e.g., the flag En-FSR has a binary 0) of applying FSR. Then, in one example, the flag En-Boundary is a logical combination of the flag En-CDEF and the flag En-FSR. Note that any suitable logical operator can be used, such as AND, OR, NOT, etc.

[0296] In some examples, the enabling / disabling of boundary processing depends on the enabling / disabling of CDEF. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing; the flag En-CDEF is used to indicate the enabling (e.g., the flag En-CDEF has a binary 1) or disabling (e.g., the flag En-CDEF has a binary 0) of applying a filter based on nonlinear mapping. Then, in one example, the flag En-Boundary can be the flag En-CDEF, or can be a logical NOT of the flag En-CDEF.

[0297] In one embodiment, the enabling / disabling of boundary processing depends on the enabling / disabling of FSR.

[0298] In some examples, the enabling / disabling of boundary processing depends on the enabling / disabling of FSR. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing; the flag En-FSR is used to indicate the enabling (e.g., the flag En-FSR has a binary 1) or disabling (e.g., the flag En-FSR has a binary 0) of applying a filter based on nonlinear mapping. Then, in one example, the flag En-Boundary can be the flag En-FSR, or can be a logical NOT of the flag En-FSR.

[0299] According to some aspects of the present invention, the pixel values ​​are clipped before the LR filter. In some examples, the pixels after applying CDEF are first clipped, which is represented as CLIP0. Then, the pixel clipping after applying the filter based on nonlinear mapping is represented as CLIP1. Note that in some examples, the pixel values ​​are clipped to a suitable range that is meaningful to the pixel values. In one example, when the pixel values ​​are represented by 8 bits, the pixel values ​​can be clipped to the range [0, 255].

[0300] Note that in this specification, filters based on nonlinear mapping are used as examples to illustrate techniques for buffering boundary pixel values ​​at different nodes along the loop filter chain and / or clipping pixel values ​​at different nodes along the loop filter chain. These techniques can be used when other coding tools (e.g., cross-component sample adaptive offset (CCSAO) tools, etc.) are applied after the deblocking filter and before the LR filter.

[0301] Fig.34 An example of a loop filter chain (3400) including a filter based on nonlinear mapping is shown, with CDEF between the input and output of the filter based on nonlinear mapping. The loop filter chain (3400) can be used in an encoding device or a decoding device instead of the loop filter chain (2900). In the loop filter chain (3400), a filter based on nonlinear mapping (labeled as SO-NLM) is applied between two intermediate nodes. The deblocking filter generates a first intermediate reconstructed sample (3411) at the first intermediate node. The first intermediate reconstructed sample is the input of CDEF and the filter based on nonlinear mapping. Based on the first intermediate reconstructed sample, CDEF is applied, and a second intermediate reconstructed sample (3412) at the second intermediate node is generated. Also based on the first intermediate reconstructed sample, a sample offset (SO) is generated based on the filter based on nonlinear mapping. The second intermediate reconstructed sample is clipped to generate CLIP0. CLIP0 is combined with the sample offset to generate a third intermediate reconstructed sample (3413) at the third intermediate node, and the third intermediate reconstructed sample is clipped to generate CLIP1. CLIP1 is provided to the LR filter for further filtering.

[0302] Fig.35An example of a loop filter chain (3500) including a filter based on nonlinear mapping is shown, with CDEF between the input and output of the filter based on nonlinear mapping. The loop filter chain (3500) can be used in an encoding device or a decoding device instead of the loop filter chain (2900). In the loop filter chain (3500), a filter based on nonlinear mapping (labeled as SO-NLM) is applied between two intermediate nodes. The deblocking filter generates a first intermediate reconstructed sample (3511) at the first intermediate node. The first intermediate reconstructed sample is the input of CDEF and the filter based on nonlinear mapping. Based on the first intermediate reconstructed sample, CDEF is applied, and a second intermediate reconstructed sample (3512) at the second intermediate node is generated. Also based on the first intermediate reconstructed sample, a sample offset (SO) is generated based on the filter based on nonlinear mapping. The second intermediate reconstructed sample is combined with the sample offset to generate a third intermediate reconstructed sample (3513) at the third intermediate node, and the third intermediate reconstructed sample is clipped to generate CLIP0. CLIP0 is provided to the LR filter for further filtering.

[0303] Fig.36 A flowchart outlining a process (3600) according to an embodiment of the present disclosure is shown. The process (3600) can be used for video filtering. When the term block is used, the block can be interpreted as a prediction block, a coding unit, a luminance block, a chrominance block, etc. In various embodiments, the process (3600) is performed by a processing circuit, for example, a processing circuit in a terminal device (310), (320), (330) and (340), a processing circuit that performs the functions of a video encoder (403), a processing circuit that performs the functions of a video decoder (410), a processing circuit that performs the functions of a video decoder (510), a processing circuit that performs the functions of a video encoder (603), etc. In some embodiments, the process (3600) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (3600). The process starts at (S3601) and proceeds to (S3610).

[0304] At (S3610), a first boundary pixel value of a subset of first reconstructed samples at a first node along the loop filter chain is buffered. The first node is associated with a nonlinear mapping-based filter applied before the loop recovery filter in the loop filter chain, wherein the subset of the first reconstructed samples may be, but is not limited to, all boundary pixel values ​​or partial pixel boundary values ​​used to indicate the first reconstructed samples, which is not limited in the present embodiment.

[0305] In some examples, the filter based on the non-linear mapping is a cross-component sample offset (CCSO) filter. In some examples, the filter based on the non-linear mapping is a local sample offset (LSO) filter.

[0306] At (S3620), a loop restoration filter is applied to the reconstructed samples to be filtered based on the buffered first boundary pixel value.

[0307] In some examples, the first reconstructed sample at the first node is an input to a filter based on a non-linear mapping.

[0308] In some examples, the first filter and the second filter are applied to first boundary pixel values ​​of a subset of the first reconstructed samples.

[0309] It should be noted that the first filter is a filter in the loop filter chain, which is used to reduce distortion and artifacts in the reconstructed image while maintaining or enhancing the details and edges of the image, and may be, but not limited to, a filter using nonlinear mapping. The second filter is a filter in the loop filter chain, which is used to reduce artifacts generated during the encoding process while maintaining the details and clarity of the image, and may be, but not limited to, a constrained directional enhancement filter.

[0310] In some examples, the second reconstructed sample at the second node is generated after applying at least one of the first filter or the second filter to the first reconstructed sample. Wherein, the application process may include: applying at least one of the first filter or the second filter to the first boundary pixel value. That is, the second reconstructed sample at the second node is generated after applying at least one of the first filter or the second filter to the first reconstructed sample, including: the second reconstructed sample at the second node is generated after applying at least one of the first filter or the second filter to the first boundary pixel value. In other words, the second reconstructed sample at the second node may be generated after directly applying at least one of the first filter or the second filter to the first reconstructed sample, or may be generated after directly applying at least one of the first filter or the second filter to the first boundary pixel value.

[0311] In some examples, the processor combines the sample offset generated by the first filter with the output of the second filter to generate the second reconstructed sample. Specifically, the first reconstructed sample is input into the first filter to generate the sample offset; the sample offset is combined with the output of the second filter to generate the second reconstructed sample.

[0312] In some examples, the processor combines the sample offset generated by the first filter with the first reconstructed sample to generate an intermediate reconstructed sample; and the processor applies the second filter to the intermediate reconstructed sample to generate the second reconstructed sample. Specifically, the first reconstructed sample is input into the first filter to generate a sample offset; the sample offset and the first reconstructed sample are combined to generate an intermediate reconstructed sample; the second filter is applied to the intermediate reconstructed sample to generate the second reconstructed sample.

[0313] exist Fig.30 , Fig.31 and Fig.33 In the example of Fig.30 , Fig.31 and Fig.33 The first intermediate node (3011) / (3111) / (3311) in the corresponding description of , the first reconstructed sample can be Fig.30 , Fig.31 and Fig.33 The first intermediate reconstructed sample in the corresponding description of .

[0314] Then, in Fig.30 and Fig.33 In the example, the second boundary pixel values ​​of the subset of the second reconstructed samples at the second node along the loop filter chain can be buffered, and the second reconstructed samples at the second node are generated after applying the sample offset generated by the filter based on the nonlinear mapping. Based on the buffered first boundary pixel values ​​and the buffered second boundary pixel values, the loop recovery filter is applied to the reconstructed samples to be filtered, wherein the subset of the second reconstructed samples can be, but is not limited to, all boundary pixel values ​​or partial pixel boundary values ​​used to indicate the second reconstructed samples, which is not limited in the present embodiment.

[0315] exist Fig.30 In the example of , the sample offsets generated by the filter based on the nonlinear mapping are consistent with the output of the constrained directional enhancement filter (e.g., Fig.30 ) to generate a second reconstructed sample (eg, Fig.30 The third intermediate reconstructed sample in the description of ).

[0316] exist Fig.33 In the example of , the sample offset generated by the filter based on the nonlinear mapping is different from the first reconstructed sample (eg, Fig.33 ) to generate an intermediate reconstructed sample (eg, Fig.33 ), and then applying the constrained directional enhancement filter to the intermediate reconstructed samples to generate second reconstructed samples (eg, Fig.33 The third intermediate reconstructed sample in the description of ).

[0317] exist Fig.31 In the example of , a second reconstructed sample at a second node along the loop filter chain is buffered (eg, Fig.33The second intermediate pixel value of the second intermediate node (3112) in the description of the present invention is a second boundary pixel value of the second intermediate reconstructed sample at the second intermediate node (3112). The second reconstructed sample at the second node is combined with the sample offset generated by the filter based on the nonlinear mapping to generate the reconstructed sample to be filtered. Based on the buffered first boundary pixel value and the buffered second boundary pixel value, a loop recovery filter is applied to the reconstructed sample to be filtered.

[0318] exist Fig.32 In the example of Fig.32 The third intermediate reconstructed sample at the third intermediate node (3213) in the description of is the result of applying the filter based on the nonlinear mapping. Then, the second reconstructed sample generated by the deblocking filter (e.g., Fig.32 The constrained directional enhancement filter is applied to the second reconstructed sample to generate an intermediate reconstructed sample (eg, Fig.32 The intermediate reconstructed samples are combined with sample offsets generated by a filter based on a nonlinear mapping to generate first reconstructed samples (eg, Fig.32 Then, a loop restoration filter may be applied based on the buffered first boundary pixel value and the buffered second boundary pixel value.

[0319] In some examples, before applying the loop restoration filter, the reconstructed samples to be filtered are clipped to a suitable range, such as [0, 255] in the 8-bit case, e.g., Fig.34 and Fig.35 In some examples (e.g., in Fig.34 ), before being combined with the sample offsets generated by the filter based on the nonlinear mapping, the intermediate reconstructed samples (e.g. Fig.34 The second intermediate reconstructed sample in the description of ) is clipped to a suitable range, for example, [0, 255] in the 8-bit case.

[0320] Process (3600) proceeds to (S3699) and terminates.

[0321] Note that in some examples, the filter based on the non-linear mapping is a cross-component sample offset (CCSO) filter, and in some other examples, the filter based on the non-linear mapping is a local sample offset (LSO) filter.

[0322] The process (3600) may be adjusted appropriately. Steps in the process (3600) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0323] Aspects of the present invention provide a method for video encoding:

[0324] Reconstruct the pixel values ​​in the current image frame to obtain reconstructed samples; apply the deblocking filter to the reconstructed samples to obtain first reconstructed samples; buffer the first boundary pixel value of a subset of the first reconstructed samples at a first node along the loop filter chain, the first boundary pixel value is the value of the pixel at the frame boundary; buffer the second boundary pixel value of a subset of the second reconstructed samples at a second node along the loop filter chain, the second reconstructed sample at the second node is generated after applying at least one of the first filter or the second filter to the first reconstructed sample, and the second boundary pixel value is the value of the pixel at the frame boundary; and apply a loop recovery filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value to obtain the reconstructed sample after filtering; perform encoding processing based on the reconstructed sample after filtering.

[0325] It should be noted that the reconstructed samples after filtering can be used as high-quality reference frames for encoding subsequent image frames, specifically, for generating the final reconstructed frames, which are entropy encoded and then successfully packaged to obtain the final bit stream.

[0326] Aspects of the present invention provide a method for video decoding:

[0327] The received coded bit stream is decoded to obtain a decoded symbol; pixels are reconstructed based on the decoded symbol to obtain a reconstructed sample; the reconstructed sample is applied to a deblocking filter to obtain a first reconstructed sample; a first boundary pixel value of a subset of the first reconstructed sample is buffered at a first node along the loop filter chain, the first boundary pixel value is the value of a pixel at a frame boundary; a second boundary pixel value of a subset of the second reconstructed sample is buffered at a second node along the loop filter chain, the second reconstructed sample at the second node is generated after applying at least one of the first filter or the second filter to the first reconstructed sample, and the second boundary pixel value is the value of a pixel at a frame boundary; and based on the buffered first boundary pixel value and the buffered second boundary pixel value, a loop recovery filter is applied to the reconstructed sample to be filtered to obtain a filtered reconstructed sample; and a decoded image frame is output based on the filtered reconstructed sample.

[0328] An aspect of the present invention provides a method for generating a video bitstream, which generates a coded bitstream of a video according to the above video coding method.

[0329] The embodiments of the present disclosure may be used alone or in any order. In addition, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0330] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.37 A computer system (3700) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0331] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.

[0332] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, IoT devices, etc.

[0333] Fig.37 The components of the computer system (3700) shown in the example are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of computer software implementing the disclosed embodiments. The configuration of components should not be interpreted as having any dependency or requirement on any one component or combination of components shown in the exemplary embodiment of the computer system (3700).

[0334] The computer system (3700) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to a person's conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0335] Input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (3701), mouse (3702), trackpad (3703), touch screen (3710), data gloves (not shown), joystick (3705), microphone (3706), scanner (3707), camera (3708).

[0336] The computer system (3700) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (3710), a data glove (not shown), or a joystick (3705), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (3709), headphones (not shown)), visual output devices (e.g., screens (3710), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through methods such as stereo output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)) and printers (not shown).

[0337] The computer system (3700) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (3720) with CD / DVD or similar media (3721), a thumb drive (3722), a removable hard drive or solid state drive (3723), traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD devices such as security dongles (not shown), and the like.

[0338] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0339] The computer system (3700) may also include an interface (3754) to one or more communication networks (3755). The network may be, for example, wireless, wired, optical. The network may also be local, wide, metropolitan, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter connected to some common data port or peripheral bus (3749) (e.g., a USB port of the computer system (3700)); others are typically integrated into the core of the computer system (3700) by connecting to a system bus as described below (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smart phone computer system). Using any of these networks, the computer system (3700) can communicate with other entities. Such communication may be one-way, receive-only (e.g., broadcast television), one-way, send-only (e.g., CANbus to certain CANbus devices), or two-way, for example, to other computer systems using local or wide area digital networks. As described above, certain protocols and protocol stacks may be used on each of these networks and network interfaces.

[0340] The aforementioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core (3740) of the computer system (3700).

[0341] The core (3740) may include one or more central processing units (CPUs) (3741), graphics processing units (GPUs) (3742), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (3743), hardware accelerators (3744) for specific tasks, graphics adapters (3750), etc. These devices, along with read-only memory (ROM) (3745), random access memory (3746), internal mass storage (3747) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (3748). In some computer systems, the system bus (3748) may be accessed in the form of one or more physical plugs to allow expansion of additional CPUs, GPUs, etc. Peripheral devices may be connected to the core's system bus (3748) directly or via a peripheral bus (3749). In one example, a display (3710) may be connected to a graphics adapter (3750). The architecture of the peripheral bus includes PCI, USB, etc.

[0342] The CPU (3741), GPU (3742), FPGA (3743) and accelerator (3744) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (3745) or RAM (3746). Transition data can also be stored in RAM (3746), while permanent data can be stored in, for example, internal mass storage (3747). Fast storage and retrieval of any storage device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (3741), GPUs (3742), mass storage (3747), ROM (3745), RAM (3746), etc.

[0343] The computer readable medium may have computer codes for performing various computer-implemented operations. The media and computer codes may be specially designed and constructed for the purposes of the present disclosure, or may be of a type well known and available to those skilled in the art of computer software.

[0344] As an example and not limitation, a computer system having an architecture (3700), in particular a core (3740), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software contained in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with a user-accessible mass storage as described above and certain memories of a core (3740) that are non-transitory, for example, a core internal mass storage (3747) or a ROM (3745). Software that implements various embodiments of the present disclosure can be stored in such a device and executed by the core (3740). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the core (3740) and in particular the processor therein (including a CPU, GPU, FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in RAM (3746) and modifying such a data structure according to a software-defined process. In addition or as an alternative, the computer system may provide functionality as a result of logic (e.g., accelerator (3744)) hardwired or otherwise contained in circuits that may operate in place of or in conjunction with software to perform specific processes or specific portions of specific processes described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits containing logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0345] Appendix A: Abbreviations

[0346] JEM: joint exploration model

[0347] VVC: versatile video coding, multifunctional video coding

[0348] BMS: benchmark set, benchmark set

[0349] MV: Motion Vector

[0350] HEVC: High Efficiency Video Coding, High Efficiency Video Coding

[0351] MPM: most probable mode, most likely mode

[0352] WAIP: Wide-Angle Intra Prediction, wide-angle intra prediction

[0353] SEI: Supplementary Enhancement Information, Supplementary Enhancement Information

[0354] VUI:Video Usability Information, video availability information

[0355] GOPs: Groups of Pictures, picture groups

[0356] TUs: Transform Units, Transformation Unit

[0357] PUs: Prediction Units, prediction units

[0358] CTUs: Coding Tree Units, coding tree unit

[0359] CTBs: Coding Tree Blocks, coding tree blocks

[0360] PBs: Prediction Blocks, prediction block

[0361] HRD:Hypothetical Reference Decoder, Hypothetical Reference Decoder

[0362] SDR: standard dynamic range, standard dynamic range

[0363] SNR:Signal Noise Ratio, signal-to-noise ratio

[0364] CPUs: Central Processing Units, central processing units

[0365] GPUs: Graphics Processing Units, Graphics Processing Units

[0366] CRT: Cathode Ray Tube, cathode ray tube

[0367] LCD:Liquid-Crystal Display, liquid crystal display

[0368] OLED: Organic Light-Emitting Diode

[0369] CD: Compact Disc, CD

[0370] DVD: Digital Video Disc, Digital Video Disc

[0371] ROM: Read-Only Memory, read-only memory

[0372] RAM: Random Access Memory, Random Access Memory

[0373] ASIC: Application-Specific Integrated Circuit

[0374] PLD:Programmable Logic Device, Programmable Logic Device

[0375] LAN: Local Area Network, local area network

[0376] GSM: Global System for Mobile communications, Global System for Mobile Communications

[0377] LTE: Long-Term Evolution

[0378] CANBus: Controller Area Network Bus, Controller Area Network Bus

[0379] USB: Universal Serial Bus, Universal Serial Bus

[0380] PCI: Peripheral Component Interconnect, Peripheral Component Interconnect

[0381] FPGA: Field Programmable Gate Areas, Field Programmable Gate Area (or, Field Programmable Gate Array)

[0382] SSD: solid-state drive, solid-state drive

[0383] IC: Integrated Circuit

[0384] CU: Coding Unit, coding unit

[0385] PDPC: Position Dependent Prediction Combination, position-dependent prediction combination

[0386] ISP: Intra Sub-Partitions, intra-frame sub-partitions

[0387] SPS:Sequence Parameter Setting, sequence parameter setting

[0388] HDR: high dynamic range

[0389] SDR: standard dynamic range, standard dynamic range

[0390] JVET: Joint Video Exploration Team

[0391] MPM: most probable mode, most likely mode

[0392] WAIP: Wide-Angle Intra Prediction, wide-angle intra prediction

[0393] CU: Coding Unit, coding unit

[0394] PU:Prediction Unit, prediction unit

[0395] TU:Transform Unit, Transformation Unit

[0396] CTU: Coding Tree Unit, coding tree unit

[0397] PDPC: Position Dependent Prediction Combination, position-dependent prediction combination

[0398] ISP: Intra Sub-Partitions, intra-frame sub-partitions

[0399] SPS:Sequence Parameter Setting, sequence parameter setting

[0400] PPS:Picture Parameter Set, picture parameter set

[0401] APS: Adaptation Parameter Set, Adaptation Parameter Set

[0402] VPS: Video Parameter Set, video parameter set

[0403] DPS: Decoding Parameter Set, decoding parameter set

[0404] ALF: Adaptive Loop Filter, Adaptive Loop Filter

[0405] SAO: Sample Adaptive Offset, sample adaptive offset

[0406] CC-ALF: Cross-Component Adaptive Loop Filter, cross-component adaptive loop filter

[0407] CDEF: Constrained Directional Enhancement Filter, constrained directional enhancement filter CCSO: Cross-Component Sample Offset, cross-component sample offset

[0408] CCSAO: Cross-Component Sample Adaptive Offset, cross-component sample adaptive offset

[0409] LSO: Local Sample Offset, local sample offset

[0410] LR: Loop Restoration Filter, loop restoration filter

[0411] FSR: Frame Super-Resolution, frame super resolution

[0412] AV1:AOMedia Video 1, AOMedia Video 1

[0413] AV2: AOMedia Video 2, AOMedia Video 2

[0414] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. A method for filtering in video coding, It is characterized in that include: buffering first boundary pixel values ​​of a subset of first reconstructed samples at a first node along the loop filter chain; buffering second boundary pixel values ​​of a subset of second reconstructed samples at a second node along the loop filter chain, the second reconstructed samples at the second node being generated after applying at least one of the first filter or the second filter to the first reconstructed samples; as well as A loop restoration filter is applied to the reconstructed samples to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value.

2. The method according to claim 1, It is characterized in that The first filter is at least one of a cross component sample offset (CCSO) filter and a local sample offset (LSO) filter.

3. The method according to claim 1, It is characterized in that The first reconstructed samples at the first node are input to the first filter.

4. The method according to claim 3, It is characterized in that Also includes: After generating a sample offset using the first filter, generating the second reconstructed sample at the second node, wherein the sample offset is obtained by applying a first boundary pixel value of a subset of the first reconstructed samples to the first filter, or the sample offset is obtained by applying a second intermediate reconstructed sample to the first filter, and the second intermediate reconstructed sample is obtained by applying the first boundary pixel value of a subset of the first reconstructed samples to the second filter.

5. The method according to claim 4, It is characterized in that Also includes: The sample offsets generated by the first filter are combined with an output of the second filter to generate the second reconstructed samples.

6. The method according to claim 4, It is characterized in that Also includes: combining the sample offsets generated by the first filter with the first reconstructed samples to generate intermediate reconstructed samples; as well as The second filter is applied to the intermediate reconstructed samples to generate the second reconstructed samples.

7. The method according to claim 3, It is characterized in that Also includes: The second reconstructed samples at the second node are combined with sample offsets generated by the first filter to generate the reconstructed samples to be filtered.

8. The method according to claim 1, It is characterized in that The first reconstructed samples are generated by a deblocking filter, and the method further comprises: applying the second filter to the first reconstructed samples to generate intermediate reconstructed samples; The intermediate reconstructed samples are combined with sample offsets generated by the first filter to generate the second reconstructed samples.

9. The method according to any one of claims 1 to 6, It is characterized in that Also includes: Before applying the loop restoration filter, the reconstructed samples to be filtered are clipped.

10. The method according to claim 6 or 8, It is characterized in that Also includes: The intermediate reconstructed samples are clipped before being combined with the sample offsets generated by the first filter.

11. A video encoding method, It is characterized in that include: Reconstruct the pixel values ​​in the current image frame to obtain a reconstructed sample; Applying a deblocking filter to the reconstructed sample to obtain a first reconstructed sample; buffering a first boundary pixel value of a subset of first reconstructed samples at a first node along the loop filter chain, the first boundary pixel value being a value of a pixel at a frame boundary; buffering second boundary pixel values ​​of a subset of second reconstructed samples at a second node along the loop filter chain, the second reconstructed samples at the second node being generated after applying at least one of the first filter or the second filter to the first reconstructed samples, and the second boundary pixel values ​​being values ​​of pixels at the frame boundary; as well as Applying a loop recovery filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value to obtain a reconstructed sample after filtering; An encoding process is performed based on the reconstructed samples after the filtering process.

12. A video decoding method, It is characterized in that include: Decoding the received coded bit stream to obtain a decoded symbol; Reconstruct pixels based on the decoded symbols to obtain reconstructed samples; Applying the reconstructed sample to a deblocking filter to obtain a first reconstructed sample; buffering first boundary pixel values ​​of a subset of first reconstructed samples at a first node along the loop filter chain; buffering second boundary pixel values ​​of a subset of second reconstructed samples at a second node along the loop filter chain, the second reconstructed samples at the second node being generated after applying at least one of the first filter or the second filter to the first reconstructed samples; Applying a loop recovery filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value to obtain a reconstructed sample after filtering; and The decoded image frame is output based on the reconstructed samples after the filtering process.

13. A method for generating a video bitstream, It is characterized in that The video encoding method according to claim 11 generates an encoded bit stream of a video.

14. A device for video encoding and decoding, It is characterized in that The method comprises a processor and a memory, wherein the memory is used to store program instructions, and the processor is used to call the program instructions stored in the memory to implement the method according to any one of claims 1 to 13.

15. A computer-readable storage medium, It is characterized in that The method comprises instructions which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 13.