METHOD AND APPARATUS FOR ADAPTIVE LOOP FILTERING AND Cross-COMPONENT ADAPTIVE LOOP FILTER

By using spatial adjacent sample points associated with the current sample points for filtering in video decoding and encoding, the problem of low adaptive loop filtering efficiency in the prior art is solved, and more efficient video data compression and quality improvement are achieved.

CN119948867APending Publication Date: 2025-05-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069053.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-10-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has problems with inefficiency in the adaptive loop filtering process, especially when dealing with chromaticity and brightness samples, it is difficult to effectively utilize information about spatially adjacent samples.

Method used

A video decoding and encoding method is proposed, by obtaining spatial adjacent samples associated with the current chromaticity or brightness samples, using these samples for filtering, and then obtaining filtered samples. This method is applicable in both decoder and encoder and supports a variety of signal sources, including chromaticity prediction signals, chroma residual signals, signals before adaptive offset of chromaticity sample points, and signals before deblocking.

Benefits of technology

By utilizing the information of spatial adjacent sample points, the efficiency of adaptive loop filtering is improved, the compression performance and quality of video data is improved, the bit rate is reduced, and the high-quality performance of video is maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948867A_ABST
    Figure CN119948867A_ABST
Patent Text Reader

Abstract

Methods and apparatus for improving coding and decoding efficiency of an adaptive loop filter (ALF) are provided. The decoder obtains one or more spatially adjacent samples associated with the current chroma sample. The one or more spatially adjacent samples are from at least one of (i) a chroma prediction signal, (ii) a chroma residual signal, (iii) a signal prior to chroma sample adaptive offset (SAO), or (iv) a signal prior to chroma deblocking. The decoder obtains a filtered chroma sample based on one or more spatially adjacent samples associated with a current chroma sample.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is filed based on and claims priority to U.S. Provisional Application No. 63 / 412,345 filed on September 30, 2022, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present application relates to video coding and compression. More specifically, the present application relates to methods and apparatus for improving the adaptive loop filtering process. Background Art

[0004] Various electronic devices (e.g., digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc.) support digital video. Electronic devices send and receive or otherwise transmit digital video data through a communication network, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited memory resources of the storage device, before the video data is transmitted or stored, video codecs can be used to compress the video data according to one or more video codec standards. For example, video codec standards include general video codecs (VVC), joint exploration test models (JEM), high efficiency video codecs (HEVC / H.265), advanced video codecs (AVC / H.264), moving picture experts group (MPEG) codecs, etc. Video codecs generally use prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that utilize the inherent redundancy in video data. Video codecs are intended to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality. Summary of the invention

[0005] Embodiments of the present disclosure provide techniques related to adaptive loop filtering.

[0006] In a first aspect, the present disclosure provides a video decoding method, comprising: obtaining, by a decoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples come from at least one of the following signals: (i) a chroma prediction signal, (ii) a chroma residual signal, (iii) a signal before chroma sample adaptive offset (SAO), or (iv) a signal before chroma deblocking; and obtaining, by the decoder, a filtered chroma sample based on the one or more spatial neighboring samples associated with the current chroma sample.

[0007] In a second aspect, the present disclosure provides a video encoding method, comprising: obtaining, by an encoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples come from at least one of the following signals: (i) a chroma prediction signal, (ii) a chroma residual signal, (iii) a signal before chroma sample adaptive offset (SAO), or (iv) a signal before chroma deblocking; and obtaining, by the encoder, a filtered chroma sample based on the one or more spatial neighboring samples associated with the current chroma sample.

[0008] In a third aspect, the present disclosure provides a video decoding method, comprising: obtaining, by a decoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples come from at least one of the following signals: (i) a luma prediction signal, (ii) a luma residual signal, (iii) a signal before luma sample adaptive offset (SAO), or (iv) a signal before luma deblocking; and obtaining, by the decoder, filtered chroma samples based on the one or more spatial neighboring samples associated with the current chroma sample.

[0009] In a fourth aspect, the present disclosure provides a video encoding method, comprising: obtaining, by an encoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples come from at least one of the following signals: (i) a luma prediction signal, (ii) a luma residual signal, (iii) a signal before luma sample adaptive offset (SAO), or (iv) a signal before luma deblocking; and obtaining, by the encoder, filtered chroma samples based on the one or more spatial neighboring samples associated with the current chroma sample.

[0010] In a fifth aspect, the present disclosure provides a video decoding method, comprising: obtaining, by a decoder, coding information associated with a coding block, wherein the coding information comprises: a first flag indicating that the coding block is encoded and decoded using a skip mode and a second flag indicating that the coding block is encoded and decoded using at least one of the following modes: intra mode, inter P mode, or inter B mode, so as to derive a new classifier for an online adaptive loop filter (ALF) process; and generating, by the decoder, a new classifier for the online adaptive ALF process based on the coding information.

[0011] In a sixth aspect, the present disclosure provides a video encoding method, comprising: obtaining, by an encoder, encoding information associated with a coding block, wherein the encoding information includes information whether the coding block is encoded and decoded using a skip mode and information whether the coding block is encoded and decoded using at least one of the following modes: intra mode, inter P mode, or inter B mode, so as to derive a new classifier for an online adaptive loop filter (ALF) process; and generating, by the encoder, a new classifier for the online ALF process based on the encoding information.

[0012] It is to be understood that both the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0014] Figure 1 is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to some embodiments of the present disclosure.

[0015] Figure 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0016] Figure 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0017] FIG. 4A to FIG. 4E is a block diagram illustrating how a frame may be recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure.

[0018] Figure 5 ALF filter shapes according to some examples of the present disclosure are shown.

[0019] Figure 6 Depicted are sub-sampled sample point gradients according to some examples of the present disclosure.

[0020] Figure 7 Geometric transformations of diamond filter shapes according to some examples of the present disclosure are shown.

[0021] Figure 8 In-line filter shapes used in ECM according to some examples of the present disclosure are shown.

[0022] Fig. 9 The CCALF architecture according to some examples of the present disclosure is shown.

[0023] Fig.10The filtered chroma samples and their supported relative positions in the luma plane are shown for a 4:2:0 chroma format with chroma position type 0.

[0024] Fig.11 A 25-tap long filter according to some examples of the present disclosure is shown.

[0025] Fig.12 Filter shapes for a prediction signal or a signal before SAO according to examples of the present disclosure are shown.

[0026] Fig.13A Adjusted ALF filter shapes according to some examples of the present disclosure are shown.

[0027] Fig. 13B Various online ALF filter inputs are shown according to some examples of the present disclosure.

[0028] Fig. 13C 1×1 and 3×3 filter shapes applied to prediction samples of ALF according to some examples of the present disclosure are shown.

[0029] Fig.14 is a flowchart illustrating a video decoding method according to some examples of the present disclosure.

[0030] Fig.15 is a flowchart illustrating a video encoding method according to some examples of the present disclosure.

[0031] Fig.16 is a flowchart illustrating a video decoding method according to some examples of the present disclosure.

[0032] Fig.17 is a flowchart illustrating a video encoding method according to some examples of the present disclosure.

[0033] Fig.18 is a flowchart illustrating a video decoding method according to some examples of the present disclosure.

[0034] Fig.19 is a flowchart illustrating a video encoding method according to some examples of the present disclosure.

[0035] Fig. 20 is a diagram illustrating a computing environment coupled to a user interface according to some implementations of the present disclosure. DETAILED DESCRIPTION

[0036] Reference will now be made in detail to specific embodiments, examples of which are shown in the accompanying drawings. In the following detailed description, a large number of non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it is apparent that those skilled in the art may use various alternatives. For example, it is apparent that those skilled in the art may implement the subject matter presented herein on various types of electronic devices with digital video capabilities.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the drawings are used to distinguish objects, but not to describe any specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those shown in the drawings or described in the present disclosure.

[0038] Filter shapes, linear filtering, and adaptive truncation

[0039] In VVC, ALF is applied to the output samples of SAO. Figure 5 As shown in Figure 1, the luminance component and chrominance component support two filter shapes: 7×7 diamond and 5×5 diamond. Figure 5 In , each square corresponds to a luma or chroma sample, and the center square corresponds to the current sample to be filtered. The filter coefficients are point symmetric, and each integer filter coefficient is represented with 7-bit decimal precision. In addition, the sum of the coefficients of a filter is equal to 128, which is the 7-bit decimal precision fixed-point representation of 1.0:

[0040]

[0041] where the number of coefficients N is equal to 13 and 7 for filter shapes of 7×7 and 5×5 respectively. The value of the filtered sample at coordinate (x, y) It is obtained by applying the coefficient c to the value of the reconstructed sample point R(x,y) i The results are as follows:

[0042]

[0043] Among them, (x+x i ,y+y i ) and (xx i ,yy i ) is related to the i-th coefficient c i The corresponding coordinates of the reconstructed sample points. Due to the constraints in equation (1), equation (2) can be written as:

[0044]

[0045] In VVC, the possibility of intercepting the difference between the neighboring sample value and the current sample to be filtered is added in equation (3), as follows:

[0046]

[0047] in,

[0048] f i =min(b i ,max(-b i ,R(x+x i ,y+y i )-R(x,y)))+

[0049] min(b i ,max(-b i ,R(xx i ,yy i )-R(x,y)))

[0050] b i is the truncation parameter of coefficient c, given by the truncation index d i Decision. i It is derived as follows:

[0051]

[0052] Where BD is the sample bit depth, and d i Can be 0, 1, 2, or 3.

[0053] Luma sub-block level filter adaptation

[0054] In VVC, sub-block level filter adaptation is only applied to the luma component. Each 4×4 luma block is classified based on its directionality and 2D Laplacian activity. First, the sample gradient values ​​in the horizontal, vertical and two diagonal directions are calculated:

[0055] H k,l =|2R(k,l)-R(k-1,l)-R(k+1,l)|,

[0056] V k,l =|2R(k,l)-R(k,l-1)-R(k,l+1)|,

[0057] D0 k,l =|2R(k,l)-R(k-1,l-1)-R(k+1,l+1)|,

[0058] D1 k,l =|2R(k,l)-R(k-1,l+1)-R(k+1,l-1)|.

[0059] Based on the sample point gradient, the sub-block horizontal gradient g h , vertical gradient g v and two diagonal gradients g d0 and g d1 is calculated as

[0060]

[0061] The indices i and j refer to the coordinates of the top left corner sample point in the 4×4 luminance block. From equation (8), it can be seen that the sum of the gradients of the samples in the 10×10 luminance window covering the target 4×4 block is used to classify the block. To reduce complexity, only the gradient of every other sample point in the 10×10 window is calculated, as Figure 6 As shown. (See Figure 6 , which depicts the sub-sampled sample gradients for the 4×4 sub-block ALF classification. The gradient values ​​of the samples marked with x are calculated. The gradient values ​​of the other samples are set to 0).

[0062] Secondly, to assign the directionality D, the ratio of the maximum to the minimum horizontal and vertical gradients of the sub-block is calculated:

[0063]

[0064] And the ratio of the maximum to minimum diagonal gradients of the two sub-blocks:

[0065]

[0066] Compared with a set of thresholds t1 and t2: Step 1: If and Then D is set to 0. Step 2: If Then calculate the directivity D in step 3, otherwise calculate it in step 4. Step 3: If Then D is set to 2, otherwise D is set to 1. Step 4: If Then D is set to 4, otherwise D is set to 3.

[0067] Each subsequent step in the calculation of D above is performed only if no value was assigned to D in the previous step. Third, the activity value A is calculated as

[0068]

[0069] A is further mapped to the range 0 to 4: Among them, {Q n Finally, each 4×4 luma block is classified into one of the following 25 classes:

[0070]

[0071] Each class can be assigned its own filter. Before filtering each 4×4 luma block, a geometric transformation is applied to the filter coefficients, such as a 90 degree rotation, diagonal or vertical flip, e.g. Figure 7 As shown (illustrating the geometric transformation of the 7×7 diamond filter shape. From left to right: diagonal flip, vertical flip, 90 degree rotation), depending on the sub-block gradient value specified in Table 1.

[0072] Sub-block gradient value Transform <![CDATA[g d1 <g d0 And g h <g v ]]> No transformation <![CDATA[g d1 <g d0 And g v ≤g h ]]> Diagonal Flip <![CDATA[g do ≤g d1 And g h <g v ]]> Flip Vertically <![CDATA[g do ≤g d1 And g v ≤g h ]]> 90 degree rotation

[0073] Table 1 Geometric transformation based on sub-block gradient values

[0074] Coding tree block level filter adaptation

[0075] In addition to luma 4×4 block-level filter adaptation, ALF also supports CTB-level filter adaptation. The luma CTB can use the filter group calculated for the current strip, or one of the filter groups calculated for the coded strip. It can also use one of 16 offline trained filter groups. In each luma CTB, which filter in the selected filter group should be applied to each 4×4 block is determined by the class C calculated for the block according to equation (12). Chroma only uses CTB-level filter adaptation. In a strip, a maximum of 8 filters can be used for chroma components. Each CTB can select one of the filters.

[0076] Syntax Design

[0077] The filter coefficients and truncation index are carried in the ALF APS. The ALF APS can include up to 8 chroma filters and a luma filter bank (which can contain up to 25 filters). Each of the 25 luma classes also includes an index i C . With the same index i C Classes of ALF coefficients share the same filter. By combining different classes, the number of bits required to represent the filter coefficients is reduced. The absolute value of the filter coefficients is represented using an exponential Golomb code of order 0, followed by a sign bit for non-zero coefficients. When truncation is enabled, a truncation index is also signaled using a two-bit fixed-length code for each filter coefficient. The storage required for the ALF coefficients and the truncation index within the APS is a maximum of 3480 bits. The decoder can use up to 8 ALF APSs simultaneously.

[0078] The filter control syntax element includes two types of information. First, the ALF on / off flag is signaled at the sequence, picture, slice, and CTB levels. Chroma ALF can be enabled at the corresponding level only when luma ALF is enabled at the picture and slice level. Second, if ALF is enabled at the picture, slice, and CTB level, the filter usage information is signaled at that level. The referenced ALF APS ID is encoded and decoded at the slice level, or at the picture level if all slices within a picture use the same APS. A maximum of 7 ALF APSs can be referenced for the luma component, while a maximum of 1 ALF APS can be referenced for the chroma component. For the luma CTB, an index is signaled that indicates which ALF APS or offline trained luma filter bank to use. For the chroma CTB, the index indicates which filter in the referenced APS to use.

[0079] Reduce line buffer

[0080] To reduce the storage requirements of ALF, VVC uses line buffer boundary processing. In VVC, the line buffer boundary is located at 4 luma samples and 2 chroma samples above the horizontal CTU boundary. When ALF is applied to samples on one side of the line buffer boundary, the samples on the other side of the line buffer boundary cannot be used.

[0081] ALF in ECM

[0082] Remove ALF Simplification

[0083] ALF gradient subsampling and ALF virtual boundary processing were removed. The block size for classification was reduced from 4×4 to 2×2. The filter size for luma and chroma (where ALF coefficients are signaled) was increased to 9×9.

[0084] ALF with fixed filter

[0085] To filter the luminance samples, three different classifiers (C0, C1, and C2) and three different filter banks (F0, F1, and F2) are used. Banks F0 and F1 contain fixed filters whose coefficients are trained for classifiers C0 and C1. The filter coefficients in F2 are signaled. i Which filter is used by the classifier C i The class C assigned to this sample point i. decided.

[0086] Filtering

[0087] First, two 13×13 diamond fixed filters F0 and F1 are applied to obtain two intermediate samples R0(x,y) and R1(x,y). Then, F2 is applied to R0(x,y), R1(x,y), neighboring samples, and samples before the deblocking filter (DBF) to obtain the filtered samples:

[0088]

[0089] Among them, f i,j is the intercept difference between the neighboring sample point and the current sample point R(x,y), g i YesR i-20 The intercept difference between (x, y) and the current sample point R(x, y), h i,j It is the intercept difference between the adjacent sample point before DBF and the current sample point R(x,y). The filter coefficient c is transmitted by the signal i , i=0,…24. The filter shape of F2 is as follows Figure 8 shown.

[0090] Classification

[0091] Based on the directionality D i and activities Assign class C to each 2×2 block i :

[0092]

[0093] Among them, M D,i Indicates directionality D i The total number of. As with VVC, the 1-D Laplacian operator is used to calculate the value of the horizontal, vertical, and two diagonal gradients of each sample point. The sum of the sample gradients within the 4×4 window covering the target 2×2 block is used for classifier C0, and the sum of the sample gradients within the 12×12 window is used for classifiers C1 and C2. The sum of the horizontal, vertical, and two diagonal gradients is expressed as and Directionality D i This is done by:

[0094]

[0095] The directionality D2 is determined by comparing it with a set of thresholds. In VVC, the directionality D2 is derived using thresholds 2 and 4.5. For D0 and D1, the horizontal / vertical edge strength is first calculated. and diagonal edge strength Threshold Th = [1.25, 1.5, 2, 3, 4.5, 8] is used. If The edge strength is 0; otherwise, is satisfied The largest integer. If The edge strength is 0; otherwise, is satisfied The maximum integer of . When horizontal / vertical edges dominate, D i is obtained using Table 2(a); otherwise, when diagonal edges dominate, D i It is obtained using Table 2.

[0096]

[0097] Table 2. and To D i Mapping

[0098] To obtain The sum of the vertical and horizontal gradients A i Mapped to the range from 0 to n, where n is equal to 4, and for and n is equal to 15. In ALF_APS, up to 4 luminance filter groups are signaled, each group can have up to 25 filters.

[0099] Alternative 2×2 ALF classifier

[0100] The classification in ALF is extended by the addition of an alternative classifier. For a signaled luma filterbank, a flag is signaled to indicate if the alternative classifier is applied. No geometric transformation is applied for the alternative band classifier. When applying a band-based classifier, the sum of the sample values ​​of the 2×2 luma block is first calculated. The class index is then calculated as: class_index = (sum*25) >> (sample bit depth + 2).

[0101] CCALF in VV

[0102] Filter shape and accuracy

[0103] CCALF uses the luma sample values ​​to refine the chroma sample values ​​in the ALF process. Fig. 9 As shown, the linear filtering operation takes the luma sample value as input and generates correction values ​​for the chroma sample values. The correction is generated independently for each chroma component i (i∈{Cb,Cr}) and can be expressed as:

[0104] Among them, (x, y) is the sample position of chrominance component i, (x C ,y C ) is the brightness sample position obtained from (x,y), (x0,y0) is (x C ,y C ) around the filter support offset, S i is the filter support region in luma for chroma component i. Luma position (x C ,y C ) is determined based on the spatial scaling factor between the luma plane and the chroma plane. The sample values ​​in the luma support region are also the input to the ALF luma stage and correspond to the output of the SAO stage.

[0105] like Fig.10 As shown, the CCALF filter has a diamond shape. Fig.10 As shown, for a 4:2:0 video sequence with chroma position type 0 (i.e., when chroma samples are co-located with even columns of luma samples horizontally and between rows of luma samples vertically), the center of the diamond is aligned with the chroma sample position.

[0106] Since symmetry constraints are not enforced, CCALF coefficients have greater flexibility than regular ALF coefficients. However, there are two limitations: (1) To maintain DC neutrality, the sum of CCALF coefficient values ​​is required to be zero. Therefore, only seven of the eight CCALF coefficients need to be signaled in the bitstream, and positions (x C ,y C ) is obtained at the decoder; (2) The absolute value of the CCALF coefficient is restricted to zero or an integer power of two, specifically {0, 1, 2, 4, 8, 16, 32, 64}. This enables the implementation to use variable shift operations instead of CCALF multiplication as needed.

[0107] Syntax Design

[0108] In the final design of VVC, the maximum number of filters for each chroma component of a picture is four. Different sets of CCALF coefficients can be selected for each CTU of a chroma component. Like regular ALF coefficients, CCALF coefficients are signaled within the ALF APS. Each ALF APS can contain up to four CCALF filters for each chroma component. Although CCALF can be enabled at the sequence level, it can only be enabled if ALF is also enabled for the sequence. Similarly, CCALF can be enabled at the corresponding level only when luma ALF is enabled at the picture and slice level.

[0109] CCALF in ECM

[0110] The CCALF process uses a linear filter to filter the luminance sample values ​​and generate residual corrections for the chrominance samples. The CCALF process uses a 25-tap large filter, such as Fig.10 For a given stripe, the encoder can collect statistics for that stripe, analyze it, and can signal up to 16 filters through APS.

[0111] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to one of ordinary skill in the art that various alternatives may be used without departing from the scope of the claims, and that the subject matter may be practiced without these specific details. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0112] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the drawings are used to distinguish objects, but not to describe any specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those shown in the drawings or described in the present disclosure.

[0113] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. Figure 1 As shown in , system 10 includes a source device 12 that generates and encodes video data to be later decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0114] In some embodiments, the target device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the target device 14. In one example, the link 16 may include a communication medium that enables the source device 12 to send the encoded video data directly to the target device 14 in real time. The encoded video data may be modulated according to a communication standard (e.g., a wireless communication protocol) and sent to the target device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form a portion of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include a router, a switch, a base station, or any other device that may be beneficial to facilitate communication from the source device 12 to the target device 14.

[0115] In some other embodiments, the encoded video data may be sent from the output interface 22 to the storage device 32. The encoded video data in the storage device 32 may then be accessed by the target device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device 32 may correspond to a file server or another intermediate storage device that may hold the encoded video data generated by the source device 12. The target device 14 may access the stored video data from the storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing encoded video data and sending the encoded video data to the target device 14. Exemplary file servers include a network server (e.g., for a website), a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Target device 14 may access the encoded video data through any standard data connection suitable for accessing encoded video data stored on a file server, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both wireless channels and wired connections. The transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both streaming and download transmissions.

[0116] like Figure 1As shown in , source device 12 includes video source 18, video encoder 20 and output interface 22. Video source 18 may include sources such as or a combination of such sources: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video. As an example, if video source 18 is a camera of a security monitoring system, source device 12 and target device 14 may form a camera phone or a video phone. However, the embodiments described in this application may be generally applicable to video encoding and decoding, and may be applied to wireless and / or wired applications.

[0117] The captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be sent directly to the target device 14 via the output interface 22 of the source device 12. The encoded video data may also (or alternatively) be stored on a storage device 32 for later access by the target device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.

[0118] Target device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 may include a receiver and / or a modem and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data sent over a communication medium, stored on a storage medium, or stored on a file server.

[0119] In some implementations, the target device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with the target device 14. The display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0120] The video encoder 20 and the video decoder 30 can operate according to a proprietary standard or an industry standard (e.g., VVC, HEVC, Part 10 of MPEG-4, AVC) or an extension of such a standard. It should be understood that the present application is not limited to a specific video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally believed that the video encoder 20 of the source device 12 can be configured to encode the video data according to any of these current standards or future standards. Similarly, it is also generally believed that the video decoder 30 of the target device 14 can be configured to decode the video data according to any of these current standards or future standards.

[0121] The video encoder 20 and the video decoder 30 can be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium, and use one or more processors to execute the instructions in the hardware to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders, and either of the encoders or decoders can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.

[0122] In some embodiments, components of source device 12 (e.g., video source 18, video encoder 20, or the like) may include: Figure 2 The components described include at least a portion of the components in the video encoder 20, and the output interface 22) and / or components of the target device 14 (e.g., the input interface 28, the video decoder 30 or the components described below). Figure 3The components described, including the video decoder 30, and at least a portion of the components in the display device 34) may be operated in a cloud computing service network such as software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS), wherein the cloud computing service network may provide software, platform, and / or infrastructure. In some embodiments, one or more components of the source device 12 and / or the target device 14 that are not included in the cloud computing service network may be set in one or more client devices, and the one or more client devices may communicate with a server computer in the cloud computing service network through a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In an embodiment, at least a portion of the operations described herein may be implemented as a cloud-based service provided by one or more server computers, wherein the one or more server computers are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in the cloud computing service network; and one or more other operations described herein may be implemented by one or more client devices. In some embodiments, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of the present disclosure, terms such as "cloud", "cloud computing", "cloud-based", etc. may be used interchangeably as appropriate. It should be understood that the present disclosure is not limited to being implemented in the above-mentioned cloud computing service network. Instead, the present disclosure may also be implemented in any other type of computing environment currently known or developed in the future.

[0123] Figure 2 is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described in the present application. The video encoder 20 can perform intra-frame prediction coding and inter-frame prediction coding on video blocks within a video frame. Intra-frame prediction coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-frame prediction coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. It should be noted that in the field of video coding and decoding, the term "frame" can be used as a synonym for the term "image" or "picture".

[0124] like Figure 2As shown in FIG. 1 , the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63 such as a deblocking filter can be located between the adder 62 and the DPB 64 to filter the block boundary to remove the block effect from the reconstructed video. In addition to the deblocking filter, another in-loop filter (such as a sample adaptive offset (SAO) filter and / or an adaptive in-loop filter (ALF)) can also be used to filter the output of the adder 62. In some examples, the in-loop filter can be omitted, and the decoded video block can be directly provided to the DPB 64 by the adder 62. Video encoder 20 may take the form of a fixed or programmable hardware unit, or may be distributed among one or more of the illustrated fixed or programmable hardware units.

[0125] Video data memory 40 may store video data to be encoded by components of video encoder 20. Figure 1 The video source 18 shown obtains video data in the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use by the video encoder 20 (e.g., in intra-frame or inter-frame prediction coding mode) when encoding the video data. The video data memory 40 and the DPB 64 can be formed by any of a variety of memory devices. In various examples, the video data memory 40 can be on-chip with other components of the video encoder 20, or off-chip relative to those components.

[0126] like Figure 2As shown in , after receiving the video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a set of video blocks) or other larger coding units (CUs) according to a predefined splitting structure (e.g., a quadtree (QT) structure) associated with the video data. A video frame is or may be considered as a two-dimensional sample array or matrix with sample values. The samples in the array may also be referred to as pixels or picture elements (pel). The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. For example, a video frame may be divided into a plurality of video blocks by using QT segmentation. A video block is again or may be considered as a two-dimensional sample array or matrix with sample values, but its dimensions are smaller than the dimensions of a video frame. The number of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. By, for example, iteratively using QT segmentation, binary tree (BT) segmentation, or ternary tree (TT) segmentation, or any combination thereof, a video block may be further segmented into one or more block partitions or sub-blocks (which may again form blocks). It should be noted that the term "block" or "video block" used herein may be a portion of a frame or picture, in particular a rectangular (square or non-square) portion. With reference to, for example, HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU) or a transform unit (TU) and / or may be or correspond to a corresponding block (e.g., a coding tree block (CTB), a coding block (CB), a prediction block (PB) or a transform block (TB)) and / or a sub-block.

[0127] The prediction processing unit 41 may select one of a plurality of feasible prediction coding modes for the current video block based on the error results (e.g., coding rate and distortion level), such as one of one or more inter-frame prediction coding modes in a plurality of intra-frame prediction coding modes. The prediction processing unit 41 may provide the resulting intra-frame prediction coding block or inter-frame prediction coding block to the adder 50 to generate a residual block, and to the adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, partition information, and other such syntax information) to the entropy coding unit 56.

[0128] To select an appropriate intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-prediction coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple encoding passes, for example, to select an appropriate coding mode for each block of video data.

[0129] In some embodiments, motion estimation unit 42 determines an inter-prediction mode for a current video frame by generating a motion vector according to a predetermined pattern within a sequence of video frames, the motion vector indicating the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating a motion vector that estimates the motion of a video block. For example, a motion vector may indicate the displacement of a video block within a current video frame or picture relative to a prediction block within a reference frame associated with a current block being encoded within the current frame. The predetermined pattern may designate a video frame in a sequence as a P frame or a B frame. Intra BC unit 48 may determine a vector (e.g., a block vector) for intra BC coding in a manner similar to the motion vector determined by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine a block vector.

[0130] In terms of pixel differences, the prediction block for the video block may be or may correspond to a block or reference block of a reference frame that is considered to closely match the video block to be encoded, and the pixel difference may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some embodiments, the video encoder 20 may calculate values ​​for sub-integer pixel positions of the reference frames stored in the DPB 64. For example, the video encoder 20 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, the motion estimation unit 42 may perform a motion search relative to full pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.

[0131] Motion estimation unit 42 calculates a motion vector for a video block in an inter-prediction coded frame by comparing the position of the video block to the position of a prediction block of a reference frame selected from a first reference frame list (list 0) or a second reference frame list (list 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.

[0132] Motion compensation performed by motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by motion estimation unit 42. After receiving the motion vector for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel difference values ​​by subtracting the pixel values ​​of the prediction block provided by motion compensation unit 44 from the pixel values ​​of the current video block being encoded.

[0133] The pixel difference values ​​forming the residual video block may include luma component differences or chroma component differences or both. Motion compensation unit 44 may also generate syntax elements associated with the video block of the video frame for use by video decoder 30 when decoding the video block of the video frame. The syntax elements may include, for example, syntax elements defining motion vectors for identifying prediction blocks, any flags indicating prediction modes, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are described separately for conceptual purposes.

[0134] In some embodiments, the intra BC unit 48 may generate vectors and obtain prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44, but these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra BC unit 48 may determine the intra prediction mode to be used to encode the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, during separate encoding passes, and test their performance through rate-distortion analysis. Next, the intra BC unit 48 may select a suitable intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values ​​for the various tested intra prediction modes using rate-distortion analysis, and select an intra prediction mode with the best rate-distortion characteristic among the tested modes as a suitable intra prediction mode to use.

[0135] The rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that is encoded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. The intra BC unit 48 may calculate ratios based on the distortion and rate for the various coded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. In other examples, the intra BC unit 48 may use the motion estimation unit 42 and the motion compensation unit 44 in whole or in part to perform such functions for intra BC prediction according to the embodiments described herein. In either case, for intra block copying, the prediction block may be a block that is considered to closely match the block to be encoded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identifying the prediction block may include calculating values ​​for sub-integer pixel positions.

[0136] Regardless of whether the prediction block is from the same frame according to intra-frame prediction or from a different frame according to inter-frame prediction, the video encoder 20 can form a residual video block by subtracting the pixel values ​​of the prediction block from the pixel values ​​of the current video block being encoded. The pixel difference values ​​forming the residual video block may include both luma component differences and chroma component differences.

[0137] As an alternative to the inter-frame prediction performed by the motion estimation unit 42 and the motion compensation unit 44 or the intra-frame block copy prediction performed by the intra BC unit 48 as described above, the intra-frame prediction processing unit 46 can perform intra-frame prediction on the current video block. Specifically, the intra-frame prediction processing unit 46 can determine the intra-frame prediction mode for encoding the current block. To this end, the intra-frame prediction processing unit 46 can use various intra-frame prediction modes to encode the current block, for example, during separate encoding passes, and the intra-frame prediction processing unit 46 (or in some examples, the mode selection unit) can select a suitable intra-frame prediction mode from the tested intra-frame prediction modes to use. The intra-frame prediction processing unit 46 can provide information indicating the intra-frame prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra-frame prediction mode into the bitstream.

[0138] After prediction processing unit 41 determines a prediction block for the current video block via inter-prediction or intra-prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform (e.g., a discrete cosine transform (DCT) or a conceptually similar transform).

[0139] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan on the matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 may perform the scan.

[0140] After quantization, entropy coding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream may then be sent to a video bitstream such as a Figure 1 The video decoder 30 shown, or archived as Figure 1 The video frame may be stored in storage device 32 as shown for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 may also entropy encode motion vectors and other syntax elements for the current video frame being encoded.

[0141] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain for use in generating reference blocks for predicting other video blocks. As noted above, motion compensation unit 44 may generate a motion compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values ​​for use in motion estimation.

[0142] Adder 62 adds the reconstructed residual block to the motion compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block may then be used as a prediction block by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 to inter-predict another video block in a subsequent video frame.

[0143] Figure 3 1 is a block diagram showing an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 may perform the above-mentioned Figure 2The encoding process is substantially the inverse of the decoding process described with respect to video encoder 20. For example, motion compensation unit 82 may generate prediction data based on motion vectors received from entropy decoding unit 80, and intra-prediction unit 84 may generate prediction data based on intra-prediction mode indicators received from entropy decoding unit 80.

[0144] In some examples, units of the video decoder 30 may be tasked to perform embodiments of the present application. In addition, in some examples, embodiments of the present disclosure may be dispersed in one or more of the units of the video decoder 30. For example, the intra BC unit 85 may perform embodiments of the present application alone or in combination with other units of the video decoder 30 (e.g., the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., the motion compensation unit 82).

[0145] The video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, from a local video source (e.g., a camera), via a wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from an encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data for use by the video decoder 30 when decoding the video data (e.g., in an intra-frame or inter-frame prediction coding mode). The video data memory 79 and the DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are shown in FIG. Figure 3 9 as two different components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 can be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 can be on-chip with other components of the video decoder 30, or off-chip relative to those components.

[0146] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes the bitstream to generate quantization coefficients, motion vectors or intra-frame prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-frame prediction mode indicators, and other syntax elements to the prediction processing unit 81.

[0147] When a video frame is encoded as an intra-prediction coded (I) frame or for intra-coded prediction blocks in other types of frames, the intra-prediction unit 84 of the prediction processing unit 81 can generate prediction data for a video block of the current video frame based on the intra-prediction mode transmitted by the signal and reference data from a previously decoded block of the current frame.

[0148] When the video frame is encoded as an inter-frame prediction coding (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the video block of the current video frame based on the motion vector and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame in one of the reference frame lists. The video decoder 30 can construct the reference frame lists, i.e., List 0 and List 1, based on the reference frames stored in the DPB 92 using a default construction technique.

[0149] In some examples, when a video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from entropy decoding unit 80. The prediction block may be within a reconstructed region of the same picture as the current video block as defined by video encoder 20.

[0150] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for a video block of the current video frame by parsing the motion vector and other syntax elements, and then use the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) for encoding a video block of a video frame, an inter prediction frame type (e.g., B or P), construction information for one or more of the reference frame lists for the frame, motion vectors for each inter prediction encoded video block of the frame, inter prediction states for each inter prediction encoded video block of the frame, and other information for decoding a video block in the current video frame.

[0151] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as flags, to determine whether the current video block is predicted using intra BC mode, construction information of which video blocks of the frame are within the reconstruction region and should be stored in the DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding the video block in the current video frame.

[0152] Motion compensation unit 82 may also perform interpolation using interpolation filters as used by video encoder 20 during encoding of the video blocks to calculate interpolated values ​​for sub-integer pixels of reference blocks. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 based on the received syntax elements and use these interpolation filters to produce the prediction blocks.

[0153] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80, using the same quantization parameters calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to reconstruct the residual block in the pixel domain.

[0154] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter 91 (e.g., a deblocking filter, an SAO filter, a CCSAO filter, and / or an ALF) may be located between the adder 90 and the DPB 92 to further process the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be provided directly to the DPB 92 by the adder 90. The decoded video blocks in a given frame are then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later presentation on a display device (e.g., Figure 1 on a display device 34).

[0155] In a typical video encoding and decoding process, a video sequence usually includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other examples, a frame may be monochrome and therefore include only a two-dimensional array of luma samples.

[0156] like Figure 4A As shown in , the video encoder 20 (or more specifically, the segmentation unit 45) generates an encoded representation of a frame by first segmenting the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and from top to bottom in a raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, one of 128×128, 64×64, 32×32, and 16×16. It should be noted, however, that the present application is not necessarily limited to a particular size. As Figure 4B As shown in , each CTU may include one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding samples of the coding tree blocks. The syntax elements describe the properties of different types of units of the coding pixel blocks and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding samples of the coding tree block. The coding tree block may be an N×N block of samples.

[0157] To achieve better performance, the video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof, on the coding tree block of the CTU and divide the CTU into smaller CUs. Figure 4C As depicted in , a 64×64 CTU 400 is first divided into four smaller CUs, each CU having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are respectively divided into four CUs with a block size of 16×16. Two 16×16 CUs 430 and CU 440 are further divided into four CUs with a block size of 8×8. Figure 4D Depicted is a diagram showing Figure 4C The quadtree data structure of the final result of the partitioning process of the CTU 400 depicted in FIG. 4 corresponds to a CU of a corresponding size ranging from 32×32 to 8×8. Figure 4BIn the CTU depicted in FIG, each CU may include a CB of luma samples and two corresponding coding blocks of chroma samples of the same size frame, and syntax elements for encoding the samples of the coding blocks. In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and syntax structures for encoding the samples of the coding block. It should be noted that Fig.10 and Fig.11 The quadtree partitioning depicted in FIG is for illustrative purposes only, and one CTU can be split into multiple CUs based on quadtree partitioning / ternary tree partitioning / binary tree partitioning to adapt to varying local characteristics. In a multi-type tree structure, one CTU is partitioned according to a quadtree structure, and each quadtree leaf CU can be further partitioned according to a binary and ternary tree structure. Figure 4E As shown, there are five possible partitioning types for a coding block with width W and height H, namely, quadruple partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.

[0158] In some embodiments, the video encoder 20 may further partition the coding block of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter or intra) is applied. The PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In a monochrome picture or a picture with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PB. The video encoder 20 may generate a predicted luma block, a predicted Cb block, and a predicted Cr block for the luma PB, Cb PB, and Cr PB of each PU of the CU.

[0159] Video encoder 20 may use intra prediction or inter prediction to generate a prediction block for a PU. If video encoder 20 uses intra prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0160] After the video encoder 20 generates a predicted luma block, a predicted Cb block, and a predicted Cr block for one or more PUs of a CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predicted luma block of the CU from the original luma coding block of the CU, so that each sample in the luma residual block of the CU indicates the difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, so that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0161] In addition, if Figure 4C As shown in , the video encoder 20 can use quadtree partitioning to decompose the luma residual block, Cb residual block and Cr residual block of the CU into one or more luma transform blocks, Cb transform blocks and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. The TU of the CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements for transforming the transform block samples. Therefore, each TU of the CU may be associated with a luma transform block, a Cb transform block and a Cr transform block. In some examples, the luma transform block associated with the TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.

[0162] The video encoder 20 may apply one or more transforms to the luma transform block of the TU to generate a luma coefficient block for the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficient may be a scalar. The video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.

[0163] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream including a sequence of bits that form a representation of an encoded frame and associated data, and the bitstream is stored in the storage device 32 or sent to the target device 14.

[0164] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 can reconstruct the frame of the video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally inverse to the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coding block of the current CU by adding the samples of the prediction block for the PU of the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the coding block for each CU of the frame, the video decoder 30 can reconstruct the frame.

[0165] As mentioned above, video coding mainly uses two modes, namely, intra-frame prediction (or intra-frame prediction) and inter-frame prediction (or inter-frame prediction) to achieve video compression. It should be noted that IBC can be regarded as intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because motion vectors are used to predict the current video block based on the reference video block.

[0166] But with the ever-improving video data capture technology and finer video block sizes for retaining details in the video data, the amount of data required to represent the motion vector for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only a set of neighboring CUs in both the spatial and temporal domains have similar video data for prediction purposes, but also the motion vectors between these neighboring CUs are similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU (which is also called the "motion vector predictor" (MVP) of the current CU) by exploring their spatial and temporal correlations.

[0167] Instead of combining as above Figure 2The actual motion vector of the current CU determined by the motion estimation unit 42 is encoded into the video bitstream, and the motion vector prediction value of the current CU is subtracted from the actual motion vector of the current CU to generate a motion vector difference (MVD) for the current CU. By doing so, the motion vector determined by the motion estimation unit 42 for each CU of the frame does not need to be encoded into the video bitstream, and the amount of data used to represent the motion information in the video bitstream can be significantly reduced.

[0168] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coding block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules for constructing a motion vector candidate list (also called a "merge list") for the current CU using those potential candidate motion vectors associated with the spatial neighboring CUs and / or temporal co-located CUs of the current CU, and then selecting a member from the motion vector candidate list as the motion vector prediction value for the current CU. By doing so, the motion vector candidate list itself does not need to be sent from the video encoder 20 to the video decoder 30, and the index of the selected motion vector prediction value within the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector prediction value within the motion vector candidate list to encode and decode the current CU.

[0169] Although ALF has been improved in ECM, its performance can still be further improved. First, the online ALF filter in ECM takes spatial neighbors, fixed ALF filter results, and spatial neighbors before deblocking filter as input. Other information (such as spatial neighbors in the prediction signal, spatial neighbors in the residual signal, or spatial neighbors before SAO) can also be used as input to the online ALF filter equation, which may benefit the codec performance.

[0170] Secondly, in ECM, the online ALF filter adaptively uses an edge-based classifier and a band-based classifier. However, these two classifiers can be further combined to provide other classifiers, which may benefit the codec performance.

[0171] Third, in ECM, the filter shape of the chroma ALF is a diamond filter shape, while the filter shape of the luma ALF is a long cross shape, and this non-uniform design may not be optimal from a standardization perspective.

[0172] Fourth, the edge-based classifier and band-based classifier in ECM only consider the pixel values ​​after SAO. However, after saving the pixel values ​​before the deblocking filter, the prediction signal, the residual signal, or before SAO as the input of the online ALF filter equation, these pixel values ​​can also be used to design new classifiers, which may benefit the encoding and decoding performance.

[0173] Fifth, the edge-based classifier and band-based classifier in ECM only consider the luma pixel values ​​after SAO. However, the chroma pixel values ​​can also be used to design new classifiers, which may benefit the encoding and decoding performance.

[0174] Sixth, similar to saving the luma pixel values ​​before the deblocking filter, in the prediction signal, in the residual signal, or before SAO as additional online luma ALF filter equation inputs, the chroma pixel values ​​before the deblocking filter, in the prediction signal, in the residual signal, or before SAO can also be saved as additional online chroma ALF filter equation inputs, which may benefit the encoding and decoding performance.

[0175] Seventh, similar to saving the luminance pixel values ​​before the deblocking filter, in the prediction signal, in the residual signal, or before SAO as additional online luminance ALF filter equation input, the luminance pixel values ​​before the deblocking filter, in the prediction signal, in the residual signal, or before SAO can also be saved as additional CCALF filter equation input, which may be beneficial to encoding and decoding performance.

[0176] Eighth, the classifier design in ECM only considers the reconstructed pixel values. However, coding mode information (such as whether the coding block is coded or decoded in skip mode, whether the coding block is coded or decoded in intra, inter P or inter B mode) can also be used to design the classifier, which may benefit the coding and decoding performance.

[0177] To address these issues, the following methods are provided to further improve the existing design of ALF. In general, the main features of the technology proposed in the present disclosure are summarized as follows: (1) The online ALF filter takes spatially neighboring pixels in the prediction signal, spatially neighboring pixels in the residual signal, or spatially neighboring pixels before SAO as additional inputs; (2) A classifier that combines the features of an edge-based classifier and the features of a band-based classifier is used as an additional classifier for the online ALF filter; (3) The filter shape of the chroma ALF is changed from a diamond to a long cross to be unified with the filter shape of the luminance ALF; (4) A classifier that utilizes pixel values ​​before the deblocking filter, in the prediction signal, in the residual signal, or before SAO is used as an additional classifier for the online ALF filter; (5) A classifier that utilizes chroma pixel values ​​is used as an additional classifier for the online ALF filter; An additional classifier for the online ALF filter; (6) the online chroma ALF filter takes the spatial neighboring pixels in the chroma prediction signal, the spatial neighboring pixels in the chroma residual signal, the spatial neighboring pixels before the chroma SAO, or the spatial neighboring pixels before the chroma deblocking as additional input; (7) the CCALF filter takes the spatial neighboring pixels in the luma prediction signal, the spatial neighboring pixels in the luma residual signal, the spatial neighboring pixels before the luma SAO, or the spatial neighboring pixels before the luma deblocking as additional input; and (8) a classifier using coding mode information (such as whether the coding block is encoded or decoded in skip mode, whether the coding block is encoded or decoded in intra-frame, inter-frame P or inter-frame B mode) is used as an additional classifier for the online ALF filter. It should be noted that the disclosed methods can be applied independently or jointly.

[0178] Prediction, residual, or information before SAO is used as additional ALF input

[0179] According to one or more embodiments of the present disclosure, prediction, residual or information before SAO is used as additional ALF equation input. Different methods can be used to achieve this goal.

[0180] In the first approach, it is proposed to use spatially neighboring pixels in the prediction signal as additional ALF equation inputs. Various filter shapes can be used to extract information from the prediction signal. For example, the filter shape can be 1×1, 3×3, or 5×5, such as Fig.12 As shown. Various equation forms can be used to extract information in the prediction signal. In one example, the intercepted difference between the surrounding pixels in the prediction signal and the current pixel is used as the input of the ALF equation. In another example, the intercepted difference between the surrounding pixels in the prediction signal and the co-located pixels in the prediction signal, and the intercepted difference between the co-located pixels in the prediction signal and the current pixel are used as the input of the ALF equation.

[0181] In the second approach, it is proposed to use spatially neighboring pixels in the residual signal as additional ALF equation inputs. Various filter shapes can be used to extract information from the residual signal. For example, the filter shape can be 1×1, 3×3, or 5×5, such as Fig.12 Various equation forms can be used to extract information from the residual signal. In one example, the interception result of the same-position pixel in the residual signal is used as the input of the ALF equation.

[0182] In the third approach, it is proposed to use spatially neighboring pixels in the signal before SAO as additional ALF equation inputs. Various filter shapes can be used to extract information in the signal before SAO. For example, the filter shape can be 1×1, 3×3, or 5×5, such as Fig.12 As shown. Various equation forms can be used to extract information in the pre-SAO signal. In one example, the intercepted differences between the surrounding pixels in the pre-SAO signal and the current pixel are used as inputs to the ALF equation. In another example, the intercepted differences between the surrounding pixels in the pre-SAO signal and the co-located pixels in the pre-SAO signal, and the intercepted differences between the co-located pixels in the pre-SAO signal and the current pixel are used as inputs to the ALF equation.

[0183] In the fourth method, it is proposed to use the information in the prediction signal, the residual signal or the signal before SAO as the input of the ALF equation. The utilization methods proposed in the first, second and third methods can be combined to implement the fourth method.

[0184] A new classifier that combines the features of edge-based classifiers and band-based classifiers

[0185] According to one or more embodiments of the present disclosure, features of edge-based classifiers and features of band-based classifiers are combined to derive a new classifier for an online ALF filter.Different approaches can be used to achieve this goal.

[0186] In the first method, it is proposed to first calculate the directionality D of the sub-block of the luminance component, then calculate the sum of the sample values ​​of the sub-block and map it to the index of the band-based classifier, and the class index of the sub-block is calculated as

[0187] C=B*M D +D (17)

[0188] Where B is the index calculated with reference to the band-based classifier, M D Denotes the total number of directivity D. In one example, for a 2×2 luma block, directivity D is calculated the same way as D2 in ECM, and B is calculated as

[0189] B = (total * 5) >> (sample bit depth + 2) (18)

[0190] In the second method, it is proposed to first calculate the activity value A of the sub-block of the luminance component, then calculate the sum of the sample values ​​of the sub-block and map it to the index of the band-based classifier, and the class index of the sub-block is calculated as

[0191] C=B*M A +A (19)

[0192] Where B is the index calculated with reference to the band-based classifier, M A Indicates the total number of activity values ​​A. In one example, for a 2×2 luminance block, the activity value A is calculated in the same way as in ECM. is the same, and B is calculated as

[0193] B = (total * 5) >> (sample bit depth + 2) (20)

[0194] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component with reference to the edge-based classifier, then calculate the sum of the sample values ​​of the sub-block and map it to the index of the band-based classifier, and the class index of the sub-block is calculated as

[0195] C=B*M E +E (21)

[0196] where B is the index calculated with reference to the band-based classifier, M E represents the total number of indices calculated with reference to the edge-based classifier, and E is the index calculated with reference to the edge-based classifier. In one example, for a 2×2 luma block, index E is calculated the same way as C2 in ECM, and B is calculated as

[0197] B = (total sum * 2) >> (sample bit depth + 2) (22)

[0198] Adjust the chroma ALF filter shape to be consistent with the luma ALF filter shape

[0199] In a third aspect of the present disclosure, it is proposed to change the chrominance ALF filter shape from a diamond to Fig.13A The long cross shown is consistent with the luminance ALF filter shape.

[0200] Input of the online ALF filter

[0201] Fig. 13B, wherein the fixed filter output samples are obtained by inputting the reconstructed samples immediately after SAO into the fixed filter trained offline. The online ALF filter may take the reconstructed samples immediately before SAO (i.e., the samples immediately before SAO) as additional input, or the predicted samples as additional input, or both the reconstructed samples immediately before SAO and the predicted samples as additional input. Fig. 13B As shown, various inputs of the online ALF filter may include not only the reconstructed samples immediately after the SAO, the fixed filter output samples, and the reconstructed samples before the DBF, but also the reconstructed samples immediately before the SAO and the predicted samples.

[0202] The filter shape applied to the prediction samples

[0203] Fig. 13C 1×1 and 3×3 filter shapes for prediction samples applied to ALF according to some examples of the present disclosure are shown. In some examples, assuming that the prediction samples are used as additional input to the online ALF filter, the filtered samples can be derived as

[0204]

[0205] Among them, R(x,y) represents the current sample point; f i,j represents the intercepted difference between the adjacent chroma sample associated with the chroma signal immediately after SAO and R(x,y); i represents the intercept difference between the fixed filter output sample and R(x,y); h i,j represents the intercepted difference between the adjacent chrominance sample associated with the chrominance signal immediately before the DBF and R(x,y). i,j is the intercepted difference between the neighboring chroma samples associated with the chroma signal (e.g., the chroma prediction signal) and the current sample R(x, y). The filter coefficient c is signaled i , i = 0, ... N. Different filter shapes can be used, for example Fig. 13C 1×1 and 3×3 rhombuses shown.

[0206] When the reconstructed samples immediately before the SAO are used as additional inputs to the online ALF filter, the prediction samples in the above equation can be directly replaced by the reconstructed samples immediately before the SAO.

[0207] New classifier using pixel values ​​before deblocking filter

[0208] According to one or more embodiments of the present disclosure, pixel values ​​before deblocking filtering are used to derive new classifiers for online ALF filtering. Different methods can be used to achieve this goal.

[0209] In the first method, it is proposed to first calculate the directionality D of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples before the deblocking filter of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0210] C=Dif*M D +D (23a)

[0211] Among them, Dif is the difference index, M D Denotes the total number of directivity D. In one example, for a 2×2 luma block, directivity D is calculated in the same way as D2 in ECM, and Dif is calculated as

[0212] Dif=sum Dif >0?2:(sum Dif <0?0:1) (24)

[0213] In the second method, it is proposed to first calculate the activity value A of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples before the deblocking filter of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0214] C=Dif*M A +A (25)

[0215] Among them, Dif is the difference index, M A Indicates the total number of activity values ​​A. In one example, for a 2×2 luminance block, the activity value A is calculated in the same way as in ECM. are the same, and Dif is calculated according to formula (24).

[0216] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component with reference to the edge-based classifier, then calculate the sum of the differences between the samples after SAO and the co-located samples before the deblocking filter of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0217] C=Dif*M E +E (26)

[0218] Where Dif is the difference index, M E represents the total number of indices calculated with reference to the edge-based classifier, and E is the index calculated with reference to the edge-based classifier. In one example, for a 2×2 luma block, the index E is calculated in the same manner as C2 in ECM, and Dif is calculated according to equation (24).

[0219] In the fourth method, it is proposed to first calculate the band index B of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples before the deblocking filter of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0220] C=Dif*M B +B (27)

[0221] Among them, Dif is the difference index, M B In one example, for a 2×2 luma block, the band index B is calculated as

[0222] B = (total sum * 8) >> (sample bit depth + 2) (28)

[0223] And Dif is calculated according to formula (24).

[0224] In the fifth method, it is proposed to calculate the sum of the differences between the samples after SAO and the co-located samples before the deblocking filter of the sub-block, and then map the sum of the differences to a difference index, and use the difference index as a class index.

[0225] In the sixth method, it is proposed to calculate an edge-based classifier or a band-based classifier based on the sample values ​​before the deblocking filter, and the calculation method is the same as the calculation method of the edge-based classifier or the band-based classifier originally calculated based on the sample values ​​after the SAO.

[0226] New classifier using pixel values ​​in prediction signal

[0227] According to one or more embodiments of the present disclosure, pixel values ​​in the prediction signal are used to derive a new classifier for an online ALF filter. Different methods can be used to achieve this goal.

[0228] In the first method, it is proposed to first calculate the directionality D of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples in the prediction signal of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0229] C=Dif*M D +D (29)

[0230] Among them, Dif is the difference index, M D Denotes the total number of directivity D. In one example, for a 2×2 luma block, directivity D is calculated in the same way as D2 in ECM, and Dif is calculated as

[0231] Dif=sum Dif >0?2:(sum Dif <0?0:1) (30)

[0232] In the second method, it is proposed to first calculate the activity value A of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples in the prediction signal of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0233] C=Dif*M A +A (31)

[0234] Among them, Dif is the difference index, M A Indicates the total number of activity values ​​A. In one example, for a 2×2 luminance block, the activity value A is calculated in the same way as in ECM. are the same, and Dif is calculated according to formula (30).

[0235] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component with reference to the edge-based classifier, then calculate the sum of the differences between the samples after SAO and the co-located samples in the prediction signal of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0236] C=Dif*M E +E (32)

[0237] Where Dif is the difference index, M E represents the total number of indices calculated with reference to the edge-based classifier, and E is the index calculated with reference to the edge-based classifier. In one example, for a 2×2 luma block, the index E is calculated in the same manner as C2 in ECM, and Dif is calculated according to equation (30).

[0238] In the fourth method, it is proposed to first calculate the band index B of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples in the prediction signal of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0239] C=Dif*M B +B (33)

[0240] Among them, Dif is the difference index, M B In one example, for a 2×2 luma block, the band index B is calculated as

[0241] B = (total * 8) >> (sample bit depth + 2) (34)

[0242] And Dif is calculated according to formula (30).

[0243] In the fifth method, it is proposed to calculate the sum of the differences between the samples after SAO and the co-located samples in the prediction signal of the sub-block, and then map the sum of the differences to a difference index, and use the difference index as a class index.

[0244] In the sixth method, it is proposed to calculate an edge-based classifier or a band-based classifier based on the sample values ​​in the prediction signal, and the calculation method is the same as the calculation method of the edge-based classifier or the band-based classifier originally calculated based on the sample values ​​after SAO.

[0245] New classifier using pixel values ​​in residual signal

[0246] According to one or more embodiments of the present disclosure, pixel values ​​in the residual signal are used to derive a new classifier for an online ALF filter. Different methods can be used to achieve this goal.

[0247] In the first method, it is proposed to first calculate the directionality D of the sub-block of the luminance component, then calculate the sum of the pixel values ​​in the residual signal of the sub-block and map it to the residual index, and the class index of the sub-block is calculated as

[0248] C=Resi*M D +D (35)

[0249] Among them, Resi is the residual index, M D Represents the total number of directivity D. In one example, for a 2×2 luma block, directivity D is calculated in the same way as D2 in ECM, and Resi is calculated as

[0250] Resi=sum Resi >0?2:(sum Resi <0?0:1) (36)

[0251] In the second method, it is proposed to first calculate the activity value A of the sub-block of the luminance component, then calculate the sum of the pixel values ​​in the residual signal of the sub-block and map it to the residual index, and the class index of the sub-block is calculated as

[0252] C=Resi*M A +A (37)

[0253] Among them, Resi is the residual index, M A Indicates the total number of activity values ​​A. In one example, for a 2×2 luminance block, the activity value A is calculated in the same way as in ECM. The same, and Resi is calculated according to formula (36).

[0254] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component with reference to the edge-based classifier, then calculate the sum of the pixel values ​​in the residual signal of the sub-block and map it to the residual index, and the class index of the sub-block is calculated as

[0255] C=Resi*M E +E (38)

[0256] Where Resi is the residual index, M E represents the total number of indices calculated with reference to the edge-based classifier, and E is the index calculated with reference to the edge-based classifier. In one example, for a 2×2 luminance block, the index E is calculated in the same manner as C2 in ECM, and Resi is calculated according to equation (36).

[0257] In the fourth method, it is proposed to first calculate the band index B of the sub-block of the luminance component, then calculate the sum of the pixel values ​​in the residual signal of the sub-block and map it to the residual index, and the class index of the sub-block is calculated as

[0258] C=Resi*M B +B (39)

[0259] Among them, Resi is the residual index, M B In one example, for a 2×2 luma block, the band index B is calculated as

[0260] B = (total sum * 8) >> (sample bit depth + 2) (40)

[0261] And Resi is calculated according to formula (36).

[0262] In the fifth method, it is proposed to calculate the sum of pixel values ​​in the residual signal of the sub-block, and then map the sum of the residual values ​​to a residual index, and use the residual index as a class index.

[0263] New classifier using pixel values ​​before SAO

[0264] According to one or more embodiments of the present disclosure, the pixel values ​​before SAO are used to derive a new classifier for the online ALF filter. Different methods can be used to achieve this goal.

[0265] In the first method, it is proposed to first calculate the directionality D of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples before SAO of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0266] C=Dif*M D +D (41)

[0267] Among them, Dif is the difference index, M D Denotes the total number of directivity D. In one example, for a 2×2 luma block, directivity D is calculated in the same way as D2 in ECM, and Dif is calculated as

[0268] Dif=sum Dif >0?2:(sum Dif <0?0:1) (42)

[0269] In the second method, it is proposed to first calculate the activity value A of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples before SAO of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0270] C=Dif*M A +A (43)

[0271] Among them, Dif is the difference index, M A Indicates the total number of activity values ​​A. In one example, for a 2×2 luminance block, the activity value A is calculated in the same way as in ECM. are the same, and Dif is calculated according to formula (42).

[0272] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component with reference to the edge-based classifier, then calculate the sum of the differences between the samples after SAO and the co-located samples before SAO of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0273] C=Dif*M E +E (44)

[0274] Where Dif is the difference index, M E represents the total number of indices calculated with reference to the edge-based classifier, and E is the index calculated with reference to the edge-based classifier. In one example, for a 2×2 luminance block, the index E is calculated in the same manner as C2 in ECM, and Dif is calculated according to equation (42).

[0275] In the fourth method, it is proposed to first calculate the band index B of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples before SAO of the sub-block and map them to the difference index, and the class index of the sub-block is calculated as

[0276] C=Dif*M B +B (45)

[0277] Among them, Dif is the difference index, M B In one example, for a 2×2 luma block, the band index B is calculated as

[0278] B = (total * 8) >> (sample bit depth + 2) (46)

[0279] And Dif is calculated according to formula (42).

[0280] In the fifth method, it is proposed to calculate the sum of the differences between the samples after SAO and the co-located samples before SAO of the sub-block, and then map the sum of the differences to a difference index, and use the difference index as a class index.

[0281] In the sixth method, it is proposed to calculate an edge-based classifier or a band-based classifier based on the sample values ​​before SAO, and the calculation method thereof is the same as the calculation method of the edge-based classifier or the band-based classifier originally calculated based on the sample values ​​after SAO.

[0282] New classifier using chrominance pixel values

[0283] According to one or more embodiments of the present disclosure, chrominance pixel values ​​are used to derive new classifiers for online ALF filters. Different methods can be used to achieve this goal.

[0284] In the first method, it is proposed to first calculate the band index B of the sub-block of the luminance component Y , and then calculate the corresponding U and V component band index B U and B V , and the class index of the sub-block is calculated as

[0285] C=B Y *M U *M V +B U *M V +B V (47)

[0286] Among them, B Y , B U and B V are the Y, U, and V indices calculated with reference to the band-based classifier, M U and M V Indicates the total number of U and V band index values. In one example, for a 2×2 luma block, B Y , B U and B V is calculated as

[0287] B Y =(sumY*6)>>(sample bit depth+2) (48)

[0288] B U =(sumU*2)>>(sample bit depth+2) (49)

[0289] B Y =(sumV*2)>>(sample bit depth+2) (50)

[0290] Chroma information before deblocking, prediction, residual, or before SAO is used as additional chroma ALF input

[0291] According to one or more embodiments of the present disclosure, chroma information before deblocking, prediction, residual, or before SAO is used as additional chroma ALF equation input. Different methods can be used to achieve this goal.

[0292] In the first approach, it is proposed to use the spatially neighboring pixels in the chrominance prediction signal as additional chrominance ALF equation inputs. Various filter shapes can be used to extract the information in the chrominance prediction signal. For example, the filter shape can be 1×1, 3×3, or 5×5, such as Fig.12 As shown. Various equation forms can be used to extract information in the chroma prediction signal. In one example, the intercepted difference between the surrounding pixels in the chroma prediction signal and the current chroma pixel is used as the input of the chroma ALF equation. In another example, the intercepted difference between the surrounding pixels in the chroma prediction signal and the co-located pixels in the chroma prediction signal, and the intercepted difference between the co-located pixels in the chroma prediction signal and the current chroma pixel are used as the input of the chroma ALF equation.

[0293] In the second approach, it is proposed to use spatially neighboring pixels in the chroma residual signal as additional chroma ALF equation inputs. Various filter shapes can be used to extract information from the chroma residual signal. For example, the filter shape can be 1×1, 3×3, or 5×5, such as Fig.12 Various equation forms can be used to extract information from the chroma residual signal. In one example, the interception result of the co-located pixels in the chroma residual signal is used as the input of the chroma ALF equation.

[0294] In the third approach, it is proposed to use spatially neighboring pixels in the signal before the chroma SAO as an additional chroma ALF equation input. Various filter shapes can be used to extract information in the signal before the chroma SAO. For example, the filter shape can be 1×1, 3×3, or 5×5, such as Fig.12 As shown. Various equation forms can be used to extract information in the signal before the chroma SAO. In one example, the intercepted difference between the surrounding pixels of the signal before the chroma SAO and the current chroma pixel is used as the input of the chroma ALF equation. In another example, the intercepted difference between the surrounding pixels in the signal before the chroma SAO and the co-located pixels in the signal before the chroma SAO, and the intercepted difference between the co-located pixels in the signal before the chroma SAO and the current chroma pixel are used as the input of the chroma ALF equation.

[0295] In the fourth method, it is proposed to use the spatially neighboring pixels in the signal before chroma deblocking as additional chroma ALF equation input. Various filter shapes can be used to extract information in the signal before chroma deblocking. For example, the filter shape can be 1×1, 3×3, or 5×5, such as Fig.12 As shown. Various equation forms can be used to extract information in the signal before chroma deblocking. In one example, the intercepted difference between the surrounding pixels of the signal before chroma deblocking and the current chroma pixel is used as the input of the chroma ALF equation. In another example, the intercepted difference between the surrounding pixels in the signal before chroma deblocking and the co-located pixels in the signal before chroma deblocking, and the intercepted difference between the co-located pixels in the signal before chroma deblocking and the current chroma pixel are used as the input of the chroma ALF equation.

[0296] In the fifth method, it is proposed to use the information in the chroma prediction, residual, pre-SAO, or pre-deblocking signal as the input of the chroma ALF equation. The utilization methods proposed in the first, second, third, and fourth methods can be combined to implement the fifth method.

[0297] Luma information before deblocking, prediction, residual, or before SAO is used as additional CCALF input

[0298] According to one or more embodiments of the present disclosure, luma information before deblocking, prediction, residual, or before SAO is used as an additional CCALF equation input.Different approaches can be used to achieve this goal.

[0299] In the first approach, it is proposed to use the spatially neighboring pixels in the luma prediction signal as additional CCALF equation inputs. Various filter shapes can be used to extract the information in the luma prediction signal. For example, the filter shape can be 3×4, such as Fig.10 As shown. Various equation forms can be used to extract information in the brightness prediction signal. In one example, the difference between the surrounding pixels in the brightness prediction signal and the currently corresponding brightness pixel is used as the input of the CCALF equation. In another example, the difference between the surrounding pixels in the brightness prediction signal and the co-located pixels in the currently corresponding brightness prediction signal, and the difference between the co-located pixels in the currently corresponding brightness prediction signal and the currently corresponding brightness pixel are used as the input of the CCALF equation.

[0300] In the second approach, it is proposed to use spatially neighboring pixels in the luma residual signal as additional CCALF equation inputs. Various filter shapes can be used to extract information from the luma residual signal. For example, the filter shape can be 3×4, such as Fig.10 Various equation forms may be used to extract information from the luma residual signal. In one or more examples, co-located pixels in the luma residual signal are used as input to the CCALF equation.

[0301] In the third approach, it is proposed to use spatially neighboring pixels in the signal before luma SAO as additional CCALF equation inputs. Various filter shapes can be used to extract information in the signal before luma SAO. For example, the filter shape can be 3×4, such as Fig.10 As shown. Various equation forms can be used to extract information in the signal before luma SAO. In one example, the difference between the surrounding pixels in the signal before luma SAO and the currently corresponding luma pixel is used as the input to the CCALF equation. In another example, the difference between the surrounding pixels in the signal before luma SAO and the co-located pixels in the currently corresponding luma pre-SAO signal, and the difference between the co-located pixels in the currently corresponding luma pre-SAO signal and the currently corresponding luma pixel are used as the input to the CCALF equation.

[0302] In the fourth method, it is proposed to use the spatially neighboring pixels in the signal before luma deblocking as additional CCALF equation inputs. Various filter shapes can be used to extract information in the signal before luma deblocking. For example, the filter shape can be 3×4, such as Fig.10 As shown. Various equation forms can be used to extract information in the signal before luma deblocking. In one example, the difference between the surrounding pixels in the signal before luma deblocking and the currently corresponding luma pixel is used as the input of the CCALF equation. In another example, the difference between the surrounding pixels in the signal before luma deblocking and the co-located pixels in the currently corresponding luma signal before deblocking, and the difference between the co-located pixels in the currently corresponding luma signal before deblocking and the currently corresponding luma pixel are used as the input of the CCALF equation.

[0303] In the fifth method, it is proposed to use information in the luma prediction, residual, signal before SAO, or before deblocking as input to the CCALF equation. The utilization methods proposed in the first, second, third and fourth methods can be combined to implement the fifth method.

[0304] New classifier using encoding pattern information

[0305] According to one or more embodiments of the present disclosure, coding mode information (such as whether a coding block is coded or decoded in skip mode, whether a coding block is coded or decoded in intra, inter P or inter B mode) is used to derive a new classifier for an online ALF filter. Different methods can be used to achieve this goal.

[0306] In the first method, it is proposed to record whether the coded block is coded or decoded in skip mode during the encoding and decoding process, and then use this information to design a new classifier. In one example, a classifier with 2 classes (corresponding to whether the skip mode is true or false) is added as a new classifier. In another example, a classifier combining the skip mode information with EO or BO is added as a new classifier.

[0307] The second method proposes to record whether the coding block is coded or decoded in intra mode, inter P mode or inter B mode during the encoding and decoding process, and then use this information to design a new classifier. In one example, a classifier with 3 classes (corresponding to intra mode, inter P mode or inter B mode, respectively) is added as a new classifier. In another example, a classifier that combines intra, inter P or inter B mode information with EO or BO is added as a new classifier.

[0308] In the third method, it is proposed to design a new classifier by considering both coding mode information (whether the coding block is coded or decoded using skip mode, and whether the coding block is coded or decoded using intra, inter P or inter B mode). The utilization methods proposed in the first and second methods can be combined to implement the third method.

[0309] Fig.14 1 is a flowchart illustrating a video decoding method 1400 according to some examples of the present disclosure. In step 1401, the method 1400 includes: obtaining, by a decoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples are from at least one of the following signals: (i) a chroma prediction signal, (ii) a chroma residual signal, (iii) a signal before chroma sample adaptive offset (SAO), or (iv) a signal before chroma deblocking. In step 1402, the method 1400 includes obtaining, by a decoder, filtered chroma samples based on the one or more spatial neighboring samples associated with the current chroma sample.

[0310] In one example, the method 1400 further includes obtaining, by the decoder, chroma samples filtered by an adaptive loop filter (ALF) based on one or more spatial neighboring samples associated with the chroma prediction signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0311] In one example, the method 1400 further includes: obtaining, by the decoder, a clipping result based on a difference between one or more spatial neighboring samples in the chroma prediction signal and the current chroma sample; and obtaining, by the decoder, a chroma ALF input based on the clipping result.

[0312] In one example, method 1400 further includes: obtaining, by the decoder, a clipping result based on the difference between surrounding samples in the chroma prediction signal and the co-located samples in the chroma prediction signal, and a clipping result based on the difference between the co-located samples in the chroma prediction signal and the current chroma sample; and obtaining, by the decoder, a chroma ALF input based on the clipping result.

[0313] In one example, the method 1400 further includes obtaining, by the decoder, chroma samples filtered by an adaptive loop filter (ALF) based on one or more spatial neighboring samples associated with the chroma residual signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0314] In an example, the method 1400 further includes: obtaining, by the decoder, a truncated result of one or more spatially adjacent samples in the chroma residual signal; and obtaining, by the decoder, a chroma ALF input based on the truncated result.

[0315] In one example, the method 1400 further includes obtaining, by the decoder, chroma samples filtered by an adaptive loop filter (ALF) based on one or more spatially neighboring samples associated with a signal before the chroma SAO and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0316] In one example, the method 1400 further includes: obtaining, by the decoder, a clipping result based on a difference between one or more spatial neighboring samples in a signal before the chroma SAO and the current chroma sample; and obtaining, by the decoder, a chroma ALF input based on the clipping result.

[0317] In one example, method 1400 further includes: obtaining, by the decoder, a clipping result based on the difference between surrounding samples in the signal before the chroma SAO and the co-located samples in the signal before the chroma SAO, and a clipping result based on the difference between the co-located samples in the signal before the chroma SAO and the current chroma sample; and obtaining, by the decoder, a chroma ALF input based on the clipping result.

[0318] In one example, method 1400 further includes obtaining, by a decoder, chroma samples filtered by an adaptive loop filter (ALF) based on one or more spatial neighboring samples associated with a signal before chroma deblocking and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0319] In one example, the method 1400 further includes: obtaining, by the decoder, a clipping result based on a difference between one or more spatial neighboring samples in the signal before chroma deblocking and the current chroma sample; and obtaining, by the decoder, a chroma ALF input based on the clipping result.

[0320] In one example, method 1400 further includes: obtaining, by the decoder, a clipping result based on the difference between one or more spatially adjacent samples in the signal before chroma deblocking and a co-located sample in the signal before chroma deblocking, and a clipping result based on the difference between the co-located sample in the signal before chroma deblocking and the current chroma sample; and obtaining, by the decoder, a chroma ALF input based on the clipping result.

[0321] In one example, the one or more surrounding samples come from a combination of: (i) a chroma prediction signal, (ii) a chroma residual signal, (iii) a signal before chroma sample adaptive offset (SAO), or (iv) a signal before chroma deblocking.

[0322] Fig.15 1 is a flowchart illustrating a video encoding method 1500 according to some examples of the present disclosure. In step 1501, the method 1500 includes: obtaining, by an encoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples are from at least one of the following signals: (i) a chroma prediction signal, (ii) a chroma residual signal, (iii) a signal before chroma sample adaptive offset (SAO), or (iv) a signal before chroma deblocking. In step 1502, the method 1500 includes obtaining, by the encoder, filtered chroma samples based on the one or more spatial neighboring samples associated with the current chroma sample.

[0323] In one example, the method 1500 further includes obtaining, by the encoder, chroma samples filtered by an adaptive loop filter (ALF) based on one or more spatial neighboring samples associated with the chroma prediction signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0324] In one example, the method 1500 further includes: obtaining, by the encoder, a clipping result based on a difference between one or more spatial neighboring samples in the chroma prediction signal and the current chroma sample; and obtaining, by the encoder, a chroma ALF input based on the clipping result.

[0325] In one example, method 1500 further includes: obtaining, by the encoder, a clipping result based on the difference between surrounding samples in the chroma prediction signal and a co-located sample in the chroma prediction signal, and a clipping result based on the difference between the co-located sample in the chroma prediction signal and the current chroma sample; and obtaining, by the encoder, a chroma ALF input based on the clipping result.

[0326] In one example, the method 1500 further includes obtaining, by the encoder, chroma samples filtered by an adaptive loop filter (ALF) based on one or more spatial neighboring samples associated with the chroma residual signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0327] In an example, the method 1500 further includes: obtaining, by the encoder, a clipping result of one or more spatially adjacent samples in the chroma residual signal; and obtaining, by the encoder, a chroma ALF input based on the clipping result.

[0328] In one example, method 1500 further includes obtaining, by the encoder, chroma samples filtered by an adaptive loop filter (ALF) based on one or more spatially neighboring samples associated with a signal before the chroma SAO and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0329] In one example, the method 1500 further includes: obtaining, by the encoder, a clipping result based on a difference between one or more spatial neighboring samples in a signal before the chroma SAO and the current chroma sample; and obtaining, by the encoder, a chroma ALF input based on the clipping result.

[0330] In one example, method 1500 further includes: obtaining, by the encoder, a clipping result based on the difference between surrounding samples in the signal before the chroma SAO and the co-located samples in the signal before the chroma SAO, and a clipping result based on the difference between the co-located samples in the signal before the chroma SAO and the current chroma sample; and obtaining, by the encoder, a chroma ALF input based on the clipping result.

[0331] In one example, method 1500 further includes: obtaining, by the encoder, chroma samples filtered by an adaptive loop filter (ALF) based on one or more spatial neighboring samples associated with the signal before chroma deblocking and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0332] In one example, the method 1500 further includes: obtaining, by the encoder, a clipping result based on a difference between one or more spatial neighboring samples in the signal before chroma deblocking and the current chroma sample; and obtaining, by the encoder, a chroma ALF input based on the clipping result.

[0333] In one example, method 1500 further includes: obtaining, by the encoder, a clipping result based on the difference between one or more spatial neighboring samples in the signal before chroma deblocking and a co-located sample in the signal before chroma deblocking, and a clipping result based on the difference between the co-located sample in the signal before chroma deblocking and the current chroma sample; and obtaining, by the encoder, a chroma ALF input based on the clipping result.

[0334] In one example, the one or more surrounding samples come from a combination of: (i) a chroma prediction signal, (ii) a chroma residual signal, (iii) a chroma signal before SAO, or (iv) a chroma signal before deblocking.

[0335] Fig.16 1 is a flowchart illustrating a video decoding method 1600 according to some examples of the present disclosure. In step 1601, the method 1600 includes: obtaining, by a decoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples are from at least one of the following signals: (i) a luma prediction signal, (ii) a luma residual signal, (iii) a signal before luma sample adaptive offset (SAO), or (iv) a signal before luma deblocking. In step 1602, the method 1600 includes obtaining, by a decoder, filtered chroma samples based on the one or more spatial neighboring samples associated with the current chroma sample.

[0336] In one example, method 1600 further includes: obtaining, by the decoder, chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on one or more spatial neighboring samples associated with the luma prediction signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0337] In one example, the method 1600 further includes: obtaining, by the decoder, a clipping result based on a difference between one or more spatially adjacent samples in the luma prediction signal and a currently corresponding luma sample; and deriving, by the decoder, a CCALF input based on the clipping result.

[0338] In one example, method 1600 further includes: obtaining, by the decoder, a clipping result based on the difference between surrounding samples in the luma prediction signal and co-located samples in the luma prediction signal, and a clipping result based on the difference between the co-located samples in the luma prediction signal and the currently corresponding luma sample; and deriving, by the decoder, a CCALF input based on the clipping result.

[0339] In one example, method 1600 further includes: obtaining, by the decoder, chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on one or more spatial neighboring samples associated with the luma residual signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0340] In an example, the method 1600 further includes: obtaining, by the decoder, a truncated result of one or more spatially adjacent samples in the luma residual signal; and deriving, by the decoder, a CCALF input based on the truncated result.

[0341] In one example, method 1600 further includes obtaining, by the decoder, chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on one or more spatially neighboring samples associated with a signal before luma SAO and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0342] In one example, the method 1600 further includes: obtaining, by the decoder, a clipping result based on a difference between one or more spatially adjacent samples in a signal before luma SAO and a currently corresponding luma sample; and deriving, by the decoder, a CCALF input based on the clipping result.

[0343] In one example, method 1600 further includes: obtaining, by the decoder, a clipping result based on the difference between surrounding samples in the signal before the luma SAO and the co-located samples in the signal before the luma SAO, and a clipping result based on the difference between the co-located samples in the signal before the luma SAO and the currently corresponding luma sample; and deriving, by the decoder, a CCALF input based on the clipping result.

[0344] In one example, method 1600 further includes: obtaining, by the decoder, chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on one or more spatial neighboring samples associated with the signal before luma deblocking and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0345] In one example, the method 1600 further includes: obtaining, by the decoder, a clipping result based on a difference between one or more spatially adjacent samples in the signal before luma deblocking and a currently corresponding luma sample; and deriving, by the decoder, a CCALF input based on the clipping result.

[0346] In one example, method 1600 further includes: obtaining, by the decoder, a clipping result based on the difference between one or more spatially adjacent samples in the signal before luma deblocking and the co-located samples in the signal before luma deblocking, and a clipping result based on the difference between the co-located samples in the signal before luma deblocking and the currently corresponding luma samples; and deriving, by the decoder, a CCALF input based on the clipping result.

[0347] In one example, the neighboring samples come from a combination of: (i) a luma prediction signal; (ii) a luma residual signal; (iii) a signal before luma SAO; or (iv) a signal before luma deblocking.

[0348] Fig.17 1 is a flowchart illustrating a video encoding method 1700 according to some examples of the present disclosure. In step 1701, the method 1700 includes: obtaining, by an encoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples are from at least one of the following signals: (i) a luma prediction signal, (ii) a luma residual signal, (iii) a signal before luma sample adaptive offset (SAO), or (iv) a signal before luma deblocking. In step 1702, the method 1700 includes obtaining, by the encoder, filtered chroma samples based on the one or more spatial neighboring samples associated with the current chroma sample.

[0349] In one example, method 1700 further includes: obtaining, by the encoder, chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on one or more spatial neighboring samples associated with the luma prediction signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0350] In one example, the method 1700 further includes: obtaining, by the encoder, a clipping result based on a difference between one or more spatial neighboring samples in the luma prediction signal and a currently corresponding luma sample; and deriving, by the encoder, a CCALF input based on the clipping result.

[0351] In one example, method 1700 further includes: obtaining, by the encoder, a clipping result based on the difference between surrounding samples in the luma prediction signal and co-located samples in the luma prediction signal, and a clipping result based on the difference between the co-located samples in the luma prediction signal and the currently corresponding luma sample; and deriving, by the encoder, a CCALF input based on the clipping result.

[0352] In one example, method 1700 further includes obtaining, by the encoder, chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on one or more spatially neighboring samples associated with the luma residual signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0353] In one example, the method 1700 further includes: obtaining, by the encoder, a truncated result of one or more spatially adjacent samples in the luma residual signal; and deriving, by the encoder, a CCALF input based on the truncated result.

[0354] In one example, method 1700 further includes obtaining, by the encoder, chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on one or more spatially neighboring samples associated with a signal before luma SAO and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0355] In one example, the method 1700 further includes: obtaining, by the encoder, a clipping result based on a difference between one or more spatially adjacent samples in a signal before luma SAO and a currently corresponding luma sample; and deriving, by the encoder, a CCALF input based on the clipping result.

[0356] In one example, method 1700 further includes: obtaining, by the encoder, a clipping result based on the difference between surrounding samples in the signal before the luma SAO and the co-located samples in the signal before the luma SAO, and a clipping result based on the difference between the co-located samples in the signal before the luma SAO and the currently corresponding luma sample; and deriving, by the encoder, a CCALF input based on the clipping result.

[0357] In one example, method 1700 further includes: obtaining, by the encoder, chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on one or more spatial neighboring samples associated with the signal before luma deblocking and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

[0358] In one example, the method 1700 further includes: obtaining, by the encoder, a clipping result based on a difference between one or more spatial neighboring samples in the signal before luma deblocking and a currently corresponding luma sample; and deriving, by the encoder, a CCALF input based on the clipping result.

[0359] In one example, method 1700 further includes: obtaining, by the encoder, a clipping result based on the difference between one or more spatially adjacent samples in the signal before luma deblocking and the co-located samples in the signal before luma deblocking, and a clipping result based on the difference between the co-located samples in the signal before luma deblocking and the currently corresponding luma samples; and deriving, by the encoder, a CCALF input based on the clipping result.

[0360] In one example, the one or more spatial neighboring samples are from a combination of: (i) a luma prediction signal; (ii) a luma residual signal; (iii) a signal before luma SAO; or (iv) a signal before luma deblocking.

[0361] Fig.18 1 is a flowchart illustrating a video decoding method 1800 according to some examples of the present disclosure. In step 1801, the method 1800 includes obtaining, by a decoder, encoding information associated with a coding block, wherein the encoding information includes: a first flag indicating that the coding block is encoded and decoded in a skip mode and a second flag indicating that the coding block is encoded and decoded in at least one of the following modes: intra mode, inter P mode, or inter B mode, to derive a new classifier for an online adaptive loop filter (ALF) process. In step 1802, the method 1800 includes generating, by a decoder, a new classifier for an online adaptive ALF process based on the encoding information.

[0362] In one example, the method 1800 further includes using the first marker to derive a new classifier for the online ALF process.

[0363] In one example, the new classifier includes two classes, corresponding to whether the coding block is encoded or decoded in skip mode.

[0364] In one example, the new classifier combines the skip mode information with at least one of: edge offset (EO) (edge-based classifier) ​​information or band offset (BO) (band-based classifier) ​​information.

[0365] In one example, method 1800 further includes: recording at a decoder that a coding block is encoded and decoded using at least one of the following modes: (i) intra mode, (ii) inter P mode, or (iii) inter B mode; and generating at the decoder a new classifier for an online ALF process based on the record.

[0366] In an example, the new classifier includes three classes, corresponding to whether the coding block is encoded or decoded in intra mode, inter P mode, and inter B mode, respectively.

[0367] In one example, the new classifier combines information that a coding block is encoded or decoded in at least one of the following modes: intra mode, inter P mode, or inter B mode, with at least one of the following: edge offset (EO) (edge-based classifier) ​​information or band offset (BO) (band-based classifier) ​​information.

[0368] In one example, method 1800 further includes: determining at a decoder whether a coding block is encoded or decoded using a skip mode; determining at a decoder whether a coding block is encoded or decoded using one of the following modes: (i) intra-frame mode; (ii) inter-frame P mode; or (iii) inter-frame B mode; and generating, at the decoder, a new classifier based on the determined mode.

[0369] Fig.19 1 is a flowchart illustrating a video encoding method 1900 according to some examples of the present disclosure. In step 1901, the method 1900 includes obtaining, by an encoder, encoding information associated with a coding block, wherein the encoding information includes information whether the coding block is encoded and decoded in a skip mode and information that the coding block is encoded and decoded in at least one of the following modes: intra mode, inter P mode, or inter B mode, to derive a new classifier for an online adaptive loop filter (ALF) process. In step 1902, the method 1900 includes generating, by the encoder, a new classifier for an online ALF process based on the encoding information.

[0370] In one example, encoding and decoding of coded blocks in skip mode during the encoding process is used to derive a new classifier for the online ALF process.

[0371] In one example, the new classifier includes two classes, corresponding to whether the coding block is encoded or decoded in skip mode.

[0372] In one example, the new classifier combines the skip mode information with edge offset (EO) (edge-based classifier) ​​information or band offset (BO) (band-based classifier) ​​information.

[0373] In one example, method 1900 includes recording at an encoder that a coding block is encoded or decoded using at least one of the following modes: (i) intra mode, (ii) inter P mode, or (iii) inter B mode; and generating, by the encoder, a new classifier for an online ALF process based on the record.

[0374] In an example, the new classifier includes three classes, corresponding to whether the coding block is encoded or decoded in intra mode, inter P mode, and inter B mode, respectively.

[0375] In one example, the new classifier combines information that a coding block is encoded or decoded in at least one of the following modes: intra mode, inter P mode, or inter B mode, with at least one of the following: edge offset (EO) (edge-based classifier) ​​information or (BO) (band-based classifier) ​​information.

[0376] In one example, method 1900 further includes: determining at the encoder whether the coding block is encoded or decoded using a skip mode; determining at the encoder whether the coding block is encoded or decoded using one of the following modes: (i) intra mode; (ii) inter P mode; or (iv) inter B mode; and generating a new classifier based on the mode.

[0377] Fig. 20 A computing environment 2010 is shown coupled to a user interface 2050. The computing environment 2010 may be part of a data processing server. The computing environment 2010 includes a processor 2020, a memory 2030, and an input / output (I / O) interface 2040.

[0378] The processor 2020 generally controls the overall operation of the computing environment 2010, such as operations associated with display, data acquisition, data communication, and image processing. The processor 2020 may include one or more processors for executing instructions to perform all or some steps in the above method. In addition, the processor 2020 may include one or more modules that facilitate interaction between the processor 2020 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, a graphics processing unit (GPU), etc.

[0379] The memory 2030 is configured to store various types of data to support the operation of the computing environment 2010. The memory 2030 may include predetermined software 2032. Examples of the above data include instructions for any application or method operating on the computing environment 2010, video data sets, image data, etc. The memory 2030 may be implemented by using any type of volatile or non-volatile memory device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk.

[0380] In one example, the memory 2030 is configured to store instructions that can be executed by the processor. When the processor executes the instructions, it is configured to execute Figure 14-19 Any of the methods described in .

[0381] The I / O interface 2040 provides an interface between the processor 2020 and peripheral interface modules (e.g., keyboard, click wheel, buttons, etc.). The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 2040 may be coupled to an encoder and a decoder.

[0382] In an embodiment, a non-transitory computer-readable storage medium including a plurality of programs (e.g., a plurality of programs in the memory 2030) is also provided, and the plurality of programs can be executed by the processor 2020 in the computing environment 2010 to perform the above method. Alternatively, the non-transitory computer-readable storage medium may store a program generated by an encoder (e.g., Figure 2 The video encoder 20 in FIG. 1 generates a video signal for a decoder (eg, Figure 3 A bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or associated one or more syntax elements, etc.) used by the video decoder 30 in the video decoder 30 when decoding video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0383] In an embodiment, a computing device is also provided, which includes: one or more processors (e.g., processor 2020); and a non-transitory computer-readable storage medium or memory 2030 in which multiple programs that can be executed by the one or more processors are stored, wherein the one or more processors are configured to perform the above-mentioned method when executing the multiple programs.

[0384] In an embodiment, a computer program product including multiple programs (e.g., multiple programs in memory 2030) is also provided, and the multiple programs can be executed by the processor 2020 in the computing environment 2010 to perform the above method. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0385] In an embodiment, the computing environment 2010 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.

[0386] The description of the present disclosure is presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.

[0387] Unless otherwise specifically stated, the order of steps of the method according to the present disclosure is intended to be illustrative only, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but can be changed according to actual conditions. In addition, at least one of the steps of the method according to the present disclosure can be adjusted, combined or deleted according to actual needs.

[0388] The examples are chosen and described in order to explain the principles of the present disclosure and to enable other persons skilled in the art to understand the various embodiments of the present disclosure and to best utilize the basic principles and various embodiments with various modifications suitable for the specific use contemplated. Therefore, it should be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed embodiments, and modifications and other embodiments are intended to be included within the scope of the present disclosure.

[0389] The above method can be implemented using an apparatus including one or more circuits, including an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor or other electronic components. The apparatus can use the circuit in combination with other hardware or software components to perform the above method. Each module, submodule, unit or subunit disclosed above can be implemented at least in part using the one or more circuits.

[0390] By considering the specification and practice disclosed herein, other embodiments of the present disclosure will be apparent to those skilled in the art. The application is intended to cover any variation, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes the difference between the present disclosure and known or customary practice in the art. The specification and embodiments should be considered as examples only. The specification and embodiments are considered to be exemplary. The application is intended to cover any variation, use or adaptation of the present disclosure.

[0391] It should be understood that the present disclosure is not limited to the specific embodiments described above and the corresponding drawings, and various modifications and changes may be made without departing from the scope thereof.

Claims

1. A video decoding method, comprising: Obtaining, by a decoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples are from at least one of the following signals: (i) a chroma prediction signal, (ii) a chroma residual signal, (iii) a signal before chroma sample adaptive offset (SAO), or (iv) a signal before chroma deblocking; and A filtered chroma sample is obtained by the decoder based on the one or more spatial neighboring samples associated with the current chroma sample.

2. The video decoding method according to claim 1, further comprising: An adaptive loop filter (ALF) filtered chroma sample is obtained by the decoder based on the one or more spatial neighboring samples associated with the chroma prediction signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

3. The video decoding method according to claim 2, further comprising: Obtaining, by the decoder, a clipping result based on a difference between one or more spatially adjacent samples in the chrominance prediction signal and the current chrominance sample; as well as The decoder obtains a chroma ALF input based on the interception result.

4. The video decoding method according to claim 2, further comprising: The decoder obtains a clipping result based on a difference between surrounding samples in the chroma prediction signal and a co-located sample in the chroma prediction signal, and a clipping result based on a difference between the co-located sample in the chroma prediction signal and a current chroma sample; as well as The decoder obtains a chroma ALF input based on the interception result.

5. The video decoding method according to claim 1, further comprising: An adaptive loop filter (ALF) filtered chroma sample is obtained by the decoder based on the one or more spatial neighboring samples associated with the chroma residual signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

6. The video decoding method according to claim 5, further comprising: The decoder obtains a truncation result of one or more spatially adjacent sample points in the chrominance residual signal; as well as The decoder obtains a chroma ALF input based on the interception result.

7. The video decoding method according to claim 1, further comprising: Chroma samples filtered by an adaptive loop filter (ALF) are obtained by the decoder based on the one or more spatially neighboring samples associated with the signal before the chroma SAO and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

8. The video decoding method according to claim 7, further comprising: Obtaining, by the decoder, a clipping result based on a difference between one or more spatial neighboring samples in the signal before the chroma SAO and the current chroma sample; as well as The decoder obtains a chroma ALF input based on the interception result.

9. The video decoding method according to claim 7, further comprising: The decoder obtains a clipping result based on a difference between surrounding samples in the signal before the chroma SAO and a co-located sample in the signal before the chroma SAO, and a clipping result based on a difference between the co-located sample in the signal before the chroma SAO and a current chroma sample; as well as The decoder obtains a chroma ALF input based on the interception result.

10. The video decoding method according to claim 1, further comprising: Chroma samples filtered by an adaptive loop filter (ALF) are obtained by the decoder based on the one or more spatial neighboring samples associated with the signal before chroma deblocking and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

11. The video decoding method according to claim 10, further comprising: Obtaining, by the decoder, a clipping result based on a difference between one or more spatial neighboring samples in the signal before chroma deblocking and the current chroma sample; as well as The decoder obtains a chroma ALF input based on the interception result.

12. The video decoding method according to claim 10, further comprising: The decoder obtains a clipping result based on a difference between one or more spatially adjacent samples in the signal before chroma deblocking and a co-located sample in the signal before chroma deblocking, and a clipping result based on a difference between a co-located sample in the signal before chroma deblocking and a current chroma sample; as well as The decoder obtains a chroma ALF input based on the interception result.

13. The video decoding method according to claim 1, wherein: The one or more surrounding samples are from a combination of: (i) the chroma prediction signal, (ii) the chroma residual signal, (iii) the signal before chroma sample adaptive offset (SAO), or (iv) the signal before chroma deblocking.

14. A video encoding method, comprising: Obtaining, by an encoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples are from chroma samples, wherein the one or more spatial neighboring samples are from at least one of the following signals: (i) a chroma prediction signal, (ii) a chroma residual signal, (iii) a signal before chroma sample adaptive offset (SAO), or (iv) a signal before chroma deblocking; and A filtered chroma sample is obtained by the encoder based on the one or more spatial neighboring samples associated with the current chroma sample.

15. The video encoding method according to claim 14, further comprising: Chroma samples filtered by an adaptive loop filter (ALF) are obtained by the encoder based on the one or more spatial neighboring samples associated with the chroma prediction signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

16. The video encoding method according to claim 15, further comprising: Obtaining, by the encoder, a clipping result based on a difference between one or more spatially adjacent samples in the chrominance prediction signal and the current chrominance sample; as well as The encoder obtains a chroma ALF input based on the clipped result.

17. The video encoding method according to claim 15, further comprising: The encoder obtains a clipping result based on a difference between surrounding samples in the chroma prediction signal and a co-located sample in the chroma prediction signal, and a clipping result based on a difference between the co-located sample in the chroma prediction signal and a current chroma sample; as well as The encoder obtains a chroma ALF input based on the clipped result.

18. The video encoding method according to claim 14, further comprising: An adaptive loop filter (ALF) filtered chroma sample is obtained by the encoder based on the one or more spatial neighboring samples associated with the chroma residual signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

19. The video encoding method according to claim 18, further comprising: The encoder obtains a truncation result of one or more spatially adjacent sample points in the chroma residual signal; as well as The encoder obtains a chroma ALF input based on the clipped result.

20. The video encoding method of claim 14, further comprising: Chroma samples filtered by an adaptive loop filter (ALF) are obtained by the encoder based on the one or more spatial neighboring samples associated with the signal before the chroma SAO and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

21. The video encoding method of claim 20, further comprising: Obtaining, by the encoder, a clipping result based on a difference between one or more spatial neighboring samples in the signal before the chroma SAO and the current chroma sample; as well as The encoder obtains a chroma ALF input based on the clipped result.

22. The video encoding method of claim 20, further comprising: The encoder obtains a clipping result based on a difference between surrounding samples in the signal before the chroma SAO and a co-located sample in the signal before the chroma SAO, and a clipping result based on a difference between the co-located sample in the signal before the chroma SAO and a current chroma sample; as well as The encoder obtains a chroma ALF input based on the clipped result.

23. The video encoding method of claim 14, further comprising: Chroma samples filtered by an adaptive loop filter (ALF) are obtained by the encoder based on the one or more spatial neighboring samples associated with the signal before chroma deblocking and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

24. The video encoding method of claim 23, further comprising: Obtaining, by the encoder, a clipping result based on a difference between one or more spatial neighboring samples in the signal before chroma deblocking and the current chroma sample; as well as The encoder obtains a chroma ALF input based on the clipped result.

25. The video encoding method of claim 23, further comprising: The encoder obtains a clipping result based on a difference between one or more spatially adjacent samples in the signal before chroma deblocking and a co-located sample in the signal before chroma deblocking, and a clipping result based on a difference between a co-located sample in the signal before chroma deblocking and a current chroma sample; as well as The encoder obtains a chroma ALF input based on the clipped result.

26. The video encoding method according to claim 14, wherein: The one or more surrounding samples are from a combination of: (i) the chroma prediction signal, (ii) the chroma residual signal, (iii) the chroma signal before SAO, or (iv) the chroma signal before deblocking.

27. A video decoding method, comprising: Obtaining, by a decoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples are from at least one of the following signals: (i) a luma prediction signal, (ii) a luma residual signal, (iii) a signal before luma sample adaptive offset (SAO), or (iv) a signal before luma deblocking; and A filtered chroma sample is obtained by the decoder based on the one or more spatial neighboring samples associated with the current chroma sample.

28. The video decoding method of claim 27, further comprising: A cross-component adaptive loop filter (CCALF) filtered chroma sample is obtained by the decoder based on the one or more spatial neighboring samples associated with the luma prediction signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

29. The video decoding method of claim 28, further comprising: The decoder obtains a clipping result based on a difference between one or more spatially adjacent samples in the luma prediction signal and the currently corresponding luma sample; as well as The CCALF input is derived by the decoder based on the interception result.

30. The video decoding method of claim 28, further comprising: The decoder obtains a clipping result based on a difference between surrounding samples in the luma prediction signal and a co-located sample in the luma prediction signal, and a clipping result based on a difference between a co-located sample in the luma prediction signal and a currently corresponding luma sample; as well as The CCALF input is derived by the decoder based on the interception result.

31. The video decoding method according to claim 27, comprising: A cross-component adaptive loop filter (CCALF) filtered chroma sample is obtained by the decoder based on the one or more spatial neighboring samples associated with the luma residual signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

32. The video decoding method of claim 31 , further comprising: The decoder obtains the truncation result of one or more spatially adjacent sample points in the luminance residual signal; as well as The CCALF input is derived by the decoder based on the interception result.

33. The video decoding method according to claim 27, comprising: The decoder obtains chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on the one or more spatially neighboring samples associated with the signal before the luma SAO and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

34. The video decoding method of claim 33, further comprising: Obtaining, by the decoder, a clipping result based on a difference between one or more spatially adjacent samples in the signal before the luma SAO and the currently corresponding luma sample; as well as The CCALF input is derived by the decoder based on the interception result.

35. The video decoding method of claim 33, further comprising: The decoder obtains a clipping result based on a difference between surrounding samples in the signal before the luma SAO and a co-located sample in the signal before the luma SAO, and a clipping result based on a difference between the co-located sample in the signal before the luma SAO and a currently corresponding luma sample; as well as The CCALF input is derived by the decoder based on the interception result.

36. The video decoding method according to claim 27, comprising: A cross-component adaptive loop filter (CCALF) filtered chroma sample is obtained by the decoder based on the one or more spatial neighboring samples associated with the signal before luma deblocking and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

37. The video decoding method of claim 36, further comprising: The decoder obtains a clipping result based on a difference between one or more spatial neighboring samples in the signal before luma deblocking and the currently corresponding luma sample; as well as The CCALF input is derived by the decoder based on the interception result.

38. The video decoding method of claim 36, further comprising: The decoder obtains a clipping result based on a difference between one or more spatially adjacent samples in the signal before luma deblocking and a co-located sample in the signal before luma deblocking, and a clipping result based on a difference between a co-located sample in the signal before luma deblocking and a currently corresponding luma sample; as well as The CCALF input is derived by the decoder based on the interception result.

39. The video decoding method according to claim 27, wherein: The one or more spatial neighboring samples are from a combination of the following signals: (i) the luma prediction signal; (ii) the luma residual signal; (iii) the luma signal before SAO; or (iv) the luma signal before deblocking.

40. A video encoding method, comprising: Obtaining, by an encoder, one or more spatial neighboring samples associated with a current chroma sample, wherein the one or more spatial neighboring samples are from at least one of the following signals: (i) a luma prediction signal, (ii) a luma residual signal, (iii) a signal before luma sample adaptive offset (SAO), or (iv) a signal before luma deblocking; and A filtered chroma sample is obtained by the encoder based on the one or more spatial neighboring samples associated with the current chroma sample.

41. The video encoding method of claim 40, further comprising: A cross-component adaptive loop filter (CCALF) filtered chroma sample is obtained by the encoder based on the one or more spatial neighboring samples associated with the luma prediction signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

42. The video encoding method of claim 41, further comprising: The encoder obtains a clipping result based on a difference between one or more spatially adjacent samples in the luma prediction signal and the currently corresponding luma sample; as well as The CCALF input is derived by the encoder based on the truncated result.

43. The video encoding method of claim 41, further comprising: The encoder obtains a clipping result based on a difference between surrounding samples in the luma prediction signal and a co-located sample in the luma prediction signal, and a clipping result based on a difference between a co-located sample in the luma prediction signal and a currently corresponding luma sample; as well as The CCALF input is derived by the encoder based on the truncated result.

44. The video encoding method of claim 40, further comprising: A cross-component adaptive loop filter (CCALF) filtered chroma sample is obtained by the encoder based on the one or more spatial neighboring samples associated with the luma residual signal and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

45. The video encoding method of claim 44, further comprising: The encoder obtains a truncation result of one or more spatially adjacent sample points in the luminance residual signal; as well as The CCALF input is derived by the encoder based on the truncated result.

46. ​​The video encoding method of claim 40, comprising: The encoder obtains chroma samples filtered by a cross-component adaptive loop filter (CCALF) based on the one or more spatially neighboring samples associated with the signal before the luma SAO and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

47. The video encoding method of claim 46, further comprising: The encoder obtains a clipping result based on a difference between one or more spatially adjacent samples in the signal before the luma SAO and the currently corresponding luma sample; as well as The CCALF input is derived by the encoder based on the truncated result.

48. The video encoding method of claim 46, further comprising: The encoder obtains a clipping result based on a difference between surrounding samples in the signal before the luma SAO and a co-located sample in the signal before the luma SAO, and a clipping result based on a difference between the co-located sample in the signal before the luma SAO and a currently corresponding luma sample; as well as The CCALF input is derived by the encoder based on the truncated result.

49. The video encoding method of claim 40, further comprising: Chroma samples filtered by a cross-component adaptive loop filter (CCALF) are obtained by the encoder based on the one or more spatial neighboring samples associated with the signal before luma deblocking and one or more filter coefficients, wherein the one or more filter coefficients are associated with different filter shapes.

50. The video encoding method of claim 49, further comprising: The encoder obtains a clipping result based on a difference between one or more spatial neighboring samples in the signal before luma deblocking and the currently corresponding luma sample; as well as The CCALF input is derived by the encoder based on the truncated result.

51. The video encoding method of claim 49, further comprising: The encoder obtains a clipping result based on a difference between one or more spatially adjacent samples in the signal before luma deblocking and a co-located sample in the signal before luma deblocking, and a clipping result based on a difference between a co-located sample in the signal before luma deblocking and a currently corresponding luma sample; as well as The CCALF input is derived by the encoder based on the truncated result.

52. The video encoding method of claim 40, wherein: The one or more spatial neighboring samples are from a combination of the following signals: (i) the luma prediction signal; (ii) the luma residual signal; (iii) the luma signal before SAO; or (iv) the luma signal before deblocking.

53. A video decoding method, comprising: Obtaining, by a decoder, encoding information associated with a coding block, wherein the encoding information includes: a first flag indicating that the coding block is coded and decoded in a skip mode and a second flag indicating that the coding block is coded and decoded in at least one of the following modes: an intra mode, an inter P mode, or an inter B mode, so as to derive a new classifier for an online adaptive loop filter (ALF) process; and A new classifier for the online adaptive ALF process is generated by the decoder based on the encoding information.

54. The video decoding method of claim 53, further comprising: The first marker is used to derive the new classifier for the online ALF process.

55. The video decoding method according to claim 54, wherein: The new classifier includes two classes, corresponding to whether the coding block adopts the skip mode for encoding and decoding.

56. The video decoding method according to claim 54, wherein: The new classifier combines the skip mode information with at least one of: edge offset (EO) (edge-based classifier) ​​information or band offset (BO) (band-based classifier) ​​information.

57. The video decoding method of claim 53, further comprising: Recording at the decoder that the coding block is coded and decoded using at least one of the following modes: (i) intra mode, (ii) inter P mode, or (iii) inter B mode; as well as The new classifier for the online ALF process is generated at the decoder based on the records.

58. The video decoding method according to claim 57, wherein: The new classifier includes three classes, which respectively correspond to the coding and decoding of the coding block using intra mode, inter P mode and inter B mode.

59. The video decoding method according to claim 57, wherein: The new classifier combines the information that the coding block is encoded and decoded in at least one of the following modes: intra-frame mode, inter-frame P mode or inter-frame B mode, with at least one of the following items: edge offset (EO) (edge-based classifier) ​​information or band offset (BO) (band-based classifier) ​​information.

60. The video decoding method of claim 53, further comprising: Determining at the decoder whether the coding block is encoded and decoded using a skip mode; Determining at the decoder that the coding block is coded or decoded in one of the following modes: (i) intra mode; (ii) inter mode (P); or (iii) inter mode (B); as well as At the decoder, the new classifier is generated based on the determined pattern.

61. A video encoding method, comprising: Obtaining, by an encoder, encoding information associated with a coding block, wherein the encoding information includes information on whether the coding block is coded and decoded in a skip mode and information on whether the coding block is coded and decoded in at least one of the following modes: an intra mode, an inter P mode, or an inter B mode, so as to derive a new classifier for an online adaptive loop filter (ALF) process; and A new classifier for the online ALF process is generated by the encoder based on the encoding information.

62. The video encoding method of claim 61, wherein: Whether the coding block is coded or decoded using the skip mode during the encoding process is used to derive the new classifier for the online ALF process.

63. The video encoding method of claim 62, wherein: The new classifier includes two classes, corresponding to whether the coding block adopts the skip mode for encoding and decoding.

64. The video encoding method of claim 62, wherein: The new classifier combines the skip mode information with the edge offset (EO) (edge-based classifier) ​​information or the band offset (BO) (band-based classifier) ​​information.

65. The video encoding method of claim 61, further comprising: Recording at the encoder that the coding block is coded and decoded using at least one of the following modes: (i) intra mode, (ii) inter P mode, or (iii) inter B mode; as well as The new classifier for the online ALF process is generated by the encoder based on the records.

66. The video encoding method of claim 65, wherein: The new classifier includes three classes, which respectively correspond to the coding and decoding of the coding block using intra mode, inter P mode and inter B mode.

67. The video encoding method of claim 65, wherein: The new classifier combines the information that the coding block is encoded and decoded in at least one of the following modes: intra-frame mode, inter-frame P mode or inter-frame B mode, with at least one of the following items: edge offset (EO) (edge-based classifier) ​​information or (BO) (band-based classifier) ​​information.

68. The video encoding method of claim 65, further comprising: Determining at the encoder whether the coding block is encoded and decoded using a skip mode; Determining, at the encoder, that the coding block is coded or decoded in one of the following modes: (i) intra mode; (ii) inter mode (P); or (iv) inter mode (B); as well as The new classifier is generated based on the pattern.

69. An apparatus for video decoding, comprising: one or more processors; as well as A memory coupled to the one or more processors, the memory being configured to store instructions executable by the one or more processors, wherein the one or more processors, when executing the instructions, are configured to perform the method as described in any one of claims 1-13, 27-39 and 53-60.

70. An apparatus for video encoding, the apparatus comprising: one or more processors; as well as A memory coupled to the one or more processors, the memory being configured to store instructions executable by the one or more processors, wherein the one or more processors, when executing the instructions, are configured to perform the method as described in any one of claims 14-26, 40-52, and 61-68.

71. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to receive a bitstream and perform the method of any one of claims 1-13, 27-39, and 53-60.

72. A non-transitory computer-readable storage medium storing computer-executable instructions which, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 14-26, 40-52, and 61-68 to generate and transmit a bitstream.

73. A bit stream to be decoded by a decoding method as claimed in any one of claims 1 to 13, 27, 39 and 53 to 60.

74. A bit stream generated by the encoding method as described in any one of claims 14-26, 40-52 and 61-68.