Video encoding and decoding method, equipment and medium
By applying a fixed filter to filter video data during the video encoding and decoding process, the problem of low encoding and decoding efficiency of ALF and CCALF in existing technologies is solved, achieving more efficient video quality and compression effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-03-27
AI Technical Summary
Existing video encoding and decoding technologies have low encoding and decoding efficiency when using adaptive loop filters (ALF) and cross-component adaptive loop filters (CCALF), making it difficult to effectively improve video quality and compression efficiency.
By applying fixed filters in the loop filter, filtering accuracy and efficiency can be improved for video data at different stages. This includes using Sample Adaptive Offset (SAO) filters, Cross-Component Sample Adaptive Offset (CCSAO) filters, and Adaptive Loop Filters (ALF) to optimize the encoding and decoding process of video blocks.
It improves the efficiency and quality of video encoding and decoding, reduces blockiness and blurring, and enhances the compression performance of video data.
Smart Images

Figure CN121750859A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is filed and claims priority to international application No. PCT / CN2024 / 121935, filed on September 27, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to video encoding / decoding and compression. More specifically, this application relates to methods and apparatus for improving the encoding / decoding efficiency of Adaptive Loop Filter (ALF) and Cross-Component Adaptive Loop Filter (CCALF). Background Technology
[0004] Various electronic devices (such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc.) support digital video. Electronic devices transmit and receive, or otherwise transmit, digital video data via communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video data can be compressed using one or more video codec standards before it is transmitted or stored. Examples of video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. Video codecs typically employ prediction methods that utilize the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs aim to compress video data to a form using a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention
[0005] Embodiments of this disclosure provide methods and apparatus for improving the encoding and decoding efficiency of adaptive loop filters (ALF) and cross-component adaptive loop filters (CCALF).
[0006] According to one aspect of this disclosure, a video decoding method is provided, comprising: determining at least one fixed filter to be applied in a loop filter, wherein the loop filter is associated with at least one stage, and the at least one fixed filter is applied to the same stage or different stages of the at least one stage; and for each of the at least one fixed filter, filtering reconstructed samples using the fixed filter to obtain filtered samples.
[0007] According to an aspect of the present disclosure, a video encoding method is provided, including: determining at least one fixed filter applied in a loop filter, wherein the loop filter is associated with at least one stage, the at least one fixed filter is applied to a same stage or different stages among the at least one stage; and for each fixed filter in the at least one fixed filter, filtering a reconstructed sample with the fixed filter to obtain a filtered sample.
[0008] According to an aspect of the present disclosure, a computing device is provided, including: one or more processors; and a memory coupled with the one or more processors, wherein the memory is configured to store instructions executable by the one or more processors, the one or more processors, when executing the instructions, causing the computing device to perform the method of any of the above aspects.
[0009] According to an aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, storing a bitstream generated by instructions, the instructions, when executed by a computing device having one or more processors, causing the one or more processors to perform the video encoding method described above.
[0010] According to an aspect of the present disclosure, a method of storing a bitstream is provided, including: generating a bitstream according to the video encoding method described above; and storing the bitstream.
[0011] According to an aspect of the present disclosure, a computer program product is provided, including instructions, the instructions, when executed by one or more processors of a computing device, causing the computing device to perform the method of any of the above aspects.
[0012] It should be understood that the foregoing general description and the following detailed description are only examples and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings, which are incorporated in and form a part of the specification, illustrate examples consistent with the present disclosure and serve to explain the principles of the present disclosure.
[0014] FIG. 1 is a block diagram illustrating an exemplary system for encoding and decoding video blocks, in accordance with some embodiments of the present disclosure.
[0015] FIG. 2 is a block diagram illustrating an exemplary video encoder, in accordance with some embodiments of the present disclosure.
[0016] FIG. 3 is a block diagram illustrating an exemplary video decoder, in accordance with some embodiments of the present disclosure.
[0017] FIG. 4A to FIG. 4Eis a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure.
[0018] FIG. 5 is a block diagram illustrating the overall of a block-based video encoder for VVC / AVS3 according to some embodiments of the present disclosure.
[0019] FIG. 6 is a block diagram illustrating block partitioning in a multi-type tree structure according to some embodiments of the present disclosure.
[0020] FIG. 7 is a block diagram illustrating the overall of a video decoder for VVC according to some embodiments of the present disclosure.
[0021] FIG. 8 is a diagram illustrating ALF filter shapes in VVC according to some embodiments of the present disclosure.
[0022] FIG. 9 is a diagram illustrating sub-sampled sample gradients for 4x4 subblock ALF classification according to some embodiments of the present disclosure.
[0023] FIG. 10 is a diagram illustrating geometry transform for 7x7 diamond filter shapes according to some embodiments of the present disclosure.
[0024] FIG. 11 is a diagram illustrating online filter shapes used in ECM according to some embodiments of the present disclosure.
[0025] FIG. 12 is a diagram illustrating CCALF architecture according to some embodiments of the present disclosure.
[0026] FIG. 13 is a diagram illustrating filtered chroma samples and their supported relative positions in the luma plane for 4:2:0 chroma format for chroma location type 0 according to some embodiments of the present disclosure.
[0027] FIG. 14 is a diagram illustrating 25-tap long filter according to some embodiments of the present disclosure.
[0028] FIG. 15 is a diagram illustrating filter shapes for a signal before prediction or SAO according to some embodiments of the present disclosure.
[0029] FIG. 16 is a diagram illustrating a computing environment coupled with a user interface according to some embodiments of the present disclosure.
[0030] FIG. 17 is a diagram illustrating an adjusted chroma ALF filter shape according to some embodiments of the present disclosure.
[0031] FIG. 18 is a diagram illustrating symmetric sample padding for luma ALF filtering when the filter shape of the residual signal (with its center position aligned with the sample to be filtered) crosses the line buffer boundary according to some embodiments of the present disclosure.
[0032] FIG. 19 is a diagram illustrating an ALF filter shape according to some embodiments of the present disclosure.
[0033] FIG. 20 is a diagram illustrating an online chroma ALF filter shape according to some embodiments of the present disclosure.
[0034] FIG. 21 is a diagram illustrating a first example of an adjusted online chroma ALF filter shape according to some embodiments of the present disclosure.
[0035] FIG. 22 is a diagram illustrating a second example of an adjusted online chroma ALF filter shape according to some embodiments of the present disclosure.
[0036] FIG. 23 is a diagram illustrating a third example of an adjusted online chroma ALF filter shape according to some embodiments of the present disclosure.
[0037] FIG. 24 is a diagram illustrating a filter shape for a luma reconstructed signal according to some embodiments of the present disclosure.
[0038] FIG. 25 is a diagram illustrating a filter shape for a luma reconstructed signal according to some embodiments of the present disclosure.
[0039] FIG. 26 is a diagram illustrating a filter shape for a luma reconstructed signal according to some embodiments of the present disclosure.
[0040] FIG. 27 is a diagram illustrating different down-sampling filter shapes used to down-sample the luma reconstructed signal to have the same resolution as the chroma reconstructed signal.
[0041] FIG. 28 is a flowchart illustrating a video decoding method according to some embodiments of the present disclosure.
[0042] FIG. 29 is a flowchart illustrating a video encoding method according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0043] Reference will now be made in detail to the specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. But various alternatives can be used without departing from the scope of the claims, and the subject matter can be practiced without these specific details. For example, the subject matter presented herein can be implemented on many classes of electronic devices with digital video capabilities.
[0044] It should be noted that the terms "first," "second," and the like used in the description and the claims of the disclosure, as well as the appended drawings, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so construed can interchange depending on the context in which it is used. Embodiments of the present disclosure described herein can be implemented in a chronological order, in parallel, or in a reverse order unless otherwise explicitly stated.
[0045] FIG. 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, in accordance with some embodiments of the present disclosure. As shown in FIG. 1 The system 10 includes a source device 12 that generates and encodes video data to be decoded at a later time by a destination device 14, as shown in
[0046] In some embodiments, the destination device 14 can receive encoded video data to be decoded via a link 16. The link 16 can comprise any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 can comprise a communication medium to enable the source device 12 to transmit encoded video data directly to the destination device 14 in real-time. The encoded video data can be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device 14. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from the source device 12 to the destination device 14.
[0047] In some other implementations, encoded video data can be sent from output interface 22 to storage device 32. The target device 14 can then access the encoded video data in storage device 32 via input interface 28. Storage device 32 can include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, digital universal discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Target device 14 can access the stored video data from storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing and sending encoded video data to target device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Target device 14 can access the encoded video data via any standard data connection suitable for accessing encoded video data stored on the file server. Standard data connections include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or a combination of both. Transfer of encoded video data from storage device 32 can be streaming, downloading, or a combination of both.
[0048] like FIG. 1 As shown, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include sources or combinations of such sources, such as: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video. As an example, if video source 18 is a camera in a security monitoring system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application are generally applicable to video encoding and decoding and can be applied to wireless and / or wired applications.
[0049] The captured, pre-captured, or computer-generated video can be encoded by video encoder 20. The encoded video data can be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data can also (or alternatively) be stored onto storage device 32 for later access by destination device 14 or other devices, for decoding and / or playback. Output interface 22 can further include a modem and / or a transmitter. The encoded video data can comprise a sequence of pictures, each picture can include one or more sample arrays, e.g., for monochrome, only luma (Y); luma and two chroma in YCbCr or YCgCo domain; or green, blue, and red in GBR (also known as RGB) domain. For ease of notation and terminology in this application, the variables and terms associated with each set of three sample arrays can be referred to as luma and chroma in some embodiments, where the two chroma arrays can be referred to as Cb and Cr, regardless of what color representation method is actually used. The video data can be in chroma format 4:0:0, chroma format 4:2:0, chroma format 4:2:2, or chroma format 4:4:4, although the application is not limited to this.
[0050] Destination device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 can include a receiver and / or a modem and receive encoded video data over link 16. The encoded video data communicated over link 16, or provided on storage device 32, can include a variety of syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements can be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0051] In some implementations, destination device 14 can include display device 34, which can be an integrated display device and an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and can include any of a variety of display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0052] Video encoder 20 and video decoder 30 can operate according to a proprietary standard or industry standard, such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that the application is not limited to a specific video coding / decoding standard and can be applicable to other video coding / decoding standards. In general, it is contemplated that video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also contemplated that video decoder 30 of destination device 14 can be configured to decode video data according to any of these current or future standards.
[0053] Video encoder 20 and video decoder 30 can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware or any combinations thereof. When implemented partially in software, an electronic device can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video coding / decoding operations disclosed in the present disclosure. Each of video encoder 20 and video decoder 30 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective device.
[0054] In some implementations, at least a portion of the components of source device 12 (e.g., video source 18, video encoder 20, or the components of video encoder 20 described below with reference to FIG. 2, and / or output interface 22), and / or the components of destination device 14 (e.g., input interface 28, video decoder 30, or the components of video decoder 30 described below with reference to FIG. 3, and / or memory 32) can be implemented as one or more integrated circuits on one or more chips. FIG. 2 In some implementations, at least a portion of the components of source device 12 (e.g., video source 18, video encoder 20, or the components of video encoder 20 described below with reference to FIG. 2, and / or output interface 22), and / or the components of destination device 14 (e.g., input interface 28, video decoder 30, or the components of video decoder 30 described below with reference to FIG. 3, and / or memory 32) can be implemented as one or more integrated circuits on one or more chips. FIG. 3At least a portion of the components described as included in video decoder 30, and display device 34, can operate in a cloud computing services network, such as a software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS), which can provide software, platforms, and / or infrastructure. In some implementations, one or more components of source device 12 and / or destination device 14 that are not included in the cloud computing services network can be disposed in one or more client devices, and the one or more client devices can communicate with server computers in the cloud computing services network through a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In embodiments, at least a portion of the operations described herein can be implemented as a cloud-based service provided by one or more server computers implemented by at least a portion of the components of source device 12 and / or at least a portion of the components of destination device 14 in the cloud computing services network; and one or more other operations described herein can be implemented by one or more client devices. In some implementations, the cloud computing services network can be a private cloud, a public cloud, or a hybrid cloud. The terms such as “cloud,” “cloud computing,” “cloud-based,” and the like can be used interchangeably herein as appropriate without departing from the scope of the disclosure. It should be understood that the disclosure is not limited to implementation in the above-described cloud computing services network. Rather, the disclosure can also be implemented in any other type of computing environment currently known or developed in the future.
[0055] FIG. 2 FIG. 1 illustrates a block diagram of an example video encoder 20 according to some embodiments described in this application. Video encoder 20 can perform intra-prediction encoding and inter-prediction encoding on video blocks within a video frame. Intra-prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in video data within neighboring video frames or pictures of a video sequence. It should be noted that in the field of video coding, the term “frame” can be used as a synonym for the term “image” or “picture.”
[0056] As FIG. 2As shown in FIG. 1, video encoder 20 receives video data from video source 30, which can include one or more video capture devices, such as a video camera, a video archive, or a video feed from a video content provider. As discussed above, the video data can be in the form of a three-dimensional video sequence comprising a series of pictures, each picture including a plurality of samples. Video encoder 20 can pre-process the video data, for example, to perform format conversions or to perform other manipulations to the video data to support coding of the video data. Video encoder 20 can then encode the video data to generate encoded video data. The encoded video data can be stored in memory 40 and / or transmitted within a video bitstream.
[0057] Video data memory 40 can store video data to be encoded by the components of video encoder 20. The video data can be obtained, for example, from video source 30 and / or from memory 40. Video data memory 40 can also store temporary copies of data used during the encoding process. FIG. 1The illustrated video source 18 obtains video data in a video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use in encoding video data by the video encoder 20 (e.g., in intra- or inter-coding modes). The video data memory 40 and the DPB 64 can be formed by any of a variety of memory devices. In various examples, the video data memory 40 can be on-chip with other components of the video encoder 20, or off-chip relative to those components.
[0058] As shown in FIG. 2 After receiving the video data, the partitioning unit 45 within the prediction processing unit 41 partitions the video data into video blocks, as shown in A video frame is or can be considered as a two-dimensional array or matrix of sample values. The samples in the array can also be referred to as pixels or pels. The number of samples in the horizontal and vertical direction (or axis) of the array or picture defines the size and / or resolution of the video frame. For example, a video frame can be divided into a plurality of video blocks by using QT partitioning. A video block is again or can be considered as a two-dimensional array or matrix of sample values, but with dimensions smaller than those of the video frame. The number of samples in the horizontal and vertical direction (or axis) of the video block defines the size of the video block. A video block can be further partitioned into one or more block partitions or sub-blocks (which can form blocks again) by, for example, iteratively using QT partitioning, binary tree (BT) partitioning or ternary tree (TT) partitioning or any combination thereof. It should be noted that the term “block” or “video block” as used herein can be a portion of a frame or picture, in particular a rectangular (square or non-square) portion. With reference to, for example, HEVC and VVC, a block or video block can be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU) or a transform unit (TU) and / or can be or correspond to a respective block (e.g., coding tree block (CTB), coding block (CB), prediction block (PB) or transform block (TB)) and / or sub-block.
[0059] The prediction processing unit 41 can select one of a plurality of possible predictive encoding modes, such as one of a plurality of intra-predictive encoding modes or one of a plurality of inter-predictive encoding modes, for the current video block based on the error results (e.g., coding rate and level of distortion). The prediction processing unit 41 can provide the resulting intra- or inter-predicted block to the summer 50 to generate a residual block, and to the summer 62 to reconstruct the encoded block for use as part of a reference frame at a later time. The prediction processing unit 41 also provides syntax elements (such as motion vectors, intra-mode indicators, partitioning information, and other such syntax information) to the entropy encoding unit 56.
[0060] To select an appropriate intra-predictive encoding mode for the current video block, the intra-prediction processing unit 46 within the prediction processing unit 41 can perform intra-predictive encoding of the current video block in relation to one or more neighboring blocks in the same frame as the current block being encoded to provide spatial prediction. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-predictive encoding of the current video block in relation to one or more predictive blocks in one or more reference frames to provide temporal prediction. The video encoder 20 can perform multiple encoding passes, e.g., to select an appropriate encoding mode for each block of video data.
[0061] In some implementations, the motion estimation unit 42 determines an inter-prediction mode for a current video frame by generating motion vectors according to a predetermined pattern within a sequence of video frames, the motion vectors indicating displacements of video blocks within the current video frame relative to predictive blocks within a reference video frame. Motion estimation performed by the motion estimation unit 42 is a process of generating motion vectors that estimate the motion for video blocks. For example, a motion vector can indicate a displacement of a video block within a current video frame or picture relative to a predictive block within a reference frame that is related to a current block being encoded within the current frame. The predetermined pattern can designate video frames in the sequence as P-frames or B-frames. The intra-BC unit 48 can determine vectors for intra-BC encoding (e.g., block vectors) in a similar manner as the motion vectors determined by the motion estimation unit 42 for inter-prediction, or can utilize the motion estimation unit 42 to determine the block vectors.
[0062] In terms of pixel differences, the prediction block for a video block can be or can correspond to a block or reference block of a reference frame deemed to closely match the video block to be encoded, and the pixel differences can be determined by a sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metric. In some implementations, video encoder 20 can calculate values for sub-integer pixel positions of reference frames stored in DPB 64. For example, video encoder 20 can interpolate values for quarter-pel positions, eighth-pel positions, or other fractional-pel positions of reference frames. Thus, motion estimation unit 42 can perform a motion search with respect to both full-pel positions and fractional-pel positions and output motion vectors with fractional-pel precision.
[0063] Motion estimation unit 42 calculates a motion vector for a video block in an inter- predicted encoded frame by comparing a location of the video block to a location of a prediction block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of the first and second reference frame lists identifying one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44, which is then sent to entropy encoding unit 56.
[0064] Motion compensation performed by motion compensation unit 44 can involve fetching or generating a prediction block based on the motion vector determined by motion estimation unit 42. Upon receiving a motion vector for a current video block, motion compensation unit 44 can locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from DPB 64, and forward the prediction block to summer 50. Summer 50 then forms a residual video block of pixel difference values by subtracting pixel values of the prediction block provided by motion compensation unit 44 from pixel values of the current video block being encoded. The pixel difference values forming the residual video block can include luma component differences or chroma component differences or both. Motion compensation unit 44 can also generate syntax elements associated with the video block of the video frame for use by video decoder 30 when decoding the video block of the video frame. The syntax elements can include, for example, syntax elements defining a motion vector identifying the prediction block, any flags indicating a prediction mode, or any other syntax information described herein. It is noted that motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are illustrated separately for conceptual purposes.
[0065] In some implementations, intra BC unit 48 can generate vectors and obtain prediction blocks in a manner similar to that described above in connection with motion estimation unit 42 and motion compensation unit 44, but the prediction blocks are in the same frame as the current block being coded, and the vectors are referred to as block vectors rather than motion vectors. In particular, intra BC unit 48 can determine an intra prediction mode to be used for coding the current block. In some examples, intra BC unit 48 can code the current block using various intra prediction modes, e.g., during a separate encoding pass, and test their performance through rate-distortion analysis. Next, intra BC unit 48 can select an appropriate intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, intra BC unit 48 can compute rate-distortion values for the various tested intra prediction modes using rate-distortion analysis, and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that was coded to produce the coded block, and the bit rate (i.e., the number of bits) used to produce the coded block. Intra BC unit 48 can compute a ratio from the distortion and rate for various coded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.
[0066] In other examples, intra BC unit 48 can perform such functions for intra BC prediction according to the implementations described herein using motion estimation unit 42 and motion compensation unit 44 in whole or in part. In either case, for intra block copy, the prediction block can be a block that is considered to closely match the block to be coded in terms of pixel difference, which can be determined by SAD, SSD, or other difference metrics, and identifying the prediction block can include computing values for sub-integer pixel positions.
[0067] Regardless of whether the prediction block is from the same frame according to intra prediction or a different frame according to inter prediction, video encoder 20 can form pixel difference values by subtracting pixel values of the prediction block from pixel values of the current video block being coded to form a residual video block. The pixel difference values forming the residual video block can include both luma component differences and chroma component differences.
[0068] As an alternative to inter prediction performed by motion estimation unit 42 and motion compensation unit 44 or intra block copy prediction performed by intra BC unit 48, as described above, intra prediction processing unit 46 can intra predict the current video block. In particular, intra prediction processing unit 46 can determine an intra prediction mode to use for encoding the current block. To this end, intra prediction processing unit 46 can encode the current block using various intra prediction modes, e.g., during a separate encoding pass, and intra prediction processing unit 46 (or, in some examples, a mode selection unit) can select an appropriate intra prediction mode to use from among the tested intra prediction modes. Intra prediction processing unit 46 can provide information indicating the selected intra prediction mode for the block to entropy encoding unit 56. Entropy encoding unit 56 can encode information indicating the selected intra prediction mode into the bitstream.
[0069] After prediction processing unit 41 determines a prediction block for the current video block via either inter prediction or intra prediction, summer 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block can be included in one or more TUs and is provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0070] Transform processing unit 52 can send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting a quantization parameter. In some examples, quantization unit 54 can then perform a scan of the matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 can perform the scan.
[0071] Following quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), Probability Interval Partitioning Entropy (PIPE) coding or another entropy encoding method or technique. The encoded bits can then be sent to video decoder 30, as shown in FIG. 3, or archived, as shown in storage device 32, for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 can also entropy encode motion vectors and other syntax elements for the current video frame being encoded. FIG. 1 FIG. 1
[0072] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain for generating a reference block used to predict other video blocks. As noted above, the motion compensation unit 44 can generate a motion compensated prediction block from one or more reference blocks of a frame stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0073] The summer 62 adds the reconstructed residual block to the motion compensated prediction block produced by the motion compensation unit 44 to produce a reference block for storage in the DPB 64. The reference block can then be used by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 as a prediction block to inter predict another video block in a subsequent video frame.
[0074] FIG. 3 FIG. 1 shows a block diagram of an example video decoder 30 in accordance with some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 can perform a decoding process substantially reciprocal to the encoding process described above in connection with the video encoder 20. FIG. 2 The decoding process described in connection with the video encoder 20 is substantially reciprocal. For example, the motion compensation unit 82 can generate prediction data based on motion vectors received from the entropy decoding unit 80, while the intra prediction unit 84 can generate prediction data based on intra prediction mode indicators received from the entropy decoding unit 80.
[0075] In some examples, the units of the video decoder 30 can be tasked to perform embodiments of the present application. Moreover, in some examples, embodiments of the present disclosure can be distributed among one or more of the units of the video decoder 30. For example, the intra BC unit 85 can perform embodiments of the present application, alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 can not include the intra BC unit 85, and the functionality of the intra BC unit 85 can be performed by other components of the prediction processing unit 81, such as the motion compensation unit 82.
[0076] Video data memory 79 can store video data to be decoded by the other components of video decoder 30, such as an encoded video bitstream. The video data stored in video data memory 79 can be obtained, for example, from storage device 32, from a local video source, such as a camera, via wired or wireless network communication of video data, or by accessing physical data storage media (e.g., a flash drive or hard disk). Video data memory 79 can include a coded picture buffer (CPB) that stores encoded video data from an encoded video bitstream. DPB 92 of video decoder 30 stores reference video data for use in decoding video data by video decoder 30 (e.g., in intra- or inter-coding modes). Video data memory 79 and DPB 92 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For FIG. 3 illustrative purposes, video data memory 79 and DPB 92 are depicted as two distinct components of video decoder 30. But it will be readily apparent to one of ordinary skill in the art that video data memory 79 and DPB 92 can be provided by same memory device or separate memory devices. In some examples, video data memory 79 can be on-chip with other components of video decoder 30, or off-chip relative to those components.
[0077] During the decoding process, video decoder 30 receives an encoded video bitstream that represents encoded video frames and associated syntax elements. Video decoder 30 can receive the syntax elements at the video frame level and / or video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors, or intra- prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors, or intra-prediction mode indicators, and other syntax elements to prediction processing unit 81.
[0078] When a video frame is coded as an intra-predicted (I) frame or intra-coded prediction blocks in other types of frames, intra-prediction unit 84 of prediction processing unit 81 can generate prediction data for a video block of the current video frame based on the signaled intra-prediction mode and reference data from previously decoded blocks of the current frame.
[0079] When a video frame is coded as an inter prediction coded (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 generates one or more prediction blocks for a video block of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame within one of the reference frame lists. Video decoder 30 can construct the reference frame lists, i.e., List 0 and List 1, using default construction techniques based on the reference frames stored in DPB 92.
[0080] In some examples, when a video block is encoded according to the intra BC modes described herein, intra BC unit 85 of prediction processing unit 81 generates a prediction block for the current video block based on block vectors and other syntax elements received from entropy decoding unit 80. The prediction block can be within a reconstructed region of the same picture as the current video block, as defined by video encoder 20.
[0081] Motion compensation unit 82 and / or intra BC unit 85 determine the prediction information for a video block of the current video frame by parsing the motion vectors and other syntax elements, and then use the prediction information to generate a prediction block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode used to code the video block of the video frame (e.g., intra prediction or inter prediction), the inter prediction frame type (e.g., B or P), the construction information for one or more of the reference frame lists for the frame, the motion vectors for each inter prediction coded video block of the frame, the inter prediction status for each inter prediction coded video block of the frame, and other information used to decode the video block in the current video frame.
[0082] Similarly, intra BC unit 85 can use some of the received syntax elements, such as flags, to determine whether the current video block is predicted using the intra BC mode, the construction information for which video blocks of the frame are within the reconstructed region and should be stored in DPB 92, the block vectors for each intra BC predicted video block of the frame, the intra BC prediction status for each intra BC predicted video block of the frame, and other information used to decode the video block in the current video frame.
[0083] Motion compensation unit 82 can also perform interpolation to calculate interpolated values for sub-integer pixels of a reference block using interpolation filters as used by video encoder 20 during encoding of the video block. In this case, motion compensation unit 82 can determine the interpolation filters used by video encoder 20 from the received syntax elements, and use these interpolation filters to generate the prediction block.
[0084] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy-decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.
[0085] After the motion compensation unit 82 or the intra-frame BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra-frame BC unit 85. A loop filter 91 (e.g., a deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 for further processing of the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a separate memory device may also store the decoded video for later presentation on a display device (e.g., ...). FIG. 1 On the display device 34).
[0086] In a typical video encoding and decoding process, a video sequence usually consists of an ordered set of frames or images. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other instances, a frame may be monochrome and therefore consist of only a two-dimensional array of luminance samples.
[0087] like FIG. 4A As shown, the video encoder 20 (or more specifically, the segmentation unit 45) generates an encoded representation of a frame by first segmenting the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set such that all CTUs in the video sequence have the same size, such as 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not limited to a specific size. FIG. 4BAs shown, each CTU may include a CTB for the luma sample, two corresponding coding tree blocks for the chroma sample, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units in the coded pixel block and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome image or an image with three separate color planes, the CTU may include a single coding tree block and syntax elements for encoding the samples of that coding tree block. The coding tree block can be an N×N sample block.
[0088] To achieve better performance, the video encoder 20 can recursively perform tree splitting on the coding tree blocks of the CTU, such as binary tree splitting, ternary tree splitting, quadtree splitting, or combinations thereof, and divide the CTU into smaller CUs. FIG. 4C As described, the 64×64 CTU 400 is first divided into four smaller CUs, each with a block size of 32×32. Of these four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. The two 16×16 CUs, 430 and CU 440, are further divided into four CUs with a block size of 8×8. FIG. 4D Depicting as shown FIG. 4C The final result of the CTU 400 partitioning process described in the figure is a quadtree data structure, where each leaf node of the quadtree corresponds to a CU of a corresponding size ranging from 32×32 to 8×8. Similar to... FIG. 4B The CTU depicted in the image can include, for example, two corresponding coded blocks (CBs) of luminance and chrominance samples of the same size frame, as well as syntax elements for encoding the samples of the coded blocks. In monochrome images or images with three separate color planes, a CU can include a single coded block and a syntax structure for encoding the samples of the coded block. It should be noted that... FIG. 4C and FIG. 4D The quadtree partitioning depicted is for illustrative purposes only, and a CTU can be split into multiple CUs based on quadtree / ternary / binary partitioning to adapt to varying local characteristics. In multi-type tree structures, a CTU is partitioned according to a quadtree structure, and each quadtree leaf CU can be further partitioned according to binary and ternary tree structures. FIG. 4E As shown, a coded block with width W and height H has five possible segmentation types: quad segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal triple segmentation, and vertical triple segmentation.
[0089] In some implementations, the video encoder 20 may further segment the coded blocks of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter-frame or intra-frame) is applied. The PU of the CU may include a PB for luma samples, two corresponding PBs for chroma samples, and syntax elements for predicting the PBs. In a monochrome image or an image with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PBs. The video encoder 20 may generate predicted luma blocks, predicted Cb blocks, and predicted Cr blocks for each PU of the CU, representing the luma PB, Cb PB, and Cr PB.
[0090] Video encoder 20 can generate prediction blocks for a PU using intra-frame prediction or inter-frame prediction. If video encoder 20 uses intra-frame prediction to generate prediction blocks for a PU, then video encoder 20 can generate prediction blocks for a PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter-frame prediction to generate prediction blocks for a PU, then video encoder 20 can generate prediction blocks for a PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0091] After the video encoder 20 generates predicted luminance blocks, predicted Cb blocks, and predicted Cr blocks for one or more PUs of the CU, the video encoder 20 can generate luminance residual blocks for the CU by subtracting the predicted luminance blocks of the CU from the original luminance coding blocks of the CU, such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample in one of the predicted luminance blocks of the CU and a corresponding sample in the original luminance coding block of the CU. Similarly, the video encoder 20 can generate Cb residual blocks and Cr residual blocks for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0092] In addition, such as FIG. 4CAs shown, the video encoder 20 can use quadtree partitioning to decompose the luminance residual block, Cb residual block, and Cr residual block of the CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. A TU of the CU can include a transform block of the luminance samples, two corresponding transform blocks of the chrominance samples, and syntax elements for transforming the transform block samples. Therefore, each TU of the CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU can be a sub-block of the CU's luminance residual block. A Cb transform block can be a sub-block of the CU's Cb residual block. A Cr transform block can be a sub-block of the CU's Cr residual block. In a monochrome image or an image with three separate color planes, a TU can include a single transform block and syntax structures for transforming the samples of that transform block.
[0093] The video encoder 20 can apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block for the TU. The coefficient block can be a two-dimensional array of transform coefficients. The transform coefficients can be scalars. The video encoder 20 can apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 can apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.
[0094] After generating coefficient blocks (e.g., luminance coefficient blocks, Cb coefficient blocks, or Cr coefficient blocks), video encoder 20 can quantize the coefficient blocks. Quantization typically refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After quantizing the coefficient blocks, video encoder 20 can entropy encode the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream comprising a bit sequence that forms a representation of coded frames and associated data; the bitstream is stored in storage device 32 or transmitted to target device 14.
[0095] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements. The video decoder 30 can reconstruct frames of video data, at least in part, based on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coded blocks of the current CU by adding samples of the predicted blocks of the PU for the current CU to corresponding samples of the transformed blocks of the TU of the current CU. After reconstructing the coded blocks for each CU of the frame, the video decoder 30 can reconstruct the frame.
[0096] As mentioned above, video coding primarily uses two modes (i.e., intra-frame prediction and inter-frame prediction) to achieve video compression. It should be noted that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block based on a reference video block.
[0097] However, with continuously improving video data capture technologies and finer video tile sizes used to preserve details in video data, the amount of data required to represent the motion vectors for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only do a set of neighboring CUs in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vectors) of the current CU (also known as the "Motion Vector Prediction" (MVP) of the current CU) by exploring their spatial and temporal correlations.
[0098] Instead of the above combination FIG. 2 The method described involves encoding the actual motion vector of the current CU, determined by the motion estimation unit 42, into the video bitstream, and subtracting the predicted motion vector value of the current CU from the actual motion vector of the current CU to produce the motion vector difference (MVD) for the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.
[0099] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coded block, both the video encoder 20 and the video decoder 30 need to employ a set of rules to construct a motion vector candidate list (also known as a "merging list") for the current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs. Then, a member is selected from the motion vector candidate list as the motion vector prediction value for the current CU. By doing so, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector prediction value within the motion vector candidate list is sufficient for both the video encoder 20 and the video decoder 30 to encode and decode the current CU using the same motion vector prediction value from the motion vector candidate list.
[0100] introduction
[0101] Various video codec technologies can be used to compress video data. Video codecs are performed according to one or more video codec standards. For example, some well-known video codec standards today include Universal Video Codec (VVC), High Efficiency Video Codec (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Codec (AVC, also known as H.264 or MPEG-4 Part 10), which were jointly developed by ISO / IEC MPEG and ITU-T VCEG. AO Media Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to its previous standard VP9. Audio and Video Codec (AVS) (which refers to the Digital Audio and Digital Video Compression Standard) is another series of video compression standards developed by the Audio and Video Coding Standard Workgroup of China. Most existing video codec standards are built on a common hybrid video codec framework, which uses block-based prediction methods (e.g., inter-frame prediction, intra-frame prediction) to reduce redundancy in video images or sequences, and transform coding to compress the energy of prediction errors. A key goal of video codec technology is to compress video data into a form that uses a lower bit rate while avoiding or minimizing video quality degradation.
[0102] The first generation of AVS standards included the Chinese national standards "Information Technology Advanced Audio and Video Coding Part 2: Video" (referred to as AVS1) and "Information Technology Advanced Audio and Video Coding Part 16: Broadcast Television Video" (referred to as AVS+). Compared to the MPEG-2 standard, the first generation of AVS standards could provide approximately 50% bitrate savings while maintaining the same perceived quality. The video portion of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second generation of AVS standards included the Chinese national standard series "Information Technology High-Efficiency Multimedia Coding" (referred to as AVS2), which was primarily aimed at the transmission of additional HD TV programs. AVS2's coding and decoding efficiency was twice that of AVS+. In May 2016, AVS2 was released as a Chinese national standard. Simultaneously, the video portion of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international application standard. The AVS3 standard is a new generation of video coding and decoding standards for UHD video applications, aiming to surpass the coding and decoding efficiency of the latest international standard HEVC. In March 2019, at the 68th AVS meeting, the AVS3-P2 baseline was completed, offering approximately 30% bit rate savings over the HEVC standard. Currently, a reference software called the High Performance Model (HPM) exists, maintained by the AVS working group to demonstrate a reference implementation of the AVS3 standard.
[0103] Similar to HEVC, the AVS3 standard is built on a block-based hybrid video codec framework. FIG. 5 A block diagram of the overall block-based hybrid video coding system is presented. The input video signal is processed block by block (called a coding unit (CU)). Unlike HEVC, which partitions blocks solely based on quadtrees, in AVS3, a coding tree unit (CTU) is divided into multiple CUs to accommodate different local characteristics based on quadtree / binary tree / extended quadtree. Furthermore, the concept of multiple partitioning unit types in HEVC is removed; that is, the splitting of CUs, prediction units (PUs), and transform units (TUs) does not exist in AVS3. Instead, each CU is always used as the basic unit for both prediction and transform, without further partitioning. In AVS3's tree partitioning structure, a CTU is first partitioned based on a quadtree structure. Then, each quadtree leaf node can be further partitioned based on binary tree and extended quadtree structures. FIG. 5In the encoder, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video picture / strip to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically represented by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is sent to identify which reference picture in the reference picture memory the temporal prediction signal originates from. After spatial and / or temporal prediction, a mode decision block in the encoder selects the optimal prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block; and the prediction residual is decorrelated using a transform and then quantized. The quantized residual coefficients are dequantized and inversely transformed to form the reconstructed residuals, which are then added back to the prediction block to form the reconstructed CU signal. Before placing the reconstructed CU in the reference image storage area and using it as a reference for encoding future video blocks, further loop filtering, such as deblocking filters, Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF), can be applied to the reconstructed CU. To form the output video bitstream, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantized residual coefficients are sent to the entropy coding unit for further compression and packing.
[0104] The first version of the HEVC standard was completed in October 2013, offering approximately 50% bitrate savings or equivalent perceived quality compared to its predecessor, H.264 / MPEG-AVC. While HEVC offered significant codec improvements over its predecessor, evidence suggested that even better codec efficiency could be achieved using additional codec tools. Based on this, both VCEG and MPEG began exploring new codec technologies for future video codec standardization. In October 2015, ITU-T VCEG and ISO / IEC MPEG established a Joint Video Exploration Team (JVET) to conduct significant research on advanced technologies capable of substantially improving codec efficiency. JVET maintains a reference software called the Joint Exploration Model (JEM) by integrating multiple additional codec tools based on the HEVC Test Model (HM).
[0105] In October 2017, ITU-T and ISO / IEC published a joint proposal (CfP) on video compression capabilities exceeding HEVC. In April 2018, the 10th JVET meeting received and evaluated 23 CfP responses, demonstrating a compression efficiency improvement of approximately 40% compared to HEVC. Based on these evaluation results, JVET launched a new project to develop a next-generation video codec standard called Universal Video Codec (VVC). That same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.
[0106] Similar to HEVC, VVC is built on a block-based hybrid video codec framework. FIG. 5 A block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (called a coding unit (CU)). In VTM-1.0, a CU can be up to 128×128 pixels. However, unlike HEVC, which partitions blocks solely based on quadtrees, in VVC, a coding tree unit (CTU) is divided into multiple CUs to accommodate different local characteristics based on quadtrees, binary trees, and ternary trees. Furthermore, the concept of multiple partitioning unit types in HEVC is removed; that is, there is no longer a division of CUs, prediction units (PUs), and transform units (TUs) in VVC. Instead, each CU is always used as the basic unit for both prediction and transform, without further partitioning. In a multi-type tree structure, a CTU is first partitioned according to a quadtree structure. Then, each quadtree leaf node can be further partitioned according to binary tree and ternary tree structures. FIG. 6 As shown, there are five partitioning types: quad partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning. FIG. 5In the encoder, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video picture / strip to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically represented by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal reference. Similarly, if multiple reference pictures are supported, an additional reference picture index is sent to identify which reference picture in the reference picture store the temporal prediction signal comes from. After spatial and / or temporal prediction, the mode decision block in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block; and the prediction residual is decorrelated and quantized using a transform. The quantized residual coefficients are dequantized and inversely transformed to form the reconstructed residuals, which are then added back to the prediction block to form the reconstructed CU signal. Further, before placing the reconstructed CU in the reference image storage area and using it to encode future video blocks, loop filtering such as deblocking filters, Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF) can be applied to the reconstructed CU. To form the output video bitstream, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantized residual coefficients are sent to the entropy coding unit for further compression and packing to form the bitstream.
[0107] FIG. 7 A general block diagram of a block-based video decoder is given. First, the video bitstream is entropy-decoded at the entropy decoding unit. The coding mode and prediction information are sent to the spatial prediction unit (in the case of intra-frame encoding / decoding) or the temporal prediction unit (in the case of inter-frame encoding / decoding) to form prediction blocks. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual blocks. Then, the prediction blocks and the residual blocks are added together. The reconstructed blocks can be further filtered through a loop and then stored in a reference image storage area. The reconstructed video from the reference image storage area is then sent to drive the display device and used to predict future video blocks.
[0108] This disclosure focuses on improving the Adaptive Loop Filter (ALF) and the Cross-Component Adaptive Loop Filter (CCALF). The relevant background information is explained in detail in the following sections.
[0109] Related research
[0110] ALF in VVC
[0111] Filter shape, linear filtering and adaptive truncation
[0112] In VVC, ALF is applied to the output samples of SAO. The luminance and chrominance components each support two filter shapes: 7×7 rhombus and 5×5 rhombus, as shown below. FIG. 8 As shown. In FIG. 8 In this model, each square corresponds to a luminance or chrominance sample, and the center square corresponds to the current sample to be filtered. The filter coefficients are point-symmetric, and each integer filter coefficient is represented with 7 decimal places of precision. Furthermore, the sum of the coefficients of a filter equals 128, which is a fixed-point representation of 1.0 with 7 decimal places of precision.
[0113]
[0114] For the 7×7 and 5×5 filter shapes, the number of coefficients N is equal to 13 and 7, respectively.
[0115] The value of the filtered sample at coordinates (x, y) By applying coefficient c to the reconstructed sample values R(x,y) i The result is shown in the following formula:
[0116]
[0117] Where (x+x) i ,y+y i ) and (xx i yy i ) is related to the i-th coefficient c i The coordinates of the corresponding reconstructed sample points. Due to the constraints in equation (1), equation (2) can be written as:
[0118]
[0119] In VVC, equation (3) adds the possibility of truncating the difference between neighboring sample values and the current sample to be filtered, as shown in the following equation:
[0120]
[0121] in
[0122] f i =min(b i ,max(-b i ,R(x+x i ,y+y i )-R(x,y)))+
[0123] min(b i ,max(-bi ,R(xx i yy i )-R(x,y))) (5)
[0125] b i It is the coefficient c i The truncation parameter is determined by the truncation index d. i Decision. i The conclusion is as follows:
[0126]
[0127] Where BD is the depth of the sample point, and d i It can be 0, 1, 2 or 3.
[0128] Luminance sub-block level filter adaptive
[0129] In VVC, sub-block-level filtering is adaptively applied only to the luminance component. Each 4×4 luminance block is classified based on its directionality and 2D Laplacian activity. First, sample gradient values in the horizontal, vertical, and two diagonal directions are calculated:
[0130] H k,l =|2R(k,l)-R(k-1,l)-R(k+1,l)|,
[0131] V k,l =|2R(k,l)-R(k,l-1)-R(k,l+1)|,
[0132] D0 k,l =|2R(k,l)-R(k-1,l-1)-R(k+1,l+1)|,
[0133] D1 k,l =|2R(k,l)-R(k-1,l+1)-R(k+1,l-1)|. (7)
[0135] Based on the sample point gradient, the sub-block horizontal gradient g h Vertical gradient g v And two diagonal gradients g d0 and g d1 Calculated as
[0136]
[0137] Indices i and j refer to the coordinates of the top-left sample point in the 4×4 brightness block. As can be seen from equation (8), the sum of the gradients of the sample points within the 10×10 brightness window covering the target 4×4 block is used to classify the block. To reduce complexity, only the gradient of every other sample point within the 10×10 window is calculated, such as... FIG. 9 As shown. The gradient values for other sample points are set to 0.
[0138] Secondly, in order to assign directionality D, the ratio of the maximum to the minimum values of the horizontal and vertical gradients of the sub-block is used:
[0139]
[0140] And the ratio of the maximum to the minimum value of the diagonal gradient of the two sub-blocks:
[0141]
[0142] Compare with a set of thresholds t1 and t2:
[0143] Step 1: If and Then D is set to 0.
[0144] Step 2: If If so, calculate the directionality D in step 3; otherwise, calculate it in step 4.
[0145] Step 3: If Then D is set to 2; otherwise, D is set to 1.
[0146] Step 4: If Then D is set to 4; otherwise, D is set to 3.
[0147] Each subsequent step in the above D calculation will only be executed if no value has been assigned to D in the preceding steps. Third, the activity value A is calculated as...
[0148]
[0149] A is further mapped to the range of 0 to 4: Among them, {Q n} = {0,1,2,2,2,2,2,3,3,3,3,3,3,3,3,3,4}. Finally, each 4×4 luminance block is categorized into one of the following 25 classes:
[0150]
[0151] Each class can be assigned its own filter.
[0152] Before filtering each 4×4 luminance block, a geometric transformation is applied to the filter coefficients, such as a 90-degree rotation, diagonal rotation, or vertical flip, as follows: FIG. 10 As shown, the specific gradient value depends on the sub-block value specified in Table 1.
[0153] Table 1
[0154] Sub-block gradient value Transform g d1 g d0 and g h g v ]]> No transform g d1 g d0 g v g h ]]> Diagonal flip g do ≤g d1 and g h <g v ]]> Vertical flip g do ≤g d1 and g v ≤g h ]]> 90-degree rotation
[0155] Coding tree block-level filter adaptive
[0156] In addition to luminance 4×4 block-level filter adaptation, ALF also supports CTB-level filter adaptation. A luminance CTB can use a filter bank computed for the current stripe, or one of a filter bank computed for an encoded stripe. It can also use one of 16 offline-trained filter banks. In each luminance CTB, which filter from the selected filter bank should be applied to each 4×4 block is determined by the class C computed for that block according to Equation (12).
[0157] Chroma is adaptively applied using only CTB-level filters. Within a single stripe, the chroma component can use up to eight filters. Each CTB can select one of these filters.
[0158] syntax design
[0159] Filter coefficients and truncation indices are carried in the ALF APS. An ALF APS can include up to eight chromaticity filters and one luma filter bank (containing up to 25 filters). Each of the 25 luma classes also includes an index i. C Having the same index i C Classes share the same filter. By merging different classes, the number of bits required to represent filter coefficients is reduced. The absolute values of filter coefficients are represented using zero-order exponent Golomb code, followed by the sign bit of the non-zero coefficient. When truncation is enabled, a two-bit fixed-length code is also used for each filter coefficient to transmit a truncation index via signal transmission. The maximum storage required for ALF coefficients and truncation indices within an APS is 3480 bits. The decoder can use up to eight ALF APSs simultaneously.
[0160] The filter control syntax elements include two types of information. First, an ALF on / off flag is signaled at the sequence, picture, strip, and CTB levels. Chroma ALF can only be enabled at the corresponding level if Luminance ALF is enabled at the picture and strip levels. Second, if ALF is enabled at the picture, strip, and CTB levels, filter usage information is signaled at that level. The referenced ALF APS ID is encoded / decoded at the strip level, or at the picture level if all strips within a picture use the same APS. A luminance component can reference up to 7 ALF APSs, while a chroma component can reference up to 1 ALF APS. For the luminance CTB, an index is signaled indicating which ALF APS or offline-trained luminance filter bank to use. For the chroma CTB, the index indicates which filter in the referenced APS is used.
[0161] Line buffer reduction
[0162] To reduce the storage requirements of ALF, VVC employs row buffer boundary handling. In VVC, the row buffer boundary is located four luma samples and two chroma samples above the horizontal CTU boundary. When ALF is applied to samples on one side of the row buffer boundary, samples on the other side of the row buffer boundary cannot be used.
[0163] ALF in ECM
[0164] Remove ALF simplification
[0165] ALF gradient subsampling and ALF virtual boundary processing have been removed. The block size used for classification has been reduced from 4×4 to 2×2. The size of the filter that transmits the luminance and chrominance coefficients as signals has been increased to 9×9.
[0166] ALF with fixed filter
[0167] To filter the luminance samples, three different classifiers (C0, C1, and C2) and three different filter banks (F0, F1, and F2) were used. Banks F0 and F1 contain fixed filters whose coefficients were trained on classifiers C0 and C1. The filter coefficients in F2 were transmitted using the signal. For a given sample, bank F1 was used... i Which filter in the classifier C is used? i Class C assigned to this sample point i Decision made.
[0168] Filtering
[0169] First, two 13×13 rhombus fixed filters, F0 and F1, are applied to obtain two intermediate samples, R0(x,y) and R1(x,y). Then, F2 is applied to R0(x,y), R1(x,y), neighboring samples, and the samples before the deblocking filter (DBF) to obtain the filtered samples.
[0170]
[0171] Among them, f i,j It is the cutoff difference between neighboring samples and the current sample R(x,y), g i It is R i-20 The difference between (x,y) and the current sample point R(x,y), h i,j It is the truncation difference between the neighboring samples before DBF and the current sample R(x,y). It is expressed using the signal transmission filter coefficients c. i , i = 0,…24. The filter shape of F2 is as follows: FIG. 11 As shown.
[0172] Classification
[0173] Based on directionality D i and activities Assign class C to each 2×2 block i :
[0174]
[0175] Among them, M D,i Indicates directionality D i The total number.
[0176] Similar to VVC, the horizontal, vertical, and two diagonal gradients for each sample are computed using a 1-D Laplacian operator. The sum of gradients from samples within a 4×4 window covering a 2×2 block of the target is used for classifier C0, and the sum of gradients from samples within a 12×12 window is used for classifiers C1 and C2. The sums of the horizontal, vertical, and two diagonal gradients are expressed as follows: and Directionality D i This is achieved by:
[0177]
[0178] It is determined by comparison with a set of thresholds. Directionality D2 in VVC is derived using thresholds 2 and 4.5. For D0 and D1, the horizontal / vertical edge strength is calculated first. and diagonal edge strength A threshold Th = [1.25, 1.5, 2, 3, 4.5, 8] was used. If... Then edge strength =0; otherwise, It is to satisfy The largest integer. If Then edge strength =0; otherwise, It is to satisfy The largest integer. When When, that is, when the horizontal / vertical edges dominate, D i This is derived using Table 2(a); otherwise, when the diagonal edge dominates, D i This was obtained using Table 2(b).
[0179] Table 2
[0180]
[0181] In order to obtain The sum of the vertical and horizontal gradients, A i Mapped to the range 0 to n, where, for n equals 4, and for and n equals 15.
[0182] In ALF_APS, a maximum of 4 luminance filter groups can be used for signal transmission, and each group can have up to 25 filters.
[0183] Alternative 2×2 ALF classifier
[0184] The classification in ALF is extended by adding an alternative classifier. For a luminance filter bank that has been signal-transmitted, a flag is transmitted to indicate whether an alternative classifier has been applied. No geometric transformation is applied to the alternative band classifier. When applying a band-based classifier, the sum of the sample values of the 2×2 luminance block is first calculated. Then the class index is calculated as follows:
[0185] class_index = (sum * 25) >> (sample point depth + 2) (16)
[0186] Residual-based classifiers
[0187] The ALF classification is extended by a third classifier based on the luminance residual sample values. For each 2×2 luminance block, the sum of the absolute values of the residual samples in the neighboring 8×8 windows is calculated, and the class index is derived as follows:
[0188] classIdx = total >> (sample point depth - 4).
[0189] The value range of classIdx is 0 to 24, the same as in ECM-8.0. This describes the use of a signal transmission classifier for each luminance filter bank in APS.
[0190] Taps for ALF based on the output of the extended fixed filter
[0191] In ALF, online trained filters include the following four types of filter taps: spatial taps, taps based on reconstruction prior to DBF, residual-based taps, and taps based on fixed filter outputs, such as... FIG. 19 As shown.
[0192] Improved fixed filter for ALF
[0193] Two Laplacian-based classifiers (one for each fixed filter) are applied to a 2×2 block. Within each classifier, activity and orientation values are derived based on the vertical, horizontal, and diagonal gradients using a window surrounding each 2×2 block. For each 2×2 block, the mean of the surrounding window is computed. Then, for each sample point within that window, the difference between that sample value and the mean is calculated. A scaling factor is determined based on the activity values derived by the Laplacian classifiers. The square root of the sum of squared differences is further quantized to C using the scaling factor. ′ The value of C′ is an integer between 0 and 7 (inclusive). For i = 0, 1, let C′ = 0. i Let C represent the classifier derived from the classifier with the i-th fixed filter in ECM-9.0. Then, the proposed class index C is... i 'Derived as
[0194] C i =C'*896+C i .
[0195] The total number of fixed filters remains constant.
[0196] The class index is then determined based on the activity and directionality values. Two diamond-shaped fixed filters are selected from two filter banks using the derived class indices. Both fixed filters are applied to the samples before DBF and the ALF input, with an additional 9×9 diamond-shaped filter used for the samples before DBF. The shape of the first fixed filter applied to the ALF input samples is reduced from 13×13 to 9×9, while the shape of the second fixed filter applied to the ALF input (i.e., 13×13) remains unchanged, as shown in Table 3.
[0197] Table 3
[0198]
[0199] Apply a fixed filter f1 to the output of f0 (instead of the ALF input) and to the samples before DBF.
[0200] Finally, the signal transmission filter is applied to the ALF input samples, the samples before the deblocking filter (DBF), the outputs of the two fixed filters, the output of the Gaussian filter, and the residual data.
[0201] Chromaticity ALF Fixed Filter
[0202] A classifier based on Laplacian value and variance is applied to the 2×2 chroma blocks. Compared to the luminance classifier with fixed filters, the sum of the vertical and horizontal Laplacian values of the chroma is multiplied by 2 before scaling when calculating the active values. Similarly, the chroma variance is multiplied by 2 before scaling. Fixed filters are then selected from the chroma filter bank using a derived class index. The chroma fixed filters use a 13×13 rhombus shape when applied to chroma ALF input samples and a 7×7 rhombus shape when applied to DBF input samples. A first luminance classifier is applied to each 2×2 chroma block. Fixed filters are then selected from the luminance fixed filter bank associated with that classifier using a derived class index. Fixed filters use a 9×9 rhombus shape when applied to chroma ALF input samples and a 9×9 rhombus shape when applied to DBF input samples. In the chroma filters used for signal transmission, 5×5 cross-shaped additional taps are introduced, which are applied to the fixed filter outputs. The updated online chroma ALF filter shape is as follows: FIG. 20 As shown.
[0203] Extended use of fixed filters
[0204] The fixed filter defined in ALF is executed immediately after DBF. The fixed filter takes the reconstructed samples before and after DBF as input to generate the filtered result for this stage. The classification and filtering logic of the fixed filter can be directly reused from ALF. The ALF process also remains unchanged, including offline and online filtering. The fixed filter in this stage uses a strip-level flag transmitted via signal transmission to implement on / off control.
[0205] CCALF in VVC
[0206] Filter shape and accuracy
[0207] CCALF uses luminance sample values to refine chrominance sample values during the ALF process. For example... FIG. 12 As shown, the linear filtering operation takes the luminance sample values as input and generates correction values for the chrominance sample values. The above correction is generated independently for each chrominance component i (i∈{Cb,Cr}) and can be expressed as:
[0208]
[0209] Where (x,y) are the sample point positions of chromaticity component i, (x C ,y C (x0, y0) is the position of the brightness sample point obtained from (x, y), and (x0, y0) is the position of the brightness sample point obtained from (x, y). C ,y C The surrounding filters support offsets, S i This is the filter support region used for the luminance of chromaticity component i. Luminance position (x) C ,y C The value is determined based on the spatial scaling factor between the luminance and chrominance planes. The sample values in the luminance support region are also the inputs to the ALF luminance stage and correspond to the outputs of the SAO stage.
[0210] like FIG. 13 As shown, the CCALF filter has a diamond shape. FIG. 13 As shown, for a 4:2:0 video sequence with a chroma position type of 0 (i.e., when the chroma sample is co-located with the even column of the luminance sample in the horizontal direction and located between the rows of the luminance sample in the vertical direction), the center of the rhombus is aligned with the position of the chroma sample.
[0211] Because it does not enforce symmetric constraints, CCALF coefficients offer greater flexibility compared to regular ALF coefficients. However, there are also two limitations:
[0212] 1) To maintain DC neutrality, the sum of the CCALF coefficient values must be zero. Therefore, only seven of the eight CCALF coefficients need to be transmitted as signals in the bitstream, and the position (x) C ,y C The coefficients at () are obtained at the decoder.
[0213] 2) The absolute values of the CCALF coefficients are restricted to zero or integer powers of two, specifically {0, 1, 2, 4, 8, 16, 32, 64}. This allows implementations to use variable shift operations instead of CCALF multiplication as needed.
[0214] syntax design
[0215] In the final VVC design, the maximum number of filters per chroma component of the image is four. Different groups of CCALF coefficients can be selected for each CTU of the chroma component. Like regular ALF coefficients, CCALF coefficients are transmitted as signals within ALF APSs. Each ALF APS can contain a maximum of four CCALF filters per chroma component. While CCALF can be enabled at the sequence level, it can only be enabled if ALF is also enabled at the sequence level. Similarly, CCALF can only be enabled at the corresponding level if Luminance ALF is enabled at both the image and stripe levels.
[0216] Line buffer reduction
[0217] As described in the "Line Buffer Reduction" section, the line buffer boundaries for luma and chroma are located four and two samples above the CTU boundary, respectively. For the 4:2:0 chroma format, this aligns the line buffer boundaries for chroma and luma. However, for the 4:2:2 and 4:4:4 chroma formats, the line buffer boundaries for chroma and luma are not aligned with each other. Due to this misalignment, for the 4:2:2 and 4:4:4 chroma formats, CC-ALF does not apply to the third and fourth line samples above the CTU boundary.
[0218] CCALF in ECM
[0219] The CCALF process uses a linear filter to filter luminance sample values and generate residual corrections for chrominance samples. A large 25-tap filter is used in the CCALF process, such as... FIG. 14 As shown. For a given strip, the encoder can collect and analyze the statistical data of that strip, and transmit the signal through up to 16 filters via the APS.
[0220] Although ALF and CCALF have been improved in ECM, there is still room for further performance enhancement.
[0221] First, the online ALF filter in ECM takes spatially neighboring pixels, the fixed ALF filter result, and the spatially neighboring pixels before the deblocking filter as inputs. However, in addition to this information, other information (such as spatially neighboring pixels in the predicted signal, spatially neighboring pixels in the residual signal, or spatially neighboring pixels before SAO) can also be used as inputs to the online ALF filter equations, which can improve encoding and decoding performance.
[0222] Secondly, in ECM, the online ALF filter adaptively uses both edge-based and band-based classifiers. However, these two classifiers can be further combined to provide other classifiers, which can improve encoding and decoding performance.
[0223] Third, in ECM, the filter shape of chroma ALF is diamond-shaped, while the filter shape of luma ALF is long cross-shaped. From a standardization perspective, this non-uniform design may not be optimal.
[0224] Fourth, edge-based and band-based classifiers in ECM only consider pixel values after SAO. However, after saving pixel values from 1) immediately before the deblocking filter, 2) the predicted signal, 3) the residual signal, and 4) immediately before SAO as inputs to the online ALF filter equations, these pixel values can also be used to design new classifiers, which can improve encoding and decoding performance.
[0225] Fifth, edge-based and band-based classifiers in ECM only consider the luminance pixel values after SAO. However, chrominance pixel values can also be used to design new classifiers, which can improve encoding and decoding performance.
[0226] Sixth, similar to saving the luminance pixel values from the stages 1) immediately before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before the SAO as additional inputs to the online luminance ALF filter equation, the chrominance pixel values from the stages 1) immediately before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before the SAO can also be saved as additional inputs to the online chrominance ALF filter equation, which can improve encoding and decoding performance.
[0227] Seventh, similar to saving the luminance pixel values from the stages immediately before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before the SAO as additional inputs to the online luminance ALF filter equation, it is also possible to save the luminance pixel values from the stages immediately before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before the SAO as additional inputs to the CCALF filter equation, which can improve encoding and decoding performance.
[0228] Eighth, the classifier design in ECM only considers the reconstructed pixel values. However, encoding mode information (such as whether the coded block uses skip mode for encoding and decoding, and whether the coded block uses intra-frame, inter-frame P, or inter-frame B modes for encoding and decoding) can also be used to design the classifier, which can improve encoding and decoding performance.
[0229] Ninth, when the online ALF filter takes samples from 1) the samples immediately before deblocking, 2) the prediction samples, 3) the residual samples, and 4) the samples immediately before SAO as additional inputs, an additional line buffer is needed to store the 4 lines of corresponding luminance samples and 2 lines of corresponding chrominance samples above the horizontal CTU boundary, depending on the current line buffer settings in VVC. This increases the implementation complexity.
[0230] Tenth, when the CCALF filter takes samples from 1) the samples immediately before deblocking, 2) the predicted samples, 3) the residual samples, and 4) the samples immediately before SAO as additional inputs, an additional line buffer is needed to store the 4 lines of corresponding luminance samples above the horizontal CTU boundary, depending on the current line buffer settings in VVC. This increases the implementation complexity.
[0231] Eleventh, when the online ALF filter uses samples from 1) the samples immediately before deblocking, 2) the predicted samples, 3) the residual samples, and 4) the samples immediately before SAO as additional inputs, sample padding is required when the filter shape of the additional input (whose center position is aligned with the sample to be filtered) crosses the virtual boundary (line buffer boundary) or the picture (strip, tile) boundary.
[0232] Twelfth, two sets of ALF fixed filters are trained using two edge-based classifiers with two different window sizes. However, in addition to edge-based classifiers, other classifiers (such as band-based classifiers, residual-based classifiers, etc.) can be used to train the corresponding ALF fixed filter sets, and the online ALF filter can use the outputs of all these trained ALF fixed filter sets as additional inputs, which can improve encoding and decoding performance.
[0233] Thirteenth, spatially neighbor reconstructed pixels are used as input to train the ALF fixed filter. However, when training the ALF fixed filter, in addition to spatially neighbor reconstructed pixels, spatially neighbor pixels immediately before the deblocking filter, spatially neighbor pixels in the predicted signal, spatially neighbor pixels in the residual signal, or spatially neighbor pixels immediately before the SAO can also be used as input to the ALF fixed filter, which can improve encoding and decoding performance.
[0234] Fourteenth, apply sub-block level filter adaptation only in the luma ALF. However, in addition to the luma ALF, the sub-block level filter adaptation can be extended to the chroma ALF, which can improve encoding and decoding performance.
[0235] Fifteenth, sub-block level filter adaptation is applied only in the luma ALF. However, in addition to the luma ALF, sub-block level filter adaptation can be extended to the CCALF, which can improve encoding and decoding performance.
[0236] Sixteenth, the online chroma ALF filter takes the spatially neighboring pixels in the chroma reconstruction signal and the fixed filter output as input. However, in addition to this information, other information, such as the spatially neighboring pixels in the luminance reconstruction signal, can also be used as input to the online chroma ALF filter, which can improve encoding and decoding performance.
[0237] Seventeenth, the CCALF filter takes spatially neighboring pixels from the original luminance reconstruction signal as input. However, in addition to this information, other information, such as spatially neighboring pixels from the downsampled luminance reconstruction signal, can also be used as input to the CCALF filter, which can improve encoding and decoding performance.
[0238] Eighteenth, the luminance fixed filter is applied once immediately after the DBF. However, besides applying it once immediately after the DBF, the luminance fixed filter can also be applied multiple times at different stages of the loop filter, which can improve encoding and decoding performance.
[0239] Nineteenth, a fixed-color filter is applied in the chroma ALF stage, where the output of the fixed-color filter is used as the input to the inline chroma ALF filter. However, besides being used as the input to the inline chroma ALF filter, the fixed-color filter can also be applied to other stages of the loop filter, which can improve encoding and decoding performance.
[0240] Twentieth, offline training is performed on the luminance-fixed and chrominance-fixed filters. When these are directly applied to different stages of the loop filter, the output can be adjusted to adapt to online conditions, which can improve encoding and decoding performance.
[0241] In this disclosure, methods for further improving the existing design of ALF are provided to address the problems identified in the "Problem Description" section. Overall, the main features of the technology proposed in this disclosure are summarized below.
[0242] 1. The online ALF filter takes spatially neighboring pixels in the predicted signal, spatially neighboring pixels in the residual signal, or spatially neighboring pixels before SAO as additional inputs.
[0243] 2. A classifier that combines features from an edge-based classifier and features from a band-based classifier is used as an additional classifier for the online ALF filter.
[0244] 3. Change the filter shape of the chroma ALF from a rhombus to a long cross shape to match the filter shape of the luminance ALF.
[0245] 4. A classifier using pixel values from 1) immediately before the deblocking filter, 2) the predicted signal, 3) the residual signal, and 4) immediately before the SAO is used as an additional classifier for the online ALF filter.
[0246] 5. A classifier using chroma pixel values is used as an additional classifier for the online ALF filter.
[0247] 6. The online chroma ALF filter takes spatial neighbor pixels from the chroma prediction signal, spatial neighbor pixels from the chroma residual signal, spatial neighbor pixels from the stage immediately preceding the chroma SAO, or spatial neighbor pixels from the stage immediately preceding the chroma deblocking as additional inputs.
[0248] 7. The CCALF filter takes spatially neighboring pixels from the luminance prediction signal, spatially neighboring pixels from the luminance residual signal, spatially neighboring pixels from the stage immediately preceding the luminance SAO, or spatially neighboring pixels from the stage immediately preceding the luminance deblocking as additional inputs.
[0249] 8. Classifiers that utilize coding mode information (e.g., whether the coding block is encoded and decoded in skip mode, and whether the coding block is encoded and decoded in intra-frame, inter-frame P, or inter-frame B mode) are used as additional classifiers for the online ALF filter.
[0250] 9. When the online ALF filter takes samples from the following stages as additional inputs: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO, the 4 rows of luminance samples and 2 rows of chrominance samples above the horizontal CTU boundary are assumed to be the default values, depending on the current line buffer settings in VVC. This saves these line buffers.
[0251] 10. When the CCALF filter takes samples from the following stages as additional inputs: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO, the 4 lines above the horizontal CTU boundary corresponding to the luminance samples are assumed to be the default value according to the current line buffer settings in VVC, which can save these line buffers.
[0252] 11. When the online ALF filter uses samples from the following stages as additional inputs: 1) samples immediately before deblocking, 2) predicted samples, 3) residual samples, and 4) samples immediately before SAO, sample padding is performed when the filter shape of the additional input (whose center position is aligned with the sample to be filtered) crosses the virtual boundary (line buffer boundary) or the picture (strip, tile) boundary.
[0253] 12. Band-based classifiers, residual-based classifiers, etc., are used to train additional ALF fixed filter banks. Then, the outputs of these additional ALF fixed filter banks, along with the outputs of the original two sets of ALF fixed filters trained based on two edge-based classifiers, are used as inputs to the online ALF filter.
[0254] 13. When training the ALF fixed filter, the spatial neighbor reconstructed pixels, along with the spatial neighbor pixels immediately preceding the deblocking filter, the spatial neighbor pixels in the predicted signal, the spatial neighbor pixels in the residual signal, or the spatial neighbor pixels immediately preceding the SAO, are used as inputs to the ALF fixed filter.
[0255] 14. Adaptively apply sub-block level filters to the chroma ALF, where edge-based, band-based, or residual-based classifiers are used in the chroma ALF.
[0256] 15. Adaptively apply sub-block level filters to CCALF, where edge-based, band-based, or residual-based classifiers are used in CCALF.
[0257] 16. The online chroma ALF filter uses spatially neighboring pixels in the luminance reconstruction signal as additional input.
[0258] 17. The CCALF filter uses spatially neighboring pixels from the downsampled luminance reconstruction signal as additional input.
[0259] 18. Apply a fixed brightness filter in each stage of the loop filter.
[0260] 19. Apply chromaticity fixed filters in each stage of the loop filter.
[0261] 20. Apply residual scaling or residual offset adjustment to the results of a luminance or chrominance fixed filter.
[0262] It should be noted that the disclosed methods can be applied independently or in combination.
[0263] Information from the forecast, residuals, or prior to SAO is used as additional ALF input.
[0264] According to one or more embodiments of this disclosure, information from the prediction, residuals, or prior to SAO is used as additional ALF equation inputs. Different methods can be used to achieve this.
[0265] In the first method, spatially neighboring pixels in the predicted signal are proposed as inputs to the additional ALF equations. Various filter shapes can be used to extract information from the predicted signal. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15As shown, various equation forms can be used to extract information from the predicted signal. In one example, the truncation difference between the surrounding pixels and the current pixel in the predicted signal is used as the input to the ALF equation. In another example, the truncation differences between the surrounding pixels and the co-pixels in the predicted signal, as well as the truncation differences between the co-pixels and the current pixel in the predicted signal, are used as the inputs to the ALF equation.
[0266] Besides applying additional online ALF filter taps directly to the predicted signal, additional online ALF filter taps can also be applied to intermediate results obtained by feeding the predicted signal to fixed filters. Various fixed filters can be applied to filter the predicted signal to obtain intermediate results, which can concentrate the predicted signal information within a larger receptive field. For example, the two 13×13 diamond fixed filters used in the ALF of ECM can be used to filter the predicted signal to obtain intermediate results. When fixed filters are applied to the predicted signal, the block-level classification result can directly utilize the block-level classification result calculated on the signal immediately following the SAO, or it can be recalculated based on the predicted signal. When fixed filters are applied to the predicted signal, an intermediate result can be obtained using a fixed filter trained on one block-level classifier, or two or more intermediate results can be obtained using two or more fixed filters trained on two or more block-level classifiers. In video codec standards, several sets of fixed filters are typically prepared, and one set of fixed filters can be selected from them via the RDO process. For example, in ECM, a set of fixed filters (containing two 13×13 diamond fixed filters) is selected from two sets via the RDO process, and the set index is transmitted to the decoder. When a fixed filter is applied to the predicted signal, the group index of the predicted signal can be the same as the group index of the signal immediately following the SAO, or different from the group index of the signal immediately following the SAO based on a predefined criterion, or the group index of the predicted signal can be determined by the RDO process. In the first and second cases, it is not necessary to transmit the group index of the predicted signal to the decoder, while in the third case, it is necessary. When additional inline filter taps are applied to the intermediate results obtained by feeding the predicted signal to the fixed filter, various filter shapes can be used to extract information from the intermediate results. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15 As shown, various equation forms can be used to extract information from the intermediate results. In one example, the truncation difference between the surrounding pixels and the current pixel in the intermediate results is used as the input to the ALF equation. In another example, the truncation differences between the surrounding pixels and the co-pixels in the intermediate results, and the truncation differences between the co-pixels and the current pixel in the intermediate results, are used as the input to the ALF equation.
[0267] It should be noted that additional online ALF filter taps can be applied only to the predicted signal, or only to the intermediate results obtained by feeding the predicted signal to a fixed filter, or simultaneously to the predicted signal and the intermediate results obtained by feeding the predicted signal to a fixed filter.
[0268] In the second method, spatially neighboring pixels in the residual signal are proposed as inputs to the additional ALF equations. Various filter shapes can be used to extract information from the residual signal. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15 As shown, various equation forms can be used to extract information from the residual signal. In one example, the cropping of co-position pixels in the residual signal is used as input to the ALF equation.
[0269] Besides applying additional online ALF filter taps directly to the residual signal, additional online ALF filter taps can also be applied to intermediate results obtained by feeding the residual signal to a fixed filter. Various fixed filters can be applied to filter the residual signal to obtain intermediate results, which can concentrate the residual signal information within a larger receptive field. For example, the two 13×13 diamond-shaped fixed filters used in the ALF of an ECM can be used to filter the residual signal to obtain intermediate results. When a fixed filter is applied to the residual signal, the filtering result can be truncated to different ranges, such as (-1024, 1024), (-512, 512), (-256, 256), (-128, 128), etc. When a fixed filter is applied to the residual signal, the block-level classification result can directly utilize the block-level classification result calculated on the signal immediately following the SAO, or it can be recalculated based on the residual signal. When applying a fixed filter to the residual signal, an intermediate result can be obtained using a single fixed filter trained on a block-level classifier, or two or more intermediate results can be obtained using two or more fixed filters trained on two or more block-level classifiers. When applying a fixed filter to the residual signal, the group index of the residual signal can be the same as the group index of the signal immediately following the SAO, or different from the group index of the signal immediately following the SAO based on a predefined criterion, or determined by the RDO process. In the first and second cases, it is not necessary to transmit the group index of the residual signal to the decoder, while in the third case, it is necessary to transmit the group index of the residual signal to the decoder. When applying additional online filter taps to the intermediate results obtained by feeding the residual signal to the fixed filter, various filter shapes can be used to extract information from the intermediate results. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15As shown, various equation forms can be used to extract information from the intermediate results. In one example, the cropping results of co-pixels in the intermediate results are used as input to the ALF equation.
[0270] It should be noted that additional online ALF filter taps can be applied only to the residual signal, or only to the intermediate results obtained by feeding the residual signal to a fixed filter, or simultaneously to the residual signal and the intermediate results obtained by feeding the residual signal to a fixed filter.
[0271] In the third method, spatially neighboring pixels in the signal prior to SAO are proposed as inputs to the additional ALF equations. Various filter shapes can be used to extract information from the signal prior to SAO. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15 As shown, various equation forms can be used to extract information from the signal prior to the SAO. In one example, the truncation difference between the surrounding pixels and the current pixel in the signal prior to the SAO is used as the input to the ALF equation. In another example, the truncation differences between the surrounding pixels and the co-pixels in the signal prior to the SAO, and the truncation differences between the co-pixels and the current pixel in the signal prior to the SAO, are used as the inputs to the ALF equation.
[0272] Besides applying additional online ALF filter taps directly to the signal immediately preceding the SAO, additional online ALF filter taps can also be applied to intermediate results obtained by feeding the signal immediately preceding the SAO to a fixed filter. Various fixed filters can be applied to filter the signal immediately preceding the SAO to obtain intermediate results, which can concentrate the information of the signal immediately preceding the SAO into a larger receptive field. For example, the two 13×13 diamond fixed filters used in the ALF of ECM can be used to filter the signal immediately preceding the SAO to obtain intermediate results. When applying a fixed filter to the signal immediately preceding the SAO, the block classification result can directly utilize the block classification result calculated on the signal immediately following the SAO, or be recalculated based on the signal immediately preceding the SAO. When applying a fixed filter to the signal immediately preceding the SAO, one intermediate result can be obtained using a fixed filter trained on a single block classifier, or two or more intermediate results can be obtained using two or more fixed filters trained on two or more block classifiers. When a fixed filter is applied to a signal immediately preceding the SAO, the group index of the signal immediately preceding the SAO can be the same as the group index of the signal immediately following the SAO, or different from the group index of the signal immediately following the SAO based on a predefined standard, or the group index of the signal immediately preceding the SAO can be determined by an RDO procedure. In the first and second cases, it is not necessary to transmit the group index of the signal immediately preceding the SAO to the decoder, while in the third case, it is necessary to transmit the group index of the signal immediately preceding the SAO to the decoder. When additional inline filter taps are applied to the intermediate result obtained by feeding the signal immediately preceding the SAO to the fixed filter, various filter shapes can be used to extract information from the intermediate result. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15 As shown, various equation forms can be used to extract information from the intermediate results. In one example, the truncation difference between the surrounding pixels and the current pixel in the intermediate results is used as the input to the ALF equation. In another example, the truncation differences between the surrounding pixels and the co-pixels in the intermediate results, and the truncation differences between the co-pixels and the current pixel in the intermediate results, are used as the input to the ALF equation.
[0273] It should be noted that the additional online ALF filter tap can be applied only to the signal immediately preceding the SAO, or only to the intermediate result obtained by feeding the signal immediately preceding the SAO to the fixed filter, or simultaneously to the signal immediately preceding the SAO and the intermediate result obtained by feeding the signal immediately preceding the SAO to the fixed filter.
[0274] The fourth method proposes using information from the predicted signal, the residual signal, or the signal prior to SAO as input to the ALF equation. The methods proposed in the first, second, and third methods can be combined to implement the fourth method.
[0275] A new classifier that combines features from an edge-based classifier and features from a band-based classifier.
[0276] According to one or more embodiments of this disclosure, features from an edge-based classifier and features from a band-based classifier are combined to derive a new classifier for an online ALF filter. Different methods can be used to achieve this goal.
[0277] In the first method, the directionality D of the sub-blocks of the luminance component is first calculated, then the sum of the sample values of the sub-blocks is calculated and mapped to indices of a band-based classifier, and the class index of the sub-block is calculated as follows:
[0278] C = B * M D +D (17)
[0279] Where B is the index calculated with reference to the band-based classifier, and M... D This represents the total number of directivity values D. In one example, for a 2×2 luma block, the directivity D is calculated in the same way as D² in ECM, and B is calculated as...
[0280] B = (Total * 5) >> (Sample depth + 2) (18)
[0281] In the second method, the activity value A of the sub-blocks of the luminance component is first calculated, then the sum of the sample values of the sub-blocks is calculated and mapped to the index of a band-based classifier, and the class index of the sub-block is calculated as follows:
[0282] C = B * M A +A (19)
[0283] Where B is the index calculated with reference to the band-based classifier, and M... A This represents the total number of activity values A. In one example, for a 2×2 luma block, the activity value A is calculated in the same way as in the ECM. The same, and B is calculated as
[0284] B = (Total * 5) >> (Sample depth + 2) (20)
[0285] In the third method, it is proposed to first calculate the index of the sub-blocks of the brightness component by referring to an edge-based classifier, then calculate the sum of the sample values of the sub-blocks and map them to the index of the band-based classifier, and the class index of the sub-block is calculated as follows:
[0286] C = B * ME +E (21)
[0287] Where B is the index calculated with reference to the band-based classifier, and M... E This represents the total number of indices calculated using the edge-based classifier, where E is the index calculated using the edge-based classifier. In one example, for a 2×2 luma block, index E is calculated in the same way as C2 in the ECM, and B is calculated as...
[0288] B = (total * 2) >> (sample point depth + 2) (22)
[0289] Adjust the shape of the chroma ALF filter to match the shape of the luminance ALF filter.
[0290] In a third aspect of this disclosure, it is proposed to change the shape of the chroma ALF filter from a rhombus shape to something like... FIG. 17 The elongated cross shape shown is consistent with the shape of the brightness ALF filter.
[0291] Consider updating the shape of the online chroma ALF filter in the new ECM version, which takes both the spatially neighboring pixels in the chroma reconstruction signal and the fixed filter output as input. To unify the shape of the chroma ALF filter with that of the luma ALF filter, it can be done as follows: FIG. 21 , FIG. 22 or FIG. 23 That's how you adjust the shape of the online chroma ALF filter.
[0292] A new classifier using pixel values immediately preceding the deblocking filter.
[0293] According to one or more embodiments of this disclosure, a new classifier for an online ALF filter is derived using pixel values immediately preceding the deblocking filter. Different methods can be used to achieve this objective.
[0294] In the first method, the directivity D of the sub-block of the luminance component is first calculated. Then, the sum of the differences between the sample immediately following the SAO and the co-located sample immediately preceding the deblocking filter in the sub-block or the neighboring NxN window surrounding the sub-block is calculated and mapped to the difference index. The class index of the sub-block is calculated as follows:
[0295] C = Dig * M D +D (23)
[0296] Where Dif is the difference index, M D This represents the total number of directivity values (D). In one example, for a 2×2 luma block, the directivity D is calculated in the same way as D2 in ECM, and Dif is calculated as...
[0297] Dif = sum Dif>0?2:(sum Dif <0?0:1) (24)
[0298] Where, sum Dif It is the sum of the differences in a 2x2 luminance block, or the sum of the differences in the neighboring NxN (e.g., 8x8) windows surrounding the 2x2 luminance block.
[0299] In the second method, the activity value A of the sub-block of the luminance component is first calculated. Then, the sum of the differences between the sample immediately following the SAO and the co-located sample immediately preceding the deblocking filter in the sub-block or the neighboring NxN window surrounding the sub-block is calculated and mapped to the difference index. The class index of the sub-block is calculated as follows:
[0300] C = Dif * M A +A (25)
[0301] Where Dif is the difference index, M A This represents the total number of activity values A. In one example, for a 2×2 luma block, the activity value A is calculated in the same way as in the ECM. The same applies, and Dif is calculated according to equation (24).
[0302] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component by referring to an edge-based classifier, then calculate the sum of the differences between the sample immediately following the SAO and the co-sampling sample immediately preceding the deblocking filter in the sub-block or the neighboring NxN window surrounding the sub-block, and map this sum to the difference index. The class index of the sub-block is then calculated as follows:
[0303] C = Dif * M E +E (26)
[0304] Where Dif is the difference index, M E Represents the total number of indices calculated with reference to the edge-based classifier, where E is the index calculated with reference to the edge-based classifier. In one example, for a 2×2 luminance block, index E is calculated in the same way as C2 in ECM, and Dif is calculated according to equation (24).
[0305] In the fourth method, it is proposed to first calculate the band index B of the sub-block of the luminance component, then calculate the sum of the differences between the sample immediately following the SAO and the co-position sample immediately preceding the deblocking filter in the sub-block or the neighboring NxN window surrounding the sub-block, and map this sum to the difference index. The class index of the sub-block is then calculated as follows:
[0306] C = Dif * M B +B (27)
[0307] Where Dif is the difference index, MB This represents the total number of band values. In one example, for a 2×2 luminance block, the band index B is calculated as...
[0308] B = (Total * 8) >> (Sample depth + 2) (28)
[0309] And Dif is calculated according to equation (24).
[0310] The fifth method proposes to calculate the sum of the differences between the sample immediately following the SAO and the co-located sample immediately preceding the deblocking filter in the sub-block or the neighboring NxN window surrounding the sub-block, then map the sum of the differences to the difference index, and use the difference index as the class index.
[0311] The sixth method proposes to compute an edge-based or band-based classifier based on the sample values immediately preceding the deblocking filter. The computation method is the same as that of the original edge-based or band-based classifier computed based on the sample values following the SAO.
[0312] A new classifier that utilizes pixel values in the predicted signal
[0313] According to one or more embodiments of this disclosure, a novel classifier for an online ALF filter is derived using pixel values in the predicted signal. Different methods can be used to achieve this objective.
[0314] In the first method, the directivity D of the sub-block of the luminance component is first calculated. Then, the sum of the differences between the samples after SAO and the co-located samples in the predicted signal within the sub-block or the neighboring NxN window surrounding the sub-block is calculated and mapped to the difference index. The class index of the sub-block is calculated as follows:
[0315] C = Dif * M D +D (29)
[0316] Where Dif is the difference index, M D This represents the total number of directivity values (D). In one example, for a 2×2 luma block, the directivity D is calculated in the same way as D2 in ECM, and Dif is calculated as...
[0317] Dif = sum Dif >0?2:(sum Dif <0?0:1) (30)
[0318] Where, sum Dif It is the sum of the differences in a 2x2 luminance block, or the sum of the differences in the neighboring NxN (e.g., 8x8) windows surrounding the 2x2 luminance block.
[0319] In the second method, the activity value A of the sub-block of the luminance component is first calculated. Then, the sum of the differences between the samples after SAO and the corresponding samples in the predicted signal of the sub-block or the neighboring NxN window surrounding the sub-block is calculated and mapped to the difference index. The class index of the sub-block is calculated as follows:
[0320] C = Dif * M A +A (31)
[0321] Where Dif is the difference index, M A This represents the total number of activity values A. In one example, for a 2×2 luma block, the activity value A is calculated in the same way as in the ECM. The same applies, and Dif is calculated according to equation (30).
[0322] In the third method, it is proposed to first calculate the index of the sub-block of the brightness component by referring to an edge-based classifier, then calculate the sum of the differences between the samples after SAO and the co-located samples in the predicted signal in the sub-block or the neighboring NxN window around the sub-block, and map this sum to the difference index. The class index of the sub-block is then calculated as follows:
[0323] C = Dif * M E +E (32)
[0324] Where Dif is the difference index, M E Represents the total number of indices calculated with reference to the edge-based classifier, where E is the index calculated with reference to the edge-based classifier. In one example, for a 2×2 luminance block, index E is calculated in the same way as C2 in ECM, and Dif is calculated according to equation (30).
[0325] In the fourth method, it is proposed to first calculate the band index B of the sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples in the predicted signal of the sub-block or the neighboring NxN window surrounding the sub-block, and map this sum to the difference index. The class index of the sub-block is then calculated as follows:
[0326] C = Dif * M B +B (33)
[0327] Where Dif is the difference index, M B This represents the total number of band values. In one example, for a 2×2 luminance block, the band index B is calculated as...
[0328] B = (Total * 8) >> (Sample depth + 2) (34)
[0329] And Dif is calculated according to equation (30).
[0330] The fifth method proposes to calculate the sum of the differences between the samples after SAO and the co-located samples in the predicted signal in the sub-block or the neighboring NxN window surrounding the sub-block, and then map the sum of the differences to the difference index, and use the difference index as the class index.
[0331] The sixth method proposes to calculate an edge-based classifier or a band-based classifier based on sample values in the predicted signal. The calculation method is the same as that of the original edge-based or band-based classifier calculated based on sample values after SAO.
[0332] A new classifier utilizing pixel values in the residual signal
[0333] According to one or more embodiments of this disclosure, a novel classifier for an online ALF filter is derived using pixel values in the residual signal. Different methods can be used to achieve this objective.
[0334] In the first method, the directionality D of the sub-block of the luminance component is first calculated, then the sum of pixel values in the residual signal of the sub-block or the neighboring NxN window surrounding the sub-block is calculated and mapped to the residual index, and the class index of the sub-block is calculated as follows:
[0335] C = Resi * M D +D (35)
[0336] Where Resi is the residual index, M D This represents the total number of directivity values (D). In one example, for a 2×2 luma block, the directivity D is calculated in the same way as D² in ECM, and Resi is calculated as...
[0337] Resi = sum Resi >0?2:(sum Resi <0?0:1) (36)
[0338] Here, is the sum of pixel values in the residual signal of the 2x2 luminance block, or the sum of pixel values in the residual signal of the neighboring NxN (e.g., 8x8) window surrounding the 2x2 luminance block.
[0339] In the second method, the activity value A of the sub-block of the luminance component is first calculated. Then, the sum of pixel values in the residual signal of the sub-block or the surrounding NxN window is calculated and mapped to the residual index. The class index of the sub-block is calculated as follows:
[0340] C = Resi * M A +A (37)
[0341] Where Resi is the residual index, M AThis represents the total number of activity values A. In one example, for a 2×2 luma block, the activity value A is calculated in the same way as in the ECM. The same applies, and Resi is calculated according to equation (36).
[0342] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component by referring to an edge-based classifier, then calculate the sum of pixel values in the residual signal of the sub-block or the neighboring NxN window surrounding the sub-block and map it to the residual index, and the class index of the sub-block is calculated as follows:
[0343] C = Resi * M E +E (38)
[0344] Where Resi is the residual index, and M is... E Represents the total number of indices calculated with reference to the edge-based classifier, where E is the index calculated with reference to the edge-based classifier. In one example, for a 2×2 luminance block, index E is calculated in the same way as C2 in ECM, and Resi is calculated according to equation (36).
[0345] In the fourth method, it is proposed to first calculate the band index B of the sub-block of the luminance component, then calculate the sum of pixel values in the residual signal of the sub-block or the neighboring NxN window surrounding the sub-block and map it to the residual index, and the class index of the sub-block is calculated as follows:
[0346] C = Resi * M B +B (39)
[0347] Where Resi is the residual index, M B This represents the total number of band values. In one example, for a 2×2 luminance block, the band index B is calculated as...
[0348] B = (Total * 8) >> (Sample depth + 2) (40)
[0349] And Resi is calculated according to equation (36).
[0350] In the fifth method, it is proposed to calculate the sum of pixel values in the residual signal in the sub-block or the neighboring NxN window surrounding the sub-block, then map the sum of residual values to the residual index, and use the residual index as the class index.
[0351] The sixth method proposes first calculating the sum of the absolute values of pixel values in the residual signal of the sub-block, or calculating the sum of the absolute values of pixel values in the residual signal of the neighboring N×N window surrounding the sub-block, and mapping it to the absolute value Resi of the residual index. AbsoThen, the sum of pixel values in the residual signal of the sub-block or the sum of pixel values in the residual signal of the neighboring N×N window around the sub-block is calculated and mapped to the symbol Resi of the residual index. Sign And the class index of the sub-block is calculated as
[0352] C = Resi Sign *M Abso +Resi Abso (41)
[0353] Among them, M Abso This represents the total number of absolute values of the residual index. In one example, for a 2×2 luminance block, the absolute value of the residual index, Reso, is... Abso Calculated as
[0354] Resi Abso =(Sum Abso *8)>>(sample point depth + 2) (42)
[0355] Among them, Sum Abso It is the sum of the absolute values of the pixel values in the residual signal of the 2×2 luma block, or the sum of the absolute values of the pixel values in the residual signal of the neighboring N×N (e.g., 8×8) windows surrounding the 2×2 luma block, and Resi Sign It is calculated according to equation (36).
[0356] The seventh method proposes a reference edge-based classifier that calculates the index of a sub-block based on the pixel values in the residual signal, and then uses that index as the class index.
[0357] The eighth method proposes a reference edge-based classifier that calculates the index of a sub-block based on the absolute value of the pixel value in the residual signal, and then uses that index as the class index.
[0358] A new classifier using the pixel values immediately preceding SAO
[0359] According to one or more embodiments of this disclosure, a new classifier for an online ALF filter is derived using the pixel values immediately preceding the SAO. Different methods can be used to achieve this objective.
[0360] In the first method, the directivity D of the sub-block of the luminance component is first calculated. Then, the sum of the differences between the sample immediately following the SAO and the co-located sample immediately preceding the SAO in the sub-block or the neighboring NxN window surrounding the sub-block is calculated and mapped to the difference index. The class index of the sub-block is calculated as follows:
[0361] C = Dig * M D +D (43)
[0362] Where Dif is the difference index, M D This represents the total number of directivity values (D). In one example, for a 2×2 luma block, the directivity D is calculated in the same way as D2 in ECM, and Dif is calculated as...
[0363] Dif = sum Dif >0?2:(sum Dif <0?0:1) (44)
[0364] Where, sum Dif It is the sum of the differences in a 2x2 luminance block, or the sum of the differences in the neighboring NxN (e.g., 8x8) windows surrounding the 2x2 luminance block.
[0365] In the second method, the activity value A of the sub-block of the luminance component is first calculated. Then, the sum of the differences between the sample immediately following the SAO and the corresponding sample immediately preceding the SAO in the sub-block or the neighboring NxN window surrounding the sub-block is calculated and mapped to the difference index. The class index of the sub-block is then calculated as follows:
[0366] C = Dif * M A +A (45)
[0367] Where Dif is the difference index, M A This represents the total number of activity values A. In one example, for a 2×2 luma block, the activity value A is calculated in the same way as in the ECM. The same applies, and Dif is calculated according to equation (42).
[0368] In the third method, the index of the sub-block of the luminance component is first calculated with reference to an edge-based classifier. Then, the sum of the differences between the sample immediately following the SAO and the co-located sample immediately preceding the SAO in the sub-block or the neighboring NxN window surrounding the sub-block is calculated and mapped to the difference index. The class index of the sub-block is then calculated as follows:
[0369] C = Dif * M E +E (46)
[0370] Where Dif is the difference index, M E Represents the total number of indices calculated with reference to the edge-based classifier, where E is the index calculated with reference to the edge-based classifier. In one example, for a 2×2 luminance block, index E is calculated in the same way as C2 in ECM, and Dif is calculated according to equation (44).
[0371] In the fourth method, it is proposed to first calculate the band index B of the sub-block of the luminance component, then calculate the sum of the differences between the sample immediately following the SAO and the co-located sample immediately preceding the SAO in the sub-block or the neighboring NxN window surrounding the sub-block, and map this sum to the difference index. The class index of the sub-block is then calculated as follows:
[0372] C = Dif * M B +B (47)
[0373] Where Dif is the difference index, M B This represents the total number of band values. In one example, for a 2×2 luminance block, the band index B is calculated as...
[0374] B = (Total * 8) >> (Sample depth + 2) (48)
[0375] And Dif is calculated according to equation (44).
[0376] The fifth method proposes to calculate the sum of the differences between the sample immediately following the SAO and the co-located sample immediately preceding the SAO in the sub-block or the neighboring NxN window surrounding the sub-block, then map the sum of the differences to the difference index, and use the difference index as the class index.
[0377] The sixth method proposes to calculate an edge-based classifier or a band-based classifier based on the sample values immediately preceding the SAO. The calculation method is the same as the original method for calculating an edge-based classifier or a band-based classifier based on the sample values following the SAO.
[0378] A new classifier using chroma pixel values
[0379] According to one or more embodiments of this disclosure, a novel classifier for an online ALF filter is derived using chroma pixel values. Different methods can be used to achieve this goal.
[0380] In the first method, it is proposed to first calculate the band index B of the sub-block of the luminance component. Y Then calculate the indexed B of the corresponding U and V components. U and B V And the class index of the sub-block is calculated as
[0381] C = B Y *M U *M V +B U *M V +B V (49)
[0382] Among them, B Y B U and BV It refers to the Y, U, and V indices calculated by a band-based classifier, M U and M V This represents the total number of U and V band index values. In one example, for a 2×2 luminance block, B Y B U and B V Calculated as
[0383] B Y =(sumY*6)>>(sample bit depth+2) (50)
[0384] B U =(sumU*2)>>(sample bit depth+2) (51)
[0385] B V =(sumV*2)>>(sample bit depth+2) (52)
[0386] Chromaticity information immediately preceding deblocking, in prediction, in residuals, or immediately preceding SAO is used as additional chromaticity ALF input.
[0387] According to one or more embodiments of this disclosure, chromaticity information immediately preceding deblocking, in prediction, in residuals, or immediately preceding SAO is used as additional chromaticity ALF equation inputs. This objective can be achieved using different methods.
[0388] In the first method, spatially neighboring pixels in the chroma prediction signal are proposed as inputs to an additional chroma ALF equation. Various filter shapes can be used to extract information from the chroma prediction signal. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15 As shown, various equation forms can be used to extract information from the chroma prediction signal. In one example, the truncation difference between the surrounding pixels and the current chroma pixel in the chroma prediction signal is used as the input to the chroma ALF equation. In another example, the truncation differences between the surrounding pixels and the co-pixels in the chroma prediction signal, as well as the truncation differences between the co-pixels and the current chroma pixel in the chroma prediction signal, are used as the input to the chroma ALF equation.
[0389] In the second method, spatially neighboring pixels in the chroma residual signal are proposed as inputs to an additional chroma ALF equation. Various filter shapes can be used to extract information from the chroma residual signal. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15As shown, information can be extracted from the chroma residual signal using various equation forms. In one example, the truncation of co-pixels in the chroma residual signal is used as input to the chroma ALF equation.
[0390] In the third method, spatially neighboring pixels in the signal immediately preceding the chroma SAO are proposed as inputs to the additional chroma ALF equation. Various filter shapes can be used to extract information from the signal immediately preceding the chroma SAO. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15 As shown, various equation forms can be used to extract information from the signal immediately preceding the chroma SAO. In one example, the truncation difference between the surrounding pixels in the signal immediately preceding the chroma SAO and the current chroma pixel is used as the input to the chroma ALF equation. In another example, the truncation differences between the surrounding pixels in the signal immediately preceding the chroma SAO and the co-pixels in the signal immediately preceding the chroma SAO, as well as the truncation differences between the co-pixels in the signal immediately preceding the chroma SAO and the current chroma pixel, are used as the input to the chroma ALF equation.
[0391] The fourth method proposes using spatially neighboring pixels in the signal immediately preceding chroma deblocking as input to an additional chroma ALF equation. Various filter shapes can be used to extract information from the signal immediately preceding chroma deblocking. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15 As shown, various equation forms can be used to extract information from the signal immediately preceding chroma deblocking. In one example, the truncation difference between the surrounding pixels of the signal immediately preceding chroma deblocking and the current chroma pixel is used as input to the chroma ALF equation. In another example, the truncation differences between the surrounding pixels of the signal immediately preceding chroma deblocking and the co-pixels of the signal immediately preceding chroma deblocking, as well as the truncation differences between the co-pixels of the signal immediately preceding chroma deblocking and the current chroma pixel, are used as input to the chroma ALF equation.
[0392] The fifth method proposes using information from the chromaticity prediction, residuals, signals prior to SAO, or signals prior to deblocking as inputs to the chromaticity ALF equation. The methods proposed in the first, second, third, and fourth methods can be combined to implement the fifth method.
[0393] Luminance information immediately before deblocking, in prediction, in residuals, or immediately before SAO is used as additional CCALF input.
[0394] According to one or more embodiments of this disclosure, luminance information immediately preceding deblocking, in prediction, in residuals, or immediately preceding SAO is used as additional CCALF equation inputs. This objective can be achieved using different methods.
[0395] In the first method, spatially neighboring pixels in the brightness prediction signal are proposed as inputs to an additional CCALF equation. Various filter shapes can be used to extract information from the brightness prediction signal. For example, the filter shape could be 3×4, such as... FIG. 13 As shown, information can be extracted from the luminance prediction signal using various equation forms. In one example, the difference between the surrounding pixels in the luminance prediction signal and the currently corresponding luminance pixel is used as the input to the CCALF equation. In another example, the differences between the surrounding pixels in the luminance prediction signal and the co-pixels in the currently corresponding luminance prediction signal, as well as the differences between the co-pixels in the currently corresponding luminance prediction signal and the currently corresponding luminance pixel, are used as the inputs to the CCALF equation.
[0396] In the second method, spatially neighboring pixels in the luminance residual signal are proposed as inputs to an additional CCALF equation. Various filter shapes can be used to extract information from the luminance residual signal. For example, the filter shape could be 3×4, such as... FIG. 13 As shown, information in the luminance residual signal can be extracted using various equation forms. In one or more examples, co-position pixels in the luminance residual signal are used as inputs to the CCALF equations.
[0397] In the third method, spatially neighboring pixels in the signal immediately preceding the Luminance SAO are proposed as input to an additional CCALF equation. Various filter shapes can be used to extract information from the signal immediately preceding the Luminance SAO. For example, the filter shape could be 3×4, such as... FIG. 13 As shown, various equation forms can be used to extract information from the signal immediately preceding the luminance SAO. In one example, the difference between the surrounding pixels in the signal immediately preceding the luminance SAO and the currently corresponding luminance pixel is used as input to the CCALF equation. In another example, the differences between the surrounding pixels in the signal immediately preceding the luminance SAO and the corresponding pixels in the signal preceding the current luminance SAO, as well as the differences between the corresponding pixels in the signal preceding the current luminance SAO and the currently corresponding luminance pixel, are used as inputs to the CCALF equation.
[0398] In the fourth method, spatially neighboring pixels in the signal immediately preceding the luma deblocking are proposed as input to an additional CCALF equation. Various filter shapes can be used to extract information from the signal immediately preceding the luma deblocking. For example, the filter shape could be 3×4, such as...FIG. 13 As shown, various equation forms can be used to extract information from the signal immediately preceding the luma deblocking. In one example, the difference between the surrounding pixels in the signal immediately preceding the luma deblocking and the currently corresponding luma pixel is used as input to the CCALF equation. In another example, the differences between the surrounding pixels in the signal immediately preceding the luma deblocking and the corresponding pixels in the signal preceding the current luma deblocking, as well as the differences between the corresponding pixels in the signal preceding the current luma deblocking and the currently corresponding luma pixel, are used as input to the CCALF equation.
[0399] The fifth method proposes using information from the brightness prediction, residuals, signals prior to SAO, or signals prior to deblocking as inputs to the CCALF equations. The methods proposed in the first, second, third, and fourth methods can be combined to implement the fifth method.
[0400] A new classifier utilizing encoded pattern information
[0401] According to one or more embodiments of this disclosure, a new classifier for an online ALF filter is derived using coding mode information, such as whether the coding block is encoded and decoded in a skip mode, or whether the coding block is encoded and decoded in an intra-frame, inter-frame P, or inter-frame B mode. Different methods can be used to achieve this objective.
[0402] In the first approach, it is proposed to record whether the encoded block uses a skip mode during encoding and decoding, and then use this information to design a new classifier. In one example, a classifier with two classes (corresponding to true or false skip mode, respectively) is added as the new classifier. In another example, a classifier that combines the skip mode information with EO or BO is added as the new classifier.
[0403] The second approach proposes recording whether a coded block is encoded or decoded using intra-frame mode, inter-frame P mode, or inter-frame B mode during the encoding and decoding process, and then using this information to design a new classifier. In one example, a new classifier is added with three classes (corresponding to intra-frame mode, inter-frame P mode, or inter-frame B mode, respectively). In another example, a new classifier is added that combines intra-frame, inter-frame P, or inter-frame B mode information with EO or BO.
[0404] The third approach proposes designing a new classifier by simultaneously considering information from both coding modes (whether the coding block uses skip mode for encoding / decoding, and whether the coding block uses intra-frame, inter-frame P, or inter-frame B modes for encoding / decoding). The methods proposed in the first and second approaches can be combined to implement the third approach.
[0405] The row buffer for additional ALF inputs is reduced.
[0406] According to one or more embodiments of this disclosure, when the online ALF filter uses samples from stages such as 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO as additional inputs, line buffers are required to store these samples. To reduce the line buffer requirements for these additional inputs, based on the current line buffer settings in the VVC, the four lines corresponding to luma samples and the two lines corresponding to chroma samples above the horizontal CTU boundary are assumed to be default values, which saves these line buffers. Different methods can be used to achieve this goal.
[0407] In the first approach, based on the current line buffer settings in the VVC, it is proposed to assume that the four lines of luminance residual samples and two lines of chrominance residual samples above the horizontal CTU boundary are zero values, and to assume that the four lines of luminance samples and two lines of chrominance samples above the horizontal CTU boundary from the following stages are equivalent sample values from the stages immediately following the SAO.
[0408] In the second method, based on the current line buffer settings in VVC, it is proposed that the four lines of luminance samples and two lines of chrominance samples above the horizontal CTU boundary from the following stages are assumed to be set in a repeating manner using the corresponding nearest sample value on the horizontal CTU boundary: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO.
[0409] In the third method, based on the current line buffer settings in VVC, it is proposed that the four rows of luminance samples and two rows of chrominance samples above the horizontal CTU boundary from the following stages are assumed to be set in a mirror manner: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO. The first row of luminance samples and the first row of chrominance samples above the horizontal CTU boundary are assumed to be the corresponding sample values on the horizontal CTU boundary, and the second row of luminance samples and the second row of chrominance samples above the horizontal CTU boundary are assumed to be the corresponding sample values in the first row of samples below the horizontal CTU boundary, and so on.
[0410] It should be noted that the 4 rows of luminance samples and 2 rows of chrominance samples above the horizontal CTU boundary are the current VVC line buffer settings, and the specific values can be adjusted according to the custom settings.
[0411] The row buffer for additional CCALF inputs is reduced.
[0412] According to one or more embodiments of this disclosure, when the CCALF filter uses samples from stages such as 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO as additional inputs, line buffers are required to store these samples. To reduce the line buffer requirements for these additional inputs, the four lines corresponding to luminance samples above the horizontal CTU boundary are assumed to be default values, based on the current line buffer settings in the VVC, which saves these line buffers. Different methods can be used to achieve this goal.
[0413] In the first approach, based on the current line buffer settings in the VVC, it is proposed to assume that the four line luminance residual samples above the horizontal CTU boundary are zero values, and assume that the four line luminance samples above the horizontal CTU boundary from the following stages are equivalent to the sample values from the stages immediately following the SAO.
[0414] In the second method, based on the current line buffer settings in VVC, it is proposed that the four line luminance samples above the horizontal CTU boundary from the following stages are assumed to be set in a repeating manner using the corresponding nearest sample value on the horizontal CTU boundary: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO.
[0415] In the third method, based on the current line buffer settings in VVC, it is proposed that the four lines of luminance samples above the horizontal CTU boundary from the following stages are assumed to be set in a mirror manner: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO. The first line of luminance samples above the horizontal CTU boundary is assumed to be the corresponding sample value on the horizontal CTU boundary, the second line of luminance samples above the horizontal CTU boundary is assumed to be the corresponding sample value in the first line of samples below the horizontal CTU boundary, and so on.
[0416] It should be noted that the four lines of luminance samples above the horizontal CTU boundary are the current VVC line buffer settings, and the specific values can be adjusted according to the custom settings.
[0417] Sample filling for additional ALF inputs
[0418] According to one or more embodiments of this disclosure, when the online ALF filter uses samples from stages such as 1) samples immediately before deblocking, 2) predicted samples, 3) residual samples, and 4) samples immediately before SAO as additional input, sample padding is performed when the filter shape of the additional input (whose center position is aligned with the sample to be filtered) crosses the virtual boundary (line buffer boundary) or the picture (strip, tile) boundary. This objective can be achieved using different methods.
[0419] In the first method, symmetrical sample padding is applied when the filter shape of the additional input (whose center is aligned with the sample to be filtered) crosses the virtual boundary (line buffer boundary) or the image (strip, tile) boundary. For example, assuming an online ALF filter uses residual samples as additional input, the filter shape of the fixed filter to be applied to the residual signal or the filter shape of the online filter directly applied to the residual signal is 7×7, and the filter shape of the residual signal (whose center is aligned with the sample to be filtered) crosses the line buffer boundary, such as... FIG. 18 The diagram shows symmetrical sample point filling, where p 12 Masking the co-position residual pixels of the sample to be filtered, from p0 to p 24 These are the original residual samples, from p′0 to p′. 24 These are the modified residual sample values. In summary, symmetric sample filling involves modifying both the additional input samples that are not on the same boundary side as the additional input samples that are in the same location ...
[0420] In the second method, when the shape of the additional input filter (whose center is aligned with the sample to be filtered) crosses the virtual boundary (line buffer boundary) or the image (strip, tile) boundary, repeating sample padding is applied. Repeating padding means that additional input samples that are not on the same boundary side as the sample to be filtered are filled in the same way as symmetrical sample padding, while additional input samples that are on the same boundary side as the sample to be filtered remain unchanged.
[0421] ALF fixed filter with additional classifier
[0422] According to one or more embodiments of this disclosure, additional ALF fixed filter banks are trained using a classifier, a residual-based classifier, or the like. The outputs of these additional ALF fixed filter banks are then used as inputs to additional online ALF filters. Different methods can be used to achieve this goal.
[0423] In the first approach, different ALF fixed filter banks are first trained using different band classifiers. Different band classifiers can be defined based on different window sizes. For example, there are two band classifiers. For the first band classifier, the sum of sample values from a 2×2 luminance block is calculated and mapped to the band classifier index, as follows:
[0424] class_index = (total * 25) >> (sample point depth + 2)
[0425] For the second band classifier, the sum of sample values in the neighboring 8×8 windows surrounding the 2×2 brightness block is calculated and mapped to the band classifier index, as follows:
[0426] class_index = (total * 25) >> (sample point depth + 6)
[0427] Different band classifiers can also be defined based on different numbers of categories. For example, there are two band classifiers. For the first band classifier, the sum of the sample values of the 2×2 brightness block is calculated and mapped to the band classifier index, as follows:
[0428] class_index = (total * 25) >> (sample point depth + 2)
[0429] The number of categories with the classifier is 25. For the second classifier, the sum of the sample values of the 2×2 brightness block is calculated and mapped to the classifier index, as shown below:
[0430] class_index = (total * 100) >> (sample point depth + 2)
[0431] The number of classes with the classifier is 100. Different ALF fixed filter taps can be used. For example, the fixed filter is a 13×13 rhombus. After training different ALF fixed filter banks based on different classifiers, intermediate results can be obtained by feeding the reconstructed pixel values into the newly trained ALF fixed filter. Then, the online ALF filter can use the intermediate results as additional input. Various filter shapes can be used to extract information from the intermediate results. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15 As shown, various equation forms can be used to extract information from intermediate results. For example, the truncation difference between surrounding pixels and the current pixel in the intermediate result is used as input to an additional online ALF filter.
[0432] In the second method, different ALF fixed filter banks are first trained using different residual-based classifiers. Different residual-based classifiers can be defined based on different window sizes. For example, there are two residual-based classifiers. For the first residual-based classifier, the sum of the absolute values of residual samples in the neighboring 8×8 windows surrounding the 2×2 brightness block is calculated and mapped to the residual-based classifier index, as follows:
[0433] classIdx = total >> (sample point depth - 4)
[0434] The value of classIdx ranges from 0 to 24. For the second residual-based classifier, the sum of the absolute values of the residual samples in the neighboring 12×12 windows surrounding the 2×2 brightness block is calculated and mapped to the residual-based classifier index, as shown below:
[0435] classIdx = total >> (sample point depth - 4)
[0436] The value of `classIdx` ranges from 0 to 24. Different residual-based classifiers can also be defined based on different numbers of classes. For example, there are two residual-based classifiers. For the first residual-based classifier, the sum of the absolute values of the residual samples in the neighboring 8×8 windows surrounding the 2×2 brightness block is calculated and mapped to the residual-based classifier index, as shown below:
[0437] classIdx = total >> (sample point depth - 4)
[0438] The value of classIdx ranges from 0 to 24. For the second residual-based classifier, the sum of the absolute values of the residual samples in the neighboring 8×8 windows surrounding the 2×2 brightness block is calculated and mapped to the residual-based classifier index, as shown below:
[0439] classIdx = total >> (sample point depth - 4)
[0440] The value of `classIdx` ranges from 0 to 49. Different ALF fixed filter taps can be used. For example, the ALF fixed filter is a 13×13 rhombus. After training different ALF fixed filter banks based on different residual-based classifiers, intermediate results can be obtained by feeding the reconstructed pixel values into the newly trained ALF fixed filter. Then, the online ALF filter can use the intermediate results as additional input. Various filter shapes can be used to extract information from the intermediate results. For example, the filter shape can be 1×1, 3×3, or 5×5, such as... FIG. 15As shown, various equation forms can be used to extract information from intermediate results. For example, the truncation difference between surrounding pixels and the current pixel in the intermediate result is used as input to an additional online ALF filter.
[0441] In the third approach, the methods proposed in the first and second approaches can be combined. For example, two sets of ALF fixed filters can be trained using a classifier with edge detection and a residual-based classifier. The outputs of these two newly trained ALF fixed filters are then used as additional inputs to the online ALF filter. It should be noted that, in addition to the newly trained ALF fixed filters, the original design of the ALF in ECM already includes two ALF fixed filters trained based on an edge-based classifier.
[0442] ALF fixed filter with additional input
[0443] According to one or more embodiments of this disclosure, when training the ALF fixed filter, spatially neighboring pixels immediately preceding the deblocking filter, spatially neighboring pixels in the predicted signal, spatially neighboring pixels in the residual signal, or spatially neighboring pixels immediately preceding the SAO are used as additional ALF fixed filter inputs. This objective can be achieved using different methods.
[0444] In the first method, when training the ALF fixed filter, the spatially neighboring pixels immediately preceding the deblocking filter are used as input to the additional ALF fixed filter. Various filter shapes can be used to extract information from the spatially neighboring pixels immediately preceding the deblocking filter. For example, the filter shape could be as follows: FIG. 15 The diagram shows a 1×1, 3×3, or 5×5 rhombus, or a 13×13 rhombus. Various equation forms can be used to extract information from spatially neighboring pixels immediately preceding the deblocking filter. For example, the truncation difference between the current pixel and the surrounding pixels immediately preceding the deblocking filter is used as input to an additional ALF fixed filter.
[0445] In the second method, when training the ALF fixed filter, spatially neighboring pixels in the predicted signal are used as inputs to the additional ALF fixed filter. Various filter shapes can be used to extract information from the predicted signal. For example, the filter shape could be as follows: FIG. 15 The diagram shows a 1×1, 3×3, or 5×5 rhombus, or a 13×13 rhombus. Various equation forms can be used to extract information from the predicted signal. For example, the truncation difference between the surrounding pixels and the current pixel in the predicted signal is used as the input to an additional ALF fixed filter.
[0446] In the third method, when training the ALF fixed filter, spatially neighboring pixels in the residual signal are used as inputs to an additional ALF fixed filter. Various filter shapes can be used to extract information from the residual signal. For example, the filter shape could be as follows: FIG. 15 The diagram shows a 1×1, 3×3, or 5×5 rhombus, or a 13×13 rhombus. Various equation forms can be used to extract information from the residual signal. For example, the truncation of surrounding pixels in the residual signal is used as input to an additional ALF fixed filter.
[0447] In the fourth method, when training the ALF fixed filter, the spatially neighboring pixels immediately preceding the SAO are used as input to the additional ALF fixed filter. Various filter shapes can be used to extract information from the spatially neighboring pixels immediately preceding the SAO. For example, the filter shape could be as follows: FIG. 15 The diagram shows a 1×1, 3×3, or 5×5 rhombus, or a 13×13 rhombus. Various equation forms can be used to extract information from the spatially neighboring pixels immediately preceding the SAO. For example, the truncation difference between the surrounding pixels immediately preceding the SAO and the current pixel is used as input to an additional ALF fixed filter.
[0448] In the fifth method, the methods proposed in the first, second, third, and fourth methods can be combined. For example, when training the ALF fixed filter, the spatially neighboring pixels immediately preceding the deblocking filter and the spatially neighboring pixels in the residual signal are used as additional ALF fixed filter inputs.
[0449] Chromatic ALF with Sub-block Level Filter Adaptation
[0450] According to one or more embodiments of this disclosure, sub-block level filters are adaptively applied to a chroma ALF, wherein an edge-based classifier, a band-based classifier, or a residual-based classifier is used in the chroma ALF. Different methods can be used to achieve this objective.
[0451] In the first approach, an edge-based classifier is used in the chroma ALF. The edge-based classifier can be defined at different sub-block levels, such as 4×4, 2×2, or 1×1 block sizes. Different methods can be used to compute the edge-based classifier. In the first example, the edge-based classifier is computed based on the reconstructed signal in the luma component. For example, an edge-based classifier used in the luma ALF and computed based on the signal immediately following the SAO in the luma component is multiplexed in the chroma ALF. In the second example, the edge-based classifier is computed based on the reconstructed signal in the chroma component. For example, first, an edge-based classifier is computed based on the signal immediately following the SAO in the Cb component, and then another edge-based classifier is computed based on the signal immediately following the SAO in the Cr component. Finally, an edge-based classifier combining the first and second edge-based classifiers is used in the chroma ALF. It should be noted that when calculating an edge-based classifier based on the signal immediately following the SAO in the Cb or Cr component, the calculation process can be the same as when calculating an edge-based classifier based on the signal immediately following the SAO in the luminance component. In the third example, an edge-based classifier combining the edge-based classifier calculated in the first example and the edge-based classifier calculated in the second example is used for the chroma ALF. It should be noted that when the edge-based classifier in the luminance ALF is reused in the chroma ALF, or when the edge-based classifier is recalculated for the chroma ALF based on the chroma component, the specific number of classes used in the chroma ALF can be the same as or different from the number of classes used in the luminance ALF. For example, the number of classes used in the luminance ALF is 25, which is obtained by combining activity (5) and directionality (5); while the number of classes used in the chroma ALF is 10, which is obtained by combining activity (2) and directionality (5).
[0452] In the second approach, a band-based classifier is used in the chroma ALF. The band-based classifier can be defined at different sub-block levels, such as 4×4, 2×2, or 1×1 block sizes. Different methods can be used to compute the band-based classifier. In the first example, the band-based classifier is computed based on the reconstructed signal in the luma component. For example, a band-based classifier used in the luma ALF and computed based on the signal immediately following the SAO in the luma component is multiplexed in the chroma ALF. In the second example, the band-based classifier is computed based on the reconstructed signal in the chroma component. For example, first, a band-based classifier is computed based on the signal immediately following the SAO in the Cb component, and then another band-based classifier is computed based on the signal immediately following the SAO in the Cr component. Finally, a band-based classifier combining the first and second band-based classifiers is used in the chroma ALF. It should be noted that when calculating a band-based classifier based on the signal immediately following the SAO in the Cb or Cr component, the calculation process can be the same as calculating a band-based classifier based on the signal immediately following the SAO in the luminance component. In the third example, a band-based classifier combining the band-based classifiers calculated in the first example and the second example is used for the chroma ALF. It should be noted that when the band-based classifier in the luminance ALF is reused in the chroma ALF, or when the band-based classifier is recalculated for the chroma ALF based on the chroma components, the specific number of classes used in the chroma ALF can be the same as or different from the number of classes used in the luminance ALF. For example, the number of classes used in the luminance ALF is 25, while the number of classes used in the chroma ALF is 10.
[0453] In the third approach, a residual-based classifier is used in the chroma ALF. The residual-based classifier can be defined at different sub-block levels, such as 4×4, 2×2, or 1×1 block sizes. Different methods can be used to compute the residual-based classifier. In the first example, the residual-based classifier is computed based on the residual signal in the luminance component. For example, a residual-based classifier used in the luminance ALF and computed based on the residual signal in the luminance component is multiplexed in the chroma ALF. In the second example, a residual-based classifier is computed based on the residual signal in the chroma component. For example, first, a residual-based classifier is computed based on the residual signal in the Cb component, and then another residual-based classifier is computed based on the residual signal in the Cr component. Finally, a residual-based classifier combining the first and second residual-based classifiers is used in the chroma ALF. It should be noted that when calculating a residual-based classifier based on the residual signals in the Cb or Cr components, the calculation process can be the same as when calculating a residual-based classifier based on the residual signals in the luminance component. In the third example, a residual-based classifier combining the residual-based classifiers calculated in the first example and the second example is used for the chroma ALF. It should be noted that when the residual-based classifier in the luminance ALF is reused in the chroma ALF, or when the residual-based classifier is recalculated for the chroma ALF based on the chroma components, the specific number of classes used in the chroma ALF can be the same as or different from the number of classes used in the luminance ALF. For example, the number of classes used in the luminance ALF is 25, while the number of classes used in the chroma ALF is 10.
[0454] In the fourth approach, the methods proposed in the first, second, and third approaches can be combined. For example, in chroma ALF, the edge-based classifier proposed in the first approach and the band-based classifier proposed in the second approach can be used simultaneously.
[0455] CCALF with Sub-block Level Filter Adaptation
[0456] According to one or more embodiments of this disclosure, sub-block level filters are adaptively applied to CCALF, wherein an edge-based classifier, a band-based classifier, or a residual-based classifier is used in CCALF. Different methods can be used to achieve this goal.
[0457] In the first approach, an edge-based classifier is used in the CCALF. The edge-based classifier can be defined at different sub-block levels, such as 4×4, 2×2, or 1×1 block sizes. Different methods can be used to compute the edge-based classifier. In the first example, the edge-based classifier is computed based on the reconstructed signal in the luminance component. For example, the edge-based classifier used in the luminance ALF and computed based on the signal immediately following the SAO in the luminance component is reused in both the Cb and Cr components of the CCALF. In the second example, the edge-based classifier is computed based on the reconstructed signal in the chrominance component. For example, for the Cb component of the CCALF, the edge-based classifier is computed based on the signal immediately following the SAO in the Cb component; for the Cr component of the CCALF, the edge-based classifier is computed based on the signal immediately following the SAO in the Cr component. It should be noted that when computed based on the signal immediately following the SAO in the Cb or Cr component, the computation process can be the same as when computed based on the signal immediately following the SAO in the luminance component. In the third example, an edge-based classifier combining the edge-based classifier calculated in the first example and the edge-based classifier calculated for the CCALF in the second example is used for the CCALF in the Cb component; an edge-based classifier combining the edge-based classifier calculated in the first example and the edge-based classifier calculated for the CCALF in the second example is used for the CCALF in the Cr component. It should be noted that when the edge-based classifier in the luminance ALF is reused in the CCALF, or when the edge-based classifier is recalculated for the CCALF based on the chrominance component, the specific number of classes used in the CCALF can be the same as or different from the number of classes used in the luminance ALF. For example, the number of classes used in the luminance ALF is 25, which is obtained by combining activity (5) and directionality (5); while the number of classes used in the CCALF is 10, which is obtained by combining activity (2) and directionality (5).
[0458] In the second approach, a band-based classifier is used in the CCALF. The band-based classifier can be defined at different sub-block levels, such as 4×4, 2×2, or 1×1 block sizes. Different methods can be used to compute the band-based classifier. In the first example, the band-based classifier is computed based on the reconstructed signal in the luminance component. For example, a band-based classifier used in the luminance ALF and computed based on the signal immediately following the SAO in the luminance component is multiplexed in both the Cb and Cr components of the CCALF. In the second example, the band-based classifier is computed based on the reconstructed signal in the chrominance component. For example, for the Cb component of the CCALF, the band-based classifier is computed based on the signal immediately following the SAO in the Cb component; for the Cr component of the CCALF, the band-based classifier is computed based on the signal immediately following the SAO in the Cr component. It should be noted that when computed based on the signal immediately following the SAO in the Cb or Cr component, the computation process can be the same as when computed based on the signal immediately following the SAO in the luminance component. In the third example, a band-based classifier combining the band-based classifier calculated in the first example and the band-based classifier calculated in the second example for the Cb component is used in the Cb component's CCALF; similarly, a band-based classifier combining the band-based classifier calculated in the first example and the band-based classifier calculated in the second example for the Cr component's CCALF is used in the Cr component's CCALF. It should be noted that when the band-based classifier in the luminance ALF is reused in the CCALF, or when the band-based classifier is recalculated for the CCALF based on the chrominance component, the specific number of classes used in the CCALF can be the same as or different from the number of classes used in the luminance ALF. For example, the luminance ALF uses 25 classes, while the CCALF uses 10.
[0459] In the third method, a residual-based classifier is used in the CCALF. The residual-based classifier can be defined at different sub-block levels, such as 4×4, 2×2, or 1×1 block sizes. Different methods can be used to compute the residual-based classifier. In the first example, the residual-based classifier is computed based on the residual signal in the luminance component. For example, a residual-based classifier used in the luminance ALF and computed based on the residual signal in the luminance component is reused in both the Cb and Cr components of the CCALF. In the second example, the residual-based classifier is computed based on the residual signal in the chrominance component. For example, for the Cb component of the CCALF, the residual-based classifier is computed based on the residual signal in the Cb component; for the Cr component of the CCALF, the residual-based classifier is computed based on the residual signal in the Cr component. It should be noted that when computed based on the residual signal in the Cb or Cr component, the computation process can be the same as when computed based on the residual signal in the luminance component. In the third example, a residual-based classifier combining the residual-based classifier calculated in the first example and the residual-based classifier calculated in the second example for the Cb component is used for the Cb component; a residual-based classifier combining the residual-based classifier calculated in the first example and the residual-based classifier calculated in the second example for the Cr component is used for the Cr component. It should be noted that when the residual-based classifier in the luminance ALF is reused in the CCALF, or when the residual-based classifier is recalculated for the CCALF based on the chrominance component, the specific number of classes used in the CCALF can be the same as or different from the number of classes used in the luminance ALF. For example, the number of classes used in the luminance ALF is 25, while the number of classes used in the CCALF is 10.
[0460] In the fourth approach, the methods proposed in the first, second, and third approaches can be combined. For example, CCALF can simultaneously use the edge-based classifier proposed in the first approach and the band-based classifier proposed in the second approach.
[0461] Information from the luminance reconstruction signal is used as additional chroma ALF input.
[0462] According to one or more embodiments of this disclosure, information in the luminance reconstruction signal is used as additional input to the chromaticity ALF equation. Different methods can be used to achieve this goal.
[0463] In the first approach, spatially neighboring pixels in the downsampled luminance reconstruction signal are proposed as additional inputs to the chrominance ALF equation. Various downsampling filters can be used to downsample the luminance reconstruction signal to achieve the same resolution as the chrominance reconstruction signal. For example, the downsampling filter is [1:2:1; 1:2:1] / 8. Various filter shapes can be used to extract information from the downsampled luminance reconstruction signal. In the first example, the filter shape could be a 1×1, 3×3, or 5×5 rhombus, such as... FIG. 15 As shown. In the second example, the filter shape can be a 5×5 cross shape, such as... FIG. 24 As shown. In the third example, the filter shape can be a 3×3 or 5×5 cross shape, such as... FIG. 25 As shown. It should be noted that in both the first and second examples, the filter coefficients are point-symmetric, meaning that two pixels that are point-symmetric to each other share the same filter coefficient, such as... FIG. 15 and FIG. 24 The numerical symbols are shown in the figure. In the third example, no symmetry constraint is enforced on the filter coefficients, meaning that each pixel has its own filter coefficients, as shown in the figure. FIG. 25 The numerical symbols are shown in the figure. Information from the downsampled luminance reconstruction signal can be extracted using various equation forms. In one example, the truncation difference between the surrounding pixels and the current chroma pixel in the downsampled luminance reconstruction signal is used as the input to the chroma ALF equation. In another example, the truncation differences between the surrounding pixels and the co-pixels in the downsampled luminance reconstruction signal, and the truncation differences between the co-pixels and the current chroma pixel in the downsampled luminance reconstruction signal are used as the input to the chroma ALF equation. In a third example, the truncation differences between the surrounding pixels and the co-pixels in the downsampled luminance reconstruction signal are used as the input to the chroma ALF equation. It should be noted that in the first and second examples, the number of filter coefficients is exactly the same as the number of filter coefficients presented in the filter shape example. In a specific example, FIG. 25 The 3×3 cross-shaped filter in the image has 5 coefficients, while... FIG. 25 The 5x5 cross-shaped filter has 9 filter coefficients. In the third example, the number of filter coefficients is the same as the number of filter coefficients shown in the filter shape example minus 1. In a specific example, FIG. 25 The 3×3 cross-shaped filter in the image has 4 coefficients, while... FIG. 25 The number of coefficients in the 5×5 cross-shaped filter is 8. FIG. 26 The modified 3×3 and 5×5 cross shapes corresponding to the third example are shown more clearly. It should be noted that... FIG. 15 , FIG. 24 ,FIG. 25 , FIG. 26 The numbers in the boxes only represent the relative relationships between different filter coefficients, such as whether filter coefficients at different positions are the same. FIG. 15 , FIG. 24 , FIG. 25 , FIG. 26 The specific value of the number within the box can be adjusted depending on the shape of the specific chroma ALF filter.
[0464] In addition to low-pass downsampling filters (e.g., {1:2:1; 1:2:1} / 8), other high-pass downsampling filters can be used to downsample the luminance reconstructed signal to achieve the same resolution as the chrominance reconstructed signal. For example, FIG. 27 Four high-pass downsampling filters are presented. Various filter shapes can be used to extract information from the downsampled brightness reconstruction signal. In the first example, the filter shape can be a 3×3 or 5×5 cross shape, such as... FIG. 25 As shown. In the second example, the filter shape can be a 3×3 or 5×5 cross shape, such as... FIG. 26 As shown. Information can be extracted from the downsampled luminance reconstruction signal using various equation forms. In one example, the cropping of surrounding pixels in the downsampled luminance reconstruction signal is used as input to the chroma ALF equation, where the corresponding filter shape can be... FIG. 25 The filter shape in the example. In another example, the truncation difference between surrounding pixels and co-pixels in the downsampled luminance reconstructed signal is used as input to the chroma ALF equation, where the corresponding filter shape can be... FIG. 26 The filter shape in the text.
[0465] Different methods can be used to process different downsampled luminance reconstruction signals. In the first example, the filter shape in the different downsampled luminance reconstruction signals is switched, and a flag is transmitted in the ALF APS to indicate which filter shape was selected. In one example, for the filter shape in the case of a luminance reconstruction signal downsampled by [1:2:1; 1:2:1] / 8, the filter shape is... FIG. 26 The 3×3 cross shape in the downsampled luminance reconstruction signal is represented by the truncation difference between surrounding pixels and co-position pixels in the downsampled luminance reconstruction signal. For [1:0:-1; 1:0:-1]( FIG. 27 The first downsampling filter in the image shows another filter shape for the downsampled luminance reconstruction signal. FIG. 25The filter shape is a 3×3 cross, and the equation is the result of truncating the surrounding pixels in the downsampled luminance reconstruction signal. These two filter shapes are switched, and a flag is transmitted in the ALF APS to indicate which filter shape is selected. In the second example, the surrounding pixels in different downsampled luminance reconstruction signals are used as additional inputs to the chroma ALF. In one example, for the filter shape in the case of a luminance reconstruction signal downsampled by [1:2:1; 1:2:1] / 8, the filter shape is... FIG. 26 The 3×3 cross shape in the downsampled luminance reconstruction signal is represented by the truncation difference between surrounding pixels and co-position pixels in the downsampled luminance reconstruction signal. For [1:0:-1; 1:0:-1]( FIG. 27 The first downsampling filter in the image shows another filter shape for the downsampled luminance reconstruction signal. FIG. 25 The filter is a 3×3 cross shape, and the equation is a truncation of surrounding pixels in the downsampled luminance reconstruction signal. These two filter shapes are cascaded and both are used as inputs to the chroma ALF.
[0466] In such FIG. 20 In the current chroma ALF filter shape shown, the absolute values of filter taps 0-23 and filter tap 24 are represented using exponential Golomb codes of different orders. When using information from the luminance reconstruction signal as additional input to the chroma ALF equation, different methods can be used to handle the newly added filter taps. In the first method, the absolute values of the newly added filter taps share the same exponential Golomb code order as filter taps 0-23. In the second method, the absolute values of the newly added filter taps share the same exponential Golomb code order as filter tap 24.
[0467] In such FIG. 20 In the current chroma ALF filter shape shown, all filter taps have their own adaptive truncation index. When using information from the luminance reconstruction signal as additional input to the chroma ALF equation, different methods can be used to handle newly added filter taps. In the first method, the newly added filter tap also has its own adaptive truncation index, where the adaptive truncation index representation is the same as that in the original chroma ALF filter tap. In the second method, the newly added filter tap has a fixed truncation index, and then the code for the corresponding truncation index is omitted for the newly added filter tap. For example, the default truncation range corresponding to the fixed truncation index could be [-1024, 1024], [-128, 128], [-32, 32], or [-8, 8].
[0468] In the second method, spatially neighboring pixels from the original luminance reconstruction signal are proposed as additional inputs to the chrominance ALF equation. Various filter shapes can be used to extract information from the original luminance reconstruction signal. For example, the filter shape could be 3×4, such as... FIG. 13 As shown. In this example, no symmetry constraint is enforced on the filter coefficients, meaning each pixel has its own filter coefficients. Various equation forms can be used to extract information from the original luminance reconstructed signal. In one example, the truncation difference between surrounding pixels and co-pixels in the original luminance reconstructed signal is used as input to the chrominance ALF equation. In another example, the truncation differences between surrounding pixels and co-pixels in the original luminance reconstructed signal, and the truncation difference between co-pixels and the current chrominance pixel in the original luminance reconstructed signal are used as input to the chrominance ALF equation. In a third example, the truncation difference between surrounding pixels and the current chrominance pixel in the original luminance reconstructed signal is used as input to the chrominance ALF equation.
[0469] In the third method, the first and second methods are combined, in which spatially neighboring pixels in both the downsampled luminance reconstruction signal and the original luminance reconstruction signal are used as additional inputs to the chromaticity ALF equation.
[0470] It should be noted that when information from the luminance reconstruction signal is used as additional input to the chrominance ALF equation, the luminance reconstruction signal can be either immediately before or immediately after the luminance ALF filter. Furthermore, the luminance reconstruction signals immediately before and immediately after the luminance ALF filter can be used individually or in combination.
[0471] Information from the downsampled luminance reconstruction signal is used as additional CCALF input.
[0472] According to one or more embodiments of this disclosure, information from the downsampled brightness reconstruction signal is used as additional input to the CCALF equation. This objective can be achieved using different methods.
[0473] Various downsampling filters can be used to downsample the luminance reconstruction signal. For example, a downsampling filter is [1:2:1; 1:2:1] / 8. Various filter shapes can be used to extract information from the downsampled luminance reconstruction signal. In the first example, the filter shape could be a 1×1, 3×3, or 5×5 rhombus, such as... FIG. 15 As shown. In the second example, the filter shape can be a 5×5 cross shape, such as... FIG. 24 As shown. In the third example, the filter shape can be a 3×3 or 5×5 cross shape, such as... FIG. 25As shown. In the fourth example, the filter shape can be a 3×3 or 5×5 cross shape, such as... FIG. 26 As shown, various equation forms can be used to extract information from the downsampled luminance reconstruction signal. In one example, the truncation difference between surrounding pixels and the current chroma pixel in the downsampled luminance reconstruction signal is used as the input to the CCALF equation. In another example, the truncation differences between surrounding pixels and co-pixels in the downsampled luminance reconstruction signal, and the truncation differences between co-pixels and the current chroma pixel in the downsampled luminance reconstruction signal, are used as the input to the CCALF equation. In a third example, the truncation differences between surrounding pixels and co-pixels in the downsampled luminance reconstruction signal are used as the input to the CCALF equation.
[0474] In addition to low-pass downsampling filters (e.g., {1:2:1; 1:2:1} / 8), other high-pass downsampling filters can be used to downsample the luminance reconstruction signal. For example, FIG. 27 Four high-pass downsampling filters are presented. Various filter shapes can be used to extract information from the downsampled brightness reconstruction signal. In the first example, the filter shape can be a 3×3 or 5×5 cross shape, such as... FIG. 25 As shown. In the second example, the filter shape can be a 3×3 or 5×5 cross shape, such as... FIG. 26 As shown, information can be extracted from the downsampled luminance reconstruction signal using various equation forms. In one example, the cropping of surrounding pixels in the downsampled luminance reconstruction signal is used as input to the CCALF equation, where the corresponding filter shape can be... FIG. 25 The filter shape in the example. In another example, the truncation difference between surrounding pixels and co-pixels in the downsampled luminance reconstructed signal is used as input to the CCALF equation, where the corresponding filter shape can be... FIG. 26 The filter shape in the text.
[0475] Different methods can be used to process different downsampled luminance reconstruction signals. In the first example, for each chroma component, the filter shape in the different downsampled luminance reconstruction signals is switched, and a flag is transmitted in the ALF APS for each chroma component to indicate which filter shape was selected. In one example, for the filter shape in the case of a luminance reconstruction signal downsampled by [1:2:1; 1:2:1] / 8, the filter shape is... FIG. 26 The 3×3 cross shape in the downsampled luminance reconstruction signal is represented by the truncation difference between surrounding pixels and co-position pixels in the downsampled luminance reconstruction signal. For [1:0:-1; 1:0:-1](FIG. 27 The first downsampling filter in the image shows another filter shape for the downsampled luminance reconstruction signal. FIG. 25 The filter shape is a 3×3 cross, and the equation is the result of truncating the surrounding pixels in the downsampled luminance reconstruction signal. For each chroma component, these two filter shapes are switched, and a flag is transmitted in the ALF APS for each chroma component to indicate which filter was selected. In the second example, the surrounding pixels in different downsampled luminance reconstruction signals are used as additional inputs to the CCALF. In one example, for the filter shape in the case of a luminance reconstruction signal downsampled by [1:2:1; 1:2:1] / 8, the filter shape is... FIG. 26 The 3×3 cross shape in the downsampled luminance reconstruction signal is represented by the truncation difference between surrounding pixels and co-position pixels in the downsampled luminance reconstruction signal. For [1:0:-1; 1:0:-1]( FIG. 27 The first downsampling filter in the image shows another filter shape for the downsampled luminance reconstruction signal. FIG. 25 The filter is a 3×3 cross shape, and the equation is a truncation of surrounding pixels in the downsampled brightness reconstruction signal. These two filter shapes are cascaded and both are used as inputs to the CCALF filter.
[0476] It should be noted that when information from the downsampled luminance reconstruction signal is used as additional input to the CCALF equation, the luminance reconstruction signal can be either the luminance reconstruction signal immediately before or immediately after the luminance ALF filter. Furthermore, the luminance reconstruction signals immediately before and immediately after the luminance ALF filter can be used individually or in combination.
[0477] Applying a fixed brightness filter in each stage of the loop filter
[0478] According to one or more embodiments of this disclosure, a brightness-fixed filter is applied in each stage of the loop filter. Different methods can be used to achieve this objective.
[0479] In the first approach, the luminance fixed filter is applied more than once in the same stage of the loop filter. In the first example, the luminance fixed filter defined in the ALF is executed more than once immediately after the DBF. For example, the luminance fixed filter defined in the ALF is executed twice immediately after the DBF. When the luminance fixed filter defined in the ALF is executed a second time immediately after the DBF, the luminance fixed filter takes the reconstructed samples before the DBF and the reconstructed samples after the first luminance fixed filter as input to generate the filtered result for this stage. When the luminance fixed filter is executed a second time immediately after the DBF, the classification and filtering logic of the fixed filter can be directly reused in the ALF. To control the use of the luminance fixed filter, the cascaded luminance fixed filters can use a strip-level flag to control the on or off of all cascaded luminance fixed filters, or each luminance fixed filter can use its own strip-level flag to control its on or off.
[0480] In the second approach, a luminance-fixed filter is applied at different stages of the loop filter. In the first example, the luminance-fixed filter defined in the ALF is executed immediately after the SAO. This filter takes the reconstructed samples before the DBF and after the SAO as input to generate the filtered result for that stage. When the luminance-fixed filter is executed immediately after the SAO, its classification and filtering logic can be directly reused from the ALF. To control the use of the luminance-fixed filter, both the luminance-fixed filter executed immediately after the DBF and the luminance-fixed filter executed immediately after the SAO can use a single stripe-level flag to enable or disable all luminance-fixed filters, or they can use their own stripe-level flags to control their on / off state. In the second example, the luminance-fixed filter defined in the ALF is executed immediately after the ALF. This filter takes the reconstructed samples before the DBF and after the ALF as input to generate the filtered result for that stage. When a luminance fixed filter is executed immediately after an ALF (Alternate Filter), its classification and filtering logic can be directly reused from the ALF. To control the use of luminance fixed filters, both those executed immediately after a DBF (Dead-of-Flight) and those executed immediately after an ALF can use a single stripe-level flag to enable or disable all luminance fixed filters, or they can use their own stripe-level flags to control their on / off state.
[0481] In the third method, the first and second methods are combined, and the brightness fixed filter is applied once or more at different stages of the loop filter.
[0482] Applying chromaticity-fixed filters in each stage of the loop filter
[0483] According to one or more embodiments of this disclosure, chromaticity-fixed filters are applied in various stages of the loop filter. Different methods can be used to achieve this objective.
[0484] In the first method, the chroma-fixed filter is applied one or more times in the same stage of the loop filter. In the first example, the chroma-fixed filter defined in the ALF is executed once immediately after the DBF, where the chroma-fixed filter takes the reconstructed samples before and after the DBF as input to generate the filtered result for this stage. The classification and filtering logic of the chroma-fixed filter directly reuses the methods in the ALF. The chroma-fixed filter in this stage uses a strip-level flag for signal transmission to implement on / off control. In the second example, the chroma-fixed filter defined in the ALF is executed more than once immediately after the DBF. For example, the chroma-fixed filter defined in the ALF is executed twice immediately after the DBF. When the chroma-fixed filter defined in the ALF is executed a second time immediately after the DBF, the chroma-fixed filter takes the reconstructed samples before the DBF and the reconstructed samples after the first chroma-fixed filtering as input to generate the filtered result for this stage. When the chroma-fixed filter is executed a second time immediately after the DBF, the classification and filtering logic of the chroma-fixed filter can directly reuse the methods in the ALF. To control the use of chroma fixed filters, cascaded chroma fixed filters can use a stripe-level flag to control the on or off of all cascaded chroma fixed filters, or each chroma fixed filter can use its own stripe-level flag to control its on or off.
[0485] In the second approach, chroma-fixed filters are applied at different stages of the loop filter. In the first example, the chroma-fixed filter defined in the ALF is executed immediately after the SAO. This filter takes the reconstructed samples from before the DBF and after the SAO as input to generate the filtered result for that stage. When the chroma-fixed filter is executed immediately after the SAO, its classification and filtering logic can be directly reused from the ALF. To control the use of the chroma-fixed filters, both the chroma-fixed filters executed immediately after the DBF and immediately after the SAO can use a single stripe-level flag to enable or disable all chroma-fixed filters, or they can use their own stripe-level flags to control their on / off state. In the second example, the chroma-fixed filter defined in the ALF is executed immediately after the ALF. This filter takes the reconstructed samples from before the DBF and after the ALF as input to generate the filtered result for that stage. When a chroma fixed filter is executed immediately following an ALF, its classification and filtering logic can be directly reused from the ALF. To control the use of chroma fixed filters, both chroma fixed filters executed immediately following a DBF and those executed immediately following an ALF can use a single stripe-level flag to enable or disable all chroma fixed filters, or they can use their own stripe-level flags to control their on / off state.
[0486] In the third method, the first and second methods are combined, and the chromaticity fixed filter is applied once or more at different stages of the loop filter.
[0487] In the fourth method, the output of the chroma-fixed filter can be directly used as the result of the chroma ALF stage, rather than just as the input of the online chroma ALF stage. In this method, the use of the chroma-fixed filter's output as the result of the chroma ALF stage can be controlled at the CTU level.
[0488] The result of adjusting the luminance / chrominance fixed filter
[0489] According to one or more embodiments of this disclosure, when the results of a luminance-fixed filter or a chrominance-fixed filter are applied in various stages of the loop filter, these results are adjusted. Different methods can be used to achieve this objective.
[0490] In the first method, a residual scaling method is applied to the results of either the luminance-fixed filter or the chrominance-fixed filter. When the luminance-fixed filter or chrominance-fixed filter is applied to the reconstructed image at each stage of the loop filter, a scaling factor is derived for each color component and transmitted as a signal in the strip header. This derivation is based on the least squares method. The difference (residual) between the input reconstructed samples and the results of the luminance-fixed filter or the chrominance-fixed filter is scaled by the scaling factor and then added to the input reconstructed samples.
[0491] In the second method, a residual offset adjustment method is applied to the results of either the luminance-fixed filter or the chrominance-fixed filter. When the luminance-fixed filter or chrominance-fixed filter is applied to the reconstructed image at each stage of the loop filter, a residual offset value is selected for each color component, and this residual offset value is transmitted as a signal in the strip header. Offset value candidates can be defined in different sets. For example, offset value candidates could be {1, 2}. The residual between the input reconstructed samples and the results of the luminance-fixed filter or the chrominance-fixed filter is adjusted by reducing the residual magnitude of each pixel by this small offset value, and then added to the input reconstructed samples.
[0492] FIG. 28 A flowchart of a video decoding method 2800 according to some embodiments of the present disclosure is shown. Method 2800 may, for example, be derived by... FIG. 3 The video decoder 30 in the middle is executed. For example... FIG. 28 As shown, method 2800 includes steps S2810 and S2820.
[0493] In step S2810, at least one fixed filter is determined to be applied in the loop filter. The loop filter is associated with at least one stage, and the at least one fixed filter is applied to the same stage or different stages of the at least one stage.
[0494] In step S2820, for each of the at least one fixed filter, the reconstructed sample points are filtered using the fixed filter to obtain filtered sample points.
[0495] According to embodiments of this disclosure, applying at least one fixed filter in the same or different stages of a loop filter is beneficial for improving encoding and decoding performance.
[0496] Loop filters are used to filter reconstructed samples to remove artifacts such as blockiness, ringing, and noise, improving the quality of the reconstructed image and making the video clearer and more natural. A loop filter is typically a multi-stage pipeline. In other words, a loop filter can include multiple filters that pipeline the initial reconstructed signal (obtained by adding the predicted signal to the residual signal) to obtain the final reconstructed signal.
[0497] According to some embodiments, the loop filter may include at least one of a deblocking filter (DBF), a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF). When the loop filter includes a DBF, a SAO filter, and an ALF, the DBF, SAO filter, and ALF may be connected in sequence.
[0498] Corresponding to the loop filter including at least one of DBF, SAO filter and ALF, at least one stage associated with the loop filter may include: a stage immediately following DBF, a stage immediately following SAO or a stage immediately following ALF.
[0499] According to embodiments of this disclosure, at least one fixed filter is applied in a loop filter to improve encoding / decoding efficiency. The at least one fixed filter can be applied to various stages of the loop filter; for example, it can be applied to the same stage or different stages among at least one stage included in the loop filter.
[0500] According to some embodiments, each of the at least one fixed filter can be a luminance fixed filter for filtering the reconstructed samples of the luminance component, or a chrominance fixed filter for filtering the reconstructed samples of the chrominance component. It should be understood that luminance fixed filters and chrominance fixed filters can be applied simultaneously at some stage of the loop filter, or only luminance fixed filters or only chrominance fixed filters can be applied at some stage.
[0501] Luminance-fixed filters can be applied at various stages of the loop filter. For example, multiple (e.g., two) cascaded luminance-fixed filters can be applied immediately after the DBF; two luminance-fixed filters can be applied, one immediately after the DBF and the other immediately after the SAO; two luminance-fixed filters can be applied, one immediately after the DBF and the other immediately after the ALF; and so on.
[0502] Chromaticity-fixed filters can be applied at various stages of the loop filter. For example, a chromaticity-fixed filter can be applied immediately after the DBF; multiple (e.g., two) cascaded chromaticity-fixed filters can be applied immediately after the DBF; two chromaticity-fixed filters can be applied, one immediately after the DBF and the other immediately after the SAO; two chromaticity-fixed filters can be applied, one immediately after the DBF and the other immediately after the ALF; and so on.
[0503] According to some embodiments, step S2820 may include steps S2821 and S2822.
[0504] In step S2821, the input reconstructed samples of the fixed filter are determined based on the stage at which the fixed filter is located. It should be understood that the input reconstructed samples are the reconstructed samples to be input into the fixed filter.
[0505] In step S2822, the input reconstructed samples are input into the fixed filter to obtain the filtered samples output by the fixed filter.
[0506] In the above embodiments, the output of the fixed filter is directly used as the filtered sample.
[0507] According to some embodiments, a fixed filter can be applied in a stage immediately preceding the first filter in a loop filter. Accordingly, the input reconstructed samples of this fixed filter can include the initial reconstructed samples of the input loop filter. Typically, these initial reconstructed samples are obtained by adding the predicted samples to the residual samples.
[0508] For example, in the case where the loop filter includes a DBF, a SAO filter and an ALF connected in sequence, the input reconstructed sample of the fixed filter immediately preceding the DBF is the initial reconstructed sample.
[0509] According to some embodiments, the fixed filter may not be applied in the stage immediately preceding the first filter in the loop filter. In this case, the input reconstructed samples of the fixed filter may include the initial reconstructed samples of the input loop filter and the reconstructed samples output by the filter immediately preceding the fixed filter.
[0510] For example, when the loop filter includes a DBF, SAO filter, and ALF connected in sequence, the initial reconstructed samples input to the loop filter are the reconstructed samples before the DBF. For two cascaded fixed filters (either two luminance fixed filters or two chrominance fixed filters) immediately following the DBF, the input reconstructed samples of the first fixed filter are the reconstructed samples before and after the DBF; the input reconstructed samples of the second fixed filter are the reconstructed samples before the DBF and the reconstructed samples after the first fixed filter. For a fixed filter (either a luminance fixed filter or a chrominance fixed filter) immediately following the SAO, its input reconstructed samples are the reconstructed samples before the DBF and the reconstructed samples after the SAO. For a fixed filter (either a luminance fixed filter or a chrominance fixed filter) immediately following the ALF, its input reconstructed samples are the reconstructed samples before the DBF and the reconstructed samples after the ALF. It should be understood that, in the embodiments of this disclosure, the reconstructed samples “before” a filter can be understood as the reconstructed samples input to the filter; and the reconstructed samples “after” a filter can be understood as the reconstructed samples output by the filter.
[0511] According to some embodiments, step S2820 may include steps S2823-S2825.
[0512] In step S2823, the input reconstructed samples of the fixed filter are determined based on the stage at which the fixed filter is located. It should be understood that the input reconstructed samples are the reconstructed samples to be input into the fixed filter.
[0513] In step S2824, the input reconstruction samples are input into a fixed filter to obtain the output result of the fixed filter.
[0514] In step S2825, the output results are adjusted to obtain filtered samples.
[0515] In the above embodiment, the output of the fixed filter is adjusted, and the adjusted output is used as the filtered sample.
[0516] According to some embodiments, in step S2825, the output of the fixed filter can be adjusted by residual scaling or residual offset.
[0517] According to some embodiments, the adjustment method for residual scaling may include: obtaining a scaling factor corresponding to the input reconstructed samples; scaling the difference (i.e., the residual) between the input reconstructed samples and the output result using the scaling factor to obtain a residual scaling result; and using the sum of the input reconstructed samples and the residual scaling result as the adjusted output result to obtain filtered samples. When the input reconstructed samples include multiple reconstructed samples (e.g., the input reconstructed samples include an initial reconstructed sample and reconstructed samples output by a filter immediately preceding a fixed filter), any one of the reconstructed samples can be added to the residual scaling result to obtain the adjusted output result; alternatively, a weighted sum of these multiple reconstructed samples can be added to the residual scaling result to obtain the adjusted output result.
[0518] According to some embodiments, a scaling factor can be derived for each color component and transmitted as a signal in the bitstream. Accordingly, on the decoder side, the color component corresponding to the input reconstructed sample can be determined, and the scaling factor corresponding to that color component can be received from the bitstream for adjusting the fixed filter output of the input reconstructed sample.
[0519] According to some embodiments, the scaling factor for each color component can be derived based on the least squares method.
[0520] According to some embodiments, the residual offset adjustment method includes: obtaining the residual offset value corresponding to the input reconstructed sample; offsetting the difference (i.e., the residual) between the input reconstructed sample and the output result using the residual offset value to obtain the residual offset result; and using the sum of the input reconstructed sample and the residual offset result as the adjusted output result to obtain the filtered sample. When the input reconstructed sample includes multiple reconstructed samples (e.g., the input reconstructed sample includes an initial reconstructed sample and reconstructed samples output by a filter immediately preceding the fixed filter), any one of the reconstructed samples can be added to the residual scaling result to obtain the adjusted output result; alternatively, the weighted sum of these multiple reconstructed samples can be added to the residual scaling result to obtain the adjusted output result.
[0521] According to some embodiments, a residual offset value can be selected for each color component, and this residual offset value can be transmitted as a signal in the bitstream. Accordingly, on the decoder side, the color component corresponding to the input reconstructed sample can be determined, and the residual offset value corresponding to that color component can be received from the bitstream for adjusting the fixed filter output of the input reconstructed sample.
[0522] According to some embodiments, for each color component, a residual offset value can be selected from a preset set of offset value candidates. The offset value candidate set can be, for example, {1, 2}. The offset value candidate sets for different color components can be the same or different.
[0523] According to some embodiments, one or more flags can be used to control the use of one or more fixed filters. One or more flags can be transmitted as signals in the bitstream. One or more flags can be stripe-level flags.
[0524] According to some embodiments, method 2800 may further include steps S2830 and S2840.
[0525] In step S2830, one or more flags are received from the bit stream, which indicate whether one or more fixed filters in the loop filter are applied.
[0526] In step S2840, based on one or more of the aforementioned flags, at least one fixed filter to be applied in the loop filter is determined from one or more fixed filters.
[0527] It should be understood that among the one or more fixed filters included in the loop filter, the fixed filter indicated by the above one or more marks as being applied is the "at least one fixed filter" in step S2810.
[0528] According to some embodiments, the aforementioned one or more flags include a first flag indicating whether the one or more fixed filters are applied simultaneously. For example, if the one or more fixed filters include two brightness fixed filters applied to different stages, a first flag can be used to control the application status of these two brightness fixed filters simultaneously. For example, a value of 1 for the first flag indicates that the two brightness fixed filters are applied, and a value of 0 for the first flag indicates that the two brightness fixed filters are not applied.
[0529] According to some embodiments, the aforementioned one or more flags include one or more second flags. These one or more second flags correspond to one or more fixed filters in the loop filter, with each second flag indicating whether the corresponding fixed filter is applied. For example, if the one or more fixed filters include two luminance fixed filters applied to different stages, a second flag flag1 can be used to control the application state of the first luminance fixed filter, and another second flag flag2 can be used to control the application state of the second luminance fixed filter. For example, flag1 = 1 indicates that the first luminance fixed filter is applied, and flag1 = 0 indicates that the first luminance fixed filter is not applied. Similarly, flag2 = 1 indicates that the second luminance fixed filter is applied, and flag2 = 0 indicates that the second luminance fixed filter is not applied.
[0530] According to some embodiments, method 2800 may also include steps S2850 and S2860.
[0531] In step S2850, a third flag is received from the bit stream, which indicates whether the output of the chroma fixed filter is used as the result of the chroma ALF stage.
[0532] In step S2860, in response to the third flag indicating that the output of the chroma fixed filter is used as the result of the chroma ALF stage, the output of the chroma fixed filter is used as the result of the chroma ALF stage.
[0533] Other implementation details of the video decoding method 2800 can be found in the sections above: “Applying a fixed luminance filter in each stage of the loop filter,” “Applying a fixed chroma filter in each stage of the loop filter,” and “Results of adjusting the fixed luminance / chroma filter.”
[0534] FIG. 29 A flowchart of a video encoding method 2900 according to some embodiments of the present disclosure is shown. Video encoding method 2900 is the video encoding method corresponding to the video decoding method 2800 described above. Method 2900 can, for example, be derived from... FIG. 2 The video encoder 20 in the middle is executed. For example... FIG. 29 As shown, method 2900 includes steps S2910 and S2920.
[0535] In step S2910, at least one fixed filter is determined to be applied in the loop filter. The loop filter is associated with at least one stage, and the at least one fixed filter is applied to the same stage or different stages of the at least one stage.
[0536] In step S2920, for each of the at least one fixed filter mentioned above, the reconstructed sample points are filtered using the fixed filter to obtain filtered sample points.
[0537] According to embodiments of this disclosure, applying at least one fixed filter in the same or different stages of a loop filter is beneficial for improving encoding and decoding performance.
[0538] According to some embodiments, the at least one stage includes: a stage immediately following the deblocking filter (DBF), a stage immediately following the sample adaptive offset (SAO), or a stage immediately following the adaptive loop filter (ALF).
[0539] According to some embodiments, filtering reconstructed samples using the fixed filter to obtain filtered samples includes: determining the input reconstructed samples of the fixed filter based on the stage at which the fixed filter is in; and inputting the input reconstructed samples into the fixed filter to obtain the filtered samples output by the fixed filter.
[0540] According to some embodiments, the fixed filter is applied in a stage immediately preceding the first filter in the loop filter, and the input reconstructed samples include the initial reconstructed samples input to the loop filter.
[0541] According to some embodiments, the fixed filter is not applied in a stage immediately preceding the first filter in the loop filter, and the input reconstructed samples include initial reconstructed samples input to the loop filter and reconstructed samples output by the filter immediately preceding the fixed filter.
[0542] According to some embodiments, method 2900 further includes: determining whether to apply one or more fixed filters in the loop filters; and encoding one or more flags indicating whether to apply the one or more fixed filters into a bitstream.
[0543] According to some embodiments, the one or more flags include a first flag indicating whether the one or more fixed filters are applied simultaneously.
[0544] According to some embodiments, the one or more flags include one or more second flags, each of the one or more second flags corresponding to one or more fixed filters, and each of the one or more second flags indicating whether the corresponding fixed filter is applied.
[0545] According to some embodiments, the at least one fixed filter includes a chroma fixed filter, and method 2900 further includes: determining whether to use the output of the chroma fixed filter as the result of the chroma ALF stage; and encoding a third flag indicating whether to use the output of the chroma fixed filter as the result of the chroma ALF stage into a bitstream.
[0546] According to some embodiments, filtering reconstructed samples using the fixed filter to obtain filtered samples includes: determining the input reconstructed samples of the fixed filter based on the stage of the fixed filter; inputting the input reconstructed samples into the fixed filter to obtain the output result of the fixed filter; and adjusting the output result to obtain the filtered samples.
[0547] According to some embodiments, adjusting the output result to obtain the filtered sample points includes: obtaining a scaling factor corresponding to the input reconstructed sample points; scaling the difference between the input reconstructed sample points and the output result using the scaling factor to obtain a residual scaling result; and using the sum of the input reconstructed sample points and the residual scaling result as the adjusted output result to obtain the filtered sample points.
[0548] According to some embodiments, method 2900 further includes: determining a scaling factor corresponding to each color component; and encoding the scaling factor corresponding to each color component into a bitstream.
[0549] According to some embodiments, adjusting the output result to obtain the filtered sample points includes: obtaining the residual offset value corresponding to the input reconstructed sample point; using the residual offset value to offset the difference between the input reconstructed sample point and the output result to obtain the residual offset result; and using the sum of the input reconstructed sample point and the residual offset result as the adjusted output result to obtain the filtered sample points.
[0550] According to some embodiments, obtaining the residual offset value corresponding to the input reconstructed sample points includes: determining the residual offset value corresponding to each color component; and encoding the residual offset value corresponding to each color component into a bit stream.
[0551] Other implementation details of video coding method 2900 can be found in the sections "Applying a fixed luminance filter in each stage of the loop filter", "Applying a fixed chrominance filter in each stage of the loop filter" and "Results of adjusting the fixed luminance / chrominance filter" above, as well as the corresponding video decoding method 2800 described above, which will not be repeated here.
[0552] FIG. 16 A computing environment 1610 coupled to a user interface 1650 is shown. The computing environment 1610 may be part of a data processing server. The computing environment 1610 includes a processor 1620, memory 1630, and input / output (I / O) interface 1640.
[0553] Processor 1620 typically controls the overall operation of computing environment 1610, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1620 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Furthermore, processor 1620 may include one or more modules that facilitate interaction between processor 1620 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, graphics processing unit (GPU), etc.
[0554] Memory 1630 is configured to store various types of data to support the operation of computing environment 1610. Memory 1630 may include predefined software 1632. Examples of such data include instructions for any application or method operating on computing environment 1610, video datasets, image data, etc. Memory 1630 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0555] I / O interface 1640 provides an interface between processor 1620 and peripheral interface modules (such as keyboards, click wheels, buttons, etc.). Buttons may include, but are not limited to, home buttons, start scan buttons, and stop scan buttons. I / O interface 1640 can be coupled to encoders and decoders.
[0556] In an embodiment, a non-transitory computer-readable storage medium is also provided, including, for example, a plurality of programs in a memory 1630 and / or a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. The plurality of programs can be executed by a processor 1620 in a computing environment 1610 to perform the above-described methods. In one example, the plurality of programs can be executed by a processor 1620 in a computing environment 1610 to (e.g., from...) FIG. 2 The video encoder 20 in the computing environment 1610 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1620 in the computing environment 1610 to perform the above-described decoding method based on the received bitstream or data stream. In another example, multiple programs can be executed by the processor 1620 in the computing environment 1610 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 1620 in the computing environment 1610 to (e.g., to...) FIG. 3 The video decoder 30 in the medium transmits the bitstream or data stream. Alternatively, a non-transitory computer-readable storage medium may store data generated by an encoder (e.g., FIG. 2 The video encoder 20 in the video is generated using, for example, the encoding method described above, for use by the decoder (e.g., FIG. 3The video decoder 30 in the video decoder uses a bitstream or data stream that includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) when decoding video data. Non-transitory computer-readable storage media may be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0557] In one embodiment, a bitstream generated by the above encoding method or a bitstream to be decoded by the above decoding method is provided. In another embodiment, a bitstream comprising encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method is provided.
[0558] In one embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 1620); and a non-transitory computer-readable storage medium or memory 1630 therein storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above when executing the plurality of programs.
[0559] In one embodiment, a computer program product having instructions for storing or transmitting a bitstream, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above, is also provided. In another embodiment, a computer program product including, for example, a plurality of programs in a memory 1630, which can be executed by a processor 1620 in a computing environment 1610 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0560] In an embodiment, the computing environment 1610 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0561] In one embodiment, a method for storing a bitstream is also provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream includes encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method.
[0562] In one embodiment, a method for transmitting a bitstream generated by the encoder described above is also provided. In another embodiment, a method for receiving a bitstream to be decoded by the decoder described above is also provided.
[0563] The description in this disclosure is presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.
[0564] Unless otherwise specified, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.
[0565] The examples were chosen and described to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure, and preferably to utilize the basic principles and various embodiments with various modifications suitable for the intended particular purpose. Therefore, it should be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.
Claims
1. A video decoding method, comprising: Determine at least one fixed filter to be applied in a loop filter, wherein the loop filter is associated with at least one stage, and the at least one fixed filter is applied to the same stage or different stages within the at least one stage; and For each of the at least one fixed filter, the reconstructed samples are filtered using the fixed filter to obtain filtered samples.
2. The method according to claim 1, wherein, The at least one stage includes: a stage immediately following the deblocking filter (DBF), a stage immediately following the sample adaptive offset (SAO), or a stage immediately following the adaptive loop filter (ALF).
3. The method according to claim 1 or 2, wherein, The reconstructed samples are filtered using the fixed filter to obtain filtered samples including: Based on the stage at which the fixed filter is located, the input reconstruction samples of the fixed filter are determined; and The input reconstructed samples are input into the fixed filter to obtain the filtered samples output by the fixed filter.
4. The method according to claim 3, wherein, The fixed filter is applied in a stage immediately preceding the first filter in the loop filter, and wherein the input reconstructed samples include the initial reconstructed samples input to the loop filter.
5. The method according to claim 3, wherein, The fixed filter is not applied to a stage immediately preceding the first filter in the loop filter, and the input reconstructed samples include initial reconstructed samples input to the loop filter and reconstructed samples output by the filter immediately preceding the fixed filter.
6. The method according to any one of claims 1-5, further comprising: Receive one or more flags from the bitstream, wherein the one or more flags indicate whether one or more fixed filters in the loop filter are applied; and Based on the one or more flags, determine the at least one fixed filter to be applied in the loop filter from the one or more fixed filters.
7. The method according to claim 6, wherein, The one or more flags include a first flag, which indicates whether the one or more fixed filters are applied simultaneously.
8. The method according to claim 6, wherein, The one or more flags include one or more second flags, each of which corresponds to one or more fixed filters, and each of the one or more second flags indicates whether the corresponding fixed filter is applied.
9. The method according to claim 1, wherein, The reconstructed samples are filtered using the fixed filter to obtain filtered samples including: Based on the stage in which the fixed filter is located, the input reconstruction sample points of the fixed filter are determined; The input reconstructed samples are input into the fixed filter to obtain the output of the fixed filter; and The output is adjusted to obtain the filtered sample points.
10. The method according to claim 9, wherein, The output is adjusted to obtain the filtered sample points, which include: Obtain the scaling factor corresponding to the input reconstructed sample points; The difference between the input reconstructed samples and the output result is scaled using the scaling factor to obtain a residual scaling result; and The sum of the input reconstructed samples and the residual scaling result is used as the adjusted output result to obtain the filtered samples.
11. The method according to claim 10, wherein, Obtaining the scaling factor corresponding to the input reconstructed sample points includes: Determine the color components corresponding to the input reconstructed sample points; and Receive the scaling factor corresponding to the color component from the bitstream.
12. The method according to claim 9, wherein, The output is adjusted to obtain the filtered sample points, which include: Obtain the residual offset value corresponding to the input reconstructed sample point; The difference between the input reconstructed samples and the output result is offset using the residual offset value to obtain the residual offset result; and The sum of the input reconstructed sample points and the residual offset result is used as the adjusted output result to obtain the filtered sample points.
13. The method according to claim 12, wherein, Obtaining the residual offset value corresponding to the input reconstructed sample points includes: Determine the color components corresponding to the input reconstructed sample points; and Receive the residual offset value corresponding to the color component from the bit stream.
14. A video coding method, comprising: Determine at least one fixed filter to be applied in a loop filter, wherein the loop filter is associated with at least one stage, and the at least one fixed filter is applied to the same stage or different stages within the at least one stage; and For each of the at least one fixed filter, the reconstructed samples are filtered using the fixed filter to obtain filtered samples.
15. The method according to claim 14, wherein, The at least one stage includes: a stage immediately following the block filter (DBF), a stage immediately following the sample adaptive offset (SAO), or a stage immediately following the adaptive loop filter (ALF).
16. The method according to claim 14 or 15, wherein, The reconstructed samples are filtered using the fixed filter to obtain filtered samples including: Based on the stage at which the fixed filter is located, the input reconstruction samples of the fixed filter are determined; and The input reconstructed samples are input into the fixed filter to obtain the filtered samples output by the fixed filter.
17. The method according to claim 16, wherein, The fixed filter is applied in a stage immediately preceding the first filter in the loop filter, and wherein the input reconstructed samples include the initial reconstructed samples input to the loop filter.
18. The method according to claim 16, wherein, The fixed filter is not applied to a stage immediately preceding the first filter in the loop filter, and the input reconstructed samples include initial reconstructed samples input to the loop filter and reconstructed samples output by the filter immediately preceding the fixed filter.
19. The method according to any one of claims 14-18, further comprising: Determine whether to apply one or more fixed filters from the loop filters; as well as One or more flags indicating whether to apply the one or more fixed filters are encoded into the bitstream.
20. The method according to claim 19, wherein, The one or more flags include a first flag, which indicates whether the one or more fixed filters are applied simultaneously.
21. The method according to claim 19, wherein, The one or more flags include one or more second flags, each of which corresponds to one or more fixed filters, and each of the one or more second flags indicates whether the corresponding fixed filter is applied.
22. The method according to claim 14, wherein, The reconstructed samples are filtered using the fixed filter to obtain filtered samples including: Based on the stage in which the fixed filter is located, the input reconstruction sample points of the fixed filter are determined; The input reconstructed samples are input into the fixed filter to obtain the output of the fixed filter; and The output is adjusted to obtain the filtered sample points.
23. The method according to claim 22, wherein, The output is adjusted to obtain the filtered sample points, which include: Obtain the scaling factor corresponding to the input reconstructed sample points; The difference between the input reconstructed samples and the output result is scaled using the scaling factor to obtain a residual scaling result; and The sum of the input reconstructed samples and the residual scaling result is used as the adjusted output result to obtain the filtered samples.
24. The method of claim 23, further comprising: Determine the scaling factor for each color component; as well as The scaling factor corresponding to each color component is encoded into the bitstream.
25. The method according to claim 22, wherein, The output is adjusted to obtain the filtered sample points, which include: Obtain the residual offset value corresponding to the input reconstructed sample point; The difference between the input reconstructed samples and the output result is offset using the residual offset value to obtain the residual offset result; and The sum of the input reconstructed sample points and the residual offset result is used as the adjusted output result to obtain the filtered sample points.
26. The method of claim 25, wherein, Obtaining the residual offset value corresponding to the input reconstructed sample points includes: Determine the residual offset value corresponding to each color component; and The residual offset value corresponding to each color component is encoded into the bitstream.
27. A computing device, comprising: One or more processors; as well as Memory coupled to the one or more processors, The memory is configured to store instructions executable by the one or more processors, which, when executing the instructions, cause the computing device to perform the method as described in any one of claims 1-26.
28. A non-transitory computer-readable storage medium storing a bit stream generated by instructions, which, when executed by a computing device having one or more processors, cause the one or more processors to perform the method as described in any one of claims 14-26.
29. A method for storing a bit stream, comprising: A bitstream is generated according to the method described in any one of claims 14-26; as well as Store the bit stream.
30. A computer program product comprising instructions that, when executed by one or more processors of a computing device, cause the computing device to perform the method as described in any one of claims 1-26.