Generalized Sample Offset

By employing CCSO and LSO techniques with adaptive loop filtering, the video coding techniques address the challenges of redundancy reduction and compression efficiency, achieving improved video data representation and reduced bandwidth.

JP7683866B2Active Publication Date: 2025-05-27TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023555221
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2022-10-27
Publication Date
2025-05-27
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

Current video coding techniques face challenges in efficiently reducing redundancy and improving compression efficiency, particularly in handling intra-picture prediction and motion compensation across different color components.

Method used

The implementation of cross-component sample offset (CCSO) and local sample offset (LSO) techniques, which utilize adaptive loop filtering (ALF) to adjust sample values between adjacent samples and different color components, enhancing compression efficiency.

Benefits of technology

These techniques effectively reduce redundancy and improve compression efficiency by adaptively filtering sample values, leading to better representation of video data and reduced bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683866000026
    Figure 0007683866000026
  • Figure 0007683866000027
    Figure 0007683866000027
  • Figure 0007683866000028
    Figure 0007683866000028
Patent Text Reader

Abstract

The present disclosure relates to adaptive loop filtering (ALF) for cross-component sample offset (CCSO) and local sample offset (LSO). The ALF uses a reconstructed sample of a first color component as an input (e.g., Y, Cb, or Cr). For CCSO, the output is applied to a second color component that is a different color component of the first color component. For LSO, the output is applied to the first color component. The joint ALF may be generalized for CCSO and LSO by considering delta values ​​between adjacent samples of the co-located (or current) sample and also considering the level value of the co-located (or current) sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Incorporation by Reference] This application claims priority to U.S. Patent Application No. 18 / 047,877, entitled "GENERALIZED SAMPLE OFFSET," filed on Oct. 19, 2022, which claims priority to U.S. Provisional Application No. 63 / 279,674, entitled "GENERALIZED SAMPLE OFFSET," filed on Nov. 15, 2021, and U.S. Provisional Application No. 63 / 289,137, entitled "GENERALIZED SAMPLE OFFSET," filed on Dec. 13, 2021, and the entire contents of all of these applications are incorporated by reference.

[0002] [Technical Field] The present disclosure relates to a set of advanced video coding techniques. More specifically, the techniques of the present disclosure include cross-component sample offset (CCSO) and local sample offset (LSO).

Background Art

[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. Works of the inventors named in this application that are within the scope of the work described in this background section, and aspects of this description that may not be prior art to the present application at the time of filing of this application in other respects, are not admitted as prior art to the present disclosure, either expressly or implicitly.

[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having spatial dimensions, for example, of 1920×1080 luminance samples and associated full or sub-sampled chrominance samples. The series of pictures can have a fixed or variable picture rate (also called frame rate), for example, a picture rate of 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, video with a pixel resolution of 1920×1080, a frame rate of 60 frames / second, and 4:2:0 chroma sub-sampling with 8 bits per pixel per color channel requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage space of over 600 GB.

[0005] One purpose of video encoding and decoding can be the reduction of redundancy in an uncompressed input video signal by compression. Compression can help reduce the above bandwidth and / or storage space requirements, possibly by more than an order of magnitude in some cases. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to a technique where, through the decoding process, an exact copy of the original signal can be reconstructed from the compressed original signal. Irreversible compression refers to an encoding / decoding process where the original video information is not fully retained during encoding and cannot be fully recovered during decoding. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to usefully render the reconstructed signal for its intended use despite some information loss. In the case of video, irreversible compression is widely used in many applications. The amount of acceptable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances, and generally, higher acceptable distortion allows encoding algorithms that result in higher loss and higher compression ratios.

[0006] Video encoders and decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy encoding.

[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially divided into blocks of samples. If all blocks of samples are coded in an intra mode, that picture can be called an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in an encoded video bitstream and video session or as a still image. Next, the samples of the blocks after intra prediction can be subjected to a transform to the frequency domain, and the transform coefficients so generated can be quantized prior to entropy coding. Intra prediction represents techniques for minimizing the sample values in the pre-transform domain. In some cases, the smaller the DC value and AC coefficients after transformation, the fewer bits are required with a given quantization step size to represent the block after entropy coding.

[0008] Traditional intra coding, such as known from MPEG-2 generation coding techniques, does not use intra prediction. However, some newer video compression techniques include techniques that attempt to encode / decoder a block based on surrounding sample data and / or metadata that precede in decoding order the data of the block being coded and / or decoded that are obtained during encoding and / or decoding of spatially adjacent ones, for example. Such techniques are hereinafter called "intra prediction" techniques. Note that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and does not use reference data from other reference pictures.

[0009] There can be various forms of intra prediction. In a given video coding technology, if two or more such technologies are available, the technology used can be called an intra prediction mode. One or more intra prediction modes may be provided in a particular codec. In certain cases, the mode can have sub - modes and / or be associated with various parameters, and the mode / sub - mode information and intra - coding parameters of a video block can be encoded individually or can be included together in a mode codeword. Which codeword to use for a given combination of mode, sub - mode and / or parameters can affect the coding efficiency gain through intra prediction, and the entropy coding technology used to convert the codeword into a bitstream can similarly have an impact.

[0010] A particular mode of intra prediction was introduced in H.264, refined in H.265, and further refined in newer coding technologies such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). Generally for intra prediction, the predictor block can be formed using the available adjacent sample values. For example, the available values of a particular set of adjacent samples along a particular direction and / or line may be copied to the predictor block. The reference to the direction used can be encoded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, in the lower right, a subset of 9 predictor directions is depicted, which corresponds to 33 of the 35 intra modes specified in H.265, namely the 33 possible intra predictor directions in H.265. The point (101) where the arrows converge represents the predicted sample. The arrows represent the directions when adjacent samples are used to predict the sample at 101. For example, arrow (102) indicates that sample (101) is predicted from the adjacent sample(s) in the upper right at an angle of 45 degrees from the horizontal direction. Similarly, arrow (103) indicates that sample (101) is predicted from the adjacent sample(s) in the lower left of sample (101) at an angle of 22.5 degrees from the horizontal direction.

[0012] Continuing to refer to FIG. 1A, in the upper left, a square block (104) of 4×4 samples is depicted (indicated by the thick dashed line). The square block (104) contains 16 samples, and each sample is labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample within block (104) in both the Y and X dimensions. Since the block is of size 4×4 samples, S44 is in the lower right. Additionally, exemplary reference samples following a similar numbering scheme are shown. The reference samples are labeled with "R" and its Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, adjacent predicted samples in the vicinity of the block being reconstructed are used.

[0013] The intra-picture prediction of block 104 may begin by copying the reference sample value from adjacent samples according to the predicted direction signaled. For example, assume that the encoded video bitstream includes signaling indicating the predicted direction of arrow (102) for this block 104. That is, the sample is predicted from the upper right predicted sample(s) at an angle of 45 degrees from the horizontal direction. In such a case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, especially when the direction is not divisible by 45 degrees, the values of multiple reference samples can be combined, for example, by interpolation, to calculate the reference sample.

[0015] As video encoding technology continues to develop, the number of possible directions has increased. In H.264 (2003), for example, nine different directions are available for intra-prediction. This increased to 33 in H.265 (2013), and JEM / VVC / BMS at the time of this disclosure can support up to 65 directions. Experiments are conducted to help identify the most appropriate intra-prediction direction, and specific techniques in entropy coding may be used to encode these most appropriate directions with a small number of bits while accepting a specific bit penalty for the direction. Further, in some cases, the direction itself can be predicted from the adjacent direction used in the intra-prediction of the decoded adjacent block.

[0016] FIG. 1B shows a schematic diagram (180) depicting 65 intra-prediction directions by JEM to show the increasing number of prediction directions in various encoding technologies developed over time.

[0017] The method of mapping intra prediction direction bits to a prediction direction in a symbolized video bitstream may vary for each video encoding technique. For example, it can range from a simple direct mapping to an intra prediction mode of the prediction direction, to a complex adaptive method related to codewords and the most probable mode, and similar techniques. However, in all cases, in video content, there may exist a specific direction of intra prediction that is statistically less likely to occur than certain other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video encoding technique, these less likely methods may be represented by a larger number of bits than the more likely directions.

[0018] Inter-picture prediction or inter prediction may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a part thereof (reference picture) is spatially shifted in the direction indicated by a motion vector (hereinafter, MV), and then used for the prediction of a newly reconstructed picture or a part thereof (e.g., a block). In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV may have two dimensions of X and Y, or three dimensions, and the third dimension is an indication of the reference picture used (similar to the temporal dimension).

[0019] In some video compression techniques, the current MV applicable to a particular region of sample data can be predicted from other MVs, for example, from other MVs related to other regions of sample data that are spatially adjacent to the region being reconstructed and that precede the current MV in decoding order. By doing so, by relying on reducing redundancy in the related MVs, the overall amount of data required for encoding the MVs can be substantially reduced, thereby increasing the compression efficiency. MV prediction can function effectively, for example, when encoding an input video signal (known as natural video) derived from a camera, because regions larger than the regions to which a single MV is applicable in the video sequence move in a similar direction, and thus, in some cases, there is a statistical likelihood that similar motion vectors derived from the MVs of adjacent regions can be used for prediction. As a result, the actual MV for a given region will be similar or identical to the MV predicted from the surrounding MVs. Such an MV may then be represented in fewer bits than would be used if the MV were directly encoded rather than predicted from adjacent MVs after entropy encoding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, the MV prediction itself may be lossy, for example, due to rounding errors when calculating predictors from some surrounding MVs.

[0020] H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016) describes various MV prediction mechanisms. Among the many MV prediction mechanisms specified by H.265, the technique hereinafter referred to as "spatial merge" will be described in this specification.

[0021] Specifically, referring to FIG. 2, the current block (201) contains samples found by the encoder during the motion search process that the current block is predictable from the previous block of the same size that has been spatially shifted. Instead of directly encoding the MV, the MV can be derived from metadata associated with one or more reference pictures, for example, from the latest reference picture (in decoding order), using an MV associated with any of five surrounding samples denoted as A0, A1, and B0, B1, B2 (202-206 respectively). In H.265, MV prediction can use predictors from the same reference pictures used by adjacent blocks.

[0022] AO Media Video 1 (AV1, AOMedia Video 1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by building on the VP9 codebase and incorporating additional technologies. The AV1 bitstream specification includes reference video codecs such as H.265 or High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC). SUMMARY OF THE INVENTION

[0023] Embodiments of the present disclosure provide methods and apparatus for cross-component sample offset (CCSO) and local sample offset (LSO). Adaptive loop filtering (ALF) uses the reconstructed samples of the first color component as input (e.g., Y, Cb, or Cr). For CCSO, the output is applied to a second color component that is a different color component from the first color component. For LSO, the output is applied to the first color component. The combined ALF may be generalized for CCSO and LSO by considering the delta values between adjacent samples of the sample at the same position (or current) and also considering the level value of the sample at the same position (or current).

[0024] In one embodiment, a method for video decoding includes decoding coding information for reconstructed samples in a current picture from a coded video bitstream, where the coding information includes a sample offset filter applied to the reconstructed samples; selecting an offset type used in the sample offset filter, where the offset type includes a gradient offset (GO) or a band offset (BO); determining an output value of the sample offset filter based on the reconstructed samples and the selected offset type. The method further includes determining a filtered sample value based on the reconstructed samples and the output value of the sample offset filter. The reconstructed samples are from a current component in the current picture. The filtered sample value is for the reconstructed samples. The selecting step further includes receiving a signal indicating the offset type. The signal includes high-level syntax transmitted in a slice header, a picture header, a frame header, a superblock header, a coding tree unit (CTU) header, or a tile header. The signal includes block-level signaling at a coding unit level, a prediction block level, a transform block level, or a filtering unit level. The signal includes a first flag indicating whether the offset is applied to one or more color components, and a second flag indicating whether GO and / or BO is applied. The selecting step includes selecting BO, selecting GO, or selecting both BO and GO. The selection of GO further includes deriving the GO using a delta value between adjacent samples and samples at the same position of a different color component. The selection of GO further includes deriving the GO using a delta value between adjacent samples and samples at the same position of the current sample being filtered. The selection of BO further includes deriving the BO using values of samples at the same position of different color components.The selection of BO further includes deriving BO using the value of the sample at the same position of the current sample to be filtered. When the step of selecting includes the step of selecting both GO and BO, the step of selecting includes the step of deriving an offset using the delta value between the color components different from the adjacent samples or the samples at the same position of any of the currently sampled samples to be filtered, and the step of deriving an offset using the value of the sample at the same position of any of the color components different from the currently sampled samples to be filtered.

[0025] In other embodiments, an apparatus for decoding a video bitstream includes a memory for storing instructions and a processor in communication with the memory. When the processor executes the instructions, the processor causes the apparatus to apply a sample offset filter to the reconstructed samples of the current component within the current picture from the video bitstream, identify an offset type for the sample offset filter, the offset type including a gradient offset (GO) or a band offset (BO), and determine the filtered sample values of the sample offset filter based on the reconstructed samples and the selected offset type. The processor is further configured to cause the apparatus to determine an output value based on the reconstructed samples and the selected offset type, and the filtered sample values are further determined based on the output value and the reconstructed samples. The processor is further configured to cause the apparatus to receive a signal indicating the offset type used for identification. The signal includes a first flag indicating whether the offset is applied to one or more color components and a second flag indicating whether GO and / or BO is applied.

[0026] In another embodiment, a non-transitory computer-readable storage medium stores instructions that, when executed by a processor, cause the processor to apply a sample offset filter to a reconstructed sample of a current component within a current picture from a video bitstream, identify an offset type for the sample offset filter, the offset type including a gradient offset (GO) or a band offset (BO), determine an output value based on the reconstructed sample and the selected offset type, and determine a filtered sample value of the sample offset filter based on the output value and the reconstructed sample. The identifying step includes using a signal having one or more flags indicating the offset type.

[0027] In some other embodiments, a device for processing video information is disclosed. The device may include circuitry configured to execute any one of the implementations of the above method.

[0028] Embodiments of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to execute a method for video decoding and / or encoding.

Brief Description of the Drawings

[0029] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19a

Figure 19b

Figure 19c

Figure 19d

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Mode for Carrying Out the Invention

[0030] Throughout the specification and claims, terms may have meanings with nuances suggested or implied in the context beyond the explicitly described meaning. As used herein, the phrases "one embodiment" or "some embodiments" do not necessarily refer to the same embodiment, and the phrases "another embodiment" or "other embodiments" as used herein do not necessarily refer to different embodiments. Similarly, the phrases "one embodiment" or "some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "another embodiment" or "other embodiments" as used herein do not necessarily refer to different embodiments. For example, the subject matter of the claims is intended to include all or part of an exemplary embodiment / combination of embodiments.

[0031] Generally, terms can be understood at least in part from their use in context. For example, terms such as "and", "or", or "and / or" as used herein may include various meanings that may depend at least in part on the context in which such terms are used. Typically, "or" is intended to mean both inclusive A, B, and C and exclusive A, B, or C as used herein when used to associate a list such as A, B, or C. Further, the terms "one or more" or "at least one" as used herein may be used to describe any feature, structure, or characteristic in a single sense or, alternatively, may be used to describe a combination of features, structures, or characteristics in a plurality of senses, depending at least in part on the context. Similarly, singular terms may be understood as conveying either singular or plural usage, depending at least in part on the context. Further, the terms "based on" or "determined by" may not be intended to convey an exclusive set of elements and, instead, may allow for the presence of additional elements that are not necessarily explicitly described, depending at least in part on the context.

[0032] Figure 3 shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) may perform unidirectional transmission of data. For example, the terminal device (310) may encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission may be implemented in a media service application or the like.

[0033] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which may be implemented, for example, during a video conferencing application. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) may encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) may receive the encoded video data transmitted by the other of the terminal devices (330) and (340), decode the encoded video data to restore the video pictures, and display the video pictures on an accessible display device according to the restored video data.

[0034] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) may be implemented as servers, personal computers, and smartphones, but the applicability of the underlying principles of the present disclosure may not be limited thereto. Embodiments of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, and / or dedicated video conferencing facilities. The network (350) represents any number or type of network that transmits encoded video data among the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit switching, packet switching, and / or another type of channel. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the discussion herein, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless explicitly described below.

[0035] FIG. 4 shows the arrangement of a video encoder and a video decoder in a video streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0036] A video streaming system can include a video source (401), such as a digital camera, and may include a video capture subsystem (413) that generates a stream (402) of, for example, uncompressed video pictures or images. In one example, the stream (402) of video pictures includes samples recorded by the digital camera of the video source 401. The stream (402) of video pictures, drawn as a thick line to emphasize the high data volume when compared to the encoded video data (404) (or encoded video bitstream), can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), drawn as a thin line to emphasize the lower data volume when compared to the stream (402) of uncompressed video pictures, can be stored in a streaming server (405) for future use or can be stored directly in a downstream video device (not shown). One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and generates an output stream (411) of uncompressed video pictures that can be rendered on a display (412) (such as a display screen) or other rendering device (not shown). The video decoder 410 may be configured to perform some or all of the various functions described in this disclosure.In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a particular video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In one example, a video encoding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC and other video encoding standards.

[0037] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can also include a video encoder (not shown).

[0038] FIG. 5 shows a block diagram of a video decoder (510) according to any embodiment of the present disclosure below. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., receiving circuitry). The video decoder (510) can be used in place of the video decoder (310) in the example of FIG. 4.

[0039] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one encoded video sequence may be decoded at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with a plurality of video frames or images. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data together with other data such as encoded audio data and / or auxiliary data streams, and these data may be transferred to respective processing circuits (not shown). The receiver (531) can separate the encoded video sequence from other data. As a countermeasure against network jitter, a buffer memory (515) may be disposed between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser"). In a specific application, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, it can exist separately outside the video decoder (510) (not shown). In still other applications, for example, to counter network jitter, a buffer memory (not shown) may exist outside the video decoder (510), and further, for example, to handle the playback timing, another additional buffer memory (515) may exist inside the video decoder (510). If the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (515) may not be required or may be small. For use in a best-effort packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and its size is relatively large.Such a buffer memory may be implemented with an adaptable size and may be implemented, at least in part, in an operating system or similar element (not shown) external to the video decoder (510).

[0040] Video decoder (510) may include a parser (520) for reconstructing symbols (521) from an encoded video sequence. The categories of these symbols include information used to manage the operation of video decoder (510) and potentially information for controlling a rendering device such as display (512) (e.g., display screen). The display may or may not be an integral part of electronic device (530) and can be coupled to electronic device (530), as shown in FIG. 5. The control information for the rendering device(s) may be in the form of Supplementary Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not shown). Parser (520) can parse / entropy decode the encoded video sequence received by parser (520). The entropy encoding of the encoded video sequence can follow video encoding techniques or standards and can follow various principles including variable length encoding, Huffman encoding, arithmetic encoding with or without context sensitivity, etc. Parser (520) can extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroups can include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. Parser (520) can also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0041] The parser (520) can perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (515), thereby generating symbols (521).

[0042] The reconstruction of the symbols (521) can involve multiple different processes or functional units depending on the type of the encoded video picture or a portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. The units involved and how they are involved may be controlled by subgroup control information parsed by the parser (520) from the encoded video sequence. Such a flow of subgroup control information between the parser (520) and the multiple processes or functional units described below is not depicted for simplicity.

[0043] In addition to the functional blocks already described, the video decoder (510) can conceptually be divided into several functional units as will be described below. In an actual implementation operating under commercial constraints, many of these functional units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, a conceptual subdivision into functional units is adopted in the following disclosure.

[0044] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive quantized transform coefficients and control information as symbols (singular or plural) (521) from the parser (520). The control information includes information indicating which type of inverse transform to use, block size, quantization coefficient / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block including sample values that can be input to the aggregator (555).

[0045] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may use surrounding block information that has already been reconstructed and stored in the current picture buffer (558) to generate blocks of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. Depending on the implementation, the aggregator (555) may add, for each sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0046] In other cases, the output samples of the scaler / inverse transform unit (551) can relate to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to retrieve the samples used for inter-picture prediction. After motion-compensating the retrieved samples according to the symbols (521) regarding the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (the output of unit 551 may be referred to as residual samples or a residual signal), thereby generating output sample information. The address in the reference picture memory (557) from which the motion compensation unit (553) retrieves the prediction samples can be controlled by the motion vectors available to the motion compensation unit (553) in the form of symbols (521). The symbols can have, for example, X, Y components (shifts), and a reference picture component (time). Motion compensation may include interpolation of the sample values fetched from the reference picture memory (557) when exact motion vectors below the sample level are used, and may also be related to motion vector prediction mechanisms and the like.

[0047] The output samples of the aggregator (555) can be subjected to various loop filtering techniques within the loop filter unit (556). Video compression techniques can include in-loop filter techniques. The in-loop filter techniques are controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but can also respond to meta-information obtained during the decoding of the previous part (in decoding order) of the coded picture or coded video sequence, and to previously reconstructed and loop-filtered sample values. As will be explained in more detail below, some types of loop filters may be included as part of the loop filter unit 556 in various orders.

[0048] The output of the loop filter unit (556) can be a sample stream, which can be output to the rendering device (512) and can also be stored in the reference picture memory (557) for use in future inter-picture prediction.

[0049] Once a particular encoded picture is fully reconstructed, it can be used as a reference picture for future inter-picture prediction. For example, when the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a fresh current picture buffer can be reallocated before starting the reconstruction of subsequent encoded pictures.

[0050] The video decoder (510) can perform a decoding operation according to a predetermined video compression technique adopted by a standard such as ITU-T Recommendation H.265. The encoded video sequence can conform to the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard, in the sense that the encoded video sequence can conform to the syntax defined by the video compression technique or standard being used. Specifically, the profile can select specific tools from all the tools available in the video compression technique or standard as the tools that are only available for use under that profile. To comply with the standard, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (e.g., measured in megasamples per second), the maximum reference picture size, etc. The limits set by the level can, in some cases, be further restricted through the virtual reference decoder (Hypothetical Reference Decoder, HRD) specifications and metadata signaled in the encoded video sequence for HRD buffer management.

[0051] In some exemplary embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0052] FIG. 6 shows a block diagram of a video encoder (603) according to an exemplary embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.

[0053] The video encoder (603) can receive video samples from a video source (601) that can capture a video image to be encoded by the video encoder (603) (which is not part of the electronic device (620) in the example of FIG. 6). In another example, the video source (601) may be implemented as part of the electronic device (620).

[0054] The video source (601) can provide a source video sequence to be encoded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 YCrCB, RGB, XYZ,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) may be a storage device capable of storing pre-prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures or images that impart motion when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0055] According to some exemplary embodiments, a video encoder (603) can encode and compress pictures of a source video sequence in real time or under any other temporal constraints required by an application to produce an encoded video sequence (643). Enforcing an appropriate encoding speed constitutes one function of a controller (650). In some embodiments, the controller (650) may be functionally coupled to and control other functional units as described below. Such couplings are not depicted for the sake of brevity. Parameters set by the controller (650) can include parameters related to rate control (picture skip, quantizer, lambda value of rate-distortion optimization techniques, …), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions related to the video encoder (603) optimized for a particular system design.

[0056] In some exemplary embodiments, the video encoder (603) may be configured to operate in an encoding loop. As a simplified and illustrative example, in one instance, the encoding loop may include a source encoder (630) (e.g., responsible for generating symbols such as a symbol stream based on an input picture and reference picture(s) to be encoded), and a (local) decoder (633) embedded within the video encoder (603). Even when the embedded decoder 633 processes the encoded video stream without entropy coding by the source coder 630, the decoder (633) reconstructs the symbols to generate sample data in a manner similar to what a (remote) decoder would also generate. (In the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols in entropy coding and the encoded video bitstream can be lossless.) The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in a bit-exact result regardless of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local encoder and the remote encoder. In other words, the prediction unit of the encoder "sees" the same sample values as reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift in the event that synchronization cannot be maintained, e.g., due to channel errors) is used to improve the encoding quality.

[0057] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder, such as the video decoder (410), which has already been described in detail above in connection with FIG. 5. However, referring briefly to FIG. 5 as well, since symbols are available and the encoding / decoding of the symbol into the encoded video sequence by the entropy encoder (645) and the parser (420) can be reversible, the entropy decoding section of the video decoder (410) including the buffer memory (415) and the parser (420) may not be fully implemented in the local decoder (633) of the encoder.

[0058] An observation that can be made at this point is that any decoder technology, except for parse / entropy decoding that may exist only within the decoder, may need to exist in substantially the same functional form within the corresponding encoder. For this reason, the disclosed subject matter may sometimes focus on decoder operation. This is the same as the decoding part of the encoder. Therefore, the description of encoder technology can be omitted since it is the reverse of the decoder technology described comprehensively. More detailed explanations are provided below only for specific areas or aspects of the encoder.

[0059] During operation, in some exemplary implementations, the source encoder (630) can perform motion-compensated predictive encoding that predictively encodes the input picture by referring to one or more previously encoded pictures from the video sequence designated as "reference pictures". In this way, the encoding engine (632) encodes the difference (or residual) in the color channels between the pixel block of the input picture and the pixel block(s) of the reference picture(s) that can be selected as the prediction reference for the input picture. The terms "residual" and its derivative form "residual of" may be used interchangeably.

[0060] The local video decoder (633) can decode the encoded video data of a picture that can be specified as a reference picture based on the symbols generated by the source encoder (630). The operation of the encoding engine (632) can advantageously be a lossy process. When the encoded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence can typically be a replica of the source video sequence with some errors. The local video decoder (633) can replicate the decoding process that can be performed on the reference picture by the video decoder and cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture that has (in the absence of transmission errors) common content as the reconstructed reference picture that would be obtained by a remote video decoder.

[0061] The predictor (635) can perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) can search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that can function as an appropriate prediction reference for the new picture. The predictor (635) can operate on a sample block-by-pixel block basis to find an appropriate prediction reference. In some cases, depending on what is determined by the search results obtained by the predictor (635), the input picture can have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (634).

[0062] The controller (650) may manage the encoding operation of the source encoder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0063] The outputs of all the above functional units can undergo entropy encoding in the entropy encoder (645). The entropy encoder (645) converts the symbols generated by various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.

[0064] The transmitter (640) can put the encoded video sequence generated by the entropy encoder (645) into a buffer and prepare it for transmission via the communication channel (660). The communication channel (660) may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) can merge the encoded video data from the video encoder (630) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0065] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) can assign a certain encoded picture type to each encoded picture. The encoded picture type can affect the encoding technique applicable to each picture. For example, a picture may often be assigned as one of the following picture types.

[0066] An intra picture (I picture) can be encoded and decoded without using other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, for example, including Independent Decoder Refresh (IDR) pictures. Those skilled in the art recognize these variations of I pictures, as well as their respective uses and characteristics.

[0067] A predicted picture (P picture) can be encoded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0068] A bi-directionally predicted picture (B picture) can be encoded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predicted picture can use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0069] A source picture is typically divided spatially into a plurality of sample-encoded blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be encoded block by block. The blocks can be encoded predictively by referring to other (already encoded) blocks as determined by the encoding assignment applied to each picture of the block. For example, blocks of an I picture may be encoded non-predictively or predictively by referring to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be encoded predictively via spatial prediction or via temporal prediction by referring to one previously encoded reference picture. Blocks of a B picture may be encoded predictively via spatial prediction or via temporal prediction by referring to one or two previously encoded reference pictures. A source picture or an intermediate processed picture may be subdivided into other types of blocks for other purposes. As will be explained in more detail below, the division of the encoded blocks and other types of blocks may or may not follow the same scheme.

[0070] The video encoder (603) can perform an encoding operation according to a predetermined video encoding technology or standard such as ITU-T Recommendation H.265. In this operation, the video encoder (603) can perform various compression operations including a predictive encoding operation that utilizes the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technology or standard used.

[0071] In some exemplary embodiments, the transmitter (640) may transmit additional data together with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0072] Video may be captured as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation in a given picture, and inter-picture prediction utilizes the temporal or other correlation between pictures. For example, a particular picture to be encoded / decoded, called the current picture, may be divided into blocks. If a block within the current picture is similar to a reference block within a reference picture that has been previously encoded and is still in the buffer in the video, it may be encoded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that specifies the reference picture when multiple reference pictures are used.

[0073] In some exemplary embodiments, bidirectional prediction techniques can be used in inter-picture prediction. According to such bidirectional prediction techniques, two reference pictures such as a first reference picture and a second reference picture that both precede the current picture in decoding order in the video (however, in display order, they may be past or future, respectively) are used. A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted together by a combination of the first reference block and the second reference block.

[0074] Furthermore, in order to improve the coding efficiency, merge mode techniques may be used in inter-picture prediction.

[0075] According to some exemplary embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and those CTUs in the picture may have the same size such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU may include three parallel coding tree blocks (CTBs) which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU or four 32×32 pixel CUs. Each of the one or more 32×32 blocks may be further divided into four 16×16 pixel CUs. In some exemplary embodiments, each CU may be analyzed during coding to determine a prediction type for that CU among various prediction types such as an inter prediction type or an intra prediction type. A CU may be divided into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. The division of a CU into PUs (or PBs of different color channels) may be performed in various division patterns. For example, a luma or chroma PB may include a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0076] FIG. 7 shows a diagram of a video encoder (703) according to another exemplary embodiment of the present disclosure. The video encoder (703) receives a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures, and is configured to encode the processing block into an encoded picture that is part of an encoded video sequence. The exemplary video encoder (703) may be used in place of the video encoder (403) in the example of FIG. 4.

[0077] For example, the video encoder (703) receives a matrix of sample values for a processing block such as a prediction block of 8×8 samples. Next, the video encoder (703) determines which of the intra mode, inter mode, or bi - directional prediction mode the processing block is best encoded using, for example, rate - distortion optimization (RDO). If it is determined that the processing block is encoded in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into the encoded picture. If it is determined that the processing block is encoded in the inter mode or bi - directional prediction mode, the video encoder (703) may use inter prediction techniques or bi - directional prediction techniques, respectively, to encode the processing block into the encoded picture. In some exemplary embodiments, the merge mode may be used as a sub - mode of inter - picture prediction where the motion vector is derived from one or more motion vector predictors but there is no benefit of the encoded motion vector components outside the predictors. In some exemplary embodiments, there may be motion vector components applicable to the target block. Thus, the video encoder (703) may include components not explicitly shown in FIG. 7, such as a mode decision module (not shown) for determining the prediction mode of the processing block.

[0078] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in the exemplary arrangement of FIG. 7.

[0079] The inter-encoder (730) receives samples of a current block (e.g., a processing block), compares the block with one or more reference blocks (e.g., blocks in previous and subsequent pictures in display order) in a reference picture, generates inter-prediction information (e.g., a description of redundant information by an inter-coding technique, a motion vector, merge mode information), and based on the inter-prediction information, is configured to calculate an inter-prediction result (e.g., a predicted block) using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded using a decoding unit 633 (shown as the residual decoder 728 of FIG. 7, described in more detail below) embedded in the exemplary encoder 620 of FIG. 6 based on the encoded video information.

[0080] The intra-encoder (722) receives samples of a current block (e.g., a processing block), compares the block with blocks already encoded in the same picture, generates quantized coefficients after transformation, and optionally also generates intra-prediction information (e.g., intra-prediction direction information by one or more intra-coding techniques). The intra-encoder (722) may also calculate an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks in the same picture.

[0081] The overall controller (721) may be configured to determine overall control data and control other components of the video encoder (703) based on the overall control data. In one example, the overall controller (721) determines a prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is the intra mode, the overall controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), selects the intra prediction information, and controls the entropy encoder (725) to include the intra prediction information in the bitstream. When the prediction mode of the block is the inter mode, the overall controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), selects the inter prediction information, and controls the entropy encoder (725) to include the inter prediction information in the bitstream.

[0082] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and the prediction result of that block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transform coefficients. For example, the residual encoder (724) may be configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. Next, the transform coefficients are subjected to quantization processing to obtain quantized transform coefficients. In various exemplary embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation to generate decoded residual data. The decoded residual data can be suitably used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is suitably processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture.

[0083] The entropy encoder (725) is configured to format the bitstream to include the encoded block and perform entropy encoding. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. When encoding a block in either the inter mode or the merge submode of the bidirectional prediction mode, the residual information may not be present.

[0084] FIG. 8 shows a diagram of an exemplary video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive an encoded picture that is part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (810) may be used in place of the video decoder (410) in the example of FIG. 4.

[0085] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) coupled together as shown in the exemplary configuration of FIG. 8.

[0086] The entropy decoder (871) can be configured to reconstruct from the encoded picture specific symbols that represent the syntax elements that the encoded picture is composed of. Such symbols can include, for example, the mode in which a block is encoded (e.g., intra mode, inter mode, bi - directional prediction mode, merge sub - mode or another sub - mode), prediction information (e.g., intra prediction information or inter prediction information) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), and residual information in the form of, for example, quantized transform coefficients. In one example, when the prediction mode is an inter or bi - directional prediction mode, the inter prediction information is provided to the inter decoder (880). When the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can undergo inverse quantization and is provided to the residual decoder (873).

[0087] The inter decoder (880) may be configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0088] The intra decoder (872) may be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0089] The residual decoder (873) may be configured to perform inverse quantization to extract the dequantized transform coefficients, process the dequantized transform coefficients, and convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize specific control information (including quantization parameter (QP)), and the information may be provided by the entropy decoder (871) (since this is only low-data-volume control information, the data path is not depicted).

[0090] The reconstruction module (874) may be configured to combine, in the spatial domain, the residual output by the residual decoder (873) and the prediction result (output by the intra or inter prediction module as appropriate) to form a reconstructed block that forms a part of the reconstructed picture as part of the reconstructed video. Note that other suitable operations such as a deblocking operation may also be performed to improve visual quality.

[0091] Note that the video encoders (403), (603), (703) and the video decoders (410), (510), (810) can be implemented using any suitable technology. In some exemplary embodiments, the video encoders (403), (603), (703) and the video decoders (410), (510), (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), (703) and the video decoders (410), (510), (810) can be implemented using one or more processors that execute software instructions.

[0092] Turning to the partitioning of blocks for symbolization and decoding, a general partition may start from a base block and may follow a predetermined set of rules, a specific pattern, a partition tree, or any partition structure or scheme. The partition may be hierarchical and recursive. After splitting or partitioning the base block according to any of the exemplary partitioning procedures shown below, other procedures, or combinations thereof, a set of final partitions or coded blocks may be obtained. Each of these partitions is one of the various partition levels within the partition hierarchy and may be of various shapes. Each partition may be referred to as a coding block (CB). For the various exemplary partition implementations described further below, each resulting CB may be of either an acceptable size or partition level. Such a partition forms a unit in which several basic encoding / decoding decisions are made and encoding / decoding parameters are optimized and determined to be signaled in the encoded video bitstream, and thus is called a coding block. The highest or deepest level of the final partition represents the depth of the coding block partition structure of the tree. The coding block may be a luma coding block or a chroma coding block. The CB tree structure for each color may be referred to as a coding block tree (CBT).

[0093] The coding blocks for all color channels may together be referred to as a coding unit (CU). The hierarchical structure for all color channels may together be referred to as a coding tree unit (CTU). The partition pattern or structure of the various color channels within a CTU may be the same or may not be the same.

[0094] In some implementations, the partition tree method or structure used for the luma channel and the chroma channel need not be the same. In other words, the luma channel and the chroma channel may have separate coding tree structures or patterns. Further, whether the luma channel and the chroma channel use the same coding partition tree structure or different coding partition tree structures, and the actual coding partition tree structure used, may depend on whether the slice being coded is a P slice, a B slice, or an I slice. For example, for an I slice, the chroma channel and the luma channel may have separate coding partition tree structures or coding partition tree structure modes, but for a P or B slice, the luma channel and the chroma channel may share the same coding partition tree structure. When separate coding partition tree structures or modes are applied, the luma channel may be divided into CB by one coding partition tree structure, and the chroma channel may be divided into chroma CB by another coding partition tree structure.

[0095] In some exemplary implementations, a predetermined partitioning pattern may be applied to the base block. As shown in FIG. 9, four exemplary partitioning trees may start from a first predetermined level (e.g., the 64×64 block level or other size as the base block size), and the base block may be hierarchically partitioned down to a predetermined lowest level (e.g., the 4×4 level). For example, the base block may be subject to four predetermined partitioning options or patterns shown as 902, 904, 906, and 908, and a partitioning shown as R allows a recursive partitioning where the same partitioning option shown in FIG. 9 can be repeated at a lower scale down to the lowest level (e.g., the 4×4 level). In some implementations, further restrictions may be applied to the partitioning method of FIG. 9. In the implementation of FIG. 9, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be allowed but may not be allowed to be recursive, while square partitions are allowed to be recursive. The recursive partitioning according to FIG. 9 generates a set of final encoded blocks as needed. An encoding tree depth may be further defined to indicate the depth of the split from the root node or root block. For example, the encoding tree depth of a root node or root block, e.g., a 64×64 block, may be set to 0, and after the root block is further split once according to FIG. 9, the encoding tree depth increases by 1. The maximum level or the deepest level from the 64×64 base block to the 4×4 minimum partition is 4 in the above method (starting from level 0). Such a partitioning method may be applied to one or more of the color channels. Each color channel may be independently partitioned according to the method of FIG. 9 (e.g., the partitioning pattern or option in a predetermined pattern may be determined independently for each color channel at each hierarchical level). Alternatively, two or more of the color channels may share the same hierarchical pattern tree of FIG. 9 (e.g., the same partitioning pattern or option in a predetermined pattern may be selected for two or more color channels at each hierarchical level).

[0096] FIG. 10 shows another exemplary predetermined partition pattern that allows recursive partitioning to form a partition tree. As shown in FIG. 10, for example, ten partition structures or patterns may be predefined. The root block may start at a predetermined level (e.g., from a 128×128 level or a 64×64 level base block). The exemplary partition structure of FIG. 10 includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. The partition type having three sub - partitions shown as 1002, 1004, 1006, and 1008 in the second row of FIG. 10 may be called a "T - type" partition. The "T - type" partitions 1002, 1004, 1006, and 1008 may be called left T - type, upper T - type, right T - type, and lower T - type, respectively. In some exemplary implementations, none of the rectangular partitions of FIG. 10 are allowed to be further subdivided. To indicate the depth of the split from the root node or root block, an encoding tree depth may be further defined. For example, in the case of a 128×128 block, the encoding tree depth of the root node or root block may be set to 0, and after the root block is further split once according to FIG. 10, the encoding tree depth increases by only 1. In some implementations, only the all - square partitions at 1010 may be allowed recursive partitioning into the next - level partition tree according to the pattern of FIG. 10. In other words, in the square partitions within the T - type patterns 1002, 1004, 1006, and 1008, recursive partitioning may not be allowed. The partitioning procedure according to FIG. 10 by recursion generates a set of final encoded blocks as needed. Such a method may be applied to one or more of the color channels. In some implementations, additional flexibility may be added to the use of partitions at levels 8×8 and below. For example, 2×2 chroma - in - ter prediction may be used in certain cases.

[0097] In some other implementations of the partitioning of the symbolization block, a quadtree structure may be used to divide the base block or the intermediate block into quadtree partitions. Such quadtree partitioning may be applied hierarchically and recursively to any square partition. Whether the base block or the intermediate block or the partition is further quadtree-divided may be adapted to various local characteristics of the base block or the intermediate block / partition. Quadtree partitioning at the picture boundary may be further adapted. For example, an implicit quadtree partitioning may be performed at the picture boundary such that the block maintains quadtree partitioning until its size matches the picture boundary.

[0098] In some other exemplary implementations, a hierarchical binary partition from a base block may be used. In such a scheme, a base block or an intermediate level block may be partitioned into two partitions. The binary partition may be either horizontal or vertical. For example, a horizontal binary partition may divide a base block or an intermediate block into equal left and right partitions. Similarly, a vertical binary partition may divide a base block or an intermediate block into equal upper and lower partitions. Such binary partitions may be hierarchical and recursive. In each of the base blocks or intermediate blocks, a decision may be made as to whether the binary partition scheme should continue, and if the scheme is to continue, whether a horizontal binary partition scheme should be used or a vertical binary partition should be used. In some implementations, further partitioning may stop at a predetermined minimum partition size (in one or both dimensions). Alternatively, further partitioning may stop when a predetermined partition level or depth from the base block is reached. In some implementations, the aspect ratio of the partitions may be limited. For example, the aspect ratio of the partitions may not be less than 1:4 (or greater than 4:1). Thus, a vertical strip partition having a vertical-to-horizontal aspect ratio of 4:1 may be further vertically binary partitioned only into upper and lower partitions each having a vertical-to-horizontal aspect ratio of 2:1.

[0099] In some further examples, as shown in FIG. 13, a three-way partitioning scheme may be used to partition the base block or any intermediate block. The three-way pattern may be implemented vertically as shown at 1302 in FIG. 13, or horizontally as shown at 1304 in FIG. 13. The exemplary split ratio in FIG. 13 is shown as 1:2:1 vertically or horizontally, although other ratios may be predefined. In some implementations, two or more different ratios may be predefined. Such a three-way partitioning scheme may be used to complement a quadtree or binary partitioning structure, and such a ternary partition can capture an object located at the center of a block within one continuous partition, while a quadtree and a binary tree always split along the block center and thus split the object into separate partitions. In some implementations, to avoid further conversions, the width and height of an exemplary ternary partition are always powers of two.

[0100] The above partitioning methods may be combined at different partition levels in any of the methods. As an example, the above quadtree and binary partition methods may be combined to partition a base block into a quadtree-binary-tree (QTBT) structure. In such a method, a base block or an intermediate block / partition may be quadtree or binary partitioned according to a set of specified conditions, if specified. A specific example is shown in FIG. 14. In the example of FIG. 14, as shown at 1402, 1404, 1406, and 1408, the base block is first quadtree partitioned into four partitions. Thereafter, each of the resulting partitions may be quadtree partitioned into four further partitions (such as 1408, etc.), or may be binary partitioned into two further partitions at the next level (such as either 1402 or 1406, etc., horizontal or vertical, for example both symmetric), or may not be partitioned (such as 1404, etc.). Binary or quadtree partitioning may be recursively allowed for square partitions, as shown in the overall exemplary partition pattern of 1410 and the corresponding tree structure / representation of 1420. Here, solid lines represent quadtree partitioning and dashed lines represent binary partitioning. A flag may be used for each binary partition node (non-leaf binary partition) to indicate whether the binary partition is horizontal or vertical. For example, as shown in 1420, according to the partition structure of 1410, the flag "0" may represent a horizontal binary partition and the flag "1" may represent a vertical binary partition. For quadtree partitioned partitions, since quadtree partitioning always partitions a block or partition both horizontally and vertically to produce four sub-blocks / partitions of the same size, there is no need to indicate the partition type. In some implementations, the flag "1" may represent a horizontal binary partition and the flag "0" may represent a vertical binary partition.

[0101] In some exemplary implementations of QTBT, the quadtree and binary rule sets may be represented by the following specified parameters and corresponding associated functions. - CTU size: The size of the root node of the quadtree (the size of the base block) -MinQTSize: Minimum allowable quadtree leaf node size -MaxBTSize: Maximum allowable binary tree root node size -MaxBTDepth: Maximum allowable binary tree depth -MinBTSize: Minimum allowable binary tree leaf node size In some exemplary implementations of the QTBT partitioning structure, the CTU size may be set as 128×128 luma samples at two corresponding 64×64 blocks of chroma samples (when exemplary chroma subsampling is considered and used), MinQTSize may be set as 16×16, MaxBTSize may be set as 64×64, MinBTSize (both width and height) may be set as 4×4, and MaxBTDepth may be set as 4. Quadtree partitioning may be first applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes may have sizes ranging from its minimum allowable size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the node is 128×128, since its size exceeds MaxBTSize (i.e., 64×64), it is not first divided by the binary tree. Otherwise, nodes not exceeding MaxBTSize may be partitioned by the binary tree. In the example of FIG. 14, the base block is 128×128. The base block can only be quad-tree partitioned according to a given set of rules. The base block has a partition depth of 0. Each of the resulting four partitions is 64×64, does not exceed MaxBTSize, and may be further quad-tree or binary-tree partitioned at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning may not be considered. If the binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal partitioning may not be considered. Similarly, if the binary tree node has a height equal to MinBTSize, further vertical partitioning is not considered.

[0102] In some exemplary implementations, the above QTBT scheme may be configured to support the flexibility that luma and chroma have the same QTBT structure or separate QTBT structures. For example, for P slices and B slices, the luma and chroma CTBs within one CTU may share the same QTBT structure. However, for I slices, the luma CTB may be divided into CUs by one QTBT structure, and the chroma CTB may be divided into chroma CUs by another QTBT structure. This means that the CU may be used to refer to different color channels within the I slice. For example, the I slice may be composed of coded blocks of the luma component or coded blocks of two chroma components, and the CUs within the P or B slice may be composed of coded blocks of all three color components.

[0103] In some other implementations, the QTBT scheme may be supplemented with the above-mentioned three-way split scheme. Such an implementation may be referred to as a multi-type-tree (MTT) structure. For example, in addition to the binary split of nodes, one of the three-way partition patterns in FIG. 13 may be selected. In some implementations, only square nodes may be subject to the three-way split. An additional flag may be used to indicate whether the three-way partition is horizontal or vertical.

[0104] The design of two-level or multi-level trees, such as QTBT implementations and QTBT implementations supplemented by three-way splits, can be motivated mainly by reducing complexity. Theoretically, the complexity of traversing the tree is T D where T represents the number of split types and D represents the depth of the tree. A trade-off may be made by using multiple types (T) while reducing the depth (D).

[0105] In some implementations, the CB may be further partitioned. For example, the CB may be further partitioned into a plurality of prediction blocks (PBs) for the purpose of intra-frame or inter-frame prediction during the encoding and decoding processes. In other words, the CB may be further divided into different sub-partitions where individual prediction decisions / configurations may be made. In parallel, the CB may be further partitioned into a plurality of transform blocks (TBs) for the purpose of describing the level at which the video data is transformed or inverse-transformed. The partitioning scheme of the CB into PBs and TBs may or may not be the same. For example, each partitioning scheme may be performed using a unique procedure based on, for example, various characteristics of the video data. In some exemplary implementations, the PB and TB partitioning schemes may be independent. In some other exemplary implementations, the PB and TB partitioning schemes and boundaries may be correlated. In some implementations, for example, the TB may be partitioned after the partitioning of the PB, and in particular, each PB may be further partitioned into one or more TBs after being determined following the partitioning of the coding block. For example, in some implementations, the PB may be divided into 1, 2, 4 or other numbers of TBs.

[0106] In some implementations, the luma and chroma channels may be treated differently in order to partition a base block into coding blocks and further partition the coding blocks into prediction blocks and / or transform blocks. For example, in some implementations, partitioning of a coding block into prediction blocks and / or transform blocks may be allowed for the luma channel, but such partitioning of a coding block into prediction blocks and / or transform blocks may not be allowed for the chroma channel. Thus, in such implementations, transformation and / or prediction of luma blocks may be performed only at the coding block level. In other examples, the minimum transform block sizes for the luma and chroma channels may be different. For example, a coding block of the luma channel may be allowed to be partitioned into smaller transform blocks and / or prediction blocks than the chroma channel. In yet other examples, the maximum depth of partitioning of a coding block into transform blocks and / or prediction blocks may be different between the luma and chroma channels. For example, a coding block of the luma channel may be allowed to be partitioned into deeper transform blocks and / or prediction blocks than the chroma channel. In a specific example, a luma coding block may be partitioned into transform blocks of multiple sizes that can be represented by a recursive partitioning up to a maximum of two levels, and transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4 and transform block sizes from 4×4 to 64×64 may be allowed. However, for chroma blocks, only the maximum possible transform blocks specified for luma blocks may be allowed.

[0107] In some exemplary implementations for partitioning a coding block into PBs, the depth, shape, and / or other characteristics of the PB partition may depend on whether the PB is intra-coded or inter-coded.

[0108] The partitioning of the symbolization block (or prediction block) into transformation blocks may be implemented in various exemplary ways, including but not limited to quadtree partitioning and predetermined pattern partitioning, either recursively or non-recursively, further considering the transformation blocks at the boundaries of the symbolization block or prediction block. Generally, the resulting transformation blocks may be at different partitioning levels, may not be of the same size, and need not be square in shape (e.g., they can be rectangular with some allowed sizes and aspect ratios). Further examples are described in more detail below in connection with FIGS. 15, 16, and 17.

[0109] However, in some other implementations, the CBs obtained via any of the above partitioning methods may be used as the basic or minimum symbolization blocks for prediction and / or transformation. In other words, no further partitioning is performed for the purposes of inter prediction / intra prediction and / or transformation. For example, the CBs obtained from the above QTBT method may be directly used as units for performing prediction. Specifically, such a QTBT structure removes the concept of multiple partition types, i.e., it removes the separation of CU, PU, and TU, and supports greater flexibility in the CU / CB partition shape as described above. In such a QTBT block structure, the CU / CB can have either a square or rectangular shape. The leaf nodes of such a QTBT are used as units for prediction and transformation processing without further partitioning. This means that in such an exemplary QTBT symbolization block structure, the CU, PU, and TU have the same block size.

[0110] Any of the various CB partitioning methods described above may be combined with further partitioning of the CB into PB and / or TB (including no PB / TB partitioning) in any manner. The following specific implementations are provided as non-limiting examples.

[0111] A specific exemplary implementation of the partitioning of the symbolization block and the conversion block will be described below. In such an exemplary implementation, a recursive quadtree partitioning, or a predetermined partitioning pattern (such as those in FIGS. 9 and 10, etc.), may be used to divide the base block into symbolization blocks. At each level, whether to continue the further quadtree partitioning of a specific partition may be determined by the local video data characteristics. The resulting CBs may be of various quadtree partitioning levels and various sizes. The decision of whether to code the picture area using inter-picture (temporal) or intra-picture (spatial) prediction may be made at the CB level (or at the CU level for all three color channels). Each CB may be further divided into 1, 2, 4, or other numbers of PBs according to a predetermined PB partitioning type. Inside one PB, the same prediction process may be applied, and the relevant information may be sent to the decoder for each PB. After obtaining the residual block by applying the prediction process based on the PB partitioning type, the CB may be partitioned into TBs according to another quadtree structure similar to the coding tree of the CB. In this specific implementation, the CB or TB may be square, but it is not necessary to be limited to this. Further, in this specific example, the PB may be square or rectangular for inter prediction and square only for intra prediction. The symbolization block may be divided, for example, into four square TBs. Each TB may be further divided recursively (using quadtree partitioning) into smaller TBs called residual quadtree (RQT).

[0112] Another exemplary implementation for partitioning the base block into CB, PB, or TB will be further described below. For example, instead of using a type of multiple partition units as shown in FIG. 9 or FIG. 10, a quadtree having a nested multi-type tree using a binary and ternary segmentation structure (e.g., QTBT as described above or QTBT by ternary segmentation) may be used. The separation of CB, PB, and TB (i.e., partitioning of CB into PB and / or TB, and partitioning of PB into TB) may be abandoned, except when necessary for a CB having a size too large for the maximum transform length and further partitioning may be required. This exemplary partitioning scheme may be designed to support greater flexibility in the CB partition shape so that both prediction and transformation can be performed at the CB level without further partitioning. In such an encoding tree structure, the CB may have either a square or rectangular shape. Specifically, the coding tree block (CTB) may first be partitioned by a quadtree structure. Then, the quadtree leaf nodes may be further partitioned by a nested multi-type tree structure. An example of a nested multi-type tree structure using binary or ternary is shown in FIG. 11. Specifically, the exemplary multi-type tree structure in FIG. 11 includes four split types called vertical binary split (SPLIT_BT_VER) (1102), horizontal binary split (SPLIT_BT_HOR) (1104), vertical ternary split (SPLIT_TT_VER) (1106), and horizontal ternary split (SPLIT_TT_HOR) (1108). Then, the CB corresponds to the leaf of the multi-type tree. In this exemplary implementation, the CB is used for both prediction and transformation processing without further partitioning, except when the CB is too large for the maximum transform length. This means that in most cases, CB, PB, and TB have the same block size in a quadtree having a nested multi-type tree encoding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color components of the CB.In some implementations, in addition to two-way or three-way partitioning, the nested pattern of FIG. 11 may further include quadtree partitioning.

[0113] One specific example of a quadtree using a nested multi-type tree coding block structure of block partitions (including quadtree, two-way, and three-way options) for one base block is shown in FIG. 12. More specifically, FIG. 12 shows that base block 1200 is quadtree partitioned into four square partitions 1202, 1204, 1206, and 1208. The decision to further use the multi-type tree structure of FIG. 11 and the quadtree for further partitioning is made for each of the quadtree partitioned partitions. In the example of FIG. 12, partition 1204 is not further partitioned. Partitions 1202 and 1208 each adopt another quadtree partition. In partition 1202, the second-level quadtree partitioned upper-left, upper-right, lower-left, and lower-right partitions each adopt a third-level partition of a quadtree, the horizontal partition 1104 of FIG. 11, non-partitioning, and the horizontal three-way partition 1108 of FIG. 11. Partition 1208 adopts another quadtree partition, and the upper-left, upper-right, lower-left, and lower-right partitions of the second-level quadtree partition each adopt a third-level partition of the vertical three-way partition 1106, non-partitioning, non-partitioning, and the horizontal two-way partition 1104 of FIG. 11. Two of the sub-partitions of the third-level upper-left partition of 1208 are further partitioned according to the horizontal two-way partition 1104 and the horizontal three-way partition 1108 of FIG. 11, respectively. Partition 1206 adopts a second-level partitioning pattern into two partitions according to the vertical two-way partition 1102 of FIG. 11, and the two partitions are further partitioned at the third level according to the horizontal three-way partition 1108 and the vertical two-way partition 1102 of FIG. 11. The fourth-level partition is further applied to one of these according to the horizontal two-way partition 1104 of FIG. 11.

[0114] In the above specific example, the maximum luma transform size may be 64×64, and the maximum supported chroma transform size may be different from the luma, for example, 32×32. The exemplary CB in FIG. 12 is generally not further divided into smaller PBs and / or TBs, but if the width or height of the luma coding block or chroma coding block is larger than the maximum transform width or height, the luma coding block or chroma coding block may be automatically divided in the horizontal and / or vertical directions to meet the limit of the transform size in that direction.

[0115] In a specific example of partitioning the base block into the above CBs, as described above, the coding tree method may support the ability for luma and chroma to have separate block tree structures. For example, for P slices and B slices, the luma CTB and chroma CTB within one CTU may share the same coding tree structure. For example, for I slices, luma and chroma may have separate coding block tree structures. When separate block tree structures are applied, the luma CTB may be partitioned into luma CBs by one coding tree structure, and the chroma CTB may be partitioned into chroma CBs by another coding tree structure. This means that the CUs within an I slice may be composed of coding blocks of the luma component or coding blocks of two chroma components, and CUs within P or B slices are always composed of coding blocks of all three color components as long as the video is not monochrome.

[0116] When the symbolization block is further partitioned into a plurality of transform blocks, the transform blocks therein may be ordered within the bitstream according to various orders or scan patterns. Exemplary implementations for partitioning the symbolization block or the prediction block into transform blocks, and the coding order of the transform blocks, will be described in further detail below. In some exemplary implementations, as described above, the partitioning of the transform may support transform block sizes from, for example, 4×4 to 64×64, and multiple shapes, for example, 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1 transform blocks. In some implementations, when the symbolization block is 64×64 or less, the partitioning of the transform block may be applied only to the luma component, and for the chroma block, the transform block size will be the same as the symbolization block size. Otherwise, when the width or height of the symbolization block is greater than 64, both the luma symbolization block and the chroma symbolization block may be implicitly divided into transform blocks that are multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32), respectively.

[0117] In some exemplary implementations of the partitioning of the transform block, for both the intra-coded block and the inter-coded block, the symbolization block may be further partitioned into a plurality of transform blocks with a partitioning depth up to a predetermined number of levels (e.g., 2 levels). The partitioning depth and size of the transform block may be related. In some exemplary implementations, the mapping from the transform size at the current depth to the transform size at the next depth is shown in Table 1 below.

Table 1

[0118] Based on the exemplary mapping in Table 1, for a 1:1 square block, the next level of transform partitioning may create four 1:1 square sub-transform blocks. The transform partition may stop, for example, at 4×4. Thus, the transform size of the current depth of 4×4 corresponds to the same size of 4×4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next level of transform partitioning may create two 1:1 square sub-transform blocks, while for a 1:4 / 4:1 non-square block, the next level of transform partitioning may create two 1:2 / 2:1 sub-transform blocks.

[0119] In some exemplary implementations, additional restrictions may apply regarding the partitioning of the transform block for the luma component of an intra-coded block. For example, for each level of the transform partition, all sub-transform blocks may be restricted to have the same size. For example, for a 32×16 coded block, the level 1 transform partition creates two 16×16 sub-transform blocks, and the level 2 transform partition creates eight 8×8 sub-transform blocks. In other words, to keep the transform units the same size, the level 2 split must be applied to all level 1 sub-blocks. An example of the partitioning of the transform block of an intra-coded square block according to Table 1 is shown in FIG. 15 along with the coding order indicated by the arrows. Specifically, 1502 shows a square coded block. The level 1 split into four equal-sized transform blocks according to Table 1 is shown at 1504 in the coding order indicated by the arrows. The level 2 split of all equal-sized blocks at level 1 into 16 equal-sized transform blocks according to Table 1 is shown at 1506 in the coding order indicated by the arrows.

[0120] In some exemplary implementations, for the luma components of the inter-coded blocks, the above restrictions for intra-coding may not need to be applied. For example, after the first level of transform splitting, any one of the sub-transform blocks may be further split independently at one or more levels. Thus, the resulting transform blocks may or may not be of the same size. An exemplary splitting of an inter-coded block into transform locks in its coding order is shown in FIG. 16. In the example of FIG. 16, the inter-coded block 1602 is split into transform blocks at two levels according to Table 1. At the first level, the inter-coded block is split into four transform blocks of equal size. Then, only one of the four transform blocks (not all) is further split into four sub-transform blocks, resulting in a total of seven transform blocks having two different sizes as shown at 1604. An exemplary coding order of these seven transform blocks is indicated by the arrows at 1604 in FIG. 16.

[0121] In some exemplary implementations, for the chroma components, some further restrictions on the transform blocks may be applied. For example, for the chroma components, the transform block size can be the same as the coding block size, but cannot be made smaller than a predetermined size, for example, smaller than 8×8.

[0122] In some other exemplary implementations, for coding blocks having a width (W) or height (H) greater than 64, both the luma coding block and the chroma coding block may be implicitly split into transform units that are multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32), respectively. Here, in the present disclosure, "min(a,b)" may return the smaller value between a and b.

[0123] FIG. 17 further shows another alternative method for partitioning an encoding block or a prediction block into transform blocks. As shown in FIG. 17, instead of using a recursive transform partition, a predetermined set of partition types may be applied to the encoding block according to the transform type of the encoding block. In the specific example shown in FIG. 17, one of six exemplary partition types may be applied to divide the encoding block into various numbers of transform blocks. A method for generating such a partition of transform blocks may be applied to either the encoding block or the prediction block.

[0124] More specifically, the partitioning method of FIG. 17 provides up to six exemplary partition types for any given transform type (the transform type indicates a type of primary transform such as ADST, etc.). In this method, a transform partition type may be assigned to all encoding blocks or prediction blocks, for example, based on a rate-distortion cost. In one example, the transform partition type assigned to an encoding block or a prediction block may be determined based on the transform type of the encoding block or the prediction block. A particular transform partition type may correspond to the split size and pattern of the transform block, as indicated by the six transform partition types shown in FIG. 17. A correspondence between various transform types and various transform partition types may be predefined. Examples that may be assigned to an encoding block or a prediction block based on a rate-distortion cost are shown below by capital letter labels indicating the transform partition type. -PARTITION_NONE: Assigns a transform size equal to the block size. -PARTITION_SPLIT: Assigns a transform size of 1 / 2 the width of the block size and 1 / 2 the height of the block size. -PARTITION_HORZ: Assigns a transform size with the same width as the block size and 1 / 2 the height of the block size. -PARTITION_VERT: Allocate a conversion size that is half the width of the block size and the same height as the block size. -PARTITION_HORZ4: Allocate a conversion size that is the same width as the block size and one-fourth the height of the block size. -PARTITION_VERT4: Allocate a conversion size that is one-fourth the width of the block size and the same height as the block size.

[0125] In the above examples, all of the conversion partition types shown in FIG. 17 include a uniform conversion size for the partitioned conversion blocks. This is not a limitation but merely an example. In some other implementations, a mixed conversion block size may be used for the partitioned conversion blocks of a particular partition type (or pattern).

[0126] A PB (or CB, also called PB if not further divided into prediction blocks) obtained from any of the above partitioning methods may become an individual block for coding via either intra prediction or inter prediction. For current inter prediction of a PB, a residual between the current block and the prediction block is generated, coded, and may be included in the coded bitstream.

[0127] Inter prediction may be implemented, for example, in a single-reference mode or a multiple-reference mode. In some implementations, a skip flag may first be included in the bitstream of the current block (or at a higher level) to indicate whether the current block is inter-coded and not skipped. If the current block is inter-coded, another flag may be further included in the bitstream as a signal to indicate whether a single-reference mode or a multiple-reference mode is used for predicting the current block. In the single-reference mode, one reference block may be used to generate the predicted block of the current block. In the multiple-reference mode, two or more reference blocks may be used to generate the predicted block, for example, by weighted averaging. The multiple-reference mode may be referred to as the multi-reference mode, two-reference mode, or multiple-reference mode. The reference block or blocks of references may be identified using a reference frame index or indices, and further using corresponding motion vector or vectors indicating a shift in position between the reference block or blocks and the current block, for example, in horizontal and vertical pixels. For example, the inter-predicted block of the current block may be generated from a single reference block identified by one motion vector in a reference frame as the predicted block in the single-reference mode, but in the multiple-reference mode, the predicted block may be generated by weighted averaging of two reference blocks in two reference frames indicated by two reference frame indices and two corresponding motion vectors. The motion vectors may be coded in various ways and may be included in the bitstream.

[0128] In some implementations, an encoding or decoding system may maintain a decoded picture buffer (DPB). Some images / pictures may be maintained in the DPB while waiting to be displayed (by the decoding system), and some images / pictures in the DPB may be used as reference frames to enable inter prediction (by the decoding or encoding system). In some implementations, the reference frames in the DPB may be tagged as either short-term or long-term references for the currently encoded or decoded picture. For example, a short-term reference frame may include a frame used for inter prediction of blocks within the current frame, or a predetermined number (e.g., two) of subsequent video frames that are closest to the current frame in decoding order. A long-term reference frame may include a frame in the DPB that can be used to predict image blocks within a frame that is a predetermined number of frames further from the current frame in decoding order. Information regarding such tags for short-term and long-term reference frames may be referred to as a Reference Picture Set (RPS) and may be added to the header of each frame in the encoded bitstream. Each frame in the encoded video stream may be identified by a Picture Order Counter (POC), which may be numbered in an absolute manner according to the playback order, or numbered in relation to a picture group starting, for example, from an I-frame

[0129] In some exemplary implementations, one or more reference picture lists including identification of short-term and long-term reference frames for inter prediction may be formed based on the information in the RPS. For example, for uni-directional inter prediction, a single picture reference list shown as L0 reference (or reference list 0) may be formed, while for bi-directional inter prediction, two picture reference lists shown as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions may be formed. The reference frames included in the L0 and L1 lists may be ordered in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. Uni-directional inter prediction may be in a single-reference mode, or in a composite-reference mode if multiple references are on the same side of the block to be predicted for generation of a predicted block by weighted averaging in a composite prediction mode. Bi-directional inter prediction may be in a bi-reference mode only in that the bi-directional inter prediction includes at least two reference blocks.

[0130] Adaptive loop filter

[0131] In VVC (Versatile Video Coding), an Adaptive Loop Filter (ALF) by block-based filter adaptation is applied. In the luma component, based on the direction and activity of the local gradient, one is selected from a number of filters for each 4×4 block. In one example, there may be 25 filters to be selected.

[0132] FIG. 18 shows the shape of an exemplary adaptive loop filter (ALF). Specifically, FIG. 18 shows two diamond filter shapes. A 7×7 diamond shape is applied to the luma component, and a 5×5 diamond shape is applied to the chroma component.

[0133] For different examples, the block classification can be calculated as follows. For the luma component, each 4×4 block is classified into one of 25 classes. The classification index C is derived as follows based on its directionality D and the quantization value of the activity

Number

Number

Number

Number

[0134] To reduce the complexity of block classification, subsampled 1-D Laplacian calculations may be applied. As shown in FIGS. 19a to 19d, the same subsampling positions may be used for gradient calculations in all directions. FIG. 19a shows the subsampling positions in the Laplacian calculation of the vertical gradient. FIG. 19b shows the subsampling positions in the Laplacian calculation of the horizontal gradient. FIG. 19c shows the subsampling positions in the Laplacian calculation of the diagonal gradient. FIG. 19d shows the subsampling positions in the Laplacian calculation of the other diagonal gradient.

[0135] Next, the D maximum and minimum values of the gradients in the horizontal and vertical directions are set as follows.

Number

[0136] There may be geometric transformations of the filter coefficients and clipping values. Before filtering each 4×4 luma block, depending on the gradient value calculated for that block, geometric transformations such as rotation or diagonal and vertical flipping may be applied to the filter coefficient f(k, l) and the corresponding filter clipping value c(k, l). This may be equivalent to applying these transformations to the samples within the filter support region. This can make the different blocks to which ALF is applied more uniform by aligning the directions. The three geometric transformations may include diagonal, vertical flipping, and rotation.

Number

Table 2

[0137] In VVC, the ALF filter parameters are signaled with an adaption parameter set (APS). In one APS, multiple sets of luma filter coefficients and clipping value indices may be used. For example, there may be 25 sets of luma filters. Further, multiple sets of chroma filter coefficients and clipping value indices may be signaled. In one example, there may be up to 8 sets of chroma filter coefficients and clipping value indices that can be signaled. To reduce the bit overhead, filter coefficients of different classifications for the luma component can be merged. In the slice header, the index of the APS used for the current slice may be signaled. The signaling of ALF may be based on a Coding Tree Unit (CTU).

[0138] The clipping value index decoded from the APS enables the determination of the clipping value using tables of luma and chroma clipping values. These clipping values may depend on the internal bit depth. More precisely, the tables of clipping values may be obtained by the following formula. [Number] When B is equal to the internal bit depth, α is a predetermined constant value equal to 2.35, and N is equal to 4, which is the number of allowable clipping values in VVC in one embodiment. Table 3 shows the output of Equation (12). [Table 3]

[0139] In an example of a slice header, up to seven APS indices can be signaled to specify the luma filter set used for the current slice. The filtering process may be further controlled at the coding tree block (CTB) level. A flag may be signaled to indicate whether ALF is applied to the luma CTB. In one example, the luma CTB can select a filter set from 16 fixed filter sets and the filter sets from APS. A filter set index is signaled for the luma CTB to indicate which filter set is applied. The 16 fixed filter sets are predefined and may be hard-coded in both the encoder and decoder. For the chroma component, an APS index is signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if there are more than one chroma filter sets in APS, a filter index is signaled for each chroma CTB. The filter coefficients may be quantized with a norm equal to 128. Bitstream confirmance is applied so that the coefficient values at non-central positions can be in the range of -27 to 27 - 1 to limit the complexity of multiplication. The coefficient at the central position is not signaled in the bitstream and is considered to be equal to 128.

[0140] In the case of VVC, the syntax and meaning of the clipping index and values can be defined as follows. alf_luma_clip_idx[sfIdx][j] specifies the clipping index of the clipping value to be used before multiplying by the j-th coefficient of the luma filter signaled by sfIdx. It may be a requirement for bitstream compliance that the values of alf_luma_clip_idx[sfIdx][j] for sfIdx = 0... alf_luma_num_filters_signalled_minus1 and j = 0..11 are in the range of 0 to 3. The luma filter clipping value AlfClipL[adaptation_parameter_set_id] with elements AlfClipL[adaptation_parameter_set_id][filtIdx][j] for filtIdx = 0... NumAlfFilters - 1 and j = 0..11 is derived as specified in Table 3, depending on the bitDepth set equal to BitDepthY and the clipIdx set equal to alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx]][j]. alf_chroma_clip_idx[altIdx][j] specifies the clipping index of the clipping value to be used before multiplying by the j-th coefficient of the alternative chroma filter at index altIdx. It is a requirement for bitstream compliance that the values of alf_chroma_clip_idx[altIdx][j] for altIdx = 0.. alf_chroma_num_alt_filters_minus1 and j = 0..5 are in the range of 0 to 3.The chroma filter clipping value AlfClipC[adaptation_parameter_set_id][altIdx] having elements AlfClipC[adaptation_parameter_set_id][altIdx][j] for altIdx = 0..alf_chroma_num_alt_filters_minus1, j = 0..5 is derived as specified in Table 3 depending on the bitDepth set equal to BitDepthC and the clipIdx set equal to alf_chroma_clip_idx[altIdx][j].

[0141] The filtering process may be performed in the following example. On the decoder side, when ALF is enabled for a CTB, each sample R(i,j) within a CU is filtered, resulting in a sample value R'(i,j). [Number] Here, f(k,l) represents the decoded filter coefficient, K(x,y) is the clipping function, and c(k,l) represents the decoded clipping parameter. The variables k and l vary between -L / 2 and L / 2, where L represents the filter length. The clipping function is K(x,y) = min(y, max(-y,x)), which corresponds to the function Clip3(-y,y,x). By incorporating this clipping function first proposed in JVET-N0242, this loop filtering method becomes a non-linear process known as non-linear ALF. The selected clipping value is coded into the "alf_data" syntax element by using the Golomb coding scheme corresponding to the index of the clipping value in Table 3. This coding scheme may be the same as the coding scheme for the filter index.

[0142] There may be a virtual boundary filtering process for line buffer reduction. To reduce the line buffer requirements of the ALF, modified block classification and filtering may be applied to samples near the horizontal CTU boundary. Thus, as shown in FIG. 20, a virtual boundary may be defined as a line by shifting the horizontal CTU boundary by "N" samples. FIG. 20 shows an example of modified block classification at the virtual boundary. In this example, N is equal to 4 for the luma component and 2 for the chroma component.

[0143] The modified block classification is applied to the luma component as shown in FIG. 20. In the 1D Laplacian gradient calculation of the 4×4 block above the virtual boundary, only the samples above the virtual boundary are used. Similarly, in the 1D Laplacian gradient calculation of the 4×4 block below the virtual boundary, only the samples below the virtual boundary are used. The quantization of the activity value A is scaled considering the reduced number of samples used in the 1D Laplacian gradient calculation.

[0144] FIG. 21 shows an example of modified adaptive loop filtering for the luma component at the virtual boundary. In the filtering process, a symmetric padding operation at the virtual boundary may be used for both the luma and chroma components. As shown in FIG. 21, when the sample being filtered is located below the virtual boundary, the adjacent sample located above the virtual boundary is padded. The corresponding sample on the opposite side may also be padded symmetrically.

[0145] FIG. 22 shows an example of picture quadtree partitioning aligned with a largest coding unit (LCU). To improve coding efficiency, an adaptive loop filter based on coded unit synchronous picture quadtree may be used. The luma picture may be divided into several multilevel quadtree partitions, and each partition boundary is aligned with the boundary of the largest coding unit (LCU). Each partition has its own filtering process and may be called a filter unit (FU). The two-pass coding flow may include the following. In the first pass, the quadtree partitioning pattern and the optimal filter for each FU are determined. The filtering distortion is estimated by FFDE during the determination process. According to the determined quadtree partitioning patterns and the selected filters of all FUs, the reconstructed picture is filtered. In the second pass, CU synchronous ALF on / off control is performed. According to the ALF on / off result, the picture first filtered is partially restored by the reconstructed picture.

[0146] A top-down partitioning strategy may be adopted to divide the picture into multilevel quadtree partitions by using a rate distortion criterion. Each partition may be called a filter unit. The partitioning process aligns the quadtree partitions with the LCU boundaries. The coding order of the FUs follows the z-scan order. For example, in FIG. 22, the picture is divided into 10 FUs, and the coding order is FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, and FU9.

[0147] FIG. 23 shows an example of a quadtree split flag encoded in z-order. To show the quadtree split pattern of a picture, the split flag is encoded and transmitted in z-order. FIG. 23 shows the quadtree split pattern corresponding to FIG. 22. The filter of each FU is selected from two filter sets based on a rate-distortion criterion. The first set has newly derived 1 / 2 symmetric square and diamond filters for the current FU. The second set is obtained from a time-delay filter buffer. The time-delay filter buffer stores the filters previously derived for the FUs of the previous picture. The filter with the minimum rate-distortion cost of these two sets is selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further split into four child FUs, the rate-distortion costs of the four child FUs are calculated. By recursively comparing the rate-distortion costs in the case of splitting and the case of non-splitting, the picture quadtree split pattern can be determined. In one example, the maximum quadtree split level is 2, which means the maximum number of FUs is 16. During the determination of the quadtree split, the correlation values for deriving the Wiener coefficients of the 16 FUs at the lower quadtree level (the smallest FUs) can be reused. The remaining FUs can derive these Wiener filters from the correlation relationships of the 16 FUs at the lower quadtree level. Therefore, there may be only one frame buffer access to derive the filter coefficients of all FUs. After the quadtree split pattern is determined, CU synchronous ALF on / off control is performed to further reduce the filtering distortion. By comparing the filtering distortion and the non-filtering distortion, the leaf CU can explicitly switch the on / off of ALF in its local area. By redesigning the filter coefficients according to the ALF on / off result, the coding efficiency can be further improved. However, the redesign process may require additional frame buffer accesses. In the modified encoder design, there may be no redesign process after the CU synchronous ALF on / off decision to minimize the number of frame buffer accesses.

[0148] Cross-Component Adaptive Loop Filter (CC-ALF)

[0149] Figure 24 shows an example of the Cross-Component Adaptive Loop Filter (CC-ALF) arrangement. The CC-ALF may utilize luma sample values to refine each chroma component. Figure 24 shows the arrangement of the CC-ALF relative to other loop filters.

[0150] Figure 25 shows an example of a diamond filter. The CC-ALF may operate by applying a linear diamond filter from Figure 25 to the luma channel for each chroma component. The filter coefficients are transmitted by the APS, scaled by a factor of 210 in one example, and rounded for fixed-point representation. The application of the filter is controlled by a variable block size and signaled by a context-coded flag received for each block of samples. The block size is received at the slice level for each chroma component, together with a CC-ALF enable flag. In one example, block sizes of 16×16, 32×32, and 64×64 (within chroma samples) are supported.

[0151] The exemplary syntax of the CC-ALF may include the following. [Table 4] The meaning of the syntax related to the CC-ALF may include the following. alf_ctb_cross_component_cb_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] equal to 0 indicates that the cross-component Cb filter is not applied to the block of Cb color component samples at the luma position (xCtb, yCtb). alf_cross_component_cb_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] not equal to 0 indicates that the cross-component Cb filter of the alf_cross_component_cb_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] is applied to the block of Cb color component samples at the luma position (xCtb, yCtb). alf_ctb_cross_component_cr_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] equal to 0 indicates that the cross-component Cr filter is not applied to the block of Cr color component samples at the luma position (xCtb, yCtb). alf_cross_component_cr_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] not equal to 0 indicates that the cross-component Cr filter of the alf_cross_component_cr_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] is applied to the block of Cr color component samples at the luma position (xCtb, yCtb).

[0152] Chrominance sampling format

[0153] Figure 26 shows an exemplary position of chroma samples relative to luma samples. Figure 26 shows the indicated relative positions of the top-left chroma samples when chroma_format_idc is equal to 1 (4:2:0 chroma format) and chroma_sample_loc_type_top_field or chroma_sample_loc_type_bottom_field is equal to the value of variable ChromaLocType. The area represented by the top-left 4:2:0 chroma sample (shown as the large square with the large dot in the center) is shown relative to the area represented by the top-left luma sample (represented as the small square with the small dot in the center). The areas represented by adjacent luma samples are shown as small hatched gray squares with small hatched gray dots in the center.

[0154] Directional Enhancement Function

[0155] One purpose of the Constrained Directional Enhancement Filter (CDEF) in the loop is to filter out coding artifacts while preserving the details of the image. In HEVC, the Sample Adaptive Offset (SAO) algorithm can achieve a similar purpose by defining signal offsets for different classes of pixels. Unlike SAO, CDEF is a non-linear spatial filter. The filter design is constrained to be easily vectorized (i.e., implementable with SIMD operations), which may not be the case for other non-linear filters such as median filters and bilateral filters. The design of CDEF is derived from the following observations. The amount of ringing artifacts in the coded image tends to be approximately proportional to the quantization step size. The amount of detail is a characteristic of the input picture, but the minimum amount of detail retained in the quantized image also tends to be proportional to the quantization step size. For a given quantization step size, generally the amplitude of ringing is smaller than the amplitude of the detail.

[0156] CDEF identifies the direction of each block, then adaptively filters along the identified direction, and filters to a lesser extent along the direction rotated 45 degrees from the identified direction. The filter strength is explicitly signaled, which enables a high degree of control over blurring. An efficient encoder search is designed for the filter strength. CDEF is based on two previously proposed in-loop filters, and a composite filter is adopted in the new AV1 codec.

[0157] Figure 27 shows an example of direction search. The direction search operates on the reconstructed pixels immediately after the deblocking filter. Since these pixels are available to the decoder, the direction does not require signaling. The search operates on 8×8 blocks, which are large enough to appropriately process non-linear edges while being large enough to reliably estimate the direction when applied to a quantized image. Having a constant direction across the 8×8 region also makes it easier to vectorize the filter. For each block, the direction that best matches the pattern within the block is determined by minimizing the sum of squared differences (SSD) between the quantized block and the closest full-direction block. A full-direction block is a block in which all the pixels along a line in one direction have the same value. Figure 27 is an example of the direction search for an 8×8 block.

[0158] There may be a non-linear low-pass direction filter. One reason to identify the direction is to align the filter taps along that direction and reduce ringing while preserving the edges or patterns in the direction. However, in some cases, direction filtering alone may not be sufficient to reduce ringing significantly. Also, it may be desirable to use filter taps for pixels not along the main direction. To reduce the risk of blurring, these extra taps are treated more conservatively. For this reason, CDEF defines primary taps and secondary taps. The complete 2D CDEF filter may be represented as follows.

Number

[0159] Loop Restoration

[0160] For use after the deblocking of video coding, beyond the conventional deblocking operation, a set of in-loop restoration methods has been proposed to generally remove noise and improve the quality of edges. These methods are switchable within the frame for tiles of appropriate size. The specific methods described are based on a separable symmetric Wiener filter and a dual self-induced filter with subspace projection. Since the content statistics can vary substantially within the frame, these tools are integrated within a switchable framework where different tools can be triggered in different regions of the frame.

[0161] There may exist a separable symmetric Wiener filter used as a restoration tool. Each pixel in a degraded frame may be reconstructed as a non-casual filtered version of the pixels within its surrounding w×w window, where w = 2r+1 is odd for an integer r. When the 2D filter taps are represented by a vector F of w2×1 elements in a column-vectorized form, simple LMMSE optimization yields filter parameters given by F = H-1M. Here, H = E[XXT] is the autocovariance of x, which is the column-vectorized version of the w2 samples within the w×w window surrounding the pixel, and M = E[YXT] is the cross-correlation between x and the scalar source sample y to be estimated. The encoder can estimate H and M from the realizations within the source and the deblocked frame and send the resulting filter F to the decoder. However, this not only incurs a significant bitrate cost in transmitting the w2 taps but also makes the filtering non-separable, which complicates the decoding significantly. Therefore, some further constraints are imposed on the nature of F. First, F is constrained to be separable so that the filtering can be implemented as separable horizontal and vertical w-tap convolutions. Second, each of the horizontal and vertical filters is constrained to be symmetric. Third, it is assumed that the sum of both the horizontal filter coefficients and the vertical filter coefficients is 1.

[0162] There may exist double self-induced filtering by subspace projection for image filtering, and the local linear model is [Number] and it is used to calculate the filtered output y from the unfiltered sample x. Here, F and G are determined based on the statistics of the degraded image and the guidance image around the filtered pixels. When the guide image is the same as the degraded image, the so-called self-guided filtering of the result has the effect of edge-preserving smoothing. The specific form of the proposed self-guided filtering depends on two parameters, the radius r and the noise parameter e, and is listed as follows. 1. Obtain the mean μ and variance σ of the pixels within the (2r + 1)×(2r + 1) window around each pixel. 2 This can be efficiently implemented by box filtering based on integral imaging. 2. For each pixel, calculate f = σ 2 / (σ 2 + e), and g = (1 - f)μ 3. Calculate F and G for each pixel as the average of the values of f and g within the 3×3 window around the pixels used. The filtering may be controlled by r and e, which means that the higher r is, the higher the spatial dispersion is, and the higher e is, the higher the range dispersion is.

[0163] Figure 28 shows an example of subspace projection. The principle of subspace projection is illustrated in Figure 28. Even if neither of the inexpensive restorations X1, X2 is close to the source Y, appropriate multipliers {α, β} can bring them quite close to the source as long as they are moving in a somewhat correct direction.

[0164] Cross-Component Sample Offset (CCSO)

[0165] The loop filtering method may include a cross-component sample offset (CCSO) for reducing distortion of reconstructed samples. In CCSO, when a processed input reconstructed sample of a first color component is given, a non-linear mapping is used to derive an output offset, and the output offset is added to the reconstructed samples of other color components in the proposed CCSO filtering process.

[0166] FIG. 29 shows an example of a filter support region. The input reconstructed sample is from a first color component located in the filter support region. As shown in FIG. 29, the filter support region includes four reconstructed samples p0, p1, p2, p3. The four input reconstructed samples follow a cross pattern in the vertical and horizontal directions. The central sample (denoted as c) in the first color component and the sample to be filtered in the second color component are at the same position. When processing the input reconstructed sample, the following steps are applied. · Step 1: Delta values between p0 to p3 and c are first calculated and denoted as m0, m1, m2, and m3. · Step 2: The delta values m0 to m3 are further quantized, and the quantization values are denoted as d0, d1, d2, d3. The quantization values can be -1, 0, 1 based on the following quantization process.

Equation

[0167] The variables d0 to d3 may be used to identify one combination of non-linear mappings. In this example, CCSO has four filter taps d0 to d3, and each filter tap may have one of three quantization values, so there are a total of 3^4 = 81 combinations. Table 4 (below) shows 81 exemplary combinations, and the last column represents the output offset value for each combination. Exemplary offset values are integers such as 0, 1, -1, 3, -3, 5, -5, -7.

Table 5

[0168] The final filtering process of CCSO is applied as follows.

Number

[0169] The local sample offset (LSO) is another exemplary offset embodiment. In LSO, a similar filtering method in CCSO is applied, but the output offset is applied to the color component that is the same color component as the reconstructed sample used as the input to the filtering process.

[0170] In an alternative embodiment, a simplified CCSO design may be employed in the reference software of AV2, namely, the AVM for CWG - B022.

[0171] Figure 30 shows an exemplary loop filter pipeline. CCSO is a loop filter process that runs in parallel with CDEF in the loop filter pipeline, that is, as shown in Figure 30, the input is the same as CDEF, and the output is applied to the samples filtered by CDEF. Note that CCSO may be applied only to the chroma color components.

[0172] Figure 31 shows an exemplary input of cross-component sample offset (CCSO). The CCSO filter is applied to the chroma reconstruction samples shown as rc. The luma reconstruction samples at the same position of rc are shown as rl. An example of the CCSO filter is shown in Figure 31. In CCSO, a set of 3-tap filters is used. The input luma reconstruction samples located at the three filter taps include the central rl and two adjacent samples p 0 and p 1 .

[0173] p i and rl(i = 0,1) are given, the following steps are applied to process the input samples. - The delta value between p i and rl is first calculated and shown as m i . - The delta value m i is quantized to d i using the following quantization process. - If m is less than -Q CCSO , d i is set to -1 - If m is greater than or equal to -Q CCSO and less than or equal to Q CCSO , d i is set to 0 - If m is greater than Q CCSO , d i is set to 1 In the above steps, Q CCSO is called the quantization step size, and Q CCSO can be 8, 16, 32, 64.

[0174] d 0 and d 1 are calculated, the offset value (shown as s) is derived using the look-up table (LUT) of CCSO. The LUT of CCSO is shown in Table 5. d 0 and d 1Each combination is used to identify a row in the LUT and obtain an offset value. The offset values are integers including 0, 1, -1, 3, -3, 5, -5, and -7. [Table 6]

[0175] Finally, the derived offset of the CCSO is applied to the chroma color component as follows. [Equation] Here, rc is the reconstructed sample filtered by the CCSO, s is the derived offset value obtained from the LUT, and the filtered sample value rc' is further clipped to the range specified by the bit depth.

[0176] Figure 32 shows an exemplary filter shape in the cross-component sample offset (CCSO). In CCSO, as shown in Figure 32, there are six optional filter shapes denoted as f i (i = 1...6). These six filter shapes are switchable at the frame level, and the selection is signaled by the syntax ext_filter_support using a 3-bit fixed-length code.

[0177] The signaling of the cross-component sample offset (CCSO) may be performed at both the frame level and the block level. At the frame level, the signal may include the following. · A 1-bit flag indicating whether the CCSO is applied · A 3-bit syntax ext_filter_support indicating the selection of the CCSO filter shape · A 2-bit index indicating the selection of the quantization step size · Nine 3-bit offset values used in the LUT At the 128×128 chroma block level, a flag is signaled to indicate whether the CCSO filter is effective or not.

[0178] Sample Adaptive Offset (SAO)

[0179] In HEVC, by using the offset value given in the slice header, a sample adaptive offset (SAO) is applied to the reconstructed signal after the deblocking filter. For luma samples, the encoder determines whether SAO is applied to the current slice. If SAO is effective, the current picture allows a recursive division into four sub-regions, and each region can select one of six SAO types as shown in Table 6. SAO classifies the reconstructed pixels into categories and reduces distortion by adding an offset to the pixels of each category in the current region. Edge properties are used for pixel classification in SAO types 1 to 4, and pixel intensity is used for pixel classification in SAO types 5 to 6.

Table 7

[0180] Band offset (BO) classifies all the pixels of a region into multiple bands, and each band contains pixels with the same intensity interval. The intensity range is evenly divided into 32 intervals from 0 to the maximum intensity value (for example, 255 for 8-bit pixels), and each interval has an offset. Next, the 32 bands are divided into two groups. One group consists of the 16 central bands, and the other group consists of the remaining 16 bands. Only the offsets of one group are transmitted. Regarding the pixel classification operation in BO, the most significant 5 bits of each pixel can be directly used as the band index.

[0181] Figure 33 shows an exemplary pixel pattern. Edge offset (EO) uses four 1-D 3-pixel patterns for pixel classification considering edge direction information as shown in Figure 33. Each region of the picture selects one pattern, and by comparing each pixel with its two adjacent pixels, the pixels can be classified into multiple categories. The selection is transmitted in the bitstream as side information. Table 7 shows the pixel classification rules for EO. [Table 8]

[0182] SAO on the decoder side may operate independently of the LCU so that the line buffer can be saved. To achieve this, when classification patterns of 90°, 135°, and 45° are selected, the pixels in the upper and lower rows within each LCU may not be SAO processed. When patterns of 0°, 135°, and 45° are selected, the pixels in the leftmost and rightmost columns within each LCU may not be SAO processed.

[0183] The following Table 8 shows an exemplary syntax that needs to be signaled for the CTU when the parameters are not merged from adjacent CTUs. [Table 9]

[0184] Cross-Component Sample Offset (CCSO) and Local Sample Offset (LSO) can utilize the values of the pixels that are filtered to select the offset values in one color component. However, further expanding these inputs for offset selection may significantly increase the signaling overhead of CCSO and LSO, which may limit / reduce the coding performance, especially for sequences with a lower resolution.

[0185] As described above, CCSO is defined as a filtering process that uses the reconstructed sample of the first color component (e.g., Y, Cb, or Cr) as an input, and the output is applied to the second color component that is a different color component from the first color component. An exemplary filter shape of CCSO is shown in FIG. 29. LSO is a filtering process that uses the reconstructed sample of the first color component (e.g., Y, Cb, or Cr) as an input, and the output is applied to the same first color component. Therefore, the difference between LSO and CCSO is the different inputs.

[0186] As described below, as shown in FIG. 34, a generalized design for CCSO and LSO is shown by considering not only the delta values between adjacent samples of the sample at the same position (or current) but also the level value of the sample itself at the same position (or current), as considered in CCSO and LSO.

[0187] FIG. 34 shows a flowchart of a method according to an exemplary embodiment of the present disclosure. At block 3402, coding information for reconstructed samples of a current component within a current picture is decoded from a coded video bitstream. The coding information indicates a sample offset filter applied to the reconstructed samples. In one example, the sample offset filter may include two types of offset values, a gradient offset (GO) and a band offset (BO). The color range of the gradient may include two or more colors, and the GO is an offset attribute where the color of the gradient starts and ends. The BO may be an offset derived using samples at the same position of different color components or the value of the current sample being filtered, as further described below, and the band is used to determine the offset value. At block 3404, the offset type used in the sample offset filter is selected. At block 3406, based on the first reconstructed sample and the selected offset type, the output value of the sample offset filter is determined. At block 3408, based on the reconstructed sample and the output value of the sample offset filter, the filtered sample value of the reconstructed sample of the current component is determined. Further embodiments are described below.

[0188] The generalized sample offset (GSO) method may include two types of offset values for CCSO and LSO, including a gradient offset (GO) and a band offset (BO). The selection of the offset type may be signaled or implicitly derived.

[0189] In one embodiment, the gradient offset may be an offset derived using a delta value between adjacent samples and samples at the same position of different color components (in the case of CCSO), or a delta value between adjacent samples and the current sample being filtered (in the case of CCSO or LSO).

[0190] In one embodiment, the band offset may be an offset derived using values of the same sample of different color components or the current sample to be filtered. The band may be used to determine the offset value. In one example, the values of the samples at the same position of different color components or the current sample to be filtered may be denoted as variable v, and the BO value is derived using v>>s, where >> represents a right shift operation and s is a predetermined value specifying the interval of the sample values using the same band offset. In one example, the value of s can be different for different color components. In another example, the values of the samples at the same position of different color components or the current sample to be filtered are denoted as variable v, the band index bi is derived using a predetermined look-up table, the input of the look-up table is v, the output value is the band index bi, and the BO value is derived using the band index bi.

[0191] In one embodiment, when a combination of GO and BO is applied (for example, when used simultaneously), the offset is derived using both 1) the delta value between adjacent samples and the samples at the same position of different color components (in the case of CCSO) or the delta value between adjacent samples and the current sample to be filtered (in the case of CCSO or LSO), and 2) the values of the samples at the same position of different color components or the current sample to be filtered.

[0192] In one embodiment, the application of GO or BO is signaled. This signaling may be applied with a high-level syntax. As some examples, the signaling may include VPS, PPS, SPS, slice header, picture header, frame header, superblock header, CTU header, tile header.

[0193] In other embodiments, regardless of whether GO or BO is signaled at the block level, the block level includes, but is not limited to, the coding unit (block) level, the prediction block level, the transform block level, or the filtering unit level. This example includes signaling at the block level for identifying GO or BO.

[0194] In other embodiments, GO or BO is signaled using flags. First, a flag indicating whether LSO and / or CCSO is applied to one or more color components is signaled, and then another flag indicating whether GO or BO is applied is signaled. For example, first, a flag indicating whether LSO and / or CCSO is applied to one or more color components is signaled, and then another flag indicating whether GO is applied together with BO is signaled, where BO is always applied regardless of whether GO is applied. In another example, first, a flag indicating whether LSO and / or CCSO is applied to one or more color components is signaled, and then another flag indicating whether BO is applied together with GO is signaled, where GO is always applied regardless of whether BO is applied.

[0195] In some embodiments, a signal for determining whether to use GO, use BO, or use a combination thereof may be derived. It may be implicitly derived using coding information including, but not limited to, the current color component and / or reconstructed samples of different color components, whether the current block is intra-coded or inter-coded, whether the current picture is a key (or intra) picture, and whether the current sample (or block) is coded by a specific prediction mode (specific intra or inter prediction mode, transform selection mode, quantization parameter, etc.).

[0196] Embodiments of the present disclosure may be used individually or combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The term block may include a prediction block, an encoded block, or an encoding unit, i.e., a CU. Embodiments of the present disclosure may be applied to luma blocks or chroma blocks.

[0197] The above-described techniques can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 35 shows a computer system (3500) suitable for implementing a particular embodiment of the disclosed subject matter.

[0198] The computer software can be coded using any suitable machine code or computer language and be the subject of assembly, compilation, linking, or similar mechanisms to create code including instructions executable directly or through interpretation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0199] The instructions can be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0200] The components shown in FIG. 35 for the computer system (3500) are of an exemplary nature and are not intended to imply any limitations on the use or functionality of the computer software implementing the embodiments of the present disclosure. Nor should the component configuration be construed as having any dependencies or requirements regarding any one or combination of the components shown in the exemplary embodiments of the computer system (3500).

[0201] The computer system (3500) can include specific human interface input devices. Such human interface input devices can respond to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, movement of a data glove), voice input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). Also, the human interface device can be used to capture certain media that are not necessarily directly related to conscious human input, such as voice (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., 2D video, 3D video including stereoscopic video).

[0202] The input human interface device may include one or more of a keyboard (3501), a mouse (3502), a trackpad (3503), a touch screen (3510), a data glove (not shown), a joystick (3505), a microphone (3506), a scanner (3507), and a camera (3508) (only one of each is shown).

[0203] The computer system (3500) may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback by a touch screen (3510), a data glove (not shown), or a joystick (3505); however, there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (3509), headphones (not shown)), visual output devices (e.g., a screen (3510) including a CRT screen, an LCD screen, a plasma screen, an OLED screen; each may or may not have a touch screen input function, each may or may not have a tactile feedback function, and some of them may be capable of outputting higher than three-dimensional output through means such as two-dimensional visual output or stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown).

[0204] The computer system (3500) may also include an optical medium including a CD / DVD ROM / RW (3520) together with a human-accessible memory device and associated media, such as a CD / DVD or similar media (3521), a thumb drive (3522), a removable hard drive or solid state drive (3523), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0205] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0206] The computer system (3500) can also include an interface (3554) to one or more communication networks (3555). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, in-vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet®, wireless LAN, GSM, 3G, 4G, 5G, LTE, etc. cellular networks, cable TV, satellite TV, TV wired or wireless wide area digital networks including terrestrial broadcast TV, in-vehicle and industrial including CAN Bus, etc. Certain networks typically require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (3549) (e.g., a USB port of the computer system (3500), etc.). Others are typically integrated into the core of the computer system (3500) by attachment to a system bus as described later (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (3500) can communicate with other entities. Such communication can be unidirectional, receive only (e.g., broadcast TV), unidirectional transmit only (e.g., CANbus to a specific CANbus device), or bidirectional to other computer systems using, for example, a local or wide area digital network. For each of the networks and network interfaces as described above, a specific protocol and protocol stack can be used.

[0207] The aforementioned human interface device, human-accessible memory device, and network interface can be attached to the core (3540) of the computer system (3500).

[0208] The core (3540) can include one or more central processing units (CPUs) (3541), graphics processing units (GPUs) (3542), specialized programmable processing devices in the form of field programmable gate arrays (FPGAs) (3543), hardware accelerators (3544) for specific tasks, graphics adapters (3550), etc. These devices can be connected through a system bus (3548) together with a read-only memory (ROM) (3545), random access memory (3546), internal mass storage devices (3547) such as internal hard drives and solid state drives (SSDs) that are not accessible to users internally. In some computer systems, the system bus (3548) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core's system bus (3548) or through a peripheral bus (3549). In one example, a screen (3510) can be connected to the graphics adapter (3550). Architectures for peripheral buses include PCI, USB, etc.

[0209] The CPU (3541), GPU (3542), FPGA (3543), and accelerator (3544) can execute specific instructions that can, in combination, constitute the above-mentioned computer code. That computer code can be stored in the ROM (3545) or RAM (3546). Temporary data can also be stored in the RAM (3546), while persistent data can be stored, for example, in the internal mass storage device (3547). By using cache memory that can be closely associated with one or more CPUs (3541), GPUs (3542), mass storage devices (3547), ROM (3545), RAM (3546), etc., fast storage and retrieval to any of the memory devices can be enabled.

[0210] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well known and available to those having skill in the computer software arts.

[0211] As a non-limiting example, a computer system having an architecture (3500), specifically a core (3540), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be media related to user-accessible mass storage as introduced above and specific storage of the core (3540) of a non-transitory nature such as a mass storage device (3547) inside the core or ROM (3545). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (3540). The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause the core (13540), particularly the processor (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in the RAM (3546) and modify such data structures according to processes defined by the software, including causing the core to execute specific processes or specific parts of specific processes described herein. Further or alternatively, the computer system can provide functionality as a result of logic being hardwired or incorporated into a circuit (e.g., an accelerator (3544)), which can operate instead of or together with software to execute specific processes or specific parts of specific processes described herein. References to software can, where appropriate, include logic, and vice versa. References to computer-readable media can, where appropriate, include circuits (such as integrated circuits (ICs), etc.) storing software for execution, circuits embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0212] Although the present disclosure describes several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be recognized by those skilled in the art that numerous systems and methods can be devised that embody the principles of the present disclosure and are thus within its spirit and scope, although not explicitly illustrated or described herein.

[0213] Appendix A: Abbreviations ALF: Adaptive Loop Filter AMVP: Advanced Motion Vector Prediction APS: Adaptation Parameter Set ASIC: Application-Specific Integrated Circuit AV1: AOMedia Video 1 AV2: AOMedia Video 2 BCW: Bi-prediction with CU-level Weights BM: Bilateral Matching BMS: benchmark set CANBus: Controller Area Network Bus CC-ALF: Cross-Component Adaptive Loop Filter CCSO: Cross-Component Sample Offset CD: Compact Disc CDEF: Constrained Directional Enhancement Filter CDF: Cumulative Density Function CfL: Chroma from Luma CIIP: Combined intra-inter prediction CPU: Central Processing Unit CRT: Cathode Ray Tube CTB: Coding Tree Block CTU: Coding Tree Unit CU: Coding Unit DMVR: Decoder-side Motion Vector Refinement DPB: Decoded Picture Buffer DPS: Decoding Parameter Set DVD: Digital Video Disc FPGA: Field Programmable Gate Areas GBI: Generalized Bi-prediction GOP: Groups of Picture GPU: Graphics Processing Unit GSM: Global System for Mobile communications HDR: high dynamic range HEVC: High Efficiency Video Coding HRD: Hypothetical Reference Decoder IBC (or IntraBC): Intra Block Copy IC: Integrated Circuit ISP: Intra Sub-Partitions JEM: joint exploration model JVET: Joint Video Exploration Team LAN: Local Area Network LCD: Liquid-Crystal Display LCU: Largest Coding Unit LR: Loop Restoration Filter LSO: Local Sample Offset LTE: Long-Term Evolution MMVD: Merge Mode with Motion Vector Difference MPM: most probable mode MV: Motion Vector MVD: Motion Vector difference MVP: Motion Vector Predictor OLED: Organic Light-Emitting Diode PB: Prediction Block PCI: Peripheral Component Interconnect PDPC: Position Dependent Prediction Combination PLD: Programmable Logic Device POC: Picture Order Count PPS: Picture Parameter Set PU: Prediction Unit RAM: Random Access Memory ROM: Read-Only Memory RPS: Reference Picture Set SAD: Sum of Absolute Difference SAO: Sample Adaptive Offset SB: Super Block SCC: Screen Content Coding SDP: Semi Decoupled Partitioning SDR: standard dynamic range SDT: Semi Decoupled Tree SEI: Supplementary Enhancement Information SNR: Signal Noise Ratio SPS: Sequence Parameter Setting SSD: solid-state drive SST: Semi Separate Tree TM: Template Matching TU: Transform Unit USB: Universal Serial Bus VPS: Video Parameter Set VUI: Video Usability Information VVC: versatile video coding WAIP: Wide-Angle Intra Prediction

Claims

Claim 1 A method for video decoding, comprising: decoding coding information for reconstructed samples within a current picture from a coded video bitstream, the coding information including a sample offset filter to be applied to the reconstructed samples; receiving a signal indicating an offset type used in the sample offset filter, the offset type including a gradient offset (GO) or a band offset (BO), the signal including a first flag indicating whether a local sample offset (LSO) and / or a cross-component sample offset (CCSO) is applied to one or more color components, and a second flag indicating whether the GO or the BO is applied; determining an output value of the sample offset filter based on the reconstructed samples and the received offset type; and a method including the above steps. Claim 2 The method according to claim 1, further comprising determining a filtered sample value based on the reconstructed samples and the output value of the sample offset filter. Claim 3 The method according to claim 1, wherein the reconstructed samples are from a current component within the current picture. Claim 4 The method according to claim 1, wherein the signal includes high-level syntax transmitted in a slice header, a picture header, a frame header, a superblock header, a coding tree unit (CTU) header, or a tile header. Claim 5 The method according to claim 1, wherein the signal includes block-level transmission at a coding unit level, a prediction block level, a transform block level, or a filtering unit level. Claim 6. The method according to claim 1, further comprising, when the offset type is the GO, deriving an offset value of the GO using a delta value between samples at the same position of different color components from adjacent samples. Claim 7. The method according to claim 1, further comprising, when the offset type is the GO, deriving an offset value of the GO using a delta value between samples at the same position of adjacent samples and the current sample to be filtered.

8. The method according to claim 1, further comprising, when the offset type is the BO, deriving an offset value of the BO using values of samples at the same position of different color components.

9. The method according to claim 1, further comprising, when the offset type is the BO, deriving an offset value of the BO using values of samples at the same position of the currently sampled sample to be filtered.

10. An apparatus for decoding a video bitstream, comprising: a memory for storing instructions; a processor communicating with the memory wherein when the processor executes the instructions, the processor is configured to cause the apparatus to execute the method according to any one of claims 1 to 9.

11. A computer program for causing a processor to execute the method according to any one of claims 1 to 9.

12. A method for video encoding, comprising: selecting an offset type used in a sample offset filter, the offset type including a gradient offset (GO) or a band offset (BO); encoding coding information for reconstructed samples within a current picture into a video bitstream, the coding information including the sample offset filter applied to the reconstructed samples; transmitting a signal indicating the offset type used in the sample offset filter, the signal including a first flag indicating whether a local sample offset (LSO) and / or a cross-component sample offset (CCSO) is applied to one or more color components, and a second flag indicating whether the GO or the BO is applied; and a method comprising the steps of.

Citation Information

Patent Citations

  • Image filter device, filtering method, and dynamic image decoding device

    JP2016201824A

  • Encoder-side decisions for sample adaptive offset filtering

    WO2015165030A1

  • Coding enhancement in cross-component sample adaptive offset

    WO2022251517A1

  • Edge offset for cross component sample adaptive offset (CCSAO) filter

    WO2023056159A1