Video decoding method, device, electronic device and computer readable storage medium

By constraining the directional enhancement filter and using intra-frame prediction mode, the recovery filtering process is optimized using directional information, which solves the problems of high computational burden and bit rate overhead in the existing technology and improves video decoding efficiency and quality.

CN115004696BActive Publication Date: 2026-03-31TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-01
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing video decoding technologies, intra-frame prediction and loop recovery filters have significant computational burdens and bit rate overheads, affecting decoding efficiency and quality.

Method used

By employing constrained directional enhancement filters (CDEF) and intra-frame prediction modes, the recovery filtering process is optimized by reusing directional information and filter parameter sets, thereby reducing computational burden and bit rate overhead.

Benefits of technology

It improves the efficiency and quality of video decoding, reduces computational complexity and bit rate overhead, and enhances the performance of intra-frame prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115004696B_ABST
    Figure CN115004696B_ABST
Patent Text Reader

Abstract

A video encoding method, apparatus, electronic device, and computer-readable storage medium are disclosed. The method includes determining directional information of a restoration filter unit included in a video frame based on at least one of a constrained directional enhancement filter (CDEF) process or an intra prediction mode, determining one of a plurality of filter parameter sets of a restoration filter process based on the directional information of the restoration filter unit, performing the restoration filter process on the restoration filter unit based on the one of the plurality of filter parameter sets, and reconstructing the video frame based on the restoration filter unit processed by the restoration filter process.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application claims priority to U.S. Patent Application No. 17 / 362,048, filed June 29, 2021, entitled “METHOD AND APPARATUS FOR VIDEO CODING,” which claims priority to U.S. Provisional Application No. 63 / 091,707, filed October 14, 2020, entitled “FEATURE INFORMATION REUSE FOR ENHANCEDRESTORATION FILTERING.” The entire disclosure of the earlier applications is incorporated herein by reference. Technical Field

[0003] This application relates to video processing technology, and more particularly to a video decoding method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0004] The purpose of the background description provided herein is to present the general context of this disclosure. To the extent that the work of the currently identified inventors is described in this background section, neither the work of the currently identified inventors nor aspects of the description that might not otherwise be considered prior art at the time of submission are expressly or implicitly acknowledged as prior art to this disclosure.

[0005] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can comprise a series of pictures, each with a spatial size of, for example, 1920×1080 luminance samples and associated chrominance samples. This series of pictures can have a fixed or variable picture rate (also informally referred to as the frame rate), such as 60 pictures per second or 60Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a 60Hz frame rate) requires close to 1.5 Gbit / s of bandwidth. One hour of such video would require over 600 gigabytes (GByte) of storage space.

[0006] One objective of video encoding and decoding is to reduce the redundancy of the input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by two or more orders of magnitude. Lossless compression and lossy compression, or combinations thereof, can be employed. Lossless compression refers to the technique of reconstructing an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may differ from the original signal, but the distortion between the original and reconstructed signals is small enough that the reconstructed signal can be used for the intended application. In the case of video, lossy compression is widely used. The amount of distortion tolerated depends on the application; for example, users of some consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio reflects the allowable / tolerable distortion: a higher allowable / tolerable distortion can result in a higher compression ratio.

[0007] Video encoders and decoders can utilize techniques from several broad categories, including motion compensation, transform, quantization, and entropy coding.

[0008] Video codec techniques can include techniques known as intra-frame coding. In intra-frame coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, the picture can be an intra-frame picture. Intra-frame pictures and their derivatives (e.g., independent decoder refresh pictures) can be used to reset the decoder state and therefore can be used as the first picture in an encoded video bitstream and video session or as a still image. Samples of an intra-frame block can be transformed, and the transform coefficients can be quantized before entropy coding. Intra-frame prediction can be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are needed to represent the entropy-coded block at a given quantization step size.

[0009] Traditional intra-frame coding—such as intra-frame coding techniques known from MPEG-2 generation coding techniques—does not use intra-frame prediction. However, some newer video compression techniques include attempts to use metadata obtained during the encoding and / or decoding of, for example, spatially adjacent data blocks that are first in the decoding order, along with / or surrounding sample data. Such techniques are hereby referred to as "intra-frame prediction" techniques. Note that, in at least some cases, intra-frame prediction uses only reference data from the current frame in the reconstruction, and not reference data from a reference frame.

[0010] Many different forms of intra-prediction can exist. When more than one such technique can be used in a given video coding technique, the technique in use can be encoded in an intra-prediction mode. In some cases, a mode can have sub-modes and / or parameters, and these sub-modes and / or parameters can be encoded separately or included in the mode codeword. Which codeword is used for a given combination of modes, sub-modes, and / or parameters can affect the coding efficiency gain through intra-prediction, and therefore the entropy coding technique used to convert the codeword into a bitstream can also affect it.

[0011] Certain patterns of intra-frame prediction were introduced with H.264, refined in H.265, and further refined in newer coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Sets (BMS). Prediction blocks can be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of neighboring samples are copied into the prediction block according to the orientation. References to the orientation in use can be encoded in the bitstream, or they can be predicted themselves.

[0012] Currently, video decoding applies three in-loop filters to the reconstructed frame in the order of deblocking, constrained directional enhancement filter (CDEF), and loop restoration filter. When applying the loop restoration filter, the target region needs to be classified and appropriate filter parameters applied to ensure the quality of the final reconstruction. However, this approach in existing technologies increases computational burden or bit rate overhead. Summary of the Invention

[0013] This disclosure provides an apparatus for video encoding / decoding. The apparatus includes a processing circuitry system that determines directional information of a recovery filter unit included in a video frame based on at least one of constrained directional enhancement filter (CDEF) processing or intra-frame prediction mode. The processing circuitry system determines one of a plurality of filter parameter sets for recovery filter processing based on the directional information of the recovery filter unit. The processing circuitry system performs recovery filter processing on the recovery filter unit based on one of the plurality of filter parameter sets. The processing circuitry system reconstructs the video frame based on the filtered recovery filter unit.

[0014] In an implementation, the recovery filter unit includes one or more directional information units, and performs at least one of CDEF processing or intra-frame prediction mode on one of the directional information units.

[0015] In an implementation, each of the multiple filter parameter sets of the recovery filter is associated with at least one directionality of the CDEF process.

[0016] In this implementation, the processing circuitry determines one of a set of multiple filter parameters for the recovery filter processing based on the block variance and directionality information of the recovery filter unit.

[0017] In this implementation, the processing circuitry determines one of a set of multiple filter parameters for the recovery filter processing based on the directional information of the recovery filter unit and the filter strength of the CDEF processing.

[0018] In one implementation, the processing circuitry determines the directionality information of the recovery filter unit based on at least one of a majority vote or a consensus check on the directionality of the recovery filter unit.

[0019] In one implementation, based on the fact that the recovery filter unit is not intra-coded and the adjacent blocks of the recovery filter unit are intra-coded, the processing circuitry determines the directionality information of the recovery filter unit based on the intra-prediction mode performed on the adjacent blocks.

[0020] In the implementation, the processing circuit system performs recovery filter processing on the recovery filter unit based on matching the directional information determined according to CDEF processing with the directional information determined according to the intra-frame prediction mode.

[0021] In the implementation, the recovery filter processing is one of the Wiener filter processing and the self-guided projection (SGRPRJ) filter processing.

[0022] In an implementation, the processing circuitry determines one of a plurality of filter parameter sets for the recovery filter processing based on a default set of filter parameters, an index of an indicator set of filter parameters indicated by a signal, or a set of filter parameters indicated by a signal.

[0023] This disclosure provides a method for video encoding / decoding. In this method, directional information of a recovery filter unit included in a video frame is determined based on at least one of CDEF processing or intra-frame prediction mode. One of a plurality of filter parameter sets for recovery filter processing is determined based on the directional information of the recovery filter unit. Recovery filter processing is performed on the recovery filter unit based on one of the plurality of filter parameter sets. The video frame is reconstructed based on the filtered recovery filter unit.

[0024] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause at least one processor to perform any one or a combination of methods for video decoding.

[0025] Therefore, the video decoding method provided in this application can enhance the performance of the restoration filtering process by reusing feature information (e.g., directional information) derived from CDEF processing and / or intra-frame prediction modes. Adaptive restoration filtering technology can effectively reuse signal features and statistical information already available at the decoder, thereby avoiding increased computational burden and bit rate overhead while ensuring the final reconstruction quality. Attached Figure Description

[0026] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, which are shown in the drawings:

[0027] Figure 1A is a schematic diagram of an exemplary subset of intra-frame prediction modes;

[0028] Figure 1B is an illustration of an exemplary intra-frame prediction direction;

[0029] Figure 1C is a schematic diagram of the current block and its surrounding space merge candidates in an example;

[0030] Figure 2 This is a simplified block diagram of a communication system according to an embodiment of this application;

[0031] Figure 3 This is a simplified block diagram of a communication system according to an embodiment of this application;

[0032] Figure 4 This is a simplified block diagram of a decoder according to an embodiment of this application;

[0033] Figure 5 This is a simplified block diagram of an encoder according to an embodiment of this application;

[0034] Figure 6 A block diagram of an encoder according to another embodiment of this application is shown;

[0035] Figure 7 A block diagram of a decoder according to another embodiment of this application is shown;

[0036] Figure 8 An exemplary nominal angle according to an embodiment of this application is shown;

[0037] Figure 9 The positions of the top, left, and upper left samples of a pixel in the current block according to an embodiment of this application are shown;

[0038] Figure 10 An exemplary intra-frame mode of a recursive filter according to an embodiment of this application is shown;

[0039] Figure 11 Some exemplary directions in constrained directional enhancement filter (CDEF) processing according to some embodiments of this application are shown;

[0040] Figure 12 Some exemplary block divisions according to some embodiments of this application are shown;

[0041] Figure 13 An example is shown in which directional unit blocks are merged into filter units according to an embodiment of this application;

[0042] Figure 14 An example is shown in another embodiment of this application in which directional unit blocks are merged into filter units;

[0043] Figure 15 An exemplary mapping of directional information between the dominant direction and the directional intra-prediction mode derived during CDEF processing, according to some embodiments of this application, is shown.

[0044] Figure 16 An exemplary flowchart according to an embodiment of this application is shown; and

[0045] Figure 17 This is a schematic diagram of a computer system according to an embodiment of this application. Detailed Implementation

[0046] I. Video Decoder and Encoder Systems

[0047] Referring to Figure 1A, the lower right corner depicts a subset of nine prediction directions known from the 33 possible prediction directions of H.265 (corresponding to 33 angular modes out of 35 intra-frame modes). The point (101) where the arrows converge represents the sample being predicted. The arrows indicate the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted based on one or more samples at a 45-degree angle to the horizontal at the upper right. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more samples at a 22.5-degree angle to the horizontal at the lower left.

[0048] Referring again to Figure 1A, a 4×4 square block (104) of samples is depicted in the upper left (indicated by a bold dashed line). The square block (104) comprises 16 samples, each labeled with “S”, its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from top) and the first sample in the X dimension (from left). Similarly, sample S44 is the fourth sample in block (104) in both the Y and X dimensions. Since the block size is 4×4 samples, S44 is in the lower right. Also shown are reference samples following a similar numbering scheme. Reference samples are labeled with R, their Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, predicted samples are adjacent to the blocks in the reconstruction; therefore, negative values ​​are not required.

[0049] Intra-frame image prediction works by appropriately copying reference sample values ​​from neighboring samples based on the prediction direction indicated by a signal. For example, suppose the encoded video bitstream includes signaling for that block, which indicates a prediction direction consistent with arrow (102)—that is, predicting samples based on one or more prediction samples at a 45-degree angle to the horizontal from the upper right. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Then sample S44 is predicted based on reference sample R08.

[0050] In some cases, the values ​​of multiple reference samples can be combined, for example, by interpolation, to calculate the reference sample; especially when the direction is not divisible by 45 degrees.

[0051] As video coding technology has developed, the number of possible directions has also increased. In H.264 (2003), nine different directions could be represented. In H.265 (2013), this increased to 33, and JEM / VVC / BMS could support up to 65 directions at the time of publication. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding have been used to represent these possible directions with a small number of bits, thus accepting some penalty for less likely directions. Furthermore, the direction itself can sometimes be predicted based on adjacent directions used in adjacent decoded blocks.

[0052] Figure 1B shows a schematic diagram (105) depicting 65 intra-frame prediction directions based on JEM to illustrate how the number of prediction directions increases over time.

[0053] The mapping of intra-predicted direction bits representing direction in an encoded video bitstream can vary depending on the video coding technique; and this mapping can range from, for example, a simple direct mapping of the predicted direction to the intra-predicted mode, to codewords, to complex adaptive schemes involving the most probable mode, and similar techniques. However, in all cases, there may be certain directions that are statistically less likely to appear in the video content than some other directions. Since the goal of video compression is to reduce redundancy, in well-functioning video coding techniques, those less probable directions will be represented by a larger number of bits compared to the more probable directions.

[0054] Motion compensation can be a lossy compression technique and can involve using blocks of sample data from a previously reconstructed image or a portion thereof (the reference image), spatially shifted in a direction indicated by a motion vector (hereinafter referred to as MV), to predict a newly reconstructed image or a portion thereof. In some cases, the reference image may be the same as the image in the current reconstruction. The MV may have two dimensions, X and Y, or three dimensions, the third dimension being an indication of the reference image in use (the third dimension may indirectly be a temporal dimension).

[0055] In some video compression techniques, an MV applicable to a specific region of the sample data can be predicted based on other MVs, such as an MV that is spatially adjacent to the region being reconstructed and precedes that MV in the decoding order. This significantly reduces the amount of data required to encode the MV, thereby eliminating redundancy and increasing compression. MV prediction works effectively, for example, because when encoding an input video signal (called natural video) derived from a camera device, there is a statistical probability that a larger region than the region applicable to a single MV may move in similar directions, and therefore, in some cases, a larger region can be predicted using similar MVs derived from neighboring regions. This makes the MV obtained for a given region similar to or the same as the MV predicted based on surrounding MVs, and can be represented by fewer bits after entropy encoding compared to the number of bits used in directly encoding the MV. In some cases, MV prediction can be an example of lossless compression of the signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when calculating the predictor based on several surrounding MVs.

[0056] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T H.265 Recommendation, “High Efficiency Video Coding”, December 2016). This article describes a technique among the many MV prediction mechanisms provided by H.265, hereinafter referred to as “spatial combining”.

[0057] Referring to Figure 1C, the current block (111) may include samples obtained by the encoder during the motion search process that can be predicted based on previous blocks of the same size that have been spatially shifted. Instead of directly encoding this MV, the MV can be derived from metadata associated with one or more reference images, for example, from the most recent (in decoding order) reference image, using the MV associated with any of the five surrounding samples denoted as A0, A1 and B0, B1, B2 (corresponding to 112 to 116, respectively). In H.265, MV prediction can use a predictor from the same reference image that is also being used in adjacent blocks.

[0058] Figure 2 A simplified block diagram of a communication system (200) according to an embodiment of the present disclosure is shown. The communication system (200) includes a plurality of terminal devices that can communicate with each other via, for example, a network (250). For example, the communication system (200) includes a first pair of terminal devices (210) and (220) interconnected via the network (250). Figure 2 In the example, the first pair of terminal devices (210) and (220) perform unidirectional data transmission. For example, terminal device (210) may encode video data (e.g., a video image stream captured by terminal device (210)) for transmission via network (250) to another terminal device (220). The encoded video data may be transmitted as one or more encoded video bitstreams. Terminal device (220) may receive the encoded video data from network (250), decode the encoded video data to recover the video images, and display the video images based on the recovered video data. Unidirectional data transmission can be common in media service applications, etc.

[0059] In another example, the communication system (200) includes a second pair of terminal devices (230) and (240) that perform bidirectional transmission of encoded video data, which may occur, for example, during a video conference. For the bidirectional transmission of data, in this example, each of the terminal devices (230) and (240) may encode video data (e.g., a stream of video images captured by the terminal device) for transmission via a network (250) to the other terminal device (230) and (240). Each of the terminal devices (230) and (240) may also receive encoded video data transmitted by the other terminal device (230) and (240), and may decode the encoded video data to recover the video images, and may display the video images at an accessible display device based on the recovered video data.

[0060] exist Figure 2 In the examples, terminal devices (210), (220), (230), and (240) may be represented as servers, personal computers, and smartphones, but the principles of this disclosure are not limited thereto. Implementations of this disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (250) refers to any number of networks that transmit encoded video data between terminal devices (210), (220), (230), and (240), including, for example, wired (connected) and / or wireless communication networks. Communication networks (250) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this discussion, unless otherwise stated below, the architecture and topology of the network (250) may be irrelevant to the operation of this disclosure.

[0061] As an example of the application of the disclosed topic, Figure 3 The placement of a video encoder and video decoder in a streaming environment is illustrated. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0062] The streaming system may include a capture subsystem (313) that may include a video source (301), such as a digital camera device, that creates, for example, an uncompressed video picture stream (302). In the example, the video picture stream (302) includes samples captured by the digital camera device. The video picture stream (302) is depicted as a thick line to emphasize the high data volume when compared with encoded video data (304) (or encoded video bitstream), which may be processed by an electronic device (320) including a video encoder (303) coupled to the video source (301). The video encoder (303) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter as described in more detail below. Encoded video data (304) (or encoded video bitstream (304)) is depicted as a thin line to emphasize the lower data volume when compared to the video picture stream (302). The encoded video data (304) can be stored on a streaming server (305) for future use. One or more streaming client subsystems (e.g., Figure 3 The client subsystems (306) and (308) can access the streaming server (305) to retrieve copies (307) and (309) of the encoded video data (304). The client subsystem (306) may include, for example, a video decoder (310) in an electronic device (330). The video decoder (310) decodes the incoming copy (307) of the encoded video data and creates an outgoing video picture stream (311) that can be displayed on a display (312) (e.g., a screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (304), (307), and (309) (e.g., video bitstreams) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed topics can be used in the context of VVC.

[0063] Note that electronic devices (320) and (330) may include other components (not shown). For example, electronic device (320) may include a video decoder (not shown), and electronic device (330) may also include a video encoder (not shown).

[0064] Figure 4 A block diagram of a video decoder (410) according to an embodiment of the present disclosure is shown. The video decoder (410) may be included in an electronic device (430). The electronic device (430) may include a receiver (431) (e.g., a receiving circuitry system). The video decoder (410) may be used in place of... Figure 3 The video decoder (310) in the example.

[0065] The receiver (431) can receive one or more encoded video sequences to be decoded by the video decoder (410); in the same or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequences can be received from a channel (401), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (431) can receive the encoded video data as well as other data such as encoded audio data and / or auxiliary data streams, which can be forwarded to their respective user entities (not depicted). The receiver (431) can separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (415) can be coupled between the receiver (431) and the entropy decoder / parser (420) (hereinafter referred to as "parser (420)"). In some applications, the buffer memory (415) is part of the video decoder (410). In other applications, the buffer memory (415) can be external to the video decoder (410) (not depicted). In other applications, a buffer memory (not depicted) may exist outside the video decoder (410) to prevent network jitter, for example, and another buffer memory (415) may exist inside the video decoder (410) to handle broadcast timing, for example. The buffer memory (415) may not be necessary, or may be small, when the receiver (431) receives data from a store / forward device with sufficient bandwidth and controllability, or from an isosynchronous network. For use on best-effort packet networks such as the Internet, a buffer memory (415) may be required, which may be relatively large and advantageously have an adaptive size, and may be implemented at least partially in the operating system or in a similar element (not depicted) outside the video decoder (410).

[0066] The video decoder (410) may include a parser (420) to reconstruct symbols (421) based on the encoded video sequence. These symbols include information for managing the operation of the video decoder (410), and may include information for controlling a presentation device such as a presentation device (412) (e.g., a display screen), which is not part of the electronic device (430) but may be coupled to the electronic device (430), such as... Figure 4As shown. Control information for one or more presentation devices may take the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (420) may parse / decode the received encoded video sequence. The encoding of the encoded video sequence may conform to video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (420) may extract a subgroup parameter set for at least one subgroup of pixel subgroups in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. Subgroups may include: Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The parser (420) can also extract information such as transform coefficients, quantizer parameter values, MV, etc. from the encoded video sequence.

[0067] The parser (420) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (415) to create symbols (421).

[0068] The reconstruction of symbol (421) may involve multiple different units depending on the type of the encoded video picture or a portion thereof (e.g., inter-frame picture and intra-frame picture, inter-frame block and intra-frame block) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (420). For the sake of brevity, such subgroup control information flow between the parser (420) and the multiple units described below is not depicted.

[0069] In addition to the functional blocks already mentioned, the video decoder (410) can be conceptually subdivided into several functional units as described below. In practical implementations operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other at least partially. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0070] The first unit is the scaler / inverse transform unit (451). The scaler / inverse transform unit (451) receives from the parser (420) quantized transform coefficients as symbols (421) and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (451) can output a block containing sample values, which can be input into the aggregator (455).

[0071] In some cases, the output samples of the scaler / inverse transform (451) may belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images but can use predictive information from previously reconstructed portions of the current image. Such predictive information may be provided by the intra-picture prediction unit (452). In some cases, the intra-picture prediction unit (452) generates blocks of the same size and shape as the blocks in the reconstruction using surrounding reconstructed information obtained from the current picture buffer (458). For example, the current picture buffer (458) buffers partially reconstructed current images and / or fully reconstructed current images. In some cases, the aggregator (455) adds the predictive information already generated by the intra-picture prediction unit (452) to the output sample information provided by the scaler / inverse transform unit (451) based on each sample.

[0072] In other cases, the output samples of the scaler / inverse transform unit (451) may belong to inter-frame encoded and possibly motion-compensated blocks. In such cases, the motion-compensated prediction unit (453) can access the reference image memory (457) to obtain samples for prediction. After motion compensation of the obtained samples according to the symbols (421) belonging to the block, these samples can be added by the aggregator (455) to the output of the scaler / inverse transform unit (451) (referred to in this case as residual samples or residual signals) to generate output sample information. The address in the reference image memory (457) from which the motion-compensated prediction unit (453) obtains the predicted samples can be controlled by the MV, which can be obtained by the motion-compensated prediction unit (453) in the form of symbols (421), which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values ​​obtained from the reference image memory (457) when using subsampled accurate MV, MV prediction mechanisms, etc.

[0073] The output samples of the aggregator (455) can undergo various loop filtering techniques in the loop filter unit (456). The video compression technique may include an in-loop filtering technique controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), which may be obtained by the loop filter unit (456) as symbols (421) from the parser (420). However, the in-loop filtering technique may also respond to metadata obtained during the decoding of previous portions of the encoded picture or encoded video sequence (in the decoding order), as well as to previously reconstructed and loop-filtered sample values.

[0074] The output of the loop filter unit (456) can be a sample stream, which can be output to the presentation device (412) and stored in the reference image memory (457) for future inter-frame image prediction.

[0075] Once fully reconstructed, certain encoded images can be used as reference images for future predictions. For example, once the encoded image corresponding to the current image has been fully reconstructed and that encoded image (by, for example, the parser (420)) is identified as the reference image, the current image buffer (458) can become part of the reference image memory (457), and a new current image buffer can be reallocated before reconstructing subsequent encoded images begins.

[0076] The video decoder (410) can perform decoding operations according to a predetermined video compression technique as specified in a standard such as ITU-T Recommendation H.265. An encoded video sequence can conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile can select certain tools from all available tools in the video compression technique or standard as tools available only under that profile. For compliance, the complexity of the encoded video sequence also needs to be within the limits defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured, for example, in megapixels per second), maximum reference picture size, etc. In some cases, the limitations set by the level can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is signaled in the encoded video sequence.

[0077] In this implementation, the receiver (431) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of one or more encoded video sequences. The additional data may be used by the video decoder (410) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0078] Figure 5 A block diagram of a video encoder (503) according to an embodiment of the present disclosure is shown. The video encoder (503) is included in an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmission circuitry system). The video encoder (503) can be used in place of... Figure 3 The example video encoder (303).

[0079] The video encoder (503) can obtain data from the video source (501) (not...). Figure 5 In one example, an electronic device (520) receives video samples, and a video source (501) can capture one or more video images to be encoded by a video encoder (503). In another example, the video source (501) is part of the electronic device (520).

[0080] A video source (501) can provide a sequence of source video samples in the form of a digital video sample stream to be encoded by a video encoder (503). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit…), any color space (e.g., BT.601 YCrCb, RGB…), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (501) can be a storage device storing previously prepared video. In a video conferencing system, the video source (501) can be a camera device capturing local image information as a video sequence. Video data can be provided as multiple individual pictures that are given motion when viewed sequentially. The pictures themselves can be organized as spatial pixel arrays, where each pixel can include one or more samples, depending on the sampling structure, color space, etc., used. Those skilled in the art will readily understand the relationship between pixels and samples. The following description focuses on samples.

[0081] According to the implementation, the video encoder (503) can encode and compress images of the source video sequence into an encoded video sequence (543) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of the controller (550). In some implementations, the controller (550) controls and is functionally coupled to other functional units as described below. Coupling is not depicted for simplicity. Parameters set by the controller (550) may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum allowed reference area for maximum video volume (MV), etc. The controller (550) can be configured to have other suitable functions belonging to the video encoder (503) optimized for a specific system design.

[0082] In some implementations, the video encoder (503) is configured to operate within an encoding loop. As a highly simplified description, in this example, the encoding loop may include a source encoder (530) (e.g., responsible for creating symbols, such as a symbol stream, based on the input picture to be encoded and one or more reference pictures) and a (local) decoder (533) embedded within the video encoder (503). The decoder (533) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (since any compression between the symbols and the encoded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory (534). Since decoding of the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference picture memory (534) are also bit-accurate between the local and remote encoders. In other words, the encoder's prediction portion "treats" the reference picture samples as if they were the exact same sample values ​​that the decoder "sees" during prediction. This fundamental principle of image synchronization (and the offset that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related techniques.

[0083] The operation of the "local" decoder (533) can be combined with the "remote" decoder, as already mentioned above. Figure 4 The operation of the video decoder (410) described in detail is the same. However, a brief additional reference is provided. Figure 4 Since symbols are available and the encoding of symbols into an encoded video sequence by the entropy encoder (545) and the decoding of symbols by the parser (420) can be lossless, the entropy decoding portion of the video decoder (410), including the buffer memory (415) and the parser (420), can be completely implemented in the local decoder (533).

[0084] It can be observed that any decoder technique other than parsing / entropy decoding, which exists in the decoder, must also exist in the corresponding encoder in essentially the same functional form. For this reason, the subject matter disclosed focuses on decoder operation. The description of encoder techniques can be simplified, as encoder techniques are the opposite of those comprehensively described decoder techniques. Only certain aspects require a more detailed description, which is provided below.

[0085] In some examples, during operation, the source encoder (530) may perform motion-compensated predictive coding, which predictively codes the input image with reference to one or more previously encoded images designated as "reference images" from the video sequence. In this way, the encoding engine (532) encodes the differences between pixel blocks of the input image and pixel blocks of one or more reference images (one or more) that can be selected as predictive references to the input image.

[0086] The local video decoder (533) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (530). The operation of the encoding engine (532) can be advantageously for lossy processing. When the encoded video data can be decoded by the video decoder (533), Figure 5 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (533) replicates the decoding processing performed on the reference image by the video decoder and can store the reconstructed reference image in a reference image buffer (534). In this way, the video encoder (503) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0087] The predictor (535) can perform a prediction search against the encoding engine (532). That is, for a new image to be encoded, the predictor (535) can search in the reference image memory (534) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference image MV, block shape, etc., that can be used as suitable prediction references for the new image. The predictor (535) can operate on a sample-by-sample block-by-pixel basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (535), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (534).

[0088] The controller (550) can manage the encoding operations of the source encoder (530), including, for example, setting parameters and subgroup parameters for encoding video data.

[0089] The outputs of all the functional units mentioned above can be entropy encoded in the entropy encoder (545). The entropy encoder (545) converts the symbols generated by the various functional units into an encoded video sequence by performing lossless compression on the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0090] The transmitter (540) can buffer one or more encoded video sequences created by the entropy encoder (545) in preparation for transmission via a communication channel (560), which may be a hardware / software link to a storage device storing the encoded video data. The transmitter (540) can combine the encoded video data from the video encoder (503) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0091] The controller (550) can manage the operation of the video encoder (503). During encoding, the controller (550) can assign a specific encoded image type to each encoded image, which may affect the encoding techniques that can be applied to the corresponding image. For example, one of the following image types can typically be assigned to an image:

[0092] An intra-frame picture (I-picture) can be a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will understand those variations of I-pictures and their corresponding applications and characteristics.

[0093] Predictive images (P-images) can be images that can be encoded and decoded using inter-frame prediction or intra-frame prediction—using at most one MV and a reference index to predict the sample values ​​of each block.

[0094] A bidirectional predictive picture (B-picture) can be an image that can be encoded and decoded using either inter-frame or intra-frame prediction—which uses up to two MVs and a reference index to predict sample values ​​for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0095] Source images are typically spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, determined by the coding assignments of the corresponding images applied to the blocks. For example, blocks of image I can be non-predictively encoded, or blocks of image I can be predictively encoded (spatial prediction or intra-frame prediction) with reference to already encoded blocks of the same image. Pixel blocks of image P can be predictively encoded with reference to a previously encoded reference image via spatial prediction or temporal prediction. Blocks of image B can be predictively encoded with reference to one or two previously encoded reference images via spatial prediction or temporal prediction.

[0096] The video encoder (503) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In its operation, the video encoder (503) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0097] In this implementation, the transmitter (540) may transmit additional data along with the encoded video. The source encoder (530) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant images and slices, SEI messages, VUI parameter set fragments, etc.

[0098] Video can be captured temporally as multiple source images (video frames). Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In the example, the specific image being encoded / decoded—referred to as the current image—is segmented into blocks. Where a block in the current image resembles a reference block in a previously encoded and buffered reference image in the video, the block in the current image can be encoded using a vector called MV. MV points to the reference block in the reference image, and when using multiple reference images, MV can have a third dimension identifying the reference images.

[0099] In some implementations, bidirectional prediction techniques can be used for inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image that both precede the current image in the video in decoding order (but may be in the past and future in display order, respectively). A block in the current image can be encoded using a first MV pointing to a first reference block in the first reference image and a second MV pointing to a second reference block in the second reference image. The block can be predicted using a combination of the first and second reference blocks.

[0100] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0101] According to some embodiments of this disclosure, predictions such as inter-frame picture prediction and intra-frame picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a video picture sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the images have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU comprises three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU can be recursively divided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In the example, each CU is analyzed to determine the prediction type for that CU, such as inter-frame prediction or intra-frame prediction. Depending on temporal and / or spatial predictability, a CU is divided into one or more prediction units (PUs). Typically, each PU includes a luminance prediction block (PB) and two chrominance PBs. In implementations, prediction operations in encoding / decoding are performed on a per-prediction-block basis. Using a luminance prediction block as an example, a prediction block includes a matrix of pixel values ​​(e.g., luminance values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0102] Figure 6 A diagram of a video encoder (603) according to another embodiment of the present disclosure is shown. The video encoder (603) is configured to receive processing blocks (e.g., prediction blocks) of sample values ​​within a current video image in a video image sequence, and to encode the processing blocks into an encoded image that is part of an encoded video sequence. In the example, the video encoder (603) is used instead of Figure 3 The video encoder (303) in the example.

[0103] In the HEVC example, the video encoder (603) receives a matrix of sample values ​​for a processing block, such as a prediction block of 8×8 samples. The video encoder (603) uses, for example, rate-distortion optimization to determine whether to best encode the processing block using intra-frame mode, inter-frame mode, or bidirectional prediction mode. When encoding the processing block in intra-frame mode, the video encoder (603) can encode the processing block into an encoded picture using intra-frame prediction techniques; and when encoding the processing block in inter-frame mode or bidirectional prediction mode, the video encoder (603) can encode the processing block into an encoded picture using inter-frame prediction or bidirectional prediction techniques, respectively. In some video coding techniques, the merging mode can be an inter-frame picture prediction sub-mode, where the MV is derived from the predicted MV components without the aid of one or more external MV predictors. In some other video coding techniques, there may be MV components applicable to the target block. In this example, the video encoder (603) includes other components, such as a mode decision module (not shown) that determines the mode of the processing block.

[0104] exist Figure 6 In the example, the video encoder (603) includes, for example, Figure 6 The inter-frame encoder (630), intra-frame encoder (622), residual calculator (623), switch (626), residual encoder (624), overall controller (621), and entropy encoder (625) are shown coupled together.

[0105] An inter-frame encoder (630) is configured to receive samples of the current block (e.g., the processing block), compare the block with one or more reference blocks in a reference image (e.g., blocks in previous and subsequent images), generate inter-frame prediction information (e.g., a description based on redundancy information, MV, and merging mode information of the inter-frame coding technique), and compute inter-frame prediction results (e.g., prediction blocks) based on the inter-frame prediction information using any suitable technique. In some examples, the reference image is a decoded reference image based on encoded video information.

[0106] The intra encoder (622) is configured to: receive samples of the current block (e.g., the processing block); in some cases compare the block with already encoded blocks in the same image; generate quantization coefficients after transformation; and in some cases also generate intra prediction information (e.g., intra prediction direction information based on one or more intra coding techniques). In the example, the intra encoder (622) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same image.

[0107] The overall controller (621) is configured to determine overall control data and control other components of the video encoder (603) based on the overall control data. In an example, the overall controller (621) determines the mode of the block and provides control signals to the switch (626) based on the mode. For example, when the mode is intra-frame mode, the overall controller (621) controls the switch (626) to select intra-frame mode results for use by the residual calculator (623) and controls the entropy encoder (625) to select intra-frame prediction information and include the intra-frame prediction information in the bitstream; and when the mode is inter-frame mode, the overall controller (621) controls the switch (626) to select inter-frame prediction results for use by the residual calculator (623) and controls the entropy encoder (625) to select inter-frame prediction information and include the inter-frame prediction information in the bitstream.

[0108] A residual calculator (623) is configured to calculate the difference (residual data) between the received block and a prediction result selected from an intra-encoder (622) or an inter-encoder (630). A residual encoder (624) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In an example, the residual encoder (624) is configured to transform the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients. In various embodiments, the video encoder (603) also includes a residual decoder (628). The residual decoder (628) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra-encoder (622) and the inter-encoder (630). For example, an inter-frame encoder (630) can generate decoded blocks based on decoded residual data and inter-frame prediction information, and an intra-frame encoder (622) can generate decoded blocks based on decoded residual data and intra-frame prediction information. In some examples, the decoded blocks are appropriately processed to generate decoded images, and these decoded images can be buffered in memory circuitry (not shown) and used as reference images.

[0109] An entropy encoder (625) is configured to format the bitstream to include encoded blocks. The entropy encoder (625) is configured to include various information according to a suitable standard such as HEVC. In this example, the entropy encoder (625) is configured to include overall control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. Note that, according to the disclosed subject matter, residual information is not present when blocks are encoded in a merged sub-mode of inter-frame mode or bidirectional prediction mode.

[0110] Figure 7A diagram of a video decoder (710) according to another embodiment of the present disclosure is shown. The video decoder (710) is configured to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed image. In the example, the video decoder (710) is used instead of Figure 3 The video decoder (310) in the example.

[0111] exist Figure 7 In the example, the video decoder (710) includes, for example, Figure 7 The entropy decoder (771), inter-frame decoder (780), residual decoder (773), reconstruction module (774), and intra-frame decoder (772) are shown coupled together.

[0112] The entropy decoder (771) can be configured to reconstruct certain symbols from the encoded picture, which represent the syntax elements constituting the encoded picture. Such symbols may include, for example, the mode encoding the block (e.g., intra-mode, inter-mode, bidirectional prediction mode, a combined sub-mode of the latter two, or another sub-mode), prediction information (e.g., intra-prediction information or inter-prediction information) that can identify certain samples or metadata used by the intra-decoder (772) or inter-decoder (780) for prediction, such as residual information in the form of quantized transform coefficients, etc. In the example, when the prediction mode is inter-mode or bidirectional prediction mode, the inter-prediction information is provided to the inter-decoder (780); and when the prediction type is intra-prediction type, the intra-prediction information is provided to the intra-decoder (772). The residual information may be inversely quantized and provided to the residual decoder (773).

[0113] The inter-frame decoder (780) is configured to receive inter-frame prediction information and generate inter-frame prediction results based on the inter-frame prediction information.

[0114] The intra-frame decoder (772) is configured to receive intra-frame prediction information and generate prediction results based on the intra-frame prediction information.

[0115] The residual decoder (773) is configured to perform inverse quantization to extract the dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require some control information (to include quantizer parameters (QP)), and this information may be provided by the entropy decoder (771) (the data path is not depicted because this may only be a small amount of control information).

[0116] The reconstruction module (774) is configured to combine the residual output by the residual decoder (773) with the prediction result (output by the inter-frame prediction module or intra-frame prediction module, depending on the situation) in the spatial domain to form a reconstructed block, which can be a part of a reconstructed image, which in turn can be a part of a reconstructed video. Note that other suitable operations, such as deblocking, can be performed to improve visual quality.

[0117] Note that any suitable technology can be used to implement the video encoders (303), (503), and (603) and the video decoders (310), (410), and (710). In one implementation, one or more integrated circuits can be used to implement the video encoders (303), (503), and (603) and the video decoders (310), (410), and (710). In another implementation, one or more processors that execute software instructions can be used to implement the video encoders (303), (503), and (603) and the video decoders (310), (410), and (710).

[0118] II. Intra-frame prediction

[0119] In some relevant examples, such as VP9, ​​eight orientation modes are supported, corresponding to angles from 45 degrees to 207 degrees. To utilize more spatial redundancy in orientation textures, in some relevant examples, such as AOMedia Video 1 (AV1), the intra-frame orientation modes are extended to a finer-grained set of angles. The original eight angles are slightly modified and referred to as nominal angles, and these eight nominal angles are named V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED.

[0120] Figure 8 An exemplary nominal angle according to an embodiment of this disclosure is shown. Each nominal angle can be associated with seven finer angles, thus in some relevant examples such as AV1, there can be a total of 56 orientation angles. The predicted angle is represented by the nominal intra-frame angle plus an angle increment, which is obtained by multiplying a factor (ranging from -3 to 3) by a step size of 3 degrees. To implement the orientation prediction mode in AV1 in a general manner, all 56 orientation intra-frame predicted angles in AV1 can be implemented using a uniform orientation predictor that projects each pixel to a reference sub-pixel position and interpolates the reference sub-pixel through a 2-tap bilinear filter.

[0121] In some relevant examples such as AV1, there are five non-directional smooth intra-prediction modes: DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. For DC prediction, the average of the left and top neighboring samples is used as the prediction value for the block to be predicted. For PAETH prediction, top, left, and top-left reference samples are first obtained, and then the closest (top + left – top-left) value is set as the prediction value for the pixel to be predicted.

[0122] Figure 9 The locations of the top, left, and upper left samples of a pixel in the current block according to an embodiment of this disclosure are shown. For SMOOTH, SMOOTH_V, and SMOOTH_H modes, the block is predicted using quadratic interpolation in the vertical or horizontal direction or the average of the two directions.

[0123] Figure 10 An exemplary recursive filter intra-frame mode according to an embodiment of this disclosure is shown.

[0124] To capture the attenuation spatial correlation with edge references, FILTER INTRA modes are designed for luma patches. Five FILTER INTRA modes are defined in AV1, each represented by a set of eight 7-tap filters, reflecting the correlation between pixels in a 4×2 patch and its seven neighboring pixels. For example, the weighting factors of the 7-tap filters are position-dependent. Figure 10 As shown, the 8×8 block is divided into eight 4×2 sub-blocks indicated by B0, B1, B2, B3, B4, B5, B6, and B7. For each sub-block, the seven neighbors of each sub-block indicated by R0 to R6 are used to predict the pixels in the corresponding sub-block. For sub-block B0, all neighbors have been reconstructed. However, for other sub-blocks, when not all neighbors have been reconstructed, the predictions of the immediate neighbors are used as reference values. For example, since none of the neighbors of sub-block B7 have been reconstructed, the prediction samples of the neighbors of sub-block B7 (i.e., B5 and B6) are used instead.

[0125] For the chroma component, the pure chroma intra-frame prediction mode (called chroma from luma (CfL) mode) models chroma pixels as a linear function of the corresponding reconstructed luma pixels. CfL prediction can be expressed as follows:

[0126] CfL(α)=α×L AC +DC Equation (1)

[0127] Among them, L ACThe AC contribution of the luminance component is represented by α, the parameter of the linear model is represented by α, and the DC contribution of the chrominance component is represented by DC. In the example, the reconstructed luminance pixels are subsampled to the chrominance resolution and then the average value is subtracted to form the AC contribution. To approximate the chrominance AC component from the AC contribution, instead of requiring the decoder to calculate the scaling parameter as in some related examples, the CfL mode in AC1 determines the parameter α based on the original chrominance pixels and signals them in the bitstream. This reduces the complexity of the decoder and produces more accurate predictions. As for the DC contribution of the chrominance component, it is calculated using the intra-frame DC mode, which is sufficient for most chrominance content and has a mature and fast implementation.

[0128] III. Loop Filtering

[0129] In some relevant examples, such as AV1, three in-loop filters can be applied to the reconstructed frame in the order of deblocking, constrained directional enhancement filter (CDEF), and loop recovery filter. The loop recovery filter includes a Wiener filter and a self-guided projection (SGRPRJ) filter, one of which can be adaptively selected.

[0130] Deblocking can be applied across transform block boundaries to reduce block artifacts caused by the quantization of transform coefficients. In some examples, 4-, 8-, and 14-tap finite impulse response (FIR) filters can be used for the luminance block, and 4- and 6-tap FIR filters can be used for the chroma block.

[0131] The filter length is initially determined by the minimum transform block size at the boundary. Variance can be used to perform condition checks to avoid blurring the actual edges. Furthermore, a flatness check can be combined to ultimately determine the filter length.

[0132] CDEF is a nonlinear deringing filter applied along the directional features detected in the target region. In some relevant examples, an 8×8 region is the cell size for which CDEF is performed. Figure 11 As shown, standardized orientation detection can be performed. Figure 11 Some exemplary directions in CDEF processing are shown.

[0133] Choose Figure 11 The candidate direction d(0~7) that minimizes the following quantities is taken as the dominant direction.

[0134]

[0135] Where, x p It is the value of pixel P, P d,k It refers to the pixels in line k along direction d, and μ. d,k P is obtained as follows d,k Average value:

[0136]

[0137] The sum of all sample values ​​within the block is a constant. Therefore, minimizing equation (2) corresponds to maximizing the following equation:

[0138]

[0139] Once the dominant direction is determined, the following filtering operation can be performed: Perform primary filtering along the selected dominant direction, and simultaneously perform secondary filtering along a secondary direction that deviates from the dominant (or primary) direction by 45 degrees.

[0140]

[0141] Among them, w p and w s These are the fixed filter coefficients for the primary and secondary filters, respectively, and the piecewise linear function f is given below:

[0142]

[0143] Here, S and D represent intensity and damping values, respectively, and a maximum of 8 preset (S, D) values ​​for luminance / chrominance are signaled per frame.

[0144] When applying filtering, each 64×64 block within a tile can select one of the presets, and filtering can be performed on each 8×8 cell within the corresponding 64×64 block.

[0145] Note that during decoding, several variables related to the signal characteristics of local regions can be derived or resolved from the bitstream. These variables include dir, var, priStr, secStr, and damping. The variable dir represents the dominant edge direction of the 8×8 block. The variable var represents the variance of the signal values ​​within the 8×8 block and is defined as the normalized difference between the cost along the direction orthogonal to the dominant direction and the cost along the dominant direction. The variable priStr represents the primary filter strength S of a 64×64 block containing 8×8 filter units. p The variable secStr represents the secondary filter strength S of a 64×64 block containing 8×8 filter units. s The variable `damping` represents the damping parameter D of a 64×64 block containing 8×8 filter units. These values ​​can be obtained separately for the luminance and chrominance channels.

[0146] After performing deblocking and CDEF processing, two types of recovery filters can be applied mutually exclusively in some relevant examples such as AV1. These two types of recovery filters include Wiener filters and SGRPRJ filters. The size of the square loop-restoration unit (LRU) can be selected from 64×64 to 256×256.

[0147] In the Wiener filter, the quality of each reconstructed pixel in the encoded frame can be improved by noncausal filtering of neighboring pixels within a W×W window surrounding the corresponding pixel. The 2D filter tap of the Wiener filter can be represented by F and determined as follows:

[0148] F = H -1 Equation (7)

[0149] Where H = E[XX] T [ is the white covariance of x, which includes the column vectorized W in the W×W window.] 2 There are 10 samples, and M = E[YX] T ] is the cross-correlation between x and the original source sample y.

[0150] In some relevant examples, such as AV1, the separability of F and the symmetry and normalization of the Wiener filter coefficients can be imposed as constraints. The Wiener filter coefficients F (formed as W) 2 A ×1 vector can be defined as:

[0151] F = column vectorization [ab] T Equation (8)

[0152] Here, a and b are W×1 vertical and horizontal filters, such that for i = 0, 1, ..., r-1, a(i) = a(W-1-i), b(i) = b(W-1-i), and ∑a(i) = ∑b(i) = 1. The coefficient vectors a and b can be searched at the encoder and encoded in the bitstream.

[0153] In SGRPRJ filtering, a simple linear filter, described by the following model, is performed to obtain a simple restored version based on the degraded image x.

[0154]

[0155] F and G can be obtained using both the guiding image and the degraded image. In some relevant examples, such as AV1, a self-guided filtering method is used, where F and G are determined using only the statistics of the degraded image itself, instead of a separate guiding image.

[0156] More specifically, the local mean (μ) and variance (σ) of the pixels within a (2r+1)×(2r+1) window surrounding the pixel can be calculated. 2 And each pixel x can be filtered as follows:

[0157]

[0158] Where r specifies the search window size, and e is a noise parameter that controls the denoising intensity.

[0159] Equation (9) gives two simple reconstructions X1 and X2 based on the degraded image X. The following subspace projection is performed to construct the final output X. r ,

[0160] X r =X + α(X1-X) + β(X2-X) Equation (11)

[0161] Using X, X1, X2, and source Y, the encoder can calculate α and β as follows:

[0162] [αβ] T =(A T A) -1 A T b. Equation (12)

[0163] Where A = {X1-X, X2-X}, and b = YX.

[0164] Then, the encoder can send a 6-tuple (r1, e1, r2, e2, α, β) for each LRU.

[0165] IV. Feature information reuse for enhanced recovery filtering

[0166] In some relevant examples, such as AV1, Wiener filtering can be performed on square cells ranging in size from 64×64 to 256×256 by uniformly dividing frames / tiles into LRUs. In the example, the filter coefficients of the Wiener filter can be obtained by assuming that the signal statistics are stationary. Therefore, it is desirable to classify the target region being filtered into one of the classification statistics types in which the stationarity assumption can reasonably hold. Possible methods for classifying the target region include using quantities such as local variance or edge information. While these quantities themselves, or the associated category information, can be computed at the decoder or signaled in the bitstream, this can be expensive in terms of computational or bit rate overhead.

[0167] In some relevant examples, such as AV1, SGRPRJ filtering can be performed on square cells ranging in size from 64×64 to 256×256 by uniformly dividing frames / tiles into LRUs. In the SGRPRJ filter, a simple edge-preserving form of filtering is performed to construct a simple reconstructed image using noise parameter pairs for each LRU and a fixed radius. Furthermore, fixed projection parameters α and β can be used as weighting factors for the error image for each LRU to form the final reconstruction. However, regions in the error image can have different statistical characteristics reflecting local signal features such as edges and texture. Therefore, using or estimating a single set of SGRPRJ filter parameters (e.g., radius, noise parameters, α, and β) over LRU regions covering pixels with widely varying signal statistics can compromise the quality of the final reconstruction. On the other hand, combining signal classification for greater adaptability can present essentially the same challenges as the Wiener filtering case in terms of additional computational burden or bit rate overhead.

[0168] This disclosure includes methods for enhancing the performance of recovery filtering techniques by reusing feature information (e.g., directional information) derived from CDEF processing and / or intra-frame prediction modes. For example, adaptive recovery filtering techniques can effectively reuse signal features and statistical information already available at the decoder.

[0169] In this disclosure, a recovery filter (or filtering) process can be defined as an operation applied to a noisy image and a filtering process that estimates a clean and original image based on the noisy image. The recovery filter process can include processing for a blurred image or inverse processing of the inverse operation for a blurred image. Examples of recovery filter processes include, but are not limited to, Wiener filtering and SGRPRJ filtering. A recovery filter (or filtering) unit is the region where the recovery filter process is performed.

[0170] In this disclosure, a directional information unit can be defined as a group of pixels with a specified shape and size, providing the dominant directionality of a feature represented by the pixel values ​​of that group. In an example such as AV1, each directional information unit in CDEF processing can be an 8×8 block. The dominant direction and the variance values ​​of the pixels in each 8×8 block can be derived in a normalized manner. In another example, the directional intra-prediction mode in AV1 can provide such information to units of different shapes and sizes corresponding to the intra-prediction blocks.

[0171] According to aspects of this disclosure, directional information derived at the decoder can be reused to infer the presence and directionality of boundary edges for a recovery filter (e.g., a Wiener filter or an SGRPRJ filter). For example, directional information can be derived from CDEF processing.

[0172] According to some implementations, the shape and size of the recovery filtering unit (e.g., a Wiener filtering unit or an SGRPRJ filtering unit) can be defined using multiple available directional information units (e.g., an 8×8 block used in CDEF direction detection and filtering in AV1). In this way, finer-grained directional adaptability can be achieved. Therefore, recovery filtering can be performed with a smaller unit size compared to one of the fixed square types 64×64, 128×128, or 256×256 used in some related examples, such as AV1.

[0173] In one implementation, the size of the recovery filter unit can be the same as, for example, the LRU size defined in AV1.

[0174] In one implementation, the size of the recovery filter unit can be the same as the size of the directional information unit (e.g., 8×8).

[0175] In one implementation, the recovery filter unit can also be partitioned from a given LRU size into square, rectangular, T-shaped, or 4-way sub-LRUs similar to or consistent with partitions in some related examples (e.g., the partitions in AV1), such as... Figure 12 As shown.

[0176] In one implementation, blocks of various sizes (e.g., 8×8) with directional information units can be merged to form a filter unit by following various scanning orders. Figure 13 An example is shown of merging blocks of 8×8 directional unit size into filter units of 32×8 size in raster scan order. For example, four 8×8 directional unit blocks (1301) to (1304) can be merged into a 32×8 filter unit block (1310), and four 8×8 directional unit blocks (1305) to (1308) can be merged into a 32×8 filter unit block (1320).

[0177] In one implementation, filtering units can be formed by merging blocks of similar size (e.g., 8×8) that each has directional information units by following various scanning orders. Figure 14 An example is shown of merging blocks, each with an 8×8 directional unit size, into a variable-size filter unit in raster scan order. The filter unit sizes include 8×8, 16×8, and 32×8. For example, one 8×8 directional unit block (1401) can be an 8×8 filter unit block (1410), two 8×8 directional unit blocks (1403) to (1404) can be merged into a 16×8 filter unit block (1420), and four 8×8 directional unit blocks (1405) to (1408) can be merged into a 32×8 filter unit block (1430).

[0178] According to some implementations, each available directivity of the CDEF process can be directly used as a category index for the signal category, and a unique set of recovery filter shapes and sizes can be defined for each signal category. That is, the selection of the recovery filter can depend on the available directivity of the CDEF process.

[0179] In one implementation, each signal category can be applied with a solution to an equation relating to the computation of the recovery filter, such as equation (7).

[0180] In one implementation, each signal category may use 2D filters with different numbers of filter taps and different shapes, with or without symmetry.

[0181] In one implementation, the separability of the 2D filter (separable or non-separable filter) can depend on the available directionality of the CDEF processing.

[0182] In one implementation, multiple directions beyond the available directions processed by CDEF can be merged into a single category, resulting in a reduction in the number of directional categories. A unique set of recovery filter shapes and sizes can be defined for each merged category.

[0183] According to some implementation methods, in addition to directionality, block variance information can be combined to further refine the directionality-based categories.

[0184] In one implementation, the 8×8 directional information units in the CDEF process can be further classified into different subclasses. Classification can be based on the variance value of the directional information units. For example, if the number of directional-based categories is 5 and the number of variance-based categories is 3, there can be 15 signal categories, and a set of recovery filters (e.g., Wiener filters or SGRPRJ filters) can be designed for each of the 15 signal categories.

[0185] According to some implementations, in addition to directionality, filter strengths can be combined to determine the signal category of the recovery filter unit. For example, the strengths of the primary CDEF filter and the secondary CDEF filter, which are signaled by a signal, can be combined to determine the signal category of the recovery filter unit. Different filter strength presets selected by the encoder can indicate different signal characteristics of the target area.

[0186] In one implementation, one of the preset primary and secondary filter strengths, indicated by signals in the bitstream, can be directly used as another dimension for the signal category index. For example, if the number of categories based on directionality is 5 and the number of categories based on preset filter strengths is 4, there can be 20 signal categories, for which a set of recovery filters (e.g., Wiener filters or SGRPRJ filters) can be designed.

[0187] According to some implementations, a directional majority vote or consensus check, including in the recovery filter unit, can be performed to determine the filter category.

[0188] In one implementation, when the size and shape of the recovery filter unit are fixed and the number of directional information units included in the recovery filter unit (e.g., 8×8 in the case of CDEF) is greater than a predefined number, a majority vote or consensus check of the directionality included in the recovery filter unit can be performed to determine the filter category.

[0189] In one implementation, for a majority vote, the most frequent direction among the available and potentially merged directions can be selected. In the example, a margin can be set between the first and second most frequent directions.

[0190] In one implementation, before a majority vote, it is determined whether the number of categories within the recovery filter unit is greater than a predefined number. If true, inconsistency can be declared, and explicit signaling for the recovery filter unit or a smaller recovery filter unit can be selected.

[0191] According to some implementations, the same mode and directionality information used for the recovery filter unit in the luminance component can be used for the recovery filter unit in the chrominance component, where such information is only available for the luminance component from the CDEF process.

[0192] According to some implementations, the recovery filter unit in the chroma component can use such information if its own filtering strength (e.g., preset values ​​including the primary filter strength and secondary filter strength in the chroma component) is available to the chroma component.

[0193] According to some implementations, the recovery filtering unit in the chromaticity component can use such information when its own variance information is available to the chromaticity component.

[0194] In one implementation, when CDEF processing is turned off and the recovery filter is turned on, the direction search process in CDEF processing, as described in Section III (Loop Filtering), can be applied to determine the existence and directionality of the boundary edges for the recovery filter unit.

[0195] In one implementation, when CDEF processing is disabled and the recovery filter is enabled, a default signal category is selected or explicit signaling of the filter category index can be executed.

[0196] According to aspects of this disclosure, directional information indicated by the intra-prediction mode available at the decoder can be reused in the recovery filtering unit (e.g., Wiener filtering unit or SGRPRJ filtering unit) as a guide for labeling signal categories, and a unique set of recovery filter shapes and sizes can be defined for each signal category. In some embodiments, such directional information can be provided by the encoder in varying unit sizes for intra-prediction.

[0197] According to some implementations, the shape and size of the reconstructive filtering unit can be defined using multiple available directional information units (e.g., 8×8 for directional intra-prediction units in AV1). In this way, finer-grained directional adaptability can be achieved. Therefore, reconstructive filtering can be performed with a smaller unit size compared to one of the fixed square types 64×64, 128×128, or 256×256 used in some relevant examples, such as AV1.

[0198] In one implementation, the size of the recovery filter unit can be the same as, for example, the LRU size defined in AV1.

[0199] In one implementation, the size of the recovery filter unit can be the same as the size of the directional information unit (e.g., 8×8).

[0200] In one implementation, the recovery filter unit can also be partitioned from a given LRU size into square, rectangular, T-shaped, or 4-way sub-LRUs similar to or consistent with partitions in some related examples (e.g., the partitions in AV1), such as... Figure 12 As shown.

[0201] In one implementation, filtering units can be formed by merging blocks of directional information units (e.g., 8×8) of a certain size by following various scanning orders, such as... Figure 13 As shown.

[0202] In one implementation, filtering units can be formed by merging directional information units of a certain size (e.g., 8×8) and blocks with similar orientations by following various scanning orders, such as... Figure 14 As shown.

[0203] According to some implementations, a fixed number of signal categories can be defined, and blocks within each signal category can have intra-prediction modes with similar directionality. A unique set of recovery filter shapes and sizes can be defined for each signal category. If non-angular intra-frame modes such as SMOOTH (including SMOOTH, SMOOTH_H, SMOOTH_V modes), Paeth predictors, or DC modes are used, each or a combination of non-angular intra-frame modes can be associated with its own signal category.

[0204] In one implementation, each of the eight nominal angles in AV1 can be combined with seven associated possible incremental angles to form a total of eight directional categories.

[0205] In one implementation, the orientation category may depend on both the nominal angle and the incremental angle associated with the nominal angle.

[0206] According to some implementations, when the recovery filter unit region is not predicted using directional intra-frame mode or intra-frame prediction mode, the signal class of the recovery filter unit can be determined based on neighboring blocks. A default signal class can be selected based on whether neighboring blocks are encoded in directional intra-frame mode or whether their directions are inconsistent, or explicit signaling of the signal class index or filter coefficients can be executed. For example, if neighboring blocks are not encoded in directional intra-frame mode or their directions are inconsistent, a default signal class can be selected, or explicit signaling of the signal class index or filter coefficients can be executed.

[0207] According to some implementations, a directional majority vote or consensus check, including in the recovery filter unit, can be performed to determine the filter category.

[0208] In one implementation, when the size and shape of the recovery filter unit are fixed and the number of directional information elements included in the recovery filter unit (e.g., 8×8 in the case of CDEF) is greater than a predefined number, a majority vote or consensus check of the directionality included in the recovery filter unit can be performed to determine the filter category. The directional information elements can have varying sizes and different directional intra-prediction modes.

[0209] In one implementation, for a majority vote, the most frequent direction among the available and potentially merged directions can be selected. In the example, a margin can be set between the first and second most frequent directions.

[0210] In one implementation, before a majority vote, it is determined whether the number of categories within the recovery filter unit is greater than a predefined number. If true, inconsistency can be declared, and explicit signaling for the recovery filter unit or a smaller recovery filter unit can be selected.

[0211] According to some implementations, the recovery filtering unit in the chroma component can use such information when its own directional information from the directional intra-prediction mode can be used alone for the chroma component.

[0212] According to aspects of this disclosure, when directional information is available from both the intra-prediction mode and CDEF processing, the directional pattern in the intra-prediction mode, together with the directional information from the CDEF processing, can be used as a guide for identifying and classifying recovery filter units (e.g., Wiener filter units or SGRPRJ filter units).

[0213] In one implementation, when directional information is available from both intra-frame predicted orientation and CDEF processing, a mapping of orientation from both sources can be introduced when checking the consistency of the directional information. Classification-based recovery filtering can only be performed if the orientation information is consistent. Figure 15 An example is shown below: where the intra-predicted orientations corresponding to the 7 incremental angles associated with one of the 8 base angles in AV1 can be mapped to a single directionality category, and thus can form a one-to-one correspondence with the 8 directions derived during CDEF processing.

[0214] In one implementation, directional information from one of the intra-prediction modes or the CDEF process can be used as a first source. The other of the intra-prediction modes or the CDEF process can only be used if a consistent direction cannot be determined using directional information from the first source.

[0215] Note that in some examples, such as AV1, the Wiener filter and the SGRPRJ filter can be adaptively selected for each LRU. For both filters, the filter parameters and filtering process can be performed in the same manner for each LRU. The difference may be the type and number of filter parameters.

[0216] In one implementation, when performing SGRPRJ filtering, the search window size r and noise parameter e can be defined for each signal category in the same manner as the Wiener filter parameters.

[0217] In one implementation, when performing SGRPRJ filtering, projection parameters α and β can be defined and used for each signal category in the same manner as Wiener filter parameters.

[0218] In one implementation, when performing SGRPRJ filtering, the search window size r, noise parameter e, and projection parameters α and β can be defined and used for each signal category in the same manner as the Wiener filter parameters.

[0219] V. Flowchart

[0220] Figure 16 A flowchart outlining an exemplary process (1600) according to an embodiment of this disclosure is shown. In various embodiments, the process (1600) is executed by a processing circuitry system, such as the processing circuitry system in terminal devices (210), (220), (230), and (240), a processing circuitry system performing the function of a video encoder (303), a processing circuitry system performing the function of a video decoder (310), a processing circuitry system performing the function of a video decoder (410), a processing circuitry system performing the function of an intra-frame prediction module (452), a processing circuitry system performing the function of a video encoder (503), a processing circuitry system performing the function of a predictor (535), a processing circuitry system performing the function of an intra-frame encoder (622), a processing circuitry system performing the function of an intra-frame decoder (772), etc. In some embodiments, the process (1600) is implemented as software instructions, so that the processing circuitry system executes the process (1600) when the processing circuitry system executes the software instructions.

[0221] The process (1600) can typically begin at step (S1610), where the process (1600) determines the directionality information of the recovery filter unit included in the video frame based on at least one of CDEF processing or intra-frame prediction mode. Then, the process (1600) proceeds to step (S1620).

[0222] At step (S1620), process (1600) determines one of a set of multiple filter parameters for the recovery filter processing based on the directional information of the recovery filter unit. Then, process (1600) proceeds to step (S1630).

[0223] At step (S1630), process (1600) performs recovery filter processing on the recovery filter unit based on one of a plurality of filter parameter sets. Then, process (1600) proceeds to step (S1640).

[0224] At step (S1640), process (1600) reconstructs the video frame based on the filtered recovery filter unit. Then, process (1600) terminates.

[0225] In an implementation, the recovery filter unit includes one or more directional information units, and performs at least one of CDEF processing or intra-frame prediction mode on one of the directional information units.

[0226] In an implementation, each of the multiple filter parameter sets of the recovery filter is associated with at least one directionality of the CDEF process.

[0227] In the implementation, process (1600) determines one of a set of multiple filter parameters for the recovery filter processing based on the block variance information and directionality information of the recovery filter unit.

[0228] In the implementation, process (1600) determines one of a set of multiple filter parameters for the recovery filter process based on the directional information of the recovery filter unit and the filter strength of the CDEF process.

[0229] In an implementation, process (1600) determines the directionality information of the recovery filter unit based on at least one of a majority vote or a consensus check on the directionality in the recovery filter unit.

[0230] In the implementation, based on the fact that the recovery filter unit is not intra-coded and the adjacent blocks of the recovery filter unit are intra-coded, the process (1600) determines the directionality information of the recovery filter unit based on the intra-prediction mode performed on the adjacent blocks.

[0231] In the implementation, process (1600) performs recovery filter processing on the recovery filter unit based on matching the directional information determined according to the CDEF processing with the directional information determined according to the intra-frame prediction mode.

[0232] In the implementation, the recovery filter processing is one of the Wiener filter processing and the SGRPRJ filter processing.

[0233] In an implementation, process (1600) determines one of a plurality of filter parameter sets for the recovery filter processing based on a default set of filter parameters, an index of an indicator set of filter parameters indicated by a signal, or a set of filter parameters indicated by a signal.

[0234] VI. Computer System

[0235] The above-described technology can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 17 A computer system (1700) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0236] Computer software can be coded using any suitable machine code or computer language, which can be subjected to mechanisms such as assembly, compilation, and linking to create code including instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or through interpretation, microcode execution, etc.

[0237] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0238] Figure 17 The components shown for the computer system (1700) are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any one or a combination of components shown in the exemplary embodiments of the computer system (1700).

[0239] The computer system (1700) may include certain human-machine interface input devices. Such human-machine interface input devices can respond to input from one or more human users via, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, tapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface device can also be used to capture certain media that are not necessarily directly related to intentional human input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from still image capturing devices), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0240] Input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1701), mouse (1702), trackpad (1703), touch screen (1710), data glove (not shown), joystick (1705), microphone (1706), scanner (1707), and camera (1708).

[0241] The computer system (1700) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback via a touchscreen (1710), data gloves (not shown), or joystick (1705), but tactile feedback devices that are not used as input devices may also exist); audio output devices (e.g., speakers (1709), headphones (not depicted)); visual output devices (e.g., screens (1710), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability—some of which may be able to output two-dimensional or more than three-dimensional visual output in a manner such as stereoscopic image output; virtual reality glasses (not depicted); holographic displays and ashtrays (not depicted)); and printers (not depicted). These visual output devices (e.g., screens (1710)) can be connected to the system bus (1748) via a graphics adapter (1750).

[0242] The computer system (1700) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1720) having media such as CD / DVD (1721), thumb drives (1722), removable hard disk drives or solid-state drives (1723), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0243] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0244] The computer system (1700) may also include a network interface (1754) for one or more communication networks (1755). The one or more communication networks (1755) may be, for example, wireless, wired, or optical. The one or more communication networks (1755) may also be local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), vehicle and industrial networks, real-time networks, latency-tolerant networks, etc. Examples of one or more communication networks (1755) include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANbus, etc. Some networks typically require external network interface adapters that attach to certain general-purpose data ports or peripheral buses (1749) (e.g., the USB port of the computer system (1700)); other networks are typically integrated into the core of the computer system (1700) via system buses as described below (e.g., integrated into a PC computer system via an Ethernet interface, or integrated into a smartphone computer system via a cellular network interface). The computer system (1700) can communicate with other entities through any of these networks. Such communication can be one-way receive-only (e.g., broadcast television), one-way send-only (e.g., to a CANbus of a certain CANbus device), or bidirectional, such as to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0245] The human-machine interface devices, human-accessible storage devices and network interfaces mentioned above can be attached to the core (1740) of the computer system (1700).

[0246] The core (1740) may include one or more central processing units (CPUs) (1741), graphics processing units (GPUs) (1742), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1743), hardware accelerators (1744) for certain tasks, graphics adapters (1750), etc. These devices, along with read-only memory (ROMs) (1745), random access memory (1746), and internal mass storage devices (1747) such as internal non-user-accessible hard disk drives, SSDs, etc., can be connected via the system bus (1748). In some computer systems, the system bus (1748) can be accessed in the form of one or more physical connectors to allow for expansion via additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1749) to the core's system bus (1748). In the example, a screen (1710) may be connected to a graphics adapter (1750). Peripheral bus architectures include PCI, USB, etc.

[0247] The CPU (1741), GPU (1742), FPGA (1743), and accelerator (1744) can execute certain instructions, which can be combined to form the computer code mentioned above. This computer code can be stored in ROM (1745) or RAM (1746). Transient data can also be stored in RAM (1746), while permanent data can be stored, for example, in an internal mass storage device (1747). Fast storage and retrieval of any storage device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1741), GPUs (1742), mass storage devices (1747), ROMs (1745), RAMs (1746), etc.

[0248] Computer-readable media may contain computer code for performing operations of various computer implementations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0249] By way of example and not limitation, a computer system (1700) with an architecture, particularly a core (1740), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage devices as described above, as well as certain storage devices of the core (1740) having non-transitory characteristics, such as mass storage devices (1747) or ROM (1745) within the core. Software implementing various embodiments of this disclosure can be stored in such devices and executed by the core (1740). Depending on specific needs, the computer-readable media may include one or more memory devices or chips. The software can cause the core (1740)—particularly the processors therein (including CPUs, GPUs, FPGAs, etc.)—to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1746) and modifying such data structures according to the processes defined by the software. Alternatively or as an alternative, a computer system may be functionalized by logic hardwired or otherwise embodied in circuitry (e.g., an accelerator (1744)), which may operate in place of or with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and conversely, references to logic may include software. Where appropriate, references to a computer-readable medium may include circuitry (e.g., an integrated circuit, IC) storing software for execution, circuitry implementing logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0250] While this disclosure has described several exemplary embodiments, variations, substitutions, and various alternative equivalents fall within the scope of this disclosure. It will therefore be appreciated that, although not expressly shown or described herein, those skilled in the art will be able to conceive of many systems and methods that implement the principles of this disclosure and thus within its spirit and scope.

[0251] Appendix A: Acronyms

[0252] ALF: Adaptive Loop Filter

[0253] AMVP: Advanced Motion Vector Prediction

[0254] APS: Adaptive Parameter Set

[0255] ASIC: Application-Specific Integrated Circuit

[0256] ATMVP: Alternative / Advanced Temporal Motion Vector Prediction

[0257] AV1: AOMedia Video 1

[0258] AV2: AOMedia Video 2

[0259] BMS: Benchmark Set

[0260] BV: Block Vector

[0261] CANBus: Controller Area Network Bus

[0262] CB: Coded Block

[0263] CC-ALF: Cross-component adaptive loop filter

[0264] CD: Compact Disc

[0265] CDEF: Constrained Directional Enhancement Filter

[0266] CPR: Current image reference

[0267] CPU: Central Processing Unit

[0268] CRT: Cathode Ray Tube

[0269] CTB: Coded Tree Block

[0270] CTU: Coding Tree Unit

[0271] CU: Encoding Unit

[0272] DPB: Decoder Image Buffer

[0273] DPCM: Differential Pulse Code Modulation

[0274] DPS: Decoding Parameter Set

[0275] DVD: Digital Video Disc

[0276] FPGA: Field Programmable Gate Domain

[0277] JCCR: Joint CbCr Residual Encoding

[0278] JVET: Joint Video Exploration Team

[0279] GOP: Image Group

[0280] GPU: Graphics Processing Unit

[0281] GSM: Global System for Mobile Communications

[0282] HDR: High Dynamic Range

[0283] HEVC: High-efficiency video encoding and decoding

[0284] HRD: Hypothetical Reference Decoder

[0285] IBC: Intra-Block Copy

[0286] IC: Integrated Circuit

[0287] ISP: Intra-Frame Sub-Partition

[0288] JEM: Joint Development Model

[0289] LAN: Local Area Network

[0290] LCD: Liquid Crystal Display

[0291] LR: Loop Recovery Filter

[0292] LRU: Loop Recovery Unit

[0293] LTE: Long Term Evolution

[0294] MPM: Most Likely Pattern

[0295] MV: Motion Vector

[0296] OLED: Organic Light Emitting Diode

[0297] PB: Prediction Block

[0298] PCI: Peripheral Component Interconnect

[0299] PDPC: Location-Related Prediction Combination

[0300] PLD: Programmable Logic Device

[0301] PPS: Image Parameter Set

[0302] PU: Prediction Unit

[0303] RAM: Random Access Memory

[0304] ROM: Read-Only Memory

[0305] SAO: Sample Adaptive Offset

[0306] SCC: Screen Content Encoding

[0307] SDR: Standard Dynamic Range

[0308] SEI: Supplemental Enhancement Information

[0309] SNR: Signal-to-noise ratio

[0310] SPS: Sequence Parameter Set

[0311] SSD: Solid State Drive

[0312] TU: Transformation Unit

[0313] USB: Universal Serial Bus

[0314] VPS: Video Parameter Set

[0315] VUI: Video Availability Information

[0316] VVC: Multi-functional Video Coding

[0317] WAIP: Wide-angle Intra-frame Prediction

Claims

1. A method of video decoding, the method comprising: comprises: determining directional information of a recovery filter unit included in a video frame based on at least one of a previously performed constrained directional enhanced filter (CDEF) process or an intra prediction mode; determining one of a plurality of filter parameter sets of a recovery filter process based on the directional information of the recovery filter unit; performing at least one of a Wiener filter process or a self-guided projection (SGRPJ) filter process as the recovery filter process on the recovery filter unit based on the one of the plurality of filter parameter sets after completion of at least one of the previously performed CDEF process or the intra prediction mode; and reconstructing the video frame based on the recovery filter unit subjected to the recovery filter process, wherein the recovery filter unit comprises one or more directional information units and at least one of the CDEF process or the intra prediction mode is performed on one of the one or more directional information units.

2. The method of claim 1, wherein, Each of the plurality of filter parameter sets is associated with at least one directionality of the CDEF process.

3. The method of claim 1, wherein, Determining the one of the plurality of filter parameter sets comprises determining the one of the plurality of filter parameter sets of the recovery filter process based on block variance information of the recovery filter unit and the directional information.

4. The method of claim 1, wherein, Determining the one of the plurality of filter parameter sets comprises determining the one of the plurality of filter parameter sets of the recovery filter process based on the directional information of the recovery filter unit and filter strength of the CDEF process.

5. The method of claim 1, wherein, Determining the one of the plurality of filter parameter sets comprises determining the one of the plurality of filter parameter sets of the recovery filter process based on a default filter parameter set, a signaled index indicating a filter parameter set, and one of the signaled filter parameter sets.

6. The method of claim 1, wherein, Determining the directional information of the recovery filter unit comprises determining the directional information of the recovery filter unit based on at least one of a majority vote on directionality in the recovery filter unit or a consistency check.

7. The method of claim 1, wherein, Determining the directional information of the recovery filter unit comprises determining the directional information of the recovery filter unit based on an intra prediction mode performed on a neighboring block of the recovery filter unit in response to the recovery filter unit not being intra coded and the neighboring block of the recovery filter unit being intra coded.

8. The method of claim 1, wherein, Performing the recovery filter process on the recovery filter unit based on the one of the plurality of filter parameter sets comprises performing the recovery filter process on the recovery filter unit based on the directional information determined according to the CDEF process matching the directional information determined according to the intra prediction mode.

9. A method of video encoding, characterized by, comprises: determining directional information of a recovery filter unit included in a video frame based on at least one of a previously performed constrained directional enhanced filter (CDEF) process or an intra prediction mode; determine one of a plurality of filter parameter sets for a restoration filter process based on directional information of the restoration filter unit; perform at least one of a Wiener filter process or a self-guided projection, SGRPRJ, filter process as the restoration filter process on the restoration filter unit based on the one of the plurality of filter parameter sets after completion of at least one of a previously performed constrained directional enhanced filter, CDEF, process or an intra prediction mode; and reconstruct the video frame based on the restoration filter unit subjected to the restoration filter process, wherein the restoration filter unit comprises one or more directional information units and at least one of the CDEF process or the intra prediction mode is performed on one of the one or more directional information units.

10. A video decoding apparatus, comprising: comprising: a direction determination module configured to determine directional information of a restoration filter unit included in a video frame based on at least one of a previously performed constrained directional enhanced filter, CDEF, process or an intra prediction mode; a parameter determination module configured to determine one of a plurality of filter parameter sets for a restoration filter process based on directional information of the restoration filter unit; a restoration filtering module configured to perform at least one of a Wiener filter process or a self-guided projection, SGRPRJ, filter process as the restoration filter process on the restoration filter unit based on the one of the plurality of filter parameter sets after completion of at least one of a previously performed CDEF process or an intra prediction mode; and a reconstruction module configured to reconstruct the video frame based on the restoration filter unit subjected to the restoration filter process, wherein the restoration filter unit comprises one or more directional information units and at least one of the CDEF process or the intra prediction mode is performed on one of the one or more directional information units.

11. The apparatus of claim 10, wherein, each of the plurality of filter parameter sets is associated with at least one directionality of the CDEF process.

12. The apparatus of claim 10, wherein, the parameter determination module is configured to: determine the one of the plurality of filter parameter sets for the restoration filter process based on block variance information and the directional information of the restoration filter unit.

13. The apparatus of claim 10, wherein, the parameter determination module is configured to: determine the one of the plurality of filter parameter sets for the restoration filter process based on the directional information of the restoration filter unit and filter strength of the CDEF process.

14. The apparatus of claim 10, wherein, the direction determination module is configured to: determine the directional information of the restoration filter unit based on at least one of a majority vote or a consensus check on directionality in the restoration filter unit.

15. The apparatus of claim 10, wherein, the direction determination module is configured to: in response to the restoration filter unit not being intra coded and neighboring blocks of the restoration filter unit being intra coded, determine the directional information of the restoration filter unit based on an intra prediction mode performed on the neighboring blocks of the restoration filter unit.

16. The apparatus of claim 10, wherein, the restoration filtering module is configured to: performing the recovery filter process on the recovery filter unit based on the directionality information determined from the CDEF process matching the directionality information determined from the intra prediction mode.

17. An electronic device comprising a processor and a memory having stored thereon computer instructions, wherein, The computer instructions, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1-9.

18. A non-transitory computer-readable storage medium storing instructions, the instructions comprising: The instructions, when executed by at least one processor, cause the at least one processor to perform the method according to any one of claims 1-9.

19. A method of storing a bitstream, the method comprising: performing the video coding method of claim 9 to generate a bitstream; and storing the bitstream.

20. A method of transmitting a bitstream, the method comprising: performing the video coding method of claim 9 to generate a bitstream; and transmitting the bitstream.

21. A computer readable storage medium having stored thereon computer programs / instructions and a bitstream, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the video coding method of claim 9 to generate the bitstream.

Citation Information

Patent Citations

  • Constrained directional enhancement filter selection for video coding

    US20190045186A1

  • Guided restoration of video data using neural networks

    US20200184603A1