Method and device for video coding

KR103014933B1Active Publication Date: 2026-09-04TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020227028504
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-31
Filing Date
2021-09-08
Publication Date
2026-09-04
Estimated Expiration
2041-09-08

Smart Images

  • Figure R1020227028504_ABST
    Figure R1020227028504_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods and apparatuses for video processing. In some examples, the apparatus for video processing includes a processing circuit. The processing circuit converts a picture in a subsampled format in color space to a non-subsampled format in color space. Then, the processing circuit clips the values ​​of the color components of the picture in the non-subsampled format before providing the picture in the non-subsampled format as input to a neural network-based filter.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present application claims the benefit of priority to U.S. Provisional Application No. 63 / 131,656, filed December 29, 2020, “APPLICATION OF CLIPPING TO IMPROVE PRE-PROCESSING IN A NEURAL NETWORK BASED IN-LOOP FILTER IN A VIDEO CODEC,” and the benefit of priority to U.S. Patent Application No. 17 / 463,352, filed August 31, 2021, “METHOD AND APPARATUS FOR VIDEO CODING.” The entire disclosures of the prior applications are incorporated herein by reference in their entirety.

[0002] The present disclosure describes embodiments generally related to video coding. More specifically, the present disclosure provides techniques for improving neural network-based in-loop filters. Background Technology

[0003] The background description provided herein is intended to provide general context for the present disclosure. The work of the inventors currently named—within the scope described in this background section—as well as modes of description that may not otherwise qualify as prior art at the time of filing, are not recognized as prior art to the present disclosure, either explicitly or implicitly.

[0004] Video coding and decoding can be performed using inter-picture prediction along with motion compensation. Uncompressed digital video may contain a series of pictures, each having, for example, a spatial dimension of 1920x1080 luminance samples and associated chrominance samples. This series of pictures may have a fixed or variable picture rate (informally also known as the frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video at 8 bits per sample (1920x1080 luminance sample resolution at a 60 Hz frame rate) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires more than 600 GBytes of storage space.

[0005] One objective of video coding and decoding may be to reduce redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage requirements by more than two orders of magnitude in some cases. Both lossless and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for the intended application. In the case of video, lossy compression is widely used. The amount of acceptable distortion depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. Achievable compression ratios may reflect the following: higher acceptable / tolerable distortion can yield a higher compression ratio.

[0006] Video encoders and decoders can utilize techniques from a wide range of categories, including, for example, motion compensation, transformation, quantization, and entropy coding.

[0007] Video codec technologies may include techniques known as intra-coding. In intra-coding, sample values ​​are represented without referencing samples from previously reconstructed reference pictures or other data. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, the picture may be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture or a still image in a coded video bitstream and video session. Samples in an intra block can be exposed to transformation, and transformation coefficients can be quantized before entropy coding. Intra-prediction may be a technique that minimizes sample values ​​in a pre-transform domain. In some cases, the smaller the DC value after conversion and the smaller the AC coefficients, the fewer bits are required at a given quantization step size to represent the block after entropy coding.

[0008] For example, traditional intra-coding, such as that known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include methods that attempt to use, for example, surrounding sample data and / or metadata acquired during the encoding / decoding of blocks of data that are spatially adjacent and precede in the decoding order. These techniques are hereinafter referred to as "intra-prediction" techniques. Note that in at least some cases, intra-prediction uses only reference data from the current picture being reconstructed, rather than from reference pictures.

[0009] There may be many different forms of intra-prediction. When more than one of these techniques can be used in a given video coding technique, the technique in use may be coded in an intra-prediction mode. In certain cases, modes may have submodes and / or parameters, which may be coded individually or included in a mode codeword. The codeword used for a given mode / submode / parameter combination can affect the coding efficiency gain through intra-prediction, and the same may apply to the entropy coding technique used to convert the codewords into a bitstream.

[0010] Specific modes of intra-prediction were introduced in H.264, improved in H.265, and further enhanced in newer coding techniques such as JEM (joint exploration model), VVC (versatile video coding), and BMS (benchmark set). A predictor block can be formed using neighbor sample values ​​belonging to already available samples. The sample values ​​of neighbor samples are copied into the predictor block according to direction. A reference to the direction in use can be encoded in the bitstream or predicted itself.

[0011] Referring to FIG. 1a, a subset of 9 known predictor directions from the 33 possible predictor directions of H.265 (corresponding to 33 angle modes of 35 intra modes) is depicted in the lower right. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample is being predicted. For example, arrow (102) indicates that the sample (101) is predicted to the upper right from the sample or samples at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that the sample (101) is predicted to the lower left from the sample or samples at an angle of 22.5 degrees from the horizontal.

[0012] Referring still to FIG. 1a, a square block (104) of 4x4 samples (indicated by dashed and bold lines) is depicted in the upper left. The square block (104) contains 16 samples, each labeled "S", its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block is 4x4 samples in size, S44 is located in the lower right. Reference samples following a similar numbering scheme are additionally depicted. The reference samples are labeled R for the block (104), its Y position (e.g., row index) and X position (column index). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed; therefore, negative values ​​do not need to be used.

[0013] Intra-picture prediction can be performed by appropriately copying reference sample values ​​from neighboring samples according to the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating a prediction direction corresponding to the arrow (102) for this block—that is, samples are predicted from the prediction sample or samples upward to the right, at a 45-degree angle from the horizontal. In that case, samples S41, S32, S23 and S14 are predicted from the same reference sample R05. Then sample S44 is predicted from reference sample R08.

[0014] In certain cases, particularly when directions cannot be uniformly divided into 45 degrees, the values ​​of multiple reference samples can be combined, for example, through interpolation to calculate a reference sample.

[0015] As video coding technology has developed, the number of possible directions has increased. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and at the time of this disclosure, JEM / VVC / BMS can support up to 65 directions. Experiments have been performed to identify the most likely directions, and specific techniques in entropy coding are used to represent such likely directions with a small number of bits, while accepting specific penalties for less likely directions. Additionally, the directions themselves can sometimes be predicted from neighboring directions used in adjacent already decoded blocks.

[0016] FIG. 1b illustrates a schematic diagram (180) depicting 65 intra-predicted directions according to JEM to illustrate a number of predicted directions that increase over time.

[0017] The mapping of intra-predicted direction bits within a coded video bitstream representing direction can vary depending on the video coding technique; for example, it can range from a simple direct mapping of the predicted direction to complex adaptive schemes involving intra-predicted modes, codewords, and most likely modes, as well as similar techniques. However, in all cases, there may be certain directions that are statistically less likely to occur in the video content than other directions. Since the goal of video compression is to reduce redundancy, in a well-functioning video coding technique, such less likely directions will be represented by a greater number of bits than more likely directions.

[0018] Motion compensation may be a lossy compression technique and may relate to techniques used for predicting a newly reconstructed picture or part of a picture, after a block of sample data from a previously reconstructed picture or part thereof (reference picture) has been spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be identical to the picture currently being reconstructed. MVs may have 2-dimensional X and Y, or 3 dimensions, with the third being an indication of the reference picture in use (the latter may, indirectly, be a time dimension).

[0019] In some video compression techniques, applicable MVs for a specific region of sample data can be predicted from other MVs, for example, from other regions of sample data spatially adjacent to the region being reconstructed, and from MVs that precede such MVs in the decoding order. By doing so, the amount of data required to code the MVs can be substantially reduced, thereby eliminating redundancy and increasing compression. MV prediction can be performed effectively, for example, when coding an input video signal derived from a camera (known as natural video), because there is a statistical probability that regions larger than the region where a single MV is applicable move in similar directions, and thus, in some cases, can be predicted using similar motion vectors derived from the MVs of neighboring regions. As a result, the MV found for a given region becomes similar or identical to the MV predicted from surrounding MVs, and it can ultimately be represented after entropy coding with fewer bits than would be used when directly coding the MVs. In some cases, MV prediction can be an example of lossless compression of signals (i.e., MVs) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when computing a predictor from multiple surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec.H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, a technique referred to hereafter as "spatial merge" is described in this specification.

[0021] Referring to FIG. 2, the current block (201) contains samples discovered by the encoder during the motion search process, which are predictable from a previous block of the same size that is spatially shifted. Instead of directly coding the MV, the MV may be derived from metadata associated with one or more reference pictures, for example, from the most recent (in decoding order) reference picture, using an MV associated with any one of five surrounding samples denoted as A0, A1, and B0, B1, B2 (202 to 206, respectively). In H.265, the MV prediction may use predictors from the same reference picture used by neighboring blocks. The problem to be solved

[0022] Aspects of the present disclosure provide methods and apparatuses for video processing. In some examples, the apparatus for video processing includes a processing circuit. The processing circuit converts a picture in a subsampled format in color space to a non-subsampled format in color space. Then, the processing circuit clips the values ​​of the color components of the picture in the non-subsampled format before providing the picture in the non-subsampled format as an input to a neural network-based filter. means of solving the problem

[0023] In some examples, the processing circuit clips the values ​​of the color components of a picture in an unsubsampled format into a valid range for the color components. In one example, the processing circuit clips the values ​​of the color components of a picture in an unsubsampled format into a range determined based on the bit depth. In another example, the processing circuit clips the values ​​of the color components of a picture in an unsubsampled format into a predetermined range.

[0024] In some examples, the processing circuit determines a range for clipping values ​​based on decoded information from a bitstream carrying a picture, and then clips the values ​​of the color components of the picture in an unsubsampled format into the determined range. In an example, the processing circuit decodes a signal indicating the range from at least one of a sequence parameter set, a picture parameter set, a slice header, and a tile header within the bitstream.

[0025] In some examples, the processing circuit reconstructs a picture in a subsampled format based on decoded information from a bitstream and applies a deblocking filter to the picture in the subsampled format. In some examples, the processing circuit applies a neural network-based filter to a picture in an unsubsampled format with clipped values ​​to generate a filtered picture in an unsubsampled format and converts the filtered picture in an unsubsampled format into a filtered picture in a subsampled format.

[0026] In some examples, a picture in an unsubsampled format with clipped values ​​is stored in storage. Then, the stored picture in an unsubsampled format with clipped values ​​can be provided as a training input to train a neural network in a neural network-based filter.

[0027] Aspects of the present disclosure also provide a non-transient computer-readable medium storing instructions that cause a computer to perform a method for video processing when executed by a computer for video decoding. Brief explanation of the drawing

[0028] Additional features, nature, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. Figure 1a is a schematic example of an exemplary subset of intra prediction modes. Figure 1b is an example of exemplary intra-predicted directions. Figure 2 is a schematic example of the current block and the surrounding space merge candidates in one example. FIG. 3 is a schematic example of a simplified block diagram of a communication system (300) according to an embodiment. FIG. 4 is a schematic example of a simplified block diagram of a communication system (400) according to an embodiment. FIG. 5 is a schematic example of a simplified block diagram of a decoder according to an embodiment. FIG. 6 is a schematic example of a simplified block diagram of an encoder according to an embodiment. FIG. 7 illustrates a block diagram of an encoder according to another embodiment. FIG. 8 illustrates a block diagram of a decoder according to another embodiment. Figure 9 shows a block diagram of a loop filter unit in some examples. Figure 10 shows a block diagram of a different loop filter unit in some examples. Figure 11 illustrates block diagrams of neural network-based filters in some examples. Figure 12 shows a block diagram of a preprocessing module in some examples. Figure 13 illustrates block diagrams of neural network structures in some examples. Figure 14 shows a block diagram of a dense residual unit. Figure 15 shows a block diagram of a post-processing module in some examples. Figure 16 shows a block diagram of a preprocessing module in some examples. Figure 17 illustrates a flowchart outlining an example of a process. FIG. 18 is a schematic example of a computer system according to an embodiment. Specific details for implementing the invention

[0029] FIG. 3 illustrates a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, through a network (350). For example, the communication system (300) includes a first pair of terminal devices (310 and 320) interconnected through the network (350). In the example of FIG. 3, the first pair of terminal devices (310 and 320) perform unidirectional transmission of data. For example, a terminal device (310) may encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) through the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) receives coded video data from the network (350), decodes the coded video data to recover video pictures, and can display the video pictures according to the recovered video data. Unidirectional data transmission may be common in media serving applications, etc.

[0030] In another example, the communication system (300) includes a second pair of terminal devices (330 and 340) that perform bidirectional transmission of coded video data that may occur, for example, during videoconferencing. For bidirectional transmission of data, in the example, each terminal device among the terminal devices (330 and 340) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to another terminal device among the terminal devices (330 and 340) via the network (350). Each terminal device among the terminal devices (330 and 340) may also receive coded video data transmitted by the other terminal device among the terminal devices (330 and 340), decode the coded video data to recover video pictures, and display the video pictures on an accessible display device according to the recovered video data.

[0031] In the example of FIG. 3, the terminal devices (310, 320, 330, and 340) may be exemplified as servers, personal computers, and smartphones, but the principles of the present disclosure are not so limited. Embodiments of the present disclosure find applications with laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that transmit coded video data between the terminal devices (310, 320, 330, and 340), including, for example, wireline and / or wireless communication networks. The communication network (350) may exchange data on circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (350) may not be important to the operation of the present disclosure unless described below in the specification.

[0032] FIG. 4 illustrates the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject. The disclosed subject may be equally applicable to other video-enabled applications, such as video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0033] A streaming system may include a capture subsystem (413) that may include, for example, a digital camera, a video source (401) that generates a stream (402) of uncompressed video pictures. In the example, the stream (402) of video pictures includes samples captured by the digital camera. The stream (402) of video pictures, depicted in bold to emphasize the large data volume compared to encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement the embodiments of the disclosed subject matter as described in more detail below. Encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize the small data volume compared to the stream (402) of video pictures, may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406 and 408) in FIG. 4, may access the streaming server (405) to retrieve copies (407 and 409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an incoming copy (407) of the encoded video data and generates an outgoing stream (411) of video pictures that can be rendered on a display (412) (e.g., a display screen) or another rendering device (not depicted). In some streaming systems, encoded video data (404, 407, and 409) (e.g., video bitstreams) may be encoded according to specific video coding / compression standards.Examples of such standards include ITU-T Recommendation H.265. In this example, the video coding standard under development is informally known as VVC (Versatile Video Coding). The disclosed subject may be used in the context of VVC.

[0034] Note that the electronic devices (420 and 430) may include other components (not shown). For example, the electronic device (420) may also include a video decoder (not shown) and the electronic device (430) may also include a video encoder (not shown).

[0035] FIG. 5 illustrates a block diagram of a video decoder (510) according to an embodiment of the present disclosure. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used instead of the video decoder (410) in the example of FIG. 4.

[0036] A receiver (531) may receive one or more coded video sequences to be decoded by a video decoder (510)—in the same or other embodiments, one coded video sequence at a time—wherein the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (531) may receive the encoded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which may be forwarded to their respective use entities (not described). The receiver (531) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (515) may be combined between the receiver (531) and the entropy decoder / parser (520) (hereinafter "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In others, it may be outside the video decoder (510) (not described). In yet others, for example, to prevent network jitter, there may be a buffer memory (not described) outside the video decoder (510), and additionally, for example to handle playout timing, there may be another buffer memory (515) inside the video decoder (510). When the receiver (531) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be needed or may be small.For use on best effort packet networks such as the Internet, a buffer memory (515) may be required, may be relatively large, advantageously may be of adaptive size, and may be implemented at least partially in an operating system or similar elements (not described) outside the video decoder (510).

[0037] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from a coded video sequence. Categories of these symbols include information used to manage the operation of the video decoder (510), and potentially, as illustrated in FIG. 5, information for controlling a rendering device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but can be coupled to the electronic device (530). The control information for the rendering device(s) may be in the form of SEI messages (Supplemental Enhancement Information) or VUI (Video Usability Information) parameter set fragments (not depicted). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be in accordance with video coding techniques or standards and may follow various principles including variable-length coding, Huffman coding, and arithmetic coding with or without context sensitivity. The parser (520) may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to a group. The subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (520) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.

[0038] The parser (520) can generate symbols (521) by performing an entropy decoding / parsing operation on a video sequence received from a buffer memory (515).

[0039] The reconstruction of the symbols (521) may involve a number of different units depending on the type of the coded video picture or parts thereof (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors. Which units are involved and how they are can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the number of units below is not described for clarity.

[0040] In addition to the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units may interact closely with one another and be at least partially integrated with one another. However, for the sake of illustrating the subject matter disclosed, the conceptual subdivision into the functional units below is appropriate.

[0041] The first unit is a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantized transform coefficients as well as control information, including which transform to use, block size, quantization factor, quantization scaling matrices, etc., as symbol(s) (521) from the parser (520). The scaler / inverse transform unit (551) can output blocks containing sample values ​​that can be input to an aggregator (555).

[0042] In some cases, the output samples of the scaler / inverse transform (551) may be related to an intra-coded block—that is, a block that can use prediction information from previously reconstructed parts of the current picture rather than prediction information from previously reconstructed pictures. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses surrounding already reconstructed information fetched from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (555) adds the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.

[0043] In other cases, the output samples of the scaler / inverse unit (551) may be intercoded and potentially associated with a motion-compensated block. In such cases, the motion compensation prediction unit (553) may access the reference picture memory (557) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (521) associated with the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse unit (551) (in this case, called residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (557) where the motion compensation prediction unit (553) fetches prediction samples may be controlled by motion vectors available to the motion compensation prediction unit (553) in the form of symbols (521) that may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​such as those fetched from the reference picture memory (557) when subsample accurate motion vectors are in use, motion vector prediction mechanisms, etc.

[0044] The output samples of the aggregator (555) may be subject to various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filter technologies that are made available to the loop filter unit (556) as symbols (521) from the parser (520) and controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream), but may be responsive to meta-information acquired during the decoding of previous (in the decoding order) parts of the coded picture or coded video sequence, as well as responsive to previously reconstructed and loop-filtered sample values.

[0045] The output of the loop filter unit (556) may be a sample stream that is not only output to the render device (512) but may also be stored in the reference picture memory (557) for use in future inter-picture predictions.

[0046] Certain coded pictures, when fully reconstructed, can be used as reference pictures for future prediction. For example, when a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before initiating the reconstruction of the next coded picture.

[0047] A video decoder (510) can perform decoding operations according to a predetermined video compression technique in a standard such as ITU-T Rec.H.265. In that the coded video sequence adheres to both the syntax of the video compression technique or standard, or profiles documented in the video compression technique or standard, the coded video sequence may comply with the syntax specified by the video compression technique or standard in use. Specifically, the profile may select specific tools from all tools available in the video compression technique or standard as the only tools available for use under that profile. Additionally, for compliance, the complexity of the coded video sequence may be within boundaries defined by the levels of the video compression technique or standard. In some cases, the levels limit the maximum picture size, maximum frame rate, maximum recovery sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the levels may, in some cases, be further restricted through HRD (Hypothetical Reference Decoder) specifications and metadata for managing HRD buffers signaled in the coded video sequence.

[0048] In an embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form, for example, time, space, or signal-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0049] FIG. 6 illustrates a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) may be used instead of the video encoder (403) in the example of FIG. 4.

[0050] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) capable of capturing video image(s) to be encoded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0051] A video source (601) may provide a source video sequence to be coded by a video encoder (603) in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that impart motion when viewed sequentially. The pictures themselves may be organized as a spatial array of pixels, and each pixel may contain one or more samples depending on the sampling structure, color space, etc. being used. A person skilled in the art can easily understand the relationship between the pixels and the samples. The following explanation focuses on the samples.

[0052] According to an embodiment, the video encoder (603) can code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints, such as those required by the application. Enforcing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls other function units and is functionally coupled to other function units as described below. The coupling is not described for clarity. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda values ​​of rate-distortion optimization techniques, ...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other suitable functions related to the video encoder (603) optimized for a specific system design.

[0053] In some embodiments, the video encoder (603) is configured to operate in a coding loop. For the sake of oversimplification, in the example, the coding loop may include a source coder (630) (responsible for generating symbols, such as a symbol stream, based on, for example, the input picture to be coded and reference picture(s)), and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in a manner similar to that which the (remote) decoder also generates (since any compression between the symbols and the coded video bitstream in the video compression techniques considered in the disclosed subject is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Because the decoding of the symbol stream leads to bit-exact results independently of the decoder location (local or remote), the content within the reference picture memory (634) is also bit exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" as reference picture samples exactly the same sample values ​​that the decoder "would see" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift, for example, when synchronization cannot be maintained due to channel errors) is also used in some related technologies.

[0054] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder, such as the video decoder (510) already described in detail above in relation to FIG. 5. However, with brief reference to FIG. 5, since symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and the parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633).

[0055] An observation that can be made at this point is that any decoder technique, excluding the parsing / entropy decoding present in the decoder, must also inevitably exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject focuses on decoder operations. Since encoder techniques are the inverse of the comprehensively described decoder techniques, their description may be condensed. Further details are required only in specific sections and are provided below.

[0056] During operation, in some examples, the source coder (630) may perform motion-compensated predictive coding, which predictively codes the input picture by referring to one or more previously coded pictures from a video sequence designated as "reference pictures." In this way, the coding engine (632) codes the differences between pixel blocks of the input picture and pixel blocks of the reference picture(s) that can be selected as predictive reference(s) for the input picture.

[0057] The local video decoder (633) can decode the coded video data of pictures that can be designated as reference pictures based on symbols generated by the source coder (630). The operations of the coding engine (632) may advantageously be lossy processes. When the coded video data can be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may be a replica of the source video sequence, typically having some errors. The local video decoder (633) can replicate the decoding processes that can be performed by the video decoder on the reference pictures and allow the reconstructed reference pictures to be stored in the reference picture cache (634). In this way, the video encoder (603) can locally store copies of reconstructed reference pictures having common content as reconstructed reference pictures to be acquired by the far-end video decoder (without transmission errors).

[0058] The predictor (635) can perform predictive searches for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for specific metadata or sample data (as candidate reference pixel blocks), such as reference picture motion vectors, block shapes, etc., which can serve as appropriate predictive references for the new pictures. The predictor (635) can operate on a sample block-by-pixel block basis to find appropriate predictive references. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have predictive references drawn from a number of reference pictures stored in the reference picture memory (634).

[0059] The controller (650) can manage the coding operations of the source coder (630), including, for example, the settings of parameters and subgroup parameters used to encode video data.

[0060] The outputs of all the aforementioned function units can be subject to entropy coding in the entropy coder (645). The entropy coder (645) converts symbols such as those generated by the various function units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.

[0061] The transmitter (640) may buffer coded video sequence(s), such as those generated by the entropy coder (645), to prepare for transmission over a communication channel (660), which may be a hardware / software link to a storage device for storing encoded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, e.g., coded audio data and / or auxiliary data streams (sources not shown).

[0062] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each coded picture, which can influence the coding techniques that can be applied to each picture. For example, pictures can often be assigned as one of the following picture types:

[0063] An Intra Picture (I Picture) may be one that can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of Intra Pictures, including, for example, "IDR (Independent Decoder Refresh) Pictures." A person skilled in the art recognizes such variations of I Pictures and their respective applications and features.

[0064] The predictive picture (P picture) may be coded and decoded using intra prediction or inter prediction, using at most one motion vector and reference index to predict sample values ​​of each block.

[0065] A bidirectionally predictive picture (B picture) may be coded and decoded using intra-prediction or inter-prediction, which uses at most two motion vectors and reference indices to predict sample values ​​of each block. Similarly, multi-prediction pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0066] Source pictures are often spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and can be coded on a block-by-block basis. Blocks can be predictively coded by referencing other (already coded) blocks determined by the coding assignment applied to each of the blocks' respective pictures. For example, blocks of pictures I can be non-predictively coded, or they can be predictively coded by referencing already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of picture P can be predictively coded via spatial prediction or temporal prediction by referencing one previously coded reference picture. Blocks of pictures B can be predictively coded via spatial prediction or temporal prediction by referencing one or two previously coded reference pictures.

[0067] The video encoder (603) can perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In its operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Thus, the coded video data can comply with the syntax specified by the video coding technique or standard in use.

[0068] In an embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as time / space / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0069] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes spatial correlations within a given picture, while inter-picture prediction utilizes (temporal or other) correlations between pictures. In the example, a specific picture being encoded / decoded, referred to as the current picture, is partitioned into blocks. When a block within the current picture is similar to a reference block within a previously encoded and still buffered reference picture in the video, the block within the current picture can be encoded by a vector referred to as a motion vector. The motion vector points to the reference block within the reference picture and, if multiple reference pictures are in use, may have a third dimension identifying the reference picture.

[0070] In some embodiments, a two-prediction technique may be used in inter-picture prediction. According to the two-prediction technique, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which precede the current picture in the video in the decoding order (however, in the display order, they may be in the past and future, respectively). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture, and a second motion vector pointing to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0071] In addition, a merge mode technique can be used in inter-picture prediction to improve coding efficiency.

[0072] According to some embodiments of the present disclosure, predictions such as inter-picture predictions and intra-picture predictions are performed in units of blocks. For example, according to the HEVC standard, a picture in a sequence of video pictures is partitioned into coding tree units (CTUs) for compression, and the CTUs in the picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU contains three coding tree blocks (CTBs), which are one luminance CTB and two chroma CTBs. Each CTU can be recursively quadtree split into one or more coding units (CUs). For example, a CTU of 64x64 pixels can be split into one CU of 64x64 pixels, four CUs of 32x32 pixels, or sixteen CUs of 16x16 pixels. In the example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. The CU is divided into one or more prediction units (PUs) based on temporal and / or spatial predictability. Typically, each PU includes a luminal prediction block (PB) and two chroma PBs. In the embodiment, the prediction operation in coding (encoding / decoding) is performed in the prediction block unit. If a luminal prediction block is used as an example of a prediction block, the prediction block includes a matrix of values ​​(e.g., luminal values) for pixels, such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0073] FIG. 7 illustrates a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values ​​within a current video picture in a sequence of video pictures, and to encode the processing block into a coded picture that is part of a coded video sequence. In the example, the video encoder (703) is used instead of the video encoder (403) in the example of FIG. 4.

[0074] In the example of HEVC, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a prediction block of 8x8 samples. The video encoder (703) determines which of the processing block is best coded using, for example, an intra mode, an inter mode, or a two-predict mode using rate-distortion optimization. When the processing block is to be coded in intra mode, the video encoder (703) may encode the processing block into the coded picture using an intra prediction technique; when the processing block is to be coded in inter mode or two-predict mode, the video encoder (703) may use an inter prediction or two-predict technique, respectively, to encode the processing block into the coded picture. In certain video coding techniques, the merge mode may be an inter-picture prediction submode in which motion vectors are derived from one or more motion vector predictors without gaining any coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the target block. In the example, the video encoder (703) includes other components such as a mode determination module (not shown) for determining the mode of the processing blocks.

[0075] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) combined together as shown in FIG. 7.

[0076] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in reference pictures (e.g., blocks in previous and later pictures), generate inter-prediction information (e.g., description of redundant information according to the inter-encoding technique, motion vectors, merge mode information), and calculate inter-prediction results (e.g., predicted blocks) based on the inter-prediction information using any suitable technique. In some examples, the reference pictures are decoded reference pictures that are decoded based on encoded video information.

[0077] The intra-encoder (722) is configured to receive samples of the current block (e.g., processing block), compare the block with already coded blocks within the same picture in some cases, generate quantized coefficients after transformation, and in some cases also receive intra-prediction information (e.g., intra-prediction direction information according to one or more intra-encoding techniques). In an example, the intra-encoder (722) also calculates intra-prediction results (e.g., prediction blocks) based on reference blocks within the same picture and intra-prediction information.

[0078] A general controller (721) is configured to determine general control data and to control other components of the video encoder (703) based on the general control data. In an example, the general controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, when the mode is intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select intra prediction information and include the intra prediction information in the bitstream; when the mode is inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter prediction information and include the inter prediction information in the bitstream.

[0079] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction results selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data and generate transformation coefficients. In the example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transformation coefficients. The transformation coefficients are then subjected to quantization processing to obtain quantized transformation coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transformation and generate decoded residual data. The decoded residual data can be suitably used by the intra-encoder (722) and the inter-encoder (730). For example, the inter encoder (730) can generate decoded blocks based on decoded residual data and inter prediction information, and the intra encoder (722) can generate decoded blocks based on decoded residual data and intra prediction information. The decoded blocks are appropriately processed to generate decoded pictures, and the decoded pictures are buffered in a memory circuit (not shown) and can be used as reference pictures in some examples.

[0080] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information according to a suitable standard, such as the HEVC standard. In the example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information within the bitstream. According to the subject matter disclosed, it is noted that when coding a block in the inter mode or the merged submode of both prediction modes, residual information is not present.

[0081] FIG. 8 illustrates a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and to decode the coded pictures to generate reconstructed pictures. In the example, the video decoder (810) is used instead of the video decoder (410) in the example of FIG. 4.

[0082] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) combined together as shown in FIG. 8.

[0083] The entropy decoder (871) may be configured to reconstruct specific symbols from the coded picture that represent the syntax elements constituting the coded picture. Such symbols may include, for example, prediction information (for example, intra prediction information or inter prediction information) capable of identifying specific samples or metadata used for prediction by the intra decoder (872) or inter decoder (880), for example, the mode in which the block is coded (e.g., intra mode, inter mode, bi-predicted mode, merged submode, or the latter two in other submodes), for example, the mode in which the block is coded (e.g., intra mode, inter mode, bi-predicted mode, merged submode, or the latter two in other submodes), for example, residual information in the form of quantized transformation coefficients. In the example, when the prediction mode is an inter or bi-predicted mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). Residual information can be subject to inverse quantization and is provided to a residual decoder (873).

[0084] The inter decoder (880) is configured to receive inter prediction information and generate inter prediction results based on the inter prediction information.

[0085] The intra decoder (872) is configured to receive intra prediction information and generate prediction results based on the intra prediction information.

[0086] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients and to process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require specific control information (including a Quantizer Parameter QP), and that information may be provided by the entropy decoder (871) (since this may only be low-capacity control information, the data path is not described).

[0087] The reconstruction module (874) is configured to form a reconstructed block by combining residuals and prediction results (as in the case, as in the case output by inter or intra prediction modules) such as those output by the residual decoder (873) in the spatial domain, and the reconstructed block may be part of a reconstructed picture, and the reconstructed picture may be part of a reconstructed video. Note that other suitable operations, such as deblocking operations, may be performed to improve visual quality.

[0088] It should be noted that the video encoders (403, 603, and 703), and video decoders (410, 510, and 810) can be implemented using any suitable technique. In an embodiment, the video encoders (403, 603, and 703), and video decoders (410, 510, and 810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403, 603, and 703), and video decoders (410, 510, and 810) can be implemented using one or more processors that execute software instructions.

[0089] Neural network technology can be used in conjunction with video coding technology, and video coding technology using neural networks can be referred to as hybrid video coding technology. For example, a loop filter unit such as a loop filter unit (556) can apply various loop filters for sample filtering. One or more of the loop filters may be implemented by a neural network. Aspects of the present disclosure provide techniques for in-loop filtering in hybrid video coding technologies for improving picture quality using a neural network. Specifically, according to an aspect of the present disclosure, techniques for clipping data before the data is fed into the kernel of a neural network-based in-loop filter may be used.

[0090] According to an aspect of the present disclosure, in-loop filters are filters that affect reference data. For example, an image filtered by a loop filter unit (556) is stored in a buffer, such as a reference picture memory (557), as a reference for further prediction. In-loop filters can improve video quality in a video codec.

[0091] FIG. 9 illustrates a block diagram of a loop filter unit (900) in some examples. The loop filter unit (900) may be used instead of the loop filter unit (556) in the example. In the example of FIG. 9, the loop filter unit (900) includes a deblocking filter (901), a sample adaptive offset (SAO) filter (902), and an adaptive loop filter (ALF) filter (903). In some examples, the ALF filter (903) may include a cross component adaptive loop filter (CCALF).

[0092] During operation, in the example, the loop filter unit (900) receives the reconstructed picture, applies various filters to the reconstructed picture, and generates an output picture in response to the reconstructed picture.

[0093] In some examples, the deblocking filter (901) and the SAO filter (902) are configured to remove blocking artifacts introduced when block coding techniques are used. The deblocking filter (901) can smooth shape edges formed when block coding techniques are used. The SAO filter (902) can be applied to specific offsets of samples to reduce distortion for other samples within a video frame. The ALF (903) can, for example, apply classification to a block of samples and then apply a filter associated with the classification to the block of samples. In some examples, the filter coefficients of the filter can be determined by the encoder and signaled to the decoder.

[0094] In some examples (e.g., JVET-T0057), an additional filter referred to as a dense residual convolutional neural network based in-loop filter (DRNLF) may be inserted between the deblocking filter (901) and the SAO filter (902). The DRNLF can further improve picture quality.

[0095] FIG. 10 illustrates a block diagram of a loop filter unit (1000) in some examples. The loop filter unit (1000) may be used instead of the loop filter unit (556) in the example. In the example of FIG. 10, the loop filter unit (1000) includes a deblocking filter (1001), an SAO filter (1002), an ALF filter (1003), and a DRNLF filter (1010) placed between the deblocking filter (1001) and the SAO filter (1002).

[0096] The deblocking filter (1001) is configured similarly to the deblocking filter (901), the SAO filter (1002) is configured similarly to the SAO filter (902), and the ALF filter (1003) is configured similarly to the ALF filter (903).

[0097] The DRNLF filter (1010) receives the output of the deblocking filter (1001) illustrated by the deblocked picture (1011) and also receives a quantization parameter (QP) map of the reconstructed picture. The QP map contains the quantization parameters of the blocks within the reconstructed picture. The DRNLF filter (1010) can output a picture illustrated by the filtered picture (1019) having improved quality, and the filtered picture (1019) is fed to the SAO filter (1002) for additional filtering processes.

[0098] According to an embodiment of the present disclosure, a neural network for video processing may include a plurality of channels for processing color components in a color space. In an example, the color space may be defined using a YCbCr model. In the YCbCr model, Y represents the luminance component, and Cb and Cr represent the chroma components. Note that in the following description, YUV is used to describe a format encoded using the YCbCr model.

[0099] According to an embodiment of the present disclosure, a plurality of channels within a neural network are configured to operate on color components of the same size. In some examples, pictures may be represented by color components of different sizes. For example, the human visual system is much more sensitive to variations in luminance than to color, and thus video systems may compress chroma components to reduce file size and save transmission time without as much visual difference as perceived by the human eye. In some examples, chroma subsampling techniques are used to achieve lower resolution for chroma information than for luminance information by utilizing the human visual system's vision for color differences rather than for luminance.

[0100] In some examples, subsampling may be expressed as a three-part ratio such as 4:4:4, 4:2:0, 4:2:2, 4:1:1, etc. For example, 4:4:4 (also referred to as YUV444) indicates that each of the YCbCr components has the same sample rate without subsampling; and 4:2:0 (also referred to as YUV420) indicates that the chroma components are subsampled, and all four pixels (or Y components) may correspond to the Cb component and the Cr component. It is noted that YUV420 is used in the following description as an example of a subsampling format to illustrate the techniques disclosed in this disclosure. The disclosed techniques may be used for other subsampling formats. For ease of explanation, a format having color components with the same sample rate without subsampling (e.g., YUV444) is referred to as a non-subsampled format. Formats having at least one color component that is subsampled (e.g., YUV420, YUV422, YUV411, etc.) are referred to as subsampled formats.

[0101] Generally, neural networks can operate on pictures in an unsubsampled format (e.g., YUV444). Therefore, for pictures in a subsampled format, the picture is converted to an unsubsampled format before being provided as input to the neural network.

[0102] FIG. 11 illustrates a block diagram of a DRNLF filter (1100) in some examples. The DRNLF filter (1100) may be used instead of the DRNLF filter (1010) in the examples. The DRNLF filter (1100) includes a QP map quantizer (1110), a preprocessing module (1120), a main processing module (1130), and a postprocessing module (1140) combined together as shown in FIG. 11. The main processing module (1130) includes a patch fetcher (1131), a patch-based DRNLF kernel processing module (1132), and a patch reassembler (1133) combined together as shown in FIG. 11.

[0103] In some examples, the QP map includes a map of QP values ​​applied to reconstruct each block within the currently reconstructed picture. The QP map quantizer (1110) can quantize the values ​​into a predetermined set of values. In an example (e.g., JVET-T0057), the QP values ​​can be quantized by the QP map quantizer (1110) into one of 22, 27, 32, and 37.

[0104] The preprocessing module (1120) receives the deblocked picture in a first format and can convert it to a second format used by the main processing module (1130). For example, the main processing module (1130) is configured to process the picture in YUV444 format. When the preprocessing module (1120) receives the deblocked picture in a format different from the YUV444 format, the preprocessing module (1120) processes the deblocked picture in a different format and can output the deblocked picture in YUV444 format. For example, the preprocessing module (1120) receives the deblocked picture in YUV420 format and then generates the deblocked picture in YUV444 format by horizontally and vertically doubling the U and V chrominance channels.

[0105] The main processing module (1130) can receive a deblocked picture in YUV444 format and a quantized QP map as inputs. The patch fetcher (1131) breaks the inputs into patches. The DRNLF kernel processing module (1132) can process each of the patches based on the DRNLF kernel. The patch reassembler (1133) can assemble the processed patches into a filtered picture in YUV444 format by the DRNLF kernel processing module (1132).

[0106] The post-processing module (1140) converts the filtered picture of the second format back to the first format. For example, the post-processing module (1140) receives a filtered picture of the YUV444 format (output from the main processing module (1130)) and outputs a filtered picture of the YUV420 format.

[0107] FIG. 12 illustrates a block diagram of a preprocessing module (1220) in some examples. In the example, the preprocessing module (1220) is used instead of the preprocessing module (1120).

[0108] The preprocessing module (1220) receives a deblocked picture in YUV420 format, converts the deblocked picture to YUV444 format, and can output a deblocked picture in YUV444 format. Specifically, the preprocessing module (1220) receives a deblocked picture from three input channels, each including a luminance input channel for the Y component and two chrominance input channels for the U (Cb) component and V (Cr) component. The preprocessing module (1220) outputs a deblocked picture from three output channels, each including a luminance output channel for the Y component and two chrominance output channels for the U (Cb) component and V (Cr) component.

[0109] In the example, when the deblocked picture has the YUV420 format, the Y component has dimensions (H,W), the U component has dimensions (H / 2,W / 2), and the V component has dimensions (H / 2,W / 2), where H represents the height of the deblocked picture (e.g., in a unit of samples) and W represents the width of the deblocked picture (e.g., in a unit of samples).

[0110] In the example of FIG. 12, the preprocessing module (1220) does not scale the Y component. The preprocessing module (1220) receives a Y component with magnitude (H,W) from a luminance input channel and outputs the Y component with magnitude (H,W) to a luminance output channel.

[0111] The preprocessing module (1220) scales the U component and the V component, respectively. The preprocessing module (1220) includes a first scaling unit (1221) and a second scaling unit (1222) for processing the U component and the V component, respectively. For example, the first scaling unit (1221) receives a U component having a size (H / 2, W / 2), scales the U component to a size (H, W), and outputs the U component having a size (H, W) to a chrominance output channel for the U component. The second scaling unit (1222) receives a V component having a size (H / 2, W / 2), scales the V component to a size (H, W), and outputs the V component having a size (H, W) to a chrominance output channel for the V component. In some examples, the first scaling unit (1221) scales the U component based on interpolation, such as using a Lanczos interpolation filter. Similarly, in some examples, the second scaling unit (1222) scales the V component based on interpolation, such as using a Lanczos interpolation filter.

[0112] In some examples, interpolation operations, such as using Ranchos interpolation filters, cannot guarantee that the output of the interpolation operations will be meaningful values, such as non-negative values ​​for the meaningful U(Cb) and V(Cr) components. In some examples, deblocked pictures in YUV444 format after preprocessing may be stored, and then the stored pictures in YUV444 format may be used in the neural network training process. Negative values ​​for the U(Cb) and V(Cr) components may have an adverse effect on the results of the neural network training process.

[0113] FIG. 13 illustrates a block diagram of a neural network structure (1300). In some examples, the neural network structure (1300) is used for a dense residual convolutional neural network-based in-loop filter (DRNLF) and can be used instead of a patch-based DRNLF kernel processing module (1132). The neural network structure (1300) includes a series of dense residual units (DRUs), such as DRUs (1301) and DRUs (1304), and the number of DRUs is denoted by N. In FIG. 13, the number of convolutional kernels is denoted by M, and M is also the number of output channels for the convolution. For example, "CONV 3×3×M" indicates a standard convolution with M convolutional kernels of kernel size 3×3, and "DSC 3×3×M" indicates a depth-directed separable convolution with M convolutional kernels of kernel size 3×3. N and M can be set for a trade-off between computational efficiency and performance. In the example (e.g., JVET-T0057), N is set to 4 and M is set to 32.

[0114] During operation, the neural network structure (1300) processes the deblocked picture by patch. For each patch of the deblocked picture in YUV444 format, the patch is normalized (e.g., divided by 1023 in the example of FIG. 13), and the average value of the deblocked picture is removed from the normalized patch to obtain a first part (1311) of the internal input (1313). A second part of the internal input (1313) is from a QP map. For example, a patch of the QP map corresponding to the patch forming the first part (1311) (referred to as a QP map patch) is obtained from the QP map. The QP map patch is normalized (e.g., divided by 51 in FIG. 13). The normalized QP map patch is the second part (1312) of the internal input (1313). The second part (1312) is connected to the first part (1311) to obtain an internal input (1313). The internal input (1313) is provided to the first normal convolution block (1351) (illustrated as CONV 3x3xM). Then, the output of the first normal convolution block (1351) is processed by N DRUs.

[0115] For each DRU, an intermediate input is received and processed. The output of the DRU is concatenated with the intermediate input to form an intermediate input for the next DRU. Using DRU (1302) as an example, DRU (1302) receives an intermediate input (1321), processes the intermediate input (1321), and generates an output (1322). The output (1322) is concatenated with the intermediate input (1321) to form an intermediate input (1323) for DRU (1303).

[0116] Note that because the intermediate input (1321) has more than M channels, a “CONV 1×1×M” convolution operation can be applied to the intermediate input (1321) to generate M channels for additional processing by the DRU (1302). Also note that the output of the first normal convolution block (1351) contains M channels, and thus the output can be processed by the DRU (1301) without using a “CONV 1×1×M” convolution operation.

[0117] The output of the last DRU is provided to the last normal convolution block (1359). The output of the last normal convolution block (1359) is converted into normal picture patch values ​​by adding the average value of the deblocked picture and multiplying by 1023, for example, as shown in FIG. 13.

[0118] FIG. 14 illustrates a block diagram of a dense residual unit (DRU) (1400). In some examples, the DRU (1400) may be used instead of each of the DRUs of FIG. 13, such as DRU (1301), DRU (1302), DRU (1303), and DRU (1304).

[0119] In the example of FIG. 14, the DRU (1400) receives an intermediate input x and propagates the intermediate input directly to a subsequent DRU via a shortcut (1401). The DRU (1400) also includes a normal processing path (1402). In some examples, the normal processing path (1402) includes a normal convolutional layer (1411), depth-separable convolutional (DSC) layers (1412 and 1414), and a rectified linear unit (ReLU) layer (1413). For example, the intermediate input x is concatenated with the output of the normal processing path (1402) to form an intermediate input for the subsequent DRU.

[0120] In some examples, DSC layers (1412 and 1414) are used to reduce computation costs.

[0121] According to an embodiment of the present disclosure, the neural network structure (1300) includes three channels corresponding to the Y, U (Cb), and V (Cr) components, respectively. In some examples, the three channels may be referred to as the Y channel, U channel, and V channel. The DRNLF filter (1100) may be applied to both intra and inter pictures. In some examples, additional flags are signaled to indicate the on / off of the DRNLF filter (1100) at the picture level and CTU level.

[0122] FIG. 15 illustrates a block diagram of a post-processing module (1540) in some examples. The post-processing module (1540) may be used instead of the post-processing module (1140) in the examples. The post-processing module (1540) includes clipping units (1541-1543) that clip the values ​​of the Y component, U component, and V component to a predetermined non-negative range [a, b], respectively. In the example, the lower limit a and upper limit b of the non-negative range may be set to a = 16 × 4 and b = 234 × 4. Additionally, the post-processing module (1540) includes scaling units (1545 and 1546) that scale the clipped U component and V component from size (H, W) to size (H / 2, W / 2), respectively, where H is the height of the original picture (e.g., the deblocked picture) and W is the width.

[0123] Aspects of the present disclosure provide preprocessing techniques. The preprocessed data can be stored and used for training a neural network, and better training and inference results can be obtained.

[0124] FIG. 16 illustrates a block diagram of a preprocessing module (1620) in some examples. In the example, the preprocessing module (1620) is used instead of the preprocessing module (1120).

[0125] The preprocessing module (1620) receives a deblocked picture in YUV420 format, converts the deblocked picture to YUV444 format, and can output a deblocked picture in YUV444 format. Specifically, the preprocessing module (1620) receives a deblocked picture in three input channels, each including a luminance input channel for the Y component and two chrominance input channels for the U (Cb) component and V (Cr) component, respectively. The preprocessing module (1620) outputs a deblocked picture by three output channels, each including a luminance output channel for the Y component and two chrominance output channels for the U component and V component, respectively.

[0126] In the example, when the deblocked picture has the YUV420 format, the Y component has dimensions (H,W), the U component has dimensions (H / 2,W / 2), and the V component has dimensions (H / 2,W / 2), where H represents the height of the deblocked picture (e.g., in a unit of samples) and W represents the width of the deblocked picture (e.g., in a unit of samples).

[0127] In the example of FIG. 16, the preprocessing module (1620) does not scale the Y component. The preprocessing module (1620) receives a Y component with magnitude (H,W) from a luminance input channel and outputs a Y component with magnitude (H,W) to a luminance output channel.

[0128] The preprocessing module (1620) scales the U component and the V component, respectively. The preprocessing module (1620) includes a first scaling unit (1621) and a second scaling unit (1622) for processing the U component and the V component, respectively. For example, the first scaling unit (1621) receives a U component with a size (H / 2, W / 2), scales the U component to a size (H, W), and outputs the U component with a size (H, W) to a chrominance output channel for the U component. The second scaling unit (1622) receives a V component with a size (H / 2, W / 2), scales the V component to a size (H, W), and outputs the V component with a size (H, W) to a chrominance output channel for the V component. In some examples, the first scaling unit (1621) scales the U component based on interpolation, such as using a Ranchos interpolation filter. Similarly, in some examples, the second scaling unit (1622) scales the V component based on interpolation, such as using a Ranchos interpolation filter.

[0129] In some examples, interpolation operations, such as using a Ranchos interpolation filter, cannot guarantee that the output of the interpolation operations will be meaningful values, such as non-negative for meaningful U(Cb) and V(Cr) components.

[0130] In the example of FIG. 16, the preprocessing module (1620) includes clipping units (1625 and 1626) that clip the values ​​of the U component and the V component, respectively, after interpolation into the range [c,d]. In some examples, the values ​​of the Y component, U component, and V component for preprocessing have a bit depth of 10, and then c and d are c=0 and d=2 bitdepth -1 can be set as 1023.

[0131] In one example, the values ​​of c and d are predefined and used. In another example, multiple pairs of c and d values ​​are predefined, and the indices of pairs of c and d values ​​for use in clipping can be signaled in the bitstream, such as in the sequence parameter set (SPS), picture parameter set (PPS), slice, or tile header.

[0132] In some examples, the clipped values ​​of the U and V components and the values ​​of the Y component may be stored as deblocked pictures in YUV444 format. In some embodiments, the stored pictures in YUV444 format may be used as inputs in the training process of a neural network, such as the neural network of the main processing module (1130). In some examples, the values ​​of the U and V components are clipped to a range that does not adversely affect the training process of the neural network. In an example, the values ​​of the U and V components are clipped so as not to be negative.

[0133] In some examples, by using pictures saved in YUV444 format with clipped values, the training of a neural network can be accelerated due to time savings by avoiding preprocessing steps (e.g., resizing, clipping) during training. Additionally, the neural network can be trained with better model parameters that can improve compression efficiency and / or picture quality.

[0134] In some examples, adding clipping units (1625 and 1626) to the preprocessing module (1620) can improve compression efficiency and / or quality, for example, to a lower BD-rate (Bjontegaard delta rate).

[0135] FIG. 17 illustrates a flowchart outlining a process (1700) according to an embodiment of the present disclosure. The process (1700) may be used in video processing. In various embodiments, the process (1700) is executed by a processing circuit within terminal devices (310, 320, 330 and 340), a processing circuit that performs the functions of a video encoder (403), a processing circuit that performs the functions of a video decoder (410), a processing circuit that performs the functions of a video decoder (510), a processing circuit that performs the functions of a video encoder (603), and the like. In some embodiments, the process (1700) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit performs the process (1700). The process starts at (S1701) and proceeds to (S1710).

[0136] In (S1710), a picture in a format subsampled in color space is converted to a format not subsampled in color space. In some examples, the conversion is performed based on interpolation and may result in invalid values. In an example, the conversion may result in negative values ​​that are invalid for the YCbCr model.

[0137] In (S1720), the values ​​of one or more color components of the picture in the unsubsampled format are clipped before the picture in the unsubsampled format is provided as input to a neural network-based filter. In some examples, one or more color components may be chroma component(s). Next, the process proceeds to (S1799).

[0138] In the example, the values ​​of the color components of the picture in the unsubsampled format are clipped within the valid range for the color components. In the example, the values ​​of the color components of the picture in the unsubsampled format are clipped so as not to be negative. In another example, the range is determined based on the bit depth. For example, the lower limit of the range is 0, and the upper limit of the range is (2 bitdepthIt is set to )-1.

[0139] In some examples, the range is predetermined. In some examples, the range is determined based on decoded information from a bitstream carrying a picture. In some examples, a signal indicating the range is decoded from at least one of a set of sequence parameters, a set of picture parameters, a slice header, and a tile header within the bitstream.

[0140] In the example, a plurality of ranges may be predetermined. Subsequently, an index indicating one of the plurality of ranges may be carried in one of the sequence parameter set, picture parameter set, slice header, and tile header within the bitstream.

[0141] In some examples, the process (1700) is used in a decoder. For example, a picture in a subsampled format is reconstructed based on decoded information from a bitstream, and a deblocking filter is applied to the picture in the subsampled format before transitioning from the subsampled format to the non-subsampled format. In another example, a neural network-based filter is applied to a picture in the non-subsampled format with clipped values ​​to generate a filtered picture in the non-subsampled format, and then the filtered picture in the non-subsampled format is transitioned to a filtered picture in the subsampled format.

[0142] In some examples, a picture in an unsubsampled format with clipped values ​​is stored in storage. Then, the stored picture in an unsubsampled format with clipped values ​​can be provided along with other pictures as training inputs to train a neural network in a neural network-based filter.

[0143] Note that the various units, blocks, and modules described above can be implemented by various technologies, such as processing circuits, processors that execute software instructions, and combinations of hardware and software.

[0144] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 18 illustrates a computer system (1800) suitable for implementing specific embodiments of the disclosed subject matter.

[0145] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to generate code containing instructions that can be executed directly or through interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0146] The commands can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0147] The components illustrated in FIG. 18 for the computer system (1800) are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be interpreted as having any dependency or requirement in relation to any one or a combination thereof of the components illustrated in the exemplary embodiments of the computer system (1800).

[0148] The computer system (1800) may include specific human interface input devices. Such human interface input devices may be responsive to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, applause), visual input (e.g., gestures), and olfactory input (not described). Human interface devices may also be used to capture specific media that do not necessarily need to be directly related to conscious input by a human, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., 2D video, 3D video including stereoscopic video).

[0149] Input human interface devices may include one or more of the following (only one of each is depicted): a keyboard (1801), a mouse (1802), a trackpad (1803), a touch screen (1810), a data glove (not shown), a joystick (1805), a microphone (1806), a scanner (1807), and a camera (1808).

[0150] The computer system (1800) may also include specific human interface output devices. Such human interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., touch-screen (1810), data glove (not shown), or haptic feedback via joystick (1805), but there may also be haptic feedback devices that do not function as input devices), audio output devices (e.g.: speakers (1809), headphones (not depicted)), visual output devices (e.g., screens (1810) including CRT screens, LCD screens, plasma screens, OLED screens—each with or without touch-screen input capability, each with or without haptic feedback capability, some of which may output two-dimensional visual output or output of three or more dimensions through means such as stereographic output; virtual reality glasses (not depicted), holographic displays and smoke tanks (not depicted)), and printers (not depicted).

[0151] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including a CD / DVD ROM / RW (1820) having a media (1821) such as a CD / DVD, a thumb-drive (1822), a removable hard drive or solid-state drive (1823), legacy magnetic media such as tape and floppy disk (not depicted), specialized ROM / ASIC / PLD-based devices such as a security dongle (not depicted).

[0152] Those skilled in the art will also understand that the term "computer readable media," as used in connection with the subject matter disclosed herein, does not include transmission media, carrier waves, or other transient signals.

[0153] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The networks may be, for example, wireless, wireline, optical. The networks may additionally be local, wide-area, metropolitan, automotive and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks, such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wireline or wireless wide-area digital networks including cable TV, satellite TV and terrestrial broadcast TV, automotive and industrial including CANBus, etc. Certain networks generally require external network interface adapters attached to certain general-purpose data ports or peripheral buses (1849) (e.g., USB ports of the computer system (1800)); Others are generally integrated into the core of the computer system (1800) by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1800) can communicate with other entities. Such communication may be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus for specific CANbus devices), or bidirectional to other computer systems using, for example, local or broadband digital networks. Specific protocols and protocol stacks may be used on each of such networks and network interfaces as described above.

[0154] The aforementioned human interface devices, human-accessible storage devices, and network interfaces can be attached to the core (1840) of the computer system (1800).

[0155] The core (1840) may include one or more central processing units (CPUs) (1841), specialized programmable processing units in the form of graphics processing units (GPUs) (1842), field programmable gate areas (FPGAs) (1843), hardware accelerators (1844) for specific tasks, graphics adapters (1850), etc. These devices may be connected via a system bus (1848), along with internal mass storage (1847), such as read-only memory (ROM) (1845), random access memory (1846), internal non-user accessible hard drives, SSDs, etc. In some computer systems, the system bus (1848) may be accessible in the form of one or more physical plugs to enable extensions by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1848) or via a peripheral bus (1849). In the example, the screen (1810) can be connected to a graphics adapter (1850). Architectures for peripheral buses include PCI, USB, etc.

[0156] CPUs (1841), GPUs (1842), FPGAs (1843), and accelerators (1844) can be combined to execute specific instructions that can constitute the aforementioned computer code. The computer code may be stored in ROM (1845) or RAM (1846). While transitional data may also be stored in RAM (1846), permanent data may be stored, for example, in internal mass storage (1847). High-speed storage and retrieval of any of the memory devices may be made possible through the use of cache memory, which may be closely associated with one or more CPUs (1841), GPUs (1842), mass storage (1847), ROM (1845), RAM (1846), etc.

[0157] A computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be those specifically designed and configured for the purposes of this disclosure, or they may be of a kind well known and available to those skilled in the field of computer software technology.

[0158] As an example rather than a limitation, a computer system (1800) having an architecture, and specifically a core (1840), may provide functionality as a result of processor(s) (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software implemented on one or more types of tangible computer-readable media. Such computer-readable media may be media associated with specific storage of the core (1840) that is of a non-transient nature, such as core-internal mass storage (1847) or ROM (1845), as well as user-accessible mass storage as described above. Software implementing various embodiments of the present disclosure may be stored on such devices and executed by the core (1840). The computer-readable media may include one or more memory devices or chips as needed. Software may enable the core (1840) and, specifically, the processors within it (including a CPU, GPU, FPGA, etc.) to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (1846) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise implemented within the circuit (e.g., accelerator (1844)), which may operate instead of or together with the software to execute specific processes or specific parts of specific processes described herein. Reference to software may, where appropriate, encompass logic, and vice versa. Reference to computer-readable medium may, where appropriate, encompass circuits storing software for execution (e.g., integrated circuits (ICs)), circuits implementing logic for execution, or both.The present disclosure covers any suitable combination of hardware and software.

[0159] Appendix A: Acronyms

[0160] JEM: joint exploration model

[0161] VVC: versatile video coding

[0162] BMS: benchmark set

[0163] MV: Motion Vector

[0164] HEVC: High Efficiency Video Coding

[0165] SEI: Supplementary Enhancement Information

[0166] VUI: Video Usability Information

[0167] GOPs: Groups of Pictures

[0168] TUs: Transform Units,

[0169] PUs: Prediction Units

[0170] CTUs: Coding Tree Units

[0171] CTBs: Coding Tree Blocks

[0172] PBs: Prediction Blocks

[0173] HRD: Hypothetical Reference Decoder

[0174] SNR: Signal Noise Ratio

[0175] CPUs: Central Processing Units

[0176] GPUs: Graphics Processing Units

[0177] CRT: Cathode Ray Tube

[0178] LCD: Liquid-Crystal Display

[0179] OLED: Organic Light-Emitting Diode

[0180] CD: Compact Disc

[0181] DVD: Digital Video Disc

[0182] ROM: Read-Only Memory

[0183] RAM: Random Access Memory

[0184] ASIC: Application-Specific Integrated Circuit

[0185] PLD: Programmable Logic Device

[0186] LAN: Local Area Network

[0187] GSM: Global System for Mobile communications

[0188] LTE: Long-Term Evolution

[0189] CANBus: Controller Area Network Bus

[0190] USB: Universal Serial Bus

[0191] PCI: Peripheral Component Interconnect

[0192] FPGA: Field Programmable Gate Areas

[0193] SSD: solid-state drive

[0194] IC: Integrated Circuit

[0195] CU: Coding Unit

[0196] Although the present disclosure describes various exemplary embodiments, there are modifications, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, a person skilled in the art will understand that numerous systems and methods can be devised that implement the principles of the present disclosure and thus fall within the spirit and scope thereof, even though they are not expressly illustrated or described herein.

Claims

Claim 1 A video processing method comprising: a step of converting a picture in a subsampled format in color space to a non-subsampled format in color space by a preprocessing unit of a dense residual convolutional neural network based in-loop filter (DRNLF) in a processing circuit; and a step of determining a range for clipping only the color components of the picture by the preprocessing unit in the processing circuit based on decoded information from a bitstream carrying the picture, and by the preprocessing unit in the processing circuit clipping the values ​​of the color components of the picture in the non-subsampled format to within the determined range without clipping the values ​​of the luminance components of the picture before providing the picture in the non-subsampled format as an input to the neural network-based filter of the DRNLF. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 A method according to claim 1, further comprising the step of decoding a signal indicating the range from at least one of a sequence parameter set, a picture parameter set, a slice header, and a tile header within the bitstream. Claim 7 A method according to claim 1, further comprising the step of reconstructing the picture into the subsampled format based on decoded information from the bitstream. Claim 8 A method according to claim 1, further comprising: a step of applying the neural network-based filter to the picture of the unsubsampled format having the clipped values ​​to generate the filtered picture of the unsubsampled format; and a step of converting the filtered picture of the unsubsampled format into the filtered picture of the subsampled format. Claim 9 A method according to claim 1, further comprising the step of storing a picture of the unsubsampled format having the clipped values. Claim 10 A method according to claim 9, further comprising the step of providing a stored picture in the unsubsampled format having the clipped values ​​as a training input for training a neural network in the neural network-based filter. Claim 11 A device for video processing comprising a processing circuit configured to perform the method of any one of claims 1 and 6 through 10. Claim 12 A non-transient computer-readable storage medium storing computer program code, wherein the computer program code is configured such that when executed by at least one processor, the at least one processor performs the method of any one of claims 1 and 6 through 10. Claim 13 delete Claim 14 delete Claim 15 delete Claim 16 delete Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete

Citation Information

Patent Citations

  • Nonlinear extensions of adaptive loop filtering for video coding

    US20200404335A1