Method, computer device, equipment and storage medium for video processing

By converting the subsampling format pictures in the color space into non-subsampling formats and cropping, the problem of low efficiency of in-loop filters based on neural networks in the prior art is solved, and more efficient video encoding and processing effects are achieved.

CN115151941BActive Publication Date: 2025-07-25TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180016162.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-31
Filing Date
2021-09-08
Publication Date
2025-07-25
Estimated Expiration
2041-09-08

AI Technical Summary

Technical Problem

The prior art lacks the technology to improve in-loop filters based on neural networks, resulting in low video encoding efficiency.

Method used

Convert pictures of subsampled formats in the color space to non-subsampled formats and crop the color components before input to a neural network-based filter.

Benefits of technology

It improves the efficiency and quality of video encoding, reduces data redundancy, and enhances the effect of video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115151941B_ABST
    Figure CN115151941B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods and devices for video processing. In some examples, a device for video processing includes processing circuitry. The processing circuitry converts a picture in a subsampling format in a color space into a non-subsampling format in the color space. Then, before providing the picture in the non-subsampling format as an input to a neural network-based filter, the processing circuitry clips values of color components of the picture in the non-subsampling format.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Reference Incorporation

[0002] This application claims the priority benefit of U.S. Patent Application No. 17 / 463,352, "METHOD AND APPARATUS FOR VIDEO CODING", filed on August 31, 2021, which claims the priority benefit of U.S. Provisional Application No. 63 / 131,656, "APPLICATION OF CLIPPING TO IMPROVE PRE-PROCESSING IN A NEURAL NETWORK BASED IN-LOOP FILTER IN A VIDEO CODEC", filed on December 29, 2020. The entire disclosure of the prior application is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure generally describes embodiments related to video coding. More specifically, the present disclosure provides techniques for improving neural network-based in-loop filters, and more particularly relates to a method for video processing, a computer device, a video processing apparatus, and a non-transitory computer-readable storage medium. Background Art

[0004] The background art description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work of the currently named inventors is described in this background art section and aspects of the specification that may not constitute prior art at the time of filing, such work is neither expressly nor implicitly admitted to be prior art to the present disclosure.

[0005] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having a spatial dimension of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (informally also called frame rate) of, for example, 60 pictures per second or 60 Hz. Uncompressed video has specific bitrate requirements. For example, a 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 gigabytes (GByte) of storage space.

[0006] One purpose of video encoding and decoding can be to reduce redundancy in an input video signal by compression. Compression can help reduce the aforementioned bandwidth or storage space requirements, in some cases by two orders of magnitude or more. Both lossless compression and lossy compression, and combinations thereof, can be employed. Lossless compression refers to techniques by which an exact copy of the original signal can be reconstructed from the compressed original signal. When lossy compression is used, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal is useful for the intended application. In the case of video, lossy compression is widely employed. The amount of distortion tolerated depends on the application; for example, users of some consumer streaming applications can tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that higher allowable / tolerable distortion can result in a higher compression ratio.

[0007] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy coding.

[0008] Video coding and decoding techniques can include techniques referred to as intra coding. In intra coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are coded in intra mode, the picture can be an intra picture. Intra pictures and their derivatives (e.g., independent decoder refresh pictures) can be used to reset the decoder state and can thus be used as the first picture in an encoded video bitstream and video session, or as a still picture. The samples of an intra block can be subjected to a transformation, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique that minimizes the sample values in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after transformation, the fewer bits are required to represent the block after entropy coding for a given quantization step size.

[0009] Traditional intra coding, such as intra coding known from, for example, MPEG-2 generation coding techniques, does not use intra prediction. However, some newer video compression techniques include techniques that attempt to obtain surrounding sample data and / or metadata from, for example, spatially adjacent and earlier in decoding order data blocks during encoding / decoding. Such techniques are hereafter referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and not from reference pictures.

[0010] There can be many different forms of intra prediction. When more than one such technique can be used in a given video coding technique, the techniques used can be coded in an intra prediction mode. In some cases, the mode can have sub-modes and / or parameters, and these sub-modes and / or parameters can be coded separately or included in the mode codeword. Which codewords to use for a given mode / sub-mode / parameter combination can affect the coding efficiency gain through intra prediction, and the same is true for the entropy coding technique used to convert the codewords into a bitstream. The prior art lacks techniques for improving neural-network-based in-loop filters. Summary of the Invention

[0011] According to an embodiment, a method for video processing is provided, characterized in that the method includes: converting, by a processing circuit, a picture in a subsampling format in a color space into a non-subsampling format in the color space; and clipping, by the processing circuit, values of color components of the picture in the non-subsampling format before providing the picture in the non-subsampling format as an input to a neural-network-based filter.

[0012] According to an embodiment, a computer device is provided, characterized in that the computer device includes: one or more non-transitory computer-readable storage media configured to store computer program code; and one or more computer processors configured to access the computer program code and execute the above-described method for video processing according to the instructions of the computer program code.

[0013] According to an embodiment, a device for video processing is provided, characterized in that the device includes a processing circuit configured to execute the above-described method for video processing.

[0014] According to an embodiment, a non-transitory computer-readable storage medium is provided, characterized in that the non-transitory computer-readable storage medium stores computer program code, and when the computer program code is executed by at least one processor, the at least one processor executes the above-described method for video processing.

[0015] The method for video processing, computer device, device for video processing, and non-transitory computer-readable storage medium of the present invention provide techniques for improving neural-network-based in-loop filters. Brief Description of the Drawings

[0016] According to the following detailed description and the drawings, other features, properties, and various advantages of the disclosed subject matter will become more apparent, in the drawings:

[0017] Figure 1A is a schematic illustration of an exemplary subset of intra prediction modes;

[0018] Figure 1B is an illustration of exemplary intra prediction directions;

[0019] Figure 2 is a schematic illustration of a current block and its surrounding spatial merge candidates in an example;

[0020] Figure 3 is a schematic illustration of a simplified block diagram of a communication system according to an embodiment;

[0021] Figure 4 is a schematic illustration of a simplified block diagram of a communication system according to another embodiment;

[0022] Figure 5 is a schematic illustration of a simplified block diagram of a decoder according to an embodiment;

[0023] Figure 6 is a schematic illustration of a simplified block diagram of an encoder according to an embodiment;

[0024] Figure 7 shows a block diagram of an encoder according to another embodiment;

[0025] Figure 8 shows a block diagram of a decoder according to another embodiment;

[0026] Figure 9 shows a block diagram of a loop filter unit in some examples.

[0027] Figure 10 shows a block diagram of another loop filter unit in some examples.

[0028] Figure 11 shows a block diagram of a neural network-based filter in some examples.

[0029] Figure 12 shows a block diagram of a preprocessing module in some examples.

[0030] Figure 13 shows a block diagram of a neural network structure in some examples.

[0031] Figure 14 shows a block diagram of a dense residual unit.

[0032] Figure 15 shows a block diagram of a postprocessing module in some examples.

[0033] Figure 16 shows a block diagram of a preprocessing module in some examples.

[0034] Figure 17 shows a flowchart outlining an example process.

[0035] Figure 18 is a schematic illustration of a computer system according to an embodiment. Detailed Embodiment

[0036] Specific patterns of intra prediction were introduced in H.264, improved in H.265, and further improved in more recent coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). Predictor blocks can be formed using adjacent sample values belonging to already available samples. The sample values of adjacent samples are copied into the predictor block according to a direction. The reference to the used direction can be encoded in the bitstream or can be predicted itself.

[0037] Referring Figure 1A , a subset of nine predictor directions known from the 33 possible predictor directions of H.265 (33 angular patterns corresponding to 35 intra modes) is depicted in the lower right. The point (101) where the arrows converge represents the sample to be predicted. The arrows indicate the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples in the upper right at a 45-degree angle to the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples in the lower left of sample (101) at a 22.5-degree angle to the horizontal.

[0038] Still referring Figure 1A , in the upper left, a block (104) of 4×4 samples (indicated by the dashed bold line) is depicted. Block (104) includes 16 samples, each labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in both the Y dimension and the X dimension in block (104). Since the size of the block is 4×4 samples, S44 is in the lower right. Reference samples following a similar numbering scheme are further shown. The reference samples are labeled with "R", their Y position (e.g., row index) relative to block (104), and X position (column index). In both H.264 and H.265, the predicted sample is adjacent to the block in the reconstruction; thus, negative values do not need to be used.

[0039] Intra picture prediction can be performed by copying the reference sample values of adjacent samples delimited by the predicted direction signaled freely. For example, assume that the encoded video bitstream includes the following signaling: for this block, the signaling indicates a predicted direction consistent with arrow (102), that is, the samples are predicted from one or more predicted samples in the upper right at a 45-degree angle to the horizontal. In this case, sample S41, sample S32, sample S23, and sample S14 are predicted from the same reference sample R05. Then sample S44 is predicted from reference sample R08.

[0040] In some cases, the values of multiple reference samples can be combined, for example, by interpolation, to calculate the reference sample; especially when the directions cannot be evenly divided by 45 degrees.

[0041] As video coding technology develops, the number of possible directions is also increasing. In H.264 (2003), nine different directions can be represented. In H.265 (2013), it increased to 33, and JEM / VVC / BMS at the time of disclosure can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and some techniques in entropy coding are used to represent those possible directions with a small number of bits, thus accepting certain penalties for the less likely directions. Additionally, sometimes the direction itself can be predicted from adjacent directions used in adjacent decoded blocks.

[0042] Figure 1B A schematic diagram (180) is shown depicting 65 intra prediction directions according to JEM to show the increase in the number of prediction directions over time.

[0043] The mapping of the intra prediction direction bits representing the direction in the encoded video bitstream can vary with different video coding technologies; and the range can be, for example, from a simple direct mapping of the predicted direction to the intra prediction mode, to codewords, to complex adaptive schemes involving the most likely modes, and similar techniques. However, in all cases, there can be certain directions that are statistically less likely to occur in the video content compared to some other directions. Since the goal of video compression is to reduce redundancy, in a well-functioning video coding technology, those less likely directions will be represented by more bits compared to the more likely directions.

[0044] Motion compensation can be a lossy compression technique and can involve techniques where a block of sample data from a previously reconstructed picture or a portion thereof (reference picture) is spatially shifted in the direction indicated by a motion vector (hereinafter referred to as MV) and then used to predict a newly reconstructed picture or picture portion. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or three dimensions, with the third dimension being an indication of the reference picture in use (the latter can indirectly be a temporal dimension).

[0045] In some video compression techniques, an MV applicable to a certain region of sample data can be predicted from other MVs, for example, from an MV associated with another region of sample data that is spatially adjacent to the region being reconstructed and that is before the MV in the decoding order. Doing so can significantly reduce the amount of data required to encode the MV, thus eliminating redundancy and increasing compression. MV prediction can work effectively, for example, because when encoding an input video signal obtained from a camera (referred to as natural video), there is a statistical likelihood that there are larger regions moving in a similar direction than the region to which a single MV applies, and thus in some cases, similar motion vectors obtained from MVs of adjacent regions can be used for prediction. This results in the MV found for a given region being similar or identical to the MV predicted from surrounding MVs and, in turn, can be represented with fewer bits than the number of bits that would be used to directly encode the MV after entropy coding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., the sample stream). In other cases, the MV prediction itself may be lossy, for example, because of rounding errors when calculating the predictor from several surrounding MVs.

[0046] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, “High Efficiency Video Coding”, December 2016). Among the many MV prediction mechanisms provided by H.265, described herein is a technique hereinafter referred to as “spatial merge”.

[0047] Referring to Figure 2 , the current block (201) includes samples that the encoder has found during the motion search process to be predictable from a previous block of the same size that has been spatially shifted. Instead of directly encoding the MV, an MV associated with any one of five surrounding samples represented as A0, A1 and B0, B1, B2 (202 to 206 respectively) can be used, and the MV can be derived from metadata associated with one or more reference pictures, for example, from the most recent (in decoding order) reference picture. In H.265, MV prediction can use a predictor from the same reference picture that adjacent blocks are using.

[0048] Figure 3 FIG. shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In Figure 3 the example, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) may encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission via the network (350) to another terminal device (320). The encoded video data may be sent in the form of one or more encoded video bitstreams. The terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to recover the video pictures, and display the video pictures based on the recovered video data. Unidirectional data transmission may be common in media service applications and the like.

[0049] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data that may occur, for example, during a video conference. For bidirectional transmission of data, in the example, each of the terminal devices (330) and (340) may encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission via the network (350) to the other of the terminal devices (330) and (340). Each of the terminal devices (330) and (340) may also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), and may decode the encoded video data to recover the video pictures, and may display the video pictures at an accessible display device based on the recovered video data.

[0050] In Figure 3In the examples, the terminal devices (310), (320), (330), and (40) may be illustrated as servers, personal computers, and smart phones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure may be applied to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) represents any number of networks that transfer encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the current discussion, unless otherwise stated herein below, the architecture and topology of the network (350) may be unimportant for the operation of the present disclosure.

[0051] Figure 4 Shows the arrangement of a video encoder and a video decoder as an example of an application of the disclosed subject matter in a streaming environment. The disclosed subject matter may equally apply to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0052] A streaming system may include a capture subsystem (413), which may include a video source (401), such as a digital camera, that creates, for example, an uncompressed video picture stream (402). In an example, the video picture stream (402) includes samples taken by the digital camera. The video picture stream (402), depicted as a thick line to emphasize the high data volume when compared to the encoded video data (404) (or encoded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to implement or realize aspects of the disclosed subject matter described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize the lower data volume when compared to the video picture stream (402), may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as Figure 4The client subsystems (406) and (408) therein can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the input copy (407) of the encoded video data and creates an output stream of video pictures (411) that can be presented on a display (412) (e.g., a display screen) or other rendering device (not depicted). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In an example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.

[0053] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can also include a video encoder (not shown).

[0054] Figure 5 A block diagram of a video decoder (510) according to an embodiment of the present disclosure is shown. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used to replace Figure 4 the video decoder (410) in the example.

[0055] A receiver (531) may receive one or more encoded video sequences to be decoded by a video decoder (510); in the same or another embodiment, one encoded video sequence at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive encoded video data with other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver (531) may separate the encoded video sequences from the other data. To counter network jitter, a buffer memory (515) may be coupled between the receiver (531) and an entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other embodiments, the buffer memory may be external to the video decoder (510) (not depicted). In other embodiments, there may be a buffer memory (not depicted) external to the video decoder (510), for example to counter network jitter, and additionally there may be another buffer memory (515) inside the video decoder (510), for example to handle playout timing. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be needed, or the buffer memory may be small. For use on a best-effort packet network such as the Internet, a buffer memory (515) may be needed, which may be relatively large, and may advantageously have an adaptive size, and may be implemented at least partially in an operating system or similar element (not depicted) external to the video decoder (510).

[0056] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510), and potentially include information for controlling a presentation device, which is a presentation device (512) such as a display screen that is not a component of the electronic device (530) but may be coupled to the electronic device (530), as Figure 5As shown. The control information for presenting the device can be in the form of supplementary enhancement information (SEI message) or a video usability information (VUI) parameter set segment (not depicted). The parser (520) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technology or standard and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup can include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser (520) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0057] The parser (520) can perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (515) to create symbols (521).

[0058] Depending on the type of the encoded video picture or a part thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) can involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed by the parser (520) from the encoded video sequence. For clarity, the flow of such subgroup control information between the parser (520) and the multiple units below is not depicted.

[0059] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated into each other. However, for the purpose of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0060] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives the quantized transform coefficients and control information as symbols (521) from the parser (520), and the control information includes what kind of transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block including sample values that can be input to the aggregator (555).

[0061] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks; that is: blocks that do not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the surrounding reconstructed information obtained from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (555) adds the predictive information that the intra-prediction unit (552) has generated to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.

[0062] In other cases, the output samples of the scaler / inverse transform unit (551) may relate to inter-coded and possibly motion-compensated blocks. In such cases, the motion-compensation prediction unit (553) may access the reference picture memory (557) to obtain samples for prediction. After the obtained samples are motion-compensated according to the symbol (521) associated with the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (referred to as residual samples or residual signals in this case) in order to generate output sample information. The address within the reference picture memory (557) from which the motion-compensation prediction unit (553) obtains the prediction samples may be controlled by a motion vector, which may be used by the motion-compensation prediction unit (553) in the form of a symbol (521) that may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values obtained from the reference picture memory (557) when sub-sample accurate motion vectors are in use, motion vector prediction mechanisms, etc.

[0063] The output samples of the aggregator (555) may be subject to various loop-filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filter techniques, which are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and are available to the loop filter unit (556) as symbols (521) from the parser (520), but may also respond to meta-information obtained during the decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0064] The output of the loop filter unit (556) may be a sample stream, which may be output to the rendering device (512) and stored in the reference picture memory (557) for use in future inter-picture prediction.

[0065] Once certain encoded pictures are fully reconstructed, they can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture has been identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting to reconstruct the subsequent encoded picture.

[0066] The video decoder (510) can perform decoding operations according to a predetermined video compression technique in a standard (e.g., ITU-T Recommendation H.265). In the sense that the encoded video sequence complies with both the syntax of the video compression technique or standard and the profile described in the video compression technique or standard, the encoded video sequence can conform to the syntax specified by the video compression technique or standard used. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available according to that profile. Also required for compliance is that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted by the Hypothetical Reference Decoder (HRD) specification and the metadata signaled in the encoded video sequence for HRD buffer management.

[0067] In an embodiment, the receiver (531) can receive additional (redundant) data of the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0068] Figure 6 A block diagram of a video encoder (603) according to an embodiment of the present disclosure is shown. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used to replace Figure 4 the video encoder (403) in the example of

[0069] The video encoder (603) can receive from a video source (601) which is notFigure 6 Part of the electronic device (620) in the example receives video samples, and the video source can capture video images to be encoded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0070] The video source (601) can provide a source video sequence to be encoded by the video encoder (603) in the form of a digital video sample stream, which can have any appropriate bit depth (e.g., 8-bit, 10-bit, 12-bit,...), any color space (e.g., BT.601 Y CrCb, RGB,...), and any appropriate sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) can be a storage device storing pre-prepared videos. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures that give the appearance of motion when viewed in sequence. The pictures themselves can be organized as a spatial array of pixels, where each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0071] According to an embodiment, the video encoder (603) can encode and compress pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Enforcing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to other functional units. For clarity, this coupling is not depicted. The parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, λ value of rate distortion optimization techniques,...), picture size, group of picture (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a specific system design.

[0072] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an overly simplified description, in an example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols in a similar manner as the (remote) decoder would also create (since any compression between the symbol stream and the encoded video bitstream is lossless in the video compression techniques contemplated in the disclosed subject matter) to create sample data. The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" exactly the same sample values as the reference picture samples that the decoder would "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and drift in case the synchronization cannot be maintained, e.g., due to channel errors) is also used in some related arts.

[0073] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder, which is, for example, the video decoder (510) already incorporated Figure 5 as detailed above. However, briefly referring to Figure 5 , since the symbols are available and encoding / decoding the symbols into an encoded video sequence by the entropy encoder (645) and the parser (520) can be lossless, the entropy decoding part of the video decoder (510) (including the buffer memory (515) and the parser (520)) may not be fully implemented in the local decoder (633).

[0074] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technologies can be simplified since they are the reverse of the decoder technologies described comprehensively. More detailed descriptions are only needed and provided in certain areas below.

[0075] During operation, in some examples, the source encoder (630) may perform motion-compensated predictive coding that predictively encodes an input picture by referring to one or more previously encoded pictures designated as "reference pictures" from a video sequence. In this way, the encoding engine (632) encodes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.

[0076] The local video decoder (633) may decode the encoded video data of pictures that may be designated as reference pictures based on the symbols created by the source encoder (630). The operation of the encoding engine (632) may advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 6 not shown), the reconstructed video sequence may generally be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference pictures, and may cause the reconstructed reference pictures to be stored in the reference picture buffer (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference pictures, which has the same content as the reconstructed reference pictures that will be obtained by the remote video decoder (in the absence of transmission errors).

[0077] The predictor (635) may perform a prediction search on the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) that may serve as an appropriate prediction reference for the new picture or specific metadata such as reference picture motion vectors, block shapes, etc. The predictor (635) may operate block by block on the sample blocks to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references obtained from multiple reference pictures stored in the reference picture memory (634).

[0078] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting the parameters and subgroups of parameters for encoding the video data.

[0079] The outputs of all the foregoing functional units may undergo entropy encoding in the entropy encoder (645). The entropy encoder (645) converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0080] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) to prepare for transmission via the communication channel (660), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) may merge the encoded video data from the video encoder (603) with other data to be transmitted (e.g., encoded audio data and / or auxiliary data streams (source not shown)).

[0081] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain type of encoded picture to each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures may generally be assigned to one of the following picture types:

[0082] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of those variants of I pictures and their respective applications and characteristics.

[0083] A predictive picture (P picture) may be a picture that is encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0084] A bi-predictive picture (B picture) may be a picture that is encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.

[0085] Source pictures may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded block by block. Blocks may be encoded predictively with reference to other (already encoded) blocks determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture may be encoded non-predictively, or they may be encoded predictively with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be encoded predictively with reference to a single previously encoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture may be encoded predictively with reference to one or two previously encoded reference pictures via spatial prediction or via temporal prediction.

[0086] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard (e.g., ITU-T Recommendation H.265). In its operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0087] In an embodiment, the transmitter (640) may send additional data together with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set segments, etc.

[0088] Video may be captured as a sequence of multiple source pictures (video pictures) in time. Intra-picture prediction (usually abbreviated as intra prediction) exploits the spatial correlation within a given picture, and inter-picture prediction exploits the (temporal or other) correlation between pictures. In an example, a particular picture being encoded / decoded, which will be referred to as the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously encoded and still buffered reference picture in the video, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, may have a third dimension identifying the reference picture.

[0089] In some embodiments, bidirectional prediction techniques may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, e.g., a first reference picture and a second reference picture, both of which are before the current picture in the video in decoding order (but may be in the past and future respectively in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0090] In addition, merge mode techniques may be used in inter-picture prediction to improve coding efficiency.

[0091] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In an example, each CU is analyzed to determine the prediction type of the CU, such as an inter-prediction type or an intra-prediction type. Depending on the temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, prediction operations in encoding (encoding / decoding) are performed in units of prediction blocks. Using a luminance prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luminance values) of pixels (e.g., 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.).

[0092] Figure 7 FIG. shows a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence and encode the processing block as an encoded picture that is part of an encoded video sequence. In an example, the video encoder (703) is used instead of Figure 4 the video encoder (403) in the example of.

[0093] In the HEVC example, a video encoder (703) receives a matrix of sample values for a processing block (e.g., a prediction block of 8×8 samples, etc.). The video encoder (703) determines, for example, using rate-distortion optimization, whether to encode the processing block optimally using an intra mode, an inter mode, or a bi-predictive mode. When the processing block is to be encoded in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded picture; and when the processing block is to be encoded in the inter mode or the bi-predictive mode, the video encoder (703) may use inter prediction or bi-prediction techniques, respectively, to encode the processing block into an encoded picture. In some video coding techniques, the merge mode may be an inter-picture prediction sub-mode, where a motion vector is obtained from one or more motion vector predictors without resorting to encoded motion vector components external to the predictor. In some other video coding techniques, there may be motion vector components applicable to the subject block. In the example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0094] In Figure 7 the example of Figure 7 shown, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as

[0095] The inter encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter prediction information (e.g., a description of redundant information according to inter coding techniques, a motion vector, merge mode information), and calculate an inter prediction result (e.g., a predicted block) based on the inter prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information.

[0096] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with blocks that have already been encoded in the same picture in some cases, generate quantization coefficients after transformation, and also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques) in some cases. In the example, the intra encoder (722) also calculates an intra prediction result (e.g., a predicted block) based on reference blocks in the same picture and the intra prediction information.

[0097] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In an example, the general controller (721) determines the mode of a block and provides a control signal to the switch (726) based on the mode. For example, when the mode is the intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; and when the mode is the inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.

[0098] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In an example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain and generate transform coefficients. Then, the transform coefficients are quantized to obtain quantized transform coefficients. In various embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the inter prediction information. In some examples, the decoded block is appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.

[0099] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information according to a suitable standard such as the HEVC standard. In an example, the entropy encoder (725) is configured to include general control data, the selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. Note that according to the disclosed subject matter, there is no residual information when encoding a block in the merge submode of the inter mode or the bi - directional prediction mode.

[0100] Figure 8FIG. shows a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (810) is used instead of Figure 4 the video decoder (410) in the example of

[0101] In Figure 8 the example of Figure 8 the video decoder (810) includes an entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) coupled together as shown.

[0102] The entropy decoder (871) may be configured to reconstruct certain symbols representing the syntax elements that make up the encoded picture from the encoded picture. Such symbols may include, for example, the mode for encoding a block (such as an intra mode, an inter mode, a bi-predictive mode, a merge sub-mode or another sub-mode of the latter two), prediction information that can identify certain samples or metadata for prediction by the intra-frame decoder (872) or the inter-frame decoder (880) respectively (such as intra prediction information or inter prediction information), residual information in the form of quantized transform coefficients, etc. In an example, when the prediction mode is an inter mode or a bi-predictive mode, the inter prediction information is provided to the inter-frame decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra-frame decoder (872). The residual information may be inverse quantized and provided to the residual decoder (873).

[0103] The inter-frame decoder (880) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.

[0104] The intra-frame decoder (872) is configured to receive the intra prediction information and generate a prediction result based on the intra prediction information.

[0105] The residual decoder (873) is configured to perform inverse quantization to extract the dequantized transform coefficients and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include Quantizer Parameter (QP)), and this information may be provided by the entropy decoder (871) (the data path is not labeled as this is only low volume control information).

[0106] The reconstruction module (874) is configured to combine, in the spatial domain, the residual output by the residual decoder (873) and the prediction result (output by the inter-frame prediction module or the intra-frame prediction module as the case may be) to form a reconstructed block, which may be part of a reconstructed picture, which in turn may be part of a reconstructed video. It should be noted that other suitable operations, such as deblocking operations, etc., may be performed to improve the visual quality.

[0107] Note that any suitable technique may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In an embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more processors that execute software instructions.

[0108] Neural network techniques can be used together with video coding techniques, and video coding techniques that utilize neural networks can be referred to as hybrid video coding techniques. For example, a loop filter unit such as the loop filter unit (556) can apply various loop filters for sample filtering. One or more loop filters can be implemented by a neural network. Aspects of the present disclosure provide in-loop filtering techniques in hybrid video coding techniques for improving picture quality using neural networks. Specifically, according to one aspect of the present disclosure, a technique of cropping data can be used before feeding the data to the core of a neural network-based in-loop filter.

[0109] According to one aspect of the present disclosure, a loop filter is a filter that affects reference data. For example, the image filtered by the loop filter unit (556) is stored in a buffer, such as the reference picture memory (557), as a reference for further prediction. The in-loop filter can improve the video quality in a video codec.

[0110] Figure 9 A block diagram of a loop filter unit (900) in some examples is shown. In the example, the loop filter unit (900) may be used instead of the loop filter unit (556). In Figure 9In an example, the loop filter unit (900) includes a deblocking filter (901), a sample adaptive offset (SAO) filter (902), and an adaptive loop filter (ALF) filter (903). In some examples, the ALF filter (903) may include a cross component adaptive loop filter (CCALF).

[0111] During operation, in an example, the loop filter unit (900) receives a reconstructed picture, applies various filters to the reconstructed picture, and generates an output picture in response to the reconstructed picture.

[0112] In some examples, the deblocking filter (901) and the SAO filter (902) are configured to remove block artifacts introduced when using block coding techniques. The deblocking filter (901) may smooth the shaped edges formed when using block coding techniques. The SAO filter (902) may apply a specific offset to samples to reduce distortion relative to other samples in a video frame. The ALF (903) may apply a classification to a block of samples, for example, and then apply a filter associated with the classification to the block of samples. In some examples, the filter coefficients of the filter may be determined by an encoder and signaled to a decoder.

[0113] In some examples (e.g., JVET-T0057), an additional filter called a dense residual convolutional neural network based in-loop filter (DRNLF) may be inserted between the deblocking filter (901) and the SAO filter (902). The DRNLF may further improve the image quality.

[0114] Figure 10 A block diagram of a loop filter unit (1000) in some examples is shown. In an example, the loop filter unit (1000) may be used instead of the loop filter unit (556). In Figure 10 an example, the loop filter unit (1000) includes a deblocking filter (1001), an SAO filter (1002), an ALF filter (1003), and a DRNLF filter (1010) located between the deblocking filter (1001) and the SAO filter (1002).

[0115] The deblocking filter (1001) is similarly configured as the deblocking filter (901), the SAO filter (1002) is similarly configured as the SAO filter (902), and the ALF filter (1003) is similarly configured as the ALF filter (903).

[0116] The DRNLF filter (1010) receives the output of the deblocking filter (1001) shown by the deblocked picture (1011), and also receives the quantization parameter (QP) map of the reconstructed picture. The QP map includes the quantization parameters of the blocks in the reconstructed picture. The DRNLF filter (1010) may output a picture shown by the filtered picture (1019) with improved quality, and the filtered picture (1019) is fed to the SAO filter (1002) for further filtering processing.

[0117] According to an aspect of the present invention, a neural network for video processing may include a plurality of channels for processing color components in a color space. In an example, the YCbCr model may be used to define the color space. In the YCbCr model, Y represents the luminance component (brightness), and Cb and Cr represent the chrominance components. It should be noted that in the following description, YUV is used to describe the format encoded using the YCbCr model.

[0118] According to an aspect of the present disclosure, a plurality of channels in a neural network are configured to operate on color components of the same size. In some examples, a picture may be represented by color components of different sizes. For example, the human visual system is much more sensitive to changes in brightness than to changes in color. Therefore, a video system may compress the chrominance components to reduce the file size and save transmission time without creating a large visual difference perceptible to the human eye. In some examples, taking advantage of the fact that the human visual system is less sensitive to color differences than to brightness, chroma subsampling techniques are used to achieve a lower resolution of chrominance information than that of luminance information.

[0119] In some examples, subsampling can be represented as a three-part ratio, such as 4:4:4, 4:2:0, 4:2:2, 4:1:1, etc. For example, 4:4:4 (also known as YUV444) indicates that each of the YCbCr components has the same sampling rate without subsampling; 4:2:0 (also known as YUV420) indicates that the chrominance components are subsampled, and every four pixels (or Y components) can correspond to a Cb component and a Cr component. It should be noted that YUV420 is used as an example of the subsampling format in the following description to illustrate the techniques disclosed in the present disclosure. The disclosed techniques can be used for other subsampling formats. For ease of description, a format of color components with the same sampling rate without subsampling (e.g., YUV444) is referred to as a non-subsampled format; and a format with at least one color component that is subsampled (e.g., YUV420, YUV422, YUV411, etc.) is referred to as a subsampled format.

[0120] Generally, a neural network can operate on pictures in a non-subsampled format (e.g., YUV444). Therefore, for pictures in a subsampled format, the pictures are converted to a non-subsampled format before being provided as input to the neural network.

[0121] Figure 11 A block diagram of a DRNLF filter (1100) in some examples is shown. In the example, the DRNLF filter (1100) can be used instead of the DRNLF filter (1010). The DRNLF filter (1100) includes a QP mapping quantizer (1110), a preprocessing module (1120), a main processing module (1130), and a postprocessing module (1140) coupled together as Figure 11 shown. The main processing module (1130) includes a block grabber (1131), a block-based DRNLF kernel processing module (1132), and a block reorganizer (1133) coupled together as Figure 11 shown.

[0122] In some examples, the QP mapping includes a mapping of QP values applied to each block in the currently reconstructed picture. The QP mapping quantizer (1110) can quantize values into a set of predetermined values. In an example (e.g., JVET-T0057), the QP value can be quantized to one of 22, 27, 32, and 37 by the QP mapping quantizer (1110).

[0123] The preprocessing module (1120) can receive the deblocked picture in the first format and convert it into the second format used by the main processing module (1130). For example, the main processing module (1130) is configured to process pictures with the YUV444 format. When the preprocessing module (1120) receives a deblocked picture in a format different from the YUV444 format, the preprocessing module (1120) can process the deblocked picture in the different format and output a deblocked picture in the YUV444 format. For example, the preprocessing module (1120) receives a deblocked picture in the YUV420 format and then interpolates the U chrominance channel and the V chrominance channel horizontally and vertically by a factor of 2 to generate a deblocked picture in the YUV444 format.

[0124] The main processing module (1130) can receive the deblocked picture in the YUV444 format and the quantized QP map as inputs. The block grabber (1131) splits the inputs into blocks. The DRNLF kernel processing module (1132) can process each block based on the DRNLF kernel. The block reassembler (1133) can assemble the blocks processed by the DRNLF kernel processing module (1132) into a filtered picture in the YUV444 format.

[0125] The postprocessing module (1140) converts the filtered picture in the second format back to the first format. For example, the postprocessing module (1140) receives the filtered picture in the YUV444 format (output from the main processing module (1130)) and outputs a filtered picture in the YUV420 format.

[0126] Figure 12 A block diagram of the preprocessing module (1220) in some examples is shown. In the example, the preprocessing module (1220) is used in place of the preprocessing module (1120).

[0127] The preprocessing module (1220) can receive a deblocked picture in the YUV420 format, convert the deblocked picture into the YUV444 format and output a deblocked picture in the YUV444 format. Specifically, the preprocessing module (1220) receives the deblocked picture in three input channels, the three input channels including a luminance input channel for the Y component and two chrominance input channels for the U (Cb) component and the V (Cr) component respectively. The preprocessing module (1220) outputs the deblocked picture through three output channels, the three output channels including a luminance output channel for the Y component and two chrominance output channels for the U (Cb) component and the V (Cr) component respectively.

[0128] In an example, when the deblocked picture has a YUV420 format, the Y component has a size of (H, W), the U component has a size of (H / 2, W / 2), and the V component has a size of (H / 2, W / 2), where H represents the height of the deblocked picture (e.g., in samples) and W represents the width of the deblocked picture (e.g., in samples).

[0129] In Figure 12 the example of, the preprocessing module (1220) does not resize the Y component. The preprocessing module (1220) receives the Y component of size (H, W) from the luminance input channel and outputs the Y component of size (H, W) to the luminance output channel.

[0130] The preprocessing module (1220) resizes the U component and the V component respectively. The preprocessing module (1220) includes a first resizing unit (1221) and a second resizing unit (1222) that process the U component and the V component respectively. For example, the first resizing unit (1221) receives the U component of size (H / 2, W / 2), resizes the U component to a size of (H, W), and outputs the U component of size (H, W) to the chrominance output channel for the U component. The second resizing unit (1222) receives the V component of size (H / 2, W / 2), resizes the V component to a size of (H, W), and outputs the V component of size (H, W) to the chrominance output channel for the V component. In some examples, the first resizing unit (1221) resizes the U component based on interpolation, for example, using a Lanczos interpolation filter. Similarly, in some examples, the second resizing unit (1222) resizes the V component based on interpolation, for example, using a Lanczos interpolation filter.

[0131] In some examples, an interpolation operation such as using a Lanczos interpolation filter cannot guarantee that the output of the interpolation operation is a meaningful value, for example, non - negative for meaningful U (Cb) and V (Cr) components. In some examples, the deblocked picture in YUV444 format after preprocessing can be stored and then the stored YUV444 - format picture can be used during the training process of the neural network. Negative values of the U (Cb) and V (Cr) components can adversely affect the result of the training process of the neural network.

[0132] Figure 13A block diagram showing a neural network structure (1300) is presented. In some examples, the neural network structure (1300) is a dense residual convolutional neural network based in-loop filter (DRNLF) and can be used to replace the block-based DRNLF kernel processing module (1132). The neural network structure

[0133] (1300) includes a series of dense residual units (DRUs), such as DRU

[0134] (1301) to DRU (1304), and the number of DRUs is represented by N. In Figure 13 it, the number of convolutional kernels is represented by M, and M is also the number of output channels used for convolution. For example,

[0135] "CONV 3×3×M" indicates a standard convolution with M convolutional kernels of kernel size 3×3,

[0136] "DSC 3×3×M" indicates a depthwise separable convolution with M convolutional kernels of kernel size 3×3. N and M can be set for the trade-off between computational efficiency and performance. In an example (e.g., JVET-T0057), N is set to 4 and M is set to 32.

[0137] During operation, the neural network structure (1300) processes the deblocked picture by patches. For each patch of the deblocked picture in YUV444 format, the patch is normalized (e.g., divided by 1023 in the Figure 13 example), and the mean of the deblocked picture is removed from the normalized patch to obtain the first part (1311) of the internal input (1313). The second part of the internal input (1313) comes from the QP mapping. For example, a patch of the QP mapping corresponding to the patch forming the first part (1311) (referred to as the QP mapping patch) is obtained from the QP mapping. The QP mapping patch is normalized (e.g., divided by 51 in Figure 13 it). The normalized QP mapping patch is the second part (1312) of the internal input (1313). The second part (1312) is concatenated with the first part (1311) to obtain the internal input (1313). The internal input (1313) is provided to the first regular convolution block (1351) (represented by CONV 3×3×M). Then, the output of the first regular convolution block (1351) is processed by N DRUs.

[0138] For each DRU, receive and process the intermediate input. Concatenate the output of the DRU with the intermediate input to form the intermediate input for the next DRU. Using DRU(1302) as an example, DRU(1302) receives the intermediate input (1321), processes the intermediate input (1321) and generates the output (1322). Concatenate the output (1322) with the intermediate input (1321) to form the intermediate input (1323) for DRU(1303).

[0139] Note that due to the reason that the intermediate input (1321) has more than M channels, a convolution operation of "CONV1×1×M" can be applied to the intermediate input (1321) to generate M channels for further processing by DRU(1302). Also note that the output of the first regular convolution block (1351) includes M channels, so this output can be processed by DRU(1301) without using the convolution operation of "CONV1×1×M".

[0140] The output of the last DRU is provided to the last regular convolution block (1359). For example, by adding the means of the deblocked pictures and multiplying by 1023, the output of the last regular convolution block (1359) is converted into a regular picture block value, as Figure 13 shown.

[0141] Figure 14 The block diagram of a dense residual unit (DRU)(1400) is shown. In some examples, DRU(1400) can be used to replace Figure 13 each DRU in, such as DRU(1301), DRU(1302), DRU(1303) and DRU(1304).

[0142] In Figure 14 the example, DRU(1400) receives the intermediate input x, and directly propagates the intermediate input x to the subsequent DRU through the shortcut (1401). DRU(1400) also includes a regular processing path (1402). In some examples, the regular processing path (1402) includes a regular convolution layer (1411), depthwise separable convolution (DSC) layers (1412) and (1414), and a rectified linear unit (ReLU) layer (1413). For example, concatenate the intermediate input x with the output of the regular processing path (1402) to form the intermediate input for the subsequent DRU.

[0143] In some examples, the DSC layers (1412) and (1414) are used to reduce the computational cost.

[0144] According to one aspect of the present disclosure, a neural network structure (1300) includes three channels corresponding to Y, U (Cb), and V (Cr) components respectively. In some examples, these three channels may be referred to as the Y channel, the U channel, and the V channel. The DRNLF filter (1100) may be applied to intra pictures and inter pictures. In some examples, an additional flag is signaled to indicate the on / off of the DRNLF filter (1100) at the picture level and the CTU level.

[0145] Figure 15 A block diagram of a post - processing module (1540) in some examples is shown. In the example, the post - processing module (1540) may be used instead of the post - processing module (1140). The post - processing module (1540) includes cropping units (1541) to (1543) that crop the values of the Y component, the U component, and the V component respectively to a predetermined non - negative range [a, b]. In the example, the lower limit a and the upper limit b of the non - negative range may be set to a = 16×4 and b = 234×4. In addition, the post - processing module (1540) includes resizing units (1545) and resizing units (1546) that resize the cropped U component and V component from size (H, W) to size (H / 2, W / 2) respectively, where H is the height of the original picture (e.g., the de - blocked picture) and W is the width of the original picture.

[0146] Aspects of the present disclosure provide pre - processing techniques. The pre - processed data can be stored and used for the training of neural networks, and better training and inference results can be obtained.

[0147] Figure 16 A block diagram of a pre - processing module (1620) in some examples is shown. In the example, the pre - processing module (1620) is used instead of the pre - processing module (1120).

[0148] The pre - processing module (1620) may receive a de - blocked picture in YUV420 format, convert the de - blocked picture to YUV444 format and output the de - blocked picture in YUV444 format. Specifically, the pre - processing module (1620) receives the de - blocked picture in three input channels, the three input channels including a luminance input channel for the Y component and two chrominance input channels for the U (Cb) component and the V (Cr) component respectively. The pre - processing module (1620) outputs the de - blocked picture through three output channels, the three output channels including a luminance output channel for the Y component and two chrominance output channels for the U component and the V component respectively.

[0149] In an example, when the deblocked picture has a YUV420 format, the Y component has a size (H, W), the U component has a size (H / 2, W / 2), and the V component has a size (H / 2, W / 2), where H represents the height of the deblocked picture (e.g., in samples) and W represents the width of the deblocked picture (e.g., in samples).

[0150] In Figure 16 the example, the preprocessing module (1620) does not resize the Y component. The preprocessing module (1620) receives a Y component of size (H, W) from the luminance input channel and outputs a Y component of size (H, W) to the luminance output channel.

[0151] The preprocessing module (1620) resizes the U component and the V component respectively. The preprocessing module 1620 includes a first resizing unit (1621) and a second resizing unit (1622) that process the U component and the V component respectively. For example, the first resizing unit (1621) receives a U component of size (H / 2, W / 2), resizes the U component to size (H, W), and outputs a U component of size (H, W) to the chrominance output channel for the U component. The second resizing unit (1622) receives a V component of size (H / 2, W / 2), resizes the V component to size (H, W), and outputs a V component of size (H, W) to the chrominance output channel for the V component. In some examples, the first resizing unit (1621) resizes the U component based on interpolation, for example, using a Lanczos interpolation filter. Similarly, in some examples, the second resizing unit (1622) resizes the V component based on interpolation, for example, using a Lanczos interpolation filter.

[0152] In some examples, an interpolation operation such as using a Lanczos interpolation filter does not guarantee that the output of the interpolation operation is a meaningful value, such as being non - negative for meaningful U (Cb) and V (Cr) components.

[0153] In Figure 16 the example, the preprocessing module (1620) includes a clipping unit (1625) and a clipping unit (1626) to clip the values of the interpolated U component and V component to the range [c, d] respectively. In some examples, if the values of the Y component, U component, and V component for preprocessing have a bit depth of 10, then c and d can be set to c = 0 and d = 2 bitdepth - 1 = 1023.

[0154] In an example, c values and d values are predefined and used. In another example, multiple pairs of c values and d values are predefined, and indices of the pairs of c values and d values for cropping can be signaled in a bitstream (e.g., in a sequence parameter set (SPS), a picture parameter set (PPS), a slice or a tile header).

[0155] In some examples, the cropped values of the U and V components and the value of the Y component can be stored as a deblocked picture in YUV444 format. In some embodiments, the stored picture in YUV444 format can be used as an input during the training process of a neural network (e.g., the neural network in the main processing module (1130)). In some examples, the values of the U and V components are cropped to a range that will not adversely affect the training process of the neural network. In an example, the values of the U and V components are cropped to be non - negative.

[0156] In some examples, in the case of using a picture with cropped values stored in YUV444 format, the training of the neural network can be accelerated because time is saved by avoiding pre - processing (e.g., resizing, cropping) steps during training. Additionally, the neural network can be trained with better model parameters that can improve compression efficiency and / or picture quality.

[0157] In some examples, adding a cropping unit (1625) and a cropping unit (1626) in a pre - processing module (1620) can improve compression efficiency and / or quality, for example, at a lower Bjontegaard delta rate (BD - rate).

[0158] Figure 17 A flowchart outlining a process (1700) according to embodiments of the present disclosure is shown. The process (1700) can be used for video processing. In various embodiments, the process (1700) is executed by a processing circuit, such as the processing circuit in terminal devices (310), (320), (330), and (340), the processing circuit that performs the function of a video encoder (403), the processing circuit that performs the function of a video decoder (410), the processing circuit that performs the function of a video decoder (510), the processing circuit that performs the function of a video encoder (603), etc. In some embodiments, the process (1700) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (1700). The process starts at (S1701) and proceeds to (S1710).

[0159] At (S1710), a picture in a subsampling format in a color space is converted to a non - subsampling format in the color space. In some examples, the conversion is performed based on interpolation and may produce invalid values. In an example, the conversion may produce negative values that are invalid for the YCbCr model.

[0160] At (S1720), before providing a picture in a non - subsampling format as input to a neural - network - based filter, the values of one or more color components of the picture in the non - subsampling format are clipped. In some examples, one or more color components may be chrominance components. Then, the process proceeds to (S1799).

[0161] In an example, the values of the color components of a picture in a non - subsampling format are clipped to the valid range of that color component. In an example, the values of the color components of a picture in a non - subsampling format are clipped to be non - negative. In another example, the range is determined based on the bit depth. For example, the lower limit of the range is 0, and the upper limit of the range is set to (2 bitdepth ) - 1.

[0162] In some examples, the range is predefined. In some examples, the range is determined based on decoded information from the bitstream carrying the picture. In some examples, a signal indicating the range is decoded from at least one of a sequence parameter set, a picture parameter set, a slice header, and a tile header in the bitstream.

[0163] In an example, multiple ranges can be predefined. Then, an index indicating one of the multiple ranges can be carried in one of a sequence parameter set, a picture parameter set, a slice header, and a tile header in the bitstream.

[0164] In some examples, process (1700) is used in a decoder. For example, a picture in a subsampling format is reconstructed based on decoded information from the bitstream, and a de - blocking filter is applied to the picture in the subsampling format before converting the picture in the subsampling format to a non - subsampling format. In another example, a neural - network - based filter is applied to a picture in a non - subsampling format with clipped values to generate a filtered picture in the non - subsampling format, and then the filtered picture in the non - subsampling format is converted to a filtered picture in the subsampling format.

[0165] In some examples, a picture in a non - subsampling format with clipped values is stored in a storage device. Then, a picture in a non - subsampling format with clipped values and other pictures can be provided as training inputs to train the neural network in the neural - network - based filter.

[0166] It should be noted that the various units, blocks, and modules in the above description can be implemented by various techniques (such as processing circuits, processors executing software instructions, combinations of hardware and software, etc.).

[0167] The above techniques can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 18 A computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0168] Computer software can be encoded using any suitable machine code or computer language, which can undergo mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), Graphics Processing Units (GPUs), etc., or executed through interpretation, microcode execution, etc.

[0169] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0170] Figure 18 The components of the computer system (1800) shown are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependencies or requirements related to any one of the components or combinations thereof shown in the exemplary embodiments of the computer system (1800).

[0171] The computer system (1800) can include certain human-machine interface input devices. Such human-machine interface input devices can respond to inputs from one or more human users through, for example, tactile inputs (such as keystrokes, swipes, data glove movements), audio inputs (such as voice, clapping), visual inputs (such as gestures), olfactory inputs (not depicted). The human-machine interface devices can also be used to capture certain media that is not necessarily directly related to a human's conscious input, such as audio (such as voice, music, ambient sound), images (such as scanned images, photographic images obtained from a still image camera), video (such as two-dimensional video, three-dimensional video including stereoscopic video).

[0172] The input human-machine interface device may include one or more of the following (each depicted only one): keyboard (1801), mouse (1802), trackpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera (1808).

[0173] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (e.g., haptic feedback through the touch screen (1810), data glove (not shown), or joystick (1805), but there may also be haptic feedback devices that do not serve as input devices), audio output devices (such as: speakers (1809), headphones (not depicted)), visual output devices (such as a screen (1810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without haptic feedback capabilities - some of which may output two-dimensional visual output or more than three-dimensional output through devices such as stereoscopic graphics output, virtual reality glasses (not depicted), holographic displays, and smokeboxes (not depicted)), and printers (not depicted).

[0174] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1820) with media such as CD / DVD etc. (1821), thumb drives (1822), removable hard disk drives or solid state drives (1823), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0175] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the subject matter disclosed in the present invention does not include transmission media, carrier waves, or other transient signals.

[0176] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, urban, in-vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks such as Ethernet; wireless LAN; cellular networks including GSM (Global System for Mobile Communication), 3G (Third Generation), 4G (Fourth Generation), 5G (Fifth Generation), LTE (Long Term Evolution), etc.; television wired connections or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television; in-vehicle and industrial networks including CAN bus, etc. Some networks typically require an external network interface adapter, which is attached to some general-purpose data port or peripheral bus (1849) (e.g., the USB port of the computer system (1800)); other networks are typically integrated into the core of the computer system (1800) by attaching to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). Using any of these networks, the computer system (1800) can communicate with other entities. Such communication may be unidirectional receive-only (e.g., broadcast television), unidirectional transmit-only (e.g., CAN bus to some CAN bus devices), or bidirectional, e.g., to other computer systems using local or wide area digital networks. Certain protocols and protocol stacks may be used on each of these networks and network interfaces as described above.

[0177] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core (1840) of the computer system (1800).

[0178] The core (1840) may include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU) (1842), a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) (1843), a hardware accelerator for certain tasks (1844), a graphics adapter (1850), etc. These devices, together with read-only memory (ROM) (1845), random access memory (1846), an internal mass storage device such as an internal non-user-accessible hard disk drive, an SSD (1847), etc., may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessible in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1849) to the system bus (1848) of the core. In an example, a screen (1810) may be connected to the graphics adapter (1850). The architecture of the peripheral bus includes PCI, USB, etc.

[0179] The CPU (1841), GPU (1842), FPGA (1843), and accelerator (1844) may execute certain instructions that, when combined, may constitute the aforementioned computer code. The computer code may be stored in the ROM (1845) or RAM (1846). Interim data may also be stored in the RAM (1846), while permanent data may be stored in, for example, the internal mass storage device (1847). Fast storage and retrieval of any memory device may be achieved by using a cache memory that may be closely associated with one or more CPUs (1841), GPUs (1842), the mass storage device (1847), ROM (1845), RAM (1846), etc.

[0180] Computer-readable media may have computer code for performing various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or they may be of the type well known and available to those skilled in the computer software art.

[0181] By way of example and not limitation, a computer system (1800) having an architecture, and in particular a core (1840), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as described above and specific storage devices of the core (1840) having non-transitoriness such as an on-core mass storage device (1847) or a ROM (1845). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1840). Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core (1840) and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in a RAM (1846) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of logic hardwired or otherwise included in a circuit (e.g., an accelerator (1844)), which can operate instead of or in conjunction with the software to execute specific processes or specific portions of specific processes described herein. In appropriate cases, references to software can include logic and vice versa. In appropriate cases, references to computer-readable media can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit including logic for execution, or both a circuit storing software for execution and a circuit including logic for execution. The present disclosure encompasses any suitable combination of hardware and software.

[0182] Appendix A: Abbreviations

[0183] JEM: Joint Exploration Model

[0184] VVC: Versatile Video Coding

[0185] BMS: Benchmark Set

[0186] MV: Motion Vector

[0187] HEVC: High Efficiency Video Coding

[0188] SEI: Supplemental Enhancement Information

[0189] VUI: Video Usability Information

[0190] GOP: Group of Pictures

[0191] TU: Transform Unit

[0192] PU: Prediction Unit

[0193] CTU: Coding Tree Unit

[0194] CTB: Coding Tree Block

[0195] PB: Prediction Block

[0196] HRD: Hypothetical Reference Decoder

[0197] SNR: Signal-to-Noise Ratio

[0198] CPU: Central Processing Unit

[0199] GPU: Graphics Processing Unit

[0200] CRT: Cathode Ray Tube

[0201] LCD: Liquid Crystal Display

[0202] OLED: Organic Light-Emitting Diode

[0203] CD: Compact Disc

[0204] DVD: Digital Video Disc

[0205] ROM: Read-Only Memory

[0206] RAM: Random Access Memory

[0207] ASIC: Application-Specific Integrated Circuit

[0208] PLD: Programmable Logic Device

[0209] LAN: Local Area Network

[0210] GSM: Global System for Mobile Communications

[0211] LTE: Long-Term Evolution

[0212] CANBus: Controller Area Network Bus

[0213] USB: Universal Serial Bus

[0214] PCI: Peripheral Component Interconnect

[0215] FPGA: Field-Programmable Gate Array

[0216] SSD: Solid State Drive

[0217] IC: Integrated Circuit

[0218] CU: Coding Unit

[0219] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method for video processing, characterized in that, The method includes: A preprocessing module of a loop filter (DRNLF) based on a dense residual convolutional neural network in a processing circuit converts a picture in a subsampled format in a color space into a non-subsampled format in the color space; and Before providing the non-subsampled format picture as an input to a neural network-based filter, a preprocessing module in the processing circuit determines a range for cropping based on information decoded from a bitstream containing the picture, and crops the values of the color components of the non-subsampled format picture to the determined range.

2. The method according to claim 1, wherein The method further includes: Cropping the values of the color components of the non-subsampled format picture to the valid range of the color components.

3. The method according to claim 1, wherein The method further includes: Cropping the values of the color components of the non-subsampled format picture to a range determined based on the bit depth.

4. The method according to claim 1, wherein The method further includes: Cropping the values of the color components of the non-subsampled format picture to a predetermined range.

5. The method according to claim 1, characterized in that, The method further includes: Decoding a signal indicating the range from at least one of a sequence parameter set, a picture parameter set, a slice header, and a tile header in the bitstream.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Reconstructing the subsampled format picture based on the decoded information from the bitstream.

7. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Applying a neural network-based filter to the non-subsampled format picture with the cropped values to generate a filtered picture in the non-subsampled format; and Converting the filtered picture in the non-subsampled format to a filtered picture in the subsampled format.

8. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Storing the non-subsampled format picture with the cropped values.

9. The method according to claim 8, characterized in that The method further includes: Providing the stored picture in the non-subsampled format with the cropped values as a training input to train the neural network in the neural network-based filter.

10. A computer device, characterized in that, The computer device includes: One or more computer-readable non-transitory storage media configured to store computer program code; and One or more computer processors configured to access the computer program code and execute the method according to any one of claims 1 to 9 according to the instructions of the computer program code.

11. A device for video processing, characterized in that, The device includes a processing circuit configured to execute the method according to any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer program code, and when the computer program code is executed by at least one processor, the at least one processor executes the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Filtering image data using a neutral network

    DE102018101030A1

  • Enhancing performance capture with real-time neural rendering

    WO2020117657A1

  • Filtering method and device, encoder and computer storage medium

    WO2020192020A1

  • Multi-parameter adaptive loop filtering in video processing

    WO2020192645A1