System and method for over-parameterized convolution by a neural-network based in-loop filter
Patent Information
- Application Number
- PCT/CN2025/083563
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2026-09-24
Smart Images

Figure CN2025083563_24092026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR OVER-PARAMETERIZED CONVOLUTION BY A NEURAL-NETWORK BASED IN-LOOP FILTERBACKGROUND
[0001] Embodiments of the present disclosure relate to video coding.
[0002] Digital video has become mainstream and is being used in a wide range of applications including digital television, video telephony, and teleconferencing. These digital video applications are feasible because of the advances in computing and communication technologies as well as efficient video coding techniques. Various video coding techniques may be used to compress video data, such that coding on the video data can be performed using one or more video coding standards. Exemplary video coding standards may include, but not limited to, versatile video coding (H. 266 / VVC) , high-efficiency video coding (H. 265 / HEVC) , advanced video coding (H. 264 / AVC) , moving picture expert group (MPEG) coding, to name a few.SUMMARY
[0003] According to one aspect of the present disclosure, a method of video decoding is provided. The method may include obtaining, by a processor, a first reference picture from a decoded picture buffer. The method may include generating, by the processor, a first feature map based on first information associated with the first reference picture. The method may include generating, by the processor, a residual feature map based on at least one Over-Parameterized Convolution (OPC) of the first feature map using a filter. The method may include generating, by the processor, a second feature map based on the residual feature map. The method may include generating, by the processor, a second reference picture based on the first reference picture and the second feature map. The method may include decoding, by the processor, a current picture region based on the second reference picture.
[0004] According to another aspect of the present disclosure, a decoder is provided. The decoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a residual feature map based on at least one OPC of the first feature map using a filter. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the residual feature map. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the first reference picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to decode a current picture region based on the second reference picture.
[0005] According to another aspect of the present disclosure, an apparatus for decoding is provided. The apparatus for decoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a residual feature map based on at least one OPC of the first feature map using a filter. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the residual feature map. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the first reference picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to decode a current picture region based on the second reference picture.
[0006] According to a further aspect of the present disclosure, a non-transitory computer readable medium storing instructions for a decoder is provided. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to obtain a first reference picture from a decoded picture buffer. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a first feature map based on first information associated with the first reference picture. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a residual feature map based on at least one OPC of the first feature map using a filter. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a second feature map based on the residual feature map. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a second reference picture based on the first reference picture and the second feature map. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to decode a current picture region based on the second reference picture.
[0007] According to one aspect of the present disclosure, a method of video encoding is provided. The method may include obtaining, by a processor, a first reference picture from a decoded picture buffer. The method may include generating, by the processor, a first feature map based on first information associated with the first reference picture. The method may include generating, by the processor, a residual feature map based on at least one OPC of the first feature map using a filter. The method may include generating, by the processor, a second feature map based on the residual feature map. The method may include generating, by the processor, a second reference picture based on the first reference picture and the second feature map. The method may include encoding, by the processor, a current picture region based on the second reference picture.
[0008] According to another aspect of the present disclosure, an encoder is provided. The encoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a residual feature map based on at least one OPC of the first feature map using a filter. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the residual feature map. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the first reference picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to encode a current picture region based on the second reference picture.
[0009] According to another aspect of the present disclosure, an apparatus for encoding is provided. The apparatus for encoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a residual feature map based on at least one OPC of the first feature map using a filter. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the residual feature map. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the first reference picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to encode a current picture region based on the second reference picture.
[0010] According to a further aspect of the present disclosure, a non-transitory computer readable medium storing instructions for an encoder is provided. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to obtain a first reference picture from a decoded picture buffer. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a first feature map based on first information associated with the first reference picture. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a residual feature map based on at least one OPC of the first feature map using a filter. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a second feature map based on the residual feature map. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a second reference picture based on the first reference picture and the second feature map. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to encode a current picture region based on the second reference picture.
[0011] According to yet a further aspect of the present disclosure, a non-transitory computer-readable medium storing a bitstream is provided. The bitstream may be generated based on one or more of the operations described herein.
[0012] These illustrative embodiments are mentioned not to limit or define the present disclosure, but to provide examples to aid understanding thereof. Additional embodiments are described in the Detailed Description, and further description is provided there.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further serve to explain the principles of the present disclosure and to enable a person skilled in the pertinent art to make and use the present disclosure.
[0014] FIG. 1 illustrates a block diagram of an exemplary video codec system, according to some embodiments of the present disclosure.
[0015] FIG. 2A illustrates a block diagram of an exemplary encoding apparatus, according to some embodiments of the present disclosure.
[0016] FIG. 2B illustrates a block diagram of an exemplary decoding apparatus, according to some embodiments of the present disclosure.
[0017] FIG. 3 illustrates a detailed block diagram of an exemplary low-operation point (LOP) in-loop filter, according to some embodiments of the present disclosure.
[0018] FIG. 4A illustrates a detailed block diagram of a first exemplary backbone block (BBB) included in the LOP filter of FIG. 3, according to some embodiments of the present disclosure.
[0019] FIG. 4B illustrates a detailed block diagram of a second exemplary BBB included in the LOP filter of FIG. 3, according to some embodiments of the present disclosure.
[0020] FIG. 5A illustrates a detailed block diagram of first exemplary OPC training strategy, according to some embodiments of the present disclosure.
[0021] FIG. 5B illustrates a detailed block diagram of a first exemplary OPC layer included in the first exemplary BBB of FIG. 4A, according to some embodiments of the present disclosure.
[0022] FIG. 5C illustrates a detailed block diagram of second exemplary OPC training strategy, according to some embodiments of the present disclosure.
[0023] FIG. 5D illustrates a detailed block diagram of a second exemplary OPC layer included in the second exemplary BBB of FIG. 4B, according to some embodiments of the present disclosure.
[0024] FIG. 6A illustrates average delta (BD) -rate (%) comparison of the proposed OPC-based LOP in-loop filter and other LOP in-loop filters under an artificial intelligence (AI) configuration, according to some embodiments of the present disclosure.
[0025] FIG. 6B illustrates average BD-rate (%) comparison of the proposed OPC-based LOP in-loop filter and other LOP in-loop filters under a random access (RA) configuration, according to some embodiments of the present disclosure.
[0026] FIG. 6C illustrates average BD-rate (%) results of the proposed OPC-based LOP in-loop filter and neural network-based video coding (NNVC) with its neural network tools (NN-tools) disabled under the AI and RA configuration, according to some embodiments of the present disclosure.
[0027] FIG. 6D illustrates a computational complexity comparison of the exemplary OPC-based LOP in-loop filter and other techniques, according to some embodiments of the present disclosure.
[0028] FIG. 7 illustrates a flow chart of an exemplary method of video decoding, according to some embodiments of the present disclosure.
[0029] FIG. 8 illustrates a flow chart of an exemplary method of video encoding, according to some embodiments of the present disclosure.
[0030] Embodiments of the present disclosure will be described with reference to the accompanying drawings.DETAILED DESCRIPTION
[0031] Although some configurations and arrangements are discussed, it should be understood that this is done for illustrative purposes only. A person skilled in the pertinent art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of the present disclosure. It will be apparent to a person skilled in the pertinent art that the present disclosure can also be employed in a variety of other applications.
[0032] It is noted that references in the specification to “one embodiment, ” “an embodiment, ” “an example embodiment, ” “some embodiments, ” “certain embodiments, ” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it would be within the knowledge of a person skilled in the pertinent art to effect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0033] In general, terminology may be understood at least in part from usage in context. For example, the term “one or more” as used herein, depending at least in part upon context, may be used to describe any feature, structure, or characteristic in a singular sense or may be used to describe combinations of features, structures or characteristics in a plural sense. Similarly, terms, such as “a, ” “an, ” or “the, ” again, may be understood to convey a singular usage or to convey a plural usage, depending at least in part upon context. In addition, the term “based on” may be understood as not necessarily intended to convey an exclusive set of factors and may, instead, allow for existence of additional factors not necessarily expressly described, again, depending at least in part on context.
[0034] Various aspects of video coding systems will now be described with reference to various apparatus and methods. These apparatus and methods will be described in the following detailed description and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as “elements” ) . These elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends upon the particular application and design constraints imposed on the overall system.
[0035] The techniques described herein may be used for various video coding applications. As described herein, video coding includes both encoding and decoding a video. Encoding and decoding of a video can be performed by the unit of block. For example, an encoding / decoding process such as transform, quantization, prediction, in-loop filtering, reconstruction, or the like may be performed on a coding block, a transform block, or a prediction block. As described herein, a block to be encoded / decoded will be referred to as a “current block. ” For example, the current block may represent a coding block, a transform block, or a prediction block according to a current encoding / decoding process. In addition, it is understood that the term “unit” used in the present disclosure indicates a basic unit for performing a specific encoding / decoding process, and the term “block” indicates a sample array of a predetermined size. Unless otherwise stated, the “block, ” “unit, ” and “component” may be used interchangeably.
[0036] With the rapid development of information technology and multimedia technology, video applications are gradually evolving towards high definition and diversification. In response to the urgent need for significantly improved coding efficiency and the ability to handle diverse video types, the Joint Video Experts Team (JVET) established the latest video coding standard with remarkable coding efficiency in July 2020, known as Versatile Video Coding (VVC / H. 266) . The VVC standard offers outstanding performance, allowing for approximately a 50%bitrate reduction compared to the High Efficiency Video Coding (HEVC / H. 265) standard at the same quality level. Furthermore, the VVC standard can accommodate a wider variety of video formats and content, providing a unified, flexible, and efficient coding compression framework for both existing and emerging video applications.
[0037] Due to the block-based hybrid coding framework still adopted by VVC, some common coding distortion effects, such as blocking effects, ringing artifacts, color deviation, and image blur still exist. To improve the quality of encoded images and provide high-quality references for subsequent images, thereby achieving better prediction results, VVC has introduced in-loop filtering techniques. These techniques include luma mapping and chroma scaling, deblocking filtering (DBF) , sample adaptive offset (SAO) , and adaptive loop filter (ALF) . Among these, deblocking filtering and sample adaptive offset continue the relevant algorithms from HEVC. Although these filters contain nonlinear elements, they are actually designed based on linear filters. Therefore, these traditionally handcrafted filters have relatively limited adaptability when dealing with diverse video content.
[0038] To address these issues, JVET established an Ad Hoc Group focused on neural network-based video coding (NNVC) . JVET NNVC is committed to exploring the application of neural network technologies in video coding standards to improve coding efficiency and visual quality. Currently, the research of JVET NNVC includes neural network-based (NN-based) loop filtering, NN-based inter prediction, and NN-based super resolution. In terms of in-loop filter, JVET NNVC introduced three operating points for the NN-based loop filter (NNLF) . The High Operating Point (HOP) aims to provide the maximum possible gain in scenarios with higher complexity, while the Low Operating Point (LOP) focuses on achieving the best performance under lower complexity constraints. The Very Low Operating Point (VLOP) seeks to offer greater potential gains while operating under the lowest complexity limits.
[0039] After rigorous research, the complexity of LOP4 was reduced to 16.83kMAC / pixel, with a parameter count of 0.2M. Under AI and RA configurations, the PSNR for the luma channel of LOP4 improved by an average of 8.63%and 7.63%, respectively. To reduce network complexity, LOP4 utilized two backbone block A (BBBA) modules for the luma and chroma channels, reducing the input patch size from the original 36x36 to 32x32. However, this process is not perfect. Exploration experiments indicate that using two BBBA modules to crop the patch size leads to some wasted complexity, making it less efficient compared to cropping the patch size with just one BBBA. Moreover, cropping the patch size often results in the input patch size for BBBA being larger than that for a subsequent backbone block B (BBBB) , leading to richer input information for BBBA under the same input channels. Therefore, increasing the input channels for BBBA would be more beneficial for enhancing the performance of the LOP4 network. Furthermore, the separable convolution layer in BBB in LOP4 has only one branch, which is not conducive to extracting multi-scale features from the input information.
[0040] To overcome these and other challenges, the present invention proposes a new backbone block (BBB) of LOP in-loop filter based on over-parameterized convolution and variable channel number. As shown in Fig. 1, the modifications of the LOP4 network in the proposed method are highlighted in green. Specifically, we have introduced a cropping channel operation on the LOP4 BBB, allowing for differing input and output channels of the BBB. Additionally, we replaced the 1x3 and 3x1 separable convolution layers in the LOP4 BBB with our Over-Parameterized Convolution (OPC) module.
[0041] The present disclosure proposes an exemplary BBB enhancement for LOP in-loop filter based on OPC and a variable number of channels. To that end, the exemplary BBB of the present disclosure includes an exemplary OPC module. The exemplary BBB described herein may increase the receptive field of the convolutional layer while better capturing the comprehensive characteristics of the input data. To obtain a richer input, the number of input channels for a first BBB (BBBA) are increased, while the number of channels for a second BBB (BBBB) are decreased accordingly. In the training phase, the 1x3 and 3x1 separable convolutional layers in the LOP BBB are replaced with OPC modules. After completing the over-parameterized training of the LOP network, the OPC module in the LOP BBB is restored to 1x5 and 5x1 separable convolutional layers in the inference phase. The proposed method improves the LOP network performance while reducing the complexity of the inference stage. The experimental results show that the proposed LOP in-loop filter reduces the average BD-rate of the LOP4 network by {0.19%(Y) , 1.57% (U) , 1.70% (V) } and {0.07% (Y) , 2.42% (U) , 2.44% (V) } respectively under AI and RA configurations. Compared with VTM-11.0 NNVC-11.0 anchor (NN tools OFF) , the average BD-rate under AI and RA configurations is reduced by {8.87% (Y) , 16.69% (U) , 16.73% (V) } and {7.73% (Y) , respectively. 16.28% (U) , 15.21% (V) } .
[0042] FIG. 1 is a block diagram of a video codec system, according to some embodiments of the present disclosure. The video codec system, according to an embodiment, may include an encoding apparatus 10 and a decoding apparatus 20. The encoding apparatus 10 may deliver encoded video and / or picture information or data to the decoding apparatus 20 in the form of a file or streaming via a digital storage medium or network.
[0043] The encoding apparatus 10, according to an embodiment, may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding apparatus 20, according to an embodiment, may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be called a video / picture encoding unit, and the decoding unit 22 may be called a video / picture decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0044] The video source generator 11 may acquire a video / picture through a process of capturing, synthesizing, or generating the video / picture. The video source generator 11 may include a video / picture capture device and / or a video / picture generating device. The video / picture capture device may include, for example, one or more cameras, video / picture archives including previously captured video / pictures, and the like. The video / picture-generating device may include, for example, computers, tablets, and smartphones, and may (electronically) generate video / pictures. For example, a virtual video / picture may be generated through a computer or the like. In this case, the video / picture capturing process may be replaced by a process of generating related data.
[0045] The encoding unit 12 may encode an input video / picture. The encoding unit 12 may perform a series of procedures such as prediction, transform, and quantization for compression and coding efficiency. The encoding unit 12 may output encoded data (encoded video / picture information) in the form of a bitstream.
[0046] The transmitter 13 may transmit the encoded video / picture information or data output in the form of a bitstream to the receiver 21 of the decoding apparatus 20 through a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage mediums such as universal serial bus (USB) , secure digital (SD) , compact disc (CD) , digital video disc (DVD) , Blu-ray, hard disk drive (HDD) , solid-state drive (SSD) , and the like. The transmitter 13 may include an element for generating a media file through a predetermined file format and may include an element for transmission through a broadcast / communication network. The receiver 21 may extract / receive the bitstream from the storage medium or network and transmit the bitstream to the decoding unit 22.
[0047] The decoding unit 22 may decode the video / picture by performing a series of procedures such as dequantization, inverse transform, and prediction corresponding to the operation of the encoding unit 12.
[0048] The renderer 23 may render the decoded video / picture. The rendered video / picture may be displayed through the display.
[0049] FIG. 2A is a schematic block diagram of an encoding apparatus, in accordance with some aspects of the present disclosure. Referring to FIG. 2A, the encoding apparatus 200 includes a picture partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 251, a filter 261, and a memory 271. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 251 may be called a reconstructor or a reconstructed block generator. The picture partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 251, and the filter 261 may be configured by at least one hardware component (e.g., an encoder chipset or processor) , according to an embodiment. In addition, the memory 271 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may further include the memory 271 as an internal / external component.
[0050] The picture partitioner 210 may partition an input picture (or a picture or a frame) input to the encoding apparatus 200 into one or more processors. For example, the processor may be called a coding unit (CU) . In this case, the coding unit may be recursively partitioned according to a quad-tree binary-tree ternary-tree (QTBTTT) structure from a coding tree unit (CTU) or a largest coding unit (LCU) . For example, one coding unit may be partitioned into a plurality of coding units of a deeper depth based on a quad tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad tree structure may be applied first, and the binary tree structure and / or ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to this invention may be performed based on the final coding unit that is no longer partitioned. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to picture characteristics, or if necessary, the coding unit may be recursively partitioned into coding units of deeper depth, and a coding unit having an optimal size may be used as the final coding unit. Here, the coding procedure may include a procedure of prediction, transform, and reconstruction, which will be described later. As another example, the processor may further include a prediction unit (PU) or a transform unit (TU) . In this case, the prediction unit and the transform unit may be split or partitioned from the aforementioned final coding unit. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0051] The unit may be used interchangeably with terms such as block or area in some cases. In a general case, an M×N block may represent a set of samples or transform coefficients composed of M columns and N rows. A sample may generally represent a pixel or a value of a pixel, may represent only a pixel / pixel value of a luma component or represent only a pixel / pixel value of a chroma component. A sample may be used as a term corresponding to one picture (or picture) for a pixel or a pel.
[0052] In the encoding apparatus 200, a prediction signal (predicted block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 is subtracted from an input picture signal (original block, original sample array) to generate a residual signal residual block, residual sample array) , and the generated residual signal is transmitted to the transformer 232. In this case, as shown, a unit for subtracting a prediction signal (predicted block, prediction sample array) from the input picture signal (original block, original sample array) in the encoding apparatus 200 may be called a subtractor 231. The predictor may perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The predictor may determine whether intra prediction or inter prediction is applied on a current block or CU basis. As described later in the description of each prediction mode, the predictor may generate various information related to prediction, such as prediction mode information, and transmit the generated information to the entropy encoder 240. The information on the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0053] The intra predictor 222 may predict the current block by referring to the samples in the current picture. The referred samples may be located in the neighborhood of the current block or may be located apart according to the prediction mode. In the intra prediction, prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional mode may include, for example, a DC mode and a planar mode. The directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of detail of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the setting. The intra predictor 222 may determine the prediction mode applied to the current block by using a prediction mode applied to a neighboring block.
[0054] The inter predictor 221 may derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the neighboring block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc. ) information. In the case of inter prediction, the neighboring block may include a spatial neighboring block present in the current picture and a temporal neighboring block present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a co-located CU (colCU) , and the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic) . For example, the inter predictor 221 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index of the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, the inter predictor 221 may use motion information of the neighboring block as motion information of the current block. In the skip mode, unlike the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling a motion vector difference.
[0055] The predictor 220 may generate a prediction signal based on various prediction methods described below. For example, the predictor may not only apply intra prediction or inter prediction to predict one block but also simultaneously apply both intra prediction and inter prediction. This may be called combined inter and intra prediction (CIIP) . In addition, the predictor may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content picture / video coding of a game or the like, for example, screen content coding (SCC) . The IBC basically performs prediction in the current picture but may be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, the IBC may use at least one of the inter prediction techniques described in the present disclosure. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, a sample value within a picture may be signaled based on information on the palette table and the palette index.
[0056] The prediction signal generated by the predictor (including the inter predictor 221 and / or the intra predictor 222) may be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT) , a discrete sine transform (DST) , a karhunen-loève transform (KLT) , a graph-based transform (GBT) , or a conditionally non-linear transform (CNT) . Here, the GBT means transform obtained from a graph when relationship information between pixels is represented by the graph. The CNT refers to the transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks having the same size or may be applied to blocks having a variable size rather than a square.
[0057] The quantizer 233 may quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information on the quantized transform coefficients) and output a bitstream. The information on the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange block type quantized transform coefficients into a one-dimensional vector form based on a coefficient scanning order and generate information on the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Information on transform coefficients may be generated. The entropy encoder 240 may perform various encoding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC) , context-adaptive binary arithmetic coding (CABAC) , and the like. The entropy encoder 240 may encode information necessary for video / picture reconstruction other than quantized transform coefficients (e.g., values of syntax elements, etc. ) together or separately. Encoded information (e.g., encoded video / picture information) may be transmitted or stored in units of NALs (network abstraction layer) in the form of a bitstream. The video / picture information may further include information on various parameter sets, such as an adaptation parameter set (APS) , a picture parameter set (PPS) , a sequence parameter set (SPS) , or a video parameter set (VPS) . In addition, the video / picture information may further include general constraint information. In the present disclosure, information and / or syntax elements transmitted / signaled from the encoding apparatus to the decoding apparatus may be included in video / picture information. The video / picture information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted over a network or may be stored in a digital storage medium. The network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. A transmitter (not shown) transmitting a signal output from the entropy encoder 240 and / or a storage unit (not shown) storing the signal may be included as internal / external element of the encoding apparatus 200, and alternatively, the transmitter may be included in the entropy encoder 240.
[0058] The quantized transform coefficients output from the quantizer 233 may be used to generate a prediction signal. For example, the residual signal (residual block or residual samples) may be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the dequantizer 234 and the inverse transformer 235. The adder 251 adds the reconstructed residual signal to the prediction signal output from the inter predictor 221 or the intra predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) . If there is no residual for the block to be processed, such as in a case where the skip mode is applied, the predicted block may be used as the reconstructed block. The adder 251 may be called a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture and may be used for inter prediction of the next picture through filtering as described below.
[0059] Meanwhile, luma mapping with chroma scaling (LMCS) may be applied during picture encoding and / or reconstruction.
[0060] The filter 261 may improve subjective / objective picture quality by applying filtering to the reconstructed signal. For example, the filter 261 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 271, specifically, a DPB of the memory 271. The various filtering methods may include, for example, deblocking filtering, a sample adaptive offset, an adaptive loop filter, a bilateral filter, and the like. The filter 261 may generate various information related to the filtering and transmit the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to the filtering may be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0061] The modified reconstructed picture transmitted to the memory 271 may be used as the reference picture in the inter predictor 221. When the inter prediction is applied through the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus 250 may be avoided, and encoding efficiency may be improved.
[0062] The DPB of the memory 271 DPB may store the modified reconstructed picture for use as a reference picture in the inter predictor 221. The memory 271 may store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks in the picture that have already been reconstructed. The stored motion information may be transmitted to the inter predictor 221 and used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 271 may store reconstructed samples of reconstructed blocks in the current picture and may transfer the reconstructed samples to the intra predictor 222.
[0063] FIG. 2B is a schematic block diagram of a decoding, in accordance with some embodiments of the present disclosure. Referring to FIG. 2B, the decoding apparatus 250 may include an entropy decoder 270, a residual processor 252, a predictor 258, an adder 264, a filter 266, and a memory 268. The predictor 258 may include an inter predictor 260 and an intra predictor 262. The residual processor 252 may include a dequantizer 254 and an inverse transformer 256. The entropy decoder 270, the residual processor 252, the predictor 258, the adder 264, and the filter 266 may be configured by a hardware component (e.g., a decoder chipset or a processor) according to an embodiment. In addition, the memory 268 may include a decoded picture buffer (DPB) or may be configured by digital storage.
[0064] When a bitstream including video / picture information is input, the decoding apparatus 250 may reconstruct a picture corresponding to a process in which the video / picture information is processed in the encoding apparatus of FIG. 2A. For example, the decoding apparatus 250 may derive units / blocks based on block partition-related information obtained from the bitstream. The decoding apparatus 250 may perform decoding using a processor applied in the encoding apparatus. Thus, the processor of decoding may be a coding unit, for example, and the coding unit may be partitioned according to a quad tree structure, binary tree structure and / or ternary tree structure from the coding tree unit or the largest coding unit. One or more transform units may be derived from the coding unit. The input picture data signal decoded and output through the decoding apparatus 250 may be reproduced through a reproducing apparatus.
[0065] The decoding apparatus 250 may receive a signal output from the encoding apparatus of FIG. 2A in the form of a bitstream, and the received signal may be decoded through the entropy decoder 270. For example, the entropy decoder 270 may parse the bitstream to derive information (e.g., video / picture information) necessary for picture reconstruction (or picture reconstruction) . The video / picture information may further include information on various parameter sets, such as an adaptation parameter set (APS) , a picture parameter set (PPS) , a sequence parameter set (SPS) , or a video parameter set (VPS) . In addition, the video / picture information may further include general constraint information. The decoding apparatus may further decode picture based on the information on the parameter set and / or the general constraint information. Signaled / received information and / or syntax elements described later in the present disclosure may be decoded by the decoding procedure and obtained from the bitstream. For example, the entropy decoder 270 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output syntax elements required for picture reconstruction and quantized values of transform coefficients for residual. More specifically, the CABAC entropy decoding method may receive a bin corresponding to each syntax element in the bitstream, determine a context model using a decoding target syntax element information, decoding information of a decoding target block or information of a symbol / bin decoded in a previous stage, and perform an arithmetic decoding on the bin by predicting a probability of occurrence of a bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method may update the context model by using the information of the decoded symbol / bin for a context model of the next symbol / bin after determining the context model. The information related to the prediction among the information decoded by the entropy decoder 270 may be provided to the predictor (the inter predictor 260 and the intra predictor 262) , and the residual value on which the entropy decoding was performed in the entropy decoder 270, that is, the quantized transform coefficients and related parameter information, may be input to the residual processor 252. The residual processor 252 may derive the residual signal (the residual block, the residual samples, and the residual sample array) . In addition, information on filtering among information decoded by the entropy decoder 270 may be provided to the filter 266. Meanwhile, a receiver (not shown) for receiving a signal output from the encoding apparatus may be further configured as an internal / external element of the decoding apparatus 250, or the receiver may be a component of the entropy decoder 270. Meanwhile, the decoding apparatus in the present disclosure may be referred to as a video / picture / picture decoding apparatus, and the decoding apparatus may be classified into an information decoder (video / picture / picture information decoder) and a sample decoder (video / picture / picture sample decoder) . The information decoder may include the entropy decoder 270, and the sample decoder may include at least one of the dequantizer 254, the inverse transformer 256, the adder 264, the filter 266, the memory 268, the inter predictor 260, and the intra predictor 262.
[0066] The dequantizer 254 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 254 may rearrange the quantized transform coefficients in the form of a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding apparatus. The dequantizer 254 may perform dequantization on the quantized transform coefficients by using a quantization parameter (e.g., quantization step size information) and obtain transform coefficients.
[0067] The inverse transformer 256 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array) .
[0068] The predictor may perform prediction on the current block and generate a predicted block including prediction samples for the current block. The predictor may determine whether intra prediction or inter prediction is applied to the current block based on the information on the prediction output from the entropy decoder 270 and may determine a specific intra / inter prediction mode.
[0069] The predictor 258 may generate a prediction signal based on various prediction methods described below. For example, the predictor may not only apply intra prediction or inter prediction to predict one block but also simultaneously apply intra prediction and inter prediction. This may be called combined inter and intra prediction (CIIP) . In addition, the predictor may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content picture / video coding of a game or the like, for example, screen content coding (SCC) . The IBC basically performs prediction in the current picture but may be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, a sample value within a picture may be signaled based on information on the palette table and the palette index.
[0070] The intra predictor 262 may predict the current block by referring to the samples in the current picture. The referred samples may be located in the neighborhood of the current block or may be located apart according to the prediction mode. In the intra prediction, prediction modes may include a plurality of nondirectional modes and a plurality of directional modes. The intra predictor 262 may determine the prediction mode applied to the current block by using a prediction mode applied to a neighboring block.
[0071] The intra predictor 262 may predict the current block by referring to the samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located apart according to the prediction mode. In intra prediction, prediction modes may include a plurality of nondirectional modes and a plurality of directional modes. The intra predictor 262 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0072] The inter predictor 260 may derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, motion information may be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the neighboring block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc. ) information. In the case of inter prediction, the neighboring block may include a spatial neighboring block present in the current picture and a temporal neighboring block present in the reference picture. For example, the inter predictor 260 may configure a motion information candidate list based on neighboring blocks and derive a motion vector of the current block and / or a reference picture index based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and the information on the prediction may include information indicating a mode of inter prediction for the current block.
[0073] The adder 264 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from the predictor (including the inter predictor 260 and / or the intra predictor 262) . If there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block may be used as the reconstructed block.
[0074] The adder 264 may be called reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter prediction of the next picture.
[0075] Meanwhile, luma mapping with chroma scaling (LMCS) may be applied in the picture decoding process.
[0076] The filter 266 may improve subjective / objective picture quality by applying filtering to the reconstructed signal. For example, the filter 266 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 268, specifically, a DPB of the memory 268. The various filtering methods may include, for example, deblocking filtering, a sample adaptive offset, an adaptive loop filter, a bilateral filter, and the like.
[0077] The (modified) reconstructed picture stored in the DPB of the memory 268 may be used as a reference picture in the inter predictor 260. The memory 268 may store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks in the picture that have already been reconstructed. The stored motion information may be transmitted to the inter predictor 221 so as to be utilized as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 268 may store reconstructed samples of reconstructed blocks in the current picture and transfer the reconstructed samples to the intra predictor 262.
[0078] In the present disclosure, the embodiments described in the filter 261, the inter predictor 221, and the intra predictor 222 of the encoding apparatus 200 may be the same as or respectively applied to correspond to the filter 266, the inter predictor 260, and the intra predictor 262 of the decoding apparatus 250. The same may also apply to the inter predictor 260 and the intra predictor 262.
[0079] In-loop filtering techniques for traditional video codecs are one of the representative components of NN-based coding models or tools. Since VVC is a block-based hybrid coding framework, it inevitably produces undesirable compression artifacts, especially at a high compression rate. To address this issue, VVC employs a series of advanced filtering techniques to eliminate or reduce compression artifacts.
[0080] FIG. 3 illustrates a detailed block diagram of an exemplary low-operation point (LOP) in-loop filter 300 (referred to hereinafter as “LOP filter 300” ) , according to some embodiments of the present disclosure. FIG. 4A illustrates a detailed block diagram of a first exemplary (BBB) 302a (referred to hereafter as “BBBA 302a” ) included in the LOP filter 300 of FIG. 3, according to some embodiments of the present disclosure. FIG. 4B illustrates a detailed block diagram of a second exemplary BBB (referred to hereafter as “BBBB 302b” ) included in the LOP filter 300 of FIG. 3, according to some embodiments of the present disclosure. FIG. 5A illustrates a detailed block diagram of first exemplary OPC training strategy 500, according to some embodiments of the present disclosure. FIG. 5B illustrates a detailed block diagram of a first exemplary OPC module (OPC1) 408 included in the BBBA 302a of FIGs. 3 and 4A, according to some embodiments of the present disclosure. FIG. 5C illustrates a detailed block diagram of second exemplary OPC training strategy 525, according to some embodiments of the present disclosure. FIG. 5D illustrates a detailed block diagram of a second exemplary OPC module (OPC2) 422 included in BBBB 302b of FIG. 4B, according to some embodiments of the present disclosure. FIGs. 3, 4A, 4B, and 5A-5D will be described together.
[0081] Referring to FIG. 3, the LOP filter 300 includes three main parts: feature extraction (head) , feature enhancement (body) , and reconstruction (tail) . Although there are subtle differences between the luma and chroma models in feature extraction and reconstruction, they share the same feature enhancement structure, namely the first BBBA 302a and second BBBB 302b, but with a different number of basic blocks (e.g., 14 BBB in the luma model and 4 BBB in the chroma model) . The inputs for the LOP filter 300 include a reconstructed picture (Rec) , a prediction picture (Pred) , a boundary strength (BS) map, a base quantization parameter (QP) map (QPbase) , a slice QP map (QPslice) , and a prediction type (IPB) .
[0082] The reconstructed picture (Rec) is formed by combining a luma reconstructed picture (RecEXTY) to which a discrete cosine transform (DCT) has been applied and a chroma reconstructed picture (RecEXTUV) . Information associated with the RecEXTY to which the DCT has been applied may be used for the Pred. Information associated with the RecEXTY after downsampling may be used for the BS map and the IPB.
[0083] During feature extraction, a first feature map (one or more feature maps) may be generated based on the Rec, the Pred, the BS map, the QPbase, the QPslice, and the IPB. For instance, a 3x3d1 convolutional (CONV) layer, a 3x3d2 CONV layer, a 1x1d3 CONV layer, two 1x1d4 CONV layers, and a 1x1d5 CONV layer may be used to extract first features from the Rec, the Pred, the BS map, the QPbase, the QPslice, and the IPB, respectively. Then, the first feature map may be generated based on the first features using a 1x1C CONV layer, a parameterized rectified linear unit (PReLU) layer, downsampling 1x3C CONV layer, a downsampling 3X1C CONV layer, a CYxCY1 CONV layer (e.g., a 1x1 CONV layer) , and a CUVxCUV1 CONV layer (e.g., a 1x1 CONV layer) . The output of the CYxCY1 CONV layer may be a luma feature map, while the output of the CUVxCUV1 CONV layer may be a chroma feature map. The luma feature map and the chroma feature map may be collectively referred to herein as “the first feature map. ”
[0084] During feature enhancement, the luma feature map is input into a plurality of backbone blocks (BBBs) (e.g., one BBBA 302a and fourteen BBBB 302b) along a luma branch of the model, and the chroma feature map is input into a plurality of BBBs (e.g., one BBBA 302a and three BBBB 302b) along a chroma branch of the model. Each BBBA 302a and BBBB 302b may include an OPC module (shown in FIGs. 4A and 4B) . The output of the last BBBB 302b from each of the luma path and the chroma path may be a luma residual feature map and a chroma residual map. In some implementations, the luma residual map and the chroma residual feature map may be collectively referred to as “a residual feature map. ” In some other implementations, the luma residual map and the chroma residual feature map may each be individually referred to as “a residual feature map. ”
[0085] For example, referring to FIG. 4A, the BBBA 302a may extract a set of output residual features 412 from the input features 402 (e.g., luma feature map) . The input features 402 may have dimensions m1, h, w, where m1 is the number of input channels and h and w is patch size. The set of output residual features 412 may be extracted by a PReLU layer 404, a 1x1 CONV layer 406 with dimensions m1xn, a first OPC module (OPC1) 408, and a 1x1 CONV layer 410 with dimensions nxm2 (where n is the number of input channels and m2 is the number of output channels) . The set of output residual features 412 may be generated based on an output of the 1x1CONV layer 410 and the input features 402 with its number of channels reduced to m2 by a cropping operation 414.
[0086] Referring to FIG. 4B, the BBBB 302b may extract a set of output residual features 426 from the input features 416 (e.g., the set of features 412 from FIG. 4A when BBBB 302b directly follows BBBA 302a or from a preceding BBBB 302b) . The input features 416 may have dimensions m1, h, w, where m1 is the number of input channels to the BBBB 302b. The set of output residual features 426 may be extracted by a PReLU layer 418, a 1x1 CONV layer 420 with dimensions m1xn, a second OPC module (OPC2) 422, and a 1x1 CONV layer 424 with dimensions nxm2 (where m2 is the number of output channels of BBBB 302b) . The output residual features 426 may be generated based on an output of the 1x1CONV layer 424 and the input features 416 with its number of channels reduced to m2 by a cropping operation 428.
[0087] In the training phase, the network structure of OPC1 408 and OPC2 422 is shown in FIGs. 5A and 5C. To limit the number of BBBA 302a to one, while at the same time increasing the receptive field and improving the prediction accuracy of the LOP filter 300, 1x5 and 5x1 sCONV layers are used. OPC1 408 uses 1x5 and 5x1 sCONV layers and only in each of the luma path and the chroma path are included to crop the input patch size from the original 36x36 to 32x32. Referring to FIGs. 5A and 5C, to fully extract the multi-scale features of the input information, the 1x5 and 5x1 sCONV layers are expanded into a multi-branch structure. As shown in FIGs. 5A and 5C, each OPC module has four branches.
[0088] In the inference phase, each OPC module (OPC1 408 and OPC2 422) is merged into the network structure, as shown in FIGs. 5B and 5D. During the merging process, the center points of the weights of each convolution branch are aligned, summed, and 1.0 is added to the weight of the center point to complete the network merging. At the same time, the corresponding biases of each convolutional branch are sequentially added to obtain the biases for the new synthesized network.
[0089] For example, referring to FIG. 5A, for training OPC1 408, input features 501 are input to a 1x5 sCONV layer 504a with dimensions CxC and padding, a 1x3 sCONV layer 504b with dimensions CxC and padding, and a 1x1 sCONV layer 504c with dimensions CxC and padding. A cropping operation 508 is applied to the input features 501, the output of the 1x3 sCONV layer 504b, and the 1x1 sCONV layer 504c. A first summation operation 510a combines the output of the 1x5 sCONV layer 504a, the cropped outputs of the 1x3 sCONV layer 504b and the 1x1 sCONV layer 504c, and the cropped input features 501. These combined features are input to a 5x1 sCONV layer 506a, a 3x1 sCONV layer 506b, and a 1x1 sCONV layer 506c. The outputs of the 3x1 sCONV layer 506b and the 1x1 sCONV layer 506c are cropped. A second summation operation 510b combines the output of the 5x1 sCONV layer 506a, the cropped outputs of the 3x1 sCONV layer 506b and the 1x1 sCONV layer 506c, and the cropped combined features output by the first summation operation 510a. Output features 503 are generated based on the output of the second summation operation 510b.
[0090] The training may be performed to generate a set of parameters (e.g., weights and biases) for use in OPC1 408 during inference. A trained OPC1 408 that is included in BBBA 302a is shown in FIG. 5B. Referring to FIG. 5B, the input features 505 are input into a 1x5 sCONV layer 502 and a 5x1 sCONV layer 512. The output features 507 are generated based on the output of the 5x1 sCONV layer 512.
[0091] Referring to FIG. 5C, for training OPC2 422, input features 509 are input to a 1x5 sCONV layer 514a with dimensions CxC, a 1x3 sCONV layer 514b with dimensions CxC, and a 1x1 sCONV layer 514c with dimensions CxC. A first summation operation 518a combines the output of the 1x5 sCONV layer 514a, the output of the 1x3 sCONV layer 514b, the output of the 1x1 sCONV layer 514c, and the input features 509. These combined features are input to a 5x1 sCONV layer 516a, a 3x1 sCONV layer 516b, and a 1x1 sCONV layer 516c. A second summation operation 518b combines the output of the 5x1 sCONV layer 516a, the output of the 3x1 sCONV layer 516b, the output of the 1x1 sCONV layer 516c, and the combined features output by the first summation operation 518a. Output features 511 are generated based on the output of the second summation operation 518b. The training may be performed to generate a set of parameters (e.g., weights and biases) for use in OPC2 422 during inference. A trained OPC2 422 that is included in BBBB 302b is shown in FIG. 5D. Referring to FIG. 5D, the input features 513 are input into a 1x5 sCONV layer 520 and a 5x1 sCONV layer 522. The 5x1 sCONV layer 522 may generate the output features 515.
[0092] As shown in FIGs. 4A and 4B, to make the LOP4 BBB combination more flexible, cropping operations are included to allow the number of input and output channels of the BBB to differ. Subsequently, to provide the LOP4 BBB with richer input information, the number of input channels is increased for BBBA 302a while correspondingly reducing the number of input channels for BBBB 302b. Non-limiting examples of the number of input and output channels for each BBB are described in Table 1. Among them, the input channel number for the LOP4 luma branch gradually decreases from 208 to 144, and the input channel number for the chroma branch gradually decreases from 144 to 128. Table 1: Number of Input and Output Channels of BBB in Luma path and Chroma path
[0093] Referring again to FIG. 3, during reconstruction, a second feature map may be generated based on the residual feature map (e.g., the luma residual map and the chroma residual map) . For instance, an output luma feature map may be generated based on the luma residual map output by the last BBBB 302b in the luma path using a 1x1 CONV layer with dimensions CY1xCY, a 1x3 sCONV layer with dimensions CYxCY21, a 3x1 sCONV layer with dimensions CY21xCY, a 1x1 CONV layer with dimensions CYxCY, a PReLU layer, a 3x3 CONV layer with dimensions CYx16, a pixel shuffle layer, and an inverse DCT (IDCT) .
[0094] An output chroma feature map may be generated based on the chroma residual map output by the last BBBB 302b in the chroma path using a 1x1 CONV layer with dimensions CUVxCUV1, a 1x3 sCONV layer with dimensions CUVxCUV21, a 3x1 sCONV layer with dimensions CUV21xCUV, a 1x1 CONV layer with dimensions CUVxCUV, a PReLU layer, a 3x3 CONV layer with dimensions CUVx8, and IDCT.
[0095] An output reconstructed luma picture may be generated based on the output luma feature map and the input luma reconstructed picture after its patch size is reduced by cropping. Similarly, a second reference chroma picture may be generated based on the output chroma feature map and the first reference chroma picture after its patch size is reduced by cropping.
[0096] Using the output reconstructed luma picture and / or the output reconstructed chroma picture a second reference picture may be generated, which can be used to code a picture region.
[0097] For the over-parameterized training, the datasets include DIV2K, BVI-DVC, and TVD the same as the dataset used for LOP4 training. The training dataset was generated based on NNVC-10.0 and contains two stages: the first stage produces AI-encoded data, and the second stage generates RA-encoded data. First, all neural network (NN) -tools are turned off during data generation, full intra-frame compression in all intra (AI) configuration is applied to the original DIV2K images, and the samples before deblocking effects are extracted to generate AI-encoded data. The QPs for compression in this stage are 19, 24, 29, 34, 39, 22, 27, 32, 37 and 42. Second, during the data generation process, NN-based in-loop filtering is activated, and RA-encoded data is generated on the BVI-DVC and TVD datasets using the LOP model provided by JVET. The QPs for compression in this stage are 22, 27, 32, 37 and 42.
[0098] Similar to the training strategy of the LOP4 network, we also utilize both AI-encoded data and RA-encoded data to train the proposed LOP in-loop filter. Our training strategy is similar to the third stage of the LOP training strategy with a learning rate of 2e-4 and a batch size of 64. The learning rate is reduced by a factor of 10 at the 41st, 51st, and 55th epochs. The total number of epochs is 60, and the loss function switches from L1 to MSE at epoch 58. In the loss function, the weighting between luma (Y) and chroma (U and V) is 12: 1: 1. During the inference phase, the model complexity is 16.81 kMAC / pixel, which is lower than the 16.83 kMAC / pixel of the LOP4 network.
[0099] In the inference stage, the trained model is converted into the float SADL framework after synthesis and integrated into VTM-11.0_NNVC-11.0 for tests according to the Common Test Conditions (CTC) in JVET. Following the CTC, the test sequence of Classes A1, A2, B, C, D, and E is used as the test set and BD-rate as the evaluation metric to assess the performance of the proposed model under AI and RA configurations.
[0100] FIG. 6A illustrates average delta (BD) -rate (%) comparison 600 of the proposed OPC-based LOP in-loop filter and other LOP in-loop filters under an artificial intelligence (AI) configuration, according to some embodiments of the present disclosure.
[0101] Referring to FIG. 6A, the BD-rates of the proposed LOP4 model and other JVET proposals model compared to the original LOP4 model with test results under the AI configuration are shown. It can be observed that the average BD-rate of the proposed method is {-0.19%, -1.57%, -1.70%} , which performs better than the results of AK0195 {-0.02%, -0.71%, -0.53%} and AK0106 {0.00%, -0.12%, -0.34%} for both luma and chroma channels.
[0102] FIG. 6B illustrates average BD-rate (%) comparison 625 of the proposed OPC-based LOP in-loop filter and other LOP in-loop filters under a random access (RA) configuration, according to some embodiments of the present disclosure.
[0103] Referring to FIG. 6B, the BD-rates of the proposed LOP4 model and other JVET proposals model compared to the original LOP4 model with test results under the RA configuration are shown. It can be observed that the average BD-rate of the proposed method is {-0.07%, -2.42%, -2.44%} , which performs better than the results of AK0195 {-0.02%, -1.49%, -1.34%} and AK0106 {0.12%, -0.28%, -0.32%} for both luma and chroma channels.
[0104] FIG. 6C illustrates average BD-rate (%) results of the proposed OPC-based LOP in-loop filter and neural network-based video coding (NNVC) with its neural network tools (NN-tools) disabled under the AI and RA configuration, according to some embodiments of the present disclosure.
[0105] Referring to FIG. 6C, the BD-rate of the proposed LOP4 network over the VTM-11.0_NNVC-11.0 anchor with NNVC tools (NN-Intra and LOP in-loop filter) disabled are shown. The results show that the proposed LOP4 network achieves average BD-rate reductions of {8.87%, 16.69%, 16.73%} and {7.73%, 16.28%, 15.21%} over the VTM-11.0_NNVC-11.0 anchor under AI and RA configurations, respectively.
[0106] FIG. 6D illustrates a computational complexity comparison 675 of the exemplary OPC-based LOP in-loop filter and other techniques, according to some embodiments of the present disclosure.
[0107] Referring to FIG. 6D, the complexity comparison among LOP4, JVET-AK0195, JVET-AK0106 and the proposed network are shown. The kMAC / pixel of the proposed network is close to the LOP4 network and AK0195 network, but the proposed network achieves higher encoding gain. The AK0106 network has the lowest KMAC, but its coding gain is lower than the other methods. Our parameters are slightly more than LOP4, AK0195 and AK0106 because the proposed network uses 1x5 and 5x1 separable convolutional layers.
[0108] FIG. 7 illustrates a flow chart of an exemplary method 700 of video decoding, according to some embodiments of the present disclosure. Method 700 may be performed by an apparatus, e.g., such as decoding apparatus 20, 250 decoding unit 22, in-loop filter 266, LOP filter 300, BBBA 302a, BBBB 302b, OPC1 408, OPC2 422, etc. Method 700 may include operations 702-716 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously, or in a different order other than shown in FIG. 7.
[0109] At 702, the apparatus may generate a set of parameters for at least one OPC using a plurality of sCONV layers of different sizes. In some implementations, the plurality of sCONV layers of different sizes may include a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer. For example, refereeing to FIGs. 5A-5D, a plurality of parameters (e.g., weights and biases) may be generated during the training phase of OPC.
[0110] At 704, the apparatus may obtain a first reference picture from a decoded picture buffer. Referring to FIG. 2B, a first picture may be obtained from DPB, and one or more reconstructed pictures, prediction pictures, partition maps, QP maps, BS maps, etc. may be generated by inter predictor 260 and / or intra predictor 262.
[0111] At 706, the apparatus may generate a first feature map based on first information associated with the first reference picture. In some implementations, the first information associated with the first reference picture may include a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a BS map, a base QP map, a slice QP map, and a prediction type. In some implementations, the apparatus may generate the first feature map based on first information associated with the first reference picture by downsampling the first reference picture to obtain a downsampled reference picture. In some implementations, the apparatus may generate the first feature map based on first information associated with the first reference picture by generating the first feature map based on a feature extraction of the downsampled reference picture. For example, the first feature map may be generated based on one or more operations described above in connection with FIG. 3.
[0112] At 708, the apparatus may generate a residual feature map based on at least one OPC of the first feature map using a filter. In some implementations, the apparatus may generate the residual feature map based on the at least one OPC of the first feature map using the filter by generating a set of features based on a first OPC of the first feature map. In some implementations, the first set of features may be associated with a first number of channels. In some implementations, the apparatus may generate the residual feature map based on the at least one OPC of the first feature map using the filter by generating the residual feature map based on a second OPC of the set of features. In some implementations, the residual feature map being associated with a second number of channels fewer than the first number of channels. In some implementations, the filter may include an NN-based filter. In some implementations, the NN-based filter may include an LOP filter. For example, the residual map may be generated based on one or more operations described above in connection with FIG. 3.
[0113] At 710, the apparatus may generate a second feature map based on the residual feature map. In some implementations, the apparatus may generate the second feature map based on the residual feature map by generating a set of features based on the residual feature map. In some implementations, the apparatus may generate the second feature map based on the residual feature map by upsampling the set of features to obtain the second feature map. For example, the second map may be generated based on one or more operations described above in connection with FIG. 3.
[0114] At 712, the apparatus may generate a second reference picture based on the first reference picture and the second feature map. For example, the second reference picture may be generated based on one or more of the operations described above in connection with FIG. 3.
[0115] At 714, the apparatus may add the second reference picture into the decoded picture buffer or replace the first reference picture in the decoded picture buffer with the second reference picture. For example, referring to FIG. 2B, in-loop filter 266 may add the second reference picture to DPB or replace the first reference picture in the DPB with the second reference picture.
[0116] At 716, the apparatus may decode a current picture region based on the second reference picture. Referring to FIG. 2B, decoding apparatus 250 may decode a current picture region based on the second reference picture added to DPB.
[0117] FIG. 8 illustrates a flow chart of an exemplary method 800 of video encoding, according to some embodiments of the present disclosure. Method 800 may be performed by an apparatus, e.g., such as encoding apparatus 10, 200, encoding unit 12, in-loop filter 261, LOP filter 300, BBBA 302a, BBBB 302b, OPC1 408, OPC2 422, etc. Method 800 may include operations 802-816 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously, or in a different order other than shown in FIG. 8.
[0118] At 802, the apparatus may generate a set of parameters for at least one OPC using a plurality of sCONV layers of different sizes. In some implementations, the plurality of sCONV layers of different sizes may include a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer. For example, refereeing to FIGs. 5A-5D, a plurality of parameters (e.g., weights and biases) may be generated during the training phase of OPC.
[0119] At 804, the apparatus may obtain a first reference picture from a decoded picture buffer. Referring to FIG. 2A, a first picture may be obtained from DPB, and one or more reconstructed pictures, prediction pictures, partition maps, QP maps, BS maps, etc. may be generated by inter predictor 221 and / or intra predictor 222.
[0120] At 806, the apparatus may generate a first feature map based on first information associated with the first reference picture. In some implementations, the first information associated with the first reference picture may include a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a BS map, a base QP map, a slice QP map, and a prediction type. In some implementations, the apparatus may generate the first feature map based on first information associated with the first reference picture by downsampling the first reference picture to obtain a downsampled reference picture. In some implementations, the apparatus may generate the first feature map based on first information associated with the first reference picture by generating the first feature map based on a feature extraction of the downsampled reference picture. For example, the first feature map may be generated based on one or more operations described above in connection with FIG. 3.
[0121] At 808, the apparatus may generate a residual feature map based on at least one OPC of the first feature map using a filter. In some implementations, the apparatus may generate the residual feature map based on the at least one OPC of the first feature map using the filter by generating a set of features based on a first OPC of the first feature map. In some implementations, the first set of features may be associated with a first number of channels. In some implementations, the apparatus may generate the residual feature map based on the at least one OPC of the first feature map using the filter by generating the residual feature map based on a second OPC of the set of features. In some implementations, the residual feature map being associated with a second number of channels fewer than the first number of channels. In some implementations, the filter may include an NN-based filter. In some implementations, the NN-based filter may include an LOP filter. For example, the residual map may be generated based on one or more operations described above in connection with FIG. 3.
[0122] At 810, the apparatus may generate a second feature map based on the residual feature map. In some implementations, the apparatus may generate the second feature map based on the residual feature map by generating a set of features based on the residual feature map. In some implementations, the apparatus may generate the second feature map based on the residual feature map by upsampling the set of features to obtain the second feature map. For example, the second map may be generated based on one or more operations described above in connection with FIG. 3.
[0123] At 812, the apparatus may generate a second reference picture based on the first reference picture and the second feature map. For example, the second reference picture may be generated based on one or more of the operations described above in connection with FIG. 3.
[0124] At 814, the apparatus may add the second reference picture into the decoded picture buffer or replace the first reference picture in the decoded picture buffer with the second reference picture. For example, referring to FIG. 2A, in-loop filter 261 may add the second reference picture to DPB or replace the first reference picture in the DPB with the second reference picture.
[0125] At 816, the apparatus may encode a current picture region based on the second reference picture. Referring to FIG. 2A, encoding apparatus 200 may encode a current picture region based on the second reference picture added to DPB.
[0126] In the present disclosure, a BBB enhancement for a LOP in-loop filter based on over-parameterized convolution (OPC) and variable channel number is proposed to reduce the complexity and improve the performance for the LOP network. Since other LOP related proposals do not conflict with the exemplary BBB described herein, the BBB of FIG. 3 can be combined with other LOP-related proposals to further improve the performance of the LOP network. That is, the proposed BBB can also be applied to AK0195, while the proposed variable number of channels can be applied to AK0106.
[0127] The very low-operation point (VLOP) network adopts a similar network structure to the LOP network, but differs from the LOP network only in the number of channels and the number of BBBs. Therefore, the proposed method can also be applied to VLOP to improve the performance of VLOP.
[0128] The proposed OPC module can improve the performance of 1x5 and 5x1 sCONV layers. Therefore, other networks applied to 1x5 and 5x1 sCONV layers or 1x3 and 3x1 sCONV layers can also consider enhancing the performance based on the proposed OPC module.
[0129] In various aspects of the present disclosure, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as instructions on a non-transitory computer-readable medium. Computer-readable media includes computer storage media. Storage media may be any available media that can be accessed by a processor, such as a processor in encoding unit 12 or decoding unit 22 in FIG. 1. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, HDD, such as magnetic disk storage or other magnetic storage devices, Flash drive, SSD, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a processing system, such as a mobile device or a computer. Disk and disc, as used herein, include CD, laser disc, optical disc, digital video disc (DVD) , and floppy disk where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0130] According to one aspect of the present disclosure, a method of video decoding is provided. The method may include obtaining, by a processor, a first reference picture from a decoded picture buffer. The method may include generating, by the processor, a first feature map based on first information associated with the first reference picture. The method may include generating, by the processor, a residual feature map based on at least one OPC of the first feature map using a filter. The method may include generating, by the processor, a second feature map based on the residual feature map. The method may include generating, by the processor, a second reference picture based on the first reference picture and the second feature map. The method may include decoding, by the processor, a current picture region based on the second reference picture.
[0131] In some implementations, the generating, by the processor, the first feature map based on first information associated with the first reference picture may include downsampling, by the processor, the first reference picture to obtain a downsampled reference picture. In some implementations, the generating, by the processor, the first feature map based on first information associated with the first reference picture may include generating, by the processor, the first feature map based on a feature extraction of the downsampled reference picture.
[0132] In some implementations, the generating, by the processor, the residual feature map based on the at least one OPC of the first feature map using the filter may include generating, by the processor, a set of features based on a first OPC of the first feature map. In some implementations, the first set of features may be associated with a first number of channels. In some implementations, the generating, by the processor, the residual feature map based on the at least one OPC of the first feature map using the filter may include generating, by the processor, the residual feature map based on a second OPC of the set of features. In some implementations, the residual feature map being associated with a second number of channels fewer than the first number of channels.
[0133] In some implementations, the method may further include generating, by the processor, a set of parameters for the at least one OPC using a plurality of sCONV layers of different sizes.
[0134] In some implementations, the plurality of sCONV layers of different sizes may include a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.
[0135] In some implementations, the generating, by the processor, the second feature map based on the residual feature map may include generating, by the processor, a set of features based on the residual feature map. In some implementations, the generating, by the processor, the second feature map based on the residual feature map may include upsampling, by the processor, the set of features to obtain the second feature map.
[0136] In some implementations, the first information associated with the first reference picture may include a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a BS map, a base QP map, a slice QP map, and a prediction type.
[0137] In some implementations, the method may further include adding, by the processor, the second reference picture into the decoded picture buffer. In some implementations, the method may further include replacing, by the processor, the first reference picture in the decoded picture buffer with the second reference picture.
[0138] In some implementations, the filter may include an NN-based filter.
[0139] In some implementations, the NN-based filter may include an LOP filter.
[0140] According to another aspect of the present disclosure, a decoder is provided. The decoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a residual feature map based on at least one OPC of the first feature map using a filter. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the residual feature map. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the first reference picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to decode a current picture region based on the second reference picture.
[0141] In some implementations, to generate the first feature map based on first information associated with the first reference picture, the memory storing instructions, which when executed by the processor, may cause the processor to downsample the first reference picture to obtain a downsampled reference picture. In some implementations, to generate the first feature map based on first information associated with the first reference picture, the memory storing instructions, which when executed by the processor, may cause the processor to generate the first feature map based on a feature extraction of the downsampled reference picture.
[0142] In some implementations, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the memory storing instructions, which when executed by the processor, may cause the processor to generate a set of features based on a first OPC of the first feature map. In some implementations, the first set of features may be associated with a first number of channels. In some implementations, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the memory storing instructions, which when executed by the processor, may cause the processor to generate the residual feature map based on a second OPC of the set of features. In some implementations, the residual feature map may be associated with a second number of channels fewer than the first number of channels.
[0143] In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to generate a set of parameters for the at least one OPC using a plurality of sCONV layers of different sizes.
[0144] In some implementations, the plurality of sCONV layers of different sizes may include a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.
[0145] In some implementations, to generate the second feature map based on the residual feature map, the memory storing instructions, which when executed by the processor, may cause the processor to generate a set of features based on the residual feature map. In some implementations, to generate the second feature map based on the residual feature map, the memory storing instructions, which when executed by the processor, may cause the processor to upsample the set of features to obtain the second feature map.
[0146] In some implementations, the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a BS map, a base QP map, a slice QP map, and a prediction type.
[0147] In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to add the second reference picture into the decoded picture buffer. In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to replace the first reference picture in the decoded picture buffer with the second reference picture
[0148] In some implementations, the filter may include an NN-based filter.
[0149] In some implementations, the NN-based filter may include an LOP filter.
[0150] According to another aspect of the present disclosure, an apparatus for decoding is provided. The apparatus for decoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a residual feature map based on at least one OPC of the first feature map using a filter. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the residual feature map. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the first reference picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to decode a current picture region based on the second reference picture.
[0151] According to a further aspect of the present disclosure, a non-transitory computer readable medium storing instructions for a decoder is provided. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to obtain a first reference picture from a decoded picture buffer. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a first feature map based on first information associated with the first reference picture. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a residual feature map based on at least one OPC of the first feature map using a filter. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a second feature map based on the residual feature map. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a second reference picture based on the first reference picture and the second feature map. The instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to decode a current picture region based on the second reference picture.
[0152] In some implementations, to generate the first feature map based on first information associated with the first reference picture, the instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to downsample the first reference picture to obtain a downsampled reference picture. In some implementations, to generate the first feature map based on first information associated with the first reference picture, the instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate the first feature map based on a feature extraction of the downsampled reference picture.
[0153] In some implementations, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a set of features based on a first OPC of the first feature map. In some implementations, the first set of features may be associated with a first number of channels. In some implementations, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate the residual feature map based on a second OPC of the set of features. In some implementations, the residual feature map may be associated with a second number of channels fewer than the first number of channels.
[0154] In some implementations, the instructions, which when executed by the processor of the decoder, may further cause the processor of the decoder to generate a set of parameters for the at least one OPC using a plurality of sCONV layers of different sizes.
[0155] In some implementations, the plurality of sCONV layers of different sizes may include a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.
[0156] In some implementations, to generate the second feature map based on the residual feature map, the instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to generate a set of features based on the residual feature map. In some implementations, to generate the second feature map based on the residual feature map, the instructions, which when executed by the processor of the decoder, may cause the processor of the decoder to upsample the set of features to obtain the second feature map.
[0157] In some implementations, the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a BS map, a base QP map, a slice QP map, and a prediction type.
[0158] In some implementations, the instructions, which when executed by the processor of the decoder, may further cause the processor of the decoder to add the second reference picture into the decoded picture buffer. In some implementations, the instructions, which when executed by the processor of the decoder, may further cause the processor of the decoder to replace the first reference picture in the decoded picture buffer with the second reference picture
[0159] In some implementations, the filter may include an NN-based filter.
[0160] In some implementations, the NN-based filter may include an LOP filter.
[0161] According to one aspect of the present disclosure, a method of video encoding is provided. The method may include obtaining, by a processor, a first reference picture from a decoded picture buffer. The method may include generating, by the processor, a first feature map based on first information associated with the first reference picture. The method may include generating, by the processor, a residual feature map based on at least one OPC of the first feature map using a filter. The method may include generating, by the processor, a second feature map based on the residual feature map. The method may include generating, by the processor, a second reference picture based on the first reference picture and the second feature map. The method may include encoding, by the processor, a current picture region based on the second reference picture.
[0162] In some implementations, the generating, by the processor, the first feature map based on first information associated with the first reference picture may include downsampling, by the processor, the first reference picture to obtain a downsampled reference picture. In some implementations, the generating, by the processor, the first feature map based on first information associated with the first reference picture may include generating, by the processor, the first feature map based on a feature extraction of the downsampled reference picture.
[0163] In some implementations, the generating, by the processor, the residual feature map based on the at least one OPC of the first feature map using the filter may include generating, by the processor, a set of features based on a first OPC of the first feature map. In some implementations, the first set of features may be associated with a first number of channels. In some implementations, the generating, by the processor, the residual feature map based on the at least one OPC of the first feature map using the filter may include generating, by the processor, the residual feature map based on a second OPC of the set of features. In some implementations, the residual feature map being associated with a second number of channels fewer than the first number of channels.
[0164] In some implementations, the method may further include generating, by the processor, a set of parameters for the at least one OPC using a plurality of sCONV layers of different sizes.
[0165] In some implementations, the plurality of sCONV layers of different sizes may include a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.
[0166] In some implementations, the generating, by the processor, the second feature map based on the residual feature map may include generating, by the processor, a set of features based on the residual feature map. In some implementations, the generating, by the processor, the second feature map based on the residual feature map may include upsampling, by the processor, the set of features to obtain the second feature map.
[0167] In some implementations, the first information associated with the first reference picture may include a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a BS map, a base QP map, a slice QP map, and a prediction type.
[0168] In some implementations, the method may further include adding, by the processor, the second reference picture into the decoded picture buffer. In some implementations, the method may further include replacing, by the processor, the first picture buffer in the decoded picture buffer with the second reference picture.
[0169] In some implementations, the filter may include an NN-based filter.
[0170] In some implementations, the NN-based filter may include an LOP filter.
[0171] According to another aspect of the present disclosure, an encoder is provided. The encoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a residual feature map based on at least one OPC of the first feature map using a filter. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the residual feature map. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the first reference picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to encode a current picture region based on the second reference picture.
[0172] In some implementations, to generate the first feature map based on first information associated with the first reference picture, the memory storing instructions, which when executed by the processor, may cause the processor to downsample the first reference picture to obtain a downsampled reference picture. In some implementations, to generate the first feature map based on first information associated with the first reference picture, the memory storing instructions, which when executed by the processor, may cause the processor to generate the first feature map based on a feature extraction of the downsampled reference picture.
[0173] In some implementations, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the memory storing instructions, which when executed by the processor, may cause the processor to generate a set of features based on a first OPC of the first feature map. In some implementations, the first set of features may be associated with a first number of channels. In some implementations, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the memory storing instructions, which when executed by the processor, may cause the processor to generate the residual feature map based on a second OPC of the set of features. In some implementations, the residual feature map may be associated with a second number of channels fewer than the first number of channels.
[0174] In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to generate a set of parameters for the at least one OPC using a plurality of sCONV layers of different sizes.
[0175] In some implementations, the plurality of sCONV layers of different sizes may include a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.
[0176] In some implementations, to generate the second feature map based on the residual feature map, the memory storing instructions, which when executed by the processor, may cause the processor to generate a set of features based on the residual feature map. In some implementations, to generate the second feature map based on the residual feature map, the memory storing instructions, which when executed by the processor, may cause the processor to upsample the set of features to obtain the second feature map.
[0177] In some implementations, the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a BS map, a base QP map, a slice QP map, and a prediction type.
[0178] In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to add the second reference picture into the decoded picture buffer. In some implementations, the memory storing instructions, which when executed by the processor, may further cause the processor to replace the first reference picture in the decoded picture buffer with the second reference picture.
[0179] In some implementations, the filter may include an NN-based filter.
[0180] In some implementations, the NN-based filter may include an LOP filter.
[0181] According to another aspect of the present disclosure, an apparatus for encoding is provided. The apparatus for encoding may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a residual feature map based on at least one OPC of the first feature map using a filter. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the residual feature map. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the first reference picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to encode a current picture region based on the second reference picture.
[0182] According to a further aspect of the present disclosure, a non-transitory computer readable medium storing instructions for an encoder is provided. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to obtain a first reference picture from a decoded picture buffer. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a first feature map based on first information associated with the first reference picture. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a residual feature map based on at least one OPC of the first feature map using a filter. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a second feature map based on the residual feature map. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a second reference picture based on the first reference picture and the second feature map. The instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to encode a current picture region based on the second reference picture.
[0183] In some implementations, to generate the first feature map based on first information associated with the first reference picture, the instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to downsample the first reference picture to obtain a downsampled reference picture. In some implementations, to generate the first feature map based on first information associated with the first reference picture, the instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate the first feature map based on a feature extraction of the downsampled reference picture.
[0184] In some implementations, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a set of features based on a first OPC of the first feature map. In some implementations, the first set of features may be associated with a first number of channels. In some implementations, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate the residual feature map based on a second OPC of the set of features. In some implementations, the residual feature map may be associated with a second number of channels fewer than the first number of channels.
[0185] In some implementations, the instructions, which when executed by the processor of the encoder, may further cause the processor of the encoder to generate a set of parameters for the at least one OPC using a plurality of sCONV layers of different sizes.
[0186] In some implementations, the plurality of sCONV layers of different sizes may include a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.
[0187] In some implementations, to generate the second feature map based on the residual feature map, the instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to generate a set of features based on the residual feature map. In some implementations, to generate the second feature map based on the residual feature map, the instructions, which when executed by the processor of the encoder, may cause the processor of the encoder to upsample the set of features to obtain the second feature map.
[0188] In some implementations, the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a BS map, a base QP map, a slice QP map, and a prediction type.
[0189] In some implementations, the instructions, which when executed by the processor of the encoder, may further cause the processor of the encoder to add the second reference picture into the decoded picture buffer. In some implementations, the instructions, which when executed by the processor of the encoder, may further cause the processor of the encoder to replace the first picture in the decoded picture buffer with the second reference picture.
[0190] In some implementations, the filter may include an NN-based filter.
[0191] In some implementations, the NN-based filter may include an LOP filter.
[0192] According to yet a further aspect of the present disclosure, a non-transitory computer-readable medium storing a bitstream is provided. The bitstream may be generated based on one or more of the operations described herein.
[0193] The foregoing description of the embodiments will so reveal the general nature of the present disclosure that others can, by applying knowledge within the skill of the art, readily modify and / or adapt for various applications such embodiments, without undue experimentation, without departing from the general concept of the present disclosure. Therefore, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed embodiments, based on the teaching and guidance presented herein. It is to be understood that the phraseology or terminology herein is for the purpose of description and not of limitation, such that the terminology or phraseology of the present specification is to be interpreted by the skilled artisan in light of the teachings and guidance.
[0194] Embodiments of the present disclosure have been described above with the aid of functional building blocks illustrating the implementation of specified functions and relationships thereof. The boundaries of these functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternate boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed.
[0195] The Summary and Abstract sections may set forth one or more but not all exemplary embodiments of the present disclosure as contemplated by the inventor (s) , and thus, are not intended to limit the present disclosure and the appended claims in any way.
[0196] Various functional blocks, modules, and steps are disclosed above. The arrangements provided are illustrative and without limitation. Accordingly, the functional blocks, modules, and steps may be reordered or combined in different ways than in the examples provided above. Likewise, some embodiments include only a subset of the functional blocks, modules, and steps, and any such subset is permitted.
[0197] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.
Claims
1.A method of video decoding, comprising:obtaining, by a processor, a first reference picture from a decoded picture buffer;generating, by the processor, a first feature map based on first information associated with the first reference picture;generating, by the processor, a residual feature map based on at least one over-parameterized convolution (OPC) of the first feature map using a filter;generating, by the processor, a second feature map based on the residual feature map;generating, by the processor, a second reference picture based on the first reference picture and the second feature map; anddecoding, by the processor, a current picture region based on the second reference picture.2.The method of claim 1, wherein the generating, by the processor, the first feature map based on first information associated with the first reference picture comprises:downsampling, by the processor, the first reference picture to obtain a downsampled reference picture; andgenerating, by the processor, the first feature map based on a feature extraction of the downsampled reference picture.3.The method of claim 1, wherein the generating, by the processor, the residual feature map based on the at least one OPC the first feature map using the filter comprises:generating, by the processor, a set of features based on a first OPC of the first feature map, the set of features being associated with a first number of channels; andgenerating, by the processor, the residual feature map based on a second OPC of the set of features, the residual feature map being associated with a second number of channels fewer than the first number of channels.4.The method of claim 1, further comprising:generating, by the processor, a set of parameters for the at least one OPC using a plurality of separable convolutional (sCONV) layers of different sizes.5.The method of claim 4, wherein the plurality of sCONV layers of different sizes comprises a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.6.The method of claim 1, wherein the generating, by the processor, the second feature map based on the residual feature map comprises:generating, by the processor, a set of features based on the residual feature map; andupsampling, by the processor, the set of features to obtain the second feature map.7.The method of claim 1, wherein the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a boundary strength (BS) map, a base quantization parameter (QP) map, a slice QP map, and a prediction type.8.The method of claim 1, further comprising:adding, by the processor, the second reference picture into the decoded picture buffer; orreplacing, by the processor, the first reference picture in decoded picture buffer with the second reference picture.9.The method of claim 1, wherein the filter comprises a neural network (NN) -based filter.10.The method of claim 9, wherein the NN-based filter comprises a low-operation point (LOP) filter.11.A decoder, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture;generate a residual feature map based on at least one over-parameterized convolution (OPC) of the first feature map using a filter;generate a second feature map based on the residual feature map;generate a second reference picture based on the first reference picture and the second feature map; anddecode a current picture region based on the second reference picture.12.The decoder of claim 11, wherein, to generate the first feature map based on first information associated with the first reference picture, the memory storing instructions, which when executed by the processor, cause the processor to:downsample the first reference picture to obtain a downsampled reference picture; andgenerate the first feature map based on a feature extraction of the downsampled reference picture.13.The decoder of claim 11, wherein, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the memory storing instructions, which when executed by the processor, cause the processor to:generate a set of features based on a first OPC of the first feature map, the set of features being associated with a first number of channels; andgenerate the residual feature map based on a second OPC of the set of features, the residual feature map being associated with a second number of channels fewer than the first number of channels.14.The decoder of claim 11, wherein the memory storing instructions, which when executed by the processor, further cause the processor to:generate a set of parameters for the at least one OPC using a plurality of separable convolutional (sCONV) layers of different sizes.15.The decoder of claim 14, wherein the plurality of sCONV layers of different sizes comprises a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.16.The decoder of claim 11, wherein, to generate the second feature map based on the residual feature map, the memory storing instructions, which when executed by the processor, cause the processor to:generate a set of features based on the residual feature map; andupsample the set of features to obtain the second feature map.17.The decoder of claim 11, wherein the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a boundary strength (BS) map, a base quantization parameter (QP) map, a slice QP map, and a prediction type.18.The decoder of claim 11, wherein the memory storing instructions, which when executed by the processor, further cause the processor to:add the second reference picture into the decoded picture buffer; orreplace the first reference picture in decoded picture buffer with the second reference picture.19.The decoder of claim 11, wherein the filter comprises a neural network (NN) -based filter.20.The decoder of claim 19, wherein the NN-based filter comprises a low-operation point (LOP) filter.21.An apparatus for decoding, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture;generate a residual feature map based on at least one over-parameterized convolution (OPC) of the first feature map using a filter;generate a second feature map based on the residual feature map;generate a second reference picture based on the first reference picture and the second feature map; anddecode a current picture region based on the second reference picture.22.A non-transitory computer-readable medium storing instructions, which when executed by a processor of a decoder, cause the processor of the decoder to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture;generate a residual feature map based on at least one over-parameterized convolution (OPC) of the first feature map using a filter;generate a second feature map based on the residual feature map;generate a second reference picture based on the first reference picture and the second feature map; anddecode a current picture region based on the second reference picture.23.The non-transitory computer-readable medium of claim 22, wherein, to generate the first feature map based on first information associated with the first reference picture, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:downsample the first reference picture to obtain a downsampled reference picture; andgenerate the first feature map based on a feature extraction of the downsampled reference picture.24.The non-transitory computer-readable medium of claim 22, wherein, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:generate a set of features based on a first OPC of the first feature map, the set of features being associated with a first number of channels; andgenerate the residual feature map based on a second OPC of the set of features, the residual feature map being associated with a second number of channels fewer than the first number of channels.25.The non-transitory computer-readable medium of claim 22, wherein the instructions, which when executed by the processor of the decoder, further cause the processor of the decoder to:generate a set of parameters for the at least one OPC using a plurality of separable convolutional (sCONV) layers of different sizes.26.The non-transitory computer-readable medium of claim 25, wherein the plurality of sCONV layers of different sizes comprises a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.27.The non-transitory computer-readable medium of claim 22, wherein, to generate the second feature map based on the residual feature map, the instructions, which when executed by the processor of the decoder, cause the processor of the decoder to:generate a set of features based on the residual feature map; andupsample the set of features to obtain the second feature map.28.The non-transitory computer-readable medium of claim 22, wherein the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a boundary strength (BS) map, a base quantization parameter (QP) map, a slice QP map, and a prediction type.29.The non-transitory computer-readable medium of claim 22, wherein the instructions, which when executed by the processor of the decoder, further cause the processor of the decoder to:add the second reference picture into the decoded picture buffer; orreplace the first reference picture in decoded picture buffer with the second reference picture.30.The non-transitory computer readable medium of claim 22, wherein the filter comprises a neural network (NN) -based filter.31.The non-transitory computer readable medium of claim 30, wherein the NN-based filter comprises a low-operation point (LOP) filter.32.A method of video encoding, comprising:obtaining, by a processor, a first reference picture from a decoded picture buffer;generating, by the processor, a first feature map based on first information associated with the first reference picture;generating, by the processor, a residual feature map based on at least one over-parameterized convolution (OPC) of the first feature map using a filter;generating, by the processor, a second feature map based on the residual feature map;generating, by the processor, a second reference picture based on the first reference picture and the second feature map; andencoding, by the processor, a current picture region based on the second reference picture.33.The method of claim 32, wherein the generating, by the processor, the first feature map based on first information associated with the first reference picture comprises:downsampling, by the processor, the first reference picture to obtain a downsampled reference picture; andgenerating, by the processor, the first feature map based on a feature extraction of the downsampled reference picture.34.The method of claim 32, wherein the generating, by the processor, the residual feature map based on the at least one OPC the first feature map using the filter comprises:generating, by the processor, a set of features based on a first OPC of the first feature map, the set of features being associated with a first number of channels; andgenerating, by the processor, the residual feature map based on a second OPC of the set of features, the residual feature map being associated with a second number of channels fewer than the first number of channels.35.The method of claim 32, further comprising:generating, by the processor, a set of parameters for the at least one OPC using a plurality of separable convolutional (sCONV) layers of different sizes.36.The method of claim 35, wherein the plurality of sCONV layers of different sizes comprises a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.37.The method of claim 32, wherein the generating, by the processor, the second feature map based on the residual feature map comprises:generating, by the processor, a set of features based on the residual feature map; andupsampling, by the processor, the set of features to obtain the second feature map.38.The method of claim 32, wherein the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a boundary strength (BS) map, a base quantization parameter (QP) map, a slice QP map, and a prediction type.39.The method of claim 32, further comprising:adding, by the processor, the second reference picture into the decoded picture buffer; orreplacing, by the processor, the first reference picture in decoded picture buffer with the second reference picture.40.The method of claim 32, wherein the filter comprises a neural network (NN) -based filter.41.The method of claim 40, wherein the NN-based filter comprises a low-operation point (LOP) filter.42.An encoder, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture;generate a residual feature map based on at least one over-parameterized convolution (OPC) of the first feature map using a filter;generate a second feature map based on the residual feature map;generate a second reference picture based on the first reference picture and the second feature map; andencode a current picture region based on the second reference picture.43.The encoder of claim 42, wherein, to generate the first feature map based on first information associated with the first reference picture, the memory storing instructions, which when executed by the processor, cause the processor to:downsample the first reference picture to obtain a downsampled reference picture; andgenerate the first feature map based on a feature extraction of the downsampled reference picture.44.The encoder of claim 42, wherein, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the memory storing instructions, which when executed by the processor, cause the processor to:generate a set of features based on a first OPC of the first feature map, the set of features being associated with a first number of channels; andgenerate the residual feature map based on a second OPC of the set of features, the residual feature map being associated with a second number of channels fewer than the first number of channels.45.The encoder of claim 42, wherein the memory storing instructions, which when executed by the processor, further cause the processor to:generate a set of parameters for the at least one OPC using a plurality of separable convolutional (sCONV) layers of different sizes.46.The encoder of claim 45, wherein the plurality of sCONV layers of different sizes comprises a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.47.The encoder of claim 42, wherein, to generate the second feature map based on the residual feature map, the memory storing instructions, which when executed by the processor, cause the processor to:generate a set of features based on the residual feature map; andupsample the set of features to obtain the second feature map.48.The encoder of claim 42, wherein the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a boundary strength (BS) map, a base quantization parameter (QP) map, a slice QP map, and a prediction type.49.The encoder of claim 42, wherein the memory storing instructions, which when executed by the processor, further cause the processor to:add the second reference picture into the decoded picture buffer; orreplace the first reference picture in decoded picture buffer with the second reference picture.50.The encoder of claim 42, wherein the filter comprises a neural network (NN) -based filter.51.The encoder of claim 50, wherein the NN-based filter comprises a low-operation point (LOP) filter.52.An apparatus for encoding, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture;generate a residual feature map based on at least one over-parameterized convolution (OPC) of the first feature map using a filter;generate a second feature map based on the residual feature map;generate a second reference picture based on the first reference picture and the second feature map; andencode a current picture region based on the second reference picture.53.A non-transitory computer-readable medium storing instructions, which when executed by a processor of An encoder, cause the processor of the encoder to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture;generate a residual feature map based on at least one over-parameterized convolution (OPC) of the first feature map using a filter;generate a second feature map based on the residual feature map;generate a second reference picture based on the first reference picture and the second feature map; andencode a current picture region based on the second reference picture.54.The non-transitory computer-readable medium of claim 53, wherein, to generate the first feature map based on first information associated with the first reference picture, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:downsample the first reference picture to obtain a downsampled reference picture; andgenerate the first feature map based on a feature extraction of the downsampled reference picture.55.The non-transitory computer-readable medium of claim 53, wherein, to generate the residual feature map based on the at least one OPC the first feature map using the filter, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:generate a set of features based on a first OPC of the first feature map, the set of features being associated with a first number of channels; andgenerate the residual feature map based on a second OPC of the set of features, the residual feature map being associated with a second number of channels fewer than the first number of channels.56.The non-transitory computer-readable medium of claim 53, wherein the instructions, which when executed by the processor of the encoder, further cause the processor of the encoder to:generate a set of parameters for the at least one OPC using a plurality of separable convolutional (sCONV) layers of different sizes.57.The non-transitory computer-readable medium of claim 56, wherein the plurality of sCONV layers of different sizes comprises a 1x5 sCONV layer, a 5x1 sCONV layer, 1x3 sCONV layer, a 3x1 sCONV layer, and a first 1x1 sCONV layer, and a second first 1x1 sCONV layer.58.The non-transitory computer-readable medium of claim 53, wherein, to generate the second feature map based on the residual feature map, the instructions, which when executed by the processor of the encoder, cause the processor of the encoder to:generate a set of features based on the residual feature map; andupsample the set of features to obtain the second feature map.59.The non-transitory computer-readable medium of claim 53, wherein the first information associated with the first reference picture comprises a luma reconstructed picture, a chroma reconstructed picture, a chroma prediction picture, a chroma partition map, a boundary strength (BS) map, a base quantization parameter (QP) map, a slice QP map, and a prediction type.60.The non-transitory computer-readable medium of claim 53, wherein the instructions, which when executed by the processor of the encoder, further cause the processor of the encoder to:add the second reference picture into the decoded picture buffer; orreplace the first reference picture in decoded picture buffer with the second reference picture.61.The non-transitory computer-readable medium of claim 53, wherein the filter comprises a neural network (NN) -based filter.62.The non-transitory computer-readable medium of claim 53, wherein the NN-based filter comprises a low-operation point (LOP) filter.63.A non-transitory computer-readable medium storing a bitstream, the bitstream being generated based on one or more of claims 32-41.