Reducing complexity in video encoding and decoding
By using different number and size convolutional layers in video encoding to process reconstruction samples and additional input data, the complexity of neural network filters is reduced, the problem of high complexity in the prior art is solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202380090182.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-03
- Filing Date
- 2023-12-04
- Publication Date
- 2025-07-29
AI Technical Summary
Existing video encoding filters based on neural networks have high complexity, making it difficult to reduce the complexity of hardware implementation while maintaining performance.
Different convolutional layer outputs are generated to reduce the computational complexity of neural network filters by assigning the components of the reconstruction samples and additional input data to different numbers and sizes of convolutional layers.
While basically maintaining or improving the performance of neural network filters, the calculation complexity is reduced, the number of multiplication accumulation operations is reduced, and the compression efficiency is improved.
Smart Images

Figure CN120391056A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to reducing the complexity of video encoding and decoding. Background Art
[0002] Video is a major form of data traffic in today's networks and is expected to continue to grow in share, as disclosed by P. Cerwall et al. in the Ericsson Mobility Report (https: / / www.ericsson.com / en / mobility-report, November 2019). One way to reduce the data traffic of video is compression. In compression, the source video is encoded into a bitstream, which can then be stored and sent to the end user. The end user can use a decoder to extract the video data and display it on a screen.
[0003] However, since the encoder may not know what type of device the encoded bitstream will be sent to, the encoder must compress the video into a standard format. In this way, all devices that support the selected standard can successfully decode the video. Compression can be lossless (i.e., the decoded video will be the same as the source video provided to the encoder), or lossy (where a certain degradation of the received content is accepted). Since factors such as noise can make lossless compression quite expensive, whether the compression is lossless or lossy has a significant impact on the bit rate (i.e., how high the compression rate is).
[0004] A video sequence contains a series of pictures. The color space commonly used in video sequences is YCbCr, where Y is the luminance component, and Cb and Cr are the chrominance components. Sometimes, the Cb and Cr components are referred to as U and V. Other color spaces are also used, such as ICtCp (also known as IPT) (where I is the luminance component, and Ct and Cp are the chrominance components), constant-luminance YCbCr (where Y is the luminance component, and Cb and Cr are the chrominance components), RGB (where R, G, and B correspond to the red, green, and blue components respectively), YCoCg (where Y is the luminance component, and Co and Cg are the chrominance components), etc.
[0005] The order in which pictures are placed in a video sequence is called the "display order". Each picture is assigned a Picture Order Count (POC) value, which is used to indicate its display order. In the present disclosure, the terms "image", "picture", or "frame" are used interchangeably.
[0006] Video compression is used to compress a video sequence into a series of encoded pictures. In many existing video codecs, a picture is divided into blocks of different sizes. A block is a two-dimensional array of samples. The block serves as the basis for encoding. Then, the video decoder decodes the encoded picture into a picture containing sample values.
[0007] Video standards are usually developed by international organizations because these international organizations represent different companies and research institutions with different areas of expertise and interests. The most widely used video compression standard currently is H.264 / AVC (Advanced Video Coding), which was jointly developed by ITU-T and ISO. The first version of H.264 / AVC was finalized in 2003 and several updates have been made in the following years. The successor of H.264 / AVC (also developed by ITU-T (International Telecommunication Union - Telecommunication) and the International Organization for Standardization (ISO)) is called H.265 / HEVC (High Efficiency Video Coding) and was finalized in 2013. MPEG and ITU-T have created the successor of HEVC within the Joint Video Exploration Team (JVET). The name of this video codec is Versatile Video Coding (VVC), and version 1 of the VVC specification has been published as Rec. ITU-T H.266|ISO / IEC (International Electrotechnical Commission) 23090-3, "Versatile Video Coding", 2020.
[0008] The VVC video coding standard is a block-based video codec and utilizes both temporal prediction and spatial prediction. Spatial prediction is achieved using intra (I) prediction within the current picture. Temporal prediction is achieved using unidirectional (P) or bidirectional inter (B) prediction at the block level based on previously decoded reference pictures. In the encoder, the difference between the original pixel data and the predicted pixel data (referred to as the residual) is transformed into the frequency domain, quantized, and then entropy encoded before being sent together with necessary prediction parameters such as prediction mode and motion vector (which can also be entropy encoded). The decoder performs entropy decoding, inverse quantization, and inverse transformation to obtain the residual, and then adds the residual to the intra prediction or inter prediction to reconstruct the picture.
[0009] The VVC video coding standard uses a block structure called Quadtree Plus Binary Tree Plus Ternary Tree (QTBT+TT), where each picture is first segmented into square blocks called Coding Tree Units (CTUs). All CTUs have the same size, and the segmentation of the picture into CTUs is done without any syntax controlling it.
[0010] Each CTU is further divided into coding units (CUs) that can have a square or rectangular shape. The CTU is first divided by a quadtree structure and can then be further divided vertically or horizontally in a binary structure with the same size of division to form coding units (CUs). Thus, the blocks can have a square or rectangular shape. The depths of the quadtree and binary tree can be set by the encoder in the bitstream. The ternary tree (TT) part increases the possibility of dividing the CU into three divisions instead of two divisions of the same size. This increases the possibility of using a block structure that is more suitable for the content structure of the picture, such as roughly following the important edges in the picture.
[0011] The blocks that are intra-coded are I blocks. The blocks that are unidirectionally predicted are P blocks, and the blocks that are bidirectionally predicted are B blocks. For some blocks, the encoder decides not to encode the residuals, possibly because the prediction is close enough to the original result. Then, the encoder signals to the decoder that the transform coding of that block should be bypassed (i.e., skipped). Such a block is called a skipped block.
[0012] At the 20th JVET meeting, it was decided to set up exploratory experiments (EEs) on neural network (NN)-based video coding. The exploratory experiments continued at the 21st and 22nd JVET meetings, where two EE tests were conducted: NN-based filtering and NN-based super-resolution. At the 23rd JVET meeting, it was decided to continue with three categories of tests: enhancement filters, super-resolution methods, and intra prediction. In the category of enhancement filters, two configurations were considered: (i) the proposed filter is used as an in-loop filter; and (ii) the proposed filter is used as a post-processing filter.
[0013] VVC includes three in-loop filters that are not based on neural networks: the deblocking filter, the sample adaptive offset (SAO) filter, and the adaptive loop filter (ALF). The deblocking filter is used to remove block artifacts by smoothing the discontinuities in the horizontal and vertical directions across block boundaries. The deblocking filter uses the block boundary strength (BS) parameter to determine the filtering strength. The BS parameter can have values of 0, 1, and 2, where a larger value indicates stronger filtering. The output of the deblocking filter is further processed by SAO, and then the output of SAO is processed by the ALF operation. Then, the output of ALF is put into the display picture buffer (DPB), which is used for the prediction of subsequent encoded (or decoded) pictures. Since the deblocking filter, SAO filter, and ALF affect the pictures used for prediction in the DPB, they are classified as in-loop filters, also known as loop filters. The decoder can further filter the image but does not send the filtered output to the DPB, but only to the display. Compared with the loop filter, this filter does not affect future predictions and is therefore classified as a post-processing filter, also known as a post-filter.
[0014] The contributions described in EE1-1.6: Combined Test of EE1-1.2 and EE1-1.4 (Y. Li, K. Zhang, L. Zhang, H. Wang, J. Chen, K. Reuze, A. M. Kotra, M. Karczewicz, JVET-X0066, October 2021) and JVET-X0066 and EE1-1.2: Test on Deep In-Loop Filter with Adaptive Parameter Selection and Residual Scaling (Y. Li, K. Zhang, L. Zhang, H. Wang, K. Reuze, A. M. Kotra, M. Karczewicz, JVET-Y0143, January 2022) are two consecutive contributions describing NN-based in-loop filtering. Both contributions use the same NN model for filtering. The NN-based in-loop filter is placed before SAO and ALF, and the deblocking filter is turned off. The purpose of using the NN-based filter is to improve the quality of the reconstructed samples. Here, it is helpful for the NN model to be non-linear. Although the deblocking filter, SAO, and ALF all contain non-linear elements such as conditions and are thus not strictly linear, all three of them are based on linear filters. In contrast, a sufficiently large NN model can in principle learn any non-linear mapping and thus can represent a wider range of functional types compared to deblocking, SAO, and ALF.
[0015] In JVET-X0066 and JVET-Y0143, there are four NN models, namely, four NN-based in-loop filters - one for intra-frame samples of luminance, one for intra-frame samples of chrominance, one for inter-frame samples of luminance, and one for inter-frame samples of chrominance. The use of NN filtering can be controlled at the block (CTU) level or the picture level. The encoder can determine whether to use NN filtering for each block or each picture.
[0016] The NN-based in-loop filters proposed in JVET-X0066 and JVET-AB0053 have significantly improved the compression efficiency of the codec, i.e., it has significantly reduced the bitrate without reducing the objective quality measured by PSNR (Peak Signal-to-Noise Ratio) based on MSE (Mean Squared Error). The improvement in compression efficiency, or simply referred to as "gain", is typically measured as the Bjontegaard-delta rate (BDR) relative to an anchor. For example, a BDR of -1% means that the same PSNR can be achieved with a 1% reduction in bits. As reported in JVET-Y0143, for the random access (RA) configuration, the BDR gain for the luminance component (Y) is -9.80%, and for the all intra (AI) configuration, the BDR gain for the luminance component is -7.39%. The complexity of the NN model used for compression is typically measured by the number of multiply-accumulate (MAC) operations per pixel. The high gain of the NN model is directly related to the high complexity of the NN model. The complexity of the luminance intra-frame model described in JVET-Y0143 is 430kMAC / pixel, i.e., 430,000 multiply-accumulate operations per pixel. There are also other measures of complexity, such as the total model size in terms of the stored parameters. SUMMARY OF THE INVENTION
[0017] There are certain challenges at present. For example, since the high complexity of the NN filter is a major challenge for practical hardware implementation, it is highly desirable to reduce the complexity of the NN filter while maintaining the performance of the NN filter (i.e., optimizing the complexity-performance trade-off). However, the above-mentioned structure of the NN filter may not be optimal in terms of the complexity-performance trade-off.
[0018] In one example, the above-mentioned NN filter (neural network loop filter) is configured to generate an improved output picture based on a number of inputs, but before these multiple inputs reach the "backbone" or "main body" of the NN filter (where most of the calculations of the NN filter occur), these inputs are processed separately and then combined. These separate processes and combinations correspond to a large part of the overall process of the NN filter. Therefore, it is necessary to reduce this part of the process of the NN filter.
[0019] Accordingly, in a first aspect of the present disclosure, there is provided a method for encoding or decoding a video. The method includes: obtaining values of components of a reconstructed sample; obtaining first additional input data; and providing the obtained values of the components of the reconstructed sample to a first set of one or more convolutional layers to generate an output of the first set of convolutional layers. The method further includes: providing the obtained first additional input data to a second set of one or more convolutional layers to generate an output of the second set of convolutional layers; and encoding or decoding the video based on the output of the first set of convolutional layers and the output of the second set of convolutional layers. The first set of convolutional layers and the second set of convolutional layers are different in terms of the number of channels, the size of the kernel filter, and / or the number of convolutional layers.
[0020] According to a second aspect of the present disclosure, there is provided a computer program including instructions that, when executed by a processing circuit, cause the processing circuit to perform the method according to the first aspect.
[0021] According to a third aspect of the present invention, there is provided a carrier containing the computer program of the above embodiment, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0022] On the other hand, there is provided an apparatus for encoding or decoding a video. The apparatus is configured to perform the method according to the first aspect.
[0023] Some embodiments of the present disclosure provide a way to reduce the complexity of an NN model while substantially maintaining or improving the performance of an NN filter. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings included herein and forming a part of the specification illustrate various embodiments.
[0025] Figure 1A A system according to some embodiments is shown
[0026] Figure 1B A system according to some embodiments is shown.
[0027] Figure 1C A system according to some embodiments is shown.
[0028] Figure 2 A schematic block diagram of an encoder according to some embodiments is shown.
[0029] Figure 3 A schematic block diagram of a decoder according to some embodiments is shown.
[0030] Figure 4 A schematic block diagram of a part of an NN filter according to some embodiments is shown.
[0031] Figure 5 A schematic block diagram showing a part of an NN filter according to some embodiments is presented.
[0032] Figure 6 The computational operations performed by the convolutional layer are shown.
[0033] Figure 7 A process according to some embodiments is shown.
[0034] Figure 8 An apparatus according to some embodiments is shown. Detailed Description
[0035] The following terms are used in the description of the following embodiments.
[0036] Neural network: A general term for an entity of simple processing units (referred to as neurons or nodes) having one or more layers, which have activation functions and interact with each other via weighted connections and biases, and together constitute a tool in the context of a non - linear transformation.
[0037] Neural network architecture (abbreviated as network architecture or architecture): The layout of a neural network that describes the placement of nodes and their connections, usually in the form of several interconnected layers, and can also specify the dimensions of the input and output and the activation functions of the nodes.
[0038] Neural network weight (or simply weight): The weight value assigned to the connections between nodes in a neural network.
[0039] Neural network model (or simply model): A transformation in the form of a trained neural network. A neural network model can be specified with a neural network architecture, activation function, bias, and / or weights.
[0040] Filter: A transformation entity. A neural network model is an implementation of a filter. The term NN filter can be used as an abbreviation for a neural - network - based filter or a neural network filter.
[0041] Neural network training (or simply training): The process of finding the weight and bias values of a neural network. Usually, a training data set is used to train the neural network, and the goal of training is to minimize a defined error. The amount of training data needs to be large enough to avoid over - training. Training a neural network is usually a time - consuming task and typically involves multiple iterations over the training data, where each iteration is called an epoch.
[0042] Figure 1AFIG. 100 shows a system 100 according to some embodiments. The system 100 includes a first entity 102, a second entity 104, and a network 110. The first entity 102 is configured to send a video stream (also known as "video bitstream", "bitstream", "encoded video") 106 to the second entity 104.
[0043] The first entity 102 can be any computing device (e.g., a network node such as a server) capable of encoding video using an encoder 112 and sending the encoded video to the second entity 104 via the network 110. The second entity 104 can be any computing device (e.g., a network node) capable of receiving the encoded video and decoding the encoded video using a decoder 114. Each of the first entity 102 and the second entity 104 can be a single physical entity or a combination of multiple physical entities. The multiple physical entities can be located at the same location or can be distributed in the cloud.
[0044] In some embodiments, as Figure 1B shown, the first entity 102 is a video stream server 132, and the second entity 104 is a user equipment (UE) 134. The UE 134 can be any one of a desktop computer, a laptop computer, a tablet computer, a mobile phone, or any other computing device. The video stream server 132 is capable of sending a video bitstream 136 (e.g., a YouTube™ video stream) to the video stream client 134. Upon receiving the video bitstream 136, the UE 134 can decode the received video bitstream 136 to generate and display a video for the video stream.
[0045] In other embodiments, as Figure 1C shown, the first entity 102 and the second entity 104 are a first UE 152 and a second UE 154. For example, the first UE 152 can be a provider of a video conference session or a caller in a video chat, and the second UE 154 can be a responder to the video conference session or a responder to the video chat. In the Figure 1C embodiment shown, the first UE 152 is capable of sending a video bitstream 156 for a video conference (e.g., Zoom™, Skype™, MS Teams™, etc.) or a video chat (e.g., Facetime™) to the second UE 154. Upon receiving the video bitstream 156, the UE 154 can decode the received video bitstream 156 to generate and display a video for the video conference session or the video chat.
[0046] Figure 2FIG. 0 shows a schematic block diagram of an encoder 112 according to some embodiments. The encoder 112 is configured to encode blocks of sample values (hereinafter referred to as "blocks") in video frames of a source video 202. In the encoder 112, the current block (e.g., a block included in a video frame of the source video 202) is predicted by performing motion estimation by a motion estimator 250 based on blocks already provided in the same frame or a previous frame. In the case of inter-frame prediction, the result of the motion estimation is a motion or displacement vector associated with a reference block. The motion compensator 250 utilizes the motion vector to output an inter-frame prediction of the block.
[0047] The intra-frame predictor 249 calculates an intra-frame prediction of the current block. The outputs from the motion estimator / compensator 250 and the intra-frame predictor 249 are input to a selector 251, which selects either the intra-frame prediction or the inter-frame prediction for the current block. The output from the selector 251 is input to an error calculator in the form of an adder 241, which also receives the sample values of the current block. The adder 241 calculates and outputs a residual as the difference in sample values between the block and its prediction. This error is transformed in a transformer 242 (such as by a discrete cosine transform) and quantized by a quantizer 243, and then encoded in an encoder 244 (such as by an entropy encoder). In inter-frame encoding, the estimated motion vector is taken to the encoder 244 to generate an encoded representation of the current block.
[0048] The transformed and quantized residual of the current block is also provided to an inverse quantizer 245 and an inverse transformer 246 to obtain the original residual. This error is added by an adder 247 to the block prediction output from the motion compensator 250 or the intra-frame predictor 249 to create a reconstructed sample block 280 that can be used for prediction and encoding of the next block. The reconstructed sample block 280 is processed by an NN filter 230 (also known as a "neural network loop filter" or "NNLF") according to an embodiment in order to perform filtering to counter any block artifacts. Then, the output from the NN filter 230 (i.e., output data 290) is temporarily stored in a frame buffer 248, in which the output can be used by the intra-frame predictor 249 and the motion estimator / compensator 250.
[0049] In some embodiments, the encoder 112 may include a SAO unit 270 and / or an ALF 272. The SAO unit 270 and the ALF 272 may be configured to: receive the output data 290 from the NN filter 230, perform additional filtering on the output data 290, and provide the filtered output data to the buffer 248.
[0050] Although in Figure 2In the illustrated embodiment, the NN filter 230 is disposed between the SAO unit 270 and the adder 247. However, in other embodiments, the NN filter 230 may replace the SAO unit 270 and / or the ALF 272. Alternatively, in other embodiments, the NN filter 230 may be disposed between the buffer 248 and the motion compensator 250. Further, in some embodiments, a deblocking filter (not shown) may be disposed between the NN filter 230 and the adder 247 such that the reconstructed sample block 280 is deblocked and then provided to the NN filter 230.
[0051] Figure 3 is a schematic block diagram of a decoder 114 according to some embodiments. The decoder 114 includes a decoder 361 (such as an entropy decoder) that decodes the encoded representation of a block to obtain a set of quantized and transformed residuals. These residuals are dequantized in the inverse quantizer 362 and inverse-transformed by the inverse transformer 363 to obtain a set of residuals. These residuals are added to the sample values of the reference block in the adder 364. Depending on whether inter-frame prediction or intra-frame prediction is performed, the reference block is determined by the motion estimator / compensator 367 or the intra-frame predictor 366.
[0052] Thereby, the selector 368 is interconnected to the adder 364 and the motion estimator / compensator 367 and the intra-frame predictor 366. The resulting decoded block 380 output from the adder 364 is input to the NN filter unit 330 according to an embodiment to filter any block artifacts. The filtered block 390 is output from the NN filter 330 and is preferably also temporarily provided to the frame buffer 365 and can be used as a reference block for subsequent blocks to be decoded.
[0053] The frame buffer (e.g., a decoded picture buffer (DPB)) 365 is thus connected to the motion estimator / compensator 367 so that the stored sample blocks are available to the motion estimator / compensator 367. The output from the adder 364 is preferably also input to the intra-frame predictor 366 to be used as an unfiltered reference block.
[0054] In some embodiments, the decoder 114 may include a SAO unit 380 and / or an ALF 372. The SAO unit 380 and the ALF 382 may be configured to receive the output data 390 from the NN filter 330, perform additional filtering on the output data 390, and provide the filtered output data to the buffer 365.
[0055] Although in Figure 3In the illustrated embodiment, the NN filter 330 is disposed between the SAO unit 380 and the adder 364, but in other embodiments, the NN filter 330 may replace the SAO unit 380 and / or the ALF 382. Alternatively, in other embodiments, the NN filter 330 may be disposed between the buffer 365 and the motion compensator 367. Additionally, in some embodiments, a deblocking filter (not shown) may be disposed between the NN filter 330 and the adder 364 such that the reconstructed sample block 380 is deblocked and then provided to the NN filter 330.
[0056] Figure 4 FIG. is a schematic block diagram of a portion of the NN filter 230 / 330 for filtering intra-frame luminance samples according to some embodiments. In the present disclosure, an intra-frame luminance (or chrominance) sample is the luminance (or chrominance) component of a sample for which intra-frame prediction is performed. Similarly, an inter-frame luminance (or chrominance) sample is the luminance (or chrominance) component of a sample for which inter-frame prediction is performed.
[0057] As Figure 4 shown, the NN filter 230 / 330 may have four inputs: (1) the value of the luminance component of the reconstructed sample (“rec”) 280 / 380; (2) the value of the luminance component of the predicted sample (“pred”) 295 / 395; (3) block boundary strength BBS information (“bs”) indicating the filtering strength applied to the boundary of the luminance component of the sample; and (4) the quantization parameter (“qp”). In some embodiments, additional inputs (e.g., segmentation information indicating how to segment the luminance component of the sample) may additionally be used as inputs to the NN filter 230 / 330.
[0058] Each of the four inputs may be respectively passed through a convolutional layer (labeled “conv3x3” in Figure 4 ), and a parametric rectified linear unit (PReLU) layer (labeled “PReLU”). Then, the four outputs from the four PReLU layers may be concatenated via a concatenation unit (labeled “concat” in Figure 4 ) and fused together to generate data (also known as “signal”) “y”. The convolutional layer “conv3x3” is a convolutional layer with a kernel size of 3x3, and the convolutional layer “conv1x1” is a convolutional layer with a kernel size of 1x1. The PReLU may form an activation layer.
[0059] In some embodiments, the qp may be a scalar value. In such embodiments, the NN filter 230 / 330 may further include a dimension manipulation unit (in Figure 4(labeled as "decompression expansion" in [description]), which can be configured to expand qp such that the expanded qp has the same dimensions as other inputs (i.e., rec, pred, and bs). However, in other embodiments, qp can be a matrix whose dimensions can be the same as the dimensions of other inputs (such as rec, pred, and bs). For example, different samples within a CTU can be associated with different qp values. In such embodiments, a dimension manipulation unit is not required.
[0060] In some embodiments, the NN filter 230 / 330 may further include a downsampler (labeled as "2↓" in [description]), which is configured to perform downsampling by a factor of 2. Figure 4 In [description], it is configured to perform downsampling by a factor of 2.
[0061] As Figure 4 shown, the data "y" can be provided to a set of N consecutive attention residual (hereinafter referred to as "AR") blocks 402. In some embodiments, the N consecutive AR blocks 402 may have the same structure, while in other embodiments, they may have different structures. N can be any integer greater than or equal to 2. For example, N can be equal to 8.
[0062] As Figure 4 shown, the first AR block 402 included in the set can be configured to receive the data "y" and generate a first output data "z0". The second AR block 402 set immediately after the first AR block 402 can be configured to receive the first output data "z0" and generate a second output data "z1".
[0063] If the set includes only two AR blocks 402 (i.e., the aforementioned first AR block and the second AR block), then the second output data "z1" can be provided to the final processing unit 550 of the NN filter 230 / 330 (as Figure 5 shown).
[0064] On the other hand, if the set includes more than two AR blocks, each AR block 402 included in the set except for the first AR block and the last AR block can be configured to receive the output data from the previous AR block 402 and provide its output data to the next AR block. The last AR block 402 can be configured to receive the output data from the previous AR block and provide its output data to the final processing unit 550 of the NN filter 230 / 330. In Figure 4 [description], N - 1 corresponds to the number of AR blocks included in the NN filter 230 / 330.
[0065] In some embodiments, some or all of the AR blocks 402 may include a spatial attention block 412, which is configured to generate an attention mask . The attention mask It may have one channel and its size may be the same as that of the data “y”. Taking the first AR block 402 as an example, the spatial attention block included in the first AR block 402 may be configured to multiply the attention mask by the residual data “r” to obtain the data “r ”. The data “rf” may be combined with the residual data “r” and then combined with the data “y” to generate the first output data “z0”.
[0066] As Figure 5 shown, the output Z of the group of AR blocks 402 N-1 may be processed by a convolutional layer 502, a PReLU 504, another convolutional layer 506, a pixel rearrangement (or actual sample rearrangement) 508, and a final scaling 510 to generate the filtered output data (“output”) 290 / 390.
[0067] Return reference Figure 4 , the NN filter 230 / 330 may include four separate first layers 452 to 458. As Figure 4 shown, each of the four first layers 452 to 458 may include 3 × 3 convolutional layers and a PReLU. Each of the four first layers 452 to 458 may be configured to receive different input data. For example, layers 452, 454, 456, and 458 may be configured to receive the input data “rec”, “pred”, “bs”, and “qp” respectively.
[0068] Each of the 3 × 3 convolutional layers included in each of the first layers 452 to 458 may include 96 channels (i.e., 96 3 × 3 kernel filters), so for a single input (which includes multiple pixel values) provided to each of the first layers 452 to 458, there will be 96 outputs (i.e., one output corresponds to each of the 96 channels). Therefore, the first layers 452 to 458 may be configured to generate 4 × 96 outputs (also known as output channels) for the four inputs (i.e., “rec”, “pred”, “bs”, and “qp”). These 4 × 96 outputs may be processed by a concatenation unit, a fusion unit, and a transformation unit to be combined into a y corresponding to the 96 outputs.
[0069] Generally, a rough way to calculate the number of multiply-accumulate (MAC) operations required for a convolutional neural network layer is to multiply the number of inputs (also known as input channels), the number of filter coefficients, and the number of outputs (also known as output channels).
[0070] For example, as Figure 6As shown, assume that the input "rec" is an image including 9 pixel values. Note that this quantity is provided only for simple illustrative purposes and does not limit the embodiments of the present disclosure in any way. For each pixel value included in the input "rec", 96 different kernel filters 602 are applied. For example, if the kernel filter 602 of the convolutional layer included in the first layer 452 is applied to the pixel value a of the input "rec" 22 , then the filtered value corresponding to the pixel value a 22 can be calculated as follows: a 11 w 11 +a 12 w 12 +a 13 w 13 +a 21 w 21 +a 22 w 22 +a 23 w 23 +a 31 w 31 +a 32 w 32 +a 33 w 33 .
[0071] Thus, for each pixel value, 9 multiplication operations (corresponding to the number of filter coefficients w 11 , w 12 ...) are performed 96 times (the number of kernel filters corresponding to the number of output channels). Thus, here, the number of MAC operations performed for each pixel of the first layer 452 (which roughly indicates the computational complexity of the first layer 452) is approximately 1 (3 3) 96 = 864 MAC operations. If all layers 452 to 458 have the same number of filtering channels (i.e., the number of kernel filters), then the total number of MAC operations required for the four layers 452 to 458 will be 864 4 = 3456.
[0072] As Figure 4 shown, the outputs of the first layers 452 to 458 are provided to the second layer 462 of the NN filters 230 / 330. The first operation of the second layer 462 (the operation performed by the fusion unit) is 1 1 Convolution, followed by a second operation of the second layer 462 (an operation performed by the conversion unit), which is a downscaling 3 3 Convolution.
[0073] For the first operation of the second layer 462, the number of MAC operations can be calculated as follows: For the first operation of the second layer 462, since all the outputs of the first layer 452 to 458 are provided to the second layer 462, the number of inputs is 96 x 4 = 384. Additionally, for the first operation of the second layer 462, the number of filter coefficients is 1 (because a 1 1 kernel filter is used in the fusion unit), and the number of outputs from the 1 1 kernel filter is 96. Therefore, the number of MAC operations performed for the first operation of the second layer 462 is 384 1 1 96 = 36864.
[0074] Similarly, the number of MAC operations for the second operation of the second layer 462 can be calculated as follows: For the second operation of the second layer 462, the number of inputs (i.e., the number of outputs from the first operation of the second layer 462) is 96, the number of filter coefficients is 9 (because a 3 3 kernel filter is used in the 3 3 convolution layer), and the number of outputs is 96. Therefore, the number of MAC operations performed for the second operation of the second layer 462 is 96 MAC operations, where two multiplications are due to the fact that the output is downsampled in the x and y directions, which means only every fourth output sample needs to be calculated.
[0075] In summary, when using 96 filter channels for each of the four inputs, the total number of MAC operations performed to generate the output y in Figure 4 is 3456 + 36864 + 20736 = 61056 MAC operations.
[0076] In the above embodiment, all inputs (e.g., "rec", "pred", "bs", and "qp") are given the same importance level. However, different inputs may not be equally important. For example, the input "rec" corresponds to the pixel values of the reconstructed image. Since this input is the input that needs to be improved (meaning this input is the focus of the filtering process performed by the NN filters 230 / 330), it makes sense to place more emphasis on this input during the filtering process compared to other inputs.
[0077] Thus, according to some embodiments, different importance levels can be assigned to different inputs so as to perform different numbers of MAC operations for different inputs. More specifically, in some embodiments, higher importance can be given to inputs that are more important to the final result of the filtering process by configuring more MAC operations for those important inputs and fewer MAC operations for those less important inputs.
[0078] In some embodiments, different importance levels are assigned to different inputs by using different numbers of filtering channels for different inputs. For example, the number of filtering channels applied to the input "rec" (the most important input among the four inputs) can be increased from 96 to 192, thereby emphasizing the input "rec" more. On the other hand, the number of filtering channels applied to the input "pred" can be reduced from 96 to 48, and the number of filtering channels applied to each of the inputs "bs" and "qp" can be reduced from 96 to 24.
[0079] Changing the number of filtering channels applied to different inputs can result in a reduction in the number of MAC operations performed for NN filtering. More specifically, in the above example, the number of MAC operations for the input "rec" is 1 (3 3) 192 = 1728, the number of MAC operations for the input "pred" is 1 (3 3) 48 = 432, and the number of MAC operations for each of the inputs "bs" and "qp" is 1 (3 3) 24 = 216. Thus, in this example, the total number of MAC operations performed by the head of the NN filter 230 / 330 (i.e., Figure 4 the top of the NN filter 230 / 330 as shown) is 1728 + 432 + 216 + 216 = 2592. The following shows the difference in the total number of MAC operations performed in the case of using the same number of filtering channels for different inputs and using different numbers of filtering channels for different inputs.
[0080]
[0081] Variations in the number of filtering channels applied to different inputs can result in greater variations on the next stage of the NN filters 230 / 330, i.e., the second layer 462. For example, in the above example, the number of inputs (i.e., the number of input channels) to the second layer 462 of the NN filters 230 / 330 is 192 + 48 + 24 + 24 = 288 (compared to (96x4) = 384). Thus, the number of MAC operations performed by the 1 1 convolutional layer included in the fusion unit of the second layer 462 is 288 1 1 96 = 27648 MAC operations (compared to 384 1 1 96). Additionally, in the above example, the number of inputs (i.e., the number of input channels) to the 3 3 convolutional layers included in the transformation unit of the second layer 462 of the NN filters 230 / 330 is 96, so the number of MAC operations performed by the 3x3 convolutional layer is 96 3 3 96 = 20736 MAC operations.
[0082] Number of filtering channels for four different inputs: 96 for rec, 96 for pred, 96 for bs, 96 for qp Number of filtering channels for four different inputs: 192 for rec, 48 for pred, 24 for bs, 24 for qp Number of MAC operations performed by the first layer 452 to 458 3456 2592 Number of MAC operations performed by the fusion unit of the second layer 462 36864 27648 Number of MAC operations performed by the conversion unit of the second layer 462 20736 20736 Total number of MAC operations performed by the head of the NN filter 230 / 330 61056 50976
[0083] Therefore, the total number of MAC operations performed by the head of the NN filters 230 / 330 is reduced from 61056 MAC operations to 50976 MAC operations - a 16.5% reduction. The entire network (including the body and the tail) involves approximately 500k MAC operations, so the change in the head corresponds to a reduction of approximately (61056 - 50976) / 500000 = 2% in the overall kMAC / pixel of the network. At the same time, the increased attention to important reconstruction inputs leads to a slight improvement in compression efficiency of -0.03%.
[0084] In some embodiments, instead of changing the number of filtering channels applied to each of the inputs "rec", "pred", "bs", and "qp" to 192, 24, 12, and 12 respectively, the number of filtering channels is changed to 192, 24, 12, and 12 for the inputs "rec", "pred", "bs", and "qp" respectively. Based on the above method of calculating MAC operations, changing the number of filtering channels used in the head of the NN filters 230 / 330 reduces the overall kMAC / pixel of the entire model by 4%.
[0085] In some embodiments, in addition to changing the number of filtering channels as discussed above, the number of channels between the fusion layer and the transformation layer can also be reduced from 96 to 48. This results in a total reduction of 9% in kMAC / pixel for the model, and a BD-rate loss of only +0.03%.
[0086] Note that in the above embodiments, the NN filters 230 / 330 are in-loop filters, which means that their outputs can be used to predict future pictures in the video. However, a similar or identical neural network can alternatively be used as a post-filter to filter the pictures after they have been used for prediction.
[0087] In some embodiments, in addition to or instead of changing the number of filtering channels, a different set of inputs can be provided to the NN filters 230 / 330. More specifically, in some embodiments, instead of four inputs ("rec", "pred", "bs", and "qp"), additional inputs such as block size, motion vector, and / or prediction mode can be provided. In other embodiments, instead of some of the inputs "pred", "bs", and "qp", any one or more additional inputs can be provided. In the above embodiments, each of these additional or alternative inputs can also have a reduced number of channels compared to the reconstructed picture input. For example, the first layer of the NN filters 230 / 330 can receive five inputs - Figure 4 the four inputs shown and a fifth input identifying the block type (I / P / B - intra-coded / uni-directional prediction / bi-directional prediction), where 192 filtering channels are used for the input "rec", 24 filtering channels are used for the input "pred", and 12 filtering channels are used for the remaining three inputs.
[0088] In some embodiments, in addition to changing the number of filtering channels applied to different inputs, the size of the kernel filter used for filtering can also be changed. For example, instead of using a 3 3 kernel filter for the input "rec", a kernel filter of a first size (e.g., 5 5) can be used to filter the input "rec", while a kernel filter of a second size (3 3) can be used to filter the other inputs, where the first size is greater than the second size.
[0089] In some embodiments, the first layer 452 used to filter the input "rec" can filter the input "rec" using kernel filters of different sizes. For example, in the case where there are 196 filtering channels in the first layer 452, the first 96 of the 196 filtering channels can each use a first size (e.g., 5 5) of the kernel filter, but the remaining filter channels among the 196 filter channels can use respective kernel filters having a second size (3 3) of the kernel filter, where the first size and the second size are different. In some embodiments, the first size is greater than the second size.
[0090] In some embodiments, some inputs (e.g., the input "rec") can be more emphasized by including more than one convolutional layer in the first layer (e.g., 452). For example, although in Figure 4 the first layer 452 includes a single 3 3 convolutional layer, followed by a single PreLU, in some embodiments, the first layer 452 includes multiple 3 3 convolutional layers and multiple PreLUs. More specifically, in such an example, the first layer 452 can include a first 3 3 convolutional layer, followed by a first PreLU, followed by a second 3 3 convolutional layer, followed by a second PreLU.
[0091] In some embodiments, some inputs (e.g., the input "rec") can be more emphasized by increasing the number of bits used for the kernel filter. For example, a first number of bits (e.g., 14) can be used to indicate the filter weights of the kernel filter of the 3 3 convolutional layer of the first layer 452 (e.g., Figure 6 the W shown 11 、W 12 …), while a second number of bits (e.g., 7) can be used to indicate the filter weights of the kernel filter of the 3 3 convolutional layers of the first layers 454, 456, and / or 458, where the first number of bits is greater than the second number of bits.
[0092] In some embodiments, the input "rec" corresponds to the pixel values of the reconstructed luminance picture or the pixel values of the reconstructed chrominance picture. However, in other embodiments, the input "rec" corresponds to both the pixel values of the reconstructed luminance picture and the pixel values of the reconstructed chrominance picture. In such embodiments, the pixel values of the reconstructed luminance picture and the pixel values of the reconstructed chrominance picture can be emphasized differently. For example, if the NN filter 230 / 330 is used to predict chrominance samples, the pixel values of the reconstructed chrominance samples can be more emphasized. The aforementioned techniques (e.g., increasing the number of filter channels) can be used to more emphasize the pixel values of the reconstructed chrominance samples.
[0093] In summary, some of the above embodiments involve: for a specific input (e.g., the input "rec"), increasing the number of filter channels, the size of the kernel filter, and / or the number of bits for indicating the filter weights of the kernel filter, and / or using additional convolutional layers, while for other inputs, reducing or maintaining the number of filter channels, the size of the kernel filter, the number of bits for indicating the filter weights of the kernel filter, and / or the number of convolutional layers. Through these selective adjustments, more important inputs can be emphasized more strongly, thereby improving or maintaining the accuracy of the filtering operation of the NN filter 230 / 330 while reducing the complexity of the filtering operation.
[0094] Figure 7 A process 700 for encoding or decoding a video according to some embodiments is shown. Figure 7 It starts at step s702. Step s702 includes: obtaining the values of the components of the reconstructed samples. Step s704 includes: obtaining first additional input data. Step s706 includes: providing the obtained values of the components of the reconstructed samples to a first group of one or more convolutional layers, thereby generating the output of the first group of convolutional layers. Step s708 includes: providing the obtained first additional input data to a second group of one or more convolutional layers, thereby generating the output of the second group of convolutional layers. Step s710 includes: encoding or decoding the video based on the output of the first group of convolutional layers and the output of the second group of convolutional layers. The first group of convolutional layers and the second group of convolutional layers are different in terms of the number of channels, the size of the kernel filter, and / or the number of convolutional layers.
[0095] In some embodiments, process 700 includes: combining the output of the first group of convolutional layers and the output of the second group of convolutional layers to generate a combined output; providing the combined output to a third group of one or more convolutional layers to generate the output of the third group of convolutional layers, wherein the video is encoded or decoded based on the output of the third group of convolutional layers.
[0096] In some embodiments, the first additional input data is any one of the following: the values of the components of the predicted samples, the segmentation information indicating how to segment the components of the samples, the block boundary strength information indicating the filtering strength applied to the boundaries of the sample components, the value of the quantization parameter, the block size information indicating the block size, the motion vector information indicating the motion vector, the prediction mode information indicating the prediction mode, and / or the block type information indicating the block type.
[0097] In some embodiments, the number of channels included in the first group of convolutional layers is more than twice the number of channels included in the second group of convolutional layers.
[0098] In some embodiments, the number of channels included in the first group of convolutional layers is 192, and the number of channels included in the second group of convolutional layers is 48, 24, or 12.
[0099] In some embodiments, process 700 includes: obtaining second additional input data; and providing the second additional input data to a fourth set of one or more convolutional layers to generate an output of the fourth set of convolutional layers, wherein the video is encoded or decoded based on the output of the fourth set of convolutional layers, and the first set of convolutional layers and the fourth set of convolutional layers are different in terms of the number of channels, the size of the kernel filters, and / or the number of convolutional layers, and the second set of convolutional layers and the fourth set of convolutional layers are different in terms of the number of channels, the size of the kernel filters, and / or the number of convolutional layers.
[0100] In some embodiments, the number of channels included in the second set of convolutional layers is 48 or 24, and the number of channels included in the fourth set of convolutional layers is 24 or 12.
[0101] In some embodiments, the size of the kernel filters of the first set of convolutional layers is different from the size of the kernel filters of the second set of convolutional layers.
[0102] In some embodiments, the size of the kernel filters of the first set of convolutional layers is 5x5, and the size of the kernel filters of the second set of convolutional layers is 3x3.
[0103] In some embodiments, the convolutional layers in the first set of convolutional layers include a set of channels, a first subset of channels included in the set of channels includes kernel filters having a first size, a second subset of channels included in the set of channels includes kernel filters having a second size, and the first size and the second size are different.
[0104] In some embodiments, the size of the kernel filters of the first subset of channels is 5x5, and the size of the kernel filters of the second subset of channels is 3x3.
[0105] In some embodiments, the first set of convolutional layers includes a first number of convolutional layers, the second set of convolutional layers includes a second number of convolutional layers, and the first number is greater than the second number.
[0106] In some embodiments, the first number is 2, and the second number is 1.
[0107] In some embodiments, the first set of convolutional layers includes a first convolutional layer and a second convolutional layer, the first convolutional layer includes kernel filters having a first size, the second convolutional layer includes kernel filters having a second size, and the first size and the second size are different.
[0108] In some embodiments, the first size is 5x5, and the second size is 3x3.
[0109] In some embodiments, the first set of convolutional layers includes kernel filters having a first plurality of filter values, the second set of convolutional layers includes kernel filters having a second plurality of filter values, and the number of bits of each filter value included in the first plurality of filter values is greater than the number of bits of each filter value included in the second plurality of filter values.
[0110] In some embodiments, the number of bits of each filter value included in the first plurality of filter values is 14 bits, and the number of bits of each filter value included in the second plurality of filter values is 7 bits.
[0111] In some embodiments, the reconstructed sample is one of a reconstructed luminance sample and a reconstructed chrominance sample. Process 700 may include: obtaining a value of a component of the other of the reconstructed luminance sample and the reconstructed chrominance sample; providing the obtained value of the component of the other of the reconstructed luminance sample and the reconstructed chrominance sample to a fifth set of one or more convolutional layers; and encoding or decoding the video based on an output of the fifth set of convolutional layers, wherein the first set of convolutional layers and the fifth set of convolutional layers are different in terms of the number of channels, the size of the kernel filters, and / or the number of convolutional layers.
[0112] Figure 8 is a block diagram of an apparatus 800 for implementing the encoder 112, the decoder 114, or components included in the encoder 112 or the decoder 114 (e.g., the NN filters 280 or 330) according to some embodiments. When the apparatus 800 implements the decoder, the apparatus 800 may be referred to as "decoding apparatus 800", and when the apparatus 800 implements the encoder, the apparatus 800 may be referred to as "encoding apparatus 800". As Figure 8As shown, device 800 may include: processing circuitry (PC) 802, which may include one or more processors (P) 855 (e.g., general microprocessors and / or one or more other processors such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc.), which may be co-located in a single housing or a single data center, or may be geographically distributed (i.e., device 800 may be a distributed computing device); at least one network interface 848, which includes a transmitter (Tx) 845 and a receiver (Rx) 847 for enabling device 800 to send data to and receive data from other nodes connected to network 110 (e.g., an Internet Protocol (IP) network), and network interface 848 is (directly or indirectly) connected to this network 110 (e.g., network interface 848 may be wirelessly connected to network 110, in which case network interface 848 is connected to an antenna arrangement); and a storage unit (also known as a "data storage system") 808, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 802 includes a programmable processor, a computer program product (CPP) 841 may be provided. CPP 841 includes a computer-readable medium (CRM) 842, which stores a computer program (CP) 843 including computer-readable instructions (CRI) 844. CRM 842 may be a non-transitory computer-readable medium such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., random access memory, flash memory), etc. In some embodiments, CRI 844 of computer program 843 is configured such that when executed by PC 802, CRI causes device 800 to perform the steps described herein (e.g., the steps described herein with reference to the flowcharts). In other embodiments, device 800 may be configured to perform the steps described herein without code. That is, for example, PC 802 may consist of only one or more ASICs. Thus, the features of the embodiments described herein may be implemented in hardware and / or software fashion.
[0113] Overview of Embodiments
[0114] A1. A method (700) for encoding or decoding video, the method comprising:
[0115] Obtaining (s702) values of components of reconstructed samples;
[0116] Obtaining (s704) first additional input data;
[0117] Providing (s706) the obtained values of components of reconstructed samples to a first set of one or more convolutional layers, thereby generating an output of the first set of convolutional layers;
[0118] Providing (s708) the obtained first additional input data to a second set of one or more convolutional layers to generate an output of the second set of convolutional layers; and
[0119] Encoding or decoding the video based on the output of the first set of convolutional layers and the output of the second set of convolutional layers (s710), wherein
[0120] the first set of convolutional layers and the second set of convolutional layers are different in terms of the number of channels, the size of the kernel filters, and / or the number of convolutional layers.
[0121] A1-1. The method according to embodiment A1, the method comprising:
[0122] Combining the output of the first set of convolutional layers and the output of the second set of convolutional layers to generate a combined output;
[0123] Providing the combined output to a third set of one or more convolutional layers to generate an output of the third set of convolutional layers, wherein
[0124] Encoding or decoding the video based on the output of the third set of convolutional layers.
[0125] A2. The method according to embodiment A1 or A1-1, wherein the first additional input data is any one of the following: the value of a component of a prediction sample, segmentation information indicating how to segment a component of a sample, block boundary strength information indicating the filtering strength applied to the boundary of a sample component, the value of a quantization parameter, block size information indicating the block size, motion vector information indicating a motion vector, prediction mode information indicating a prediction mode, and / or block type information indicating a block type.
[0126] A3. The method according to any one of embodiments A1 to A1-1, wherein the number of channels included in the first set of convolutional layers is more than twice the number of channels included in the second set of convolutional layers.
[0127] A4. The method according to embodiment A3, wherein
[0128] the number of channels included in the first set of convolutional layers is 192, and
[0129] the number of channels included in the second set of convolutional layers is 48, 24, or 12.
[0130] A5. The method according to any one of embodiments A1 to A4, the method comprising:
[0131] Obtaining second additional input data; and
[0132] Providing the second additional input data to a third set of one or more convolutional layers to generate an output of the third set of convolutional layers, wherein
[0133] Encoding or decoding a video based on the output of a third set of convolutional layers
[0134] The first set of convolutional layers and the third set of convolutional layers are different in terms of the number of channels, the size of the kernel filters, and / or the number of convolutional layers, and
[0135] The second set of convolutional layers and the third set of convolutional layers are different in terms of the number of channels, the size of the kernel filters, and / or the number of convolutional layers.
[0136] A6. The method according to embodiment A5 (when embodiment A5 is subordinate to embodiment A4), wherein
[0137] The number of channels included in the second set of convolutional layers is 48 or 24, and
[0138] The number of channels included in the third set of convolutional layers is 24 or 12.
[0139] A7. The method according to any one of embodiments A1 to A6, wherein the size of the kernel filters of the first set of convolutional layers is different from the size of the kernel filters of the second set of convolutional layers.
[0140] A8. The method according to embodiment A7, wherein
[0141] The size of the kernel filters of the first set of convolutional layers is 5x5, and
[0142] The size of the kernel filters of the second set of convolutional layers is 3x3.
[0143] A9. The method according to any one of embodiments A1 to A8, wherein
[0144] The convolutional layers of the first set of convolutional layers include a channel set
[0145] The first channel subset included in the channel set includes kernel filters having a first size
[0146] The second channel subset included in the channel set includes kernel filters having a second size, and
[0147] The first size and the second size are different.
[0148] A10. The method according to embodiment A9, wherein
[0149] The size of the kernel filters of the first channel subset is 5x5, and
[0150] The size of the kernel filters of the second channel subset is 3x3.
[0151] A11. The method according to any one of embodiments A1 to A10, wherein
[0152] The first set of convolutional layers includes a first number of convolutional layers,
[0153] the second set of convolutional layers includes a second number of convolutional layers, and
[0154] the first number is greater than the second number.
[0155] A12. The method according to embodiment A11, wherein the first number is 2 and the second number is 1.
[0156] A13. The method according to embodiment A11 or A12, wherein
[0157] the first set of convolutional layers includes a first convolutional layer and a second convolutional layer,
[0158] the first convolutional layer includes a kernel filter having a first size,
[0159] the second convolutional layer includes a kernel filter having a second size, and
[0160] the first size and the second size are different.
[0161] A14. The method according to embodiment A13, wherein
[0162] the first size is 5x5, and
[0163] the second size is 3x3.
[0164] A15. The method according to any one of embodiments A1 to A14, wherein
[0165] the first set of convolutional layers includes kernel filters having a first plurality of filter values,
[0166] the second set of convolutional layers includes kernel filters having a second plurality of filter values, and
[0167] the number of bits of each filter value included in the first plurality of filter values is greater than the number of bits of each filter value included in the second plurality of filter values.
[0168] A16. The method according to embodiment A15, wherein
[0169] the number of bits of each filter value included in the first plurality of filter values is 14 bits, and
[0170] the number of bits of each filter value included in the second plurality of filter values is 7 bits.
[0171] A17. The method according to any one of embodiments A1 to A16, wherein
[0172] The reconstructed sample is one of a reconstructed luminance sample and a reconstructed chrominance sample,
[0173] The method includes:
[0174] obtaining a value of a component of the other one of the reconstructed luminance sample and the reconstructed chrominance sample;
[0175] providing the obtained value of the component of the other one of the reconstructed luminance sample and the reconstructed chrominance sample to a third set of one or more convolutional layers; and
[0176] encoding or decoding the video based on an output of the third set of convolutional layers, and
[0177] a first set of convolutional layers and the third set of convolutional layers are different in terms of a number of channels, a size of a kernel filter, and / or a number of convolutional layers.
[0178] B1. A computer program (800) comprising instructions (844) which, when executed by a processing circuit (802), cause the processing circuit to perform the method according to any one of embodiments A1 to A17.
[0179] B2. A carrier containing the computer program according to embodiment B1, wherein the carrier is one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0180] C1. An apparatus (800) for encoding or decoding a video, the apparatus being configured to:
[0181] obtain (s702) a value of a component of the reconstructed sample;
[0182] obtain (s704) first additional input data;
[0183] provide (s706) the obtained value of the component of the reconstructed sample to a first set of one or more convolutional layers, thereby generating an output of the first set of convolutional layers;
[0184] provide (s708) the obtained first additional input data to a second set of one or more convolutional layers, thereby generating an output of the second set of convolutional layers; and
[0185] encode or decode the video based on the output of the first set of convolutional layers and the output of the second set of convolutional layers (s710), wherein
[0186] the first set of convolutional layers and the second set of convolutional layers are different in terms of a number of channels, a size of a kernel filter, and / or a number of convolutional layers.
[0187] C2. The apparatus according to embodiment C1, wherein the apparatus is further configured to perform the method according to any one of embodiments A2 to A17.
[0188] D1. An apparatus (800) comprising:
[0189] processing circuitry (802); and
[0190] a memory (841) containing instructions executable by the processing circuitry, whereby the apparatus is operable to perform a method according to any one of embodiments A1 to A17.
[0191] Conclusion
[0192] While various embodiments have been described herein, it should be understood that these embodiments are presented by way of example only and not limitation. Accordingly, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments. Additionally, unless otherwise specified herein or clearly contradicted by context, the present disclosure encompasses any combination of the above elements in all possible variations thereof.
[0193] Furthermore, while the processes shown above and in the figures are shown as a sequence of steps, this is for illustration only. Thus, it is contemplated that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be performed in parallel.
Claims
1. A method (700) for encoding or decoding a video, the method comprising: obtaining (s702) values of components of a reconstructed sample; obtaining (s704) first additional input data; providing (s706) the obtained values of components of the reconstructed sample to a first group of one or more convolutional layers, thereby generating an output of the first group of convolutional layers; providing (s708) the obtained first additional input data to a second group of one or more convolutional layers, thereby generating an output of the second group of convolutional layers; and encoding or decoding the video based on the output of the first group of convolutional layers and the output of the second group of convolutional layers (s710), wherein, the first group of convolutional layers and the second group of convolutional layers are different in terms of the number of channels, the size of the kernel filter, and / or the number of convolutional layers.
2. The method according to claim 1, further comprising: combining the output of the first group of convolutional layers and the output of the second group of convolutional layers, thereby generating a combined output; providing the combined output to a third group of one or more convolutional layers, thereby generating an output of the third group of convolutional layers, wherein, encoding or decoding the video based on the output of the third group of convolutional layers.
3. The method according to claim 1 or 2, wherein The first additional input data is any one of the following: values of components of a predicted sample, segmentation information indicating how to segment components of a sample, block boundary strength information indicating the filtering strength applied to boundaries of components of a sample, a value of a quantization parameter, block size information indicating a block size, motion vector information indicating a motion vector, prediction mode information indicating a prediction mode, and / or block type information indicating a block type.
4. The method according to any one of claims 1 to 2, wherein The number of channels included in the first group of convolutional layers is more than twice the number of channels included in the second group of convolutional layers.
5. The method according to claim 4, wherein, the number of channels included in the first group of convolutional layers is 192, and the number of channels included in the second group of convolutional layers is 48, 24, or 12.
6. The method according to any one of claims 1 to 5, the method comprising: obtaining second additional input data; and providing the second additional input data to a fourth group of one or more convolutional layers, thereby generating an output of the fourth group of convolutional layers, wherein, encoding or decoding the video based on the output of the fourth group of convolutional layers, the first group of convolutional layers and the fourth group of convolutional layers are different in terms of the number of channels, the size of the kernel filter, and / or the number of convolutional layers, and the second group of convolutional layers and the fourth group of convolutional layers are different in terms of the number of channels, the size of the kernel filter, and / or the number of convolutional layers.
7. The method according to claim 6 which depends on claim 5, wherein, the number of channels included in the second group of convolutional layers is 48 or 24, and the number of channels included in the fourth group of convolutional layers is 24 or 12.
8. The method according to any one of claims 1 to 7, wherein, The size of the kernel filter of the first group of convolutional layers is different from the size of the kernel filter of the second group of convolutional layers.
9. The method according to claim 8, wherein, the size of the kernel filter of the first group of convolutional layers is 5x5, and The size of the kernel filter of the second set of convolutional layers is 3x3.
10. The method according to any one of claims 1 to 9, wherein the convolutional layers of the first set of convolutional layers include a set of channels, a first subset of channels included in the set of channels includes kernel filters having a first size, a second subset of channels included in the set of channels includes kernel filters having a second size, and the first size and the second size are different.
11. The method according to claim 10, wherein the size of the kernel filter of the first subset of channels is 5x5, and the size of the kernel filter of the second subset of channels is 3x3.
12. The method according to any one of claims 1 to 11, wherein the first set of convolutional layers includes a first number of convolutional layers, the second set of convolutional layers includes a second number of convolutional layers, and the first number is greater than the second number.
13. The method according to claim 12, wherein, The first number is 2, and the second number is 1.
14. The method according to claim 12 or 13, wherein the first set of convolutional layers includes a first convolutional layer and a second convolutional layer, the first convolutional layer includes a kernel filter having a first size, the second convolutional layer includes a kernel filter having a second size, and the first size and the second size are different.
15. The method according to claim 14, wherein the first size is 5x5, and the second size is 3x3.
16. The method according to any one of claims 1 to 15, wherein the first set of convolutional layers includes kernel filters having a first plurality of filter values, the second set of convolutional layers includes kernel filters having a second plurality of filter values, and the number of bits of each filter value included in the first plurality of filter values is greater than the number of bits of each filter value included in the second plurality of filter values.
17. The method according to claim 16, wherein the number of bits of each filter value included in the first plurality of filter values is 14 bits, and the number of bits of each filter value included in the second plurality of filter values is 7 bits.
18. The method according to any one of embodiments 1 to 17, wherein the reconstructed sample is one of a reconstructed luminance sample and a reconstructed chrominance sample, the method includes: obtaining the value of a component of the other of the reconstructed luminance sample and the reconstructed chrominance sample; providing the obtained value of the component of the other of the reconstructed luminance sample and the reconstructed chrominance sample to a fifth set of one or more convolutional layers; and encoding or decoding the video based on the output of the fifth set of convolutional layers, and the first set of convolutional layers and the fifth set of convolutional layers are different in terms of the number of channels, the size of the kernel filter, and / or the number of convolutional layers.
19. A computer program (800) comprising instructions (844) which, when executed by a processing circuit (802), cause the processing circuit to perform the method according to any one of claims 1 to 18.
20. A carrier containing the computer program according to claim 19, wherein, The carrier is one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium.
21. An apparatus (800) for encoding or decoding a video, the apparatus being configured to perform the method according to any one of claims 1 to 18.